Sound signal processing method and sound signal processing device

The sound signal processing method allows performers to hear their own performance sounds from designated positions in a virtual space by applying binaural processing based on positional information, improving the realism and uniqueness of virtual performances.

JP2026055179APending Publication Date: 2026-03-31YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing sound signal processing devices do not allow the sound of one's own performance to be audible from an arbitrary position in a virtual space, limiting the realism of virtual performances.

Method used

A sound signal processing method that involves receiving acoustic space information and performer positions, applying binaural processing to sound signals based on these positions, and outputting the processed signals to headphones, allowing the sound of one's own performance to be localized from a designated position in the virtual space.

Benefits of technology

Enables performers to perceive their own performance sounds from a specific location in the virtual space, enhancing the realism of virtual performances and providing a unique customer experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026055179000001_ABST
    Figure 2026055179000001_ABST
Patent Text Reader

Abstract

This provides a sound processing method that makes your own performance sound as if it were coming from a predetermined location within the space. [Solution] In the sound signal processing method, the sound signal processing device, which is a general-purpose information processing device, receives acoustic space information, first position information of a first performer in the acoustic space, and second position information of a second performer in the acoustic space, receives a first sound signal related to the performance of the first performer from an audio interface, receives a second sound signal related to the performance of the second performer via a network, applies first binaural processing to the first sound signal based on the first position information, applies second binaural processing to the second sound signal based on the second position information, and outputs the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing from the audio interface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One embodiment of the present invention relates to a sound signal processing method and a sound signal processing device.

Background Art

[0002] Patent Document 1 discloses an information processing device including an acoustic processing unit that performs acoustic processing of convolution of a transfer characteristic according to a positional relationship between respective users in a virtual space on an acoustic signal obtained by sound collection in a space where each of a plurality of co-performing users is present, and an output control unit that causes sound based on a signal generated by the acoustic processing to be output from output devices used by the respective users.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Since the device of Patent Document 1 convolves the transfer characteristic so that the sound of other performers can be heard at a relative position based on the position of the first performer, it does not make the sound of one's own performance audible from an arbitrary position in the space.

[0005] One of the present embodiments aims to provide a sound processing method capable of making the sound of one's own performance audible from a predetermined position in the space.

Means for Solving the Problems

[0006] The sound signal processing method involves a first terminal receiving acoustic space information, first position information of a first performer in the acoustic space, and second position information of a second performer in the acoustic space; the first terminal receiving a first sound signal related to the performance of the first performer from an audio interface; the first terminal outputting the first sound signal to the second terminal of the second performer via a network; the first terminal receiving a second sound signal related to the performance of the second performer via the network; the first terminal applying first binaural processing to the first sound signal based on the first position information; applying second binaural processing to the second sound signal based on the second position information; and the first terminal outputting the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing from the audio interface. [Effects of the Invention]

[0007] According to one embodiment of the present invention, it is possible to make the sound of one's own performance audible from a predetermined location within the space. [Brief explanation of the drawing]

[0008] [Figure 1] This is a block diagram showing the configuration of an audio signal processing system. [Figure 2] This is a block diagram showing the configuration of the sound signal processing device 1. [Figure 3] This is a block diagram showing the configuration of the sound signal processing device 1A. [Figure 4] This is a plan view showing the configuration of the virtual space 101. [Figure 5] This is a flowchart showing the operation of the sound signal processing device 1. [Figure 6] This is a plan view showing the configuration of virtual space 101. [Figure 7] This is a flowchart showing the operation of the second terminal. [Figure 8] This is a plan view showing the configuration of virtual space 101. [Figure 9]This is a block diagram showing the configuration of the sound signal processing system according to modified example 6. [Figure 10] This is a block diagram showing the configuration of an audio signal processing device 1B, which is an example of a viewer terminal. [Figure 11] This is a plan view showing the configuration of virtual space 101. [Figure 12] This is a flowchart showing the operation of the sound signal processing device 1B. [Figure 13] This is a block diagram showing the configuration of the sound signal processing system according to modified example 9. [Figure 14] This block diagram shows the configuration of an audio signal processing device 1C, which is an example of an administrator terminal. [Figure 15] This is a block diagram showing the configuration of the sound signal processing device 1 according to modified example 10. [Figure 16] This is a block diagram showing the configuration of the sound signal processing system according to modified example 12. [Figure 17] This is a conceptual diagram illustrating the configuration of an audio signal processing system when providing individual environments. [Modes for carrying out the invention]

[0009] Figure 1 is a block diagram showing the configuration of the sound signal processing system of this embodiment. The sound signal processing system has a first terminal and a second terminal connected via a network. Figure 2 is a block diagram showing the configuration of the sound signal processing device 1. Figure 3 is a block diagram showing the configuration of the sound signal processing device 1A. The sound signal processing device 1 is, for example, a general-purpose information processing device and is an example of the first terminal. The sound signal processing device 1A is, for example, a general-purpose information processing device and is an example of the second terminal. The user connects a head-mounted display (HMD) 51 and an audio I / F 71 to the sound signal processing device 1. The sound signal processing device 1 includes a display I / F 31, a user I / F 32, flash memory 33, a CPU 34, RAM 35, a communication I / F 36, a USB I / F 37, and a DSP 38.

[0010] The display I / F 31 is connected to a display. The display I / F 31 of the audio signal processing device 1 is connected to an HMD 51 which is an example of a display. Note that in this embodiment, an example where the HMD 51 is connected to the audio signal processing device 1 is shown, but a display such as an LCD may be connected to the audio signal processing device 1 instead.

[0011] The user I / F 32 is a keyboard, a mouse, or a touch panel laminated on a display. When the user I / F 32 is a touch panel, the user I / F 32 constitutes a GUI (Graphical User Interface) together with the display.

[0012] The communication I / F 36 is connected to a network such as a LAN or the Internet.

[0013] The USB I / F 37 is connected to the audio I / F 71. The audio I / F 71 has a USB terminal and an audio terminal. The audio I / F 71 is connected to an audio device via an audio cable. In this embodiment, the audio I / F 71 has a headphone 81 and a guitar amplifier 82 connected thereto. A guitar 83 is connected to the guitar amplifier 82. The guitar amplifier 82 inputs an audio signal from the guitar 83.

[0014] The guitar amplifier 82 inputs an audio signal related to the performance sound of the guitar 83 to the audio I / F 71. The audio I / F 71 outputs an audio signal to the headphone 81. Note that the audio I / F 71 may be connected to the audio signal processing device 1 via a LAN. In this case, the USB I / F 37 may be connected to the audio I / F 71 only for power supply.

[0015] The CPU 34 is a general-purpose processor. The CPU 34 reads out a program stored in the flash memory 33 which is a storage medium to the RAM 35 and controls each component of the audio signal processing device 1.

[0016] DSP38 is a processor dedicated to signal processing. DSP38 decodes audio data received from other devices via the communication interface 36. The audio data transmitted by other devices includes sound signals from sound sources located in a virtual space and location information of those sound sources. DSP38 performs signal processing on the decoded sound signals and the sound signals input from the audio interface 71. Signal processing may also be performed by the CPU 34. Furthermore, signal processing does not need to be performed by the processor inside the sound signal processing device 1; it may be performed by the processor of another device connected via, for example, the USB interface 37 or the communication interface 36.

[0017] The DSP38 performs binaural processing by convolving, for example, an HRTF (Head Related Transfer Function) into the sound signal. The HRTF represents the transfer function from a virtual sound source location to the user's right and left ears. This allows the user to perceive the sound as if it were emanating from a virtual sound source location. The DSP38 may also perform reverb processing to correspond to the reverberation of the virtual space. The reverb processing may have parameters corresponding to the room size of the virtual space.

[0018] Figure 4 is a plan view showing the configuration of a virtual space 101. Figure 5 is a flowchart showing the operation of the sound signal processing device 1. The first performer A1, a user of the sound signal processing device 1, is located at a first location and uses the sound signal processing device 1, which is the first terminal, to conduct a session with the second performer A2, who is located at a second location. As an example of a performance, the first performer A1 plays the guitar 83, and the second performer A2 sings. The sound signal processing device 1A has the same configuration as the sound signal processing device 1 shown in Figure 2. However, the sound signal processing device 1A is connected to the microphone 90 via the audio I / F 71, and is not connected to the guitar amplifier 82 and guitar 83.

[0019] The sound signal processing device 1 receives, for example via a GUI, the acoustic space information of the virtual space 101, the first position information of the first performer A1, and the second position information of the second performer A2 from the first performer A1 (S11). The virtual space information includes at least information relating to the size and shape of the virtual space. The virtual space information may also include information indicating the material of the wall surfaces constituting the virtual space (information corresponding to reflectance, sound absorption coefficient, etc.). The virtual space information, the first position information of the first performer A1, and the second position information of the second performer A2 are represented by two-dimensional or three-dimensional coordinates with a certain position as the origin.

[0020] In the example shown in Figure 4, the first position information includes information indicating the coordinates of the first performer A1 at the center of the virtual space 101, and information indicating the coordinates of the virtual guitar amplifier 82V to the right rear of the virtual space 101 as seen from the first performer A1. The second position information includes information indicating the coordinates of the second performer A2 at the front center of the virtual space 101. Note that the acoustic space information, the first position information, or the second position information may be received from a second terminal via the network.

[0021] The sound signal processing device 1 receives a first sound signal related to the performance of the first performer A1 from the audio I / F 71 (S12). In the example shown in Figure 4, the sound signal processing device 1 receives the first sound signal related to the performance of the first performer A1, which is output from the guitar amplifier 82, from the audio I / F 71. The sound signal processing device 1 outputs the first sound signal to the second terminal of the second performer A2 via the network through the communication I / F 36 (S13). The sound signal processing device 1 also receives a second sound signal related to the singing of the second performer A2 via the network (S14). The singing sound of the second performer is picked up by a microphone 90 connected to the sound signal processing device 1A, which is the second terminal. The second sound signal related to the singing sound picked up by the microphone 90 is input to the sound signal processing device 1A via the audio I / F 71. The sound signal processing device 1A outputs the second sound signal to the sound signal processing device 1, which is the first terminal, via the network through the communication I / F 36. Note that the order in which S12 or S14 are performed is not limited; either process can be performed first.

[0022] The sound signal processing device 1 may also render CG images of the virtual space and objects such as performers to generate a video of the virtual space viewed in a predetermined direction from the viewpoint of the first performer A1. The sound signal processing device 1 may also output the rendered video to the HMD51 via the display I / F31.

[0023] The sound signal processing device 1 applies first binaural processing to the first sound signal based on first position information, and applies second binaural processing to the second sound signal based on second position information (S15). As first binaural processing, the sound signal processing device 1 convolves an HRTF into the first sound signal such that it is localized to a position to the right rear of the user (relative coordinates of the virtual guitar amplifier 82V with the coordinates of the first performer A1 as the origin). As second binaural processing, the sound signal processing device 1 convolves an HRTF into the second sound signal such that it is localized to a position in front of the user (relative coordinates of the second performer A2 with the coordinates of the first performer A1 as the origin).

[0024] The sound signal processing device 1 outputs the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing to the headphones 81 via the audio I / F 71 (S16). At this time, the sound signal processing device 1 outputs the first binaural signal and the second binaural signal as a mixed sound signal to the audio I / F 71. The mixing balance of the first binaural signal and the second binaural signal may be adjusted, for example, by the first performer A1, for example, via a GUI.

[0025] The sound signal processing device 1 may also perform reverb processing on the first binaural signal, the second binaural signal, or the mixed sound signal to correspond to the reverberation of the virtual space 101.

[0026] First performer A1, by listening to the first and second binaural signals through headphones 81, can perceive themselves as being in the center of the virtual space 101 and hearing the singing of second performer A2 coming from directly in front of them. Furthermore, first performer A1 can perceive their own performance sounds coming from a different location (to the right rear) than their own position in the virtual space 101.

[0027] In a performance in a real space, the first performer A1 hears the sound of their own performance emanating from sound equipment such as a guitar amplifier 82, or from an instrument such as a guitar 83. If the sound of their own performance, such as the sound of the guitar amplifier 82 or guitar 83, were localized inside the performer's head, it would cause discomfort. In contrast, the sound signal processing device 1 of this embodiment convolves the transmission characteristics so that the sound is heard from a designated position for each of the first and second performers in the virtual space. Therefore, the first performer A1, while in the virtual space 101, can perceive the sound of their own guitar 83 as coming from the guitar amplifier 82 located to their right and behind them, providing a new customer experience that allows for a realistic session in a virtual space, something that was not possible with conventional methods.

[0028] However, the position from which the sound of one's own performance is localized is not limited to the position of sound equipment or instruments installed in the actual space. The sound signal processing device 1 of this embodiment can also hear the sound of one's own performance from any position. For example, in the example shown in Figure 6, the sound signal processing device 1 receives an instruction from the first performer A1 via the GUI to change the coordinates of the virtual guitar amplifier 82V to the left rear of the first performer A1. As a first binaural processing, the sound signal processing device 1 convolves an HRTF into the first sound signal that localizes it to the left rear of the user. As a result, the first performer A1 can perceive that they are in the center of the virtual space 101 and are hearing the sound of their own performance from the left rear of themselves.

[0029] The operation of the sound signal processing device 1 described above may be performed not only at the first terminal used by the first presenter A1, but also at the second terminal (sound signal processing device 1A) used by the second presenter A2. Figure 7 is a flowchart showing the operation of the sound signal processing device 1A, which is the second terminal. Since the sound signal processing device 1A, which is the second terminal, has the same configuration as the sound signal processing device 1, the illustration and explanation of its configuration are omitted.

[0030] The second terminal, the sound signal processing device 1A, receives the specification of acoustic space information for the virtual space 101, the first position information of the first performer A1, and the second position information of the second performer A2 (S21). The acoustic space information and the first position information are received from the first terminal via the network. The second position information is received from the second performer A2, for example, via a GUI. In the example shown in Figure 6, the first position information includes information indicating the coordinates of the first performer A1 at the center of the virtual space 101, and information indicating the coordinates of the virtual guitar amplifier 82V to the left front of the virtual space 101 as seen from the second performer A2. The second position information includes information indicating the coordinates of the second performer A2 to the front center of the virtual space 101 as seen from the first performer A1. The second position information may also be received from the first terminal via the network.

[0031] The second terminal, the sound signal processing device 1A, receives the second sound signal related to the singing of the second performer A2 from the audio I / F 71 (S22). The second terminal, the sound signal processing device 1A, outputs the second sound signal to the first terminal of the first performer A1 via the network (S23). The second terminal, the sound signal processing device 1A, also receives the first sound signal related to the performance of the first performer A1 via the network (S24). Note that the order of processing in S22 or S24 is not limited and can be performed in any order.

[0032] Furthermore, the second terminal, the sound signal processing device 1A, may also render CG images of the virtual space and objects such as performers to generate video of the virtual space viewed in a predetermined direction from the viewpoint of the second performer A2. The second terminal, the sound signal processing device 1A, may also output the rendered video to the HMD51 via the display I / F31.

[0033] The second terminal, the sound signal processing device 1A, applies first binaural processing to the first sound signal based on first position information, and applies second binaural processing to the second sound signal based on second position information (S25). As first binaural processing, the second terminal, the sound signal processing device 1A, convolves an HRTF into the first sound signal that localizes it to the left front of the user. As second binaural processing, the second terminal, the sound signal processing device 1A, convolves an HRTF into the second sound signal that localizes it to the user's position.

[0034] The second terminal, the sound signal processing device 1A, outputs the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing from the audio I / F 71 (S26). At this time, the second terminal, the sound signal processing device 1A, outputs the first binaural signal and the second binaural signal as a mixed sound signal to the audio I / F 71. The level balance between the first binaural signal and the second binaural signal may be adjusted by the second performer A2, for example, via a GUI. The level balance may be the same as or different from the level balance of the first terminal.

[0035] Furthermore, the second terminal, the sound signal processing device 1A, may perform reverb processing on the first binaural signal, the second binaural signal, or the mixed sound signal to correspond to the reverberation of the virtual space 101.

[0036] As a result, the second performer A2 can perceive themselves as being in the center of virtual space 101 and hearing the performance sounds of the first performer A1 from their left front. Therefore, the second performer A2 can perceive themselves as being in the same space as the first performer A1 and conducting the session.

[0037] Furthermore, the second position information may include not only the coordinates of the second performer A2 but also the coordinates of sound equipment, etc. For example, in the example in Figure 8, the second position information includes information indicating the coordinates of the second performer A2 at the front center of the virtual space 101 as seen from the perspective of the first performer A1, and information indicating the coordinates of the virtual monitor speaker 90V at the front of the virtual space 101 as seen from the perspective of the second performer A2. The second terminal, as a second binaural processing, convolves an HRTF into the second sound signal such that it is localized at the position of the virtual monitor speaker 90V. In this case, the second performer A2 can perceive that their own singing sound is coming from a different position (front) than their own position in the virtual space 101.

[0038] In a session in a real-world space, the singer may hear their own singing sound emanating from a monitor speaker. Therefore, if their own singing sound is localized inside their head, it may cause discomfort for the singer. In contrast, the second terminal of this embodiment convolves the transmission characteristics so that the sound is heard from a designated position for each of the first and second performers in the virtual space. As a result, the second performer A2, while in the virtual space 101, can perceive their own singing sound as emanating from a virtual monitor speaker 90V located in front of them, providing a new customer experience that enables a realistic session in a virtual space, something that was not possible with conventional methods.

[0039] The first and second position information within the virtual space 101 may be the same or different for the first and second terminals. For example, the second position information at the first terminal may include information indicating the coordinates of the second performer A2, and the second position information at the second terminal may include information indicating the coordinates of the virtual monitor speaker 90V located in front of the virtual space 101 as seen from the perspective of the second performer A2, as shown in Figure 8. In this case, the first performer A1 hears the singing sound from the position of the second performer A2, and the second performer A2 hears their own singing sound from a different position (in front) than their own position within the virtual space 101. In this case, the sound signal processing device 1 can also realize a session unique to the virtual space that cannot be realized in a real space, by having each performer set their own localization position.

[0040] Furthermore, the acoustic space information may be the same or different for the first and second terminals. In this case as well, the sound signal processing device 1 can enable each performer to set up their own virtual space, thereby realizing a session unique to the virtual space that cannot be achieved in a real space.

[0041] Furthermore, the sound signal processing device 1 may store or learn information about past sound settings received from the first performer A1 or the second performer A2, and automatically set them. The sound settings include the level balance between the first binaural signal and the second binaural signal, gain settings, reverb settings, or equalizer settings.

[0042] As described above, the sound signal processing device 1 of this embodiment can enable each performer to create their own unique sound settings, thereby realizing sessions that are unique to virtual spaces and cannot be achieved in real-world spaces.

[0043] (Variation 1) In the modified example 1, the first binaural processing includes the process of localizing a first indirect sound to a first sound signal based on first position information, and the second binaural processing includes the process of localizing a second indirect sound to a second sound signal based on the first and second position information. Indirect sound refers to sound that reaches the listening position after the sound from the sound source reflects off the walls of the virtual space, and includes early reflections with a fixed direction and phase, or reverberations with random direction and phase. Binaural processing of indirect sound localization is performed only on the early reflection system, and localization processing may be performed using effects such as reverb in the reverberation processing system. The sound signal processing device 1 uses the coordinates of the first performer A1 in the virtual space 101, which are included in the first position information, as the listening position, and the coordinates of the virtual guitar amplifier 82V as the sound source position, to acquire an impulse response corresponding to the reverberation of the virtual space 101. The impulse response is acquired by measurement, for example, by placing a dummy head at the listening position in the real space corresponding to the virtual space 101. The sound signal processing device 1 performs a process to localize the first indirect sound by convolving the acquired impulse response into the first sound signal.

[0044] Alternatively, the impulse response may be obtained by simulation based on, for example, the ray method or the virtual image method. The ray method is a technique that tracks the trajectory (ray) of sound radiated from a sound source and calculates the time pattern of the energy of the ray passing through the listening position. A simulation using the ray method determines the direction, arrival time, and arrival level from each virtual sound source at the listening position, based on the energy of the ray in the listening region, assuming that each ray is a virtual sound image of reverberation. The virtual image method is a technique that creates a virtual image (virtual sound source) of the sound source against the wall surface of the space as a virtual sound source, and determines the direction, arrival time, and arrival level from each virtual sound source at the listening position. The sound signal processing device 1 may generate an impulse response of a head-related transfer function that represents the direction, arrival time, and arrival level of each virtual sound source obtained by simulation, and perform a process to localize the first indirect sound by convolving the impulse response into the first sound signal.

[0045] Alternatively, the sound signal processing device 1 may perform a process to localize the first indirect sound by applying a level delay filter to the first sound signal, which has a delay amount and attenuation amount corresponding to each virtual sound source determined by simulation.

[0046] Similarly, the sound signal processing device 1 uses the coordinates of the first performer A1 in the virtual space 101 included in the first position information as the listening position, and the coordinates of the second performer A2 included in the second position information as the sound source position, to obtain an impulse response corresponding to the reverberation of the virtual space 101. As described above, the impulse response may be obtained by measurement or by simulation. The sound signal processing device 1 performs a process to localize the second indirect sound by convolving the obtained impulse response into the second sound signal. Alternatively, the sound signal processing device 1 may perform a process to localize the second indirect sound by applying a level delay filter process to the second sound signal, which has delay and attenuation amounts corresponding to each virtual sound source obtained by simulation.

[0047] This allows the first performer A1 to perceive not only direct sounds but also indirect sounds that occur in a real space, enabling them to have a more realistic session with the second performer A2 while remaining in the virtual space 101.

[0048] The operation of the sound signal processing device 1 in Modification 1 may be performed not only in the first terminal used by the first presenter A1, but also in the second terminal, the sound signal processing device 1A, used by the second presenter A2.

[0049] (Modification 2) The first position information may include the first orientation information of the first performer, and the second position information may include the second orientation information of the second performer. In this case, the first binaural processing further includes localization processing based on the first orientation information, and the second binaural processing further includes localization processing based on the first and second orientation information.

[0050] The first orientation information includes the direction information of the first performer A1 and the direction information of the virtual guitar amplifier 82V. The second orientation information includes the direction information of the second performer A2. The sound signal processing device 1 convolves the impulse response of the head-related transfer function corresponding to the direction of the virtual guitar amplifier 82V relative to the direction of the first performer A1 into the first sound signal. The sound signal processing device 1 convolves the impulse response of the head-related transfer function corresponding to the direction of the second performer A2 relative to the direction of the first performer A1 into the second sound signal.

[0051] The sound levels of performances and vocals are highest when the front of the sound source and the front of the listener are facing each other, and decrease as the left-right angle increases. Also, as the left-right angle increases, the high frequencies are attenuated more than the low frequencies. Therefore, the sound signal processing device 1 may perform gain correction such that the sound signal level decreases as the difference (angle difference) between the direction of the first performer A1 and the direction of the second performer A2 increases. Alternatively, the sound signal processing device 1 may perform equalizer processing such that the high-frequency level decreases or the low-frequency level increases as the difference (angle difference) between the direction of the first performer A1 and the direction of the second performer A2 increases. The sound signal processing device 1 may also perform gain correction such that the sound signal level decreases as the difference (angle difference) between the direction of the first performer A1 and the direction of the virtual guitar amplifier 82V increases. Furthermore, the sound signal processing device 1 may perform equalizer processing such that the higher the difference (angle difference) between the orientation of the first performer A1 and the orientation of the virtual guitar amplifier 82V, the lower the level of high frequencies or the higher the level of low frequencies.

[0052] This allows the sound signal processing device 1 to perceive changes in the relative positions of the performer and the sound source in real time. The sound signal processing device 1 can represent the relative positions of the performer and the sound source with greater precision. Therefore, the first performer A1 can be in the virtual space 101 and conduct a more realistic session with the second performer A2.

[0053] The operation of the sound signal processing device 1 in Modification 2 may be performed not only on the first terminal used by the first presenter A1, but also on the second terminal, the sound signal processing device 1A, used by the second presenter A2.

[0054] (Variation 3) The sound signal processing device 1 may receive first sensing information, which senses the performance of the first performer, from the sensor interface, and second sensing information, which senses the performance of the second performer, via the network.

[0055] The first sensing information and the second sensing information are acquired by sensors, respectively. For example, the sound signal processing device 1 estimates information such as the position, orientation, and motion of the first performer A1 based on an image of the first performer A1 taken by an image sensor such as a camera, and acquires the first sensing information. Alternatively, the sound signal processing device 1 may estimate information such as the position, orientation, and motion of the first performer A1 based on sensors such as a position sensor, motion sensor, or head tracker worn by the first performer A1, and acquire the first sensing information.

[0056] The sound signal processing device 1 synchronizes and plays back video related to the performances of the first and second performers based on the first and second sensing information. Specifically, the sound signal processing device 1 operates 3D models of the first performer A1 and the second performer A2 based on information such as their respective positions, orientations, and motions.

[0057] This allows the first presenter, A1, to be in the virtual space 101 and visually perceive the performance of the second presenter, A2, enabling them to conduct a realistic session.

[0058] The operation of the sound signal processing device 1 in Modification 3 may be performed not only in the first terminal used by the first presenter A1, but also in the second terminal, the sound signal processing device 1A, used by the second presenter A2.

[0059] (Modification 4) The sound signal processing device 1 according to Modification 4 corrects the sound environment at the first location. For example, the sound signal processing device 1 performs processing to reduce the reverberation of the actual space in which the first performer A1 is located. Information on the reverberation of the actual space in which the first performer A1 is located is obtained in advance by measuring the impulse response. In the first binaural processing and the second binaural processing, the sound signal processing device 1 reduces the reverberation of the actual space in which the first performer A1 is located by convolving the inverse characteristics of the impulse response measured in advance.

[0060] Furthermore, for example, the sound signal processing device 1 may perform a process to cancel out noise in the actual space where the first performer A1 is located. Also, the sound signal processing device 1 may perform a process to cancel out the sound of the guitar that is generated in the actual space where the first performer A1 is located.

[0061] This allows the first presenter, A1, to conduct the session with a greater sense of immersion, perceiving themselves as being in virtual space 101.

[0062] The operation of the sound signal processing device 1 in Modification 4 may be performed not only in the first terminal used by the first presenter A1, but also in the second terminal, the sound signal processing device 1A, used by the second presenter A2.

[0063] (Variation 5) The sound signal processing device 1 according to Modification 5 further overlays and displays information used by the performer on the rendered image. The sound signal processing device 1 overlays and displays information such as musical scores, lyrics, and dialogue. The sound signal processing device 1 may also estimate the current performance position in the session based on the musical score information and the first or second sound signal, and display the current performance position on the musical score information. Furthermore, the sound signal processing device 1 may display information indicating the timing of chord progressions and rhythm changes in the next measure based on the estimated performance position.

[0064] (Experimental variation 6) Figure 9 is a block diagram showing the configuration of the sound signal processing system according to Modification 6. The sound signal processing system according to Modification 6 further includes multiple viewer terminals connected via a network.

[0065] Figure 10 is a block diagram showing the configuration of an audio signal processing device 1B, which is an example of a viewer terminal. Components common to the audio signal processing device 1 in Figure 2 are denoted by the same reference numerals and their descriptions are omitted. The audio signal processing device 1B differs from the audio signal processing device 1 in that it does not have a DSP 38. Other components are the same as those of the audio signal processing device 1. However, only headphones 81 are connected to the audio I / F 71.

[0066] Figure 11 is a plan view showing the configuration of the virtual space 101. Figure 12 is a flowchart showing the operation of the sound signal processing device 1B. Viewer L1, who is a user of the sound signal processing device 1B, is located at a third location and watches the performances of the first presenter A1 and the second presenter A2, such as their sessions.

[0067] The audio signal processing device 1B, which is the viewer terminal, receives acoustic space information, first position information, and second position information (S31). For example, the audio signal processing device 1B receives acoustic space information, first position information, and second position information from the first terminal and the second terminal via the network. Furthermore, the audio signal processing device 1B receives the viewer's third position information (S32). As shown in Figure 11, the third position information includes information indicating the coordinates of viewer L1 at the left edge center of the virtual space 101.

[0068] Next, the sound signal processing device 1B receives the first sound signal and the second sound signal via the network (S33). In other words, the first terminal distributes the first sound signal, and the second terminal distributes the second sound signal. Note that the processing in S31 to S33 can be performed in any order, and the order of processing is not limited.

[0069] The sound signal processing device 1B applies first binaural processing to the first sound signal based on first position information and third position information, and applies second binaural processing to the second sound signal based on second position information and third position information (S34). As first binaural processing, the sound signal processing device 1B convolves an HRTF into the first sound signal such that it is localized to a position to the right front of the viewer (relative coordinates of the virtual guitar amplifier 82V with the coordinates of viewer L1 as the origin). As second binaural processing, the sound signal processing device 1B convolves an HRTF into the second sound signal such that it is localized to a position to the left front of the user (relative coordinates of the second performer A2 with the coordinates of viewer L1 as the origin).

[0070] In this modified example 7, the sound signal processing device 1B performs binaural processing by software processing of the CPU 34 instead of a DSP. However, in the present invention, the listener terminal may perform binaural processing using a DSP.

[0071] The sound signal processing device 1B then outputs the first binaural signal and the second binaural signal to headphones 81, which is an example of the listener's audio equipment (S35). At this time, the sound signal processing device 1B outputs the mixed sound signal of the first binaural signal and the second binaural signal to the audio I / F 71. The mixing balance of the first binaural signal and the second binaural signal may be adjusted by the listener L1, for example, via a GUI.

[0072] Furthermore, the sound signal processing device 1B may render CG images of the virtual space and objects such as performers to generate video of the virtual space viewed in a predetermined direction from the viewpoint position of the viewer L1. The sound signal processing device 1 may also output the rendered video to the HMD51 via the display I / F31.

[0073] Through the above actions, viewer L1 can perceive that they are in the virtual space 101 and that the sound of performer A1's guitar 83 is coming from the guitar amplifier 82, providing a new customer experience that allows them to view a realistic performance in a virtual space, something that was not possible before.

[0074] (Example 7) In the modified example 7, the sound signal processing device 1 relating to the first terminal or the sound signal processing device 1A relating to the second terminal receives reaction sound signals relating to the viewer's reactions from the viewer terminal via the network. The sound signal processing device 1 applies a third binaural processing to the reaction sound signals.

[0075] The reaction sound signal is, for example, the viewer's voice or sound effect such as applause, acquired by a microphone (not shown) connected to the sound signal processing device 1B, which is a viewer terminal. Alternatively, the sound signal processing device 1B displays icon images such as "cheers," "applause," "calls," and "murmuring" on a display. The sound signal processing device 1B receives reaction information (text information) such as "cheers," "applause," "calls," and "murmuring" from the user by accepting selection operations for these icon images via the user I / F 32. The sound signal processing device 1B transmits the received reaction information to the sound signal processing device 1. The sound signal processing device 1 generates a reaction sound signal corresponding to the received reaction information. Alternatively, the sound signal processing device 1B may generate a reaction sound signal corresponding to the received reaction. When sending and receiving reaction information that contains less information than the reaction sound signal, the sound signal processing system can reduce the communication load.

[0076] The sound signal processing device 1B transmits the viewer's reaction sound signal to the sound signal processing device 1. The sound signal processing device 1B also transmits the viewer's third location information to the sound signal processing device 1.

[0077] The sound signal processing unit 1 receives a reaction sound signal and third position information via the network. Based on the third position information, the sound signal processing unit 1 applies third binaural processing to the reaction sound signal. As third binaural processing, the sound signal processing unit 1 convolves an HRTF into the reaction sound signal such that it is localized to the left position of the first performer A1 (relative coordinates of viewer L1 with the coordinates of the first performer A1 as the origin). At this time, the sound signal processing unit 1 may perform binaural processing by software processing of the CPU 34 instead of the DSP 38. In other words, the first and second binaural processing may be performed by the DSP 28, and the third binaural processing may be performed by software processing of the CPU 34. By doing so, the sound signal processing unit 1 can prioritize binaural processing of the first and second sound signals, which are related to performance, using the DSP 38, which has relatively low latency, and reduce the resources allocated to the third binaural processing, thereby minimizing the impact of latency on performance.

[0078] In this way, the first performer A1 or the second performer A2 can perform as if viewer L1 were also in the same virtual space 101. Furthermore, the first performer A1 or the second performer A2 can perform while referring to viewer L1's reactions.

[0079] (Variation 8) In the modified example 8, the sound signal processing device 1 receives third sensing information, which is a sensed viewer reaction, via the network. Based on the third sensing information, the sound signal processing device 1 plays back the video related to the viewer's reaction.

[0080] The sound signal processing device 1B estimates information such as the position, orientation, and motion of viewer L1 based on an image of viewer L1 captured by an image sensor such as a camera, and acquires third sensing information. Alternatively, the sound signal processing device 1B may estimate information such as the position, orientation, and motion of viewer L1 based on sensors such as a position sensor and a motion sensor worn by viewer L1, and acquire third sensing information.

[0081] The sound signal processing device 1 operates a 3D model of the viewer L1, including its position, orientation, and motion, based on the third sensing information.

[0082] This allows the first performer A1 or the second performer A2 to be in the virtual space 101 and visually perceive the reactions of the viewer L1, enabling them to perform more realistic sessions and other performances.

[0083] (Extreme variation 9) Figure 13 is a block diagram showing the configuration of the sound signal processing system according to Modification 9. The sound signal processing system according to Modification 9 has an administrator terminal in addition to the configuration shown in Figure 10. The administrator terminal is connected to the first terminal, the second terminal, and the viewer terminal via a network.

[0084] In the sound signal processing system according to Modification 9, the administrator terminal distributes acoustic space information, first location information, and second location information to the viewer terminal. In addition, in the sound signal processing system according to Modification 9, the first terminal transmits a first sound signal to the administrator terminal, and the second terminal transmits a second sound signal to the administrator terminal.

[0085] The administrator terminal performs the first binaural processing on the first audio signal received from the first terminal, and the second binaural processing on the second audio signal received from the second terminal. The administrator terminal then distributes the first binaural signal and the second binaural signal to the viewer terminal.

[0086] In this case, the audio signal delivered to the viewer's terminal undergoes high-processing binaural processing on the administrator terminal. Therefore, the administrator terminal can provide a realistic performance experience in the virtual space, regardless of the processing power of the viewer's terminal.

[0087] Figure 14 is a block diagram showing the configuration of an audio signal processing device 1C, which is an example of an administrator terminal. The configuration of the audio signal processing device 1C is the same as the configuration of the audio signal processing device 1 shown in Figure 2. However, headphones 81 and a microphone 90 are connected to the audio I / F 71, and an LCD 501 is connected to the display I / F 31.

[0088] Microphone 90 acquires audio signals related to the voice of the administrator of the sound signal processing system. The sound signal processing device 1C transmits the audio signals related to the administrator's voice acquired by microphone 90 to the terminals of other users, such as the first terminal or the second terminal. Microphone 90 may be a microphone built into the headphones 81.

[0089] The administrator terminal may accept the specification of the destination for the audio signal related to the administrator's voice acquired by microphone 90. The destination specification may include individual specification of the first terminal or the second terminal, or a specification of all terminals including the first and second terminals.

[0090] This allows administrators to provide talkback to any performer or all performers, enabling them to manage performance and provide services in the virtual space while communicating with the performers.

[0091] (Variation 10) Figure 15 is a block diagram showing the configuration of the sound signal processing device 1 according to modified example 10. The sound signal processing device 1 shown in Figure 15 has the same configuration as shown in Figure 2. However, the sound signal processing device 1 is connected to the microphone 90 via the audio I / F 71. The microphone 90 acquires sound signals related to the voice of the first presenter A1. The sound signal processing device 1 transmits the sound signals related to the voice of the first presenter A1 acquired by the microphone 90 to the terminals of other users, such as the second terminal. Note that the microphone 90 may be a microphone built into the headphones 81.

[0092] The audio signal processing unit 1 of the first terminal receives audio signals related to the voice of other users from the second terminal or other user terminals such as viewer terminals via the network. The audio signal processing unit 1 receives specifications for sound processing conditions, including the localization position (conditions related to the fourth binaural processing) of the audio signals related to the voice of the other users. Based on the received conditions, the audio signal processing unit 1 applies sound processing, including the fourth binaural processing, to the audio signal and outputs the processed audio signal from the audio I / F 71.

[0093] For example, in the example in Figure 8, the second binaural processing convolves an HRTF that localizes the virtual monitor speaker 90V to its relative coordinates, with the coordinates of the first performer A1 as the origin. The fourth binaural processing convolves an HRTF that localizes the second performer A2 to its relative coordinates, with the coordinates of the first performer A1 as the origin.

[0094] Furthermore, the sound processing conditions may include, for example, level balance conditions. Level balance includes the level balance of the first binaural signal, the second binaural signal, the third binaural signal, and the above-mentioned audio signal.

[0095] Alternatively, the sound processing conditions may be accepted via a GUI, or automatically according to the orientation (gaze) of each user detected by various sensors such as motion sensors. The sound signal processing device 1 may perform processing to maximize the level of the sound corresponding to other performers or viewers located in the direction that the first performer A1 is facing, or to emphasize that sound.

[0096] In real space, people unconsciously concentrate on listening to sounds coming from the direction they are facing (sounds they are paying attention to), creating a cocktail party effect. However, it is difficult to create a cocktail party effect in virtual space. In contrast, the sound signal processing device 1 according to Modification 10 can achieve a cocktail party effect in virtual space by making the level of the sound corresponding to other performers or viewers located in the direction that the first performer A1 is facing the highest, or by processing to emphasize that sound. This makes it possible to communicate among any participants, including performers or viewers.

[0097] (Variation 11) The sound signal processing system may perform sound processing individually at each terminal, as shown in Figure 9, or it may perform some sound processing at a server such as an administrator terminal, as shown in Figure 13, and then distribute the processed sound signal to the viewer terminal. Alternatively, the sound signal processing system may perform some sound processing at an edge server and then distribute the processed sound signal to the viewer terminal. Figure 16 is a block diagram showing the configuration of the sound signal processing system according to Modification 11. The sound signal processing system according to Modification 11 has an edge server 1D in addition to the configuration shown in Figure 13. The edge server 1D is connected to the first terminal, the second terminal, the administrator terminal, and the viewer terminal via a network. The configuration of the edge server 1D is the same as that of the sound signal processing device 1. However, the edge server 1D does not need to have a display I / F 31 and a USB I / F 37.

[0098] In the sound signal processing system according to Modification 11, the edge server 1D performs some or all of the sound processing. For example, the administrator terminal transmits acoustic space information, first location information, and second location information to the edge server. The first terminal also transmits a first sound signal to the edge server 1D, and the second terminal transmits a second sound signal to the edge server 1D.

[0099] Edge server 1D performs first binaural processing on the first audio signal received from the first terminal and second binaural processing on the second audio signal received from the second terminal. Edge server 1D distributes the first binaural signal and the second binaural signal to the viewer terminal.

[0100] In this case, the computationally intensive binaural processing is performed on edge server 1D, which is closer to the viewer's device. As a result, viewers can enjoy a more lag-free viewing experience.

[0101] Next, Figure 17 is a conceptual diagram illustrating the configuration of an audio signal processing system when providing individual environments. The positions of the performers and viewers within the virtual space 101 may be the same or different for all terminals. For example, in the example in Figure 17, at the first point of the first performer A1, the first performer A1 is located at the center of the virtual space 101, the virtual guitar amplifier 82V is located to the right rear of the first performer A1, the second performer A2 is located in the front center, and the viewer L1 is located to the left. At the second point of the second performer A2, the second performer A2 is located at the center of the virtual space 101, the virtual monitor speaker 90V is located in front of the second performer A2, the first performer A1 is located behind it, the virtual guitar amplifier 82V is located to the left front, and the viewer L1 is located to the right. At viewer L1's third location, viewer L1 is at the center of virtual space 101, virtual right main speaker 90RV is located to the right and in front of viewer L1, virtual left main speaker 90LV is located to the left and in front of viewer L1, virtual guitar amplifier 82V is located to the right of virtual left main speaker 90LV, the first performer is located in virtual guitar amplifier 82V, and the second performer A2 is located to the left of virtual right main speaker 90RV.

[0102] In this case, viewer L1 sees the first performer A1 and the second performer A2 in front of them, with the sound of the first performer's guitar playing localized at the position of the virtual guitar amplifier 82V, and the sound of the second performer A2's singing localized by panning using the virtual right main speaker 90RV and the virtual left main speaker 90LV.

[0103] The audio settings for viewer L1, such as localization position and level balance, may be specified by viewer L1 or by the administrator of the administrator terminal. Alternatively, the audio settings for each viewer may be automatically configured by the viewer terminal or administrator terminal by storing or learning past information.

[0104] In other words, the sound signal processing system (sound signal processing device or sound signal processing method) in the example shown in Figure 17 has the following technical concept. The first terminal receives acoustic space information, first position information of the first performer in the acoustic space, and second position information of the second performer in the acoustic space. In the first terminal, the first sound signal relating to the performance of the first performer is received from the audio interface. The first terminal outputs the first sound signal to the second terminal of the second performer via the network, and the first terminal receives the second sound signal related to the second performer's performance via the network. In the first terminal, first binaural processing is applied to the first sound signal based on the first position information, and second binaural processing is applied to the second sound signal based on the second position information. In the first terminal, the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing are output from the audio interface. The second terminal receives the acoustic space information, the first location information, and the second location information, In the second terminal, the aforementioned two audio signals are received from the audio interface. In the second terminal, the second sound signal is output to the first terminal via the network, and the first sound signal is received via the network. In the second terminal, the first binaural processing and the second binaural processing are performed. In the second terminal, the first binaural signal and the second binaural signal are output from the audio interface. A method for processing sound signals, The acoustic processing parameters for the first binaural processing and the second binaural processing in the first terminal are different from those for the first binaural processing and the second binaural processing in the second terminal.

[0105] Furthermore, the sound signal processing system (sound signal processing device or sound signal processing method) may also have the following technical concepts. The viewer terminal receives the acoustic space information, the first location information, and the second location information, In the viewer terminal, the first sound signal and the second sound signal are distributed to the viewer terminal via the network. The viewer terminal performs the first binaural processing and the second binaural processing, In the viewer terminal, the first binaural signal and the second binaural signal are output to the viewer's audio equipment. The acoustic processing parameters differ between the first binaural processing and the second binaural processing in the first terminal, the first binaural processing and the second binaural processing in the second terminal, and the first binaural processing and the second binaural processing in the viewer terminal.

[0106] By performing individual sound settings for each performer and viewer as described above, it becomes possible to realize sessions and performances that are unique to virtual spaces and cannot be achieved in real-world spaces.

[0107] In the above embodiment, binaural processing was shown as an example of localization processing, but localization processing is not limited to binaural processing. For example, panning processing and beamforming processing using multiple speakers are also examples of localization processing. Furthermore, the sound signal after binaural processing may be played back not only through headphones, but also through multiple speakers. When playing back the sound signal after binaural processing through multiple speakers, it is preferable to perform crosstalk cancellation processing.

[0108] In the above embodiment, a session by a first performer A1 and a second performer A2 was used as an example, but performances via a network at different locations are not limited to this example. The sound signal processing system may, for example, process the sound of a performance by even more performers. The sound signal processing system may also process the sound of a group performance, such as an orchestra, taking place in a real space, and a performance by another performer in a remote location. In this case as well, the sound signal processing system processes the sound of the other performer's performance to localize it to a predetermined position, so that a viewer in a remote location can perceive that a group performance, such as an orchestra, taking place in a real space is simultaneously being performed by an even more remote performer.

[0109] In the above embodiment, an example was shown in which a general-purpose information processing device (sound signal processing device 1) performs the first binaural processing and the second binaural processing. However, if, for example, an audio I / F 71 is connected to a network, the audio I / F 71 may perform the first binaural processing and the second binaural processing. In other words, the sound signal processing device of the present invention can also be realized by an audio I / F that can be connected to a network.

[0110] The description of this embodiment should be considered in all respects to be illustrative and not restrictive. The scope of the invention is indicated by the claims, rather than by the embodiments described above. Furthermore, the scope of the invention includes the scope equivalent to the claims. [Explanation of Symbols]

[0111] 1,1A,1B,1C: Audio signal processing unit, 1D: Edge server, 31: Display I / F, 32: User I / F, 33: Flash memory, 34: CPU, 35: RAM, 36: Communication I / F, 37: USB I / F, 51: HMD, 71: Audio I / F, 81: Headphones, 82: Guitar amplifier, 82V: Virtual guitar amplifier, 83: Guitar, 90: Microphone, 90LV: Virtual left main speaker, 90RV: Virtual right main speaker, 90V: Virtual monitor speaker, 101: Virtual space, 501: LCD

Claims

1. The first terminal receives acoustic space information, first position information of the first performer in the acoustic space, and second position information of the second performer in the acoustic space. In the first terminal, a first sound signal relating to the performance of the first performer is received from the audio interface. The first terminal outputs the first sound signal to the second performer's second terminal via the network, and the first terminal receives the second sound signal related to the second performer's performance via the network. In the first terminal, first binaural processing is applied to the first sound signal based on the first position information, and second binaural processing is applied to the second sound signal based on the second position information. In the first terminal, the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing are output from the audio interface. Audio signal processing method.

2. The second terminal receives the acoustic space information, the first location information, and the second location information, In the second terminal, the second audio signal is received from the audio interface, The second terminal outputs the second sound signal to the first terminal via the network, and receives the first sound signal via the network. In the second terminal, the first binaural processing and the second binaural processing are performed. In the second terminal, the first binaural signal and the second binaural signal are output from the audio interface. The sound signal processing method according to claim 1.

3. The first binaural processing includes processing to localize the first indirect sound to the first sound signal based on the first position information, The second binaural processing includes processing to localize the second indirect sound to the second sound signal based on the first position information and the second position information. The sound signal processing method according to claim 1 or claim 2.

4. The first position information includes the first orientation information of the first performer, The second position information includes the second orientation information of the second performer, The first binaural processing further includes localization processing based on the first orientation information, The second binaural processing further includes localization processing based on the first orientation information and the second orientation information. The sound signal processing method according to claim 1 or claim 2.

5. The first terminal receives first sensing information, which senses the performance of the first performer, from the sensor interface. The first terminal receives second sensing information, which is obtained by sensing the performance of the second performer, via the network. In the first terminal, based on the first sensing information and the second sensing information, the video related to the performances of the first performer and the second performer is played back in sync. The sound signal processing method according to claim 1 or claim 2.

6. The viewer terminal receives the acoustic space information, the first position information, and the second position information, The first sound signal and the second sound signal are distributed to the viewer terminal via the network. The viewer terminal performs the first binaural processing and the second binaural processing, The viewer terminal outputs the first binaural signal and the second binaural signal to the viewer's audio equipment. The sound signal processing method according to claim 1 or claim 2.

7. The first terminal receives a reaction sound signal related to the viewer's reaction from the viewer terminal via the network. In the first terminal, a third binaural processing is applied to the reaction sound signal. The sound signal processing method according to claim 6.

8. The first terminal receives the third sensing information, which is the viewer's reaction, via the network. In the first terminal, based on the third sensing information, video related to the viewer's reaction is played back. The sound signal processing method according to claim 7.

9. The administrator terminal distributes the acoustic space information, the first location information, and the second location information. The first sound signal and the second sound signal are distributed to the administrator terminal via the network. The administrator terminal performs the first binaural processing and the second binaural processing, The administrator terminal outputs the first binaural signal and the second binaural signal to the administrator's audio equipment. The sound signal processing method according to claim 1 or claim 2.

10. The first terminal receives an audio signal related to the voice of another user from another user's terminal via the network, The first terminal receives a specification of sound processing conditions including the fourth binaural processing, In the first terminal, the sound processing is applied to the audio signal based on the conditions described above. In the first terminal, the audio signal after sound processing is output from the audio interface. The sound signal processing method according to claim 1 or claim 2.

11. The system receives acoustic space information, first position information of the first performer in the acoustic space, and second position information of the second performer in the acoustic space. The first sound signal related to the performance of the first performer is received from the audio interface, The first sound signal is output to the second terminal of the second performer via the network, and the second sound signal related to the second performer's performance is received via the network. Based on the first position information, the first sound signal is subjected to first binaural processing, and based on the second position information, the second sound signal is subjected to second binaural processing. The first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing are output from the audio interface. A sound signal processing device equipped with a processor.

Citation Information

Patent Citations

  • Information processing system, information processing method, and program

    WO2022196073A1