Sound signal processing method and sound signal processing device
The sound signal processing method allows performers to hear their own sound from a designated position in a virtual space by applying binaural processing based on positional information, enhancing the realism and immersion of virtual performances.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2026-03-26
AI Technical Summary
Existing sound signal processing methods fail to make the sound of one's own performance audible from a predetermined position in a virtual space, leading to an unrealistic and uncomfortable user experience.
A sound signal processing method that involves receiving acoustic space information and performer positions, applying binaural processing to sound signals based on these positions, and outputting the processed signals to headphones, allowing the sound of one's own performance to be localized to a designated position in the virtual space.
Enables each performer to perceive their own sound from a specific location in the virtual space, creating a realistic and immersive performance experience that was not possible with conventional methods.
Smart Images

Figure JP2025029648_26032026_PF_FP_ABST
Abstract
Description
Sound signal processing method and sound signal processing apparatus
[0001] One embodiment of the present invention relates to a sound signal processing method and a sound signal processing apparatus.
[0002] In Patent Document 1, for an acoustic signal obtained by collecting sound in a space where each of a plurality of co-performing users is located, acoustic processing is performed that convolves the transmission characteristics of sound according to the positional relationship between the respective users in a virtual space. An information processing apparatus is disclosed that includes an acoustic processing unit and an output control unit that outputs sound based on the signal generated by the acoustic processing from an output device used by each of the users.
[0003] International Publication No. 2022 / 196073
[0004] Since the apparatus of Patent Document 1 convolves the transmission characteristics so that the sound of other performers can be heard at a relative position based on the position of the first performer, it does not make the sound of one's own performance audible from an arbitrary position in the space.
[0005] One of the purposes of this embodiment is to provide a sound processing method that can make the sound of one's own performance audible from a predetermined position in the space.
[0006] The sound signal processing method includes, at a first terminal, receiving acoustic space information, first position information of a first performer in the acoustic space, and second position information of a second performer in the acoustic space. At the first terminal, receiving a first sound signal related to the performance of the first performer from an audio interface. At the first terminal, outputting the first sound signal to a second terminal of the second performer via a network. At the first terminal, receiving a second sound signal related to the performance of the second performer via the network. At the first terminal, performing first binaural processing on the first sound signal based on the first position information, and performing second binaural processing on the second sound signal based on the second position information. At the first terminal, outputting the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing from the audio interface.
[0007] According to one embodiment of the present invention, it is possible to make the sound of one's own performance audible from a predetermined location within the space.
[0008] This is a block diagram showing the configuration of an audio signal processing system. This is a block diagram showing the configuration of audio signal processing device 1. This is a block diagram showing the configuration of audio signal processing device 1A. This is a flowchart showing the operation of audio signal processing device 1, which is a plan view showing the configuration of virtual space 101. This is a plan view showing the configuration of virtual space 101. This is a flowchart showing the operation of the second terminal. This is a plan view showing the configuration of virtual space 101. This is a block diagram showing the configuration of an audio signal processing system according to Modification 6. This is a block diagram showing the configuration of audio signal processing device 1B, which is an example of a viewer terminal. This is a plan view showing the configuration of virtual space 101. This is a flowchart showing the operation of audio signal processing device 1B. This is a block diagram showing the configuration of an audio signal processing system according to Modification 9. This is a block diagram showing the configuration of audio signal processing device 1C, which is an example of an administrator terminal. This is a block diagram showing the configuration of audio signal processing device 1 according to Modification 10. This is a block diagram showing the configuration of an audio signal processing system according to Modification 12. This is a conceptual diagram illustrating the configuration of an audio signal processing system when providing an individual environment.
[0009] Figure 1 is a block diagram showing the configuration of the sound signal processing system of this embodiment. The sound signal processing system has a first terminal and a second terminal connected via a network. Figure 2 is a block diagram showing the configuration of the sound signal processing device 1. Figure 3 is a block diagram showing the configuration of the sound signal processing device 1A. The sound signal processing device 1 is, for example, a general-purpose information processing device and is an example of the first terminal. The sound signal processing device 1A is, for example, a general-purpose information processing device and is an example of the second terminal. The user connects a head-mounted display (HMD) 51 and an audio I / F 71 to the sound signal processing device 1. The sound signal processing device 1 includes a display I / F 31, a user I / F 32, flash memory 33, CPU 34, RAM 35, communication I / F 36, USB I / F 37, and DSP 38.
[0010] The display interface 31 is connected to the display. The display interface 31 of the sound signal processing device 1 is connected to an HMD 51, which is an example of a display. In this embodiment, an example is shown in which an HMD 51 is connected to the sound signal processing device 1, but a display such as an LCD may also be connected to the sound signal processing device 1.
[0011] The user interface 32 is a keyboard, mouse, or a touch panel stacked on a display. When the user interface 32 is a touch panel, it, together with the display, constitutes a GUI (Graphical User Interface).
[0012] Communication I / F36 is connected to a network such as a LAN or the Internet.
[0013] The USB I / F 37 is connected to the audio I / F 71. The audio I / F 71 has a USB terminal and an audio terminal. The audio I / F 71 is connected to audio equipment via an audio cable. In this embodiment, the audio I / F 71 is connected to headphones 81 and a guitar amplifier 82. A guitar 83 is connected to the guitar amplifier 82. The guitar amplifier 82 receives an audio signal from the guitar 83.
[0014] The guitar amplifier 82 inputs an audio signal related to the sound of the guitar 83 being played to the audio interface 71. The audio interface 71 outputs the audio signal to the headphones 81. The audio interface 71 may also be connected to the audio signal processing device 1 via a LAN. In this case, the USB interface 37 may be connected to the audio interface 71 solely for power supply.
[0015] The CPU 34 is a general-purpose processor. The CPU 34 reads the program stored in the flash memory 33, which is a storage medium, into the RAM 35 and controls each component of the sound signal processing device 1.
[0016] DSP38 is a processor dedicated to signal processing. DSP38 decodes audio data received from other devices via the communication interface 36. The audio data transmitted by other devices includes sound signals from sound sources located in a virtual space and location information of those sound sources. DSP38 performs signal processing on the decoded sound signals and the sound signals input from the audio interface 71. Signal processing may also be performed by the CPU 34. Furthermore, signal processing does not need to be performed by the processor inside the sound signal processing device 1; it may be performed by the processor of another device connected via, for example, the USB interface 37 or the communication interface 36.
[0017] The DSP38 performs binaural processing by convolving, for example, an HRTF (Head Related Transfer Function) into the sound signal. The HRTF represents the transfer function from a virtual sound source location to the user's right and left ears. This allows the user to perceive the sound as if it were emanating from a virtual sound source location. The DSP38 may also perform reverb processing to correspond to the reverberation of the virtual space. The reverb processing may have parameters corresponding to the room size of the virtual space.
[0018] Figure 4 is a plan view showing the configuration of a virtual space 101. Figure 5 is a flowchart showing the operation of the sound signal processing device 1. The first performer A1, who is a user of the sound signal processing device 1, is at a first location and uses the sound signal processing device 1, which is the first terminal, to conduct a session with the second performer A2, who is at a second location. As an example of a performance, the first performer A1 plays the guitar 83, and the second performer A2 sings. The sound signal processing device 1A has the same configuration as the sound signal processing device 1 shown in Figure 2. However, the sound signal processing device 1A is connected to the microphone 90 via the audio I / F 71, and is not connected to the guitar amplifier 82 and the guitar 83.
[0019] The sound signal processing device 1 receives, for example via a GUI, the specification of acoustic space information of the virtual space 101, the first position information of the first performer A1, and the second position information of the second performer A2 from the first performer A1 (S11). The virtual space information includes at least information relating to the size and shape of the virtual space. The virtual space information may also include information indicating the material of the wall surfaces constituting the virtual space (information corresponding to reflectance, sound absorption coefficient, etc.). The virtual space information, the first position information of the first performer A1, and the second position information of the second performer A2 are represented by two-dimensional or three-dimensional coordinates with a certain position as the origin.
[0020] In the example shown in Figure 4, the first position information includes information indicating the coordinates of the first performer A1 at the center of the virtual space 101, and information indicating the coordinates of the virtual guitar amplifier 82V to the right rear of the virtual space 101 as seen from the first performer A1. The second position information includes information indicating the coordinates of the second performer A2 at the front center of the virtual space 101. The acoustic space information, the first position information, or the second position information may be received from a second terminal via the network.
[0021] The sound signal processing device 1 receives a first sound signal related to the performance of the first performer A1 from the audio I / F 71 (S12). In the example shown in Figure 4, the sound signal processing device 1 receives the first sound signal related to the performance of the first performer A1, which is output from the guitar amplifier 82, from the audio I / F 71. The sound signal processing device 1 outputs the first sound signal to the second terminal of the second performer A2 via the network through the communication I / F 36 (S13). The sound signal processing device 1 also receives a second sound signal related to the singing of the second performer A2 via the network (S14). The singing sound of the second performer is picked up by a microphone 90 connected to the sound signal processing device 1A, which is the second terminal. The second sound signal related to the singing sound picked up by the microphone 90 is input to the sound signal processing device 1A via the audio I / F 71. The sound signal processing device 1A outputs the second sound signal to the sound signal processing device 1, which is the first terminal, via the network through the communication I / F 36. Note that the order in which processing S12 or S14 is performed is not limited; either processing can be performed first.
[0022] The sound signal processing device 1 may also render CG images of the virtual space and objects such as performers to generate an image of the virtual space viewed in a predetermined direction from the viewpoint of the first performer A1. The sound signal processing device 1 may also output the rendered image to the HMD 51 via the display I / F 31.
[0023] The sound signal processing device 1 applies first binaural processing to the first sound signal based on first position information, and applies second binaural processing to the second sound signal based on second position information (S15). As first binaural processing, the sound signal processing device 1 convolves an HRTF into the first sound signal such that it is localized to a position to the right rear of the user (relative coordinates of the virtual guitar amplifier 82V with the coordinates of the first performer A1 as the origin). As second binaural processing, the sound signal processing device 1 convolves an HRTF into the second sound signal such that it is localized to a position in front of the user in the center (relative coordinates of the second performer A2 with the coordinates of the first performer A1 as the origin).
[0024] The sound signal processing device 1 outputs the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing to the headphones 81 via the audio I / F 71 (S16). At this time, the sound signal processing device 1 outputs the first binaural signal and the second binaural signal as a mixed sound signal to the audio I / F 71. The mixing balance of the first binaural signal and the second binaural signal may be adjusted, for example, by the first performer A1, for example, via a GUI.
[0025] The sound signal processing device 1 may also perform reverb processing on the first binaural signal, the second binaural signal, or the mixed sound signal to correspond to the reverberation of the virtual space 101.
[0026] First performer A1, by listening to the first and second binaural signals through headphones 81, can perceive that they are located in the center of the virtual space 101 and are hearing the singing of second performer A2 from directly in front of them. Furthermore, first performer A1 can perceive that the sound of their own performance is coming from a different location (to the right rear) than their own location in the virtual space 101.
[0027] In a performance in a real space, the first performer A1 hears the sound of their own performance emanating from sound equipment such as a guitar amplifier 82, or from an instrument such as a guitar 83. If the sound of their own performance, such as the sound of the guitar amplifier 82 or guitar 83, were localized inside the performer's head, it would create a sense of unease. In contrast, the sound signal processing device 1 of this embodiment convolves the transmission characteristics so that the sound is heard from a designated position for each of the first and second performers in the virtual space. Therefore, the first performer A1, while in the virtual space 101, can perceive the sound of their own guitar 83 as coming from the guitar amplifier 82 located to their right and behind them, providing a new customer experience that allows for a realistic session in a virtual space, something that was not possible with conventional methods.
[0028] However, the position from which the sound of one's own performance is localized is not limited to the position of sound equipment or instruments installed in the actual space. The sound signal processing device 1 of this embodiment can also hear the sound of one's own performance from any position. For example, in the example shown in Figure 6, the sound signal processing device 1 receives an instruction from the first performer A1 via the GUI to change the coordinates of the virtual guitar amplifier 82V to the left rear of the first performer A1. As a first binaural process, the sound signal processing device 1 convolves an HRTF into the first sound signal that localizes it to the left rear of the user. As a result, the first performer A1 can perceive that they are in the center of the virtual space 101 and are hearing the sound of their own performance from the left rear of themselves.
[0029] The operation of the sound signal processing device 1 described above may be performed not only at the first terminal used by the first performer A1, but also at the second terminal (sound signal processing device 1A) used by the second performer A2. Figure 7 is a flowchart showing the operation of the sound signal processing device 1A, which is the second terminal. Since the sound signal processing device 1A, which is the second terminal, has the same configuration as the sound signal processing device 1, the illustration and explanation of its configuration are omitted.
[0030] The second terminal, the sound signal processing device 1A, receives the specification of acoustic space information for the virtual space 101, the first position information of the first performer A1, and the second position information of the second performer A2 (S21). The acoustic space information and the first position information are received from the first terminal via the network. The second position information is received from the second performer A2, for example, via a GUI. In the example shown in Figure 6, the first position information includes information indicating the coordinates of the first performer A1 at the center of the virtual space 101, and information indicating the coordinates of the virtual guitar amplifier 82V to the left front of the virtual space 101 as seen from the second performer A2. The second position information includes information indicating the coordinates of the second performer A2 to the front center of the virtual space 101 as seen from the first performer A1. The second position information may also be received from the first terminal via the network.
[0031] The second terminal, the sound signal processing device 1A, receives a second sound signal related to the singing of the second performer A2 from the audio I / F 71 (S22). The second terminal, the sound signal processing device 1A, outputs the second sound signal to the first terminal of the first performer A1 via the network (S23). The second terminal, the sound signal processing device 1A, also receives a first sound signal related to the performance of the first performer A1 via the network (S24). Note that the processing in S22 or S24 may be performed in any order, and the order of processing is not limited.
[0032] Furthermore, the second terminal, the sound signal processing device 1A, may also render CG images of the virtual space and objects such as performers to generate an image of the virtual space viewed in a predetermined direction from the viewpoint of the second performer A2. The second terminal, the sound signal processing device 1A, may also output the rendered image to the HMD 51 via the display I / F 31.
[0033] The second terminal, the sound signal processing device 1A, applies first binaural processing to the first sound signal based on first position information, and applies second binaural processing to the second sound signal based on second position information (S25). As first binaural processing, the second terminal, the sound signal processing device 1A, convolves an HRTF into the first sound signal such that it is localized to the left front position of the user. As second binaural processing, the second terminal, the sound signal processing device 1A, convolves an HRTF into the second sound signal such that it is localized to the user's position.
[0034] The second terminal, the sound signal processing device 1A, outputs the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing from the audio I / F 71 (S26). At this time, the second terminal, the sound signal processing device 1A, outputs the first binaural signal and the second binaural signal as a mixed sound signal to the audio I / F 71. The level balance between the first binaural signal and the second binaural signal may be adjusted by the second performer A2, for example, via a GUI. The level balance may be the same as or different from the level balance of the first terminal.
[0035] Furthermore, the second terminal, the sound signal processing device 1A, may perform reverb processing on the first binaural signal, the second binaural signal, or the mixed sound signal to correspond to the reverberation of the virtual space 101.
[0036] As a result, the second performer A2 can perceive that they are located in the center of the virtual space 101 and are hearing the performance sounds of the first performer A1 from their left front. Therefore, the second performer A2 can perceive that they are in the same space as the first performer A1 and are conducting a session together.
[0037] The second position information may include not only the coordinates of the second performer A2 but also the coordinates of sound equipment, etc. For example, in the example of Figure 8, the second position information includes information indicating the coordinates of the second performer A2 in the front center of the virtual space 101 as seen from the perspective of the first performer A1, and information indicating the coordinates of the virtual monitor speaker 90V in front of the virtual space 101 as seen from the perspective of the second performer A2. The second terminal, as a second binaural processing, convolves an HRTF into the second sound signal such that it is localized to the position of the virtual monitor speaker 90V. In this case, the second performer A2 can perceive that the sound of their own singing is coming from a different position (in front) than their own position in the virtual space 101.
[0038] In a session in a real space, the singer may hear their own singing sound emanating from a monitor speaker. Therefore, if their own singing sound is localized inside their head, it may cause discomfort for the singer. In contrast, the second terminal of this embodiment convolves the transmission characteristics so that the sound is heard from a designated position for each of the first and second performers in the virtual space. As a result, the second performer A2, while in the virtual space 101, can perceive their own singing sound as emanating from a virtual monitor speaker 90V located in front of them, providing a new customer experience that enables a realistic session in a virtual space, something that was not possible with conventional methods.
[0039] The first and second position information within the virtual space 101 may be the same or different for the first and second terminals. For example, the second position information at the first terminal may include information indicating the coordinates of the second performer A2, and the second position information at the second terminal may include information indicating the coordinates of the virtual monitor speaker 90V located in front of the virtual space 101 as seen from the perspective of the second performer A2, as shown in Figure 8. In this case, the first performer A1 hears the singing sound from the position of the second performer A2, and the second performer A2 hears their own singing sound from a position (in front) different from their own position within the virtual space 101. In this case, the sound signal processing device 1 can also realize a session unique to the virtual space that cannot be realized in a real space, by having each performer set their own localization position.
[0040] Furthermore, the acoustic space information may be the same or different for the first and second terminals. In this case as well, the sound signal processing device 1 can enable each performer to set up their own virtual space, thereby realizing a session unique to the virtual space that cannot be achieved in a real space.
[0041] Furthermore, the sound signal processing device 1 may store or learn information about past sound settings received from the first performer A1 or the second performer A2, and set them automatically. The sound settings include the level balance between the first binaural signal and the second binaural signal, gain settings, reverb settings, or equalizer settings.
[0042] As described above, the sound signal processing device 1 of this embodiment can enable each performer to create their own unique sound settings, thereby realizing sessions that are unique to virtual space and cannot be achieved in a real-world space.
[0043] (Modification 1) The first binaural processing in Modification 1 includes processing to localize a first indirect sound to a first sound signal based on first position information, and the second binaural processing includes processing to localize a second indirect sound to a second sound signal based on the first and second position information. Indirect sound is sound that reaches the listening position after the sound from the sound source is reflected off the walls of the virtual space, and includes early reflections with a fixed direction and phase, or reverberations with a random direction and phase. The localization processing of indirect sound by binaural processing is performed only on the early reflection system, and localization processing may be performed using effects such as reverb in the reverberation processing system. The sound signal processing device 1 uses the coordinates of the first performer A1 in the virtual space 101, which are included in the first position information, as the listening position, and the coordinates of the virtual guitar amplifier 82V as the sound source position, and acquires an impulse response corresponding to the reverberation of the virtual space 101. The impulse response is acquired by measurement, for example, by placing a dummy head at the listening position in the real space corresponding to the virtual space 101. The sound signal processing device 1 performs a process to localize the first indirect sound by convolving the acquired impulse response into the first sound signal.
[0044] Alternatively, the impulse response may be obtained by simulation based on, for example, the ray acoustics method or the virtual image method. The ray acoustics method is a technique for tracking the trajectory (sound rays) of sound radiated from a sound source and calculating the temporal pattern of the energy of the sound rays passing through the listening position. In the simulation using the ray acoustics method, based on the energy of the sound rays in the sound reception area, when each sound ray is regarded as a virtual sound image of the reverberant sound, the direction, arrival time, and arrival level from each virtual sound source at the listening position are determined. The virtual image method is a technique for creating virtual sound sources (virtual sound images of the sound source) with respect to the wall surface of the space and determining the direction, arrival time, and arrival level from each virtual sound source at the listening position. The sound signal processing device 1 may perform a process of localizing the first indirect sound by generating an impulse response of the head-related transfer function representing the direction, arrival time, and arrival level of each virtual sound source obtained by simulation and convolving the impulse response with the first sound signal.
[0045] Alternatively, the sound signal processing device 1 may perform a process of localizing the first indirect sound by subjecting the first sound signal to level delay filter processing having a delay amount and an attenuation amount corresponding to each virtual sound source obtained by simulation.
[0046] Similarly, the sound signal processing device 1 uses the coordinates of the first performer A1 in the virtual space 101 included in the first position information as the listening position and the coordinates of the second performer A2 included in the second position information as the sound source position to obtain an impulse response corresponding to the reverberation of the virtual space 101. As described above, the impulse response may be obtained by measurement or by simulation. The sound signal processing device 1 performs a process of localizing the second indirect sound by convolving the obtained impulse response with the second sound signal. Alternatively, the sound signal processing device 1 may perform a process of localizing the second indirect sound by subjecting the second sound signal to level delay filter processing having a delay amount and an attenuation amount corresponding to each virtual sound source obtained by simulation.
[0047] Thereby, the first performer A1 can perceive not only the direct sound but also the indirect sound generated in the real space, and can be in the virtual space 101 and conduct a more realistic session with the second performer A2.
[0048] The operation of the sound signal processing device 1 of the first modification example may be performed not only on the first terminal used by the first performer A1 but also on the sound signal processing device 1A which is the second terminal used by the second performer A2.
[0049] (Second Modification Example) The first position information may include the first orientation information of the first performer, and the second position information may include the second orientation information of the second performer. In this case, the first binaural processing further includes a localization process based on the first orientation information, and the second binaural processing further includes a localization process based on the first orientation information and the second orientation information.
[0050] The first orientation information includes the direction information in which the first performer A1 is facing and the direction information in which the virtual guitar amplifier 82V is facing. The second orientation information includes the direction information in which the second performer A2 is facing. The sound signal processing device 1 convolves the impulse response of the head-related transfer function corresponding to the orientation of the virtual guitar amplifier 82V with respect to the orientation of the first performer A1 with the first sound signal. The sound signal processing device 1 convolves the impulse response of the head-related transfer function corresponding to the orientation of the second performer A2 with respect to the orientation of the first performer A1 with the second sound signal.
[0051] The performance sound and the singing sound show the highest level when the front direction of the sound source and the front direction of the listener face each other, and attenuate as the left and right directions increase. Also, as the left and right directions increase, the high frequency range attenuates more than the low frequency range. Therefore, the sound signal processing device 1 may perform gain correction such that the level of the sound signal becomes lower as the difference (angle difference) between the orientation of the first performer A1 and the orientation of the second performer A2 becomes larger. Also, the sound signal processing device 1 may perform equalizer processing such that the level of the high frequency range becomes lower or the level of the low frequency range becomes higher as the difference (angle difference) between the orientation of the first performer A1 and the orientation of the second performer A2 becomes larger. The sound signal processing device 1 may perform gain correction such that the level of the sound signal becomes lower as the difference (angle difference) between the orientation of the first performer A1 and the orientation of the virtual guitar amplifier 82V becomes larger. Also, the sound signal processing device 1 may perform equalizer processing such that the level of the high frequency range becomes lower or the level of the low frequency range becomes higher as the difference (angle difference) between the orientation of the first performer A1 and the orientation of the virtual guitar amplifier 82V becomes larger.
[0052] This allows the sound signal processing device 1 to perceive changes in the relative positions of the performer and the sound source in real time. The sound signal processing device 1 can represent the relative positions of the performer and the sound source with greater precision. Therefore, the first performer A1 can be in the virtual space 101 and conduct a more realistic session with the second performer A2.
[0053] The operation of the sound signal processing device 1 in Modification 2 may be performed not only on the first terminal used by the first performer A1, but also on the second terminal, the sound signal processing device 1A, used by the second performer A2.
[0054] (Modification 3) The sound signal processing device 1 may receive first sensing information, which senses the performance of the first performer, from the sensor interface, and second sensing information, which senses the performance of the second performer, via the network.
[0055] The first sensing information and the second sensing information are acquired by sensors, respectively. For example, the sound signal processing device 1 estimates information such as the position, orientation, and motion of the first performer A1 based on an image of the first performer A1 taken by an image sensor such as a camera, and acquires the first sensing information. Alternatively, the sound signal processing device 1 may estimate information such as the position, orientation, and motion of the first performer A1 based on sensors such as a position sensor, motion sensor, and head tracker worn by the first performer A1, and acquire the first sensing information.
[0056] The sound signal processing device 1 synchronizes and plays back video related to the performances of the first performer and the second performer based on the first sensing information and the second sensing information. Specifically, the sound signal processing device 1 operates 3D models of the first performer A1 and the second performer A2 based on information such as their respective positions, orientations, and motions.
[0057] This allows the first performer A1 to be in the virtual space 101 and visually perceive the performance of the second performer A2, enabling them to conduct a realistic session.
[0058] The operation of the sound signal processing device 1 in Modification 3 may be performed not only on the first terminal used by the first performer A1, but also on the second terminal, the sound signal processing device 1A, used by the second performer A2.
[0059] (Modification 4) The sound signal processing device 1 according to Modification 4 corrects the sound environment at the first location. For example, the sound signal processing device 1 performs processing to reduce the reverberation of the actual space in which the first performer A1 is located. Information on the reverberation of the actual space in which the first performer A1 is located is obtained in advance by measuring the impulse response. In the first binaural processing and the second binaural processing, the sound signal processing device 1 reduces the reverberation of the actual space in which the first performer A1 is located by convolving the inverse characteristics of the impulse response measured in advance.
[0060] Furthermore, for example, the sound signal processing device 1 may perform a process to cancel out noise in the actual space where the first performer A1 is located. Also, the sound signal processing device 1 may perform a process to cancel out the sound of the guitar that is generated in the actual space where the first performer A1 is located.
[0061] This allows the first performer A1 to conduct the session with a greater sense of immersion, perceiving themselves as being in the virtual space 101.
[0062] The operation of the sound signal processing device 1 in Modification 4 may be performed not only on the first terminal used by the first performer A1, but also on the second terminal, the sound signal processing device 1A, used by the second performer A2.
[0063] (Modification 5) The sound signal processing device 1 according to Modification 5 further superimposes and displays information used by the performer onto the rendered image. The sound signal processing device 1 superimposes and displays information such as musical scores, lyrics, and dialogue. The sound signal processing device 1 may also estimate the current performance position in the session based on the musical score information and the first or second sound signal, and display the current performance position on the musical score information. The sound signal processing device 1 may also display information indicating the timing of chord progressions and rhythm changes in the next measure based on the estimated performance position.
[0064] (Modification 6) Figure 9 is a block diagram showing the configuration of the sound signal processing system according to Modification 6. The sound signal processing system according to Modification 6 further includes a plurality of viewer terminals connected via a network.
[0065] Figure 10 is a block diagram showing the configuration of an audio signal processing device 1B, which is an example of a viewer terminal. Components common to the audio signal processing device 1 in Figure 2 are denoted by the same reference numerals, and their explanations are omitted. The audio signal processing device 1B differs from the audio signal processing device 1 in that it does not have a DSP 38. Other components are the same as those of the audio signal processing device 1. However, only headphones 81 are connected to the audio I / F 71.
[0066] Figure 11 is a plan view showing the configuration of the virtual space 101. Figure 12 is a flowchart showing the operation of the sound signal processing device 1B. The viewer L1, who is a user of the sound signal processing device 1B, is located at a third location and watches the performances of the first performer A1 and the second performer A2, such as their sessions.
[0067] The audio signal processing device 1B, which is the viewer terminal, receives acoustic space information, first position information, and second position information (S31). For example, the audio signal processing device 1B receives acoustic space information, first position information, and second position information from the first terminal and second terminal via the network. Furthermore, the audio signal processing device 1B receives the viewer's third position information (S32). As shown in Figure 11, the third position information includes information indicating the coordinates of viewer L1 at the left edge center of the virtual space 101.
[0068] Next, the sound signal processing device 1B receives the first sound signal and the second sound signal via the network (S33). In other words, the first terminal distributes the first sound signal, and the second terminal distributes the second sound signal. Note that the processing in S31 to S33 can be performed in any order, and the order of processing is not limited.
[0069] The sound signal processing device 1B applies first binaural processing to the first sound signal based on first position information and third position information, and applies second binaural processing to the second sound signal based on second position information and third position information (S34). As first binaural processing, the sound signal processing device 1B convolves an HRTF into the first sound signal such that it is localized to a position to the right front of the viewer (relative coordinates of the virtual guitar amplifier 82V with the viewer L1's coordinates as the origin). As second binaural processing, the sound signal processing device 1B convolves an HRTF into the second sound signal such that it is localized to a position to the left front of the user (relative coordinates of the second performer A2 with the viewer L1's coordinates as the origin).
[0070] In this modified example 7, the sound signal processing device 1B performs binaural processing by software processing of the CPU 34 instead of a DSP. However, in the present invention, the viewer terminal may perform binaural processing using a DSP.
[0071] The sound signal processing device 1B then outputs the first binaural signal and the second binaural signal to headphones 81, which is an example of the listener's audio equipment (S35). At this time, the sound signal processing device 1B outputs the mixed sound signal of the first binaural signal and the second binaural signal to the audio I / F 71. The mixing balance of the first binaural signal and the second binaural signal may be adjusted by the listener L1, for example, via a GUI.
[0072] Furthermore, the sound signal processing device 1B may render CG images of the virtual space and objects such as performers to generate an image of the virtual space viewed in a predetermined direction from the viewpoint of the viewer L1. The sound signal processing device 1 may also output the rendered image to the HMD 51 via the display I / F 31.
[0073] Through the above operations, viewer L1 can perceive that they are in the virtual space 101 and that the sound of the first performer A1's guitar 83 is coming from the guitar amplifier 82, thus providing a new customer experience that allows them to view a realistic performance in a virtual space, something that was not possible before.
[0074] (Modification 7) The sound signal processing device 1 relating to the first terminal or the sound signal processing device 1A relating to the second terminal in Modification 7 receives reaction sound signals relating to the viewer's reactions from the viewer terminal via the network. The sound signal processing device 1 applies a third binaural processing to the reaction sound signals.
[0075] The reaction sound signal is, for example, the viewer's voice or sound effect such as applause, acquired by a microphone (not shown) connected to the sound signal processing device 1B, which is a viewer terminal. Alternatively, the sound signal processing device 1B displays icon images such as "cheers," "applause," "calls," and "buzz" on a display. The sound signal processing device 1B receives reaction information (text information) such as "cheers," "applause," "calls," and "buzz" from the user by receiving selection operations for these icon images via the user I / F 32. The sound signal processing device 1B transmits the received reaction information to the sound signal processing device 1. The sound signal processing device 1 generates a reaction sound signal corresponding to the received reaction information. Alternatively, the sound signal processing device 1B may generate a reaction sound signal corresponding to the received reaction. When sending and receiving reaction information that contains less information than the reaction sound signal, the sound signal processing system can reduce the communication load.
[0076] The sound signal processing device 1B transmits the viewer's reaction sound signal to the sound signal processing device 1. The sound signal processing device 1B also transmits the viewer's third location information to the sound signal processing device 1.
[0077] The sound signal processing device 1 receives the reaction sound signal and third position information via the network. Based on the third position information, the sound signal processing device 1 applies third binaural processing to the reaction sound signal. As third binaural processing, the sound signal processing device 1 convolves an HRTF into the reaction sound signal such that it is localized to the left position of the first performer A1 (relative coordinates of viewer L1 when the coordinates of the first performer A1 are taken as the origin). At this time, the sound signal processing device 1 may perform binaural processing by software processing of the CPU 34 instead of the DSP 38. In other words, the first and second binaural processing may be performed by the DSP 28, and the third binaural processing may be performed by software processing of the CPU 34. As a result, the sound signal processing device 1 can prioritize binaural processing of the first and second sound signals related to performance using the DSP 38, which has relatively low latency, and reduce the resources allocated to the third binaural processing, thereby minimizing the impact of latency on performance.
[0078] In this manner, the first performer A1 or the second performer A2 can perform while perceiving that viewer L1 is also in the same virtual space 101. Furthermore, the first performer A1 or the second performer A2 can perform while referring to viewer L1's reactions.
[0079] (Modification 8) The sound signal processing device 1 in Modification 8 receives third sensing information, which senses the viewer's reaction, via the network. Based on the third sensing information, the sound signal processing device 1 plays back the video related to the viewer's reaction.
[0080] The sound signal processing device 1B estimates information such as the position, orientation, and motion of viewer L1 based on an image of viewer L1 captured by an image sensor such as a camera, and acquires third sensing information. Alternatively, the sound signal processing device 1B may estimate information such as the position, orientation, and motion of viewer L1 based on sensors such as a position sensor and a motion sensor worn by viewer L1, and acquire third sensing information.
[0081] The sound signal processing device 1 operates a 3D model of the viewer L1, including its position, orientation, and motion, based on the third sensing information.
[0082] This allows the first performer A1 or the second performer A2 to be in the virtual space 101 and visually perceive the reactions of the viewer L1, enabling them to perform more realistic sessions and other performances.
[0083] (Modification 9) Figure 13 is a block diagram showing the configuration of the sound signal processing system according to Modification 9. The sound signal processing system according to Modification 9 has an administrator terminal in addition to the configuration shown in Figure 10. The administrator terminal is connected to the first terminal, the second terminal, and the viewer terminal via a network.
[0084] In the sound signal processing system according to Modification 9, the administrator terminal distributes acoustic space information, first location information, and second location information to the viewer terminal. In addition, in the sound signal processing system according to Modification 9, the first terminal transmits a first sound signal to the administrator terminal, and the second terminal transmits a second sound signal to the administrator terminal.
[0085] The administrator terminal performs first binaural processing on the first audio signal received from the first terminal, and second binaural processing on the second audio signal received from the second terminal. The administrator terminal then distributes the first binaural signal and the second binaural signal to the viewer terminal.
[0086] In this case, the audio signal delivered to the viewer's terminal undergoes high-processing binaural processing on the administrator terminal. Therefore, the administrator terminal can provide a realistic performance experience in the virtual space, regardless of the processing power of the viewer's terminal.
[0087] Figure 14 is a block diagram showing the configuration of an audio signal processing device 1C, which is an example of an administrator terminal. The configuration of the audio signal processing device 1C is the same as the configuration of the audio signal processing device 1 shown in Figure 2. However, headphones 81 and a microphone 90 are connected to the audio I / F 71, and an LCD 501 is connected to the display I / F 31.
[0088] Microphone 90 acquires an audio signal related to the voice of the administrator of the sound signal processing system. The sound signal processing device 1C transmits the audio signal related to the administrator's voice acquired by microphone 90 to the terminal of another user, such as the first terminal or the second terminal. Microphone 90 may be a microphone built into the headphones 81.
[0089] The administrator terminal may accept a destination specification for the audio signal related to the administrator's voice acquired by microphone 90. The destination specification may include individual specification of the first terminal or the second terminal, or a specification of all terminals including the first terminal and the second terminal.
[0090] This allows administrators to provide talkback to any performer or all performers, enabling them to manage performance and provide services in the virtual space while communicating with the performers.
[0091] (Modification 10) Figure 15 is a block diagram showing the configuration of the sound signal processing device 1 according to Modification 10. The sound signal processing device 1 shown in Figure 15 has the same configuration as shown in Figure 2. However, the sound signal processing device 1 is connected to the microphone 90 via the audio I / F 71. The microphone 90 acquires sound signals related to the voice of the first performer A1. The sound signal processing device 1 transmits the sound signals related to the voice of the first performer A1 acquired by the microphone 90 to the terminals of other users, such as the second terminal. Note that the microphone 90 may be a microphone built into the headphones 81.
[0092] The audio signal processing unit 1 of the first terminal receives audio signals related to the voice of other users from the second terminal or other user terminals such as viewer terminals via the network. The audio signal processing unit 1 receives specifications for sound processing conditions, including the localization position (conditions related to the fourth binaural processing) of the audio signals related to the voice of the other users. Based on the received conditions, the audio signal processing unit 1 applies sound processing, including the fourth binaural processing, to the audio signals and outputs the processed audio signals from the audio I / F 71.
[0093] For example, in the example shown in Figure 8, the second binaural processing convolves an HRTF that localizes the virtual monitor speaker 90V to its relative coordinates, with the coordinates of the first performer A1 as the origin. The fourth binaural processing convolves an HRTF that localizes the second performer A2 to its relative coordinates, with the coordinates of the first performer A1 as the origin.
[0094] Furthermore, the sound processing conditions may include, for example, level balance conditions. Level balance includes the level balance of the first binaural signal, the second binaural signal, the third binaural signal, and the above-mentioned audio signal.
[0095] Alternatively, the sound processing conditions may be accepted via a GUI, or automatically according to the orientation (gaze) of each user detected by various sensors such as motion sensors. The sound signal processing device 1 may perform processing to maximize the level of the sound corresponding to other performers or viewers located in the direction that the first performer A1 is facing, or to emphasize that sound.
[0096] In real space, people unconsciously concentrate on listening to sounds coming from the direction they are facing (sounds they are paying attention to), creating a cocktail party effect. However, it is difficult to create a cocktail party effect in virtual space. In contrast, the sound signal processing device 1 according to Modification 10 can achieve a cocktail party effect in virtual space by making the level of the sound corresponding to other performers or viewers located in the direction that the first performer A1 is facing the highest, or by processing to emphasize that sound. This makes it possible to communicate among any participants, including performers or viewers.
[0097] (Modification 11) The sound signal processing system may perform sound processing individually at each terminal, as shown in Figure 9, or it may perform some sound processing at a server such as an administrator terminal, as shown in Figure 13, and then distribute the processed sound signal to the viewer terminal. Alternatively, the sound signal processing system may perform some sound processing at an edge server and then distribute the processed sound signal to the viewer terminal. Figure 16 is a block diagram showing the configuration of the sound signal processing system according to Modification 11. The sound signal processing system according to Modification 11 has an edge server 1D in addition to the configuration shown in Figure 13. The edge server 1D is connected to the first terminal, the second terminal, the administrator terminal, and the viewer terminal via a network. The configuration of the edge server 1D is the same as that of the sound signal processing device 1. However, the edge server 1D does not need to have a display I / F 31 and a USB I / F 37.
[0098] In the sound signal processing system according to Modification 11, the edge server 1D performs some or all of the sound processing. For example, the administrator terminal transmits acoustic space information, first location information, and second location information to the edge server. The first terminal also transmits a first sound signal to the edge server 1D, and the second terminal transmits a second sound signal to the edge server 1D.
[0099] Edge server 1D performs first binaural processing on the first sound signal received from the first terminal and second binaural processing on the second sound signal received from the second terminal. Edge server 1D distributes the first binaural signal and the second binaural signal to the viewer terminal.
[0100] In this case, the computationally intensive binaural processing is performed on edge server 1D, which is closer to the viewer's terminal. Therefore, viewers can enjoy a more lag-free viewing experience.
[0101] Next, Figure 17 is a conceptual diagram illustrating the configuration of an audio signal processing system when providing individual environments. The positions of the performers and viewers in the virtual space 101 may be the same or different for all terminals. For example, in the example of Figure 17, at the first point of the first performer A1, the first performer A1 is located at the center of the virtual space 101, the virtual guitar amplifier 82V is located to the right rear of the first performer A1, the second performer A2 is located in the front center, and the viewer L1 is located to the left. At the second point of the second performer A2, the second performer A2 is located at the center of the virtual space 101, the virtual monitor speaker 90V is located in front of the second performer A2, the first performer A1 is located behind it, the virtual guitar amplifier 82V is located to the left front, and the viewer L1 is located to the right. At viewer L1's third location, viewer L1 is at the center of the virtual space 101, the virtual right main speaker 90RV is located to the right and in front of viewer L1, the virtual left main speaker 90LV is located to the left and in front of viewer L1, the virtual guitar amplifier 82V is located to the right of the virtual left main speaker 90LV, the first performer is located in the virtual guitar amplifier 82V, and the second performer A2 is located to the left of the virtual right main speaker 90RV.
[0102] In this case, viewer L1 sees the first performer A1 and the second performer A2 in front of them, with the sound of the first performer's guitar playing localized at the position of the virtual guitar amplifier 82V, and the sound of the second performer A2's singing localized by panning using the virtual right main speaker 90RV and the virtual left main speaker 90LV.
[0103] The sound settings for viewer L1, such as localization position and level balance, may be specified by viewer L1 or by the administrator of the administrator terminal. Alternatively, the sound settings for each viewer may be automatically configured by storing or learning past information on the viewer terminal or administrator terminal.
[0104] In other words, the sound signal processing system (sound signal processing device or sound signal processing method) in the example shown in Figure 17 has the following technical concept. The first terminal receives acoustic space information, first position information of the first performer in the acoustic space, and second position information of the second performer in the acoustic space; the first terminal receives a first sound signal related to the performance of the first performer from the audio interface; the first terminal outputs the first sound signal to the second performer's second terminal via the network; the first terminal receives a second sound signal related to the performance of the second performer via the network; the first terminal applies first binaural processing to the first sound signal based on the first position information, and applies second binaural processing to the second sound signal based on the second position information; the first terminal outputs the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing from the audio interface; the second terminal receives the acoustic space information, the first position information, and the second position information; the second terminal receives the second sound signal from the audio interface; A sound signal processing method comprising: a second terminal outputting the second sound signal to the first terminal via a network, receiving the first sound signal via the network, performing the first binaural processing and the second binaural processing at the second terminal, and outputting the first binaural signal and the second binaural signal from the audio interface at the second terminal, wherein the acoustic processing parameters of the first binaural processing and the second binaural processing at the first terminal and the first binaural processing and the second binaural processing at the second terminal are different.
[0105] Furthermore, the sound signal processing system (sound signal processing device or sound signal processing method) may also have the following technical concepts: The viewer terminal receives the acoustic space information, the first position information, and the second position information; the viewer terminal distributes the first sound signal and the second sound signal to the viewer terminal via the network; the viewer terminal performs the first binaural processing and the second binaural processing; the viewer terminal outputs the first binaural signal and the second binaural signal to the viewer's audio equipment; and the acoustic processing parameters of the first binaural processing and the second binaural processing in the first terminal, the first binaural processing and the second binaural processing in the second terminal, and the first binaural processing and the second binaural processing in the viewer terminal are different.
[0106] By performing individual sound settings for each performer and viewer as described above, it becomes possible to realize sessions and performances that are unique to virtual spaces and cannot be achieved in real-world spaces.
[0107] In the above embodiment, binaural processing was shown as an example of localization processing, but localization processing is not limited to binaural processing. For example, panning processing and beamforming processing using multiple speakers are also examples of localization processing. Furthermore, the sound signal after binaural processing may be played back not only through headphones, but also through multiple speakers. When playing back the sound signal after binaural processing through multiple speakers, it is preferable to perform crosstalk cancellation processing.
[0108] In the above embodiment, a session by a first performer A1 and a second performer A2 was used as an example, but performances via a network at different locations are not limited to this example. The sound signal processing system may, for example, process the sound of a performance by even more performers. The sound signal processing system may also process the sound of a group performance, such as an orchestra, taking place in a real space, and a performance by another performer in a remote location. In this case as well, the sound signal processing system processes the sound of the other performer's performance to localize it to a predetermined position, so that a viewer in a remote location can perceive that a group performance, such as an orchestra, taking place in a real space is simultaneously being performed by another performer in a remote location.
[0109] In the above embodiment, an example was shown in which a general-purpose information processing device (sound signal processing device 1) performs the first binaural processing and the second binaural processing. However, if, for example, an audio I / F 71 is connected to a network, the audio I / F 71 may perform the first binaural processing and the second binaural processing. In other words, the sound signal processing device of the present invention can also be realized by an audio I / F that can be connected to a network.
[0110] The description of these embodiments should be considered in all respects to be illustrative and not restrictive. The scope of the invention is indicated by the claims, rather than by the embodiments described above. Furthermore, the scope of the invention includes the claims and equivalents.
[0111] 1, 1A, 1B, 1C: Audio signal processing unit, 1D: Edge server, 31: Display I / F, 32: User I / F, 33: Flash memory, 34: CPU, 35: RAM, 36: Communication I / F, 37: USB I / F, 51: HMD, 71: Audio I / F, 81: Headphones, 82: Guitar amplifier, 82V: Virtual guitar amplifier, 83: Guitar, 90: Microphone, 90LV: Virtual left main speaker, 90RV: Virtual right main speaker, 90V: Virtual monitor speaker, 101: Virtual space, 501: LCD
Claims
1. An audio signal processing method comprising: a first terminal receiving acoustic space information, first position information of a first performer in the acoustic space, and second position information of a second performer in the acoustic space; a first terminal receiving a first sound signal relating to the performance of the first performer from an audio interface; a first terminal outputting the first sound signal to the second performer's second terminal via a network; a first terminal receiving a second sound signal relating to the performance of the second performer via the network; a first terminal applying first binaural processing to the first sound signal based on the first position information, and a second binaural processing to the second sound signal based on the second position information; and a first terminal outputting the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing from the audio interface.
2. The sound signal processing method according to claim 1, wherein the second terminal receives the acoustic space information, the first position information, and the second position information; the second terminal receives the second sound signal from the audio interface; the second terminal outputs the second sound signal to the first terminal via the network, receives the first sound signal via the network; the second terminal performs the first binaural processing and the second binaural processing; and the second terminal outputs the first binaural signal and the second binaural signal from the audio interface.
3. The sound signal processing method according to claim 1 or claim 2, wherein the first binaural processing includes processing to localize a first indirect sound to the first sound signal based on the first position information, and the second binaural processing includes processing to localize a second indirect sound to the second sound signal based on the first position information and the second position information.
4. The sound signal processing method according to claim 1 or claim 2, wherein the first position information includes first orientation information of the first performer, the second position information includes second orientation information of the second performer, the first binaural processing further includes localization processing based on the first orientation information, and the second binaural processing further includes localization processing based on the first orientation information and the second orientation information.
5. The sound signal processing method according to claim 1 or 2, wherein the first terminal receives first sensing information of the performance of the first performer from a sensor interface, the first terminal receives second sensing information of the performance of the second performer via the network, and the first terminal synchronously plays back video related to the performances of the first and second performers based on the first sensing information and the second sensing information.
6. The sound signal processing method according to claim 1 or 2, wherein the viewer terminal receives the acoustic space information, the first position information, and the second position information, distributes the first sound signal and the second sound signal to the viewer terminal via the network, performs the first binaural processing and the second binaural processing at the viewer terminal, and outputs the first binaural signal and the second binaural signal to the viewer's audio equipment at the viewer terminal.
7. The sound signal processing method according to claim 6, wherein the first terminal receives a reaction sound signal relating to the viewer's reaction from the viewer terminal via the network, and the first terminal applies a third binaural processing to the reaction sound signal.
8. The sound signal processing method according to claim 7, wherein the first terminal receives third sensing information, which senses the viewer's reaction, via the network, and the first terminal plays back video related to the viewer's reaction based on the third sensing information.
9. The sound signal processing method according to claim 1 or claim 2, comprising: distributing the acoustic space information, the first location information, and the second location information from an administrator terminal; distributing the first sound signal and the second sound signal to the administrator terminal via the network; performing the first binaural processing and the second binaural processing at the administrator terminal; and outputting the first binaural signal and the second binaural signal to the administrator's sound equipment at the administrator terminal.
10. The audio signal processing method according to claim 1 or 2, wherein the first terminal receives an audio signal relating to the voice of another user from another user's terminal via the network, the first terminal receives a specification of conditions for sound processing including a fourth binaural processing, the first terminal applies the sound processing to the audio signal based on the conditions, and the first terminal outputs the processed audio signal from the audio interface.
11. An audio signal processing device comprising a processor that receives acoustic space information, first position information of a first performer in the acoustic space, and second position information of a second performer in the acoustic space; receives a first sound signal relating to the performance of the first performer from an audio interface; outputs the first sound signal to the second terminal of the second performer via a network; receives a second sound signal relating to the performance of the second performer via the network; applies first binaural processing to the first sound signal based on the first position information; applies second binaural processing to the second sound signal based on the second position information; and outputs the first binaural signal after the first binaural processing and the second binaural signal after the second binaural processing from the audio interface.
Citation Information
Patent Citations
Live data delivery method, live data delivery system, live data delivery device, live data reproduction device, and live data reproduction method
WO2022113289A1
Information processing system, information processing method, and program
WO2022196073A1
Sound signal processing method, terminal, sound signal processing system, and management device
WO2023042671A1