Sound signal processing method and sound signal processing device
By classifying ambient sounds into groups based on their positions and applying tailored sound processing, the method effectively reduces processing load while maintaining natural ambient sound representation in virtual spaces, enhancing the auditory experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-03-31
AI Technical Summary
Existing sound signal processing methods fail to naturally represent ambient sounds such as environmental sounds and cheers in virtual spaces due to high processing loads when dealing with multiple sound sources.
A sound signal processing method that classifies ambient sounds into groups based on their positions relative to a listener, applying tailored sound processing to each group to reduce processing load while maintaining natural representation.
Achieves natural ambient sound representation in virtual spaces with reduced processing load by grouping and processing ambient sounds according to their positional characteristics, providing a realistic auditory experience.
Smart Images

Figure 2026055180000001_ABST
Abstract
Description
Technical Field
[0001] One embodiment of the present invention relates to a sound signal processing method and a sound signal processing apparatus.
Background Art
[0002] Patent Document 1 discloses an information processing apparatus including: an acoustic processing unit that performs acoustic processing of convolution of an acoustic signal obtained by collecting sound in a space where each of a plurality of co-performing users is located, with a sound transmission characteristic corresponding to a positional relationship between the respective users in a virtual space; and an output control unit that causes sound based on the signal generated by the acoustic processing to be output from output devices used by the respective users.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The apparatus of Patent Document 1 does not consider ambient sounds such as environmental sounds and cheers.
[0005] In order to naturally express ambient sounds such as environmental sounds and cheers in a virtual space, it is necessary to process the sounds of a large number of sound sources, and the processing load increases significantly.
[0006] One of the present embodiments aims to provide a sound signal processing method capable of realizing natural ambient sounds in a virtual space while suppressing the processing load.
Means for Solving the Problems
[0007] The sound signal processing method receives acoustic space information, first position information indicating the listening position in the acoustic space, and second position information indicating the positions of multiple ambient sounds in the acoustic space, receives a first sound signal of content sound and second sound signals relating to the multiple ambient sounds, classifies each of the multiple ambient sounds into one of a plurality of groups based on the first position information and the second position information, applies sound processing to each of the second sound signals classified into the plurality of groups, outputs the first sound signal and the second sound signal after sound processing. [Effects of the Invention]
[0008] According to one embodiment of the present invention, natural ambient sound can be achieved in a virtual space while reducing the processing load. [Brief explanation of the drawing]
[0009] [Figure 1] This is a block diagram showing the configuration of an audio signal processing system. [Figure 2] This is a block diagram showing the configuration of the sound signal processing device 1. [Figure 3] This is a block diagram showing the configuration of the sound signal processing device 1A. [Figure 4] This is a plan view showing the configuration of virtual space 101. [Figure 5] This is a flowchart showing the operation of the sound signal processing device 1. [Figure 6] This is a conceptual diagram illustrating an example of group classification. [Figure 7] This is a conceptual diagram illustrating other examples of group classification. [Figure 8] This is a plan view showing the configuration of the virtual space 101 according to modified example 5. [Figure 9] This is a block diagram showing the configuration of the sound signal processing system according to the modified example 7. [Figure 10] This is a plan view showing the configuration of the virtual space 101 according to the modified example 7. [Figure 11]This is a block diagram showing the configuration of the sound signal processing system according to modified example 8. [Figure 12] This is a block diagram showing the configuration of the sound signal processing system according to modified example 10. [Figure 13] This is a conceptual diagram illustrating the configuration of an audio signal processing system when providing individual environments. [Modes for carrying out the invention]
[0010] Figure 1 is a block diagram showing the configuration of the sound signal processing system of this embodiment. The sound signal processing system has performer terminals and listener terminals connected via a network. Figure 2 is a block diagram showing the configuration of an example of a performer terminal, the sound signal processing device 1. Figure 3 is a block diagram showing the configuration of an example of a listener terminal, the sound signal processing device 1A.
[0011] The sound signal processing device 1 is, for example, a general-purpose information processing device. The user connects a head-mounted display (HMD) 51 and an audio interface 71 to the sound signal processing device 1. The sound signal processing device 1 is equipped with a display interface 31, a user interface 32, flash memory 33, a CPU 34, RAM 35, a communication interface 36, a USB interface 37, and a DSP 38. The sound signal processing device 1A is also, for example, a general-purpose information processing device. The sound signal processing device 1A has the same configuration as the sound signal processing device 1 except that it does not have a DSP 38. However, in the sound signal processing device 1A, headphones 81 and a microphone 90 are connected to the audio interface 71.
[0012] The display interface 31 is connected to a display. The display interface 31 of the sound signal processing device 1 is connected to an HMD51, which is an example of a display. In this embodiment, an example is shown in which an HMD51 is connected to the sound signal processing device 1, but a display such as an LCD may also be connected to the sound signal processing device 1.
[0013] The user I / F 32 is a keyboard, a mouse, or a touch panel laminated on a display. When the user I / F 32 is a touch panel, the user I / F 32 and the display together constitute a GUI (Graphical User Interface).
[0014] The communication I / F 36 is connected to a network such as a LAN or the Internet.
[0015] The USB I / F 37 is connected to the audio I / F 71. The audio I / F 71 has a USB terminal and an audio terminal. The audio I / F 71 is connected to an audio device via an audio cable. In this embodiment, the audio I / F 71 has a headphone 81 and a guitar amplifier 82 connected thereto. A guitar 83 is connected to the guitar amplifier 82. The guitar amplifier 82 inputs an audio signal from the guitar 83.
[0016] The guitar amplifier 82 inputs an audio signal related to the performance sound of the guitar 83 to the audio I / F 71. The audio I / F 71 outputs an audio signal to the headphone 81. Note that the audio I / F 71 may be connected to the audio signal processing device 1 via a LAN. In this case, the USB I / F 37 may be connected to the audio I / F 71 only for power supply.
[0017] The microphone 90 acquires an audio signal related to the voice of the viewer. Note that the microphone 90 may be a microphone built into the headphone 81. The audio signal processing device 1A transmits the audio signal related to the voice of the viewer acquired by the microphone 90 to the terminal of another user such as a performer terminal. The voice of the viewer is an example of ambient sound. Ambient sound includes not only the voice of the viewer but also effect sounds emitted at unspecified timings such as applause, the sound of the wind, and the chirping of birds. Alternatively, ambient sound includes environmental sounds that are constantly generated such as the sound of rain, the rustling of a river, and the sound of air conditioning.
[0018] The CPU 34 is a general-purpose processor. The CPU 34 reads the program stored in the flash memory 33, which is a storage medium, into the RAM 35 and controls each component of the sound signal processing device 1.
[0019] DSP38 is a processor dedicated to signal processing. DSP38 decodes audio data received from other devices via the communication interface 36. The audio data transmitted by other devices includes sound signals from sound sources located in a virtual space and location information of those sound sources. DSP38 performs signal processing on the decoded sound signals and the sound signals input from the audio interface 71. Signal processing may also be performed by the CPU 34. Furthermore, signal processing does not need to be performed by the processor inside the sound signal processing device 1; it may be performed by the processor of another device connected via, for example, the USB interface 37 or the communication interface 36.
[0020] The DSP38 performs binaural processing by convolving, for example, an HRTF (Head Related Transfer Function) into the sound signal. The HRTF represents the transfer function from a virtual sound source location to the user's right and left ears. This allows the user to perceive the sound as if it were emanating from a virtual sound source location. The DSP38 may also perform reverb processing to correspond to the reverberation of the virtual space. The reverb processing may have parameters corresponding to the room size of the virtual space.
[0021] Figure 4 is a plan view showing the configuration of a virtual space 101. Figure 5 is a flowchart showing the operation of the sound signal processing device 1. The first performer A1, who is a user of the sound signal processing device 1, is located at a first location and uses the sound signal processing device 1, which is the performer terminal, to distribute performance-related content to multiple viewers located at other locations. As an example of a performance, the first performer A1 plays guitar 83.
[0022] The sound signal processing device 1 receives from the first performer A1 acoustic space information of the virtual space 101, first position information indicating the position of the first performer A1 in the virtual space 101 (listening position), and second position information indicating the positions of multiple ambient sounds in the virtual space 101 (S11). The virtual space information includes at least information relating to the size and shape of the virtual space. The virtual space information may also include information indicating the material of the walls constituting the virtual space (information corresponding to reflectance, sound absorption coefficient, etc.). The virtual space information, first position information, and second position information are represented by two-dimensional or three-dimensional coordinates with a certain position as the origin.
[0023] Acoustic space information, first position information, and second position information are received, for example, from the first performer A1 via a GUI. In the example shown in Figure 4, the first position information includes information indicating the coordinates of the first performer A1 at the center of the virtual space 101. The second position information includes information indicating the coordinates of multiple viewers L1 to L14. The second position information may also be received from each viewer terminal via the network.
[0024] The sound signal processing device 1 receives a first sound signal related to the performance (content sound) of the first performer A1 from the audio I / F 71 (S12). In the example shown in Figure 1, the sound signal processing device 1 receives a first sound signal related to the performance of the first performer A1 output from the guitar amplifier 82 from the audio I / F 71. The sound signal processing device 1 transmits the first sound signal to the viewer terminal via the network through the communication I / F 36 (S13). The sound signal processing device 1 also receives a second sound signal corresponding to the viewer's voice via the network (S14). Note that the order of processing in S12 or S14 is not limited and can be performed in any order.
[0025] The sound signal processing device 1 may also render CG images of objects such as the virtual space, performers, and viewers to generate video of the virtual space viewed in a predetermined direction from the viewpoint of the first performer A1. The sound signal processing device 1 may also output the rendered video to the HMD51 via the display I / F31.
[0026] The sound signal processing device 1 classifies each of the multiple ambient sounds into one of several groups (S15). Figure 6 is a conceptual diagram showing an example of group classification. In the example in Figure 6, the sound signal processing device 1 sets up the regions of the first group G1, the second group G2, and the third group G3, determines which region each viewer's position belongs to, and performs group classification. The region of the first group G1 is set around the first performer A1 and close to the first performer A1. The second group G2 is set in front of or to the side of the first performer A1. The region of the third group G3 is set far from the first performer A1 and in a direction other than in front of the first performer A1. The sound signal processing device 1 classifies viewers L8, L9, L12, and L13 into the first group G1, viewers L3, L4, L7, L10, L11, and L14 into the second group G2, and viewers L1, L2, L5, and L6 into the third group G3.
[0027] The sound signal processing device 1 applies the respective sound processing to each of the second sound signals classified into multiple groups (S16).
[0028] For example, for viewers L8, L9, L12, and L13 classified into the first group G1, the sound signal processing device 1 applies binaural processing to each viewer's second sound signal based on their respective positions. For example, the sound signal processing device 1 convolves an HRTF into the second sound signal of viewer L8 such that the second sound signal of viewer L8 is localized at the relative coordinates of viewer L8, with the coordinates of the first performer A1 as the origin. The sound signal processing device 1 convolves an HRTF into the second sound signal of viewer L9 such that the second sound signal of viewer L9 is localized at the relative coordinates of viewer L9, with the coordinates of the first performer A1 as the origin. The sound signal processing device 1 convolves an HRTF into the second sound signal of viewer L12 such that the second sound signal of viewer L12 is localized at the relative coordinates of viewer L12, with the coordinates of the first performer A1 as the origin. The sound signal processing device 1 convolves an HRTF (Head-Related Transformer) onto the second sound signal of viewer L13 such that the second sound signal of viewer L13 is localized to the relative coordinates of viewer L13, with the coordinates of the first performer A1 as the origin.
[0029] As a result, viewers L8, L9, L12, and L13, who are classified as Group 1 G1, can hear the audio clearly enough to identify its content and perceive its direction.
[0030] Furthermore, the sound signal processing device 1 mixes the second audio signals of viewers L3, L4, L7, L10, L11, and L14, who are classified as the second group G2, and performs binaural processing on the mixed second audio signals. The sound signal processing device 1 may also perform reverb processing on the mixed second audio signals. By performing reverb processing, the first performer A1 can perceive the voices of viewers classified as the second group G2 as being further away than the voices of viewers classified as the first group G1. In addition, the sound signal processing device 1 applies a high-pass filter to the mixed second audio signals. A band-pass filter, such as a filter or low-pass filter, may be applied. Even with filtering, the first presenter A1 can perceive the voices of viewers classified as group G2 as being further away than the voices of viewers classified as group G1.
[0031] The sound signal processing device 1 mixes the second audio signals of viewers L3 and L4 from among viewers L3, L4, L7, L10, L11, and L14, who are classified into the second group G2. The sound signal processing device 1 performs binaural processing on the mixed second audio signals. At this time, the sound signal processing device 1 averages the coordinates of viewers L3 and L4 and rounds them into a single coordinate system, and then convolves an HRTF such that the mixed second audio signals are localized to the averaged relative coordinates of viewers L3 and L4, with the coordinates of the first performer A1 as the origin.
[0032] The sound signal processing device 1 mixes the second audio signals of viewers L7 and L11. Furthermore, the sound signal processing device 1 performs binaural processing on the mixed second audio signals. At this time, the sound signal processing device 1 averages the coordinates of viewers L7 and L11 and rounds them into a single coordinate system, and convolves an HRTF such that the mixed second audio signal is localized to the averaged relative coordinates of viewers L7 and L11, with the coordinates of the first performer A1 as the origin. The sound signal processing device 1 mixes the second audio signals of viewers L10 and L14. The sound signal processing device 1 performs binaural processing on the mixed second audio signals. At this time, the sound signal processing device 1 averages the coordinates of viewers L10 and L14 and rounds them into a single coordinate system, and then convolves an HRTF such that the mixed second sound signal is localized to the averaged relative coordinates of viewers L3 and L4, with the coordinates of the first performer A1 as the origin.
[0033] As a result of the above sound processing, the voices of the audience in the second group G2 are perceived as being further away than those of the audience in the first group G1, and the intelligibility of their speech is reduced, as is their sense of direction. Therefore, the first presenter A1 can perceive a sense of distance to the audience at a medium distance in real space. In addition, the sound signal processing device 1 can reduce the processing load by mixing the voices of multiple audiences and processing them, compared to processing the voices of all audiences individually.
[0034] The sound signal processing device 1 performs binaural processing on the second audio signals of viewers L1, L2, L5, and L6, classified as the third group G3, to localize them at a more distant position. The sound signal processing device 1 may, for example, mix the second audio signals of viewers L1, L2, L5, and L6, classified as the third group G3, and then perform filtering on the mixed second audio signals. The filtering on the second audio signals of the third group G3 is set to stronger parameters (e.g., a higher cutoff frequency) than the filtering on the second audio signals of the second group G2. The sound signal processing device 1 may also perform reverb processing on the mixed second audio signals. The reverb processing on the second audio signals of the third group G3 is set to stronger parameters (e.g., a longer decay time) than the reverb processing on the second audio signals of the second group G2. In addition, the sound signal processing device 1 may perform level reduction processing on the mixed second audio signals. Level reduction processing is a process that reduces the level inversely proportional to the square of the distance between the listening position and the sound source.
[0035] As a result of the above sound processing, the voices of the audience in the third group G3 are perceived as being even further away than those of the audience in the second group G2, reducing the intelligibility of their speech and diminishing their sense of direction. Therefore, the first presenter A1 can perceive the distance to distant audiences in real space. Furthermore, the sound signal processing device 1 can significantly reduce the processing load for distant audiences by mixing multiple voices and performing sound processing on them.
[0036] The sound signal processing device 1 outputs the first sound signal and the processed second sound signal from the audio interface 71 to the headphones 81 worn by the first performer (S17). At this time, the sound signal processing device 1 outputs the mixed sound signal of the first sound signal and the second sound signal to the audio interface 71. The mixing balance of the first sound signal and the second sound signal may be adjusted, for example, by the first performer A1 via a GUI.
[0037] Furthermore, the sound signal processing device 1 may perform reverb processing on the first sound signal, the second sound signal, or the mixed sound signal to correspond to the reverberation of the virtual space 101.
[0038] In real space, ambient sounds such as environmental sounds and cheers are each individual sound sources. To naturally represent these ambient sounds in a virtual space, ideally, binaural processing should be applied to the sound signals of each sound source. However, applying binaural processing to the sounds of numerous sound sources significantly increases the processing load. In contrast, the sound signal processing device 1 of this embodiment applies sound processing to each of the multiple groups classified according to the location of the ambient sounds, according to the positional characteristics of each group. This allows for a natural representation of ambient sounds while keeping the processing load down. As a result, users can experience a new customer experience in a virtual space where they can feel natural ambient sounds similar to those in real space.
[0039] Group classification may be based on distance. That is, multiple ambient sounds may be classified into one of several groups based on the distance of the second position information to the first position information. Figure 7 is a conceptual diagram showing another example of group classification. In the example in Figure 7, the sound signal processing device 1 sets the range of concentric circles at a first distance (short distance) centered on the first performer A1 as the first group G1, the range of concentric circles at a second distance (medium distance) as the second group G2, and the range at a third distance (long distance) as the third group G3.
[0040] In this case as well, the first speaker A1 can perceive the identifiability and directional nature of the audio content according to the distance, and can experience natural ambient sound similar to that of a real space.
[0041] (Variation 1) The sound signal processing device 1 according to Modification 1 analyzes the performance sound of the first performer A1, which is the content sound, and performs sound processing on the first or second sound signal based on the analysis results. The sound signal processing device 1 analyzes changes in musical elements of the performance sound, such as changes in rhythm or changes from the A section to the chorus. When the sound signal processing device 1 detects, for example, the timing of the rhythm speeding up or the timing of the change to the chorus, it increases the level of the sound signal related to the ambient sound. Alternatively, when the sound signal processing device 1 detects, for example, the timing of the rhythm speeding up or the timing of the change to the chorus, it may change the type of ambient sound of the cheers. The timing of changes in musical elements is detected, for example, by inputting the performance sound into a trained model that has been trained on the relationship between the performance sound (content sound) and the content information. The sound signal processing device 1 acquires a large number of performance sounds and content information corresponding to each performance sound in advance, and trains a predetermined model using a predetermined algorithm to train the relationship between the performance sound and the content information. Furthermore, the algorithm used to train the model is not limited; any machine training algorithm such as CNN (Convolutional Neural Network) or RNN (Recurrent Neural Network) can be used. The machine training algorithm may be supervised training, unsupervised training, semi-supervised training, reinforcement training, inverse reinforcement training, active training, or transfer training, etc.
[0042] The sound signal processing device 1, as an example, uses the pre-trained model described above to estimate the song title of the content corresponding to the input performance sound. The sound signal processing device 1 uses a pre-trained model that has been trained on the relationship between performance sounds and song titles to estimate the song title corresponding to the input performance sound. The sound signal processing device 1 acquires information about the content corresponding to the estimated song title. The content information may be, for example, sheet music information or audio data. The sound signal processing device 1 estimates the current performance position based on the input performance sound and content information. The sound signal processing device 1 applies sound processing to the first or second sound signal based on the estimated performance position. The sound signal processing device 1, for example, increases the level of the sound signal related to ambient sound at the timing when the song changes to the chorus.
[0043] In this case, the first performer A1 can realistically perceive the excitement generated by their own performance and experience a more natural ambient sound that resembles that of a real space. Furthermore, the sound signal processing device 1 may increase the level of the sound signal related to the ambient sound, for example, as the degree of agreement with the chord progression and rhythm detected from the content information increases. In this case, the first performer A1 can feel the excitement as the degree of agreement with the example increases, and can perform in a game-like manner.
[0044] (Modification 2) In the above embodiment, the sound signal processing device 1 acquires the user's voice, obtained by the microphone 90 of each viewer terminal, as a reaction sound signal. In the modified example 2, the sound signal processing device 1 receives reaction information from multiple viewer terminals, for example. The sound signal processing device 1A, which is a viewer terminal, displays icon images such as "cheers," "applause," "calls," and "murmuring" on a display. The sound signal processing device 1A receives reaction information (text information) such as "cheers," "applause," "calls," and "murmuring" from the user via the user I / F 32 by accepting selection operations for these icon images. The sound signal processing device 1A transmits the received reaction information to the sound signal processing device 1. The sound signal processing device 1 generates a reaction sound signal corresponding to the received reaction information. The sound signal processing device 1 classifies the acquired reaction sound signals into one of several groups as ambient sounds, applies sound processing to the reaction sound signals classified into one of the several groups, and outputs the reaction sound signal after sound processing.
[0045] As a result, the sound signal processing system can send and receive reaction information with less information than the reaction sound signal, thus providing natural ambient sound that resembles real space while keeping the communication load down.
[0046] Reaction information may also be obtained as sensing information obtained by sensing the viewer's reaction. The sound signal processing device 1A may estimate information such as the viewer's position, orientation, and motion based on an image of the viewer captured by an image sensor such as a camera, and obtain reaction information. Alternatively, the sound signal processing device 1A may estimate information such as the viewer's position, orientation, and motion based on sensors such as a position sensor and motion sensor worn by the viewer, and obtain reaction information.
[0047] Furthermore, the sound signal processing device 1 may operate the 3D model based on the viewer's position, orientation, motion, etc., according to the acquired reaction information. This allows the first performer A1, while in the virtual space 101, to visually recognize the viewer's reactions and perform a more realistic session or other performance.
[0048] (Variation 3) The sound signal processing device 1 according to Modification 3 acquires the reaction sound signal using a trained model that has been trained on the relationship between the content sound and the reaction sound signal.
[0049] As a preparatory step, the sound signal processing device 1 acquires a large number of content sounds and reaction sound signals corresponding to each content sound, and trains a predetermined model using a predetermined algorithm to understand the relationship between the content sounds and the reaction sound signals.
[0050] The sound signal processing device 1 uses the trained model described above to acquire reaction sound signals corresponding to the input performance sound. For example, at the timing when the music changes to the chorus, the sound signal processing device 1 acquires reaction sound signals related to the ambient sound of numerous claps.
[0051] In this case, the sound signal processing device 1 does not need to receive reaction sound signals or reaction information from a large number of viewer terminals. Therefore, the sound signal processing system can provide natural ambient sound that resembles real space while keeping the communication load down.
[0052] (Modification 4) The sound signal processing device 1 according to Modification 4 acquires reaction sound signals corresponding to acoustic space information. The sound signal processing device 1 acquires reaction sound signals that match the magnitude of the acoustic space information, for example. If the acoustic space information indicates a small acoustic space, the sound signal processing device 1 acquires reaction sound signals related to the ambient sound of a small number of claps. If the acoustic space information indicates a large acoustic space, the sound signal processing device 1 acquires reaction sound signals related to the ambient sound of a large number of claps.
[0053] The relationship between acoustic spatial information and reaction sound signals may be stored in the flash memory 33 as a table beforehand, but the sound signal processing device 1 may acquire reaction sound signals using a predetermined trained model. As a preparatory step, the sound signal processing device 1 acquires a large amount of acoustic spatial information and the reaction sound signals corresponding to each piece of acoustic spatial information, and trains a predetermined model using a predetermined algorithm to understand the relationship between acoustic spatial information and reaction sound signals. The sound signal processing device 1 then uses the trained model to acquire reaction sound signals corresponding to the input acoustic spatial information.
[0054] In this case as well, the sound signal processing device 1 does not need to receive reaction sound signals or reaction information from a large number of viewer terminals. Therefore, the sound signal processing system can provide natural ambient sound that resembles real space while keeping the communication load down.
[0055] (Variation 5) Figure 8 is a plan view showing the configuration of the virtual space 101 according to Modification 5. The sound signal processing device 1 of Modification 5 acquires third position information of the sound source included in the first sound signal and performs sound processing on the first sound signal based on the first position information and the third position information. In the example of Figure 8, the third position information includes information indicating the coordinates of the virtual guitar amplifier 82V located to the right rear of the virtual space 101 as seen from the perspective of the first performer A1.
[0056] The sound signal processing device 1 applies binaural processing to the first sound signal based on the first position information and the third position information. As binaural processing, the sound signal processing device 1 convolves an HRTF into the first sound signal such that it is localized to a position to the right rear of the first performer A1 (relative coordinates of the virtual guitar amplifier 82V with the coordinates of the first performer A1 as the origin).
[0057] First performer A1, by listening to the first sound signal after binaural processing with headphones 81, can perceive that they are located in the center of the virtual space 101 and that their own performance sounds are coming from a different location (to the right rear) than their actual position in the virtual space 101.
[0058] In a performance in a real space, the first performer A1 hears the sound of their own performance emanating from sound equipment such as a guitar amplifier 82, or from an instrument such as a guitar 83. If the sound of their own performance, such as the sound of the guitar amplifier 82 or guitar 83, were localized inside the performer's head, it would create a sense of unease. In contrast, the sound signal processing device 1 of Modification 5 convolves the transmission characteristics so that the sound is heard from a different position than the first performer A1 in the virtual space. Therefore, the first performer A1, while in the virtual space 101, can perceive the sound of their own guitar 83 as coming from the guitar amplifier 82 located to their right and behind them, providing a new customer experience that allows for a realistic session in a virtual space, something that was not possible with conventional methods.
[0059] However, the position from which the sound of one's own performance is localized is not limited to the position of sound equipment or instruments installed in the actual space. The sound signal processing device 1 of this embodiment can also hear the sound of one's own performance from any position. For example, the sound signal processing device 1 receives an instruction from the first performer A1 via the GUI to change the coordinates of the virtual guitar amplifier 82V to the left rear of the first performer A1. As a first binaural process, the sound signal processing device 1 convolves an HRTF into the first sound signal that localizes it to the left rear of the user. As a result, the first performer A1 can perceive that they are in the center of the virtual space 101 and are hearing the sound of their own performance from the left rear of themselves.
[0060] (Experimental variation 6) The sound signal processing device 1 according to Modification 6 estimates the type of sound source. The sound signal processing device 1 detects this by inputting the performance sound (content sound) into a trained model that has been trained to recognize the relationship between the performance sound and the type of sound source. The sound signal processing device 1 acquires a large number of performance sounds and the type of sound source information corresponding to each performance sound in advance, and trains a predetermined model using a predetermined algorithm to recognize the relationship between the performance sound and the type of sound source information.
[0061] The sound signal processing device 1 uses the trained model to estimate the type of sound source corresponding to the input performance sound. Based on the estimated type of sound source, the sound signal processing device 1 performs sound processing on the first sound signal. For example, if the sound signal processing device 1 estimates that the input first sound signal is an electric guitar, it convolves an HRTF into the first sound signal such that it is localized to a position to the right rear of the user (relative coordinates of the virtual guitar amplifier 82V with the coordinates of the first performer A1 as the origin), as shown in Figure 8. For example, if the sound signal processing device 1 estimates that the input first sound signal is an acoustic guitar, it convolves an HRTF into the first sound signal such that it is localized to the relative position of the acoustic guitar with the coordinates of the first performer A1 as the origin (for example, the position of the user's chest).
[0062] This allows the sound signal processing device 1 to localize the sound to the optimal position for each type of sound source.
[0063] The localization position for each type of sound source may be determined by referring to a table predetermined for each type of sound source, or it may be determined based on the coordinates of musical instruments and sound equipment objects placed in the virtual space 101. Alternatively, the sound signal processing device 1 may determine the localization position for each type of sound source based on images of performers, musical instruments, or sound equipment captured by, for example, an image sensor such as a camera. Or, the sound signal processing device 1 may estimate information about the performer's motion and determine the localization position for each type of sound source based on the estimated motion information. For example, when a performer makes a motion such as clapping or snapping their fingers, the sound signal processing device 1 determines the localization position of the sound source at the position of the performer's hands. Also, when a performer makes a motion such as moving their mouth, the sound signal processing device 1 determines the localization position of the sound source at the position of the performer's mouth.
[0064] Furthermore, the sound signal processing device 1 may apply a filter to the first sound signal that simulates the acoustic characteristics of estimated audio equipment (effects pedals, amplifiers, speakers, etc.). For example, the sound signal processing device 1 applies digital signal processing to the first sound signal that simulates the output characteristics for each audio device as a digital filter. This allows the sound signal processing device 1 to reproduce the input and output characteristics of actual audio equipment for the sound being played.
[0065] (Example 7) The operation of the sound signal processing device 1 described above may be performed not only on the performer terminal used by the first performer A1, but also on the performer terminals used by other performers. Figure 9 is a block diagram showing the configuration of the sound signal processing system according to Modification 7, which further includes other performer terminals. Figure 10 is a plan view showing the configuration of the virtual space 101 according to Modification 7. In Figure 10, as an example of a performance, the first performer A1 plays the guitar 83, and the second performer A2 sings.
[0066] The sound signal processing system according to Modification 7 includes a first presenter terminal used by the first presenter A1 and a second presenter terminal used by the second presenter A2. The configuration of the second presenter terminal used by the second presenter A2 is the same as that of the sound signal processing device 1 shown in Figure 2.
[0067] The first presenter terminal and the second presenter terminal each receive location information of other presenters. The location information of other presenters may be received, for example, via a GUI, or it may be received from the first presenter terminal via the network from the other presenter terminals. In the example in Figure 10, the location information of the second presenter A2 includes information indicating the coordinates of the left center of the virtual space 101.
[0068] The first performer terminal and the second performer terminal transmit audio signals related to their respective performances. The first performer terminal receives audio signals related to the singing of the second performer A2 from the second performer terminal. The second performer terminal receives audio signals related to the singing of the first performer A1 from the first performer terminal.
[0069] The first performer terminal and the second performer terminal each apply binaural processing to the audio signals received from the audio I / F71 and the audio signals received via the network, respectively.
[0070] The first and second performer terminals each output the audio signal, after binaural processing, from the audio I / F71.
[0071] Furthermore, the first and second performer terminals may each render CG images of the virtual space and objects such as performers to generate video of the virtual space viewed in a predetermined direction from each performer's viewpoint. The first and second performer terminals may each output the rendered video to the HMD51 via the display I / F31.
[0072] As a result, the first performer A1 and the second performer A2 can perceive themselves as being in the virtual space 101, listening to the sounds of each other's performances, and perceiving themselves as being in the same space and conducting a session.
[0073] The binaural processing described above may include processing to localize indirect sounds. For example, the sound signal processing device 1 uses the coordinates of the first performer A1 in the virtual space 101 as the listening position and the coordinates of the virtual guitar amplifier 82V as the sound source position to acquire an impulse response corresponding to the reverberation of the virtual space 101. The impulse response is acquired by measurement, for example, by placing a dummy head at the listening position in a real space corresponding to the virtual space 101. The sound signal processing device 1 performs processing to localize indirect sounds by convolving the acquired impulse response into the first sound signal. Indirect sounds are sounds that reach the listening position after the sound from the sound source reflects off the walls of the virtual space, and include early reflections with a fixed direction and phase, or reverberations with random direction and phase. The localization processing of indirect sounds by binaural processing may only be performed on the early reflection system, and localization processing may be performed on the reverberation system using effects such as reverb.
[0074] Alternatively, the impulse response may be obtained by simulation based on, for example, the ray method or the virtual image method. The ray method is a technique that tracks the trajectory (ray) of sound radiated from a sound source and calculates the time pattern of the energy of the ray passing through the listening position. A simulation using the ray method determines the direction, arrival time, and arrival level from each virtual sound source at the listening position, based on the energy of the ray in the listening region, assuming that each ray is a virtual sound image of reverberation. The virtual image method is a technique that creates a virtual image (virtual sound source) of the sound source against the wall surface of the space as a virtual sound source, and determines the direction, arrival time, and arrival level from each virtual sound source at the listening position. The sound signal processing device 1 may generate an impulse response of a head-related transfer function that represents the direction, arrival time, and arrival level of each virtual sound source obtained by simulation, and perform a process to localize indirect sound by convolving the impulse response into the first sound signal.
[0075] Alternatively, the sound signal processing device 1 may perform a process to localize indirect sound by applying a level delay filter to the first sound signal, which has a delay amount and attenuation amount corresponding to each virtual sound source determined by simulation.
[0076] Similarly, the sound signal processing device 1 may acquire an impulse response corresponding to the reverberation of the virtual space 101, using the coordinates of the first performer A1 in the virtual space 101 as the listening position and the coordinates of the second performer A2 as the sound source position. As described above, the impulse response may be acquired by measurement or by simulation. The sound signal processing device 1 performs a process to localize the indirect sound by convolving the acquired impulse response with the sound signal related to the performance of the second performer. Alternatively, the sound signal processing device 1 may perform a process to localize the second indirect sound by applying a level delay filter process to the sound signal related to the performance of the second performer, which has delay and attenuation amounts corresponding to each virtual sound source obtained by simulation.
[0077] This allows the first performer A1 to perceive not only direct sounds but also indirect sounds occurring in a real space, enabling them to have a more realistic session with the second performer A2 while remaining in the virtual space 101.
[0078] The operation to localize indirect sound may be performed not only on the first presenter's terminal used by the first presenter A1, but also on the second presenter's terminal used by the second presenter A2.
[0079] The positional information for each performer and viewer may include orientation information. In this case, binaural processing further includes localization processing based on orientation information.
[0080] The orientation information may include the direction each performer is facing, as well as the direction the audience is facing, the direction the instrument is facing, or the direction the sound equipment is facing. The sound signal processing device 1 convolves the impulse response of the head transfer function corresponding to the orientation of the audience, instrument, or sound equipment relative to the orientation of the first performer A1 into the first sound signal.
[0081] The sound levels of performances and vocals are highest when the front of the sound source and the front of the listener are facing each other, and decrease as the left-right angle increases. Also, as the left-right angle increases, the high frequencies are attenuated more than the low frequencies. Therefore, the sound signal processing device 1 may perform gain correction such that the greater the difference (angle difference) between the direction of the first performer A1 and the direction the listener is facing, the direction the instrument is facing, or the direction the sound equipment is facing, the lower the level of the sound signal. The sound signal processing device 1 may also perform equalizer processing such that the greater the difference (angle difference) between the direction of the first performer A1 and the direction the listener is facing, the direction the instrument is facing, or the direction the sound equipment is facing, the lower the level of the high frequencies or the higher the level of the low frequencies. The sound signal processing device 1 may also perform gain correction such that the greater the difference (angle difference) between the direction of the first performer A1 and the direction the listener is facing, the direction the instrument is facing, or the direction the sound equipment is facing, the lower the level of the sound signal. Furthermore, the sound signal processing device 1 may perform equalizer processing such that the greater the difference (angle difference) between the orientation of the first performer A1 and the orientation of the audience, the orientation of the instrument, or the orientation of the sound equipment, the lower the level of high frequencies or the higher the level of low frequencies.
[0082] This allows the sound signal processing device 1 to perceive changes in the relative positions of the performer and the sound source in real time. The sound signal processing device 1 can represent the relative positions of the performer and the sound source with greater precision. Therefore, the first performer A1 can perceive ambient sounds more realistically while in the virtual space 101. Alternatively, the first performer A1 can have a more realistic session with the second performer A2 while in the virtual space 101.
[0083] The operation of such an audio signal processing device 1 may be performed not only on the first presenter's terminal used by the first presenter A1, but also on the second presenter's terminal used by the second presenter A2.
[0084] The sound signal processing device 1 may receive sensing information that senses the performance of the first performer, the second performer, or the audience.
[0085] Sensing information is acquired by each sensor. For example, the sound signal processing device 1 estimates information such as the position, orientation, and motion of the first performer A1 based on an image of the first performer A1 taken by an image sensor such as a camera, and acquires first sensing information. Alternatively, the sound signal processing device 1 may estimate information such as the position, orientation, and motion of the first performer A1 based on sensors such as a position sensor, motion sensor, and head tracker worn by the first performer A1, and acquire first sensing information.
[0086] The sound signal processing device 1 synchronizes and plays back video related to the performance of the first performer, the second performer, or the viewer based on the acquired sensing information. Specifically, the sound signal processing device 1 operates the 3D models of the first performer A1, the second performer A2, or the viewers L1 to L14 based on information such as their respective positions, orientations, and motions.
[0087] As a result, the first performer A1 is located in the virtual space 101, and the second performer A2, or viewers L1 through L14, can visually perceive the motion of the second performer A2 and conduct a realistic session.
[0088] The operation of the sound signal processing device 1 described above may be performed not only on the first presenter's terminal used by the first presenter A1, but also on the second presenter's terminal used by the second presenter A2.
[0089] The sound signal processing device 1 may perform correction of the sound environment at each point. For example, the sound signal processing device 1 performs processing to reduce the reverberation of the actual space in which the first performer A1 is located. Information on the reverberation of the actual space in which the first performer A1 is located is obtained in advance by measuring the impulse response. In binaural processing, the sound signal processing device 1 reduces the reverberation of the actual space in which the first performer A1 is located by convolving the inverse characteristics of the impulse response measured in advance.
[0090] Furthermore, for example, the sound signal processing device 1 may perform a process to cancel out noise in the actual space where the first performer A1 is located. Also, the sound signal processing device 1 may perform a process to cancel out the sound of the guitar that is generated in the actual space where the first performer A1 is located.
[0091] This allows the first presenter, A1, to conduct the session with a greater sense of immersion, perceiving themselves as being in virtual space 101.
[0092] The operation of the sound signal processing device 1 described above may be performed not only on the first presenter's terminal used by the first presenter A1, but also on the second presenter's terminal used by the second presenter A2.
[0093] (Variation 8) Figure 11 is a block diagram showing the configuration of the sound signal processing system according to Modification 8. The sound signal processing system according to Modification 8 has an administrator terminal in addition to the configuration shown in Figure 9. The administrator terminal is connected to the first performer terminal, the second performer terminal, and the audience terminal via a network. The administrator terminal has the same configuration as the sound signal processing device 1 shown in Figure 2.
[0094] In the audio signal processing system according to Modification 8, the first performer terminal and the second performer terminal each transmit audio signals related to their performance to the administrator terminal. The administrator terminal generates audio signals to be distributed to viewer terminals based on the audio signals related to the performance received from the first and second performer terminals. The administrator terminal also receives information indicating the viewing position from multiple viewer terminals.
[0095] The administrator terminal processes the audio signals related to the performance received from the first performer terminal and the second performer terminal. For example, the administrator terminal performs binaural processing on the audio signals received from the first performer terminal and the audio signals received from the second performer terminal. The administrator terminal also acquires audio signals related to ambient sound and performs binaural processing on the acquired ambient sound signals.
[0096] The administrator terminal mixes the binaurally processed audio signals and distributes them to the viewer terminals. In this case, the audio signals distributed to the viewer terminals undergo high-processing binaural processing on the administrator terminal. Therefore, the administrator terminal can provide a realistic performance viewing experience in a virtual space, regardless of the processing power of the viewer terminals.
[0097] The administrator terminal may synchronize the audio signals related to the performances of multiple performers before mixing them. For example, the administrator terminal estimates information about the content being played based on the audio signals related to the performance received from the first performer terminal or the second performer terminal. The content information may be, for example, musical score information or audio data. The administrator terminal estimates the current performance position of each performer based on the audio signals related to the performance received from the first and second performer terminals. Based on the estimated performance positions, the administrator terminal delays the audio signals related to the performance received from the first and second performer terminals to match the performance timing and mix them.
[0098] As a result, the audio signals related to the performances of multiple performers are synchronized before being broadcast, so even if multiple performers are performing at different locations, viewers can perceive them as being in the same space and performing together.
[0099] (Extreme variation 9) The sound processing conditions for each terminal may be accepted via a GUI, or they may be accepted automatically according to the orientation (gaze) of each user detected by various sensors such as motion sensors. The sound signal processing device 1 may perform processing to make the level of the sound corresponding to other performers or viewers located in the direction that the first performer A1 is facing the highest, or to emphasize that sound.
[0100] In real space, people unconsciously concentrate on listening to sounds coming from the direction they are facing (sounds they are paying attention to), creating a cocktail party effect. However, it is difficult to create a cocktail party effect in virtual space. In contrast, the sound signal processing device 1 of this embodiment can realize a cocktail party effect in virtual space by processing to maximize the level of sound corresponding to other performers or viewers located in the direction that the first performer A1 is facing, or by emphasizing such sound. This enables communication among any participants, including performers or viewers.
[0101] (Variation 10) The audio signal processing system may perform audio processing individually at each terminal, as shown in Figure 1, or it may perform some audio processing at a server such as an administrator terminal, as shown in Figure 11, and then distribute the processed audio signal to the viewer terminals. Alternatively, the audio signal processing system may perform some audio processing at an edge server and then distribute the processed audio signal to the viewer terminals. Figure 12 is a block diagram showing the configuration of the audio signal processing system according to Modification 10. The audio signal processing system according to Modification 10 has an edge server 1D in addition to the configuration shown in Figure 11. The edge server 1D is connected to the first performer terminal, the second performer terminal, the administrator terminal, and the viewer terminals via a network. The configuration of the edge server 1D is the same as that of the audio signal processing device 1. However, the edge server 1D does not need to have a display I / F 31 and a USB I / F 37.
[0102] In the sound signal processing system according to Modification 10, the edge server 1D performs some or all of the sound processing. For example, the administrator terminal transmits acoustic space information, first location information, and second location information to the edge server. The first performer terminal transmits a first sound signal to the edge server 1D, and the second performer terminal transmits a second sound signal to the edge server 1D.
[0103] Edge server 1D performs first binaural processing on the first audio signal received from the first performer terminal and second binaural processing on the second audio signal received from the second performer terminal. Edge server 1D distributes the first binaural signal and the second binaural signal to the viewer terminal.
[0104] In this case, the computationally intensive binaural processing is performed on edge server 1D, which is closer to the viewer's device. As a result, viewers can enjoy a more lag-free viewing experience.
[0105] Next, Figure 13 is a conceptual diagram illustrating the configuration of an audio signal processing system when providing individual environments. The positions of the performers and viewers within the virtual space 101 may be the same or different for all terminals. For example, in the example in Figure 13, at the first point of the first performer A1, the first performer A1 is located at the center of the virtual space 101, the virtual guitar amplifier 82V is located to the right rear of the first performer A1, the second performer A2 is located in the front center, and the viewer L1 is located to the left. At the second point of the second performer A2, the second performer A2 is located at the center of the virtual space 101, the virtual monitor speaker 90V is located in front of the second performer A2, the first performer A1 is located behind it, the virtual guitar amplifier 82V is located to the left front, and there is no viewer. At viewer L1's third location, viewer L1 is at the center of virtual space 101, virtual right main speaker 90RV is located to viewer L1's right front, virtual left main speaker 90LV is located to viewer L1's left front, virtual guitar amplifier 82V is located to the right of virtual left main speaker 90LV, second performer A2 is located to the right of virtual guitar amplifier 82V, and first performer A1 is located to the left of virtual right main speaker 90RV. Additionally, at viewer L1's third location, other viewers are located to viewer L1's left, right, and behind.
[0106] In this case, viewer L1 sees the first performer A1 and the second performer A2 in front of them, with the sound of the first performer's guitar playing localized at the position of the virtual guitar amplifier 82V, and the sound of the second performer A2's singing localized by panning using the virtual right main speaker 90RV and the virtual left main speaker 90LV. In addition, viewer L1 localizes the sounds related to the reactions of other viewers to their left, right, and behind themselves.
[0107] The audio settings for viewer L1, such as localization position and level balance, may be specified by viewer L1 or by the administrator of the administrator terminal. Alternatively, the audio settings for each viewer may be automatically configured by the viewer terminal or administrator terminal by storing or learning past information.
[0108] In other words, the sound signal processing system (sound signal processing device or sound signal processing method) in the example shown in Figure 13 has the following technical concept. The performer's terminal receives acoustic space information, first position information indicating the listening position in the acoustic space, and second position information indicating the positions of multiple ambient sounds in the acoustic space. The system receives a first audio signal of the content sound and a second audio signal relating to the plurality of ambient sounds, Based on the first and second position information, each of the plurality of ambient sounds is classified into one of the plurality of groups. Each of the second sound signals classified into the aforementioned multiple groups is subjected to the respective sound processing for each classified group. The first sound signal and the second sound signal after sound processing are output. The viewer terminal receives the acoustic space information, first position information indicating the listening position in the acoustic space, and second position information indicating the positions of multiple ambient sounds in the acoustic space. Upon receiving the first sound signal and the second sound signal, The first signal or the second signal is subjected to sound processing, The first sound signal and the second sound signal after sound processing are output. The parameters for sound processing in the performer's terminal and the parameters for sound processing in the viewer's terminal are different. Furthermore, if there are multiple performer terminals, the sound processing parameters may differ for each performer terminal.
[0109] By performing individual sound settings for each performer and viewer as described above, it becomes possible to realize sessions and performances that are unique to virtual spaces and cannot be achieved in real-world spaces.
[0110] In the above embodiment, binaural processing was shown as an example of localization processing, but localization processing is not limited to binaural processing. For example, panning processing and beamforming processing using multiple speakers are also examples of localization processing. Furthermore, the sound signal after binaural processing may be played back not only through headphones, but also through multiple speakers. When playing back the sound signal after binaural processing through multiple speakers, it is preferable to perform crosstalk cancellation processing.
[0111] In the above embodiment, a session by a first performer A1 and a second performer A2 was used as an example, but performances via a network at different locations are not limited to this example. The sound signal processing system may, for example, process the sound of a performance by even more performers. The sound signal processing system may also process the sound of a group performance, such as an orchestra, taking place in a real space, and a performance by another performer in a remote location. In this case as well, the sound signal processing system processes the sound of the other performer's performance to localize it to a predetermined position, so that a viewer in a remote location can perceive that a group performance, such as an orchestra, taking place in a real space is simultaneously being performed by an even more remote performer.
[0112] In the above embodiment, an example was shown in which a general-purpose information processing device (sound signal processing device 1) performs binaural processing. However, if, for example, an audio I / F 71 is connected to a network, the audio I / F 71 may perform binaural processing. In other words, the sound signal processing device of the present invention can also be realized by an audio I / F that can be connected to a network.
[0113] The description of this embodiment should be considered in all respects to be illustrative and not restrictive. The scope of the invention is indicated by the claims, rather than by the embodiments described above. Furthermore, the scope of the invention includes the scope equivalent to the claims. [Explanation of Symbols]
[0114] 1,1A: Audio signal processing unit, 1D: Edge server, 31: Display I / F, 32: User I / F, 33: Flash memory, 34: CPU, 35: RAM, 36: Communication I / F, 37: USB I / F, 51: HMD, 71: Audio I / F, 81: Headphones, 82: Guitar amplifier, 82V: Virtual guitar amplifier, 83: Guitar, 90: Microphone, 90LV: Virtual left main speaker, 90RV: Virtual right main speaker, 90V: Virtual monitor speaker, 101: Virtual space
Claims
1. The system receives acoustic space information, first position information indicating the listening position in the acoustic space, and second position information indicating the positions of multiple ambient sounds in the acoustic space. The system receives a first audio signal of the content sound and a second audio signal relating to the plurality of ambient sounds, Based on the first and second position information, each of the plurality of ambient sounds is classified into one of the plurality of groups. Each of the second sound signals classified into the aforementioned multiple groups is subjected to the respective sound processing for each classified group. The first sound signal and the second sound signal after sound processing are output. Audio signal processing method.
2. Each of the aforementioned ambient sounds is classified into one of the aforementioned groups based on the distance of the second position information to the first position information. The sound signal processing method according to claim 1.
3. The aforementioned multiple ambient sounds include, The sound signal processing method according to claim 1 or claim 2.
4. The aforementioned content sound is analyzed, Based on the analysis results, the aforementioned sound processing is performed. The sound signal processing method according to claim 1 or claim 2.
5. We acquire reaction sound signals related to viewer reactions, The reaction sound signals are classified into one of the plurality of groups as ambient sounds. The reaction sound signals, classified into one of the aforementioned groups, are subjected to sound processing. Outputs the reaction sound signal after the aforementioned sound processing. The sound signal processing method according to claim 1 or claim 2.
6. The reaction sound signal is acquired using a trained model that has been trained to understand the relationship between the content sound and the reaction sound signal. The sound signal processing method according to claim 5.
7. The reaction sound signal corresponding to the acoustic spatial information is acquired. The sound signal processing method according to claim 5.
8. The third position information of the sound source included in the first sound signal is acquired, The first sound signal is subjected to sound processing based on the first position information and the third position information. The sound signal processing method according to claim 1 or claim 2.
9. The type of sound source is estimated, The first sound signal is subjected to sound processing based on the estimated type. The sound signal processing method according to claim 8.
10. The system receives acoustic space information, first position information indicating the listening position in the acoustic space, and second position information indicating the positions of multiple ambient sounds in the acoustic space. The system receives a first audio signal of the content sound and a second audio signal relating to the plurality of ambient sounds, Based on the first and second position information, each of the plurality of ambient sounds is classified into one of the plurality of groups. Each of the second sound signals classified into the aforementioned multiple groups is subjected to the respective sound processing for each classified group. The first sound signal and the second sound signal after sound processing are output. A sound signal processing device equipped with a processor.
Citation Information
Patent Citations
Information processing system, information processing method, and program
WO2022196073A1