Signal processing device, signal processing method, and program

The signal processing apparatus enhances sound source separation by generating and processing acoustic signals of virtual sound sources based on power differences, addressing the challenge of multiple sound sources in environments where microphones cannot be placed close, thereby improving the sense of presence and realism in sound reproduction.

WO2026070343A1PCT designated stage Publication Date: 2026-04-02SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In environments where microphones cannot be placed close to sound sources, such as during sports viewing, the acoustic signals collected include multiple sound sources, making it difficult to separate and reproduce a desired object sound source accurately, which impairs the sense of presence during stereophonic reproduction.

Method used

A signal processing apparatus and method that generates acoustic signals of virtual sound sources at multiple positions, emphasizing and/or suppressing these signals based on their power differences to enhance separation and reproduction.

Benefits of technology

Improves the sense of presence by enhancing the separation between virtual sound sources, improving the sense of distance and spaciousness of the sound image, and reducing the loss of realism in both surround and binaural sound systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025031750_02042026_PF_FP_ABST
    Figure JP2025031750_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to a signal processing device, a signal processing method, and a program that make it possible to prevent loss of a sense of realism. A virtual sound source generation unit generates acoustic signals of virtual sound sources at three or more virtual sound source positions from an input acoustic signal, and an enhancement / suppression unit enhances and / or suppresses the acoustic signals of the virtual sound sources to increase the difference in power on the basis of the power of the acoustic signals of the virtual sound sources. The present technology is capable of being applied, for example, when stereophonic sound is reproduced.
Need to check novelty before this filing date? Find Prior Art

Description

Signal processing apparatus, signal processing method, and program

[0001] The present technology relates to a signal processing apparatus, a signal processing method, and a program, and more particularly to a signal processing apparatus, a signal processing method, and a program that can suppress a loss of presence feeling, for example.

[0002] There is a technique for accurately separating a desired object (an acoustic signal emitted from a specific sound source) by extracting an acoustic signal (audio signal) of a selected video object by fixed beamforming (see, for example, Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2022-036998

[0004] An object sound source, that is, a sound source obtained by collecting sound emitted from each individual sound source with a microphone close to each individual sound source does not include an acoustic signal of another sound source (an acoustic signal corresponding to the sound emitted from another sound source).

[0005] However, in an environment where it is difficult to place a microphone close to a sound source, for example, in sound collection during sports viewing, the acoustic signal obtained by sound collection includes acoustic signals of a plurality of sound sources, and it is difficult to obtain an object sound source.

[0006] When performing stereophonic reproduction using, as it is, an acoustic signal that is not object-source-formed and includes an acoustic signal of another sound source in addition to the acoustic signal of a certain sound source, the sense of presence may be impaired.

[0007] The present technology has been made in view of such a situation, and can suppress a loss of presence feeling.

[0008] The signal processing apparatus or program of the present technology includes a virtual sound source generation unit that generates acoustic signals of virtual sound sources at three or more virtual sound source positions from an input acoustic signal, and based on the power of the acoustic signals of the virtual sound sources, emphasizes and / or suppresses the acoustic signals of the virtual sound sources so that the difference in power becomes large. It is a signal processing apparatus including an emphasizing / suppressing unit, or a program for causing a computer to function as such a signal processing apparatus.

[0009] The signal processing method of this technology includes generating acoustic signals of three or more virtual sound sources at virtual sound source positions from an input acoustic signal, and enhancing and / or suppressing the acoustic signals of the virtual sound sources based on the power of the acoustic signals of the virtual sound sources so that the power differences become larger.

[0010] In this technology, acoustic signals from three or more virtual sound sources at virtual sound source positions are generated from the input acoustic signal, and based on the power of the acoustic signals of the virtual sound sources, the acoustic signals of the virtual sound sources are enhanced and / or suppressed so that the power differences become larger.

[0011] A signal processing device may be a single, independent device, or it may be an internal block constituting a single independent device. Furthermore, a signal processing device can be composed of multiple independent devices.

[0012] The program can be provided by recording it on a recording medium or by transmitting it via a transmission medium.

[0013] This is a block diagram illustrating an example configuration of one embodiment of a signal processing system to which this technology is applied. This is a block diagram illustrating an example configuration of the sound pickup unit 11. This diagram illustrates the overview of signal processing performed in the signal processing unit 12 when 3D sound is reproduced using the ambisonic method. This diagram illustrates the overview of ambisonic sound pickup and reproduction by the signal processing system 10. This diagram illustrates how sound is heard during ambisonic playback when the acoustic signal of the virtual sound source at the virtual sound source location is an ideal acoustic signal. This diagram illustrates how sound is heard during ambisonic playback when the acoustic signal of the virtual sound source at the virtual sound source location is not an ideal acoustic signal. This diagram illustrates how sound is heard during ambisonic playback when the acoustic signal of the virtual sound source is enhanced and / or suppressed in the signal processing unit 12. This is a block diagram illustrating an example configuration of the signal processing unit 12. This is a flowchart illustrating the processing of the signal processing unit 12. This is a block diagram illustrating a first example configuration of the enhancement / suppression unit 33. This is a block diagram illustrating a second example configuration of the enhancement / suppression unit 33. This is a block diagram illustrating an example configuration of one embodiment of a computer to which this technology is applied.

[0014] <One embodiment of a signal processing system applying this technology>

[0015] Figure 1 is a block diagram showing an example configuration of one embodiment of a signal processing system to which this technology is applied.

[0016] In Figure 1, the signal processing system 10 includes a sound pickup unit 11, a signal processing unit 12, and an output unit 13. Each of the sound pickup unit 11 to the output unit 13 can be configured as a separate device.

[0017] The sound collection unit 11 has a device such as a microphone that senses sound, collects sound, and outputs an acoustic signal as an electrical signal.

[0018] The signal processing unit 12 functions as a signal processing device that processes the acoustic signal (input acoustic signal) output by the sound collection unit 11 and outputs the resulting acoustic signal. The signal processing unit 12 can be configured with, for example, a CPU or a DSP.

[0019] The output unit 13 consists of devices such as speakers and headphones that output sound, and outputs sound corresponding to the sound signal from the signal processing unit 12. Headphones include sound output devices that are used in contact with a person's ear, such as earphones and neck speakers, and sound output devices that are used in close proximity to a person's ear. If the output unit 13 consists of speakers, the output unit 13 consists of N speakers, which are 2 or 3 channels (or more).

[0020] In the signal processing system 10 configured as described above, the acoustic signal obtained by the sound collection unit 11 is processed by the signal processing unit 12, and the sound corresponding to the processed acoustic signal is output from the output unit 13, thereby reproducing three-dimensional sound.

[0021] Figure 2 is a block diagram showing an example configuration of the sound collection unit 11 in Figure 1.

[0022] The sound pickup unit 11 has two, three, or four or more M-channel microphones 21 and an A / D (analog to digital) converter 22. The microphones 21 pick up sound by sensing it and output the corresponding analog sound signal. The analog sound signal output by the microphones 21 is amplified by a microphone amplifier (not shown) and output to the subsequent A / D converter 22. The A / D converter 22 performs A / D conversion of the analog sound signal output by the microphones 21 and outputs a digital sound signal.

[0023] In the sound pickup unit 11, sound pickup and A / D conversion are performed by the M-channel microphone 21 and the A / D converter, resulting in the output of an M-channel (digital) acoustic signal.

[0024] In this embodiment, the signal processing system 10 reproduces three-dimensional sound using, for example, the ambisonics method. In this case, the sound pickup unit 11 employs a microphone for the ambisonics method. When a microphone for the ambisonics method is employed in the sound pickup unit 11, the number of channels (number) M of the microphone 21 is 4 for the first-order ambisonics method, 9 for the second-order ambisonics method, and (D+1)^2 for the D-order ambisonics method.

[0025] <Ambisonics method>

[0026] Figure 3 is a diagram illustrating the overview of the signal processing performed by the signal processing unit 12 when 3D sound reproduction is performed using the ambisonics method.

[0027] When 3D sound reproduction is performed using the ambisonics method, ambisonics sound recording is performed. That is, the sound recording unit 11 employs a microphone designed for the ambisonics method for sound recording. The acoustic signal obtained by ambisonics sound recording, i.e., the acoustic signal output by the ambisonics microphone, is called an A-format acoustic signal.

[0028] For example, if we adopt a first-order ambisonics system, then four microphones (=(1+1)^2) channels, ch1, ch2, ch3, and ch4, will be used for the ambisonics system. The four channels, ch1 to ch4, can be described as follows based on the placement (directivity) of the microphones in each channel.

[0029] Ch1 = FLU (Front Left Up) Ch2 = FRD (Front Right Down) Ch3 = BLD (Back Left Down) Ch4 = BRU (Back Right Up)

[0030] FLU, FRD, BLD, and BRU can refer to a channel, as well as the microphone associated with that channel or the audio signal output by that microphone.

[0031] The signal processing unit 12 converts the A-format FLU, FRD, BLD, and BRU channel acoustic signals into ambisonic B-format W, X, Y, and Z channel acoustic signals according to the following formula. W, X, Y, and Z may refer to channels as well as the acoustic signals of those channels.

[0032] W = FLU + FRD + BLD + BRU X = FLU + FRD - BLD - BRU Y = FLU - FRD + BLD - BRU Z = FLU - FRD - BLD + BRU

[0033] The W channel acoustic signal has an omnidirectional component, the X channel acoustic signal has a front-to-back spreading component, the Y channel acoustic signal has a left-to-right spreading component, and the Z channel acoustic signal has an up-to-down spreading component.

[0034] The signal processing unit 12 generates an acoustic signal of a virtual sound source at a virtual sound source position as an intermediate format acoustic signal from the acoustic signals of the W channel, X channel, Y channel, and Z channel.

[0035] The virtual sound source position is the location where the virtual sound source (a virtual sound source) is placed. This can be set in advance or set according to user operations (such as a listener or the operator of the signal processing system 10). The position of the speaker of the output unit 13 can also be set as the virtual sound source position.

[0036] The signal processing unit 12 generates, for example, acoustic signals S(θ, φ) for virtual sound sources at virtual sound source positions for 3 or more channels from the acoustic signals W, X, Y, Z of the W channel, X channel, Y channel, and Z channel, according to the following equation.

[0037] S(θ, φ) = fp ・ W + (1-fp) ・{ X sinθcosφ + Y sinθsinφ Y + Z cosθ}

[0038] φ represents the horizontal angle of the virtual sound source position relative to the front direction, and θ represents the vertical angle of the virtual sound source position relative to the direction directly above. The front direction and the direction directly above refer to, for example, the front direction and the direction directly above the microphone used for ambisonics recording. fp is a coefficient that adjusts the balance between omnidirectional and bidirectional patterns, and is set to a value between 0 and 1. When fp = 0.5, the directivity will be cardioid.

[0039] The signal processing unit 12 generates an acoustic signal S(θ, φ) for each of the 3 or more K virtual sound source positions, based on the angles φ and θ that represent the direction of the virtual sound source position.

[0040] The signal processing unit 12 uses the acoustic signals of virtual sound sources at the virtual sound source positions of K channels to perform surround sound playback such as 5.1ch or 7.1ch, and / or binaural sound playback. In other words, the signal processing unit 12 generates surround signals for surround sound systems such as 5.1ch or 7.1ch, and / or binaural signals for binaural sound systems from the acoustic signals of virtual sound sources at the virtual sound source positions of K channels.

[0041] For example, the signal processing unit 12 performs rendering to generate, as surround signals, the acoustic signals to be supplied to each of the N-channel (number) speakers constituting the output unit 13 from the acoustic signals of the virtual sound sources at the virtual sound source positions of K channels. The signal processing unit 12 outputs each of the N-channel surround signals obtained by rendering to the corresponding speaker. The rendering is performed so that when a listener (at an ideal position) listens to the sound from the speaker, it sounds as if there is a sound source at the virtual sound source position.

[0042] For example, the signal processing unit 12 applies the HRTF (head-related transfer function) corresponding to the virtual sound source position (convolves the HRIR (head-related impulse response)) to the acoustic signal of each virtual sound source at the virtual sound source positions of K channels, thereby generating a two-channel acoustic signal as a binaural signal in the binaural method.

[0043] Note that the signal processing unit 12 can generate, in addition to the surround signal in the surround method and the binaural signal in the binaural method, for example, an acoustic signal as a transcranial signal in the transcranial method.

[0044] FIG. 4 is a diagram for explaining an overview of ambisonics method sound collection and reproduction by the signal processing system 10.

[0045] In the sound collection by the ambisonics method, sound collection is performed by a microphone for the ambisonics method as the sound collection unit 11, and the acoustic signal obtained by the sound collection is output.

[0046] In FIG. 4, as seen from the microphone for the ambisonics method, an electric guitar, a vocal, an acoustic guitar, and a drum are respectively positioned at the front left, front right, rear left, and rear right. In the microphone for the ambisonics method, the sounds of the electric guitar, the vocal, the acoustic guitar, and the drum respectively positioned at the front left, front right, rear left, and rear right are collected, and the corresponding acoustic signals are output.

[0047] In the playback of the ambisonics method, in the signal processing unit 12, the directivity in the direction of the speaker constituting the output unit 13 is generated by signal processing from the acoustic signal output from the microphone for the ambisonics method as the sound collection unit 11. That is, in the signal processing unit 12, acoustic signals of virtual sound sources at virtual sound source positions of K channels (acoustic signals of virtual sound sources that reproduce the directivity at each of the virtual sound source positions of K channels) are generated by signal processing.

[0048] In FIG. 4, the output unit 13 is composed of 4-channel (unit) speakers and is arranged at the front left, front right, rear left, and rear right as viewed from the listener (at an ideal position). Further, in FIG. 4, as the virtual sound source positions of K channels, 4-channel virtual sound source positions are adopted, and the 4-channel virtual sound source positions respectively coincide with the positions of the 4-channel speakers. And, as the acoustic signal of the virtual sound source at the virtual sound source position that is the position of the front left speaker, the acoustic signal of an electric guitar as a sound source is generated, and as the acoustic signal of the virtual sound source at the virtual sound source position that is the position of the front right speaker, the acoustic signal of a vocal as a sound source is generated. As the acoustic signal of the virtual sound source at the virtual sound source position that is the position of the rear left speaker, the acoustic signal of an acoustic guitar as a sound source is generated, and as the acoustic signal of the virtual sound source at the virtual sound source position that is the position of the rear right speaker, the acoustic signal of a drum as a sound source is generated.

[0049] In FIG. 4, the acoustic signal of the electric guitar as the acoustic signal of the virtual sound source at the virtual sound source position that is the position of the front left speaker is output to the front left speaker. The acoustic signal of the vocal as the virtual sound source at the virtual sound source position that is the position of the front right speaker is output to the front right speaker. The acoustic signal of the acoustic guitar as the virtual sound source at the virtual sound source position that is the position of the rear left speaker is output to the rear left speaker. The acoustic signal of the drum as the virtual sound source at the virtual sound source position that is the position of the rear right speaker is output to the rear right speaker.

[0050] Therefore, the sound of an electric guitar is output from the front left speaker, and the sound of vocals is output from the front right speaker. The sound of an acoustic guitar is output from the rear left speaker, and the sound of drums is output from the rear right speaker. As a result, the listener can feel a sense of presence as if the sound sources were located at the virtual sound source locations. For example, the listener can enjoy the sensation that the sound of the electric guitar is coming from the front left speaker, the sound of vocals is coming from the front right speaker, the sound of the acoustic guitar is coming from the rear left speaker, and the sound of drums is coming from the rear right speaker.

[0051] Incidentally, in lower-order ambisonics systems such as first-order ambisonics, the virtual sound sources (and their acoustic signals) at each virtual sound source location are not sufficiently separated and are correlated with one another. As a result, the sound image (sound source) during playback is placed in an intermediate position due to the phantom center effect, which can impair the sense of presence. For example, during surround sound playback, the sense of distance and breadth of the sound image is impaired, resulting in a loss of presence. During binaural playback, the degree of external localization is significantly reduced, resulting in a loss of presence.

[0052] Figure 5 illustrates how sound is perceived during ambisonic playback when the acoustic signal of the virtual sound source at the virtual sound source location is an ideal acoustic signal.

[0053] For the acoustic signals of virtual sound sources at virtual sound source locations to be ideal acoustic signals, it means that the separation between virtual sound sources at each virtual sound source location is perfect and they are not correlated with one another. In other words, for the acoustic signals of virtual sound sources at virtual sound source locations to be ideal acoustic signals means that the acoustic signals of virtual sound sources at virtual sound source locations are object-based sound sources.

[0054] Figure 5A shows the sound pickup situation using the ambisonics method. In Figure 5A, a man is speaking to the left front and a woman is speaking to the right front, as viewed from the ambisonics microphone.

[0055] As in the case of Figure 4, the output unit 13 is assumed to consist of four channels of speakers positioned to the front left, front right, rear left, and rear right, respectively, as viewed from the listener. Furthermore, the positions of the speakers positioned to the front left, front right, rear left, and rear right are assumed to be the virtual sound source positions. The positions of the male and female voices correspond to the positions of the front left and front right speakers, respectively.

[0056] Figure 5B shows the sound output from the speaker when the acoustic signal of the virtual sound source at the virtual sound source location is an ideal acoustic signal.

[0057] An ideal acoustic signal for a virtual sound source at a virtual sound source location is when, for example, the signal processing unit 12 generates an acoustic signal corresponding to the sound emitted by the sound source at the time of sound recording (corresponding to the sound) as the acoustic signal for the virtual sound source at the virtual sound source location corresponding to the sound source location at the time of sound recording. Therefore, an ideal acoustic signal for a virtual sound source at a virtual sound source location is when, in the signal processing unit 12, an acoustic signal corresponding to the male voice is generated as the acoustic signal for the virtual sound source at the virtual sound source location corresponding to the female voice is generated as the acoustic signal for the virtual sound source at the virtual sound source location corresponding to the female voice at the right front speaker location, and silent acoustic signals are generated as the acoustic signals for the virtual sound source at the left rear and right rear speaker locations where no sound source is located.

[0058] If the acoustic signal of the virtual sound source at the virtual sound source position is an ideal acoustic signal, then the acoustic signal of the virtual sound source at the virtual sound source position corresponding to the front left speaker corresponds to a male voice. Furthermore, the acoustic signal of the virtual sound source at the virtual sound source position corresponding to the front right speaker corresponds to a female voice. In addition, the acoustic signals of the virtual sound sources at the virtual sound source positions corresponding to the rear left and rear right speakers are silent acoustic signals.

[0059] Therefore, only male voices (indicated by solid arrows in the diagram) are output from the front left speaker, and only female voices (indicated by dotted arrows in the diagram) are output from the front right speaker. No sound is output from the rear left and rear right speakers (silence is output).

[0060] Figure 5C shows how a listener perceives sound when the acoustic signal of a virtual sound source is an ideal acoustic signal.

[0061] As explained in Figure 5B, only male voices are output from the front left speaker, and only female voices are output from the front right speaker. No sound is output from the rear left and rear right speakers. As a result, the listener can experience a sense of presence as if the sound source were located at the virtual sound source location. For example, the listener can enjoy the sensation that the male voice is coming from the front left speaker and the female voice is coming from the front right speaker.

[0062] Figure 6 illustrates how sound is perceived during ambisonic playback when the acoustic signal of the virtual sound source at the virtual sound source location is not an ideal acoustic signal.

[0063] Figure 6A shows the sound pickup situation using the ambisonics method. Figure 6A is similar to Figure 5A, and from the perspective of the ambisonics microphone, a man is speaking to the left front and a woman is speaking to the right front.

[0064] As in the case of Figure 4, the output unit 13 is assumed to consist of four channels of speakers positioned to the front left, front right, rear left, and rear right, respectively, as viewed from the listener. Furthermore, the positions of the speakers positioned to the front left, front right, rear left, and rear right are assumed to be the virtual sound source positions. The positions of the male and female voices correspond to the positions of the front left and front right speakers, respectively.

[0065] Figure 6B shows the situation of the sound output from the speaker when the acoustic signal of the virtual sound source at the virtual sound source location is not an ideal acoustic signal.

[0066] In low-order ambisonics methods such as first-order ambisonics, it is difficult to express sharp directivity. Therefore, the virtual sound sources at each virtual sound source location generated in the signal processing unit 12 have insufficient separation and are correlated with each other. In other words, the acoustic signals of the virtual sound sources at each virtual sound source location are not object-based, but rather include ideal acoustic signals, such as the acoustic signals of a specific sound source as well as the acoustic signals of other sound sources.

[0067] For example, the acoustic signal of the virtual sound source at the virtual sound source position, which corresponds to the position of the speaker to the left front of the man's position, is not objectified as an acoustic source. Instead, it contains the acoustic signal of a specific sound source as an ideal acoustic signal, namely, the acoustic signal corresponding to the man's voice, as well as the acoustic signal corresponding to the woman's voice as another sound source. This acoustic signal corresponding to the woman's voice has less power than the acoustic signal corresponding to the woman's voice contained in the acoustic signal of the virtual sound source at the virtual sound source position, which corresponds to the position of the speaker to the right front of the woman's position.

[0068] For example, the acoustic signal of the virtual sound source at the virtual sound source position, which corresponds to the position of the speaker to the front right corresponding to the woman's position, is not objectified as an acoustic source. Instead, it includes the acoustic signal of a specific sound source as an ideal acoustic signal, namely, the acoustic signal corresponding to the woman's voice, as well as the acoustic signal corresponding to the man's voice as another sound source. This acoustic signal corresponding to the man's voice has less power than the acoustic signal corresponding to the man's voice included in the acoustic signal of the virtual sound source at the virtual sound source position, which corresponds to the position of the speaker to the front left corresponding to the man's position.

[0069] For example, the acoustic signals of the virtual sound sources at the virtual sound source locations, which are the positions of the left rear and right rear speakers where no actual sound source is located, are not converted into object sound sources. Instead, they include not only a silent acoustic signal as an ideal acoustic signal, but also acoustic signals from other virtual sound source locations, such as an acoustic signal corresponding to a male voice and an acoustic signal corresponding to a female voice. These acoustic signals corresponding to the male voice and the acoustic signals corresponding to the female voice have low power, as in the case described above.

[0070] As described above, in the low-order ambisonics method, the acoustic signal of the virtual sound source at the virtual sound source location is not object-based (it is not an ideal acoustic signal), and the acoustic signal of the virtual sound source at the virtual sound source location contains not only the ideal acoustic signal but also acoustic signals from other sound sources at low power.

[0071] Therefore, the front left speaker outputs a male voice (indicated by the thick solid arrow in the diagram) and a female voice (indicated by the thin dotted arrow in the diagram) at low power (volume). The front right speaker outputs a female voice (indicated by the thick dotted arrow in the diagram) and a male voice (indicated by the thin solid arrow in the diagram) at low power. The rear left and rear right speakers output male and female voices at low power.

[0072] Figure 6C shows how a listener perceives sound when the acoustic signal of a virtual sound source is not an ideal acoustic signal.

[0073] As explained in Figure 6B, the front left speaker outputs both male and female voices. The front right speaker outputs both female and male voices. The rear left and rear right speakers output both male and female voices, respectively. In other words, the male voice is output with a certain amount of power from the front left speaker, and also with lower power from the front right, rear left, and rear right speakers. The female voice is output with a certain amount of power from the front right speaker, and also with lower power from the front left, rear left, and rear right speakers. In this case, due to the phantom center effect, the listener perceives the positions of the male and female voices as sound sources to be intermediate positions closer to the listener than the virtual sound source positions of the front left and front right speakers, respectively. As a result, the sense of distance and breadth of the sound image is impaired, and the sense of presence is lost. When using the binaural playback method, the degree of external localization is significantly reduced, resulting in a loss of realism.

[0074] Therefore, the signal processing unit 12 enhances and / or suppresses the acoustic signals of the virtual sound sources so that the power differences between the acoustic signals of the virtual sound sources become larger, thereby substantially improving the separation between the virtual sound sources. As a result, the sense of distance and spaciousness of the sound image is improved, and the listener can enjoy the feeling that the distance of the sound source is expanding farther away, and that the sense of localization is clearer. Furthermore, when playing in a binaural system, the sense of external localization can be improved. Consequently, even with a low-order ambisonic system such as a first-order system, the listener can enjoy the same experience as with a high-order ambisonic system, and the loss of a sense of presence can be suppressed.

[0075] Figure 7 illustrates how the sound is perceived during ambisonic playback when the signal processing unit 12 enhances and / or suppresses the acoustic signal of the virtual sound source.

[0076] Figure 7A shows the sound pickup situation using the ambisonics method. Figure 7A is the same as Figure 5A, and from the perspective of the ambisonics microphone, a man is speaking to the left front and a woman is speaking to the right front.

[0077] As in the case of Figure 4, the output unit 13 is assumed to consist of four channels of speakers positioned to the front left, front right, rear left, and rear right, respectively, as viewed from the listener. Furthermore, the positions of the speakers positioned to the front left, front right, rear left, and rear right are assumed to be the virtual sound source positions. The positions of the male and female voices correspond to the positions of the front left and front right speakers, respectively.

[0078] Figure 7B shows the state of the sound output from the speaker when the signal processing unit 12 performs emphasis and / or suppression of the acoustic signal of the virtual sound source.

[0079] As explained in Figure 6, in low-order ambisonics schemes such as first-order ambisonics, the acoustic signals of the virtual sound sources at each virtual sound source location generated by the signal processing unit 12 include not only the ideal acoustic signal but also the acoustic signals of other sound sources.

[0080] For example, the acoustic signal of the virtual sound source at the virtual sound source position, which is the position of the speaker to the left front corresponding to the man's position, includes not only the acoustic signal corresponding to the man's voice, but also a low-power acoustic signal corresponding to the woman's voice as an acoustic signal from another sound source.

[0081] For example, the acoustic signal of the virtual sound source at the virtual sound source position, which is the position of the speaker to the right front corresponding to the woman's position, includes not only the acoustic signal corresponding to the woman's voice, but also a low-power acoustic signal corresponding to the man's voice as an acoustic signal from another sound source.

[0082] For example, the acoustic signals of the virtual sound sources at the virtual sound source locations, which are the positions of the left rear and right rear speakers where no actual sound source is located, include low-power acoustic signals corresponding to a male voice as an acoustic signal of another sound source, and low-power acoustic signals corresponding to a female voice.

[0083] The signal processing unit 12 enhances and / or suppresses the acoustic signals of virtual sound sources so that the power differences between the acoustic signals of the virtual sound sources are increased. For example, to simplify the explanation, in order to increase the power differences between the acoustic signals of virtual sound sources, the power of the acoustic signal of the virtual sound source with low power is set to 0 as suppression of the acoustic signals of virtual sound sources.

[0084] Here, the virtual sound source positions, which are the positions of the four channel speakers positioned to the left front, right front, left rear, and right rear respectively, are also referred to as the left front, right front, left rear, and right rear virtual sound source positions.

[0085] For example, if the power of the acoustic signal from the virtual sound source at the left front position is greater than the power of the acoustic signals from the other virtual sound source positions, the signal processing unit 12 suppresses the acoustic signals from the right front, left rear, and right rear virtual sound source positions, respectively, so that the difference between the power of the acoustic signal from the left front position and the power of the acoustic signals from the other virtual sound source positions (right front, left rear, and right rear) becomes larger, and the power of each of these virtual sound sources is reduced to zero.

[0086] In this case, the power of the acoustic signal corresponding to the male voice contained in the acoustic signal of the virtual sound source at the virtual sound source position to the front right becomes 0. Furthermore, the power of the acoustic signal corresponding to the male voice contained in the acoustic signal of the virtual sound source at the virtual sound source positions to the rear left and rear right becomes 0.

[0087] Therefore, the front left speaker outputs a male voice (indicated by a solid arrow in the diagram) and a female voice (indicated by a dotted arrow in the diagram) at low power. However, the other speakers, namely the front right, rear left, and rear right speakers, do not output a male voice, as indicated by the crossed-out solid arrows representing the male voice in the diagram.

[0088] On the other hand, for example, if the power of the acoustic signal of the virtual sound source at the front right virtual sound source position is greater than the power of the acoustic signals of the virtual sound sources at the other virtual sound source positions, the signal processing unit 12 suppresses the acoustic signals of the virtual sound sources at the front left, rear left, and rear right virtual sound source positions, so that the difference between the power of the acoustic signal of the virtual sound source at the front right virtual sound source position and the power of the acoustic signals of the other virtual sound source positions, namely the front left, rear left, and rear right virtual sound source positions, becomes larger, and the power of each of these virtual sound sources is reduced to zero.

[0089] In this case, the power of the acoustic signal corresponding to the woman's voice contained in the acoustic signal of the virtual sound source at the left front position becomes 0. Furthermore, the power of the acoustic signal corresponding to the woman's voice contained in the acoustic signal of the virtual sound source at the left rear and right rear positions becomes 0.

[0090] Therefore, the speaker on the front right outputs a female voice (indicated by the dotted arrow in the diagram) and a male voice (indicated by the solid arrow in the diagram) at low power. However, the other speakers, namely the speakers on the front left, rear left, and rear right, do not output a female voice, as indicated by the crossed-out dotted arrow representing the female voice in the diagram.

[0091] Figure 7C shows how the sound is perceived by the listener when the signal processing unit 12 enhances and / or suppresses the acoustic signal of the virtual sound source.

[0092] As explained in Figure 7B, among the virtual sound source positions of left front, right front, left rear, and right rear, if the power of the acoustic signal from the virtual sound source at the left front position is greater than the power of the acoustic signals from the virtual sound sources at the other virtual sound source positions, the male voice will be output from the left front speaker and not from the other speakers. Therefore, the phantom center effect does not occur, and the listener can perceive the position of the male voice as the sound source at the position of the left front speaker, which is the virtual sound source position.

[0093] Furthermore, as explained in Figure 7B, if the power of the acoustic signal from the virtual sound source at the front right virtual sound source location is greater than the power of the acoustic signals from the other virtual sound source locations, the woman's voice will be output from the front right speaker and not from the other speakers. Therefore, the phantom center effect does not occur, and the listener can perceive the location of the woman's voice as the sound source at the location of the front right speaker, which is the virtual sound source location.

[0094] As described above, listeners can enjoy the sensation that the male voice is coming from the front left speaker and the female voice is coming from the front right speaker, creating a sense of presence as if there are sound sources at both the front left and front right speaker locations.

[0095] <Example of the configuration of the signal processing unit 12>

[0096] Figure 8 is a block diagram showing an example configuration of the signal processing unit 12 in Figure 1.

[0097] In Figure 8, the signal processing unit 12 includes a format conversion unit 31, a virtual sound source generation unit 32, an emphasis / suppression unit 33, a surround signal generation unit 34, and a binaural signal generation unit 35.

[0098] The format conversion unit 31 is supplied with an M-channel acoustic signal as an A-format acoustic signal output by the sound pickup unit 11 (Figure 1). The format conversion unit 31 converts the M-channel acoustic signal as an A-format acoustic signal from the sound pickup unit 11 into an M-channel acoustic signal as a B-format acoustic signal and outputs it. For example, the format conversion unit 31 converts the FLU, FRD, BLD, BRU channel acoustic signals as an A-format acoustic signal output by the ambisonic microphone, which is the sound pickup unit 11, into W-channel, X-channel, Y-channel, Z-channel acoustic signals as a B-format acoustic signal and outputs them.

[0099] The virtual sound source generation unit 32 generates and outputs acoustic signals for virtual sound sources at three or more virtual sound source positions from the B-format acoustic signals output by the format conversion unit 31. For example, the virtual sound source generation unit 32 generates and outputs acoustic signals S(θ, φ) for virtual sound sources at three or more virtual sound source positions, such as 13 or 32 channels, from the W-channel, X-channel, Y-channel, and Z-channel acoustic signals output by the format conversion unit 31, according to the formula described above.

[0100] The enhancement / suppression unit 33 enhances / suppresses (enhances and / or suppresses) the acoustic signals of virtual sound sources based on the power of the acoustic signals of the virtual sound sources at the virtual sound source positions of the K channels output by the virtual sound source generation unit 32, and outputs them in such a way that the power differences are increased. Enhancement / suppression of the acoustic signals of virtual sound sources that increases the power differences can be achieved, for example, by increasing the degree of emphasis on the acoustic signals of virtual sound sources with high power and decreasing the degree of emphasis on the acoustic signals of virtual sound sources with low power. Alternatively, enhancement / suppression of the acoustic signals of virtual sound sources that increases the power differences can be achieved, for example, by decreasing the degree of suppression on the acoustic signals of virtual sound sources with high power and increasing the degree of suppression on the acoustic signals of virtual sound sources with low power. Furthermore, enhancement / suppression of the acoustic signals of virtual sound sources that increases the power differences can be achieved, for example, by emphasizing the acoustic signals of virtual sound sources with high power and / or suppressing the acoustic signals of virtual sound sources with low power.

[0101] The surround signal generation unit 34 generates surround signals in a surround sound format from the acoustic signals of virtual sound sources at virtual sound source positions for K channels output by the enhancement / suppression unit 33, and outputs them to the output unit 13 (Figure 1).

[0102] The binaural signal generation unit 35 generates a binaural signal using a binaural method from the acoustic signals of virtual sound sources at virtual sound source positions for K channels output by the enhancement / suppression unit 33, and outputs it to the output unit 13.

[0103] Figure 9 is a flowchart illustrating the processing of the signal processing unit 12.

[0104] In step S11, the format conversion unit 31 converts the M-channel A-format audio signal output by the sound pickup unit 11 into an M-channel B-format audio signal and outputs it.

[0105] In step S12, the virtual sound source generation unit 32 generates and outputs virtual sound source audio signals for 3 or more channels at virtual sound source positions from the M-channel B-format audio signal output by the format conversion unit 31.

[0106] In step S13, the enhancement / suppression unit 33 enhances / suppresses the acoustic signals of the virtual sound sources and outputs them based on the power of the acoustic signals of the virtual sound sources at the virtual sound source positions of the K channels output by the virtual sound source generation unit 32, so as to increase the power difference. That is, the enhancement / suppression unit 33 enhances / suppresses the acoustic signals of the virtual sound sources and outputs them so as to increase the power difference between the acoustic signals of the virtual sound sources, in particular, between the acoustic signals of virtual sound sources with high power and the acoustic signals of virtual sound sources with low power.

[0107] In step S14, the surround signal generation unit 34 generates a surround signal in a surround format from the acoustic signals of virtual sound sources at the virtual sound source positions of K channels output by the enhancement / suppression unit 33, and outputs it to the output unit 13. And / or, the binaural signal generation unit 35 generates a binaural signal in a binaural format from the acoustic signals of virtual sound sources at the virtual sound source positions of K channels output by the enhancement / suppression unit 33, and outputs it to the output unit 13.

[0108] <First example of the configuration of the emphasis / suppression unit 33>

[0109] Figure 10 is a block diagram showing a first configuration example of the enhancement / suppression unit 33 in Figure 8.

[0110] In Figure 10, the enhancement / suppression unit 33 includes a power calculation unit 41 and a power adjustment unit 42.

[0111] The power calculation unit 41 is supplied with the acoustic signals of the virtual sound sources at the virtual sound source positions for the K channels output by the virtual sound source generation unit 32. The power calculation unit 41 calculates the power of the acoustic signals of the virtual sound sources at the virtual sound source positions for the K channels output by the virtual sound source generation unit 32 in units of frames, which are acoustic signals of a predetermined time length, and outputs them to the power adjustment unit 42 along with the acoustic signals of the virtual sound sources.

[0112] The power adjustment unit 42 outputs the enhancement / suppression of each of the virtual sound source audio signals at the K virtual sound source positions from the power calculation unit 41 on a frame-by-frame basis, based on the power ratio between the power of each of the virtual sound source audio signals at the K virtual sound source positions from the power calculation unit 41 and the power of the audio signal of the virtual sound source with the maximum power (maximum power).

[0113] For example, the power adjustment unit 42 calculates a gain for each frame based on the power ratio between the power of the acoustic signal of each virtual sound source and the maximum power, and uses this gain to adjust the power of the acoustic signal of each virtual sound source, thereby performing emphasis / suppression of the acoustic signal of each virtual sound source on a frame-by-frame basis. In other words, the power adjustment unit 42 adjusts the power of the acoustic signal of each virtual sound source by multiplying the acoustic signal of each virtual sound source by the gain for each frame.

[0114] The gain, which is based on the power ratio between the power of the virtual sound source's audio signal and its maximum power, can be a value based on a power of the power ratio, for example. For example, in a frame, if we represent the power of the k-th virtual sound source's audio signal as Pk and the maximum power as Pmax, then with a non-negative value V as the exponent, the gain G(k) of the k-th virtual sound source's audio signal can be G(k) = (Pk / Pmax)^V. The exponent V can be set in advance or set according to user operation.

[0115] When the gain G(k) = (Pk / Pmax)^V is adopted, the gain G(k) of the acoustic signals of virtual sound sources with power close to the maximum power Pmax (including the acoustic signals of virtual sound sources with maximum power Pmax) will be close to 1. On the other hand, the gain G(k) of the acoustic signals of virtual sound sources with power relatively small compared to the maximum power Pmax will be close to 0. As a result, the power of the acoustic signals of each virtual sound source is adjusted so that the power difference between the acoustic signals of virtual sound sources with high power close to the maximum power Pmax and the acoustic signals of virtual sound sources with low power relative to the maximum power Pmax becomes large.

[0116] If we use G(k) = (Pk / Pmax)^V as the gain G(k), the gain G(k) will be 1 or less, so the acoustic signal of the virtual sound source will not be amplified.

[0117] To prevent the gain G(k) from becoming too small, a threshold value such as -20dB can be set, and if the gain G(k) falls below this threshold, the gain G(k) can be limited to the threshold value.

[0118] Furthermore, the gain G(k), which is a power of the power ratio, can be expressed as, for example, G(k) = a * (Pk / Pmax)^V + b. a and b are coefficients that can be pre-set or adjusted according to user input. By setting coefficients a and b, it is possible to perform only emphasis, only suppression, or both, and the degree of emphasis or suppression can be adjusted.

[0119] <Example of the second configuration of the emphasis / suppression unit 33>

[0120] Figure 11 is a block diagram showing a second configuration example of the enhancement / suppression unit 33 in Figure 8.

[0121] In Figure 11, the enhancement / suppression unit 33 includes a splitting unit 51, a power calculation unit 52, a power adjustment unit 53, and a synthesis unit 54, and enhances / suppresses the acoustic signal of the virtual sound source for each frequency band based on the power of each frequency band of the acoustic signal of the virtual sound source.

[0122] Here, the enhancement / suppression unit 33 in Figure 10 enhances / suppresses the entire frequency band of the virtual sound source's acoustic signal based on the power of the entire frequency band of the virtual sound source's acoustic signal, regardless of the frequency band. In contrast, the enhancement / suppression unit 33 in Figure 11 differs from the case in Figure 10 in that it enhances / suppresses the acoustic signal of the virtual sound source for each frequency band based on the power of each frequency band of the virtual sound source's acoustic signal.

[0123] The splitting unit 51 is supplied with the acoustic signals of the virtual sound sources at the virtual sound source positions for K channels, which are output by the virtual sound source generation unit 32. The splitting unit 51 divides each of the acoustic signals of the virtual sound sources at the virtual sound source positions for K channels into frequency components for each frequency band of a predetermined bandwidth and outputs them.

[0124] For example, the splitting unit 51 performs an STFT (short-term Fourier transform) on each of the acoustic signals of the virtual sound sources at the virtual sound source positions of the K channels, converting them into frequency domain signals on a frame-by-frame basis and outputting them. In the STFT, the (time-domain) acoustic signals of the virtual sound sources at the virtual sound source positions of each channel are applied to a predetermined window function that is shifted in the time direction, converting them into a series of overlapping frames. Each frame is then subjected to a (discrete) Fourier transform, converting them into frequency domain signals for each frame.

[0125] Here, the frequency-domain signal obtained by Fourier transforming a time-domain acoustic signal corresponds to the frequency components of the time-domain acoustic signal in each frequency band with a frequency resolution of Δf. Therefore, it can be said that the time-domain acoustic signal is band-divided according to the Fourier transform, with each frequency band having a frequency resolution of Δf. As a method for band-dividing a time-domain acoustic signal, in addition to STFT, methods such as QMF (Quadrature Mirror Filter) or DFT filter banks can be employed.

[0126] The power calculation unit 52 uses the frequency domain signals for each of the K virtual sound source positions output by the division unit 51 to calculate the power of the acoustic signals of the virtual sound sources at each of the K virtual sound source positions for each frequency band on a frame-by-frame basis. The power calculation unit 52 outputs the power of the acoustic signals of the virtual sound sources at each of the K virtual sound source positions for each frequency band on a frame-by-frame basis, along with the frequency domain signals for each of the K virtual sound source positions, to the power adjustment unit 53.

[0127] The power adjustment unit 53 uses the power of the acoustic signals of the virtual sound sources at the K channel virtual sound source positions from the power calculation unit 52 for each frequency band, and based on the power ratio for each frequency band between the power of each of the acoustic signals of the virtual sound sources at the K channel virtual sound source positions and the power of the acoustic signal of the maximum power virtual sound source, it performs emphasis / suppression for each frequency band of the acoustic signals of the virtual sound sources at the K channel virtual sound source positions on a frame-by-frame basis and outputs the result.

[0128] For example, the power adjustment unit 53 calculates a gain for each frequency band on a frame-by-frame basis, based on the power ratio between the power of the acoustic signal of each virtual sound source and the maximum power. Then, the power adjustment unit 53 adjusts the power of the acoustic signal of each virtual sound source for each frequency band using the gain for each frequency band, thereby performing emphasis / suppression of the acoustic signal of each virtual sound source for each frequency band on a frame-by-frame basis. In other words, the power adjustment unit 53 adjusts the power of the acoustic signal of each virtual sound source for each frequency band by multiplying the frequency domain signal (frequency components for each frequency band) for each of the K channel virtual sound source positions by the gain for each frequency band on a frame-by-frame basis.

[0129] As the gain, a value based on the power ratio between the power of the virtual sound source's acoustic signal and its maximum power can be, for example, a value based on a power of the power ratio can be adopted. For example, in a frame, Pk(ω) represents the power of the frequency band of the acoustic signal of the k-th virtual sound source, centered at angular frequency ω and with a frequency resolution of △f, among the acoustic signals of the K virtual sound sources at each of the K channel virtual sound source positions. Also, in a frame, Pmax(ω) represents the maximum power among the powers P1(ω), P2(ω), ..., PK(ω) of the frequency bands centered at angular frequency ω of each of the acoustic signals of the K channel virtual sound sources. Also, as in the case of Figure 10, V represents an exponent of a value of 0 or more. In this case, as the gain G(k,ω) of the frequency band centered at angular frequency ω of the acoustic signal of the k-th virtual sound source, G(k,ω) = (Pk(ω) / Pmax(ω))^V can be adopted. Here, the frequency band centered on the angular frequency ω will also be referred to as the frequency band ω.

[0130] When the gain G(k,ω) is set to G(k,ω) = (Pk(ω) / Pmax(ω))^V, the gain G(k,ω) in the frequency band ω of the acoustic signal of a virtual sound source whose power Pk(ω) in the frequency band ω is close to the maximum power Pmax(ω) will be close to 1. On the other hand, the gain G(k,ω) in the frequency band ω of the acoustic signal of a virtual sound source whose power Pk(ω) in the frequency band ω is relatively small compared to the maximum power Pmax(ω) will be close to 0. As a result, for the acoustic signals of virtual sound sources at K channel virtual sound source positions, the power in the frequency band ω of each virtual sound source's acoustic signal is adjusted so that the difference in power in the frequency band ω is large between the acoustic signal of a virtual sound source with high power (close to the maximum power Pmax(ω)) and the acoustic signal of a virtual sound source with low power (smaller than the maximum power Pmax(ω)).

[0131] Regarding the gain G(k,ω), similar to the gain G(k) explained in Figure 10, if the gain G(k,ω) falls below the threshold value, it can be limited to the threshold value to prevent it from becoming too small.

[0132] Furthermore, as the gain G(k, ω), which is a power of the power ratio, a value calculated according to the formula G(k, ω) = a(ω) * (Pk(ω) / Pmax(ω))^V + b(ω), using coefficients a(ω) and b(ω), can be adopted, for example, similar to the gain G(k) explained in Figure 10. The coefficients a(ω) and b(ω) are coefficients for each frequency band ω, and can be set in advance or set according to user operation.

[0133] The combining unit 54 synthesizes the frequency domain signals, which have been emphasized / suppressed for each frequency band, for each of the K virtual sound source positions output by the power adjustment unit 53, thereby returning them to the frequency band signals before the band division was performed by the splitting unit 51.

[0134] For example, the synthesis unit 54 performs an inverse FFT (fast Fourier transformation) on the frequency domain signal (frequency components for each frequency band) for each of the K virtual sound source positions output by the power adjustment unit 53, after the emphasis / suppression for each frequency band has been applied. The synthesis unit 54 synthesizes the time domain signal obtained by the inverse FFT, which is the time series of frames of the acoustic signals of each of the K virtual sound source positions, by overlapping the overlapping portions from the band division and adding them together.

[0135] The synthesis unit 54 outputs the acoustic signals of the virtual sound sources at each of the K virtual sound source positions obtained by synthesis to the surround signal generation unit 34 and the binaural signal generation unit 35 (Figure 8) as acoustic signals of virtual sound sources that have been enhanced / suppressed for each frequency band.

[0136] The signal processing system 10 generates acoustic signals for three or more virtual sound sources at different positions from the acoustic signal (input acoustic signal). Based on the power of the acoustic signals of the virtual sound sources, it can enhance or suppress the acoustic signals of the virtual sound sources so that the power differences are large, thereby suppressing any loss of realism. For example, when 3D sound is reproduced using a low-order ambisonics method such as a first-order ambisonics method, the listener can enjoy the spatial resolution and sense of distance that would be possible with a higher-order ambisonics method.

[0137] Furthermore, the signal processing system 10 can reproduce immersive 3D sound for UGC (User Generated Content), for example. For instance, many (virtual) surround sounds, such as music, are realized by applying an HRTF (Human-Related Transformer) to object sound sources (acoustic signals) recorded for each sound source. However, obtaining object sound sources is often difficult in UGC production. For example, even if sound is recorded using multiple microphones in a real-world environment such as a sports event, the resulting recording will be an acoustic signal containing various sound sources. Therefore, since individual sound sources (object sound sources) are not obtained, it is difficult to realize surround sound by applying an HRTF to object sound sources. In contrast, the signal processing system 10 can reproduce immersive 3D sound from an acoustic signal containing various sound sources.

[0138] Furthermore, the signal processing system 10 can robustly separate sound sources (and their acoustic signals) from each other using simple processing with minimal computational complexity (improving separation accuracy). In other words, for example, sound source separation processes such as independent component analysis and beamforming require a large amount of computation, and attempting to separate more sound sources than the number of microphones used for sound collection can lead to processing failure. In contrast, the signal processing system 10 can easily and robustly separate any number of virtual sound sources from each other (making the listener perceive the separated virtual sound sources).

[0139] Furthermore, in the signal processing system 10, the signal processing unit 12 includes a format conversion unit 31 to a binaural signal generation unit 35. Therefore, the user can easily obtain surround signals or binaural signals simply by inputting an A-format audio signal to the signal processing unit 12. Note that the functions of each of the format conversion unit 31 to the binaural signal generation unit 35 that constitute the signal processing unit 12 can be provided individually. When the functions of each of the format conversion unit 31 to the binaural signal generation unit 35 are provided individually, the user can manually generate surround signals or binaural signals using the functions of each of the format conversion unit 31 to the binaural signal generation unit 35 that are provided individually. However, when manually generating surround or binaural signals using the functions of the individually provided format conversion unit 31 to binaural signal generation unit 35, the user must perform the following tasks: convert an A-format audio signal to a B-format audio signal using the function of the format conversion unit 31; generate virtual sound source audio signals for K channels at virtual sound source positions from the B-format audio signal using the function of the virtual sound source generation unit 32; perform emphasis / suppression of the virtual sound source audio signals for K channels at virtual sound source positions using the function of the emphasis / suppression unit 33; and further generate a surround signal from the emphasized / suppressed virtual sound source audio signals using the function of the surround signal generation unit 34, or generate a binaural signal from the emphasized / suppressed virtual sound source audio signals using the function of the binaural signal generation unit 35. In contrast, when using the signal processing unit 12, surround or binaural signals can be obtained simply by inputting an A-format audio signal.

[0140] Furthermore, this technology can be applied even when the ambisonics method is not employed in the reproduction of 3D sound. For example, this technology can be applied when it is possible to generate audio signals for three or more virtual sound sources at different virtual sound source positions from an audio signal by some method.

[0141] <Description of a computer using this technology>

[0142] The series of processes described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up that software are installed on a computer. Here, a computer includes computers built into dedicated hardware, as well as general-purpose personal computers, for example, that can perform various functions by installing various programs.

[0143] Figure 12 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above using a program.

[0144] In a computer, the processing circuit 901, ROM (Read Only Memory) 902, and RAM (Random Access Memory) 903 are interconnected by a bus 904.

[0145] An input / output interface 905 is further connected to the bus 904. An input unit 906, an output unit 907, a storage unit 908, a communication unit 909, and a drive 910 are connected to the input / output interface 905.

[0146] The input unit 906 may include physical or virtual operating means that the user operates to input information, such as a keyboard, mouse, or touch panel, as well as means that the user inputs information through voice, eye gaze, etc. Furthermore, the input unit 906 may include sensors that acquire physical quantities such as light or sound, such as a camera or microphone. The output unit 907 may include means that present information to the user by stimulating the user's perception, such as a display, speaker, or haptic device. The storage unit 908 is composed of a hard disk, non-volatile or volatile memory, etc., and stores various types of information (including programs). The communication unit 909 is a network interface, etc., and performs wired or wireless communication with the outside. The drive 910 drives removable media 911 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0147] The processing circuit 901 includes a processor that executes programs such as a CPU (Central Processing Unit) and a DSP (Digital Signal Processor). The processing circuit 901 (its processor) performs the above-described series of processes by loading the program stored in the storage unit 908 into the RAM 903 via the input / output interface 905 and the bus 904 and executing it. The processing circuit 901 can output the processing results of the series of processes from the output unit 907 via the bus 904 and the input / output interface 905 as needed. The processing circuit 901 can also store the processing results in the storage unit 908 or transmit them from the communication unit 909.

[0148] The program executed by the computer (processing circuit 901) can be provided by recording it on a removable medium 911, such as a package medium. The program can also be provided via wired or wireless transmission media, such as a local area network, the internet, or digital satellite broadcasting.

[0149] In a computer, a program can be installed in the storage unit 908 via the input / output interface 905 by inserting the removable media 911 into the drive 910. Alternatively, a program can be received by the communication unit 909 via a wired or wireless transmission medium and installed in the storage unit 908. Furthermore, programs can be pre-installed in the ROM 902 or the storage unit 908.

[0150] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.

[0151] The processes that a computer performs according to a program do not necessarily have to follow the order described in the flowchart. In other words, the processes that a computer performs according to a program include processes that are executed in parallel or individually (e.g., parallel processing and object-based processing).

[0152] The program may be processed by a single computer (processor), or it may be processed in a distributed manner by multiple computers. Furthermore, the program may be transferred to a remote computer and executed there.

[0153] For example, the sound collection unit 11 in Figure 1 corresponds to the input unit 906, and the format conversion unit 31 to the binaural signal generation unit 35 of the signal processing unit 12 in Figure 8 each correspond to the processing circuit 901 (or its processor) that executes the program. The output unit 13 in Figure 1 corresponds to the output unit 907.

[0154] In this specification, a system means one component or a collection of multiple components (devices, modules (parts), etc.), and in the case of a collection of multiple components, it is not necessary whether all components are located in the same enclosure or not. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device in which multiple modules are housed in one enclosure, are both systems. Furthermore, for example, the entire computer described above, or a combination of a computer and other devices such as a server (not shown), are also systems. One or more components of a computer, for example, only the processing circuit 901, or a combination of the processing circuit 901 to the bus 904, are also systems.

[0155] Furthermore, the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the spirit of this technology.

[0156] For example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.

[0157] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.

[0158] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.

[0159] Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also occur.

[0160] Furthermore, this technology can take the following configuration.

[0161] <1> A signal processing device comprising: a virtual sound source generation unit that generates sound signals of three or more virtual sound source locations from an input sound signal; and an enhancement / suppression unit that enhances and / or suppresses the sound signals of the virtual sound sources based on the power of the sound signals of the virtual sound sources so that the power differences are large. <2> The signal processing device according to <1>, wherein the enhancement / suppression unit enhances and / or suppresses the sound signals of each virtual sound source based on the power ratio of the power of the sound signals of each virtual sound source to the power of the sound signals of the virtual sound source with the maximum power. <3> The signal processing device according to <2>, wherein the enhancement / suppression unit adjusts the power of the sound signals of each virtual sound source using a value based on the power ratio as a gain. <4> The signal processing device according to <3>, wherein the enhancement / suppression unit adjusts the power of the sound signals of each virtual sound source using a value based on a power of the power ratio as a gain. <5> The enhancement / suppression unit calculates the power of the acoustic signal of each virtual sound source, calculates a gain based on the power of the power ratio between the power of the acoustic signal of each virtual sound source and the power of the acoustic signal of the virtual sound source with maximum power, and adjusts the power of the acoustic signal of each virtual sound source by multiplying the acoustic signal of each virtual sound source by the gain, as described in <4>. <6> The enhancement / suppression unit enhances and / or suppresses the acoustic signal of each virtual sound source for each frequency band based on the power of the acoustic signal of each virtual sound source for each frequency band, as described in <1>. <7> The enhancement / suppression unit enhances and / or suppresses the acoustic signal of each virtual sound source for each frequency band based on the power ratio between the power of the acoustic signal of each virtual sound source and the power of the acoustic signal of the virtual sound source with maximum power, as described in <6>. <8> The enhancement / suppression unit adjusts the power of the acoustic signal of each virtual sound source for each frequency band using the value based on the power ratio as the gain, as described in <7>. <9> The signal processing device described in <8>, wherein the enhancement / suppression unit uses a value based on a power of the power ratio as the gain to adjust the power of the acoustic signal of each virtual sound source for each frequency band.<10> The enhancement / suppression unit calculates the power of the acoustic signal of each virtual sound source for each frequency band, calculates a gain for each frequency band based on the power of the power ratio between the power of the acoustic signal of each virtual sound source and the power of the acoustic signal of the virtual sound source with the maximum power, and adjusts the power of the acoustic signal of each virtual sound source by multiplying the acoustic signal of each virtual sound source by the gain for each frequency band. The signal processing device according to <9>. <11> The enhancement / suppression unit enhances the acoustic signal of a virtual sound source with high power and / or suppresses the acoustic signal of a virtual sound source with low power. The signal processing device according to <1> to <11>. <12> The input acoustic signal is an acoustic signal obtained by ambisonics sound collection. The signal processing device according to any one of <1> to <11>. <13> The ambisonics method is a first-order ambisonics method. The signal processing device according to <12>. <14> The signal processing device according to <13>, wherein the virtual sound source generation unit generates an acoustic signal S(θ, φ) of the virtual sound source at the virtual sound source position in the direction of angles φ and θ, according to the formula S(θ, φ) = fp ・ W + (1-fp) ・{ X sinθcosφ + Y sinθsinφ Y + Z cosθ}, where φ is the horizontal angle of the virtual sound source position with respect to the front direction, θ is the vertical angle of the virtual sound source position with respect to the upward direction, fp is a coefficient between 0 and 1, and W, X, Y, and Z are the acoustic signals of the W channel, X channel, Y channel, and Z channel of the B format of the ambisonics method, respectively. <15> The signal processing device according to <14>, further comprising a format conversion unit that converts the acoustic signal of the A format of the ambisonics method into an acoustic signal of the B format. <16> The signal processing device according to any one of <1> to <15>, wherein the virtual sound source position is the position of a speaker. <17> The signal processing device according to any one of <1> to <16>, further comprising a surround signal generation unit that generates a surround signal in a surround system from the acoustic signals of three or more virtual sound sources at the virtual sound source positions after enhancement and / or suppression.<18> The signal processing apparatus according to any one of <1> to <17>, further comprising a binaural signal generation unit that generates a binaural signal in a binaural manner from the acoustic signals of three or more virtual sound sources at the virtual sound source positions after emphasis and / or suppression. <19> A signal processing method comprising generating acoustic signals of three or more virtual sound sources at the virtual sound source positions from an input acoustic signal, and emphasizing and / or suppressing the acoustic signals of the virtual sound sources so as to increase the power differences based on the power of the acoustic signals of the virtual sound sources. <20> A program for causing a computer to function as a virtual sound source generation unit that generates acoustic signals of three or more virtual sound sources at the virtual sound source positions from an input acoustic signal, and an emphasis / suppression unit that emphasizes and / or suppresses the acoustic signals of the virtual sound sources so as to increase the power differences based on the power of the acoustic signals of the virtual sound sources.

[0162] 10 Signal processing system, 11 Sound collection unit, 12 Signal processing unit, 13 Output unit, 31 Format conversion unit, 32 Virtual sound source generation unit, 33 Emphasis / suppression unit, 34 Surround signal generation unit, 35 Binaural signal generation unit, 41 Power calculation unit, 42 Power adjustment unit, 51 Splitting unit, 52 Power calculation unit, 53 Power adjustment unit, 54 Synthesis unit, 901 Processing circuit, 902 ROM, 903 RAM, 904 Bus, 905 Input / Output interface, 906 Input unit, 907 Output unit, 208 Storage unit, Communication unit, 210 Drive, 211 Removable media

Claims

1. A signal processing device comprising: a virtual sound source generation unit that generates acoustic signals of three or more virtual sound sources at virtual sound source positions from an input acoustic signal; and an enhancement / suppression unit that enhances and / or suppresses the acoustic signals of the virtual sound sources based on the power of the acoustic signals of the virtual sound sources so that the power differences are large.

2. The signal processing apparatus according to claim 1, wherein the enhancement / suppression unit enhances and / or suppresses the acoustic signals of each virtual sound source based on the power ratio between the power of the acoustic signals of each virtual sound source and the power of the acoustic signals of the virtual sound source with the maximum power.

3. The signal processing device according to claim 2, wherein the enhancement / suppression unit adjusts the power of the acoustic signal of each virtual sound source using a value based on the power ratio as the gain.

4. The signal processing device according to claim 3, wherein the enhancement / suppression unit adjusts the power of the acoustic signal of each virtual sound source using a value based on a power of the power ratio as the gain.

5. The signal processing device according to claim 4, wherein the enhancement / suppression unit calculates the power of the acoustic signal of each virtual sound source, calculates a gain based on the power of the power ratio between the power of the acoustic signal of each virtual sound source and the power of the acoustic signal of the virtual sound source with the maximum power, and adjusts the power of the acoustic signal of each virtual sound source by multiplying the acoustic signal of each virtual sound source by the gain.

6. The signal processing apparatus according to claim 1, wherein the enhancement / suppression unit enhances and / or suppresses the acoustic signal of the virtual sound source for each frequency band based on the power of each frequency band of the acoustic signal of the virtual sound source.

7. The signal processing apparatus according to claim 6, wherein the enhancement / suppression unit enhances and / or suppresses the acoustic signals of each virtual sound source for each frequency band based on the power ratio for each frequency band between the power of the acoustic signal of each virtual sound source and the power of the acoustic signal of the virtual sound source with maximum power.

8. The signal processing device according to claim 7, wherein the enhancement / suppression unit adjusts the power of the acoustic signal of each virtual sound source for each frequency band using a value based on the power ratio as the gain.

9. The signal processing device according to claim 8, wherein the enhancement / suppression unit uses a value based on a power of the power ratio as the gain to adjust the power of the acoustic signal of each virtual sound source for each frequency band.

10. The signal processing device according to claim 9, wherein the enhancement / suppression unit calculates the power of the acoustic signal of each virtual sound source for each frequency band, calculates a gain for each frequency band based on a power of the power ratio between the power of the acoustic signal of each virtual sound source and the power of the acoustic signal of the virtual sound source with the maximum power, and adjusts the power of the acoustic signal of each virtual sound source by multiplying the acoustic signal of each virtual sound source by the gain for each frequency band.

11. The signal processing apparatus according to claim 1, wherein the enhancement / suppression unit enhances the acoustic signal of a virtual sound source with high power and / or suppresses the acoustic signal of a virtual sound source with low power.

12. The signal processing device according to claim 1, wherein the input acoustic signal is an acoustic signal obtained by ambisonic sound collection.

13. The signal processing apparatus according to claim 12, wherein the ambisonics method is a first-order ambisonics method.

14. The signal processing device according to claim 13, wherein the virtual sound source generation unit generates an acoustic signal S(θ, φ) of the virtual sound source at the virtual sound source position in the directions of angles φ and θ, according to the formula S(θ, φ) = fp ・ W + (1-fp) ・{ X sinθcosφ + Y sinθsinφ Y + Z cosθ}, where φ is the horizontal angle of the virtual sound source position with respect to the front direction, θ is the vertical angle of the virtual sound source position with respect to the upward direction, fp is a coefficient between 0 and 1, and W, X, Y, and Z are the acoustic signals of the W channel, X channel, Y channel, and Z channel of the B format of the ambisonics method, respectively.

15. The signal processing apparatus according to claim 14, further comprising a format conversion unit that converts an A-format acoustic signal of the ambisonics method into an B-format acoustic signal.

16. The signal processing apparatus according to claim 1, wherein the virtual sound source position is the position of a speaker.

17. The signal processing apparatus according to claim 1, further comprising a surround signal generation unit that generates a surround signal in a surround format from the acoustic signals of three or more virtual sound sources at the virtual sound source positions after enhancement and / or suppression.

18. The signal processing apparatus according to claim 1, further comprising a binaural signal generation unit that generates a binaural signal in a binaural manner from the acoustic signals of three or more virtual sound sources at the virtual sound source positions after emphasis and / or suppression.

19. A signal processing method comprising: generating acoustic signals of three or more virtual sound sources at virtual sound source positions from an input acoustic signal; and enhancing and / or suppressing the acoustic signals of the virtual sound sources based on the power of the acoustic signals of the virtual sound sources so that the power differences become larger.

20. A program for causing a computer to function as a virtual sound source generation unit that generates acoustic signals of virtual sound sources at three or more virtual sound source positions from an input acoustic signal, and an enhancement / suppression unit that enhances and / or suppresses the acoustic signals of the virtual sound sources based on the power of the acoustic signals of the virtual sound sources so that the power differences become larger.

Citation Information

Patent Citations

  • Information processing unit, sound image localization enhancement method, and sound image localization enhancement program

    JP2014090293A

  • Hybrid spatial audio decoder

    WO2020247033A1