Sound pickup apparatus, sound pickup program, and sound pickup method
The sound collection device corrects frequency-dependent directivity issues by using a null former and correction gain multiplication to enhance sound collection in the target direction, addressing distortion of high-frequency components in beamformers with two microphones.
Patent Information
- Application Number
- JP2024025637
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-22
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2044-02-22
AI Technical Summary
Conventional beamformers using spectral subtraction with two microphones form frequency-dependent directivity, leading to suppression and distortion of high-frequency components of the target sound when the blind spot direction differs from the target sound direction.
A sound collection device and method that includes a null former forming means to suppress sounds in the target direction, a correction gain multiplication means to adjust the null former spectrum, and a target direction acoustic enhancement means to enhance sound collection in the target direction, correcting frequency-dependent issues by using a predetermined correction gain value.
The solution achieves improved sound quality by correcting frequency-dependent directivity, reducing distortion of high-frequency components and enhancing sound collection in the target direction, even when the blind spot direction differs from the target sound direction.
Smart Images

Figure 2025128750000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sound collection device, a sound collection program, and a sound collection method, and can be applied to a sound collection device that, for example, emphasizes only sounds coming from a specific direction and suppresses sounds coming from other directions. [Background technology]
[0002] When using a speech recognition system in a noisy environment, ambient noise that mixes in with the required target sound is a nuisance that reduces the speech recognition rate of the recorded speech.
[0003] In a conventional environment with multiple sound sources, beamforming using a microphone array is a technique for capturing only sounds from a specific direction and obtaining the desired target sound while avoiding unwanted sounds. Beamforming is a technology for forming directivity by utilizing the time difference between signals arriving at each microphone (see Non-Patent Document 1). A microphone array that performs beamforming is also called a beamformer. Basic beamformer configurations include additive and subtractive beamformers, which consist of a delay unit and an adder / subtractor. First, the delay unit calculates the arrival time difference between the sound from the target sound direction and the sound arriving at each microphone, and then adds a delay to align the phase of the target sound. Additive beamformers create directivity with high sensitivity to the target sound direction by adding phase-aligned microphone input signals. Subtractive beamformers create directivity with a sharp null in a specific direction by subtracting one phase-aligned microphone input signal from the other. When the target sound and noise source are in close proximity to the beamformer, a sharper directivity is required to avoid unwanted sound contamination and obtain only the desired target sound. To achieve a sharper directivity, both additive and subtractive beamformers require a large number of microphones. For example, the Minimum Variance Distortionless Response (MVDR) method requires one more microphone than the number of noise sources to be suppressed. While beamforming with such a large number of microphones can be experimentally achieved in a laboratory, it is not suitable for commercial products due to component costs and installation space requirements.
[0004] To overcome this problem, a spectral subtraction method as disclosed in Patent Document 1 is available as a method for forming sharp directivity using a small number of microphones.
[0005] The technology described in Patent Document 1 achieves sharp directivity with a small number of microphones by using spectral subtraction in the frequency domain (a method of subtracting amplitude spectra from each other, but replacing any negative results with zero or a small number). A null former is created that forms a blind spot in the direction of the target area, and sharp directivity is created in the direction of the target area by spectrally subtracting the null former amplitude spectrum from the microphone input amplitude spectrum. Using this method, sharp directivity can be created with just two microphones. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-183358 [Non-patent literature]
[0007] [Non-Patent Document 1] Futoshi Asano, "Acoustic Technology Series 16: Array Signal Processing of Sound - Localization, Tracking and Separation of Sound Sources", edited by the Acoustical Society of Japan, Corona Publishing, published February 25, 2011 Summary of the Invention [Problem to be solved by the invention]
[0008] As described above, the technology described in Patent Document 1 can form sharp directivity with just two microphones. However, the directivity formed by the technology described in Patent Document 1 is frequency-dependent. The directivity depends on the microphone spacing relative to the wavelength, and the directivity formed by the technology described in Patent Document 1 becomes wider at low frequencies and narrower at high frequencies, based on the microphone array spacing. This frequency-dependency problem is also a similar issue for other beamformers.
[0009] Therefore, with conventional beamformers using spectral subtraction, if the direction of the blind spot forming the null is different from the direction of the target sound, there was a problem in that the high-frequency components of the target sound would be suppressed and distorted due to the frequency dependency of the directivity.
[0010] In view of the above problems, it is possible to provide a sound collection device, a sound collection program, and a sound collection method that suppress distortion of sound collected when collecting sound from a target direction. [Means for solving the problem]
[0011] The first invention is characterized by comprising: a null former forming means for forming a null former that directs a blind spot in a target direction and suppresses only sounds emitted in the target direction based on an input spectrum of an acoustic signal supplied from a microphone array, thereby obtaining a null former spectrum; a correction gain multiplication means for multiplying the null former spectrum obtained by the null former forming means by a predetermined correction gain value to calculate a corrected null former spectrum; and a target direction acoustic enhancement means for obtaining a beamformer output sound that has formed directivity in the target direction based on the corrected null former spectrum calculated by the correction gain multiplication means and the input spectrum.
[0012] The sound collection program of the second invention is characterized in that it causes a computer to function as: a null former forming means that, based on the input spectrum of an acoustic signal supplied from a microphone array, forms a null former that directs a blind spot in a target direction and suppresses only sounds emitted in the target direction, thereby obtaining a null former spectrum; a correction gain multiplication means that multiplies the null former spectrum obtained by the null former forming means by a predetermined correction gain value to calculate a corrected null former spectrum; and a target direction sound enhancement means that obtains a beamformer output sound that has formed directivity in the target direction based on the corrected null former spectrum calculated by the correction gain multiplication means and the input spectrum.
[0013] The third aspect of the present invention is a sound collection method performed by a sound collection device, the sound collection device having a null former forming means, a correction gain multiplication means, and a target direction acoustic enhancement means, wherein the null former forming means forms a null former that directs a blind spot toward a target direction and suppresses only sounds emitted in the target direction based on an input spectrum of an acoustic signal supplied from a microphone array to obtain a null former spectrum, the correction gain multiplication means multiplies the null former spectrum obtained by the null former forming means by a predetermined correction gain value to calculate a corrected null former spectrum, and the target direction acoustic enhancement means obtains a beamformer output sound that forms directivity in the target direction based on the corrected null former spectrum calculated by the correction gain multiplication means and the input spectrum. [Effects of the Invention]
[0014] According to the present invention, it is possible to provide a sound collection device, a sound collection program, and a sound collection method that suppress distortion of sound collected when performing area sound collection processing to collect sound in a target area. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a block diagram showing a functional configuration of a sound collection device according to an embodiment. [Figure 2] 1 is a block diagram illustrating an example of a hardware configuration of a sound collection device according to an embodiment. [Figure 3] 10A and 10B are diagrams illustrating an example of the positional relationship between a microphone array and a target area according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0016] (A) Main embodiment Hereinafter, a first embodiment of a sound collection device, a sound collection program, and a sound collection method according to the present invention will be described in detail with reference to the drawings.
[0017] (A-1) Configuration of the embodiment FIG. 1 is a block diagram showing the functional configuration of a sound collection device 1 according to this embodiment.
[0018] The sound collection device 1 suppresses sounds emitted in directions other than the front of the microphone array MA based on acoustic signals captured by one microphone array MA having two microphones M11 and M12, and collects sounds emitted in the front of the microphone array MA.
[0019] The two microphones M11 and M12 constitute a microphone array MA. In this embodiment, the number of microphones constituting the microphone array is two, but the number of microphones may be increased to two or more.
[0020] The sound collection device 1 includes a frequency analysis unit 101 , a null former formation unit 102 , a correction gain supply unit 103 , a correction gain multiplication unit 104 , and a target direction sound enhancement unit 105 .
[0021] The frequency analysis unit 101 converts acoustic signals captured by two microphones M11 and M12 (each microphone constituting the microphone array MA) into a frequency domain signal, and obtains the signal as an "input spectrum."
[0022] Based on the input spectrum obtained by the frequency analysis unit 101, the null former forming unit 102 forms a null former with a blind spot facing in the front direction of the microphone array MA, and obtains a "null former spectrum" that suppresses sounds emitted in the front direction.
[0023] The correction gain supplying unit 103 stores a correction gain value for each predetermined frequency, and supplies the correction gain value to the correction gain multiplying unit 104 .
[0024] The correction gain multiplication unit 104 multiplies the null former spectrum obtained by the null former formation unit 102 by the correction gain value supplied from the correction gain supply unit 103 to obtain a "corrected null former spectrum."
[0025] The target direction acoustic enhancement unit 105 forms directivity in the front direction of the microphone array MA based on the corrected null former spectrum obtained by the corrected gain multiplication unit 104 and the input spectrum of the microphone array MA, and obtains a beamformer output sound.
[0026] FIG. 2 is a block diagram showing an example of the hardware configuration of the sound collection device 1. As shown in FIG.
[0027] FIG. 2 shows an example of a hardware configuration when the sound collection device 1 is configured using software (computer).
[0028] 2 includes, as a hardware component, a computer 300 on which a program (including the sound collection program of the embodiment) is installed. The computer 300 may be a computer dedicated to the sound collection program, or may be configured to be shared with programs for other functions.
[0029] The computer 300 shown in FIG. 2 includes a processor 301, a primary storage unit 302, and a secondary storage unit 303. The primary storage unit 302 is a storage unit that functions as a working memory for the processor 301, and may be, for example, a high-speed memory such as a dynamic random access memory (DRAM). The secondary storage unit 303 is a storage unit that records various data such as an operating system (OS) and program data (including data of the sound collection program according to the embodiment), and may be, for example, a non-volatile memory such as a flash (registered trademark) memory, HDD, or SSD. In the computer 300 of this embodiment, when the processor 301 starts up, the OS and programs (including the sound collection program according to the embodiment) recorded in the secondary storage unit 303 are read, deployed on the primary storage unit 302, and executed.
[0030] Note that the specific configuration of the computer 300 is not limited to the configuration in Fig. 2, and various configurations can be applied. For example, if the primary storage unit 302 is a non-volatile memory (e.g., a flash memory), the secondary storage unit 303 may be excluded from the configuration.
[0031] (A-2) Operation of the embodiment Next, the operation of the sound collection device 1 of this embodiment having the above-described configuration (sound collection method according to the embodiment) will be described.
[0032] The microphone array MA converts the input acoustic signal from an analog signal to a digital signal and supplies it to the frequency analysis unit 101 as an input signal.
[0033] The frequency analysis unit 101 performs an arbitrary frequency analysis on a given time-domain input signal, and supplies the obtained input spectrum to the null former formation unit 102. There is no limitation on the frequency analysis method used by the frequency analysis unit 101, and for example, a fast Fourier transform or a wavelet transform may be used.
[0034] The null former forming unit 102 calculates a null former that forms a blind spot in the front direction of the microphone array based on the supplied input spectrum, and supplies the obtained null former spectrum to the correction gain multiplication unit 104.
[0035] Here, the operation of the null former forming unit 102 will be described in detail.
[0036] FIG. 3 is a diagram showing an example of the positional relationship between the microphone array MA and the target area TA according to the embodiment.
[0037] Here, the front direction of the microphone array MA (the direction of 0 degrees / 0 rad) is defined as the direction perpendicular to the line connecting the two microphones. Furthermore, the null former forming unit 102 first forms a blind spot in the front direction to emphasize the sound in the front direction. That is, a null former is formed in the blind spot direction φ=0 rad. Note that, since the target area TA exists in the front direction of the microphone array MA, the following description will be given assuming that the front direction of the microphone array MA coincides with the target direction.
[0038] Here, since the microphone array MA is made up of two microphones, if a blind spot is formed at φ=0 rad, a blind spot is also formed at φ=π rad.
[0039] The input spectrum of microphones M11 and M12 is X M11 (ω), X M12 (ω), the null former forming unit 102 calculates the null former spectrum Y(ω) using the following equation (1): where ω (omega) is the angular frequency, i is the imaginary unit, d is the microphone spacing, and c is the speed of sound.
number
[0040] The correction gain supplying unit 103 stores a correction gain value that corrects the characteristics of the null former spectrum formed by the null former forming unit 102 based on a null former spectrum of a predetermined frequency, and supplies the correction gain value to the correction gain multiplying unit 104.
[0041] Here, any known method can be used to determine the correction gain value, but the following determination method is preferable. null In this case, when the sound source arrives from the direction of arrival θ, the ratio of the input spectrum to the null former spectrum (hereinafter referred to as "null former gain") is θ null (ω, θ) can be expressed by the following equation (2).
number
[0042] As can be seen from equation (2), the null former gain differs for each frequency and for each direction of arrival of the sound source. Therefore, to correct the frequency characteristics of the null former gain using the null former gain of a specific frequency as a reference, it is necessary to determine the reference direction to use as the reference for correcting the frequency characteristics of the null former gain.
[0043] Here, the reference direction is the direction θ where the null former gain of a specific frequency becomes a predetermined value TG [db]. bw The correction gain value is calculated using the above formula. Here, the description is given assuming that TG=-6 dB. TG is not limited to -6 dB, and a more suitable value may be applied based on experiments, simulations, etc.
[0044] The correction gain value G(ω) is the reference direction θ bw For a particular frequency ω ref By taking the ratio of the null former gain of the frequency band to the null former gain for each frequency band, it can be calculated using the following equation (3).
number
[0045] The correction gain multiplication unit 104 obtains a corrected null former spectrum by multiplying the null former spectrum supplied from the null former formation unit 102 by the correction gain value supplied from the correction gain supply unit 103. Specifically, the correction gain multiplication unit 104 corrects the gain of the null former spectrum using the following equation (4) to obtain a corrected null former spectrum Y'(ω). Y´(ω)=Y(ω)G(ω)…(4)
[0046] The target direction acoustic enhancement unit 105 calculates a beamformer spectrum for the target area direction for each frequency based on the supplied corrected null former spectrum and the input spectrum, and outputs the obtained beamformer spectrum (beamformer output sound). Specifically, the target direction acoustic enhancement unit 105 calculates the beamformer spectrum B(ω) for the target area direction based on the following equation (5): B(ω)=X M11 (ω)-Y´(ω)…(5)
[0047] Regarding the calculation of the beamformer spectrum B(ω), in equation (5), the input spectrum X M11 The corrected null former spectrum Y'(ω) is subtracted from (ω), but any value equivalent to the microphone input spectrum is sufficient. M12 (ω) and X M11 (ω) and X M12 The geometric mean or arithmetic mean of (ω) may also be used.
[0048] (A-3) Effects of the embodiment According to this embodiment, the following effects can be achieved.
[0049] In the sound collection device 1 of this embodiment, the frequency-dependent null former spectrum is corrected for each frequency based on the direction where the null former gain of a specific frequency becomes TG, thereby making it possible to make the frequency characteristics near the blind spot direction where the null is formed closer to flat. As a result, in the sound collection device 1, the frequency characteristics of the directionality formed by spectral subtraction can also be made closer to flat, and even if the blind spot direction where the null is directed is different from the direction of the target sound, it is possible to improve the deterioration of sound quality caused by the loss of high-frequency components of the target sound.
[0050] (B) Other embodiments The present invention is not limited to the above-described embodiments, and may include modified embodiments such as those exemplified below.
[0051] (B-1) In the above embodiment, the sound collection device 1 is configured such that a digital signal is supplied from the microphone array MA. However, the sound collection device 1 may be configured such that an analog signal is supplied from the microphone array MA, and the analog signal (the acoustic signal captured by each microphone of each microphone array) is converted into a digital signal on the sound collection device 1 side.
[0052] (B-2) In the sound collection device 1 of the above embodiment, processing is performed to collect target direction sound in the target area direction based on the acoustic signal of one microphone array, but it is also possible to perform area sound collection processing to collect sound in the target area based on the acoustic signals of two or more microphone arrays. [Explanation of symbols]
[0053] 1...sound collection device, 101...frequency analysis unit, 102...null former formation unit, 103...correction gain supply unit, 104...correction gain multiplication unit, 105...target direction sound enhancement unit, M11, M12...microphone, MA...microphone array
Claims
1. a null former generating means for generating a null former spectrum by directing a blind spot in a target direction based on an input spectrum of an acoustic signal supplied from the microphone array and suppressing only sounds emitted in the target direction; a correction gain multiplication means for multiplying the null former spectrum obtained by the null former generation means by a predetermined correction gain value to calculate a corrected null former spectrum; a target direction acoustic enhancement means for acquiring a beamformer output sound in which directivity is formed in the target direction based on the corrected null former spectrum calculated by the corrected gain multiplication means and the input spectrum; A sound collection device comprising:
2. further comprising correction gain holding means for holding the correction gain value of a predetermined frequency; The correction gain multiplication means calculates the corrected null former spectrum using the correction gain value supplied from the correction gain holding means.
2. The sound pickup device according to claim 1.
3. The sound collection device described in claim 2, characterized in that the correction gain value stored in the correction gain storage means is a value obtained by taking the ratio of the null former gain of the specified frequency to the null former gain for each frequency for a specified direction that serves as a reference for correcting the frequency characteristics of the null former spectrum.
4. 2. The sound collection device according to claim 1, wherein the target direction sound enhancement means acquires the beamformer output sound by subtracting the corrected null former spectrum from the input spectrum.
5. Computer, a null former generating means for generating a null former spectrum by directing a blind spot in a target direction based on an input spectrum of an acoustic signal supplied from the microphone array and suppressing only sounds emitted in the target direction; a correction gain multiplication means for multiplying the null former spectrum obtained by the null former generation means by a predetermined correction gain value to calculate a corrected null former spectrum; a target direction acoustic enhancement means for acquiring a beamformer output sound in which directivity is formed in the target direction based on the corrected null former spectrum calculated by the corrected gain multiplication means and the input spectrum; A sound collection program characterized by functioning as follows.
6. In the sound collection method performed by the sound collection device, the sound collection device includes a null former forming means, a correction gain multiplying means, and a target direction sound enhancement means; the null former forming means forms a null former that directs a blind spot in a target direction and suppresses only sounds emitted in the target direction based on an input spectrum of an acoustic signal supplied from a microphone array, thereby obtaining a null former spectrum; the correction gain multiplication means multiplies the null former spectrum obtained by the null former generation means by a predetermined correction gain value to calculate a corrected null former spectrum; The target direction sound enhancement means acquires a beamformer output sound in which directivity is formed in the target direction based on the corrected nullformer spectrum calculated by the corrected gain multiplication means and the input spectrum. A sound collection method characterized by:
Citation Information
Patent Citations
Noise suppressing processor and method therefor
JP2000047699A
Sound processing device and program
JP2010217551A
Noise cancellation apparatus and noise cancellation method
JP2012058360A
Sound processing unit and sound processing method
JP2015118284A
Non-target sound suppression device, method and program
JP2018142826A