Sound collecting device, sound collecting program, and sound collecting method
The sound collection device addresses the issue of frequency-dependent directivity in conventional beamformers by using a null-former forming mechanism and correction gain multiplication to enhance target-direction acoustic signals and reduce distortion.
Patent Information
- Application Number
- JP2024025637
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-02-22
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-02-22
AI Technical Summary
Conventional beamformers using spectral subtraction suffer from frequency-dependent directivity, leading to distortion and suppression of high-frequency components of target sounds when the dead angle direction and the target sound direction are misaligned.
A sound collection device and method that employs a null-former forming mechanism to create a null-former spectrum, followed by correction gain multiplication to correct frequency characteristics, thereby enhancing target-direction acoustic signals while suppressing unwanted sounds.
The proposed solution effectively reduces distortion and improves sound quality by flattening the frequency characteristics of the null-former spectrum, ensuring accurate collection of target sounds even when the target and null directions are not perfectly aligned.
Smart Images

Figure 0007687466000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sound collection device, a sound collection program, and a sound collection method, and can be applied to, for example, a sound collection device that emphasizes only sounds arriving from a specific direction and suppresses sounds arriving from other directions.
Background Art
[0002] When using a speech recognition system in a noisy environment, ambient noise that is mixed in simultaneously with the necessary target sound is a troublesome existence that causes a decrease in the speech recognition rate of the recorded speech.
[0003] Conventionally, in an environment where there are such multiple sound sources, as a technique for obtaining necessary target sounds while avoiding the mixing of unwanted sounds by picking up only the sound in a specific direction, there is beamforming using a microphone array. Beamforming is a technique for forming directivity by utilizing the time difference of signals reaching each microphone (see Non-Patent Document 1). A microphone array that performs beamforming is also called a beamformer. As the basic configuration of a beamformer, there are addition-type and subtraction-type beamformers, which are composed of a delay device and an adder / subtractor. First, the delay device calculates the arrival time difference when the sound existing in the target sound direction arrives at each microphone, and adds a delay to align the phase of the target sound. In the addition type, by adding the input signals of each microphone with the phases aligned, a directivity having high sensitivity with respect to the target sound direction is formed. In the subtraction type, by subtracting the input signal of one microphone from the input signal of the other microphone with the phases aligned, a directivity having a sharp dead angle (null) in a specific direction is formed. When the target sound and the noise source are in an azimuth close to each other as seen from the beamformer, in order to avoid the mixing of unwanted sounds and obtain only the necessary target sounds, it is necessary to form a sharper directivity. To form a sharp directivity, a large number of microphones are required in both addition-type and subtraction-type beamformers. For example, in the Minimum Variance Distortionless Response (MVDR) method, the number of microphones equal to the number of noise sources to be suppressed + 1 is required. Such beamforming with a large number of microphones can be experimentally realized in a laboratory, but it is not suitable for commercial products in terms of component cost and installation space.
[0004] In response to such problems, as a method for forming a sharp directivity using a small number of microphones, there is a spectrum subtraction method such as that in Patent Document 1.
[0005] In the technology described in Patent Document 1, by using spectral subtraction in the frequency domain (subtracting the amplitude spectra from each other, and replacing with zero or a small number when the subtraction result is negative), sharp directivity is achieved with a small number of microphones. A nullformer that forms a dead zone in the direction of the target area is created, and by spectrally subtracting the nullformer amplitude spectrum from the microphone input amplitude spectrum, sharp directivity is formed in the direction of the target area. By using this method, sharp directivity can be formed with only two microphones.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Non - Patent Documents
[0007]
Non - Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] As described above, according to the technology described in Patent Document 1, sharp directivity can be formed with only two microphones. However, the directivity formed by the technology described in Patent Document 1 has frequency dependence. The directivity depends on the microphone interval with respect to the wavelength, and the directivity formed by the technology described in Patent Document 1 becomes wider at lower frequencies and narrower at higher frequencies based on the microphone array interval. This problem of frequency dependence also exists in other beamformers with similar issues.
[0009] Therefore, in a conventional beamformer using spectral subtraction, when the dead angle direction for forming a null and the direction of the target sound are misaligned, due to the problem of the frequency dependence of directivity, there is a problem that the high-frequency component of the target sound is suppressed and distorted.
[0010] In view of the above problems, it is possible to provide a sound collection device, a sound collection program, and a sound collection method that suppress distortion of the sound collected when collecting the sound in the target direction.
Means for Solving the Problems
[0011] A first aspect of the present invention is a null-former forming means for forming a null-former that faces a dead angle in a target direction and suppresses only the sound emitted in the target direction based on an input spectrum of an acoustic signal supplied from a microphone array to obtain a null-former spectrum, a correction gain multiplying means for multiplying the null-former spectrum obtained by the null-former forming means by a predetermined correction gain value to calculate a corrected null-former spectrum, and a target-direction acoustic enhancement means for obtaining a beamformer output sound having directivity in the target direction based on the corrected null-former spectrum calculated by the correction gain multiplying means and the input spectrum , function as correction gain holding means for holding the correction gain value of a predetermined frequency. The correction gain multiplication means calculates the corrected null-formation spectrum using the correction gain value supplied from the correction gain holding means. The correction gain value held by the correction gain holding means is a value obtained by taking the ratio of the null-formation gain of the predetermined frequency to the null-formation gain for each frequency in a predetermined direction that serves as a reference for correcting the frequency characteristics of the null-formation spectrum. characterized by the above.
[0012] A sound collection program according to a second aspect of the present invention causes a computer to perform null-former forming means for forming a null-former that faces a dead angle in a target direction and suppresses only the sound emitted in the target direction based on an input spectrum of an acoustic signal supplied from a microphone array to obtain a null-former spectrum, correction gain multiplying means for multiplying the null-former spectrum obtained by the null-former forming means by a predetermined correction gain value to calculate a corrected null-former spectrum, and target-direction acoustic enhancement means for obtaining a beamformer output sound having directivity in the target direction based on the corrected null-former spectrum calculated by the correction gain multiplying means and the input spectrum , function as correction gain holding means for holding the correction gain value of a predetermined frequency. The correction gain multiplication means calculates the corrected null-formation spectrum using the correction gain value supplied from the correction gain holding means. The correction gain value held by the correction gain holding means is a value obtained by taking the ratio of the null-formation gain of the predetermined frequency to the null-formation gain for each frequency in a predetermined direction that serves as a reference for correcting the frequency characteristics of the null-formation spectrum. characterized by the above.
[0013] In the third invention, in the sound collection method performed by the sound collection device, the sound collection device includes a nullformer formation means and a correction gain multiplication means 、 Target direction acoustic enhancement means and correction gain holding means and the nullformer formation means forms a nullformer that suppresses only the sound emitted in the target direction by directing a dead angle in the target direction based on the input spectrum of the acoustic signal supplied from the microphone array, and obtains a nullformer spectrum. The correction gain multiplication means multiplies the nullformer spectrum obtained by the nullformer formation means by a predetermined correction gain value to calculate a corrected nullformer spectrum, and the target direction acoustic enhancement means obtains a beamformer output sound having directivity in the target direction based on the corrected nullformer spectrum calculated by the correction gain multiplication means and the input spectrum , wherein the correction gain holding means holds the correction gain value of a predetermined frequency, the correction gain multiplication means calculates the corrected null-formation spectrum using the correction gain value supplied from the correction gain holding means, and the correction gain value held by the correction gain holding means is a value obtained by taking the ratio of the null-formation gain of the predetermined frequency to the null-formation gain for each frequency in a predetermined direction that serves as a reference for correcting the frequency characteristics of the null-formation spectrum.
Advantages of the Invention
[0014] According to the present invention, it is possible to provide a sound collection device, a sound collection program, and a sound collection method that suppress distortion of the sound collected when performing area sound collection processing for collecting sound in a target area
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Modes for Carrying Out the Invention
[0016] (A) Main Embodiment Hereinafter, a first embodiment of a sound collection device, a sound collection program, and a sound collection method according to the present invention will be described in detail with reference to the drawings
[0017] (A-1) Configuration of the Embodiment FIG. 1 is a block diagram showing the functional configuration of the sound collection device 1 according to this embodiment.
[0018] The sound collection device 1 suppresses sound emitted outside the front direction of the microphone array MA and collects sound emitted in the front direction of the microphone array MA based on the acoustic signals captured by one microphone array MA having two microphones M11 and M12.
[0019] The two microphones M11 and M12 constitute the microphone array MA. In this embodiment, the number of microphones constituting the microphone array is two, but the number of microphones may be increased as long as it is two or more.
[0020] The sound collection device 1 includes a frequency analysis unit 101, a nullformer formation unit 102, a correction gain supply unit 103, a correction gain multiplication unit 104, and a target direction acoustic enhancement unit 105.
[0021] The frequency analysis unit 101 obtains, as an "input spectrum", a conversion of the acoustic signals captured by the two microphones M11 and M12 (each microphone constituting the microphone array MA) into frequency domain signals.
[0022] The nullformer formation unit 102 forms a nullformer with a dead angle facing the front direction of the microphone array MA based on the input spectrum obtained by the frequency analysis unit 101, and obtains, as a "nullformer spectrum", a suppression of the sound emitted in the front direction.
[0023] The correction gain supply unit 103 stores correction gain values for each predetermined frequency and supplies the correction gain values to the correction gain multiplication unit 104.
[0024] The null-former gain multiplication unit 104 obtains, as the "corrected null-former spectrum," the product of the null-former spectrum obtained by the null-former forming unit 102 and the correction gain value supplied from the correction gain supply unit 103.
[0025] Based on the corrected null-former spectrum obtained by the corrected gain multiplication unit 104 and the input spectrum of the microphone array MA, the target direction acoustic enhancement unit 105 forms directivity in the front direction of the microphone array MA to obtain the beamformer output sound.
[0026] FIG. 2 is a block diagram showing an example of the hardware configuration of the sound collection device 1.
[0027] FIG. 2 shows an example of the hardware configuration when the sound collection device 1 is configured using software (computer).
[0028] The sound collection device 1 shown in FIG. 2 has a computer 300 in which a program (including the sound collection program of the embodiment) is installed as a hardware component. Further, the computer 300 may be a computer dedicated to the sound collection program or may be configured to be shared with programs having other functions.
[0029] The computer 300 shown in FIG. 2 has a processor 301, a primary storage unit 302, and a secondary storage unit 303. The primary storage unit 302 is a storage means that functions as a working memory (work memory) of the processor 301. For example, a high-speed operating memory such as DRAM (Dynamic Random Access Memory) can be applied. The secondary storage unit 303 is a storage means for recording various data such as an OS (Operating System) and program data (including data of the sound collection program according to the embodiment). For example, a non-volatile memory such as a FLASH (registered trademark) memory, an HDD, or an SSD can be applied. In the computer 300 of this embodiment, when the processor 301 is activated, the OS and programs (including the sound collection program according to the embodiment) recorded in the secondary storage unit 303 are read and expanded and executed on the primary storage unit 302.
[0030] Note that the specific configuration of the computer 300 is not limited to the configuration in FIG. 2, and various configurations can be applied. For example, if the primary storage unit 302 is a non-volatile memory (e.g., FLASH memory, etc.), the configuration excluding the secondary storage unit 303 may be used.
[0031] (A-2) Operations of the Embodiment Next, the operation of the sound collection device 1 of this embodiment (sound collection method according to the embodiment) having the above configuration will be described.
[0032] The microphone array MA converts the input acoustic signal from an analog signal to a digital signal and supplies it as an input signal to the frequency analysis unit 101.
[0033] The frequency analysis unit 101 performs arbitrary frequency analysis on the input signal in the given time domain and supplies the obtained input spectrum to the null-former formation unit 102. In the frequency analysis unit 101, there is no limitation on the frequency analysis method. For example, a fast Fourier transform or a wavelet transform may be used.
[0034] Based on the supplied input spectrum, the nullformer forming unit 102 calculates a nullformer that forms a dead angle in the front direction of the microphone array, and supplies the obtained nullformer spectrum to the correction gain multiplication unit 104.
[0035] Here, the operation of the nullformer forming unit 102 will be described in detail.
[0036] FIG. 3 is a diagram showing an example of the positional relationship between the microphone array MA and the target area TA according to the embodiment.
[0037] Here, the front direction (the direction of 0 degrees / 0 rad) of the microphone array MA is defined as the direction orthogonal to the straight line connecting the two microphones. Also, in the nullformer forming unit 102, in order to emphasize the sound in the front direction, first, a dead angle is formed in the front direction. That is, a nullformer is formed in the dead angle direction φ = 0 rad. Here, since the target area TA exists in the front direction of the microphone array MA, the description will be made assuming that the front direction of the microphone array MA coincides with the target direction.
[0038] Here, since there are two microphones constituting the microphone array MA, when a dead angle is formed at φ = 0 rad, a dead angle is also formed at φ = π rad.
[0039] Let the input spectra of the microphones M11 and M12 be X M11 (ω) and X M12 (ω). Then, the nullformer forming unit 102 calculates the nullformer spectrum Y(ω) using the following equation (1). Here, ω (omega) is the angular frequency, i is the imaginary unit, d is the microphone interval, and c is the speed of sound.
Equation
[0040] In the correction gain supply unit 103, a correction gain value for correcting the characteristics of the null-form spectrum formed by the null-form forming unit 102 is stored with reference to a null-form spectrum of a predetermined frequency determined in advance, and the correction gain value is supplied to the correction gain multiplication unit 104.
[0041] Here, any known method can be used to determine the correction gain value, but the following determination method is preferable. Hereinafter, the dead angle direction of the null-form is set to be directed to θ null At this time, when the sound source arrives from the arrival direction θ, the ratio of the input spectrum to the null-form spectrum (hereinafter referred to as "null-form gain") θ null (ω, θ) can be expressed by the following equation (2).
Equation
[0042] As can be seen from equation (2), it can be seen that the null-form gain varies for each frequency and for each arrival direction of the sound source. Therefore, in order to correct the frequency characteristics of the null-form gain based on the null-form gain of a specific frequency, it is necessary to determine the direction serving as the reference for correcting the frequency characteristics of the null-form gain and the reference direction.
[0043] Here, the reference direction is set to be the direction θ bw in which the null-form gain of a specific frequency becomes a predetermined value TG [dB], and the correction gain value is calculated. Here, it is assumed that TG = -6 dB for explanation. TG is not limited to -6 dB, and a suitable value may be applied by experiments, simulations, etc.
[0044] The correction gain value G(ω) is the ratio of the null-form gain of a specific frequency ω bw for the direction θ ref serving as a reference for correcting the frequency characteristics to the null-form gain for each frequency, and can be calculated by the following equation (3). [Number]
[0045] The null-former gain multiplication unit 104 multiplies the null-former spectrum supplied from the null-former formation unit 102 by the correction gain value supplied from the correction gain supply unit 103 to obtain a corrected null-former spectrum. Specifically, the null-former gain multiplication unit 104 corrects the gain of the null-former spectrum using the following equation (4) to obtain a corrected null-former spectrum Y′(ω). Y′(ω) = Y(ω)G(ω) … (4)
[0046] The target direction acoustic enhancement unit 105 calculates a beamformer spectrum in the target area direction for each frequency based on the supplied corrected null-former spectrum and the input spectrum, and outputs the obtained beamformer spectrum (beamformer output sound). Specifically, the target direction acoustic enhancement unit 105 calculates a beamformer spectrum B(ω) in the target area direction based on the following equation (5). B(ω) = X M11 (ω) − Y′(ω) … (5)
[0047] Regarding the calculation of the beamformer spectrum B(ω), in equation (5), the corrected null-former spectrum Y′(ω) is subtracted from the input spectrum X M11 (ω), but any value corresponding to the input spectrum of the microphone may be used, such as X M12 (ω), or the geometric mean or arithmetic mean of X M11 (ω) and X M12 (ω).
[0048] (A-3) Effects of the Embodiment According to this embodiment, the following effects can be achieved.
[0049] In the sound collection device 1 of this embodiment, by correcting the null-form gain for each frequency with reference to the direction in which the null-form spectrum having frequency dependence has a null-form gain of TG, it is possible to make the frequency characteristics near the dead angle direction where the null is formed closer to flat. As a result, in the sound collection device 1, it is also possible to make the frequency characteristics of the directivity formed by spectral subtraction closer to flat, and even when the dead angle direction facing the null and the direction of the target sound are misaligned, it is possible to improve the sound quality degradation caused by the loss of high-frequency components of the target sound or the like.
[0050] (B) Other embodiments The present invention is not limited to the above-described embodiments, and modified embodiments as exemplified below can also be cited.
[0051] (B-1) In the sound collection device 1 of the above-described embodiment, a configuration in which a digital signal is supplied from the microphone array MA has been shown. However, an analog signal may be supplied from the microphone array MA to the sound collection device 1, and the sound collection device 1 may be configured to convert the analog signal (the acoustic signal captured by each microphone of each microphone array) into a digital signal.
[0052] (B-2) In the sound collection device 1 of the above-described embodiment, a process of collecting the target direction sound in the target area direction based on the acoustic signal of one microphone array has been performed. However, an area sound collection process of collecting the sound in the target area based on the acoustic signals of two or more microphone arrays may be performed.
Explanation of reference numerals
[0053] 1... Sound collection device, 101... Frequency analysis unit, 102... Null-form formation unit, 103... Correction gain supply unit, 104... Correction gain multiplication unit, 105... Target direction sound enhancement unit, M11, M12... Microphones, MA... Microphone array
Claims
1. a null former forming means for forming a null former that directs a blind spot in a target direction based on an input spectrum of an acoustic signal supplied from a microphone array and suppresses only sounds emitted in the target direction to obtain a null former spectrum; a correction gain multiplication means for multiplying the null former spectrum obtained by the null former generation means by a predetermined correction gain value to calculate a corrected null former spectrum; a target direction sound enhancement unit that acquires a beamformer output sound having directivity formed in the target direction based on the corrected null former spectrum calculated by the corrected gain multiplication unit and the input spectrum; a correction gain holding means for holding the correction gain value of a predetermined frequency, the correction gain multiplication means calculates the corrected null former spectrum using the correction gain value supplied from the correction gain holding means; The correction gain value held by the correction gain holding means is a value obtained by taking a ratio between the null former gain of the predetermined frequency and the null former gain for each frequency in a predetermined direction that is a reference for correcting the frequency characteristic of the null former spectrum. A sound collecting device characterized by the above.
2. 2. The sound collection device according to claim 1, wherein the target direction sound emphasis means obtains the beamformer output sound by subtracting the corrected null former spectrum from the input spectrum.
3. Computer, a null former forming means for forming a null former that directs a blind spot in a target direction based on an input spectrum of an acoustic signal supplied from a microphone array and suppresses only sounds emitted in the target direction to obtain a null former spectrum; a correction gain multiplication means for multiplying the null former spectrum obtained by the null former generation means by a predetermined correction gain value to calculate a corrected null former spectrum; a target direction sound enhancement unit that acquires a beamformer output sound having directivity formed in the target direction based on the corrected null former spectrum calculated by the corrected gain multiplication unit and the input spectrum; functioning as a correction gain holding means for holding the correction gain value of a predetermined frequency; the correction gain multiplication means calculates the corrected null former spectrum using the correction gain value supplied from the correction gain holding means; The correction gain value held by the correction gain holding means is a value obtained by taking a ratio between the null former gain of the predetermined frequency and the null former gain for each frequency in a predetermined direction that is a reference for correcting the frequency characteristic of the null former spectrum. A sound recording program characterized by:
4. In the sound collection method performed by the sound collection device, The sound collection device includes a null former forming means, a correction gain multiplying means, a target direction sound enhancement means, and a correction gain holding means, the null former forming means forms a null former that directs a blind spot in a target direction and suppresses only sounds emitted in the target direction based on an input spectrum of an acoustic signal supplied from a microphone array, thereby obtaining a null former spectrum; the correction gain multiplication means multiplies the null former spectrum obtained by the null former formation means by a predetermined correction gain value to calculate a corrected null former spectrum; the target direction sound emphasis means acquires a beamformer output sound in which directivity is formed in the target direction based on the corrected null former spectrum calculated by the corrected gain multiplication means and the input spectrum, the correction gain holding means holds the correction gain value for a predetermined frequency; the correction gain multiplication means calculates the corrected null former spectrum using the correction gain value supplied from the correction gain holding means; The correction gain value held by the correction gain holding means is a value obtained by taking a ratio between the null former gain of the predetermined frequency and the null former gain for each frequency in a predetermined direction that is a reference for correcting the frequency characteristic of the null former spectrum. A sound collection method comprising:
Citation Information
Patent Citations
Noise suppressing processor and method therefor
JP2000047699A
Sound processing device and program
JP2010217551A
Noise cancellation apparatus and noise cancellation method
JP2012058360A
Sound pickup device and program
JP2013183358A
Sound processing unit and sound processing method
JP2015118284A