Sound pickup device, sound pickup program, and sound pickup method
The sound collection device uses null formers to suppress unwanted sounds and correct frequency characteristics, addressing the need for fewer microphones and distortion in sound collection, achieving efficient and compact sound collection.
Patent Information
- Application Number
- JP2024122375
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2044-07-29
AI Technical Summary
Conventional sound collection methods require a large number of microphones for sharp directivity, making equipment bulky, and spectral subtraction with fewer microphones leads to sound distortion.
A sound collection device and method that uses multiple null formers to form blind spots outside the target area, correcting frequency characteristics to suppress unwanted sounds, and selecting the least distorted output to collect target area sounds without spectral subtraction.
Achieves sharp directivity with fewer microphones, suppressing unwanted sounds while minimizing distortion, thus enhancing sound collection efficiency and reducing equipment size.
Smart Images

Figure 2026020809000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sound collection device, a program, and a method, and can be applied to, for example, a sound collection device that emphasizes only sounds emitted from a specific area and suppresses sounds from other areas. [Background technology]
[0002] When using a speech recognition system in a noisy environment, ambient noise that mixes in with the required target sound is a nuisance that reduces the speech recognition rate of the recorded speech.
[0003] In a conventional environment where multiple sound sources exist, beamforming using a microphone array is a technique for capturing only sounds from a specific direction and obtaining the desired sound while avoiding the inclusion of unwanted sounds. Beamforming is a technique for forming directivity by utilizing the time difference between signals arriving at each microphone (see Non-Patent Document 1). A microphone array that performs beamforming is also called a beamformer.
[0004] On the other hand, because a beamformer collects all sounds from a specific direction, if you want to collect only a specific area (also called a target area), it will also collect sounds outside the target area that are in the same direction as the beamformer. Therefore, Patent Document 1 proposes a method (hereinafter referred to as "area sound collection") of collecting target sounds by using multiple beamformers, each pointing its directivity toward the target area from a different direction and having the directivities intersect at the target area.
[0005] Area pickup requires a beamformer with sharp directionality to form the target area. The beam width of the beamformer should be approximately 60 degrees or less. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-183358 [Non-patent literature]
[0007] [Non-Patent Document 1] Futoshi Asano, "Acoustic Technology Series 16: Array Signal Processing of Sound - Localization, Tracking and Separation of Sound Sources", edited by the Acoustical Society of Japan, Corona Publishing, published February 25, 2011 Summary of the Invention [Problem to be solved by the invention]
[0008] In both additive and subtractive beamformers, a large number of microphones are required to form a sharp directivity pattern. For example, the Minimum Variance Distortionless Response (MVDR) method requires one more microphone than the number of noise sources to be suppressed. While area pickup with such a large number of microphones can be experimentally achieved in a laboratory, it is not suitable for commercial products due to component costs and installation space requirements.
[0009] On the other hand, the technology described in Patent Document 1 achieves sharp directivity with a small number of microphones by using spectral subtraction in the frequency domain (a method of subtracting amplitude spectra from each other, but replacing any negative results with zero or a small number). A null former that forms a blind spot in the direction of the target area is created, and sharp directivity is created in the direction of the target area by spectrally subtracting the null former amplitude spectrum from the microphone input amplitude spectrum. While this method can create sharp directivity with just two microphones, it may result in distortions such as musical noise and loss of area sound components.
[0010] Therefore, conventional technologies have had issues such as the need for a large number of microphones to achieve area sound collection, which makes the equipment large-scale, and the use of spectral subtraction with a small number of microphones can result in distortion of the extracted area sound.
[0011] In view of the above problems, there is a need for a sound collection device, a sound collection program, and a sound collection method that suppress distortion of sounds collected when performing area sound collection processing to collect sounds in a target area. [Means for solving the problem]
[0012] The sound collection device of the first invention comprises: a null former forming means for forming a plurality of null formers for each microphone array based on an input spectrum of an acoustic signal supplied from the plurality of microphone arrays, the null former forming means forming a plurality of null formers directed toward a blind spot outside a target area, and calculating a null former spectrum; a target area direction estimating means for performing a target area direction estimation process for each of the microphone arrays based on the input spectrum and a currently set target area direction, and setting and updating the set target area direction based on the result of the target area direction estimation process; and a null former forming means for calculating a frequency characteristic of the set target area direction set and updated by the target area direction estimating means for each of the null former spectra. The system is characterized by comprising: a null-correction gain value calculation means for calculating and storing a null-correction gain value that brings the signal closer to flat; a null characteristic correction means for acquiring a corrected null former spectrum corrected using the corresponding null-correction gain value for each of the null former spectra; an acoustic enhancement means for selecting, for each of the microphone arrays, the corrected null former spectrum with the smallest output from the corresponding plurality of corrected null former spectra, and acquiring the selected corrected null former spectrum as a beamformer output; and a target area sound acquisition means for extracting and acquiring target area sound whose sound source is the target area, using the beamformer output for each of the microphone arrays acquired by the acoustic enhancement means.
[0013] A sound collection program according to a second aspect of the present invention includes a computer including: a null former forming means for forming, for each microphone array, a plurality of null formers each directed toward a blind spot outside a target area based on an input spectrum of an acoustic signal supplied from the plurality of microphone arrays, and calculating a null former spectrum; a target area direction estimating means for performing, for each of the microphone arrays, a target area direction estimating process for estimating a target area direction based on the input spectrum and a currently set target area direction, and setting and updating the set target area direction based on a result of the target area direction estimation process; and a frequency characteristic of the set target area direction set and updated by the target area direction estimating means for each of the null former spectra. and a null correction gain value calculation means for calculating and storing a null correction gain value that brings the characteristic closer to flatter; a null characteristic correction means for acquiring a corrected null former spectrum corrected using the corresponding null correction gain value for each of the null former spectra; an acoustic enhancement means for selecting, for each of the microphone arrays, the corrected null former spectrum with the smallest output from the corresponding plurality of corrected null former spectra, and acquiring the selected corrected null former spectrum as a beamformer output; and a target area sound acquisition means for extracting and acquiring target area sound having a sound source in the target area, using the beamformer output for each of the microphone arrays acquired by the acoustic enhancement means.
[0014] The third aspect of the present invention is a sound collection method performed by a sound collection device, the sound collection device having null former forming means, target area direction estimating means, null correction gain value calculating means, null characteristic correcting means, audio enhancement means and target area sound acquiring means, wherein the null former forming means forms a plurality of null formers for each of the microphone arrays based on input spectra of audio signals supplied from the plurality of microphone arrays, with blind spots directed to outside the target area, and calculates a null former spectrum, the target area direction estimating means performs a target area direction estimation process for estimating a target area direction for each of the microphone arrays based on the input spectrum and a currently set target area direction, and updates the set target area direction based on a result of the target area direction estimation process, and the null correction gain value calculating means calculates a null former spectrum for each of the null formers. and calculating and storing a null correction gain value for each of the null former spectra that makes the frequency characteristics of the set target area direction set and updated by the target area direction estimation means closer to flatter, the null characteristic correction means acquiring a corrected null former spectrum corrected using the corresponding null correction gain value for each of the null former spectra, the acoustic enhancement means selecting the corrected null former spectrum with the smallest output from the corresponding plurality of corrected null former spectra for each of the microphone arrays and acquiring the selected corrected null former spectrum as a beamformer output, and the target area sound acquisition means extracting and acquiring the target area sound that originates from the target area as a sound source using the beamformer output for each of the microphone arrays acquired by the acoustic enhancement means. [Effects of the Invention]
[0015] According to the present invention, it is possible to provide a sound collection device, a sound collection program, and a sound collection method that suppress distortion of sound collected when performing area sound collection processing to collect sound in a target area. [Brief explanation of the drawings]
[0016] [Figure 1]1 is a block diagram showing a functional configuration of an area sound collection device according to an embodiment. [Figure 2] 1 is a diagram showing an example of the arrangement of microphones (microphone arrays) that supply acoustic signals (input signals) to an area sound collection device according to an embodiment. [Figure 3] 1 is a block diagram showing an example of a hardware configuration of an area sound collection device according to an embodiment. [Figure 4] 10A and 10B are diagrams illustrating examples of directions of null formers in an arbitrary microphone array according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0017] (A) Main embodiment Hereinafter, an embodiment of a sound collection device, a sound collection program, and a sound collection method according to the present invention will be described in detail with reference to the drawings. In this embodiment, an example in which the sound collection device, the sound collection program, and the sound collection method according to the present invention are applied to an area sound collection device will be described.
[0018] (A-1) Configuration of the embodiment FIG. 1 is a block diagram showing the functional configuration of an area sound collection device 1 according to the first embodiment.
[0019] The area sound collection device 1 suppresses sounds emitted outside the target area and collects sounds emitted within the target area based on acoustic signals captured by two microphone arrays having two microphones.
[0020] FIG. 2 is a diagram showing an example of the arrangement of microphones (microphone arrays) that supply acoustic signals (input signals) to the area sound collection device 1. As shown in FIG.
[0021] In FIG. 2, the target area TA from which sound is to be collected is hatched (with a diagonal line pattern).
[0022] As shown in FIG. 2, the area sound collection device 1 of this embodiment performs processing to collect sound from the target area TA as a sound source based on acoustic signals (input signals) supplied from four microphones M11, M12, M21, and M22.
[0023] Here, it is assumed that a first microphone array MA1 is made up of two microphones M11 and M12, and a second microphone array MA2 is made up of microphones M21 and M22. In this embodiment, each microphone array is made up of two microphones (2 channels), and the number of microphone arrays is two (2-microphone array), but the number of microphones and microphone arrays, and the number of microphones per microphone array (number of channels), may be increased as long as there are two or more of each.
[0024] Next, the internal configuration of the area sound collection device 1 will be described with reference to FIG.
[0025] As shown in FIG. 1, the area sound collection device 1 has a frequency analysis unit 101, a null former forming unit 102, a null characteristic correction unit 103, an acoustic enhancement unit 104, a beamformer selection unit 105, a null correction gain calculation unit 106, and a target area direction estimation unit 107.
[0026] The frequency analysis unit 101 converts the acoustic signals (input signals) captured by the two microphone arrays MA1 and MA2 into frequency domain signals to obtain an input spectrum.
[0027] The null former forming unit 102 forms a plurality of null formers, each of which has a blind spot directed outside the target area, based on the input spectrum obtained by the frequency analysis unit 101, and obtains a plurality of null former spectra.
[0028] The destination area direction estimation unit 107 uses the input spectrum for each microphone array obtained by the frequency analysis unit 101 to perform a process of estimating the direction of the destination area for each microphone array (hereinafter referred to as the "destination area direction estimation process") and obtains the resulting direction.
[0029] The null correction gain calculation unit 106 calculates a correction value (hereinafter referred to as a "null correction gain value") for each of the multiple null formers formed by the null former formation unit 102, which corrects the frequency characteristics in the direction obtained as a result of the destination area direction estimation process by the destination area direction estimation unit 107 to be flat, and supplies the null correction gain value to the null characteristic correction unit 103.
[0030] The null characteristic correction unit 103 multiplies each of the multiple null formers calculated by the null former formation unit 102 by multiple null correction gain values (null correction gain values corresponding to each null former spectrum) supplied from the null correction gain calculation unit 106, and obtains the resulting "corrected null former spectrum."
[0031] The acoustic enhancement unit 104 obtains, for each microphone array, a beamformer output in which directivity is formed in the direction of the target area (beamformer output) based on the corrected null former spectrum calculated by the null characteristic correction unit 103.
[0032] The beamformer selection unit 105 extracts and obtains the spectrum of the target area sound (hereinafter referred to as the "area sound spectrum") whose sound source is the target area, based on the beamformer output sound calculated for each microphone array by the acoustic enhancement unit 104. The means and data format by which the beamformer selection unit 105 outputs the area sound spectrum (target area sound) are not limited, and various means and data formats can be applied.
[0033] FIG. 3 is a block diagram showing an example of the hardware configuration of the area sound collection device 1. As shown in FIG.
[0034] FIG. 3 shows an example of a hardware configuration when the area sound collection device 1 is configured using software (computer).
[0035] 3 has, as a hardware component, a computer 300 on which a program (including the sound collection program of the embodiment) is installed. The computer 300 may be a computer dedicated to the sound collection program, or may be configured to be shared with programs for other functions.
[0036] The computer 300 shown in FIG. 3 includes a processor 301, a primary storage unit 302, and a secondary storage unit 303. The primary storage unit 302 is a storage unit that functions as a working memory for the processor 301, and may be, for example, a high-speed memory such as a dynamic random access memory (DRAM). The secondary storage unit 303 is a storage unit that records various data such as an operating system (OS) and program data (including data of the sound collection program according to the embodiment), and may be, for example, a non-volatile memory such as a FLASH (registered trademark) memory, HDD, or SSD. In the computer 300 of this embodiment, when the processor 301 starts up, the OS and programs (including the sound collection program according to the embodiment) recorded in the secondary storage unit 303 are read, deployed on the primary storage unit 302, and executed.
[0037] Note that the specific configuration of the computer 300 is not limited to the configuration in Fig. 3, and various configurations can be applied. For example, if the primary storage unit 302 is a non-volatile memory (e.g., a flash memory), the secondary storage unit 303 may be excluded from the configuration.
[0038] (A-2) Operation of the embodiment Next, the operation of the area sound collecting device 1 of this embodiment having the above-mentioned configuration will be described.
[0039] The microphone arrays MA1 and MA2 convert the input acoustic signals from analog signals into digital signals and supply them to the frequency analysis unit 101 as input signals.
[0040] The frequency analysis unit 101 performs an arbitrary frequency analysis on a given time-domain input signal, and supplies the obtained input spectrum to the null former formation unit 102 and the destination area direction estimation unit 107. There are no limitations on the frequency analysis method used by the frequency analysis unit 101, and for example, fast Fourier transform or wavelet transform may be used.
[0041] Based on the supplied input spectrum, the null former forming unit 102 calculates, for each microphone array, multiple null formers that form blind spots in predetermined directions other than the target area, and supplies the obtained null former spectra to the null characteristic correcting unit 103. The operation of the null former forming unit 102 will be explained in detail. Since the target area is not a point but has a range, the target area direction also has a range.
[0042] Next, the positional relationship between each microphone array and the target area TA will be described.
[0043] FIG. 4 is a diagram showing an example of the direction of null formers in an arbitrary microphone array MAa.
[0044] Hereinafter, the "a" in "MAa" etc. indicates the microphone array number (identifier; ID). Here, the microphone array MA1 is numbered 1, and the microphone array MA2 is numbered 2. In other words, the microphone array MAa shown in FIG. 4 indicates either the microphone array MA1 or MA2 (any microphone array). FIG. 4 illustrates that the microphone array MAa includes microphones ML and MR. For example, if the microphone array MAa is the microphone array MA1, the microphones ML and MR are microphones M11 and M12, respectively, and if the microphone array MAa is the microphone array MA2, the microphones ML and MR are microphones M21 and M22, respectively. In the image diagram of FIG. 4, the shape of the target area TA is illustrated as a circle for convenience of illustration, but in this embodiment, the shape of the target area TA is actually the shape shown in FIG. 2.
[0045] Here, the front direction (0 degree direction) of the microphone array MAa is defined as the direction perpendicular to the line connecting the positions (center positions) of the two microphones ML and MR, and forming the smallest angle with the direction from the center point of the two microphones ML and MR toward the center point (which can also be the center of gravity) of the target area TA (for example, toward the center point of the target area).
[0046] 2, since the target area TA is not a point but has a range, the apparent area direction from each microphone array also has a range. Therefore, here, the target area direction (range of the target area direction) for microphone array MA1 is set to θ1L to θ1R (rad) (-π / 2<θ1L<0<θ1R<π / 2), and the target area direction (range of the target area direction) for microphone array MA2 is set to θ2L to θ2R (rad) (-π / 2<θ2L<0<θ2R<π / 2).
[0047] As described above, the predetermined blind spot direction applied to each null former of the null former forming unit 102 must be outside the target area direction (a direction that does not include the range of the target area direction). Therefore, the n-th blind spot direction φan(rad) (n=1, ..., N) of the microphone array MAa (a=1, 2) must satisfy -π-θaL<φan<θaL or θaR<φan<π-θaR.
[0048] Here, since the microphone array MAa consists of two microphones, if a blind spot is formed in φan, a blind spot will also be formed in π-φan (=-π-φan). Furthermore, since it is difficult to know in advance from which direction noise will arrive, the predetermined blind spot direction may be determined in advance by design (applying a pre-designed fixed value).
[0049] For the above reasons, it is preferable to determine the same number of predetermined blind spot directions to be applied to each null former of the null former forming unit 102 within the two ranges of -π / 2≦φan<θaL and θaR<φan≦π / 2. The number of predetermined blind spot directions set for each microphone array in the null former forming unit 102 (the number of null formers set for each microphone array MAa in the null former forming unit 102) and the combination of blind spot directions are not limited. For example, if the number of predetermined blind spot directions N is 6 (N=6), the following can be selected, and this selection method is preferable. Note that N is not limited to 6 and may be increased or decreased depending on the target area TA, the expected number of interfering sound sources (the number of interfering sound sources and their arrival directions), etc.
[0050] Here, the input spectrum of microphone array a is expressed as X aL (ω), X aR (ω), the null former generating unit 102 generates a plurality of null former spectra Y an Calculate (ω), where ω is the angular frequency, i is the imaginary unit, d is the microphone spacing, and c is the speed of sound.
number
[0051] Next, a specific example of the destination area direction estimation process performed by the destination area direction estimation unit 107 will be described.
[0052] The destination area direction estimation unit 107 determines a new destination area center point direction for each microphone array based on the input spectrum supplied from the frequency analysis unit 101 and the current destination area center point direction θa (rad) (hereinafter, this θa will also be referred to as the "set destination area direction") stored in the destination area direction estimation unit 107, and updates (sets and updates) the direction as the new destination area center point direction θa (rad). The destination area direction estimation unit 107 can use various methods to update the destination area center point direction θa (rad), but the updating method described below is preferred, for example. The initial value of θa (rad) is not limited, and a pre-stored value may be set. For example, the initial value of θa (rad) may be set to the front direction of each microphone array (0 degree direction).
[0053] First, as shown in the following equations (2) and (3), the destination area direction estimation unit 107 estimates three directions, θa, θa+δθ, and θa-δθ, for each microphone array, with the direction of the center point of the current destination area stored in the destination area direction estimation unit 107 as a reference (here, these three directions are generalized as θ ak ) for the null former spectrum ^Y ak Here, δθ is a preset shift value, and any value can be set. In this embodiment, the destination area direction estimation unit 107 calculates the null former spectrum ^Y in three directions, θa, θa+δθ, and θa-δθ, with the center point direction as the reference. ak However, any other combination of directions (for example, any four or more directions based on θa) may be applied. In other words, the null former spectrum ^Y ak The directions used to calculate (ω) preferably include at least a clockwise direction and a counterclockwise direction with θa as the reference (that is, three or more directions including θa).
number
[0054] Next, the destination area direction estimation unit 107 calculates each null former spectrum ^Y ak For (ω), the sum P ak Take.
number
[0055] Then, the destination area direction estimation unit 107 calculates the sum P ak The direction for which the value of is smallest is updated as the new center point direction of the destination area. For example, if the sum value of the null former spectrum for θa+δθ is smallest, the destination area direction estimation unit 107 updates θa+δθ as the new center point direction of the destination area. Then, the destination area direction estimation unit 107 supplies the updated center point direction of the destination area to the null correction gain calculation unit 106.
[0056] The null correction gain calculation unit 106 calculates a null correction gain value for each of the multiple null former spectra formed in a predetermined blind spot direction other than the target area, which null correction gain value flattens the frequency characteristics of the null former gain (ratio between the input spectrum and the null former spectrum) in the direction of the center point of the target area supplied from the target area direction estimation unit 107, and supplies the calculated value to the null characteristic correction unit 103.
[0057] Here, the null-correction gain calculation unit 106 can use various methods to determine the null-correction gain value, but for example, the following determination method is preferable.
[0058] Here, the predetermined blind spot direction φ an Consider the null former of (rad). In this case, the null former gain H an (ω, θ) can be expressed by the following equation (5), where θ is the direction of arrival of the sound source.
number
[0059] From equation (5), we can see that the null former gain differs for each angular frequency and for each direction from which the sound source arrives. Therefore, in order to flatten the frequency characteristics of the null former gain in the direction of the center point of the target area, it is necessary to correct the null former gain, which differs for each frequency, using the direction of the center point of the target area (θa (rad)) as the reference. The null correction gain value can be calculated by taking the reciprocal of the null former gain when the target sound arrives from the direction of the center point of the target area, θa (rad), as shown in the following equation (6).
number
[0060] The null characteristic corrector 103 obtains a corrected null former spectrum by multiplying the multiple null former spectra supplied from the null former forming unit 102 by multiple null correction gain values supplied from the null correction gain calculating unit 106. Specifically, the null characteristic corrector 103 obtains the corrected null former spectrum using the following equation (7).
number
[0061] The acoustic enhancement unit 104 selects, for each microphone array and for each frequency, the corrected null former spectrum that has the smallest amplitude based on the multiple corrected null former spectra supplied from the null characteristic correction unit 103, and supplies the obtained beamformer spectrum (beamformer output sound) to the beamformer selection unit 105. Specifically, the acoustic enhancement unit 104 calculates the beamformer spectrum Ba(ω) of the microphone array a using the following equations (8) and (9): where v(ω) is the index number of the blind spot angle at which the corrected null former spectrum is smallest.
number
[0062] The beamformer selection unit 105 selects, for each frequency based on the two supplied beamformer spectra, the one that minimizes the amplitude of the corrected beamformer spectrum, thereby selecting it as the spectrum of the target area sound (area sound spectrum) whose sound source is the target area, and outputs the obtained area sound spectrum. Specifically, the beamformer selection unit 105 calculates the area sound spectrum Z(ω) based on the following equations (10) and (11).
number
[0063] (A-3) Effects of the embodiment According to this embodiment, the following effects can be achieved.
[0064] In the area sound collection device 1 of this embodiment, multiple null formers are calculated in parallel, the characteristics are corrected for each null former so that the frequency characteristics in the direction of the center point of the target area are flat, and the smallest output from among these outputs is selected to suppress sounds arriving from directions other than the target area direction, forming a beamformer that collects sound only in the direction of the target area. In other words, the area sound collection device 1 of this embodiment achieves area sound collection processing based on the outputs of multiple beamformers without using spectral subtraction. As a result, the area sound collection device 1 of this embodiment can collect only area sounds without causing distortion that occurs with spectral subtraction.
[0065] (B) Other embodiments The present invention is not limited to the above-described embodiments, and may include modified embodiments such as those exemplified below.
[0066] (B-1) In the area sound collection device 1 of each of the above embodiments, a configuration has been shown in which digital signals are supplied from the microphone arrays MA1 and MA2, but the area sound collection device 1 may also be configured to supply analog signals from the microphone arrays MA1 and MA2, and convert the analog signals (acoustic signals captured by each microphone of each microphone array) into digital signals on the area sound collection device 1 side. [Explanation of symbols]
[0067] 1...area sound collection device, 101...frequency analysis unit, 102...null former formation unit, 103...filter gain calculation unit, 104...acoustic enhancement unit, 105...beam former selection unit, 106...null correction gain supply unit, 107...target area direction estimation unit, M11, M12, M21, M22, MR, ML...microphone, MA1, MA2, MAa...microphone array, TA...target area
Claims
1. a null former generating means for generating a plurality of null formers each having a blind spot outside a target area based on an input spectrum of an acoustic signal supplied from the plurality of microphone arrays, and calculating a null former spectrum; a destination area direction estimation means for performing a destination area direction estimation process for estimating a destination area direction based on the input spectrum and a currently set destination area direction for each of the microphone arrays, and setting and updating the set destination area direction based on a result of the destination area direction estimation process; a null-correction gain value calculation means for calculating and storing a null-correction gain value for each of the null former spectra, the null-correction gain value being used to make the frequency characteristics of the set target area direction, which have been set and updated by the target area direction estimation means, closer to flat; a null characteristic correction means for correcting each of the null former spectra using the corresponding null-correction gain value to obtain a corrected null former spectrum; an acoustic enhancement means for selecting, for each of the microphone arrays, the correction null former spectrum having the smallest output from the corresponding plurality of correction null former spectra, and acquiring the selected correction null former spectrum as a beamformer output; a target area sound acquisition means for extracting and acquiring a target area sound having a sound source in the target area using the beamformer output for each of the microphone arrays acquired by the sound enhancement means; A sound collection device comprising:
2. The sound collection device according to claim 1, wherein the null correction gain value for each null former spectrum calculated by the null correction gain value calculation means is obtained by taking the reciprocal of the null former gain for the set target area direction.
3. Computer, a null former generating means for generating a plurality of null formers each having a blind spot outside a target area based on an input spectrum of an acoustic signal supplied from the plurality of microphone arrays, and calculating a null former spectrum; a destination area direction estimation means for performing a destination area direction estimation process for estimating a destination area direction based on the input spectrum and a currently set destination area direction for each of the microphone arrays, and setting and updating the set destination area direction based on a result of the destination area direction estimation process; a null-correction gain value calculation means for calculating and storing a null-correction gain value for each of the null former spectra, the null-correction gain value being used to make the frequency characteristics of the set target area direction, which have been set and updated by the target area direction estimation means, closer to flat; a null characteristic correction means for correcting each of the null former spectra using the corresponding null-correction gain value to obtain a corrected null former spectrum; an acoustic enhancement means for selecting, for each of the microphone arrays, the correction null former spectrum having the smallest output from the corresponding plurality of correction null former spectra, and acquiring the selected correction null former spectrum as a beamformer output; a target area sound acquisition means for extracting and acquiring a target area sound having a sound source in the target area using the beamformer output for each of the microphone arrays acquired by the sound enhancement means; A sound collection program characterized by functioning as follows.
4. In the sound collection method performed by the sound collection device, The sound collection device includes a null former forming means, a target area direction estimating means, a null correction gain value calculating means, a null characteristic correcting means, an audio enhancing means, and a target area sound acquiring means, the null former generating means generates a plurality of null formers each having a blind spot outside a target area based on an input spectrum of an acoustic signal supplied from a plurality of microphone arrays, and calculates a null former spectrum; the destination area direction estimation means performs a destination area direction estimation process for estimating a destination area direction for each of the microphone arrays based on the input spectrum and a currently set destination area direction, and updates the set destination area direction based on a result of the destination area direction estimation process; the null-correction gain value calculation means calculates and stores a null-correction gain value for each of the null former spectra that makes the frequency characteristics of the set target area direction set and updated by the target area direction estimation means closer to flatter; the null characteristic correction means obtains a corrected null former spectrum by correcting each of the null former spectra using the corresponding null correction gain value; the acoustic enhancement means selects, for each of the microphone arrays, the correction null former spectrum having the smallest output from the corresponding plurality of correction null former spectra, and acquires the selected correction null former spectrum as a beamformer output; The target area sound acquisition means extracts and acquires target area sound having a sound source in the target area using the beamformer output for each microphone array acquired by the sound enhancement means. A sound collection method characterized by:
Citation Information
Patent Citations
Sound pickup device and program
JP2013183358A