Sound collection device, sound collection program, and sound collection method
The sound collection device forms null formers to suppress interference and collect target area sound without distortion, addressing the need for fewer microphones and reducing equipment scale.
Patent Information
- Application Number
- JP2024017413
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-07
- Publication Date
- 2025-08-20
- Estimated Expiration
- 2044-02-07
AI Technical Summary
Conventional sound collection methods require a large number of microphones for sharp directivity, leading to large-scale equipment, and using spectral subtraction with a small number of microphones results in sound distortion.
A sound collection device and method that forms multiple null formers directed towards blind spots outside the target area, calculates a filter gain for a frequency filter to pass components in the target area direction, and enhances acoustic signals to suppress interference and collect sound without distortion.
The method effectively suppresses interference and collects sound from a target area without distortion, using a reduced number of microphones and minimizing equipment size.
Smart Images

Figure 2025121743000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sound collection device, a program, and a method, and can be applied to, for example, a sound collection device that emphasizes only sounds emitted from a specific area and suppresses sounds from other areas. [Background technology]
[0002] When using a speech recognition system in a noisy environment, ambient noise that mixes in with the required target sound is a nuisance that reduces the speech recognition rate of the recorded speech.
[0003] In a conventional environment where multiple sound sources exist, beamforming using a microphone array is a technique for capturing only sounds from a specific direction and obtaining the desired sound while avoiding the inclusion of unwanted sounds. Beamforming is a technique for forming directivity by utilizing the time difference between signals arriving at each microphone (see Non-Patent Document 1). A microphone array that performs beamforming is also called a beamformer.
[0004] On the other hand, because a beamformer collects all sounds from a specific direction, if you want to collect only a specific area (also called a target area), it will also collect sounds outside the target area that are in the same direction as the beamformer. Therefore, Patent Document 1 proposes a method (hereinafter referred to as "area sound collection") of collecting target sounds by using multiple beamformers, each pointing its directivity toward the target area from a different direction and having the directivities intersect at the target area.
[0005] Area pickup requires a beamformer with sharp directionality to form the target area. The beam width of the beamformer should be approximately 60 degrees or less. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-183358 [Non-patent literature]
[0007] [Non-Patent Document 1] Futoshi Asano, "Acoustic Technology Series 16: Array Signal Processing of Sound - Localization, Tracking and Separation of Sound Sources", edited by the Acoustical Society of Japan, Corona Publishing, published February 25, 2011 Summary of the Invention [Problem to be solved by the invention]
[0008] In both additive and subtractive beamformers, a large number of microphones are required to form a sharp directivity pattern. For example, the Minimum Variance Distortionless Response (MVDR) method requires one more microphone than the number of noise sources to be suppressed. While area pickup with such a large number of microphones can be experimentally achieved in a laboratory, it is not suitable for commercial products due to component costs and installation space requirements.
[0009] On the other hand, the technology described in Patent Document 1 achieves sharp directivity with a small number of microphones by using spectral subtraction in the frequency domain (a method of subtracting amplitude spectra from each other, but replacing any negative results with zero or a small number). A null former that forms a blind spot in the direction of the target area is created, and sharp directivity is created in the direction of the target area by spectrally subtracting the null former amplitude spectrum from the microphone input amplitude spectrum. While this method can create sharp directivity with just two microphones, it may result in distortions such as musical noise and loss of area sound components.
[0010] Therefore, conventional technologies have had issues such as the need for a large number of microphones to achieve area sound collection, which makes the equipment large-scale, and the use of spectral subtraction with a small number of microphones can result in distortion of the extracted area sound.
[0011] In view of the above problems, there is a need for a sound collection device, a sound collection program, and a sound collection method that suppress distortion of sounds collected when performing area sound collection processing to collect sounds in a target area. [Means for solving the problem]
[0012] The first sound collection device of the present invention is characterized by having: a null former forming means that forms a plurality of null formers directed toward blind spots outside a target area for each microphone array based on the input spectrum of acoustic signals supplied from the plurality of microphone arrays and calculates a null former spectrum; a filter gain calculation means that calculates a filter gain of a frequency filter that passes components in the target area direction for each microphone array based on the null former spectrum calculated by the null former forming means; an acoustic enhancement means that calculates, for each microphone array, the filter gain calculated by the filter gain calculation means and a beamformer output sound with directivity formed in the target area direction based on the input spectrum; and a target area sound extraction means that forms a target area sound based on the beamformer output sound calculated by the acoustic enhancement means.
[0013] The sound collection program of the second invention is characterized in that it causes a computer to function as: a null former forming means that forms, for each microphone array, a plurality of null formers directed toward blind spots outside a target area based on the input spectrum of acoustic signals supplied from the plurality of microphone arrays and calculates a null former spectrum; a filter gain calculation means that calculates, for each microphone array, a filter gain of a frequency filter that passes components in the target area direction based on the null former spectrum calculated by the null former forming means; an acoustic enhancement means that calculates, for each microphone array, the filter gain calculated by the filter gain calculation means and a beamformer output sound with directivity formed in the target area direction based on the input spectrum; and a target area sound extraction means that forms a target area sound based on the beamformer output sound calculated by the acoustic enhancement means.
[0014] A third sound collection method of the present invention is a sound collection method performed by a sound collection device, wherein the sound collection device has null former forming means, filter gain calculation means, acoustic enhancement means, and target area sound extraction means, wherein the null former forming means forms a plurality of null formers facing blind spots outside the target area for each microphone array based on input spectra of acoustic signals supplied from the plurality of microphone arrays to calculate a null former spectrum, the filter gain calculation means calculates a filter gain of a frequency filter that passes components in the target area direction for each microphone array based on the null former spectrum calculated by the null former forming means, the acoustic enhancement means calculates a beamformer output sound with directivity formed in the target area direction based on the filter gain calculated by the filter gain calculation means and the input spectrum for each microphone array, and the target area sound extraction means forms the target area sound based on the beamformer output sound calculated by the acoustic enhancement means. [Effects of the Invention]
[0015] According to the present invention, it is possible to provide a sound collection device, a sound collection program, and a sound collection method that suppress distortion of sound collected when performing area sound collection processing to collect sound in a target area. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a block diagram showing the functional configuration of an area sound collection device according to a first embodiment. [Figure 2] 1 is a diagram showing an example of the arrangement of microphones (microphone arrays) that supply acoustic signals (input signals) to an area sound collection device according to a first embodiment. [Figure 3] 1 is a block diagram showing an example of a hardware configuration of an area sound collection device according to a first embodiment. [Figure 4] 3A to 3C are diagrams showing examples of directions of null formers in an arbitrary microphone array according to the first embodiment. [Figure 5] FIG. 10 is a block diagram showing the functional configuration of an area sound collection device according to a second embodiment. [Figure 6] FIG. 10 is a block diagram showing the functional configuration of an area sound collection device according to a third embodiment. [Figure 7] FIG. 10 is a block diagram showing the functional configuration of an area sound collection device according to a fourth embodiment. [Figure 8] FIG. 10 is a block diagram showing the functional configuration of an area sound collection device according to a fifth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0017] (A) First embodiment A first embodiment of a sound collection device, a sound collection program, and a sound collection method according to the present invention will be described below in detail with reference to the drawings. In this embodiment, an example in which the sound collection device, the sound collection program, and the sound collection method according to the present invention are applied to an area sound collection device will be described.
[0018] (A-1) Configuration of the First Embodiment FIG. 1 is a block diagram showing the functional configuration of an area sound collection device 1 according to the first embodiment.
[0019] The area sound collection device 1 suppresses sounds emitted outside the target area and collects sounds emitted within the target area based on acoustic signals captured by two microphone arrays having two microphones.
[0020] FIG. 2 is a diagram showing an example of the arrangement of microphones (microphone arrays) that supply acoustic signals (input signals) to the area sound collection device 1. As shown in FIG.
[0021] In FIG. 2, the target area TA from which sound is to be collected is hatched (with a diagonal line pattern).
[0022] As shown in FIG. 2, the area sound collection device 1 of this embodiment performs processing to collect sound from the target area TA as a sound source based on acoustic signals (input signals) supplied from four microphones M11, M12, M21, and M22.
[0023] Here, it is assumed that a first microphone array MA1 is made up of two microphones M11 and M12, and a second microphone array MA2 is made up of microphones M21 and M22. In this embodiment, each microphone array is made up of two microphones (2 channels), and the number of microphone arrays is two (2-microphone array), but the number of microphones and microphone arrays, and the number of microphones per microphone array (number of channels), may be increased as long as there are two or more of each.
[0024] Next, the internal configuration of the area sound collection device 1 will be described with reference to FIG.
[0025] As shown in FIG. 1, the area sound collection device 1 includes a frequency analysis unit 101, a null former forming unit 102, a filter gain calculation unit 103, an audio enhancement unit 104, and a beamformer selection unit 105.
[0026] The frequency analysis unit 101 converts the acoustic signals (input signals) captured by the two microphone arrays MA1 and MA2 into frequency domain signals to obtain an input spectrum.
[0027] The null former forming unit 102 forms a plurality of null formers, each of which has a blind spot directed outside the target area, based on the input spectrum obtained by the frequency analysis unit 101, and obtains a plurality of null former spectra.
[0028] The filter gain calculation unit 103 calculates, for each microphone array, the gain of a frequency filter that passes only components in the direction of the target area, based on the multiple null former spectra calculated by the null former formation unit 102 .
[0029] The acoustic enhancement unit 104 obtains a beamformer output sound in which directivity is formed in the direction of the target area for each microphone array based on the filter gain calculated for each microphone array by the filter gain calculation unit 103 and the input spectrum.
[0030] The beamformer selection unit 105, which serves as a target area sound extraction means, obtains the area sound based on the beamformer output sound calculated for each microphone array by the sound enhancement unit 104.
[0031] FIG. 3 is a block diagram showing an example of the hardware configuration of the area sound collection device 1. As shown in FIG.
[0032] FIG. 3 shows an example of a hardware configuration when the area sound collection device 1 is configured using software (computer).
[0033] 3 has, as a hardware component, a computer 300 on which a program (including the sound collection program of the embodiment) is installed. The computer 300 may be a computer dedicated to the sound collection program, or may be configured to be shared with programs for other functions.
[0034] The computer 300 shown in FIG. 3 includes a processor 301, a primary storage unit 302, and a secondary storage unit 303. The primary storage unit 302 is a storage unit that functions as a working memory for the processor 301, and may be, for example, a high-speed memory such as a dynamic random access memory (DRAM). The secondary storage unit 303 is a storage unit that records various data such as an operating system (OS) and program data (including data of the sound collection program according to the embodiment), and may be, for example, a non-volatile memory such as a FLASH (registered trademark) memory, HDD, or SSD. In the computer 300 of this embodiment, when the processor 301 starts up, the OS and programs (including the sound collection program according to the embodiment) recorded in the secondary storage unit 303 are read, deployed on the primary storage unit 302, and executed.
[0035] Note that the specific configuration of the computer 300 is not limited to the configuration in Fig. 3, and various configurations can be applied. For example, if the primary storage unit 302 is a non-volatile memory (e.g., a flash memory), the secondary storage unit 303 may be excluded from the configuration.
[0036] (A-2) Operation of the First Embodiment Next, the operation of the area sound collection device 1 of the first embodiment having the above-mentioned configuration (sound collection method according to the embodiment) will be described.
[0037] The microphone arrays MA1 and MA2 convert the input acoustic signals from analog signals to digital signals and supply them to the frequency analysis unit 101 as input signals.
[0038] The frequency analysis unit 101 performs an arbitrary frequency analysis on a given time-domain input signal, and supplies the obtained input spectrum to the null former formation unit 102. There is no limitation on the frequency analysis method used by the frequency analysis unit 101, and for example, a fast Fourier transform or a wavelet transform may be used.
[0039] Based on the supplied input spectrum, the null former forming unit 102 calculates, for each microphone array, multiple null formers that form blind spots in predetermined directions other than the target area, and supplies the obtained null former spectra to the filter gain calculation unit 103. The operation of the null former forming unit 102 will be described in detail. Since the target area is not a point but has a range, the target area direction also has a range.
[0040] Next, the positional relationship between each microphone array and the target area TA will be described.
[0041] FIG. 4 is a diagram showing an example of the direction of null formers in an arbitrary microphone array MAa.
[0042] Hereinafter, the "a" in "MAa" etc. indicates the microphone array number (identifier; ID). Here, the microphone array MA1 is numbered 1, and the microphone array MA2 is numbered 2. In other words, the microphone array MAa shown in FIG. 4 indicates either the microphone array MA1 or MA2 (any microphone array). FIG. 4 illustrates that the microphone array MAa includes microphones ML and MR. For example, if the microphone array MAa is the microphone array MA1, the microphones ML and MR are microphones M11 and M12, respectively, and if the microphone array MAa is the microphone array MA2, the microphones ML and MR are microphones M21 and M22, respectively. In the image diagram of FIG. 4, the shape of the target area TA is illustrated as a circle for convenience of illustration, but in this embodiment, the shape of the target area TA is actually the shape shown in FIG. 2.
[0043] Here, the front direction (0 degree direction) of the microphone array MAa is defined as the direction that is perpendicular to the line connecting the positions (center positions) of the two microphones ML and MR and that forms the smallest angle with the direction from the center point of the two microphones ML and MR toward the center point (which can also be the center of gravity) of the target area TA.
[0044] 2, since the target area TA is not a point but has a range, the apparent area direction from each microphone array also has a range. Therefore, here, the target area direction (range of the target area direction) for microphone array MA1 is set to θ1L to θ1R (rad) (-π / 2<θ1L<0<θ1R<π / 2), and the target area direction (range of the target area direction) for microphone array MA2 is set to θ2L to θ2R (rad) (-π / 2<θ2L<0<θ2R<π / 2).
[0045] As described above, the predetermined blind spot direction applied to each null former of the null former forming unit 102 must be outside the target area direction (a direction that does not include the range of the target area direction). Therefore, the n-th blind spot direction φan(rad) (n=1, ..., N) of the microphone array MAa (a=1, 2) must satisfy -π-θaL<φan<θaL or θaR<φan<π-θaL.
[0046] Here, since the microphone array MAa consists of two microphones, if a blind spot is formed in φan, a blind spot will also be formed in π-φan (=-π-φan). Furthermore, since it is difficult to know in advance from which direction noise will arrive, the predetermined blind spot direction may be determined in advance by design (applying a pre-designed fixed value).
[0047] For the above reasons, it is preferable to determine the same number of predetermined blind spot directions to be applied to each null former of the null former forming unit 102 within the two ranges of -π / 2≦φan<θaL and θaR<φan≦π / 2. The number of predetermined blind spot directions set for each microphone array in the null former forming unit 102 (the number of null formers set for each microphone array MAa in the null former forming unit 102) and the combination of blind spot directions are not limited. For example, if the number of predetermined blind spot directions N is 6 (N=6), the following can be selected, and this selection method is preferable. Note that N is not limited to 6 and may be increased or decreased depending on the target area TA, the expected number of interfering sound sources (the number of interfering sound sources and their arrival directions), etc.
[0048] Next, an example of specific processing of each null former in the null former generating unit 102 will be described.
[0049] The input spectrum of the microphone array MAa is X aL (ω), X aR (ω), the null former generating unit 102 generates a plurality of null former spectra Y an Calculate (ω), where ω (omega) is the angular frequency, i is the imaginary unit, d is the microphone spacing, and c is the speed of sound.
number
[0050] The filter gain calculation unit 103 calculates a filter gain for each microphone array and for each frequency based on the supplied null former spectra and the input spectrum, and supplies the obtained filter gain to the acoustic enhancement unit 104. Specifically, the filter gain calculation unit 103 calculates a filter gain G that passes only the component in the direction of the target area. a (ω) is calculated based on the following equation (2).
number
[0051] Filter Gain G a (ω) is calculated by multiplying the ratio of the null former spectrum to the input spectrum of the microphone array by the number of null formers.
[0052] Filter Gain G a Regarding the calculation method of (ω), in equation (2), the denominator is the input spectrum X aL (ω), but it is sufficient if it is a value equivalent to the microphone input spectrum, and X aR (ω) and X aL (ω) and X aR The geometric mean of (ω), X aL (ω) and X aR The acoustic enhancement unit 104 calculates a beamformer spectrum in the target area direction for each microphone array and for each frequency based on the supplied filter gain and input spectrum, and supplies the obtained beamformer spectrum (beamformer output sound) to the beamformer selection unit 105. Specifically, the acoustic enhancement unit 104 calculates the beamformer spectrum B in the target area direction. a (ω) is calculated based on the following equation (3).
number
[0053] The beamformer selection unit 105 selects, for each frequency, the beamformer spectrum with the smallest amplitude based on the two supplied beamformer spectra, and outputs the obtained area sound spectrum (area sound). Specifically, the beamformer selection unit 105 calculates the area sound spectrum Z(ω) based on the following equations (4) and (5).
number
[0054] (A-3) Effects of the First Embodiment According to the first embodiment, the following effects can be achieved.
[0055] In the area sound collection device 1 of the first embodiment, multiple null formers are calculated in parallel, and a filter gain that extracts only the component in the direction of the target area is calculated based on the null former spectrum and the input spectrum, thereby suppressing sounds arriving from directions other than the direction of the target area and forming a beamformer that collects sound only in the direction of the target area. As a result, in the area sound collection device 1, by using this beamformer output, it is possible to collect only the area sound without causing distortion such as occurs in spectral subtraction.
[0056] (B) Second embodiment A second embodiment of the sound collection device, sound collection program, and sound collection method according to the present invention will be described below in detail with reference to the drawings. In this embodiment, an example in which the sound collection device, sound collection program, and sound collection method according to the present invention are applied to an area sound collection device will be described.
[0057] (B-1) Configuration of the second embodiment FIG. 5 is a block diagram showing the functional configuration of an area sound collection device 1A according to the second embodiment, and the same or corresponding parts as those in FIG. 1 above are denoted by the same or corresponding reference numerals.
[0058] The following describes the area sound collection device 1A of the second embodiment, focusing on the differences from the first embodiment.
[0059] The area sound collecting device 1A of the second embodiment is different from the first embodiment in that a blind spot direction updating unit 106 is inserted between the frequency analysis unit 101 and the null former forming unit 102.
[0060] In the area sound collection device 1 of the first embodiment, when forming a blind spot using the null former forming unit 102, it is difficult to know in advance from which direction the noise will come, so the direction of the predetermined blind spot is determined in advance.
[0061] In contrast to this, the area sound collection device 1A of the second embodiment adds a blind spot direction updating unit 106 that updates the blind spot directions of the null former forming unit 102 in real time based on the value of the null former spectrum (hereinafter, this update processing will be referred to as "blind spot direction updating processing"), making it possible to adaptively and accurately point the blind spot directions toward the direction of interfering sounds arriving from outside the target area direction. As a result, the area sound collection device 1A of the second embodiment can more accurately suppress interfering sounds without being affected by the position.
[0062] (B-2) Operation of the Second Embodiment Next, the operation of the area sound collection device 1A of the second embodiment having the above-mentioned configuration (sound collection method according to the embodiment) will be described.
[0063] The operation of the area sound collecting device 1A of the second embodiment will be described below, focusing on the operation of the blind spot direction updating unit 106, which is a difference from the first embodiment.
[0064] The blind spot direction update unit 106 determines the optimum blind spot direction for suppressing the interfering sound based on the input spectrum supplied from the frequency analysis unit 101 and the current blind spot direction φan(rad) stored in the blind spot direction update unit 106, and updates the blind spot direction φan(rad) to that direction.
[0065] Any known method can be used to update the blind spot direction φan(rad) in the blind spot direction update unit 106, but the update method described below is preferred.
[0066] First, the null direction update unit 106 calculates the null former spectrum Y for each current null direction φan (rad) stored in the directivity direction formation unit for each microphone array, for the three directions φan, φan+δφ, and φan-δφ, using the following equations (6) and (7). ank (ω) is calculated, where δφ is a preset shift value in the blind spot direction.
[0067] Next, the blind spot direction update unit 106 calculates the null former spectrum Y ank For (ω), the sum P ank Take.
number
[0068] Then, the blind spot direction update unit 106 calculates the sum P ank The blind spot direction for which the value of is smallest is updated as the new blind spot direction. For example, if the sum value of the null former spectrum for φan+δφ is smallest, the blind spot direction updating unit 106 updates φan+δφ as the new blind spot direction. Then, the blind spot direction updating unit 106 supplies the updated blind spot direction value to the null former forming unit 102.
[0069] It is desirable that the shift value δφ be the angle between adjacent null formers (for example, an intermediate angle between adjacent null formers). In the above example, for each null former, evaluation values (sum P ank ) and calculate the most highly evaluated update angle candidate (total P ankHowever, the number of candidate update angles is not limited and may be more than three. For example, the candidate update angles may be five, namely, φan, φan+δφ / 2, φan+δφ, φan-δφ / 2, and φan-δφ. The blind spot direction update unit 106 may perform the blind spot direction update process for some null formers, or may update the blind spot directions for all null formers. In other words, the number and combination of null formers for which the blind spot direction update process is performed in the blind spot direction update unit 106 is not limited. Note that the shift value δφ is preferably a small angle (e.g., π / 180) so that it can be aligned with the direction of the interfering sound, and it is desirable to set the angle between adjacent null formers (e.g., an angle approximately midway between adjacent null formers) as the upper and lower limit values for the blind spot direction.
[0070] (B-3) Effects of the Second Embodiment In addition to the effects of the first embodiment, the second embodiment can achieve the following effects.
[0071] In the area sound collection device 1A of the second embodiment, the blind spot direction of the null former forming section 102 is updated in real time based on the value of the null former spectrum, so that the blind spot direction can be directed toward the direction of the interfering sound. This makes it possible for the area sound collection device 1A to further suppress the interfering sound regardless of the position of the interfering sound.
[0072] (C) Third embodiment A third embodiment of the sound collection device, sound collection program, and sound collection method according to the present invention will be described below in detail with reference to the drawings. In this embodiment, an example in which the sound collection device, sound collection program, and sound collection method according to the present invention are applied to an area sound collection device will be described.
[0073] (C-1) Configuration of the third embodiment FIG. 6 is a block diagram showing the functional configuration of an area sound collection device 1B according to the third embodiment, and the same or corresponding parts as those in FIG. 1 above are denoted by the same or corresponding reference numerals.
[0074] The following describes the differences between the area sound collection device 1B of the third embodiment and the first embodiment.
[0075] The area sound collecting device 1B of the third embodiment is different from the first embodiment in that a null former selecting section 107 is inserted between the null former forming section 102 and the filter gain calculating section 103.
[0076] In the first embodiment, when the filter gain calculation unit 103 calculates the filter gain, the filter gain is calculated using all the null formers supplied from the null former formation unit 102.
[0077] When there are multiple interfering sounds, increasing the number of null formers makes it possible to suppress the interfering sounds more effectively. However, when there are fewer interfering sounds than the number of null formers, null formers with blind spots in the direction of the interfering sounds contribute significantly to suppressing the interfering sounds, but other null formers do not contribute to suppressing the interfering sounds and are therefore unnecessary. In fact, using unnecessary null formers to calculate the filter gain may result in excessive suppression of the target sound, resulting in degradation of the target sound.
[0078] Therefore, the area sound collection device 1B of the third embodiment adds a null former selection unit 107 that selects only a specific null former from the multiple null formers supplied from the null former formation unit 102 and supplies the selected null former to the filter gain calculation unit 103. This makes it possible for the area sound collection device 1B of the third embodiment to exclude null formers that do not contribute to suppressing interfering sounds and to form filter gains with fewer null formers, thereby reducing the influence of audio distortion that increases when the number of null formers to be multiplied is increased.
[0079] (C-2) Operation of the Third Embodiment Next, the operation of the area sound collection device 1B of the third embodiment having the above-mentioned configuration (sound collection method according to the embodiment) will be described.
[0080] The operation of the area sound collection device 1B of the third embodiment will be described below, focusing on the operation of the null former selection unit 107, which is a difference from the first embodiment.
[0081] The null former selection unit 107 selects only a specific null former based on the values of the null former spectra of the multiple null formers supplied from the null former generation unit, and supplies the null former spectrum of the selected null former to the filter gain calculation unit. The null former selection unit 107 can use any known method to select a specific null former, but the update method described below is preferred.
[0082] First, the null former selection unit 107 selects a plurality of null former spectra Y an For (ω), the sum P an Take.
number
[0083] Next, the null former selection unit 107 selects the sum P an The null former selection unit 107 extracts only a predetermined specific number of null formers in ascending order of the value of P an For example, the null former selection unit 107 may dynamically determine the number of null formers to be extracted depending on P an Compare with a predetermined threshold, P an is smaller than a threshold value, the null former is supplied to the filter gain calculation unit 103.
[0084] (C-3) Effects of the Third Embodiment According to the third embodiment, the following effects can be achieved compared to the first embodiment.
[0085] In the area sound collection device 1B of the third embodiment, the null former selection unit 107 can exclude null formers that do not contribute to the suppression of interfering sounds from the calculation of the filter gain, making it possible to further reduce distortion of the target sound.
[0086] (D) Fourth embodiment FIG. 7 is a block diagram showing the functional configuration of an area sound collection device 1C according to the fourth embodiment, and the same or corresponding parts as those in FIG. 1 above are denoted by the same or corresponding reference numerals.
[0087] The following describes the differences between the area sound collection device 1C of the fourth embodiment and the first embodiment.
[0088] The area sound collection device 1C of the fourth embodiment differs from the first embodiment in that an amplitude correction unit 108 is inserted after the frequency analysis unit 101 (between the frequency analysis unit 101 and the null former formation unit 102).
[0089] The amplitude correction unit 108 corrects the amplitude ratio between the two microphones that make up each microphone array. The amplitude correction unit 108 calculates the amplitude ratio between the two microphones for each microphone array using any known method, and corrects the amplitude ratio so that it approaches 1. This makes it possible for the area sound collection device 1C to further suppress interfering sounds when the null former forming unit 102 calculates the null former spectrum.
[0090] It should be noted that the area sound collecting devices 1, 1A, and 1B of the first to third embodiments may also be configured to have an amplitude correcting unit 108 added thereto.
[0091] (E) Fifth embodiment FIG. 8 is a block diagram showing the functional configuration of an area sound collection device 1D according to the fifth embodiment, and the same or corresponding parts as those in FIG. 1 above are denoted by the same or corresponding reference numerals.
[0092] The following describes the differences between the area sound collection device 1D of the fifth embodiment and the first embodiment.
[0093] The area sound collection device 1D of the fifth embodiment differs from the first embodiment in that a noise suppression unit 109 is inserted between the frequency analysis unit 101 and the null former formation unit 102 (after the frequency analysis unit 101).
[0094] The noise suppression unit 109 removes noise from the input spectrum using a noise suppression method such as a Wiener filter, for example. This makes it possible for the area sound collection device 1D to remove diffuse noise that cannot be removed by beamforming, and allows it to be used in a wider scene.
[0095] It should be noted that the area sound collecting devices 1, 1A, 1B, and 1C of the first to fourth embodiments may also be configured to have a noise suppressing section 109 added thereto.
[0096] (F) Other embodiments The present invention is not limited to the above-described embodiments, and may include modified embodiments such as those exemplified below.
[0097] (F-1) In the processing of the filter gain calculation unit 103 in each of the above embodiments, the filter gain G a (ω) is calculated by multiplying the ratio of the null former spectrum to the input spectrum of the microphone array by the number of null formers, but it may also be multiplied by the root of the number of null formers N of the ratio of the null former spectrum to the input spectrum of the microphone array, as in the following equation (11):
number
[0098] (F-2) In the above embodiments, the area sound collection devices 1 to 1D are configured such that digital signals are supplied from the microphone arrays MA1 and MA2. However, the area sound collection devices 1 to 1D may be configured such that analog signals are supplied from the microphone arrays MA1 and MA2, and the analog signals (acoustic signals captured by each microphone in each microphone array) are converted into digital signals on the area sound collection device 1 to 1D side. [Explanation of symbols]
[0099] 1, 1A, 1B, 1C, 1D...area sound collection device, 101...frequency analysis unit, 102...null former formation unit, 103...filter gain calculation unit, 104...acoustic enhancement unit, 105...beam former selection unit, 106...blind spot direction update unit, 107...null former selection unit, 108...amplitude correction unit, 109...noise suppression unit, M11, M12, M21, M22, MR, ML...microphone, MA1, MA2, MAa...microphone array, TA...target area
Claims
1. a null former generating means for generating a plurality of null formers each having a blind spot outside a target area based on an input spectrum of an acoustic signal supplied from the plurality of microphone arrays, and calculating a null former spectrum; a filter gain calculation means for calculating a filter gain of a frequency filter that passes a component in a direction of a target area for each of the microphone arrays based on the null former spectrum calculated by the null former generation means; an acoustic enhancement unit that calculates, for each of the microphone arrays, the filter gain calculated by the filter gain calculation unit and a beamformer output sound having directivity formed in the direction of the target area based on the input spectrum; a target area sound extraction means for forming a target area sound based on the beamformer output sound calculated by the acoustic enhancement means; A sound collection device comprising:
2. The sound collection device according to claim 1, further comprising a blind spot direction update means for evaluating the blind spot directions of some or all of the null formers for each of a plurality of update angle candidates selected from angles within a predetermined range including the blind spot directions of the null formers, and updating the blind spot directions of the null formers to the update angle candidate with the highest evaluation.
3. The blind spot direction update means calculates a sum value in the frequency direction of the null former spectrum for each of the update angle candidates, and selects the update angle candidate with the smallest sum value as the one with the highest evaluation.
4. further comprising a null former selection means for selecting some of the null formers based on the null former spectrum of each of the null formers; The filter gain calculation means calculates the gain of the frequency filter based only on the null former spectrum of the null former selected by the null former selection means.
2. The sound pickup device according to claim 1.
5. The microphone array further includes an amplitude correction unit that corrects the amplitude ratio of the acoustic signals of each microphone so that the amplitude ratio approaches 1, The null former generating means calculates each of the null formers based on the acoustic signal corrected by the amplitude correcting means.
2. The sound pickup device according to claim 1.
6. Further, for each of the microphone arrays, a noise suppression means is provided for suppressing noise in the acoustic signal of each microphone; The null former generating means calculates the null former based on the acoustic signal from which noise has been suppressed by the noise suppressing means.
2. The sound pickup device according to claim 1.
7. Computer, a null former generating means for generating a plurality of null formers each having a blind spot outside a target area based on an input spectrum of an acoustic signal supplied from the plurality of microphone arrays, and calculating a null former spectrum; a filter gain calculation means for calculating a filter gain of a frequency filter that passes a component in a direction of a target area for each of the microphone arrays based on the null former spectrum calculated by the null former generation means; an acoustic enhancement unit that calculates, for each of the microphone arrays, the filter gain calculated by the filter gain calculation unit and a beamformer output sound having directivity formed in the direction of the target area based on the input spectrum; a target area sound extraction means for forming a target area sound based on the beamformer output sound calculated by the acoustic enhancement means; A sound collection program characterized by functioning as follows.
8. In the sound collection method performed by the sound collection device, The sound collection device includes a null former forming means, a filter gain calculating means, an audio enhancement means, and a target area sound extracting means, the null former generating means generates a plurality of null formers each having a blind spot outside a target area based on an input spectrum of an acoustic signal supplied from a plurality of microphone arrays, and calculates a null former spectrum; the filter gain calculation means calculates, for each of the microphone arrays, a filter gain of a frequency filter that passes a component in a direction of a target area, based on the null former spectrum calculated by the null former formation means; the acoustic enhancement means calculates, for each of the microphone arrays, the filter gain calculated by the filter gain calculation means and a beamformer output sound in which directivity is formed in the direction of the target area based on the input spectrum; The target area sound extraction means forms a target area sound based on the beamformer output sound calculated by the sound enhancement means. A sound collection method characterized by:
Citation Information
Patent Citations
Signal processor
JP1998207490A
Earhole attachment-type sound pickup device, signal processing device, and sound pickup method
JP2013121106A
Sound source separating device, sound source separating program, sound collecting device, and sound collecting program
JP2015050558A
Sound pickup device, program, and method
JP2018170617A
Sound collection device, program, and method
JP2019176328A