Sound collecting device, sound collecting program, and sound collecting method
The sound collection device addresses the challenge of large-scale and distorted area sound collection by using null-formers and filter gains to suppress interference, enabling efficient sound collection from a target area with fewer microphones.
Patent Information
- Application Number
- JP2024017413
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-02-07
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2044-02-07
AI Technical Summary
Conventional sound collection devices require a large number of microphones to achieve sharp directivity for area sound collection, leading to large-scale devices, or when using fewer microphones, they suffer from distortion and musical noise.
The sound collection device employs a null-former forming unit to create null-formers directed to dead angles outside the target area, a filter gain calculation unit to pass components in the target area direction, and an acoustic enhancement unit to form beamformer output sound, thereby suppressing distortion and collecting sound only from the target area.
The device effectively collects sound from a target area without distortion, using a reduced number of microphones, by forming null-formers and calculating filter gains to suppress interfering sounds.
Smart Images

Figure 0007715222000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sound collection device, a program, and a method, and can be applied to, for example, a sound collection device that emphasizes only the sound emitted from a specific area and suppresses the sound from other areas.
Background Art
[0002] When using a speech recognition system in a noisy environment, the ambient noise that is mixed in simultaneously with the necessary target sound is a troublesome existence that causes a decrease in the speech recognition rate of the recorded speech.
[0003] Conventionally, in an environment where such multiple sound sources exist, as a technique for obtaining a necessary target sound by avoiding the mixing of unwanted sounds by collecting only the sound in a specific direction, there is beamforming using a microphone array. Beamforming is a technique for forming directivity by utilizing the time difference of signals reaching each microphone (see Non-Patent Document 1). A microphone array that performs beamforming is also called a beamformer.
[0004] On the other hand, since a beamformer collects all the sounds in a specific direction, when it is desired to collect only a specific area (also called a target area), sounds outside the target area that are in the same direction as seen from the beamformer are also collected. Therefore, Patent Document 1 proposes a method (hereinafter referred to as "area sound collection") of using a plurality of beamformers, directing the directivity from different directions to the target area, and intersecting the directivities in the target area to collect the target sound.
[0005] Area sound collection requires the beamformer to have a sharp directivity in order to form the target area. It is desirable that the beam width of the beamformer is generally 60 degrees or less.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Non-Patent Document
[0007]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] In any of the addition-type and subtraction-type beamformers, a large number of microphones are required to form a sharp directivity. For example, in the minimum variance distortionless response (MVDR) method, the number of microphones equal to the number of interfering sound sources to be suppressed + 1 is required. Although area sound collection with such a large number of microphones can be experimentally realized in a laboratory, it is not suitable for commercial products in terms of component cost and installation space.
[0009] On the other hand, in the technique described in Patent Document 1, by using spectral subtraction (subtracting the amplitude spectra from each other, and replacing with zero or a small number when the subtraction result becomes negative) in the frequency domain, a sharp directivity is realized with a small number of microphones. A nullformer that forms a dead angle in the direction of the target area is created, and by spectrally subtracting the nullformer amplitude spectrum from the microphone input amplitude spectrum, a sharp directivity is formed in the direction of the target area. By using this method, a sharp directivity can be formed with only 2 microphones, but there is a risk of generating musical noise and distortion such as missing area sound components.
[0010] Therefore, in the conventional technology, there are problems that the device becomes large-scale because a large number of microphones are required to realize area sound collection, or the area sound extracted by using spectral subtraction with a small number of microphones is distorted.
[0011] In view of the above problems, there is a need for a sound collection device, a sound collection program, and a sound collection method that suppress distortion of sound collected during area sound collection processing for collecting sound in a target area.
Means for Solving the Problem
[0012] The sound collection device of the first aspect of the present invention includes, for each microphone array, null-former forming means for forming a plurality of null-formers directed to dead angles outside the target area based on the input spectrum of the acoustic signals supplied from a plurality of microphone arrays and calculating a null-former spectrum; filter gain calculating means for calculating the filter gain of a frequency filter that passes components in the target area direction for each microphone array based on the null-former spectrum calculated by the null-former forming means; acoustic enhancement means for calculating a beamformer output sound having directivity in the target area direction based on the filter gain calculated by the filter gain calculating means and the input spectrum for each microphone array; and target area sound extraction means for forming a target area sound based on the beamformer output sound calculated by the acoustic enhancement means.
[0013] The sound collection program of the second aspect of the present invention causes a computer to function as null-former forming means for forming a plurality of null-formers directed to dead angles outside the target area based on the input spectrum of the acoustic signals supplied from a plurality of microphone arrays and calculating a null-former spectrum for each microphone array; filter gain calculating means for calculating the filter gain of a frequency filter that passes components in the target area direction for each microphone array based on the null-former spectrum calculated by the null-former forming means; acoustic enhancement means for calculating a beamformer output sound having directivity in the target area direction based on the filter gain calculated by the filter gain calculating means and the input spectrum for each microphone array; and target area sound extraction means for forming a target area sound based on the beamformer output sound calculated by the acoustic enhancement means.
[0014] The third sound collection method of the present invention is a sound collection method performed by a sound collection device. The sound collection device includes a null-former forming means, a filter gain calculation means, an acoustic enhancement means, and a target area sound extraction means. The null-former forming means forms a plurality of null-formers directed to blind spots outside the target area for each microphone array based on the input spectrum of the acoustic signals supplied from a plurality of microphone arrays, and calculates a null-former spectrum. The filter gain calculation means calculates the filter gain of a frequency filter that passes components in the direction of the target area for each microphone array based on the null-former spectrum calculated by the null-former forming means. The acoustic enhancement means calculates a beamformer output sound with directivity formed in the direction of the target area for each microphone array based on the filter gain calculated by the filter gain calculation means and the input spectrum. The target area sound extraction means forms a target area sound based on the beamformer output sound calculated by the acoustic enhancement means.
Effect of the Invention
[0015] According to the present invention, it is possible to provide a sound collection device, a sound collection program, and a sound collection method that suppress distortion of the sound collected when performing area sound collection processing for collecting the sound of a target area.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Mode for Carrying Out the Invention
[0017] (A) First Embodiment Hereinafter, a first embodiment of the sound collection device, the sound collection program, and the sound collection method according to the present invention will be described in detail with reference to the drawings. In this embodiment, an example in which the sound collection device, the sound collection program, and the sound collection method of the present invention are applied to an area sound collection device will be described.
[0018] (A-1) Configuration of the First Embodiment FIG. 1 is a block diagram showing the functional configuration of the area sound collection device 1 according to the first embodiment.
[0019] The area sound collection device 1 suppresses the sound emitted outside the target area and collects the sound emitted inside the target area based on the acoustic signals captured by two microphone arrays each having two microphones.
[0020] FIG. 2 is a diagram showing an example of the arrangement configuration of the microphones (microphone arrays) that supply acoustic signals (input signals) to the area sound collection device 1.
[0021] In FIG. 2, the target area TA to be sound-collected is hatched (with a diagonal pattern).
[0022] As shown in Fig. 2, the area sound collection device 1 of this embodiment performs a process of area sound collection of the sound having the target area TA as the sound source based on the acoustic signals (input signals) supplied from the four microphones M11, M12, M21, and M22.
[0023] Here, it is assumed that the two microphones M11 and M12 constitute the first microphone array MA1, and the microphones M21 and M22 constitute the second microphone array MA2. In this embodiment, the number of microphones constituting each microphone array is two (2ch), and the number of microphone arrays is two (two microphone arrays). However, as long as each is two or more, the number of microphones and microphone arrays and the number of microphones (ch number) for each microphone array may be increased.
[0024] Next, the internal configuration of the area sound collection device 1 will be described with reference to Fig. 1.
[0025] As shown in Fig. 1, the area sound collection device 1 includes a frequency analysis unit 101, a nullformer formation unit 102, a filter gain calculation unit 103, an acoustic enhancement unit 104, and a beamformer selection unit 105.
[0026] The frequency analysis unit 101 converts the acoustic signals (input signals) captured by the two microphone arrays MA1 and MA2 into frequency domain signals to obtain input spectra.
[0027] The nullformer formation unit 102 forms a plurality of nullformers each having a dead angle directed outside the target area based on the input spectra obtained by the frequency analysis unit 101, and obtains a plurality of nullformer spectra.
[0028] The filter gain calculation unit 103 calculates the gain of a frequency filter that allows only the components in the target area direction to pass through for each microphone array based on the plurality of nullformer spectra calculated by the nullformer formation unit 102.
[0029] In the acoustic enhancement unit 104, based on the filter gain calculated for each microphone array by the filter gain calculation unit 103 and the input spectrum, a beamformer output sound with directivity formed in the target area direction is obtained for each microphone array.
[0030] The beamformer selection unit 105 as the target area sound extraction means obtains the area sound based on the beamformer output sound calculated for each microphone array by the acoustic enhancement unit 104.
[0031] FIG. 3 is a block diagram showing an example of the hardware configuration of the area sound collection device 1.
[0032] FIG. 3 shows an example of the hardware configuration when the area sound collection device 1 is configured using software (computer).
[0033] The area sound collection device 1 shown in FIG. 3 has a computer 300 in which a program (including the sound collection program of the embodiment) is installed as a hardware component. Also, the computer 300 may be a computer dedicated to the sound collection program or may be configured to be shared with programs of other functions.
[0034] The computer 300 shown in FIG. 3 has a processor 301, a primary storage unit 302, and a secondary storage unit 303. The primary storage unit 302 is a storage means that functions as a working memory (work memory) of the processor 301. For example, a high-speed operating memory such as a DRAM (Dynamic Random Access Memory) can be applied. The secondary storage unit 303 is a storage means for recording various data such as an OS (Operating System) and program data (including data of the sound collection program according to the embodiment). For example, a non-volatile memory such as a FLASH (registered trademark) memory, an HDD, or an SSD can be applied. In the computer 300 of this embodiment, when the processor 301 is activated, the OS and programs (including the sound collection program according to the embodiment) recorded in the secondary storage unit 303 are read, expanded, and executed on the primary storage unit 302.
[0035] Note that the specific configuration of the computer 300 is not limited to the configuration of FIG. 3, and various configurations can be applied. For example, if the primary storage unit 302 is a non-volatile memory (for example, a FLASH memory or the like), the configuration excluding the secondary storage unit 303 may be used.
[0036] (A-2) Operation of the First Embodiment Next, the operation of the area sound collection device 1 of the first embodiment (the sound collection method according to the embodiment) having the above configuration will be described.
[0037] The microphone arrays MA1 and MA2 convert the input acoustic signal from an analog signal into a digital signal, and supply it as an input signal to the frequency analysis unit 101.
[0038] The frequency analysis unit 101 performs arbitrary frequency analysis on the given input signal in the time domain, and supplies the obtained input spectrum to the null-former formation unit 102. In the frequency analysis unit 101, there is no limitation on the frequency analysis method. For example, a fast Fourier transform or a wavelet transform may be used.
[0039] Based on the supplied input spectrum, the null-former forming unit 102 calculates, for each microphone array, a plurality of null-formers that form dead zones in a predetermined direction other than the target area, and supplies the obtained null-former spectrum to the filter gain calculation unit 103. The operation of the null-former forming unit 102 will be described in detail. Since the target area has a range rather than a point, the target area direction also has a range.
[0040] Next, the positional relationship between each microphone array and the target area TA will be described.
[0041] FIG. 4 is a diagram showing an example of the direction of the null-former in an arbitrary microphone array MAa.
[0042] Hereinafter, "a" in "MAa" etc. shall indicate the number (identifier; ID) of the microphone array. Here, let the number of the microphone array MA1 be 1 and the number of the microphone array MA2 be 2. That is, the microphone array MAa shown in FIG. 4 shall indicate either the microphone array MA1 or MA2 (an arbitrary microphone array). In FIG. 4, it is illustrated that the microphone array MAa includes microphones ML and MR. For example, when the microphone array MAa is the microphone array MA1, the microphones ML and MR are the microphones M11 and M12 respectively, and when the microphone array MAa is the microphone array MA2, the microphones ML and MR are the microphones M21 and M22 respectively. In the image diagram of FIG. 4, for the sake of illustration, the shape of the target area TA is shown as circular, but in this embodiment, the actual shape of the target area TA is as shown in FIG. 2.
[0043] Here, the front direction (0-degree direction) of the microphone array MAa is defined as the direction that is orthogonal to the straight line connecting the positions (center positions) of the two microphones ML and MR and has the smallest angle with the direction from the center point of the two microphones ML and MR to the center point (which may be the centroid) of the target area TA.
[0044] Incidentally, as shown in FIG. 2, since the target area TA has a range rather than a point, the direction of the target area as seen from each microphone array also has a range. Therefore, here, the direction of the target area (range of the target area direction) in the microphone array MA1 is set to θ1L to θ1R (rad) (-π / 2 < θ1L < 0 < θ1R < π / 2), and the direction of the target area (range of the target area direction) in the microphone array MA2 is set to θ2L to θ2R (rad) (-π / 2 < θ2L < 0 < θ2R < π / 2).
[0045] As described above, the predetermined dead angle direction applied to each null-form in the null-form forming unit 102 must be outside the target area direction (a direction not including the range of the target area direction). Therefore, the n-th dead angle direction φan (rad) (n = 1,..., N) of the microphone array MAa (a = 1, 2) needs to satisfy -π - θaL < φan < θaL, or θaR < φan < π - θaL.
[0046] Here, since there are two microphones constituting the microphone array MAa, when a dead angle is formed at φan, a dead angle is also formed at π - φan (=-π - φan). Also, since it is difficult to know in advance from which direction noise arrives, the predetermined dead angle direction may be determined in advance by design (applying a fixed value designed in advance).
[0047] From the above, it is preferable to determine the same number of predetermined dead angle directions to be applied to each null form of the null form forming unit 102 from within two ranges of -π / 2 ≦ φan < θaL and θaR < φan ≦ π / 2, respectively. In the null form forming unit 102, the number of predetermined dead angle directions set for each microphone array (the number of null forms set for each microphone array MAa in the null form forming unit 102) and the combination of dead angle directions are not limited. For example, when the number N of predetermined dead angle directions is 6 (when N = 6), φa1 = -π / 2, φa2 = -π / 3, φa3 = -π / 6, φa4 = π / 6, φa5 = π / 3, φa6 = π / 2 can be selected, and this selection is preferable. Note that N is not limited to 6, and it may be designed to increase or decrease according to the target area TA, the assumption of interfering sound sources (the number and arrival directions of interfering sound sources), etc.
[0048] Next, an example of the specific processing of each null form in the null form forming unit 102 will be described.
[0049] Let the input spectrum of the microphone array MAa be X aL (ω), X aR (ω). When this is the case, the null form forming unit 102 calculates a plurality of null form spectra Y an (ω) using equation (1). Here, ω (omega) is the angular frequency, i is the imaginary unit, d is the microphone interval, and c is the speed of sound.
Equation
[0050] Based on the supplied plurality of null form spectra and the input spectrum, the filter gain calculation unit 103 calculates the filter gain for each microphone array and for each frequency, and supplies the obtained filter gain to the acoustic enhancement unit 104. Specifically, the filter gain calculation unit 103 calculates the filter gain G a (ω) that allows only the component in the target area direction to pass through based on the following equation (2). [Number]
[0051] Filter gain G a (ω) is calculated by multiplying the ratio of the null form spectrum to the input spectrum of the microphone array by the number of null forms.
[0052] Filter gain G a Regarding the calculation method of (ω), in Equation (2), the denominator is the input spectrum X aL (ω), but it can be any value corresponding to the input spectrum of the microphone, such as X aR (ω), or the geometric mean of X aL (ω) and X aR (ω), or the arithmetic mean of X aL (ω) and X aR (ω). In the acoustic enhancement unit 104, based on the supplied filter gain and the input spectrum, for each microphone array and for each frequency, the beamformer spectrum in the target area direction is calculated, and the obtained beamformer spectrum (beamformer output sound) is supplied to the beamformer selection unit 105. Specifically, in the acoustic enhancement unit 104, the beamformer spectrum B a (ω) in the target area direction is calculated based on the following Equation (3). [Number]
[0053] Based on the two supplied beamformer spectra, the beamformer selection unit 105 selects, for each frequency, the one with the minimum amplitude of the beamformer spectrum and outputs the obtained area sound spectrum (area sound). Specifically, the beamformer selection unit 105 calculates the area sound spectrum Z(ω) based on the following Equations (4) and (5). [Number]
[0054] (A-3) Effects of the First Embodiment According to the first embodiment, the following effects can be achieved.
[0055] In the area sound collection device 1 of the first embodiment, by calculating a plurality of null formers in parallel and calculating a filter gain that extracts only the components in the target area direction based on the null former spectrum and the input spectrum, the sound coming from directions other than the target area direction is suppressed, and a beamformer that collects sound only in the target area direction can be formed. As a result, the area sound collection device 1 can collect only the area sound without generating distortion such as that caused by spectral subtraction by using this beamformer output.
[0056] (B) Second Embodiment Hereinafter, a second embodiment of the sound collection device, sound collection program, and sound collection method according to the present invention will be described in detail with reference to the drawings. In this embodiment, an example in which the sound collection device, sound collection program, and sound collection method of the present invention are applied to an area sound collection device will be described.
[0057] (B-1) Configuration of the Second Embodiment FIG. 5 is a block diagram showing the functional configuration of the area sound collection device 1A according to the second embodiment, and the same or corresponding parts as those in FIG. 1 described above are denoted by the same or corresponding reference numerals.
[0058] Hereinafter, the differences between the area sound collection device 1A of the second embodiment and the first embodiment will be described.
[0059] The area sound collection device 1A of the second embodiment is different from the first embodiment in that a dead angle direction update unit 106 is inserted between the frequency analysis unit 101 and the null former formation unit 102.
[0060] In the area sound collection device 1 of the first embodiment, when forming a dead angle in the null former formation unit 102, it is difficult to know in advance from which direction the noise comes, so a predetermined dead angle direction is determined by a fixed decision.
[0061] On the other hand, in the area sound collection device 1A of the second embodiment, a dead angle direction update unit 106 that performs real-time update processing (hereinafter, this update processing is referred to as "dead angle direction update processing") on the dead angle direction of the null form formation unit 102 based on the value of the null form spectrum is added. As a result, it becomes possible to adaptively and accurately direct the dead angle direction toward the direction of the interfering sound arriving from outside the target area direction. Thereby, the area sound collection device 1A of the second embodiment can suppress the interfering sound more accurately without being affected by the position.
[0062] (B-2) Operation of the Second Embodiment Next, the operation of the area sound collection device 1A of the second embodiment (the sound collection method according to the embodiment) having the above configuration will be described.
[0063] Hereinafter, the operation of the area sound collection device 1A of the second embodiment will be mainly described with respect to the operation of the dead angle direction update unit 106, which is the difference from the first embodiment.
[0064] The dead angle direction update unit 106 obtains an optimal dead angle direction for suppressing the interfering sound based on the input spectrum supplied from the frequency analysis unit 101 and the current dead angle direction φan (rad) stored in the dead angle direction update unit 106, and updates the dead angle direction φan (rad) in that direction.
[0065] For the update of the dead angle direction φan (rad) in the dead angle direction update unit 106, any known method can be used, but the update method described below is preferable.
[0066] First, the dead angle direction update unit 106 calculates the null form spectrum Y ank (ω) for three directions of φan, φan + δφ, and φan - δφ for each microphone array and for each current dead angle direction φan (rad) stored in the directivity formation unit according to the following equations (6) and (7). Here, δφ is a preset shift value of the dead angle direction.
[0067] Next, the blind spot direction update unit 106 calculates the sum P in the frequency direction for the null form spectrum Y ank (ω) according to the following equations (8) and (9). ank Take it. [Number]
[0068] Then, the blind spot direction update unit 106 updates the blind spot direction at which the value of this sum P ank is the smallest as the new blind spot direction. For example, when the sum value of the null form spectrum for φan + δφ is the smallest, the blind spot direction update unit 106 updates φan + δφ as the new blind spot direction. Then, the blind spot direction update unit 106 supplies the updated blind spot direction value to the null form formation unit 102.
[0069] Note that for the above shift value δφ, it is desirable to set it as the angle between adjacent null forms (for example, an angle around the middle between adjacent null forms). Also, in the above example, for each null form, evaluation values (sum P ank ) are calculated for the three update angle candidates of φan, φan + δφ, and φan - δφ from the angles within a predetermined range including φan (angles within the range of "φan - δφ" to "φan + δφ"), and the update angle candidate with the highest evaluation (sum P ankThe smallest angle (the smallest angle) is updated as the new dead angle direction, but the update angle candidates are not limited and may be more than three. For example, the update angle candidates may be five, namely φan, φan + δφ / 2, φan + δφ, φan - δφ / 2, and φan - δφ. Also, the dead angle direction update unit 106 may perform the dead angle direction update process for some null formers, or may update the dead angle direction for all null formers. That is, in the dead angle direction update unit 106, the number and combination of null formers for which the dead angle direction update process is performed are not limited either. Note that the above shift value δφ is preferably a minute angle (for example, π / 180) so that it can be adjusted according to the direction of the interfering sound, and it is desirable to set an upper limit value and a lower limit value as the angle between adjacent null formers (for example, an angle around the middle between adjacent null formers) for the dead angle direction.
[0070] (B-3) Effects of the Second Embodiment In the second embodiment, in addition to the effects of the first embodiment, the following effects can be achieved.
[0071] In the area sound collection device 1A of the second embodiment, by updating the dead angle direction of the null former formation unit 102 in real time based on the value of the null former spectrum, it becomes possible to direct the dead angle direction toward the direction of the interfering sound. Thereby, in the area sound collection device 1A, it is possible to suppress the interfering sound more regardless of the position of the interfering sound.
[0072] (C) Third Embodiment Hereinafter, a third embodiment of the sound collection device, the sound collection program, and the sound collection method according to the present invention will be described in detail with reference to the drawings. In this embodiment, an example in which the sound collection device, the sound collection program, and the sound collection method of the present invention are applied to an area sound collection device will be described.
[0073] (C-1) Configuration of the Third Embodiment FIG. 6 is a block diagram showing the functional configuration of the area sound collection device 1B according to the third embodiment, and the same reference numerals or corresponding reference numerals are assigned to the same parts or corresponding parts as those in FIG. 1 described above.
[0074] Hereinafter, the differences between the area sound collection device 1B of the third embodiment and the first embodiment will be described.
[0075] The area sound collection device 1B of the third embodiment is different from the first embodiment in that a nullformer selection unit 107 is inserted between the nullformer formation unit 102 and the filter gain calculation unit 103.
[0076] In the first embodiment, when calculating the filter gain by the filter gain calculation unit 103, the filter gain was calculated using all the nullformers supplied from the nullformer formation unit 102.
[0077] When there are multiple interfering sounds, it is possible to suppress more interfering sounds by increasing the number of nullformers. However, when the number of interfering sounds is less than the number of nullformers, the nullformers having a dead angle in the direction of the interfering sound greatly contribute to the suppression of the interfering sound, while the other nullformers do not contribute to the suppression of the interfering sound and are thus unnecessary. On the contrary, using unnecessary nullformers in the calculation of the filter gain may cause the target sound to be excessively suppressed and the target voice to deteriorate.
[0078] Therefore, in the area sound collection device 1B of the third embodiment, a nullformer selection unit 107 is added to select only specific nullformers from the plurality of nullformers supplied from the nullformer formation unit 102 and supply them to the filter gain calculation unit 103. As a result, in the area sound collection device 1B of the third embodiment, it is possible to exclude the nullformers that do not contribute to the suppression of the interfering sound and form the filter gain with fewer nullformers, and it is possible to reduce the influence of the increased voice distortion caused by increasing the number of nullformers to be multiplied.
[0079] (C-2) Operations of the Third Embodiment Next, the operation of the area sound collection device 1B according to the third embodiment (the sound collection method according to the embodiment) having the above configuration will be described.
[0080] Hereinafter, the operation of the area sound collection device 1B according to the third embodiment will be described centering on the operation of the nullformer selection unit 107, which is the difference from the first embodiment.
[0081] The nullformer selection unit 107 selects only a specific nullformer based on the values of the nullformer spectra of a plurality of nullformers supplied from the nullformer formation unit, and supplies the nullformer spectrum of that nullformer to the filter gain calculation unit. For the selection of a specific nullformer in the nullformer selection unit 107, any known method can be used, but the update method described below is preferable.
[0082] First, the nullformer selection unit 107 calculates the sum P an in the frequency direction for a plurality of nullformer spectra Y an ().
Equation
[0083] Next, the nullformer selection unit 107 extracts only a predetermined number of nullformers in ascending order of the value of this sum P an and supplies them to the filter gain calculation unit 103. Here, when the nullformer selection unit 107 extracts a predetermined number of nullformers, it is not necessary to determine the number to be extracted in advance, and the number to be extracted may be dynamically determined according to P an . For example, the nullformer selection unit 107 may compare P an with a predetermined threshold value, and supply to the filter gain calculation unit 103 the nullformers for which P an is smaller than the threshold value.
[0084] (C-3) Effects of the Third Embodiment According to the third embodiment, compared with the first embodiment, the following effects can be achieved.
[0085] In the area sound collection device 1B of the third embodiment, the nullformer selection unit 107 can exclude the nullformers that do not contribute to the suppression of interfering sound from the calculation of the filter gain, and it becomes possible to further reduce the distortion of the target voice.
[0086] (D) Fourth Embodiment FIG. 7 is a block diagram showing the functional configuration of the area sound collection device 1C according to the fourth embodiment, and the same reference numerals or corresponding reference numerals are assigned to the same parts or corresponding parts as those in FIG. 1 described above.
[0087] Hereinafter, the differences between the area sound collection device 1C of the fourth embodiment and the first embodiment will be described.
[0088] The area sound collection device 1C of the fourth embodiment is different from the first embodiment in that an amplitude correction unit 108 is inserted after the frequency analysis unit 101 (between the frequency analysis unit 101 and the nullformer formation unit 102).
[0089] The amplitude correction unit 108 corrects the amplitude ratio of the two microphones constituting each microphone array. In the amplitude correction unit 108, for each microphone array, the amplitude ratio of the two microphones is obtained using an arbitrary known method, and the correction is performed so that the amplitude ratio approaches 1. As a result, in the area sound collection device 1C, it becomes possible to suppress more interfering sound when the nullformer formation unit 102 calculates the nullformer spectrum.
[0090] Note that, also in the area sound collection devices 1, 1A, and 1B of the first to third embodiments described above, the amplitude correction unit 108 may be added.
[0091] (E) Fifth Embodiment FIG. 8 is a block diagram showing the functional configuration of the area sound collection device 1D according to the fifth embodiment, and the same reference numerals or corresponding reference numerals are assigned to the same parts or corresponding parts as those in FIG. 1 described above.
[0092] Hereinafter, the differences between the area sound collection device 1D of the fifth embodiment and the first embodiment will be described.
[0093] The area sound collection device 1D of the fifth embodiment is different from the first embodiment in that a noise suppression unit 109 is inserted between the frequency analysis unit 101 and the nullformer formation unit 102 (the subsequent stage of the frequency analysis unit 101).
[0094] As the noise suppression unit 109, for example, noise is removed from the input spectrum by a noise suppression method such as a Wiener filter. As a result, in the area sound collection device 1D, it is possible to remove diffusive noise that cannot be removed by beamforming, and it becomes possible to use it in a wider scene.
[0095] Note that, in the area sound collection devices 1, 1A, 1B, and 1C of the first to fourth embodiments described above, a configuration in which the noise suppression unit 109 is added may also be adopted.
[0096] (F) Other Embodiments The present invention is not limited to the above-described embodiments, and modified embodiments as exemplified below can also be cited.
[0097] (F-1) In the process of the filter gain calculation unit 103 of each of the above embodiments, the filter gain G a (ω) is calculated by multiplying the ratio of the nullformer spectrum to the input spectrum of the microphone array by the number of nullformers, but as shown in the following equation (11), it may be multiplied by the square root of the power of the ratio of the nullformer spectrum to the input spectrum of the microphone array with respect to the number N of nullformers.
Equation
[0098] (F-2) In the area sound collection devices 1 to 1D of each of the above embodiments, a configuration in which digital signals are supplied from the microphone arrays MA1 and MA2 has been shown. However, an analog signal may be supplied from the microphone arrays MA1 and MA2 to the area sound collection devices 1 to 1D, and the area sound collection devices 1 to 1D may be configured to convert the analog signal (the acoustic signal captured by each microphone of each microphone array) into a digital signal.
Explanation of Signs
[0099] 1, 1A, 1B, 1C, 1D... Area sound collection devices, 101... Frequency analysis unit, 102... Null-former formation unit, 103... Filter gain calculation unit, 104... Acoustic enhancement unit, 105... Beamformer selection unit, 106... Dead angle direction update unit, 107... Null-former selection unit, 108... Amplitude correction unit, 109... Noise suppression unit, M11, M12, M21, M22, MR, ML... Microphones, MA1, MA2, MAa... Microphone arrays, TA... Target area
Claims
1. Null-former forming means for calculating a null-former spectrum by forming a plurality of null-formers each having a blind spot facing outside the target area for each microphone array based on the input spectrum of acoustic signals supplied from a plurality of microphone arrays; Filter gain calculating means for calculating a filter gain of a frequency filter that passes components in the direction of the target area for each microphone array based on the null-former spectrum calculated by the null-former forming means; Acoustic enhancement means for calculating a beamformer output sound having directivity formed in the direction of the target area for each microphone array based on the filter gain calculated by the filter gain calculating means and the input spectrum; Target area sound extraction means for forming a target area sound based on the beamformer output sound calculated by the acoustic enhancement means A sound collection device characterized by comprising:
2. The sound collection device according to claim 1, further comprising blind spot direction updating means for evaluating each of a plurality of update angle candidates selected from angles within a predetermined range including the blind spot direction of the null-former for some or all of the blind spot directions of the null-former, and updating the blind spot direction of the null-former to the update angle candidate with the highest evaluation.
3. The sound collection device according to claim 2, wherein the blind spot direction updating means calculates a total value in the frequency direction of the null-former spectrum for each of the update angle candidates, and selects the update angle candidate having the smallest total value as the one with the highest evaluation.
4. The sound collection device according to claim 1, further comprising null-former selection means for selecting some of the null-formers based on the null-former spectrum of each of the null-formers, wherein the filter gain calculating means calculates the gain of the frequency filter based only on the null-former spectrum of the null-former selected by the null-former selection means.
5.
5. The sound collection device according to claim 1, further comprising amplitude correction means for correcting each microphone array so that the amplitude ratio of the acoustic signals of the respective microphones approaches 1, wherein the null-former forming means calculates each of the null-formers based on the acoustic signals corrected by the amplitude correction means.
6.
6. For each of the microphone arrays, further comprising noise suppression means for suppressing noise of the acoustic signal of each microphone. The null-former forming means calculates the null-former based on the acoustic signal whose noise has been suppressed by the noise suppression means. The sound collection device according to claim 1, characterized in that.
7. A computer, Based on the input spectrum of the acoustic signals supplied from a plurality of microphone arrays, for each microphone array, null-former forming means for forming a plurality of null-formers facing dead angles outside the target area and calculating a null-former spectrum; Based on the null-former spectrum calculated by the null-former forming means, for each microphone array, filter gain calculating means for calculating the filter gain of a frequency filter that passes components in the direction of the target area; For each microphone array, acoustic enhancement means for calculating a beamformer output sound having directivity in the direction of the target area based on the filter gain calculated by the filter gain calculating means and the input spectrum; Target area sound extraction means for forming a target area sound based on the beamformer output sound calculated by the acoustic enhancement means A sound collection program characterized by functioning as such.
8. In the sound collection method performed by the sound collection device, The sound collection device includes null-former forming means, filter gain calculating means, acoustic enhancement means, and target area sound extraction means. The null-former forming means forms a plurality of null-formers facing dead angles outside the target area for each microphone array based on the input spectrum of the acoustic signals supplied from a plurality of microphone arrays, and calculates a null-former spectrum. The filter gain calculating means calculates the filter gain of a frequency filter that passes components in the direction of the target area for each microphone array based on the null-former spectrum calculated by the null-former forming means. The acoustic enhancement means calculates a beamformer output sound having directivity in the direction of the target area for each microphone array based on the filter gain calculated by the filter gain calculating means and the input spectrum. The target area sound extraction means forms a target area sound based on the beamformer output sound calculated by the acoustic enhancement means. A sound collection method characterized by that.
Citation Information
Patent Citations
Signal processor
JP1998207490A
Earhole attachment-type sound pickup device, signal processing device, and sound pickup method
JP2013121106A
Sound pickup device and program
JP2013183358A
Sound source separating device, sound source separating program, sound collecting device, and sound collecting program
JP2015050558A
Sound pickup device, program, and method
JP2018170617A