Sound collecting device, sound collecting program, and sound collecting method
By employing multiple null formers to create dead angles outside the target area and selecting the minimum amplitude outputs for area sound extraction, the sound collection device achieves sharp directivity and reduced distortion with fewer microphones, addressing the limitations of conventional technologies.
Patent Information
- Application Number
- JP2023129522
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-08-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-08-08
AI Technical Summary
Conventional sound collection devices face challenges in achieving sharp directivity for area sound collection while minimizing the number of microphones, as they either require a large number of microphones or suffer from distortion using spectral subtraction with fewer microphones.
The solution involves using a microphone group of three or more microphones to calculate multiple null formers that create dead angles outside the target area. These null formers are used to select the output with the minimum amplitude for each frequency, and the beamformer selection unit chooses the output with the minimum amplitude across all microphone arrays to extract the area sound without distortion.
This approach allows for effective area sound collection with reduced distortion, achieving sharp directivity using fewer microphones compared to traditional methods, while maintaining the clarity of the target area sound.
Smart Images

Figure 0007697497000005 
Figure 0007697497000006 
Figure 0007697497000007
Abstract
Description
Technical Field
[0001] The present invention relates to a sound collection device, a sound collection program, and a sound collection method, and can be applied to area sound collection processing for suppressing sounds other than those in a specific area and emphasizing only the sounds in the specific area in order to perform voice recognition only on the voices emitted from the specific area, for example.
Background Art
[0002] When using a voice recognition system in a noisy environment, the surrounding noise that is mixed in simultaneously with the necessary target sound is a troublesome presence that causes a decrease in the voice recognition rate of the recorded voice.
[0003] Conventionally, in an environment where such a plurality of sound sources exist, as a technique for obtaining a necessary target sound by avoiding the mixing of unnecessary sounds by collecting only the sound in a specific direction, there is beamforming using a microphone array. Beamforming is a technique for forming directivity by utilizing the time difference of signals reaching each microphone (see Non-Patent Document 1). A microphone array that performs beamforming is also called a beamformer.
[0004] On the other hand, since a beamformer collects all the sounds in a specific direction, when it is desired to collect only a specific area (hereinafter also referred to as the "target area"), sounds outside the target area that are in the same direction as viewed from the beamformer are also collected. Therefore, Patent Document 1 proposes a method (area sound collection) of using a plurality of beamformers, directing directivity to the target area from different directions, and intersecting the directivities in the target area to collect the target sound.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Non-Patent Documents
[0006]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] By the way, when collecting human voices in area pickup, it is desirable for the beamformer to have a sharp directivity in order to form the target area. For example, when collecting human voices, the beam width of the beamformer is desirably approximately 60 degrees or less.
[0008] And in both additive and subtractive beamformers, a large number of microphones are required to form a sharp directivity. For example, in the Minimum Variance Distortionless Response (MVDR) method, the number of microphones equal to the number of interfering sound sources to be suppressed + 1 is required. Such area pickup with a large number of microphones can be experimentally realized in a laboratory, but it is not suitable for commercial products in terms of component cost and installation space.
[0009] On the other hand, in the technology described in Patent Document 1, by using spectral subtraction (subtracting the amplitude spectra from each other, and replacing with zero or a small number when the subtraction result is negative) in the frequency domain, a sharp directivity is realized with a small number of microphones. A nullformer that forms a dead angle in the target area direction is created, and by spectrally subtracting the nullformer amplitude spectrum from the microphone input amplitude spectrum, a sharp directivity is formed in the target area direction. By using this method, a sharp directivity can be formed with only 2 microphones, but it causes distortions such as the generation of musical noise and the omission of area sound components.
[0010] Therefore, in the conventional technology, there are problems that the device becomes large-scale because a large number of microphones are required to realize area pickup, or the area sound extracted by using spectral subtraction with a small number of microphones is distorted.
[0011] In view of the above problems, there is a need for a sound collection device, a sound collection program, and a sound collection method that suppress distortion of sound collected during area sound collection processing for collecting sound in a target area.
Means for Solving the Problem
[0012] A first aspect of the present invention is based on acoustic signals supplied from a microphone group composed of three or more microphones, and for each of a plurality of microphone arrays that can be configured by combinations of the microphones in the microphone group, calculates a plurality of null formers that form dead angles outside the target area to obtain null former output sounds, and null former processing means; and area sound extraction means for obtaining an area sound obtained by extracting the sound of the target area based on the null former output sounds of the respective null formers. The area sound extraction means includes, for each of the microphone arrays, a null former selection unit that selects, for each frequency, the null former output sound with the minimum output from the plurality of null former output sounds as the beamformer output sound of the microphone array, and a beamformer selection unit that selects the beamformer output sound with the minimum output from the plurality of beamformer output sounds as the area sound. Then, the nullformer processing means acquires the same number of nullformer output sounds from each of a first azimuth angle range from -π / 2 to 0 and excluding the target area direction and a second azimuth angle range from π / 2 to 0 and excluding the target area direction for each of the microphone arrays. It is characterized by this.
[0013] The second sound collection program of the present invention causes a computer to calculate a plurality of null formers that form dead angles outside a target area for each of a plurality of microphone arrays that can be configured by combinations of the microphones in the microphone group, based on acoustic signals supplied from a microphone group composed of three or more microphones, and to obtain null former output sounds, and functions as area sound extraction means for obtaining area sounds in which the sounds of the target area are extracted based on the null former output sounds of the respective null formers. The area sound extraction means includes, for each of the microphone arrays, a null former selection unit that selects, for each frequency, the null former output sound having the minimum output from among the plurality of null former output sounds as the beamformer output sound of the microphone array, and a beamformer selection unit that selects the beamformer output sound having the minimum output from among the plurality of beamformer output sounds as the area sound. Then, the nullformer processing means acquires the same number of nullformer output sounds from each of a first azimuth angle range from -π / 2 to 0 and excluding the target area direction and a second azimuth angle range from π / 2 to 0 and excluding the target area direction for each of the microphone arrays. It is characterized by this.
[0014] A third aspect of the present invention is a sound collection method performed by a sound collection device. The sound collection device includes null former processing means and area sound extraction means. The area sound extraction means includes a null former selection unit and a beamformer selection unit. The null former processing means calculates a plurality of null formers that form dead angles outside a target area for each of a plurality of microphone arrays that can be configured by combinations of the microphones in the microphone group, based on acoustic signals supplied from a microphone group composed of three or more microphones, and obtains null former output sounds. The area sound extraction means obtains area sounds in which the sounds of the target area are extracted based on the null former output sounds of the respective null formers. The null former selection unit selects, for each frequency, the null former output sound having the minimum output from among the plurality of null former output sounds as the beamformer output sound of the microphone array, and the beamformer selection unit selects the beamformer output sound having the minimum output from among the plurality of beamformer output sounds as the area sound. Then, the nullformer processing means acquires the same number of nullformer output sounds from each of a first azimuth angle range from -π / 2 to 0 and excluding the target area direction and a second azimuth angle range from π / 2 to 0 and excluding the target area direction for each of the microphone arrays. It is characterized by this.
Effect of the Invention
[0015] According to the present invention, it is possible to provide a sound collection device, a sound collection program, and a sound collection method that suppress distortion of sound collected when performing area sound collection processing for collecting sound in a target area.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Modes for Carrying Out the Invention
[0017] (A) First Embodiment Hereinafter, a first embodiment of the sound collection device, the sound collection program, and the sound collection method according to the present invention will be described in detail with reference to the drawings.
[0018] (A-1) Configuration of the First Embodiment FIG. 1 is a block diagram showing the functional configuration of the sound collection device 10 according to the first embodiment.
[0019] The sound collection device 10 suppresses the sound emitted outside the target area and collects the sound emitted within the target area based on the acoustic signals captured by the two microphone arrays MA (MA1, MA2).
[0020] In this embodiment, it is assumed that the microphone arrays MA1 and MA2 are each a microphone array including two microphones. The microphone array MA1 includes two microphones M (M1, M2), and the microphone array MA2 also includes two microphones M (M3, M4). That is, in the first embodiment, two microphone arrays MA1 and MA2 are configured using a microphone group composed of four microphones M1 to M4.
[0021] Note that the sound collection device 10 of this embodiment will be described as corresponding to a configuration having two microphone arrays each including two microphones, but it may be configured to correspond to three or more microphone arrays, or each microphone array may include three or more microphones. However, from the viewpoints of component cost and installation space, it is desirable to minimize the number of microphone arrays / microphones that the sound collection device 10 of this embodiment corresponds to. Therefore, the configuration of FIG. 1 (a configuration of two 2ch microphone arrays) with the minimum configuration is most preferable.
[0022] FIG. 2 is a diagram showing an example of the positional relationship between the two microphone arrays (microphone arrays MA1, MA2) and the target area TA.
[0023] In FIG. 2, an area where the front directions of the microphone array MA1 and the microphone array MA2 intersect is illustrated as a target area TA. In FIG. 2, the target area TA is hatched (diagonal lines). In this embodiment, as shown in FIG. 2, the number of microphones constituting each microphone array is two, and the number of microphone arrays is two. However, if each is two or more, the number of microphones and microphone arrays may be increased. For example, as shown in FIG. 3, a configuration in which three microphone arrays each including four microphones are arranged (constituting three microphone arrays from a microphone group composed of 12 microphones) may be used. In FIG. 3, a microphone array MA1 including microphones M1 to M4, a microphone array MA2 including microphones M5 to M8, and a microphone array MA3 including microphones M9 to M12 are arranged around the target area TA. In FIG. 3, an area where the front directions of the microphone arrays MA1 to MA3 intersect is defined as the target area TA and is hatched (diagonal lines).
[0024] In FIG. 3, the target area TA is hatched (diagonal lines).
[0025] Next, based on FIG. 1, the internal configuration of the sound collection device 10 will be described.
[0026] The sound collection device 10 includes a nullformer processing unit 101, a nullformer selection unit 102, and a beamformer selection unit 103.
[0027] Based on the acoustic signals captured by the two microphone arrays MA1 and MA2, the nullformer processing unit 101 forms a plurality of nullformers each having a dead angle facing outside the target area (a direction not including the target area), and obtains a plurality of nullformer output sounds. In the first embodiment, it is assumed that the nullformer processing unit 101 includes means (hereinafter referred to as "signal input means") for converting the acoustic signals captured by the microphone arrays MA1 and MA2 from analog signals to digital signals and further converting them into the frequency domain.
[0028] The null-former selection unit 102 obtains, for each microphone array, a sound with directivity formed in the direction of the target area (hereinafter referred to as "beamformer output sound") based on a plurality of null-former output sounds calculated by the null-former processing unit 101.
[0029] The beamformer selection unit 103 obtains the result of extracting the target area sound (hereinafter referred to as "area sound") based on the beamformer output sound calculated by the null-former selection unit 102 for each microphone array.
[0030] As described above, in the first embodiment, the null-former selection unit 102 and the beamformer selection unit 103 constitute area sound extraction means for performing area sound collection processing using a null-former.
[0031] FIG. 4 is a block diagram showing an example of the hardware configuration of the sound collection device 10.
[0032] FIG. 4 shows an example of the hardware configuration when the sound collection device 10 is configured using software (computer).
[0033] The sound collection device 10 shown in FIG. 4 has, as a hardware component, a computer 300 in which a program (including the sound collection program of the embodiment) is installed. Also, the computer 300 may be a computer dedicated to the sound collection program or may be configured to be shared with programs of other functions.
[0034] The computer 300 shown in FIG. 4 has a processor 301, a primary storage unit 302, and a secondary storage unit 303. The primary storage unit 302 is a storage means that functions as a working memory (work memory) for the processor 301. For example, a high-speed operating memory such as a DRAM (Dynamic Random Access Memory) can be applied. The secondary storage unit 303 is a storage means for recording various data such as an OS (Operating System) and program data (including data of the sound collection program according to the embodiment). For example, a non-volatile memory such as a FLASH (registered trademark) memory, an HDD, or an SSD can be applied. In the computer 300 of this embodiment, when the processor 301 is activated, the OS and programs (including the sound collection program according to the embodiment) recorded in the secondary storage unit 303 are read and expanded and executed on the primary storage unit 302.
[0035] Note that the specific configuration of the computer 300 is not limited to the configuration of FIG. 4, and various configurations can be applied. For example, if the primary storage unit 302 is a non-volatile memory (for example, a FLASH memory or the like), the configuration excluding the secondary storage unit 303 may be used.
[0036] (A-2) Operation of the First Embodiment Next, the operation of the sound collection device 10 of the first embodiment (the sound collection method according to the embodiment) having the above configuration will be described.
[0037] In the null-former processing unit 101, the acoustic signals captured by the microphone arrays MA1 and MA2 (microphones M1 to M4) are converted from analog signals to digital signals by the signal input means, arbitrary frequency conversion is performed, and the obtained frequency spectrum (hereinafter referred to as the "input spectrum") is acquired. The frequency conversion method for obtaining the input spectrum by the signal input means is most preferably the fast Fourier transform, but is not limited thereto, and discrete Gabor transform, wavelet transform, filter bank, etc. may be used.
[0038] Based on the supplied input spectrum, for each microphone array, the nullformer processing unit 101 calculates a plurality of nullformers that form dead zones in a predetermined direction other than the target area, and supplies the obtained nullformer spectrum (hereinafter referred to as "nullformer output sound") to the nullformer selection unit 102.
[0039] Next, the details of the processing of the nullformer processing unit 101 will be described.
[0040] FIG. 5 is a diagram showing an example of the direction of the nullformer in an arbitrary microphone array MAa.
[0041] Hereinafter, "a" in "MAa" etc. shall indicate the number (identifier; ID) of the microphone array. Here, the number of the microphone array MA1 is 1, and the number of the microphone array MA2 is 2.
[0042] That is, the microphone array MAa shown in FIG. 5 shall indicate either the microphone array MA1 or MA2 (arbitrary microphone array). In FIG. 5, it is illustrated that the microphone array MAa includes microphones ML and MR. For example, when the microphone array MAa is the microphone array MA1, the microphones ML and MR are the microphones M1 and M2 respectively, and when the microphone array MAa is the microphone array MA2, the microphones ML and MR are the microphones M3 and M4 respectively. In the image diagram of FIG. 5, for the sake of illustration, the shape of the target area TA is shown as circular, but in this embodiment, the actual shape of the target area TA is as shown in FIG. 2.
[0043] Here, the front direction (0-degree direction) of the microphone array MAa is defined as the direction orthogonal to the straight line connecting the positions (center positions) of the two microphones ML and MR, and the angle formed with the direction from the center point of the two microphones ML and MR to the center point (which may be the center of gravity) of the target area TA is the smallest.
[0044] Incidentally, as shown in FIG. 2, since the target area TA has a range rather than a point, the direction of the target area as seen from each microphone array also has a range. Therefore, here, the direction of the target area (the range of the direction of the target area) in the microphone array MA1 is set as θ1L to θ1R (rad) (-π / 2 < θ1L < 0 < θ1R < π / 2), and the direction of the target area (the range of the direction of the target area) in the microphone array MA2 is set as θ2L to θ2R (rad) (-π / 2 < θ2L < 0 < θ2R < π / 2).
[0045] As described above, the predetermined dead angle direction applied to each nullformer of the nullformer processing unit 101 must be outside the target area direction (a direction that does not include the range of the target area direction). Therefore, the n-th dead angle direction φan (rad) (n = 1, …, N) of the microphone array MAa (a = 1, 2) needs to satisfy -π - θaL < φan < θaL, or θaR < φan < π - θaL.
[0046] Here, since there are two microphones constituting the microphone array MAa, if a dead angle is formed in φan, a dead angle is also formed in π - φan (=-π - φan). Also, since it is difficult to know in advance from which direction the noise arrives, it is desirable to determine the predetermined dead angle direction in advance by design (apply a fixed value designed in advance).
[0047] From the above, it is preferable to determine the same number of predetermined dead angle directions applied to each nullformer of the nullformer processing unit 101 from within the two ranges of -π / 2 ≦ φan < θaL and θaR < φan ≦ π / 2, respectively. In the nullformer processing unit 101, the number of predetermined dead angle directions set for each microphone array (the number of nullformers set for each microphone array MAa in the nullformer processing unit 101) and the combination of dead angle directions are not limited. For example, when the number N of predetermined dead angle directions is 6 (when N = 6), φa1 = -π / 2, φa2 = -π / 3, φa3 = -π / 6, φa4 = π / 6, φa5 = π / 3, φa6 = π / 2 can be selected, and this selection is preferable.
[0048] Next, an example of the specific processing of each nullformer in the nullformer processing unit 101 will be described.
[0049] Let the input spectrum of the microphone array MAa be X aL (ω), X aR (ω). When this is done, the nullformer processing unit 101 calculates a plurality of nullformer spectra Y an (ω) using Equation (1). Here, ω (omega) is the angular frequency, i is the imaginary unit, d is the microphone interval, and c is the speed of sound. [Equation]
[0050] Next, the specific processing of the nullformer selection unit 102 will be described.
[0051] Based on the plurality of supplied nullformer spectra, the nullformer selection unit 102 selects, for each microphone array and for each frequency, the one with the minimum amplitude of the nullformer spectrum, and obtains the resulting spectrum (hereinafter referred to as the "beamformer spectrum") as the beamformer output sound, and supplies it to the beamformer selection unit 103. Specifically, the nullformer selection unit 102 calculates the beamformer spectrum (beamformer output sound) B a (ω) using Equations (2) and (3). Here, ν(ω) (ν: new) is the index number of the dead angle at which the nullformer spectrum is minimum, and argmin_n{} is an operator that outputs n that minimizes the inside of {}.
[0052] That is, the beamformer spectrum B1(ω) is the sound in the target area direction (range of the target area direction) in the microphone array MA1, and the beamformer spectrum B2(ω) is the sound in the target area direction (range of the target area direction) in the microphone array MA2. The nullformer selection unit 102 supplies the beamformer spectrum B1(ω) of the microphone array MA1 and the beamformer spectrum B2(ω) of the microphone array MA2 to the beamformer selection unit 103.
Number
[0053] Next, the specific processing of the beamformer selection unit 103 will be described.
[0054] The beamformer selection unit 103 obtains, as the area sound, the spectrum (hereinafter referred to as the "area sound spectrum") obtained by selecting, for each frequency, the one with the minimum amplitude from the supplied beamformer spectra B1(ω) and B2(ω). Specifically, the beamformer selection unit 103 calculates the area sound spectrum Z(ω) using equations (4) and (5).
Number
[0055] In Patent Document 1, the area sound spectrum is extracted by performing two spectral subtractions on the amplitude of the beamformer spectrum. However, in this embodiment, the same processing can be realized using equations (4) and (5).
[0056] FIG. 6 is a diagram showing in tabular form the properties of the sound selected by the minimum selection of the beamformer selection unit 103.
[0057] In FIG. 6, within the range of the first beamformer is within the directivity range of the beamformer spectrum B1(ω) of the microphone array MA1, and within the range of the second beamformer is within the directivity range of the beamformer spectrum B2(ω) of the microphone array MA2.
[0058] As shown in FIG. 6, for a certain frequency, when sound exists within the range of both beamformers (when sound exists within the target area), regardless of which beamformer's sound (amplitude) is selected, the component of the target area sound will be selected. Also, as shown in FIG. 6, when sound exists only within the range of one of the beamformers, the beamformer on the side where no sound exists will be selected by minimum selection. Furthermore, as shown in FIG. 6, when sound exists outside the range of either beamformer, regardless of which beamformer's sound (amplitude) is selected, the sound outside the target area will not be selected. Therefore, as shown in FIG. 6, by the minimum selection process in the beamformer selection unit 103, sound is picked up only when sound exists within the range of both beamformers (that is, when area sound is emitted within the target area), and in other cases (that is, when sound is emitted outside the target area), the sound is not picked up.
[0059] (A-3) Effects of the First Embodiment According to the first embodiment, the following effects can be achieved.
[0060] In the sound pickup device 10 of the first embodiment, a beamformer is formed that suppresses sound arriving from directions other than the target area direction and picks up sound only from the target area direction, and the target area sound is picked up using this beamformer output. Thereby, in the sound pickup device 10 of the first embodiment, it is possible to pick up only the area sound without generating distortion such as that caused by spectral subtraction.
[0061] (B) Second Embodiment Hereinafter, a second embodiment of the sound pickup device, sound pickup program, and sound pickup method according to the present invention will be described in detail with reference to the drawings.
[0062] (B-1) Configuration of the Second Embodiment FIG. 7 is a block diagram showing the functional configuration of the sound pickup device 10A according to the second embodiment, and the same parts or corresponding parts as those in FIG. 1 described above are denoted by the same reference numerals or corresponding reference numerals.
[0063] In the first embodiment, it was an explicit constraint to prepare microphone arrays (two 2-channel microphone arrays in the example of the first embodiment) when arranging microphones. However, looking at equations (2) to (5), it can be seen that, regardless of which microphone array the components included in the target area sound finally originate from, the area sound spectrum can be obtained by selecting the one with the minimum output from all the nullform spectra.
[0064] Therefore, the sound collection device 10A of the second embodiment sets two or more microphone arrays using a microphone group including three or more microphones freely arranged so as to surround the target area, and extracts the area sound spectrum by selecting the one with the minimum output from the nullform spectra of all the microphone arrays.
[0065] Next, the internal configuration of the sound collection device 10A will be described with reference to FIG. 7.
[0066] The sound collection device 10A acquires the nullforms of two or more microphone arrays based on the acoustic signals captured by M microphones M1 to MM (M is 3 or more), suppresses the sound emitted outside the target area based on each acquired nullform, and collects the sound emitted inside the target area.
[0067] FIGS. 8 and 9 are diagrams showing configuration examples of the microphone arrays in the second embodiment.
[0068] In FIG. 8, three microphones M1 to M3 are arranged so as to surround the target area TA (M = 3), and an example in which two microphone arrays MA1 and MA2 are constituted by the three microphones M1 to M3 is shown. Further, in FIG. 9, four microphones M1 to M4 are arranged so as to surround the target area TA (M = 4), and an example in which three microphone arrays MA1 to MA3 are constituted by the four microphones M1 to M4 is shown. In FIGS. 8 and 9, the area where the front directions of the respective microphone arrays intersect is defined as the target area TA, and it is hatched (diagonal lines). The target area TA is hatched (diagonal lines). In the second embodiment, as shown in FIGS. 8 and 9, microphones (microphone groups) that are not configured as microphone arrays in advance are arranged around the target area TA, and an arbitrary combination of microphones is used as a microphone array on the sound collection device 10A side. It is assumed that the configuration is such that it is used. FIGS. 8 and 9 show an example in which the closest microphones to each other around the target area TA are combined to form a microphone array, but it is not always necessary to form a microphone array with the closest microphones to each other. For example, in FIG. 9, a microphone array may be formed by the microphones M1 and M3.
[0069] As shown in FIG. 7, the sound collection device 10A includes a microphone array selection unit 201, a nullformer processing unit 101A, and a nullformer selection unit 202.
[0070] The microphone array selection unit 201 selects a plurality of combinations from the M microphones M1 to MM and configures (functions) each combination (a combination of a plurality of microphones) as a microphone array. In the selection of microphones for each microphone array, it is free as long as different microphone arrays do not have exactly the same combination of microphones, but it must be selected so that the target area TA can be formed (in other words, the common part of the directivities of a plurality of beamformers formed later can be formed into a closed shape). As shown in FIGS. 8 and 9, a combination in which some microphones are common between different microphone arrays may be used.
[0071] In the second embodiment, the most preferable configuration is the one shown in FIG. 2 (a configuration of two sets of 2ch microphone arrays) that allows the target area TA to be designed relatively freely while reducing the number of microphones. However, for example, as shown in FIG. 3, a configuration of three microphone arrays composed of 12 microphones (a configuration of three sets of 4ch microphone arrays) may be used.
[0072] In the second embodiment, it is assumed that the microphone array selection unit 201 includes signal input means for converting the acoustic signals captured by the microphones M1 to MM from analog signals to digital signals and further converting them into the frequency domain.
[0073] The nullformer processing unit 101A forms a plurality of nullformers each having a dead angle facing outside the target area based on the acoustic signals captured by the microphones selected by each microphone array, and obtains a plurality of nullformer output sounds. That is, for each of the microphone arrays configured by the microphone array selection unit 201, the nullformer processing unit 101A forms a plurality of nullformers each having a dead angle facing outside the target area to obtain nullformer output sounds.
[0074] The nullformer selection unit 202 obtains area sounds based on the plurality of nullformer output sounds calculated by the nullformer processing unit 101A.
[0075] As described above, in the second embodiment, the nullformer selection unit 202 constitutes area sound extraction means for performing area sound collection processing using nullformers.
[0076] (B-2) Operation of the Second Embodiment Next, the operation of the sound collection device 10A of the second embodiment having the above configuration (the sound collection method according to the embodiment) will be described.
[0077] In the microphone array selection unit 201, the acoustic signals captured by the microphones M1 to MM are converted from analog signals to digital signals by the signal input means, and further subjected to arbitrary frequency conversion, and the obtained frequency spectrum is acquired as the input spectrum. Since the same processing as that in the first embodiment can be applied to the processing of the signal input means, a detailed description thereof is omitted.
[0078] The microphone array selection unit 201 assigns microphones to a predetermined R (R is an integer of 2 or more) microphone arrays according to a predetermined microphone selection list or the like, and supplies the input spectra supplied from the microphones M1 to MM to the nullformer processing unit 101A for each configured microphone array. Note that the number of microphone arrays set in the microphone array selection unit 201 and the selection process of the microphones associated with each microphone array may be performed based on preset (designed) parameters or programs, or may be changed dynamically.
[0079] Based on the supplied input spectrum, the nullformer processing unit 101A calculates a plurality of nullformers that form dead zones in a predetermined direction other than the target area for each microphone array, and supplies the obtained nullformer spectrum (nullformer output sound) to the nullformer selection unit 202.
[0080] Since the detailed operation of the nullformer processing unit 101A is the same as that of the nullformer processing unit 101 according to the first embodiment, a detailed description thereof is omitted. Here, the plurality of nullformer spectra obtained by the nullformer processing unit 101A are the same as those in the first embodiment, and are denoted as Y an (ω).
[0081] Based on the supplied plurality of nullformer spectra, the nullformer selection unit 202 selects, for each frequency, the one with the minimum amplitude of the nullformer spectrum, and outputs the obtained area sound spectrum (area sound). Specifically, the nullformer selection unit 202 calculates the area sound spectrum Z(ω) using equations (6) and (7).
Equation
[0082] (B-3) Effects of the Second Embodiment In the second embodiment, in addition to the effects of the first embodiment, the following effects can be achieved.
[0083] In the sound collection device 10A of the second embodiment, since the microphone array selection unit 201 can set a microphone array from the microphone group in an arbitrary combination, it is possible to more flexibly select the microphone arrangement and the configuration of the microphone array than in the first embodiment.
[0084] (C) Other Embodiments The present invention is not limited to the above-described embodiments, and modified embodiments as exemplified below can also be cited.
[0085] (C-1) In each of the above embodiments, the configuration in which the sound collection devices 10 and 10A include signal input means (means for converting the acoustic signals captured by the respective microphones from analog signals to digital signals and further converting them into the frequency domain) has been shown. However, the signal input means may be arranged outside the sound collection devices 10 and 10A (arranged between each microphone and the sound collection devices 10 and 10A).
Explanation of Reference Numerals
[0086] 10, 10A... Sound collection device, 101... Nullformer processing unit, 101A... Nullformer processing unit, 102... Nullformer selection unit, 103... Beamformer selection unit, 201... Microphone array selection unit, 202... Nullformer selection unit, M1 to MM... Microphones, MA1 to MA3... Microphone arrays.
Claims
1. Nullformer processing means for calculating a plurality of nullformers that form dead zones outside a target area for each of a plurality of microphone arrays that can be configured by combinations of the microphones in the microphone group, based on acoustic signals supplied from a microphone group composed of three or more microphones, and obtaining nullformer output sounds; Area sound extraction means for obtaining an area sound obtained by extracting the sound of the target area based on the nullformer output sound of each of the nullformers; The area sound extraction means For each of the microphone arrays, a nullformer selection unit that selects the nullformer output sound with the minimum output from among the plurality of nullformer output sounds for each frequency and uses it as the beamformer output sound of the microphone array; A beamformer selection unit that selects the beamformer output sound with the minimum output from among the plurality of beamformer output sounds and uses it as the area sound; The nullformer processing means obtains the same number of nullformer output sounds from each of a first azimuth range from -π / 2 to 0 and excluding the target area direction and a second azimuth range from π / 2 to 0 and excluding the target area direction for each of the microphone arrays. A sound collection device characterized by the above.
2. Further comprising microphone array selection means for selecting a plurality of combinations of the microphones from the microphone group and configuring a microphone array for each selected combination; The nullformer processing means calculates a plurality of nullformers for each of the microphone arrays configured by the microphone array selection means and obtains the nullformer output sounds; The area sound extraction means selects the nullformer output sound with the minimum output from among the plurality of nullformer output sounds for each frequency and uses it as the area sound. The sound collection device according to claim 1, characterized by the above.
3. A computer Nullformer processing means for calculating a plurality of nullformers that form dead zones outside a target area for each of a plurality of microphone arrays that can be configured by combinations of the microphones in the microphone group, based on acoustic signals supplied from a microphone group composed of three or more microphones, and obtaining nullformer output sounds; Function as area sound extraction means for obtaining an area sound obtained by extracting the sound of the target area based on the nullformer output sound of each of the nullformers; The area sound extraction means For each of the respective microphone arrays, a nullformer selection unit that selects, for each frequency, the nullformer output sound with the minimum output from among the plurality of nullformer output sounds and sets it as the beamformer output sound of the microphone array, and a beamformer selection unit that selects the beamformer output sound with the minimum output from among the plurality of beamformer output sounds and sets it as the area sound. The nullformer processing means acquires the same number of nullformer output sounds from each of a first azimuth angle range from -π / 2 to 0 and excluding the target area direction and a second azimuth angle range from π / 2 to 0 and excluding the target area direction for each of the respective microphone arrays. A sound collection program characterized by this.
4. In a sound collection method performed by a sound collection device, the sound collection device includes nullformer processing means and area sound extraction means, the area sound extraction means includes a nullformer selection unit and a beamformer selection unit, the nullformer processing means calculates a plurality of nullformers that form dead angles outside the target area for each of a plurality of microphone arrays that can be configured by combinations of the microphones in the microphone group based on an acoustic signal supplied from a microphone group composed of 3 or more microphones, and acquires nullformer output sounds, the area sound extraction means acquires an area sound obtained by extracting the sound of the target area based on the nullformer output sounds of the respective nullformers, the nullformer selection unit selects, for each frequency, the nullformer output sound with the minimum output from among the plurality of nullformer output sounds for each of the respective microphone arrays and sets it as the beamformer output sound of the microphone array, the beamformer selection unit selects the beamformer output sound with the minimum output from among the plurality of beamformer output sounds and sets it as the area sound, the nullformer processing means acquires the same number of nullformer output sounds from each of a first azimuth angle range from -π / 2 to 0 and excluding the target area direction and a second azimuth angle range from π / 2 to 0 and excluding the target area direction for each of the respective microphone arrays. A sound collection method characterized by this.
Citation Information
Patent Citations
Sound pickup device and program
JP2013183358A
Sound information processing device and programs
JP2020141160A
Sound collection device, sound collection program, sound collection method, and keyboard
JP2022169998A