Sound receiving device, sound receiving method, and storage medium storing a sound receiving program
By introducing a sensitivity correction unit and a target sound detection unit into the audio device, the problem of low noise suppression performance in the directional synthesis in the prior art is solved, and a more efficient noise suppression and a high signal-to-noise ratio target sound recording effect is achieved.
Patent Information
- Application Number
- CN202011268693.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-13
- Filing Date
- 2020-11-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-11-13
AI Technical Summary
The prior art has low noise suppression performance in directivity synthesis, especially when there is a sensitivity difference between multiple microphone elements, making it difficult to effectively separate the target sound.
By introducing a sensitivity correction unit into the audio recording device, the sensitivity difference is corrected by multiplying the output signals of the plurality of microphone elements by gain, and the gain is updated when the target sound detection unit detects the speech of the speaker to improve the noise suppression performance.
More efficient noise suppression in directional synthesis is achieved, reducing the leakage of target sound into the noise reference signal, thereby improving the signal-to-noise ratio sound effect.
Smart Images

Figure CN112822578B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for receiving a target sound using a plurality of microphone elements. Background Art
[0002] Conventionally, a beamformer that controls directivity using output signals from at least two microphone elements has been known. Further, there is a sound receiving device that suppresses ambient noise using the beamformer and separates a target sound from the ambient noise for reception. Due to differences in sensitivity between at least two microphone elements, the noise suppression performance of the beamformer may deteriorate.
[0003] For example, Japanese Patent No. 4734070 discloses a beamformer that combines an automatic calibration processing with a General Sidelobe Canceller (hereinafter referred to as GSC). Japanese Patent No. 4734070 discloses correcting differences in sensitivity between a plurality of microphones using ambient noise.
[0004] However, in the above-described conventional technology, there is a risk of poor noise suppression performance in directivity synthesis, and thus further improvement is required. Summary of the Invention
[0005] The present invention has been made to solve the above problems, and an object thereof is to provide a technique that can improve noise suppression performance in directivity synthesis and can receive a target sound with a high signal-to-noise ratio.
[0006] A sound receiving device according to an embodiment of the present invention includes: a plurality of microphone elements; a sensitivity correction unit that corrects a sensitivity difference between the plurality of microphone elements by multiplying an output signal of the plurality of microphone elements by a gain; a target sound detection unit that detects speech of a speaker as a target sound; a gain control unit that controls the gain based on a detection result of the target sound detection unit; and a directivity synthesis unit that emphasizes and receives the target sound arriving from a specified direction using the output signals of the plurality of microphone elements corrected by the sensitivity correction unit, wherein the gain control unit updates the gain based on the output signals of the plurality of microphone elements when the speech of the speaker is detected by the target sound detection unit, and does not update the gain when the speech of the speaker is not detected by the target sound detection unit. Brief Description of the Drawings
[0007] Figure 1 It is a block diagram showing the configuration of a sound receiving device according to a first embodiment of the present invention.
[0008] Figure 2 It is a schematic diagram showing an example of the installation position of the microphone array according to the first embodiment of the present invention.
[0009] Figure 3 It is a schematic diagram showing an example of the arrangement of microphone elements of the microphone array according to the first embodiment of the present invention.
[0010] Figure 4 It is a block diagram showing the configuration of the target sound detection unit of the sound collection device according to the first embodiment of the present invention.
[0011] Figure 5 It is a block diagram showing the configuration of the voice judgment unit of the sound collection device according to the first embodiment of the present invention.
[0012] Figure 6 It is a block diagram showing the configuration of the sensitivity correction control unit of the sound collection device according to the first embodiment of the present invention.
[0013] Figure 7 It is a flowchart for explaining the operation of the sound collection device according to the first embodiment of the present invention.
[0014] Figure 8 It is a block diagram showing the configuration of the sound collection device according to the second embodiment of the present invention.
[0015] Figure 9 It is a block diagram showing the configuration of the target sound detection unit of the sound collection device according to the second embodiment of the present invention.
[0016] Figure 10 It is a block diagram showing the configuration of the target sound direction judgment unit of the sound collection device according to the second embodiment of the present invention.
[0017] Figure 11 It is a block diagram showing the configuration of the target sound direction judgment unit of the sound collection device according to a modified example of the second embodiment of the present invention. Detailed Embodiments
[0018] (Basic Knowledge of the Present Invention)
[0019] As described above, in the prior art, the difference in sensitivity between multiple microphones is corrected using ambient noise.
[0020] However, in the case where the noise source is near a microphone array composed of a plurality of microphone elements, the distance differences between the noise source and the respective microphone elements cannot be ignored, and the noise generated from the noise source appears as a sound pressure difference at the positions of the respective microphone elements. When performing sensitivity correction or auto-calibration on the plurality of microphone elements using the noise generated from a noise source near such a microphone array, the sensitivity correction or auto-calibration cannot be performed correctly, and instead, the output performance of the beamformer in a subsequent stage may deteriorate.
[0021] Particularly in GSC, a blocking matrix creates a noise reference signal having a dead angle of sensitivity in the target sound direction. However, if there are sensitivity differences between the plurality of microphone elements, a dead angle of sensitivity cannot be formed in the target sound direction, and the target sound will leak into the noise reference signal. In this case, by subtracting the noise reference signal in which the target sound leaking through the subsequent-stage adaptive noise cancellation is included from the output of the weighted sum beamformer, distortion may be caused to the target sound of the output signal. In order to prevent the target sound from leaking into the noise reference signal, it is necessary to make the sensitivities of at least the plurality of microphone elements consistent.
[0022] To solve the above problems, a sound collection device according to an embodiment of the present invention includes: a plurality of microphone elements; a sensitivity correction unit that corrects a sensitivity difference between the plurality of microphone elements by multiplying output signals of the plurality of microphone elements by a gain; a target sound detection unit that detects speech of a speaker as a target sound; a gain control unit that controls the gain based on a detection result of the target sound detection unit; and a directional synthesis unit that emphasizes and collects the target sound arriving from a specified direction using the output signals of the plurality of microphone elements corrected by the sensitivity correction unit, wherein the gain control unit updates the gain based on the output signals of the plurality of microphone elements when the speech of the speaker is detected by the target sound detection unit, and does not update the gain when the speech of the speaker is not detected by the target sound detection unit.
[0023] According to this configuration, the sensitivity difference between the plurality of microphone elements is corrected by multiplying the output signals of the plurality of microphone elements by a gain. At this time, when the speech of the speaker is detected, the gain is updated based on the output signals of the plurality of microphone elements, and when the speech of the speaker is not detected, the gain is not updated. Moreover, the target sound arriving from a specified direction is emphasized and collected using the output signals of the plurality of microphone elements whose sensitivity differences have been corrected.
[0024] Therefore, when the target sound, i.e., the speech of the speaker, is detected, since the gain for correcting the sensitivity difference between the multiple microphone elements is updated, the sensitivity difference for the target sound between the multiple microphone elements can be corrected. In the subsequent directional synthesis, the leakage amount of the target sound leaking into the noise reference signal in the dead angle with sensitivity in the target sound direction can be reduced. As a result, the noise suppression performance in the directional synthesis can be improved and the target sound can be received with a higher signal-to-noise ratio.
[0025] Moreover, in the above-described sound receiving device, it is also possible that the target sound detection unit includes a voice determination unit that determines whether the output signal of one of the multiple microphone elements is the voice or non-voice other than the voice.
[0026] According to this configuration, by determining whether the output signal of one of the multiple microphone elements is the voice or non-voice other than the voice, the speech of the speaker can be easily detected.
[0027] Moreover, in the above-described sound receiving device, it is also possible that the target sound detection unit includes a first extraction unit that extracts a signal in a specific frequency band from the output signal of the one microphone element, and the voice determination unit determines whether the signal extracted by the first extraction unit is the voice or the non-voice.
[0028] According to this configuration, since it is determined whether the signal in the specific frequency band extracted from the output signal of one microphone element is the voice or the non-voice, the speech of the speaker can be detected with higher accuracy.
[0029] Moreover, in the above-described sound receiving device, it is also possible that the target sound detection unit includes: a target sound direction determination unit that determines whether the target sound arrives from a predetermined target sound direction by using the output signals of the multiple microphone elements; and a target sound determination unit that determines that the target sound is detected when it is determined by the target sound direction determination unit that the target sound arrives from the target sound direction and it is determined by the voice determination unit that the output signal of the one microphone element is the voice.
[0030] If only it is determined whether the voice is detected, there may be a case where it is determined that the voice is detected when there is speech from a direction other than the target sound direction. However, according to the above-described configuration, since it is determined that the target sound is detected and the gain is updated only when the voice is detected and the target sound arrives from the target sound direction, the sensitivity difference can be corrected with higher accuracy using the target sound.
[0031] Furthermore, the sound collection device may also be such that the target sound detection unit includes a second extraction unit that extracts signals in a specific frequency band from the output signals of the plurality of microphone elements, and the target sound direction determination unit determines whether the target sound arrives from the target sound direction based on the signals extracted by the second extraction unit.
[0032] According to this configuration, since it is determined whether the target sound arriving from the target sound direction is based on the signals in the specific frequency band extracted from the output signals of the plurality of microphone elements, it is possible to determine with higher accuracy whether the target sound arrives from the target sound direction.
[0033] Furthermore, the sound collection device may also be such that the target sound direction determination unit includes: a direction estimation unit that estimates the direction of arrival of the target sound using the phase difference of the output signals of the plurality of microphone elements; and a direction determination unit that determines whether the direction estimated by the direction estimation unit is the predetermined target sound direction.
[0034] The direction of arrival of the target sound can be easily estimated using the phase difference of the output signals of the plurality of microphone elements. Therefore, the determination of whether the target sound arrives from the target sound direction can be easily made based on the estimation result of the direction of arrival of the target sound if the target sound direction is known in advance.
[0035] Furthermore, the sound collection device may also be such that the target sound direction determination unit includes: a first directivity synthesis unit that forms directivity in the target sound direction by emphasizing the signals in the target sound direction using the output signals of the plurality of microphone elements; a second directivity synthesis unit that forms a dead angle of sensitivity in the target sound direction using the output signals of the plurality of microphone elements; and a level comparison determination unit that compares the output level of the output signal from the first directivity synthesis unit with the output level of the output signal from the second directivity synthesis unit to determine whether the target sound arrives from the target sound direction.
[0036] When the target sound arrives from the target sound direction, the output level of the output signal from the first directional synthesis unit is greater than the output level of the output signal from the second directional synthesis unit. Therefore, when the output level of the output signal from the first directional synthesis unit is greater than the output level of the output signal from the second directional synthesis unit, it can be determined that the target sound arrives from the target sound direction. On the other hand, when the target sound does not arrive from the target sound direction, only ambient noise is included in the output signals of the first directional synthesis unit and the second directional synthesis unit. Therefore, the output level of the output signal from the first directional synthesis unit is almost equal to or less than the output level of the output signal from the second directional synthesis unit. Therefore, when the output level of the output signal from the first directional synthesis unit is less than or equal to the output level of the output signal from the second directional synthesis unit, it can be determined that the target sound does not arrive from the target sound direction.
[0037] Moreover, the sound collection device may further be such that the gain control unit includes: a level detection unit that detects the output levels of the output signals of the plurality of microphone elements; a time-average level calculation unit that calculates the time-average levels of the output levels detected by the level detection unit when the voice of the speaker is detected by the target sound detection unit; and a correction gain calculation unit that calculates a correction gain with the updated gain based on the time-average levels calculated by the time-average level calculation unit.
[0038] According to this configuration, when the voice of the speaker is detected, the time-average levels of the output levels of the output signals of the plurality of microphone elements are calculated. Moreover, since the correction gain with the updated gain is calculated based on the calculated time-average levels, the sensitivity difference between the plurality of microphone elements can be corrected by multiplying the output signals of the plurality of microphone elements by the calculated correction gain.
[0039] Moreover, the sound collection device may further be such that the correction gain calculation unit calculates the correction gain of the other microphone elements in such a way that the time-average levels of the other microphone elements except the predetermined one microphone element among the plurality of microphone elements are the same as the time-average level of the one microphone element.
[0040] According to this configuration, the sensitivity difference between the plurality of microphone elements can be corrected in such a way that the output level of the predetermined one microphone element among the plurality of microphone elements is the same as the output levels of the other microphone elements.
[0041] Moreover, the sound collection device may be such that the correction gain calculation unit calculates the correction gain of the plurality of microphone elements based on the average value of the time-averaged levels of at least two predetermined microphone elements among the plurality of microphone elements, in such a way that the time-averaged levels of the plurality of microphone elements are the same as the average value of the time-averaged levels of the at least two microphone elements.
[0042] According to this configuration, the sensitivity difference between the plurality of microphone elements can be corrected by making the average output level of at least two predetermined microphone elements among the plurality of microphone elements the same as the output levels of the plurality of microphone elements.
[0043] Moreover, the sound collection device may be such that the gain control unit includes a third extraction unit that extracts signals in a specific frequency band from the output signals of the plurality of microphone elements, and the level detection unit detects the output levels of the signals extracted by the third extraction unit.
[0044] According to this configuration, since the output levels of the signals in the specific frequency band extracted from the output signals of the plurality of microphone elements are detected, the influence caused by noise other than the target sound can be reduced.
[0045] Moreover, the sound collection device may be such that the specific frequency band is a frequency band from 200 Hz to 500 Hz.
[0046] According to this configuration, the output levels of the signals in the frequency band from 200 Hz to 500 Hz extracted from the output signals of the plurality of microphone elements are detected. Therefore, the influence of low-frequency band noise can be reduced by removing the low-frequency band noise below 200 Hz. Moreover, by removing the frequency band above 500 Hz and limiting the sound to a long-wave sound that is sufficiently longer than the size of the microphone array, the difference in sound pressure caused by the positions of the microphone elements constituting the microphone array can be made smaller. Thus, highly accurate sensitivity correction can be performed.
[0047] A sound collection method according to another embodiment of the present invention causes a computer to perform the following steps: correcting the sensitivity difference between a plurality of microphone elements by multiplying the output signals of the plurality of microphone elements by a gain; detecting the speech of a speaker as a target sound; controlling the gain based on the detection result of the target sound; emphasizing and collecting the target sound arriving from a specified direction using the output signals of the plurality of corrected microphone elements. In the control of the gain, when the speech of the speaker is detected, the gain is updated based on the output signals of the plurality of microphone elements, and when the speech of the speaker is not detected, the gain is not updated.
[0048] According to this configuration, the sensitivity difference between multiple microphone elements is corrected by multiplying the output signals of the multiple microphone elements by a gain. At this time, when the speech of the speaker is detected, the gain is updated based on the output signals of the multiple microphone elements, and when the speech of the speaker is not detected, the gain is not updated. Moreover, using the output signals of the multiple microphone elements whose sensitivity difference has been corrected, the target sound arriving from a specified direction is emphasized and picked up.
[0049] Therefore, when the target sound, that is, the speech of the speaker, is detected, since the gain for correcting the sensitivity difference between the multiple microphone elements is updated, the sensitivity difference between the multiple microphone elements with respect to the target sound can be corrected, and in the subsequent directional synthesis, the leakage amount of the target sound leaking into the noise reference signal having sensitivity in the direction of the target sound can be reduced. As a result, the noise suppression performance in the directional synthesis can be improved and the target sound can be picked up with a higher signal-to-noise ratio.
[0050] A storage medium according to another embodiment of the present invention is a non-transitory computer-readable storage medium storing a sound pickup program, which causes a computer to function as the following components: a sensitivity correction unit that corrects the sensitivity difference between the multiple microphone elements by multiplying the output signals of the multiple microphone elements by a gain; a target sound detection unit that detects the speech of the speaker as the target sound; a gain control unit that controls the gain based on the detection result of the target sound detection unit; and a directional synthesis unit that uses the output signals of the multiple microphone elements corrected by the sensitivity correction unit to emphasize and pick up the target sound arriving from a specified direction, wherein the gain control unit updates the gain based on the output signals of the multiple microphone elements when the speech of the speaker is detected by the target sound detection unit, and does not update the gain when the speech of the speaker is not detected by the target sound detection unit.
[0051] According to this configuration, the sensitivity difference between multiple microphone elements is corrected by multiplying the output signals of the multiple microphone elements by a gain. At this time, when the speech of the speaker is detected, the gain is updated based on the output signals of the multiple microphone elements, and when the speech of the speaker is not detected, the gain is not updated. Moreover, using the output signals of the multiple microphone elements whose sensitivity difference has been corrected, the target sound arriving from a specified direction is emphasized and picked up.
[0052] Therefore, when the target sound, i.e., the speech of the speaker, is detected, since the gain for correcting the sensitivity difference between multiple microphone elements is updated, the sensitivity difference between multiple microphone elements for the target sound can be corrected. In the subsequent directional synthesis stage, the leakage amount of the target sound leaking into the noise reference signal with sensitivity in the dead angle in the target sound direction can be reduced. As a result, the noise suppression performance in the directional synthesis can be improved and the target sound can be received with a higher signal-to-noise ratio.
[0053] Embodiments of the present invention will be described below with reference to the drawings. In addition, the embodiments described below are all specific examples of the present invention and are not used to limit the technical protection scope of the present invention.
[0054] (First Embodiment)
[0055] Figure 1 It is a block diagram showing the configuration of a sound collection device according to the first embodiment of the present invention.
[0056] Figure 1 The illustrated sound collection device 101 includes a microphone array 1, a sensitivity correction unit 2, a target sound detection unit 3, a sensitivity correction control unit (gain control unit) 4, and a directional synthesis unit 5.
[0057] The microphone array 1 includes n (n is a natural number) microphone elements 11, 12,..., 1n that convert sound signals into electrical signals. The microphone array 1 includes multiple microphone elements.
[0058] Figure 2 It is a schematic diagram showing an example of the installation position of the microphone array according to the first embodiment of the present invention. Figure 3 It is a schematic diagram showing an example of the arrangement of the microphone elements of the microphone array according to the first embodiment of the present invention.
[0059] As Figure 2 shown, the microphone array 1 of the first embodiment is arranged near the display 201 in the vehicle. The display 201 is a component of a car navigation system. Moreover, an air outlet 202 of an air conditioner is provided below the display 201. Cooled air or heated air is blown out from the air outlet 202.
[0060] Moreover, Figure 3The microphone array 1 shown, for example, includes four microphone elements 11, 12, 13, and 14. The microphone elements 11, 12, 13, and 14 are respectively arranged at the four corners of a quadrangular substrate. The horizontal interval between the microphone elements 11 and 12 arranged at the lower part of the substrate is, for example, 2 cm. Moreover, the horizontal interval between the microphone elements 13 and 14 arranged at the upper part of the substrate is, for example, 2 cm. In addition, the vertical interval between the microphone elements 11 and 13 is, for example, 2 cm, and the vertical interval between the microphone elements 12 and 14 is, for example, 2 cm.
[0061] The interval between the microphone array 1 and the air outlet 202 is, for example, 2 cm. The microphone array 1 acquires the voice of the speaker sitting in the driver's seat as the target sound. At this time, the sound of the air blown out from the air outlet 202 is included in the target sound as noise. The interval between the microphone element 13 closest to the air outlet 202 and the air outlet 202 is 2 cm, and the interval between the microphone element 11 farthest from the air outlet 202 and the air outlet 202 is 4 cm. The interval between the microphone element 11 and the air outlet 202 is twice the interval between the microphone element 13 and the air outlet 202.
[0062] In this case, the distance difference between the air outlet 202, which is the noise source, and each of the microphone elements 11 and 13 cannot be ignored, and the noise generated from the air outlet 202 appears as a sound pressure difference at the positions of the respective microphone elements 11 and 13. When correcting the sensitivities of the microphone elements 11 to 14 using the noise generated from a noise source near the microphone array 1 like this, the sensitivity correction cannot be performed correctly, and instead, the performance of the output of the subsequent directivity synthesis unit (beamforming) may deteriorate. Therefore, the sound collection device 101 of the first embodiment corrects the sensitivities of the microphone elements 11 to 14 using the target sound.
[0063] In addition, the number of microphone elements included in the microphone array 1 is not limited to four. Moreover, the arrangement positions of the multiple microphone elements are not limited to Figure 3 the arrangement positions shown.
[0064] The output signal of one microphone element 11 among the microphone elements 11, 12,..., 1n is input to the target sound detection unit 3. Moreover, the output signals of the respective microphone elements 11, 12,..., 1n are input to the sensitivity correction unit 2 and the sensitivity correction control unit 4.
[0065] The sensitivity correction unit 2 corrects the sensitivity difference between the plurality of microphone elements 11, 12, ……, 1n by multiplying the gain by the output signals of the plurality of microphone elements 11, 12, ……, 1n. The sensitivity correction unit 2 corrects the difference in sensitivity of each of the microphone elements 11, 12, ……, 1n by multiplying the specified gain by the output signals of the respective microphone elements 11, 12, ……, 1n. The sensitivity correction unit 2 makes the sensitivities of the plurality of microphone elements 11, 12, ……, 1n consistent.
[0066] The target sound detection unit 3 detects the speech of the speaker as the target sound. The target sound detection unit 3 acquires the output signal of one of the microphone elements 11 among the microphone elements 11, 12, ……, 1n, and detects whether there is a target sound received by the microphone array 1. In addition, in the first embodiment, the target sound detection unit 3 uses the output signal of the microphone element 11 to detect whether there is a target sound, but the present invention is not particularly limited thereto. The target sound detection unit 3 may also use the output signal of any one of the microphone elements 11, 12, ……, 1n to detect whether there is a target sound.
[0067] In addition, use Figure 4 and Figure 5 to describe the configuration of the target sound detection unit 3 in detail.
[0068] The sensitivity correction control unit 4 controls the gain based on the detection result of the target sound detection unit 3. When the sensitivity correction control unit 4 acquires the output signals of the respective microphone elements 11, 12, ……, 1n and detects a target sound through the target sound detection unit 3, it calculates the sensitivity correction gain for the output signals from the respective microphone elements 11, 12, ……, 1n in the sensitivity correction unit 2.
[0069] When the sensitivity correction control unit 4 detects the speech of the speaker through the target sound detection unit 3, it updates the gain based on the output signals of the plurality of microphone elements. When the speech of the speaker is not detected through the target sound detection unit 3, it does not update the gain. In addition, use Figure 6 to describe the configuration of the sensitivity correction control unit 4 in detail.
[0070] The directivity synthesis unit (beamforming) 5 uses the output signals of the plurality of microphone elements corrected by the sensitivity correction unit 2 to perform sound collection in a manner that emphasizes the target sound coming from a specified direction. The directivity synthesis unit 5 acquires the output signals of the respective microphone elements 11, 12, ……, 1n corrected by the sensitivity correction unit 2, and improves the S / N ratio (signal-to-noise ratio) of the target sound.
[0071] Next, for Figure 1The configuration of the target sound detection unit 3 shown will be further described.
[0072] Figure 4 It is a block diagram showing the configuration of the target sound detection unit of the sound collection device according to the first embodiment of the present invention.
[0073] Figure 4 The target sound detection unit 3 shown includes a band-pass filter unit (first extraction unit) 31 and a voice determination unit 32.
[0074] The band-pass filter unit 31 extracts a signal of a specific frequency band from the output signal of one microphone element 11 among the plurality of microphone elements 11, 12, ……, 1n. The band-pass filter unit 31 extracts, for example, a signal in the frequency band from 200 Hz to 500 Hz from the output signal of the microphone element 11. The band-pass filter unit 31 extracts a signal in a frequency band capable of extracting the voice of a person speaking from the output signal of the microphone element 11.
[0075] The voice determination unit 32 determines whether the output signal of one microphone element 11 among the plurality of microphone elements 11, 12, ……, 1n is voice or non-voice other than voice. The voice determination unit 32 determines whether the signal extracted by the band-pass filter unit 31 is voice or non-voice.
[0076] Next, Figure 4 The configuration of the voice determination unit 32 shown will be further described.
[0077] Figure 5 It is a block diagram showing the configuration of the voice determination unit of the sound collection device according to the first embodiment of the present invention.
[0078] The voice determination unit 32 includes a level detection unit 321, a noise level detection unit 322, a comparison unit 323, a time-frequency transformation unit 324, a voice feature amount extraction unit 325, and a determination unit 326.
[0079] The level detection unit 321 detects the signal level of the output signal of the microphone element 11.
[0080] The noise level detection unit 322 detects the noise level by maintaining the minimum value of the signal level detected by the level detection unit 321.
[0081] A comparison unit 323 compares the output of the level detection unit 321 with the output of the noise level detection unit 322 to determine whether there is speech in terms of waveform degree. For example, the comparison unit 323 sets a value twice the noise level detected by the noise level detection unit 322 as a threshold. Then, the comparison unit 323 determines whether the signal level detected by the level detection unit 321 is above the threshold. When the signal level detected by the level detection unit 321 is above the threshold, the comparison unit 323 determines that the output signal of the microphone element 11 contains speech. On the other hand, when the signal level detected by the level detection unit 321 is less than the threshold, the comparison unit 323 determines that the output signal of the microphone element 11 does not contain speech.
[0082] A time-frequency transformation unit 324 transforms the output signal of the microphone element 11 in the time domain into an output signal in the frequency domain.
[0083] A speech feature quantity extraction unit 325 extracts speech feature quantities from the output signal in the frequency domain. The speech feature quantity is a quantity representing the features of speech. The speech feature quantity extraction unit 325 may also adopt the method of extracting speech feature quantities using voice pitch disclosed in the specification of Japanese Patent No. 5450298 or the method of extracting speech feature quantities using the properties of the harmonic structure as feature quantities disclosed in the specification of Japanese Patent No. 3849116. When the radio receiving device 101 is mounted on a vehicle, as Figure 2 shown, the microphone array 1 is assembled around a display 201 embedded in a console. For this reason, the air outlet 202 of the air conditioner becomes a noise source. In this case, since the spectrum of the noise is relatively monotonous, the speech feature quantity extraction unit 325 may also extract the AC component of the amplitude spectrum or the ratio of the peak to the trough of the amplitude spectrum as the speech feature quantity. Thereby, the noise generated from the air outlet 202 of the air conditioner and the speech can be distinguished.
[0084] When the comparison unit 323 determines that the output signal of the microphone element 11 contains speech and the speech feature quantity extraction unit 325 extracts speech feature quantities from the output signal of the microphone element 11, a determination unit 326 determines that the output signal of the microphone element 11 is speech. On the other hand, when the comparison unit 323 determines that the output signal of the microphone element 11 does not contain speech, or when the speech feature quantity extraction unit 325 does not extract speech feature quantities from the output signal of the microphone element 11, the determination unit 326 determines that the output signal of the microphone element 11 is non-speech. The determination unit 326 outputs a determination result signal Odet(j) indicating whether it is speech or non-speech to the sensitivity correction control unit 4. Here, j represents the sampling number corresponding to time.
[0085] As a result, when the target sound detection unit 3 determines that the output signal of the microphone element 11 is speech, it outputs a determination result signal Odet(j) = 1, and when it determines that the output signal of the microphone element 11 is non-speech, it outputs a determination result signal Odet(j) = 0.
[0086] Next, Figure 1 the configuration of the sensitivity correction control unit 4 shown in
[0087] Figure 6 is further described.
[0088] The sensitivity correction control unit 4 includes first to nth band-pass filter units (third extraction units) 411 to 41n, first to nth level detection units 421 to 42n, first to nth average level calculation units (time average level calculation units) 431 to 43n, and a correction gain calculation unit 44. The first to nth band-pass filter units 411 to 41n, the first to nth level detection units 421 to 42n, and the first to nth average level calculation units 431 to 43n are respectively provided according to the number of microphone elements 11 to 1n. For example, the output signal x(1, j) of the microphone element 11 is input to the first band-pass filter unit 411.
[0089] The first to nth band-pass filter units 411 to 41n extract signals in a specific frequency band from the output signals of the respective microphone elements 11 to 1n. Here, the specific frequency band is a frequency band from 200 Hz to 500 Hz.
[0090] The first to nth level detection units 421 to 42n detect the output levels of the output signals of the respective microphone elements 11 to 1n.
[0091] The first to nth level detection units 421 to 42n detect the output level Lx(i, j) of the output signal x(i, j) of each microphone element using the following general amplitude smoothing formula (1).
[0092] Lx(i, j) = beta1 · |x(i, j)| + (1 - beta1) · Lx(i, j - 1)......(1)
[0093] In formula (1), i represents the microphone element number, and j represents the sampling number corresponding to time. Moreover, in formula (1), beta1 is a parameter representing a weighting coefficient and determining the averaging speed.
[0094] Moreover, in the first embodiment, the output signals xbpf(i, j) that have passed through the first to nth band-pass filter units 411 to 41n are input to the first to nth level detection units 421 to 42n. For this purpose, the first to nth level detection units 421 to 42n detect the output level Lx(i, j) of the output signal xbpf(i, j) of each microphone element extracted by the first to nth band-pass filter units 411 to 41n using the following general amplitude smoothing formula (2).
[0095] Lx(i, j) = beta1 · |xbp(i, j)| + (1 - beta1) · Lx(i, j - 1)…(2)
[0096] When the target sound detection unit 3 detects the speech of the speaker, the first to nth average level calculation units 431 to 43n calculate the time-average level Avex(i, j) of each output level Lx(i, j) detected by the first to nth level detection units 421 to 42n.
[0097] The first to nth average level calculation units 431 to 43n calculate the long-term average value (time-average level Avex(i, j)) of the output levels Lx(i, j) of the respective microphone elements using the following formula (3) only during the period when the target sound detection unit 3 detects the target sound (judgment result signal Odet(j) = 1). Moreover, the first to nth average level calculation units 431 to 43n calculate the time-average level Avex(i, j) using the following formula (4) during the period when the target sound detection unit 3 does not detect the target sound (judgment result signal Odet(j) = 0). That is, when the target sound detection unit 3 does not detect the speech of the speaker, the first to nth average level calculation units 431 to 43n calculate the previous calculated time-average level Avex(i, j - 1) as the current time-average level Avex(i, j).
[0098] Avex(i, j) = beta2 · |Lx(i, j)| + (1 - beta2) · Avex(i, j - 1)
[0099] if Odet(j) = 1……(3)
[0100] Avex(i, j) = Avex(i, j - 1) if Odet(j) = 0 …… (4)
[0101] In Formula (3) and Formula (4), i represents the microphone element number, and j represents the sampling number corresponding to time. Moreover, in Formula (3), beta2 is a weighting coefficient, which is a parameter determining the averaging speed. And beta1 >> beta2. For example, when the sampling frequency is 16 kHz, beta1 is set to 0.000625 to obtain the average level in 100 milliseconds, and beta2 is set to 0.0000125 to obtain the average level in 5 seconds. By using the average signal level over a long time, the average signal level used for the sensitivity correction of the microphone element can correctly calculate the sensitivity correction gain.
[0102] The correction gain calculation unit 44 calculates the sensitivity correction gain with the updated gain based on the time-average levels calculated by the first to n average level calculation units 431 to 43n.
[0103] The correction gain calculation unit 44 calculates the sensitivity correction gains of the other microphone elements 12 to 1n other than one microphone element 11 among the plurality of microphone elements 11 to 1n so that the time-average levels of the other microphone elements 12 to 1n are the same as the time-average level of one microphone element 11. That is, the correction gain calculation unit 44 calculates the sensitivity correction gain G(i, j) through the following Formula (5) by using the time-average levels Avex(i, j) of each of the microphone elements 11 to 1n calculated by the first to n average level calculation units 431 to 43n and the time-average level Avex(1, j) of the microphone element 11.
[0104] G(i, j) = Avex(1, j) / Avex(i, j) …… (5)
[0105] In the case of the sensitivity correction gain using the above Formula (5), the sensitivity correction is performed to make the output levels of the other microphone elements 12 to 1n consistent with the reference of the microphone element 11.
[0106] In addition, in the above Formula (5), the correction gain calculation unit 44 calculates the sensitivity correction gain based on the time-average level of a predetermined one microphone element 11. However, the present invention is not particularly limited thereto. The correction gain calculation unit 44 may also calculate the sensitivity correction gain based on the time-average level of another microphone element different from the microphone element 11.
[0107] Furthermore, the calibration gain calculation unit 44 may also calculate the sensitivity calibration gains of the plurality of microphone elements 11 to 1n in such a way that the time-averaged levels of the plurality of microphone elements 11 to 1n are the same as the average of the time-averaged levels of at least two predetermined microphone elements among the plurality of microphone elements 11 to 1n. That is, the calibration gain calculation unit 44 may also calculate the sensitivity calibration gain G(i, j) by using the time-averaged level Avex(i, j) of each of the microphone elements 11 to 1n calculated by the first to n average level calculation units 431 to 43n and the average of the time-averaged levels Avex(i, j) according to the following formula (6).
[0108] G(i, j) = {Avex(1, j) + Avex(2, j) + … + Avex(n, j)} / n / Avex(i, j) ……(6)
[0109] In addition, in the above formula (6), the calibration gain calculation unit 44 calculates the sensitivity calibration gain based on the average of the time-averaged levels of all the microphone elements 11 to 1n among the microphone elements 11 to 1n. However, the present invention is not particularly limited thereto. The calibration gain calculation unit 44 may also calculate the sensitivity calibration gain based on the average of the time-averaged levels of at least two microphone elements among the microphone elements 11 to 1n.
[0110] The sensitivity calibration unit 2 performs sensitivity calibration by multiplying the sensitivity calibration gain G(i, j) corresponding to each of the microphone elements 11 to 1n calculated by the sensitivity calibration control unit 4 by the output signal x(i, j) of each of the microphone elements 11 to 1n.
[0111] The directivity synthesis unit 5 performs directivity synthesis (beamforming) by GSC shown in the specification of Japanese Patent No. 4734070 using the output signal G(i, j)·x(i, j) corrected by the sensitivity calibration unit 2. Moreover, the directivity synthesis unit 5 may also perform beamforming by existing beamforming processes other than GSC, such as the Maximum Likelihood method or the Minimum Variance method.
[0112] Next, the operation of the sound collecting device 101 according to the first embodiment of the present invention will be described.
[0113] Figure 7 It is a flowchart for explaining the operation of the sound collecting device according to the first embodiment of the present invention.
[0114] First, in step S1, the target sound detection unit 3 obtains an output signal from the microphone element 11, and the sensitivity correction unit 2 and the sensitivity correction control unit 4 obtain output signals from each of the microphone elements 11 to 1n.
[0115] Next, in step S2, the target sound detection unit 3 determines whether a target sound (voice) has been detected from the output signal of the microphone element 11. The target sound detection unit 3 outputs a determination result signal indicating whether a target sound has been detected from the output signal of the microphone element 11 to the sensitivity correction control unit 4.
[0116] Here, in the case where it is determined that a target sound has been detected from the output signal of the microphone element 11 (Yes in step S2), in step S3, the sensitivity correction control unit 4 updates the sensitivity correction gain based on the output signals of the plurality of microphone elements 11 to 1n.
[0117] On the other hand, in the case where it is determined that a target sound has not been detected from the output signal of the microphone element 11 (No in step S2), the sensitivity correction gain is not updated, and the process proceeds to step S4.
[0118] Next, in step S4, the sensitivity correction unit 2 corrects the sensitivity difference between the respective microphone elements by multiplying the output signals of the respective microphone elements 11 to 1n by the sensitivity correction gain.
[0119] Next, in step S5, the directivity synthesis unit 5 synthesizes the directivity using the output signals of the respective microphone elements 11 to 1n corrected by the sensitivity correction unit 2. By synthesizing the directivity, it is possible to emphasize and pick up the target sound coming from a specified direction.
[0120] As described above, the sensitivity difference between the plurality of microphone elements 11 to 1n is corrected by multiplying the output signals of the plurality of microphone elements 11 to 1n by a gain. At this time, in the case where the speech of the speaker is detected, the gain is updated based on the output signals of the plurality of microphone elements 11 to 1n, and in the case where the speech of the speaker is not detected, the gain is not updated. Moreover, the output signals of the plurality of microphone elements 11 to 1n whose sensitivity difference has been corrected are used to emphasize and pick up the target sound coming from a specified direction.
[0121] Therefore, when the target sound, i.e., the speech of the speaker, is detected, since the gain for correcting the sensitivity difference between the plurality of microphone elements 11 to 1n is updated, the sensitivity difference for the target sound between the plurality of microphone elements 11 to 1n can be corrected. In the subsequent directivity synthesis, the leakage amount of the target sound leaking into the noise reference signal having a dead angle of sensitivity in the target sound direction can be reduced. As a result, the noise suppression performance in the directivity synthesis can be improved and the target sound can be received with a higher signal-to-noise ratio.
[0122] (Second Embodiment)
[0123] In the above-described first embodiment, the target sound detection unit 3 determines whether the output signal of one microphone element is speech or non-speech. Correspondingly, in the second embodiment, the target sound detection unit further determines whether the target sound arrives from the target sound direction predetermined using the output signals of the plurality of microphone elements.
[0124] Figure 8 It is a block diagram showing the configuration of the sound receiving device according to the second embodiment of the present invention.
[0125] Figure 8 The shown sound receiving device 102 includes a microphone array 1, a sensitivity correction unit 2, a sensitivity correction control unit 4, a directivity synthesis unit 5, and a target sound detection unit 6. The difference from the sound receiving device 101 of the first embodiment is that the output signals from the plurality of microphone elements 11, 12,..., 1n are input to the target sound detection unit 6. In addition, in the second embodiment, the same components as those in the first embodiment are given the same reference numerals and their descriptions are omitted.
[0126] Figure 9 It is a block diagram showing the configuration of the target sound detection unit of the sound receiving device according to the second embodiment of the present invention.
[0127] Figure 9 The shown target sound detection unit 6 includes a band-pass filter unit 31, a speech determination unit 32, a band-pass filter unit (second extraction unit) 63, a target sound direction determination unit 64, and a target sound determination unit 65. Compared with the target sound detection unit 3 of the first embodiment, the band-pass filter unit 63, the target sound direction determination unit 64, and the target sound determination unit 65 are added to the target sound detection unit 6 in the second embodiment.
[0128] The band-pass filter unit 63 extracts signals in a specific frequency band from the output signals of the plurality of microphone elements. The band-pass filter unit 63 extracts signals in a frequency band from, for example, 200 Hz to 500 Hz from the output signals of each of the microphone elements 11 to 1n.
[0129] The target sound direction determination unit 64 determines whether the target sound arrives from a predetermined target sound direction using the output signals of a plurality of microphone elements. The target sound direction determination unit 64 determines whether the target sound arrives from the target sound direction for the signal extracted by the band-pass filter unit 63. Here, when the sound collecting device 102 in the vehicle collects the speech of the driver, the angle at which the driver's speech enters the microphone array 1 is predetermined. For this reason, the target sound direction determination unit 64 stores in advance the incident angle of the speech. In addition, using Figure 10 and Figure 11 the configuration of the target sound direction determination unit 64 will be described in more detail.
[0130] The target sound determination unit 65 determines whether there is a target sound using the two determination results of the speech determination unit 32 and the target sound direction determination unit 64. The target sound determination unit 65 determines that a target sound has been detected when it is determined by the target sound direction determination unit 64 that the target sound arrives from the target sound direction and it is determined by the speech determination unit 32 that the output signal of one microphone element is speech. Moreover, the target sound determination unit 65 determines that no target sound has been detected when it is determined by the target sound direction determination unit 64 that the target sound does not arrive from the target sound direction, or when it is determined by the speech determination unit 32 that the output signal of one microphone element is not speech.
[0131] Next, Figure 9 the configuration of the target sound direction determination unit 64 shown will be further described.
[0132] Figure 10 is a block diagram showing the configuration of the target sound direction determination unit of the sound collecting device according to the second embodiment of the present invention. In addition, in Figure 10 for ease of explanation, an example in which the output signals from two microphone elements 11 and 12 are input to the target sound direction determination unit 64 will be described.
[0133] The target sound direction determination unit 64 includes a delay and directivity synthesis unit (delay and beamformer) (first directivity synthesis unit) 641, an inclined directivity synthesis unit (inclined beamformer) (second directivity synthesis unit) 642, a target sound level detection unit 643, a non-target sound level detection unit 644, and a level comparison determination unit 645.
[0134] The delay and directivity synthesis unit 641 forms directivity in the target sound direction by emphasizing the signal in the target sound direction using the output signals of a plurality of microphone elements 11 to 1n. The delay and directivity synthesis unit 641 has a high directivity sensitivity in the target sound direction. Figure 10The shown directivity characteristic 6411 represents the directivity characteristic of the delay and directivity synthesis unit 641. The directivity characteristic 6411 of the delay and directivity synthesis unit 641 has directivity in the target sound direction and emphasizes the signal in the target sound direction.
[0135] For the delay and directivity synthesis unit 641, if the distance between the microphone element 11 and the microphone element 12 is d and the incident angle from the target sound direction is θ, it delays the output signal from the microphone element 11 by the path difference Δ (Δ = dsinθ). Moreover, the delay and directivity synthesis unit 641 adds the delayed output signal from the microphone element 11 and the output signal from the microphone element 12. Additionally, the distance d and the incident angle θ are stored in a memory (not shown) in advance.
[0136] The inclined directivity synthesis unit 642 forms a dead angle of sensitivity in the target sound direction using the output signals of the plurality of microphone elements 11 and 12. Figure 10 The shown directivity characteristic 6421 represents the directivity characteristic of the inclined directivity synthesis unit 642. The directivity characteristic 6421 of the inclined directivity synthesis unit 642 has a dead angle in the target sound direction and emphasizes the signal (noise) in the direction perpendicular to the target sound direction.
[0137] For the inclined directivity synthesis unit 642, if the distance between the microphone element 11 and the microphone element 12 is d and the incident angle of the sound from the target sound direction is θ, it delays the output signal from the microphone element 11 by the path difference Δ (Δ = dsinθ). Moreover, the inclined directivity synthesis unit 642 subtracts the output signal from the microphone element 12 from the delayed output signal from the microphone element 11. Additionally, the distance d and the incident angle θ are stored in advance.
[0138] The target sound level detection unit 643 detects the output signal level of the delay and directivity synthesis unit 641.
[0139] The non-target sound level detection unit 644 detects the output signal level of the inclined directivity synthesis unit 642.
[0140] The level comparison and determination unit 645 compares the output level of the output signal from the delay and directivity synthesis unit 641 with the output level of the output signal from the inclined directivity synthesis unit 642 to determine whether the target sound arrives from the target sound direction. The level comparison and determination unit 645 compares the output signal level detected by the target sound level detection unit 643 with the output signal level detected by the non-target sound level detection unit 644 to determine whether the target sound arrives from the target sound direction.
[0141] The delay and directivity synthesis unit 641 has directivity in the target sound direction. For this reason, the speech of the speaker who is the target sound is included in the output of the delay and directivity synthesis unit 641. On the other hand, the inclined directivity synthesis unit 642 has a dead angle in the target sound direction. For this reason, the speech of the speaker who is the target sound is hardly included in the output of the inclined directivity synthesis unit 642. Therefore, when the target sound arrives from the target sound direction, the output signal level detected by the target sound level detection unit 643 becomes large, and the output signal level detected by the non-target sound level detection unit 644 becomes small. The level comparison and determination unit 645 determines that the target sound arrives from the target sound direction when the output signal level (target sound level) detected by the target sound level detection unit 643 is greater than the output signal level (non-target sound level) detected by the non-target sound level detection unit 644.
[0142] On the other hand, when the target sound does not arrive from the target sound direction, the outputs of the delay and directivity synthesis unit 641 and the inclined directivity synthesis unit 642 only include ambient noise. Therefore, the output signal level detected by the target sound level detection unit 643 is approximately equal to or less than the output signal level detected by the non-target sound level detection unit 644. The level comparison and determination unit 645 determines that the target sound does not arrive from the target sound direction when the output signal level (target sound level) detected by the target sound level detection unit 643 is less than or equal to the output signal level (non-target sound level) detected by the non-target sound level detection unit 644.
[0143] In the first embodiment, since it is determined that the target sound is detected if speech is detected, even when there is a voice from a direction other than the target sound direction, it is determined that the target sound is detected and sensitivity correction is performed. However, in the second embodiment, it is determined that the target sound is detected only when speech is detected and the target sound arrives from the target sound direction. Therefore, the sound collecting device 102 of the second embodiment can perform sensitivity correction using a target sound with higher accuracy than the sound collecting device 101 of the first embodiment.
[0144] Next, the configuration of the target sound direction determination unit in the modification of the second embodiment will be further described.
[0145] Figure 11 It is a block diagram showing the configuration of the target sound direction determination unit of the sound collecting device in the modification of the second embodiment of the present invention. In addition, in Figure 11In the following, for the sake of convenience of explanation, an example in which output signals from two microphone elements 11 and 12 are input to the target sound direction determination unit 64A will be described. Moreover, Figure 9 the target sound detection unit 6 shown Figure 11 may also be provided with Figure 9 the target sound direction determination unit 64A shown
[0146] instead of the target sound direction determination unit 64 shown
[0147] The target sound direction determination unit 64A includes a target sound direction estimation unit (direction estimation unit) 646 and a direction determination unit 647.
[0148] The target sound direction estimation unit 646 estimates the direction from which the target sound arrives by using the phase difference of the output signals of a plurality of microphone elements. A memory (not shown) stores in advance the distance d between the microphone element 11 and the microphone element 12. The target sound direction estimation unit 646 estimates the incident angle θ of the sound from the target sound direction based on the phase difference between the microphone element 11 and the microphone element 12 and the distance d between the microphone element 11 and the microphone element 12.
[0149] In addition, in each of the above-described embodiments, each component may be configured by dedicated hardware or may be implemented by executing a software program suitable for each component. Each component may also be implemented by having a program execution unit such as a CPU or a processor read a software program recorded in a recording medium such as a hard disk or a semiconductor memory.
[0150] Part or all of the functions of the device according to the embodiments of the present invention can typically be implemented as an integrated circuit LSI (Large Scale Integration). Part or all of these functions can be separately chip-formed or formed to include part or all of the chip-formation. Moreover, the integrated circuit is not limited to LSI and can also be implemented using an application-specific circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after manufacturing the LSI or a reconfigurable processor that can reconstruct the connection or setting of circuit units inside the LSI can also be used.
[0151] Moreover, part or all of the functions of the device according to the embodiments of the present invention can also be implemented by having a processor such as a CPU execute a program.
[0152] Moreover, the numbers used above are examples given for specifically illustrating the present invention, and the present invention is not limited to these exemplified numbers.
[0153] Moreover, the order in which the steps shown in the above flowchart are executed is only an example given for specifically illustrating the present invention, and it can also be an order other than the above within the range where the same effect can be obtained. Moreover, a part of the above steps can also be executed simultaneously (in parallel) with other steps.
[0154] The technology related to the present invention has practical value as a technology for receiving a target sound using a plurality of microphone elements because it can improve the noise suppression performance in directional synthesis and can receive the target sound with a high signal-to-noise ratio.
Claims
1. A radio receiving device, characterized in that, Comprising: A plurality of microphone elements; A sensitivity correction unit that corrects the sensitivity difference between the plurality of microphone elements by multiplying the output signals of the plurality of microphone elements by a gain; A target sound detection unit that detects the speech of a speaker as a target sound; A gain control unit that controls the gain based on the detection result of the target sound detection unit; And A directivity synthesis unit that emphasizes and picks up the target sound coming from a specified direction by using the output signals of the plurality of microphone elements corrected by the sensitivity correction unit, wherein The gain control unit, when the speech of the speaker is detected by the target sound detection unit, updates the gain based on the output signals of the plurality of microphone elements, and when the speech of the speaker is not detected by the target sound detection unit, does not update the gain, The gain control unit includes: A level detection unit that detects the output levels of the output signals of the plurality of microphone elements respectively; A time-averaged level calculation unit that calculates the time-averaged levels of the respective output levels detected by the level detection unit when the speech of the speaker is detected by the target sound detection unit; and A corrected gain calculation unit that calculates a corrected gain with the gain updated based on the time-averaged levels calculated by the time-averaged level calculation unit.
2. The sound pickup device according to claim 1, wherein The target sound detection unit includes a voice determination unit that determines whether the output signal of one of the plurality of microphone elements is the voice or non-voice other than the voice.
3. The sound pickup device according to claim 2, wherein The target sound detection unit includes a first extraction unit that extracts a signal in a specific frequency band from the output signal of the one microphone element, The voice determination unit determines whether the signal extracted by the first extraction unit is the voice or the non-voice.
4. The radio device according to claim 2, characterized in that, The target sound detection unit includes: A target sound direction determination unit that determines whether the target sound comes from a predetermined target sound direction by using the output signals of the plurality of microphone elements; And A target sound determination unit that determines that the target sound is detected when it is determined by the target sound direction determination unit that the target sound comes from the target sound direction and it is determined by the voice determination unit that the output signal of the one microphone element is the voice.
5. The sound pickup device according to claim 4, wherein The target sound detection unit includes a second extraction unit that extracts a signal in a specific frequency band from the output signals of the plurality of microphone elements, The target sound direction determination unit determines whether the target sound comes from the target sound direction for the signal extracted by the second extraction unit.
6. The sound receiving device according to claim 4 or 5, characterized in that, The target sound direction determination unit includes: A direction estimation unit that estimates the direction of arrival of the target sound by using the phase difference of the output signals of the plurality of microphone elements; And A direction determination unit that determines whether the direction estimated by the direction estimation unit is the predetermined target sound direction.
7. The sound receiving device according to claim 4 or 5, characterized in that, The target sound direction determination unit includes: A first directivity synthesis unit that forms a directivity in the target sound direction by emphasizing a signal in the target sound direction using output signals of the plurality of microphone elements; A second directivity synthesis unit that forms a dead angle of sensitivity in the target sound direction using output signals of the plurality of microphone elements; And A level comparison and determination unit that compares the output level of the output signal from the first directivity synthesis unit with the output level of the output signal from the second directivity synthesis unit, and determines whether the target sound arrives from the target sound direction.
8. The sound collection device according to any one of claims 1 to 4, characterized in that The correction gain calculation unit calculates the correction gain of the other microphone elements in such a manner that the time-average level of the other microphone elements becomes the same as the time-average level of a predetermined one of the plurality of microphone elements, based on the time-average level of the predetermined one microphone element among the plurality of microphone elements.
9. The sound collection device according to any one of claims 1 to 4, characterized in that The correction gain calculation unit calculates the correction gain of the plurality of microphone elements in such a manner that the time-average levels of the plurality of microphone elements become the same as the average of the time-average levels of at least two predetermined microphone elements among the plurality of microphone elements, based on the average of the time-average levels of the at least two predetermined microphone elements among the plurality of microphone elements.
10. The sound collection device according to any one of claims 1 to 4, characterized in that The gain control unit includes a third extraction unit that extracts signals in a specific frequency band from the output signals of the plurality of microphone elements respectively, A level detection unit that detects the output levels of the signals extracted by the third extraction unit.
11. The sound collection device according to claim 10, characterized in that The specific frequency band is a frequency band from 200 Hz to 500 Hz.
12. A sound receiving method, characterized in that, Cause a computer to perform the following steps: Correct a sensitivity difference between the plurality of microphone elements by multiplying output signals of the plurality of microphone elements by a gain; Detect speech of a speaker as a target sound; Control the gain based on a detection result of the target sound; Use the output signals of the plurality of microphone elements that have been corrected to emphasize and collect the target sound arriving from a specified direction, In the control of the gain, when the speech of the speaker is detected, update the gain based on the output signals of the plurality of microphone elements, and when the speech of the speaker is not detected, do not update the gain; In the control of the gain, Detect output levels of the output signals of the plurality of microphone elements respectively, When the speech of the speaker is detected, calculate a time-average level of each detected output level. A correction gain with the updated gain is calculated based on the calculated time-average level.
13. A storage medium is a non-transitory computer-readable storage medium storing a radio receiving program, characterized in that, The computer functions as the following components: A sensitivity correction unit that corrects the sensitivity difference between the plurality of microphone elements by multiplying the output signals of the plurality of microphone elements by a gain; A target sound detection unit that detects the speech of a speaker as a target sound; A gain control unit that controls the gain based on the detection result of the target sound detection unit; And A directivity synthesis unit that emphasizes and picks up the target sound arriving from a specified direction by using the output signals of the plurality of microphone elements corrected by the sensitivity correction unit, wherein the gain control unit updates the gain based on the output signals of the plurality of microphone elements when the speech of the speaker is detected by the target sound detection unit, and does not update the gain when the speech of the speaker is not detected by the target sound detection unit, the gain control unit includes: A level detection unit that detects the output levels of the output signals of the respective microphone elements; A time-average level calculation unit that calculates the time-average level of each output level detected by the level detection unit when the speech of the speaker is detected by the target sound detection unit; and A correction gain calculation unit that calculates a correction gain with the updated gain based on the time-average level calculated by the time-average level calculation unit.
Citation Information
Patent Citations
JP1972034070U
Ultrasonic vehicle detector
JP1979050298A
Amplifying conversation method, device and program
JP2011172081A
Pickup signal processing apparatus, method, and program product
US20110313763A1