Signal processing apparatus and signal processing method
The signal processing apparatus addresses the challenge of suppressing non-target voice noise in in-vehicle systems by using delayed signals from multiple microphones to estimate and eliminate crosstalk noise, thereby enhancing voice recognition accuracy.
Patent Information
- Application Number
- JP2024002391
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2039-03-06
AI Technical Summary
Existing in-vehicle sound collection technologies struggle to suppress noise from non-target voices without inadvertently suppressing the target voice, leading to poor noise cancellation performance.
A signal processing apparatus and method that utilize multiple microphones installed in different vehicle seats, with delayed signals from these microphones used to estimate and eliminate crosstalk noise, thereby isolating and preserving the target voice.
Effectively suppresses interference noise from non-target voices without suppressing the target voice, thereby improving the accuracy of voice recognition and reducing unwanted noise in in-vehicle infotainment systems.
Smart Images

Figure 0007692069000001 
Figure 0007692069000002 
Figure 0007692069000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a signal processing apparatus and a signal processing method.
Background Art
[0002] As in-vehicle equipment, for the purpose of acquiring the voices of drivers etc. inside the vehicle, a sound collection device is mainly provided at the driver's seat etc. As a result, when the driver etc. use in-vehicle infotainment etc., they can perform operations etc. by voice. For example, Patent Document 1 discloses an in-vehicle sound collection device and a sound collection method that can estimate and suppress noise such as voices of people other than the driver that are mixed into a target voice such as the voice of the driver etc.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the technology disclosed in Patent Document 1, when a part of the target voice is erased, the target voice may be suppressed.
[0005] An object of the present disclosure is to suppress noise mixed into a target voice and to prevent suppression of the target voice.
Means for Solving the Problems
[0006] For example, a signal processing apparatus according to the present disclosure includes a first acquisition unit that acquires a first signal output from a first microphone installed in a driver's seat of a vehicle, a second acquisition unit that acquires a second signal output from a second microphone installed in a passenger seat of the vehicle, a third acquisition unit that acquires a third signal output from a third microphone installed in a rear seat of the vehicle, a first delay unit that delays the second signal with a delay time within 2 msec, a second delay unit that delays the third signal with the same delay time as the delay time by which the first delay unit delays the second signal, a crosstalk noise estimation unit that estimates first noise mixed into the first signal based on the second signal delayed by the first delay unit and the third signal delayed by the second delay unit, and an elimination unit that eliminates the first noise estimated by the crosstalk noise estimation unit from the first signal.
[0007] Also, for example, the signal processing apparatus according to the present disclosure includes a first acquisition unit that acquires a first signal output from a first microphone installed in the driver's seat of a vehicle, a second acquisition unit that acquires a second signal output from a second microphone installed in the passenger seat of the vehicle, a third acquisition unit that acquires a third signal output from a third microphone installed in the rear seat of the vehicle, a first delay unit that delays the first signal with a delay time within 2 msec, a second delay unit that delays the second signal with the same delay time as the delay time by which the first delay unit delays the first signal, a third delay unit that delays the third signal with the same delay time, a noise mixing estimation unit that estimates a first noise mixed into the first signal based on the second signal delayed by the second delay unit and the third signal delayed by the third delay unit, estimates a second noise mixed into the second signal based on the first signal delayed by the first delay unit and the third signal delayed by the third delay unit, and estimates a third noise mixed into the third signal based on the first signal delayed by the first delay unit and the second signal delayed by the second delay unit, and an elimination unit that eliminates the first noise estimated by the noise mixing estimation unit from the first signal, eliminates the second noise estimated by the noise mixing estimation unit from the second signal, and eliminates the third noise estimated by the noise mixing estimation unit from the third signal.
[0008] Also, for example, the signal processing apparatus according to the present disclosure includes a first acquisition unit that acquires a first signal output from a first microphone, a second acquisition unit that acquires a second signal output from a second microphone installed at a position different from the first microphone, and a third acquisition unit that acquires a third signal output from a third microphone installed at a position different from the first microphone and the second microphone. 3An acquisition unit, a first delay unit that delays the first signal, a second delay unit that delays the second signal by the same delay time as the delay time by which the first delay unit delays the first signal, a third delay unit that delays the third signal by the same delay time, a noise mixing estimation unit that estimates first noise mixed into the first signal based on the second signal delayed by the second delay unit and the third signal delayed by the third delay unit, estimates second noise mixed into the second signal based on the first signal delayed by the first delay unit and the third signal delayed by the third delay unit, and estimates third noise mixed into the third signal based on the first signal delayed by the first delay unit and the second signal delayed by the second delay unit, and an elimination unit that eliminates the first noise estimated by the noise mixing estimation unit from the first signal, eliminates the second noise estimated by the noise mixing estimation unit from the second signal, and eliminates the third noise estimated by the noise mixing estimation unit from the third signal.
[0009] Also, for example, the signal processing method according to the present disclosure includes a first acquisition step of acquiring a first signal output from a first microphone installed in a driver's seat of a vehicle, a second acquisition step of acquiring a second signal output from a second microphone installed in a passenger seat of the vehicle, a third acquisition step of acquiring a third signal output from a third microphone installed in a rear seat of the vehicle, a first delay step of delaying the second signal by a delay time within 2 msec, a second delay step of delaying the third signal by the same delay time as the delay time by which the second signal is delayed in the first delay step, a noise mixing estimation step of estimating first noise mixed into the first signal based on the second signal delayed in the first delay step and the third signal delayed in the second delay step, and an elimination step of eliminating the first noise estimated in the noise mixing estimation step from the first signal.
[0010] Also, for example, the signal processing method according to the present disclosure includes a first acquisition step of acquiring a first signal output from a first microphone installed in the driver's seat of a vehicle, a second acquisition step of acquiring a second signal output from a second microphone installed in the passenger seat of the vehicle, a third acquisition step of acquiring a third signal output from a third microphone installed in the rear seat of the vehicle, a first delay step of delaying the first signal by a delay time within 2 msec, a second delay step of delaying the second signal by the same delay time as the delay time for delaying the first signal in the first delay step, a third delay step of delaying the third signal by the same delay time, a noise mixing estimation step of estimating a first noise mixed into the first signal based on the second signal delayed in the second delay step and the third signal delayed in the third delay step, estimating a second noise mixed into the second signal based on the first signal delayed in the first delay step and the third signal delayed in the third delay step, and estimating a third noise mixed into the third signal based on the first signal delayed in the first delay step and the second signal delayed in the second delay step, and an elimination step of eliminating the first noise estimated in the noise mixing estimation step from the first signal, eliminating the second noise estimated in the noise mixing estimation step from the second signal, and eliminating the third noise estimated in the noise mixing estimation step from the third signal.
[0011] Also, for example, the signal processing method according to the present disclosure includes a first acquisition step of acquiring a first signal output from a first microphone, a second acquisition step of acquiring a second signal output from a second microphone installed at a position different from the first microphone, and a third acquisition step of acquiring a third signal output from a third microphone installed at a position different from the first microphone and the second microphone. 3An acquisition step, a first delay step for delaying the first signal, a second delay step for delaying the second signal by the same delay time as the delay time for delaying the first signal in the first delay step, a third delay step for delaying the third signal by the same delay time, an interference noise estimation step for estimating first noise mixed into the first signal based on the second signal delayed in the second delay step and the third signal delayed in the third delay step, estimating second noise mixed into the second signal based on the first signal delayed in the first delay step and the third signal delayed in the third delay step, and estimating third noise mixed into the third signal based on the first signal delayed in the first delay step and the second signal delayed in the second delay step, and an elimination step for eliminating the first noise estimated in the interference noise estimation step from the first signal, eliminating the second noise estimated in the interference noise estimation step from the second signal, and eliminating the third noise estimated in the interference noise estimation step from the third signal.
Effect of the Invention
[0012] The signal processing apparatus and signal processing method according to the present disclosure can suppress interference noise such as non-target voices mixed into the target voice and prevent the target voice from being suppressed.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
DETAILED DESCRIPTION OF THE INVENTION
[0014] (Knowledge underlying the present disclosure) Conventionally, in order to estimate noise from a target voice, noise mixed into the microphone that acquires the target voice has been estimated using a signal acquired from a microphone placed at a position away from the microphone that acquires the target voice, and a technique for suppressing the noise mixed into the target voice has been used. For example, when it is desired to acquire the driver's voice from a microphone installed in the driver's seat, it is necessary to identify the voices of passengers in seats other than the driver's seat included in the voice acquired by the microphone installed in the driver's seat. Therefore, voices of passengers in seats other than the driver's seat mixed into the voice acquired from the microphone installed in the driver's seat have been identified using the voices respectively acquired from microphones installed in seats other than the driver's seat. However, in order to use this technique, it has been necessary to control the operation of the sound collection devices such as microphones installed in each seat in accordance with the timing of the speech of each passenger. However, it is difficult to determine whether or not a passenger in each seat is speaking, and if control is performed based on an incorrect determination as to whether or not a passenger in each seat is speaking, there is a problem that the target voice is suppressed. Also, if control is performed based on an incorrect determination as to whether or not a passenger in each seat is speaking, there is also a problem that the performance of suppressing noise such as the voices of other passengers mixed into the target voice deteriorates.
[0015] Therefore, a signal processing apparatus according to an aspect of the present disclosure may include a first acquisition unit that acquires a first signal output from a first microphone, a second acquisition unit that acquires a second signal output from a second microphone installed at a position different from the first microphone, a delay unit that delays the second signal, a mixed sound estimation unit that estimates noise mixed into the first signal based on the second signal delayed by the delay unit, and an elimination unit that eliminates the noise estimated by the mixed sound estimation unit from the first signal.
[0016] As a result, the signal processing apparatus according to one aspect of the present disclosure can suppress a signal representing voice, which is noise mixed into the target voice, without suppressing the signal representing the target voice. Therefore, the signal processing apparatus according to one aspect of the present disclosure can more accurately recognize the target voice.
[0017] Further, for example, the delay unit may delay the second signal by a time determined based on the positional relationship between the first microphone and the second microphone.
[0018] As a result, the signal processing apparatus according to one aspect of the present disclosure can prevent the second signal representing noise from arriving at the cancellation unit later than the first signal representing the target voice. Therefore, the signal processing apparatus according to one aspect of the present disclosure can process the first signal using the second signal that has already arrived at the cancellation unit.
[0019] Further, for example, the delay unit may delay the second signal based on the frequency components included in the second signal.
[0020] As a result, the signal processing apparatus according to one aspect of the present disclosure can set a more appropriate delay time for the second signal representing noise. Therefore, the signal processing apparatus according to one aspect of the present disclosure can more effectively suppress the noise mixed into the first signal representing the target voice.
[0021] Further, the signal processing method according to one aspect of the present disclosure may include a first acquisition step of acquiring a first signal output from a first microphone, a second acquisition step of acquiring a second signal output from a second microphone installed at a position different from the first microphone, a delay step of delaying the second signal, a mixing noise estimation step of estimating noise mixed into the first signal based on the second signal delayed by the delay step, and a cancellation step of canceling the noise estimated in the mixing noise estimation step from the first signal.
[0022] Accordingly, the signal processing method according to one aspect of the present disclosure can suppress a signal representing voice that is noise mixed into the target voice without suppressing the signal representing the target voice. Therefore, the signal processing apparatus according to one aspect of the present disclosure can more accurately recognize the target voice.
[0023] (Embodiment) Hereinafter, embodiments will be specifically described with reference to the drawings.
[0024] FIG. 1 is a configuration diagram of a signal processing apparatus according to an embodiment of the present disclosure. The signal processing apparatus 1 according to the embodiment of the present disclosure includes a microphone 11, a microphone 12, a microphone 13, a microphone 14, a first acquisition unit 15, a second acquisition unit 16, a third acquisition unit 17, a fourth acquisition unit 18, a delay unit 4a, a delay unit 4b, a delay unit 4c, a mixed sound estimation unit 2 including an adaptive filter 2a, an adaptive filter 2b, and an adaptive filter 2c, and an elimination unit 3.
[0025] The microphones 11, 12, 13, and 14 acquire sound such as voice and convert it into a signal. The microphone may be a moving coil microphone or a ribbon microphone. Further, the microphone may be a condenser microphone or a laser optical microphone or the like.
[0026] The first acquisition unit 15 is electrically connected to the microphone 11 by wire or wirelessly. The first acquisition unit 15 receives a signal obtained by converting the voice acquired by the microphone 11 from the microphone 11. The first acquisition unit 15 does not have a delay unit.
[0027] The second acquisition unit 16 is electrically connected to the microphone 12 by wire or wirelessly. The second acquisition unit 16 receives a signal obtained by converting the voice acquired by the microphone 12 from the microphone 12.
[0028] The third acquisition unit 17 is electrically connected to the microphone 13 either wired or wirelessly. The third acquisition unit 17 receives, from the microphone 13, a signal obtained by converting the sound captured by the microphone 13.
[0029] The fourth acquisition unit 18 is electrically connected to the microphone 14 either wired or wirelessly. The fourth acquisition unit 18 receives, from the microphone 14, a signal obtained by converting the sound captured by the microphone 14.
[0030] The noise estimation unit 2 including the first acquisition unit 15, the second acquisition unit 16, the third acquisition unit 17, the fourth acquisition unit 18, the delay unit 4a, the delay unit 4b, the delay unit 4c, the adaptive filter 2a, the adaptive filter 2b, and the adaptive filter 2c, and the cancellation unit 3 are realized by a processor and a memory. The functions of the processor and the memory may utilize those provided by cloud computing. Further, the first acquisition unit 15, the second acquisition unit 16, the third acquisition unit 17, the fourth acquisition unit 18, the adaptive filter 2a, the adaptive filter 2b, and the adaptive filter 2c may each be realized by a dedicated circuit.
[0031] The delay unit 4a is electrically connected to the second acquisition unit 16, the delay unit 4b is electrically connected to the third acquisition unit 17, and the delay unit 4c is electrically connected to the fourth acquisition unit 18, either by wire or wirelessly. The delay unit 4a, the delay unit 4b, and the delay unit 4c receive the second signal, the third signal, and the fourth signal respectively acquired by the second acquisition unit 16, the third acquisition unit 17, and the fourth acquisition unit 18, and delay the received signals by a predetermined time. Here, delaying the signal means that the signals received by the delay unit 4a, the delay unit 4b, and the delay unit 4c are transmitted to the ambient noise estimation unit 2 after a certain period of time. Also, for example, the delay unit 4a, the delay unit 4b, and the delay unit 4c are a memory group that realizes a stack by a plurality of continuously connected memories for delaying the signal by a predetermined time. The memory group may output the acquired signal in a First In First Out (FIFO) manner to delay the signal. Also, for example, the time when the delay unit 4a, the delay unit 4b, and the delay unit 4c output the signal is not more than the value obtained by dividing the respective distances between the microphone 11 and the microphones 12, 13, and 14 by the speed of sound.
[0032] The ambient noise estimation unit 2 includes an adaptive filter 2a, an adaptive filter 2b, and an adaptive filter 2c. The ambient noise estimation unit 2 is electrically connected to the delay unit 4a, the delay unit 4b, and the delay unit 4c, either by wire or wirelessly. The ambient noise estimation unit 2 receives the second signal, the third signal, and the fourth signal delayed by the delay unit 4a, the delay unit 4b, and the delay unit 4c. The ambient noise estimation unit 2 estimates the noise mixed in the first signal acquired by the first acquisition unit 15 based on the second signal, the third signal, and the fourth signal.
[0033] Specifically, the ambient noise estimation unit 2 corrects the filter coefficients of the adaptive filter 2a, the adaptive filter 2b, and the adaptive filter 2c using a predetermined adaptive algorithm so that the signal SO (an example of the output signal) output from the cancellation unit 3 is uncorrelated or independent of the inputs of the adaptive filter 2a, the adaptive filter 2b, and the adaptive filter 2c. The signal SO is a signal obtained by subtracting the ambient noise signal S2' from the signal S1 (an example of the first signal) acquired by the microphone 11. Therefore, when the filter coefficient of the adaptive filter 2a is corrected so that the signal SO is uncorrelated or independent of the input of the adaptive filter 2a, the signal output from the adaptive filter 2a indicates the ambient noise signal S2' that is the ambient noise in which the voice emitted by the passenger P2 included in the signal S1 is mixed with the voice generated by the passenger P1.
[0034] Note that the ambient noise estimation unit 2 may execute the correction process of the filter coefficients periodically, or may execute it each time the microphones 12, 13, and 14 acquire signals of a certain level or higher. Here, as the predetermined adaptive algorithm, an LMS (The least-mean-square) algorithm, an ICA (Independent Component Analisys) algorithm, or the like can be adopted.
[0035] The adaptive filter 2a, the adaptive filter 2b, and the adaptive filter 2c extract necessary signals from the received signals through a mathematical filter whose coefficients are variable. Specifically, as described above, the adaptive filter 2a, the adaptive filter 2b, and the adaptive filter 2c can calculate new coefficients by calculation at any time and change the coefficients used for the filter. For the calculation of the coefficients, dynamic non-linear feedback control or the like that feeds back and uses the outputs of the adaptive filter 2a, the adaptive filter 2b, and the adaptive filter 2c respectively can be performed. Also, the adaptive filter 2a, the adaptive filter 2b, and the adaptive filter 2c can change the magnitude (gain) of the output of the received signal. As the adaptive filter, an LMS filter or the like can be adopted.
[0036] The cancellation unit 3 is electrically connected to the first acquisition unit 15 and the ambient noise estimation unit 2, either wired or wirelessly. The cancellation unit 3 suppresses the noise estimated by the ambient noise estimation unit 2 in the first signal acquired by the first acquisition unit 15.
[0037] For example, in the signal processing apparatus 1, a microphone 11 may be installed in the driver's seat of the automobile, a microphone 12 may be installed in the passenger seat of the automobile, and microphones 13 and 14 may be installed in the rear seats of the automobile. In this case, the signal processing apparatus 1 operates to suppress noise from the voice of the passenger (driver) in the driver's seat of the automobile. A signal processing system symmetric to the above-described components may be mounted on the automobile to suppress noise from the voices of passengers in seats other than the driver's seat of the automobile. For example, in order to suppress noise from the voice of the passenger in the passenger seat of the automobile, a microphone 11 may be installed in the passenger seat of the automobile, a microphone 12 may be installed in the driver's seat of the automobile, and microphones 13 and 14 may be installed in the rear seats of the automobile. Further, for example, in order to suppress noise from the voice of the passenger in the left rear seat of the automobile, a microphone 11 may be installed in the left rear seat of the automobile, a microphone 12 may be installed in the driver's seat of the automobile, a microphone 13 may be installed in the passenger seat of the automobile, and a microphone 14 may be installed in the right rear seat of the automobile. Further, for example, in order to suppress noise from the voice of the passenger in the right rear seat of the automobile, a microphone 11 may be installed in the right rear seat of the automobile, a microphone 12 may be installed in the driver's seat of the automobile, a microphone 13 may be installed in the passenger seat of the automobile, and a microphone 14 may be installed in the left rear seat of the automobile. Further, a plurality of these symmetric signal processing systems may be installed.
[0038] The connection relationships of the microphones 11, 12, 13, and 14 with other components of the signal processing device 1 shall be the same as those described above. Also, the installation locations of the microphones 11, 12, 13, and 14 are not limited to the locations shown above. Further, the number of seats in the automobile where the microphones are installed is not limited to four. Four or more microphones may be installed at four or more locations. For example, it may be six seats in a passenger car with a three-row seat, or there may be six or more locations. Also, the number of installation locations of the microphones may be three or less.
[0039] Also, a plurality of each of the components described here may be installed in the signal processing device 1.
[0040] FIG. 2 is a flowchart showing the operation of the signal processing device according to an embodiment of the present disclosure. Hereinafter, the overall flow of the system will be described using the flowchart.
[0041] First, in the signal processing device 1, the first acquisition unit 15 and the second acquisition unit 16 acquire a first signal and a second signal from the microphones 11 and 12, respectively (step S101). At this time, further, the third acquisition unit 17 and the fourth acquisition unit 18 may acquire a third signal and a fourth signal from the microphones 13 and 14, respectively.
[0042] Next, in the signal processing device 1, the delay unit 4a delays the second signal received from the second acquisition unit 16 by a predetermined time (step S102). At this time, further, the delay unit 4b and the delay unit 4c may delay the third signal received from the third acquisition unit 17 and the fourth signal received from the fourth acquisition unit 18 by a predetermined time.
[0043] The time for the delay unit 4a, the delay unit 4b, and the delay unit 4c to delay the second signal, the third signal, and the fourth signal may all be the same or may be different from each other. The time for the delay unit 4a, the delay unit 4b, and the delay unit 4c to delay the second signal, the third signal, and the fourth signal may be determined from their respective positional relationships with the microphone 11 and the microphones 12, 13, and 14. Here, the positional relationship means, for example, the distance from the microphone 11 and the like. For example, the longer the distance from the microphone 11, the longer the time for delaying the signal may be set.
[0044] Also, the time for the delay unit 4a, the delay unit 4b, and the delay unit 4c to delay the second signal, the third signal, and the fourth signal may be determined individually according to the frequencies of the second signal, the third signal, and the fourth signal respectively. For example, the delay unit 4a identifies the frequency of the second signal by frequency analysis. Then, the delay unit 4a may determine the time for delaying the signal based on a table prepared in advance that defines the delay time according to the frequency. Also, the delay unit 4a may determine the time for delaying the signal according to the frequency components even within the same signal. For low-frequency signals, since the wavelength is long, it is difficult for the delay of the signal to act effectively. For example, it has been found from experiments that when the time for delaying a low-frequency signal is set to a certain time or more, the degree of suppressing the first signal decreases. Therefore, for example, the time for delaying a signal of a low-frequency component may be set longer than a certain time, and the time for delaying a signal of a high-frequency component may be set shorter than a certain time.
[0045] Also, the time for the delay unit 4a, the delay unit 4b, and the delay unit 4c to delay the second signal, the third signal, and the fourth signal may be determined by the temperature in the vehicle interior or the like.
[0046] The time for the delay unit 4a, the delay unit 4b, and the delay unit 4c to delay the second signal, the third signal, and the fourth signal is set by the delay unit 4a, the delay unit 4b, and the delay unit 4c so as not to exceed the time it takes for sound to travel between the microphone 11 and the microphones 12, 13, and 14. For example, if the distance between the microphone 11 and the microphone 12 is 1 m and the time it takes for sound to travel between the microphone 11 and the microphone 12 is 3 msec, the delay unit 4a sets a delay time within 3 msec for the second signal.
[0047] Subsequently, in the signal processing device 1, the ambient noise estimation unit 2 estimates the noise mixed into the first signal acquired by the first acquisition unit 15 based on the second signal delayed by the delay unit 4a (step S103). At this time, the ambient noise estimation unit 2 may further estimate the noise mixed into the first signal acquired by the first acquisition unit 15 based on the third signal and the fourth signal delayed by the delay unit 4b and the delay unit 4c.
[0048] Then, in the signal processing device 1, the cancellation unit 3 suppresses the noise estimated by the ambient noise estimation unit 2 in the first signal (step S104).
[0049] Next, the state of the signal processed in the signal processing device 1 will be described with reference to FIGS. 3 to 8.
[0050] In the signal processing apparatus 1, by appropriately delaying the signals acquired by the second acquisition unit 16, the third acquisition unit 17, and the fourth acquisition unit 18, it is possible to prevent estimating, as noise, a signal representing a target voice included in the signals representing voices or the like acquired from the second acquisition unit 16, the third acquisition unit 17, and the fourth acquisition unit 18. A signal representing a target voice included in the signals representing voices or the like acquired by the second acquisition unit 16, the third acquisition unit 17, and the fourth acquisition unit 18 reaches the elimination unit 3 after being delayed from the signal representing the target voice acquired from the first acquisition unit 15. For this reason, it does not happen that the elimination unit 3 suppresses the target voice acquired by the first acquisition unit 15 based on the signal representing the target voice included in the signals of voices or the like acquired from the second acquisition unit 16, the third acquisition unit 17, and the fourth acquisition unit 18.
[0051] FIG. 3 is a diagram showing the frequency characteristics before and after processing by the elimination unit of a signal representing the voice of a driver acquired from the microphone 11 installed in the driver's seat when the delay unit delays the signal by 0 msec in the embodiment of the present disclosure. Line 100 is the frequency characteristic of the first signal that has not been processed by the elimination unit 3 when the delay unit 4a delays the second signal by 0 msec. Line 101 is the frequency characteristic of the first signal that has been processed by the elimination unit 3 when the delay unit 4a delays the second signal by 0 msec. Between approximately 100 Hz and approximately 10,000 Hz, line 101 shows a lower value than line 100. That is, when a 0 msec delay is added to the second signal in the delay unit 4a and the elimination unit 3 is activated, the elimination unit 3 suppresses the first signal. This is because the second signal includes a signal having the same waveform as the first signal, and thus the first signal is suppressed by eliminating the noise estimated using the second signal.
[0052] For example, consider a case where the voice obtained by the microphone 12 installed on the passenger seat is used to suppress noise from the voice obtained by the microphone 11 installed on the driver's seat. The voice emitted by the driver is mixed into the voice obtained by the microphone 12 installed on the passenger seat through reflections in the vehicle interior or the like. Based on the voice of the driver mixed into the voice obtained by the microphone 12 installed on the passenger seat, the voice of the driver included in the voice obtained by the microphone 11 installed on the driver's seat is suppressed.
[0053] FIG. 4 is a diagram showing the frequency characteristics before and after the processing by the cancellation unit of the signal representing the voice of the driver obtained from the microphone 11 installed on the driver's seat when the delay unit delays the signal by 2 msec in the embodiment of the present disclosure. Line 102 is the frequency characteristic of the first signal that has not been processed by the cancellation unit 3 when the delay unit 4a delays the second signal by 2 msec. Line 103 is the frequency characteristic of the first signal that has been processed by the adaptive filter 2a when the delay unit 4a delays the second signal by 2 msec. Between approximately 100 Hz and approximately 10,000 Hz, line 102 shows a lower value than line 103, but the difference is reduced compared to the difference between lines 100 and 101 shown in FIG. 3.
[0054] That is, when the cancellation unit 3 is operated when the delay unit 4a delays the second signal by 2 msec, the amount by which the cancellation unit 3 suppresses the first signal is reduced. This is because delaying the second signal by 2 msec reduces the signal having the same waveform as the first signal included in the second signal. Therefore, suppressing the noise estimated using the second signal reduces the suppression of the first signal compared to the case shown in FIG. 3.
[0055] Here, since the speed of sound in air is approximately 340 m / s, a delay of 2 msec corresponds to a distance that the sound travels, which is approximately 60 cm to approximately 70 cm. That is, when the microphones 11 and 12 are installed approximately 60 cm to approximately 70 cm apart, the time it takes for the driver's voice to reach the microphone 12 is approximately 2 msec.
[0056] For example, consider a case where the voice obtained by the microphone 12 installed on the passenger seat is used to suppress noise from the voice obtained by the microphone 11 installed on the driver's seat. The voice emitted by the driver is mixed into the voice obtained by the microphone 12 installed on the passenger seat through reflections in the vehicle interior and the like. However, by delaying the signal representing the voice obtained by the microphone 12 installed on the passenger seat by 2 msec, the degree to which the voice of the driver included in the voice obtained by the microphone 12 installed on the passenger seat suppresses the voice of the driver included in the voice obtained by the microphone 11 installed on the driver's seat is reduced. In order for the elimination unit 3 to suppress the voice of the driver included in the voice obtained by the microphone 11 installed on the driver's seat based on the voice of the driver included in the voice obtained by the microphone 12 installed on the passenger seat, the voice obtained by the microphone 12 installed on the passenger seat must arrive at the elimination unit 3 before the voice obtained by the microphone 11 installed on the driver's seat. However, by delaying the voice obtained by the microphone 12 installed on the passenger seat, the voice obtained by the microphone 12 installed on the passenger seat cannot arrive at the elimination unit 3 before the voice obtained by the microphone 11 installed on the driver's seat. Therefore, the degree to which the elimination unit 3 suppresses the voice of the driver included in the voice obtained by the microphone 11 installed on the driver's seat based on the voice of the driver included in the voice obtained by the microphone 12 installed on the passenger seat is reduced.
[0057] FIG. 5 is a diagram showing the frequency characteristics before and after the processing by the cancellation unit of the signal representing the driver's voice acquired from the microphone 11 installed in the driver's seat when the signal is delayed by 6 msec in the delay unit in the embodiment of the present disclosure. Line 104 is the frequency characteristic of the first signal without the processing by the cancellation unit 3 when the first signal is delayed by 6 msec. Line 105 is the frequency characteristic of the first signal after the processing by the adaptive filter 2a when the first signal is delayed by 6 sec. Between approximately 100 Hz and approximately 10,000 Hz, line 104 shows a lower value than line 105, but the difference between line 104 and line 105 is reduced compared to the difference between line 102 and line 103 shown in FIG. 4.
[0058] That is, when the second signal is delayed by 6 msec in the delay unit 4a and the cancellation unit 3 is activated, the amount by which the cancellation unit 3 suppresses the first signal is reduced. This is because the signal having the same waveform as the first signal included in the second signal is decreased by delaying the second signal by 6 msec. Therefore, suppressing the noise estimated using the second signal reduces the suppression of the first signal to a level lower than the difference between line 102 and line 103 shown in FIG. 4.
[0059] For example, consider a case where the voice acquired by the microphone 12 installed in the passenger seat is used to suppress noise from the voice acquired by the microphone 11 installed in the driver's seat. The voice emitted by the driver is mixed into the voice acquired by the microphone 12 installed in the passenger seat through reflections in the vehicle interior and the like. However, by delaying the signal representing the voice acquired by the microphone 12 installed in the passenger seat by 6 msec, based on the voice of the driver included in the voice acquired by the microphone 12 installed in the passenger seat, the degree to which the voice of the driver included in the voice acquired by the microphone 11 installed in the driver's seat is suppressed is further reduced compared to when it is delayed by 2 msec.
[0060] As can be seen from FIGS. 3, 4, and 5, the signal representing the target voice is less suppressed by delaying it in the delay section.
[0061] FIG. 6 is a diagram showing the frequency characteristics before and after processing by the cancellation unit of the signal representing the voice of the passenger in the passenger seat acquired from the microphone 11 installed in the driver's seat when the signal is delayed by 0 msec in the delay section in the embodiment of the present disclosure. Line 106 is the frequency characteristic of the first signal without the processing by the cancellation unit 3 when the first signal is delayed by 0 msec. Line 107 is the frequency characteristic of the first signal after the processing by the adaptive filter 2a when the first signal is delayed by 0 msec. Between approximately 100 Hz and approximately 10,000 Hz, line 104 shows a value approximately 20 dB lower than line 105. That is, the voice of the passenger in the passenger seat mixed into the microphone 11 installed in the driver's seat is suppressed.
[0062] For example, consider the case of using the voice acquired by the microphone 12 installed in the passenger seat to suppress noise from the voice acquired by the microphone 11 installed in the driver's seat. The voice emitted by the passenger in the passenger seat is mixed into the voice acquired by the microphone 11 installed in the driver's seat directly or via reflections in the vehicle interior or the like. By estimating this mixed noise using the waveform of the signal of the voice acquired by the microphone 12 installed in the passenger seat, the noise is removed from the voice acquired by the microphone 11 installed in the driver's seat. As shown in FIG. 6, when the signal representing the voice acquired by the microphone 12 installed in the passenger seat is delayed by 0 msec, the voice acquired by the microphone 11 installed in the driver's seat is suppressed by approximately 20 dB.
[0063] FIG. 7 is a diagram showing the frequency characteristics before and after the processing by an erasure unit of a signal representing the voice of a passenger in the passenger seat acquired from a microphone 11 installed in the driver's seat when the signal is delayed by 2 msec in the delay unit in an embodiment of the present disclosure. A line 108 is the frequency characteristic of a first signal before the processing by an erasure unit 3 when the first signal is delayed by 2 msec. A line 109 is the frequency characteristic of the first signal after the processing by an adaptive filter 2a when a delay of 2 msec is added to the first signal. Over the range from about 100 Hz to about 10,000 Hz, the line 108 shows a lower value than the line 109. However, the suppression amount of the first signal shown in FIG. 7 is reduced compared to the suppression amount of the first signal shown in FIG. 6.
[0064] For example, consider a case where the voice acquired by a microphone 12 installed in the passenger seat is used to suppress noise from the voice acquired by a microphone 11 installed in the driver's seat. As described above, the voice uttered by the passenger in the passenger seat is mixed into the voice acquired by the microphone 11 installed in the driver's seat directly or via reflection in the vehicle interior or the like. By estimating this mixed noise using the waveform of the signal of the voice acquired by the microphone 12 installed in the passenger seat, the noise is removed from the voice acquired by the microphone 11 installed in the driver's seat. As shown in FIG. 7, when the signal representing the voice acquired by the microphone 12 installed in the passenger seat is delayed by 2 msec, the degree to which the voice acquired by the microphone 11 installed in the driver's seat is suppressed is reduced compared to the case where the signal representing the voice acquired by the microphone 12 installed in the passenger seat is not delayed. From the viewpoint of estimating the voice of the passenger in the passenger seat mixed into the voice acquired from the microphone 11 installed in the driver's seat as noise and suppressing it from the voice acquired from the microphone 11 installed in the driver's seat, it is desirable not to delay the signal representing the voice acquired from the microphone 12 installed in the passenger seat.
[0065] FIG. 8 is a diagram showing the frequency characteristics before and after processing by an erasure unit of a signal representing the voice of a passenger in the passenger seat acquired from a microphone 11 installed in the driver's seat when the signal is delayed by 6 msec in the delay unit in an embodiment of the present disclosure. Line 110 is the frequency characteristic of the first signal without the processing by the erasure unit 3 when a delay of 6 msec is added to the first signal. Line 111 is the frequency characteristic of the first signal subjected to the processing by the adaptive filter 2a when the first signal is delayed by 6 msec. Between approximately 100 Hz and approximately 10,000 Hz, line 110 shows a lower value than line 111. However, the suppression amount shown in FIG. 8 is reduced compared to the suppression amount shown in FIG. 7.
[0066] For example, consider a case where the voice acquired by the microphone 12 installed in the passenger seat is used to suppress noise from the voice acquired by the microphone 11 installed in the driver's seat. As described above, the voice uttered by the passenger in the passenger seat is mixed into the voice acquired by the microphone 11 installed in the driver's seat, either directly or via reflection in the vehicle interior or the like. By estimating this mixed noise using the waveform of the signal of the voice acquired by the microphone 12 installed in the passenger seat, the noise is removed from the voice acquired by the microphone 11 installed in the driver's seat. As shown in FIG. 7, when the signal representing the voice acquired by the microphone 12 installed in the passenger seat is delayed by 6 msec, the degree to which the voice acquired by the microphone 11 installed in the driver's seat is suppressed is reduced compared to the case where the signal representing the voice acquired by the microphone 12 installed in the passenger seat is delayed by 0 msec and the case where the signal representing the voice acquired by the microphone 12 installed in the passenger seat is delayed by 2 msec. From the viewpoint of estimating the voice of the passenger in the passenger seat mixed into the voice acquired from the microphone 11 installed in the driver's seat as noise and suppressing the noise from the voice acquired from the microphone 11 installed in the driver's seat, it is desirable not to delay the signal representing the voice acquired from the microphone 12 installed in the passenger seat.
[0067] As can be seen from FIGS. 6, 7, and 8, the signal representing the voice which is noise is less suppressed by delaying it in the delay unit 4a. Therefore, in order to suppress the signal representing the voice which is noise, it is desirable not to delay it more than a predetermined time in the delay unit 4a.
[0068] Conversely, as shown in FIGS. 3, 4, and 5, in order not to suppress the voice which is the target, it is desirable to delay the signal for a predetermined time or more. However, as described above, the time for delaying the signal is set by the delay unit 4a, the delay unit 4b, and the delay unit 4c so as not to exceed the time taken for the voice to travel between the microphone 11 and the microphones 12, 13, and 14.
[0069] The delay time of the signal set by the delay unit 4a, the delay unit 4b, and the delay unit 4c may be determined based on the tendency of voice suppression in the signal processing device 1 as shown in FIGS. 3 to 8.
[0070] (Modification example) The above-described signal processing device 1 can also be applied to a translation system 20 that recognizes the voice of a speaker, translates it into another language, and outputs the translation. FIG. 9 is a diagram of a translation system to which the signal processing system is applied in a modification example of the present disclosure. As shown in FIG. 9, the translation system 20 includes a microphone 21a, a microphone 21b, a first acquisition unit 25, a second acquisition unit 26, a delay unit 24 including an adaptive filter 22, and an elimination unit 23. The translation system 20 may further include an information processing unit that recognizes the voice with suppressed noise and performs translation, and an output unit that outputs the translation result.
[0071] The first acquisition unit 25 is electrically connected to the microphone 21a by wire or wirelessly. The first acquisition unit 25 acquires a first signal such as voice from the microphone 21a. The first acquisition unit 25 receives a signal obtained by converting the voice acquired by the microphone 21a from the microphone 21a.
[0072] The second acquisition unit 26 is electrically connected to the microphone 21b either wired or wirelessly. The second acquisition unit 26 acquires a second signal such as audio from the microphone 21b. The second acquisition unit 26 receives, from the microphone 21b, a signal obtained by converting the audio acquired by the microphone 21b.
[0073] The first acquisition unit 25, the second acquisition unit 26, the delay unit 24 including the adaptive filter 22, and the cancellation unit 23 are realized by a processor and a memory. The functions of the processor and the memory may utilize those provided by cloud computing. Also, the first acquisition unit 25, the second acquisition unit 26, and the adaptive filter 22 may each be realized by a dedicated circuit.
[0074] The delay unit 24 is electrically connected to the second acquisition unit 26 either wired or wirelessly. The delay unit 24 receives the fifth signal acquired by the second acquisition unit 26 and delays the received signal by a predetermined time.
[0075] The adaptive filter 22 is electrically connected to the delay unit 24 either wired or wirelessly. The adaptive filter 22 receives the fifth signal delayed by the delay unit 24. The adaptive filter 22 estimates the noise mixed into the fifth signal acquired by the first acquisition unit 25 based on the fifth signal.
[0076] The adaptive filter 22 extracts a necessary signal from the received signal through a mathematical filter with variable coefficients. The adaptive filter 22 can calculate new coefficients by calculation at any time and change the coefficients used for the filter.
[0077] The cancellation unit 23 is electrically connected to the first acquisition unit 25 and the adaptive filter 22 either wired or wirelessly. The cancellation unit 23 suppresses the noise estimated by the adaptive filter 22 from the fifth signal acquired by the first acquisition unit 25.
[0078] In a modification of the present disclosure, the translation system 20 is assumed to be used face-to-face by the person 30a and the person 30b. In the translation system 20, the first acquisition unit 25 acquires the voice uttered by the person 30a through the microphone 21a. Also, in the translation system 20, the second acquisition unit 26 acquires the voice uttered by the person 30b through the microphone 21b. The voice signal acquired by the second acquisition unit 26 is delayed by a certain time in the delay unit 24 and processed by the adaptive filter 22. Then, the information on the noise estimated by the adaptive filter 22 arrives at the suppression unit 23. The suppression unit 23 suppresses the noise estimated by the adaptive filter 22 from the voice signal acquired by the first acquisition unit 25.
[0079] Translation processing is performed on the voice signal with the noise suppressed, and the translation result is output. As described above, the translation system 20 can suppress noise such as the voice uttered by the person 30b from the voice uttered by the person 30a. Note that the number of each component included in the translation system 20 may be increased compared to that shown above. In the translation system 20 in the modification of the present disclosure, although it is assumed that two persons use it, the number of users is not limited to two. The translation system 20 in the modification of the present disclosure may be configured to be used by three or more persons.
[0080] These general or specific aspects may be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0081] As described above, the signal processing apparatus 1 and the signal processing method have been described based on the embodiments. However, the signal processing system and the signal processing method are not limited to these embodiments. As long as the gist of the present disclosure is not deviated from, forms obtained by applying various modifications conceived by those skilled in the art to these embodiments or forms constructed by combining components in different embodiments may also be included within the scope of one or more aspects.
Industrial Applicability
[0082] The present disclosure is applicable to an in-vehicle sound collection system or a translation system.
Explanation of Reference Numerals
[0083] 1 Signal processing device 2 Background noise estimation unit 2a, 2b, 2c, 22 Adaptive filter 3, 23 Canceling unit 4a, 4b, 4c, 24 Delay unit 11, 12, 13, 14, 21a, 21b Microphone 15, 25 First acquisition unit 16, 26 Second acquisition unit 17 Third acquisition unit 18 Fourth acquisition unit 20 Translation system 30a, 30b Person 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111 Line
Claims
1. A first acquisition unit that acquires a first signal output from a first microphone installed in a driver's seat of the vehicle; a second acquisition unit that acquires a second signal output from a second microphone installed in a passenger seat of the vehicle; a third acquisition unit that acquires a third signal output from a third microphone installed in a rear seat of the vehicle; a first delay unit that delays the second signal by a delay time of 2 msec or less; a second delay unit that delays the third signal by the same delay time as the first delay unit delays the second signal; a mixed sound estimation unit that estimates a first noise mixed into the first signal based on the second signal delayed by the first delay unit and the third signal delayed by the second delay unit; and an erasure unit that erases the first noise estimated by the mixed sound estimation unit from the first signal. Signal processing device.
2. a third delay unit that delays the first signal by the same delay time, The mixed sound estimation unit estimates a second noise mixed into the second signal based on the first signal delayed by the third delay unit and the third signal delayed by the second delay unit, The elimination unit eliminates the second noise estimated by the mixed sound estimation unit from the second signal. The signal processing device according to claim 1 .
3. The mixed sound estimation unit estimates a third noise mixed into the third signal based on the first signal delayed by the third delay unit and the second signal delayed by the first delay unit, The elimination unit eliminates the third noise estimated by the mixed sound estimation unit from the third signal. The signal processing device according to claim 2 .
4. A first acquisition unit that acquires a first signal output from a first microphone installed in a driver's seat of the vehicle; a second acquisition unit that acquires a second signal output from a second microphone installed in a passenger seat of the vehicle; a third acquisition unit that acquires a third signal output from a third microphone installed in a rear seat of the vehicle; a first delay unit that delays the first signal by a delay time of 2 msec or less; a second delay unit that delays the second signal by the same delay time as the first delay unit delays the first signal; a third delay unit that delays the third signal by the same delay time; a mixed sound estimation unit that estimates a first noise mixed into the first signal based on the second signal delayed by the second delay unit and the third signal delayed by the third delay unit, estimates a second noise mixed into the second signal based on the first signal delayed by the first delay unit and the third signal delayed by the third delay unit, and estimates a third noise mixed into the third signal based on the first signal delayed by the first delay unit and the second signal delayed by the second delay unit; and an elimination unit that eliminates the first noise estimated by the mixed sound estimation unit from the first signal, eliminates the second noise estimated by the mixed sound estimation unit from the second signal, and eliminates the third noise estimated by the mixed sound estimation unit from the third signal. Signal processing device.
5. a first acquisition unit that acquires a first signal output from a first microphone; a second acquisition unit that acquires a second signal output from a second microphone that is installed at a position different from that of the first microphone; a third acquisition unit that acquires a third signal output from a third microphone that is installed at a position different from the first microphone and the second microphone; a first delay unit that delays the first signal; a second delay unit that delays the second signal by the same delay time as the first delay unit delays the first signal; a third delay unit that delays the third signal by the same delay time; a mixed sound estimation unit that estimates a first noise mixed into the first signal based on the second signal delayed by the second delay unit and the third signal delayed by the third delay unit, estimates a second noise mixed into the second signal based on the first signal delayed by the first delay unit and the third signal delayed by the third delay unit, and estimates a third noise mixed into the third signal based on the first signal delayed by the first delay unit and the second signal delayed by the second delay unit; and an elimination unit that eliminates the first noise estimated by the mixed sound estimation unit from the first signal, eliminates the second noise estimated by the mixed sound estimation unit from the second signal, and eliminates the third noise estimated by the mixed sound estimation unit from the third signal. Signal processing device.
6. a first acquisition step of acquiring a first signal output from a first microphone installed in a driver's seat of the vehicle; a second acquisition step of acquiring a second signal output from a second microphone installed in a passenger seat of the vehicle; a third acquisition step of acquiring a third signal output from a third microphone installed in a rear seat of the vehicle; a first delay step for delaying the second signal by a delay time of 2 msec or less; a second delay step for delaying the third signal by the same delay time as the delay time for delaying the second signal by the first delay step; a mixing sound estimating step of estimating a first noise mixed into the first signal based on the second signal delayed by the first delay step and the third signal delayed by the second delay step; and a canceling step of canceling the first noise estimated in the mixing sound estimating step from the first signal. Signal processing methods.
7. a first acquisition step of acquiring a first signal output from a first microphone installed in a driver's seat of the vehicle; a second acquisition step of acquiring a second signal output from a second microphone installed in a passenger seat of the vehicle; a third acquisition step of acquiring a third signal output from a third microphone installed in a rear seat of the vehicle; a first delay step of delaying the first signal by a delay time of 2 msec or less; a second delay step for delaying the second signal by the same delay time as the first delay step for delaying the first signal; a third delay step of delaying the third signal by the same delay time; a mixing sound estimating step of estimating a first noise mixed into the first signal based on the second signal delayed by the second delay step and the third signal delayed by the third delay step, estimating a second noise mixed into the second signal based on the first signal delayed by the first delay step and the third signal delayed by the third delay step, and estimating a third noise mixed into the third signal based on the first signal delayed by the first delay step and the second signal delayed by the second delay step; an elimination step of eliminating the first noise estimated in the mixed sound estimation step from the first signal, eliminating the second noise estimated in the mixed sound estimation step from the second signal, and eliminating the third noise estimated in the mixed sound estimation step from the third signal, Signal processing methods.
8. A first acquisition step of acquiring a first signal output from a first microphone; a second acquisition step of acquiring a second signal output from a second microphone installed at a position different from that of the first microphone; a third acquisition step of acquiring a third signal output from a third microphone installed at a position different from the first microphone and the second microphone; a first delay step for delaying the first signal; a second delay step for delaying the second signal by the same delay time as the first delay step for delaying the first signal; a third delay step of delaying the third signal by the same delay time; a mixing sound estimating step of estimating a first noise mixed into the first signal based on the second signal delayed by the second delay step and the third signal delayed by the third delay step, estimating a second noise mixed into the second signal based on the first signal delayed by the first delay step and the third signal delayed by the third delay step, and estimating a third noise mixed into the third signal based on the first signal delayed by the first delay step and the second signal delayed by the second delay step; an elimination step of eliminating the first noise estimated in the mixed sound estimation step from the first signal, eliminating the second noise estimated in the mixed sound estimation step from the second signal, and eliminating the third noise estimated in the mixed sound estimation step from the third signal, Signal processing methods.
Citation Information
Patent Citations
Voice input device
JP1994075591A
Microphone array unit
JP1999018194A
System and method for eliminating noise
JP2003323194A
Audio equipment including means for de-noising speech signal by fractional delay filtering, in particular for "hands-free" telephony system
JP2012253771A
Speech recognition control device
JP2014203031A