Audio processing system and audio processing device

The audio processing system effectively suppresses uncorrelated noise using adaptive filters and selective coefficient updates, ensuring high accuracy in target sound extraction and improved speech recognition.

DE112020004700B4Active Publication Date: 2025-12-24PANASONIC AUTOMOTIVE SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE112020004700
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-30
Filing Date
2020-06-02
Publication Date
2025-12-24
Estimated Expiration
2040-06-02

AI Technical Summary

Technical Problem

Existing audio processing systems struggle to accurately extract target sounds when ambient noise contains uncorrelated background noise, leading to inefficiencies in echo cancellation and speech recognition.

Method used

An audio processing system utilizing multiple microphones, adaptive filters, and a control module to identify and suppress uncorrelated noise, setting the intensity of noise-containing signals to zero and selectively updating filter coefficients to enhance target sound extraction.

Benefits of technology

Accurately preserves target sounds by reducing processing overhead and improving signal-to-noise ratio, enhancing speech recognition accuracy even in the presence of uncorrelated background noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000017_0000
    Figure 00000017_0000
  • Figure 00000018_0000
    Figure 00000018_0000
  • Figure 00000019_0000
    Figure 00000019_0000
Patent Text Reader

Abstract

Audio processing system (5), comprising: a first microphone (MC1) designed to capture a first audio signal containing a first audio component and to output a first signal based on the first audio signal; one or more microphones (MC2-MC4), each of the one or more microphones being designed to detect an audio signal containing an audio component different from the first audio component and to output a microphone signal based on the audio signal; one or more adaptive filters (F1A-F1C) designed to receive the microphone signals from the one or more microphones (MC2-MC4) and to output pass-through signals based on the microphone signals; a determination module (30) designed to determine whether the individual microphone signals contain uncorrelated noise, wherein the noise is one that has no correlation between the audio signals; a control module (28; 28A) designed to control one or more filter coefficients of one or more adaptive filters (F1A-F1C); and an addition module (27: 27A) designed to subtract a subtraction signal from the first signal based on the through signals, wherein the one or more microphones (MC2-MC4) include a second microphone designed to capture a second audio signal containing a second audio component different from the first audio component, and to output a second signal based on the second audio signal, and If the determination module (30) determines that the second signal contains the uncorrelated noise, the control module (28; 28A) is designed to set the level of the second signal, which is input into the appropriate adaptive filter, to zero.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field

[0001] The present disclosure relates to an audio processing system and an audio processing device. State of the art

[0002] DE 603 ​​10 725 T2 describes a method and a system for processing subband signals with adaptive filters. The system is implemented on an oversampled WOLA filter bank. The input signals are oversampled. The system includes an adaptive filter for each subband and improves its convergence characteristics. For example, convergence is improved by brightening the spectra of the oversampled subband signals and / or by using an affine projection algorithm. The system is suitable for echo and / or noise suppression. Adaptive step size control and adaptation process control with double-speaker detection can be implemented. The system can also implement non-adaptive processing for reducing uncorrelated noise and / or for crosstalk-resistant adaptive noise suppression.

[0003] US 2009 / 0299742A1 describes systems, methods and devices for spectral contrast enhancement of speech signals based on information from a noise reference derived from a multi-channel acquired audio signal by a spatially selective processing filter.

[0004] In a vehicle-mounted speech recognition device and a hands-free system, an echo suppressor for removing ambient noise in order to recognize only the sound of a speaker is known. Patent specification 1 discloses an echo suppressor that switches the number of adaptive filters operated and the number of taps according to the number of sound sources. Bibliography Patent literature

[0005] Patent specification 1: JP 4 889 810 B2 Summary of the invention

[0006] When echo cancellation is performed using adaptive filters, the ambient noise is fed into the adaptive filters as a reference signal. Even if the ambient noise contains uncorrelated background noise, it is advantageous to remove the ambient noise to preserve the target sound.

[0007] One aspect of the present disclosure creates an audio processing system capable of obtaining a target sound with high accuracy, even when the ambient noise contains uncorrelated background noise.

[0008] An audio processing system according to the present disclosure comprises a first microphone, one or more microphones, one or more adaptive filters, a determination module, a control module, and an addition module. The first microphone is configured to detect a first audio signal containing a first audio component and to output a first signal based on the first audio signal. Each of the one or more microphones is configured to detect an audio signal containing an audio component distinct from the first audio component and to output a microphone signal based on the audio signal. The one or more adaptive filters are configured to receive the microphone signals from the one or more microphones and to output pass-through signals based on the microphone signals.The detection module is designed to determine whether the individual microphone signals contain uncorrelated noise, which is noise that exhibits no correlation between the audio signals. The control module is designed to control one or more filter coefficients of the one or more adaptive filters. The addition module is designed to subtract a subtraction signal from the first signal based on the pass-through signals. The one or more microphones include a second microphone designed to capture a second audio signal containing a second audio component distinct from the first, and to output a second signal based on this second audio signal.If the determination module determines that the second signal contains the uncorrelated noise, the control module is designed to set the level of the second signal, which is fed into the appropriate adaptive filter, to zero.

[0009] It should be noted that these comprehensive or specific aspects may be realized by a system, a method, an integrated circuit, a computer program or a data carrier, or by any combination of a system, a device, a method, an integrated circuit, a computer program and a data carrier.

[0010] According to one aspect of the present disclosure, an audio processing system capable of obtaining a target sound with high accuracy, even when the ambient noise contains uncorrelated background noise, has been created. Brief description of the drawing Fig. Figure 1 is a diagram that shows an example of a schematic setup of an audio processing system according to a first embodiment. Fig. Figure 2 is a block diagram illustrating the structure of an audio processing device according to the first embodiment. Fig. Figure 3 is a flowchart illustrating an operating procedure of the audio processing device according to the first embodiment. Fig. Figure 4 is a diagram representing an output result of the audio processing device. Fig. Figure 5 is a diagram that shows an example of a schematic setup of an audio processing system according to a second embodiment. Fig. Figure 6 is a block diagram illustrating the structure of an audio processing device according to the second embodiment. Fig. Figure 7 is a flowchart illustrating an operating procedure of the audio processing device according to the second embodiment. Description of embodiments

[0011] Knowledge underlying the present disclosure In a case where the ambient noise contains uncorrelated background noise, it can be difficult to remove the ambient noise and obtain a target noise, even when echo cancellation is performed using the ambient noise.

[0012] Embodiments of the present disclosure are described in detail below, possibly with reference to the drawing. However, unnecessarily detailed descriptions may have been omitted. It should be noted that the accompanying drawing and the following description are intended for those skilled in the art to enable them to thoroughly understand the present disclosure and are not intended to limit the subject matter described in the claims. First embodiment

[0013] Fig. Figure 1 is a diagram illustrating a schematic diagram of an audio processing system 5 according to a first embodiment. The audio processing system 5 is, for example, mounted on a vehicle 10. An example is described below in which the audio processing system 5 is mounted on the vehicle 10. A plurality of seats is provided in the passenger compartment of the vehicle 10. These plurality of seats are, for example, four seats comprising a driver's seat, a front passenger seat, and left and right rear seats. The number of seats is not limited to this. The audio processing system 5 includes a microphone MC1, a microphone MC2, a microphone MC3, a microphone MC4, and an audio processing device 20. In this example, the number of seats corresponds to the number of microphones, but the number of microphones can also differ from the number of seats.The output of the audio processing device 20 is fed into a speech recognition machine (not shown).

[0014] The speech recognition result from the speech recognition machine is entered into an electronic device 50.

[0015] Microphone MC1 picks up the sound emitted by a driver hm1. In other words, microphone MC1 captures an audio signal containing the audio component emitted by driver hm1. Microphone MC1 is located, for example, on an auxiliary handle on the right side of the driver's seat. Microphone MC2 picks up the sound emitted by a passenger hm2. In other words, microphone MC2 captures an audio signal containing the audio component emitted by passenger hm2. Microphone MC2 is located, for example, on an auxiliary handle on the left side of the front passenger seat. Microphone MC3 picks up the sound emitted by a passenger hm3. In other words, microphone MC3 captures an audio signal containing the audio component emitted by passenger hm3. Microphone MC3 is located, for example, on an auxiliary handle on the left side of the rear seat. Microphone MC4 picks up the sound emitted by a passenger hm4.In other words, microphone MC4 captures an audio signal containing the audio component uttered by occupant hm4. Microphone MC4 is, for example, located on an auxiliary grab handle on the right side of the rear seat.

[0016] The placement positions of microphones MC1, MC2, MC3, and MC4 are not limited to the example described above. For instance, microphone MC1 can be located on the right front surface of the dashboard. Microphone MC2 can be located on the left front surface of the dashboard. Microphone MC3 can be located on the backrest of the passenger seat. Microphone MC4 can be located on the backrest of the driver's seat.

[0017] Any microphone can be a directional microphone or an omnidirectional microphone. Any microphone can be a MEMS microphone (Small Micro Electro Mechanical System) or an electret condenser microphone (ECM). Any microphone can be a directional microphone (beamforming microphone). For example, any microphone can be a microphone array that has a directional effect in the direction of a given location and is capable of capturing sound using the directional method.

[0018] In the present embodiment, the audio processing system 5 comprises a plurality of audio processing devices 20, each corresponding to a specific microphone. More precisely, the audio processing system 5 comprises an audio processing device 21, an audio processing device 22, an audio processing device 23, and an audio processing device 24. Audio processing device 21 corresponds to microphone MC1. Audio processing device 22 corresponds to microphone MC2. Audio processing device 23 corresponds to microphone MC3. Audio processing device 24 corresponds to microphone MC4. Hereinafter, audio processing device 21, audio processing device 22, audio processing device 23, and audio processing device 24 may be collectively referred to as the audio processing device 20.

[0019] In the Fig. In the setup shown in Figure 1, the audio processing device 21, the audio processing device 22, the audio processing device 23, and the audio processing device 24 are each implemented by different hardware, but the functions of the audio processing device 21, the audio processing device 22, the audio processing device 23, and the audio processing device 24 can also be implemented by a single audio processing device 20. Alternatively, some of the functions of the audio processing device 21, the audio processing device 22, the audio processing device 23, and the audio processing device 24 can be implemented by common hardware, and the others can be implemented by different hardware.

[0020] In the present embodiment, each audio processing device 20 is arranged in a respective seat near the respective corresponding microphone. Each audio processing device 20 can be arranged in a dashboard. Fig. Figure 2 is a block diagram illustrating the structure of the audio processing system 5 and the structure of the audio processing device 21. As shown in Fig. As shown in Figure 2, the audio processing system 5 comprises, in addition to the audio processing device 21, the audio processing device 22, the audio processing device 23, and the audio processing device 24, a speech recognition machine 40 and the electronic device 50. The output of the audio processing device 20 is input to the speech recognition machine 40. The speech recognition machine 40 detects the noise contained in the output signal of at least one audio processing device 20 and outputs a speech recognition result. The speech recognition machine 40 generates a speech recognition result and a signal based on the speech recognition result. The signal based on the speech recognition result is, for example, an operating signal for the electronic device 50. The speech recognition result from the speech recognition machine 40 is input to the electronic device 50.The speech recognition machine 40 can be a device distinct from the audio processing device 20. For example, the speech recognition machine 40 is located within a dashboard. The speech recognition machine 40 can be housed and arranged within the seat. Alternatively, the speech recognition machine 40 can be an integrated device incorporated into the audio processing device 20.

[0021] A signal output by the speech recognition machine 40 is input into the electronic device 50. The electronic device 50 then performs an operating process according to the operating signal. The electronic device 50 is, for example, located on the dashboard of the vehicle 10. The electronic device 50 is, for example, a vehicle navigation device. The electronic device 50 could be a display device, a television, or a mobile device. Fig. Paragraph 1 represents a case where there are four people in the vehicle, but the number of occupants is not limited to this. The number of occupants can be less than or equal to the vehicle's maximum passenger capacity. For example, in a case where the vehicle's maximum passenger capacity is six, the number of occupants can be six, five, or fewer.

[0022] The audio processing device 21, audio processing device 22, audio processing device 23, and audio processing device 24 have similar configurations and functions, except for the configuration of part of the filter unit, which is described below. Here, audio processing device 21 is described. Audio processing device 21 sets the noise uttered by the driver hm1 as a target component. Here, setting a noise as a target component is synonymous with setting a noise as an audio signal for detection. Audio processing device 21 outputs an audio signal as an output signal, which is obtained by suppressing a crosstalk component from the audio signal picked up by microphone MC1.Here, the crosstalk component is a noise component that contains a sound from an occupant other than the occupant who is making the sound set as the target component.

[0023] As in Fig. As shown in Figure 2, the audio processing device 21 includes an audio input module 29, a noise detection module 30, a filter unit F1 containing a plurality of adaptive filters, a control module 28 controlling filter coefficients of the plurality of adaptive filters, and an addition module 27.

[0024] An audio signal of the sound picked up by microphones MC1, MC2, MC3, and MC4 is input to audio input module 29. In other words, each microphone MC1, MC2, MC3, and MC4 outputs a signal based on the recorded sound to audio input module 29. Microphone MC1 outputs an audio signal A to audio input module 29. Audio signal A contains the sound of the driver hm1 and the background noise, which contains the sound of an occupant other than the driver hm1. Here, in the audio processing device 21, the sound of the driver hm1 is a target component, and the background noise, which contains the sound of the occupant other than the driver hm1, is a crosstalk component. Microphone MC1 corresponds to the first microphone. The sound picked up by microphone MC1 corresponds to the first audio signal.The sound of the driver hm1 corresponds to the first audio component. The sound of the other occupant (hm1) corresponds to the second audio component. Audio signal A corresponds to the first signal. Microphone MC2 outputs an audio signal B to audio input module 29. Audio signal B contains the sound of occupant hm2 and background noise, which includes the sound of an occupant other than occupant hm2. Microphone MC3 outputs an audio signal C to audio input module 29. Audio signal C contains the sound of occupant hm3 and background noise, which includes the sound of an occupant other than occupant hm3. Microphone MC4 outputs an audio signal D to audio input module 29. Audio signal D contains the sound of occupant hm4 and background noise, which includes the sound of an occupant other than occupant hm4.Microphones MC2, MC3, and MC4 correspond to the second microphone. The sound picked up by microphones MC2, MC3, and MC4 corresponds to the second audio signal. Audio signals B, C, and D correspond to the second signal. Audio input module 29 outputs audio signals A, B, C, and D. Audio input module 29 corresponds to a receiver.

[0025] In the present embodiment, the audio processing device 21 includes an audio input module 29 into which audio signals from all microphones are input. However, it can also include the audio input module 29 into which a corresponding audio signal for each microphone is input. For example, the setup can be such that an audio signal of the sound picked up by microphone MC1 is input to an audio input module corresponding to microphone MC1, an audio signal of the sound picked up by microphone MC2 is input to another audio input module corresponding to microphone MC2, an audio signal of the sound picked up by microphone MC3 is input to another audio input module corresponding to microphone MC3, and an audio signal of the sound picked up by microphone MC4 is input to another audio input module corresponding to microphone MC4.

[0026] The audio signals A, B, C, and D, output by the audio input module 29, are fed into the noise detection module 30. The noise detection module 30 determines whether the individual audio signals contain uncorrelated noise. Uncorrelated noise is noise that exhibits no correlation between the audio signals. Examples of uncorrelated noise include wind noise, circuit noise, or contact noise from a microphone. Uncorrelated noise is also referred to as non-acoustic noise. For example, if the intensity of a particular audio signal is greater than or equal to a predefined value, the noise detection module 30 determines that the audio signal contains uncorrelated noise.Alternatively, the noise detection module 30 can compare the intensity of a specific audio signal with the intensity of another audio signal and, if the intensity of the specific audio signal is greater than the intensity of the other audio signal by a predetermined value or more, determine that the audio signal contains uncorrelated noise. Furthermore, the noise detection module 30 can determine, based on vehicle information, that a specific audio signal contains uncorrelated noise.For example, the noise detection module 30 can receive vehicle information such as the vehicle speed and the open / closed status of the window. If the vehicle speed is greater than or equal to a certain value and the rear window is open, it can determine that audio signals C and D contain uncorrelated noise. The noise detection module 30 then outputs the result of its determination—whether the individual audio signals contain uncorrelated noise—to the control module 28. The noise detection module 30 outputs this result to the control module 28, for example, as a label. This label indicates a value of "1" or "0" for each audio signal."1" means that the audio signal contains uncorrelated noise, and "0" means that the audio signal does not contain uncorrelated noise. For example, in a case where it is determined that audio signals A and B do not contain uncorrelated noise, and audio signals C and D do contain uncorrelated noise, the noise detection module 30 outputs a code "0, 0, 1, 1" as a determination result to the control module 28. After determining whether the uncorrelated noise is present, the noise detection module 30 outputs audio signal A to the addition module 27 and outputs audio signals B, C, and D to the filter unit F1. Here, the noise detection module 30 corresponds to the determination module.

[0027] In the present embodiment, the audio processing device 21 includes a noise detection module 30 into which all audio signals are input. However, it can also include the noise detection module 30 into which a corresponding audio signal is input for each audio signal. For example, audio signal A can be input to a noise detection module 301, audio signal B can be input to a noise detection module 302, audio signal C can be input to a noise detection module 303, and audio signal D can be input to a noise detection module 304.

[0028] The filter unit F1 contains an adaptive filter F1A, an adaptive filter F1B, and an adaptive filter F1C. An adaptive filter is a filter that has a function for modifying properties in a signal processing process. The filter unit F1 is used for processing to suppress a crosstalk component that is distinct from the sound of the driver hm1 and is present in the sound picked up by the microphone MC1. In the present embodiment, the filter unit F1 contains three adaptive filters, but the number of adaptive filters is appropriately set based on the number of input audio signals and the processing effort of the crosstalk suppression process. The crosstalk suppression method is described in detail below. Here, the filter unit F1 corresponds to a first filter unit.

[0029] Audio signal B is input as a reference signal into adaptive filter F1A. Adaptive filter F1A outputs a pass-through signal PB based on a filter coefficient CB and audio signal B. Audio signal C is input as a reference signal into adaptive filter F1B. Adaptive filter F1B outputs a pass-through signal PC based on a filter coefficient CC and audio signal C. Audio signal D is input as a reference signal into adaptive filter F1C. Adaptive filter F1C outputs a pass-through signal PD based on a filter coefficient CD and audio signal D. Filter unit F1 adds the pass-through signal PB, the pass-through signal PC, and the pass-through signal PD and outputs them. In the present embodiment, adaptive filters F1A, F1B, and F1C are implemented by a processor that executes a program stored in memory.The adaptive filter F1A, the adaptive filter F1B and the adaptive filter F1C can have separate hardware configurations that are spatially separated from each other.

[0030] Here is an overview of how the adaptive filter works. The adaptive filter is a filter used to suppress a crosstalk component. For example, in a case where a least mean square (LMS) is used as the filter coefficient update algorithm, the adaptive filter is a filter that minimizes a cost function defined by the mean square of the error signal. The error signal here is a difference between the output signal and the target component.

[0031] Here, an FIR filter (Finite Impulse Response Filter) is shown as an example of an adaptive filter. Other types of adaptive filters can also be used. For example, an IIR filter (Infinite Impulse Response Filter) could be used.

[0032] In a case where the audio processing device 21 uses an FIR filter as the adaptive filter, the error signal, which is the difference between the output signal of the audio processing device 21 and the target component, is expressed by the following formula (1). e(n)=d(n)−∑i=1i−1wix(n−i)

[0033] Here, n is time, e(n) is an error signal, d(n) is a target component, wi is a filter coefficient, x(n) is a reference signal, and 1 is a sampling length. With a larger sampling length 1, the adaptive filter can accurately reproduce the acoustic properties of the audio signal. In a case where there is no reverberation, the sampling length 1 can be 1. For example, the sampling length 1 is set to a constant value. In a case where the target component is the driver's noise hm1, the reference signal x(n) is, for example, audio signal B, audio signal C, and audio signal D.

[0034] The addition module 27 generates an output signal by subtracting the subtraction signal from the target audio signal output by the audio input module 29. In the present embodiment, the subtraction signal is a signal obtained by adding the pass-through signal PB, the pass-through signal PC, and the pass-through signal PD output by the filter unit F1. The addition module 27 outputs the resulting signal to the control module 28.

[0035] The control module 28 outputs the output signal provided by the addition module 27. The output signal of the control module 28 is input into the speech recognition machine 40. Alternatively, the output signal can also be input directly from the control module 28 into the electronic device 50. In the case where an output signal is input directly from the control module 28 into the electronic device 50, the control module 28 and the electronic device 50 can be connected by wire or wirelessly. For example, the electronic device 50 can be a mobile device, and an output signal can be input directly from the control module 28 into the mobile device via a wireless communication network. The output signal input into the mobile device can be output as a sound via a speaker of the mobile device.

[0036] Furthermore, the control module 28 references the output signal issued by the addition module 27 and the designation as the determination result issued by the noise detection module 30 and updates the filter coefficient of each adaptive filter.

[0037] First, based on the determination result, control module 28 determines an adaptive filter as the update target of the filter coefficient. More precisely, control module 28 sets the adaptive filter, into which the audio signal is fed (which noise detection module 30 has determined does not contain uncorrelated noise), as the update target of the filter coefficient. Furthermore, control module 28 does not set the adaptive filter, into which the audio signal is fed (which noise detection module 30 has determined does not contain uncorrelated noise), as the update target of the filter coefficient.For example, in a case where a signal “0, 0, 1, 1” is received from the noise detection module 30, the control module 28 determines that audio signal A and audio signal B contain no uncorrelated noise, and audio signal C and audio signal D contain uncorrelated noise. The control module 28 then sets the adaptive filter F1A as the update target of the filter coefficient and does not set the adaptive filters F1B and F1C as update targets of the filter coefficient. In this case, adaptive filter F1A corresponds to a second adaptive filter, and adaptive filters F1B and F1C correspond to a first adaptive filter.

[0038] Then the control module 28 updates the filter coefficient in such a way that the value of the error signal in formula (1) approaches 0 for the adaptive filter set as the update target of the filter coefficient.

[0039] The following describes the updating of the filter coefficient in a case where the LMS is used as the update algorithm. In a case where a filter coefficient w(n) at time n is updated to a filter coefficient w(n + 1) at time n + 1, the relationship between w(n + 1) and w(n) is expressed by the following formula (2). w(n+1)=w(n)−αx(n)e(n)

[0040] Here, α is a correction coefficient of the filter coefficient. An expression αx(n)e(n) corresponds to the update amount.

[0041] It should be noted that the algorithm is not limited to LMS at the time of the filter coefficient update and other algorithms can also be used. For example, an algorithm such as Independent Component Analysis (ICA) or Normalized Least Mean Square (NLMS) can be used.

[0042] At the time of the filter coefficient update, the control module 28 sets the intensity of the input reference signal for the adaptive filter, which is not set as the update target of the filter coefficient, to zero. For example, if a label “0, 0, 1, 1” is received from the noise detection module 30, the control module 28 leaves the audio signal B, which is input to the adaptive filter F1A as the reference signal and is input with the intensity output by the noise detection module 30, unchanged and sets the intensities of the audio signal C, which is input to the adaptive filter F1B as the reference signal, and the audio signal D, which is input to the adaptive filter F1C as the reference signal, to zero.Here, "setting the intensity of the reference signal input to the adaptive filter to zero" means suppressing the intensity of the reference signal input to the adaptive filter to a value close to zero. Furthermore, "setting the intensity of the reference signal input to the adaptive filter to zero" also means setting the reference signal not to be input to the adaptive filter. In a case where the intensity of the audio signal input as the reference signal is not set to zero, the audio signal containing uncorrelated noise will be fed into the adaptive filter, even if it is not set as the update target of the filter coefficient. For example, if an audio signal containing strong wind noise as uncorrelated noise is used as a reference signal, it can be difficult to accurately preserve a target component.Setting the intensity input to the adaptive filter to zero for an audio signal containing uncorrelated noise is equivalent to not using that signal as a reference signal. Even if the crosstalk component contains uncorrelated noise, the target component can therefore be accurately preserved. Adaptive filtering cannot be performed in an adaptive filter where the intensity of the input reference signal is set to zero. Consequently, the processing overhead of crosstalk suppression can be reduced when using the adaptive filter.

[0043] Control module 28 then updates the filter coefficient only for the adaptive filter that is set as the filter coefficient update target, and does not update the filter coefficient for the adaptive filter that is not set as the filter coefficient update target. As a result, the processing overhead of crosstalk suppression processing can be reduced when using the adaptive filter.

[0044] For example, consider a case where the target seat is the driver's seat and there is no utterance from the driver hm1, but utterances from occupants hm2, hm3, and hm4 are present. At this time, the utterance of the occupant other than the driver hm1 is transmitted into the audio signal of the noise picked up by microphone MC1. In other words, the audio signal A contains the crosstalk component. The audio processing device 21 can suppress the crosstalk component and update the adaptive filter to minimize the error signal. In this case, the error signal ideally becomes a silent signal, since there is no utterance from the driver's seat. Furthermore, in the case described above, if there is an utterance from the driver hm1, the utterance from the driver hm1 is transmitted to a microphone other than microphone MC1.Furthermore, in this case, the utterance by driver hm1 is not suppressed by the processing of the audio processing device 21. The reason for this is that the utterance by driver hm1, which is contained in audio signal A, precedes the utterance by driver hm1 contained in the other audio signals. This is due to the law of causality. Therefore, the audio processing device 21 can reduce the crosstalk component contained in audio signal A by updating the adaptive filter to minimize the error signal, regardless of whether the audio signal of the target component is present or not.

[0045] In the present embodiment, the functions of the audio input module 29, the noise detection module 30, the filter unit F1, the control module 28, and the addition module 27 are implemented by a processor that executes a program stored in memory. Alternatively, the audio input module 29, the noise detection module 30, the filter unit F1, the control module 28, and the addition module 27 can be implemented by separate hardware.

[0046] Although the audio processing device 21 has been described above, the audio processing devices 22, 23, and 24 also have essentially similar configurations, with the exception of the filter unit. The audio processing device 22 sets the noise uttered by occupant hm2 as a target component. The audio processing device 22 outputs an audio signal as an output signal, which is obtained by suppressing a crosstalk component from the audio signal picked up by microphone MC2. Therefore, the audio processing device 22 differs from the audio processing device 21 in that it includes the filter unit into which audio signal A, audio signal C, and audio signal D are input. Similarly, the audio processing device 23 sets the noise uttered by occupant hm3 as a target component.The audio processing device 23 outputs an audio signal obtained by suppressing a crosstalk component from the audio signal picked up by microphone MC3. Therefore, the audio processing device 23 differs from the audio processing device 21 in that it includes the filter unit into which audio signal A, audio signal B, and audio signal D are input. The audio processing device 24 uses the noise uttered by occupant hm4 as a target component. The audio processing device 24 outputs an audio signal obtained by suppressing a crosstalk component from the audio signal picked up by microphone MC4. Therefore, the audio processing device 24 differs from the audio processing device 21 in that it includes the filter unit into which audio signal A, audio signal B, and audio signal C are input.

[0047] Fig. Figure 3 is a flowchart illustrating an operating procedure of the audio processing device 21. First, audio signals A, B, C, and D are input into the audio input module 29 (S1). Next, the noise detection module 30 determines whether each audio signal contains uncorrelated noise (S2). The noise detection module 30 outputs this result as a marker to the control module 28. If none of the audio signals contains uncorrelated noise, the filter unit F1 generates a subtraction signal as follows (S3). The adaptive filter F1A passes audio signal B and outputs the pass-through signal PB. The adaptive filter F1B passes audio signal C and outputs the pass-through signal PC. The adaptive filter F1C passes audio signal D and outputs the pass-through signal PD.Filter unit F1 adds the pass-through signal PB, the pass-through signal PC, and the pass-through signal PD and outputs it as a subtraction signal. Addition module 27 subtracts the subtraction signal from the audio signal A, generating and outputting an output signal (S4). The output signal is input to control module 28 and output by control module 28. Next, control module 28 references the label as the determination result output by noise detection module 30 and updates the filter coefficients of adaptive filter F1A, adaptive filter F1B, and adaptive filter F1C based on the output signal in such a way as to maximize the target component contained in the output signal (S5). Then, audio processing device 21 repeats step S1.

[0048] In a case where, in step S2, it is determined that any one of the audio signals contains uncorrelated noise, the noise detection module 30 determines whether the audio signal containing the uncorrelated noise is a target component or not (S6). Specifically, it determines whether the audio signal containing the uncorrelated noise is audio signal A. If the audio signal containing the uncorrelated noise is the target component, the control module 28 sets the intensity of audio signal A to zero and outputs audio signal A as the output signal (S7). At this time, the control module 28 does not update the filter coefficients of adaptive filter F1A, adaptive filter F1B, and adaptive filter F1C. Then, the audio processing device 21 performs step S1 again.

[0049] In step S6, if the audio signal containing uncorrelated noise is not the target component, control module 28 sets the intensity of the audio signal containing uncorrelated noise, which is input to filter unit F1, to zero. For example, consider a case where audio signals C and D contain uncorrelated noise, and audio signal B does not. In this case, control module 28 sets the intensities of audio signals C and D, which are input to filter unit F1, to zero and does not change the intensity of audio signal B. Then, filter unit F1 generates a subtraction signal through an operation similar to that in step S3 (S8). Similar to step S4, addition module 27 subtracts the subtraction signal from audio signal A, generates an output signal, and outputs it (S9).Next, the control module 28 updates the filter coefficient of the adaptive filter, into which the signal containing no uncorrelated noise is input, based on the output signal, in such a way as to maximize the target component contained in the output signal (S10). For example, consider a case in which audio signal C and audio signal D contain uncorrelated noise, and audio signal B does not. In this case, the control module 28 updates the filter coefficient of adaptive filter F1A but does not update the filter coefficients of adaptive filters F1B and F1C. Then, the audio processing device 21 performs step S1 again.

[0050] As described above, in the audio processing system 5 according to the first embodiment, a plurality of audio signals are acquired by a plurality of microphones, and a subtraction signal, generated using an adaptive filter, is subtracted from a specific audio signal using another audio signal as a reference signal. Thus, the sound of a specific speaker is obtained with high accuracy. In the first embodiment, when a subtraction signal is generated using the adaptive filter, the intensity of the audio signal containing uncorrelated background noise and fed into the adaptive filter is set to zero. For example, there is a case where wind blows into the area of ​​the rear seat, and a strong wind noise is picked up by a microphone near the rear seat.If the audio signal received from the rear seat is used as the reference signal, it can be difficult to identify the sound of a specific speaker. However, in the present embodiment, by setting the intensity of the audio signal containing uncorrelated noise, which is fed into the adaptive filter, to zero, the audio signal of the target component can be accurately preserved even if the uncorrelated noise occurs at a seat other than the target seat. Furthermore, in the first embodiment, the filter coefficient is not updated for the adaptive filter into which the audio signal containing uncorrelated noise is fed. As a result, the processing overhead for suppressing the crosstalk component can be reduced.

[0051] It should be noted that in a case where each microphone is part of a microphone array, the microphone array can exhibit a directional effect towards a specific occupant at the time of sound recording, thus capturing the sound, i.e., performing beamforming. As a result, the signal-to-noise ratio of the audio signal input to the respective microphone is improved. This, in turn, enhances the accuracy of the crosstalk suppression performed by the audio processing system 5.

[0052] Fig. Figure 4 represents an output result of the audio processing device 20. Fig. 4 represents output results of the audio processing devices 20 when a strong wind noise is recorded by the microphone MC3 and the microphone MC4 in a state in which utterances are made by the driver hm1, the occupant hm2, the occupant hm3 and the occupant hm4. Fig. 4(a), Fig. 4(b), Fig. 4(c) and Fig. 4(d) represent output results of the audio processing devices 20 when the input intensities of audio signal C and audio signal D are not set to zero and the updating of adaptive filter F1B and adaptive filter F1C is not stopped. Fig. 4(a) corresponds to the output result of the audio processing device 21, Fig. 4(b) corresponds to the output result of the audio processing device 22, Fig. 4(c) corresponds to the output result of the audio processing device 23 and Fig. 4(d) corresponds to the output result of the audio processing device 24. Fig. 4(e), Fig. 4(f), Fig. 4(g) and Fig. 4(h) represent output results of the audio processing devices 20 when the input intensities of audio signal C and audio signal D are set to zero and the updating of adaptive filter F1B and adaptive filter F1C is stopped. Fig. 4(e) corresponds to the output result of the audio processing device 21, Fig. 4(f) corresponds to the output result of the audio processing device 22, Fig. 4(g) corresponds to the output result of the audio processing device 23 and Fig. 4(h) corresponds to the output result of the audio processing device 24.

[0053] Out of Fig. 4(a), Fig. 4(b), Fig. 4(c) and Fig. 4(d) It is evident that when using the audio signal containing uncorrelated noise as the reference signal, the output signals of the audio processing device 21 and the audio processing device 22 are signals containing a very large amount of noise. Even if the output signals of the audio processing device 21 and the audio processing device 22 are used for speech recognition in this case, the recognition accuracy is considered to be low. On the other hand, it is evident that the output signals of the audio processing device 21 and the audio processing device 22, which are used in Fig. 4(e) and Fig. 4(f) show less background noise than those in Fig. 4(a) and Fig. 4(b). Therefore, the output signals of the audio processing device 21 and the audio processing device 22 can, in this case, be noises that are detected with high accuracy. As in Fig. 4(g) and Fig. In addition, as shown in Figure 4(h), the intensities of the output signals of the audio processing device 23 and the audio processing device 24 are zero. Second embodiment

[0054] Fig. Figure 5 is a diagram illustrating a schematic diagram of an audio processing system 5A according to a second embodiment. The audio processing system 5A according to the second embodiment differs from the audio processing system 5 according to the first embodiment in that it includes an audio processing device 20A instead of the audio processing device 20. The audio processing device 20A according to the second embodiment differs from the audio processing device 20 according to the first embodiment in that it includes an additional filter unit. In the present embodiment, the audio processing system 5A includes a plurality of audio processing devices 20A corresponding to the respective microphones.More precisely, the audio processing system 5A includes an audio processing device 21A, an audio processing device 22A, an audio processing device 23A, and an audio processing device 24A. Audio processing device 20A is described below with reference to... Fig. 6 and Fig. 7 described. The same components and operating procedures as those described in the first embodiment are designated by the same reference numerals, and the description is omitted or simplified.

[0055] Fig. Figure 6 is a block diagram illustrating the setup of the audio processing device 21A. The audio processing devices 21A, 22A, 23A, and 24A have similar configurations and functions, with the exception of a portion of a filter unit described below. Here, the audio processing device 21A is described. The audio processing device 21A sets the noise uttered by the driver hm1 as a target. The audio processing device 21A outputs an audio signal as an output signal, which is obtained by suppressing a crosstalk component from the audio signal picked up by microphone MC1.

[0056] The audio processing device 21A includes the audio input module 29, a noise detection module 30A, the filter unit F1 which contains a plurality of adaptive filters, a filter unit F2 which contains one or more adaptive filters, a control module 28A which controls a filter coefficient of the adaptive filter of the filter unit F1, and an addition module 27A.

[0057] The filter unit F2 contains one or more adaptive filters. In the present embodiment, the filter unit F2 contains one adaptive filter F2A. The filter unit F2 is used for processing to suppress a crosstalk component that is different from the driver's noise hm1 and is present in the noise picked up by the microphone MC1. The number of adaptive filters contained in the filter unit F2 is less than the number of adaptive filters contained in the filter unit F1. In the present embodiment, the filter unit F2 contains one adaptive filter, but the number of adaptive filters is appropriately set based on the number of input audio signals and the processing overhead of the crosstalk suppression. The crosstalk suppression method is described in detail below. Here, the filter unit F2 corresponds to a second filter unit.

[0058] The audio signal B is input as a reference signal into the adaptive filter F2A. The adaptive filter F2A outputs a pass signal PB2 based on a unique filter coefficient CB2 and the audio signal B. In the present embodiment, the function of the adaptive filter F2A is implemented by software processing. The adaptive filter F2A can have a separate hardware configuration that is physically separate from each adaptive filter in the filter unit F1. Here, the adaptive filter F1A corresponds to the second adaptive filter, the adaptive filters F1B and F1C correspond to the first adaptive filter, and the adaptive filter F2A corresponds to a third adaptive filter. Furthermore, the audio signal B corresponds to a third signal.

[0059] The adaptive filter F2A can be an FIR filter, an IIR filter, or another type of adaptive filter. It is desirable for the adaptive filter F2A to be the same type of adaptive filter as the adaptive filters F1A, F1B, and F1C, as this reduces processing overhead compared to using different types of adaptive filters. Here, we describe a case where an FIR filter is used as the adaptive filter F2A.

[0060] The addition module 27A generates an output signal by subtracting the subtraction signal from the target audio signal output by the audio input module 29. In the present embodiment, the subtraction signal is a signal obtained by adding the pass-through signal PB, the pass-through signal PC and the pass-through signal PD output by the filter unit F1, or the pass-through signal PB2 output by the filter unit F2. The addition module 27A outputs the output signal to the control module 28A.

[0061] Control module 28A outputs the output signal that is output by the adder module 27A. The output signal of control module 28A is input to the speech recognition machine 40. Alternatively, the output signal can be input directly from control module 28A to the electronic device 50. In the case where an output signal is input directly from control module 28A to the electronic device 50, the control module 28A and the electronic device 50 can be connected by wire or wirelessly. For example, the electronic device 50 can be a mobile device, and an output signal can be input directly from control module 28A to the mobile device via a wireless communication network. The output signal input to the mobile device can be played back as sound through a speaker on the mobile device.

[0062] In addition to the function of the noise detection module 30, the noise detection module 30A determines whether the individual audio signals contain an audio component due to an utterance. The noise detection module 30A outputs the result of this determination to the control module 28. The noise detection module 30A outputs this result to the control module 28, for example, as a label. The label indicates a value of "1" or "0" for each audio signal. "1" means that the audio signal contains an audio component due to an utterance, and "0" means that the audio signal does not contain an audio component due to an utterance.In a case where it is determined that audio signals A and B contain an audio component due to an utterance, and audio signals C and D do not, the noise detection module 30 outputs, for example, a label "1, 1, 0, 0" as a determination result to the control module 28A. Here, the audio component due to an utterance corresponds to a first component derived from the utterance. Then, based on the result of the determination, the control module 28A determines whether the individual audio signals contain an audio component due to an utterance, which is used from filter unit F1 and filter unit F2 to generate the subtraction signal.For example, there is a case in which an adaptive filter, into which the audio signal determined to be free of the utterance component is input, is contained in filter unit F1 and not in filter unit F2. In this case, the control module 28A determines that the subtraction signal is to be generated using filter unit F2. The audio processing devices 21A may contain an utterance determination module, separate from the noise detection module 30A, which determines whether the individual audio signals contain an utterance component. In this case, the utterance determination module is connected between the audio input module 29 and the noise detection module 30A, or between the noise detection module 30A and filter units F1 and F2.The function of the utterance determination module is implemented, for example, by a processor that executes a program stored in memory. The function of the utterance determination module can also be implemented by hardware.

[0063] For example, consider a case where audio signal B contains an audio component due to occupant hm2's utterance, audio signal C contains no audio component due to occupant hm3's utterance, and audio signal D contains no audio component due to occupant hm4's utterance. At this time, an adaptive filter, into which audio signals C and D are input, is contained in filter unit F1 and not in filter unit F2. The filter coefficient of each adaptive filter contained in filter unit F1 is updated, for example, in a case where the reference signal is input to each of the adaptive filters, to minimize the error signal. On the other hand, the filter coefficient of adaptive filter F2A, contained in filter unit F2, has a unique value, assuming that only audio signal B is used as the reference signal.Therefore, when comparing a case in which only the audio signal B is input into each filter unit as the reference signal, there is a possibility that the error signal can be made smaller using filter unit F2 than using filter unit F1.

[0064] In a case where the number of adaptive filters contained in filter unit F2 is less than the number of adaptive filters contained in filter unit F1, the processing effort can be reduced by generating the subtraction signal using filter unit F2 compared to generating a subtraction signal using filter unit F1.

[0065] Alternatively, an adaptive filter, into which the audio signal, which the noise detection module 30 determines contains uncorrelated noise, is fed, is contained in filter unit F1 and not in filter unit F2. In this case as well, control module 28A determines that a subtraction signal is to be generated using filter unit F2.

[0066] The filter coefficient of each adaptive filter contained in filter unit F1 is updated, for example, when the reference signal is input to each of the adaptive filters, in order to minimize the error signal. On the other hand, the filter coefficient of adaptive filter F2A, contained in filter unit F2, has a unique value when only audio signal B is used as the reference signal. In a case where audio signal C and audio signal D contain uncorrelated noise, the intensities of audio signal C and audio signal D input to filter unit F1 are set to zero.In this case, the error signal can in some cases be reduced by using filter unit F2, which uses only audio signal B as the reference signal, instead of using filter unit F1, which assumes that all of the audio signals from B, C and D are used as the reference signal.

[0067] Furthermore, the control module 28A updates the filter coefficient of each adaptive filter of filter unit F1 in a case where a subtraction signal is generated using filter unit F1, based on the output signal output by the addition module 27A and the determination result output by the noise detection module 30. The method for updating the filter coefficient is similar to that of the first embodiment.

[0068] In the present embodiment, the functions of the audio input module 29, the noise detection module 30, the filter unit F1, the filter unit F2, the control module 28A, and the addition module 27A are implemented by a processor that executes a program stored in memory. Alternatively, the audio input module 29, the noise detection module 30, the filter unit F1, the filter unit F2, the control module 28A, and the addition module 27A can be implemented by separate hardware.

[0069] Fig.Figure 7 is a flowchart illustrating an operating procedure of the audio processing device 21A. First, audio signal A, audio signal B, audio signal C, and audio signal D are input into the audio input module 29 (S11). Next, the noise detection module 30 determines whether each audio signal contains uncorrelated noise (S12). If none of the audio signals contains uncorrelated noise, the control module 28A determines which filter unit is used to generate a subtraction signal (S13). If the control module 28A determines that filter unit F1 is to be used, filter unit F1 generates and outputs a subtraction signal, similar to step S3 of the first embodiment (S14). The addition module 27A subtracts the subtraction signal from audio signal A and generates and outputs an output signal (S15).The output signal is input to and output by control module 28. Next, control module 28 updates the filter coefficients of adaptive filter F1A, adaptive filter F1B, and adaptive filter F1C based on the output signal in such a way as to maximize the target component contained in the output signal (S16). Then, audio processing device 21A performs step S11 again.

[0070] In step S13, if control module 28A specifies that filter unit F2 is to be used, the filter unit F2 generates a subtraction signal as follows (S17). The adaptive filter F2A passes the audio signal B and outputs the pass-through signal PB2. Filter unit F2 outputs the pass-through signal PB2 as a subtraction signal. The addition module 27A subtracts the subtraction signal from the audio signal A, generating and outputting an output signal (S18). The output signal is input to control module 28 and output by control module 28. The audio processing device 21 then performs step S11 again.

[0071] In a case where, in step S2, it is determined that any one of the audio signals contains uncorrelated noise, the noise detection module 30 determines whether the audio signal containing the uncorrelated noise is a target component (S19). Specifically, it determines whether the audio signal containing the uncorrelated noise is audio signal A. If the audio signal containing the uncorrelated noise is the target component, the control module 28 sets the intensity of audio signal A to zero and outputs audio signal A as the output signal (S20). At this time, the control module 28 does not update the filter coefficients of adaptive filter F1A, adaptive filter F1B, and adaptive filter F1C. Then, the audio processing device 21A performs step S11 again.

[0072] If the audio signal containing uncorrelated noise is not the target component, control module 28A determines in step S19 which filter unit is used to generate a subtraction signal (S21). In a case where control module 28A determines that filter unit F1 is to be used, control module 28 sets the intensity of the audio signal containing uncorrelated noise and input to filter unit F1 to zero. For example, consider a case where audio signal B contains uncorrelated noise, and audio signals C and D do not. In this case, control module 28 sets the intensity of audio signal B, input to filter unit F1, to zero and does not change the intensities of audio signals C and D.The filter unit F1 then generates a subtraction signal through an operation similar to that in step S3 of the first embodiment (S22). Similar to step S4 of the first embodiment, the addition module 27A subtracts the subtraction signal from the audio signal A, generating and outputting an output signal (S23). Next, based on the output signal, the control module 28A updates the filter coefficient of the adaptive filter, into which the signal containing no uncorrelated noise is input, such that the target component contained in the output signal is maximized (S24). For example, consider a case where audio signal B contains uncorrelated noise, while audio signals C and D do not.In this case, control module 28 updates the filter coefficients of adaptive filter F1B and adaptive filter F1C, but does not update the filter coefficient of adaptive filter F1A. Then, audio processing device 21A performs step S11 again.

[0073] In step S21, if the control module 28A determines that filter unit F2 is to be used, the filter unit F2 generates a subtraction signal similar to that in step S17 (S25). The addition module 27A subtracts the subtraction signal from the audio signal A, generating and outputting an output signal (S26). The output signal is input to and output by the control module 28. The audio processing device 21A then performs step S11 again.

[0074] As described above, in the audio processing system 5A according to the second embodiment, the audio signal of the target component can also be accurately preserved, similar to audio processing system 5, even if an uncorrelated noise occurs at a different location than the target location. Furthermore, in the first embodiment, the filter coefficient is not updated for the adaptive filter into which the audio signal containing the uncorrelated noise is fed. Consequently, the processing overhead for suppressing the crosstalk component can be reduced.

[0075] Furthermore, the audio processing system 5A additionally includes the filter unit F2, which has a smaller number of adaptive filters than the filter unit F1, and the control module 28A determines which of the filters from filter unit F1 and filter unit F2 is to be used. As a result, the processing overhead can be reduced compared to a case where the subtraction signal is always generated using filter unit F1.

[0076] In the present embodiment, the case is described in which the filter unit F2 contains an adaptive filter with a unique filter coefficient, but the filter unit F2 can also contain two or more adaptive filters. Furthermore, the coefficient of the adaptive filter contained in the filter unit F2 need not be unique and can be controlled by the control module 28A. In a case in which the filter unit F2 contains an adaptive filter capable of controlling the filter coefficient, the control module 28A can update the filter coefficient of the adaptive filter, into which the audio signal, which contains no uncorrelated noise, is input, after step S18 or after step S26. Explanation of reference symbols 5 Audio processing system 10 vehicles 20, 21, 22, 23, 24 Audio processing device 27 Addition module 28 Control module 29 Audio input module 30 Noise Detection Module F1 filter unit F1A, F1B, F1C Adaptive Filter 40 speech recognition machine 50 Electronic device

Claims

[1] Audio processing system (5), comprising: a first microphone (MC1) designed to capture a first audio signal containing a first audio component and to output a first signal based on the first audio signal; one or more microphones (MC2-MC4), each of the one or more microphones being designed to detect an audio signal containing an audio component different from the first audio component and to output a microphone signal based on the audio signal; one or more adaptive filters (F1A-F1C) designed to receive the microphone signals from the one or more microphones (MC2-MC4) and to output pass-through signals based on the microphone signals; a determination module (30) designed to determine whether the individual microphone signals contain uncorrelated noise, wherein the noise is one that has no correlation between the audio signals; a control module (28; 28A) designed to control one or more filter coefficients of one or more adaptive filters (F1A-F1C); and an addition module (27: 27A) designed to subtract a subtraction signal from the first signal based on the through signals, wherein the one or more microphones (MC2-MC4) include a second microphone designed to capture a second audio signal containing a second audio component different from the first audio component, and to output a second signal based on the second audio signal, and If the determination module (30) determines that the second signal contains the uncorrelated noise, the control module (28; 28A) is designed to set the level of the second signal, which is input into the appropriate adaptive filter, to zero. [2] Audio processing system according to claim 1, wherein the one or more microphones (MC2-MC4) include a third microphone designed to capture a third audio signal containing a third audio component that is different from the first and second audio components, and to output a third signal based on the third audio signal, the one or more adaptive filters (F1A-F1C) comprise a first adaptive filter into which the second signal is input, and a second adaptive filter into which the third signal is input, and, If the determination module (30) determines that the second signal contains the uncorrelated noise and the third signal does not contain the uncorrelated noise, the control module (28; 28A) is designed to change the filter coefficient of the second adaptive filter without changing the filter coefficient of the first adaptive filter. [3] Audio processing system (5) according to claim 1, wherein the one or more microphones (MC2-MC4) include a third microphone designed to capture a third audio signal containing a third audio component that is different from the first and second audio components, and to output a third signal based on the third audio signal, the audio processing system (5) further includes: a first filter unit (F1) containing one or more adaptive filters (F1A-F1C), wherein the one or more adaptive filters (F1A-F1C) comprise a first adaptive filter into which the second signal is input, and a second adaptive filter into which the third signal is input, and a second filter unit (F2) containing one or more adaptive filters (F2A), wherein the second filter unit (F2) contains a third adaptive filter that is different from the second adaptive filter and is designed to receive the third signal and output a first pass signal based on the third signal, a number of one or more adaptive filters contained in the second filter unit is less than a number of adaptive filters contained in the first filter unit and The control module (28; 28A) is designed to determine, based on the second signal and the third signal, which from the first filter unit and the second filter unit is used to generate the subtraction signal. [4] Audio processing system (5) according to claim 3, wherein the second filter unit contains only the third adaptive filter and, If it is determined that the second signal does not contain a first component derived from an utterance, and the third signal does contain the first component, the control module (28; 28A) is designed to generate the subtraction signal by using the second filter unit. [5] Audio processing system (5) according to any one of claims 1 to 4, wherein the determination module (30) is designed to determine that the microphone signal contains the uncorrelated noise when an intensity of the microphone signal is greater than or equal to a predetermined value. [6] Audio processing system (5) according to any one of claims 1 to 4, wherein the one or more microphones (MC2-MC4) comprise a first target microphone designed to output a first microphone signal and a second target microphone designed to output a second microphone signal, wherein the second target microphone is different from the first target microphone, and the determination module (30) is designed to determine that the first microphone signal contains the uncorrelated noise when the intensity of the first microphone signal is greater than the intensity of the second microphone signal by a predetermined value or more. [7] Audio processing system (5) according to any one of claims 1 to 4, wherein the determination module (30) is designed to determine, based on vehicle information, that the microphone signal contains the uncorrelated background noise. [8] Audio processing device comprising: a first receiving unit designed to receive a first signal based on a first audio signal containing a first audio component; one or more receiving units, each of the one or more receiving units being designed to receive a microphone signal based on an audio signal containing an audio component that is different from the first audio component; one or more adaptive filters (F1A-F1C) designed to each receive the microphone signals from the one or more receiving units and to output pass-through signals based on the microphone signals; a determination module (30) designed to determine whether the individual microphone signals contain uncorrelated background noise; a control module (28; 28A) designed to control one or more filter coefficients of one or more adaptive filters (F1A-F1C); and an addition module (27: 27A) designed to subtract a subtraction signal from the first signal based on the through signals, wherein the one or more receiving units include a second receiving unit designed to receive a second signal based on a second audio signal containing a second audio component that is different from the first audio component, and, If the determination module (30) determines that the second signal contains the uncorrelated noise, the control module (28; 28A) is designed to set the level of the second signal, which is input into the appropriate adaptive filter, to zero. [9] Audio processing device according to claim 8, wherein which are designed to receive one or more receiving units based on a third audio signal containing a third audio component that is different from the first audio component and the second audio component, which include one or more filters (F1A-F1C) comprising a first adaptive filter from which the second signal is output, and a second adaptive filter from which the third signal is output, and If the determination module (30) determines that the second signal contains the uncorrelated noise and the third signal does not contain the uncorrelated noise, the control module (28; 28A) is designed to change the filter coefficient of the second adaptive filter without changing the filter coefficient of the first adaptive filter. [10] Audio processing device according to claim 8, wherein which are designed to receive one or more receiving units based on a third audio signal containing a third audio component that is different from the first audio component and the second audio component, The audio processing device further includes: a first filter unit (F1) containing one or more adaptive filters (F1A-F1C), wherein the one or more adaptive filters comprise a first adaptive filter into which the second signal is input, and a second adaptive filter into which the third signal is input, and a second filter unit (F2) containing one or more adaptive filters (F2A), wherein the second filter unit contains a third adaptive filter that is different from the second adaptive filter and is designed to receive the third signal and output a first pass-through signal based on the third signal, a number of one or more adaptive filters (F2A) contained in the second filter unit is less than a number of adaptive filters (F1A-F1C) contained in the first filter unit and The control module (28; 28A) is designed to determine, based on the second signal and the third signal, which from the first filter unit and the second filter unit is used to generate the subtraction signal. [11] Audio processing device according to claim 10, wherein the second filter unit contains only the third adaptive filter and, If it is determined that the second signal does not contain a first component derived from an utterance, and the third signal does contain the first component, the control module (28; 28A) is designed to generate the subtraction signal by using the second filter unit. [12] Audio processing device according to any one of claims 8 to 11, wherein the determination module (30) is designed to determine that the microphone signal contains the uncorrelated noise when an intensity of the microphone signal is greater than or equal to a predetermined value. [13] Audio processing device according to any one of claims 8 to 11, wherein the first receiving unit receives a first microphone signal and the second receiving unit receives a second microphone signal that is different from the first microphone signal, and the determination module (30) is designed to determine that the first microphone signal contains the uncorrelated noise when the intensity of the first microphone signal is greater than the intensity of the second microphone signal by a predetermined value or more. [14] Audio processing device according to any one of claims 8 to 11, wherein the determination module (30) is designed to determine, based on vehicle information, that the microphone signal contains the uncorrelated background noise.

Citation Information

Patent Citations

  • METHOD AND DEVICE FOR PROCESSING SUB-BAND SIGNALS BY MEANS OF ADAPTIVE FILTERS

    DE60310725T2

  • Systems, methods, apparatus, and computer program products for spectral contrast enhancement

    US20090299742A1