Residual echo suppression based on autoregression
By using an autoregressive model based on signal energy attenuation in a vehicle hands-free telecommunications system, the power spectral density of the residual echo signal is estimated, which solves the problem that residual echoes are difficult to eliminate in the reverberation environment, and achieves higher call quality and clarity.
Patent Information
- Application Number
- CN202110531463.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-28
- Filing Date
- 2021-05-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-05-14
AI Technical Summary
In a reverberation environment, it is difficult for the prior art to effectively eliminate residual echoes, affecting the call clarity of the vehicle's hands-free telecommunications system.
The autoregressive model based on signal energy attenuation is used to estimate the power spectral density of the residual echo signal through an autoregressive algorithm in the short-time Fourier transform (STFT) domain, thereby isolating the speech signal from the speech signal and the microphone signal.
It effectively reduces the impact of residual echoes in the vehicle, improves the call quality and clarity of the hands-free telecommunications system, and avoids distortion of the expected signal.
Smart Images

Figure CN114333872B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to programming automotive vehicle audio and sound suppression systems. More specifically, aspects of the present invention relate to systems, methods and apparatus for autoregressive-based residual echo suppression in the short-time Fourier transform (STFT) domain by executing an algorithm for suppressing residual echo in a reverberant environment based on an autoregressive model of signal energy decay. Background Art
[0002] Vehicle subsystems often include noise and vibration constraints to reduce cabin noise and improve occupant comfort. For example, a rear differential control module may have noise and vibration constraints that lock the clutch to avoid gear rattle or limit torque to avoid hypoid gear noise. In addition, vehicle cabin systems may monitor cabin noise and then generate noise cancellation sound waves or noise masking sounds to cover unwanted vehicle noise. The volume of a vehicle audio system may increase as vehicle speed increases. Fans and other heating and air conditioning systems may change speed when a mobile phone call is received by the vehicle communication system.
[0003] In order to provide an effective hands-free telecommunication system in a vehicle, the cabin noise must be reduced as much as possible to improve the clarity of telephone calls. In addition, an acoustic echo canceller (AEC) is an essential component of a full-duplex hands-free communication system to eliminate unwanted echo signals caused by the acoustic coupling between the speaker and the microphone. In a reverberant sound environment, linear echo cancellation cannot eliminate all echo components due to the long tail of the indoor impulse response and computational limitations. In order to obtain an acceptable echo level, a nonlinear residual echo suppressor is also required. The above information disclosed in this background section is only for enhancing the understanding of the background of the invention and therefore it may contain information that does not constitute the prior art known in this country to a person of ordinary skill in the art. Summary of the invention
[0004] Disclosed herein are vehicle sensor systems, vehicle user interface systems and associated control logic for providing vehicle user interface systems, methods of manufacturing and operating such systems, and motor vehicles equipped with user interface systems. By way of example and not limitation, a user interface system is provided for predicting neighboring vehicle behavior, predicting an increased probability of a vehicle-to-vehicle contact event, and displaying an indication of an increased probability of contact with an associated neighboring vehicle.
[0005] According to one aspect of the present disclosure, a device includes a telecommunications processor for generating a speaker signal, a speaker for receiving the speaker signal from the telecommunications processor and broadcasting the speaker signal within a vehicle cabin, a microphone for detecting a microphone signal within the vehicle cabin, wherein the microphone signal includes a speech signal and a residual echo signal, a signal processor configured to receive the speaker signal from the telecommunications processor, receive the microphone signal from the microphone, generate an estimated power spectral density of the residual echo signal in response to a previous power spectral density of a previous residual echo signal, isolate the speech signal from the microphone signal by mixing the estimated power spectral density of the residual echo signal with the microphone signal; and couple the speech signal to the telecommunications processor.
[0006] According to another aspect of the present disclosure, an autoregressive algorithm is used to estimate the power spectral density of the residual echo signal.
[0007] According to another aspect of the present disclosure, wherein the microphone signal further includes reflections of the speaker signal, and wherein the acoustic echo cancellation algorithm is used to suppress the reflections of the speaker signal in the microphone signal in response to the speaker signal.
[0008] According to another aspect of the present disclosure, the telecommunication processor and the signal processor form part of a vehicle cabin hands-free telecommunication system.
[0009] According to another aspect of the present disclosure, a high-order autoregressive model is used to generate an estimated power spectral density of the residual echo signal.
[0010] According to another aspect of the present disclosure, the residual echo signal is generated in response to reverberation of a speaker signal in a vehicle cabin.
[0011] According to another aspect of the present disclosure, wherein the speech signal is isolated in response to the estimated power spectral density of the loudspeaker signal and the estimated power spectral density of the residual echo signal and at least one previous value of the estimated power spectral density of the residual echo signal.
[0012] According to another aspect of the present disclosure, a method includes receiving a speaker signal for coupling to a speaker from a communication processor; receiving a microphone signal from a microphone, wherein the microphone signal includes a speech signal and a residual echo signal; generating an estimated power spectral density of the residual echo signal in response to a previous power spectral density of a previous residual echo signal; estimating a gain by using the estimated power spectral density of the residual echo signal, which is multiplied by the microphone signal to isolate the speech signal from the microphone signal; and coupling the speech signal to the communication processor.
[0013] According to another aspect of the present disclosure, an autoregressive algorithm is used to estimate the power spectral density of the residual echo signal.
[0014] According to another aspect of the present disclosure, wherein the microphone signal further includes reflections of the speaker signal, and wherein the acoustic echo cancellation algorithm is used to suppress the reflections of the speaker signal in the microphone signal in response to the speaker signal.
[0015] According to another aspect of the present disclosure, the communication processor forms part of a vehicle cabin hands-free telecommunication system.
[0016] According to another aspect of the present disclosure, a high-order autoregressive model is used to generate an estimated power spectral density of the residual echo signal.
[0017] According to another aspect of the present disclosure, the residual echo signal is generated in response to reverberation of a speaker signal in a vehicle cabin.
[0018] According to another aspect of the present disclosure, the estimated power spectral density of the residual echo signal is determined in response to the loudspeaker signal and the microphone signal and at least one previous value of the estimated power spectral density of the residual echo signal.
[0019] According to another aspect of the present disclosure, the speech signal is isolated in response to the estimated power spectral density of the loudspeaker signal and the estimated power spectral density of the residual echo signal and at least one previous value of the estimated power spectral density of the residual echo signal.
[0020] According to another aspect of the present disclosure, the method is performed by a digital signal processor as part of a two-way telecommunication system.
[0021] According to another aspect of the present disclosure, the speech signal is not used to estimate the power spectral density of the residual echo signal.
[0022] According to another aspect of the present disclosure, a hands-free telecommunication system within a vehicle cabin includes: a telecommunication processor for generating a speaker signal for broadcast by a speaker within the vehicle cabin; a microphone for generating a microphone signal in response to sound within the vehicle cabin; and a processor for estimating an estimated power spectral density of a residual echo signal in response to a previous power spectral density of a previous residual echo signal, estimating a gain by using the estimated power spectral density of the residual echo signal, multiplying the microphone signal, isolating a speech signal from the microphone signal, and coupling the speech signal to the telecommunication processor.
[0023] According to another aspect of the present disclosure, the estimated power spectral density of the residual echo signal is determined in response to the loudspeaker signal and the microphone signal and at least one previous value of the estimated power spectral density of the residual echo signal.
[0024] According to another aspect of the present disclosure, the speech signal is isolated in response to the estimated power spectral density of the loudspeaker signal and the estimated power spectral density of the residual echo signal and at least one previous value of the estimated power spectral density of the residual echo signal.
[0025] The above advantages and other advantages and features of the present disclosure will be apparent from the following detailed description of preferred embodiments taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The above and other features and advantages of the present invention and the methods for achieving the same will become more apparent, and the present invention will be better understood by referring to the following description of the embodiments of the present invention in conjunction with the accompanying drawings.
[0027] Figure 1A An operating environment for providing autoregression-based residual echo suppression in a two-way vehicle communication system is shown in accordance with an exemplary embodiment.
[0028] Figure 1B is an illustration of an exemplary reverberation signal in a vehicle cabin according to an exemplary embodiment.
[0029] Figure 2 A block diagram of an exemplary system for providing autoregression-based residual echo suppression in a two-way vehicle communication system is shown in accordance with an exemplary embodiment.
[0030] Figure 3 A flow chart is shown of an exemplary method for providing autoregression-based residual echo suppression in a two-way vehicle communication system in accordance with another exemplary embodiment.
[0031] Figure 4 A block diagram of another exemplary system for providing autoregression-based residual echo suppression in a two-way vehicle communication system is shown in accordance with another exemplary embodiment.
[0032] Figure 5 A flow chart is shown of another exemplary method for providing autoregression-based residual echo suppression in a two-way vehicle communication system in accordance with another exemplary embodiment.
[0033] The exemplifications set out herein illustrate preferred embodiments of the invention, and such exemplifications are not to be construed as limiting the scope of the invention in any way. DETAILED DESCRIPTION
[0034] Embodiments of the present disclosure are described herein. However, it should be understood that the disclosed embodiments are merely examples, and other embodiments may take various and alternative forms. These figures are not necessarily to scale; some features may be exaggerated or minimized to show the details of a particular component. Therefore, the specific structural and functional details disclosed herein should not be interpreted as restrictive, but merely representative. The various features illustrated and described with reference to any one of the accompanying drawings may be combined with the features illustrated in one or more other drawings to produce embodiments that are not explicitly illustrated or described. The combination of features shown provides representative embodiments for typical applications. However, for specific applications or implementations, various combinations and modifications of features consistent with the teachings of the present disclosure may be required.
[0035] Figure 1A 1 is an illustration of an exemplary environment 100 for providing autoregression-based residual echo suppression in a two-way vehicle communication system. In an exemplary vehicle cabin 110, a speaker 120 is configured to receive an incoming portion of an audio conversation from a communication system 140 and broadcast the incoming portion within the vehicle cabin 110. A microphone 130 is used to receive an outgoing portion of the audio conversation and send the outgoing portion to the communication system 140 for transmission via a wireless network (e.g., a cellular telephone network, etc.).
[0036] Feedback or echo suppression is an important feature of two-way communication systems to reduce feedback from the speaker 120 to the microphone 130, as well as to increase the intelligibility of the conversation to the listener. Echo cancellation generally involves identifying the transmitted signal as it is received by the microphone 130 in the outgoing signal path, and then subtracting the signal from the incoming signal path before it is coupled to the transmitter communication system 140. This configuration is used to prevent the audio signal broadcast by the speaker 120 from being recorded by the microphone 130, thereby reducing the echo and / or feedback sent by the system back to the communication system 140.
[0037] In hands-free telecommunication systems, residual echo suppression is desirable in order to achieve acceptable echo levels in hands-free telecommunication conferencing systems. To eliminate the residual echo, spectral subtraction is usually applied in the STFT domain. The residual echo power spectral density (PSD) is usually adopted in order to be applied to the spectral subtraction algorithm. The residual echo PSD is estimated by measuring the cross-correlation between the reference signal and the residual echo signal. When the desired speaker is also present on the input signal, it can introduce a bias to the cross-correlation metric, resulting in distortion in the desired signal.
[0038] Traditionally, the correlation between the reference and residual echo signals is measured and used to estimate the level of the residual echo level in the noise signal. Since the desired signal can also appear during the correlation estimation, this method may introduce distortion when the desired signal is present. The proposed system and method are configured to model the energy decay of the residual signal in the STFT domain as an autoregressive process and rely only on past estimated values to estimate the current value, thereby avoiding the desired signal distortion problem. In signal processing, an autoregressive process uses previous values of a variable to estimate the current value of the variable.
[0039] Figure 1B is a graph 150 showing a typical reverberation impulse response between a speaker 120 and a microphone 130 in a vehicle cabin. Graph 150 is an illustration of an acoustic amplitude level 160 at the microphone shown as amplitude 161 at time 162. Additionally, graph 150 is an illustration of a typical frame size 165 with respect to time and a late reverberation processing time 170. In previous systems, the late reverberation response may not be processed. In an exemplary embodiment, the presently proposed system is configured to account for the late reverberation response by estimating the late reverberation response energy in the vehicle cabin using an autoregressive technique to reduce un-cancelled echo on the audio system input without distorting the desired driver signal.
[0040] Now go to Figure 2 , a block diagram showing an exemplary implementation of a system 200 for providing autoregression-based residual echo suppression in the STFT domain in a two-way vehicle communication system is shown. The system 200 may include a microphone 210, a speaker 220, a first mixer 215, a second mixer 225, an acoustic echo cancellation (AEC) block 230, a power spectral density (PSD) estimator 240, a residual PSD estimation block 260, G 270, and an autoregression (AR) coefficient estimator 250. The exemplary system 200 is configured to estimate the energy of the uncancelled echo in the car, and the energy estimate is used to reduce the uncancelled echo on the microphone line without distorting the desired driver signal. The energy decay of the uncancelled echo is modeled as an autoregressive process in the time / frequency domain, thereby allowing for more accurate estimation in the absence of a direct relationship between desired signals, such as drive signals. Therefore, the desired signal is not compromised when removing the undesired echo. Improve the quality of calls between passengers and callers outside the vehicle.
[0041] In this exemplary embodiment, microphone 210 and speaker 220 are located in the vehicle cabin as part of a two-way communication system, such as a hands-free telecommunication system. In some instances, microphone 210 may receive not only the desired speech of the speaker, but also any sound emitted by speaker 220, thereby generating undesirable echo and / or feedback. To address this problem, exemplary system 200 includes an STFT domain AEC block 230 (including an echo estimation block 231) for eliminating the initial or linear portion of the echo received at the microphone, and a residual echo suppression block (RES) 235 for removing residual echo to improve performance stability. The echo signal may be caused by a distant speaker, or a speech signal received by a telecommunication processor that is broadcast in the vehicle cabin through a speaker and then picked up by a microphone. Therefore, the echo suppression system is configured to attenuate the echo without interfering with the in-vehicle intercom signal, and then transmit it back to the distant speaker for hearing.
[0042] 3) AEC (this is a prior art block) input is the microphone signal and the loudspeaker signal. It uses these two signals to estimate the echo signal, which is then subtracted from the microphone signal to generate a residual echo signal, which is input to the RES block. The RES block 235 handles the late reverberation present during the later part of the processing time period.
[0043] The RES block 235 may be composed of four main processing blocks, a PSD estimator 240, an EST RES PSD block 260, a gain block (G) 270, and an AR Coef estimator 250. The PSD estimator 240 is configured to receive the AEC input signal (Y(n, k)) from the microphone 210, the AEC error signal (E(n, k)) from the first mixer 215, and the AEC estimated echo signal from the AEC block 230. PSD estimator 240 uses these signals to generate and where φ xx Represents the PSD of the x signal:
[0044] φ ee (n, k) = λφ ee (n-1, k)+(1-λ)|E| 2
[0045] φ yy (n, k) = λφ yy (n-1, k)+(1-λ)|Y| 2
[0046]
[0047] The reverberation room impulse response model shows that the signal decays exponentially:
[0048] h(n)=b(n)c -ρn / f s
[0049] This can be used to represent the residual signal energy decay in the STFT domain:
[0050]
[0051] Y-AEC input signal and Estimated echo signal. Because the residual echo Can be used
[0052]
[0053] in
[0054] The AR parameters can be assumed to vary slowly over time and can therefore be estimated when the desired speaker is not present in the input signal. The residual echo PSD is estimated using only the AR model parameters and the estimated previous values. Since there is no direct connection to the desired speech, no distortion is expected.
[0055] The AR coefficient estimator is configured to receive
[0056] φ ee (n, k) = λφ ee (n-1, k)+(1-λ)|E| 2
[0057] φ yy (n, k) = λφ yy (n-1, k)+(1-λ)|Y| 2
[0058] and estimated
[0059]
[0060] The residual estimation PSD block 260 is configured to receive
[0061] as well as
[0062]
[0063] and the iteration value
[0064]
[0065] To estimate the residual PSDφ δδ , where φ δδ equal
[0066]
[0067] The gain block then receives φ ee and φ δδ By estimate
[0068] in
[0069]
[0070] Where Out(n, k) is output to the telecommunication voice processor. In an exemplary embodiment, the output of the RES block 235 may be a scalar gain that is multiplied by the residual echo signal to produce a clean signal. The gain may be calculated by using the estimated residual echo PSD and the residual echo PSD using an algorithm such as Wiener filtering in the STFT domain.
[0071] In an exemplary embodiment, autoregression based residual echo suppression is performed in the short time Fourier transform (STFT) domain. An exemplary system may include STFT blocks 211, 222 to convert incoming time domain signals to the STFT domain, and ISTFT blocks 212, 221 to convert outgoing signals back to the time domain.
[0072] Now go to Figure 3 , a flow chart of an exemplary implementation of a method 300 for providing autoregression-based residual echo suppression in a two-way vehicle communication system is shown in the STFT domain. Initially, the method is first used to receive 310 a speaker signal intended to be broadcast into a vehicle cabin and a microphone signal 330 transmitted back to the communication system and convert them to the STFT domain. The speaker signal may be generated by a telecommunications system processor or other communication processor. The speaker signal may convey information as part of a two-way electronic voice communication.
[0073] The method next determines 320 an AEC estimated echo signal from the speaker and microphone signals. The AEC estimated echo signal may indicate a portion of the speaker signal to be attenuated if coupled back into the microphone of the telecommunications system, a linear portion of the echo and reverberation. The method then receives 330 microphone signals indicative of sounds from within the vehicle cabin, including speech signals and reverberant speaker signals. The reverberant speaker signals may include directly coupled representations of the speaker signals broadcast by the speakers within the vehicle cabin and reflected representations of the speaker signals reflected from one or more surfaces within the vehicle cabin. These reflected representations may be delayed and nonlinear compared to the directly coupled representations.
[0074] Next, an AEC error output signal is generated in response to the microphone signal and the AEC estimated echo signal. The AEC estimated echo signal is used to identify the speaker signal detected in the microphone and attenuate the speaker signal so that mainly only the desired speech signal is coupled to the telecommunications transmitter. The AEC error output signal may contain only the speech signal and the reverberant portion of the speaker signal. It is desirable to remove the reverberant portion of the speaker signal to further increase the clarity of the speech signal for transmission to the telecommunications system.
[0075] To suppress the residual loudspeaker signal from the microphone signal, the method is configured to estimate the residual echo energy by using an autoregressive model of energy decay. For example, the method can employ a high-order autoregressive model to suppress the residual echo energy. To achieve this, the method is first used to estimate 350 a power spectral density (PSD) in response to the unchanged microphone signal, the AEC error output signal, and the generated AEC estimated echo signal. The estimated residual echo PSD is then used to estimate 360 autoregressive coefficients.
[0076] A residual echo PSD is then estimated 370 in response to the estimated autoregressive coefficients, the estimated PSD, and one or more previous estimates of the residual PSD. The residual echo PSD may be estimated as follows:
[0077]
[0078]
[0079] Using autoregressive techniques employing one or more previous estimates of the PSD, the energy decay of the residual echo can be modeled in the time / frequency domain as an autoregressive process to improve the residual PSD estimate without direct relationship to the desired output signal.
[0080] The residual PSD along with the estimated PSD is applied to a Wiener filtering algorithm to produce an estimated residual echo gain (G) signal 270. The estimated residual echo gain (G) signal is then applied to the AEC error output signal via a second mixer 225 to generate an output signal to be coupled to a telecommunications system processor and transmitter.
[0081] The output signal is then converted back to the time domain via an ISTFT block to produce the telecommunication system input
[0082] Now go to Figure 4 , a block diagram of another exemplary implementation of a system 400 for providing autoregression-based residual echo suppression in a two-way vehicle communication system is shown. The exemplary system 400 may form a vehicle cabin hands-free telecommunication system including a telecommunication processor 410, a signal processor 420, a speaker 430, and a microphone 440. The telecommunication processor and the signal processor form a part.
[0083] The telecommunications processor 410 is used to process signals received via a wireless network, such as signals received from a cellular phone or a data network, and prepare signals for transmission via the wireless network. The telecommunications processor 410 is also configured to generate a speaker signal representing an audio signal for coupling to a speaker 430 or the like for broadcasting within the vehicle cabin. The telecommunications processor 410 may also receive an audio signal from a microphone 440 or a signal processor 420 coupled to a microphone 440.
[0084] The speaker 430 is configured to receive a speaker signal from the telecommunication processor 410 and broadcast the speaker signal in the vehicle cabin. The microphone 440 is used to detect a microphone signal in the vehicle cabin, wherein the microphone signal includes a voice signal and an echo signal generated by reverberation or reflection of the speaker signal broadcast by the speaker 430 in the vehicle cabin. The microphone signal may include a reflection of the speaker signal, wherein, in response to the speaker signal, an acoustic echo cancellation algorithm is used to suppress the reflection of the speaker signal in the microphone signal.
[0085] The signal processor 420 may be a digital signal processor, an audio processor, etc., which is configured to receive a speaker signal from the telecommunications processor 410, a microphone signal from the microphone 430, and generate an estimated power spectrum density of the residual echo signal in response to a previous power spectrum density of the previous residual echo signal. In one example, an autoregressive algorithm is used to estimate the power spectrum density of the residual echo signal. For example, a high-order autoregressive model may be used to generate an estimated power spectrum density of the residual echo signal. The residual echo signal may be generated in response to the reverberation of the speaker signal in the vehicle cabin.
[0086] Then, the signal processor 420 isolates the speech signal from the microphone signal by multiplying the microphone signal with an estimated residual echo gain generated using the estimated power spectral density of the residual echo signal, and couples the speech signal to the telecommunications processor 410. In one example, the speech signal may be isolated in response to the estimated power spectral density of the speaker signal and the estimated power spectral density of the residual echo signal and at least one previous value of the estimated power spectral density of the residual echo signal.
[0087] In an exemplary embodiment, the system can be a hands-free telecommunication system in a vehicle cabin, comprising: a telecommunication processor 410 for generating a speaker signal for broadcasting by a speaker 430 in the vehicle cabin; a microphone 430 for generating a microphone signal in response to sound in the vehicle cabin; and a processor 420 for estimating an estimated power spectral density of a residual echo signal in response to a previous power spectral density of a previous residual echo signal, isolating a speech signal from the microphone signal by multiplying an estimated residual echo gain generated using the estimated power spectral density of the residual echo signal with the microphone signal; and coupling the speech signal to the telecommunication processor. In this exemplary embodiment, the estimated power spectral density of the residual echo signal is determined in response to at least one previous value of the speaker signal and the microphone signal and the estimated power spectral density of the residual echo signal. In addition, the speech signal can be isolated in response to the estimated power spectral density of the speaker signal and the estimated power spectral density of the residual echo signal and at least one previous value of the estimated power spectral density of the residual echo signal.
[0088] Now go to Figure 5 , a flow chart illustrating an exemplary implementation of a system 500 for providing autoregression-based residual echo suppression in a two-way vehicle communication system is shown. The exemplary method 500 is first used to receive 510 a speaker signal for coupling to a speaker from a communication processor. The speaker signal may represent an incoming portion of a two-way telecommunication communication. The method is next used to receive 520 a microphone signal from a microphone, wherein the microphone signal includes a speech signal and a residual echo signal. The residual echo signal may be generated in response to reverberation of the speaker signal within a vehicle cabin. Therefore, the microphone signal may include reflections of the speaker signal, and wherein an acoustic echo cancellation algorithm in the STFT domain is used to suppress the reflections of the speaker signal in the microphone signal in response to the speaker signal.
[0089] Then, the exemplary method generates 530 an estimated power spectral density of the residual echo signal in response to a previous power spectral density of the previous residual echo signal. The power spectral density of the residual echo signal can be estimated using an autoregressive algorithm. An autoregressive model of order or more is used to generate the estimated power spectral density of the residual echo signal. The estimated power spectral density of the residual echo signal can also be determined in response to at least one previous value of the estimated power spectral density of the speaker signal and the microphone signal and the estimated power spectral density of the residual echo signal. In an exemplary embodiment, the speech signal can be isolated in response to the estimated power spectral density of the speaker signal and the estimated power spectral density of the residual echo signal and at least one previous value of the estimated power spectral density of the residual echo signal. In the autoregressive configuration, the speech signal is not used to estimate the power spectral density of the residual echo signal.
[0090] The method is then configured to isolate the speech signal from the microphone signal by multiplying the microphone signal by an estimated residual echo gain generated using the estimated power spectral density of the residual echo signal 540. The isolated speech signal is then transformed back to the time domain by an I S T F T transform and then coupled to the communication processor 550. The communication processor forms part of the vehicle cabin hands-free telecommunication system.
[0091] Although at least one exemplary embodiment has been given in the foregoing detailed description, it should be understood that there are a large number of variations. It should also be understood that one or more exemplary embodiments are only examples and are not intended to limit the scope, applicability or configuration of the present disclosure in any way. On the contrary, the above detailed description will provide a convenient roadmap for implementing one or more exemplary embodiments to those skilled in the art. It should be understood that various changes may be made to the function and arrangement of elements without departing from the scope of the disclosure described in the attached claims and their legal equivalents.
Claims
1. A residual echo suppression device based on autoregression, comprising: a telecommunications processor for generating a speaker signal; a speaker for receiving a speaker signal from the telecommunications processor and broadcasting the speaker signal within the vehicle cabin; A microphone for detecting a microphone signal in the vehicle compartment, wherein the microphone signal includes a speech signal and a residual echo signal, wherein the residual echo signal is generated in response to reverberation of a speaker signal in the vehicle compartment; and A signal processor configured to receive a speaker signal from a telecommunications processor, receive a microphone signal from a microphone, generate an estimated power spectral density of the residual echo signal in response to a previous power spectral density of a previous residual echo signal, isolate the speech signal from the microphone signal by mixing the estimated power spectral density of the residual echo signal with the microphone signal; and couple the speech signal to the telecommunications processor.
2. The device according to claim 1, wherein: The power spectral density of the residual echo signal is estimated using an autoregressive algorithm.
3. The device according to claim 1, wherein: The microphone signal also includes reflections of the speaker signal, wherein an acoustic echo cancellation algorithm is used to suppress the reflections of the speaker signal in the microphone signal responsive to the speaker signal.
4. The device according to claim 1, wherein: The signal processor is further configured to convert the speaker signal and the microphone signal into a short-time Fourier transform domain speaker signal and a short-time Fourier transform microphone signal.
5. The device according to claim 1, wherein: A high-order autoregressive model is used to generate an estimated power spectral density of the residual echo signal.
6. The device according to claim 1, wherein: The speech signal is isolated in response to the estimated power spectral density of the loudspeaker signal and the estimated power spectral density of the residual echo signal and at least one previous value of the estimated power spectral density of the residual echo signal.
7. A residual echo suppression method based on autoregression, comprising: receiving a speaker signal from the communication processor for coupling to a speaker; receiving a microphone signal from a microphone, wherein the microphone signal includes a speech signal and a residual echo signal, wherein the residual echo signal is generated in response to reverberation of a speaker signal in a vehicle cabin; generating an estimated power spectral density of the residual echo signal in response to a previous power spectral density of a previous residual echo signal; isolating the speech signal from the microphone signal by mixing an estimated power spectral density of the residual echo signal with the microphone signal; and The voice signal is coupled to the communications processor.
8. The method according to claim 7, wherein: The power spectral density of the residual echo signal is estimated using an autoregressive algorithm.
9. The method according to claim 7, wherein: The microphone signal also includes reflections of the speaker signal, and wherein an acoustic echo cancellation algorithm is used to suppress reflections of the speaker signal in the microphone signal in response to the speaker signal.
Citation Information
Patent Citations
System and method for performing speech enhancement using a deep neural network-based signal
US20180040333A1
Methods and apparatus for providing echo suppression using frequency domain nonlinear processing
US6658107B1