Echo cancellation method, device, equipment and storage medium
By introducing an NLMS filter bank into the Kalman filter for parallel filtering, the problem of energy difference between near-end signals and echo signals is solved, and the efficiency and accuracy of echo cancellation is achieved.
Patent Information
- Application Number
- CN202411973190.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing echo cancellation methods are difficult to effectively distinguish the energy of the proximal signal and the echo signal, resulting in inaccurate echo cancellation.
An NLMS filter bank is introduced into the Kalman filter, and the parallel filtering is performed through multiple NLMS filters, the energy of the near-end signal is estimated, and echo cancellation is performed in combination with the Kalman filter.
Improves the efficiency and accuracy of echo cancellation, especially when echo path changes, which can quickly converge.
Smart Images

Figure CN119626242B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio technology, and in particular to an echo cancellation method, apparatus, device, and storage medium. Background Art
[0002] Acoustic Echo Cancellation (AEC) is widely used in voice calls, video conferencing, and smart devices. It aims to eliminate the echo interference caused by the speaker signal returning to the microphone through the acoustic coupling path. The core of AEC is to build an adaptive filter, such as one based on algorithms such as NLMS (Normalized Least Mean Squares), RLS (Recursive Least Squares), and Kalman filtering, to estimate and eliminate the echo component of the far-end signal in the microphone.
[0003] Currently, in the field of echo cancellation, Kalman filters demonstrate significant advantages over traditional NLMS algorithms, especially in complex double talk scenarios and rapidly changing echo paths. Kalman filters can continuously estimate and track echo path coefficients, while NLMS algorithms need to suspend coefficient updates to avoid performance degradation. This continuous tracking capability makes Kalman filters more robust in dynamic acoustic environments. However, the accuracy of Kalman filters depends on accurate estimation of near-end signal energy (i.e., measurement noise in the Kalman filter). Accurately estimating near-end signal energy is crucial to preventing divergence of adaptive filter coefficients and preventing the filter from deviating from the true echo path coefficients.
[0004] However, current adaptive filters require the energy of the near-end signal to accurately estimate the echo path coefficients. Since near-end speech and / or background noise are mixed with the echo signal, it is difficult to effectively distinguish the energy of the near-end signal from the echo signal. Therefore, how to effectively distinguish the energy of the near-end signal from the echo signal to cancel the echo remains an unresolved issue in the field. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide an echo cancellation method, apparatus, device, and storage medium that can improve the efficiency and accuracy of echo cancellation. The specific solution is as follows:
[0006] In a first aspect, the present application discloses an echo cancellation method, comprising:
[0007] Acquire the audio signal played by the current sound playback device and the audio signal collected by the sound collection device to obtain a far-end audio signal and a near-end audio signal;
[0008] Obtaining a historical residual signal obtained in a previous iteration of a Kalman filter and a historical remote reference signal played by the sound playback device, and inputting the historical residual signal and the historical remote reference signal into an NLMS filter bank located in the Kalman filter for NLMS filtering to obtain a plurality of filtered residual signals;
[0009] Calculating a current measurement noise estimate of the Kalman filter based on the plurality of filtered residual signals, and calculating a Kalman gain based on the current measurement noise estimate, the historical remote reference signal, and process noise of the Kalman filter;
[0010] updating the Kalman filter coefficient estimate obtained in the previous iteration of the Kalman filter based on the Kalman gain and the historical residual signal to obtain a current filter coefficient estimate;
[0011] Echoes generated on an echo path between the sound collection device and the sound playback device are eliminated based on the current filter coefficient estimate, the far-end audio signal, and the near-end audio signal to obtain an echo-cancelled signal.
[0012] Optionally, the step of obtaining a historical residual signal obtained in a previous iteration of the Kalman filter and a historical remote reference signal played by the sound playback device, and inputting the historical residual signal and the historical remote reference signal into an NLMS filter bank located in the Kalman filter for NLMS filtering, to obtain multiple filtered residual signals, includes:
[0013] Acquire a frequency domain residual signal obtained in a previous iteration of a Kalman filter pre-created based on an NLMS filter bank and a frequency domain far-end reference signal played by the sound playback device to obtain a historical residual signal and a historical far-end reference signal;
[0014] The historical residual signal and the historical far-end reference signal are input into the NLMS filter bank for NLMS filtering to obtain a plurality of filtered residual signals.
[0015] Optionally, obtaining a frequency domain residual signal obtained in a previous iteration of a Kalman filter pre-created based on an NLMS filter bank and a frequency domain far-end reference signal played by the sound playback device to obtain a historical residual signal and a historical far-end reference signal includes:
[0016] Performing frequency domain transformation on the near-end audio signal and the far-end audio signal respectively to obtain a near-end frequency domain signal and a far-end frequency domain signal;
[0017] A frequency domain residual signal obtained in the last iteration of the Kalman filter pre-created based on the NLMS filter bank and a frequency domain far-end reference signal played by the sound playing device are obtained to obtain a historical residual signal and a historical far-end reference signal.
[0018] Optionally, the canceling an echo generated on an echo path between the sound collection device and the sound playback device based on the current filter coefficient estimate, the far-end audio signal, and the near-end audio signal to obtain an echo-canceled signal includes:
[0019] updating the far-end frequency domain signal based on the historical far-end reference signal to obtain a current audio signal, and calculating a residual signal of a current frame based on the current filter coefficient estimate and the near-end frequency domain signal to obtain a current residual signal;
[0020] The current residual signal is transformed in the time domain to obtain an audio signal after the echo generated on the echo path between the sound collection device and the sound playback device is eliminated, thereby obtaining an echo-cancelled signal.
[0021] Optionally, the inputting the historical residual signal and the historical remote reference signal into the NLMS filter bank for NLMS filtering to obtain a plurality of filtered residual signals includes:
[0022] The frequency domain buffer of the historical remote reference signal is subjected to time block processing to obtain multiple frequency domain signal blocks, and the historical residual signal and each of the frequency domain signal blocks are input into the NLMS filter bank for NLMS filtering to obtain multiple filtered residual signals.
[0023] Optionally, calculating a current measurement noise estimate of the Kalman filter based on the plurality of filtered residual signals includes:
[0024] Recursively smoothing each of the filtered residual signals to obtain a plurality of smoothed residuals;
[0025] The minimum value of all the smoothed residuals is calculated to obtain a current measurement noise estimate of the Kalman filter.
[0026] Optionally, the calculating the Kalman gain based on the current measurement noise estimate, the historical remote reference signal, and the process noise of the Kalman filter includes:
[0027] Using process noise to update the posterior offset matrix obtained in the previous iteration of the Kalman filter to obtain a priori offset matrix;
[0028] A Kalman gain is calculated based on the prior offset matrix, the current measurement noise estimate, and the historical far-end reference signal.
[0029] In a second aspect, the present application discloses an echo cancellation device, comprising:
[0030] A first signal acquisition module is used to acquire the audio signal played by the current sound playback device and the audio signal collected by the sound collection device to obtain a far-end audio signal and a near-end audio signal;
[0031] A second signal acquisition module is used to obtain a historical residual signal obtained in the last iteration of the Kalman filter and a historical remote reference signal played by the sound playback device;
[0032] A filtering module, configured to input the historical residual signal and the historical remote reference signal into an NLMS filter bank located in the Kalman filter to perform NLMS filtering to obtain a plurality of filtered residual signals;
[0033] a calculation module, configured to calculate a current measurement noise estimate of the Kalman filter based on the plurality of filtered residual signals, and calculate a Kalman gain based on the current measurement noise estimate, the historical remote reference signal, and the process noise of the Kalman filter;
[0034] An updating module is configured to update the Kalman filter coefficient estimate obtained in the previous iteration of the Kalman filter based on the Kalman gain and the historical residual signal to obtain a current filter coefficient estimate;
[0035] The echo cancellation module is used to cancel the echo generated on the echo path between the sound collection device and the sound playback device based on the current filter coefficient estimate, the far-end audio signal and the near-end audio signal to obtain an echo-canceled signal.
[0036] In a third aspect, the present application discloses an electronic device, comprising a processor and a memory; wherein the processor implements the aforementioned echo cancellation method when executing a computer program stored in the memory.
[0037] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned echo cancellation method is implemented.
[0038] As can be seen, in this application, the audio signal currently played by the sound playback device and the audio signal collected by the sound collection device are first obtained to obtain a far-end audio signal and a near-end audio signal. Then, a historical residual signal obtained in the previous iteration of the Kalman filter and a historical far-end reference signal played by the sound playback device are obtained, and the historical residual signal and the historical far-end reference signal are input into the NLMS filter bank located in the Kalman filter for NLMS filtering to obtain multiple filtered residual signals. Then, a current measurement noise estimate of the Kalman filter is calculated based on the multiple filtered residual signals, and a Kalman gain is calculated based on the current measurement noise estimate, the historical far-end reference signal, and the process noise of the Kalman filter. Then, the Kalman filter coefficient estimate obtained in the previous iteration of the Kalman filter is updated based on the Kalman gain and the historical residual signal to obtain a current filter coefficient estimate. Finally, based on the current filter coefficient estimate, the far-end audio signal, and the near-end audio signal, the echo generated in the echo path between the sound collection device and the sound playback device is cancelled to obtain an echo-cancelled signal. The present application first obtains the near-end audio signal collected by the sound collection device and the far-end reference signal played by the sound playback device, and then inputs the historical residual signal and the historical far-end reference signal into the NLMS filter group of the Kalman filter, so as to perform NLMS filtering through multiple NLMS filters in the NLMS filter group to obtain multiple filtered residual signals, and calculates the filter coefficient estimation value of the current Kalman filter based on the multiple filtered residual signals, and finally obtains the echo-cancelled signal based on the filter coefficient estimation value, the far-end audio signal and the near-end audio signal. It can be seen that the present application adds an NLMS filter group to the Kalman filter, and obtains the echo-cancelled signal through multiple NLMS filters in the NLMS filter group. The device performs parallel filtering processing to estimate the measurement noise (i.e., the energy of the near-end signal) in the Kalman filter, and cancels the echo on the echo path between the sound collection device and the sound playback device based on the estimated measurement noise. This solution introduces a group of parallel NLMS filters into the NLMS filter and performs filtering processing through multiple NLMS filters in the group. This allows each NLMS filter to estimate only a small number of coefficients, so that it can converge quickly when the echo path changes, thereby improving the efficiency of echo cancellation. At the same time, by combining the NLMS filter with the Kalman filter, it can effectively distinguish the energy of the near-end signal and the echo signal, thereby improving the accuracy of echo cancellation. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0040] Figure 1 This is a flow chart of an echo cancellation method disclosed in this application;
[0041] Figure 2 This is a block diagram of a specific echo cancellation method disclosed in this application;
[0042] Figure 3 This is a schematic structural diagram of an echo cancellation device disclosed in this application;
[0043] Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0045] The present application discloses an echo cancellation method. Figure 1 As shown, the method includes:
[0046] Step S11: Acquire the audio signal currently played by the sound playing device and the audio signal collected by the sound collecting device to obtain a far-end audio signal and a near-end audio signal.
[0047] In this embodiment, when it is necessary to eliminate the echo generated on the echo path between the sound playback device (such as a speaker) and the sound collection device (such as a microphone), the audio signal played by the current sound playback device (such as a speaker) and the audio signal collected by the sound collection device (such as a microphone) can be obtained respectively to obtain the corresponding remote audio signal. and near-end audio signals For details, see Figure 2 As shown, the near-end audio signal Contains all sounds collected by the near-end microphone, which can be expressed as:
[0048] ;
[0049] Where, A mixed signal representing speech and / or background noise picked up by the near-end microphone.
[0050] Specifically, the far-end audio signal Refers to the signal to be played from the local speaker. The signal collected by the microphone after being reflected by the speaker and the room is , Specifically, it can be expressed as:
[0051] ;
[0052] Where, Indicates the echo path.
[0053] Step S12: Obtain a historical residual signal obtained in the previous iteration of the Kalman filter and a historical remote reference signal played by the sound playback device, and input the historical residual signal and the historical remote reference signal into an NLMS filter group located in the Kalman filter for NLMS filtering to obtain multiple filtered residual signals.
[0054] It is understandable that if the echo path changes, the estimation of the echo signal is inaccurate, when the estimated value of near-end energy / measurement noise is To address this issue, the present application introduces a set of parallel NLMS filters after the sub-band Kalman filter. This filter bank takes the residual error signal (Residual Error) of the Kalman filter as input and processes the block-based far-end signal (i.e., the audio signal played by the sound playback device). Since each NLMS filter only needs to estimate a small number of filter coefficients, it can converge quickly when the echo path changes.
[0055] In this embodiment, the remote audio signal played by the current sound playback device is obtained. And the near-end audio signal collected by the sound collection device Afterwards, the historical residual signal obtained from the previous iteration of the Kalman filter pre-created based on the NLMS filter bank and the historical far-end reference signal played by the sound playback device are obtained, and then the above historical residual signal and the above historical far-end reference signal are input into the multiple NLMS filters in the NLMS filter bank for parallel NLMS filtering, thereby obtaining multiple filtered residual signals. It should be noted that the present application pre-creates a Kalman filter containing multiple NLMS filters. The multiple NLMS filters can not only perform parallel filtering operations, but also improve the accuracy of echo cancellation by combining the Kalman filter with the multiple NLMS filters.
[0056] Specifically, the method of obtaining a historical residual signal obtained in a previous iteration of a Kalman filter and a historical far-end reference signal played by the sound playback device, and inputting the historical residual signal and the historical far-end reference signal into an NLMS filter group located in the Kalman filter for NLMS filtering to obtain multiple filtered residual signals may include: obtaining a frequency domain residual signal obtained in a previous iteration of a Kalman filter pre-created based on the NLMS filter group and a frequency domain far-end reference signal played by the sound playback device to obtain a historical residual signal and a historical far-end reference signal; and inputting the historical residual signal and the historical far-end reference signal into the NLMS filter group for NLMS filtering to obtain multiple filtered residual signals. In this embodiment, the residual signal obtained in the previous iteration of the Kalman filter created in advance based on the NLMS filter group and the audio signal played by the sound playback device can be obtained to obtain the corresponding historical residual signal and the historical far-end reference signal. Then, the historical residual signal and the historical far-end reference signal are input into the NLMS filter group located in the Kalman filter to perform parallel NLMS filtering through multiple NLMS filters in the NLMS filter group to obtain corresponding multiple filtered residual signals.
[0057] In this embodiment, the frequency domain residual signal obtained in the previous iteration of the Kalman filter created in advance based on the NLMS filter group and the frequency domain far-end reference signal played by the sound playback device are obtained to obtain the historical residual signal and the historical far-end reference signal. Specifically, it may include: performing frequency domain transformation on the near-end audio signal and the far-end audio signal respectively to obtain the near-end frequency domain signal and the far-end frequency domain signal; obtaining the frequency domain residual signal obtained in the previous iteration of the Kalman filter created in advance based on the NLMS filter group and the frequency domain far-end reference signal played by the sound playback device to obtain the historical residual signal and the historical far-end reference signal. In this embodiment, in order to improve the echo cancellation speed, the time domain audio signal is converted into a frequency domain audio signal. For example, the near-end audio signal collected by the microphone is converted into a frequency domain audio signal. signal and the far-end audio signal played by the loudspeaker The near-end frequency domain signal is obtained by performing frame division, mixed and superimposed window processing respectively, and M-point short-time Fourier transform (STFT) and far-end frequency domain signals , where n is the frame index and k is the frequency index. Then, the residual signal obtained in the last iteration of the Kalman filter and the audio signal played by the sound playback device are obtained to obtain the historical residual signal and historical remote reference signal Among them, the historical remote reference signal Specifically, it can be expressed as:
[0058] ;
[0059] Where L is the maximum number of taps in the Kalman filter and T represents the transpose.
[0060] It can be understood that framing is to divide an infinitely long speech signal into segments, because the speech signal has short-term stability and is easy to process. Windowing is to make the framed speech signal more stable.
[0061] In this embodiment, the inputting of the historical residual signal and the historical remote reference signal into the NLMS filter bank for NLMS filtering to obtain multiple filtered residual signals may specifically include: performing time block processing on the frequency domain buffer of the historical remote reference signal to obtain multiple frequency domain signal blocks, and inputting the historical residual signal and each of the frequency domain signal blocks into the NLMS filter bank for NLMS filtering to obtain multiple filtered residual signals. In this embodiment, the historical remote audio signal played by the speaker is first processed. After block processing, p blocks are obtained. The length of each block is l. The i-th block can be expressed as:
[0062] ;
[0063] in, .
[0064] Next, the historical residual signal and each frequency domain signal block Input into the NLMS filter bank for NLMS filtering to obtain multiple filtered residual signals Specifically, we can first obtain the filter coefficients of the i-th filter in the (n-1)th iteration of the Kalman filter. , where the filter coefficients It can be expressed as:
[0065] ;
[0066] Further, As a remote reference signal, As the microphone signal, it traverses p NLMS filters, that is, , and finally obtain multiple filtered residual signals , Specifically, it can be expressed as:
[0067] .
[0068] In addition, the residual signal can be filtered Calculates the current filter coefficients of the NLMS filter , the specific calculation formula is:
[0069] ;
[0070] Where μ is the filter step size, This is to avoid the denominator being a stabilizing factor of 0.
[0071] Step S13: calculating a current measurement noise estimate of the Kalman filter based on the plurality of filtered residual signals, and calculating a Kalman gain based on the current measurement noise estimate, the historical far-end reference signal, and the process noise of the Kalman filter.
[0072] In this embodiment, after NLMS filtering is performed by multiple filters in the Kalman filter, the residual signal after the multiple filtering can be Calculate the current measurement noise estimate of the Kalman filter , then based on the above current measurement noise estimate , the above historical remote reference signal and process noise of Kalman filter Calculate the Kalman gain of the current frame .
[0073] Specifically, the calculation of the current measurement noise estimate of the Kalman filter based on the multiple filtered residual signals may include: recursively smoothing each of the filtered residual signals to obtain multiple smoothed residuals; and calculating the minimum value of all the smoothed residuals to obtain the current measurement noise estimate of the Kalman filter. In this embodiment, when calculating the current measurement noise estimate of the Kalman filter, When all the residual signals after filtering are Perform recursive smoothing to obtain multiple smoothed residuals, then calculate the minimum value among all the smoothed residuals and use this minimum value as the current measurement noise estimate of the Kalman filter. The specific calculation steps include:
[0074] ;
[0075] ;
[0076] Where, is the smoothing coefficient;
[0077] Next, calculate all Minimum :
[0078] ;
[0079] Finally, the minimum As That is, after calculating the energy of the output residual signal of the NLMS filter bank, the minimum energy value is selected as the estimated value of the near-end signal energy , The specific calculation formula is:
[0080] .
[0081] In this embodiment, the calculation of the Kalman gain based on the current measurement noise estimate, the historical remote reference signal and the process noise of the Kalman filter may specifically include: using the process noise to update the posterior offset matrix obtained in the previous iteration of the Kalman filter to obtain a priori offset matrix; and calculating the Kalman gain based on the priori offset matrix, the current measurement noise estimate and the historical remote reference signal. For example, first using the process noise of the Kalman filter Update the posterior misalignment matrix obtained in the previous iteration to obtain the prior misalignment matrix (posteriori misalignment) , The specific calculation formula is:
[0082] ;
[0083] Where, represents the posterior misalignment matrix obtained at the (n - 1)th iteration, and I is the identity matrix.
[0084] Further, based on and Calculate the Kalman gain. The specific calculation formula is:
[0085] ;
[0086] .
[0087] Step S14: updating the Kalman filter coefficient estimation value obtained in the previous iteration of the Kalman filter based on the Kalman gain and the historical residual signal to obtain the current filter coefficient estimation value.
[0088] In this embodiment, based on the above Kalman gain And the above historical residual signal Update the Kalman filter coefficient estimate obtained in the previous iteration of the Kalman filter to obtain the current filter coefficient estimate The current filter coefficient estimate The specific calculation process includes:
[0089] First, using the Kalman gain Calculate the posterior misalignment matrix for the current frame , the specific calculation formula is:
[0090] ;
[0091] Furthermore, the filter coefficients obtained by the (n - 1)th iteration are obtained. , and combined with Update the filter coefficient of the Kalman filter to obtain the current filter coefficient estimate. The specific calculation formula is:
[0092] ;
[0093] .
[0094] Step S15: canceling the echo generated on the echo path between the sound collection device and the sound playback device based on the current filter coefficient estimate, the far-end audio signal, and the near-end audio signal to obtain an echo-canceled signal.
[0095] In this embodiment, the echo generated on the echo path between the sound collection device and the sound playback device may be eliminated based on the current filter coefficient estimation value, the far-end audio signal, and the near-end audio signal.
[0096] In this embodiment, the echo generated on the echo path between the sound collection device and the sound playback device is eliminated based on the current filter coefficient estimate, the far-end audio signal and the near-end audio signal to obtain the echo-cancelled signal. Specifically, it can include: updating the far-end frequency domain signal based on the historical far-end reference signal to obtain the current audio signal, and calculating the residual signal of the current frame based on the current filter coefficient estimate and the near-end frequency domain signal to obtain the current residual signal; performing a time domain transform on the current residual signal to obtain the audio signal after eliminating the echo generated on the echo path between the sound collection device and the sound playback device to obtain the echo-cancelled signal. Specifically, the far-end frequency domain signal of the nth frame can be used. renew , get the updated audio signal , the specific update formula is:
[0097] ;
[0098] Where, express The first L-1 elements of .
[0099] Next, calculate the current residual signal , the specific calculation formula is:
[0100] .
[0101] Finally, the current residual signal is transformed by the inverse short-time Fourier transform (ISTFT) Perform inverse Fourier transform and merge the frames together by overlap-add (OLA) to obtain the time domain signal .
[0102] As can be seen, the embodiment of the present application first obtains the audio signal currently played by the sound playback device and the audio signal collected by the sound collection device to obtain a far-end audio signal and a near-end audio signal. Then, a historical residual signal obtained in the previous iteration of the Kalman filter and a historical far-end reference signal played by the sound playback device are obtained. The historical residual signal and the historical far-end reference signal are input into the NLMS filter bank located in the Kalman filter for NLMS filtering to obtain multiple filtered residual signals. Then, a current measurement noise estimate of the Kalman filter is calculated based on the multiple filtered residual signals. A Kalman gain is calculated based on the current measurement noise estimate, the historical far-end reference signal, and the process noise of the Kalman filter. Then, the Kalman filter coefficient estimate obtained in the previous iteration of the Kalman filter is updated based on the Kalman gain and the historical residual signal to obtain a current filter coefficient estimate. Finally, the echo generated in the echo path between the sound collection device and the sound playback device is cancelled based on the current filter coefficient estimate, the far-end audio signal, and the near-end audio signal to obtain an echo-cancelled signal. In an embodiment of the present application, the near-end audio signal collected by the sound collection device and the far-end reference signal played by the sound playback device are first acquired, and then the historical residual signal and the historical far-end reference signal are input into the NLMS filter group of the Kalman filter, so as to perform NLMS filtering through multiple NLMS filters in the NLMS filter group to obtain multiple filtered residual signals, and calculate the filter coefficient estimation value of the current Kalman filter based on the multiple filtered residual signals, and finally obtain the echo-cancelled signal based on the filter coefficient estimation value, the far-end audio signal and the near-end audio signal. It can be seen that the present application adds an NLMS filter group to the Kalman filter, and uses multiple NLMS in the NLMS filter group to obtain multiple filtered residual signals. The filter performs parallel filtering processing to estimate the measurement noise (i.e., the energy of the near-end signal) in the Kalman filter, and cancels the echo on the echo path between the sound collection device and the sound playback device based on the estimated measurement noise. This solution introduces a group of parallel NLMS filters into the NLMS filter and performs filtering processing through multiple NLMS filters in the group. This allows each NLMS filter to estimate only a small number of coefficients, enabling rapid convergence when the echo path changes, thereby improving the efficiency of echo cancellation. At the same time, by combining the NLMS filter with the Kalman filter, the energy of the near-end signal and the echo signal can be effectively distinguished, thereby improving the accuracy of echo cancellation.
[0103] Correspondingly, the embodiment of the present application also discloses an echo cancellation device, see Figure 3 As shown, the device includes:
[0104] The first signal acquisition module 11 is used to acquire the audio signal currently played by the sound playback device and the audio signal collected by the sound collection device to obtain a far-end audio signal and a near-end audio signal;
[0105] A second signal acquisition module 12 is configured to acquire a historical residual signal obtained in a previous iteration of the Kalman filter and a historical remote reference signal played by the sound playback device;
[0106] A filtering module 13 is configured to input the historical residual signal and the historical remote reference signal into an NLMS filter bank located in the Kalman filter to perform NLMS filtering to obtain a plurality of filtered residual signals;
[0107] a calculation module 14, configured to calculate a current measurement noise estimate of the Kalman filter based on the plurality of filtered residual signals, and calculate a Kalman gain based on the current measurement noise estimate, the historical remote reference signal, and the process noise of the Kalman filter;
[0108] An updating module 15 is configured to update the Kalman filter coefficient estimation value obtained in the previous iteration of the Kalman filter based on the Kalman gain and the historical residual signal to obtain a current filter coefficient estimation value;
[0109] The echo cancellation module 16 is configured to cancel the echo generated on the echo path between the sound collection device and the sound playback device based on the current filter coefficient estimate, the far-end audio signal, and the near-end audio signal to obtain an echo-canceled signal.
[0110] Among them, the specific work processes of the above modules can refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0111] It can be seen that in the embodiment of the present application, the audio signal currently played by the sound playback device and the audio signal collected by the sound collection device are first obtained to obtain a far-end audio signal and a near-end audio signal. Then, a historical residual signal obtained in the previous iteration of the Kalman filter and a historical far-end reference signal played by the sound playback device are obtained, and the historical residual signal and the historical far-end reference signal are input into the NLMS filter bank located in the Kalman filter for NLMS filtering to obtain multiple filtered residual signals. Then, a current measurement noise estimate of the Kalman filter is calculated based on the multiple filtered residual signals, and a Kalman gain is calculated based on the current measurement noise estimate, the historical far-end reference signal, and the process noise of the Kalman filter. Then, the Kalman filter coefficient estimate obtained in the previous iteration of the Kalman filter is updated based on the Kalman gain and the historical residual signal to obtain a current filter coefficient estimate. Finally, based on the current filter coefficient estimate, the far-end audio signal, and the near-end audio signal, the echo generated in the echo path between the sound collection device and the sound playback device is cancelled to obtain an echo-cancelled signal. In an embodiment of the present application, the near-end audio signal collected by the sound collection device and the far-end audio signal played by the sound playback device are first acquired, and then the historical residual signal and the historical far-end reference signal are input into the NLMS filter group of the Kalman filter, so as to perform NLMS filtering through multiple NLMS filters in the NLMS filter group to obtain multiple filtered residual signals, and calculate the filter coefficient estimation value of the current Kalman filter based on the multiple filtered residual signals, and finally obtain the echo-cancelled signal based on the filter coefficient estimation value, the far-end audio signal and the near-end audio signal. It can be seen that the present application adds an NLMS filter group to the Kalman filter, and uses multiple NLMS in the NLMS filter group to obtain multiple filtered residual signals. The filter performs parallel filtering processing to estimate the measurement noise (i.e., the energy of the near-end signal) in the Kalman filter, and cancels the echo on the echo path between the sound collection device and the sound playback device based on the estimated measurement noise. This solution introduces a group of parallel NLMS filters into the NLMS filter and performs filtering processing through multiple NLMS filters in the group. This allows each NLMS filter to estimate only a small number of coefficients, enabling rapid convergence when the echo path changes, thereby improving the efficiency of echo cancellation. At the same time, by combining the NLMS filter with the Kalman filter, the energy of the near-end signal and the echo signal can be effectively distinguished, thereby improving the accuracy of echo cancellation.
[0112] In some specific embodiments, the second signal acquisition module 12 may specifically include:
[0113] A first signal acquisition unit is configured to acquire a frequency domain residual signal obtained in a previous iteration of a Kalman filter pre-created based on an NLMS filter bank and a frequency domain far-end reference signal played by the sound playback device, to obtain a historical residual signal and a historical far-end reference signal;
[0114] Accordingly, the filtering module 13 may specifically include:
[0115] The first filtering unit is configured to input the historical residual signal and the historical remote reference signal into the NLMS filter bank for NLMS filtering to obtain a plurality of filtered residual signals.
[0116] In some specific embodiments, the first signal acquisition unit may specifically include:
[0117] a frequency domain transform unit, configured to perform frequency domain transform on the near-end audio signal and the far-end audio signal respectively to obtain a near-end frequency domain signal and a far-end frequency domain signal;
[0118] The second signal acquisition unit is used to obtain the frequency domain residual signal obtained in the previous iteration of the Kalman filter created in advance based on the NLMS filter group and the frequency domain far-end reference signal played by the sound playback device to obtain the historical residual signal and the historical far-end reference signal.
[0119] In some specific embodiments, the echo cancellation module 16 may specifically include:
[0120] A first updating unit, configured to update the remote frequency domain signal based on the historical remote reference signal to obtain a current audio signal;
[0121] a first calculation unit, configured to calculate a residual signal of a current frame based on the current filter coefficient estimate and the near-end frequency domain signal to obtain a current residual signal;
[0122] The time domain transform unit is configured to perform a time domain transform on the current residual signal to obtain an audio signal after canceling the echo generated on the echo path between the sound collection device and the sound playback device, thereby obtaining an echo-canceled signal.
[0123] In some specific embodiments, the first filtering unit may specifically include:
[0124] A blocking unit, configured to perform time blocking processing on the frequency domain buffer of the historical remote reference signal to obtain a plurality of frequency domain signal blocks;
[0125] The second filtering unit is configured to input the historical residual signal and each of the frequency domain signal blocks into the NLMS filter bank for NLMS filtering to obtain a plurality of filtered residual signals.
[0126] In some specific embodiments, the calculation module 14 may specifically include:
[0127] a smoothing processing unit, configured to perform recursive smoothing processing on each of the filtered residual signals to obtain a plurality of smoothed residuals;
[0128] The second calculation unit is used to calculate the minimum value of all the smoothed residuals to obtain the current measurement noise estimation value of the Kalman filter.
[0129] In some specific embodiments, the calculation module 14 may specifically include:
[0130] A second updating unit is configured to update the posterior offset matrix obtained in the previous iteration of the Kalman filter using process noise to obtain a priori offset matrix;
[0131] A third calculation unit is configured to calculate a Kalman gain based on the prior offset matrix, the current measurement noise estimate, and the historical far-end reference signal.
[0132] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.
[0133] Figure 4 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the echo cancellation method disclosed in any of the aforementioned embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0134] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0135] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0136] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20. The operating system 221 can be Windows Server, NetWare, Unix, Linux, etc. In addition to including a computer program capable of implementing the echo cancellation method performed by the electronic device 20 as disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program capable of performing other specific tasks.
[0137] Furthermore, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned echo cancellation method. The specific steps of this method can be referred to the corresponding contents disclosed in the aforementioned embodiments and will not be repeated here.
[0138] Furthermore, an embodiment of the present application also discloses a computer program product, including a computer program / instruction, which implements the steps of the echo cancellation method disclosed above when executed by a processor.
[0139] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0140] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0141] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0142] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0143] The above is a detailed introduction to the echo cancellation method, device, equipment and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. At the same time, for those skilled in the art, based on the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. An echo cancellation method, characterized in that: include: Acquire the audio signal played by the current sound playback device and the audio signal collected by the sound collection device to obtain a far-end audio signal and a near-end audio signal; Obtaining a historical residual signal obtained in a previous iteration of a Kalman filter and a historical remote reference signal played by the sound playback device, and inputting the historical residual signal and the historical remote reference signal into an NLMS filter bank located in the Kalman filter for NLMS filtering to obtain a plurality of filtered residual signals; Calculating a current measurement noise estimate of the Kalman filter based on the plurality of filtered residual signals, and calculating a Kalman gain based on the current measurement noise estimate, the historical remote reference signal, and process noise of the Kalman filter; updating the Kalman filter coefficient estimate obtained in the previous iteration of the Kalman filter based on the Kalman gain and the historical residual signal to obtain a current filter coefficient estimate; Echoes generated on an echo path between the sound collection device and the sound playback device are eliminated based on the current filter coefficient estimate, the far-end audio signal, and the near-end audio signal to obtain an echo-cancelled signal.
2. The echo cancellation method according to claim 1, wherein: The method further comprises: obtaining a historical residual signal obtained in a previous iteration of the Kalman filter and a historical remote reference signal played by the sound playback device, and inputting the historical residual signal and the historical remote reference signal into an NLMS filter bank located in the Kalman filter for NLMS filtering to obtain a plurality of filtered residual signals, including: Acquire a frequency domain residual signal obtained in a previous iteration of a Kalman filter pre-created based on an NLMS filter bank and a frequency domain far-end reference signal played by the sound playback device to obtain a historical residual signal and a historical far-end reference signal; The historical residual signal and the historical far-end reference signal are input into the NLMS filter bank for NLMS filtering to obtain a plurality of filtered residual signals.
3. The echo cancellation method according to claim 2, wherein: The step of obtaining a frequency domain residual signal obtained in a previous iteration of a Kalman filter pre-created based on an NLMS filter bank and a frequency domain far-end reference signal played by the sound playback device to obtain a historical residual signal and a historical far-end reference signal includes: Performing frequency domain transformation on the near-end audio signal and the far-end audio signal respectively to obtain a near-end frequency domain signal and a far-end frequency domain signal; A frequency domain residual signal obtained in the last iteration of the Kalman filter pre-created based on the NLMS filter bank and a frequency domain far-end reference signal played by the sound playing device are obtained to obtain a historical residual signal and a historical far-end reference signal.
4. The echo cancellation method according to claim 3, wherein: The canceling of the echo generated on the echo path between the sound collection device and the sound playback device based on the current filter coefficient estimate, the far-end audio signal, and the near-end audio signal to obtain an echo-canceled signal includes: updating the far-end frequency domain signal based on the historical far-end reference signal to obtain a current audio signal, and calculating a residual signal of a current frame based on the current filter coefficient estimate and the near-end frequency domain signal to obtain a current residual signal; The current residual signal is transformed in the time domain to obtain an audio signal after the echo generated on the echo path between the sound collection device and the sound playback device is eliminated, thereby obtaining an echo-cancelled signal.
5. The echo cancellation method according to claim 4, wherein: The step of inputting the historical residual signal and the historical remote reference signal into the NLMS filter bank for NLMS filtering to obtain a plurality of filtered residual signals includes: The frequency domain buffer of the historical remote reference signal is subjected to time block processing to obtain multiple frequency domain signal blocks, and the historical residual signal and each of the frequency domain signal blocks are input into the NLMS filter bank for NLMS filtering to obtain multiple filtered residual signals.
6. The echo cancellation method according to claim 1, wherein: The calculating a current measurement noise estimate of the Kalman filter based on the plurality of filtered residual signals comprises: Recursively smoothing each of the filtered residual signals to obtain a plurality of smoothed residuals; The minimum value of all the smoothed residuals is calculated to obtain a current measurement noise estimate of the Kalman filter.
7. The echo cancellation method according to any one of claims 1 to 6, characterized in that: The calculating the Kalman gain based on the current measurement noise estimate, the historical remote reference signal, and the process noise of the Kalman filter includes: Using process noise to update the posterior offset matrix obtained in the previous iteration of the Kalman filter to obtain a priori offset matrix; A Kalman gain is calculated based on the prior offset matrix, the current measurement noise estimate, and the historical far-end reference signal.
8. An echo cancellation device, characterized in that: include: A first signal acquisition module is used to acquire the audio signal played by the current sound playback device and the audio signal collected by the sound collection device to obtain a far-end audio signal and a near-end audio signal; A second signal acquisition module is used to obtain a historical residual signal obtained in the last iteration of the Kalman filter and a historical remote reference signal played by the sound playback device; A filtering module, configured to input the historical residual signal and the historical remote reference signal into an NLMS filter bank located in the Kalman filter to perform NLMS filtering to obtain a plurality of filtered residual signals; a calculation module, configured to calculate a current measurement noise estimate of the Kalman filter based on the plurality of filtered residual signals, and calculate a Kalman gain based on the current measurement noise estimate, the historical remote reference signal, and the process noise of the Kalman filter; An updating module is configured to update the Kalman filter coefficient estimate obtained in the previous iteration of the Kalman filter based on the Kalman gain and the historical residual signal to obtain a current filter coefficient estimate; The echo cancellation module is used to cancel the echo generated on the echo path between the sound collection device and the sound playback device based on the current filter coefficient estimate, the far-end audio signal and the near-end audio signal to obtain an echo-canceled signal.
9. An electronic device, characterized in that: The method comprises a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the echo cancellation method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that Used to store a computer program; wherein, when the computer program is executed by a processor, the echo cancellation method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Broadband acoustics echo eliminating method
CN101227537A
Echo cancellation method and device
CN112687285A