A dual-filter-based echo suppression method
Through the echo suppression method combining dual filters, the error signal and noise energy smoothing estimation are calculated, and the correction coefficient is generated for echo suppression. This solves the residual echo problem of the PBFDKF algorithm in rapidly changing scenarios and achieves effective echo suppression effect.
Patent Information
- Application Number
- CN202310145777.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-21
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-02-21
AI Technical Summary
The existing PBFDKF algorithm has insufficient tracking speed in scenarios where the echo path changes rapidly, such as when the device moves, resulting in a large amount of residual echo leakage and an inability to effectively suppress echo noise.
An echo suppression method based on dual filters is adopted. By combining a fast filter and a slow filter, the error signal and noise energy smoothing estimation are calculated, and correction coefficients are generated for echo suppression. The method includes error signal calculation, noise energy smoothing estimation and coefficient updating of the fast and slow filters.
In duplex communication, it suppresses residual echo to the greatest extent, reduces echo noise in the output signal, and improves the effect and stability of echo suppression.
Smart Images

Figure CN116168715B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech enhancement, and in particular to an echo suppression method based on dual filters. Background Art
[0002] The existing PBFDKF (partitioned block frequency-domain Kalman filter) is an improved MDF (multi-delay block frequency-domain adaptive filter). It uses the Kalman state-space model to optimize the filter step size in both single- and double-mode modes, thereby outputting a signal with less residual echo. However, in scenarios where the echo path changes rapidly, such as when devices are moving, PBFDKF's tracking speed is significantly insufficient, resulting in slow or even non-convergence, and a high level of residual echo leakage. Therefore, reducing residual echo in such time-varying communication scenarios is an urgent problem that needs to be addressed. Summary of the Invention
[0003] The embodiment of the present invention provides an echo suppression method based on a dual filter, which can suppress residual echo to the greatest extent in duplex communication and reduce echo noise contained in the output signal.
[0004] An embodiment of the present invention provides an echo suppression method based on a dual filter, comprising:
[0005] Acquire an input signal; wherein the input signal includes: a signal to be extracted and an interference signal;
[0006] Performing echo suppression processing on each frame of input signal to generate a first output signal corresponding to each frame of input signal;
[0007] generating a final output signal according to the first output signal of each frame;
[0008] The echo suppression process includes:
[0009] Calculating a first error signal of the input signal of the current frame in the fast filter according to the input signal of the current frame and the fast filter coefficient of the previous frame; calculating a second error signal of the input signal of the current frame in the slow filter according to the input signal of the current frame and the slow filter coefficient of the previous frame;
[0010] Calculating a fast filter noise energy smoothing estimate for the current frame based on the first error signal and a preset fast filter noise estimation smoothing factor; calculating a slow filter noise energy smoothing estimate for the current frame based on the second error signal and a preset slow filter noise estimation smoothing factor;
[0011] Calculate the post-filter coefficient of the fast filter in the current frame according to the preset attenuation factor, the fast filter Kalman gain of the current frame and the input signal of the current frame;
[0012] Calculating a correction coefficient of the first error signal based on a fast filter noise energy smoothing estimate in the current frame, a slow filter noise energy smoothing estimate in the current frame, a smoothing speed, and a preset error ratio threshold;
[0013] The error signal of the fast filter is subjected to echo suppression by using the correction coefficient of the first error signal and the post-filtering coefficient to generate a first output signal corresponding to the input signal of the current frame.
[0014] Furthermore, the step of calculating a first error signal of the input signal of the current frame in the fast filter according to the input signal of the current frame and the fast filter coefficient of the previous frame; and calculating a second error signal of the input signal of the current frame in the slow filter according to the input signal of the current frame and the slow filter coefficient of the previous frame includes:
[0015] The first error signal is calculated by the following formula:
[0016]
[0017] Among them, e1 represents the first error signal, d1 represents the signal to be extracted in the current frame, S1 represents the total number of blocks of the fast filter, s represents the sth block of the fast filter, W1 上s Indicates the fast filter coefficient of the sth block in the previous frame; X s represents the interference signal of the current frame under the sth block;
[0018] The second error signal is calculated by the following formula:
[0019]
[0020] Wherein, e2 represents the second error signal, S2 represents the total number of slow filter blocks, s represents the sth slow filter block, W2 上s Represents the slow filter coefficients of the sth block in the previous frame.
[0021] Further, the calculating of a fast filter noise energy smoothing estimate in the current frame according to the first error signal and a preset fast filter noise estimation smoothing factor; and the calculating of a slow filter noise energy smoothing estimate in the current frame according to the second error signal and a preset slow filter noise estimation smoothing factor include:
[0022] The fast filter noise energy smoothing estimate for the current frame is calculated using the following formula:
[0023] E1=fft(F 01 e1)
[0024] VV1=lam1*VV1+(1-lam1)*E1*conj(E1)
[0025] Wherein, VV1 represents the fast filter noise energy smoothing estimation in the current frame, lam1 represents the preset fast filter noise estimation smoothing factor, and E1 represents the first error signal in the frequency domain state;
[0026] The slow filter noise energy smoothing estimate for the current frame is calculated using the following formula:
[0027] E2=fft(F 01 e2)
[0028] VV2=lam2*VV2+(1-lam2)*E2*conj(E2)
[0029] Among them, VV2 represents the slow filter noise energy smoothing estimation in the current frame, lam2 represents the preset slow filter noise estimation smoothing factor, and E2 represents the second error signal in the frequency domain state.
[0030] Furthermore, the generation of the fast filter coefficients of the current frame includes:
[0031] Generate a smoothed estimate of the interference signal of the current frame according to a preset interference energy smoothing factor and the interference signal of the current frame;
[0032] When the smoothed estimate of the interference signal of the current frame is greater than the preset interference energy threshold, the fast filter coefficient error covariance of the current frame is calculated based on the fast filter coefficient of the previous frame, the fast filter coefficient error covariance of the previous frame, the fast filter Kalman gain of the current frame and the preset fast filter control factor; the fast filter Kalman gain of the current frame is calculated based on the smoothed estimate of the fast filter noise energy of the current frame, the fast filter coefficient of the previous frame, the fast filter coefficient error covariance of the previous frame and the interference signal; the fast filter coefficient of the current frame is generated based on the fast filter coefficient of the previous frame and the fast filter Kalman gain of the current frame.
[0033] Furthermore, the generation of the slow filter coefficients of the current frame includes:
[0034] When the smoothed estimate of the interference signal of the current frame is greater than the preset interference energy threshold, the slow filter coefficient error covariance of the current frame is calculated based on the slow filter coefficient of the previous frame, the error covariance of the slow filter coefficient of the previous frame, the slow filter Kalman gain of the current frame and the preset slow filter control factor; the slow filter Kalman gain of the current frame is calculated based on the smoothed estimate of the slow filter noise energy of the current frame, the slow filter coefficient of the previous frame, the error covariance of the slow filter coefficient of the previous frame and the interference signal; the slow filter coefficient of the current frame is generated based on the slow filter coefficient of the previous frame and the slow filter Kalman gain of the current frame.
[0035] Furthermore, after generating the fast filter coefficient of the current frame and the slow filter coefficient of the current frame, the method further includes:
[0036] Calculating the sum of all frequency points of the fast filter noise energy smoothing estimate of the current frame to obtain a first frequency point value; calculating the sum of all frequency points of the slow filter noise energy smoothing estimate of the current frame to obtain a second frequency point value; calculating the ratio between the first frequency point value and the second frequency point value to obtain a total error energy ratio; generating a coefficient of variation based on the fast filter coefficients and slow filter coefficients of several historical frames;
[0037] The relationship between the total energy ratio of the comparison error, the coefficient of variation and the set threshold;
[0038] When the total error energy ratio is greater than the first threshold and less than the second threshold, the slow filter coefficient of the current frame is assigned to the fast filter coefficient of the current frame to generate an updated fast filter coefficient of the current frame.
[0039] When the total error energy ratio is less than or equal to a first threshold, assigning the slow filter coefficient of the current frame to the fast filter coefficient of the current frame to generate an updated fast filter coefficient of the current frame, and initializing the fast filter coefficient error covariance and the slow filter coefficient error covariance of the current frame;
[0040] When the total error energy ratio is greater than the third threshold and the coefficient of variation is less than the fourth threshold, the fast filter coefficient of the current frame is assigned to the slow filter coefficient of the current frame to generate an updated slow filter coefficient of the current frame.
[0041] Furthermore, the calculation of the post-filter coefficient of the fast filter in the current frame according to the preset attenuation factor, the Kalman gain of the fast filter in the current frame, and the input signal of the current frame includes:
[0042] The post-filter coefficient of the fast filter in the current frame is calculated using the following formula:
[0043]
[0044] H = max(H, 0)
[0045] Among them, H represents the post-filter coefficient of the fast filter of the current frame, Q represents the preset attenuation factor, K1 represents the Kalman gain of the fast filter of the current frame, X s Represents the input signal of the current frame under the sth block, K1 s Represents the fast filter Kalman gain of the current frame under the s-th block.
[0046] Furthermore, the calculation of the correction coefficient of the first error signal according to the fast filter noise energy smoothing estimate in the current frame, the slow filter noise energy smoothing estimate in the current frame, the smoothing speed, and a preset error ratio threshold includes:
[0047] Grouping the frequency points in the fast filter noise energy smoothing estimate for the current frame into pairs to generate a plurality of first frequency point groups; wherein each first frequency point group consists of an odd-numbered frequency point and an adjacent even-numbered frequency point; assigning an average value of each first frequency point group to the odd-numbered frequency points and the even-numbered frequency points in the corresponding first frequency point group to generate a second frequency point group; and generating a second fast filter noise energy smoothing estimate for the current frame based on each second frequency point group;
[0048] Grouping the frequency points in the slow filter noise energy smoothing estimate of the current frame into pairs to generate third frequency point groups; wherein each third frequency point group consists of an odd-numbered frequency point and an adjacent even-numbered frequency point; assigning the average value of each third frequency point group to the odd-numbered frequency points and the even-numbered frequency points in the corresponding third frequency point group to generate a fourth frequency point group; generating a second slow filter noise energy smoothing estimate of the current frame based on each fourth frequency point group;
[0049] Calculating a ratio between a noise energy smoothing estimate of the second fast filter and a noise energy smoothing estimate of the second slow filter of the current frame to obtain a noise energy ratio of the current frame;
[0050] Determine the corresponding smoothing speed in the sigmoid function according to the noise energy ratio of the current frame;
[0051] The error signal correction coefficient of the current frame is calculated using the following formula:
[0052]
[0053] Wherein, H1 represents the error signal correction coefficient of the current frame, B represents the smoothing speed, dt_P0 represents the preset error ratio threshold, and VV21 represents the noise energy ratio.
[0054] Furthermore, performing echo suppression on the error signal of the fast filter by using the correction coefficient of the first error signal and the post-filter coefficient to generate a first output signal corresponding to the input signal of the current frame includes:
[0055] Calculating the product of the correction coefficient of the first error signal and the post-filtering coefficient to obtain the correction coefficient of the second error signal;
[0056] The product of the correction coefficients of the first error signal and the second error signal in the frequency domain is calculated to generate a first output signal corresponding to the input signal of the current frame.
[0057] Furthermore, before generating the final output signal according to the first output signal of each frame, the method further includes:
[0058] Performing residual echo detection on the first output signal of each frame, and performing secondary echo suppression on the first output signal with residual echo;
[0059] Wherein, the secondary echo suppression includes:
[0060] Calculate the sum of the absolute values of the frequency points of the low-frequency part of the first output signal of the current frame to obtain a third frequency point value;
[0061] Calculate the sum of the absolute values of the frequency points of the high frequency part of the first output signal of the current frame to obtain a fourth frequency point value;
[0062] Calculating a ratio between the third frequency point value and the fourth frequency point value to obtain a first frequency point ratio;
[0063] When the first frequency point ratio is less than a fifth threshold, setting the correction coefficient of the second error signal to zero to generate a correction coefficient of the third error signal;
[0064] The product of the first error signal and the correction coefficient of the third error signal in the frequency domain is calculated to generate a corresponding updated first output signal.
[0065] The following beneficial effects are achieved by implementing the present invention:
[0066] The present invention provides an echo suppression method based on a dual filter. The echo suppression method obtains an input signal, wherein the input signal includes: a signal to be extracted and an interference signal; performs echo suppression processing on each frame of the input signal to generate a first output signal corresponding to each frame of the input signal; and generates a final output signal based on the first output signal of each frame. The method calculates an error signal contained in the input signal, calculates corresponding correction coefficients using two filters, a fast filter and a slow filter, performs echo suppression on the error signal using the correction coefficients, and finally outputs an echo-suppressed signal. The present invention suppresses the error signal by calculating the correction coefficients to suppress the residual echo to the greatest extent in duplex communication, thereby reducing the echo noise contained in the output signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is a flowchart of a dual-filter-based echo suppression method provided by one embodiment of the present invention.
[0068] Figure 2 The figure is a schematic diagram of an echo suppression process provided by an embodiment of the present invention.
[0069] Figure 3 It is a flowchart of a noise energy processing method provided by one embodiment of the present invention.
[0070] Figure 4This is a schematic diagram of a smooth mapping of a sigmoid function provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0071] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0072] like Figure 1 As shown, an embodiment of the present invention provides an echo suppression method based on a dual filter, comprising:
[0073] Step S1: obtaining an input signal; wherein the input signal includes: a signal to be extracted and an interference signal;
[0074] Step S2: performing echo suppression processing on each frame of input signal to generate a first output signal corresponding to each frame of input signal;
[0075] Step S3: generating a final output signal according to the first output signal of each frame;
[0076] For step S1, the signal to be extracted in the input signal obtained is generally a microphone signal, and the interference signal in the input signal is generally a speaker signal. Specifically, a notch filter is used to remove DC from the signal to be extracted in the input signal to eliminate invalid signal components in the signal to be extracted, thereby avoiding interference caused by the invalid components on the filter convergence characteristics, thereby enhancing the stability of the filter convergence; the signal to be extracted is updated to the signal to be extracted after the invalid signal is eliminated to perform subsequent steps; the DC removal operation can be based on the system function of the notch filter DC removal method to process the signal to be processed. The following is the system function of the notch filter DC removal method:
[0077]
[0078] Wherein, z represents the signal to be processed, and H(z) represents the signal to be processed generated after DC removal.
[0079] For step S2, specifically, the input signal obtained in step S1 is divided into frames to obtain the input signal of each frame, and the input signal of each frame is subjected to echo suppression processing;
[0080] like Figure 2 As shown, a schematic diagram of an echo suppression process flow provided by an embodiment of the present invention includes:
[0081] Step S2.1: Calculate a first error signal of the input signal of the current frame in the fast filter based on the input signal of the current frame and the fast filter coefficients of the previous frame; and calculate a second error signal of the input signal of the current frame in the slow filter based on the input signal of the current frame and the slow filter coefficients of the previous frame;
[0082] Step S2.2: Calculate a fast filter noise energy smoothing estimate for the current frame based on the first error signal and a preset fast filter noise estimation smoothing factor; calculate a slow filter noise energy smoothing estimate for the current frame based on the second error signal and a preset slow filter noise estimation smoothing factor;
[0083] Step S2.3: Calculate the post-filter coefficient of the fast filter in the current frame according to the preset attenuation factor, the Kalman gain of the fast filter in the current frame, and the input signal of the current frame;
[0084] Step S2.4: Calculating a correction coefficient for the first error signal based on a fast filter noise energy smoothing estimate for the current frame, a slow filter noise energy smoothing estimate for the current frame, a smoothing speed, and a preset error ratio threshold;
[0085] Step S2.5: performing echo suppression on the error signal of the fast filter using the correction coefficient of the first error signal and the post-filter coefficient, and generating a first output signal corresponding to the input signal of the current frame.
[0086] Regarding step S2.1, in a preferred embodiment, calculating a first error signal of the input signal of the current frame in a fast filter based on the input signal of the current frame and the fast filter coefficients of the previous frame; and calculating a second error signal of the input signal of the current frame in a slow filter based on the input signal of the current frame and the slow filter coefficients of the previous frame, comprises:
[0087] The first error signal is calculated by the following formula:
[0088]
[0089] Among them, e1 represents the first error signal, d1 represents the signal to be extracted in the current frame, S1 represents the total number of blocks of the fast filter, s represents the sth block of the fast filter, W1 上s Indicates the fast filter coefficient of the sth block in the previous frame; X s represents the interference signal of the current frame under the sth block;
[0090] The second error signal is calculated by the following formula:
[0091]
[0092] Wherein, e2 represents the second error signal, S2 represents the total number of slow filter blocks, s represents the sth slow filter block, W2 上s Represents the slow filter coefficient of the sth block in the previous frame;
[0093] It should be noted that the total number of blocks in the above formula is calculated by the ratio between the total length of the filter and the moving block length. The G used in the formula is 01 Indicates that the current result only retains the last L elements, F 01 Indicates adding L zero elements in front of the current result, where L is the moving block length of the current block number.
[0094] For step S2.2, in a preferred embodiment, calculating a fast filter noise energy smoothing estimate for the current frame based on the first error signal and a preset fast filter noise estimation smoothing factor; calculating a slow filter noise energy smoothing estimate for the current frame based on the second error signal and a preset slow filter noise estimation smoothing factor includes:
[0095] The fast filter noise energy smoothing estimate for the current frame is calculated using the following formula:
[0096] E1=fft(F 01 e1)
[0097] VV1=lam1*VV1+(1-lam1)*E1*conj(E1)
[0098] Wherein, VV1 represents the fast filter noise energy smoothing estimation in the current frame, lam1 represents the preset fast filter noise estimation smoothing factor, and E1 represents the first error signal in the frequency domain state;
[0099] The slow filter noise energy smoothing estimate for the current frame is calculated using the following formula:
[0100] E2=fft(F 01 e2)
[0101] VV2=lam2*VV2+(1-lam2)*E2*conj(E2)
[0102] Wherein, VV2 represents the slow filter noise energy smoothing estimation in the current frame, lam2 represents the preset slow filter noise estimation smoothing factor, and E2 represents the second error signal in the frequency domain state;
[0103] It should be noted that the preset fast filter noise estimation smoothing factor and the preset slow filter noise estimation smoothing factor are usually 0.985, which is an empirical value obtained based on hundreds of groups of parameter adjustment tests that conform to actual scenarios.
[0104] In a preferred embodiment, the generation of the fast filter coefficients of the current frame includes:
[0105] Generate a smoothed estimate of the interference signal of the current frame according to a preset interference energy smoothing factor and the interference signal of the current frame;
[0106] When the smoothed estimate of the interference signal of the current frame is greater than a preset interference energy threshold, the fast filter coefficient error covariance of the current frame is calculated based on the fast filter coefficient of the previous frame, the fast filter coefficient error covariance of the previous frame, the fast filter Kalman gain of the current frame and the preset fast filter control factor; the fast filter Kalman gain of the current frame is calculated based on the smoothed estimate of the fast filter noise energy of the current frame, the fast filter coefficient of the previous frame, the fast filter coefficient error covariance of the previous frame and the interference signal; the fast filter coefficient of the current frame is generated based on the fast filter coefficient of the previous frame and the fast filter Kalman gain of the current frame;
[0107] Specifically, the input signal of the current frame is converted from a time domain signal to a frequency domain signal using Fourier transform. Converting the time domain signal to a frequency domain signal can speed up the calculation. It should be noted that the dimension of the interference signal is twice the moving block length; the dimension of the signal to be extracted is one times the moving block length.
[0108] The smoothed estimate of the interference signal in the current frame is calculated using the following formula:
[0109] P0=alp*P0+(1-alp)*(X1*conj(X1))
[0110] Wherein, P0 represents the smoothed estimate of the interference signal of the current frame, alp represents the preset interference energy smoothing factor, and X1 represents the interference signal of the current frame; wherein, the preset interference energy smoothing factor is generally set to 0.7, which is an empirical value obtained based on hundreds of sets of parameter adjustment tests and conforms to actual scenarios;
[0111] To enhance the robustness of the fast filter, when generating the fast filter coefficients, it is determined whether the fast filter coefficients of the current frame need to be updated. Specifically, when the smoothed estimate of the interference signal of the current frame is greater than a preset interference energy threshold, the fast filter coefficient generation operation of the current frame is executed. The fast filter coefficient generation operation of the current frame includes:
[0112] The fast filter Kalman gain of the current frame is calculated by the following formula:
[0113]
[0114] Among them, K1 represents the fast filter Kalman gain of the current frame, P1 上 Indicates the fast filter coefficient error covariance of the previous frame, P1 上srepresents the error covariance of the fast filter coefficient of the sth block in the previous frame;
[0115] The coefficient error covariance of the fast filter in the current frame is calculated using the following formula;
[0116] P1=[I-0.5*K1*conj(X)]*P1 上 +[1-(A1) 2 ]*W1 上 *conj(W1 上 )
[0117] Where P1 represents the coefficient error covariance of the fast filter in the current frame, A1 represents the control factor of the preset fast filter, I represents the unit matrix, and P1 上 Indicates the fast filter coefficient error covariance of the previous frame, W1 上 Indicates the fast filter coefficient of the previous frame;
[0118] The coefficients of the fast filter in the current frame are calculated using the following formula:
[0119] W1=W1 上 +fft(F 10 G 10 ifft(K1*E1))
[0120] Among them, W1 represents the coefficient of the fast filter in the current frame, W1 上 Indicates the coefficients of the fast filter of the previous frame.
[0121] In another preferred embodiment, the generation of the slow filter coefficients of the current frame includes:
[0122] When the smoothed estimate of the interference signal of the current frame is greater than a preset interference energy threshold, the slow filter coefficient error covariance of the current frame is calculated based on the slow filter coefficient of the previous frame, the error covariance of the slow filter coefficient of the previous frame, the slow filter Kalman gain of the current frame and the preset slow filter control factor; the slow filter Kalman gain of the current frame is calculated based on the smoothed estimate of the slow filter noise energy of the current frame, the slow filter coefficient of the previous frame, the error covariance of the slow filter coefficient of the previous frame and the interference signal; the slow filter coefficient of the current frame is generated based on the slow filter coefficient of the previous frame and the slow filter Kalman gain of the current frame;
[0123] Specifically, to enhance the robustness of the slow filter, when generating the slow filter coefficient, it is determined whether the slow filter coefficient of the current frame needs to be updated; specifically, when the smoothed estimate of the interference signal of the current frame is greater than a preset interference energy threshold, the slow filter coefficient update operation of the current frame is performed; the slow filter coefficient update operation of the current frame includes:
[0124] The slow filter Kalman gain of the current frame is calculated by the following formula:
[0125]
[0126] Among them, K2 represents the slow filter Kalman gain of the current frame, P2 上 Indicates the error covariance of the slow filter coefficient of the previous frame, P2 上s represents the error covariance of the slow filter coefficient of the sth block in the previous frame;
[0127] The coefficient error covariance of the slow filter in the current frame is calculated using the following formula;
[0128] P2=[I-0.5*K2*conj(X)]*P2 上 +[1-(A2) 2 ]*W2 上 *conj(W2 上 )
[0129] Among them, P2 represents the coefficient error covariance of the slow filter in the current frame, A2 represents the control factor of the preset slow filter, I represents the unit matrix, and P2 上 Indicates the fast filter coefficient error covariance of the previous frame, W2 上 Indicates the fast filter coefficient of the previous frame;
[0130] The coefficients of the slow filter in the current frame are calculated using the following formula:
[0131] W2=W2 上 +fft(F 10 G 10 ifft(K2*E2))
[0132] Among them, W2 represents the coefficient of the slow filter in the current frame, W2 上 Indicates the coefficients of the slow filter of the previous frame.
[0133] It should be noted that the control factor of the preset fast filter is usually 0.9995, and the control factor of the preset slow filter is usually 1. The coefficient error covariance of the corresponding filter is controlled by the relevant terms of the coefficient error covariance equation of the corresponding filter, and then the Kalman gain of the corresponding filter is controlled; the fast filter used in this application is a Kalman filter with a certain time-varying tracking capability, and the slow filter used is a Kalman filter without tracking capability.
[0134] In another preferred embodiment, after generating the fast filter coefficients of the current frame and the slow filter coefficients of the current frame, the method further includes:
[0135] Calculating the sum of all frequency points of the fast filter noise energy smoothing estimate of the current frame to obtain a first frequency point value; calculating the sum of all frequency points of the slow filter noise energy smoothing estimate of the current frame to obtain a second frequency point value; calculating the ratio between the first frequency point value and the second frequency point value to obtain a total error energy ratio; generating a coefficient of variation based on the fast filter coefficients and slow filter coefficients of several historical frames;
[0136] The relationship between the total energy ratio of the comparison error, the coefficient of variation and the set threshold;
[0137] When the total error energy ratio is greater than the first threshold and less than the second threshold, the slow filter coefficient of the current frame is assigned to the fast filter coefficient of the current frame to generate an updated fast filter coefficient of the current frame.
[0138] When the total error energy ratio is less than or equal to a first threshold, assigning the slow filter coefficient of the current frame to the fast filter coefficient of the current frame to generate an updated fast filter coefficient of the current frame, and initializing the fast filter coefficient error covariance and the slow filter coefficient error covariance of the current frame;
[0139] When the total error energy ratio is greater than a third threshold and the coefficient of variation is less than a fourth threshold, assigning the fast filter coefficient of the current frame to the slow filter coefficient of the current frame to generate an updated slow filter coefficient of the current frame;
[0140] Specifically, the main function of the total error energy ratio is to detect the difference in error magnitude between the current fast filter and the slow filter, and to serve as an evaluation parameter for judging the convergence effect of the current filter; in the steady-state case, the errors of the two filters are similar, so the ratio is close to 1, when the slow filter is better, the ratio is less than 1, and when the fast filter is better, the ratio is greater than 1; based on the fast filter coefficients and slow filter coefficients of several historical frames, a coefficient of variation is generated; for example: by counting the smoothed estimate of the Euclidean distance of the fast filter coefficients and the smoothed estimate of the Euclidean distance of the slow filter coefficients in the past 200 frames, and calculating the smoothed estimate of the Euclidean distance of the fast filter coefficients and the Euclidean distance of the slow filter coefficients The filter ratio is generated by calculating the ratio between the smoothed estimates of the fast and slow filter coefficients. The coefficient of variation is the ratio between the standard deviation of the filter ratio and the average value of the filter ratio. Normally, the operation of calculating the coefficient of variation is performed every 2 seconds to reduce the computational complexity. The coefficient of variation reflects the fluctuation magnitude of the current fast filter coefficient and the slow filter coefficient. As an evaluation parameter for judging the stability of the filter coefficient, when the coefficient of variation is small, it indicates that the filter has entered a steady-state scenario. When the coefficient of variation is large, it indicates that the filter is not stable and may be in a time-varying scenario. According to the relationship between the total error energy ratio, the coefficient of variation and the set threshold, the following steps are performed:
[0141] When the total error energy ratio is greater than the first threshold and less than the second threshold, the slow filter coefficient of the current frame is assigned to the fast filter coefficient to generate an updated fast filter coefficient of the current frame, and the slow filter coefficient of the current frame remains unchanged; usually the value of the first threshold is 0.3, and the value of the second threshold is 0.95; the purpose of this assignment operation is that since subsequent operations need to be based on the fast filter coefficient to calculate the relevant correction coefficient, when the fast filter coefficient is worse than the slow filter coefficient, it will lead to poor correction coefficients calculated by the fast filter coefficient. Therefore, in order to ensure that the fast filter coefficient has a better output, when the total error energy ratio is greater than the first threshold and less than the second threshold, the slow filter coefficient of the current frame is assigned to the fast filter coefficient of the current frame;
[0142] When the total error energy ratio is less than or equal to the first threshold, the slow filter coefficient of the current frame is assigned to the fast filter coefficient of the current frame to generate an updated fast filter coefficient of the current frame, the slow filter coefficient of the current frame remains unchanged, and the fast filter coefficient error covariance and the slow filter coefficient error covariance of the current frame are initialized and reset to one; the purpose of this assignment operation is to reset the filter coefficient error covariance after the filter diverges, and the large initial filter coefficient error covariance can ensure that the filter starts to converge quickly again;
[0143] When the total error energy ratio is greater than the third threshold and the coefficient of variation is less than the fourth threshold, the fast filter coefficient of the current frame is assigned to the slow filter coefficient of the current frame to generate an updated slow filter coefficient of the current frame, and the fast filter coefficient of the current frame remains unchanged; under normal circumstances, the value of the third threshold is 1.02, and the value of the fourth threshold is 0.0035; since the slow filter has almost no tracking ability, once a scenario occurs where the steady-state environmental coefficient changes significantly, the slow filter at this time will no longer have a reference function and may cause signal cancellation problems. Therefore, in order to solve the problem that the slow filter does not have a reference function and signal cancellation problems, this assignment operation corrects the slow filter coefficient in real time by determining the stable state of the fast filter to ensure that the slow filter coefficient has a good reference, and at the same time, the corrected slow filter coefficient will not cancel the signal.
[0144] It should be noted that the usual values of the above thresholds are reference data obtained through hundreds of parameter configurations and regular summary experiences of simulation data sets and actual data sets.
[0145] Regarding step S2.3, in a preferred embodiment, the post-filter coefficient of the fast filter in the current frame is calculated based on the preset attenuation factor, the Kalman gain of the fast filter in the current frame, and the input signal of the current frame, including:
[0146] The post-filter coefficient of the fast filter in the current frame is calculated using the following formula:
[0147]
[0148] H = max(H, 0)
[0149] Among them, H represents the post-filter coefficient of the fast filter of the current frame, Q represents the preset attenuation factor, K1 represents the Kalman gain of the fast filter of the current frame, X s Represents the input signal of the current frame under the sth block, K1 s represents the fast filter Kalman gain of the current frame under the sth block;
[0150] It should be noted that the post-filter coefficient of the fast filter of the current frame is jointly determined by the Kalman gain of the fast filter of the current frame and the input signal of the current frame, and the value range is 0 to 1. The post-filter coefficient of the fast filter of the current frame plays the role of echo elimination based on the Wiener filtering principle; usually, the value range of the post-filter coefficient of the fast filter of the current frame is 1 to 2. Within this value range, the degree of echo suppression can be further controlled. This value range is a reference data obtained through hundreds of parameter adjustment configurations and regular summary experiences of simulation data sets and actual data sets.
[0151] Regarding step S2.4, in a preferred embodiment, calculating the correction coefficient of the first error signal based on the fast filter noise energy smoothing estimate in the current frame, the slow filter noise energy smoothing estimate in the current frame, the smoothing speed, and a preset error ratio threshold includes:
[0152] Grouping the frequency points in the fast filter noise energy smoothing estimate for the current frame into pairs to generate a plurality of first frequency point groups; wherein each first frequency point group consists of an odd-numbered frequency point and an adjacent even-numbered frequency point; assigning an average value of each first frequency point group to the odd-numbered frequency points and the even-numbered frequency points in the corresponding first frequency point group to generate a second frequency point group; and generating a second fast filter noise energy smoothing estimate for the current frame based on each second frequency point group;
[0153] Grouping the frequency points in the slow filter noise energy smoothing estimate of the current frame into pairs to generate third frequency point groups; wherein each third frequency point group consists of an odd-numbered frequency point and an adjacent even-numbered frequency point; assigning the average value of each third frequency point group to the odd-numbered frequency points and the even-numbered frequency points in the corresponding third frequency point group to generate a fourth frequency point group; generating a second slow filter noise energy smoothing estimate of the current frame based on each fourth frequency point group;
[0154] Calculating a ratio between a noise energy smoothing estimate of the second fast filter and a noise energy smoothing estimate of the second slow filter of the current frame to obtain a noise energy ratio of the current frame;
[0155] Determine the corresponding smoothing speed in the sigmoid function according to the noise energy ratio of the current frame;
[0156] The correction coefficient of the first error signal is calculated by the following formula:
[0157]
[0158] Wherein, H1 represents the correction coefficient of the first error signal, B represents the smoothing speed, dt_P0 represents the preset error ratio threshold, and VV21 represents the noise energy ratio;
[0159] Specifically, such as Figure 3 As shown, the fast filter noise energy smoothing estimate and the slow filter noise energy smoothing estimate of the current frame are respectively calculated as follows Figure 2 The processing is performed in the manner shown; the following is an example of the processing of the fast filter noise energy smoothing estimation of the current frame; each frequency point in the fast filter noise energy smoothing estimation of the current frame is grouped into two groups to generate multiple first frequency point groups, wherein each first frequency point group consists of an odd-numbered frequency point and an even-numbered frequency point adjacent to it; for example: Figure 3 As shown, the fast filter noise energy smoothing estimate of the current frame includes the first frequency point, the second frequency point, the third frequency point and the fourth frequency point. The first frequency point and the second frequency point are divided into one group, and the third frequency point and the fourth frequency point are divided into one group. The average value of the corresponding frequency point values of the two groups of frequency points is obtained respectively, and the obtained average value is assigned to the odd-numbered frequency points and the even-numbered frequency points in the corresponding group to generate the noise energy smoothing estimate of the second fast filter. Similarly, the noise energy smoothing estimate of the second slow filter can be obtained. The ratio between the noise energy of the second fast filter and the noise energy of the second slow filter is calculated, and the ratio is substituted into Figure 4 The smoothing mapping is performed in the sigmoid function shown in the figure to obtain the corresponding sigmoid function smoothing speed, and the correction coefficient of the first error signal is calculated by the following formula:
[0160]
[0161] Wherein, H1 represents the correction coefficient of the first error signal, B represents the smoothing speed of the sigmoid function corresponding to the noise energy ratio, dt_P0 represents the preset error ratio threshold, and VV21 represents the noise energy ratio;
[0162] It should be noted that under normal circumstances, the preset error ratio threshold is 1.5. This value is a reference data obtained through hundreds of parameter adjustment configurations and regular summary experiences of simulation data sets and actual data sets.
[0163] With respect to step S2.5, in a preferred embodiment, echo suppression is performed on the error signal of the fast filter using the correction coefficient of the first error signal and the post-filter coefficient to generate a first output signal corresponding to the input signal of the current frame, including: calculating the product of the correction coefficient of the first error signal and the post-filter coefficient to obtain the correction coefficient of the second error signal; calculating the product of the correction coefficients of the first error signal and the second error signal in the frequency domain to generate the first output signal corresponding to the input signal of the current frame;
[0164] Specifically, before performing echo suppression, it is necessary to perform a windowing operation, such as a Hanning window, on the first error signal in the frequency domain. The windowed first error signal is more stable. Most of the residual echoes in the first output signal corresponding to the input signal of the current frame generated by the method in this preferred embodiment have been eliminated.
[0165] In a preferred embodiment, before generating a final output signal based on the first output signal of each frame, the method further includes: performing residual echo detection on the first output signal of each frame, and performing secondary echo suppression on the first output signal containing residual echoes; wherein the secondary echo suppression includes: calculating the sum of the absolute values of each frequency point of the low-frequency portion of the first output signal of the current frame to obtain a third frequency point value; calculating the sum of the absolute values of each frequency point of the high-frequency portion of the first output signal of the current frame to obtain a fourth frequency point value; calculating the ratio between the third frequency point value and the fourth frequency point value to obtain a first frequency point ratio; when the first frequency point ratio is less than a fifth threshold, setting the correction coefficient of the second error signal to zero to generate the correction coefficient of the third error signal; and calculating the product of the correction coefficients of the first error signal and the third error signal in the frequency domain to generate the corresponding updated first output signal.
[0166] Specifically, residual echo detection is performed on the mid-high frequency parts of the first output signal of each frame. When it is detected that there is no residual echo in the mid-high frequency parts of the first output signal of the current frame, the corresponding first output signal of the current frame is output; when it is monitored that there is a residual echo in the mid-high frequency parts of the first output signal of the current frame, secondary echo suppression is performed on the first output signal of the current frame; secondary echo suppression utilizes the diffraction principle of waves. The wavelength of low-frequency sound waves is longer than that of mid-high frequency waves and has stronger penetrating power, so it is less susceptible to changes in the echo path caused by obstacles such as physical movement; according to the above-mentioned diffraction principle of waves, the absolute values of the frequency points of the low-frequency part of the first output signal are summed to obtain a third frequency point value, and the absolute values of the frequency points of the mid-high frequency part of the first output signal are summed to obtain a fourth frequency point value. value; it should be noted that the low-frequency part usually refers to 250Hz~850Hz, and the medium and high-frequency part usually refers to 2000Hz~4000Hz; the third frequency point value is divided by the fourth frequency point value to obtain the first frequency point ratio; when the first frequency point ratio is less than the fifth threshold value, it can be confirmed that the current frame is a time-varying frame. At this time, the correction coefficient of the second error signal is set to zero to generate the correction coefficient of the third error signal; under normal circumstances, the value of the fifth threshold is 0.001, which is a reference data obtained through hundreds of parameter adjustment configurations and regular summary experiences of simulation data sets and actual data sets; a windowing operation is performed on the first error signal in the frequency domain state, and then the product between the correction coefficients of the first error signal and the third error signal in the frequency domain state is calculated to generate the corresponding updated first output signal.
[0167] In step S3, a final output signal is generated according to the first output signal of each frame processed in step S2.
[0168] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A dual-filter-based echo suppression method, characterized in that: include: Acquire an input signal; wherein the input signal includes: a signal to be extracted and an interference signal; Performing echo suppression processing on each frame of input signal to generate a first output signal corresponding to each frame of input signal; generating a final output signal according to the first output signal of each frame; The echo suppression process includes: Calculating a first error signal of the input signal of the current frame in the fast filter according to the input signal of the current frame and the fast filter coefficient of the previous frame; calculating a second error signal of the input signal of the current frame in the slow filter according to the input signal of the current frame and the slow filter coefficient of the previous frame; Calculating a fast filter noise energy smoothing estimate for the current frame based on the first error signal and a preset fast filter noise estimation smoothing factor; calculating a slow filter noise energy smoothing estimate for the current frame based on the second error signal and a preset slow filter noise estimation smoothing factor; Calculate the post-filter coefficient of the fast filter in the current frame according to the preset attenuation factor, the fast filter Kalman gain of the current frame and the input signal of the current frame; Calculating a correction coefficient of the first error signal based on a fast filter noise energy smoothing estimate in the current frame, a slow filter noise energy smoothing estimate in the current frame, a smoothing speed, and a preset error ratio threshold; The error signal of the fast filter is subjected to echo suppression by using the correction coefficient of the first error signal and the post-filtering coefficient to generate a first output signal corresponding to the input signal of the current frame.
2. The dual-filter echo suppression method according to claim 1, wherein: The first error signal of the input signal of the current frame in the fast filter is calculated according to the input signal of the current frame and the fast filter coefficient of the previous frame; Calculating a second error signal of the input signal of the current frame in the slow filter according to the input signal of the current frame and the slow filter coefficient of the previous frame, including: The first error signal is calculated by the following formula: in, represents the first error signal, Indicates the signal to be extracted in the current frame, represents the total number of fast filter blocks, represents the sth block of the fast filter, Represents the fast filter coefficient of the sth block in the previous frame; represents the interference signal of the current frame under the sth block; Indicates that the current result only retains the last L elements; The second error signal is calculated by the following formula: in, represents the second error signal, represents the total number of slow filter blocks, represents the sth block of the slow filter, , Indicates that the current result only retains the last L elements, where L is the moving block length of the current block number.
3. The dual-filter echo suppression method according to claim 2, wherein: The step of calculating a fast filter noise energy smoothing estimate for a current frame based on the first error signal and a preset fast filter noise estimation smoothing factor; and calculating a slow filter noise energy smoothing estimate for a current frame based on the second error signal and a preset slow filter noise estimation smoothing factor, comprising: The fast filter noise energy smoothing estimate for the current frame is calculated using the following formula: in, represents the fast filter noise energy smoothing estimate in the current frame, represents the preset fast filter noise estimation smoothing factor, represents a first error signal in a frequency domain state; Indicates adding L zero elements in front of the current result, where L represents the moving block length of the current block number; The slow filter noise energy smoothing estimate for the current frame is calculated using the following formula: in, represents the slow filter noise energy smoothing estimate in the current frame, represents the preset slow filter noise estimation smoothing factor, A second error signal representing a frequency domain state; It means adding L zero elements in front of the current result, where L is the moving block length of the current block number.
4. The dual-filter echo suppression method according to claim 3, wherein: The generation of fast filter coefficients for the current frame includes: Generate a smoothed estimate of the interference signal of the current frame according to a preset interference energy smoothing factor and the interference signal of the current frame; When the smoothed estimate of the interference signal of the current frame is greater than the preset interference energy threshold, the fast filter coefficient error covariance of the current frame is calculated based on the fast filter coefficient of the previous frame, the fast filter coefficient error covariance of the previous frame, the fast filter Kalman gain of the current frame and the preset fast filter control factor; the fast filter Kalman gain of the current frame is calculated based on the smoothed estimate of the fast filter noise energy of the current frame, the fast filter coefficient of the previous frame, the fast filter coefficient error covariance of the previous frame and the interference signal; the fast filter coefficient of the current frame is generated based on the fast filter coefficient of the previous frame and the fast filter Kalman gain of the current frame.
5. The dual-filter echo suppression method according to claim 4, wherein: The generation of the slow filter coefficients of the current frame includes: When the smoothed estimate of the interference signal of the current frame is greater than the preset interference energy threshold, the slow filter coefficient error covariance of the current frame is calculated based on the slow filter coefficient of the previous frame, the error covariance of the slow filter coefficient of the previous frame, the slow filter Kalman gain of the current frame and the preset slow filter control factor; the slow filter Kalman gain of the current frame is calculated based on the smoothed estimate of the slow filter noise energy of the current frame, the slow filter coefficient of the previous frame, the error covariance of the slow filter coefficient of the previous frame and the interference signal; the slow filter coefficient of the current frame is generated based on the slow filter coefficient of the previous frame and the slow filter Kalman gain of the current frame.
6. The dual-filter-based echo suppression method according to any one of claims 3 or 4, characterized in that: After generating the fast filter coefficients of the current frame and the slow filter coefficients of the current frame, it also includes: Calculating the sum of all frequency points of the fast filter noise energy smoothing estimate of the current frame to obtain a first frequency point value; calculating the sum of all frequency points of the slow filter noise energy smoothing estimate of the current frame to obtain a second frequency point value; calculating the ratio between the first frequency point value and the second frequency point value to obtain a total error energy ratio; generating a coefficient of variation based on the fast filter coefficients and slow filter coefficients of several historical frames; The relationship between the total energy ratio of the comparison error, the coefficient of variation and the set threshold; When the total error energy ratio is greater than the first threshold and less than the second threshold, the slow filter coefficient of the current frame is assigned to the fast filter coefficient of the current frame to generate an updated fast filter coefficient of the current frame. When the total error energy ratio is less than or equal to a first threshold, assigning the slow filter coefficient of the current frame to the fast filter coefficient of the current frame to generate an updated fast filter coefficient of the current frame, and initializing the fast filter coefficient error covariance and the slow filter coefficient error covariance of the current frame; When the total error energy ratio is greater than the third threshold and the coefficient of variation is less than the fourth threshold, the fast filter coefficient of the current frame is assigned to the slow filter coefficient of the current frame to generate an updated slow filter coefficient of the current frame.
7. The dual-filter-based echo suppression method according to claim 6, wherein: The step of calculating the post-filter coefficient of the fast filter in the current frame according to the preset attenuation factor, the Kalman gain of the fast filter in the current frame, and the input signal of the current frame includes: The post-filter coefficient of the fast filter in the current frame is calculated using the following formula: in, Represents the post-filter coefficient of the fast filter of the current frame, Indicates the preset attenuation factor, represents the fast filter Kalman gain of the current frame, represents the input signal of the current frame under the sth block, Represents the fast filter Kalman gain of the current frame under the s-th block.
8. The dual-filter-based echo suppression method according to claim 3, wherein: Calculating the correction coefficient of the first error signal according to the fast filter noise energy smoothing estimate in the current frame, the slow filter noise energy smoothing estimate in the current frame, the smoothing speed, and a preset error ratio threshold includes: Grouping the frequency points in the fast filter noise energy smoothing estimate for the current frame into pairs to generate a plurality of first frequency point groups; wherein each first frequency point group consists of an odd-numbered frequency point and an adjacent even-numbered frequency point; assigning an average value of each first frequency point group to the odd-numbered frequency points and the even-numbered frequency points in the corresponding first frequency point group to generate a second frequency point group; and generating a second fast filter noise energy smoothing estimate for the current frame based on each second frequency point group; Grouping the frequency points in the slow filter noise energy smoothing estimate of the current frame into pairs to generate third frequency point groups; wherein each third frequency point group consists of an odd-numbered frequency point and an adjacent even-numbered frequency point; assigning the average value of each third frequency point group to the odd-numbered frequency points and the even-numbered frequency points in the corresponding third frequency point group to generate a fourth frequency point group; generating a second slow filter noise energy smoothing estimate of the current frame based on each fourth frequency point group; Calculating a ratio between a noise energy smoothing estimate of the second fast filter and a noise energy smoothing estimate of the second slow filter of the current frame to obtain a noise energy ratio of the current frame; Determine the corresponding smoothing speed in the sigmoid function according to the noise energy ratio of the current frame; The correction coefficient of the first error signal is calculated by the following formula: in, represents the correction coefficient of the first error signal, Indicates the smoothing speed, Indicates the preset error ratio threshold, Represents the noise energy ratio.
9. The dual-filter-based echo suppression method according to claim 8, wherein: The method of performing echo suppression on the error signal of the fast filter by using the correction coefficient of the first error signal and the post-filter coefficient to generate a first output signal corresponding to the input signal of the current frame includes: Calculating the product of the correction coefficient of the first error signal and the post-filtering coefficient to obtain the correction coefficient of the second error signal; The product of the correction coefficients of the first error signal and the second error signal in the frequency domain is calculated to generate a first output signal corresponding to the input signal of the current frame.
10. The dual-filter-based echo suppression method according to claim 9, wherein: Before generating the final output signal according to the first output signal of each frame, the method further includes: Performing residual echo detection on the first output signal of each frame, and performing secondary echo suppression on the first output signal with residual echo; Wherein, the secondary echo suppression includes: Calculate the sum of the absolute values of the frequency points of the low-frequency part of the first output signal of the current frame to obtain a third frequency point value; Calculate the sum of the absolute values of the frequency points of the high frequency part of the first output signal of the current frame to obtain a fourth frequency point value; Calculating a ratio between the third frequency point value and the fourth frequency point value to obtain a first frequency point ratio; When the first frequency point ratio is less than a fifth threshold, setting the correction coefficient of the second error signal to zero to generate a correction coefficient of the third error signal; The product of the first error signal and the correction coefficient of the third error signal in the frequency domain is calculated to generate a corresponding updated first output signal.
Citation Information
Patent Citations
Systems and methods for echo cancellation and echo suppression
CN102164210A
Method for optimizing convergence characteristic of acoustic echo canceller
CN106533500A