An echo cancellation system and method
By processing near-end and far-end signals through frequency domain sequences and dual-talk detection modules, and updating the filter coefficients under different states using two-stage filters, the problems of filter coefficient divergence and slow convergence speed in echo cancellation systems are solved, achieving more effective echo cancellation.
Patent Information
- Application Number
- CN202411444985.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2044-10-16
AI Technical Summary
In existing technologies, the filter coefficients of echo cancellation systems tend to diverge under conditions of two-way speech and echo nonlinear distortion, and the convergence speed is slow.
The frequency domain sequence module is used to process near-end and far-end signals. The coherence coefficient is obtained through the dual-talk detection module. Different thresholds are selected for the first and second filters to perform dual-talk detection. The filter coefficients are updated under different states using two-stage linear filters to eliminate echo.
It effectively reduces the steady-state misalignment of the filter under dual-talk mode and echo nonlinear distortion, and improves the echo cancellation effect and convergence speed.
Smart Images

Figure CN119229839B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information communication technology, in particular to an echo cancellation system and method. BACKGROUND
[0002] Generally, mobile terminals such as mobile phones will produce echo when talking in speaker mode. An echo canceller generally uses an adaptive linear filter to cancel the echo. Common adaptive filter coefficient update algorithms include least mean square (NLMS), recursive least squares (RLS), kalman filter, etc. In echo cancellation, the problem of double talk often occurs. When double talk occurs, the filter coefficients continue to update and diverge. Another problem is nonlinear distortion. When the nonlinear distortion of the echo is large, the filter is also prone to divergence. The conventional method is to use a double talk detection (DTD) module to detect single talk and double talk. When double talk is detected, the filter coefficients stop updating.
[0003] The Speex scheme uses double adaptive filters. One of the filters is a main filter that uses a variable step size adaptive algorithm to update the coefficients, and the other is a backup filter. The topology is shown in Figure 4 This scheme can avoid filter divergence to some extent when double talk occurs, but the steady-state distortion significantly increases when the nonlinear distortion of the echo is large, and the convergence speed is slow when the echo path changes.
[0004] Therefore, it is an urgent problem for those skilled in the art to provide an echo cancellation system and method that effectively solve the above problems. SUMMARY
[0005] The purpose of the present application is to provide an echo cancellation system that is simple in structure, safe, effective, reliable, easy to operate, and can reduce the steady-state distortion of the filter coefficients in the case of double talk and strong nonlinear distortion of the echo in echo cancellation, and ensure a faster convergence speed when the echo path changes.
[0006] To achieve the above purpose, the technical scheme provided by the present application is as follows:
[0007] An echo cancellation system, comprising:
[0008] A frequency domain sequence module for processing a near-end signal and a far-end signal respectively to obtain a near-end frequency domain sequence and a far-end frequency domain sequence;
[0009] A double talk detection module for obtaining a coherence coefficient according to a preset reference signal, the near-end frequency domain sequence and the far-end frequency domain sequence, and a double talk detection algorithm.
[0010] The double-talk detection module is further configured to select different thresholds for a first filter and a second filter in a filtering module according to the coherence coefficient to perform double-talk detection, so as to obtain a first double-talk detection result and a second double-talk detection result.
[0011] The filtering module comprises the first filter and the second filter, the first filter is configured to update a first filter coefficient according to the first double-talk detection result in a filtering process, and output a first filtering result to the second filter, and the second filter is configured to update a second filter coefficient according to the first filtering result and the second double-talk detection result in the filtering process, so as to obtain a two-stage linear filtering result.
[0012] The filtering module is further configured to, based on the two-stage linear filtering result, converge the second filter in a first state to eliminate echo, and adjust a learning rate of the first filter according to the second filter coefficient in a second state to converge the first filter to eliminate echo.
[0013] An echo cancellation method is implemented based on the echo cancellation system as described above, and comprises the following steps:
[0014] The near-end signal and the far-end signal are processed respectively to obtain a near-end frequency domain sequence and a far-end frequency domain sequence.
[0015] A coherence coefficient is obtained according to a preset reference signal, the near-end frequency domain sequence, the far-end frequency domain sequence and a double-talk detection algorithm.
[0016] Different thresholds are selected for a first filter and a second filter according to the coherence coefficient to perform double-talk detection, so as to obtain a first double-talk detection result and a second double-talk detection result.
[0017] In a filtering process, a first filter coefficient is updated according to the first double-talk detection result, and a first filtering result is output to the second filter, and the second filter is configured to update a second filter coefficient according to the first filtering result and the second double-talk detection result in the filtering process, so as to obtain a two-stage linear filtering result.
[0018] Based on the two-stage linear filtering result, the second filter is converged in a first state to eliminate echo, and a learning rate of the first filter is adjusted according to the second filter coefficient in a second state to converge the first filter to eliminate echo.
[0019] Preferably, the near-end signal and the far-end signal are processed respectively to obtain a near-end frequency domain sequence and a far-end frequency domain sequence, and specifically:
[0020] The near-end signal and the far-end signal are obtained.
[0021] frame, overlap-add window and M-point discrete Fourier transform are sequentially performed on the near-end signal and the far-end signal respectively to obtain the near-end frequency domain sequence and the far-end frequency domain sequence.
[0022] Preferably, the coherence coefficient is obtained according to a preset reference signal, the near-end frequency domain sequence and the far-end frequency domain sequence and a double-talk detection algorithm, and specifically:
[0023] ;
[0024] wherein, and are the near-end frequency domain sequence and the far-end frequency domain sequence respectively, is a conjugate operator, is a near-end signal and far-end signal frequency domain cross-correlation value, is a near-end signal and far-end signal cross-power spectrum, is a near-end signal self-power spectrum, is a far-end signal self-power spectrum, is a coherence coefficient of each frequency point.
[0025] Preferably, different threshold values are selected for the first filter and the second filter according to the coherence coefficient for double-talk detection to obtain a first double-talk detection result and a second double-talk detection result, and specifically:
[0026] ,
[0027] ,
[0028] wherein, and are the first double-talk detection result and the second double-talk detection result respectively, and are different threshold values, and 1 is single-talk and 0 is double-talk.
[0029] Preferably, the first filter filtering process is specifically:
[0030] ,
[0031] ,
[0032] ;
[0033] wherein, is a reference signal frequency domain sequence, is a first filter coefficient, is an order of the first filter, is a first filtering result.
[0034] Preferably, in the filtering process, the first filtering coefficient is updated according to the first double-talk detection result, specifically:
[0035] ,
[0036] ,
[0037] + ,
[0038] wherein, is a reference signal autocorrelation power spectrum, is a sliding average coefficient, is a coefficient update step, is a learning rate parameter calculated in a previous frame.
[0039] Preferably, the second filter filtering process and the second filtering coefficient updating process are specifically:
[0040] ,
[0041] ,
[0042] ,
[0043] ,
[0044] ,
[0045] ,
[0046] ,
[0047] ,
[0048] ,
[0049] ,
[0050] ,
[0051] = ;
[0052] wherein, is a filtered estimate of the echo, is a filter order, is a preset maximum smoothing coefficient, is a sliding average coefficient, is a second filter output, is a cross power spectrum of the first filter output and the far-end signal, is a self power spectrum of the far-end signal, is a linear regression coefficient between the first filter output and the far-end signal, i.e., a second filter coefficient.
[0053] Preferably, the learning rate of the first filter is adjusted according to the second filter coefficient, in particular:
[0054] a coefficient vector of the second filter is defined , is a conjugate transpose parameter, and a first filter learning rate parameter is obtained as follows:
[0055] ,
[0056] wherein, is a constant less than 1, is a coefficient of the first filter.
[0057] The echo cancellation system provided by the application processes the obtained near-end signal and far-end signal through a frequency domain sequence module to obtain a near-end frequency domain sequence and a far-end frequency domain sequence; a reference signal is preset through a double-talk detection module, and the near-end frequency domain sequence and the far-end frequency domain sequence are combined with a double-talk detection algorithm to obtain a coherence coefficient, and different threshold values are selected for a first filter and a second filter for double-talk detection to obtain first and second double-talk detection results; the corresponding filter parameters are updated in the filtering process of the two filters through a filtering module, and two-stage linear filtering results are output; based on the two-stage linear filtering results, in a first state, the second filter is converged, and in a second state, the learning rate of the first filter is adjusted, the first filter is converged, and the echo is eliminated. Compared with the prior art, the application obtains the double-talk detection results of the filters, updates the filter coefficients, outputs the filtering results, and converges different filters in different states, which can effectively solve the steady-state misadjustment and improve the echo cancellation effect.
[0058] The application further provides an echo cancellation method, which solves the same technical problem as the system, belongs to the same inventive concept, and should have the same beneficial effects, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description only constitute some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0060] Figure 1 A structural schematic diagram of an echo cancellation system provided by the embodiment of the present application is provided.
[0061] Figure 2 A flowchart of an echo cancellation method provided by the embodiment of the present application is provided.
[0062] Figure 3 A flowchart of step S1 provided by the embodiment of the present application is provided.
[0063] Figure 4 A topological structure diagram of an existing speex provided by the embodiment of the present application is provided. DETAILED DESCRIPTION
[0064] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort fall within the protection scope of the present application.
[0065] The embodiments of the present application are written in a progressive manner.
[0066] The embodiments of the present application provide an echo cancellation system and method. The technical problem of the steady-state misadjustment of filter coefficients in the two cases of strong nonlinear distortion of double-talk state and echo in echo cancellation in the prior art is solved.
[0067] An echo cancellation system comprises:
[0068] A frequency domain sequence module is configured to process a near-end signal and a far-end signal respectively to obtain a near-end frequency domain sequence and a far-end frequency domain sequence.
[0069] A double-talk detection module is configured to obtain a coherence coefficient according to a preset reference signal, the near-end frequency domain sequence and the far-end frequency domain sequence, and a double-talk detection algorithm.
[0070] The double-talk detection module is also configured to select different thresholds for the first filter and the second filter in the filtering module according to the coherence coefficient to perform double-talk detection, so as to obtain a first double-talk detection result and a second double-talk detection result.
[0071] The filtering module comprises a first filter and a second filter, the first filter is configured to update a first filter coefficient according to the first double-talk detection result in a filtering process, and output a first filter result to the second filter; and the second filter is configured to update a second filter coefficient according to the first filter result and the second double-talk detection result in the filtering process, so as to obtain a two-stage linear filter result.
[0072] The filtering module is further configured to, based on the two-stage linear filter result, converge the second filter in a first state to eliminate echo, and adjust a learning rate of the first filter according to the second filter coefficient in a second state to converge the first filter to eliminate echo.
[0073] The echo cancellation system provided by the application obtains a near-end frequency domain sequence and a far-end frequency domain sequence by processing the obtained near-end signal and far-end signal through the frequency domain sequence module; obtains a coherence coefficient by presetting a reference signal, combining the near-end frequency domain sequence and the far-end frequency domain sequence, and using a double-talk detection algorithm, and selecting different thresholds for the first filter and the second filter to perform double-talk detection, so as to obtain a first double-talk detection result and a second double-talk detection result; updates corresponding filter parameters in the filtering process of the two filters through the filtering module, and outputs a two-stage linear filter result; based on the two-stage linear filter result, converges the second filter in a first state, and adjusts the learning rate of the first filter in a second state to converge the first filter, so as to eliminate echo.
[0074] In the embodiment, the first filter specifically selects an NLMS filter, and the second filter specifically selects an SLR filter.
[0075] The LMS algorithm is a linear adaptive filtering algorithm, which does not need to calculate a correlation function or a matrix inversion. Due to its simplicity (small amount of calculation), the LMS algorithm becomes a reference standard for other linear adaptive filtering algorithms. The LMS generally comprises two basic processes: a filtering process: a. calculating a response of a linear filter output to an input signal; and b. generating an estimated error by comparing an output result with an expected response; and an adaptive process: automatically adjusting filter parameters according to the estimated error.
[0076] NLMS is the normalized LMS (Normalized LMS) algorithm, which is an improved version of the LMS algorithm. In the LMS algorithm, the adjustment amount of the tap weight is proportional to the tap input vector. When the tap input vector is relatively large, the LMS algorithm will encounter the problem of gradient noise amplification. In order to overcome this problem, the NLMS algorithm can be used. The reason why it is called "Normalized" is that it uses the squared Euclidean norm of the tap input vector to normalize the tap weight adjustment amount.
[0077] The SLR filter is a filter based on the linear regression algorithm structure. The Stage Linear Regression (SLR) algorithm is a method that can achieve fast convergence of the filter. Considering that the energy of the echo is concentrated in a short continuous period of time, the filter in the frequency domain can use fewer orders, and the filter coefficients are iteratively solved by using the continuous reference signal through the stage linear regression method.
[0078] An echo cancellation method is implemented based on the echo cancellation system as described above, comprising the following steps:
[0079] S1. Process the near-end signal and the far-end signal respectively to obtain a near-end frequency domain sequence and a far-end frequency domain sequence;
[0080] S2. Obtain a coherence coefficient according to a preset reference signal, the near-end frequency domain sequence, the far-end frequency domain sequence, and a double-talk detection algorithm;
[0081] S3. Select different threshold values for the first filter and the second filter according to the coherence coefficient for double-talk detection to obtain a first double-talk detection result and a second double-talk detection result;
[0082] S4. In the filtering process, update the first filter coefficient according to the first double-talk detection result, and output the first filtering result to the second filter; the second filter is used to update the second filter coefficient according to the first filtering result and the second double-talk detection result in the filtering process, to obtain a two-stage linear filtering result;
[0083] S5. Based on the two-stage linear filtering result, in the first state, converge the second filter to eliminate the echo; in the second state, adjust the learning rate of the first filter according to the second filter coefficient, and converge the first filter to eliminate the echo.
[0084] In steps S1 to S3, the near-end frequency domain sequence and the far-end frequency domain sequence are obtained by processing the near-end signal and the far-end signal respectively, and the coherence coefficient of each frequency point is calculated and obtained by combining the double-talk detection (DTD) algorithm. According to the coherence coefficient of each frequency point, different threshold values are selected for double-talk detection of the NLMS filter and the SLR filter to obtain the corresponding double-talk detection results.
[0085] In steps S4 to S5, the coefficients of the NLMS filter are updated according to the first double-talk detection result, and the filtering result of the NLMS filter is output to the SLR filter. In the filtering process, the SLR filter updates the coefficients of the SLR filter according to the filtering result of the NLMS filter and the second double-talk detection result, to obtain a two-stage linear filtering result. When the NLMS convergence efficiency is poor, the residual echo can be eliminated by fast convergence of the SLR filter based on the two-stage linear filtering result. When the SLR filter converges well but the NLMS does not converge, the residual echo can be eliminated by adjusting the learning rate of the first filter to make the NLMS filter converge quickly.
[0086] Preferably, S1, specifically:
[0087] A1. Obtain a near-end signal and a far-end signal;
[0088] A2. Frame, overlap and window, and perform M-point discrete Fourier transform on the near-end signal and the far-end signal, respectively, to obtain a near-end frequency domain sequence and a far-end frequency domain sequence.
[0089] In steps A1 to A2, the near-end signal and the far-end signal are collected by a conventional communication system, and the near-end signal and the far-end signal are framed, overlapped and windowed, and subjected to M-point discrete Fourier transform to obtain corresponding near-end frequency domain sequences and far-end frequency domain sequences.
[0090] In this embodiment, before frequency domain analysis, the signal is preprocessed as in the above steps. In the framing process, there is a certain overlap between adjacent two frames. Because the speech signal is time-varying, the characteristics change little in a short time range, so it can be analyzed as a steady state. However, if it exceeds this short time range, the signal will change, and the corresponding pitch of the two adjacent frames may change. If it is exactly in the middle of two syllables, or exactly in the transition from initial to final, etc., the characteristic parameters of the two frames may change greatly. However, in order to make the characteristic parameter change more smoothly, some frames are inserted between two non-overlapping frames to extract the characteristic parameters, forming an overlapping part between adjacent frames.
[0091] Preferably, the coherence coefficient is obtained according to a preset reference signal, the near-end frequency domain sequence and the far-end frequency domain sequence, and a double-talk detection algorithm, specifically:
[0092] ;
[0093] wherein, and are the near-end frequency domain sequence and the far-end frequency domain sequence, respectively, is a conjugate operator, a cross-power spectrum of the near-end signal and the far-end signal, a cross-power spectrum of the near-end signal and the far-end signal, a self-power spectrum of the near-end signal, a self-power spectrum of the far-end signal, a coherence coefficient of each frequency point.
[0094] In actual application, the near-end frequency domain sequence and the far-end frequency domain sequence are input into the above-mentioned double-talk detection algorithm formula, and the coherence coefficient of each frequency point is obtained in combination with a preset reference signal.
[0095] Preferably, different threshold values are selected for the first filter and the second filter according to the coherence coefficient for double-talk detection to obtain a first double-talk detection result and a second double-talk detection result, specifically as follows:
[0096]
[0097]
[0098] wherein, and are the first double-talk detection result and the second double-talk detection result respectively, and are different threshold values, and 1 is defined as single talk and 0 is defined as double talk.
[0099] In actual application, after the coherence coefficient is determined, different threshold values are selected for the NLMS filter and the SLR filter respectively, the threshold value of the NLMS filter is greater than that of the SLR filter, when the double-talk detection result is 1, it is defined as single talk, and when the result is 0, it is defined as double talk.
[0100] Preferably, the filtering process of the first filter is specifically as follows:
[0101]
[0102]
[0103]
[0104] wherein, is a reference signal frequency domain sequence, is a first filter coefficient, is an order of the first filter, is a first filtering result.
[0105] Preferably, in the filtering process, the first filtering coefficient is updated according to the first double-talk detection result, specifically as follows:
[0106] ,
[0107] ,
[0108] + ,
[0109] wherein, is a reference signal autocorrelation power spectrum, is a sliding average coefficient, is a coefficient update step size, is a learning rate parameter calculated in a previous frame.
[0110] In actual use, in the filtering process of the NLMS filter, the filtering coefficients of the NLMS filter are updated according to the first double-talk detection result, and after the updating is completed, the first filtering result is output to the SLR filter;
[0111] In this embodiment, the sliding average coefficient is in a range of 0.8 to 0.99.
[0112] Preferably, the second filter filtering process and the second filter coefficient updating process are specifically as follows:
[0113]
[0114] ,
[0115] ,
[0116] ,
[0117] ,
[0118] ,
[0119] ,
[0120] ,
[0121] ,
[0122] ,
[0123] ,
[0124] = ;
[0125] wherein, is the filtered estimate of the echo, is the filter order, is a preset maximum smoothing coefficient, is a sliding average coefficient, is the second filter output, is the cross power spectrum of the first filter output and the far-end signal, is the auto power spectrum of the far-end signal, is the linear regression coefficient between the first filter output and the far-end signal, i.e. the coefficient of the second filter.
[0126] In actual application, after the SLR filter receives the first filtering result , the SLR filter controls updating and filtering through double-talk detection, and this process is hierarchical, and the filter order is , the filtering and coefficient updating are performed according to the above formula, and finally the second filtering result is obtained, and the two-stage linear filtering result is output.
[0127] Preferably, in step S5, the learning rate of the first filter is adjusted according to the second filtering coefficient, specifically:
[0128] The coefficient vector of the second filter is defined as , is the conjugate transpose parameter, and the first filter learning rate parameter
[0129] ,
[0130] wherein, is a constant less than 1, is the coefficient of the first filter.
[0131] In actual application, when the SLR filter converges well and the NLMS does not converge, the learning rate of the NLMS has a relatively large value, so the convergence of the NLMS filtering can be accelerated.
[0132] In the embodiments of the present application, it should be understood that the disclosed method and system can be implemented in other manners. The above described system embodiments are merely schematic, for example, the division of the modules is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be indirect coupling or communication connection through some interfaces, devices or modules, and can be electrical, mechanical or other forms.
[0133] In addition, each of the functional modules in the embodiments of the present application can be integrated in one processor, or each of the modules can be a separate device, or two or more modules can be integrated in one device; each of the functional modules in the embodiments of the present application can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0134] Those skilled in the art can understand that all or part of the steps of the above method embodiments can be completed by program instructions and related hardware, and the above program instructions can be stored in a computer readable storage medium, and the program instructions are executed to perform the steps of the above method embodiments; and the above storage medium includes mobile storage devices, read only memory (ROM), magnetic discs or optical discs, and various storage media that can store program codes.
[0135] It should be understood that if "system", "device", "unit" and / or "module" are used in the present application, it is only a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the words can be replaced by other expressions.
[0136] As shown in the present application and claims, unless the context clearly indicates otherwise, "one", "a", "an" and / or "the" do not refer to the singular, but also include the plural. Generally, the terms "comprise" and "include" only indicate the inclusion of the steps and elements explicitly identified, and these steps and elements do not constitute an exclusive list, and the method or device can also include other steps or elements. The element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, product or device comprising the element.
[0137] In the description of the embodiments of the present application, "a plurality of" means two or more than two.
[0138] Hereinafter, the terms "first", "second", etc. are used only for the purpose of description, and are not to be construed as indicating or implying relative importance or a specific number of the technical features indicated. Thus, the features defined with "first", "second", etc. can explicitly or implicitly include one or more of the features.
[0139] If a flowchart is used in the present application, the flowchart is used to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or subsequent operations are not necessarily performed in sequence. Instead, each step can be processed in reverse order or simultaneously. Meanwhile, other operations can be added to these processes, or one or more steps of operations can be removed from these processes.
[0140] The above has been a detailed introduction to the echo cancellation system and method provided by the present application. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An echo cancellation system, characterized in that, include: The frequency domain sequence module is used to process near-end signals and far-end signals separately to obtain near-end frequency domain sequences and far-end frequency domain sequences; The dual-talk detection module is used to obtain the coherence coefficient based on the preset reference signal, the near-end frequency domain sequence, the far-end frequency domain sequence, and the dual-talk detection algorithm; The dual-talk detection module is also used to select different thresholds for the first filter and the second filter in the filtering module according to the coherence coefficient to perform dual-talk detection, so as to obtain the first dual-talk detection result and the second dual-talk detection result. The filtering module includes: a first filter and a second filter. The first filter is used to update the first filtering coefficients according to the first dual-talk detection result during the filtering process and output the first filtering result to the second filter. The second filter is used to update the second filtering coefficients according to the first filtering result and the second dual-talk detection result during the filtering process to obtain a two-stage linear filtering result. The filtering module is further configured to, based on the two-stage linear filtering results, converge the second filter in the first state to eliminate echo; and in the second state, adjust the learning rate of the first filter according to the second filtering coefficients to converge the first filter to eliminate echo.
2. An echo cancellation method, implemented based on the echo cancellation system as described in claim 1, characterized in that, Includes the following steps: The near-end signal and the far-end signal are processed separately to obtain the near-end frequency domain sequence and the far-end frequency domain sequence; The coherence coefficient is obtained based on the preset reference signal, the near-end frequency domain sequence, the far-end frequency domain sequence, and the dual-talk detection algorithm. Based on the coherence coefficient, different thresholds are selected for the first filter and the second filter to perform dual-talk detection, so as to obtain the first dual-talk detection result and the second dual-talk detection result. During the filtering process, the first filtering coefficients are updated based on the first dual-channel detection result, and the first filtering result is output to the second filter; The second filter is used to update the second filter coefficients based on the first filter result and the second double-pass detection result during the filtering process, so as to obtain a two-stage linear filtering result; Based on the two-stage linear filtering results, in the first state, the second filter is converged to eliminate echo; in the second state, the learning rate of the first filter is adjusted according to the second filter coefficients to converge the first filter to eliminate echo.
3. The echo cancellation method as described in claim 2, characterized in that, The process of processing the near-end and far-end signals separately to obtain near-end frequency domain sequences and far-end frequency domain sequences specifically involves: Acquire the near-end signal and the far-end signal; The near-end signal and the far-end signal are sequentially subjected to frame division, mixed superposition windowing, and M-point discrete Fourier transform to obtain the near-end frequency domain sequence and the far-end frequency domain sequence.
4. The echo cancellation method as described in claim 3, characterized in that, The process of obtaining the coherence coefficient based on the preset reference signal, the near-end frequency domain sequence, the far-end frequency domain sequence, and the dual-talk detection algorithm is as follows: , in, and These are the near-end frequency domain sequence and the far-end frequency domain sequence, respectively. The conjugate operator. The frequency domain cross-correlation value of the near-end and far-end signals. The cross-power spectrum of the near-end and far-end signals. The power spectrum of the near-end signal. The power spectrum of the remote signal. The coherence coefficient at each frequency point.
5. The echo cancellation method as described in claim 4, characterized in that, The step of selecting different thresholds for the first and second filters based on the coherence coefficient to perform dual-talk detection, in order to obtain the first and second dual-talk detection results, specifically involves: , , in, and These are the results of the first and second double-talk tests, respectively. and For different thresholds, and Define 1 as single-talk and 0 as double-talk.
6. The echo cancellation method as described in claim 5, characterized in that, The filtering process of the first filter is as follows: , , ; in, For the reference signal frequency domain sequence, These are the coefficients of the first filter. Let be the order of the first filter. This is the result of the first filtering.
7. The echo cancellation method as described in claim 6, characterized in that, In the filtering process, the first filtering coefficients are updated based on the first dual-talk detection results, specifically as follows: , , + , in, For the autocorrelation power spectrum of the reference signal, The moving average coefficient, To update the step size of the coefficients, The learning rate parameter is calculated from the previous frame.
8. The echo cancellation method as described in claim 7, characterized in that, The filtering process of the second filter and the process of updating the second filter coefficients are as follows: in, The echo is the result of filtering estimation. Let the filter order be . The maximum value of the preset smoothing coefficient. The moving average coefficient, This is the output of the second filter. The cross-power spectrum of the first filter output and the far-end signal is shown. The auto-power spectrum of the far-end signal, The coefficients are the linear regression coefficients between the output of the first filter and the remote signal, which are also the coefficients of the second filter.
9. The echo cancellation method as described in claim 8, characterized in that, The learning rate of the first filter is adjusted according to the second filter coefficient, specifically as follows: Define the coefficient vector of the second filter , Using the conjugate transpose parameter, we obtain the learning rate parameter of the first filter. ,as follows: , in, A constant less than 1 These are the coefficients of the first filter.
Citation Information
Patent Citations
Echo cancellation method, echo cancellation device, conference tablet computer and computer storage medium
CN107123430A
Echo residue judgment method
CN111968663A