Narrow-band minimum mean square error echo suppression method introducing complex neural network

By introducing complex deep neural networks of the STFT domain into the adaptive filtering algorithm, the performance degradation and nonlinear distortion of traditional algorithms when the input signal is highly correlated, achieving more efficient echo signal estimation and steady-state performance improvement of narrowband filters.

CN119920262AActive Publication Date: 2025-05-02CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510092270.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-02
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The performance of traditional adaptive filtering algorithms significantly deteriorates when the input signal is highly correlated, and it is difficult to effectively deal with the nonlinear distortion problem caused by remote voice signals after being played by speakers.

Method used

The minimum mean square error algorithm of complex deep neural network (CDNN) based on the STFT domain is adopted to suppress nonlinear distortion of the input signal through complex DNN, and adaptive control of step size is realized, and the error is time series modeled through complex DNN to accurately estimate the error signal between the actual echo and the filter output echo.

Benefits of technology

It significantly improves the accuracy of echo signal estimation, enhances the steady-state performance of narrowband filters, improves the convergence speed and stability of the algorithm, and has a smaller parameter quantity and better echo cancellation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920262A_ABST
    Figure CN119920262A_ABST
Patent Text Reader

Abstract

The invention relates to a narrowband minimum mean square error echo suppression method introducing a complex neural network, and belongs to the field of adaptive echo cancellation. An existing neural network method still has the problem of inaccurate priori knowledge modeling, and phase information in complex signals is not fully utilized. According to the method, a plurality of DNNs are adopted to carry out modeling and estimation on step length, loudspeaker nonlinear distortion and error signals respectively. And the adaptive capacity to an echo channel in a complex environment is improved by utilizing the efficient parameter learning capacity of a plurality of neural networks. According to the method, the phase information of complex data is fully utilized, the nonlinear distortion problem of the loudspeaker can be effectively suppressed, meanwhile, the small model parameter quantity is kept, efficient echo suppression is achieved, high practicability and hardware adaptability are achieved, and the method is suitable for various practical application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of adaptive echo cancellation and relates to a narrowband minimum mean square error echo suppression method introducing a complex neural network. Background Art

[0002] In a duplex communication system where microphones and speakers interact with each other, the voice of the far-end speaker is first collected by the far-end microphone and transmitted to the far-end communication device, and then reaches the near-end communication device after communication transmission and is played through the near-end speaker. The sound played by the speaker will reflect off walls and other objects to form echo signals, which are collected by the near-end microphone, transmitted through the network, and then returned to the far-end communication device and played through the far-end speaker, causing the far-end speaker to hear his own echo, seriously affecting the communication quality.

[0003] To solve this problem, echo suppression is required, that is, the echo signal in the near-end microphone signal is effectively offset. Adaptive filtering algorithm is the key to echo cancellation. The commonly used algorithm is the adaptive algorithm of the least mean squares (LMS) type. Its core idea is to eliminate echo interference by estimating the echo signal and using the difference between the microphone signal and the estimated signal. The LMS algorithm was first proposed by Widrow and Hoff. The core idea of ​​the LMS algorithm is derived from the Wiener filtering theory, and combined with the concept of steepest gradient descent, the weight coefficient of the filter is updated by minimizing the energy of the error signal. The NLMS (Normalized LMS) algorithm is improved on the basis of the LMS algorithm, taking into account the energy change of the input signal, so that the step size can be adaptively adjusted according to the change of the input signal. This improvement not only improves the convergence speed of the algorithm, but also enhances its stability, so it has been widely used in the field of echo cancellation. However, when the input signals are highly correlated, the performance of these two algorithms will deteriorate significantly. To this end, many scholars have proposed improved algorithms on this basis, such as adaptive optimal step size adjustment.

[0004] However, the modeling of traditional adaptive filters usually requires accurate estimation of expected values, which often depend on unobservable signals. For example, it is usually assumed that the near-end speech signal and the far-end speech signal come from different sound sources and are independent of each other, but in actual applications, there may be a certain correlation between the two in some scenarios. Usually, the power estimation of the far-end speech signal assumes that the signal is stationary, but the signal in the actual environment usually only has short-term stationarity within 20-40ms. When the filter order is too high, there will be a significant difference between the power estimate and the actual value. In addition, many assumptions about the optimal step size are usually valid under theoretical models, but it is often difficult to achieve good results in actual scenarios.

[0005] With the development of deep learning technology, some scholars have tried to combine the advantages and disadvantages of traditional adaptive filtering algorithms and fully deep learning-driven echo cancellation models, while retaining the traditional adaptive filtering framework, using neural networks to estimate parameters that are difficult to accurately estimate by traditional methods. For example, Thomas Haubner proposed a method in 2023 to estimate the step size adaptive control in the NLMS algorithm through a deep neural network (DNN). This method achieved good performance using only a very small number of network parameters. However, this method still relies on a mathematical modeling-based method to calculate the far-end signal power in step size control, and actual tests have shown that this calculation method has a significant impact on the convergence of the algorithm.

[0006] In addition, this method does not fully consider the nonlinear distortion problem caused by the far-end speech signal after it is played by the speaker. This problem is amplified in the frequency domain narrowband adaptive filtering, which further adversely affects the update process of the NLMS algorithm. More importantly, the update of the NLMS algorithm for echo cancellation is usually based on the assumption that the far-end signal and the near-end signal are uncorrelated with each other. However, in actual scenarios, this assumption is often difficult to hold, resulting in interference in error calculation and gradient update, thereby affecting the overall performance of the algorithm. Summary of the invention

[0007] In view of this, the purpose of the present invention is to propose a complex deep neural network CDNN minimum mean square error algorithm based on the STFT domain for echo cancellation. The method uses a complex DNN to suppress nonlinear distortion of the input signal, uses the complex DNN to realize adaptive control of the step size, and uses the complex DNN to perform time series modeling of the error, so as to accurately estimate the error signal between the actual echo and the filter output echo. Since each DNN independently estimates different parameters, it is possible to use a smaller model to effectively and accurately model the actual environment, and the complex neural network can better utilize the phase information of the data. Compared with the echo cancellation method of the traditional adaptive filter, this method has significant advantages in echo signal estimation. At the same time, compared with the algorithm of adjusting the LMS parameters of the existing deep learning method, this method has a smaller number of parameters and provides a better echo cancellation effect.

[0008] In order to achieve the above object, the present invention provides the following technical solutions:

[0009] A narrowband minimum mean square error echo suppression method using a complex neural network is provided, wherein the specific steps are as follows:

[0010] Step 1: In the echo suppression system, the far-end speech signal x(n) is played back by the near-end speaker to generate an echo signal The signal d(n) collected by the near-end microphone is composed of the near-end speech signal s(n) and the far-end speech echo signal Superposition composition.

[0011]

[0012] Step 2: Perform short-time Fourier transform (STFT) on the far-end speech signal x(n) and the near-end microphone signal d(n), where the number of Fourier transform points is N and the step size is p, to obtain the far-end speech signal X in the time-frequency domain k,t , and the near-end speech signal D k,t .

[0013] X k,t =STFT(x(n))

[0014] X k,t =STFT(x(n))

[0015] Step 3: Near-end speech signal D k,t As the reference signal of the narrowband adaptive filter, the far-end speech signal X k,t As the input signal of the narrowband adaptive filter.

[0016] Step 4: The signal in the STFT domain has a total of N / 2+1 frequency points, corresponding to a total of N / 2+1 narrowband filters of order L. The filters at each frequency point have the same structure. f,t Constructed as the input vector to the narrowband filter:

[0017] X f,t =(X f,t ,X f,t-1 …X f,t-L+2 ,X f,t-L+1 )

[0018] The filter coefficient vector of the narrowband filter at the fth frequency point is:

[0019]

[0020] According to the filter input vector and coefficient vector of frequency f, the echo estimation output signal y of frequency f is calculated. f,t :

[0021]

[0022] The reference signal D at frequency f f,t The filter output signal y f,t Perform difference calculation to obtain the error signal:

[0023] e f,t =D f,t -y f,t

[0024] The output y of each frequency point narrowband filter f,t and error e f,t The output data and error data of the speech frame at time t are respectively combined, and the data represent the estimated echo signal and the near-end speech signal in the STFT domain respectively:

[0025]

[0026] Step 5: Apply the error signal e k,t , near-end speech signal D k,t , far-end voice signal X k,t , the gradient value Δh of the last filter update f,t As the input feature of the complex deep neural network (CDNN), the neural network is used to adaptively control the step size parameter, input signal mask and error signal mask to guide the estimated filter coefficient H(k,t) to be iteratively updated in the gradient direction. The update method is as follows:

[0027] h f,τ =h f,t-1 +Δh f,t

[0028] Where Δh f,t is the gradient vector of the filter coefficients, The calculation function of the i-th gradient coefficient is defined as

[0029] Δh i,f,t =μ i,f,t *(mask_X i,f,t *mask_e f,t )

[0030] Step 6: Error signal e k,t Perform inverse short-time Fourier transform as the estimated near-end speech time-frequency domain signal Complete the suppression of echo signals.

[0031] As a further improvement of the present invention, the narrowband adaptive filter adopts a method combining the LMS algorithm and a neural network, wherein the neural network is used to adaptively estimate key parameters required for updating the LMS filter coefficients, thereby being able to effectively cope with complex actual environments.

[0032] As a further improvement of the present invention, the filter updating method is to use the input vector X of the narrowband filter t time frame f,t , reference signal D f,t , error signal e f,t and the gradient Δh of the previous time frame l,f,t-1 Splicing as a step-size adaptive control network The input features of the network output step vector, the step vector is The update formula is as follows

[0033]

[0034] The input vector X of the narrowband filter t time frame f,t As a nonlinear distortion suppression network The input features of the network are mask_X, which is a vector that suppresses nonlinear distortion. f,t ,

[0035]

[0036] The error data e of the narrowband filter t time frame f,t As an error estimation network Input features, network output error mask_e f,t ,

[0037]

[0038] The filter coefficient gradient vector is as follows:

[0039]

[0040] The gradient update formula of the i-th coefficient is:

[0041] Δh i,f,t =μ i,f,t *(mask_X i,f,t *mask_e f,t )

[0042] According to the update criterion of the LMS algorithm, the coefficients of each narrowband filter are iteratively updated:

[0043] h f,τ =h f,t-1 +Δh f,t

[0044] As a further improvement of the present invention, the step size adaptive control network The invention is characterized in that the power of the far-end speech signal and the near-end speech signal is estimated more accurately and combined with the previous gradient information to realize the adaptive adjustment of the network output step vector, so as to independently control the update step size of each filter coefficient.

[0045] As a further improvement of the present invention, the step size adaptive control network The specific network structure is to f,t , D f,t , e f,t and Δh l,f,t-1 Splice into vector Z f,t∈(2L+2) as the step size adaptive control neural network The input features of The architecture consists of at least one complex fully connected layer with input feature dimension 2L+2 and output dimension >4L+4 with RELU activation, at least one complex gated recurrent unit (Complex GRU) with dimension >4L, one or more complex fully connected layers with output dimension >4L with RELU activation, and one complex fully connected layer with output dimension L, with output vector μ f,t .

[0046] As a further improvement of the present invention, the remote signal nonlinear distortion estimation network The characteristic is that in practical applications, nonlinear distortion will be generated after the near-end speaker plays the far-end voice signal, and this distortion is usually particularly obvious in the frequency domain. Through the nonlinear distortion suppression network, these distortion components can be accurately suppressed, thereby effectively improving the steady-state performance of the narrowband filter.

[0047] As a further improvement of the present invention, the remote signal nonlinear distortion estimation network The specific network structure is at least one complex fully connected layer with input feature dimension L, the output dimension of this layer is greater than 2L, and has a ReLU activation function; at least one complex gated recurrent unit (Complex GRU) with hidden dimension greater than 2L; at least one complex fully connected layer with output dimension greater than 2L, with a ReLU activation function; a complex fully connected layer with output dimension L is connected in series, and the output vector is mask_X f,t .

[0048] As a further improvement of the present invention, the error estimation network The characteristic is that the derivation process of the LMS algorithm in echo cancellation is based on the assumption that there is no correlation between the far-end speech signal and the near-end speech signal, so that the gradient update will not be affected in the expected calculation process. The network can effectively solve the influence of the relevant part of the error on the accuracy of gradient update when there is partial correlation between the far-end signal and the near-end speech signal.

[0049] As a further improvement of the present invention, the error estimation network The specific network structure is a complex fully connected layer with an input feature dimension of 1, an output dimension of this layer greater than 4, and a ReLU activation function; at least one complex gated recurrent unit (Complex GRU) with a hidden dimension greater than 4; at least one complex fully connected layer with an output dimension greater than 4, with a ReLU activation function; a complex fully connected layer with an output dimension of 1 connected in series, and the output error mask mask_e f,t .

[0050] As a further improvement of the present invention, during the parameter training process of the complex neural network, the loss function uses a time domain negative log return loss enhancement (TD-ERLE) function to guide model training.

[0051]

[0052] where δ LOSS >0 is a regularization factor to avoid abnormal values, s(n) is the time domain signal of the near-end speech, and s'(n) is the near-end speech time domain signal s'(n) estimated by the adaptive filtering algorithm.

[0053] (1) The present invention adopts complex DNN to model and estimate the step size, speaker nonlinear distortion and error signal respectively, making full use of the phase information of complex data to more accurately capture the echo channel characteristics, thereby improving the accuracy of echo estimation.

[0054] (2) The present invention combines CDNN with the LMS algorithm and uses CDNN to estimate the key parameters required for updating the LMS filter coefficients, which effectively solves the performance problem of the traditional LMS algorithm in complex environments and further improves the echo suppression effect.

[0055] (3) The step size adaptive control network dynamically adjusts the update step size of each filter coefficient according to the input signal and gradient information, thereby improving the convergence speed and stability of the algorithm and adapting to the echo suppression requirements in different scenarios.

[0056] (4) The nonlinear distortion estimation network can effectively suppress the nonlinear distortion generated when the speaker plays the far-end voice signal, improve the steady-state performance of the narrowband filter, and avoid the degradation of the echo suppression effect due to distortion.

[0057] (5) The present invention uses multiple independent DNNs to estimate different parameters respectively. The structure of each DNN is relatively simple and the number of parameters is small, which reduces the computational complexity and improves the efficiency of the algorithm.

[0058] (6) Using STFT domain processing can effectively reduce the amount of calculation and facilitate frequency domain analysis and filtering, further improving the efficiency of the algorithm.

[0059] (7) The DNN structure adopted in the present invention is relatively simple, easy to implement and deploy, can adapt to different hardware platforms, and has broad application prospects.

[0060] (8) The present invention has achieved good echo suppression effects on multiple test sets, indicating that the algorithm has good generalization ability and can adapt to different environments and scenarios.

[0061] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0063] Figure 1 It is a structural schematic diagram of the present invention;

[0064] Figure 2 The model structure and calculation method diagram of the step-size adaptive control network of the present invention;

[0065] Figure 3 The model structure and calculation method diagram of the far-end signal nonlinear distortion estimation network of the present invention;

[0066] Figure 4 The model structure and calculation method diagram of the error estimation network of the present invention;

[0067] Figure 5 is the average ERLE curve of Benfaming on test set 1;

[0068] Figure 6 is the average ERLE curve of this method on test set 2;

[0069] Figure 7 This is a spectrogram after the present invention implements echo suppression in a real echo scene. DETAILED DESCRIPTION

[0070] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0071] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0072] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0073] The embodiment of the present invention discloses a narrowband minimum mean square error echo suppression method that introduces a complex neural network. The core idea of ​​the echo suppression method is to suppress the echo through an adaptive filter, such as Figure 1 As shown, the basic structure of the method is demonstrated. Figure 2 The model structure and calculation method diagram of the step-size adaptive control network of the present invention; Figure 3 The model structure and calculation method diagram of the far-end signal nonlinear distortion estimation network of the present invention; Figure 4 It is a diagram of the model structure and calculation method of the error estimation network of the present invention.

[0074] In the echo suppression system, the far-end speech signal x(n) generates an echo signal after being played by the near-end speaker. The signal d(n) collected by the near-end microphone is composed of the near-end speech signal s(n) and the far-end speech echo signal Superposition composition.

[0075]

[0076] The far-end speech signal x(n) and the near-end microphone signal d(n) are subjected to short-time Fourier transform (STFT), where the number of Fourier transform points is 512 and the step size is 128, to obtain the time-frequency domain complex far-end speech signal X k,t , and the near-end speech signal D k,t , the frequency point number of the signal is 257.

[0077] X k,t =STFT(x(n))

[0078] Xk,t =STFT(x(n))

[0079] The STFT domain signal is processed in the time dimension of each frequency point using a narrowband adaptive filter with consistent parameters, where the near-end speech signal D k,t As the reference signal of the narrowband adaptive filter, the far-end speech signal X k,t As the input signal of the narrowband adaptive filter.

[0080] The signal in the STFT domain has a total of N / 2+1 frequency points, corresponding to a total of N / 2+1 narrowband filters of order L. The filters at each frequency point have the same structure. f,t Constructed as the input vector to the narrowband filter:

[0081] X f,t =(X f,t ,X f,t-1 …X f,t-L+2 ,X f,t-L+1 )

[0082] The filter coefficient vector of the narrowband filter at the fth frequency point is:

[0083]

[0084] According to the filter input vector and coefficient vector of frequency f, the echo estimation output signal y of frequency f is calculated. f,t :

[0085]

[0086] The reference signal D at frequency f f,t The filter output signal y f,t Perform difference calculation to obtain the error signal:

[0087] e f,t =D f,t -y f,t

[0088] The output y of each frequency point narrowband filter f,t and error e f,t The output data and error data of the speech frame at time t are respectively combined, and the data represent the estimated echo signal and the near-end speech signal in the STFT domain respectively:

[0089]

[0090] Using a narrowband filter, the input vector X of time frame t f,t , reference signal D f,t , error signal e f,tand the gradient Δh of the previous time frame l,f,t-1 Splice to Z f,t ∈18 as the step size adaptive control network The input features of the network output step vector, the step vector is By more accurately estimating the power of the far-end speech signal and the near-end speech signal and combining the previous gradient information, the network output step vector is adaptively adjusted, thereby independently controlling the update step size of each filter coefficient.

[0091] Step-size adaptive control network The structure consists of a complex fully connected layer with input feature dimension of 18 and output dimension of 48 with RELU activation, a complex gated recurrent unit (Complex GRU) with hidden dimension of 48, a complex fully connected layer with output dimension of 48 and RELU activation, and a complex fully connected layer with output dimension of 8. The output vector is μ f,t ,

[0092]

[0093] The input vector X of the narrowband filter t time frame f,t As a nonlinear distortion suppression network The input feature has a dimension of 8. Through the nonlinear distortion suppression network, the nonlinear distortion components of the loudspeaker can be accurately suppressed, thereby effectively improving the steady-state performance of the narrowband filter. The network structure consists of a complex fully connected layer with an input feature dimension of 8, an output dimension greater than 18, and a ReLU activation function; a complex gated recurrent unit (Complex GRU) with a hidden dimension of 18; a complex fully connected layer with an output dimension of 18, with a ReLU activation function; a complex fully connected layer with an output dimension of 8, and the output vector is mask_X f,t .

[0094] Error data e of time frame t of narrowband filter f,t As an error estimation network The input features solve the problem that the derivation process of the LMS algorithm in echo cancellation is based on the assumption that there is no correlation between the far-end speech signal and the near-end speech signal, which is inconsistent with the actual situation, so that the gradient update will not be affected in the expected calculation process. The specific network structure is a complex fully connected layer with an input feature dimension of 1, an output dimension of 5, and a ReLU activation function; a complex gated recurrent unit (Complex GRU) with a hidden dimension of 5; a complex fully connected layer with an output dimension of 5, with a ReLU activation function; a complex fully connected layer with an output dimension of 5 connected in series, and the output error mask mask_e f,t .

[0095]

[0096] The filter coefficient gradient vector is as follows:

[0097]

[0098] The gradient update formula of the i-th coefficient is:

[0099] Δh i,f,t =μ i,f,t *(mask_X i,f,t *mask_e f,t )

[0100] According to the update criterion of the LMS algorithm, the coefficients of each narrowband filter are iteratively updated:

[0101] h f,τ =h f,t-1 +Δh f,t

[0102] Error and signal vector in time frame Perform inverse short-time Fourier transform as the estimated near-end speech time domain signal s'(n) to suppress the echo signal.

[0103] During the training of the neural network, the loss function uses the time-domain negative log return loss enhancement (TD-ERLE) function to guide the model training.

[0104]

[0105] where δ LOSS >0 is a regularization factor to avoid abnormal values, s(n) is the time domain signal of the near-end speech, and s'(n) is the near-end speech time domain signal s'(n) estimated by the adaptive filtering algorithm.

[0106] The model training dataset uses the dataset from the Microsoft AEC Challenge as the training audio library, from which 10,000 speech clips are extracted, each of which is 10 seconds long. During the training process, far-end and near-end signals are randomly extracted from the corpus, and the voice liveness detection method (VAD) is used to extract 1s speech clips from each audio signal. At the same time, in order to simulate the nonlinear distortion of the near-end speaker playback signal, a nonlinear function is used for processing. The function is,

[0107]

[0108] Where far(n) is the far-end signal and β is randomly selected from {0,5,1,10,99}. The resulting nonlinear distortion signal far'(n) is convolved with the room impulse response. The RIR is modeled as a Gaussian white noise with exponential decay modulation to achieve a reverberation time (RT60) randomly sampled in a continuous range of [20,192] milliseconds. The signal-to-echo SER of each training data is randomly sampled in the range of {-10db, 10db}. Using this method, a total of 2.77 hours of training data were generated.

[0109] During the training process, the ADAM optimizer was used to train the CDNN network for a total of 80 epochs with an initial learning rate of 0.001. When the performance of the training set did not improve for three consecutive times, the learning rate was reduced to 70% of the current value. In addition, during the back propagation of the model, an l2 norm gradient clipping mechanism with a threshold of 1 was introduced to stabilize the training process. Finally, the best training model was saved according to the minimum value of the loss function on the validation set.

[0110] In order to demonstrate the performance of the present invention in the field of echo cancellation, two traditional adaptive algorithms based on mathematical modeling, frequency domain narrowband NLMS (STFT-NLMS) and frequency domain Kalman filter (PFDKF), and an adaptive filtering algorithm of a hybrid neural network (E2E-NB-DNN, hereinafter referred to as NB-DNN) are compared. All methods use the same window length, covering the corresponding room impulse response duration of about 88ms, and use three test sets to evaluate the performance. Test set 1 is a synthetic data set of far-end single talk, with a total of 200 audio clips, each audio clip is 10s, and the far-end voice signal uses the clean voice in the DNS challenge as the corpus. In order to test the generalization ability of the model, the nonlinear distortion in the test set is different from that in the training data set, and the function is

[0111]

[0112] Where α = 10 4 The image source method is used to simulate the real room impulse response RIR. The reverberation length of the room impulse response is truncated to [64ms-92ms] to test the performance of the algorithm in the range that the window length can cover. It is convolved with the nonlinear distortion signal and output as an echo audio signal. For far-end single talk, the scoring standard is tested using the objective evaluation standard echo loss enhancement (ERLE). Figure 5 is the average ERLE of all test samples in the test set. The calculation formula of ERLE is as follows:

[0113]

[0114] Test set 2 simulates the performance degradation problem when the reverberation time exceeds the filter window coverage length in a real environment. It is produced in the same way as data set 1, where the reverberation length of the room impulse response is truncated to [128ms-192ms]. Figure 6 It is the average ERLE of all test samples in the test set. It can be seen from Table 1 that even when the reverberation length is about 2 times the window length, the performance degradation of the algorithm proposed in the present invention is minimal and a good echo suppression effect can still be maintained.

[0115] Test set 3 is a test set recorded in an actual environment, including clean blind double-talk audio in the ICASSP 2021 AEC Challenge. The data set contains 200 audio clips, half of which are echo RIR changes. For the double-talk scenario, since there is no clean near-end speech signal in the clean blind test set of the AEC challenge, the non-intrusive perceptual objective metric AECMOS proposed by Microsoft is used to evaluate the performance of AEC in the real recorded test set. Among them, AECMOS-Echo in the score represents the echo quality, which evaluates the system's echo suppression effect and perceptual quality. The higher the value, the better the echo suppression performance. AECMOS-Other represents the near-end speech quality. The higher the value, the better the near-end speech retention. Through these two indicators, it can be judged whether the near-end speech is distorted due to excessive suppression while achieving good echo suppression (higher AECMOS-Echo score). Table 1 is a performance comparison of different echo cancellation models. According to Table 1, the comparison with the classic model and the current advanced model shows that the proposed method can have excellent echo cancellation performance in real scenarios, and can ensure that the near-end speech distortion is small, while having a lower number of model parameters. Figure 7 The spectrogram is shown after the method of the present invention achieves echo suppression in a real echo scene.

[0116] Table 1

[0117]

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A narrowband minimum mean square error echo suppression method using a complex neural network, characterized in that: The method comprises the following steps: Performing short-time Fourier transform (STFT) on the far-end speech signal x(n) and the near-end microphone signal d(n) in the echo suppression system; The far-end speech signal X in the STFT domain k,t , and the near-end speech signal D k,t In the time dimension of each frequency point, a narrowband adaptive filter with the same parameters is used for filtering; For the narrowband adaptive filter, a method combining the LMS algorithm and the complex deep neural network CDNN is adopted, wherein the complex deep neural network CDNN is used to adaptively estimate the parameters required for updating the LMS filter coefficients, so as to cope with complex actual environments; The key parameters required for updating the LMS filter coefficients include a step vector, a speaker signal nonlinear distortion suppression vector and an error estimate, and the three parameters are ultimately used to calculate a filter gradient vector; The step vector estimates the input vector X using a narrowband filter for t time frames f,t , reference signal D f,t , error signal e f,t and the gradient Δh of the previous time frame l,f,t-1 Splicing as a step-size adaptive control network The input features of the network output step vector, the step vector is The update formula is as follows: The nonlinear distortion suppression vector of the loudspeaker signal is converted into the input vector X of the narrowband filter t time frame. f,t As a nonlinear distortion suppression network The input features of the network are mask_X, which is a vector that suppresses nonlinear distortion. f,t , The error estimate is the error data e of the narrowband filter in time frame t f,t As an error estimation network Input features, network output error mask_e f,t , The key parameters required for updating the LMS filter coefficients described above are the gradient vectors of the filter coefficients , where the gradient update formula of the i-th coefficient is: Dh i,f,t =μ i,f,t *(mask_X i,f,t *mask_e f,t ) Finally, the gradient vector calculated by the complex deep neural network CDNN is used to iteratively update the coefficients of each narrowband filter: h f,τ =h f,t-1 +Δh f,t 。 2. The narrowband minimum mean square error echo suppression method using a complex neural network according to claim 1, characterized in that: The method of combining the LMS algorithm with the complex deep neural network CDNN is specifically as follows: For the frequency point f, input the far-end voice signal X f,t Constructed as the input vector to the narrowband filter: X f,t =(X f,t ,X f,t-1 …X f,t-L+2 ,X f,t-L+1 ) The filter coefficient vector of the narrowband filter at the fth frequency point is: According to the filter input vector and coefficient vector of frequency f, the echo estimation output signal y of frequency f is calculated. f,t : The reference signal D at the frequency point f f,t The filter output signal y f,t Perform difference calculation to obtain the error signal: e f,t =D f,t -y f,t The output y of each frequency point narrowband filter f,t and error e f,t The output data and error data of the speech frame at time t are respectively combined, and the data represent the estimated echo signal and the near-end speech signal in the STFT domain respectively: The output of the complex deep neural network CDNN is used to calculate the gradient vector and iteratively update the filter coefficients H(k,t); Error signal e k,t Perform inverse short-time Fourier transform as the estimated near-end speech time-frequency domain signal s'(n) to complete the suppression of the echo signal.

3. The narrowband minimum mean square error echo suppression method using a complex neural network as claimed in claim 1, characterized in that: The step size adaptive control network Specifically: By estimating the power of the far-end speech signal and the near-end speech signal and combining the previous gradient information, the network output step vector is adaptively adjusted to independently control the update step size of each filter coefficient; Step-size adaptive control network The network structure is: X f,t , D f,t , e f,t and Δh l,f,t-1 Splice into vector Z f,t ∈(2L+2) as the step size adaptive control neural network The input features of The architecture consists of at least one complex fully connected layer with input feature dimension 2L+2 and output dimension >4L+4 with RELU activation, at least one complex gated recurrent unit with dimension >4L, one or more complex fully connected layers with output dimension >4L with RELU activation, and one complex fully connected layer with output dimension L, with output vector μ f,t .

4. The narrowband minimum mean square error echo suppression method using a complex neural network as claimed in claim 1, characterized in that: The nonlinear distortion estimation network The specific network structure is at least one complex fully connected layer with input feature dimension L, the output dimension of this layer is greater than 2L, and has a ReLU activation function; at least one layer of complex gated recurrent unit with hidden dimension greater than 2L; at least one layer of complex fully connected layer with output dimension greater than 2L, with a ReLU activation function; a layer of complex fully connected layer with output dimension L is connected in series, and the output vector is mask_X f,t .

5. The narrowband minimum mean square error echo suppression method using a complex neural network as claimed in claim 1, characterized in that: The error estimation network The network structure is: A complex fully connected layer with an input feature dimension of 1, an output dimension greater than 4, and a ReLU activation function; at least one complex gated recurrent unit with a hidden dimension greater than 4; at least one complex fully connected layer with an output dimension greater than 4, with a ReLU activation function; a complex fully connected layer with an output dimension of 1 connected in series, outputting an error mask mask_e f,t .

6. The narrowband minimum mean square error echo suppression method using a complex neural network as claimed in claim 1, characterized in that: The complex deep neural network CDNN is used to adaptively estimate the parameters required for updating the LMS filter coefficients. During the training process of the complex deep neural network, the loss function uses a time-domain negative logarithmic return loss enhancement function to guide the model training. where δ LOSS >0 is a regularization factor to avoid abnormal values, s(n) is the time domain signal of the near-end speech, and s'(n) is the near-end speech time domain signal s'(n) estimated by the adaptive filtering algorithm.

Citation Information

Patent Citations

  • Low-complexity residual echo suppression method combining signal processing and deep neural network

    CN115312073A

  • DOA estimation method based on double-channel complex convolutional neural network

    CN118051719A

  • Full-duplex digital self-interference suppression method based on complex neural network

    CN118869097A

  • Echo cancellation and noise reduction method and device

    CN119091900A

  • Echo cancellation method and device based on deep learning and readable storage medium

    CN119107964A

Cited By

  • Building talkback echo elimination method and system based on multi-level self-adaption and neural network fusion, and building talkback terminal

    CN121583273A

  • Active noise reduction earphone control method and system based on deep learning

    CN121617378A