Multi-frame filtering amplitude-phase decoupled deep neural network stereo acoustic echo cancellation method and system based on human ear hearing

By using a deep neural network based on multi-frame filtering amplitude and phase decoupling of human hearing, the problem of poor echo cancellation effect of traditional stereo echo cancellation algorithm in complex noise scenarios is solved, and more efficient echo suppression and near-end speech quality improvement are achieved.

CN120089148BActive Publication Date: 2026-03-24INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Traditional stereo echo cancellation algorithms struggle to effectively remove interfering echoes in complex noise scenarios, and existing deep learning algorithms neglect the harmonic structure of speech signals, resulting in poor echo cancellation performance.

Method used

A multi-frame filtering amplitude-phase decoupling deep neural network based on human hearing is adopted. The amplitude spectrum and phase information of near-end speech are recovered through a multi-stage neural network. The network is trained using the human hearing perception cost function to process linear echo and noise in steps.

Benefits of technology

It improves echo cancellation performance and near-end speech quality in complex noise scenarios, and enhances speech intelligibility and echo suppression capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089148B_ABST
    Figure CN120089148B_ABST
Patent Text Reader

Abstract

The application discloses a multi-frame filtering amplitude-phase decoupling deep neural network stereo echo cancellation method and system based on human ear hearing. The method comprises the following steps: using a microphone to receive a signal and two-channel reference signal to estimate the real part and the imaginary part of a linear filter, and using a multi-frame filter structure to estimate a linear echo signal in real time to obtain a residual echo complex spectrum; estimating a near-end speech amplitude spectrum, training a neural network for estimating the near-end speech amplitude spectrum through a human ear hearing related cost function; performing secondary estimation on the near-end speech complex spectrum, and training the neural network through the human ear hearing related cost function; and performing stereo echo cancellation through the trained neural network. The application decouples the mapping of the near-end speech complex spectrum in the stereo echo cancellation system into amplitude spectrum mapping and phase mapping, and finely recovers the near-end speech complex spectrum through secondary estimation, thereby improving the quality and intelligibility of the near-end speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of stereo echo cancellation based on deep neural networks, specifically relating to a stereo echo cancellation method and system based on human hearing, using multi-frame filtering amplitude and phase decoupling deep neural networks. Background Technology

[0002] Due to the high correlation between the two reference signals, traditional stereo echo cancellation algorithms based on adaptive filtering suffer from non-uniqueness, slow filter convergence, and significant susceptibility to background noise, failing to effectively remove interfering echoes. Decorrelation algorithms can reduce the correlation between the two reference signals, but they also compromise the speech quality and stereo image of the reference signals. In recent years, echo cancellation algorithms based on deep learning have attracted widespread attention both domestically and internationally. These algorithms utilize deep neural networks to estimate near-end speech from the microphone received signal, establishing a nonlinear mapping relationship between the microphone received signal and the near-end speech signal. This avoids decorrelation processing and fully preserves the stereo image.

[0003] Existing deep learning-based stereo echo cancellation algorithms struggle to directly recover the amplitude and phase information of the target signal from the interference signal in complex noise scenarios. This is because the amplitude spectrum exhibits a clear and regular harmonic structure, while the phase spectrum lacks this regularity, making phase information more difficult to extract than amplitude information. Furthermore, most existing deep learning algorithms train deep neural networks based on mean squared error (MSE), neglecting the impact of the cost function on the recovery of the harmonic structure of the speech signal. The magnitude of the MSE is not entirely correlated with speech quality. Summary of the Invention

[0004] The purpose of this invention is to overcome the problem of stereo echo cancellation in complex noise scenarios. This invention provides a stereo echo cancellation method and system based on multi-frame filtering amplitude-phase decoupling using an auditory perception cost function. It utilizes a multi-stage deep neural network to decouple the amplitude and phase mappings of the near-end speech, decoupling the complex spectrum mapping of the near-end speech into amplitude spectrum mapping and phase mapping. First, the deep neural network recovers the amplitude spectrum of the near-end speech with obvious harmonic structure, and combines it with the phase of the microphone received signal to perform a preliminary estimation of the complex spectrum of the near-end speech. Then, the deep neural network further optimizes the complex spectrum of the near-end speech. Furthermore, this invention, starting from the principle of echo generation, constructs a multi-frame filtering structure and uses a neural network to estimate the room impulse response in the frequency domain in real time to obtain the linear echo in the microphone received signal. Then, another neural network is used to remove residual echoes and noise, processing linear echoes and residual echoes and noise separately, reducing the difficulty of network modeling and improving echo cancellation performance. Meanwhile, a neural network cost function based on human hearing is constructed. The cost function is weighted using the clean speech amplitude spectrum. Considering the masking effect of human hearing, the power exponent is used to control the degree of noise suppression and speech preservation of the deep neural network, thereby further improving the near-end speech quality.

[0005] To achieve the above objectives, this invention provides a stereo echo cancellation method based on human hearing, using a multi-frame filtering amplitude-phase decoupling deep neural network. The method includes:

[0006] Step 1): Using the microphone received signal and the complex spectrum of the two-channel reference signal of the stereo echo system, estimate the real and imaginary parts of the linear filter, and design a multi-frame filter structure to estimate the linear echo signal online in real time and obtain the residual echo complex spectrum.

[0007] Step 2): Estimate the near-end speech amplitude spectrum using the residual echo amplitude spectrum, the microphone received signal amplitude spectrum, and the two-channel reference signal amplitude spectrum. Train the neural network for estimating the near-end speech amplitude spectrum using the auditory perception related cost function.

[0008] Step 3): Using the amplitude spectrum of the near-end speech estimated in Step 2) and the phase of the microphone received signal, a preliminary estimate of the complex spectrum of the near-end speech is obtained. This spectrum, along with the complex spectrum of the microphone received signal, is input into the neural network to perform a second estimation of the complex spectrum of the near-end speech. The neural network is then trained using the human auditory correlation cost function to improve speech quality.

[0009] Step 4): Input the microphone received signal and the two-channel reference signal into the trained neural network to perform stereo echo cancellation.

[0010] As an improvement to the above method, step 1) specifically includes:

[0011] Step 1-1): Perform a short-time Fourier transform on the microphone received signal and the two-channel reference signal to obtain the real and imaginary parts of its complex spectrum;

[0012] Step 1-2): Input the real and imaginary parts of the complex spectrum of the microphone received signal and the two-channel reference signal obtained in Step 1-1) into the convolutional recurrent neural network. The decoding module includes four parts, which are the real and imaginary parts of the complex spectrum of the impulse response of the two rooms in the near-end room.

[0013] Steps 1-3): The real and imaginary parts of the complex spectrum of the room impulse response output from the convolutional recurrent neural network in Steps 1-2) are compared with the real and imaginary parts of the microphone received signal using a set multi-frame filtering method to obtain the estimated linear echo signal; the multi-frame filtering calculation formula is shown below, where... , , , These are the real and imaginary parts of the complex spectrum of the room's impulse response. , , , These are the real and imaginary parts of the complex spectrum of the two-channel reference signal. and These are the estimated real and imaginary parts of the linear echo complex spectrum. It is a multi-frame filtering function. It is a frame index. It's a frequency point. These are the indices of the time-frequency points used in the calculation of the multi-frame filtering function. L It is the total frame length used in the calculation of the multi-frame filtering function;

[0014]

[0015]

[0016]

[0017] Steps 1-4): Subtract the real and imaginary parts of the linear echo complex spectrum estimated in step 1-3) from the real and imaginary parts of the complex spectrum of the received signal from the microphone to obtain the real and imaginary parts of the residual echo complex spectrum.

[0018] As an improvement to the above method, step 2) specifically includes:

[0019] Step 2-1): Extract the amplitude spectrum features of the residual echo complex spectrum, the microphone received signal complex spectrum, and the reference signal complex spectrum, input them into the convolutional recurrent neural network, and estimate the near-end speech amplitude spectrum;

[0020] Step 2-2): Train the neural networks in Steps 2-1) and 1-2) using the auditory perception-related cost function, as shown below, where... It is the estimated near-end speech amplitude spectrum. It is the true near-end speech amplitude spectrum. It is the power-law parameter of the cost function related to auditory perception. .

[0021]

[0022] As an improvement to the above method, step 3) specifically includes:

[0023] Step 3-1): Combine the near-end speech amplitude spectrum estimated in Step 2) with the phase of the microphone received signal to obtain the preliminary estimated real and imaginary parts of the near-end speech complex spectrum;

[0024] Step 3-2): Input the real and imaginary parts of the complex spectrum of the near-end speech obtained in Step 3-1) together with the real and imaginary parts of the complex spectrum of the microphone received signal into the convolutional recurrent neural network. The decoding module includes two parts: re-estimation of the real part of the complex spectrum of the near-end speech and re-estimation of the imaginary part of the complex spectrum.

[0025] Steps 3-4): After the neural network training in Step 2-2) is completed, the convolutional recurrent neural network in Step 3-2) is trained using the human auditory cost function, and the weight parameters of the neural network in Step 2-2) are fine-tuned. As shown below, where , It is the real and imaginary parts of the complex spectrum of near-end speech estimated in a quadratic manner. and These are the real and imaginary parts of the complex spectrum of actual near-end speech. It is the complex spectrum of actual near-end speech. It is the estimated complex spectrum of the near-end speech;

[0026]

[0027]

[0028] The convolutional recurrent neural network in steps 1-2), 2-1), and 3-2) is only one type of neural network structure for implementing the present invention, and the method of the present invention is also applicable to other neural network structures.

[0029] The present invention also provides a stereo echo cancellation system based on human hearing and multi-frame filtering amplitude-phase decoupling deep neural network, employing the above method, comprising:

[0030] The linear echo cancellation module is used to estimate the real and imaginary parts of the linear filter using the microphone received signal and the complex spectrum of the two-channel reference signal of the stereo echo system, and to estimate the linear echo signal online in real time using a multi-frame filter structure to obtain the residual echo complex spectrum.

[0031] The near-end speech amplitude spectrum estimation module is used to estimate the near-end speech amplitude spectrum using the residual echo amplitude spectrum, the microphone received signal amplitude spectrum and the two-channel reference signal amplitude spectrum. The neural network for estimating the near-end speech amplitude spectrum is trained using the human auditory correlation cost function.

[0032] The near-end speech complex spectrum estimation module is used to obtain a preliminary estimate of the near-end speech complex spectrum by using the estimated near-end speech amplitude spectrum and the phase of the microphone received signal. This estimate, along with the microphone received signal complex spectrum, is then input into a neural network to perform a secondary estimation of the near-end speech complex spectrum. The neural network is then trained using the human auditory correlation cost function to improve speech quality.

[0033] The stereo echo cancellation module is used to input the microphone received signal and two-channel reference signals into a trained neural network to perform stereo echo cancellation.

[0034] The advantages of this invention are:

[0035] 1. The method provided by this invention decouples the mapping of the complex spectrum of near-end speech in a stereo echo cancellation system into amplitude spectrum mapping and phase mapping, and performs fine recovery of the complex spectrum of near-end speech through a quadratic estimation method, thereby improving the quality and intelligibility of near-end speech.

[0036] 2. The method provided by this invention combines deep neural networks with traditional adaptive filtering. It uses neural networks to estimate the parameters of the adaptive filter in real time and designs a multi-frame filter structure suitable for neural networks to estimate linear echoes in order to eliminate residual echoes and noise. It breaks down echo cancellation in complex noise scenarios into steps to further improve the echo suppression.

[0037] 3. This invention introduces the characteristics of human auditory perception into a neural network-based stereo echo cancellation system through a cost function, and utilizes the masking effect of the human ear to improve the neural network's ability to recover the spectral peaks and valleys of near-end speech signals, thereby improving the quality and intelligibility of near-end speech. Attached Figure Description

[0038] Figure 1 This is a flowchart of a stereo echo cancellation algorithm based on human hearing, using multi-frame filtering amplitude and phase decoupling deep neural networks.

[0039] Figure 2 This is a flowchart of an algorithm that combines amplitude adjustment and feature compensation.

[0040] Figure 3 These are the time-domain amplitude adjustment results of different algorithms.

[0041] Figure 4 This is a comparison of PESQ and STOI scores of different algorithms in a mechanical noise scenario on the training set.

[0042] Figure 5 The PESQ and STOI scores (average results under different noise levels) are for different compensation algorithms in a noise-free scenario. Detailed Implementation

[0043] The method of the present invention will now be described in detail with reference to the accompanying drawings.

[0044] The method of this invention can improve the echo suppression and near-end speech quality of deep learning-based stereo echo cancellation algorithms in complex noise scenarios. First, a multi-frame filter for deep learning is designed. A first-stage convolutional recurrent neural network is used to estimate the multi-frame filter parameters to eliminate linear echo. Then, a second-stage convolutional neural network is used to estimate the near-end speech amplitude spectrum to suppress residual echo and noise. Finally, the estimated near-end speech amplitude spectrum is combined with the phase of the microphone received signal to obtain a preliminary estimated near-end speech complex spectrum. A third-stage convolutional recurrent network is then used to perform a second estimation of the near-end speech complex spectrum, improving the quality of near-end speech complex spectrum recovery.

[0045] like Figure 1 As shown, a stereo echo cancellation method based on human hearing, using multi-frame filtering amplitude and phase decoupling deep neural networks, includes the following steps:

[0046] Step 1) Linear echo cancellation:

[0047] First, short-time Fourier transforms are performed on the microphone received signal and the two-channel reference signals to obtain their complex spectra. The real and imaginary parts of the complex spectra of these three signals are then used to construct a 6-channel time-frequency input feature, which is input into the first-stage convolutional recurrent neural network (RNN) to estimate the multi-frame filter parameters. The RNN consists of an encoding / decoding module and a timing modeling module. The timing modeling module includes a two-layer GRU network, the encoding module includes a multi-layer 2D convolutional neural network, and the decoding module consists of four parts, each composed of a multi-layer deconvolutional neural network, representing the real and imaginary parts of the complex spectrum of the impulse response from the two rooms in the near-end room, respectively. The estimated filter parameters are convolved with the two-channel reference signals to obtain the estimated linear echo signal. The multi-frame filtering calculation formula is shown below, where... , , , These are the real and imaginary parts of the complex spectrum of the room's impulse response. , , , These are the real and imaginary parts of the complex spectrum of the two-channel reference signal. and These are the estimated real and imaginary parts of the linear echo.

[0048]

[0049]

[0050] Represents a multi-frame filtering function, with For example, the calculation formula is shown below. of L Each feature map is equivalent to an order length of... L The filter vector consists of feature maps corresponding to the same time-frequency point. continuous L The vectors composed of time-frequency points in each frame are summed to obtain the multi-frame filtered output for the current time-frequency point:

[0051]

[0052] The real and imaginary parts of the complex spectrum of the signal received from the microphone. , Subtracting the estimated real and imaginary parts of the linear echo complex spectrum from the original data yields the real and imaginary parts of the residual echo complex spectrum. , }

[0053] Step 2) Near-end speech amplitude spectrum estimation:

[0054] The residual echo amplitude spectrum is calculated based on the real and imaginary parts of the residual echo complex spectrum obtained in step 1). And calculate their respective amplitude spectra based on the microphone received signal and the complex spectra of the two channel reference signals. , , The amplitude spectra of the four signals are combined to form a 4-channel time-frequency feature, which is then input into the second-stage convolutional recurrent network. The second-stage convolutional recurrent network has the same structure as the first-stage convolutional recurrent network, including an encoding / decoding module and a temporal modeling module. The decoding module includes one part, which is the estimation of the amplitude spectrum of the near-end speech.

[0055] During network training, a first-stage convolutional recurrent neural network (CRN) and a second-stage CRN are cascaded, and the two-stage networks are jointly trained using the amplitude spectrum of real near-end speech as a constraint. The MSE function is a commonly used cost function in deep neural network training; however, the magnitude of the MSE error is not entirely correlated with speech quality. This invention introduces the influence of human auditory characteristics into the network training process, using a weighted Euclidean distance cost function as the network training cost function, as shown in the following equation. The amplitude spectrum of clean speech is used to weight the MSE error function. It is the power exponent parameter. Based on the masking effect of human hearing, in the time-frequency range where speech energy is high, most quantization noise has lower energy compared to speech and is difficult to detect. By weighting it using the amplitude spectrum of clean speech, the noise can be modified... Adjusting the focus on errors at speech spectrum peaks and valleys, when The smaller the value, the more the network training focuses on suppressing interference at spectral valleys. As the amplitude increases, network training focuses more on recovering information at speech spectrum peaks. Based on experimental verification, this invention selects... These are the optimal parameters.

[0056]

[0057] Step 3) Complex spectrum estimation of near-end speech:

[0058] The near-end speech amplitude spectrum estimation obtained in step 2) Phase with the microphone received signal The preliminary estimated complex spectrum of the near-end speech is calculated. The real and imaginary parts of this complex spectrum are then combined with the real and imaginary parts of the microphone received signal to form a 4-channel time-frequency feature, which is input into the third-stage convolutional recurrent network. The third-stage convolutional recurrent network has the same structure as the first and second-stage networks, including an encoding / decoding module and a timing modeling module. The decoding module consists of two parts: residual estimation of the real part and residual estimation of the imaginary part of the near-end speech complex spectrum. The output of the decoding module is then processed. , Compared with the preliminary estimates of the real and imaginary parts of the complex spectrum of proximal speech , Each corresponding expression is added together to obtain the final estimated complex spectrum of the near-end speech. , After the first and second stage neural networks are trained, the third stage neural network is trained, while the weight parameters of the first two stage neural networks are fine-tuned to achieve better echo cancellation and speech enhancement performance. The cost function of the third stage also incorporates the influence of human auditory characteristics, as shown below.

[0059]

[0060]

[0061] The clean speech used in the training set came from the Timit database. 120 speaker pairs were randomly selected from the Timit training set as near-end and far-end speakers (30 male-male pairs, 30 male-female pairs, 30 male-female pairs, and 30 female-male pairs). Each speaker's speech consisted of 10 sentences, each approximately 3 seconds long. Seven sentences were randomly selected to construct the neural network training set, and the remaining three sentences were used to construct the test set to evaluate the algorithm's performance. The mirror method was used to generate the RIR. Three near-end and far-end room sizes were set when constructing the training set: [4, 3, 3] m, [6, 4, 3] m, and [8, 7, 3] m. Room reverberation time... Set to [0.3, 0.6, 0.9] s. In the reverberant room, the simulated RIR length is set to... The loudspeaker and microphone positions in the two rooms are fixed. The loudspeaker positions are (1, 2, 1.2) m and (3, 2, 1.2) m, and the microphone positions are (1.8, 1, 1.2) m and (2.2, 1, 1.2) m, respectively. The distance between the speaker in the far room and the two microphones is set to [0.3, 0.7, 1.1] m. For each distance, a speaker position is selected every 20°, so the training set contains a total of 54 speaker positions in the far room. The SER in the training set is set to [0, 5, 10, 15] dB. The microphone received signals also include noise under different SNR conditions, with 115 noise types. There are a total of 37,800 near-end to far-end signals, approximately 95 hours in length. 80% of these signals are randomly selected for model training, and the remaining 20% ​​is used as a cross-validation set to determine if the model has converged.

[0062] In the multi-frame filter structure, the order of the multi-frame filter is set. , Figure 2 The performance of the proposed algorithm improves with increasing multi-frame filter order, and the background noise is devoid of babble noise. This experiment does not consider the impact of amplitude-phase decoupling and the human auditory cost function on the algorithm; therefore, this experiment only includes a two-stage convolutional recurrent neural network. The output of the second-stage convolutional recurrent neural network is the complex spectrum of near-end speech. Figure 2 (a) Figure 2 (b) Figure 2 (c) represents the improvement in PESQ score of near-end speech under different frame lengths of multi-frame filter settings when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s (the difference between the PESQ score of the microphone received signal and the actual improvement). Figure 2 (d) Figure 2 (e) Figure 2 (f) represents the improvement of the near-end speech ETOI score (the difference between the ETOI score of the microphone received signal) under different frame lengths of the multi-frame filter when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s. Figure 2 (g) Figure 2 (h) Figure 2 (i) Echo suppression (ERLE) values ​​for the multi-frame filter with different frame lengths when the near-room reverberation times are 0.3s, 0.6s, and 0.9s. The horizontal axis represents the signal-to-noise ratio (SNR). PESQ and STOI scores vary with filter frame length. The increase in the value indicates that the longer the frame length of the multi-frame filter, the better the near-end speech quality. Specifically, when the frame length... The advantage is more pronounced when the reverberation time is longer. This is because the frame length of the multi-frame filter represents the duration of the estimated room impulse response. Therefore, when the reverberation time is short, increasing the frame length can only improve the near-end speech quality to a limited extent. However, as the reverberation time increases, the longer the frame length, the more accurate the estimation of the room impulse response, and thus the better the near-end speech quality can be improved.

[0063] Figure 3 Under the same conditions, with different frame lengths for multiple filter frames, this section describes the algorithm's improvement in CSIG, CBAK, and COVL scores for near-end speech. CSIG represents the assessment of speech quality, CBAK represents the assessment of echo persistence, and COVL represents the overall assessment of both speech quality and echo persistence. Higher scores for all three categories indicate better algorithm performance. Figure 3 (a) Figure 3 (b) Figure 3 (c) represents the improvement of the near-end speech CSIG score (the difference between the CSIG score of the microphone received signal) under different frame lengths of the multi-frame filter when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s. Figure 3 (d) Figure 3 (e) Figure 3 (f) shows the improvement in the near-end speech CBAK score (the difference between the CBAK score and the microphone received signal) under different frame lengths of the multi-frame filter when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s. The horizontal axis in the figure represents the signal-to-noise ratio. Figure 3 (g) Figure 3 (h) Figure 3 (i) represents the improvement in COVL score of near-end speech (the difference between the COVL score and the microphone received signal) under different frame lengths of the multi-frame filter when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s, respectively. As shown in the figure, CSIG, CBAK, and COVL scores all increase with the increase of filter frame length.

[0064] Figure 4 The comparison results of the proposed algorithm with other algorithms in terms of PESQ, ESTOI, and ERLE scores are shown under the condition of a fixed filter order of 12. The background noise is excluding baby noise. The algorithms compared include a single-stage convolutional recurrent neural network (CRN) algorithm and a two-stage convolutional recurrent neural network (TS_CCRN) algorithm. In the figure, the proposed (No WE) algorithm is the result of training the proposed algorithm using the MSE cost function, and the proposed (WE) algorithm is the result of training the proposed algorithm using the human auditory correlation cost function. Figure 4 (a) Figure 4 (b) Figure 4 (c) shows the improvement of the PESQ score of the near-end speech by different algorithms when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s (the difference between the PESQ score of the microphone received signal and the improvement of the near-end speech PESQ score). Figure 4 (d) Figure 4 (e) Figure 4 (f) shows the improvement of the near-end speech ETOI score by different algorithms when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s (the difference between the ETOI score of the microphone received signal and the improvement of the near-end speech ETOI score). Figure 4 (g) Figure 4 (h) Figure 4 (i) Echo suppression ratio (ERLE) of different algorithms when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s, respectively. The horizontal axis in the figure represents the signal-to-noise ratio. It is clear from the figure that the algorithm proposed in this invention significantly improves the PESQ, ESTOI, and ERLE scores compared to the CRN and TS_CCRN algorithms, demonstrating a significant improvement in near-end speech quality and echo suppression. The auditory correlation cost function utilizes the auditory masking effect, significantly improving near-end speech quality and intelligibility. While its echo suppression is slightly lower than the MSE cost function, both achieve echo suppression levels above 70dB, far exceeding the echo suppression levels of the CRN and TS_CCRN algorithms.

[0065] Figure 5 The results show the comparison of the proposed algorithm with CRN, TS_CCRN, Proposed (No WE) and Proposed (WE) algorithms in terms of CSIG, CBAK and COVL scores under the condition of a fixed filter order of 12. The background noise is the absence of babble noise. Figure 5 In the table, (a), (b), and (c) show the improvement of the near-end speech CSIG score by different algorithms (the difference between the CSIG score and the microphone received signal CISG score) when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s, respectively. Figure 5In the diagram, (d), (e), and (f) represent the improvements in the near-end speech CBAK score (the difference between the CBAK score of the microphone received signal) by different algorithms when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s, respectively. Figure 5 In the figure, (g), (h), and (i) represent the improvement in near-end speech COVL score (the difference between the COVL score and the microphone received signal) of different algorithms when the near-end room reverberation time is 0.3s, 0.6s, and 0.9s, respectively. The horizontal axis in the figure represents the signal-to-noise ratio. It is clear from the figure that the algorithm proposed in this invention has a significant improvement over the CRN and TS_CCRN algorithms in terms of CSIG, CBAK, and COVL scores, which can significantly improve near-end speech quality and reduce echo reverberation. The human auditory correlation cost function utilizes the human auditory masking effect, which significantly further improves near-end speech quality and enhances the overall performance of near-end speech.

[0066] Another embodiment of the present invention provides a multi-frame filtering amplitude-phase decoupling deep neural network stereo echo cancellation system based on human hearing, comprising:

[0067] The linear echo cancellation module is used to estimate the real and imaginary parts of the linear filter using the microphone received signal and the complex spectrum of the two-channel reference signal of the stereo echo system, and to estimate the linear echo signal online in real time using a multi-frame filter structure to obtain the residual echo complex spectrum.

[0068] The near-end speech amplitude spectrum estimation module is used to estimate the near-end speech amplitude spectrum using the residual echo amplitude spectrum, the microphone received signal amplitude spectrum and the two-channel reference signal amplitude spectrum. The neural network for estimating the near-end speech amplitude spectrum is trained by the auditory perception related cost function.

[0069] The near-end speech complex spectrum estimation module is used to obtain a preliminary estimate of the near-end speech complex spectrum by using the estimated near-end speech amplitude spectrum and the phase of the microphone received signal. This estimate, along with the microphone received signal complex spectrum, is then input into a neural network to perform a secondary estimation of the near-end speech complex spectrum. The neural network is then trained using the human auditory correlation cost function to improve speech quality.

[0070] The stereo echo cancellation module is used to input the microphone received signal and two-channel reference signals into a trained neural network to perform stereo echo cancellation.

[0071] For the specific implementation process of each module, please refer to the description of the method of the present invention above.

[0072] Another embodiment of the present invention provides a computer device (computer, server, smartphone, etc.) including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the steps of the method of the present invention.

[0073] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk) that stores a computer program, which, when executed by a computer, implements the various steps of the method of the present invention.

[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A stereo echo cancellation method based on human hearing, using multi-frame filtering amplitude and phase decoupling deep neural networks, characterized in that... Includes the following steps: 1) The real and imaginary parts of the linear filter are estimated using the microphone received signal and the complex spectrum of the two-channel reference signal of the stereo echo system, and the linear echo signal is estimated online in real time using a multi-frame filter structure to obtain the residual echo complex spectrum. 2) The near-end speech amplitude spectrum is estimated using the residual echo amplitude spectrum, the microphone received signal amplitude spectrum, and the two-channel reference signal amplitude spectrum. The neural network for estimating the near-end speech amplitude spectrum is trained using an auditory perception correlation cost function, as shown below: in It is the estimated near-end speech amplitude spectrum. It is the true near-end speech amplitude spectrum. It is the power exponent parameter. ; 3) A preliminary estimate of the complex spectrum of the near-end speech is obtained by using the estimated amplitude spectrum of the near-end speech and the phase of the microphone received signal. This estimate, along with the complex spectrum of the microphone received signal, is input into a neural network to perform a secondary estimation of the complex spectrum of the near-end speech. The neural network is then trained using a human auditory correlation cost function to improve speech quality. The human auditory correlation cost function is: in , It is the real and imaginary parts of the complex spectrum of near-end speech estimated in a quadratic manner. and These are the real and imaginary parts of the complex spectrum of actual near-end speech. It is the complex spectrum of actual near-end speech. It is the estimated complex spectrum of near-terminal speech. It is the power exponent parameter. ; 4) Input the microphone received signal and the two-channel reference signal into the trained neural network to perform stereo echo cancellation.

2. The method according to claim 1, characterized in that, Step 1) includes: 1-1) Perform a short-time Fourier transform on the microphone received signal and the two-channel reference signal to obtain the real and imaginary parts of its complex spectrum; 1-2) Input the real and imaginary parts of the complex spectrum of the microphone received signal and the two-channel reference signal obtained in step 1-1) into the convolutional recurrent neural network. The decoding module includes four parts, namely the real and imaginary parts of the complex spectrum of the impulse response of the two rooms in the near room. 1-3) The real and imaginary parts of the complex spectrum of the room impulse response output by the convolutional recurrent neural network in step 1-2) are calculated with the real and imaginary parts of the microphone received signal through a set multi-frame filtering method to obtain the estimated linear echo signal. 1-4) Subtract the real and imaginary parts of the linear echo complex spectrum estimated in step 1-3) from the real and imaginary parts of the complex spectrum of the received signal from the microphone to obtain the real and imaginary parts of the residual echo complex spectrum.

3. The method according to claim 2, characterized in that, The calculation formula for the multi-frame filtering method is as follows: in , , , These are the real and imaginary parts of the complex spectrum of the room's impulse response. , , , These are the real and imaginary parts of the complex spectrum of the two-channel reference signal. and These are the estimated real and imaginary parts of the linear echo complex spectrum. It is a multi-frame filtering function. It is a frame index. It's a frequency point. These are the indices of the time-frequency points used in the calculation of the multi-frame filtering function. L It is the total frame length used in the calculation of the multi-frame filtering function.

4. The method according to claim 3, characterized in that, Step 2) includes: 2-1) Extract the amplitude spectrum features of the residual echo complex spectrum, the microphone received signal complex spectrum, and the reference signal complex spectrum, input them into a convolutional recurrent neural network, and estimate the near-end speech amplitude spectrum; 2-2) The neural networks in steps 2-1) and 1-2) are trained using the auditory perception-related cost function.

5. The method according to claim 1, characterized in that, Step 3) includes: 3-1): Combine the near-end speech amplitude spectrum estimated in step 2) with the phase of the microphone received signal to obtain the preliminary estimated real and imaginary parts of the near-end speech complex spectrum; 3-2): The real and imaginary parts of the complex spectrum of the near-end speech obtained in step 3-1) are input together with the real and imaginary parts of the complex spectrum of the microphone received signal into a convolutional recurrent neural network. The decoding module includes two parts: re-estimation of the real part of the complex spectrum of the near-end speech and re-estimation of the imaginary part of the complex spectrum. 3-4): After the neural network training in step 2-2) is completed, the convolutional recurrent neural network in step 3-2) is trained using the human auditory cost function, and the weight parameters of the neural network in step 2-2) are fine-tuned.

6. A stereo echo cancellation system based on human hearing, employing the method described in any one of claims 1 to 5, comprising multi-frame filtering amplitude-phase decoupling deep neural network, characterized in that, include: The linear echo cancellation module is used to estimate the real and imaginary parts of the linear filter using the microphone received signal and the complex spectrum of the two-channel reference signal of the stereo echo system, and to estimate the linear echo signal online in real time using a multi-frame filter structure to obtain the residual echo complex spectrum. The near-end speech amplitude spectrum estimation module is used to estimate the near-end speech amplitude spectrum using the residual echo amplitude spectrum, the microphone received signal amplitude spectrum and the two-channel reference signal amplitude spectrum. The neural network for estimating the near-end speech amplitude spectrum is trained by the auditory perception related cost function. The near-end speech complex spectrum estimation module is used to obtain a preliminary estimate of the near-end speech complex spectrum by using the estimated near-end speech amplitude spectrum and the phase of the microphone received signal. This estimate, along with the microphone received signal complex spectrum, is then input into a neural network to perform a secondary estimation of the near-end speech complex spectrum. The neural network is then trained using the human auditory correlation cost function to improve speech quality. The stereo echo cancellation module is used to input the microphone received signal and two-channel reference signals into a trained neural network to perform stereo echo cancellation.

7. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Stereo echo cancellation method and system based on neural network

    CN111292759A

  • Method and device for eliminating echo, electronic equipment and storage medium

    CN113192527A