A method and apparatus for echo cancellation

By using an end-to-end echo cancellation model based on target knowledge distillation, the performance and robustness issues of linear echo cancellation algorithms under extremely low signal-to-return ratios and variable echo propagation paths are solved, enabling effective deployment on low-computing-power devices and improving voice interaction performance.

CN119091901BActive Publication Date: 2025-11-18ZHEJIANG FUTURE ELF ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411307928.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-11-18
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

Existing linear echo cancellation algorithms suffer from poor voice interaction performance and robustness under conditions of extremely low signal-to-return ratio, unstable latency, and variable echo propagation paths, making them difficult to deploy effectively on low-computing-power devices.

Method used

An end-to-end echo cancellation model based on target knowledge distillation is adopted. Through preprocessing and feature extraction, combined with training using the joint loss function of teacher and student branches, an end-to-end echo cancellation model based on knowledge distillation is generated, and the model is compressed to achieve echo cancellation.

Benefits of technology

It improves the voice interaction performance and robustness of terminal devices under conditions of extremely low signal-to-return ratio, unstable latency, and variable echo propagation paths, and can be effectively deployed on low-computing-power devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091901B_ABST
    Figure CN119091901B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an echo cancellation method and device. In the embodiments of the present application, a first near-end speech and a first far-end speech are acquired; the first near-end speech and the first far-end speech are preprocessed to generate a second near-end speech and a second far-end speech after preprocessing; the second near-end speech and the second far-end speech are respectively input into a knowledge distillation end-to-end echo cancellation model to be trained to determine a joint loss function; the knowledge distillation end-to-end echo cancellation model to be trained is adjusted according to the joint loss function to generate a knowledge distillation end-to-end echo cancellation model; and the knowledge distillation end-to-end echo cancellation model is compressed to generate a target knowledge distillation end-to-end echo cancellation model, wherein the target knowledge distillation end-to-end echo cancellation model is used to implement echo cancellation. Through the above method, acoustic echo cancellation can be better implemented under a very low signal-to-echo ratio, and the speech interaction performance and robustness of a terminal device are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and more specifically, to a method and apparatus for echo cancellation. Background Technology

[0002] Acoustic echo cancellation (AEC) has wide applications in conferencing systems, in-vehicle voice platforms, and smart speaker interaction. Traditional linear echo cancellation algorithms are highly robust when the acoustic echo path is linear and time-invariant, and exhibit low nonlinear distortion, resulting in low sound quality loss. However, when there are large time-domain delay fluctuations in far-end and near-end speech, significant nonlinear distortion of the speaker, variable acoustic echo propagation paths, and significant noise interference, the performance of these traditional linear echo cancellation algorithms deteriorates.

[0003] In existing technologies, two new schemes have been proposed to replace common linear echo cancellation algorithms. Scheme 1 combines signal processing and neural networks. Specifically, linear echo suppression is first performed during signal processing, and then nonlinear echo suppression, noise reduction, dereverberation, and automatic gain control are performed through neural networks. Scheme 2 establishes an end-to-end multi-task learning network architecture such as DeepVQE to simultaneously perform acoustic echo cancellation, noise suppression, and automatic gain control, featuring low latency, high operability, and simple link. However, when the signal-to-echo ratio reaches below -30dB, the voice interaction performance and robustness of both schemes are poor.

[0004] In summary, how to achieve better acoustic echo cancellation and improve the voice interaction performance and robustness of terminal devices under conditions of extremely low signal-to-return ratio, unstable latency, and variable echo propagation paths is a problem that needs to be solved. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method and apparatus for echo cancellation, which adopts an end-to-end echo cancellation model based on target knowledge distillation. Under conditions of extremely low signal-to-echo ratio, unstable delay, and variable echo propagation path, it can achieve better acoustic echo cancellation and improve the voice interaction performance and robustness of terminal devices.

[0006] In a first aspect, embodiments of the present invention provide an echo cancellation method, the method comprising:

[0007] Acquire the first near-end speech and the first far-end speech;

[0008] The first near-end speech and the first far-end speech are preprocessed to generate the preprocessed second near-end speech and the second far-end speech.

[0009] The second near-end speech and the second far-end speech are respectively input into the knowledge distillation end-to-end echo cancellation model to be trained, and the joint loss function is determined.

[0010] The knowledge distillation end-to-end echo cancellation model to be trained is adjusted according to the joint loss function to generate a knowledge distillation end-to-end echo cancellation model.

[0011] The knowledge distillation end-to-end echo cancellation model is compressed to generate a target knowledge distillation end-to-end echo cancellation model, wherein the target knowledge distillation end-to-end echo cancellation model is used to achieve echo cancellation.

[0012] Optionally, the method further includes:

[0013] The target knowledge distillation end-to-end echo cancellation model is deployed to the terminal side.

[0014] Optionally, the step of preprocessing the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech specifically includes:

[0015] The first near-end speech and the first far-end speech are pre-emphasized, power frequency interference is removed, windowing is applied, and short-time Fourier transform is performed to generate the second near-end speech and the second far-end speech.

[0016] Optionally, the step of inputting the second near-end speech and the second far-end speech into the knowledge distillation end-to-end echo cancellation model to be trained to determine the joint loss function specifically includes:

[0017] The second near-end speech and the second far-end speech are respectively input into the teacher branch and student branch of the knowledge distillation end-to-end echo cancellation model to be trained, and the training loss functions of the teacher branch and student branch are obtained respectively.

[0018] The joint loss function is determined based on the training loss function of the teacher branch and the training loss function of the student branch.

[0019] Optionally, the step of inputting the second near-end speech and the second far-end speech into the teacher branch of the knowledge distillation end-to-end echo cancellation model to be trained, and obtaining the training loss function of the teacher branch, specifically includes:

[0020] Feature extraction is performed on the second proximal speech and the second distal speech to obtain the features of the second proximal speech and the features of the second distal speech;

[0021] Based on the features of the second near-end speech and the features of the second far-end speech, the first delay loss and the first speech enhancement loss of the teacher branch are determined;

[0022] The training loss function for the teacher branch is determined based on the first delay loss and the first speech enhancement loss.

[0023] Optionally, determining the first delay loss of the teacher branch based on the features of the second near-end speech and the features of the second far-end speech specifically includes:

[0024] The features of the second near-end speech and the features of the second far-end speech are input into the first time delay detection module of the teacher branch to generate the first predicted near-end speech signal of the first time delay detection module.

[0025] The first delay loss of the teacher branch is determined based on the first predicted near-end speech signal.

[0026] Optionally, determining the first speech enhancement loss of the teacher branch based on the features of the second proximal speech and the features of the second distal speech specifically includes:

[0027] The features of the second near-end speech and the features of the second far-end speech are input into the first time delay detection module of the teacher branch to generate the first predicted near-end speech signal.

[0028] The first predicted near-end speech signal is input into the coding layer, bottleneck layer, decoding layer, complex spectrum convolutional mapping layer and data frame filter to generate the first predicted target speech signal;

[0029] The first speech enhancement loss of the teacher branch is determined based on the first predicted target speech signal.

[0030] Optionally, determining the training loss function for the teacher branch based on the first delay loss and the first speech enhancement loss specifically includes:

[0031] Determine the first product of the first delay loss and the first weight, and the second product of the first speech enhancement loss and the second weight;

[0032] The sum of the first product and the second product is determined as the training loss function of the teacher branch.

[0033] Optionally, the step of inputting the second near-end speech and the second far-end speech into the student branch of the knowledge distillation end-to-end echo cancellation model to be trained, and obtaining the training loss function of the student branch, specifically includes:

[0034] Feature extraction is performed on the second proximal speech and the second distal speech to obtain the features of the second proximal speech and the features of the second distal speech;

[0035] Based on the features of the second near-end speech and the features of the second far-end speech, the second delay loss and the second speech enhancement loss of the student branch are determined;

[0036] The training loss function for the student branch is determined based on the second delay loss and the second speech enhancement loss.

[0037] Optionally, determining the second delay loss of the student branch based on the features of the second near-end speech and the features of the second far-end speech specifically includes:

[0038] The features of the second near-end speech and the features of the second far-end speech are input into the second time delay detection module of the student branch to generate the second predicted near-end speech signal.

[0039] The second delay loss of the student branch is determined based on the second predicted near-end speech signal.

[0040] Optionally, determining the first speech enhancement loss of the student branch based on the features of the second proximal speech and the features of the second distal speech specifically includes:

[0041] The features of the second near-end speech and the features of the second far-end speech are input into the second time delay detection module of the student branch to generate the second predicted near-end speech signal of the second time delay detection module.

[0042] The second predicted near-end speech signal is input into the coding layer, bottleneck layer, decoding layer, complex spectrum convolutional mapping layer and data frame filter to generate the second predicted target speech signal;

[0043] The second speech enhancement loss of the student branch is determined based on the second predicted target speech signal.

[0044] Optionally, determining the training loss function for the student branch based on the second delay loss and the second speech enhancement loss specifically includes:

[0045] Determine the third product of the second delay loss and the third weight, and the fourth product of the second speech enhancement loss and the fourth weight;

[0046] The sum of the third product and the fourth product is determined as the training loss function of the student branch.

[0047] Optionally, determining the joint loss function based on the training loss function of the teacher branch and the training loss function of the student branch specifically includes:

[0048] Determine the product of the training loss function of the student branch and a set parameter, wherein the set parameter represents the KL divergence between the teacher branch and the student branch;

[0049] The sum of the product and the training loss function of the teacher branch is determined as the joint loss function.

[0050] Optionally, the step of compressing the knowledge distillation end-to-end echo cancellation model to generate the target knowledge distillation end-to-end echo cancellation model specifically includes:

[0051] The knowledge distillation end-to-end echo cancellation model is quantized online to generate a knowledge distillation end-to-end echo cancellation model with a Torch model structure.

[0052] The knowledge distillation end-to-end echo cancellation model of the Torch model structure is transformed into ONNX through an open neural network to generate a knowledge distillation end-to-end echo cancellation model of the ONNX model structure.

[0053] The knowledge distillation end-to-end echo cancellation model of the ONNX model structure is passed through the MNN inference engine to generate the target knowledge distillation end-to-end echo cancellation model.

[0054] Secondly, embodiments of the present invention provide an echo cancellation method, the method comprising:

[0055] Acquire the signal to be processed;

[0056] The signal to be processed is input into the target knowledge distillation end-to-end echo cancellation model to generate the target signal, wherein the target knowledge distillation end-to-end echo cancellation model is deployed on the terminal side to achieve echo cancellation.

[0057] Thirdly, embodiments of the present invention provide an echo cancellation device, the device comprising:

[0058] An acquisition unit is used to acquire the first near-end speech and the first far-end speech;

[0059] The processing unit is configured to preprocess the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech.

[0060] The determination unit is used to input the second proximal speech and the second far-end speech into the knowledge distillation end-to-end echo cancellation model to be trained, and determine the joint loss function.

[0061] The generation unit is used to adjust the knowledge distillation end-to-end echo cancellation model to be trained according to the joint loss function, and generate the knowledge distillation end-to-end echo cancellation model.

[0062] The generation unit is further configured to compress the knowledge distillation end-to-end echo cancellation model to generate a target knowledge distillation end-to-end echo cancellation model, wherein the target knowledge distillation end-to-end echo cancellation model is used to achieve echo cancellation.

[0063] Optionally, the device further includes:

[0064] The deployment unit is used to deploy the target knowledge distillation end-to-end echo cancellation model to the terminal side.

[0065] Optionally, the processing unit is specifically used for:

[0066] The first near-end speech and the first far-end speech are pre-emphasized, power frequency interference is removed, windowing is applied, and short-time Fourier transform is performed to generate the second near-end speech and the second far-end speech.

[0067] Optionally, the determining unit is specifically used for:

[0068] The second near-end speech and the second far-end speech are respectively input into the teacher branch and student branch of the knowledge distillation end-to-end echo cancellation model to be trained, and the training loss functions of the teacher branch and student branch are obtained respectively.

[0069] The joint loss function is determined based on the training loss function of the teacher branch and the training loss function of the student branch.

[0070] Optionally, the determining unit is further configured to:

[0071] Feature extraction is performed on the second proximal speech and the second distal speech to obtain the features of the second proximal speech and the features of the second distal speech;

[0072] Based on the features of the second near-end speech and the features of the second far-end speech, the first delay loss and the first speech enhancement loss of the teacher branch are determined;

[0073] The training loss function for the teacher branch is determined based on the first delay loss and the first speech enhancement loss.

[0074] Optionally, the determining unit is further configured to: input the features of the second proximal speech and the features of the second distal speech into the first time delay detection module of the teacher branch to generate the first predicted proximal speech signal of the first time delay detection module;

[0075] The first delay loss of the teacher branch is determined based on the first predicted near-end speech signal.

[0076] Optionally, the determining unit is further configured to: input the features of the second near-end speech and the features of the second far-end speech into the first time delay detection module of the teacher branch to generate a first predicted near-end speech signal;

[0077] The first predicted near-end speech signal is input into the coding layer, bottleneck layer, decoding layer, complex spectrum convolutional mapping layer and data frame filter to generate the first predicted target speech signal;

[0078] The first speech enhancement loss of the teacher branch is determined based on the first predicted target speech signal.

[0079] Optionally, the determining unit is further configured to: determine a first product of the first delay loss and the first weight, and a second product of the first speech enhancement loss and the second weight;

[0080] The sum of the first product and the second product is determined as the training loss function of the teacher branch.

[0081] Optionally, the determining unit is further configured to: extract features from the second proximal speech and the second distal speech to obtain features of the second proximal speech and features of the second distal speech;

[0082] Based on the features of the second near-end speech and the features of the second far-end speech, the second delay loss and the second speech enhancement loss of the student branch are determined;

[0083] The training loss function for the student branch is determined based on the second delay loss and the second speech enhancement loss.

[0084] Optionally, the determining unit is further configured to: input the features of the second near-end speech and the features of the second far-end speech into the second time delay detection module of the student branch to generate a second predicted near-end speech signal;

[0085] The second delay loss of the student branch is determined based on the second predicted near-end speech signal.

[0086] Optionally, the determining unit is further configured to: input the features of the second near-end speech and the features of the second far-end speech into the second time delay detection module of the student branch to generate the second predicted near-end speech signal of the second time delay detection module;

[0087] The second predicted near-end speech signal is input into the coding layer, bottleneck layer, decoding layer, complex spectrum convolutional mapping layer and data frame filter to generate the second predicted target speech signal;

[0088] The second speech enhancement loss of the student branch is determined based on the second predicted target speech signal.

[0089] Optionally, the determining unit is further configured to: determine the third product of the second delay loss and the third weight, and the fourth product of the second speech enhancement loss and the fourth weight;

[0090] The sum of the third product and the fourth product is determined as the training loss function of the student branch.

[0091] Optionally, the determining unit is further configured to: determine the product of the training loss function of the student branch and a set parameter, wherein the set parameter represents the KL divergence between the teacher branch and the student branch;

[0092] The sum of the product and the training loss function of the teacher branch is determined as the joint loss function.

[0093] Optionally, the generation unit is specifically used for:

[0094] The knowledge distillation end-to-end echo cancellation model is quantized online to generate a knowledge distillation end-to-end echo cancellation model with a Torch model structure.

[0095] The knowledge distillation end-to-end echo cancellation model of the Torch model structure is transformed into ONNX through an open neural network to generate a knowledge distillation end-to-end echo cancellation model of the ONNX model structure.

[0096] The knowledge distillation end-to-end echo cancellation model of the ONNX model structure is passed through the MNN inference engine to generate the target knowledge distillation end-to-end echo cancellation model.

[0097] Optionally, the acquisition unit is further configured to: acquire the signal to be processed;

[0098] The generation unit is further configured to: input the signal to be processed into the target knowledge distillation end-to-end echo cancellation model to generate the target signal.

[0099] Fourthly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect, any possible method of the first aspect, or any one of the second aspects.

[0100] Fifthly, embodiments of the present invention provide a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the method as described in the first aspect, any possible form of the first aspect, or any of the second aspects.

[0101] In this embodiment of the invention, a first near-end speech and a first far-end speech are acquired; the first near-end speech and the first far-end speech are preprocessed to generate preprocessed second near-end speech and second far-end speech; the second near-end speech and the second far-end speech are respectively input into a knowledge distillation end-to-end echo cancellation model to be trained, and a joint loss function is determined; the knowledge distillation end-to-end echo cancellation model to be trained is adjusted according to the joint loss function to generate a knowledge distillation end-to-end echo cancellation model; the knowledge distillation end-to-end echo cancellation model is compressed to generate a target knowledge distillation end-to-end echo cancellation model, wherein the target knowledge distillation end-to-end echo cancellation model is used to achieve echo cancellation. Through the above method, a target knowledge distillation end-to-end echo cancellation model is generated. This target knowledge distillation end-to-end echo cancellation model can achieve better acoustic echo cancellation under conditions of extremely low signal-to-return ratio, unstable latency, and variable echo propagation paths, thereby improving the voice interaction performance and robustness of terminal devices. Attached Figure Description

[0102] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:

[0103] Figure 1 This is a flowchart of an echo cancellation method according to an embodiment of the present invention;

[0104] Figure 2 This is a schematic diagram of an end-to-end echo cancellation model for knowledge distillation to be trained in an embodiment of the present invention;

[0105] Figure 3 This is a flowchart of a method for obtaining the training loss function of the teacher branch in an embodiment of the present invention;

[0106] Figure 4 This is a schematic diagram of the structure of a teacher branch in an embodiment of the present invention;

[0107] Figure 5 This is a schematic diagram of the structure of a training loss function for generating teacher branches in an embodiment of the present invention;

[0108] Figure 6 This is a flowchart of a method for obtaining the training loss function of the student branch in an embodiment of the present invention;

[0109] Figure 7 This is a schematic diagram of the structure of a student branch in an embodiment of the present invention;

[0110] Figure 8 This is a schematic diagram of the structure of a training loss function for generating student branches in an embodiment of the present invention;

[0111] Figure 9 This is a schematic diagram of the structure of a first TDE module in an embodiment of the present invention;

[0112] Figure 10 This is a flowchart of a model compression method according to an embodiment of the present invention;

[0113] Figure 11 This is a flowchart of another echo cancellation method in an embodiment of the present invention;

[0114] Figure 12 This is a flowchart of another echo cancellation method in an embodiment of the present invention;

[0115] Figure 13 This is a schematic diagram of a system structure in an embodiment of the present invention;

[0116] Figure 14 This is a schematic diagram of an echo cancellation device according to an embodiment of the present invention;

[0117] Figure 15 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0118] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.

[0119] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0120] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".

[0121] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0122] To address the poor performance of traditional linear acoustic echo cancellation (LAEC) and neural network acoustic echo cancellation (NNAEC) algorithms in scenarios with large delay fluctuations, significant nonlinear distortion of loudspeakers, variable acoustic echo propagation paths, and significant noise interference, two new schemes are proposed to replace common linear echo cancellation algorithms, as detailed below:

[0123] Option 1 combines signal processing and neural networks. Specifically, linear echo suppression is first performed during signal processing, and then nonlinear echo suppression, noise reduction, dereverberation, and automatic gain control are performed through neural networks. Examples include gated convolutional FT-LSTM neural networks (GFTNN), band-split recurrent neural networks (BSRNN), and neural Kalman filtering (NKF). GFTNN and BSRNN can also be referred to as an end-to-end model integrating echo cancellation, noise suppression, and dereverberation.

[0124] Option 2: Establish an end-to-end multi-task learning network architecture to simultaneously perform functions such as acoustic echo cancellation (AEC), noise suppression (NS), and automatic gain control (AGC). This approach features low latency, high operability, and a simple link structure. Examples include the Deep Variational Quantum Eigensolver (DEEPVQE) algorithm and a network for repairing and denoising speech signals (RadNET).

[0125] However, in smart speakers, large conference screens, and in-vehicle voice interaction, situations often involve playing music and videos at high volumes, as well as text-to-speech (TTS) operations, causing the signal-to-return ratio (SRR) to drop below -30dB. Under these conditions, the voice interaction performance of both solutions is poor. Furthermore, the aforementioned solutions align far-end and near-end speech by increasing the number of delay taps in the frequency domain block adaptive filter, increasing computational complexity and causing divergence when latency jitter occurs, resulting in poor robustness. Moreover, the model parameters of these solutions are enormous, making them unsuitable for deployment on low-computing-power devices. Therefore, how to effectively achieve acoustic echo cancellation in full-duplex environments with extremely low SRR, unstable latency, and variable echo propagation paths, thereby improving the voice interaction performance and robustness of terminal devices, while also being deployable on low-computing-power devices, is a problem that needs to be solved.

[0126] In this embodiment of the invention, the signal-to-return ratio refers to the ratio of the power of the near-end speech signal received from the terminal side to the power of the echo signal received from the speaker, wherein the echo signal is generated by the voice of the user on the opposite side of the terminal side being played through the speaker on the terminal side; full-duplex is a communication method in which data can be transmitted simultaneously in two directions, that is, both parties can send and receive data at the same time, and the transmission in these two directions does not interfere with each other. For example, real-time communication (RTC) and live streaming scenarios are full-duplex communication.

[0127] In this embodiment of the invention, to solve the above problems, an echo cancellation method is proposed, specifically as follows: Figure 1 As shown, the method includes:

[0128] Step S101: Obtain the first near-end speech and the first far-end speech.

[0129] Specifically, both the first near-end speech and the first far-end speech are training data. Before acquiring the first near-end speech and the first far-end speech, a simulation dataset needs to be constructed, and then the first near-end speech and the first far-end speech are acquired from the simulation dataset.

[0130] In this embodiment of the invention, the near-end speech and the far-end speech are affected by factors such as asynchronous sampling clocks, network latency instability, and audio effect algorithm processing. The near-end speech and the far-end speech are not strictly aligned in the time dimension and have a certain degree of fluctuation. The audio effect algorithm processing refers to Dolby sound effects, surround sound effects, etc.

[0131] In one possible implementation, when constructing the simulation dataset, data acquisition is first performed. The acquired data can come from public datasets such as Librispeech, Aishell, AEC-CHALLENGE and recorded echo data, OPENSIL and 100,000 simulated room impact responses, as well as DNS-CHALLENGE and recorded data. The public datasets Librispeech and Aishell are used as clean speech data; the AEC-CHALLENGE and recorded echo data are used as far-end speech data for synthesis; the OPENSIL and 100,000 simulated room impact responses are used as transfer functions; and the noise data comes from DNS-CHALLENGE and recorded data. The acquired data is used to synthesize 500 hours of training data with a signal-to-return ratio (SRR) range of [-35, 15] and a signal-to-noise ratio (SNR) range of [-5, 15] dB. The number of simulated room impact responses, the SRR range, the SNR range, and the training data duration are all determined based on actual conditions and are only illustrative examples here.

[0132] In this embodiment of the invention, a target near-end speech and a clean speech are acquired simultaneously with the first near-end speech and the first far-end speech. The target near-end speech and the clean speech are used to calculate the joint loss function of the knowledge distillation end-to-end echo cancellation model to be trained.

[0133] Step S102: Preprocess the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech.

[0134] Specifically, the first near-end speech and the first far-end speech are pre-emphasized, power frequency interference is removed, windowing is applied, and short-time Fourier transform (STFT) is performed to generate the second near-end speech and the second far-end speech.

[0135] In one possible implementation, the parameters used in the short-time Fourier transform are shown in Table 1, as detailed below:

[0136] Table 1

[0137] parameter numerical values Sampling rate 16000 Window length 32ms (512 points) Frame shift 8ms (128 points) Fast Fourier Transform (FFT) Length 512 points

[0138] The window function used in the short-time Fourier transform is the Hanning window.

[0139] Step S103: Input the second near-end speech and the second far-end speech into the knowledge distillation end-to-end echo cancellation model to be trained, and determine the joint loss function.

[0140] In this embodiment of the invention, the end-to-end echo cancellation model to be trained by knowledge distillation includes a teacher model and a student model, which can also be referred to as a teacher branch and a student branch. Knowledge distillation is a machine learning technique used to transfer knowledge from a complex, well-trained teacher model to a smaller, simpler student model. Knowledge distillation can also be referred to as teacher-student learning (TS-Learning).

[0141] In one possible implementation, the step of inputting the second proximal speech and the second distal speech into the knowledge distillation end-to-end echo cancellation model to be trained to determine the joint loss function specifically includes: inputting the second proximal speech and the second distal speech into the teacher branch and student branch of the knowledge distillation end-to-end echo cancellation model to be trained, respectively, and obtaining the training loss functions of the teacher branch and the student branch respectively; determining the joint loss function based on the training loss function of the teacher branch and the training loss function of the student branch, and performing joint training; specifically as follows... Figure 2 As shown, the Figure 2 The diagram illustrates the knowledge distillation end-to-end echo cancellation model to be trained, including a teacher branch and a student branch. The first near-end speech and the first far-end speech are preprocessed, including pre-emphasis, removal of power line interference, windowing, and STFT. The preprocessed second near-end speech and second far-end speech are then input into the teacher branch and student branch of the knowledge distillation end-to-end echo cancellation model to be trained, respectively. The outputs from the student branch and the teacher branch are then subjected to Inverse Short-Time Fourier Transform (ISTFT) to generate the training loss function L for the teacher branch. dt and the training loss function L of the student branch kd The joint loss function is determined based on the training loss function of the teacher branch and the training loss function of the student branch.

[0142] The following two specific examples illustrate in detail the methods for obtaining the training loss functions for the teacher branch and the student branch: Specific Implementation Example 1

[0144] The second near-end speech and the second far-end speech are respectively input into the teacher branch of the knowledge distillation end-to-end echo cancellation model to be trained, and the training loss function of the teacher branch is obtained, such as... Figure 3 As shown, it specifically includes:

[0145] Step S301: Extract features from the second proximal speech and the second distal speech to obtain the features of the second proximal speech and the features of the second distal speech.

[0146] In one possible implementation, features of the second near-end speech and the second far-end speech are extracted using an Equivalent Rectangular Bandwidth (ERB) filter to obtain features of the second near-end speech and the second far-end speech.

[0147] Step S302: Determine the first delay loss and the first speech enhancement loss of the teacher branch based on the features of the second near-end speech and the features of the second far-end speech.

[0148] Specifically, the features of the second near-end speech and the features of the second far-end speech are input into the first time delay estimation (TDE) module of the teacher branch to generate a first predicted near-end speech signal; the first delay loss of the teacher branch is determined based on the first predicted near-end speech signal. Simultaneously, the generated first predicted near-end speech signal is input into an encoder layer, a bottleneck layer, a decoder layer, a complex convolving mask block (CCM) layer, and a data frame filter (DF filter) to generate a first predicted target speech signal; the first speech enhancement loss of the teacher branch is determined based on the first predicted target speech signal.

[0149] In one possible implementation, a first delay loss is determined based on the first predicted near-end speech signal and the pre-acquired target near-end speech; simultaneously, a first speech enhancement loss is determined based on the first predicted target speech signal and the pre-acquired clean speech.

[0150] In this embodiment of the invention, the structural diagram of the teacher branch is as follows: Figure 4As shown, assuming that the second near-end speech (Mic) and the second far-end speech (Ref) are input to the first feature extraction module 401 of the teacher branch, and then sequentially input to the first TDE module 402, the first encoding layer 403, the first bottleneck layer 404, the first decoding layer 405, the first complex spectrum convolutional mapping (CCM) layer 406, and the first data frame filter 407, wherein the output of the first TDE module is transformed by inverse short-time Fourier transform to generate the first predicted near-end speech signal; the output of the first data frame filter is transformed by inverse short-time Fourier transform to generate the first predicted target speech signal; wherein the first encoding layer can also be called the first encoding module, the first bottleneck layer can also be called the first bottleneck module, the first decoding layer can also be called the first decoding module, and the first complex spectrum convolutional mapping (CCM) layer can also be called the first complex spectrum convolutional mapping (CCM) module.

[0151] In one possible implementation, the network structures of the first encoding layer and the first decoding layer are based on open-source network structures such as Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement (DCCRN) and DEEPVQE. The first encoding layer consists of four small encoding modules, each comprising a Conv2d convolutional layer, a batch normalization layer, an exponential linear unit (ELU) unit, and grouped temporal convolution blocks (GT-Conv blocks). The first CCM layer consists of two stages: convolving the output features of the first decoding layer with the input second near-end speech mic, and dividing the convolution output into three weight components to construct a complex mask, where each component is the weight of a 120° rotation vector in the complex plane, resulting in the enhanced complex domain mask features; finally, the output of the first CCM layer is passed through a first data frame filter (DF). The filter performs deep convolution filtering, and then the output of the first data frame filter is subjected to inverse short-time Fourier transform (ISTFT) to generate the first predicted target speech signal.

[0152] In this embodiment of the invention, a schematic diagram of determining the training loss function of the teacher branch is shown below. Figure 5 As mentioned above, in Figure 4Based on this, two Inverse Short Time Fourier Transform (ISTFT) functions are added, along with a diagram illustrating the determination of a first delay loss (i.e., the first TDE Loss) based on the first predicted near-end speech signal (Est Mic) and the target near-end speech (Target Mic), a first speech enhancement loss (i.e., the first Se Loss) based on the first predicted target speech signal (Est Speech) and the pre-acquired clean speech, and the determination of the training loss function of the teacher branch based on the aforementioned first delay loss and first speech enhancement loss.

[0153] In one possible implementation, assume that the clean speech is s(n), and the corresponding frequency domain signal is S(k,f); the target near-end speech (Target Mic) of the first TDE module is m(n), and the corresponding frequency domain signal is M(k,f); and the first predicted near-end speech signal (Est MIC) of the first TDE module is... The corresponding frequency domain signal is The first predicted target speech signal (EST SPEECH) predicted by the teacher branch is The corresponding frequency domain signal is Using commonly used speech denoising and separation domain loss functions such as mean square error (MSE) and scale-invariant source-to-noise ratio (SISNR), the first delay loss L1 and the first speech enhancement loss L2 are determined as follows:

[0154]

[0155] Wherein, α can be set to 0.5.

[0156] Step S303: Determine the training loss function for the teacher branch based on the first delay loss and the first speech enhancement loss.

[0157] Specifically, the first product of the first delay loss L1 and the first weight, and the second product of the first speech enhancement loss L2 and the second weight are determined; the sum of the first product and the second product is determined as the training loss function L of the teacher branch. dt .

[0158] Assuming the first weight is 0.3, the second weight is 0.7, and the training loss function L for the teacher branch... dt The formula is as follows:

[0159] L dt =0.3*L1+0.7*L2 Specific Implementation Example 2

[0161] The second near-end speech and the second far-end speech are respectively input into the student branch of the knowledge distillation end-to-end echo cancellation model to be trained, and the training loss function of the student branch is obtained, such as... Figure 6 As shown, it specifically includes:

[0162] Step S601: Extract features from the second proximal speech and the second distal speech to obtain the features of the second proximal speech and the features of the second distal speech.

[0163] In one possible implementation, features of the second near-end speech and the second far-end speech are extracted using an Equivalent Rectangular Bandwidth (ERB) filter to obtain features of the second near-end speech and the second far-end speech.

[0164] Step S602: Determine the second delay loss and the second speech enhancement loss of the student branch based on the features of the second near-end speech and the features of the second far-end speech.

[0165] Specifically, the features of the second near-end speech and the features of the second far-end speech are input into the second time delay estimation (TDE) module of the student branch to generate the second predicted near-end speech signal of the second time delay estimation module; the second delay loss of the student branch is determined based on the second predicted near-end speech signal. Simultaneously, the generated second predicted near-end speech signal is input into the encoder layer, bottleneck layer, decoder layer, complex convolving mask block (CCM) layer, and data frame filter (DF filter) to generate the second predicted target speech signal; the second speech enhancement loss of the student branch is determined based on the second predicted target speech signal.

[0166] In one possible implementation, a second delay loss is determined based on the second predicted near-end speech signal and the pre-acquired target near-end speech; simultaneously, a second speech enhancement loss is determined based on the second predicted target speech signal and the pre-acquired clean speech.

[0167] In this embodiment of the invention, the structural diagram of the student branch is as follows: Figure 7As shown, assuming the second near-end speech (Mic) and the second far-end speech (Ref) are input to the second feature extraction module 701 of the student branch, and then sequentially input to the second TDE module 702, the second coding layer 703, the second bottleneck layer 704, the second decoding layer 705, the second complex spectrum convolutional mapping (CCM) layer 706, and the second data frame filter 707, wherein the output of the second TDE module is transformed by inverse short-time Fourier transform to generate the second predicted near-end speech signal; the output of the second data frame filter is transformed by inverse short-time Fourier transform to generate the second predicted target speech signal; wherein the second coding layer can also be called the second coding module, the second bottleneck layer can also be called the second bottleneck module, the second decoding layer can also be called the second decoding module, and the second complex spectrum convolutional mapping (CCM) layer can also be called the second complex spectrum convolutional mapping (CCM) module.

[0168] In one possible implementation, the network structures of the second encoding layer and the second decoding layer reference open-source network structures such as Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement (DCCRN) and DEEPVQE. The encoding layer consists of four small encoding modules, each comprising a Conv2d convolutional layer, a BatchNorm regularization module, an exponential linear unit (ELU) unit, and grouped temporal convolution blocks (GT-Conv blocks). The CCM layer consists of two stages: first, by convolving the output features of the decoding layer with the input second near-end speech mic, the convolution output is divided into three weight components to construct a complex mask, where each component is the weight of a 120° rotation vector in the complex plane, resulting in the enhanced complex domain mask features; finally, the output of the CCM layer is passed through a second data frame filter (DF). The filter performs deep convolution filtering, and then the output of the second data frame filter is subjected to inverse short-time Fourier transform (ISTFT) to generate the second predicted target speech signal.

[0169] In this embodiment of the invention, a schematic diagram of determining the training loss function of the student branch is shown below. Figure 8 As mentioned above, in Figure 7Based on this, two Inverse Short Time Fourier Transforms (ISTFTs) are added, along with a diagram illustrating the second delay loss (i.e., the second TDE Loss) determined based on the second predicted near-end speech signal and the target near-end speech, the second speech enhancement loss (i.e., the second Se Loss) determined based on the second predicted target speech signal and the pre-acquired clean speech, and the training loss function of the student branch determined based on the aforementioned second delay loss and the second speech enhancement loss.

[0170] In one possible implementation, assume that the clean speech is s(n), and the corresponding frequency domain signal is S(k,f); the target near-end speech (Target Mic) of the second TDE module is m(n), and the corresponding frequency domain signal is M(k,f); and the second predicted near-end speech signal (Est MIC) of the second TDE module is... The corresponding frequency domain signal is The second predicted target speech signal (EST SPEECH) predicted by the teacher branch is The corresponding frequency domain signal is Using commonly used speech denoising and separation domain loss functions such as mean square error (MSE) and scale-invariant source-to-noise ratio (SISNR), the second delay loss L1 and the second speech enhancement loss L2 are determined as follows:

[0171]

[0172] Wherein, α can be set to 0.5.

[0173] Step S603: Determine the training loss function for the student branch based on the second delay loss and the second speech enhancement loss.

[0174] Specifically, the third product of the second delay loss L1 and the third weight, and the fourth product of the second speech enhancement loss L2 and the fourth weight are determined; the sum of the third product and the fourth product is determined as the training loss function L of the student branch. kt .

[0175] Assume the third weight is 0.3, the fourth weight is 0.7, and the training loss function L for the student branch is... kt The formula is as follows:

[0176] L kt =0.3*L1+0.7*L2

[0177] In this embodiment of the invention, the L1 of the teacher branch and the L1 of the student branch can be the same value or different data, and the L2 of the teacher branch and the L2 of the student branch can be the same value or different data, depending on the actual situation.

[0178] In this embodiment of the invention, the training loss function L of the teacher branch determined according to the above-described specific embodiment one is used. dt The training loss function L for the student branch determined in Specific Implementation Example 2 kt The joint loss function is determined together, specifically including:

[0179] Determine the product of the training loss function of the student branch and a set parameter, wherein the set parameter represents the KL (Kullback-Leibler) divergence between the teacher branch and the student branch; and determine the sum of the product and the training loss function of the teacher branch as the joint loss function.

[0180] In one possible implementation, the joint loss function is formulated as follows:

[0181] L = L dt +γL kd

[0182] Wherein, γ is a set parameter, determined based on the KL divergence between the teacher branch and the student branch.

[0183] In this embodiment of the invention, the network structures of the teacher branch (i.e., the teacher model) and the student branch (i.e., the student model) are completely identical. The teacher model is mainly responsible for training noisy echo data with a latency within a set range, for example, 500 milliseconds (ms). The student model is jointly trained with the teacher model using the same data during training. When the teacher model's performance is determined to be optimal through objective evaluation data, the decision space of the teacher model is mapped (initialized) to the decision space of the student model using distillation learning. The decision space includes model parameters. Simultaneous training and parameter updates of the student models help narrow the performance gap between the teacher and student models. Furthermore, training both models simultaneously using data with a large latency dynamic range from the AEC-Challenge and simulation data enhances their robustness under conditions of significant latency jitter. Additionally, knowledge distillation between teacher and student branches effectively improves the performance of the end-to-end model. Joint training reduces excessive speech nonlinear distortion during end-to-end model denoising and offers advantages such as low latency and better speech recovery.

[0184] In one possible implementation, the first TDE module and the second TDE module have the same structure. Taking the first TDE module as an example, specifically as follows: Figure 9 As shown, it includes: determining the features X of the first near-end speech. mic ∈R c×t×f The features X of the first distant speech ref ∈R c×t×f Where R represents the dataset, c represents the number of channels, t represents the time dimension, and f represents the frequency dimension; a convolutional network (Conv) is used to perform point-to-point convolution processing on the two features at the channel layer, and the processed feature is Q. mic ∈R h×t×f and K ref ∈R h×t×f Where h represents the current number of convolutional layers; then, for each Q... mic ∈R h×t×f and K ref ∈R h×t×f Process K ref ∈R h×t×f Unfolding in the time domain, the feature K_Delay after increasing the time length is K. ref ∈R h×t×d×f Wherein, d is the delay fluctuation range, which can be set to 120 frames; simultaneously, for Q... mic ∈R h×t×f Perform a transpose operation to generate Q. mic ∈R h×f×t Q mic ∈R h×f×t and K ref ∈R c×t×d×f Perform matrix multiplication (matmul) on the frequency dimension to obtain the differential feature K on the implicit time delay dimension. ref ∈R h×d×t , will the K ref ∈R h×d×t The input is fed into the Softmax layer to compute the delay probability distribution D∈R d×t Then the Q mic ∈R h×t×f and the D∈R d×t Weighted calculations are performed to obtain the near-end alignment mapping feature X. mic ∈R h×t×d×f Then, a Max operation is performed on the delay channel dimension, followed by a deconvolution operation to obtain the aligned output X. mic_align ∈R c×t×fThat is, the output of the first TDE module; the first TDE module adds a multi-layer attention mechanism and convolution operation alignment signal delay, and directly outputs X. mic_align The signal is used to calculate the training loss function of the teacher branch and has strong robustness.

[0185] In this embodiment of the invention, the second TDE module has the same structure as the first TDE module, and the training loss function for the student branch is generated according to the above process, which also has strong robustness.

[0186] Step S104: Adjust the knowledge distillation end-to-end echo cancellation model to be trained according to the joint loss function to generate the knowledge distillation end-to-end echo cancellation model.

[0187] Step S105: Compress the knowledge distillation end-to-end echo cancellation model to generate a target knowledge distillation end-to-end echo cancellation model, wherein the target knowledge distillation end-to-end echo cancellation model is used to achieve echo cancellation.

[0188] Specifically, the model compression process is as follows: Figure 10 As shown, it includes:

[0189] Step S1001: Quantize the knowledge distillation end-to-end echo cancellation model online to generate a knowledge distillation end-to-end echo cancellation model with Torch model structure.

[0190] Specifically, the knowledge distillation end-to-end echo cancellation model is quantized online using QAT, and the weights are quantized to INT_8 or Float_16 to generate a knowledge distillation end-to-end echo cancellation model with a Torch model structure.

[0191] Step S1002: The knowledge distillation end-to-end echo cancellation model of the Torch model structure is converted into ONNX through an open neural network to generate a knowledge distillation end-to-end echo cancellation model of the ONNX model structure.

[0192] Step S1003: The knowledge distillation end-to-end echo cancellation model of the ONNX model structure is processed by the MNN inference engine to generate the target knowledge distillation end-to-end echo cancellation model.

[0193] Through the above embodiments, firstly, during the data simulation process, the delays of the near-end signal and the far-end signal are randomly set, then the delay magnitude is predicted by the time delay detection module and the completion is completed; finally, the echo cancellation training is completed by the target knowledge distillation end-to-end echo cancellation model, so as to maximize the echo cancellation performance and delay detection performance of the target knowledge distillation end-to-end echo cancellation model.

[0194] In this embodiment of the invention, after step S105, the method further includes other steps, specifically as follows: Figure 11 As shown, the details are as follows:

[0195] Step S106: Deploy the target knowledge distillation end-to-end echo cancellation model to the terminal side.

[0196] Specifically, the target knowledge distillation end-to-end echo cancellation model is deployed on an Advanced RISC Machine (ARM) or Digital Signal Processing (DSP) platform on the terminal side, and then the entire voice link is integrated through engineering integration to form an end-to-end voice signal streaming processing system. Since the parameters of the target knowledge distillation end-to-end echo cancellation model are relatively small, it can be deployed on a terminal side with low computing power.

[0197] In this embodiment of the invention, the target knowledge distillation end-to-end echo cancellation model is deployed on the terminal side, which can realize functions such as delay detection, echo cancellation, speech noise reduction, and dereverberation, and complete the speech signal processing flow.

[0198] In one possible implementation, after completing step S106 of deploying the target knowledge distillation end-to-end echo cancellation model to the terminal side, the method further includes other steps, specifically as follows: Figure 12 As shown:

[0199] Step S107: Obtain the signal to be processed.

[0200] The signal to be processed is a voice signal.

[0201] Step S108: Input the signal to be processed into the target knowledge distillation end-to-end echo cancellation model to generate the target signal.

[0202] Specifically, by inputting the signal to be processed into the target knowledge distillation end-to-end echo cancellation model, the signal to be processed can be directly converted into the target signal. This improves the performance of delay detection, echo cancellation, speech noise reduction, and dereverberation of the signal to be processed under conditions of extremely low signal-to-return ratio, unstable delay, and variable echo propagation path, thereby enhancing the voice interaction performance and robustness of the terminal device.

[0203] In this embodiment of the invention, the system structure for training the end-to-end echo cancellation model of the target knowledge distillation is briefly described, as follows: Figure 13As shown, it includes: a data acquisition and simulation dataset construction module 1301, a model training module 1302, and a model compression and edge deployment module 1303. This is only an example illustration, and the specific system structure is constructed according to the actual situation.

[0204] In this embodiment of the invention, an echo cancellation device is provided, such as... Figure 14 As shown, it specifically includes: an acquisition unit 1401, a processing unit 1402, a determination unit 1403, and a generation unit 1404;

[0205] The acquisition unit 1401 is used to acquire a first near-end speech and a first far-end speech; the processing unit 1402 is used to preprocess the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech; the determining unit 1403 is used to input the second near-end speech and the second far-end speech into the knowledge distillation end-to-end echo cancellation model to be trained, respectively, and determine the joint loss function; the generating unit 1404 is used to adjust the knowledge distillation end-to-end echo cancellation model to be trained according to the joint loss function to generate a knowledge distillation end-to-end echo cancellation model; the generating unit 1404 is further used to compress the knowledge distillation end-to-end echo cancellation model to generate a target knowledge distillation end-to-end echo cancellation model, wherein the target knowledge distillation end-to-end echo cancellation model is used to achieve echo cancellation.

[0206] Furthermore, the device also includes:

[0207] The deployment unit is used to deploy the target knowledge distillation end-to-end echo cancellation model to the terminal side.

[0208] Furthermore, the processing unit is specifically used for:

[0209] The first near-end speech and the first far-end speech are pre-emphasized, power frequency interference is removed, windowing is applied, and short-time Fourier transform is performed to generate the second near-end speech and the second far-end speech.

[0210] Furthermore, the determining unit is specifically used for:

[0211] The second near-end speech and the second far-end speech are respectively input into the teacher branch and student branch of the knowledge distillation end-to-end echo cancellation model to be trained, and the training loss functions of the teacher branch and student branch are obtained respectively.

[0212] The joint loss function is determined based on the training loss function of the teacher branch and the training loss function of the student branch.

[0213] Furthermore, the determining unit is specifically used for:

[0214] Feature extraction is performed on the second proximal speech and the second distal speech to obtain the features of the second proximal speech and the features of the second distal speech;

[0215] Based on the features of the second near-end speech and the features of the second far-end speech, the first delay loss and the first speech enhancement loss of the teacher branch are determined;

[0216] The training loss function for the teacher branch is determined based on the first delay loss and the first speech enhancement loss.

[0217] Furthermore, the determining unit is specifically used to: input the features of the second near-end speech and the features of the second far-end speech into the first time delay detection module of the teacher branch to generate the first predicted near-end speech signal of the first time delay detection module;

[0218] The first delay loss of the teacher branch is determined based on the first predicted near-end speech signal.

[0219] Furthermore, the determining unit is specifically used to: input the features of the second near-end speech and the features of the second far-end speech into the first time delay detection module of the teacher branch to generate a first predicted near-end speech signal;

[0220] The first predicted near-end speech signal is input into the coding layer, bottleneck layer, decoding layer, complex spectrum convolutional mapping layer and data frame filter to generate the first predicted target speech signal;

[0221] The first speech enhancement loss of the teacher branch is determined based on the first predicted target speech signal.

[0222] Furthermore, the determining unit is specifically used to: determine a first product of the first delay loss and the first weight, and a second product of the first speech enhancement loss and the second weight;

[0223] The sum of the first product and the second product is determined as the training loss function of the teacher branch.

[0224] Furthermore, the determining unit is specifically used to: extract features from the second proximal speech and the second distal speech to obtain features of the second proximal speech and features of the second distal speech;

[0225] Based on the features of the second near-end speech and the features of the second far-end speech, the second delay loss and the second speech enhancement loss of the student branch are determined;

[0226] The training loss function for the student branch is determined based on the second delay loss and the second speech enhancement loss.

[0227] Furthermore, the determining unit is specifically used to: input the features of the second near-end speech and the features of the second far-end speech into the second time delay detection module of the student branch to generate a second predicted near-end speech signal;

[0228] The second delay loss of the student branch is determined based on the second predicted near-end speech signal.

[0229] Furthermore, the determining unit is specifically used to: input the features of the second near-end speech and the features of the second far-end speech into the second time delay detection module of the student branch to generate the second predicted near-end speech signal of the second time delay detection module;

[0230] The second predicted near-end speech signal is input into the coding layer, bottleneck layer, decoding layer, complex spectrum convolutional mapping layer and data frame filter to generate the second predicted target speech signal;

[0231] The second speech enhancement loss of the student branch is determined based on the second predicted target speech signal.

[0232] Furthermore, the determining unit is specifically used to: determine the third product of the second delay loss and the third weight, and the fourth product of the second speech enhancement loss and the fourth weight;

[0233] The sum of the third product and the fourth product is determined as the training loss function of the student branch.

[0234] Furthermore, the determining unit is specifically used to: determine the product of the training loss function of the student branch and a set parameter, wherein the set parameter represents the KL divergence between the teacher branch and the student branch;

[0235] The sum of the product and the training loss function of the teacher branch is determined as the joint loss function.

[0236] Furthermore, the generation unit is specifically used for:

[0237] The knowledge distillation end-to-end echo cancellation model is quantized online to generate a knowledge distillation end-to-end echo cancellation model with a Torch model structure.

[0238] The knowledge distillation end-to-end echo cancellation model of the Torch model structure is transformed into ONNX through an open neural network to generate a knowledge distillation end-to-end echo cancellation model of the ONNX model structure.

[0239] The knowledge distillation end-to-end echo cancellation model of the ONNX model structure is passed through the MNN inference engine to generate the target knowledge distillation end-to-end echo cancellation model.

[0240] Furthermore, the acquisition unit is also used to: acquire the signal to be processed;

[0241] The generation unit is further configured to: input the signal to be processed into the target knowledge distillation end-to-end echo cancellation model to generate the target signal.

[0242] Figure 15 This is a schematic diagram of the structure of the electronic device described in an embodiment of the present invention. Figure 15 As shown, it includes a general computer hardware architecture, which includes at least a processor 1501 and a memory 1502. The processor 1501 and the memory 1502 are connected via a bus 1503. The memory 1502 is adapted to store instructions or programs executable by the processor 1501. The processor 1501 can be a standalone microprocessor or a collection of one or more microprocessors. Thus, the processor 1501 executes the instructions stored in the memory 1502 to perform the method flow of the embodiments of the present invention as described above, thereby realizing data processing and control of other devices. The bus 1503 connects the above-mentioned components together, and also connects the above-mentioned components to the display controller 1504, the display device, and the input / output (I / O) device 1505. The input / output (I / O) device 1505 can be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, the input / output device 1505 is connected to the system via an input / output (I / O) controller 1506.

[0243] The instructions stored in memory 1502 are executed by at least one processor 1501 to achieve the following: acquiring a first near-end speech and a first far-end speech; preprocessing the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech; inputting the second near-end speech and the second far-end speech into a knowledge distillation end-to-end echo cancellation model to be trained, and determining a joint loss function; adjusting the knowledge distillation end-to-end echo cancellation model to be trained according to the joint loss function to generate a knowledge distillation end-to-end echo cancellation model; and compressing the knowledge distillation end-to-end echo cancellation model to generate a target knowledge distillation end-to-end echo cancellation model, wherein the target knowledge distillation end-to-end echo cancellation model is used to implement echo cancellation.

[0244] Specifically, the electronic device includes: one or more processors 1501 and a memory 1502. Figure 15Take processor 1501 as an example. Processor 1501 and memory 1502 can be connected via a bus or other means. Figure 15 Taking a bus connection as an example, memory 1502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Processor 1501 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in memory 1502, thereby implementing the aforementioned method for determining echo cancellation.

[0245] Memory 1502 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store an option list, etc. Furthermore, memory 1502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 1502 may optionally include memory remotely located relative to processor 1501, and these remote memories may be connected to external devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0246] One or more modules are stored in memory 1502 and, when executed by one or more processors 1501, perform the echo cancellation method in any of the above method embodiments.

[0247] As those skilled in the art will recognize, various aspects of the embodiments of the present invention can be implemented as a system, method, or computer program product. Therefore, various aspects of the embodiments of the present invention can take the form of a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software and hardware aspects, which may generally be referred to herein as a "circuit," "module," or "system." Furthermore, various aspects of the embodiments of the present invention can take the form of a computer program product implemented in one or more computer-readable media having computer-readable program code implemented thereon.

[0248] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, (but not limited to) an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media will include: an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the context of embodiments of the present invention, a computer-readable storage medium can be any tangible medium capable of containing or storing a program used by or in conjunction with an instruction execution system, device, or apparatus.

[0249] Computer-readable signal media may include propagated digital signals having computer-readable program code implemented therein, such as in baseband or as part of a carrier wave. Such propagated signals may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and can communicate, propagate, or transmit a program used by or in conjunction with an instruction execution system, device, or apparatus.

[0250] Program code implemented on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, or any suitable combination thereof.

[0251] Computer program code for performing operations relating to various aspects of embodiments of the present invention can be written in any combination of one or more programming languages, including: object-oriented programming languages ​​such as Java, Smalltalk, C++, etc.; and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The program code can be executed as a standalone software package entirely on the user's computer, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet provided by an Internet service provider).

[0252] The flowchart illustrations and / or block diagrams of the methods, apparatus (systems), and computer program products according to embodiments of the present invention describe various aspects of the embodiments of the present invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions (executed via the processor of the computer or other programmable data processing apparatus) create means for implementing the functions / actions specified in the flowchart and / or block diagram blocks or blocks.

[0253] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus or other means to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of writing that includes instructions that implement the functions / actions specified in flowchart and / or block diagram blocks or blocks.

[0254] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operable steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide for implementing the functions / actions specified in flowchart and / or block diagram blocks or blocks.

[0255] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0256] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of such data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding access points are provided for users to choose to authorize or refuse processing. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

Claims

1. A method for echo cancellation, characterized in that, The method includes: Acquire the first near-end speech and the first far-end speech; The first near-end speech and the first far-end speech are preprocessed to generate the preprocessed second near-end speech and the second far-end speech. The second near-end speech and the second far-end speech are respectively input into the knowledge distillation end-to-end echo cancellation model to be trained, and the joint loss function is determined. The knowledge distillation end-to-end echo cancellation model to be trained is adjusted according to the joint loss function to generate a knowledge distillation end-to-end echo cancellation model. The knowledge distillation end-to-end echo cancellation model is compressed to generate a target knowledge distillation end-to-end echo cancellation model, wherein the target knowledge distillation end-to-end echo cancellation model is used to achieve echo cancellation.

2. The method according to claim 1, characterized in that, The method further includes: The target knowledge distillation end-to-end echo cancellation model is deployed to the terminal side.

3. The method according to claim 1, characterized in that, The step of preprocessing the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech specifically includes: The first near-end speech and the first far-end speech are pre-emphasized, power frequency interference is removed, windowing is applied, and short-time Fourier transform is performed to generate the second near-end speech and the second far-end speech.

4. The method according to claim 1, characterized in that, The step of inputting the second near-end speech and the second far-end speech into the knowledge distillation end-to-end echo cancellation model to be trained to determine the joint loss function specifically includes: The second near-end speech and the second far-end speech are respectively input into the teacher branch and student branch of the knowledge distillation end-to-end echo cancellation model to be trained, and the training loss functions of the teacher branch and student branch are obtained respectively. The joint loss function is determined based on the training loss function of the teacher branch and the training loss function of the student branch.

5. The method according to claim 4, characterized in that, The step of inputting the second near-end speech and the second far-end speech into the teacher branch of the knowledge distillation end-to-end echo cancellation model to be trained, and obtaining the training loss function of the teacher branch, specifically includes: Feature extraction is performed on the second proximal speech and the second distal speech to obtain the features of the second proximal speech and the features of the second distal speech; Based on the features of the second near-end speech and the features of the second far-end speech, the first delay loss and the first speech enhancement loss of the teacher branch are determined; The training loss function for the teacher branch is determined based on the first delay loss and the first speech enhancement loss.

6. The method according to claim 5, characterized in that, The step of determining the first delay loss of the teacher branch based on the features of the second near-end speech and the features of the second far-end speech specifically includes: The features of the second near-end speech and the features of the second far-end speech are input into the first time delay detection module of the teacher branch to generate the first predicted near-end speech signal of the first time delay detection module. The first delay loss of the teacher branch is determined based on the first predicted near-end speech signal.

7. The method according to claim 5, characterized in that, The step of determining the first speech enhancement loss of the teacher branch based on the features of the second proximal speech and the features of the second distal speech specifically includes: The features of the second near-end speech and the features of the second far-end speech are input into the first time delay detection module of the teacher branch to generate the first predicted near-end speech signal. The first predicted near-end speech signal is input into the coding layer, bottleneck layer, decoding layer, complex spectrum convolutional mapping layer and data frame filter to generate the first predicted target speech signal; The first speech enhancement loss of the teacher branch is determined based on the first predicted target speech signal.

8. The method according to claim 5, characterized in that, The step of determining the training loss function for the teacher branch based on the first delay loss and the first speech enhancement loss specifically includes: Determine the first product of the first delay loss and the first weight, and the second product of the first speech enhancement loss and the second weight; The sum of the first product and the second product is determined as the training loss function of the teacher branch.

9. The method according to claim 4, characterized in that, The step of inputting the second near-end speech and the second far-end speech into the student branch of the knowledge distillation end-to-end echo cancellation model to be trained, and obtaining the training loss function of the student branch, specifically includes: Feature extraction is performed on the second proximal speech and the second distal speech to obtain the features of the second proximal speech and the features of the second distal speech; Based on the features of the second near-end speech and the features of the second far-end speech, the second delay loss and the second speech enhancement loss of the student branch are determined; The training loss function for the student branch is determined based on the second delay loss and the second speech enhancement loss.

10. The method according to claim 9, characterized in that, The step of determining the second delay loss of the student branch based on the features of the second near-end speech and the features of the second far-end speech specifically includes: The features of the second near-end speech and the features of the second far-end speech are input into the second time delay detection module of the student branch to generate the second predicted near-end speech signal. The second delay loss of the student branch is determined based on the second predicted near-end speech signal.

11. The method according to claim 9, characterized in that, The step of determining the first speech enhancement loss for the student branch based on the features of the second near-end speech and the features of the second far-end speech specifically includes: The features of the second near-end speech and the features of the second far-end speech are input into the second time delay detection module of the student branch to generate the second predicted near-end speech signal of the second time delay detection module. The second predicted near-end speech signal is input into the coding layer, bottleneck layer, decoding layer, complex spectrum convolutional mapping layer and data frame filter to generate the second predicted target speech signal; The second speech enhancement loss of the student branch is determined based on the second predicted target speech signal.

12. The method according to claim 9, characterized in that, The step of determining the training loss function for the student branch based on the second delay loss and the second speech enhancement loss specifically includes: Determine the third product of the second delay loss and the third weight, and the fourth product of the second speech enhancement loss and the fourth weight; The sum of the third product and the fourth product is determined as the training loss function of the student branch.

13. The method according to claim 4, characterized in that, The step of determining the joint loss function based on the training loss function of the teacher branch and the training loss function of the student branch specifically includes: Determine the product of the training loss function of the student branch and a set parameter, wherein the set parameter represents the KL divergence between the teacher branch and the student branch; The sum of the product and the training loss function of the teacher branch is determined as the joint loss function.

14. The method according to claim 1, characterized in that, The step of compressing the knowledge distillation end-to-end echo cancellation model to generate the target knowledge distillation end-to-end echo cancellation model specifically includes: The knowledge distillation end-to-end echo cancellation model is quantized online to generate a knowledge distillation end-to-end echo cancellation model with a Torch model structure. The knowledge distillation end-to-end echo cancellation model of the Torch model structure is transformed into ONNX through an open neural network to generate a knowledge distillation end-to-end echo cancellation model of the ONNX model structure. The knowledge distillation end-to-end echo cancellation model of the ONNX model structure is passed through the MNN inference engine to generate the target knowledge distillation end-to-end echo cancellation model.

15. A method for echo cancellation, characterized in that, The method includes: Acquire the signal to be processed; The signal to be processed is input into the target knowledge distillation end-to-end echo cancellation model to generate the target signal, wherein the target knowledge distillation end-to-end echo cancellation model is obtained by the method of any one of claims 1-14, and the target knowledge distillation end-to-end echo cancellation model is deployed on the terminal side to achieve echo cancellation.

16. An echo cancellation device, characterized in that, The device includes: An acquisition unit is used to acquire the first near-end speech and the first far-end speech; The processing unit is configured to preprocess the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech. The determination unit is used to input the second proximal speech and the second far-end speech into the knowledge distillation end-to-end echo cancellation model to be trained, and determine the joint loss function. The generation unit is used to adjust the knowledge distillation end-to-end echo cancellation model to be trained according to the joint loss function, and generate the knowledge distillation end-to-end echo cancellation model. The generation unit is further configured to compress the knowledge distillation end-to-end echo cancellation model to generate a target knowledge distillation end-to-end echo cancellation model, wherein the target knowledge distillation end-to-end echo cancellation model is used to achieve echo cancellation.

17. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-15.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-15.

Citation Information

Patent Citations

  • Model training method, echo cancellation method, system and device and storage medium

    CN114530160A

  • Server, terminal equipment and model compression method

    CN117892778A