A method and apparatus for echo cancellation and noise reduction

By using a two-stage neural network multi-task fusion model, the problem of echo cancellation and noise reduction in voice communication and recognition systems with extremely low signal-to-return ratios is solved, thereby improving the voice interaction and communication performance of terminal devices.

CN119091900BActive Publication Date: 2025-11-18ZHEJIANG FUTURE ELF ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411303715.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-11-18
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

Under extremely low signal-to-return ratio conditions, existing technical solutions perform poorly in voice communication and recognition systems, failing to effectively achieve acoustic echo cancellation and noise reduction, thus affecting the voice interaction and communication performance of terminal devices.

Method used

A two-stage neural network multi-task fusion model is adopted. Through preprocessing, delay detection, short-time Fourier transform and linear echo cancellation, combined with the target two-stage neural network multi-task fusion model, a target two-stage neural network multi-task fusion model is generated to achieve echo cancellation and noise reduction.

Benefits of technology

With an extremely low signal-to-return ratio, it significantly improves the voice interaction and communication performance of terminal devices, and enhances the accuracy of voice recognition and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091900B_ABST
    Figure CN119091900B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method and device for echo cancellation and noise reduction. In the embodiments of the present application, a first near-end speech and a first far-end speech are obtained, and after pre-processing, delay detection and short-time Fourier transform, a fourth near-end speech and a fourth far-end speech are generated, and linear echo cancellation is performed to generate an error signal and a linear echo signal of the linear echo cancellation; the fourth near-end speech, the fourth far-end speech, the error signal and the linear echo signal of the linear echo cancellation are input into a two-stage neural network multi-task fusion model to be trained to determine a joint loss function; the two-stage neural network multi-task fusion model is generated according to the joint loss function, and then model compression is performed to generate a target two-stage neural network multi-task fusion model, which is used to realize echo cancellation and noise reduction. Through the above method, acoustic echo cancellation and noise reduction are realized under the condition of extremely low signal-to-echo ratio, and the speech interaction and communication performance of a terminal device are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and more specifically, to a method and apparatus for echo cancellation and noise reduction. Background Technology

[0002] Intelligent voice technology is mainly used in voice communication and interaction systems, such as smart tablets, smart speakers, and mobile conference screens. These voice communication and interaction systems suffer from problems such as noise interference, large echo residue at a distance, and long pickup distance, which severely limit the performance of voice communication and recognition systems.

[0003] In existing technologies, two schemes have been proposed to improve the performance of voice communication and recognition systems. Scheme 1 combines signal processing and neural networks. Specifically, linear echo suppression is first performed during signal processing, and then nonlinear echo suppression, noise reduction, dereverberation, and automatic gain control are achieved through neural networks. Scheme 2 establishes an end-to-end multi-task learning network architecture such as DeepVQE to simultaneously perform acoustic echo cancellation, noise suppression, and automatic gain control, featuring low latency, high operability, and simple link characteristics. However, when the signal-to-return ratio reaches below -30dB, the voice interaction and communication performance of both schemes is poor.

[0004] In summary, how to achieve better acoustic echo cancellation and noise reduction under extremely low signal-to-return ratio conditions, and improve the voice interaction and communication performance of terminal devices, is a problem that needs to be solved. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method and apparatus for echo cancellation and noise reduction, which adopts a target two-stage neural network multi-task fusion model to achieve better acoustic echo cancellation and noise reduction under extremely low signal-to-return ratio conditions, thereby improving the voice interaction and communication performance of terminal devices.

[0006] In a first aspect, embodiments of the present invention provide a method for echo cancellation and noise reduction, the method comprising:

[0007] Acquire the first near-end speech and the first far-end speech;

[0008] The first near-end speech and the first far-end speech are preprocessed to generate the preprocessed second near-end speech and the second far-end speech.

[0009] Delay detection is performed on the second near-end speech and the second far-end speech to generate aligned third near-end speech and third far-end speech;

[0010] The third proximal speech and the third distal speech are respectively subjected to short-time Fourier transform to generate the fourth proximal speech and the fourth distal speech;

[0011] Linear echo cancellation is performed on the fourth near-end speech and the fourth far-end speech to generate an error signal for linear echo cancellation and a linear echo signal.

[0012] The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are input into the two-stage neural network multi-task fusion model to be trained, and the joint loss function is determined.

[0013] The two-stage neural network multi-task fusion model to be trained is adjusted according to the joint loss function to generate the two-stage neural network multi-task fusion model.

[0014] The two-stage neural network multi-task fusion model is compressed to generate a target two-stage neural network multi-task fusion model, wherein the two-stage neural network multi-task fusion model is used to achieve echo cancellation and noise reduction.

[0015] Optionally, the method further includes:

[0016] The two-stage neural network multi-task fusion model is deployed to the terminal side.

[0017] Optionally, the step of preprocessing the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech specifically includes:

[0018] The first near-end speech and the first far-end speech are pre-emphasized, power frequency interference is removed, and windowing is applied to generate the second near-end speech and the second far-end speech.

[0019] Optionally, the step of inputting the fourth proximal speech, the fourth distal speech, the error signal of linear echo cancellation, and the linear echo signal into the two-stage neural network multi-task fusion model to be trained, and determining the joint loss function, specifically includes:

[0020] The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are input into the first stage of the two-stage neural network multi-task fusion model to be trained, generating the first stage training loss function and generating the first enhanced near-end speech and the first enhanced linear echo cancellation signal. The first stage of the two-stage neural network multi-task fusion model is used for preliminary noise reduction, preliminary echo cancellation, and removal of late reverberation.

[0021] The fourth near-end speech, the error signal of the linear echo cancellation, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal are input into the second stage of the two-stage neural network multi-task fusion model to be trained, generating the second-stage training loss function. The second stage of the two-stage neural network multi-task fusion model is used to remove residual noise, residual echo, and early reverberation.

[0022] The joint loss function is determined based on the training loss function of the first stage and the training loss function of the second stage.

[0023] Optionally, the step of inputting the fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal into the first stage of the two-stage neural network multi-task fusion model to be trained, and generating the first-stage training loss function, specifically includes:

[0024] The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are subjected to feature extraction and feature splicing to obtain the first speech feature;

[0025] The first speech feature is input into the first encoding layer, the first bottleneck layer and the first decoding layer to generate the first intermediate speech data;

[0026] The first intermediate speech data and the fourth near-end speech data are input into the first complex spectrum convolutional mapping layer to generate the first enhanced near-end speech data, and the first intermediate speech data and the fourth far-end speech data are input into the second complex spectrum convolutional mapping layer to generate the first enhanced linear echo cancellation signal.

[0027] The first enhanced near-end speech and the first enhanced linear echo cancellation signal are input into the first fusion module to generate the first stage output features;

[0028] The training loss function for the first stage is determined based on the output features of the first stage.

[0029] Optionally, determining the training loss function for the first stage based on the output features of the first stage specifically includes:

[0030] The first stage output features are subjected to inverse short-time Fourier transform to generate the first stage target speech;

[0031] The training loss function for the first stage is determined based on the target speech in the first stage and the pre-acquired clean speech.

[0032] Optionally, the step of inputting the fourth near-end speech, the linear echo cancellation error signal, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal into the second stage of the two-stage neural network multi-task fusion model to be trained, and generating the second-stage training loss function, specifically includes:

[0033] The second speech feature is obtained by extracting and concatenating features from the fourth near-end speech, the linear echo cancellation error signal, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal.

[0034] The second speech feature is input into the second coding layer, the second bottleneck layer, and the second decoding layer to generate the second intermediate speech data.

[0035] The second intermediate speech data and the first enhanced near-end speech are input into the third complex spectrum convolutional mapping layer to generate the second enhanced near-end speech, and the second intermediate speech data and the first enhanced linear echo cancellation signal are input into the fourth complex spectrum convolutional mapping layer to generate the second enhanced linear echo cancellation signal.

[0036] The second enhanced near-end speech and the second enhanced linear echo cancellation signal are input into the second fusion module to generate the second-stage output features;

[0037] The training loss function for the second stage is determined based on the output features of the second stage.

[0038] Optionally, determining the second-stage training loss function based on the second-stage output features specifically includes:

[0039] The output features of the second stage are subjected to inverse short-time Fourier transform to generate the target speech of the second stage;

[0040] The second-stage training loss function is determined based on the target speech in the second stage and the pre-acquired clean speech.

[0041] Optionally, determining the joint loss function based on the first-stage training loss function and the second-stage training loss function specifically includes:

[0042] Obtain the first weight of the first stage training loss function and the second weight of the first stage training loss function;

[0043] Determine the first product of the first weight and the first stage training loss function, and the second product of the second weight and the second stage training loss function;

[0044] The sum of the first product and the second product is determined as the joint loss function.

[0045] Optionally, the step of compressing the two-stage neural network multi-task fusion model to generate the target two-stage neural network multi-task fusion model specifically includes:

[0046] The two-stage neural network multi-task fusion model is quantized online to generate a two-stage neural network multi-task fusion model with a Torch model structure.

[0047] The two-stage neural network multi-task fusion model of the Torch model structure is converted into ONNX through open neural network transformation to generate a two-stage neural network multi-task fusion model of the ONNX model structure.

[0048] The two-stage neural network multi-task fusion model of the ONNX model structure is passed through the MNN inference engine to generate the target two-stage neural network multi-task fusion model.

[0049] Secondly, embodiments of the present invention provide a method for echo cancellation and noise reduction, the method comprising:

[0050] Acquire the signal to be processed;

[0051] The signal to be processed is input into the target two-stage neural network multi-task fusion model to generate the target signal. The target two-stage neural network multi-task fusion model is deployed on the terminal side to achieve echo cancellation and noise reduction.

[0052] Thirdly, embodiments of the present invention provide an echo cancellation and noise reduction apparatus, the apparatus comprising:

[0053] An acquisition unit is used to acquire the first near-end speech and the first far-end speech;

[0054] The processing unit is configured to preprocess the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech.

[0055] The detection unit is used to perform delay detection on the second near-end speech and the second far-end speech to generate aligned third near-end speech and third far-end speech;

[0056] The processing unit is further configured to perform short-time Fourier transform on the third proximal speech and the third distal speech respectively to generate a fourth proximal speech and a fourth distal speech.

[0057] The generation unit is used to perform linear echo cancellation on the fourth near-end speech and the fourth far-end speech, and generate an error signal for linear echo cancellation and a linear echo signal.

[0058] The determining unit is used to input the fourth proximal speech, the fourth distal speech, the error signal of the linear echo cancellation, and the linear echo signal into the two-stage neural network multi-task fusion model to be trained, and to determine the joint loss function;

[0059] The generation unit is also used to adjust the two-stage neural network multi-task fusion model to be trained according to the joint loss function, and generate a two-stage neural network multi-task fusion model.

[0060] The generation unit is further configured to compress the two-stage neural network multi-task fusion model to generate a target two-stage neural network multi-task fusion model, wherein the two-stage neural network multi-task fusion model is used to achieve echo cancellation and noise reduction.

[0061] Optionally, the device further includes:

[0062] The deployment unit is used to deploy the two-stage neural network multi-task fusion model to the terminal side.

[0063] Optionally, the processing unit is specifically used for:

[0064] The first near-end speech and the first far-end speech are pre-emphasized, power frequency interference is removed, and windowing is applied to generate the second near-end speech and the second far-end speech.

[0065] Optionally, the determining unit is specifically used for:

[0066] The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are input into the first stage of the two-stage neural network multi-task fusion model to be trained, generating the first stage training loss function and generating the first enhanced near-end speech and the first enhanced linear echo cancellation signal. The first stage of the two-stage neural network multi-task fusion model is used for preliminary noise reduction, preliminary echo cancellation, and removal of late reverberation.

[0067] The fourth near-end speech, the error signal of the linear echo cancellation, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal are input into the second stage of the two-stage neural network multi-task fusion model to be trained, generating the second-stage training loss function. The second stage of the two-stage neural network multi-task fusion model is used to remove residual noise, residual echo, and early reverberation.

[0068] The joint loss function is determined based on the training loss function of the first stage and the training loss function of the second stage.

[0069] Optionally, the determining unit is specifically used for:

[0070] The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are subjected to feature extraction and feature splicing to obtain the first speech feature;

[0071] The first speech feature is input into the first encoding layer, the first bottleneck layer and the first decoding layer to generate the first intermediate speech data;

[0072] The first intermediate speech data and the fourth near-end speech data are input into the first complex spectrum convolutional mapping layer to generate the first enhanced near-end speech data, and the first intermediate speech data and the fourth far-end speech data are input into the second complex spectrum convolutional mapping layer to generate the first enhanced linear echo cancellation signal.

[0073] The first enhanced near-end speech and the first enhanced linear echo cancellation signal are input into the first fusion module to generate the first stage output features;

[0074] The training loss function for the first stage is determined based on the output features of the first stage.

[0075] Optionally, the determining unit is specifically used for:

[0076] The first stage output features are subjected to inverse short-time Fourier transform to generate the first stage target speech;

[0077] The training loss function for the first stage is determined based on the target speech in the first stage and the pre-acquired clean speech.

[0078] Optionally, the determining unit is specifically used for:

[0079] The second speech feature is obtained by extracting and concatenating features from the fourth near-end speech, the linear echo cancellation error signal, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal.

[0080] The second speech feature is input into the second coding layer, the second bottleneck layer, and the second decoding layer to generate the second intermediate speech data.

[0081] The second intermediate speech data and the first enhanced near-end speech are input into the third complex spectrum convolutional mapping layer to generate the second enhanced near-end speech, and the second intermediate speech data and the first enhanced linear echo cancellation signal are input into the fourth complex spectrum convolutional mapping layer to generate the second enhanced linear echo cancellation signal.

[0082] The second enhanced near-end speech and the second enhanced linear echo cancellation signal are input into the second fusion module to generate the second-stage output features;

[0083] The training loss function for the second stage is determined based on the output features of the second stage.

[0084] Optionally, the determining unit is specifically used for:

[0085] The output features of the second stage are subjected to inverse short-time Fourier transform to generate the target speech of the second stage;

[0086] The second-stage training loss function is determined based on the target speech in the second stage and the pre-acquired clean speech.

[0087] Optionally, the determining unit is specifically used for:

[0088] Obtain the first weight of the first stage training loss function and the second weight of the first stage training loss function;

[0089] Determine the first product of the first weight and the first stage training loss function, and the second product of the second weight and the second stage training loss function;

[0090] The sum of the first product and the second product is determined as the joint loss function.

[0091] Optionally, the generation unit is specifically used for:

[0092] The two-stage neural network multi-task fusion model is quantized online to generate a two-stage neural network multi-task fusion model with a Torch model structure.

[0093] The two-stage neural network multi-task fusion model of the Torch model structure is converted into ONNX through open neural network transformation to generate a two-stage neural network multi-task fusion model of the ONNX model structure.

[0094] The two-stage neural network multi-task fusion model of the ONNX model structure is passed through the MNN inference engine to generate the target two-stage neural network multi-task fusion model.

[0095] Optionally, the acquisition unit is further configured to: acquire the signal to be processed;

[0096] The generation unit is further configured to: input the signal to be processed into the target two-stage neural network multi-task fusion model to generate the target signal.

[0097] Fourthly, embodiments of the present invention provide an electronic device, including a memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect or any one of the possible methods of the first aspect.

[0098] Fifthly, embodiments of the present invention provide a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the method as described in the first aspect or any one of the possible methods described in the first aspect.

[0099] In this embodiment of the invention, a first near-end speech and a first far-end speech are acquired; the first near-end speech and the first far-end speech are preprocessed to generate preprocessed second near-end speech and second far-end speech; the second near-end speech and the second far-end speech are subjected to delay detection to generate aligned third near-end speech and third far-end speech; the third near-end speech and the third far-end speech are subjected to short-time Fourier transform to generate fourth near-end speech and fourth far-end speech; the fourth near-end speech and the fourth far-end speech are subjected to linear echo cancellation to generate an error signal for linear echo cancellation and a linear echo signal; the fourth near-end speech, the fourth far-end speech, the error signal for linear echo cancellation, and the linear echo signal are input into a two-stage neural network multi-task fusion model to be trained to determine a joint loss function; the two-stage neural network multi-task fusion model to be trained is adjusted according to the joint loss function to generate a two-stage neural network multi-task fusion model; the two-stage neural network multi-task fusion model is compressed to generate a target two-stage neural network multi-task fusion model, wherein the two-stage neural network multi-task fusion model is used to achieve echo cancellation and noise reduction. Using the above method, a target two-stage neural network multi-task fusion model is generated. This target two-stage neural network multi-task fusion model can achieve acoustic echo cancellation and noise reduction under extremely low signal-to-return ratio, thereby improving the voice interaction and communication performance of terminal devices. Attached Figure Description

[0100] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:

[0101] Figure 1 This is a flowchart of an echo cancellation and noise reduction method according to an embodiment of the present invention;

[0102] Figure 2 This is a schematic diagram of a preliminary data processing structure in an embodiment of the present invention;

[0103] Figure 3 This is a flowchart of a method for obtaining the first-stage training loss function according to an embodiment of the present invention;

[0104] Figure 4 This is a schematic diagram of the structure of a first fusion module in an embodiment of the present invention;

[0105] Figure 5 This is a schematic diagram of the structure of a first stage in an embodiment of the present invention;

[0106] Figure 6 This is a schematic diagram of a structure for obtaining the first-stage training loss function in an embodiment of the present invention;

[0107] Figure 7 This is a flowchart of a method for training a loss function in the second stage according to an embodiment of the present invention;

[0108] Figure 8 This is a schematic diagram of the structure of the second stage in an embodiment of the present invention;

[0109] Figure 9 This is a schematic diagram of a structure for obtaining the second-stage training loss function in an embodiment of the present invention;

[0110] Figure 10 This is a schematic diagram of a structure for generating the joint loss function in an embodiment of the present invention;

[0111] Figure 11 This is a flowchart of a model compression method according to an embodiment of the present invention;

[0112] Figure 12 This is a flowchart of another echo cancellation method in an embodiment of the present invention;

[0113] Figure 13 This is a flowchart of another echo cancellation method in an embodiment of the present invention;

[0114] Figure 14 This is a schematic diagram of a system structure in an embodiment of the present invention;

[0115] Figure 15 This is a schematic diagram of another echo cancellation and noise reduction device in an embodiment of the present invention;

[0116] Figure 16 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0117] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.

[0118] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0119] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".

[0120] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0121] With the upgrading of the intelligent voice industry in terms of acoustic hardware, application scenarios, user groups, and software functions, full-duplex communication (i.e., voice communication) and interactive systems face challenges such as diverse and complex acoustic environmental noise, prominent nonlinear distortion of speakers in high-volume playback scenarios, large latency fluctuations between far-end and near-end speech, extremely low signal-to-return ratio, significant distortion and reverberation in far-end interactive speech, impaired speech clarity, large echo residue at far-end points, and long pickup distances. These problems severely limit the performance of voice communication and recognition systems. Two solutions have been proposed in existing technologies to improve the performance of voice communication and recognition systems.

[0122] Option 1 combines signal processing and neural networks. Specifically, linear echo suppression is first performed during signal processing, and then nonlinear echo suppression, noise reduction, dereverberation, and automatic gain control are performed through neural networks. Examples include gated convolutional FT-LSTM neural networks (GFTNN), band-split recurrent neural networks (BSRNN), and neural Kalman filtering (NKF). GFTNN and BSRNN can also be referred to as an end-to-end model integrating echo cancellation, noise suppression, and dereverberation.

[0123] Option 2: Establish an end-to-end multi-task learning network architecture to simultaneously perform functions such as acoustic echo cancellation (AEC), noise suppression (NS), and automatic gain control (AGC). This approach features low latency, high operability, and a simple link structure. Examples include the Deep Variational Quantum Eigensolver (DEEPVQE) algorithm and a network for repairing and denoising speech signals (RadNET).

[0124] However, in smart speakers, large conference screens, and in-vehicle voice interaction, situations often involve playing music and videos at high volumes, as well as text-to-speech (TTS) operations, causing the signal-to-return ratio (SRR) to drop below -30dB. At this point, the voice interaction performance of both solutions is poor. Furthermore, the robustness of these solutions is poor, making them unsuitable for deployment on low-computing-power devices. Therefore, how to effectively achieve acoustic echo cancellation in full-duplex environments with extremely low SRRs, thereby improving the voice communication performance, speech recognition performance, and robustness of terminal devices, while also enabling deployment on low-computing-power devices, is a problem that needs to be solved.

[0125] In this embodiment of the invention, the signal-to-return ratio refers to the ratio of the power of the near-end speech signal received from the terminal side to the power of the echo signal received from the speaker, wherein the echo signal is generated by the voice of the user on the opposite side of the terminal side being played through the speaker on the terminal side; full-duplex is a communication method in which data can be transmitted simultaneously in two directions, that is, both parties can send and receive data at the same time, and the transmission in these two directions does not interfere with each other. For example, real-time communication (RTC) and live streaming scenarios are full-duplex communication.

[0126] In this embodiment of the invention, to solve the above problems, an echo cancellation and noise reduction method is proposed, specifically as follows: Figure 1 As shown, the method includes:

[0127] Step S101: Obtain the first near-end speech and the first far-end speech.

[0128] Specifically, the first near-end speech x(n) and the first far-end speech r(n) are both training data. Before acquiring the first near-end speech and the first far-end speech, it is necessary to construct a simulation dataset, and then acquire the first near-end speech and the first far-end speech from the simulation dataset.

[0129] In one possible implementation, when constructing the simulation dataset, data acquisition is first performed. The acquired data can come from public datasets such as Librispeech, Aishell, AEC-CHALLENGE and recorded echo data, OPENSIL and 100,000 simulated room impact responses, as well as DNS-CHALLENGE and recorded data. The public datasets Librispeech and Aishell are used as clean speech data; the AEC-CHALLENGE and recorded echo data are used as far-end speech data for synthesis; the OPENSIL and 100,000 simulated room impact responses are used as transfer functions; and the noise data comes from DNS-CHALLENGE and recorded data. The acquired data is used to synthesize 500 hours of training data with a signal-to-return ratio (SRR) range of [-35, 15] and a signal-to-noise ratio (SNR) range of [-5, 15] dB. The number of simulated room impact responses, the SRR range, the SNR range, and the training data duration are all determined based on actual conditions and are only illustrative examples here.

[0130] In this embodiment of the invention, clean speech is also acquired while acquiring the first near-end speech and the first far-end speech. The clean speech is used to calculate the joint loss function of the two-stage neural network multi-task fusion model to be trained.

[0131] Step S102: Preprocess the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech.

[0132] Specifically, the first near-end speech and the first far-end speech are pre-emphasized, power frequency interference is removed, and windowing is applied to generate pre-processed second near-end speech and second far-end speech.

[0133] Step S103: Perform delay detection on the second near-end speech and the second far-end speech to generate aligned third near-end speech and third far-end speech.

[0134] Specifically, the delays of the second proximal speech and the second distal speech are estimated and aligned using a generalized cross-correlation transformation algorithm (e.g., the GCC-PHAT function), generating aligned third proximal speech and third distal speech.

[0135] In one possible implementation, the Time Delay Estimation (TDE) can improve the convergence speed and echo cancellation capability of the subsequent Linear Acoustic Echo Cancellation (LAEC) algorithm.

[0136] In this embodiment of the invention, the second near-end speech and the second far-end speech are affected by factors such as asynchronous sampling clocks, network latency instability, and audio effect algorithm processing. The second near-end speech and the second far-end speech are not strictly aligned in the time dimension and have a certain degree of fluctuation. The audio effect algorithm processing refers to Dolby audio, surround sound, etc.

[0137] Step S104: Perform short-time Fourier transform on the third proximal speech and the third distal speech to generate the fourth proximal speech and the fourth distal speech, respectively.

[0138] In one possible implementation, the parameters used in the short-time Fourier transform (STFT) are shown in Table 1, as detailed below:

[0139] Table 1

[0140] parameter numerical values Sampling rate 16000 Window length 32ms (512 points) Frame shift 8ms (128 points) Fast Fourier Transform (FFT) Length 512 points

[0141] The window function used in the short-time Fourier transform is the Hanning window.

[0142] Step S105: Perform linear echo cancellation on the fourth near-end speech and the fourth far-end speech to generate an error signal for linear echo cancellation and a linear echo signal.

[0143] Specifically, the fourth near-end speech and the fourth far-end speech are input into the linear echo cancellation (LACE) module. The LACE module generates an error signal and a linear echo signal based on the filters in the LACE module. The LACE module uses a frequency domain block Kalman filter (FDBKF) and a power-normalized least mean squares (PNLMS) filter as the primary and secondary adaptive filters, respectively. When the FDBKF diverges, the PNLMS is used as the filter for the FDBKF to improve the overall echo cancellation performance.

[0144] In one possible implementation, the filter can also be a Partitioned Block Frequeney Domain Kalman Filter (PBFDKF), a Partitioned Block Frequeney Domain Adaptive Filter (PBDAF), a Partitioned Frequency-domain Block least-mean-square (PFBLMS) filter, or a Recursive Least-Squares (RLS) adaptive filter. During data simulation, the above filters can also be used to dynamically expand the echo cancellation effect under different conditions, enhance the suppression of residual echoes by nonlinear echo cancellation, and improve the robustness of the two-stage neural network multi-task fusion model.

[0145] In this embodiment of the invention, assuming r(n) is the far-end speech and x(n) is the near-end speech, when the far-end speech passes through an unknown echo path w(n), it generates an echo signal y(n) = r(n) * w(n). The signal received by the near-end microphone is d(n) = y(n) + x(n). The adaptive filter w^(n) in the near-end linear echo cancellation module estimates the linear echo signal y^(n) with reference to the far-end speech r(n), and subtracts it from the signal d(n) received by the near-end microphone to obtain an error signal e(n) = d(n) - y^(n). Here, e(n) should be close to the near-end speech x(n). Without considering the near-end speech, the smaller the value of the error signal, the closer the echo path estimated by the adaptive filter is to the real echo path.

[0146] In this embodiment of the invention, steps S101 to S105 above complete the preliminary data processing, and the specific structural diagram is as follows. Figure 2 As shown, after the first near-end speech x(n) and the first far-end speech r(n) are preprocessed, they are then subjected to delay detection (TDE) to generate the second near-end speech and the second far-end speech. The second near-end speech and the second far-end speech are then subjected to short-time Fourier transform (STFT) to generate the third near-end speech and the third far-end speech. The third near-end speech and the third far-end speech are then input into the linear echo cancellation (LACE) module to generate the linear echo cancellation error signal and the linear echo signal.

[0147] Step S106: Input the fourth proximal speech, the fourth distal speech, the error signal of linear echo cancellation, and the linear echo signal into the two-stage neural network multi-task fusion model to be trained, and determine the joint loss function.

[0148] In one possible implementation, the fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are input into the first stage of the two-stage neural network multi-task fusion model to be trained to generate a first-stage training loss function and generate a first enhanced near-end speech and a first enhanced linear echo cancellation signal. The first stage of the two-stage neural network multi-task fusion model is used for preliminary noise reduction, preliminary echo cancellation, and removal of late reverberation. The fourth near-end speech, the error signal of the linear echo cancellation, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal are input into the second stage of the two-stage neural network multi-task fusion model to be trained to generate a second-stage training loss function. The second stage of the two-stage neural network multi-task fusion model is used to remove residual noise, residual echo, and early reverberation. A joint loss function is determined based on the first-stage training loss function and the second-stage training loss function.

[0149] In this embodiment of the invention, the first stage of the two-stage neural network multi-task fusion model to be trained is also used to reduce the discontinuity distortion and amplitude variation problems of the fourth proximal speech and the fourth distal speech. The first stage of the two-stage neural network multi-task fusion model to be trained can also be called a wake-up and automatic speech recognition (ASR) model. Traditional deep learning-based noise reduction models impair speech and cannot accurately identify speech content when speech is impaired. In the first stage, only 80% of the noise is removed to adapt to the wake-up and ASR model, that is, to enable the first stage to accurately identify speech content. The second stage of the two-stage neural network multi-task fusion model to be trained is used to reduce the distortion and overfitting problems generated by the first stage model, improve the sound quality of the output speech, and thus improve the user's auditory experience during voice communication.

[0150] The following two specific examples illustrate in detail how to obtain the first-stage training loss function and the second-stage training loss function: Specific Implementation Example 1

[0152] The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are input into the first stage of the two-stage neural network multi-task fusion model to be trained, generating the first-stage training loss function, such as... Figure 3 As shown, it specifically includes:

[0153] Step S301: Extract and concatenate features from the fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal to obtain the first speech feature.

[0154] Specifically, the fourth near-end speech, the fourth far-end speech, the error signal from linear echo cancellation, and the linear echo signal are subjected to short-time Fourier transform, and the real and imaginary parts of the above four signals are retained. Feature splicing is performed on the channel dimension to determine the feature dimension of the first speech feature as (B, T, C, F), where B represents the size of each batch, T represents the total number of frames of the training audio, C represents the number of feature channels (e.g., 8 channels), and F represents the number of frequency points (e.g., 257 items). This is only an illustrative example and should be determined according to the actual situation.

[0155] Step S302: Input the first speech features into the first encoding layer, the first bottleneck layer and the first decoding layer to generate the first intermediate speech data.

[0156] In one possible implementation, the network structures of the first coding layer and the first decoding layer are based on open-source network structures such as Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement (DCCRN), Dual-Path Convolution Recurrent Network (DPCRN), Grouped Temporal Convolutional Recurrent Network (GTCRN), and DEEPVQE. The first coding layer consists of four small coding modules, each of which consists of 2-convolution (Conv2d), batch normalization (Batch Norm), exponential linear unit (ELU) units, and grouped temporal convolution blocks (GT-Conv blocks).

[0157] Step S303: Input the first intermediate speech data and the fourth near-end speech data into the first complex spectrum convolutional mapping layer to generate the first enhanced near-end speech data, and input the first intermediate speech data and the fourth far-end speech data into the second complex spectrum convolutional mapping layer to generate the first enhanced linear echo cancellation signal.

[0158] Specifically, the first complex spectral convolving mapping The first CCM layer and the second complex spectral convolutional mapping layer each consist of two stages. First, taking the first CCM layer as an example, a complex mask is constructed by convolving the output features of the first decoding layer with the input fourth near-end speech. The convolution output is divided into three weighted components, each representing the weight of a 120° rotation vector in the complex plane, resulting in enhanced complex domain mask features. These complex domain mask features are then processed using a multi-kernel convolutional filter in both time and frequency dimensions to generate the first enhanced near-end speech. Second, taking the second CCM layer as an example, a complex mask is constructed by convolving the output features of the first decoding layer with the input fourth far-end speech. The convolution output is divided into three weighted components, each representing the weight of a 120° rotation vector in the complex plane, resulting in enhanced complex domain mask features. These complex domain mask features are then processed using a multi-kernel convolutional filter in both time and frequency dimensions to generate the first enhanced linear echo cancellation signal.

[0159] In one possible implementation, the first CCM layer may also be represented as a first MIC CCM layer, and the second CCM layer may also be represented as a first LAEC CCM layer.

[0160] Step S304: Input the first enhanced near-end speech and the first enhanced linear echo cancellation signal into the first fusion module to generate the first stage output features.

[0161] Specifically, the first enhanced near-end speech representation output by the first MIC CCM layer is X. mic The first enhanced linear echo cancellation signal output by the first LAEC CCM layer is represented as X. laec The first enhanced near-end speech X mic With the first enhanced linear echo cancellation signal X laec Fusion is performed along the channel dimension to generate a fused feature X. concate =(X laec ,X mic) The fusion feature X concate Perform convolution, and generate X after one convolution module. fusion =conv1(X concate After passing through a second convolution module, w = sigmoid(conv2(X) is generated. concate In this context, w represents the weight, and the parameters of the first and second convolutions are the same; the first enhanced near-end speech is generated after processing. The enhanced linear echo cancellation signal is generated after processing. For the Y LAEC After performing average pooling and max pooling, the concatenation yields the Y. LAEC Adaptive weight mask, thus determining Y MIC The adaptive weight is 1-Mask, and the weight is... The first-stage output feature identified as the output of the first fusion module is illustrated in the following diagram. Figure 4 As shown.

[0162] Step S305: Determine the training loss function for the first stage based on the output features of the first stage.

[0163] Specifically, the output features of the first stage are subjected to inverse short-time Fourier transform to generate the target speech of the first stage; the training loss function of the first stage is determined based on the target speech of the first stage and the pre-acquired clean speech.

[0164] In one possible implementation, assume that the clean speech is s(n), and the corresponding frequency domain signal is S(k,f); the target speech in the first stage is t(n), and the corresponding frequency domain signal is T(k,f). The target speech in the first stage is a first near-end speech with 80% noise reduction, and the clean speech is a first near-end speech with full noise reduction. This is only for illustrative purposes. The target speech in the first stage is... The corresponding frequency domain signal is The mean square error (MSE) and scale-invariant source-to-noise ratio (SISNR) loss functions commonly used in speech denoising and separation domains are employed to determine the L1 training loss function for the first stage, as follows:

[0165]

[0166] Wherein, α can be set to 0.5.

[0167] In this embodiment of the invention, the structural diagram of the first stage is as follows: Figure 5As shown, assuming that the fourth near-end speech (Mic), the fourth far-end speech (Far end), the error signal of linear echo cancellation (Error), and the linear echo signal (Lref) are input into the first feature extraction module 501 of the two-stage neural network multi-task fusion model to be trained, and then sequentially input into the first encoding layer 502, the first bottleneck layer 503, and the first decoding layer 504 to generate first intermediate speech data, the first intermediate speech data and the fourth near-end speech are input into the first CCM layer 505 to generate the first enhanced near-end speech, and the first intermediate speech data and the fourth far-end speech are input into the second complex spectrum convolution CCM layer 506 to generate the first enhanced linear echo cancellation signal; the first enhanced near-end speech and the first enhanced linear echo cancellation signal are input into the first fusion module 507 to generate the first stage output features.

[0168] Wherein, the first encoding layer can also be called the first encoding module, the first bottleneck layer can also be called the first bottleneck module, the first decoding layer can also be called the first decoding module, the first complex spectrum convolution mapping (CCM) layer can also be called the first complex spectrum convolution mapping (CCM) module, and the second complex spectrum convolution mapping (CCM) layer can also be called the second complex spectrum convolution mapping (CCM) module.

[0169] In this embodiment of the invention, a schematic diagram of determining the training loss function in the first stage is shown below. Figure 6 As mentioned above, in Figure 5 Based on this, an Inverse Short Time Fourier Transform (ISTFT) is added. Specifically, the output features of the first stage are subjected to ISTFT to generate the first stage target speech (FS SPEECH). The first stage training loss function (FS LOSS) is determined based on the first stage target speech (FS SPEECH) and the pre-acquired clean speech (Clean Speech). Specific Implementation Example 2

[0171] The process involves inputting the fourth near-end speech, the linear echo cancellation error signal, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal into the second stage of the two-stage neural network multi-task fusion model to be trained, generating the second-stage training loss function, such as... Figure 7 As shown, it specifically includes:

[0172] Step S701: Extract and concatenate features from the fourth near-end speech, the linear echo cancellation error signal, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal to obtain the second speech features.

[0173] Specifically, the fourth near-end speech, the error signal of the linear echo cancellation, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal are subjected to short-time Fourier transform, and the real and imaginary parts of the above four signals are retained. Feature splicing is performed on the channel dimension to determine the feature dimension of the first speech feature as (B, T, C, F), where B represents the size of each batch, T represents the total number of frames of the training audio, C represents the number of feature channels, for example, the number of channels is 8, and F represents the number of frequency points, for example, the number of single items is 257. This is only an illustrative example and the specific determination should be made according to the actual situation.

[0174] Step S702: Input the second speech features into the second coding layer, the second bottleneck layer and the second decoding layer to generate the second intermediate speech data.

[0175] In one possible implementation, the network structures of the second coding layer and the second decoding layer are based on open-source network structures such as Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement (DCCRN), Dual-Path Convolution Recurrent Network (DPCRN), Grouped Temporal Convolutional Recurrent Network (GTCRN), and DEEPVQE. The first coding layer consists of four small coding modules, each of which consists of 2-convolution (Conv2d), batch normalization (Batch Norm), exponential linear unit (ELU) units, and grouped temporal convolution blocks (GT-Conv blocks).

[0176] Step S703: Input the second intermediate speech data and the first enhanced near-end speech into the third complex spectrum convolutional mapping layer to generate the second enhanced near-end speech, and input the second intermediate speech data and the first enhanced linear echo cancellation signal into the fourth complex spectrum convolutional mapping layer to generate the second enhanced linear echo cancellation signal.

[0177] Specifically, the third complex spectral convolving mapping... The mask (CCM) layer and the fourth complex spectral convolutional mapping layer, the third CCM layer and the fourth CCM layer are both composed of two stages. First, taking the third CCM layer as an example, it is described in detail that by convolving the output features of the second decoding layer with the first enhanced near-end speech, the convolution output is divided into three weight components to construct a complex mask. Each component is the weight of a 120° rotation vector in the complex plane, resulting in the enhanced complex domain mask features. The complex domain mask features are then processed by a multi-kernel convolutional filter in the time and frequency dimensions to generate the second enhanced near-end speech. Second, taking the fourth CCM layer as an example, it is described in detail that by convolving the output features of the second decoding layer with the first enhanced linear echo cancellation signal, the convolution output is divided into three weight components to construct a complex mask. Each component is the weight of a 120° rotation vector in the complex plane, resulting in the enhanced complex domain mask features. The complex domain mask features are then processed by a multi-kernel convolutional filter in the time and frequency dimensions to generate the second enhanced linear echo cancellation signal.

[0178] In one possible implementation, the third CCM layer can also be represented as a second MIC CCM layer, and the fourth CCM layer can also be represented as a second LAEC CCM layer.

[0179] Step S704: Input the second enhanced near-end speech and the second enhanced linear echo cancellation signal into the second fusion module to generate the second-stage output features.

[0180] Specifically, the second enhanced near-end speech representation output by the second MIC CCM layer is X. mic The enhanced linear echo cancellation signal output by the second LAEC CCM layer is represented as X. laec The second enhanced near-end speech X mic With the second enhanced linear echo cancellation signal X laec Fusion is performed along the channel dimension to generate a fused feature X. concate =(X laec ,X mic) The fusion feature X concate Perform convolution, and generate X after one convolution. fusion =conv1(X concate After a second convolution, w = sigmoid(conv2(X) is generated. concate In this context, w represents the weight, and the parameters of the first and second convolutions are the same; the second enhanced near-end speech is generated after processing. The enhanced linear echo cancellation signal is generated after processing. For the Y LAEC After performing average pooling and max pooling, the concatenation yields the Y. LAEC Adaptive weight mask, thus determining Y MIC The adaptive weight is 1-Mask, and the weight is... The second-stage output feature identified as the output of the second fusion module is illustrated in the schematic diagram of the specific processing procedure. Figure 4 They are the same, only the input is different, which will not be elaborated here.

[0181] Step S705: Determine the training loss function for the second stage based on the output features of the second stage.

[0182] Specifically, the output features of the second stage are subjected to inverse short-time Fourier transform to generate the target speech of the second stage; the training loss function of the second stage is determined based on the target speech of the second stage and the pre-acquired clean speech.

[0183] In one possible implementation, suppose the clean speech is s(n), the corresponding frequency domain signal is S(k,f), and the target speech in the second stage is... Corresponding frequency domain signal The first-stage training loss function L1 is determined using commonly used loss functions for speech denoising and separation, namely mean square error (MSE) and scale-invariant source-to-noise ratio (SISNR), as follows:

[0184]

[0185] Wherein, α can be set to 0.5.

[0186] In this embodiment of the invention, the structural schematic diagram of the second stage is as follows: Figure 8As shown, assuming that the fourth near-end speech Mic, the linear echo cancellation error signal Error, the first enhanced near-end speech Mic CCM, and the first enhanced linear echo cancellation signal Lace CCM are input into the second feature extraction module 801 of the two-stage neural network multi-task fusion model to be trained, and then sequentially input into the second encoding layer 802, the second bottleneck layer 803, and the second decoding layer 804 to generate second intermediate speech data, the second intermediate speech data and the first enhanced near-end speech are input into the third CCM layer 805 to generate the second enhanced near-end speech, and the second intermediate speech data and the first enhanced linear echo cancellation signal are input into the fourth complex spectrum convolution CCM layer 806 to generate the second enhanced linear echo cancellation signal; the second enhanced near-end speech and the second enhanced linear echo cancellation signal are input into the second fusion module 807 to generate the second-stage output features.

[0187] The second encoding layer can also be called the second encoding module, the second bottleneck layer can also be called the second bottleneck module, the second decoding layer can also be called the second decoding module, the third complex spectrum convolution mapping (CCM) layer can also be called the third complex spectrum convolution mapping (CCM) module, and the fourth complex spectrum convolution mapping (CCM) layer can also be called the fourth complex spectrum convolution mapping (CCM) module.

[0188] In this embodiment of the invention, a schematic diagram of determining the training loss function in the second stage is shown below. Figure 9 As mentioned above, in Figure 8 Based on this, an Inverse Short Time Fourier Transform (ISTFT) is added. Specifically, the output features of the second stage are subjected to ISTFT to generate the second stage target speech (TS SPEECH). The second stage training loss function (TS LOSS) is determined based on the second stage target speech (TS SPEECH) and the pre-acquired clean speech (Clean Speech).

[0189] In this embodiment of the invention, the first-stage training loss function FS LOSS and the second-stage training loss function TS LOSS are determined through the above-described specific embodiments one and two.

[0190] Step S107: Adjust the two-stage neural network multi-task fusion model to be trained according to the joint loss function to generate a two-stage neural network multi-task fusion model.

[0191] Specifically, the first weight and the second weight of the first stage training loss function are obtained; the first product of the first weight and the first stage training loss function are determined, and the second product of the second weight and the second stage training loss function are determined; the sum of the first product and the second product is determined as the joint loss function.

[0192] In one possible implementation, assuming the first weight is 0.3, the second weight is 0.7, the first-stage training loss function is L1, and the second-stage training loss function is L2, then the joint loss function is as follows:

[0193] L loss =0.3*L1+0.7*L2

[0194] The values ​​of the first weight and the second weight are determined according to the actual situation, and are only illustrative here.

[0195] In this embodiment of the invention, in conjunction with the above Figure 6 and stated Figure 9 The structural diagram for generating the joint loss function TOTALLOSS is determined, as follows: Figure 10 As shown.

[0196] Step S108: Compress the two-stage neural network multi-task fusion model to generate a target two-stage neural network multi-task fusion model, wherein the two-stage neural network multi-task fusion model is used to achieve echo cancellation and noise reduction.

[0197] Specifically, the model compression process is as follows: Figure 11 As shown, it includes:

[0198] Step S1101: Quantize the two-stage neural network multi-task fusion model online to generate a two-stage neural network multi-task fusion model with Torch model structure.

[0199] Specifically, the two-stage neural network multi-task fusion model is quantized online using QAT, and the weights are quantized to INT_8 or Float_16 to generate a two-stage neural network multi-task fusion model with a Torch model structure.

[0200] Step S1102: The two-stage neural network multi-task fusion model of the Torch model structure is converted to ONNX through an open neural network to generate a two-stage neural network multi-task fusion model of the ONNX model structure.

[0201] Step S1103: The two-stage neural network multi-task fusion model of the ONNX model structure is processed by the MNN inference engine to generate the target two-stage neural network multi-task fusion model.

[0202] Through the above embodiments, the first stage of the two-stage neural network multi-task fusion model performs preliminary noise reduction, preliminary echo cancellation, and removal of late reverberation to avoid excessive noise reduction causing speech distortion and inaccurate speech recognition. Then, the second stage removes residual noise, residual echo, and early reverberation to obtain a speech signal for voice communication. This achieves acoustic echo cancellation and noise reduction even with extremely low signal-to-noise ratio, improving the voice interaction and communication performance of terminal devices. Furthermore, the two-stage neural network multi-task fusion model also features low latency, weak dependence on linear AEC, and a greater improvement in signal-to-noise ratio. Through model compression, quantization, and other strategies, the model is deployed on the ARM platform using the MNN framework, giving the two-stage neural network multi-task fusion model the characteristics of low latency and low computational cost.

[0203] In this embodiment of the invention, after step S108, the method further includes other steps, specifically as follows: Figure 12 As shown, the details are as follows:

[0204] Step S109: Deploy the two-stage neural network multi-task fusion model to the terminal side.

[0205] Specifically, the two-stage neural network multi-task fusion model is deployed on an Advanced RISC Machine (ARM) or Digital Signal Processing (DSP) platform on the terminal side, and then the entire voice link is integrated through engineering integration to form an end-to-end voice signal streaming processing system. Since the parameters of the target two-stage neural network multi-task fusion model are relatively small, it can be deployed on a terminal side with low computing power.

[0206] In this embodiment of the invention, the target two-stage neural network multi-task fusion model is deployed on the terminal side, which can realize functions such as echo cancellation, speech noise reduction, nonlinear echo suppression, and dereverberation, and complete the speech signal processing flow.

[0207] In one possible implementation, after completing step S109 of deploying the target two-stage neural network multi-task fusion model to the terminal side, the method further includes other steps, specifically as follows: Figure 13 As shown:

[0208] Step S110: Obtain the signal to be processed.

[0209] The signal to be processed is a voice signal.

[0210] Step S111: Input the signal to be processed into the target two-stage neural network multi-task fusion model to generate the target signal.

[0211] Specifically, by inputting the signal to be processed into the target two-stage neural network multi-task fusion model, the signal to be processed can be directly generated into the target signal. This improves the ability to perform echo cancellation, speech noise reduction, and dereverberation on the signal to be processed even under conditions of extremely low signal-to-return ratio, unstable delay, and variable echo propagation paths, thereby enhancing the voice interaction performance and robustness of the terminal device.

[0212] In this embodiment of the invention, the system structure for training the target two-stage neural network multi-task fusion model is briefly described, as follows: Figure 14 As shown, it includes: a data acquisition and simulation dataset construction module 1401, a model training module 1402, and a model compression and edge deployment module 1403. This is only an illustrative example, and the specific system structure should be constructed according to the actual situation.

[0213] In this embodiment of the invention, an echo cancellation and noise reduction device is provided, such as... Figure 15 As shown, it specifically includes: an acquisition unit 1501, a processing unit 1502, a detection unit 1503, a generation unit 1504, and a determination unit 1505;

[0214] The acquisition unit 1501 is used to acquire a first near-end speech and a first far-end speech; the processing unit 1502 is used to preprocess the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech; the detection unit 1503 is used to perform delay detection on the second near-end speech and the second far-end speech to generate aligned third near-end speech and third far-end speech; the processing unit 1502 is further used to perform short-time Fourier transform on the third near-end speech and the third far-end speech to generate fourth near-end speech and fourth far-end speech; the generation unit 1504 is used to perform linear echo cancellation on the fourth near-end speech and the fourth far-end speech to generate linear echo cancellation... The determination unit 1505 is used to input the fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal into the two-stage neural network multi-task fusion model to be trained, and determine the joint loss function; the generation unit 1504 is further used to adjust the two-stage neural network multi-task fusion model to be trained according to the joint loss function, and generate a two-stage neural network multi-task fusion model; the generation unit 1504 is further used to compress the two-stage neural network multi-task fusion model to generate a target two-stage neural network multi-task fusion model, wherein the two-stage neural network multi-task fusion model is used to achieve echo cancellation and noise reduction.

[0215] Furthermore, the device also includes:

[0216] The deployment unit is used to deploy the two-stage neural network multi-task fusion model to the terminal side.

[0217] Furthermore, the processing unit is specifically used for:

[0218] The first near-end speech and the first far-end speech are pre-emphasized, power frequency interference is removed, and windowing is applied to generate the second near-end speech and the second far-end speech.

[0219] Furthermore, the determining unit is specifically used for:

[0220] The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are input into the first stage of the two-stage neural network multi-task fusion model to be trained, generating the first stage training loss function and generating the first enhanced near-end speech and the first enhanced linear echo cancellation signal. The first stage of the two-stage neural network multi-task fusion model is used for preliminary noise reduction, preliminary echo cancellation, and removal of late reverberation.

[0221] The fourth near-end speech, the error signal of the linear echo cancellation, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal are input into the second stage of the two-stage neural network multi-task fusion model to be trained, generating the second-stage training loss function. The second stage of the two-stage neural network multi-task fusion model is used to remove residual noise, residual echo, and early reverberation.

[0222] The joint loss function is determined based on the training loss function of the first stage and the training loss function of the second stage.

[0223] Furthermore, the determining unit is specifically used for:

[0224] The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are subjected to feature extraction and feature splicing to obtain the first speech feature;

[0225] The first speech feature is input into the first encoding layer, the first bottleneck layer and the first decoding layer to generate the first intermediate speech data;

[0226] The first intermediate speech data and the fourth near-end speech data are input into the first complex spectrum convolutional mapping layer to generate the first enhanced near-end speech data, and the first intermediate speech data and the fourth far-end speech data are input into the second complex spectrum convolutional mapping layer to generate the first enhanced linear echo cancellation signal.

[0227] The first enhanced near-end speech and the first enhanced linear echo cancellation signal are input into the first fusion module to generate the first stage output features;

[0228] The training loss function for the first stage is determined based on the output features of the first stage.

[0229] Furthermore, the determining unit is specifically used for:

[0230] The first stage output features are subjected to inverse short-time Fourier transform to generate the first stage target speech;

[0231] The training loss function for the first stage is determined based on the target speech in the first stage and the pre-acquired clean speech.

[0232] Furthermore, the determining unit is specifically used for:

[0233] The second speech feature is obtained by extracting and concatenating features from the fourth near-end speech, the linear echo cancellation error signal, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal.

[0234] The second speech feature is input into the second coding layer, the second bottleneck layer, and the second decoding layer to generate the second intermediate speech data.

[0235] The second intermediate speech data and the first enhanced near-end speech are input into the third complex spectrum convolutional mapping layer to generate the second enhanced near-end speech, and the second intermediate speech data and the first enhanced linear echo cancellation signal are input into the fourth complex spectrum convolutional mapping layer to generate the second enhanced linear echo cancellation signal.

[0236] The second enhanced near-end speech and the second enhanced linear echo cancellation signal are input into the second fusion module to generate the second-stage output features;

[0237] The training loss function for the second stage is determined based on the output features of the second stage.

[0238] Furthermore, the determining unit is specifically used for:

[0239] The output features of the second stage are subjected to inverse short-time Fourier transform to generate the target speech of the second stage;

[0240] The second-stage training loss function is determined based on the target speech in the second stage and the pre-acquired clean speech.

[0241] Furthermore, the determining unit is specifically used for:

[0242] Obtain the first weight of the first stage training loss function and the second weight of the first stage training loss function;

[0243] Determine the first product of the first weight and the first stage training loss function, and the second product of the second weight and the second stage training loss function;

[0244] The sum of the first product and the second product is determined as the joint loss function.

[0245] Furthermore, the generation unit is specifically used for:

[0246] The two-stage neural network multi-task fusion model is quantized online to generate a two-stage neural network multi-task fusion model with a Torch model structure.

[0247] The two-stage neural network multi-task fusion model of the Torch model structure is converted into ONNX through open neural network transformation to generate a two-stage neural network multi-task fusion model of the ONNX model structure.

[0248] The two-stage neural network multi-task fusion model of the ONNX model structure is passed through the MNN inference engine to generate the target two-stage neural network multi-task fusion model.

[0249] Furthermore, the acquisition unit is also used to: acquire the signal to be processed;

[0250] The generation unit is further configured to: input the signal to be processed into the target two-stage neural network multi-task fusion model to generate the target signal.

[0251] Figure 16 This is a schematic diagram of the structure of the electronic device described in an embodiment of the present invention. Figure 16 As shown, it includes a general computer hardware architecture, which includes at least a processor 1601 and a memory 1602. The processor 1601 and the memory 1602 are connected via a bus 1603. The memory 1602 is adapted to store instructions or programs executable by the processor 1601. The processor 1601 can be a standalone microprocessor or a collection of one or more microprocessors. Thus, the processor 1601 executes the instructions stored in the memory 1602 to perform the method flow of the embodiments of the present invention as described above, thereby realizing data processing and control of other devices. The bus 1603 connects the above-mentioned components together, and also connects the above-mentioned components to the display controller 1604, the display device, and the input / output (I / O) device 1605. The input / output (I / O) device 1605 can be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, the input / output device 1605 is connected to the system via an input / output (I / O) controller 1606.

[0252] The instructions stored in memory 1602 are executed by at least one processor 1601 to achieve the following: acquiring a first near-end speech and a first far-end speech; preprocessing the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech; performing delay detection on the second near-end speech and the second far-end speech to generate aligned third near-end speech and third far-end speech; performing short-time Fourier transform on the third near-end speech and the third far-end speech to generate fourth near-end speech and fourth far-end speech; and performing linear echo cancellation on the fourth near-end speech and the fourth far-end speech to generate linear echo... The error signal of sound cancellation and the linear echo signal are input into a two-stage neural network multi-task fusion model to be trained, and a joint loss function is determined. The two-stage neural network multi-task fusion model to be trained is adjusted according to the joint loss function to generate a two-stage neural network multi-task fusion model. The two-stage neural network multi-task fusion model is compressed to generate a target two-stage neural network multi-task fusion model, wherein the two-stage neural network multi-task fusion model is used to achieve echo cancellation and noise reduction.

[0253] Specifically, the electronic device includes: one or more processors 1601 and memory 1602. Figure 16 Take processor 1601 as an example. Processor 1601 and memory 1602 can be connected via a bus or other means. Figure 16 Taking a bus connection as an example, memory 1602, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Processor 1601 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in memory 1602, thereby implementing the aforementioned method for determining echo cancellation and noise reduction.

[0254] Memory 1602 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store an option list, etc. Furthermore, memory 1602 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 1602 may optionally include memory remotely located relative to processor 1601, and these remote memories may be connected to external devices via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0255] One or more modules are stored in memory 1602 and, when executed by one or more processors 1601, perform the echo cancellation and noise reduction methods in any of the above method embodiments.

[0256] As those skilled in the art will recognize, various aspects of the embodiments of the present invention can be implemented as a system, method, or computer program product. Therefore, various aspects of the embodiments of the present invention can take the form of a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or an implementation combining software and hardware aspects, which may generally be referred to herein as a "circuit," "module," or "system." Furthermore, various aspects of the embodiments of the present invention can take the form of a computer program product implemented in one or more computer-readable media having computer-readable program code implemented thereon.

[0257] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, (but not limited to) an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media will include: an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the context of embodiments of the present invention, a computer-readable storage medium can be any tangible medium capable of containing or storing a program used by or in conjunction with an instruction execution system, device, or apparatus.

[0258] Computer-readable signal media may include propagated digital signals having computer-readable program code implemented therein, such as in baseband or as part of a carrier wave. Such propagated signals may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and can communicate, propagate, or transmit a program used by or in conjunction with an instruction execution system, device, or apparatus.

[0259] Program code implemented on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, or any suitable combination thereof.

[0260] Computer program code for performing operations relating to various aspects of embodiments of the present invention can be written in any combination of one or more programming languages, including: object-oriented programming languages ​​such as Java, Smalltalk, C++, etc.; and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The program code can be executed as a standalone software package entirely on the user's computer, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet provided by an Internet service provider).

[0261] The flowchart illustrations and / or block diagrams of the methods, apparatus (systems), and computer program products according to embodiments of the present invention describe various aspects of the embodiments of the present invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions (executed via the processor of the computer or other programmable data processing apparatus) create means for implementing the functions / actions specified in the flowchart and / or block diagram blocks or blocks.

[0262] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing apparatus or other means to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of writing that includes instructions that implement the functions / actions specified in flowchart and / or block diagram blocks or blocks.

[0263] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operable steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide for implementing the functions / actions specified in flowchart and / or block diagram blocks or blocks.

[0264] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for echo cancellation and noise reduction, characterized in that, The method includes: Acquire the first near-end speech and the first far-end speech; The first near-end speech and the first far-end speech are preprocessed to generate the preprocessed second near-end speech and the second far-end speech. Delay detection is performed on the second near-end speech and the second far-end speech to generate aligned third near-end speech and third far-end speech; The third proximal speech and the third distal speech are respectively subjected to short-time Fourier transform to generate the fourth proximal speech and the fourth distal speech; Linear echo cancellation is performed on the fourth near-end speech and the fourth far-end speech to generate an error signal for linear echo cancellation and a linear echo signal. The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are input into the two-stage neural network multi-task fusion model to be trained, and the joint loss function is determined. The two-stage neural network multi-task fusion model to be trained is adjusted according to the joint loss function to generate the two-stage neural network multi-task fusion model. The two-stage neural network multi-task fusion model is compressed to generate a target two-stage neural network multi-task fusion model, wherein the two-stage neural network multi-task fusion model is used to achieve echo cancellation and noise reduction.

2. The method according to claim 1, characterized in that, The method further includes: The two-stage neural network multi-task fusion model is deployed to the terminal side.

3. The method according to claim 1, characterized in that, The step of preprocessing the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech specifically includes: The first near-end speech and the first far-end speech are pre-emphasized, power frequency interference is removed, and windowing is applied to generate the second near-end speech and the second far-end speech.

4. The method according to claim 1, characterized in that, The step of inputting the fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal into the two-stage neural network multi-task fusion model to be trained, and determining the joint loss function, specifically includes: The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are input into the first stage of the two-stage neural network multi-task fusion model to be trained, generating the first stage training loss function and generating the first enhanced near-end speech and the first enhanced linear echo cancellation signal. The first stage of the two-stage neural network multi-task fusion model is used for preliminary noise reduction, preliminary echo cancellation, and removal of late reverberation. The fourth near-end speech, the error signal of the linear echo cancellation, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal are input into the second stage of the two-stage neural network multi-task fusion model to be trained, generating the second-stage training loss function. The second stage of the two-stage neural network multi-task fusion model is used to remove residual noise, residual echo, and early reverberation. The joint loss function is determined based on the training loss function of the first stage and the training loss function of the second stage.

5. The method according to claim 4, characterized in that, The step of inputting the fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal into the first stage of the two-stage neural network multi-task fusion model to be trained, and generating the first-stage training loss function, specifically includes: The fourth near-end speech, the fourth far-end speech, the error signal of the linear echo cancellation, and the linear echo signal are subjected to feature extraction and feature splicing to obtain the first speech feature; The first speech feature is input into the first encoding layer, the first bottleneck layer and the first decoding layer to generate the first intermediate speech data; The first intermediate speech data and the fourth near-end speech data are input into the first complex spectrum convolutional mapping layer to generate the first enhanced near-end speech data, and the first intermediate speech data and the fourth far-end speech data are input into the second complex spectrum convolutional mapping layer to generate the first enhanced linear echo cancellation signal. The first enhanced near-end speech and the first enhanced linear echo cancellation signal are input into the first fusion module to generate the first stage output features; The training loss function for the first stage is determined based on the output features of the first stage.

6. The method according to claim 5, characterized in that, The step of determining the training loss function for the first stage based on the output features of the first stage specifically includes: The first stage output features are subjected to inverse short-time Fourier transform to generate the first stage target speech; The training loss function for the first stage is determined based on the target speech in the first stage and the pre-acquired clean speech.

7. The method according to claim 4, characterized in that, The step of inputting the fourth near-end speech, the linear echo cancellation error signal, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal into the second stage of the two-stage neural network multi-task fusion model to be trained, and generating the second-stage training loss function, specifically includes: The second speech feature is obtained by extracting and concatenating features from the fourth near-end speech, the linear echo cancellation error signal, the first enhanced near-end speech, and the first enhanced linear echo cancellation signal. The second speech feature is input into the second coding layer, the second bottleneck layer, and the second decoding layer to generate the second intermediate speech data. The second intermediate speech data and the first enhanced near-end speech are input into the third complex spectrum convolutional mapping layer to generate the second enhanced near-end speech, and the second intermediate speech data and the first enhanced linear echo cancellation signal are input into the fourth complex spectrum convolutional mapping layer to generate the second enhanced linear echo cancellation signal. The second enhanced near-end speech and the second enhanced linear echo cancellation signal are input into the second fusion module to generate the second-stage output features; The training loss function for the second stage is determined based on the output features of the second stage.

8. The method according to claim 7, characterized in that, The step of determining the second-stage training loss function based on the second-stage output features specifically includes: The output features of the second stage are subjected to inverse short-time Fourier transform to generate the target speech of the second stage; The second-stage training loss function is determined based on the target speech in the second stage and the pre-acquired clean speech.

9. The method according to claim 4, characterized in that, The step of determining the joint loss function based on the first-stage training loss function and the second-stage training loss function specifically includes: Obtain the first weight of the first stage training loss function and the second weight of the first stage training loss function; Determine the first product of the first weight and the first stage training loss function, and the second product of the second weight and the second stage training loss function; The sum of the first product and the second product is determined as the joint loss function.

10. The method according to claim 1, characterized in that, The step of compressing the two-stage neural network multi-task fusion model to generate the target two-stage neural network multi-task fusion model specifically includes: The two-stage neural network multi-task fusion model is quantized online to generate a two-stage neural network multi-task fusion model with a Torch model structure. The two-stage neural network multi-task fusion model of the Torch model structure is converted into ONNX through open neural network transformation to generate a two-stage neural network multi-task fusion model of the ONNX model structure. The two-stage neural network multi-task fusion model of the ONNX model structure is passed through the MNN inference engine to generate the target two-stage neural network multi-task fusion model.

11. A method for echo cancellation and noise reduction, characterized in that, The method further includes: Acquire the signal to be processed; The signal to be processed is input into the target two-stage neural network multi-task fusion model to generate the target signal. The target two-stage neural network multi-task fusion model is obtained by the method of any one of claims 1-10. The target two-stage neural network multi-task fusion model is deployed on the terminal side to achieve echo cancellation and noise reduction.

12. An echo cancellation and noise reduction device, characterized in that, The device includes: An acquisition unit is used to acquire the first near-end speech and the first far-end speech; The processing unit is configured to preprocess the first near-end speech and the first far-end speech to generate preprocessed second near-end speech and second far-end speech. The detection unit is used to perform delay detection on the second near-end speech and the second far-end speech to generate aligned third near-end speech and third far-end speech; The processing unit is further configured to perform short-time Fourier transform on the third proximal speech and the third distal speech respectively to generate a fourth proximal speech and a fourth distal speech. The generation unit is used to perform linear echo cancellation on the fourth near-end speech and the fourth far-end speech, and generate an error signal for linear echo cancellation and a linear echo signal. The determining unit is used to input the fourth proximal speech, the fourth distal speech, the error signal of the linear echo cancellation, and the linear echo signal into the two-stage neural network multi-task fusion model to be trained, and to determine the joint loss function; The generation unit is also used to adjust the two-stage neural network multi-task fusion model to be trained according to the joint loss function, and generate a two-stage neural network multi-task fusion model. The generation unit is further configured to compress the two-stage neural network multi-task fusion model to generate a target two-stage neural network multi-task fusion model, wherein the two-stage neural network multi-task fusion model is used to achieve echo cancellation and noise reduction.

13. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage The medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Unified deep neural network model for acoustic echo cancellation and residual echo suppression

    CN116547750A

  • Echo cancellation model training method and device, equipment and storage medium

    CN117219107A