Method, apparatus, electronic device and computer-readable medium for echo cancellation

By combining traditional signal processing and deep learning methods, using pre-trained filters and target models to eliminate linear and nonlinear echoes in audio signals, solving the shortcomings of the prior art in nonlinear echo scenarios, and achieving more efficient echo cancellation and noise suppression effects.

CN115083431BActive Publication Date: 2025-06-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210701124.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2025-06-10
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

In the prior art, when dealing with nonlinear echo scenes, the echo cancellation effect is not thorough enough, affecting the audio quality.

Method used

By obtaining the original audio signal and reference signal, using a pre-trained filter for linear echo filtering, the residual signal is obtained, and the residual signal is eliminated and the effective signal enhancement is enhanced by using the pre-trained target model to obtain the target audio signal.

Benefits of technology

While ensuring audio quality, it effectively eliminates echo and noise interference, improves user experience, and is suitable for a variety of scenarios, including karaoke and regular voice echo cancellation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115083431B_ABST
    Figure CN115083431B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, electronic device, and computer-readable medium for echo cancellation, belonging to the technical field of signal processing. The method includes: obtaining an original audio signal and a reference signal required for echo cancellation by a filter; inputting the original audio signal and the reference signal into the filter for linear echo filtering processing to obtain a residual signal; and performing interference signal cancellation and effective signal enhancement on the residual signal based on a target model to obtain a target audio signal, where the interference signal includes a non-linear echo signal and a noise signal. The present disclosure can improve the echo cancellation effect at a lower model complexity and ensure a higher audio quality by using a signal processing method to cancel the linear echo signal in the original audio signal and then using a deep learning model to cancel the non-linear echo signal and the noise signal in the remaining residual signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of signal processing, and in particular, to a method for eliminating echo, an apparatus for eliminating echo, an electronic device, and a computer-readable medium. Background Art

[0002] In recent years, with the rapid development of communication technologies, audio-visual products have continuously enriched the diversity of life and social interactions, such as online conferencing systems, karaoke systems, etc. In these application systems, the smoothness, integrity, and intelligibility of audio information content directly determine the communication quality between users, and the optimization and innovation of echo cancellation technology are indispensable behind these technologies.

[0003] The echo phenomenon is caused by the sound of the speaker being fed back to the microphone. If the echo cannot be effectively suppressed, the user can hear their own delayed sound, which will directly affect the intelligibility of the speech and bring a bad experience to the user. Therefore, AEC (Acoustic echo cancellation) technology plays a crucial role in improving audio quality.

[0004] Currently, the main method for echo cancellation is the adaptive filtering algorithm. However, in the scenario of non-linear echo, this method is not sufficient to completely eliminate the echo.

[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The purpose of the present disclosure is to provide a method for eliminating echo, an apparatus for eliminating echo, an electronic device, and a computer-readable medium, so as to at least to some extent improve the echo cancellation effect and ensure high audio quality.

[0007] According to a first aspect of the present disclosure, there is provided a method for eliminating echo, including:

[0008] Obtaining an original audio signal and a reference signal required for echo cancellation by a filter;

[0009] Inputting the original audio signal and the reference signal into the filter for linear echo filtering processing to obtain a residual signal;

[0010] Based on a target model, performing interference signal cancellation and effective signal enhancement on the residual signal to obtain a target audio signal, where the interference signal includes a non-linear echo signal and a noise signal.

[0011] In an exemplary embodiment of the present disclosure, the step of inputting the original audio signal and the reference signal into the filter for linear echo filtering to obtain a residual signal includes:

[0012] Convolving the reference signal with the filter weights of the filter to obtain an estimated value of the linear echo signal;

[0013] Obtaining the remaining residual signal in the original audio signal according to the difference between the original audio signal and the estimated value of the linear echo signal.

[0014] In an exemplary embodiment of the present disclosure, the step of performing interference signal cancellation and effective signal enhancement on the residual signal based on the target model to obtain the target audio signal includes:

[0015] Obtaining the frequency-domain features corresponding to the original audio signal according to the residual signal and the reference signal;

[0016] Inputting the frequency-domain features into a pre-trained target model to obtain an audio signal time-frequency mask corresponding to the target audio signal;

[0017] Obtaining the spectrum of the target audio signal according to the spectrum of the residual signal and the audio signal time-frequency mask;

[0018] Converting the spectrum of the target audio signal into a time-domain signal to obtain the enhanced target audio signal.

[0019] In an exemplary embodiment of the present disclosure, the training method of the filter includes:

[0020] Obtaining multiple sets of training data, where each set of training data includes original audio training data, reference training data corresponding to the original audio training data, and residual training data in the original audio training data;

[0021] Convolving the reference training data with the initial filter weights of the filter to obtain linear echo prediction data;

[0022] Obtaining residual prediction data according to the difference between the original audio training data and the linear echo prediction data;

[0023] Iterating the initial filter weights of the filter according to the difference between the residual prediction data and the residual training data to train the filter.

[0024] In an exemplary embodiment of the present disclosure, the training method of the target model includes:

[0025] Use the residual training data and the reference training data in each group of the training data as the first training data, and train an initial neural network model according to the first training data to obtain a first neural network model;

[0026] Use the original audio training data and the reference training data in each group of the training data as the second training data, and obtain mixed training data according to the first training data and the second training data according to a preset data ratio;

[0027] Train the first neural network model according to the mixed training data to obtain a target model.

[0028] In an exemplary embodiment of the present disclosure, the method further includes:

[0029] Obtain test data according to the reference training data and the original audio training data in each group of the training data, and test the target model according to the test data.

[0030] In an exemplary embodiment of the present disclosure, the training the initial neural network model according to the first training data to obtain a first neural network model includes:

[0031] Obtain corresponding frequency-domain feature training data according to the first training data, and input the frequency-domain feature training data into the initial neural network model to obtain an audio sample time-frequency mask and an interference sample time-frequency mask;

[0032] Obtain the sample spectrum of the enhanced target audio signal according to the sample spectrum of the residual training data and the audio sample time-frequency mask, and obtain the sample spectrum of the interference signal according to the sample spectrum of the residual training data and the interference sample time-frequency mask;

[0033] Determine a spectrum distance loss according to the sample spectrum of the enhanced target audio signal and the sample spectrum of the interference signal, and determine a signal-to-noise ratio loss according to the signal-to-noise ratio parameter of the enhanced target audio signal and the signal-to-noise ratio parameter of the interference signal;

[0034] Obtain an overall loss according to the spectrum distance loss and the signal-to-noise ratio loss, and iterate the neural network parameters in the initial neural network model according to the overall loss to obtain the trained first neural network model.

[0035] In an exemplary embodiment of the present disclosure, the determining the spectrum distance loss according to the sample spectrum of the enhanced target audio signal and the sample spectrum of the interference signal includes:

[0036] Determine an audio spectrum distance loss based on the sample spectrum of the residual training data and the sample spectrum of the enhanced target audio signal;

[0037] Obtain interference training data in the training data, and determine an interference spectrum distance loss based on the sample spectrum of the interference training data and the sample spectrum of the interference signal;

[0038] Obtain a spectrum distance loss based on the audio spectrum distance loss and the interference spectrum distance loss.

[0039] In an exemplary embodiment of the present disclosure, the method for generating the residual training data includes:

[0040] Obtain the original audio training data collected by the user - end microphone, where the original audio training data includes original speech training data and original music training data;

[0041] Perform variable - speed processing, reverberation processing, and delay processing on the original audio training data to obtain simulated echo data corresponding to the original audio training data;

[0042] Obtain the residual training data corresponding to the original audio training data according to the difference between the original audio training data and the corresponding simulated echo data.

[0043] According to a second aspect of the present disclosure, there is provided an echo cancellation device, including:

[0044] An audio signal acquisition module, configured to execute obtaining an original audio signal and a reference signal required for echo cancellation by a filter;

[0045] A linear echo cancellation module, configured to execute inputting the original audio signal and the reference signal into the filter for linear echo filtering processing to obtain a residual signal;

[0046] An interference signal cancellation module, configured to execute canceling interference signals and enhancing effective signals on the residual signal based on a target model to obtain a target audio signal, where the interference signals include non - linear echo signals and noise signals.

[0047] In an exemplary embodiment of the present disclosure, the linear echo cancellation module includes:

[0048] A linear echo estimation unit, configured to execute convolving the reference signal with the filter weights of the filter to obtain an estimated value of the linear echo signal;

[0049] A linear echo cancellation unit, configured to obtain a residual signal remaining in the original audio signal according to a difference between the original audio signal and an estimated value of the linear echo signal.

[0050] In an exemplary embodiment of the present disclosure, the interference signal cancellation module includes:

[0051] A frequency-domain feature determination unit, configured to obtain frequency-domain features corresponding to the original audio signal according to the residual signal and the reference signal;

[0052] A time-frequency mask determination unit, configured to input the frequency-domain features into a pre-trained target model to obtain an audio signal time-frequency mask corresponding to the target audio signal;

[0053] A signal spectrum determination unit, configured to obtain a spectrum of the target audio signal according to the spectrum of the residual signal and the audio signal time-frequency mask;

[0054] A signal spectrum conversion unit, configured to convert the spectrum of the target audio signal into a time-domain signal to obtain an enhanced target audio signal.

[0055] In an exemplary embodiment of the present disclosure, the echo cancellation device further includes a filter training module, and the filter training module includes:

[0056] A training data acquisition unit, configured to acquire multiple groups of training data, where each group of the training data includes original audio training data, reference training data corresponding to the original audio training data, and residual training data in the original audio training data;

[0057] A linear echo prediction data determination unit, configured to perform convolution on the reference training data and an initial filter weight of the filter to obtain linear echo prediction data;

[0058] A residual prediction data determination unit, configured to obtain residual prediction data according to a difference between the original audio training data and the linear echo prediction data;

[0059] A filter weight iteration unit, configured to iterate the initial filter weight of the filter according to a difference between the residual prediction data and the residual training data to train the filter.

[0060] In an exemplary embodiment of the present disclosure, the echo cancellation device further includes a target model training module, and the target model training module includes:

[0061] An initial neural network model training unit, configured to execute taking the residual training data and the reference training data in each group of the training data as first training data, and training an initial neural network model according to the first training data to obtain a first neural network model;

[0062] A mixed training data determination unit, configured to execute taking the original audio training data and the reference training data in each group of the training data as second training data, and obtaining mixed training data according to the first training data and the second training data according to a preset data ratio;

[0063] A first neural network model training unit, configured to execute training the first neural network model according to the mixed training data to obtain a target model.

[0064] In an exemplary embodiment of the present disclosure, the target model training module further includes:

[0065] A target model testing unit, configured to execute obtaining test data according to the reference training data and the original audio training data in each group of the training data, and testing the target model according to the test data.

[0066] In an exemplary embodiment of the present disclosure, the initial neural network model training unit includes:

[0067] A sample time-frequency mask determination unit, configured to execute obtaining corresponding frequency-domain feature training data according to the first training data, and inputting the frequency-domain feature training data into the initial neural network model to obtain an audio sample time-frequency mask and an interference sample time-frequency mask;

[0068] A sample spectrum determination unit, configured to execute obtaining the sample spectrum of the enhanced target audio signal according to the sample spectrum of the residual training data and the audio sample time-frequency mask, and obtaining the sample spectrum of the interference signal according to the sample spectrum of the residual training data and the interference sample time-frequency mask;

[0069] A network loss determination unit, configured to execute determining a spectrum distance loss according to the sample spectrum of the enhanced target audio signal and the sample spectrum of the interference signal, and determining a signal-to-noise ratio loss according to the signal-to-noise ratio parameter of the enhanced target audio signal and the signal-to-noise ratio parameter of the interference signal;

[0070] A neural network parameter iteration unit, configured to execute obtaining an overall loss according to the spectrum distance loss and the signal-to-noise ratio loss, and iterating the neural network parameters in the initial neural network model according to the overall loss to obtain the trained first neural network model.

[0071] In an exemplary embodiment of the present disclosure, the network loss determination unit includes:

[0072] An audio spectrum distance loss determination unit configured to determine an audio spectrum distance loss according to the sample spectrum of the residual training data and the sample spectrum of the enhanced target audio signal;

[0073] An interference spectrum distance loss determination unit configured to obtain interference training data in the training data and determine an interference spectrum distance loss according to the sample spectrum of the interference training data and the sample spectrum of the interference signal;

[0074] A spectrum distance loss determination unit configured to obtain a spectrum distance loss according to the audio spectrum distance loss and the interference spectrum distance loss.

[0075] In an exemplary embodiment of the present disclosure, the training data acquisition unit includes:

[0076] A raw audio training data acquisition unit configured to obtain the raw audio training data collected by a user-side microphone, where the raw audio training data includes raw speech training data and raw music training data;

[0077] A simulated echo data generation unit configured to perform variable speed processing, reverberation processing, and delay processing on the raw audio training data to obtain simulated echo data corresponding to the raw audio training data;

[0078] A residual training data generation unit configured to obtain residual training data corresponding to the raw audio training data according to the difference between the raw audio training data and the corresponding simulated echo data.

[0079] According to a third aspect of the present disclosure, there is provided an electronic device including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the echo cancellation method according to any one of the above.

[0080] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the echo cancellation method according to any one of the above.

[0081] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program implements the echo cancellation method according to any one of the above when executed by a processor.

[0082] The exemplary embodiments of the present disclosure may have the following beneficial effects:

[0083] In the echo cancellation method of the exemplary embodiments of the present disclosure, the linear echo signal in the original audio signal is cancelled by using a pre-trained filter, and then the non-linear echo signal and the noise signal in the remaining residual signal are cancelled by using a pre-trained target model to obtain an enhanced target audio signal. On the one hand, in the echo cancellation method of the exemplary embodiments of the present disclosure, by combining traditional signal processing with deep learning methods, the echo cancellation effect can be improved on the premise of a relatively small system complexity, ensuring a relatively high audio quality. On the other hand, not only can echoes be cancelled, but also noise interference can be cancelled, which can maximize the intelligibility of the audio and thus improve the user experience. The echo cancellation method in the exemplary embodiments of the present disclosure is not only applicable to the KTV scenario, but also supports voice echo cancellation in conventional scenarios, and can achieve a relatively good sound quality.

[0084] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0086] Figure 1 It shows a schematic flow chart of the echo cancellation method in a related embodiment of the present disclosure;

[0087] Figure 2 It shows a schematic diagram of an echo cancellation system according to a specific embodiment of the present disclosure.

[0088] Figure 3 It shows a schematic flow chart of cancelling the linear echo signal in the original audio signal in the exemplary embodiments of the present disclosure;

[0089] Figure 4 It shows a schematic flow chart of cancelling the interference signal in the residual signal in the exemplary embodiments of the present disclosure;

[0090] Figure 5 It shows a schematic flow chart of the training method of the filter in the exemplary embodiments of the present disclosure;

[0091] Figure 6 It shows a schematic flow chart of the training method of the target model in the exemplary embodiments of the present disclosure;

[0092] Figure 7Shows a schematic diagram of a neural network model according to a specific embodiment of the present disclosure.

[0093] Figure 8 Shows a schematic flow chart of training an initial neural network model with the first training data of the exemplary embodiment of the present disclosure;

[0094] Figure 9 Shows a block diagram of an echo cancellation device of the exemplary embodiment of the present disclosure;

[0095] Figure 10 Shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure. Specific Embodiments

[0096] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0097] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein.

[0098] The following exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be more complete and comprehensive, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to give a full understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0099] In addition, the accompanying drawings are only schematic diagrams of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0100] In some related embodiments, an adaptive filtering algorithm can be used to eliminate echo in an audio signal, such as the LMS (Least Mean Square) algorithm, the NLMS (Normalized Least Mean Square) algorithm, etc. The adaptive filtering algorithm estimates an approximate echo path to approximate the real echo path by adjusting the weight vector of the filter, and then obtains an estimated value of the echo signal. By subtracting the estimated echo signal from the original signal at the receiving end, an ideal clean signal can be obtained. However, for scenarios with non-linear echo, the adaptive filtering algorithm is not thorough enough in eliminating echo.

[0101] In some other related embodiments, echo cancellation of an audio signal can also be performed based on a neural network. The neural network has the ability of non-linear fitting and can directly eliminate linear echo and non-linear echo. By extracting effective features from the signal collected by the proximal microphone and the distal reference signal and serially inputting them into the neural network, the clean signal can be restored. However, this method places a greater test on the modeling ability of the model. Usually, a more complex or refined model is required to achieve better results. Therefore, it is difficult to be applied in actual scenarios.

[0102] This exemplary embodiment first provides a method for eliminating echo. Refer to Figure 1 As shown, the method for eliminating echo can include the following steps:

[0103] Step S110. Obtain the original audio signal and the reference signal required for echo cancellation by the filter.

[0104] Step S120. Input the original audio signal and the reference signal into the filter for linear echo filtering processing to obtain a residual signal.

[0105] Step S130. Based on the target model, perform interference signal cancellation and effective signal enhancement on the residual signal to obtain the target audio signal, where the interference signals include non-linear echo signals and noise signals.

[0106] In the echo cancellation method of the exemplary embodiment of the present disclosure, a pre-trained filter is used to cancel the linear echo signal in the original audio signal, and then a pre-trained target model is used to cancel the non-linear echo signal and the noise signal in the remaining residual signal, so as to obtain an enhanced target audio signal. On the one hand, the echo cancellation method in the exemplary embodiment of the present disclosure can improve the echo cancellation effect on the premise of a small system complexity by combining traditional signal processing and deep learning methods, and ensure high audio quality. On the other hand, it can not only cancel echoes, but also cancel noise interference, maximize the audio intelligibility, and thus improve the user experience. The echo cancellation method in the exemplary embodiment of the present disclosure is not only applicable to the KTV scenario, but also supports voice echo cancellation in conventional scenarios, and can achieve good sound quality.

[0107] Next, Figures 2 to 8 The above steps of the exemplary embodiment will be described in more detail.

[0108] In step S110, the original audio signal and the reference signal required for echo cancellation by the filter are obtained.

[0109] Figure 2 FIG. is a schematic diagram of an echo cancellation system according to a specific embodiment of the present disclosure. In this echo cancellation system, r(t) is the far-end reference signal at time t, s(t) is the proximal target audio signal at time t, n(t) is the noise, and h(t) is the impulse response corresponding to the echo path of the far-end reference signal.

[0110] For example, when performing real-time communication, the signal transmitted from the other end is the far-end signal and also the reference signal for echo cancellation. In a KTV system, the far-end reference signal may be the audio data transmitted from the other end during real-time communication, or the background music played by the system during KTV, or a mixture of the two signals.

[0111] The echo signal mainly includes a linear echo signal and a non-linear echo signal. Among them, the linear echo signal y(t) is the echo signal directly received by the proximal microphone, and is defined as:

[0112] y(t) = r(t) * h(t)

[0113] where * represents the convolution operation. The non-linear echo signal is the echo signal received by the proximal microphone after the far-end reference signal propagates through multiple paths, and is denoted as v(t). The original audio signal actually collected by the proximal microphone is d(t), and is defined as:

[0114] d(t) = s(t) + n(t) + r(t) * h(t) + v(t)

[0115] In order to be able to receive a pure target audio signal s(t), in the present exemplary embodiment, a traditional adaptive filtering method can be used to first eliminate the linear echo signal y(t) in d(t) to obtain a residual signal e(t), and then the neural network in the RES (Residual Echo Suppression) module is used to eliminate the remaining non-linear echo v(t) and noise n(t), so that the audio signal received by the remote end is closer to the target audio signal s(t) at the proximal end.

[0116] In step S120, the original audio signal and the reference signal are input into a filter for linear echo filtering processing to obtain a residual signal.

[0117] In the present exemplary embodiment, the filter refers to an adaptive filter. The elimination of the linear echo signal aims to remove the r(t)*h(t) part in d(t) described in the above echo cancellation system to obtain a residual signal e(t). As Figure 3 shown, inputting the original audio signal and the reference signal into a filter for linear echo filtering processing to obtain a residual signal may specifically include the following steps:

[0118] Step S310. Convolve the reference signal with the filter weights of the filter to obtain an estimated value of the linear echo signal.

[0119] For the input remote reference signal r(t), convolve it with the filter weights h(t) of the adaptive filter to obtain an estimated value of the linear echo signal r(t)*h(t).

[0120] Step S320. Obtain the remaining residual signal in the original audio signal according to the difference between the original audio signal and the estimated value of the linear echo signal.

[0121] According to the difference between the original audio signal d(t) and the estimated value of the linear echo signal r(t)*h(t), the remaining residual signal e(t) in the original audio signal can be obtained. After linear AEC, the components of the residual signal e(t) are represented by the following formula:

[0122] e(t) = s(t) + n(t) + v(t)

[0123] In step S130, based on the target model, interference signal cancellation and effective signal enhancement are performed on the residual signal to obtain a target audio signal, where the interference signals include non-linear echo signals and noise signals.

[0124] In this exemplary embodiment, the target model can be a neural network model, such as CrossNet (Cross Network), which can be used to eliminate interference signals in the residual signal, where the interference signals include non-linear echo signals and noise signals. As Figure 4 shown, based on the target model, interference signal elimination and effective signal enhancement are performed on the residual signal to obtain the target audio signal. Specifically, it can include the following steps:

[0125] Step S410. Obtain the frequency domain features corresponding to the original audio signal according to the residual signal and the reference signal.

[0126] First, extract the frequency domain features of the output signal of the linear AEC, that is, the residual signal e(t), and the far-end reference signal r(t), and splice them into X∈R corresponding to the original audio signal T×F×C , and the formula is as follows:

[0127] X = concat(FEAT[e(t)], FEAT[r(t)])

[0128] Among them, what FEAT[*] obtains is the logarithmic power spectrum corresponding to the audio signal, T is the frame index, F is the frequency index, C represents two splicing channels, and concat is the splicing function.

[0129] Step S420. Input the frequency domain features into the pre-trained target model to obtain the audio signal time-frequency mask corresponding to the target audio signal.

[0130] Input the frequency domain features extracted in the above steps into the pre-trained target model, and predict the audio signal time-frequency mask mask speech .

[0131] Step S430. Obtain the spectrum of the target audio signal according to the spectrum of the residual signal and the audio signal time-frequency mask.

[0132] According to the spectrum of the residual signal and the audio signal time-frequency mask mask speech Perform calculations to obtain the spectrum of the enhanced target audio signal. The specific formula is as follows:

[0133] Enhanced audio spectrum = mask speech * Residual signal spectrum

[0134] Step S440. Convert the spectrum of the target audio signal into a time-domain signal to obtain the enhanced target audio signal.

[0135] Finally, by using the phase of the target audio signal, the spectrum of the enhanced target audio signal is converted back to the time domain using the ISTFT (Inverse Short-Time Fourier Transform) function, and the enhanced target audio signal can be obtained.

[0136] In the exemplary embodiment, by combining traditional signal processing and deep learning methods, the echo cancellation effect can be improved on the premise of a relatively small system complexity, ensuring a high audio quality. At the same time, not only can echoes be cancelled, but also noise interference can be eliminated, maximizing the audio intelligibility and thus enhancing the user experience.

[0137] In addition, in the echo cancellation method provided in the exemplary embodiment, the training processes of the filter and the target model may also be included.

[0138] In the exemplary embodiment, as Figure 5 shown, the training method of the filter may specifically include the following steps:

[0139] Step S510. Obtain multiple groups of training data, where each group of training data includes original audio training data, reference training data corresponding to the original audio training data, and residual training data in the original audio training data.

[0140] First, multiple groups of training data required for training the adaptive filter and the neural network model need to be prepared according to requirements. Each group of training data includes original audio training data, reference training data corresponding to the original audio training data, and residual training data included in the original audio training data.

[0141] Among them, the generation method of the residual training data includes: obtaining the original audio training data collected by the user-side microphone, where the original audio training data includes original speech training data and original music training data; performing speed change processing, reverberation processing, and delay processing on the original audio training data to obtain the simulated echo data corresponding to the original audio training data; obtaining the residual training data corresponding to the original audio training data according to the difference between the original audio training data and the corresponding simulated echo data.

[0142] The far-end echo data in the training data may include background music and voice echo signals, while the original audio training data is mainly composed of voice and singing data, mainly including two parts: original speech training data and original music training data. To obtain better audio quality, data processing is an indispensable part. By processing the original audio training data through various data processing methods for data augmentation, the training data can be made closer to the actual distribution, laying a good foundation for obtaining higher sound quality.

[0143] For example, in this exemplary embodiment, the following several data augmentation methods can be mainly used:

[0144] (1) Add a large number of karaoke music to the training dataset, and keep the ratio of the original speech training data to the original music training data in the training dataset at about 1:1.

[0145] (2) Add data with variable speed and reverberation to the original speech training data in the original audio training data to simulate the corresponding echo data.

[0146] (3) Add delay to the original audio training data.

[0147] (4) Adjust the signal-to-noise ratio of the noise and echo within a certain reasonable range. The signal-to-noise ratio can generally be set from 0 dB to 20 dB, and the signal-to-echo ratio can generally be set from -20 dB to 20 dB.

[0148] Step S520. Convolve the reference training data with the initial filter weights of the filter to obtain linear echo prediction data.

[0149] In this exemplary embodiment, the NLMS-based adaptive filter algorithm can be used. First, convolve the reference training data with the initial filter weights of the adaptive filter to obtain linear echo prediction data.

[0150] Step S530. Obtain residual prediction data according to the difference between the original audio training data and the linear echo prediction data.

[0151] Subtract the above linear echo prediction data from the original audio training data to obtain the corresponding residual prediction data.

[0152] Step S540. Iterate the initial filter weights of the filter according to the difference between the residual prediction data and the residual training data to train the filter.

[0153] Apply the least mean square error principle to the difference between the residual prediction data and the residual training data. Through continuous iteration, finally make the initial filter weights of the adaptive filter converge to a path closest to the linear echo, and the training process of the adaptive filter can be completed.

[0154] In this exemplary embodiment, as Figure 6 shown, the training method of the target model can specifically include the following steps:

[0155] Step S610. Use the residual training data and the reference training data in each group of training data as the first training data, and train the initial neural network model according to the first training data to obtain the first neural network model.

[0156] Figure 7 It is a schematic diagram of a neural network model in a specific embodiment according to the present disclosure. This neural network model uses an intersection network and can be used for non-linear echo cancellation and noise suppression. The intersection network is mainly composed of multiple convolutional modules and GRU (Gated Recurrent Unit) modules. The modules are cross-connected, and the output of the previous layer is cross-spliced as the input of the next layer to simultaneously utilize the potential relationship between the two tasks. Among them, one task estimates the speech or singing signal that the proximal end wants to retain, and the output parameter is the time-frequency mask mask of the audio signal spee%h , and the other task estimates the echo and noise signals, and the output parameter is the time-frequency mask mask of the interference signal interference .

[0157] As Figure 7 shown, each branch in the intersection network respectively includes 4 convolutional modules (Conv block) and 3 GRU modules (GRU block). Each convolutional module is composed of a Conv2D (two-dimensional convolutional module), a batch normalization module BN (Batch Normalization), and an activation function (ReLU, Linear rectification function). The output of the GRU layer of the last branch is used as the input of the Dense layer, and finally, after passing through another GRU module, the time-frequency mask mask value is output

[0158] In this exemplary embodiment, the linear echo cancellation part can be combined with the training process of the neural network model, and the training process of the neural network model is divided into two stages

[0159] In the first stage, linear echo cancellation can be added to train the neural network model to make the model converge, and the first neural network model is obtained. Therefore, the training data in the first stage is the residual training data and the reference training data in each group of training data

[0160] In this exemplary embodiment, as Figure 8 shown, training the initial neural network model according to the first training data to obtain the first neural network model may specifically include the following steps

[0161] Step S810. Obtain the corresponding frequency-domain feature training data according to the first training data, and input the frequency-domain feature training data into the initial neural network model to obtain the time-frequency mask of the audio sample and the time-frequency mask of the interference sample

[0162] First, extract frequency-domain features from the residual training data and reference training data in the first training data that do not contain linear echo to obtain frequency-domain feature training data. Then, input the frequency-domain feature training data into the initial neural network model to obtain two time-frequency mask sample data, namely the audio sample time-frequency mask mask speech and the interference sample time-frequency mask mask interference .

[0163] Step S820. Obtain the sample spectrum of the enhanced target audio signal according to the sample spectrum of the residual training data and the audio sample time-frequency mask, and obtain the sample spectrum of the interference signal according to the sample spectrum of the residual training data and the interference sample time-frequency mask.

[0164] Convolve the sample spectrum of the residual training data and the audio sample time-frequency mask to obtain the sample spectrum of the enhanced target audio signal. Convolve the sample spectrum of the residual training data and the interference sample time-frequency mask to obtain the sample spectrum of the enhanced interference signal. The specific formula is as follows:

[0165] Sample spectrum of the target audio signal = mask speech * Sample spectrum of the residual training data Sample spectrum of the interference signal = mask interference * Sample spectrum of the residual training data

[0166] Step S830. Determine the spectral distance loss according to the sample spectra of the enhanced target audio signal and the interference signal, and determine the signal-to-noise ratio loss according to the signal-to-noise ratio parameters of the enhanced target audio signal and the interference signal.

[0167] In this exemplary embodiment, the audio spectral distance loss can be determined according to the sample spectra of the residual training data and the enhanced target audio signal. Then, obtain the interference training data in the training data, and determine the interference spectral distance loss according to the sample spectra of the interference training data and the interference signal. Finally, obtain the spectral distance loss according to the audio spectral distance loss and the interference spectral distance loss.

[0168] In this exemplary embodiment, two loss functions are selected, namely the OSISNR (Optimal scale-invariant signal-to-noise ratio) function and the CSD (Compressedspectrum distance) function. Among them, the OSISNR function is used to calculate the distortion degree between the original clean audio and the enhanced audio, and the signal-to-noise ratio loss can be obtained based on the OSISNR function; the CSD function is used to calculate the mismatch degree of the spectra of the original clean audio and the enhanced audio, and the spectral distance loss can be obtained based on the CSD function. Through these two loss functions, the effects of acoustic echo cancellation and noise suppression can be achieved. The specific formulas are as follows:

[0169]

[0170]

[0171]

[0172]

[0173] Among them, L OSISNR is the signal-to-noise ratio loss, L CSD is the spectral distance loss, and represent the OSISNR and spectral distance losses of the audio part, and represent the OSISNR and spectral distance losses of the interference part. S(t,f) represents the audio amplitude spectrum, t represents the time index, f represents the frequency index, and the exponential factor α = 0.3.

[0174] Step S840. Obtain the overall loss according to the spectral distance loss and the signal-to-noise ratio loss, and iterate the neural network parameters in the initial neural network model according to the overall loss to obtain the trained first neural network model.

[0175] The overall loss L = L OSISNR + γL CSD can be obtained according to the spectral distance loss and the signal-to-noise ratio loss, where the weight factor γ = 10. Iterating the neural network parameters in the initial neural network model according to the overall loss L can obtain the first neural network model.

[0176] Step S620. Use the original audio training data and the reference training data in each group of training data as the second training data, and obtain the mixed training data according to the first training data and the second training data according to the preset data ratio.

[0177] For the second stage of neural network model training, the original audio training data and reference training data in each group of training data can be used as the second training data, and then the first training data and the second training data are mixed according to a preset data ratio to obtain the mixed training data required for the second stage of training. For example, the residual training data without linear echo can be used with a probability of 10%, and the original audio training data with linear echo can be used with a probability of 90%, plus the reference training data corresponding to the above two types of data respectively, to form the mixed training data required for the second stage of training.

[0178] Step S630. Train the first neural network model according to the mixed training data to obtain the target model.

[0179] Continue to train the first neural network model obtained in the first stage based on the mixed training data until the model converges again to obtain the final target model. The specific training process is similar to that in steps S810 to S840, except that the training data used is different, which will not be elaborated here.

[0180] After the neural network model training is completed, the test data can also be obtained according to the reference training data and the original audio training data in each group of training data, and the target model can be tested according to the test data.

[0181] In the test stage of the neural network model, the original audio training data with linear echo can be used as the test data to test the neural network model, verify the effect of the neural network model in eliminating interference signals, and optimize the model.

[0182] In the embodiment of this example, in the first stage of neural network model training, based on linear AEC, the residual training data without linear echo is used, which can clarify the target eliminated by the neural network model. However, since linear AEC will damage the sound quality of music and affect the user experience, linear AEC is not used in the test stage either. However, for model matching, it is necessary to continue the second stage of training on the basis that the model has converged in the first stage, so that the model converges again to obtain the final neural network model. By using linear AEC to guide the training of the neural network model in stages, the convergence speed of the model and the model complexity can be further balanced, achieving the effect of being able to eliminate echoes to the greatest extent without damaging the sound quality of music.

[0183] It should be noted that although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0184] Furthermore, the present disclosure also provides an echo cancellation device. Refer to Figure 9 As shown, the echo cancellation device may include an audio signal acquisition module 910, a linear echo cancellation module 920, and an interference signal cancellation module 930. Among them:

[0185] The audio signal acquisition module 910 is configured to acquire the original audio signal and the reference signal required for echo cancellation by the filter;

[0186] The linear echo cancellation module 920 is configured to input the original audio signal and the reference signal into the filter for linear echo filtering processing to obtain a residual signal;

[0187] The interference signal cancellation module 930 is configured to perform interference signal cancellation and effective signal enhancement on the residual signal based on the target model to obtain the target audio signal, where the interference signals include non-linear echo signals and noise signals.

[0188] In some exemplary embodiments of the present disclosure, the linear echo cancellation module 920 may include a linear echo estimation unit and a linear echo cancellation unit. Among them:

[0189] The linear echo prediction unit is configured to convolve the reference signal with the filter weights of the filter to obtain an estimated value of the linear echo signal;

[0190] The linear echo cancellation unit is configured to obtain the remaining residual signal in the original audio signal according to the difference between the original audio signal and the estimated value of the linear echo signal.

[0191] In some exemplary embodiments of the present disclosure, the interference signal cancellation module 930 may include a frequency domain feature determination unit, a time-frequency mask determination unit, a signal spectrum determination unit, and a signal spectrum conversion unit. Among them:

[0192] The frequency domain feature determination unit is configured to obtain the frequency domain features corresponding to the original audio signal according to the residual signal and the reference signal;

[0193] The time-frequency mask determination unit is configured to input the frequency domain features into the pre-trained target model to obtain the audio signal time-frequency mask corresponding to the target audio signal;

[0194] The signal spectrum determination unit is configured to obtain the spectrum of the target audio signal according to the spectrum of the residual signal and the audio signal time-frequency mask;

[0195] The signal spectrum conversion unit is configured to convert the spectrum of the target audio signal into a time domain signal to obtain the enhanced target audio signal.

[0196] In some exemplary embodiments of the present disclosure, an echo cancellation device provided by the present disclosure may further include a filter training module, and the filter training module may include a training data acquisition unit, a linear echo prediction data determination unit, a residual prediction data determination unit, and a filter weight iteration unit. Among them:

[0197] The training data acquisition unit is configured to acquire multiple groups of training data, and each group of training data includes original audio training data, reference training data corresponding to the original audio training data, and residual training data in the original audio training data;

[0198] The linear echo prediction data determination unit is configured to convolve the reference training data with the initial filter weights of the filter to obtain linear echo prediction data;

[0199] The residual prediction data determination unit is configured to obtain residual prediction data according to the difference between the original audio training data and the linear echo prediction data;

[0200] The filter weight iteration unit is configured to iterate the initial filter weights of the filter according to the difference between the residual prediction data and the residual training data to train the filter.

[0201] In some exemplary embodiments of the present disclosure, an echo cancellation device provided by the present disclosure may further include a target model training module, and the target model training module may include an initial neural network model training unit, a mixed training data determination unit, and a first neural network model training unit. Among them:

[0202] The initial neural network model training unit is configured to use the residual training data and the reference training data in each group of training data as the first training data, and train the initial neural network model according to the first training data to obtain the first neural network model;

[0203] The mixed training data determination unit is configured to use the original audio training data and the reference training data in each group of training data as the second training data, and obtain mixed training data according to the first training data and the second training data according to a preset data ratio;

[0204] The first neural network model training unit is configured to train the first neural network model according to the mixed training data to obtain the target model.

[0205] In some exemplary embodiments of the present disclosure, the target model training module may further include a target model testing unit, and the target model testing unit is configured to obtain test data according to the reference training data and the original audio training data in each group of training data, and test the target model according to the test data.

[0206] In some exemplary embodiments of the present disclosure, the initial neural network model training unit may include a sample time-frequency mask determination unit, a sample spectrum determination unit, a network loss determination unit, and a neural network parameter iteration unit. Among them:

[0207] The sample time-frequency mask determination unit is configured to obtain corresponding frequency-domain feature training data according to the first training data, and input the frequency-domain feature training data into the initial neural network model to obtain an audio sample time-frequency mask and an interference sample time-frequency mask;

[0208] The sample spectrum determination unit is configured to obtain the sample spectrum of the enhanced target audio signal according to the sample spectrum of the residual training data and the audio sample time-frequency mask, and obtain the sample spectrum of the interference signal according to the sample spectrum of the residual training data and the interference sample time-frequency mask;

[0209] The network loss determination unit is configured to determine a spectral distance loss according to the sample spectra of the enhanced target audio signal and the interference signal, and determine a signal-to-noise ratio loss according to the signal-to-noise ratio parameter of the enhanced target audio signal and the signal-to-noise ratio parameter of the interference signal;

[0210] The neural network parameter iteration unit is configured to obtain an overall loss according to the spectral distance loss and the signal-to-noise ratio loss, and iterate the neural network parameters in the initial neural network model according to the overall loss to obtain the trained first neural network model.

[0211] In some exemplary embodiments of the present disclosure, the network loss determination unit may include an audio spectral distance loss determination unit, an interference spectral distance loss determination unit, and a spectral distance loss determination unit. Among them:

[0212] The audio spectral distance loss determination unit is configured to determine an audio spectral distance loss according to the sample spectrum of the residual training data and the sample spectrum of the enhanced target audio signal;

[0213] The interference spectral distance loss determination unit is configured to obtain interference training data in the training data, and determine an interference spectral distance loss according to the sample spectrum of the interference training data and the sample spectrum of the interference signal;

[0214] The spectral distance loss determination unit is configured to obtain a spectral distance loss according to the audio spectral distance loss and the interference spectral distance loss.

[0215] In some exemplary embodiments of the present disclosure, the training data acquisition unit may include an original audio training data acquisition unit, a simulated echo data generation unit, and a residual training data generation unit. Among them:

[0216] The original audio training data acquisition unit is configured to acquire the original audio training data collected by the client microphone, where the original audio training data includes original speech training data and original music training data;

[0217] The simulated echo data generation unit is configured to perform variable speed processing, reverberation processing, and delay processing on the original audio training data to obtain the simulated echo data corresponding to the original audio training data;

[0218] The residual training data generation unit is configured to obtain the residual training data corresponding to the original audio training data according to the difference between the original audio training data and the corresponding simulated echo data.

[0219] The specific details of each module / unit in the above echo cancellation device have been described in detail in the corresponding method embodiment section, and will not be elaborated here.

[0220] Figure 10 The structural diagram of the computer system of the electronic device suitable for implementing the embodiments of the present disclosure is shown.

[0221] It should be noted that Figure 10 The computer system 1000 of the electronic device shown is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0222] As Figure 10 shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage section 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.

[0223] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. The drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed, so that the computer program read from it can be installed into the storage section 1008 as needed.

[0224] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, various functions defined in the system of the present application are performed.

[0225] It should be noted that the computer-readable medium shown in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0226] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0227] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device is caused to implement the method as described in the above embodiments.

[0228] It should be noted that although several modules of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules described above may be embodied in one module. Conversely, the features and functions of one module described above may be further divided and embodied by multiple modules.

[0229] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure.

[0230] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An echo cancellation method, characterized in that, comprising: obtaining an original audio signal and a reference signal required for echo cancellation by a filter; inputting the original audio signal and the reference signal into the filter for linear echo filtering processing to obtain a residual signal; performing interference signal cancellation and effective signal enhancement on the residual signal based on a target model to obtain a target audio signal, wherein the interference signal includes a non-linear echo signal and a noise signal; wherein, the training method of the target model includes: obtaining multiple groups of training data, wherein each group of the training data includes original audio training data, reference training data corresponding to the original audio training data, and residual training data in the original audio training data; using the residual training data and the reference training data in each group of the training data as first training data; obtaining corresponding frequency-domain feature training data according to the first training data, and inputting the frequency-domain feature training data into an initial neural network model to obtain an audio sample time-frequency mask and an interference sample time-frequency mask; obtaining a sample spectrum of an enhanced target audio signal according to a sample spectrum of the residual training data and the audio sample time-frequency mask, and obtaining a sample spectrum of an interference signal according to the sample spectrum of the residual training data and the interference sample time-frequency mask; determining a spectrum distance loss according to the sample spectrum of the enhanced target audio signal and the sample spectrum of the interference signal, and determining a signal-to-noise ratio loss according to a signal-to-noise ratio parameter of the enhanced target audio signal and a signal-to-noise ratio parameter of the interference signal; obtaining an overall loss according to the spectrum distance loss and the signal-to-noise ratio loss, and iterating neural network parameters in the initial neural network model according to the overall loss to obtain a trained first neural network model; using the original audio training data and the reference training data in each group of the training data as second training data, and obtaining mixed training data according to the first training data and the second training data according to a preset data ratio; training the first neural network model according to the mixed training data to obtain a target model.

2. The echo cancellation method according to claim 1, characterized in that, the step of inputting the original audio signal and the reference signal into the filter for linear echo filtering processing to obtain a residual signal includes: convolving the reference signal with filter weights of the filter to obtain an estimated value of a linear echo signal; obtaining a remaining residual signal in the original audio signal according to a difference between the original audio signal and the estimated value of the linear echo signal.

3. The echo cancellation method according to claim 1, characterized in that, the step of performing interference signal cancellation and effective signal enhancement on the residual signal based on a target model to obtain a target audio signal includes: obtaining a frequency-domain feature corresponding to the original audio signal according to the residual signal and the reference signal; inputting the frequency-domain feature into a pre-trained target model to obtain an audio signal time-frequency mask corresponding to the target audio signal; Obtain the spectrum of the target audio signal based on the spectrum of the residual signal and the time-frequency mask of the audio signal; Convert the spectrum of the target audio signal into a time-domain signal to obtain the enhanced target audio signal.

4. The echo cancellation method according to claim 1, characterized in that the training method of the filter includes: Convolve the reference training data with the initial filter weights of the filter to obtain linear echo prediction data; Obtain residual prediction data according to the difference between the original audio training data and the linear echo prediction data; Iteratively update the initial filter weights of the filter according to the difference between the residual prediction data and the residual training data to train the filter.

5. The echo cancellation method according to claim 1, characterized in that the method further includes: Obtain test data according to the reference training data and the original audio training data in each group of the training data, and test the target model according to the test data.

6. The echo cancellation method according to claim 1, characterized in that the determination of the spectral distance loss according to the sample spectrum of the enhanced target audio signal and the sample spectrum of the interference signal includes: Determine the audio spectral distance loss according to the sample spectrum of the residual training data and the sample spectrum of the enhanced target audio signal; Obtain the interference training data in the training data, and determine the interference spectral distance loss according to the sample spectrum of the interference training data and the sample spectrum of the interference signal; Obtain the spectral distance loss according to the audio spectral distance loss and the interference spectral distance loss.

7. The echo cancellation method according to claim 1, characterized in that the generation method of the residual training data includes: Obtain the original audio training data collected by the user-side microphone, wherein the original audio training data includes original speech training data and original music training data; Perform variable-speed processing, reverberation processing, and delay processing on the original audio training data to obtain the simulated echo data corresponding to the original audio training data; Obtain the residual training data corresponding to the original audio training data according to the difference between the original audio training data and the corresponding simulated echo data.

8. An echo cancellation device, characterized in that it includes: An audio signal acquisition module configured to acquire the original audio signal and the reference signal required for the filter to cancel the echo; A linear echo cancellation module configured to input the original audio signal and the reference signal into the filter for linear echo filtering processing to obtain a residual signal; An interference signal cancellation module configured to perform interference signal cancellation and effective signal enhancement on the residual signal based on a target model to obtain a target audio signal, wherein the interference signal includes a non-linear echo signal and a noise signal; A target model training module, configured to execute obtaining multiple sets of training data, where each set of the training data includes original audio training data, reference training data corresponding to the original audio training data, and residual training data in the original audio training data; using the residual training data and the reference training data in each set of the training data as first training data; obtaining corresponding frequency-domain feature training data according to the first training data, and inputting the frequency-domain feature training data into an initial neural network model to obtain an audio sample time-frequency mask and an interference sample time-frequency mask; obtaining a sample spectrum of an enhanced target audio signal according to the sample spectrum of the residual training data and the audio sample time-frequency mask, and obtaining a sample spectrum of an interference signal according to the sample spectrum of the residual training data and the interference sample time-frequency mask; determining a spectrum distance loss according to the sample spectrum of the enhanced target audio signal and the sample spectrum of the interference signal, and determining a signal-to-noise ratio loss according to the signal-to-noise ratio parameter of the enhanced target audio signal and the signal-to-noise ratio parameter of the interference signal; obtaining an overall loss according to the spectrum distance loss and the signal-to-noise ratio loss, and iterating the neural network parameters in the initial neural network model according to the overall loss to obtain a trained first neural network model; using the original audio training data and the reference training data in each set of the training data as second training data, and obtaining mixed training data according to the first training data and the second training data according to a preset data ratio; training the first neural network model according to the mixed training data to obtain a target model.

9. The echo cancellation device according to claim 8, wherein, the linear echo cancellation module includes: a linear echo estimation unit, configured to execute convolving the reference signal with the filter weights of the filter to obtain an estimated value of the linear echo signal; a linear echo cancellation unit, configured to execute obtaining a residual signal remaining in the original audio signal according to the difference between the original audio signal and the estimated value of the linear echo signal.

10. The echo cancellation device according to claim 8, wherein, the interference signal cancellation module includes: a frequency-domain feature determination unit, configured to execute obtaining the frequency-domain features corresponding to the original audio signal according to the residual signal and the reference signal; a time-frequency mask determination unit, configured to execute inputting the frequency-domain features into a pre-trained target model to obtain an audio signal time-frequency mask corresponding to the target audio signal; a signal spectrum determination unit, configured to execute obtaining the spectrum of the target audio signal according to the spectrum of the residual signal and the audio signal time-frequency mask; a signal spectrum conversion unit, configured to execute converting the spectrum of the target audio signal into a time-domain signal to obtain an enhanced target audio signal.

11. The echo cancellation device according to claim 8, wherein, the echo cancellation device further includes a filter training module, and the filter training module includes: A linear echo prediction data determination unit, configured to perform convolution of the reference training data with the initial filter weights of the filter to obtain linear echo prediction data; A residual prediction data determination unit, configured to perform obtaining residual prediction data according to the difference between the original audio training data and the linear echo prediction data; A filter weight iteration unit, configured to perform iterating the initial filter weights of the filter according to the difference between the residual prediction data and the residual training data to train the filter.

12. The echo cancellation device according to claim 8, wherein, the target model training module further includes: A target model testing unit, configured to perform obtaining test data according to the reference training data and the original audio training data in each group of the training data, and testing the target model according to the test data.

13. The echo cancellation device according to claim 8, wherein, the target model training module includes: An audio spectrum distance loss determination unit, configured to perform determining an audio spectrum distance loss according to the sample spectrum of the residual training data and the sample spectrum of the enhanced target audio signal; An interference spectrum distance loss determination unit, configured to perform obtaining interference training data in the training data, and determining an interference spectrum distance loss according to the sample spectrum of the interference training data and the sample spectrum of the interference signal; A spectrum distance loss determination unit, configured to perform obtaining a spectrum distance loss according to the audio spectrum distance loss and the interference spectrum distance loss.

14. The echo cancellation device according to claim 8, wherein, the target model training module includes: An original audio training data acquisition unit, configured to perform obtaining the original audio training data collected by a user-side microphone, wherein the original audio training data includes original speech training data and original music training data; A simulated echo data generation unit, configured to perform performing speed change processing, reverberation processing, and delay processing on the original audio training data to obtain simulated echo data corresponding to the original audio training data; A residual training data generation unit, configured to perform obtaining residual training data corresponding to the original audio training data according to the difference between the original audio training data and the corresponding simulated echo data.

15. An electronic device, wherein, it includes: A processor; A memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the echo cancellation method according to any one of claims 1 to 7.

16. A computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the echo cancellation method according to any one of claims 1 to 7.

17. A computer program product, including a computer program, wherein, When the computer program is executed by a processor, it implements the method for eliminating echo as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Audio signal processing method and device thereof

    CN113362843A

  • Method and device for eliminating echo signal, computing equipment and storage medium

    CN113763977A