Training Method and Device for Echo Cancellation Model and Echo Cancellation Method and Device

By calculating the prediction loss and frequency domain mapping vectors of the echo cancellation model, and adjusting parameters to separate residual echo and speech signals, the problems of insufficient echo cancellation and leakage echo in the prior art are solved, and higher signal-to-echo ratio and echo return loss are achieved, ensuring sound quality.

CN114743558BActive Publication Date: 2025-08-05BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210473053.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-08-05
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

In the prior art, the acoustic echo cancellation system that integrates deep learning methods has limited echo cancellation and is prone to leakage of echo, especially in dual-talk scenarios.

Method used

By calculating the predicted loss of the echo cancellation model, adjusting the model parameters based on the residual echo component and speech signal component, using deep neural network for training, using frequency domain mapping vectors and signal-echo ratio to separate the speech signal and residual echo, and designing a suitable loss function to improve the echo cancellation effect.

Benefits of technology

Without shearing the voice, the echo cancellation amount is significantly increased, the signal-to-echo ratio and echo return loss are enhanced, and the problem of insufficient echo cancellation and leakage of echo in the prior art is solved, ensuring that the sound quality is not damaged.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114743558B_ABST
    Figure CN114743558B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method and device for an echo cancellation model and an echo cancellation method and device. The training method includes: obtaining a first signal based on a near-end signal; inputting the first signal into an echo cancellation model, and obtaining a second signal based on the output of the echo cancellation model, wherein the second signal is a predicted signal of the first signal after echo cancellation; determining a prediction loss of the echo cancellation model based on the first signal and the second signal, wherein the prediction loss includes a first prediction loss, which is calculated based on a residual echo component and a speech signal component separated from the second signal; and adjusting the parameters of the echo cancellation model based on the prediction loss. According to the training method and device for the echo cancellation model and the echo cancellation method and device of the present disclosure, it is possible to perform additional suppression on the residual echo and improve the amount of echo cancellation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of signal processing, and in particular to a method and device for training an echo cancellation model, and a method and device for echo cancellation. Background Art

[0002] When communicating with a remote user indoors in real time, the sound from the remote end can be played by the near-end speaker, reflected and propagated through the indoor space, and re-collected by the near-end microphone, forming an acoustic echo. The Acoustic Echo Cancellation (AEC) system can identify and suppress echo signals, preventing the remote user from hearing their own echo, greatly improving the audio experience. In related technologies, echo cancellation can be performed using AEC that incorporates deep learning methods. However, the amount of echo cancellation it can achieve is limited, and echo leakage is prone to occur. Summary of the Invention

[0003] The present disclosure provides a method and device for training an echo cancellation model, as well as an echo cancellation method and device, to at least solve the problems in the above-mentioned related technologies, or may not solve any of the above-mentioned problems. The technical solutions of the present disclosure are as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, a training method for an echo cancellation model is provided, comprising: obtaining a first signal based on a near-end signal; inputting the first signal into an echo cancellation model, and obtaining a second signal based on an output of the echo cancellation model, wherein the second signal is a predicted signal of the first signal after echo cancellation; determining a prediction loss of the echo cancellation model based on the first signal and the second signal, wherein the prediction loss includes a first prediction loss, which is calculated based on a residual echo component and a speech signal component separated from the second signal; and adjusting parameters of the echo cancellation model based on the prediction loss.

[0005] Optionally, the output of the echo cancellation model includes a first masking matrix; obtaining the second signal based on the output of the echo cancellation model includes: performing time-frequency transformation on the near-end signal to obtain a second near-end signal; and obtaining the second signal according to the product of the first masking matrix and the second near-end signal.

[0006] Optionally, obtaining the first signal based on the near-end signal includes: performing linear echo cancellation on the near-end signal to obtain the first near-end signal; performing time-frequency transformation on the first near-end signal to obtain the first signal; wherein the near-end signal includes a speech signal and an echo signal; the first near-end signal includes the speech signal and the residual echo signal; the first signal includes a frequency domain speech signal and a frequency domain residual echo signal.

[0007] Optionally, the first prediction loss is determined by the following steps: mapping the second signal in the direction of the frequency domain residual echo signal to obtain a frequency domain residual echo signal mapping vector, and using the frequency domain residual echo signal mapping vector as the residual echo component separated from the second signal; mapping the second signal in the direction of the frequency domain speech signal to obtain a frequency domain speech signal mapping vector, and using the frequency domain speech signal mapping vector as the speech signal component separated from the second signal; determining the first prediction loss based on the frequency domain residual echo signal mapping vector and the frequency domain speech signal mapping vector.

[0008] Optionally, determining the first prediction loss based on the frequency domain residual echo signal mapping vector and the frequency domain speech signal mapping vector includes: obtaining the average signal-to-back ratio of the second signal based on the frequency domain residual echo signal mapping vector and the frequency domain speech signal mapping vector; and determining the first prediction loss based on the inverse of the average signal-to-back ratio of the second signal.

[0009] Optionally, the first prediction loss is expressed as:

[0010]

[0011] Among them, loss res represents the first prediction loss, C proj (n,k) represents the frequency domain speech signal mapping vector at the time-frequency point (n,k), R proj (n,k) represents the frequency domain residual echo signal mapping vector at the time-frequency point (n,k), the time-frequency point (n,k) represents the kth frequency point of the nth frame, N represents the number of frames, K represents the number of frequency points, 1≤n≤N, 1≤k≤K, represents the average signal-to-back ratio of the second signal.

[0012] Optionally, the predicted loss further includes a second predicted loss, where the second predicted loss includes at least one loss value related to echo cancellation, wherein the predicted loss is obtained according to the first predicted loss and the second predicted loss.

[0013] According to a second aspect of an embodiment of the present disclosure, an echo cancellation method is provided, comprising: obtaining a third signal based on an acquired near-end acquisition signal; inputting the third signal into an echo cancellation model trained by the echo cancellation model training method of the present disclosure, and obtaining a fourth signal based on the output of the trained echo cancellation model, wherein the fourth signal is a predicted signal of the third signal after echo cancellation; and obtaining a near-end acquisition signal after echo cancellation based on the fourth signal.

[0014] Optionally, obtaining the third signal based on the acquired near-end acquisition signal includes: performing linear echo cancellation on the near-end acquisition signal to obtain the first near-end acquisition signal; and performing time-frequency transformation on the first near-end acquisition signal to obtain the third signal.

[0015] Optionally, the output of the trained echo cancellation model includes a second masking matrix; the fourth signal is obtained based on the output of the trained echo cancellation model, including: performing time-frequency transformation on the near-end acquisition signal to obtain a second near-end acquisition signal; and obtaining the fourth signal according to the product of the second masking matrix and the second near-end acquisition signal.

[0016] Optionally, obtaining the near-end acquisition signal after echo cancellation according to the fourth signal includes: performing an inverse time-frequency transform on the fourth signal to obtain the near-end acquisition signal after echo cancellation.

[0017] According to a third aspect of an embodiment of the present disclosure, a training device for an echo cancellation model is provided, comprising: a first signal determination unit, configured to obtain a first signal based on a near-end signal; a first model prediction unit, configured to input the first signal into the echo cancellation model, and obtain a second signal based on the output of the echo cancellation model, wherein the second signal is a predicted signal of the first signal after echo cancellation; a loss determination unit, configured to determine a prediction loss of the echo cancellation model based on the first signal and the second signal, wherein the prediction loss includes a first prediction loss, which is calculated based on a residual echo component and a speech signal component separated from the second signal; and a parameter adjustment unit, configured to adjust parameters of the echo cancellation model based on the prediction loss.

[0018] Optionally, the output of the echo cancellation model includes a first masking matrix; the first model prediction unit is configured to: perform time-frequency transformation on the near-end signal to obtain a second near-end signal; and obtain the second signal based on the product of the first masking matrix and the second near-end signal.

[0019] Optionally, the first signal determination unit is configured to: perform linear echo cancellation on the near-end signal to obtain a first near-end signal; perform time-frequency transformation on the first near-end signal to obtain a first signal; wherein the near-end signal includes a speech signal and an echo signal; the first near-end signal includes the speech signal and a residual echo signal; the first signal includes a frequency domain speech signal and a frequency domain residual echo signal.

[0020] Optionally, the loss determination unit is configured to: map the second signal in the direction of the frequency domain residual echo signal to obtain a frequency domain residual echo signal mapping vector, and use the frequency domain residual echo signal mapping vector as the residual echo component separated from the second signal; map the second signal in the direction of the frequency domain speech signal to obtain a frequency domain speech signal mapping vector, and use the frequency domain speech signal mapping vector as the speech signal component separated from the second signal; determine the first prediction loss based on the frequency domain residual echo signal mapping vector and the frequency domain speech signal mapping vector.

[0021] Optionally, the loss determination unit is configured to: obtain the average signal-to-back ratio of the second signal based on the frequency domain residual echo signal mapping vector and the frequency domain speech signal mapping vector; and determine the first prediction loss based on the inverse of the average signal-to-back ratio of the second signal.

[0022] Optionally, the first prediction loss is expressed as:

[0023]

[0024] Among them, loss res represents the first prediction loss, C proj (n,k) represents the frequency domain speech signal mapping vector at the time-frequency point (n,k), R proj (n,k) represents the frequency domain residual echo signal mapping vector at the time-frequency point (n,k), the time-frequency point (n,k) represents the kth frequency point of the nth frame, N represents the number of frames, K represents the number of frequency points, 1≤n≤N, 1≤k≤K, represents the average signal-to-back ratio of the second signal.

[0025] Optionally, the predicted loss further includes a second predicted loss, where the second predicted loss includes at least one loss value related to echo cancellation, wherein the predicted loss is obtained according to the first predicted loss and the second predicted loss.

[0026] According to a fourth aspect of an embodiment of the present disclosure, an echo cancellation device is provided, including: a third signal determination unit, configured to obtain a third signal based on an acquired near-end acquisition signal; a second model prediction unit, configured to input the third signal into an echo cancellation model trained by the training method of the echo cancellation model of the present disclosure, and obtain a fourth signal based on the output of the trained echo cancellation model, wherein the fourth signal is a predicted signal of the third signal after echo cancellation; and an echo cancellation signal unit, configured to obtain the near-end acquisition signal after echo cancellation based on the fourth signal.

[0027] Optionally, the third signal determination unit is configured to: perform linear echo cancellation on the near-end collected signal to obtain a first near-end collected signal; and perform time-frequency transformation on the first near-end collected signal to obtain a third signal.

[0028] Optionally, the output of the trained echo cancellation model includes a second masking matrix; the second model prediction unit is configured to: perform time-frequency transformation on the near-end acquisition signal to obtain a second near-end acquisition signal; and obtain the fourth signal based on the product of the second masking matrix and the second near-end acquisition signal.

[0029] Optionally, the echo cancellation signal unit is configured to: perform time-frequency inverse transform on the fourth signal to obtain a near-end acquisition signal after echo cancellation.

[0030] According to a fifth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer-executable instructions, wherein the computer-executable instructions, when executed by the at least one processor, prompt the at least one processor to execute the training method of the echo cancellation model or the echo cancellation method according to the present disclosure.

[0031] According to a sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium storing instructions is provided, which, when executed by at least one processor, prompts the at least one processor to execute the training method of the echo cancellation model or the echo cancellation method according to the present disclosure.

[0032] According to a seventh aspect of an embodiment of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by at least one processor, implement the echo cancellation model training method or echo cancellation method according to the present disclosure.

[0033] According to an eighth aspect of an embodiment of the present disclosure, a smart speaker is provided, comprising the echo cancellation device of the present disclosure.

[0034] According to a ninth aspect of an embodiment of the present disclosure, a smart speaker is provided, comprising: at least one audio signal processor; at least one memory for storing instructions; at least one audio signal collector; and at least one audio signal outputter, wherein the instructions, when executed by the at least one audio signal processor, prompt the at least one audio signal processor to execute the echo cancellation method according to the present disclosure.

[0035] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:

[0036] According to the training method and device of the echo cancellation model and the echo cancellation method and device disclosed in the present invention, the parameters of the echo cancellation model can be adjusted based on the first prediction loss calculated from the residual echo component and the speech signal component separated from the second signal predicted by the echo cancellation model, so as to train the echo cancellation model. This can help guide the training process of the echo cancellation model, and perform additional suppression on the residual echo without introducing speech clipping and ensuring the speech quality, thereby improving the echo cancellation amount, solving the problems of insufficient echo cancellation and double-speaking echo leakage in existing solutions, and achieving a higher signal-to-return ratio and echo return loss enhancement (ERLE).

[0037] In addition, according to the echo cancellation model training method and device and the echo cancellation method and device disclosed in the present invention, the mapping principle can be used to separate the correctly predicted speech signal component and the residual echo component in the second signal, and the signal-to-echo ratio can be calculated to provide more penalty for the residual echo.

[0038] Furthermore, according to the disclosed echo cancellation model training method and apparatus, as well as the echo cancellation method and apparatus, the first prediction loss is scale-invariant, meaning the magnitude of the prediction loss is unaffected by signal amplitude. This forces the model to optimize the prediction loss from the perspective of suppressing residual echo rather than overall signal amplitude. Even for small-amplitude speech audio signals, to ensure a high signal-to-response ratio, speech clipping is avoided, thereby preserving sound quality. Furthermore, the echo leakage that occurs in dual-talk scenarios in existing solutions is eliminated, preventing excessive suppression of the speech in the dual-talk portion.

[0039] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0041] Figure 1 The figure is a schematic structural diagram of an echo cancellation model according to an exemplary embodiment.

[0042] Figure 2 The figure is a flowchart of a method for training an echo cancellation model according to an exemplary embodiment.

[0043] Figure 3 The figure is a flow chart showing an echo cancellation method according to an exemplary embodiment.

[0044] Figure 4 The figure is a block diagram of a training device for an echo cancellation model according to an exemplary embodiment.

[0045] Figure 5 The figure is a block diagram of an echo cancellation device according to an exemplary embodiment.

[0046] Figure 6 is a block diagram of an electronic device 600 according to an exemplary embodiment.

[0047] Figure 7 is a block diagram of a smart speaker 700 according to an exemplary embodiment of the present disclosure.

[0048] Figure 8 is a block diagram of a smart speaker 800 according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0049] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0050] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation methods described in the following examples do not represent all implementation methods consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0051] It should be noted that the phrase "at least one of the items" in this disclosure includes three types of parallel situations: "any one of the items", "a combination of any multiple items of the items", and "all of the items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. For another example, "performing at least one of step 1 and step 2" includes the following three parallel situations: (1) performing step 1; (2) performing step 2; and (3) performing steps 1 and 2.

[0052] When communicating with a remote user indoors in real time, the sound from the remote end can be played back by the local speaker. After being reflected and propagated through the room, it is re-collected by the local microphone, forming an acoustic echo. The Acoustic Echo Cancellation (AEC) system can identify and suppress the echo signal, preventing the remote user from hearing their own echo, greatly improving the audio experience.

[0053] In recent years, deep learning methods have begun to be applied to the field of signal processing, such as AEC. Compared with traditional AEC based on signal correlation, Deep AEC that integrates deep learning methods has shown better performance in speech preservation and sound quality improvement. In the process of using Deep AEC for echo cancellation, the selection of loss function is crucial. It provides guidance for model optimization and iteration and greatly affects the performance of the model. In related technologies, loss functions derived from fields such as sound source separation and noise suppression are often used to train Deep AEC. For example, loss functions in related technologies can be divided into two types. One is a loss function based on the signal-to-noise ratio, such as the scale-invariant signal-to-noise ratio loss function SI-SNRloss (scale-invariant source-to-noise ratio); the other is a loss function based on the mean squared error (MSE), such as spectral MSE loss. However, the above loss function does not distinguish between echo and noise, but instead adopts the same optimization strategy for all interference signals. Therefore, the above loss function does not impose additional penalties on echoes. This causes the trained model to pay insufficient attention to echoes. As a result, in practical applications, especially in double-talk scenarios, the amount of echo cancellation is limited, and echo leakage is prone to occur.

[0054] In order to solve the problems existing in the above-mentioned related technologies, the present disclosure proposes a training method and device for an echo cancellation model, as well as an echo cancellation method and device. The parameters of the echo cancellation model can be adjusted based on the first prediction loss calculated from the residual echo component and the speech signal component separated from the second signal predicted by the echo cancellation model, thereby training the echo cancellation model. This can help guide the training process of the echo cancellation model, and perform additional suppression on the residual echo without introducing speech clipping and ensuring speech quality, thereby improving the echo cancellation amount, solving the problems of insufficient echo cancellation and double-speaking echo leakage in existing solutions, and achieving a higher signal-to-echo ratio and ERLE.

[0055] Below, we will refer to Figures 1 to 8 The following describes in detail the echo cancellation model training method and device and the echo cancellation method and device according to the present disclosure.

[0056] Figure 1 FIG. 1 is a schematic diagram showing the structure of an echo cancellation model according to an exemplary embodiment. Figure 1 The echo cancellation model can be a Deep AEC network structure, which can be constructed using a deep neural network (Deep Neural Networks, DNN).

[0057] First, a near-end microphone signal d(t) and a far-end reference signal f(t) may be obtained. The near-end microphone signal may be a collected near-end microphone signal or a synthesized near-end microphone signal. It should be noted that the signal in the exemplary embodiment of the present disclosure may be an audio signal.

[0058] Here, the collected near-end microphone signal may include, but is not limited to, a user's voice signal collected by the near-end microphone and an audio signal collected by the near-end microphone and transmitted from the far-end via a network link and played by the near-end speaker. The far-end reference audio signal may be an audio signal transmitted from the far-end that has not been played by the near-end speaker.

[0059] Then, linear echo cancellation may be performed on the near-end microphone signal d(t) based on the far-end reference signal f(t) to obtain a first near-end microphone signal l(t) after linear echo cancellation.

[0060] Next, the first near-end microphone signal l(t) and the far-end reference signal f(t) can be transformed in time-frequency to extract complex spectral features, and the two signals after time-frequency transformation can be input into the echo cancellation model to further suppress residual echo (linear and nonlinear) and noise.

[0061] Finally, a spectrum-based masking matrix mask can be output speech , and apply it to the near-end microphone signal D(n,k) after time-frequency transformation to obtain the near-end microphone signal S(n,k) after echo cancellation.

[0062] Next, combine Figure 1 The model structure shown will describe the training method of the echo cancellation model and the echo cancellation method of the present disclosure from the training stage and the application reasoning stage of the echo cancellation model respectively.

[0063] Figure 2 FIG. 1 is a flow chart of a method for training an echo cancellation model according to an exemplary embodiment. Figure 2 In step 201, a first signal may be obtained based on a near-end signal.

[0064] According to an exemplary embodiment of the present disclosure, the near-end signal may be a near-end microphone signal. The exemplary embodiment of the present disclosure may perform echo cancellation on the near-end signal to obtain a first signal. For example, the first signal here may be obtained by the following two steps: first, linear echo cancellation may be performed on the near-end signal to obtain a first near-end signal. Then, time-frequency transformation may be performed on the first near-end signal to obtain the first signal. It should be noted that here only the setting of first performing linear echo cancellation and then performing time-frequency transformation is taken as an example, and the present disclosure also protects the scheme of first performing time-frequency transformation and then performing linear echo cancellation. It should also be noted that performing linear echo cancellation on the near-end signal may be performing linear echo cancellation on the near-end signal based on the first far-end reference signal. The time-frequency transformation in the exemplary embodiment of the present disclosure may be a short-time Fourier transform (STFT).

[0065] The near-end signal in the exemplary embodiment of the present disclosure may use a synthesized near-end microphone signal. Based on the characteristics of the near-end microphone signal, the near-end signal may be configured to include a voice signal and an echo signal. Thus, the first near-end signal after linear echo cancellation may include the voice signal and a residual echo signal. The residual echo signal may be the residual echo signal after linear echo cancellation of the near-end signal. The residual echo signal in the exemplary embodiment of the present disclosure may be obtained by calculation, and since linear echo cancellation has little effect on signals such as voice, the exemplary embodiment of the present disclosure assumes that the voice signal does not change after linear echo cancellation, that is, the residual echo signal may be obtained by the difference between the first near-end signal and the voice signal. Next, the first near-end signal may be converted into a first signal through time-frequency transformation. The first signal may include a frequency-domain voice signal and a frequency-domain residual echo signal.

[0066] Furthermore, the near-end signal in the exemplary embodiment of the present disclosure may also include a noise signal, that is, the near-end signal includes a speech signal, an echo signal, and a noise signal. Then, the near-end signal d(t) can be expressed as the following formula (1):

[0067] d(t)=c(t)+n(t)+e(t) (1)

[0068] Where d(t) represents the near-end signal, c(t) represents the speech signal, n(t) represents the noise signal, e(t) represents the echo signal, and t represents time.

[0069] The first near-end signal after linear echo cancellation may include the above-mentioned speech signal, the above-mentioned noise signal, and a residual echo signal. The residual echo signal may be the residual echo signal after linear echo cancellation is performed on the above-mentioned near-end signal. The residual echo signal in the exemplary embodiment of the present disclosure can be obtained by calculation. Since linear echo cancellation has little effect on signals such as speech and noise, the exemplary embodiment of the present disclosure assumes that the speech signal and noise signal after linear echo cancellation are unchanged, that is, the residual echo signal r(t) can be expressed as the following formula (2):

[0070] r(t)=l(t)-c(t)-n(t) (2)

[0071] Wherein, r(t) represents the residual echo signal, l(t) represents the first near-end signal, c(t) represents the speech signal, n(t) represents the noise signal, and t represents time.

[0072] The first near-end signal may be converted into a first signal through time-frequency transformation. The first signal may include a frequency-domain speech signal C(n,k), a frequency-domain noise signal, and a frequency-domain residual echo signal R(n,k).

[0073] After obtaining the first signal, exemplary embodiments of the present disclosure may input the first signal into an echo cancellation model for subsequent model training steps. Specifically, in step 202, the first signal may be input into the echo cancellation model, and a second signal may be obtained based on the output of the echo cancellation model. The second signal is a predicted signal of the first signal after echo cancellation.

[0074] Here, in addition to inputting the first signal into the echo cancellation model, the first far-end reference signal after time-frequency transformation is also input into the echo cancellation model. That is, the input of the echo cancellation model includes the first signal and the first far-end reference signal after time-frequency transformation.

[0075] Based on the above Figure 1 As can be seen from the description of , the output of the echo cancellation model may include a first masking matrix. Then, the second signal can be obtained by applying the first masking matrix to the near-end signal after time-frequency transformation. For example, the second signal can be obtained by the following steps: First, the near-end signal can be time-frequency transformed to obtain a second near-end signal. Then, the second signal can be obtained by multiplying the first masking matrix and the second near-end signal. For example, the second signal can be expressed as the following formula (3):

[0076] S(n,k)=D(n,k)·mask speech (3)

[0077] Among them, S(n,k) represents the second signal at the time-frequency point (n,k), D(n,k) represents the second near-end signal, and mask speechRepresents the first masking matrix, time-frequency point (n, k) represents the kth frequency point of the nth frame, N represents the number of frames, K represents the number of frequency points, 1≤n≤N, 1≤k≤K.

[0078] Next, the prediction loss of the echo cancellation model in this training can be determined. In order to address the deficiencies of the loss function in the prior art, an exemplary embodiment of the present disclosure can select a more appropriate target design loss function, increase the suppression of the echo through a first prediction loss, and penalize the residual echo signal in the above-mentioned second signal. Specifically, it can be: in step 203, the prediction loss of the echo cancellation model can be determined based on the first signal and the second signal, wherein the prediction loss includes a first prediction loss, and the first prediction loss is calculated based on the residual echo component and the speech signal component separated from the second signal.

[0079] In an exemplary embodiment of the present disclosure, the frequency-domain speech signal C(n, k) and the frequency-domain residual echo signal R(n, k) are selected as targets, and the two are compared with the second signal to obtain a residual echo component and a speech signal component separated from the second signal. A first prediction loss is then obtained based on these two components. Specifically, the first prediction loss can be determined by the following steps:

[0080] First, the second signal may be mapped toward the frequency domain residual echo signal to obtain a frequency domain residual echo signal mapping vector, and the frequency domain residual echo signal mapping vector is used as the residual echo component separated from the second signal.

[0081] For example, the frequency domain residual echo signal mapping vector can be expressed as the following equation (4):

[0082]

[0083] Among them, R proj (n,k) represents the frequency domain residual echo signal mapping vector at the time-frequency point (n,k), the time-frequency point (n,k) represents the kth frequency point of the nth frame, N represents the number of frames, K represents the number of frequency points, 1≤n≤N, 1≤k≤K, S(n,k) represents the second signal at the time-frequency point (n,k), R(n,k) represents the frequency domain residual echo signal at the time-frequency point (n,k), <S(n,k), R(n,k)> represents the dot product of S(n,k) and R(n,k), Represents the square of the L2 norm of R(n,k).

[0084] Then, the second signal may be mapped in the direction of the frequency domain speech signal to obtain a frequency domain speech signal mapping vector, and the frequency domain speech signal mapping vector is used as the speech signal component separated from the second signal.

[0085] For example, the frequency domain speech signal mapping vector can be expressed as the following formula (5):

[0086]

[0087] Among them, C proj (n,k) represents the frequency domain speech signal mapping vector at the time-frequency point (n,k), the time-frequency point (n,k) represents the kth frequency point of the nth frame, N represents the number of frames, K represents the number of frequency points, 1≤n≤N, 1≤k≤K, S(n,k) represents the second signal at the time-frequency point (n,k), C(n,k) represents the frequency domain speech signal at the time-frequency point (n,k), <S(n,k), C(n,k)> represents the dot product of S(n,k) and C(n,k), Represents the square of the L2 norm of C(n,k).

[0088] Finally, a first prediction loss may be determined according to the frequency-domain residual echo signal mapping vector and the frequency-domain speech signal mapping vector.

[0089] Here, first, an average signal-to-echo ratio of the second signal can be obtained based on the frequency-domain residual echo signal mapping vector and the frequency-domain speech signal mapping vector. For example, the signal-to-echo ratio of each frame of the second signal can be obtained based on the frequency-domain residual echo signal mapping vector and the frequency-domain speech signal mapping vector, and the signal-to-echo ratios of each frame can be averaged to obtain the average signal-to-echo ratio of the second signal.

[0090] The first prediction loss can then be determined based on the average signal-to-echo ratio of the second signal. For example, the first prediction loss can be determined based on the inverse of the average signal-to-echo ratio of the second signal. It should be noted that exemplary embodiments of the present disclosure require the second signal to have the largest average signal-to-echo ratio, thereby imposing a greater penalty on residual echo. Therefore, the first prediction loss is set to be the inverse of the average signal-to-echo ratio of the second signal.

[0091] For example, the first prediction loss can be expressed as the following equation (6):

[0092]

[0093] Among them, loss res represents the first prediction loss, C proj (n,k) represents the frequency domain speech signal mapping vector at the time-frequency point (n,k), R proj (n,k) represents the frequency domain residual echo signal mapping vector at the time-frequency point (n,k), the time-frequency point (n,k) represents the kth frequency point of the nth frame, N represents the number of frames, K represents the number of frequency points, 1≤n≤N, 1≤k≤K, Represents the average signal-to-back ratio of the second signal.

[0094] It should be noted that before the calculations of (4) to (6), R(n,k), C(n,k), and S(n,k) can all be normalized to their mean values. The first prediction loss represented by (6) above is scale invariant.

[0095] According to an exemplary embodiment of the present disclosure, the predicted loss may further include a second predicted loss, which includes at least one loss value related to echo cancellation. Here, the loss value related to echo cancellation may include, but is not limited to, SI-SNR loss and spectral MSE loss. Based on this, the predicted loss may be obtained based on the first predicted loss and the second predicted loss. For example, the predicted loss may be obtained based on the sum of the first predicted loss and the second predicted loss.

[0096] For example, the prediction loss can be expressed as the following equation (7):

[0097] loss=loss SI_SNR +α·loss res (7)

[0098] Among them, loss represents the prediction loss, loss SI_SNR represents the second prediction loss (SI-SNR loss), α represents the parameter used to adjust the ratio of the first prediction loss and the second prediction loss, loss res The first prediction loss and the second prediction loss can be adjusted to the same order of magnitude.

[0099] Back to Figure 2 In step 204, the parameters of the echo cancellation model may be adjusted according to the predicted loss.

[0100] above Figure 2 This describes the process of training the echo cancellation model. Figure 3 , and expand the application reasoning process of the echo cancellation model. Figure 3 FIG. 1 is a flow chart showing a method for echo cancellation according to an exemplary embodiment. Figure 3 In step 301, a third signal may be obtained based on the acquired near-end acquisition signal.

[0101] According to an exemplary embodiment of the present disclosure, the near-end acquisition signal may be a collected near-end microphone signal. The exemplary embodiment of the present disclosure may perform echo cancellation on the near-end acquisition signal to obtain a third signal. For example, the third signal here may be obtained by the following two steps: first, linear echo cancellation may be performed on the near-end acquisition signal to obtain a first near-end acquisition signal. Then, time-frequency transformation may be performed on the first near-end acquisition signal to obtain a third signal. It should be noted that here only the setting of first performing linear echo cancellation and then performing time-frequency transformation is taken as an example. The present disclosure also protects the scheme of first performing time-frequency transformation and then performing linear echo cancellation. It should also be noted that performing linear echo cancellation on the near-end acquisition signal may be performing linear echo cancellation on the near-end acquisition signal based on the collected second far-end reference signal. The time-frequency transformation in the exemplary embodiment of the present disclosure may be a short-time Fourier transform (STFT).

[0102] In step 302, the third signal may be input into a trained echo cancellation model, and a fourth signal may be obtained based on an output of the trained echo cancellation model, wherein the fourth signal is a predicted signal of the third signal after echo cancellation.

[0103] Here, in addition to inputting the third signal into the trained echo cancellation model, the second far-end reference signal after time-frequency transformation is also input into the trained echo cancellation model. That is, the input of the trained echo cancellation model includes the third signal and the second far-end reference signal after time-frequency transformation.

[0104] Based on the above Figure 1 As can be seen from the description, the output of the trained echo cancellation model may include a second masking matrix. Therefore, the fourth signal can be obtained by applying the second masking matrix to the near-end acquired signal after time-frequency transformation. For example, the fourth signal can be obtained by the following steps: first, time-frequency transformation can be performed on the near-end acquired signal to obtain a second near-end acquired signal. Then, the fourth signal can be obtained by multiplying the second masking matrix and the second near-end acquired signal.

[0105] In step 303, a near-end collected signal after echo cancellation may be obtained according to the fourth signal.

[0106] According to an exemplary embodiment of the present disclosure, a fourth signal in the time domain can be obtained as the near-end acquisition signal after echo cancellation. For example, an inverse time-frequency transform can be performed on the fourth signal to obtain the near-end acquisition signal after echo cancellation. The inverse time-frequency transform in the exemplary embodiment of the present disclosure can be a short-time inverse Fourier transform (ISTFT).

[0107] Figure 4 FIG1 is a block diagram of a training device for an echo cancellation model according to an exemplary embodiment. Figure 4The training device 400 includes a first signal determination unit 401, a first model prediction unit 402, a loss determination unit 403 and a parameter adjustment unit 404.

[0108] The first signal determining unit 401 may obtain a first signal according to the near-end signal.

[0109] According to an exemplary embodiment of the present disclosure, the near-end signal may be a near-end microphone signal. The first signal determination unit 401 may perform echo cancellation on the near-end signal to obtain a first signal. For example, the first signal here may be obtained by configuring the first signal determination unit 401 as follows: first, the first signal determination unit 401 may perform linear echo cancellation on the near-end signal to obtain a first near-end signal. Then, the first signal determination unit 401 may perform time-frequency transformation on the first near-end signal to obtain the first signal. It should be noted that here, only the example of the first signal determination unit 401 first performing linear echo cancellation and then performing time-frequency transformation is taken as an example. The present disclosure also protects the solution in which the first signal determination unit 401 first performs time-frequency transformation and then performs linear echo cancellation. It should be noted that performing linear echo cancellation on the near-end signal may be performing linear echo cancellation on the near-end signal based on the first far-end reference signal. The time-frequency transformation in the exemplary embodiment of the present disclosure may be a short-time Fourier transform (STFT).

[0110] The near-end signal in the exemplary embodiment of the present disclosure may use a synthesized near-end microphone signal. Based on the characteristics of the near-end microphone signal, the near-end signal may be configured to include a voice signal and an echo signal. Thus, the first near-end signal after linear echo cancellation may include the voice signal and a residual echo signal. The residual echo signal may be the residual echo signal after linear echo cancellation of the near-end signal. The residual echo signal in the exemplary embodiment of the present disclosure may be obtained by calculation, and since linear echo cancellation has little effect on signals such as voice, the exemplary embodiment of the present disclosure assumes that the voice signal does not change after linear echo cancellation, that is, the residual echo signal may be obtained by the difference between the first near-end signal and the voice signal. Next, the first near-end signal may be converted into a first signal through time-frequency transformation. The first signal may include a frequency-domain voice signal and a frequency-domain residual echo signal.

[0111] Furthermore, the near-end signal in the exemplary embodiment of the present disclosure may also include a noise signal, that is, the near-end signal includes a speech signal, an echo signal, and a noise signal. Then, the near-end signal d(t) may be expressed as the above formula (1).

[0112] The first near-end signal after linear echo cancellation may include the above-mentioned voice signal, the above-mentioned noise signal and the residual echo signal. The residual echo signal may be the residual echo signal after linear echo cancellation is performed on the above-mentioned near-end signal. The residual echo signal in the exemplary embodiment of the present disclosure can be obtained by calculation, and since linear echo cancellation has little effect on signals such as voice and noise, the exemplary embodiment of the present disclosure assumes that the voice signal and noise signal do not change after linear echo cancellation, that is, the residual echo signal r(t) can be expressed as the above formula (2).

[0113] The first near-end signal may be converted into a first signal through time-frequency transformation. The first signal may include a frequency-domain speech signal C(n,k), a frequency-domain noise signal, and a frequency-domain residual echo signal R(n,k).

[0114] The first model prediction unit 402 may input the first signal into the echo cancellation model and obtain a second signal based on the output of the echo cancellation model, wherein the second signal is a predicted signal of the first signal after echo cancellation.

[0115] Here, in addition to inputting the first signal into the echo cancellation model, the first far-end reference signal after time-frequency transformation is also input into the echo cancellation model. That is, the input of the echo cancellation model includes the first signal and the first far-end reference signal after time-frequency transformation.

[0116] Based on the above Figure 1 As can be seen from the description, the output of the echo cancellation model may include a first masking matrix, so the second signal can be obtained by the first model prediction unit 402 applying the first masking matrix to the near-end signal after time-frequency transformation. For example, the first model prediction unit 402 can first perform time-frequency transformation on the near-end signal to obtain the second near-end signal; then the second signal can be obtained based on the product of the first masking matrix and the second near-end signal.

[0117] For example, the second signal can be expressed as the above equation (3).

[0118] The loss determination unit 403 may determine a prediction loss of the echo cancellation model according to the first signal and the second signal, wherein the prediction loss includes a first prediction loss calculated based on the residual echo component and the speech signal component separated from the second signal.

[0119] In an exemplary embodiment of the present disclosure, the frequency-domain speech signal C(n, k) and the frequency-domain residual echo signal R(n, k) may be selected as targets, and the two may be compared with the second signal to obtain a residual echo component and a speech signal component separated from the second signal. The first prediction loss may then be obtained based on these two components. Specifically, the loss determination unit 403 may determine the first prediction loss using the following configuration:

[0120] First, the loss determination unit 403 may map the second signal toward the frequency domain residual echo signal to obtain a frequency domain residual echo signal mapping vector, and use the frequency domain residual echo signal mapping vector as the residual echo component separated from the second signal.

[0121] For example, the frequency-domain residual echo signal mapping vector can be expressed as the above equation (4).

[0122] Then, the loss determination unit 403 may map the second signal toward the frequency domain speech signal to obtain a frequency domain speech signal mapping vector, and use the frequency domain speech signal mapping vector as the speech signal component separated from the second signal.

[0123] For example, the frequency domain speech signal mapping vector can be expressed as the above formula (5).

[0124] Finally, the loss determining unit 403 may determine a first prediction loss according to the frequency-domain residual echo signal mapping vector and the frequency-domain speech signal mapping vector.

[0125] Here, first, the loss determination unit 403 may obtain an average signal-to-echo ratio (SER) of the second signal based on the frequency-domain residual echo signal mapping vector and the frequency-domain speech signal mapping vector. For example, the loss determination unit 403 may obtain the SER of each frame of the second signal based on the frequency-domain residual echo signal mapping vector and the frequency-domain speech signal mapping vector, and average the SERs of each frame to obtain the average SER of the second signal.

[0126] The loss determination unit 403 may then determine a first predicted loss based on the average signal-to-echo ratio of the second signal. For example, the loss determination unit 403 may determine the first predicted loss based on the inverse of the average signal-to-echo ratio of the second signal. It should be noted that exemplary embodiments of the present disclosure require the second signal to have the highest average signal-to-echo ratio, thereby imposing a greater penalty on residual echo. Therefore, the first predicted loss is set to be the inverse of the average signal-to-echo ratio of the second signal.

[0127] For example, the first prediction loss can be expressed as the above formula (6).

[0128] It should be noted that before the calculations of (4) to (6), R(n,k), C(n,k), and S(n,k) can all be normalized to their mean values. The first prediction loss represented by (6) above is scale invariant.

[0129] According to an exemplary embodiment of the present disclosure, the predicted loss may further include a second predicted loss, which includes at least one loss value related to echo cancellation. Here, the loss value related to echo cancellation may include, but is not limited to, SI-SNR loss and spectral MSE loss. Based on this, the predicted loss may be obtained based on the first predicted loss and the second predicted loss. For example, the loss determination unit 403 may obtain the predicted loss based on the sum of the first predicted loss and the second predicted loss.

[0130] For example, the prediction loss can be expressed as the above equation (7).

[0131] The parameter adjustment unit 404 may adjust the parameters of the echo cancellation model according to the prediction loss.

[0132] Figure 5 FIG1 is a block diagram of an echo cancellation device according to an exemplary embodiment. Figure 5 The echo cancellation device 500 includes a third signal determination unit 501 , a second model prediction unit 502 and an echo cancellation signal unit 503 .

[0133] The third signal determining unit 501 may obtain a third signal according to the acquired near-end acquisition signal.

[0134] According to an exemplary embodiment of the present disclosure, the near-end acquisition signal may be a collected near-end microphone signal. The third signal determination unit 501 may perform echo cancellation on the near-end acquisition signal to obtain a third signal. For example, the third signal determination unit 501 may first perform linear echo cancellation on the near-end acquisition signal to obtain a first near-end acquisition signal. Then the third signal determination unit 501 may perform time-frequency transformation on the first near-end acquisition signal to obtain a third signal. It should be noted that here only the example of the third signal determination unit 501 first performing linear echo cancellation and then performing time-frequency transformation is taken. The present disclosure also protects the solution in which the third signal determination unit 501 first performs time-frequency transformation and then performs linear echo cancellation. It should also be noted that performing linear echo cancellation on the near-end acquisition signal may be performing linear echo cancellation on the near-end acquisition signal based on the collected second far-end reference signal. The time-frequency transformation in the exemplary embodiment of the present disclosure may be a short-time Fourier transform (STFT).

[0135] The second model prediction unit 502 may input the third signal into the trained echo cancellation model, and obtain a fourth signal based on the output of the trained echo cancellation model, wherein the fourth signal is a predicted signal of the third signal after echo cancellation.

[0136] Here, in addition to inputting the third signal into the trained echo cancellation model, the second far-end reference signal after time-frequency transformation is also input into the trained echo cancellation model. That is, the input of the trained echo cancellation model includes the third signal and the second far-end reference signal after time-frequency transformation.

[0137] Based on the above Figure 1 As can be seen from the description, the output of the trained echo cancellation model may include a second masking matrix. Therefore, the fourth signal can be obtained by the second model prediction unit 502 applying the second masking matrix to the near-end collected signal after time-frequency transformation. For example, the second model prediction unit 502 may first perform time-frequency transformation on the near-end collected signal to obtain the second near-end collected signal. The second model prediction unit 502 may then obtain the fourth signal based on the product of the second masking matrix and the second near-end collected signal.

[0138] The echo cancellation signal unit 503 may obtain a near-end collected signal after echo cancellation according to the fourth signal.

[0139] According to an exemplary embodiment of the present disclosure, the echo cancellation signal unit 503 may obtain a fourth signal in the time domain as the near-end acquisition signal after echo cancellation. For example, the echo cancellation signal unit 503 may perform an inverse time-frequency transform on the fourth signal to obtain the near-end acquisition signal after echo cancellation. The inverse time-frequency transform in the exemplary embodiment of the present disclosure may be a short-time inverse Fourier transform (ISTFT).

[0140] Figure 6 is a block diagram of an electronic device 600 according to an exemplary embodiment.

[0141] Reference Figure 6 The electronic device 600 includes at least one memory 601 and at least one processor 602, wherein the at least one memory 601 stores a set of computer-executable instructions. When the computer-executable instruction set is executed by the at least one processor 602, the training method of the echo cancellation model or the echo cancellation method according to the present disclosure is executed.

[0142] As an example, the electronic device 600 may be a PC, a tablet device, a personal digital assistant, a smart phone, or other device capable of executing the above-mentioned instruction set. Here, the electronic device 600 is not necessarily a single electronic device, but may also be any device or circuit collection capable of executing the above-mentioned instructions (or instruction set) individually or in combination. The electronic device 600 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device interconnected with a local or remote (e.g., via wireless transmission) interface.

[0143] In electronic device 600, processor 602 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0144] The processor 602 can execute instructions or codes stored in the memory 601, wherein the memory 601 can also store data. Instructions and data can also be sent and received over the network via the network interface device, wherein the network interface device can use any known transmission protocol.

[0145] The memory 601 may be integrated with the processor 602, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. Furthermore, the memory 601 may comprise a separate device, such as an external disk drive, a storage array, or any other storage device usable by a database system. The memory 601 and the processor 602 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, or the like, such that the processor 602 can access files stored in the memory.

[0146] In addition, the electronic device 600 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 600 may be connected to each other via a bus and / or a network.

[0147] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are executed by at least one processor, the at least one processor is prompted to perform the training method of the echo cancellation model or the echo cancellation method according to the present disclosure. Examples of computer-readable storage media here include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or ultra-fast digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device configured to store the computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.

[0148] According to exemplary embodiments of the present disclosure, a computer program product may also be provided. Instructions in the computer program product may be executed by a processor of a computer device to implement the echo cancellation model training method or the echo cancellation method according to the present disclosure.

[0149] Figure 7 is a block diagram of a smart speaker 700 according to an exemplary embodiment of the present disclosure.

[0150] Reference Figure 7According to the exemplary embodiment of the present disclosure, the smart speaker 700 includes the echo cancellation device 500 shown in the present disclosure. In a specific implementation process, the smart speaker 700 can be applied to video conferencing, for example. In this scenario, the smart speaker 700 may include a signal acquisition module, a signal processing module and a signal output module, wherein the signal acquisition module can collect audio signals in the environment (for example, the signal acquisition module can collect, but is not limited to, microphone signals, which may include the near-end user sound signal collected by the near-end microphone and the audio signal collected by the near-end microphone and transmitted from the far end through the network link and played through the near-end speaker), the signal acquisition module may include, but is not limited to, a microphone, and the signal processing module may process the audio signal collected by the signal acquisition module (for example, including using the echo cancellation method according to the exemplary embodiment of the present disclosure). The smart speaker 700 is a device that can process audio signals and output the audio signals processed by the signal processing module (for example, the signal output module can output the audio signals processed by the signal processing module in the environment through a speaker). Of course, the smart speaker 700 can also be applied to other scenarios, such as home environments, etc., and there is no limitation on this. In different usage environments, the composition structure of the smart speaker 700 may be different. It should be clear that as long as the smart speaker uses the echo cancellation method shown in the present disclosure to perform echo cancellation, it falls within the scope of protection of the present disclosure.

[0151] Figure 8 is a block diagram of a smart speaker 800 according to an exemplary embodiment of the present disclosure.

[0152] Reference Figure 8 According to an exemplary embodiment of the present disclosure, the smart speaker 800 includes at least one memory 801 for storing instructions, at least one audio signal processor 802, at least one audio signal collector 803 and at least one audio signal outputter 804, wherein the instructions, when executed by the at least one audio signal processor 802, prompt the at least one audio signal processor 802 to perform the echo cancellation method according to the present disclosure.

[0153] According to the training method and device of the echo cancellation model and the echo cancellation method and device disclosed in the present invention, the parameters of the echo cancellation model can be adjusted based on the first prediction loss calculated from the residual echo component and the speech signal component separated from the second signal predicted by the echo cancellation model, so as to train the echo cancellation model. This can help guide the training process of the echo cancellation model, and perform additional suppression on the residual echo without introducing speech clipping and ensuring the speech quality, thereby improving the echo cancellation amount, solving the problems of insufficient echo cancellation and double-speaking echo leakage in existing solutions, and achieving a higher signal-to-response ratio and ERLE.

[0154] In addition, according to the echo cancellation model training method and device and the echo cancellation method and device disclosed in the present invention, the mapping principle can be used to separate the correctly predicted speech signal component and the residual echo component in the second signal, and the signal-to-echo ratio can be calculated to provide more penalty for the residual echo.

[0155] Furthermore, according to the disclosed echo cancellation model training method and apparatus, as well as the echo cancellation method and apparatus, the first prediction loss is scale-invariant, meaning the magnitude of the prediction loss is unaffected by signal amplitude. This forces the model to optimize the prediction loss from the perspective of suppressing residual echo rather than overall signal amplitude. Even for small-amplitude speech audio signals, to ensure a high signal-to-response ratio, speech clipping is avoided, thereby preserving sound quality. Furthermore, the echo leakage that occurs in dual-talk scenarios in existing solutions is eliminated, preventing excessive suppression of the speech in the dual-talk portion.

[0156] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0157] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A training method for an echo cancellation model, characterized in that: include: Obtaining a first signal according to the near-end signal; Inputting the first signal into an echo cancellation model, and obtaining a second signal based on an output of the echo cancellation model, wherein the second signal is a predicted signal of the first signal after echo cancellation; Determining a prediction loss of an echo cancellation model based on the first signal and the second signal, wherein the prediction loss includes a first prediction loss calculated based on a residual echo component and a speech signal component separated from the second signal; Parameters of the echo cancellation model are adjusted according to the prediction loss.

2. The training method according to claim 1, wherein: The output of the echo cancellation model includes a first masking matrix; The obtaining of the second signal based on the output of the echo cancellation model includes: Performing time-frequency transformation on the near-end signal to obtain a second near-end signal; The second signal is obtained according to the product of the first masking matrix and the second near-end signal.

3. The training method according to claim 1, wherein: The obtaining of the first signal according to the near-end signal includes: performing linear echo cancellation on the near-end signal to obtain a first near-end signal; Performing time-frequency transformation on the first near-end signal to obtain a first signal; The near-end signal includes a speech signal and an echo signal; the first near-end signal includes the speech signal and a residual echo signal; and the first signal includes a frequency-domain speech signal and a frequency-domain residual echo signal.

4. The training method according to claim 3, wherein: The first prediction loss is determined by the following steps: Mapping the second signal toward the direction of the frequency domain residual echo signal to obtain a frequency domain residual echo signal mapping vector, and using the frequency domain residual echo signal mapping vector as the residual echo component separated from the second signal; Mapping the second signal to the direction of the frequency domain speech signal to obtain a frequency domain speech signal mapping vector, and using the frequency domain speech signal mapping vector as a speech signal component separated from the second signal; The first prediction loss is determined according to the frequency-domain residual echo signal mapping vector and the frequency-domain speech signal mapping vector.

5. The training method according to claim 4, wherein: The determining the first prediction loss according to the frequency domain residual echo signal mapping vector and the frequency domain speech signal mapping vector includes: Obtaining an average signal-to-echo ratio of the second signal according to the frequency-domain residual echo signal mapping vector and the frequency-domain speech signal mapping vector; The first prediction loss is determined according to the inverse of an average signal-to-back ratio of the second signal.

6. The training method according to claim 5, wherein: The first prediction loss is expressed as: Among them, loss res represents the first prediction loss, C proj (n, k) represents the frequency domain speech signal mapping vector at the time-frequency point (n, k), R proj (n, k) represents the frequency domain residual echo signal mapping vector at the time-frequency point (n, k), the time-frequency point (n, k) represents the kth frequency point of the nth frame, N represents the number of frames, K represents the number of frequency points, 1≤n≤N, 1≤k≤K, represents the average signal-to-back ratio of the second signal.

7. The training method according to claim 1, wherein: The predicted loss further includes a second predicted loss, which includes at least one loss value related to echo cancellation, wherein the predicted loss is obtained according to the first predicted loss and the second predicted loss.

8. An echo cancellation method, characterized in that: include: Obtaining a third signal according to the acquired near-end acquisition signal; Inputting the third signal into an echo cancellation model trained by the echo cancellation model training method according to any one of claims 1 to 7, and obtaining a fourth signal based on an output of the trained echo cancellation model, wherein the fourth signal is a predicted signal of the third signal after echo cancellation; A near-end acquisition signal after echo cancellation is obtained according to the fourth signal.

9. The echo cancellation method according to claim 8, wherein: The obtaining of the third signal according to the acquired near-end acquisition signal includes: Performing linear echo cancellation on the near-end collected signal to obtain a first near-end collected signal; Perform time-frequency transformation on the first near-end collected signal to obtain a third signal.

10. The echo cancellation method according to claim 8, wherein: The output of the trained echo cancellation model includes a second masking matrix; The obtaining of a fourth signal based on the output of the trained echo cancellation model includes: Performing time-frequency transformation on the near-end acquisition signal to obtain a second near-end acquisition signal; The fourth signal is obtained according to the product of the second masking matrix and the second near-end acquisition signal.

11. The echo cancellation method according to claim 8, wherein: The step of obtaining the near-end collected signal after echo cancellation according to the fourth signal includes: Performing an inverse time-frequency transform on the fourth signal to obtain a near-end acquisition signal after echo cancellation.

12. A training device for an echo cancellation model, characterized in that: include: A first signal determining unit is configured to: obtain a first signal according to a near-end signal; a first model prediction unit configured to: input the first signal into an echo cancellation model, and obtain a second signal based on an output of the echo cancellation model, wherein the second signal is a predicted signal of the first signal after echo cancellation; a loss determining unit configured to: determine a prediction loss of an echo cancellation model based on the first signal and the second signal, wherein the prediction loss includes a first prediction loss, and the first prediction loss is calculated based on a residual echo component and a speech signal component separated from the second signal; The parameter adjustment unit is configured to adjust the parameters of the echo cancellation model according to the prediction loss.

13. The training device according to claim 12, wherein: The output of the echo cancellation model includes a first masking matrix; The first model prediction unit is configured as follows: Performing time-frequency transformation on the near-end signal to obtain a second near-end signal; The second signal is obtained according to the product of the first masking matrix and the second near-end signal.

14. The training device according to claim 12, wherein: The first signal determining unit is configured to: performing linear echo cancellation on the near-end signal to obtain a first near-end signal; Performing time-frequency transformation on the first near-end signal to obtain a first signal; The near-end signal includes a speech signal and an echo signal; the first near-end signal includes the speech signal and a residual echo signal; and the first signal includes a frequency-domain speech signal and a frequency-domain residual echo signal.

15. The training device according to claim 14, characterized in that The loss determination unit is configured to: Mapping the second signal toward the direction of the frequency domain residual echo signal to obtain a frequency domain residual echo signal mapping vector, and using the frequency domain residual echo signal mapping vector as the residual echo component separated from the second signal; Mapping the second signal to the direction of the frequency domain speech signal to obtain a frequency domain speech signal mapping vector, and using the frequency domain speech signal mapping vector as a speech signal component separated from the second signal; The first prediction loss is determined according to the frequency-domain residual echo signal mapping vector and the frequency-domain speech signal mapping vector.

16. The training device according to claim 15, characterized in that The loss determination unit is configured to: Obtaining an average signal-to-echo ratio of the second signal according to the frequency-domain residual echo signal mapping vector and the frequency-domain speech signal mapping vector; The first prediction loss is determined according to the inverse of an average signal-to-back ratio of the second signal.

17. The training device according to claim 16, characterized in that The first prediction loss is expressed as: Among them, Loss res represents the first prediction loss, C proj (n, k) represents the frequency domain speech signal mapping vector at the time-frequency point (n, k), R proj (n, k) represents the frequency domain residual echo signal mapping vector at the time-frequency point (n, k), the time-frequency point (n, k) represents the kth frequency point of the nth frame, N represents the number of frames, K represents the number of frequency points, 1≤n≤N, 1≤k≤K, represents the average signal-to-back ratio of the second signal.

18. The training device according to claim 12, wherein: The predicted loss further includes a second predicted loss, which includes at least one loss value related to echo cancellation, wherein the predicted loss is obtained according to the first predicted loss and the second predicted loss.

19. An echo cancellation device, characterized in that: include: A third signal determining unit is configured to: obtain a third signal according to the acquired near-end acquisition signal; a second model prediction unit configured to: input the third signal into an echo cancellation model trained by the echo cancellation model training method according to any one of claims 1 to 7, and obtain a fourth signal based on an output of the trained echo cancellation model, wherein the fourth signal is a predicted signal of the third signal after echo cancellation; The echo cancellation signal unit is configured to obtain a near-end collected signal after echo cancellation according to the fourth signal.

20. The echo cancellation device according to claim 19, wherein: The third signal determining unit is configured to: Performing linear echo cancellation on the near-end collected signal to obtain a first near-end collected signal; Perform time-frequency transformation on the first near-end collected signal to obtain a third signal.

21. The echo cancellation device according to claim 19, wherein: The output of the trained echo cancellation model includes a second masking matrix; The second model prediction unit is configured as follows: Performing time-frequency transformation on the near-end acquisition signal to obtain a second near-end acquisition signal; The fourth signal is obtained according to the product of the second masking matrix and the second near-end acquisition signal.

22. The echo cancellation device according to claim 19, wherein: The echo cancellation signal unit is configured as follows: Performing an inverse time-frequency transform on the fourth signal to obtain a near-end acquisition signal after echo cancellation.

23. An electronic device, characterized in that: include: at least one processor; at least one memory storing computer-executable instructions, When the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to execute the echo cancellation model training method according to any one of claims 1 to 7 or the echo cancellation method according to any one of claims 8 to 11.

24. A computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by at least one processor, the at least one processor is prompted to perform the echo cancellation model training method according to any one of claims 1 to 7 or the echo cancellation method according to any one of claims 8 to 11.

25. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by at least one processor, the method for training an echo cancellation model according to any one of claims 1 to 7 or the echo cancellation method according to any one of claims 8 to 11 is implemented.

26. A smart speaker, characterized in that: The method comprises the echo cancelling device according to any one of claims 19 to 22.

27. A smart speaker, characterized in that: include: at least one audio signal processor; at least one memory storing instructions; at least one audio signal collector; at least one audio signal output device, Wherein, when the instructions are executed by the at least one audio signal processor, the at least one audio signal processor is prompted to perform the echo cancellation method according to any one of claims 8 to 11.

Citation Information

Patent Citations

  • Echo cancellation method and device and electronic equipment

    CN113113038A

  • Interference signal elimination model training method and interference signal elimination method and device

    CN113257267A

  • Network model training method, echo cancellation method and equipment

    CN113744748A