Echo cancellation method, device, electronic device and storage medium
By pre-training the echo cancellation model to extract nonlinear features on different acoustic devices, the problem of insufficient equipment generalization capabilities of deep learning methods is solved, and the echo cancellation effect with low complexity is achieved.
Patent Information
- Application Number
- CN202210787555.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-07-04
AI Technical Summary
The existing deep learning methods lack equipment generalization capabilities in echo cancellation, and data admission and training are required for a single model, resulting in high complexity and inapplicable to multiple devices.
The pre-trained echo cancellation model is used to assist in echo cancellation through nonlinear echo feature extraction task, and the neural network module is used to extract nonlinear features on different machines to achieve low complexity echo cancellation.
Echo cancellation with strong generalization capabilities on different acoustic devices is achieved, reducing the complexity of model training and deployment, and improving the echo cancellation effect.
Smart Images

Figure CN115050383B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of data processing technology, and in particular to an echo cancellation method, device, electronic device, and storage medium. Background Art
[0002] At present, acoustic echo cancellation is mainly divided into two parts: linear echo elimination and nonlinear echo suppression. The former mainly relies on traditional signal processing methods, while deep learning methods are gradually replacing traditional algorithms and becoming the mainstream research direction of nonlinear echo suppression.
[0003] Existing deep learning methods lack the ability to generalize across data from different aircraft models. They require sufficient data collected for a single aircraft model to train and adjust the model accordingly, or they employ complex models to perform detailed training on a sufficient number of aircraft model data to meet the echo cancellation requirements of the trained aircraft model. Therefore, enhancing the universality of echo cancellation models is particularly important. Summary of the Invention
[0004] The embodiments of the present disclosure provide an echo cancellation method, apparatus, electronic device, and storage medium to implement a low-complexity echo cancellation model with strong device generalization capability.
[0005] In a first aspect, an embodiment of the present disclosure provides an echo cancellation method, the method comprising:
[0006] Determining a speech signal to be processed by a pre-trained echo cancellation model;
[0007] Using the pre-trained echo cancellation model to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed; wherein the nonlinear echo feature extraction task is used to assist in nonlinear echo suppression during the echo cancellation process;
[0008] The nonlinear echo feature extraction task and the echo cancellation task are performed by the pre-trained echo cancellation model, and an echo-cancelled speech signal corresponding to the speech signal to be processed is output.
[0009] In a second aspect, an embodiment of the present disclosure further provides an echo cancellation device, the device comprising:
[0010] A speech signal determination module, used to determine the speech signal to be processed by the pre-trained echo cancellation model;
[0011] an echo cancellation execution module, configured to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed using the pre-trained echo cancellation model; wherein the nonlinear echo feature extraction task is used to assist in nonlinear echo suppression during the echo cancellation process;
[0012] The echo cancellation output module is used to perform the nonlinear echo feature extraction task and the echo cancellation task through the pre-trained echo cancellation model, and output the echo-cancelled speech signal corresponding to the speech signal to be processed.
[0013] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the echo cancellation method according to any one of the above embodiments.
[0017] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable medium, wherein the computer-readable medium stores computer instructions, and the computer instructions are used to enable a processor to implement the echo cancellation method described in any one of the above embodiments when executed.
[0018] The technical solution of the disclosed embodiments first determines a pre-trained echo cancellation model's speech signal to be processed. The pre-trained echo cancellation model then performs nonlinear echo feature extraction and echo cancellation on the speech signal. Finally, the pre-trained echo cancellation model performs these nonlinear echo feature extraction and echo cancellation tasks, outputting an echo-cancelled speech signal corresponding to the speech signal to be processed. This technical solution processes the speech signal using a low-complexity echo cancellation model with strong device generalization capabilities, resulting in a speech signal with excellent echo cancellation.
[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 A flowchart of an echo cancellation method provided by an embodiment of the present disclosure;
[0022] Figure 2A flowchart of a method for training an echo cancellation model provided in an embodiment of the present disclosure;
[0023] Figure 3 A flowchart of another echo cancellation model training method provided in an embodiment of the present disclosure;
[0024] Figure 4 Schematic diagram of the principle of determining the characteristics of training samples provided by an embodiment of the present disclosure;
[0025] Figure 5 A flowchart of another echo cancellation model training method provided in an embodiment of the present disclosure;
[0026] Figure 6 A schematic diagram of the principle of echo cancellation model training provided by an embodiment of the present disclosure;
[0027] Figure 7 A flowchart of another echo cancellation method provided by an embodiment of the present disclosure;
[0028] Figure 8 This is a structural block diagram of an echo cancellation device provided by an embodiment of the present disclosure;
[0029] Figure 9 A structural block diagram of an electronic device for implementing the echo cancellation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0031] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0032] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0034] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0035] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0036] In the following embodiments, each embodiment provides optional features and examples. The various features described in the embodiments can be combined to form multiple optional solutions. Each numbered embodiment should not be regarded as just one technical solution. In addition, the embodiments and features in the embodiments of this disclosure can be combined with each other unless there is a conflict.
[0037] Figure 1 This is a flowchart of an echo cancellation method provided by an embodiment of the present disclosure. The technical solution of this embodiment can be applied to the case where an echo cancellation model is applied. The method can be executed by an echo cancellation device, which can be implemented in software and / or hardware and integrated into any electronic device with network communication function. Figure 1 As shown, the echo cancellation method of this embodiment may include the following steps S110-S130:
[0038] S110: Determine a speech signal to be processed for a pre-trained echo cancellation model.
[0039] In this solution, the pre-trained echo cancellation model does not need to go through the data collection, separate training, or online training process for each model. Instead, it can extract the nonlinear feature vectors of the device through the neural network module on different machines to achieve better echo cancellation effects. It is a universal, low-complexity pre-trained echo cancellation model.
[0040] The speech signal to be processed may be composed of various forms of original microphone signals and playback signals of an acoustic device; or may be various forms of signal combinations of the original microphone signals and playback signals after various signal processing steps.
[0041] Specifically, the signals from the microphone and the loudspeaker in the acoustic device may be collected to obtain the speech signal to be processed.
[0042] S120. Use the pre-trained echo cancellation model to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed; wherein the nonlinear echo feature extraction task is used to assist in nonlinear echo suppression during the echo cancellation process.
[0043] In this scheme, the nonlinear echo feature extraction task is used to learn the nonlinear characteristics of different acoustic devices in order to integrate them into the echo cancellation task. In the process of executing the echo cancellation task based on the voice signal to be processed, it can assist the echo cancellation task in performing nonlinear echo suppression during echo cancellation, thereby achieving better echo cancellation effect.
[0044] S130: Execute the nonlinear echo feature extraction task and the echo cancellation task through the pre-trained echo cancellation model, and output the echo-cancelled speech signal corresponding to the speech signal to be processed.
[0045] According to the technical solutions of the embodiments of this disclosure, a pre-trained echo cancellation model processes the speech signal to be processed, outputting an echo-cancelled speech signal corresponding to the speech signal to be processed. This low-complexity, highly generalizable echo cancellation model processes the speech signal to be processed, resulting in a speech signal with excellent echo cancellation.
[0046] Figure 2 This is a flow chart of a method for training an echo cancellation model provided by an embodiment of the present disclosure. The technical solution of this embodiment can be applied to the case of training an echo cancellation model to be trained, and realizing a low-complexity pre-trained echo cancellation model with strong device generalization capability. Figure 2 As shown, the training method of the echo cancellation model of this embodiment may include the following steps S210-S230:
[0047] S210: Determine the training sample features used by the echo cancellation model to be trained.
[0048] In this solution, the echo cancellation model to be trained does not need to go through the process of data collection, separate training or online training for each model. Instead, the nonlinear feature vectors of the device can be extracted through the neural network module on different machines to achieve better echo cancellation effect.
[0049] The training sample features can be features composed of various forms of the original microphone signal and playback signal of the acoustic device, or can be features composed of various forms of signals obtained by combining the original microphone signal and playback signal after various signal processing steps. Specifically, the training sample features can be obtained by performing signal processing on the speech signal collected by the microphone of the acoustic device and the speech signal before playback by the speaker.
[0050] S220 : Control the to-be-trained echo cancellation model to perform a nonlinear echo feature extraction training task and an echo cancellation training task based on the training sample features.
[0051] The nonlinear echo feature extraction training task is used to assist in nonlinear echo suppression during the execution of the echo cancellation training task.
[0052] In this scheme, the nonlinear echo feature extraction training task is used to learn the nonlinear characteristics of different acoustic devices in order to integrate them into the echo cancellation training task. In the process of executing the echo cancellation training task based on the training sample feature control, it can assist the echo cancellation training task in performing nonlinear echo suppression during echo cancellation, thereby achieving better echo cancellation effect.
[0053] S230: Adjust the echo cancellation model to be trained according to the nonlinear echo feature extraction training task and the echo cancellation training task to obtain a pre-trained echo cancellation model after training and updating.
[0054] According to the technical solution of the embodiment of the present disclosure, when the echo cancellation model to be trained is trained, the echo cancellation training task is assisted in training by introducing a nonlinear echo feature extraction training task, thereby obtaining a universal low-complexity pre-trained echo cancellation model to achieve better echo cancellation effect. The present application solution does not need to carry out data acquisition, separate training or online training processes for each model, but can extract the nonlinear feature vector of the device through the neural network module on different machines to achieve better echo cancellation effect. The model can also learn nonlinear features offline to assist in achieving the optimal echo cancellation effect.
[0055] Figure 3 This is a flowchart of another method for training an echo cancellation model provided by an embodiment of the present disclosure. The technical solution of this embodiment further optimizes the process of determining the characteristics of the training samples used by the echo cancellation model to be trained in the above embodiment. This embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 3 As shown, the training method of the echo cancellation model of this embodiment may include the following steps S310-S340:
[0056] S310: Determine a preset amount of voice data obtained through different acoustic devices respectively; the voice data includes a voice signal collected by a microphone in the acoustic device and a voice signal before being played by a speaker.
[0057] In this solution, a small amount of voice data can be collected for each of at least a preset number of acoustic device models (for example, at least 30 widely varying acoustic device models). The amount and length of voice data collected for each acoustic device model can be adjusted based on actual needs and model performance. For example, each acoustic device can be set to collect 10 voice data items, each of which can be 2 minutes long.
[0058] In this embodiment, there can be multiple types of acoustic devices, which can be adjusted according to needs and model performance. By setting requirements for the type of acoustic device, each acoustic device only needs a small amount of voice data.
[0059] S320: Perform preset feature processing on the speech data obtained through different acoustic devices to obtain speech feature signals corresponding to the different acoustic devices, which constitute training sample features used by the echo cancellation model to be trained.
[0060] For example, Figure 4 FIG. 1 is a schematic diagram of the principle of determining the characteristics of training samples provided by an embodiment of the present disclosure, such as Figure 4 As shown, the speech data obtained by different acoustic devices is input into the feature processing M. The feature processing M processes the speech data and outputs speech feature signals corresponding to the acoustic devices, which constitute the training sample features used by the echo cancellation model to be trained.
[0061] Optionally, the speech feature signal is represented in the following forms: time domain, short-time Fourier transform domain, Mel feature domain and Bark feature domain.
[0062] The independent variable in the time domain is time, meaning the horizontal axis is time and the vertical axis is the signal's change. Its dynamic signal is a function that describes the signal's value at different moments. The independent variable in the frequency domain is frequency, meaning the horizontal axis is frequency and the vertical axis is the amplitude of that frequency signal. The short-time Fourier transform is a mathematical transformation related to the Fourier transform, used to determine the frequency and phase of a localized sinusoidal wave in a time-varying signal. The Mel feature domain is used to describe the nonlinear mapping of pitch perception. The Bark feature domain, measured in Hz, is used to convert physical frequencies to psychoacoustic frequencies.
[0063] In this solution, the feature processing M can output speech feature signals in different forms, which constitute the training sample features used by the echo cancellation model to be trained.
[0064] By combining speech signals of different forms into training sample features used by the echo cancellation model to be trained, the learning ability of the echo cancellation model can be improved, thereby achieving better echo cancellation effects.
[0065] Optionally, the preset feature processing includes at least one of the following: delay estimation processing, linear echo cancellation processing, signal domain transformation processing and signal splicing processing; the signal splicing processing is used to splice and combine speech feature signals in different signal domains.
[0066] In this solution, delay estimation processing is used to roughly align the time of the voice signal collected by the microphone in the acoustic device and the voice signal before being played by the speaker when the two are delayed significantly; linear echo cancellation processing is used to eliminate linear echo from the voice signal through a linear filter; and signal domain transformation processing is used to transform the time domain voice signal into the short-time Fourier transform domain, Mel feature domain, Bark feature domain, etc.
[0067] Optionally, the preset feature processing may also include signal amplitude processing and signal state adjustment processing. Signal amplitude processing is used to adjust the amplitude of the voice signal collected by the microphone and the voice signal before being played by the speaker. The signal state adjustment component can classify the voice signal into a far-end single-talk signal (without a near-end signal), a near-end single-talk signal (without an echo or playback signal), a dual-talk signal, and a headphone-type signal (without an echo signal) according to a specific ratio.
[0068] In this embodiment, the speech feature signal can be various forms (time domain, short-time Fourier transform domain, Mel feature domain, and Bark feature domain, etc.) including the speech signal collected by the microphone and the speech signal before being played by the speaker, or it can be various forms of a combination of signals that have undergone various feature processing (delay estimation processing, linear echo cancellation processing, signal domain transformation processing, signal splicing processing, signal amplitude processing, and signal state adjustment processing, etc.).
[0069] By performing feature processing on the voice data obtained from different acoustic devices, the voice feature signals corresponding to the different acoustic devices are obtained to form the training sample features used by the echo cancellation model to be trained. The amount of data required for each model is small, and there is no need to transmit the data to the server, avoiding the occurrence of data security and privacy issues.
[0070] S330. Control the to-be-trained echo cancellation model to execute a nonlinear echo feature extraction training task and an echo cancellation training task based on the training sample features; wherein the nonlinear echo feature extraction training task is used to assist in nonlinear echo suppression during the execution of the echo cancellation training task.
[0071] S340: Adjust the echo cancellation model to be trained according to the nonlinear echo feature extraction training task and the echo cancellation training task to obtain a pre-trained echo cancellation model after training and updating.
[0072] According to the technical solution of the embodiment of the present disclosure, a small amount of voice data is obtained for each type of acoustic device, and the voice data is feature processed to obtain the corresponding voice feature signal to form the training sample features used by the echo cancellation model to be trained. When the echo cancellation model to be trained is trained, the echo cancellation training task is supplemented by introducing a nonlinear echo feature extraction training task to obtain a universal low-complexity pre-trained echo cancellation model to achieve better echo cancellation effect. The present application solution does not need to conduct data collection, separate training or online training processes for each model. Instead, the nonlinear feature vector of the device can be extracted through the neural network module on different types of machines to achieve better echo cancellation effect.
[0073] Figure 5 This is a flowchart of another echo cancellation model training method provided by the embodiment of the present disclosure. The technical solution of this embodiment further optimizes the process of controlling the trained echo cancellation model to perform nonlinear echo feature extraction training tasks and echo cancellation training tasks based on training sample features in the above embodiment. This embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 5 As shown, the training method of the speech translation model of this embodiment may include the following steps S510-S520:
[0074] S510: Input the first sample feature in the training sample features into the echo cancellation model to be trained, and perform a nonlinear echo feature extraction training task.
[0075] In this embodiment, first and second sample features obtained by feature processing speech signals of different contents collected by the same acoustic device can be used. The first sample features are then input into the echo cancellation model to be trained to perform a nonlinear echo feature extraction training task, thereby obtaining a nonlinear echo feature expression. The nonlinear echo feature expression may refer to a representation of the nonlinear echo feature.
[0076] S520: synchronously inputting the second sample feature in the training sample feature into the echo cancellation model to be trained, and controlling the echo cancellation model to be trained to perform the echo cancellation training task based on the nonlinear echo feature expression obtained by performing the nonlinear echo feature extraction training task;
[0077] The first sample feature and the second sample feature are obtained by performing feature processing on speech signals of different contents collected by the same acoustic device.
[0078] In this solution, the nonlinear echo feature expression obtained from the nonlinear echo feature extraction training task and the second sample feature are combined and used as input to the echo cancellation training task in the to-be-trained echo cancellation model, thereby training the to-be-trained echo cancellation model. The nonlinear echo feature expression and the second sample feature can be combined by concatenating a dimension of features from a node in the network or by performing mathematical operations. The specific method of combination depends primarily on the to-be-trained echo cancellation model.
[0079] Optionally, the first sample feature and the second sample feature are obtained by processing the same or different preset features respectively; the first sample feature includes at least echo data, which is used to extract the nonlinear echo features of the acoustic device from the echo data through a nonlinear echo feature extraction training task.
[0080] Specifically, speech data collected by the same acoustic device can be processed based on feature processing M to obtain first sample features and second sample features. If feature processing M uses different processing methods to process the speech data, different first sample features and second sample features can be obtained; if feature processing M uses the same processing method to process the speech data, the same first sample features and second sample features can be obtained.
[0081] By processing the sample features, the nonlinear echo characteristics of the acoustic device can be extracted from the echo data, achieving better echo cancellation effect.
[0082] Optionally, adjusting the echo cancellation model to be trained according to the nonlinear echo feature extraction training task and the echo cancellation training task includes:
[0083] Determining a first loss function value corresponding to the nonlinear echo feature extraction training task and a second loss function corresponding to the echo cancellation training task;
[0084] Adjusting network parameters of the echo cancellation model to be trained according to the first loss function value and the second loss function;
[0085] The first loss function value can prompt the echo cancellation model to be trained to learn the nonlinear characteristics of the acoustic device when performing the nonlinear echo feature extraction training task.
[0086] The first loss function and the second loss function may be mean square error loss (mean square error loss function) or l1 loss (regression loss function).
[0087] Specifically, the first sample feature can be used as input to perform feature extraction through a nonlinear feature extraction network in the echo cancellation model to be trained, and a first loss function value corresponding to the nonlinear echo feature extraction training task can be determined. The nonlinear feature extraction network can be any network, including CNNs (Convolutional Neural Networks), DNNs (Deep Neural Networks), RNNs (Recurrent Neural Networks), and combinations of various structures thereof.
[0088] In this embodiment, after performing echo cancellation on the second sample feature using an echo cancellation network in the to-be-trained echo cancellation model, a second loss function corresponding to the echo cancellation training task can be determined. The echo cancellation network can also be any network, including a CNN, a DNN, an RNN, and combinations of various structures thereof.
[0089] In this solution, the first loss function and the second loss function can be weighted together to obtain a weighted result, and the network parameters of the echo cancellation model to be trained can be adjusted based on the weighted result. Alternatively, the network parameters of the echo cancellation model to be trained can be adjusted based on the first loss function and the second loss function, respectively, to optimize the first loss function and the second loss function as much as possible.
[0090] By determining the first loss function value and the second loss function, the training of the to-be-trained echo cancellation model can be optimized, thereby realizing a to-be-trained echo cancellation model with low complexity and strong device generalization capability.
[0091] Optionally, determining a first loss function value corresponding to the nonlinear echo feature extraction training task includes:
[0092] After performing nonlinear echo feature extraction on the first sample feature through the nonlinear feature extraction network in the echo cancellation model to be trained, performing nonlinear feature conversion on the extracted nonlinear echo feature;
[0093] Determining a first loss function value corresponding to the nonlinear echo feature extraction training task based on the nonlinear degree expression obtained by the nonlinear feature conversion and the pre-labeled nonlinear degree expression corresponding to the first sample feature;
[0094] The nonlinear degree expression is used to describe the nonlinear degree of the acoustic device.
[0095] In this embodiment, the nonlinear echo feature can be converted into a nonlinear feature based on the pre-labeled nonlinear degree expression corresponding to the first sample feature to obtain a nonlinear degree expression that matches the pre-labeled nonlinear degree expression corresponding to the first sample feature. The linear echo feature extraction training task can be trained based on the nonlinear degree expression obtained by the nonlinear feature conversion and the pre-labeled nonlinear degree expression corresponding to the first sample feature to determine the first loss function value corresponding to the nonlinear echo feature extraction training task.
[0096] By determining the first loss function value and the second loss function, the training of the echo cancellation model to be trained can be optimized.
[0097] Optionally, determining a second loss function corresponding to the echo cancellation training task includes:
[0098] After performing echo cancellation on the second sample feature through the echo cancellation network in the to-be-trained echo cancellation model, performing feature post-processing on the echo cancellation result to obtain an echo cancellation result after feature post-processing;
[0099] A second loss function value corresponding to the echo cancellation training task is determined based on the echo cancellation result after feature post-processing and the pre-labeled echo cancellation result corresponding to the second sample feature.
[0100] Feature post-processing may include spectrum smoothing or energy threshold determination and limitation processing.
[0101] In this embodiment, the echo cancellation result after feature post-processing may be a time domain waveform signal, a short-time Fourier transform domain signal, or a short-time Fourier transform domain masking value.
[0102] Specifically, the second sample feature can be used as the input of the echo cancellation network in the echo cancellation model to be trained to obtain the echo cancellation result, and the echo cancellation result is subjected to feature post-processing. The echo cancellation result and the pre-labeled echo cancellation result corresponding to the second sample feature are kept in the same structure to obtain the echo cancellation result after feature post-processing.
[0103] By determining the second loss function value corresponding to the echo cancellation training task, the training of the echo cancellation model to be trained can be optimized.
[0104] Optionally, determining a second loss function corresponding to the echo cancellation training task includes:
[0105] A second loss function value corresponding to the echo cancellation training task is determined based on an echo cancellation result after the echo cancellation network in the to-be-trained echo cancellation model performs echo cancellation on the second sample feature and a pre-labeled echo cancellation result corresponding to the second sample feature.
[0106] In this scheme, the echo cancellation result may not be subjected to feature post-processing, and the second loss function value corresponding to the echo cancellation training task may be determined directly based on the echo cancellation result after the echo cancellation network in the echo cancellation model to be trained performs echo cancellation on the second sample feature and the pre-labeled echo cancellation result corresponding to the second sample feature.
[0107] By determining the second loss function value corresponding to the echo cancellation training task, the training of the echo cancellation model to be trained can be optimized.
[0108] For example, Figure 6 A schematic diagram of the principle of echo cancellation model training provided by the embodiment of the present disclosure, such as Figure 6 As shown, the echo cancellation model to be trained includes an echo cancellation network A, feature post-processing B, a second loss function calculation C, a nonlinear feature extraction network D, a nonlinear feature conversion network E, and a first loss function calculation F. The nonlinear feature extraction network D, the nonlinear feature conversion network E, and the first loss function calculation F are used to perform the nonlinear echo feature extraction training task; the echo cancellation network A, feature post-processing B, and the second loss function calculation C are used to perform the echo cancellation training task.
[0109] The first sample feature serves as the input to the nonlinear feature extraction network D, which then extracts the nonlinear echo feature. The nonlinear feature conversion network E then performs nonlinear feature conversion on the extracted nonlinear echo feature to obtain a nonlinear degree representation. This nonlinear degree representation is then used as the input to the first loss function calculation F to determine the first loss function value corresponding to the nonlinear echo feature extraction training task. In this case, the entire nonlinear extraction network performs supervised learning. If ideal nonlinear degree measurement and calculation are not possible, the nonlinear feature conversion network E and the first loss function calculation F can be omitted, resulting in unsupervised learning for the entire nonlinear extraction network.
[0110] The second sample feature serves as the input to echo cancellation network A. The output of nonlinear feature extraction network D is combined with a portion of echo cancellation network A. After the comprehensive action of the network model, this is used as the input to feature post-processing B. This post-processed echo cancellation result is obtained. A second loss function calculation C is then used to determine the second loss function value corresponding to the echo cancellation training task based on the post-processed echo cancellation result and the pre-labeled echo cancellation result corresponding to the second sample feature. Feature post-processing B can also be omitted, and the final input to the second loss function calculation C can be obtained directly from the echo cancellation network A.
[0111] The technical solution of the disclosed embodiments first determines the training sample features used by the to-be-trained echo cancellation model. Based on the training sample features, the to-be-trained echo cancellation model is then controlled to perform nonlinear echo feature extraction and echo cancellation training tasks. Furthermore, the to-be-trained echo cancellation model is adjusted based on these tasks to obtain a pre-trained echo cancellation model after training and updating. This technical solution enables the implementation of a low-complexity pre-trained echo cancellation model with strong device generalization capabilities.
[0112] Figure 7 This is a flowchart of another echo cancellation method provided by the embodiment of the present disclosure. The technical solution of this embodiment further optimizes the process of performing nonlinear echo feature extraction and echo cancellation tasks on the processed speech signal using the pre-trained echo cancellation model in the above embodiment. This embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 7 As shown, the echo cancellation method of this embodiment may include the following steps S710-S730:
[0113] S710: Determine a reference voice signal for the voice signal to be processed; the voice signal to be processed and the reference voice signal are derived from the same acoustic device, and the reference voice signal includes echo data.
[0114] The reference speech signal of the speech signal to be processed can be obtained by collecting speech signals from a microphone and a loudspeaker in the acoustic device.
[0115] Optionally, the speech signal to be processed and the reference speech signal are represented in the following forms: time domain, short-time Fourier transform domain, Mel feature domain and Bark feature domain.
[0116] In this embodiment, the speech signal to be processed and the reference speech signal can be in various forms (such as time domain, short-time Fourier transform domain, Mel feature domain, and Bark feature domain) including the speech signal collected by the microphone and the speech signal before being played by the speaker.
[0117] By processing the speech signal to be processed and the reference speech signal obtained by different acoustic devices, a speech signal with better echo cancellation effect can be obtained.
[0118] S720: Input the reference speech signal into a pre-trained echo cancellation model to perform a nonlinear echo feature extraction task.
[0119] In this solution, the pre-trained echo cancellation model can be obtained by using the training method of the echo cancellation model to be trained described in the above embodiment.
[0120] S730: synchronously input the speech signal to be processed into a pre-trained echo cancellation model, and based on the nonlinear echo feature expression obtained by performing the nonlinear echo feature extraction task, use the pre-trained echo cancellation model to perform the echo cancellation task.
[0121] The technical solution of the disclosed embodiments first determines a reference speech signal for the speech signal to be processed. Based on the reference speech signal, the trained echo cancellation model is then controlled to perform nonlinear echo feature extraction. Furthermore, the pre-trained echo cancellation model is controlled to perform echo cancellation based on the speech signal to be processed. By implementing this technical solution, the speech signal to be processed is processed using a low-complexity echo cancellation model with strong device generalization capabilities, resulting in a speech signal with excellent echo cancellation.
[0122] Figure 8 This is a structural block diagram of an echo cancellation device provided by an embodiment of the present disclosure. The technical solution of this embodiment can be applied to the case of applying a pre-trained echo cancellation model. The device can be implemented by software and / or hardware and is generally integrated into any electronic device with network communication function, including but not limited to computers, personal digital assistants and other devices. Figure 8 As shown, the echo cancellation device of this embodiment may include: a speech signal determination module 810, an echo cancellation execution module 820, and an echo cancellation output module 830. Among them:
[0123] A speech signal determination module 810 is used to determine a speech signal to be processed by a pre-trained echo cancellation model;
[0124] The echo cancellation execution module 820 is configured to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed using the pre-trained echo cancellation model; wherein the nonlinear echo feature extraction task is used to assist in nonlinear echo suppression during the echo cancellation process;
[0125] The echo cancellation output module 830 is configured to perform the nonlinear echo feature extraction task and the echo cancellation task through the pre-trained echo cancellation model, and output an echo-cancelled speech signal corresponding to the speech signal to be processed.
[0126] Based on the above embodiment, optionally, the device further includes:
[0127] A training sample feature determination module, used to determine the training sample features used by the echo cancellation model to be trained;
[0128] A training task execution module, configured to control the to-be-trained echo cancellation model to execute a nonlinear echo feature extraction training task and an echo cancellation training task based on the training sample features; wherein the nonlinear echo feature extraction training task is used to assist in performing nonlinear echo suppression during the execution of the echo cancellation training task;
[0129] The echo cancellation model obtaining module is used to adjust the echo cancellation model to be trained according to the nonlinear echo feature extraction training task and the echo cancellation training task to obtain a pre-trained echo cancellation model after training and updating.
[0130] Based on the above embodiment, optionally, the training sample feature determination module is specifically configured to:
[0131] Determining a preset amount of voice data obtained by different acoustic devices respectively; the voice data includes a voice signal collected by a microphone in the acoustic device and a voice signal before being played by a speaker;
[0132] The speech data obtained by different acoustic devices are processed with preset features to obtain speech feature signals corresponding to different acoustic devices, which constitute the training sample features used by the echo cancellation model to be trained.
[0133] Based on the above embodiment, optionally, the speech feature signal is represented in the following forms: time domain, short-time Fourier transform domain, Mel feature domain and Bark feature domain.
[0134] Based on the above embodiment, optionally, the preset feature processing includes at least one of the following: delay estimation processing, linear echo cancellation processing, signal domain transformation processing and signal splicing processing; the signal splicing processing is used to splice and combine speech feature signals in different signal domains.
[0135] Based on the above embodiment, optionally, the training task execution module is specifically configured to:
[0136] Inputting the first sample feature in the training sample features into the echo cancellation model to be trained to perform a nonlinear echo feature extraction training task;
[0137] Synchronously inputting the second sample feature in the training sample features into the echo cancellation model to be trained, and controlling the echo cancellation model to be trained to perform the echo cancellation training task based on the nonlinear echo feature expression obtained by performing the nonlinear echo feature extraction training task;
[0138] The first sample feature and the second sample feature are obtained by performing feature processing on speech signals of different contents collected by the same acoustic device.
[0139] Based on the above embodiment, optionally, the first sample feature and the second sample feature are obtained by processing the same or different preset features respectively; the first sample feature includes at least echo data, which is used to extract the nonlinear echo features of the acoustic device from the echo data through a nonlinear echo feature extraction training task.
[0140] Based on the above embodiment, optionally, the echo cancellation model obtaining module includes:
[0141] a loss function determining unit, configured to determine a first loss function value corresponding to the nonlinear echo feature extraction training task and a second loss function corresponding to the echo cancellation training task;
[0142] a network parameter adjustment unit, configured to adjust the network parameters of the echo cancellation model to be trained according to the first loss function value and the second loss function;
[0143] The first loss function value can prompt the echo cancellation model to be trained to learn the nonlinear characteristics of the acoustic device when performing the nonlinear echo feature extraction training task.
[0144] Based on the above embodiment, optionally, the loss function determining unit is specifically configured to:
[0145] After performing nonlinear echo feature extraction on the first sample feature through the nonlinear feature extraction network in the echo cancellation model to be trained, performing nonlinear feature conversion on the extracted nonlinear echo feature;
[0146] Determining a first loss function value corresponding to the nonlinear echo feature extraction training task based on the nonlinear degree expression obtained by the nonlinear feature conversion and the pre-labeled nonlinear degree expression corresponding to the first sample feature;
[0147] The nonlinear degree expression is used to describe the nonlinear degree of the acoustic device.
[0148] Based on the above embodiment, optionally, the loss function determining unit is further configured to:
[0149] After performing echo cancellation on the second sample feature through the echo cancellation network in the to-be-trained echo cancellation model, performing feature post-processing on the echo cancellation result to obtain an echo cancellation result after feature post-processing;
[0150] A second loss function value corresponding to the echo cancellation training task is determined based on the echo cancellation result after feature post-processing and the pre-labeled echo cancellation result corresponding to the second sample feature.
[0151] Based on the above embodiment, optionally, the loss function determining unit is further configured to:
[0152] A second loss function value corresponding to the echo cancellation training task is determined based on an echo cancellation result after the echo cancellation network in the to-be-trained echo cancellation model performs echo cancellation on the second sample feature and a pre-labeled echo cancellation result corresponding to the second sample feature.
[0153] Based on the above embodiment, optionally, the echo cancellation execution module 820 is specifically configured to:
[0154] Determining a reference voice signal for the voice signal to be processed; the voice signal to be processed and the reference voice signal are derived from the same acoustic device, and the reference voice signal includes echo data;
[0155] Inputting the reference speech signal into a pre-trained echo cancellation model to perform a nonlinear echo feature extraction task;
[0156] The speech signal to be processed is synchronously input into a pre-trained echo cancellation model, and based on the nonlinear echo feature expression obtained by performing the nonlinear echo feature extraction task, the pre-trained echo cancellation model is used to perform the echo cancellation task.
[0157] Based on the above embodiment, optionally, the speech signal to be processed and the reference speech signal are represented in the following forms: time domain, short-time Fourier transform domain, Mel feature domain and Bark feature domain.
[0158] The echo cancellation device provided in the embodiment of the present invention can execute the echo cancellation method provided in any embodiment of the present invention, and has the corresponding functions and beneficial effects of executing the echo cancellation method. For detailed processes, please refer to the relevant operations of the echo cancellation method in the above embodiments.
[0159] Reference below Figure 9 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0160] like Figure 9As shown, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 906 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0161] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 906 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 9 The electronic device 900 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0162] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the echo cancellation method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 909, or installed from the storage device 906, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the echo cancellation method of the embodiment of the present disclosure are performed.
[0163] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0164] The electronic device provided in the embodiment of the present disclosure and the echo cancellation method provided in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment. This embodiment and the echo cancellation method in the above embodiment have the same beneficial effects.
[0165] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the echo cancellation method provided by the above embodiment is implemented.
[0166] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0167] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0168] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0169] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: determines a speech signal to be processed of a pre-trained echo cancellation model; uses the pre-trained echo cancellation model to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed; wherein the nonlinear echo feature extraction task is used to assist in nonlinear echo suppression during the echo cancellation process; and performs the nonlinear echo feature extraction task and the echo cancellation task through the pre-trained echo cancellation model to output an echo-cancelled speech signal corresponding to the speech signal to be processed.
[0170] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0171] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0172] The units described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, a training sample feature determination module may also be described as "determining training sample features used by an echo cancellation model."
[0173] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0174] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0175] According to one or more embodiments of the present disclosure, Example 1 provides an echo cancellation method, the method including:
[0176] Determining a speech signal to be processed by a pre-trained echo cancellation model;
[0177] Using the pre-trained echo cancellation model to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed; wherein the nonlinear echo feature extraction task is used to assist in nonlinear echo suppression during the echo cancellation process;
[0178] The nonlinear echo feature extraction task and the echo cancellation task are performed by the pre-trained echo cancellation model, and an echo-cancelled speech signal corresponding to the speech signal to be processed is output.
[0179] According to one or more embodiments of the present disclosure, Example 2, according to the method of Example 1, the process of determining the pre-trained echo cancellation model includes: determining training sample features used by the echo cancellation model to be trained;
[0180] Controlling the to-be-trained echo cancellation model to perform a nonlinear echo feature extraction training task and an echo cancellation training task based on the training sample features; wherein the nonlinear echo feature extraction training task is used to assist in nonlinear echo suppression during the execution of the echo cancellation training task;
[0181] The echo cancellation model to be trained is adjusted according to the nonlinear echo feature extraction training task and the echo cancellation training task to obtain a pre-trained echo cancellation model after training and updating.
[0182] According to one or more embodiments of the present disclosure, Example 3, based on the method of Example 2, determines the training sample features used by the echo cancellation model to be trained, including:
[0183] Determining a preset amount of voice data obtained by different acoustic devices respectively; the voice data includes a voice signal collected by a microphone in the acoustic device and a voice signal before being played by a speaker;
[0184] The speech data obtained by different acoustic devices are processed with preset features to obtain speech feature signals corresponding to different acoustic devices, which constitute the training sample features used by the echo cancellation model to be trained.
[0185] According to one or more embodiments of the present disclosure, Example 4 is the method according to Example 3, wherein the speech feature signal is represented in the following forms: time domain, short-time Fourier transform domain, Mel feature domain, and Bark feature domain.
[0186] According to one or more embodiments of the present disclosure, Example 5 is based on the method described in Example 3 or 4, and the preset feature processing includes at least one of the following: delay estimation processing, linear echo cancellation processing, signal domain transformation processing, and signal splicing processing; the signal splicing processing is used to splice and combine speech feature signals in different signal domains.
[0187] According to one or more embodiments of the present disclosure, Example 6, according to the method of Example 2, controls the to-be-trained echo cancellation model to perform nonlinear echo feature extraction training tasks and echo cancellation training tasks based on training sample features, including:
[0188] Inputting the first sample feature in the training sample features into the echo cancellation model to be trained to perform a nonlinear echo feature extraction training task;
[0189] Synchronously inputting the second sample feature in the training sample features into the echo cancellation model to be trained, and controlling the echo cancellation model to be trained to perform the echo cancellation training task based on the nonlinear echo feature expression obtained by performing the nonlinear echo feature extraction training task;
[0190] The first sample feature and the second sample feature are obtained by performing feature processing on speech signals of different contents collected by the same acoustic device.
[0191] According to one or more embodiments of the present disclosure, Example 7 is based on the method described in Example 6, wherein the first sample feature and the second sample feature are obtained by processing the same or different preset features respectively; the first sample feature includes at least echo data, which is used to extract the nonlinear echo features of the acoustic device from the echo data through a nonlinear echo feature extraction training task.
[0192] According to one or more embodiments of the present disclosure, Example 8, according to the method of Example 6, adjusts the to-be-trained echo cancellation model based on the nonlinear echo feature extraction training task and the echo cancellation training task, including:
[0193] Determining a first loss function value corresponding to the nonlinear echo feature extraction training task and a second loss function corresponding to the echo cancellation training task;
[0194] Adjusting network parameters of the echo cancellation model to be trained according to the first loss function value and the second loss function;
[0195] The first loss function value can prompt the echo cancellation model to be trained to learn the nonlinear characteristics of the acoustic device when performing the nonlinear echo feature extraction training task.
[0196] According to one or more embodiments of the present disclosure, Example 9, based on the method of Example 8, determines a first loss function value corresponding to the nonlinear echo feature extraction training task, including:
[0197] After performing nonlinear echo feature extraction on the first sample feature through the nonlinear feature extraction network in the echo cancellation model to be trained, performing nonlinear feature conversion on the extracted nonlinear echo feature;
[0198] Determining a first loss function value corresponding to the nonlinear echo feature extraction training task based on the nonlinear degree expression obtained by the nonlinear feature conversion and the pre-labeled nonlinear degree expression corresponding to the first sample feature;
[0199] The nonlinear degree expression is used to describe the nonlinear degree of the acoustic device.
[0200] According to one or more embodiments of the present disclosure, Example 10, according to the method of Example 8, determines a second loss function corresponding to the echo cancellation training task, including:
[0201] After performing echo cancellation on the second sample feature through the echo cancellation network in the to-be-trained echo cancellation model, performing feature post-processing on the echo cancellation result to obtain an echo cancellation result after feature post-processing;
[0202] A second loss function value corresponding to the echo cancellation training task is determined based on the echo cancellation result after feature post-processing and the pre-labeled echo cancellation result corresponding to the second sample feature.
[0203] According to one or more embodiments of the present disclosure, Example 11 determines a second loss function corresponding to the echo cancellation training task according to the method of Example 8, including:
[0204] A second loss function value corresponding to the echo cancellation training task is determined based on an echo cancellation result after the echo cancellation network in the to-be-trained echo cancellation model performs echo cancellation on the second sample feature and a pre-labeled echo cancellation result corresponding to the second sample feature.
[0205] According to one or more embodiments of the present disclosure, Example 12, according to the method of Example 1, uses the pre-trained echo cancellation model to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed, including:
[0206] Determining a reference voice signal for the voice signal to be processed; the voice signal to be processed and the reference voice signal are derived from the same acoustic device, and the reference voice signal includes echo data;
[0207] Inputting the reference speech signal into a pre-trained echo cancellation model to perform a nonlinear echo feature extraction task;
[0208] The speech signal to be processed is synchronously input into a pre-trained echo cancellation model, and based on the nonlinear echo feature expression obtained by performing the nonlinear echo feature extraction task, the pre-trained echo cancellation model is used to perform the echo cancellation task.
[0209] According to one or more embodiments of the present disclosure, Example 13 is the method according to Example 12, wherein the speech signal to be processed and the reference speech signal are represented in the following forms: time domain, short-time Fourier transform domain, Mel feature domain and Bark feature domain.
[0210] According to one or more embodiments of the present disclosure, Example 14 provides an echo cancellation device, the device comprising:
[0211] A speech signal determination module, used to determine the speech signal to be processed by the pre-trained echo cancellation model;
[0212] an echo cancellation execution module, configured to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed using the pre-trained echo cancellation model; wherein the nonlinear echo feature extraction task is used to assist in nonlinear echo suppression during the echo cancellation process;
[0213] The echo cancellation output module is used to perform the nonlinear echo feature extraction task and the echo cancellation task through the pre-trained echo cancellation model, and output the echo-cancelled speech signal corresponding to the speech signal to be processed.
[0214] According to one or more embodiments of the present disclosure, Example 15 provides an electronic device, the electronic device including:
[0215] at least one processor; and
[0216] a memory communicatively connected to the at least one processor; wherein,
[0217] The memory stores a computer program executable by the at least one processor, where the computer program is executed by the at least one processor to enable the at least one processor to perform the echo cancellation method according to any one of Examples 1-13.
[0218] According to one or more embodiments of the present disclosure, Example 16 provides a computer-readable medium storing computer instructions, which are used to enable a processor to implement the echo cancellation method described in any one of Examples 1-13 when executed.
[0219] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0220] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0221] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An echo cancellation method, characterized in that: The method comprises: Determining a speech signal to be processed by a pre-trained echo cancellation model; Using the pre-trained echo cancellation model to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed; wherein the nonlinear echo feature extraction task is used to assist in nonlinear echo suppression during the echo cancellation process; Performing the nonlinear echo feature extraction task and the echo cancellation task through the pre-trained echo cancellation model, and outputting an echo-cancelled speech signal corresponding to the speech signal to be processed; The pre-trained echo cancellation model is used to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed, including: Determining a reference voice signal for the voice signal to be processed; the voice signal to be processed and the reference voice signal are derived from the same acoustic device, and the reference voice signal includes echo data; Inputting the reference speech signal into a pre-trained echo cancellation model to perform a nonlinear echo feature extraction task; The speech signal to be processed is synchronously input into a pre-trained echo cancellation model, and based on the nonlinear echo feature expression obtained by performing the nonlinear echo feature extraction task, the pre-trained echo cancellation model is used to perform the echo cancellation task.
2. The method according to claim 1, characterized in that The process of determining the pre-trained echo cancellation model includes: Determining the characteristics of the training samples used by the echo cancellation model to be trained; Controlling the to-be-trained echo cancellation model to perform a nonlinear echo feature extraction training task and an echo cancellation training task based on the training sample features; wherein the nonlinear echo feature extraction training task is used to assist in nonlinear echo suppression during the execution of the echo cancellation training task; The echo cancellation model to be trained is adjusted according to the nonlinear echo feature extraction training task and the echo cancellation training task to obtain a pre-trained echo cancellation model after training and updating.
3. The method according to claim 2, characterized in that Determine the training sample features used by the echo cancellation model to be trained, including: Determining a preset amount of voice data obtained by different acoustic devices respectively; the voice data includes a voice signal collected by a microphone in the acoustic device and a voice signal before being played by a speaker; The speech data obtained by different acoustic devices are processed with preset features to obtain speech feature signals corresponding to different acoustic devices, which constitute the training sample features used by the echo cancellation model to be trained.
4. The method according to claim 3, characterized in that The speech feature signal is represented in the following forms: time domain, short-time Fourier transform domain, Mel feature domain and Bark feature domain.
5. The method according to claim 3 or 4, characterized in that The preset feature processing includes at least one of the following: delay estimation processing, linear echo cancellation processing, signal domain transformation processing and signal splicing processing; the signal splicing processing is used to splice and combine speech feature signals in different signal domains.
6. The method according to claim 2, characterized in that Based on the training sample features, the echo cancellation model to be trained is controlled to perform nonlinear echo feature extraction training tasks and echo cancellation training tasks, including: Inputting the first sample feature in the training sample features into the echo cancellation model to be trained to perform a nonlinear echo feature extraction training task; Synchronously inputting the second sample feature in the training sample features into the echo cancellation model to be trained, and controlling the echo cancellation model to be trained to perform the echo cancellation training task based on the nonlinear echo feature expression obtained by performing the nonlinear echo feature extraction training task; The first sample feature and the second sample feature are obtained by performing feature processing on speech signals of different contents collected by the same acoustic device.
7. The method according to claim 6, characterized in that The first sample feature and the second sample feature are obtained by processing the same or different preset features respectively; the first sample feature includes at least echo data, which is used to extract the nonlinear echo feature of the acoustic device from the echo data through a nonlinear echo feature extraction training task.
8. The method according to claim 6, characterized in that The adjusting the echo cancellation model to be trained according to the nonlinear echo feature extraction training task and the echo cancellation training task includes: Determining a first loss function value corresponding to the nonlinear echo feature extraction training task and a second loss function corresponding to the echo cancellation training task; Adjusting network parameters of the echo cancellation model to be trained according to the first loss function value and the second loss function; The first loss function value can prompt the echo cancellation model to be trained to learn the nonlinear characteristics of the acoustic device when performing the nonlinear echo feature extraction training task.
9. The method according to claim 8, characterized in that Determining a first loss function value corresponding to the nonlinear echo feature extraction training task includes: After performing nonlinear echo feature extraction on the first sample feature through the nonlinear feature extraction network in the echo cancellation model to be trained, performing nonlinear feature conversion on the extracted nonlinear echo feature; Determining a first loss function value corresponding to the nonlinear echo feature extraction training task based on the nonlinear degree expression obtained by the nonlinear feature conversion and the pre-labeled nonlinear degree expression corresponding to the first sample feature; The nonlinear degree expression is used to describe the nonlinear degree of the acoustic device.
10. The method according to claim 8, characterized in that Determining a second loss function corresponding to the echo cancellation training task includes: After performing echo cancellation on the second sample feature through the echo cancellation network in the to-be-trained echo cancellation model, performing feature post-processing on the echo cancellation result to obtain an echo cancellation result after feature post-processing; A second loss function value corresponding to the echo cancellation training task is determined based on the echo cancellation result after feature post-processing and the pre-labeled echo cancellation result corresponding to the second sample feature.
11. The method according to claim 8, characterized in that Determining a second loss function corresponding to the echo cancellation training task includes: A second loss function value corresponding to the echo cancellation training task is determined based on an echo cancellation result after the echo cancellation network in the to-be-trained echo cancellation model performs echo cancellation on the second sample feature and a pre-labeled echo cancellation result corresponding to the second sample feature.
12. The method according to claim 1, characterized in that The speech signal to be processed and the reference speech signal are represented in the following forms: time domain, short-time Fourier transform domain, Mel feature domain and Bark feature domain.
13. An echo cancellation device, characterized in that: The device comprises: A speech signal determination module, used to determine the speech signal to be processed by the pre-trained echo cancellation model; an echo cancellation execution module, configured to perform a nonlinear echo feature extraction task and an echo cancellation task on the speech signal to be processed using the pre-trained echo cancellation model; wherein the nonlinear echo feature extraction task is used to assist in nonlinear echo suppression during the echo cancellation process; An echo cancellation output module, configured to perform the nonlinear echo feature extraction task and the echo cancellation task through the pre-trained echo cancellation model, and output an echo-cancelled speech signal corresponding to the speech signal to be processed; The echo cancellation execution module is specifically used to: Determining a reference voice signal for the voice signal to be processed; the voice signal to be processed and the reference voice signal are derived from the same acoustic device, and the reference voice signal includes echo data; Inputting the reference speech signal into a pre-trained echo cancellation model to perform a nonlinear echo feature extraction task; The speech signal to be processed is synchronously input into a pre-trained echo cancellation model, and based on the nonlinear echo feature expression obtained by performing the nonlinear echo feature extraction task, the pre-trained echo cancellation model is used to perform the echo cancellation task.
14. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the echo cancellation method according to any one of claims 1 to 12.
15. A computer-readable medium, characterized in that The computer-readable medium stores computer instructions, and the computer instructions are used to enable a processor to implement the echo cancellation method according to any one of claims 1 to 12 when executed.
Citation Information
Patent Citations
Audio signal processing method and device, training method and device, equipment and storage medium
CN114242100A