Audio signal processing method, echo cancellation model training method, and speech recognition method
By acquiring the audio signal to be processed and the reference audio signal, and combining the linear echo cancellation algorithm and the echo cancellation model, the problem of poor echo cancellation effect under complex audio conditions is solved, and high-quality communication effect is achieved.
Patent Information
- Application Number
- PCT/CN2025/101285
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-22
- Filing Date
- 2025-06-16
- Publication Date
- 2026-01-29
AI Technical Summary
Under complex audio conditions, traditional echo cancellation technology is difficult to effectively eliminate echo interference, affecting communication quality and clarity, especially in full-duplex communication scenarios.
By acquiring the audio signal to be processed and the reference audio signal, the linear echo cancellation algorithm is used to initially eliminate the linear echo component, and then the echo cancellation model is used for deeper elimination, including signal alignment and Fourier transform processing, to ensure the purity of the audio signal.
It significantly reduces echo components, improves communication quality, ensures that both parties in a full-duplex interactive scenario can clearly and accurately understand each other's content, and enhances the call experience.
Smart Images

Figure CN2025101285_29012026_PF_FP_ABST
Abstract
Description
Method for processing audio signal, method for training echo cancellation model, and method for speech recognition TECHNICAL FIELD
[0001] The present disclosure relates to the field of artificial intelligence, and in particular, to a method for processing audio signal, a method for training echo cancellation model, and a method for speech recognition. BACKGROUND
[0002] Currently, mobile communication has become an indispensable part of people's daily life. However, as the use scenarios of mobile devices become diversified, for example, when using mobile devices to communicate in noisy outdoor environments, inside vehicles, or multimedia conference scenarios, the communication quality is likely to be affected by background noise signals, echoes, and other interferences. Especially when communicating in a full-duplex communication scenario, the interference of echoes poses a serious challenge to the clarity and understanding of the conversation. The traditional echo cancellation technology has great limitations in processing real-time communication and complex audio conditions, and it may also be difficult to adapt to mobile device hardware configurations and various noise signal environments, resulting in poor echo cancellation effect when using the echo cancellation technology to cancel echoes in the audio to be processed under complex audio conditions.
[0003] At present, there is no effective solution to the above problems. SUMMARY
[0004] The embodiments of the present disclosure provide a method for processing audio signal, a method for training echo cancellation model, and a method for speech recognition to at least solve the technical problem of echo cancellation effect for audio to be processed under complex audio conditions in the prior art.
[0005] According to an aspect of an embodiment of the present disclosure, a method for processing audio signal is provided, comprising: obtaining an audio signal to be processed and a reference audio signal, wherein the reference audio signal corresponds to an echo component in the audio signal to be processed; performing echo cancellation on a linear echo component in the audio signal to be processed based on the reference audio signal to obtain a processed audio signal; and performing echo cancellation on the processed audio signal based on a preset audio signal using an echo cancellation model to obtain a target audio signal, wherein the preset audio signal comprises at least one of the reference audio signal and the audio signal to be processed.
[0006] According to another aspect of an embodiment of the present disclosure, a method for training an echo cancellation model is also provided, comprising: obtaining original training data, wherein the original training data comprises a first training audio signal and a training speech signal corresponding to the first training audio signal; performing data augmentation on the first training audio signal to obtain a second training audio signal; and performing progressive learning on an initial cancellation model using the second training audio signal to obtain an echo cancellation model, wherein the echo cancellation model is used to execute any of the above methods.
[0007] According to a further aspect of the embodiments of the present disclosure, a voice recognition method is also provided, including: obtaining a to-be-processed audio signal and a reference audio signal, wherein the reference audio signal corresponds to an echo component in the to-be-processed audio signal; performing echo cancellation on a linear echo component in the to-be-processed audio signal based on the reference audio signal to obtain a processed audio signal; performing echo cancellation on the processed audio signal based on a preset audio signal by using an echo cancellation model to obtain a target audio signal, wherein the preset audio signal contains at least one of the following: the reference audio signal and the to-be-processed audio signal; and performing voice recognition on the target audio signal to obtain a voice recognition result of the to-be-processed audio signal.
[0008] According to a further aspect of the embodiments of the present disclosure, a processing apparatus of an audio signal is also provided, including: a first obtaining component configured to obtain a to-be-processed audio signal and a reference audio signal, wherein the reference audio signal corresponds to an echo component in the to-be-processed audio signal; a first cancelling component configured to perform echo cancellation on a linear echo component in the to-be-processed audio signal based on the reference audio signal to obtain a processed audio signal; and a second cancelling component configured to perform echo cancellation on the processed audio signal based on a preset audio signal by using an echo cancellation model to obtain a target audio signal, wherein the preset audio signal contains at least one of the following: the reference audio signal and the to-be-processed audio signal.
[0009] According to a further aspect of the embodiments of the present disclosure, a training apparatus of an echo cancellation model is also provided, including: a data obtaining component configured to obtain original training data, wherein the original training data contains a first training audio signal and a training voice signal corresponding to the first training audio signal; a data augmenting component configured to perform data augmentation on the first training audio signal to obtain a second training audio signal; and a model learning component configured to perform progressive learning on an initial cancellation model by using the second training audio signal to obtain the echo cancellation model, wherein the echo cancellation model is configured to perform the method of any one of the above.
[0010] According to a further aspect of the embodiments of the present disclosure, a voice recognition apparatus is also provided, comprising: a second obtaining component configured to obtain a to-be-processed audio signal and a reference audio signal, wherein the reference audio signal corresponds to an echo component in the to-be-processed audio signal; a third canceling component configured to perform echo cancellation on a linear echo component in the to-be-processed audio signal based on the reference audio signal, to obtain a processed audio signal; a fourth canceling component configured to perform echo cancellation on the processed audio signal based on a preset audio signal by using an echo cancellation model, to obtain a target audio signal, wherein the preset audio signal comprises at least one of the following: the reference audio signal and the to-be-processed audio signal; a voice recognition component configured to perform voice recognition on the target audio signal, to obtain a voice recognition result of the to-be-processed audio signal.
[0011] According to a further aspect of the embodiments of the present disclosure, a computer terminal is also provided, comprising: a memory storing an executable program; and a processor configured to run the program, wherein the program performs the method in the various embodiments of the present disclosure when running.
[0012] According to a further aspect of the embodiments of the present disclosure, a computer readable storage medium is also provided, comprising a stored executable program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to perform the method in the various embodiments of the present disclosure when the executable program runs.
[0013] According to a further aspect of the embodiments of the present disclosure, a computer program product is also provided, comprising a computer program, which, when executed by a processor, implements the method in the various embodiments of the present disclosure.
[0014] According to a further aspect of the embodiments of the present disclosure, a computer program product is also provided, comprising a non-volatile computer readable storage medium storing a computer program, which, when executed by a processor, implements the method in the various embodiments of the present disclosure.
[0015] According to a further aspect of the embodiments of the present disclosure, a computer program is also provided, which, when executed by a processor, implements the method in the various embodiments of the present disclosure.
[0016] In the embodiments of the present disclosure, the method comprises the following steps: obtaining a to-be-processed audio signal and a reference audio signal; performing echo cancellation on a linear echo component in the to-be-processed audio signal based on the reference audio signal to obtain a processed audio signal; and performing echo cancellation on the processed audio signal based on a preset audio signal by using an echo cancellation model to obtain a target audio signal. In the method, first, the linear echo cancellation algorithm is used to quickly eliminate most simple echoes contained in the to-be-processed audio signal based on the reference audio signal to obtain the processed audio signal, and then the echo cancellation model is used to perform deeper echo cancellation processing on the processed audio signal to obtain the target audio signal. In this way, the echo components contained in the to-be-processed audio signal can be greatly reduced, the purity of the obtained target audio signal is improved, the communication quality of different users in a full-duplex interaction scene is ensured, and the technical problem of poor echo cancellation effect of the to-be-processed audio under complex audio conditions in the related art is solved.
[0017] It should be noted that the general description above and the detailed description below are merely intended to exemplify and explain the present disclosure, and do not constitute a limitation on the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings, which are included to provide a further understanding of the present disclosure, constitute a part of the present disclosure and illustrate the exemplary embodiments of the present disclosure and their descriptions serve to explain the present disclosure, and do not constitute an improper limitation on the present disclosure. In the drawings:
[0019] FIG. 1 is a schematic diagram of an audio signal processing scene according to an embodiment of the present disclosure;
[0020] FIG. 2 is a structural block diagram of a computing environment for audio signal processing according to an embodiment of the present disclosure;
[0021] FIG. 3 is a flowchart of a processing method of an audio signal according to an embodiment of the present disclosure;
[0022] FIG. 4 is a schematic diagram of a processing process of an audio signal according to an embodiment of the present disclosure;
[0023] FIG. 5 is a flowchart of a training method of an echo cancellation model according to an embodiment of the present disclosure;
[0024] FIG. 6 is a schematic diagram of an audio signal masking result according to an embodiment of the present disclosure;
[0025] FIG. 7 is a schematic diagram of an audio signal splicing result according to an embodiment of the present disclosure;
[0026] FIG. 8 is a schematic diagram of an echo cancellation process according to an embodiment of the present disclosure;
[0027] FIG. 9 is a flowchart of a speech recognition method according to an embodiment of the present disclosure;
[0028] FIG. 10 is a schematic diagram of a speech recognition process according to an embodiment of the present disclosure;
[0029] FIG. 11 is a schematic diagram of a processing device of an audio signal according to an embodiment of the present disclosure;
[0030] FIG. 12 is a schematic diagram of a training device of an echo cancellation model according to an embodiment of the present disclosure;
[0031] FIG. 13 is a schematic diagram of a speech recognition device according to an embodiment of the present disclosure;
[0032] FIG. 14 is a structural block diagram of a computer terminal according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] In order to make the person skilled in the art better understand the present disclosure scheme, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, not all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by a person skilled in the art without creative labor should be within the scope of protection of the present disclosure.
[0034] It should be noted that the terms "first", "second" and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or components does not have to be limited to those steps or components clearly listed, but can include other steps or components not clearly listed or inherent to these processes, methods, products or devices.
[0035] First, some nouns or terms that appear in the process of describing the embodiments of the present disclosure are applicable to the following explanations:
[0036] AEC: Acoustic Echo Cancellation, acoustic echo cancellation.
[0037] ASR: Automatic Speech Recognition, automatic speech recognition.
[0038] DA: Data Augmentation, data augmentation.
[0039] DFSMN: Deep Feedforward Sequential Memory Network.
[0040] ERLE: Echo Return Loss Enhancement.
[0041] FAR: False Alarm Rate.
[0042] LAEC: Linear Acoustic Echo Cancellation.
[0043] MAR: Miss Alarm Rate.
[0044] NES: Neural Echo Suppressor.
[0045] NN: Neural Network.
[0046] PESQ: Perceptual Evaluation of Speech Quality.
[0047] PL: Progressive Learning.
[0048] SER: Signal-to-Echo Ratio.
[0049] SNR: Signal-to-Noise Ratio.
[0050] STFT: Short-Time Fourier Transform.
[0051] TDE: Time Delay Estimation.
[0052] VAD: Voice Activity Detection.
[0053] WER: Word Error Rate.
[0054] According to embodiments of this disclosure, an audio signal processing method is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0055] Figure 1 is a schematic diagram of an audio signal processing scenario according to an embodiment of the present disclosure. Considering that machine learning models consume a large amount of computing resources on mobile terminals, the method provided by the embodiments of the present disclosure can be applied to the application scenario shown in Figure 1, but is not limited thereto. In the application scenario shown in Figure 1, the machine learning model is deployed on server 11. Server 11 can connect to one or more clients 20 through a local area network (LAN), a wide area network (WAN), an internet connection, or other types of data networks. Clients 20 may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Clients 20 can interact with users through a graphical user interface to invoke the large model, thereby implementing the method provided by the embodiments of the present disclosure.
[0056] It should be noted that, provided that the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiments disclosed herein can be performed on the client device.
[0057] Figure 2 is a structural block diagram of a computing environment for an audio signal processing method according to an embodiment of the present disclosure. As shown in Figure 2, the computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (shown as 211-1, 211-2, ..., in the figure). Each computing node contains local processing and memory resources, and the end user 202 can remotely run applications or store data within the computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 within the computing environment 201, representing services "A", "D", "E", and "H", respectively.
[0058] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).
[0059] The services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.
[0060] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, as shown in Figure 2, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers in a Pod handle requests related to one or more corresponding functions of the service. The proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be equipped with similar Pods.
[0061] During operation, executing a user request from end user 202 may require calling one or more services in computing environment 201. Executing one or more functions of one service requires calling one or more functions of another service. As shown in Figure 2, service "A" 220-1 receives a user request from end user 202 from ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to execute one or more functions.
[0062] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.
[0063] In the above operating environment, this disclosure provides an audio signal processing method as shown in Figure 3. It should be noted that the audio signal processing method of this embodiment can be executed by the server in the embodiment shown in Figure 1. Figure 3 is a flowchart of an audio signal processing method according to an embodiment of this disclosure. As shown in Figure 3, the method may include the following steps:
[0064] Step S302: Obtain the audio signal to be processed and the reference audio signal.
[0065] The reference audio signal corresponds to the echo component in the audio signal to be processed.
[0066] The aforementioned audio signal to be processed can refer to the audio signal that currently requires acoustic echo cancellation (AEC) processing, and can be an audio signal received by an audio acquisition device. The aforementioned reference audio signal can refer to the audio signal output by an audio playback device. The audio acquisition device and the audio playback device can be devices deployed on the same device. For example, when user A makes a voice call with user B through mobile phone A, the microphone device on mobile phone A can be considered as the audio acquisition device, and the audio signal received by the microphone in the scenario where user A is located can be considered as the aforementioned audio signal to be processed. The speaker on mobile phone A can be considered as the audio playback device, and the audio signal output by the speaker in the scenario where user B is located can be considered as the aforementioned reference audio signal.
[0067] In one optional embodiment, considering that the audio acquisition device receives the sound of the environment in which the device is located, in some scenarios, such as when the noise level is high in the environment where mobile phone A is located, the audio signal received by the audio acquisition device may contain a large amount of environmental noise signals. This may cause the receiver of the audio signal to be unable to clearly identify the target signal contained in the audio signal when viewing the audio signal. For example, user B may not be able to clearly hear the content sent by user A in the voice. Therefore, in order to improve the clarity of the transmitted audio signal, the audio signal processing system (hereinafter referred to as the processing system) can acquire the audio signal received by the audio acquisition device in real time, that is, the audio signal to be processed mentioned above, and perform noise reduction processing on the noise signal contained in the audio signal to be processed by echo cancellation, so as to reduce the noise signal contained in the audio signal to be processed, thereby ensuring that the receiver of the audio signal can clearly obtain the main content contained in the audio signal to be processed.
[0068] Furthermore, if the environment where the audio acquisition device is located is open and has a large echo, the received audio signal may contain not only ambient noise but also echo components corresponding to the audio signal output by the audio playback device. For example, when user A is having a real-time call with user B via mobile phone A, if user A plays user B's words aloud at a high volume, the audio acquisition device in mobile phone A may receive the audio signal output by the audio playback device, i.e., user B's words. Consequently, when mobile phone A sends audio signal A to user B, it may retransmit user B's words, causing user B to hear them repeatedly. The echo of user B's own voice can significantly impact their call experience. Simply performing echo cancellation on audio signal A might not be sufficient to remove user B's voice from it. Therefore, to improve call quality in full-duplex interactive scenarios where two-way communication is possible, the processing system acquires not only the audio signal to be processed received by the audio acquisition device but also the audio signal output by the audio playback device associated with the audio acquisition device—the aforementioned reference audio signal. This allows the processing system to perform noise reduction on the audio signal to be processed based on the reference audio signal, thereby improving the quality of the audio signal and enhancing the experience for both parties in the call.
[0069] Step S304: Based on the reference audio signal, perform echo cancellation on the linear echo component in the audio signal to be processed to obtain the processed audio signal.
[0070] In one optional embodiment, to improve the communication quality of different users in a full-duplex interaction scenario, the processing system can first remove any reference audio signals that may be contained in the audio signal to be processed, in order to avoid users repeatedly hearing their own voices, which would affect the user's communication experience. Based on this, the processing system can use AEC algorithms, such as linear echo cancellation algorithms or nonlinear echo cancellation algorithms, to cancel the linear echo components in the audio signal to be processed based on the obtained reference audio signal. For example, the processing system can first analyze the audio signal to be processed, extract the signal features of the signal, such as frequency and amplitude, and determine the linear echo components in the audio signal to be processed by analyzing the signal features. Then, it can remove the linear echo components from the audio signal to be processed as much as possible to obtain the processed audio signal, thereby reducing the amount of data of the reference audio signal in the audio signal to be processed and reducing the impact of the reference audio signal on both parties in the full-duplex interaction scenario.
[0071] Step S306: Use an echo cancellation model to cancel the echo of the processed audio signal based on a preset audio signal to obtain the target audio signal.
[0072] The preset audio signal includes at least one of the following: a reference audio signal and an audio signal to be processed.
[0073] The echo cancellation model mentioned above can refer to a model used to remove noise and echo components from audio signals. It can be a neural echo suppressor, i.e., a NES model. The target audio signal mentioned above can refer to the clean audio signal obtained after echo cancellation. The target audio signal has a high signal-to-noise ratio (SNR) and can clearly and accurately reflect the actual content expressed by both parties in the communication.
[0074] In one optional embodiment, as described above, considering that the echo cancellation algorithm may not be able to completely eliminate the echo components in the audio signal to be processed, that is, there may be some incompletely eliminated reference audio signals in the processed audio signal, and the audio signal to be processed may also contain environmental noise signals from the environment where the audio acquisition device is located, in order to further improve the purity of the audio signals received by both parties in a full-duplex interactive scenario, the processing system can also use the above-mentioned echo cancellation model to perform echo cancellation on the processed audio signal again based on the above-mentioned preset audio signal, that is, the reference audio signal or the audio signal to be processed, in order to further remove the echo components corresponding to the reference audio signal that may remain in the processed audio signal, the environmental noise signals contained in the audio signal to be processed, and other signals that may affect the quality of the audio signal, so as to obtain a pure audio signal, that is, to obtain the above-mentioned target audio signal, thereby enabling both parties to clearly and accurately understand the actual communication content expressed by the other party, and ensuring the communication quality of both parties in a full-duplex interactive scenario.
[0075] In this embodiment, the method involves acquiring an audio signal to be processed and a reference audio signal; performing echo cancellation on the linear echo components in the audio signal to be processed based on the reference audio signal to obtain a processed audio signal; and then using an echo cancellation model to perform echo cancellation on the processed audio signal based on a preset audio signal to obtain a target audio signal. By first using a linear echo cancellation algorithm to quickly eliminate most of the simple echoes in the audio signal to be processed for the reference audio signal to obtain a processed audio signal, and then using an echo cancellation model to perform deeper echo cancellation processing on the processed audio signal to obtain the target audio signal, the method can significantly reduce the echo components in the audio signal to be processed, improve the purity of the obtained target audio signal, ensure the communication quality of different users communicating in full-duplex interactive scenarios, and thus solve the technical problem of echo cancellation effect on audio under complex audio conditions in the relevant technology.
[0076] In this embodiment, the echo cancellation model is used to cancel the echo of the processed audio signal based on a preset audio signal to obtain the target audio signal. This includes: performing a Fourier transform on the processed audio signal to obtain a first transform feature; performing a Fourier transform on the preset audio signal to obtain a second transform feature; using the echo cancellation model to cancel the echo of the first transform feature based on the second transform feature to obtain the target transform feature; and performing an inverse Fourier transform on the target transform feature to obtain the target audio signal.
[0077] In one optional embodiment, to improve the accuracy of echo cancellation on the processed audio signal, the processing system can first perform a Fourier Transform (STFT) on the processed audio signal to obtain Fourier features of the processed audio signal in multiple dimensions, i.e., the first transform feature mentioned above. At the same time, a Fourier transform is also performed on a preset audio signal to obtain Fourier features of the preset audio signal in multiple dimensions, i.e., the second transform feature mentioned above. Then, the first transform feature and the second transform feature are input into the echo cancellation model to perform echo cancellation on the first transform feature based on the second transform feature. For example, the echo cancellation model removes the Fourier features corresponding to the reference audio signal and the Fourier features corresponding to the environmental noise signal from the first transform feature to obtain Fourier features that reflect a pure audio signal, i.e., the target transform feature mentioned above. Then, an inverse Fourier transform is performed on the target transform feature to obtain the target audio signal mentioned above, thereby ensuring the accuracy when performing echo cancellation on the processed audio signal again, and thus ensuring the accuracy of the obtained target audio signal.
[0078] In this embodiment, the echo cancellation model is used to cancel the echo of the first transformation feature based on the second transformation feature to obtain the target transformation feature. This includes: inputting the first transformation feature and the second transformation feature into the echo cancellation model to obtain the signal mask output by the echo cancellation model, wherein the signal mask is used to cancel the echo component in the audio signal to be processed; and obtaining the product of the signal mask and the first transformation feature to obtain the target transformation feature.
[0079] In one optional embodiment, when using the second transform feature to cancel the echo of the first transform feature, the processing system can first input the first transform feature and the second transform feature into the echo cancellation model to determine the signal mask of the second transform feature relative to the first transform feature. Taking the reference audio signal as an example, if the second transform feature is the Fourier feature of the reference audio signal, it means that the part in the first transform feature that is the same as the second transform feature represents the part of the reference audio signal corresponding to the echo component remaining in the processed audio signal. At this time, this part of the signal can be regarded as the signal that needs to be echo cancelled. Correspondingly, the signal mask can be constructed by using 0 to represent the feature position of the Fourier feature corresponding to the same part of the signal and 1 to represent other positions besides the feature position. After obtaining the signal mask, the processing system can multiply the signal mask and the first transform feature, and then use the product of the two as the target transform feature, thereby avoiding the situation where the target transform feature is determined incorrectly, resulting in the inability to accurately cancel the echo in the processed audio signal or damage the target audio signal contained in the processed audio signal, thereby improving the accuracy of the determined target audio signal.
[0080] In this embodiment, the echo cancellation model includes: a first fully connected layer, multiple signal synthesis layers, a second fully connected layer, and an activation function layer. The first transform feature and the second transform feature are input into the echo cancellation model to obtain the signal mask output by the echo cancellation model. This includes: using the first fully connected layer to perform feature mapping on the first transform feature and the second transform feature respectively, obtaining a first mapped feature corresponding to the first transform feature and a second mapped feature corresponding to the second transform feature; using multiple signal synthesis layers to synthesize the first mapped feature and the second mapped feature to obtain a synthesized feature; using the second fully connected layer to perform feature mapping on the synthesized feature to obtain a third mapped feature; and using the activation function layer to process the third mapped feature to obtain the signal mask.
[0081] In one optional embodiment, the echo cancellation model can include at least the first fully connected layer, multiple signal synthesis layers, a second fully connected layer, and an activation function layer. Correspondingly, when constructing the signal mask, the first fully connected layer can be used to map the first transform feature and the second transform feature to obtain the first mapped feature corresponding to the first transform feature and the second mapped feature corresponding to the second transform feature. Then, the multiple signal synthesis layers can be used to synthesize the first and second mapped features to obtain the synthesized features. Then, the second fully connected layer can be used to perform feature mapping on the synthesized features to obtain the third mapped feature. Finally, the activation function layer can be used to activate the third mapped feature to obtain the signal mask.
[0082] In this embodiment, echo cancellation is performed on the linear echo component in the audio signal to be processed based on the reference audio signal to obtain the processed audio signal. This includes: aligning the audio signal to be processed and the reference audio signal to obtain the aligned audio signal to be processed and the aligned reference audio signal; and canceling the linear echo component in the aligned audio signal to be processed to obtain the processed audio signal.
[0083] In one optional embodiment, to improve the accuracy of echo cancellation of the signal to be processed, when using the linear echo cancellation algorithm to cancel the echo of the audio signal to be processed, the processing system can first align the audio signal to be processed and the reference audio signal to obtain the aligned audio signal to be processed and the aligned reference audio signal. Then, the linear echo cancellation algorithm is used to perform echo cancellation on the aligned audio signal to be processed based on the aligned reference audio signal, thereby obtaining an audio signal with higher accuracy.
[0084] In this embodiment, signal alignment is performed between the audio signal to be processed and the reference audio signal to obtain aligned audio signal to be processed and aligned reference audio signal. This includes: estimating the time delay difference between the audio signal to be processed and the reference audio signal to obtain a target time delay difference; using the target time delay difference to delay the output of the reference audio signal to obtain aligned reference audio signal; and determining that the audio signal to be processed is the aligned audio signal to be processed.
[0085] In one optional embodiment, when aligning the audio signal to be processed and the reference audio signal, the two signals can be analyzed first, for example, by performing time delay estimation (TDE) to estimate the time delay difference between them, thus obtaining the target time delay difference. Then, the reference audio signal is delayed using the target time delay difference to obtain the aligned reference audio signal. This can compensate for the time delay difference between the audio signal to be processed and the reference audio signal, thereby avoiding the failure of alignment between the audio signal to be processed and the reference audio signal, which would lead to the failure of echo cancellation of the audio signal to be processed based on the reference audio signal. Correspondingly, the audio signal to be processed can be directly determined as the aligned audio signal to be processed.
[0086] In this embodiment, acquiring the audio signal to be processed and the reference audio signal includes: acquiring the audio signal to be processed through an audio acquisition device on an electronic device; and acquiring the reference audio signal played through an audio playback device on an electronic device, wherein the echo cancellation model is deployed on the electronic device.
[0087] In one optional embodiment, to facilitate echo cancellation of the audio signal to be processed, the audio acquisition device, the audio playback device, and the echo cancellation model for echo cancellation of the audio signal to be processed can all be installed on the corresponding electronic device, such as a mobile phone commonly used by users. The corresponding processing system can directly acquire the audio signal to be processed acquired by the audio acquisition device and the reference audio signal played by the audio playback device. Then, the echo cancellation model is used on the electronic device to quickly perform echo cancellation on the audio signal to be processed based on the reference audio signal, thereby improving the efficiency of echo cancellation and ensuring the timeliness of communication between different users in a full-duplex interactive scenario.
[0088] To facilitate understanding of the above-described process of processing the audio signal to be processed, Figure 4 is a schematic diagram of an audio signal processing process according to an embodiment of this disclosure. In this diagram, mic represents the audio signal to be processed, ref represents the reference audio signal, and res represents the processing of the reference audio signal using a linear echo cancellation algorithm. As shown in Figure 4, the processing system can first acquire the audio signal to be processed that requires echo cancellation, such as the SDK (Software Development Kit). The software development kit (SDK) first processes the audio signal, then performs noise reduction on the audio signal to be processed, and performs VAD detection on the denoised audio signal. Finally, the cloud-based ASR module is used to identify the detected audio signal, thereby ensuring that the processed audio signal is as noise-free as possible. During the noise reduction process, the TED delay estimation compensation and data simulation can be performed on the mic and ref signals to ensure the accuracy of the mic signal processing. Then, the linear echo cancellation algorithm (LACE) is used to process the mic signal based on the ref signal to obtain the processed audio signal res. Finally, the ref, mic, and res signals are input into the echo cancellation model MAEC (Model-AEC model) for echo cancellation to obtain a clean target audio signal. The echo cancellation model can include two parts: a neural network module and a loss function module. The neural network module... The system can include a first fully connected layer FC1, multiple signal synthesis layers DFSMN, a second fully connected layer FC2, and an activation function layer Pred. In the corresponding NN neural network module, the processing system first extracts the res and ref signals for Fourier feature extraction to obtain the corresponding first and second transform features. Then, the first fully connected layer FC1 maps the first and second transform features to obtain the corresponding first and second mapped features. Multiple signal synthesis layers sequentially synthesize the first and second mapped features to obtain the synthesized features. The second fully connected layer performs feature mapping on the synthesized features to obtain the third mapped features. The activation function layer processes the third mapped features to obtain the signal mask. Finally, the signal mask is multiplied by the first transform feature of the res signal to obtain the target transform feature Pred. Performing an inverse Fourier transform on the target transform feature yields the target audio signal.
[0089] By performing echo cancellation on the audio signal to be processed through the above process, the effects of Echo Feedback Loss Enhancement (ERLE) and Perceptual Speech Quality Evaluation (PESQ) can be significantly improved, the word error rate (WER) can be reduced, the effect of Voice Activity Detection (VAD) can be improved, and the false alarm rate (FAR) and false negative rate (MAR) can be reduced.
[0090] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0091] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this disclosure.
[0092] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solutions of this disclosure, in essence, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this disclosure.
[0093] According to embodiments of this disclosure, a training method for an echo cancellation model is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0094] Figure 5 is a flowchart of a training method for an echo cancellation model according to an embodiment of the present disclosure. As shown in Figure 5, the method may include the following steps:
[0095] Step S502: Obtain the raw training data.
[0096] The original training data includes: a first training audio signal and a training speech signal corresponding to the first training audio signal.
[0097] The aforementioned first training audio signal can refer to the audio signal collected by the audio acquisition device used to train the echo cancellation model, and can correspond to the audio signal to be processed in the above embodiments. The aforementioned training speech signal can refer to the clean audio signal used to train the echo cancellation model, and can correspond to the target audio signal in the above embodiments. For example, assuming that user A and user B are currently talking in a full-duplex interactive scenario, user A inputs speech A through the audio acquisition device on mobile phone A, and at the same time, the audio playback device on mobile phone A plays speech B input by user B. Then the aforementioned first training audio signal can refer to the audio signal currently collected by the audio acquisition device, which can simultaneously contain speech A and speech B, and the corresponding training speech signal actually refers to speech A.
[0098] In one alternative embodiment, in order to stably train the initial elimination model, the first training audio signal and the corresponding training speech signal can be obtained first, and the original training data can be constructed based on these two signals.
[0099] Step S504: Perform data amplification on the first training audio signal to obtain the second training audio signal.
[0100] In one optional embodiment, considering that the echo components or noise signals contained in the audio signals acquired by the audio acquisition device may be different in different full-duplex interaction scenarios, the dataset can be enhanced by data amplification (DA) to significantly improve the recognition effect of the automatic speech recognition (ASR) model. Therefore, in order to improve the accuracy of training the initial elimination model, after obtaining the original training data, the first training audio signal can be amplified to expand the amount and types of data in the first training audio signal to obtain a more comprehensive second training audio signal.
[0101] Step S506: The initial cancellation model is progressively learned using the second training audio signal to obtain the echo cancellation model.
[0102] The echo cancellation model is used to perform the echo cancellation processing method in the above embodiments.
[0103] In one alternative embodiment, considering that traditional deep learning methods for speech enhancement typically rely on processing noisy spectral inputs to generate clear spectra, it is difficult to achieve accurate mapping between these complex relationships within the neural network. Therefore, in the actual training process, a progressive learning (PL) approach can be adopted. The initial cancellation model is trained using the amplified second training audio signal to obtain the predicted speech signal corresponding to the second training audio signal. Then, by comparing the predicted speech signal with the training speech signal, the current loss function of the initial cancellation model is constructed. Finally, the model parameters of the initial cancellation model are adjusted using the constructed loss function to obtain a more accurate echo cancellation model.
[0104] In this embodiment, data amplification is performed on the first training audio signal to obtain the second training audio signal, including at least one of the following: randomly masking the first training audio signal according to the time dimension and / or frequency dimension to obtain a masked audio signal, and using the masked audio signal and the first training audio signal as the second training audio signal; randomly extracting multiple signal segments from the first training audio signal and splicing the multiple signal segments to obtain a spliced signal, and using the spliced signal and the first training audio signal as the second training audio signal; using multiple linear echo cancellation algorithms to cancel the echo in the first training audio signal to obtain an echo-cancelled signal, and using the echo-cancelled signal and the first training audio signal as the second training audio signal, wherein the parameters of different linear echo cancellation algorithms are different.
[0105] In one optional embodiment, to ensure the rationality of the amplified second training audio signal, the first training audio signal can be amplified in one or more of the following ways. For example, the first training audio signal can be randomly masked according to a preset dimension, such as time dimension or frequency dimension, to obtain a masked audio signal. Then, the masked audio signal and the first training audio signal can be used as the first training audio signal. Alternatively, multiple signal segments can be randomly extracted from the first training audio signal, and then the multiple signal segments can be spliced together to obtain multiple spliced signals. The processing system can use these multiple spliced signals and the first training audio signal as the second training audio signal. Alternatively, multiple linear echo cancellation algorithms configured with different parameters, such as different time orders, can be used to perform echo cancellation on the first training audio signal to obtain an echo-cancelled signal. Then, the echo-cancelled signal and the first training audio signal can be used as the second training audio signal.
[0106] For ease of understanding, Figure 6 is a schematic diagram of an audio signal masking result according to an embodiment of the present disclosure. The upper region of Figure 6 represents the original audio signal, and the lower region represents the masked audio signal. As shown in Figure 6, when masking the audio signal, a rectangular block can be randomly selected and used to mask a portion of the signal region in the original audio signal to form the corresponding masked audio signal. Figure 7 is a schematic diagram of an audio signal splicing result according to an embodiment of the present disclosure. The upper region of Figure 7 represents the original audio signal, which may be the voice input by the user to the voice acquisition device. The lower region represents the spliced signal obtained by splicing multiple selected segments. By splicing the audio signal, the overlapping and continuity of audio that may occur during natural user dialogue can be reflected, thereby ensuring that the constructed second training audio signal is more consistent with the display scenario, and thus improving the accuracy of the echo cancellation model trained based on the second training audio signal.
[0107] In this embodiment, the initial cancellation model is progressively learned using the second training audio signal to obtain the echo cancellation model. This includes: dividing the multiple network layers of the initial cancellation model to obtain multiple sub-networks corresponding to multiple learning stages; and training the sub-networks corresponding to the multiple learning stages sequentially using the second training audio signal to obtain the echo cancellation model. The input of any sub-network corresponding to a learning stage is the output of the sub-network corresponding to the previous learning stage.
[0108] In one optional implementation, when progressively learning the initial echo cancellation model, the initial model's multiple network layers can be divided into sub-networks corresponding to different learning stages. These sub-networks are connected sequentially according to the learning stages. The input to any sub-network at any learning stage can be the output of the sub-network at the previous learning stage. After obtaining multiple sub-networks, the constructed second training audio signal can be used to train each sub-network at each learning stage sequentially, gradually increasing the SER (Signal Echo Ratio) from the lowest value to progressively eliminate echoes in the second training audio signal, thereby obtaining a predicted speech signal with good echo cancellation performance. Similarly, when using the echo cancellation model to perform echo cancellation on the audio signal to be processed, the audio signal can be processed sequentially using the aforementioned multiple sub-networks to achieve echo cancellation.
[0109] For ease of understanding, Figure 8 is a schematic diagram of an echo cancellation process according to an embodiment of the present disclosure, wherein a represents the input audio signal to be processed, b and c represent the stage results obtained by sequentially performing echo cancellation on the audio signal to be processed using a sub-network, and d represents the clean target audio signal obtained after echo cancellation.
[0110] In this embodiment, the initial cancellation model is progressively learned using the second training audio signal to obtain the echo cancellation model, including: progressively learning the initial cancellation model using the second training audio signal to obtain the predicted speech signal corresponding to the second training audio signal; constructing a loss function for the initial cancellation model based on the training speech signal and the predicted speech signal; and adjusting the model parameters of the initial cancellation model based on the loss function to obtain the echo cancellation model.
[0111] In one optional embodiment, during the actual training process, a progressive learning (PL) approach can be adopted. The initial cancellation model is trained using the amplified second training audio signal to obtain the predicted speech signal corresponding to the second training audio signal. Then, by comparing the predicted speech signal with the training speech signal, the current loss function of the initial cancellation model is constructed. Finally, the model parameters of the initial cancellation model are adjusted using the constructed loss function to obtain a high-precision echo cancellation model.
[0112] In this embodiment, the initial cancellation model is progressively learned using the second training audio signal to obtain the predicted speech signal corresponding to the second training audio signal. This includes: performing a Fourier transform on the second training audio signal to obtain a first training feature; performing a Fourier transform on a preset audio signal to obtain a second training feature; using the initial cancellation model to perform echo cancellation on the first training feature based on the second training feature to obtain a target training feature; and performing an inverse Fourier transform on the target training feature to obtain the predicted speech signal.
[0113] In this embodiment, the initial cancellation model performs echo cancellation on the first training feature based on the second training feature to obtain the target training feature. This includes: inputting the first training feature and the second training feature into the initial cancellation model to obtain the training mask output by the initial cancellation model, wherein the training mask is used to cancel the echo component in the second training audio signal; and obtaining the product of the training mask and the first training feature to obtain the target training feature.
[0114] In this embodiment, the initial elimination model includes: a first fully connected layer, multiple signal synthesis layers, a second fully connected layer, and an activation function layer. The first training feature and the second training feature are input into the initial elimination model to obtain a training mask output by the initial elimination model. This includes: using the first fully connected layer to perform feature mapping on the first and second training features respectively, obtaining a first mapped training feature corresponding to the first training feature and a second mapped training feature corresponding to the second training feature; using multiple signal synthesis layers to synthesize the first mapped training feature and the second mapped training feature, obtaining a synthesized training feature; using the second fully connected layer to perform feature mapping on the synthesized training feature, obtaining a third training feature; and using the activation function layer to process the third training feature, obtaining the training mask.
[0115] It should be noted that the preferred embodiments involved in the above embodiments of this disclosure are the same as the solutions, application scenarios and implementation processes provided in the above embodiments, but are not limited to the solutions provided in the above embodiments.
[0116] According to embodiments of this disclosure, a speech recognition method is also provided. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0117] Figure 9 is a flowchart of a speech recognition method according to an embodiment of the present disclosure. As shown in Figure 9, the method may include the following steps:
[0118] Step S902: Obtain the audio signal to be processed and the reference audio signal.
[0119] The reference audio signal corresponds to the echo component in the audio signal to be processed.
[0120] The aforementioned audio signal to be processed can refer to the audio signal that needs to be identified, or it can be an audio signal received by an audio acquisition device. The aforementioned reference audio signal can refer to the audio signal output by an audio playback device, and it can correspond to the echo component that needs to be eliminated in the audio signal to be processed. For example, suppose the audio acquisition device is currently acquiring what user A is saying, and the corresponding acquired audio signal can be used as the aforementioned audio signal to be processed. If, at the same time that the audio acquisition device is acquiring the audio signal, the audio playback device is playing the reference audio signal, then the reference audio signal will appear in the audio signal to be processed. In this case, the reference audio signal is not the audio signal that needs to be identified, and the reference audio signal appearing in the audio signal to be processed can be regarded as the echo component that needs to be eliminated.
[0121] In one optional embodiment, in order to improve the accuracy of speech recognition, a reference audio signal corresponding to the audio signal to be recognized can be acquired at the same time as the audio signal to be recognized. The audio signal to be recognized is then recognized based on the reference audio signal to avoid identifying incorrect results from the audio signal to be recognized.
[0122] Step S904: Based on the reference audio signal, perform echo cancellation on the linear echo component in the audio signal to be processed to obtain the processed audio signal.
[0123] In one optional embodiment, after acquiring the audio signal to be processed and the reference audio signal, the processing system can use an AEC algorithm, such as a linear echo cancellation algorithm or a nonlinear echo cancellation algorithm, to cancel the linear echo component in the audio signal to be processed based on the acquired reference audio signal. For example, the processing system can first analyze the audio signal to be processed, extract the signal features of the signal, such as frequency and amplitude, and determine the linear echo component in the audio signal to be processed by analyzing the signal features. Then, the linear echo component is removed from the audio signal to be processed as much as possible to obtain the processed audio signal, thereby reducing the amount of data of the reference audio signal in the audio signal to be processed and ensuring the purity of the processed audio signal.
[0124] Step S906: Use an echo cancellation model to cancel the echo of the processed audio signal based on a preset audio signal to obtain the target audio signal.
[0125] The preset audio signal includes at least one of the following: a reference audio signal and an audio signal to be processed.
[0126] In one optional embodiment, considering that the echo cancellation algorithm may not be able to completely eliminate the echo components in the audio signal to be processed, that is, there may be some incompletely eliminated reference audio signals in the processed audio signal, and the audio signal to be processed may also contain environmental noise signals from the environment where the audio acquisition device is located, in order to further improve the accuracy of identifying the audio signal to be processed, the processing system can also use the above-mentioned echo cancellation model to perform echo cancellation on the processed audio signal again based on the above-mentioned preset audio signal, that is, the reference audio signal or the audio signal to be processed, so as to further remove the echo components corresponding to the reference audio signal that may remain in the processed audio signal, the environmental noise signals contained in the audio signal to be processed, and other signals that may affect the quality of the audio signal, so as to obtain a pure audio signal, that is, to obtain the above-mentioned target audio signal.
[0127] Step S908: Perform speech recognition on the target audio signal to obtain the speech recognition result of the audio signal to be processed.
[0128] In one optional embodiment, after the target audio signal is obtained, the processing system can perform speech recognition on the target audio signal to determine the speech recognition result of the audio signal to be processed, and output the speech recognition result in a preset operation interface for the user to view.
[0129] In this embodiment, speech recognition is performed on the target audio signal to obtain the speech recognition result of the audio signal to be processed, including: using a speech activity detection model to detect speech activity in the target audio signal to obtain the speech signal to be recognized; and using a speech recognition model to perform speech recognition on the speech signal to be recognized to obtain the speech recognition result.
[0130] In one optional embodiment, when performing speech recognition on the target audio signal, the processing system can first use a speech activity detection model to perform speech activity detection on the target audio signal to determine the part of the target audio signal that needs to be recognized, that is, to determine the speech signal to be recognized, and then use a speech recognition model to perform speech recognition on the speech signal to be recognized to obtain the corresponding speech recognition result.
[0131] In this embodiment, a speech recognition model is used to perform speech recognition on the speech signal to be recognized to obtain a speech recognition result. This includes: using a confidence model to predict the confidence of the speech signal to be recognized to obtain the confidence of the speech signal to be recognized, wherein the confidence is used to describe the probability that the speech signal to be recognized belongs to the speech signal emitted by the target object; when the confidence of the speech signal to be recognized is greater than a preset confidence, the speech recognition model is used to perform speech recognition on the speech signal to be recognized to obtain a speech recognition result.
[0132] In one optional embodiment, when using a speech recognition model to perform speech recognition on the speech to be recognized, a confidence model can first be used to predict the confidence of the speech signal to be recognized in order to determine the probability that the speech signal to be recognized is a speech signal emitted by the target object, that is, to obtain the aforementioned confidence. For example, assuming that user A inputs speech A into the speech acquisition device, if user C is speaking next to user A, then the speech signal to be recognized may contain both speech A and speech C at the same time. If user A is the target object, the processing system can determine the confidence of speech A and speech C relative to user A respectively. The aforementioned confidence can be determined based on the speech parameters of different speech signals in the speech signal to be recognized, such as volume, pitch variation, etc., which are not limited here. Once the confidence level of the speech signal to be recognized is determined and is greater than the preset confidence level, the processing system can determine that the speech signal to be recognized is the speech signal emitted by the target object. At this time, the speech recognition model described above can be used to perform speech recognition on the speech signal to be recognized, avoiding misidentification of speech signals emitted by other objects as speech signals emitted by the target object, thereby obtaining a speech recognition result with high accuracy.
[0133] In this embodiment, the training data for the speech activity detection model and the speech recognition model are constructed using the training data for the echo cancellation model.
[0134] In one optional embodiment, in order to ensure the accuracy of recognizing the speech signal to be recognized and reduce the influence of environmental noise, echoes and other signals, the data used to train the speech activity detection model and the speech recognition model can refer to the training data used to train the echo cancellation model, thereby ensuring the accuracy of the constructed speech activity detection model and speech recognition model and ensuring the accuracy of the speech recognition result.
[0135] For ease of understanding, Figure 10 is a schematic diagram of a speech recognition process according to an embodiment of the present disclosure. As shown in Figure 10, the processing system can first acquire the audio signal to be processed, and then use a front-end algorithm, such as a linear echo cancellation algorithm, to perform noise reduction on the audio signal to be processed to obtain the processed audio signal. Then, the VAD model, i.e., the speech activity detection model, is used to perform speech activity detection on the noise-reduced audio signal to obtain the speech signal to be recognized. The confidence model is then used to predict the confidence of the speech signal to be recognized to obtain the confidence of the speech signal to be recognized. Finally, the speech recognition model ASR is used to detect the speech signal to be recognized that has a confidence greater than a preset confidence level to obtain the speech recognition result and avoid false detection.
[0136] According to embodiments of this disclosure, an audio signal processing apparatus for implementing the above-described audio signal processing method is also provided. This apparatus can be deployed in a target client. FIG11 is a schematic diagram of an audio signal processing apparatus according to an embodiment of this disclosure. As shown in FIG11, the apparatus 1100 includes: a first acquisition component 1102, a first cancellation component 1104, and a second cancellation component 1106.
[0137] The first acquisition component 1102 is configured to acquire an audio signal to be processed and a reference audio signal, wherein the reference audio signal corresponds to the echo component in the audio signal to be processed; the first cancellation component 1104 is configured to perform echo cancellation on the linear echo component in the audio signal to be processed using the reference audio signal to obtain a processed audio signal; the second cancellation component 1106 is configured to perform echo cancellation on the processed audio signal based on a preset audio signal using an echo cancellation model to obtain a target audio signal, wherein the preset audio signal includes at least one of the following: a reference audio signal and an audio signal to be processed.
[0138] In this embodiment, the second cancellation component 1106 includes: a first transformation component configured to perform a Fourier transform on the processed audio signal to obtain a first transformation feature; a second transformation component configured to perform a Fourier transform on a preset audio signal to obtain a second transformation feature; an echo cancellation component configured to use an echo cancellation model to perform echo cancellation on the first transformation feature based on the second transformation feature to obtain a target transformation feature; and an inverse transformation component configured to perform an inverse Fourier transform on the target transformation feature to obtain a target audio signal.
[0139] In this embodiment, the echo cancellation component is further configured to: input the first transform feature and the second transform feature into the echo cancellation model to obtain the signal mask output by the echo cancellation model, wherein the signal mask is configured to cancel the echo component in the audio signal to be processed; and obtain the product of the signal mask and the first transform feature to obtain the target transform feature.
[0140] In this embodiment, the echo cancellation model includes: a first fully connected layer, multiple signal synthesis layers, a second fully connected layer, and an activation function layer. The echo cancellation component is further configured to: use the first fully connected layer to perform feature mapping on the first transform feature and the second transform feature respectively, to obtain a first mapped feature corresponding to the first transform feature and a second mapped feature corresponding to the second transform feature; use the multiple signal synthesis layers to synthesize the first mapped feature and the second mapped feature to obtain a synthesized feature; use the second fully connected layer to perform feature mapping on the synthesized feature to obtain a third mapped feature; and use the activation function layer to process the third mapped feature to obtain a signal mask.
[0141] In this embodiment, the first cancellation component 1104 includes: a signal alignment component configured to perform signal alignment on the audio signal to be processed and a reference audio signal to obtain an aligned audio signal to be processed and an aligned reference audio signal; and a signal cancellation component configured to perform echo cancellation on the aligned audio signal to be processed based on the aligned reference audio signal to obtain a processed audio signal.
[0142] In this embodiment, the signal alignment component is further configured to: estimate the time delay difference between the audio signal to be processed and the reference audio signal to obtain a target time delay difference; use the target time delay difference to delay the output of the reference audio signal to obtain an aligned reference audio signal; and determine that the audio signal to be processed is the aligned audio signal to be processed.
[0143] In this embodiment, the first acquisition component 1102 includes: a first acquisition component configured to acquire an audio signal to be processed via an audio acquisition device on an electronic device; and a second acquisition component configured to acquire a reference audio signal played via an audio playback device on an electronic device, wherein the echo cancellation model is deployed on the electronic device.
[0144] It should be noted that the first acquisition component 1102, the first elimination component 1104, and the second elimination component 1106 mentioned above correspond to steps S302 to S306 in the above embodiments. The three components and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above components may be hardware or software components stored in memory and processed by one or more processors. The above components may also be part of the device and run in the server 11 provided in the above embodiments.
[0145] It should be noted that the preferred embodiments involved in the above embodiments of this disclosure are the same as the solutions, application scenarios and implementation processes provided in the above embodiments, but are not limited to the solutions provided in the above embodiments.
[0146] According to embodiments of this disclosure, a training apparatus for an echo cancellation model used to implement the training method of the above-described echo cancellation model is also provided. This apparatus can be deployed in a target client. Figure 12 is a schematic diagram of a training apparatus for an echo cancellation model according to an embodiment of this disclosure. As shown in Figure 12, the apparatus 1200 includes: a data acquisition component 1202, a data augmentation component 1204, and a model learning component 1206.
[0147] The data acquisition component 1202 is configured to acquire raw training data, which includes a first training audio signal and a training speech signal corresponding to the first training audio signal. The data amplification component 1204 is configured to amplify the first training audio signal to obtain a second training audio signal. The model learning component 1206 is configured to progressively learn the initial cancellation model using the second training audio signal to obtain an echo cancellation model. The echo cancellation model is configured to execute any of the above methods.
[0148] In this embodiment, the data augmentation component 1204 includes at least one of the following: a first construction component configured to randomly mask a first training audio signal according to a time dimension and / or a frequency dimension to obtain a masked audio signal, and to use the masked audio signal and the first training audio signal as a second training audio signal; a second construction component configured to randomly extract multiple signal segments from the first training audio signal and splice the multiple signal segments to obtain a spliced signal, and to use the spliced signal and the first training audio signal as the second training audio signal; and a third construction component configured to perform echo cancellation on the first training audio signal using multiple linear echo cancellation algorithms to obtain an echo cancellation signal, and to use the echo cancellation signal and the first training audio signal as the second training audio signal, wherein the parameters of different linear echo cancellation algorithms are different.
[0149] In this embodiment, the model learning component 1206 includes: a model partitioning component configured to partition multiple network layers of the initial cancellation model to obtain multiple sub-networks corresponding to multiple learning stages; and a network training component configured to sequentially train the sub-networks corresponding to the multiple learning stages using a second training audio signal to obtain an echo cancellation model, wherein the input of any sub-network corresponding to a learning stage is the output of the sub-network corresponding to the previous learning stage.
[0150] In this embodiment, the model learning component 1206 includes: a model learning component configured to progressively learn an initial cancellation model using a second training audio signal to obtain a predicted speech signal corresponding to the second training audio signal; a loss function construction component configured to construct a loss function for the initial cancellation model based on the training speech signal and the predicted speech signal; and a parameter adjustment component configured to adjust the model parameters of the initial cancellation model based on the loss function to obtain an echo cancellation model.
[0151] In this embodiment, the model learning component includes: a first transform sub-component configured to perform a Fourier transform on a second training audio signal to obtain a first training feature; a second transform sub-component configured to perform a Fourier transform on a preset audio signal to obtain a second training feature; an echo cancellation sub-component configured to use an initial cancellation model to perform echo cancellation on the first training feature based on the second training feature to obtain a target training feature; and an inverse transform sub-component configured to perform an inverse Fourier transform on the target training feature to obtain a predicted speech signal.
[0152] In this embodiment, the echo cancellation sub-component is further configured to: input the first training feature and the second training feature into the initial cancellation model to obtain the training mask output by the initial cancellation model, wherein the training mask is configured to cancel the echo component in the second training audio signal; and obtain the product of the training mask and the first training feature to obtain the target training feature.
[0153] In this embodiment, the echo cancellation model includes: a first fully connected layer, multiple signal synthesis layers, a second fully connected layer, and an activation function layer. The echo cancellation sub-component is further configured to: use the first fully connected layer to perform feature mapping on the first training feature and the second training feature respectively, to obtain a first mapped training feature corresponding to the first training feature and a second mapped training feature corresponding to the second training feature; use the multiple signal synthesis layers to synthesize the first mapped training feature and the second mapped training feature to obtain a synthesized training feature; use the second fully connected layer to perform feature mapping on the synthesized training feature to obtain a third training feature; and use the activation function layer to process the third training feature to obtain a training mask.
[0154] It should be noted that the data acquisition component 1202, data augmentation component 1204, and model learning component 1206 correspond to steps S802 to S806 in the above embodiments. The three components and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above components can be hardware or software components stored in memory and processed by one or more processors. These components can also run as part of the device in the server 11 provided in the above embodiments.
[0155] It should be noted that the preferred embodiments involved in the above embodiments of this disclosure are the same as the solutions, application scenarios and implementation processes provided in the above embodiments, but are not limited to the solutions provided in the above embodiments.
[0156] According to embodiments of this disclosure, a speech recognition device for implementing the above-described speech recognition method is also provided, which can be deployed in a target client. Figure 13 is a schematic diagram of a speech recognition device according to an embodiment of this disclosure. As shown in Figure 13, the device 1300 includes: a second acquisition component 1302, a third cancellation component 1304, a fourth cancellation component 1306, and a speech recognition component 1308.
[0157] The second acquisition component 1302 acquires the audio signal to be processed and the reference audio signal, wherein the reference audio signal corresponds to the echo component in the audio signal to be processed; the third cancellation component 1304 is configured to perform echo cancellation on the linear echo component in the audio signal to be processed based on the reference audio signal to obtain the processed audio signal; the fourth cancellation component 1306 is configured to perform echo cancellation on the processed audio signal based on a preset audio signal using an echo cancellation model to obtain the target audio signal, wherein the preset audio signal includes at least one of the following: the reference audio signal and the audio signal to be processed; and the speech recognition component 1308 performs speech recognition on the target audio signal to obtain the speech recognition result of the audio signal to be processed.
[0158] In this embodiment, the speech recognition component 1308 is further configured to: perform speech activity detection on the target audio signal using a speech activity detection model to obtain a speech signal to be recognized; and perform speech recognition on the speech signal to be recognized using a speech recognition model to obtain a speech recognition result.
[0159] In this embodiment, the speech recognition component 1308 is further configured to: use a confidence model to predict the confidence of the speech signal to be recognized, thereby obtaining the confidence of the speech signal to be recognized, wherein the confidence is used to describe the probability that the speech signal to be recognized belongs to the speech signal emitted by the target object; when the confidence of the speech signal to be recognized is greater than a preset confidence, use a speech recognition model to perform speech recognition on the speech signal to be recognized, thereby obtaining the speech recognition result.
[0160] In this embodiment, the training data for the speech activity detection model and the speech recognition model are constructed using the training data for the echo cancellation model.
[0161] It should be noted that the second acquisition component 1302, the third elimination component 1304, the fourth elimination component 1306, and the speech recognition component 1308 mentioned above correspond to steps S902 to S908 in the above embodiments. The four components and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above components may be hardware or software components stored in memory and processed by one or more processors. The above components may also be part of the device and run in the server 11 provided in the above embodiments.
[0162] It should be noted that the preferred embodiments involved in the above embodiments of this disclosure are the same as the solutions, application scenarios and implementation processes provided in the above embodiments, but are not limited to the solutions provided in the above embodiments.
[0163] According to embodiments of this disclosure, a computer terminal is also provided. The computer terminal may include a server and a client. The server may be any one of the servers in a server device group or a cloud server.
[0164] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.
[0165] In this embodiment, the computer terminal described above can execute the program code in the method.
[0166] Optionally, FIG14 is a structural block diagram of a computer terminal according to an embodiment of the present disclosure. As shown in FIG14, the computer terminal A may include: one or more (only one is shown in the figure) processors 1402, memory 1404, memory controller, and peripheral interface, wherein the peripheral interface is connected to radio frequency components, audio components, and a display.
[0167] The memory can be used to store software programs and components, such as the program instructions / components corresponding to the methods and apparatus in the embodiments of this disclosure. The processor executes various functional applications and data processing by running the software programs and components stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to computer terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0168] The processor can invoke information and application programs stored in the memory through a transmission device to perform the following steps: acquiring an audio signal to be processed and a reference audio signal, wherein the reference audio signal corresponds to the echo component in the audio signal to be processed; performing echo cancellation on the linear echo component in the audio signal to be processed based on the reference audio signal to obtain a processed audio signal; and performing echo cancellation on the processed audio signal based on a preset audio signal using an echo cancellation model to obtain a target audio signal, wherein the preset audio signal includes at least one of the following: a reference audio signal and an audio signal to be processed.
[0169] Optionally, the processor may also execute program code for the following steps: using an echo cancellation model to cancel the echo of the processed audio signal based on a preset audio signal to obtain a target audio signal, including: performing a Fourier transform on the processed audio signal to obtain a first transform feature; performing a Fourier transform on the preset audio signal to obtain a second transform feature; using an echo cancellation model to cancel the echo of the first transform feature based on the second transform feature to obtain a target transform feature; and performing an inverse Fourier transform on the target transform feature to obtain the target audio signal.
[0170] Optionally, the processor may also execute program code with the following steps: using an echo cancellation model to cancel the echo of a first transform feature based on a second transform feature to obtain a target transform feature, including: inputting the first transform feature and the second transform feature into an echo cancellation model to obtain a signal mask output by the echo cancellation model, wherein the signal mask is used to cancel the echo component in the audio signal to be processed; and obtaining the product of the signal mask and the first transform feature to obtain the target transform feature.
[0171] Optionally, the processor may also execute program code with the following steps: The echo cancellation model includes a first fully connected layer, multiple signal synthesis layers, a second fully connected layer, and an activation function layer. The first transform feature and the second transform feature are input into the echo cancellation model to obtain the signal mask output by the echo cancellation model. This includes: using the first fully connected layer to perform feature mapping on the first transform feature and the second transform feature respectively, obtaining a first mapped feature corresponding to the first transform feature and a second mapped feature corresponding to the second transform feature; using multiple signal synthesis layers to synthesize the first mapped feature and the second mapped feature, obtaining the synthesized feature; using the second fully connected layer to perform feature mapping on the synthesized feature, obtaining a third mapped feature; and using the activation function layer to process the third mapped feature, obtaining the signal mask.
[0172] Optionally, the processor may also execute program code that performs the following steps: aligns the audio signal to be processed and the reference audio signal to obtain the aligned audio signal to be processed and the aligned reference audio signal; and performs echo cancellation on the linear echo component in the aligned audio signal to be processed based on the aligned reference audio signal to obtain the processed audio signal.
[0173] Optionally, the processor may also execute program code that performs the following steps: estimating the time delay difference between the audio signal to be processed and the reference audio signal to obtain a target time delay difference; using the target time delay difference to delay the output of the reference audio signal to obtain an aligned reference audio signal; and determining that the audio signal to be processed is the aligned audio signal to be processed.
[0174] Optionally, the processor may also execute program code for the following steps: acquiring the audio signal to be processed and the reference audio signal includes: acquiring the audio signal to be processed through an audio acquisition device on an electronic device; acquiring the reference audio signal played through an audio playback device on an electronic device, wherein the echo cancellation model is deployed on the electronic device.
[0175] Those skilled in the art will understand that the structure shown in Figure 14 above is merely illustrative, and computer terminal A can also be a smartphone (such as an Android phone, iOS phone, etc.), tablet computer, PDA, mobile internet device (MID), PAD, and other terminal devices. Figure 14 above does not limit the structure of the aforementioned electronic device. For example, computer terminal A may also include more or fewer components (such as network interfaces, display devices, etc.) than shown in Figure 14 above, or have a different configuration than that shown in Figure 14 above.
[0176] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0177] According to embodiments of this disclosure, a computer-readable storage medium is also provided. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the method provided in the above embodiments.
[0178] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a computer terminal cluster, or in any mobile terminal in a mobile terminal cluster.
[0179] According to embodiments of this disclosure, a computer program product is also provided. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.
[0180] According to embodiments of this disclosure, a computer program product is also provided. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which can be used to store a computer program that, when executed by a processor, implements the methods provided in the above embodiments.
[0181] According to embodiments of this disclosure, a computer program is also provided. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the method provided in the above embodiments.
[0182] In the above embodiments of this disclosure, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0183] In the several embodiments provided in this disclosure, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of components is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interface; the indirect coupling or communication connection of components may be electrical or other forms.
[0184] The components described as separate parts may or may not be physically separate. The components shown as components may or may not be physical components; that is, they may be located in one place or distributed across multiple network components. Some or all of the components can be selected to achieve the purpose of this embodiment according to actual needs.
[0185] Furthermore, the functional components in the various embodiments of this disclosure can be integrated into a single processing component, or each component can exist physically separately, or two or more components can be integrated into a single component. The integrated components described above can be implemented in hardware or as software functional components.
[0186] If integrated components are implemented as software functional components and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of this disclosure, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0187] The above are merely preferred embodiments of this disclosure. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this disclosure, and these improvements and modifications should also be considered within the scope of protection of this disclosure. Industrial applicability
[0188] The solution provided in this disclosure can be applied to the processing of audio signals and speech recognition. It involves acquiring an audio signal to be processed and a reference audio signal; performing echo cancellation on the linear echo components in the audio signal to be processed based on the reference audio signal to obtain a processed audio signal; and then using an echo cancellation model to perform echo cancellation on the processed audio signal based on a preset audio signal to obtain a target audio signal. By first using a linear echo cancellation algorithm to quickly eliminate most of the simple echoes in the audio signal to be processed, obtaining a processed audio signal, and then using an echo cancellation model to perform deeper echo cancellation processing on the processed audio signal to obtain the target audio signal, the solution can significantly reduce the echo components in the audio signal to be processed, improve the purity of the obtained target audio signal, and ensure the communication quality of different users communicating in full-duplex interactive scenarios. This solves the technical problem of poor echo cancellation performance for audio signals under complex audio conditions in related technologies.
Claims
1. A method for processing an audio signal, comprising: obtaining a to-be-processed audio signal and a reference audio signal, wherein the reference audio signal corresponds to an echo component in the to-be-processed audio signal; performing echo cancellation on a linear echo component in the to-be-processed audio signal based on the reference audio signal to obtain a processed audio signal; performing echo cancellation on the processed audio signal based on a preset audio signal by using an echo cancellation model to obtain a target audio signal, wherein the preset audio signal comprises at least one of the reference audio signal and the to-be-processed audio signal.
2. The method of claim 1, wherein, The performing echo cancellation on the processed audio signal based on a preset audio signal by using an echo cancellation model to obtain a target audio signal comprises: performing Fourier transform on the processed audio signal to obtain a first transformed feature; performing Fourier transform on the preset audio signal to obtain a second transformed feature; performing echo cancellation on the first transformed feature based on the second transformed feature by using the echo cancellation model to obtain a target transformed feature; performing inverse Fourier transform on the target transformed feature to obtain the target audio signal.
3. The method of claim 2, wherein, The performing echo cancellation on the first transformed feature based on the second transformed feature by using the echo cancellation model to obtain a target transformed feature comprises: inputting the first transformed feature and the second transformed feature into the echo cancellation model to obtain a signal mask output by the echo cancellation model, wherein the signal mask is used to cancel the echo component in the to-be-processed audio signal; obtaining a product of the signal mask and the first transformed feature to obtain the target transformed feature.
4. The method of claim 3, wherein, The echo cancellation model comprises a first fully connected layer, a plurality of signal synthesis layers, a second fully connected layer, and an activation function layer, and the inputting the first transformed feature and the second transformed feature into the echo cancellation model to obtain a signal mask output by the echo cancellation model comprises: performing feature mapping on the first transformed feature and the second transformed feature respectively by using the first fully connected layer to obtain a first mapping feature corresponding to the first transformed feature and a second mapping feature corresponding to the second transformed feature; synthesizing the first mapping feature and the second mapping feature by using the plurality of signal synthesis layers to obtain a synthesized feature; performing feature mapping on the synthesized feature by using the second fully connected layer to obtain a third mapping feature; processing the third mapping feature by using the activation function layer to obtain the signal mask.
5. The method according to any one of claims 1 to 4, wherein, The performing echo cancellation on a linear echo component in the to-be-processed audio signal based on the reference audio signal to obtain a processed audio signal comprises: performing signal alignment on the to-be-processed audio signal and the reference audio signal to obtain an aligned to-be-processed audio signal and an aligned reference audio signal; performing echo cancellation on the aligned to-be-processed audio signal based on the aligned reference audio signal by using the linear echo cancellation algorithm to obtain the processed audio signal.
6. The method of claim 5, wherein, The signal alignment on the to-be-processed audio signal and the reference audio signal obtains an aligned to-be-processed audio signal and an aligned reference audio signal, and comprises: estimating a time delay difference between the to-be-processed audio signal and the reference audio signal to obtain a target time delay difference; delaying the reference audio signal by using the target time delay difference to obtain the aligned reference audio signal; determining that the to-be-processed audio signal is the aligned to-be-processed audio signal.
7. A training method of an echo cancellation model, comprising: obtaining original training data, wherein the original training data comprises a first training audio signal and a training speech signal corresponding to the first training audio signal; performing data augmentation on the first training audio signal to obtain a second training audio signal; performing progressive learning on an initial cancellation model by using the second training audio signal to obtain an echo cancellation model, wherein the echo cancellation model is used to execute the method in any one of claims 1 to 6.
8. The method of claim 7, wherein, The data augmentation on the first training audio signal to obtain a second training audio signal comprises at least one of the following: randomly masking the first training audio signal in a time dimension and / or a frequency dimension to obtain a masked audio signal, and taking the masked audio signal and the first training audio signal as the second training audio signal; randomly extracting a plurality of signal segments from the first training audio signal and splicing the plurality of signal segments to obtain a spliced signal, and taking the spliced signal and the first training audio signal as the second training audio signal; performing echo cancellation on the first training audio signal by using a plurality of linear echo cancellation algorithms to obtain an echo cancellation signal, and taking the echo cancellation signal and the first training audio signal as the second training audio signal, wherein parameters of different linear echo cancellation algorithms are different.
9. The method of claim 7, wherein, The progressive learning on the initial cancellation model by using the second training audio signal to obtain an echo cancellation model comprises: dividing a plurality of network layers of the initial cancellation model to obtain sub-networks corresponding to a plurality of learning stages; training the sub-networks corresponding to the plurality of learning stages in sequence by using the second training audio signal to obtain the echo cancellation model, wherein an input of a sub-network corresponding to any one learning stage is an output of a sub-network corresponding to a previous learning stage.
10. The method of any one of claims 7 to 9, wherein, The progressive learning on the initial cancellation model by using the second training audio signal to obtain an echo cancellation model comprises: performing progressive learning on the initial cancellation model by using the second training audio signal to obtain a predicted speech signal corresponding to the second training audio signal; constructing a loss function of the initial cancellation model based on the training speech signal and the predicted speech signal; adjusting model parameters of the initial cancellation model based on the loss function to obtain the echo cancellation model.
11. The method of claim 10, wherein, The progressive learning on the initial cancellation model by using the second training audio signal to obtain a predicted speech signal corresponding to the second training audio signal comprises: performing Fourier transform on the second training audio signal to obtain a first training feature; performing Fourier transform on the preset audio signal to obtain a second training feature; performing echo cancellation on the first training feature based on the second training feature by using the initial cancellation model to obtain a target training feature; performing inverse Fourier transform on the target training feature to obtain the predicted speech signal.
12. The method of claim 11, wherein, The method further includes: inputting the first training feature and the second training feature into the initial cancellation model to obtain a training mask output by the initial cancellation model, wherein the training mask is used to cancel the echo component in the second training audio signal; obtaining a product of the training mask and the first training feature to obtain the target training feature.
13. The method of claim 12, wherein, The initial cancellation model includes a first fully connected layer, a plurality of signal synthesis layers, a second fully connected layer, and an activation function layer. The method further includes: performing feature mapping on the first training feature and the second training feature respectively by using the first fully connected layer to obtain a first mapping training feature corresponding to the first training feature and a second mapping training feature corresponding to the second training feature; synthesizing the first mapping training feature and the second mapping training feature by using the plurality of signal synthesis layers to obtain a synthesized training feature; performing feature mapping on the synthesized training feature by using the second fully connected layer to obtain a third training feature; processing the third training feature by using the activation function layer to obtain the training mask.
14. A speech recognition method, comprising: obtaining a to-be-processed audio signal and a reference audio signal, wherein the reference audio signal corresponds to an echo component in the to-be-processed audio signal; performing echo cancellation on a linear echo component in the to-be-processed audio signal based on the reference audio signal to obtain a processed audio signal; performing echo cancellation on the processed audio signal based on a preset audio signal by using an echo cancellation model to obtain a target audio signal, wherein the preset audio signal includes at least one of the reference audio signal and the to-be-processed audio signal; 15. The method of claim 14, wherein, performing speech recognition on the target audio signal to obtain a speech recognition result of the to-be-processed audio signal. The method further includes: performing speech activity detection on the target audio signal by using a speech activity detection model to obtain a to-be-recognized speech signal; 16. The method of claim 15, wherein, performing speech recognition on the to-be-recognized speech signal by using a speech recognition model to obtain the speech recognition result. The method further includes: The confidence model is used to perform confidence prediction on the to-be-recognized voice signal, to obtain a confidence of the to-be-recognized voice signal, wherein the confidence is used to indicate a probability that the to-be-recognized voice signal belongs to a voice signal emitted by a target object. In a case where the confidence of the to-be-recognized voice signal is greater than a preset confidence, the voice recognition model is used to perform voice recognition on the to-be-recognized voice signal, to obtain the voice recognition result.
17. The method of claim 15 or 16, wherein, The training data of the voice activity detection model and the voice recognition model are constructed by using the training data of the echo cancellation model.
18. A computer terminal, comprising: a memory storing an executable program; a processor configured to run the program, wherein the program performs the method of any one of claims 1 to 17 when running.
19. A computer readable storage medium comprising a stored executable program, wherein, controlling a device in which the computer readable storage medium is located to perform the method of any one of claims 1 to 17 when the executable program is running.
20. A computer program product, comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 17.
Citation Information
Patent Citations
Model training method, echo cancellation method, system and device and storage medium
CN114530160A
Echo elimination method and device, electronic equipment and computer readable medium
CN115083431A
Joint noise and echo suppression for two-way audio communication enhancement
US11924367B1
Noise suppression for speech processing based on machine-learning mask estimation
US9640194B1
Multi-feature fusion echo cancellation method and system based on self-attention transform network
WO2023044961A1