Audio Optimization Method and Related Devices, Electronic Devices, Storage Media
By extracting and interacting the audio representation of speakers and microphones, the echo and noise problems caused by coupling speakers and microphones are solved, and better audio optimization results are achieved, improving the audio quality of communications and smart hardware devices.
Patent Information
- Application Number
- CN202111088953.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-09-16
AI Technical Summary
In the prior art, the coupling of speakers and microphones causes echo and noise to affect audio quality, especially in communication and smart hardware devices to seriously affect the call experience and interactive recognition effects, and the existing audio optimization algorithms are not ideal.
By extracting the representations of collected audio and reference audio, echo, voice and noise representations are extracted respectively, and interactively processed, including echo suppression, noise suppression and voice enhancement, the statistical characteristics of different signals are processed in parallel to improve the audio optimization effect.
Effectively suppress echo and noise, enhance voice quality, improve audio optimization, improve call experience and interactive recognition performance.
Smart Images

Figure CN113990337B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and in particular, to an audio optimization method, related devices, electronic devices, and storage media. Background Art
[0002] In real-world scenarios, due to the coupling between the speaker and the microphone in electronic devices such as mobile phones, smart speakers, and televisions, the microphone will collect the signals emitted by the speaker, thus forming an echo system, and there is inevitably noise in the environment. As a result, the audio quality will be greatly affected.
[0003] Taking the communication field as an example, if the voice of the proximal speaker and the voice of the distal speaker played by the speaker are transmitted to the distal end simultaneously, a time delay will occur after network transmission. The distal speaker will hear their own echo and environmental noise, seriously affecting the call experience and even causing communication barriers; or, taking the field of smart hardware as an example, when interacting with electronic devices such as smart TVs and smart speakers that have both playback and interaction functions, the sound playback source is usually much closer to the microphone than the speaker, thus affecting the interaction recognition. Although several audio optimization algorithms have been proposed currently, the optimization effects are not ideal. In view of this, how to improve the audio optimization effect has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by the present application is to provide an audio optimization method, related devices, electronic devices, and storage media, which can improve the audio optimization effect.
[0005] To solve the above technical problem, a first aspect of the present application provides an audio optimization method, including: extracting a first audio representation of the collected audio, and extracting a second audio representation of the reference audio; respectively extracting a first echo representation, a first speech representation, and a first noise representation based on the first audio representation and the second audio representation; performing interaction processing on the first speech representation with the first echo representation and the first noise representation respectively to obtain a second speech representation, a second echo representation, and a second noise representation; where the interaction processing includes: echo suppression, noise suppression, and speech enhancement; obtaining the optimized target audio based on at least one of the second speech representation, the second echo representation, and the second noise representation.
[0006] To solve the above technical problems, the second aspect of the present application provides an audio optimization device, including: an audio feature extraction module, a voice representation extraction module, a voice representation interaction module, and a target audio acquisition module. The audio feature extraction module is configured to extract a first audio representation of the collected audio and a second audio representation of the reference audio; the voice representation extraction module is configured to respectively extract a first echo representation, a first voice representation, and a first noise representation based on the first audio representation and the second audio representation; the voice representation interaction module is configured to perform interaction processing on the first voice representation with the first echo representation and the first noise representation respectively to obtain a second voice representation, a second echo representation, and a second noise representation; wherein the interaction processing includes: echo suppression, noise suppression, and voice enhancement; the target audio acquisition module is configured to obtain an optimized target audio based on at least one of the second voice representation, the second echo representation, and the second noise representation.
[0007] To solve the above technical problems, the third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other. Program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the audio optimization method in the first aspect above.
[0008] To solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are configured to implement the audio optimization method in the first aspect above.
[0009] In the above solution, a first audio representation of the collected audio is extracted, and a second audio representation of the reference audio is extracted. Based on the first audio representation and the second audio representation, a first echo representation, a first voice representation, and a first noise representation are respectively extracted. On this basis, the first voice representation is respectively subjected to interaction processing with the first echo representation and the first noise representation to obtain a second voice representation, a second echo representation, and a second noise representation. The interaction processing includes: echo suppression, noise suppression, and voice enhancement, and an optimized target audio is obtained based on at least one of the second voice representation, the second echo representation, and the second noise representation. Since the first voice representation and the first echo representation are subjected to interaction processing, it is beneficial to suppress echoes and enhance the voice, and the first voice representation and the first noise representation are subjected to interaction processing, which is beneficial to suppressing noise and enhancing the voice. Therefore, in the audio optimization process, the statistical characteristics of different signals can be considered, and the first voice representation and the first echo representation, as well as the first voice representation and the first noise representation, are processed in parallel by interaction, which is beneficial to improving the audio optimization effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic flowchart of an embodiment of the audio optimization method of the present application;
[0011] Figure 2 It is a schematic framework diagram of an embodiment of an audio optimization model;
[0012] Figure 3 It is a schematic framework diagram of an embodiment of an interaction network;
[0013] Figure 4 It is a schematic flowchart of another embodiment of the audio optimization method of the present application;
[0014] Figure 5 It is a schematic flowchart of an embodiment of training an audio optimization model;
[0015] Figure 6 It is a schematic framework diagram of an embodiment of the audio optimization device of the present application;
[0016] Figure 7 It is a schematic framework diagram of an embodiment of the electronic device of the present application;
[0017] Figure 8 It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners
[0018] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.
[0019] In the following description, specific details such as specific system architectures, interfaces, and technologies are set forth for the purpose of illustration and not limitation in order to provide a thorough understanding of the present application.
[0020] The terms "system" and "network" are often used interchangeably herein. The term "and / or" in this document merely describes an association relationship between associated objects and means that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after. In addition, "plurality" in this document means two or more than two.
[0021] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the audio optimization method of the present application. Specifically, it may include the following steps:
[0022] Step S11: Extract the first audio representation of the collected audio and extract the second audio representation of the reference audio.
[0023] In an implementation scenario, the collected audio can be obtained by the microphone of an electronic device, and the reference audio can be the audio played by the speaker of the electronic device. Taking the voice interaction scenario as an example, the speaker of an electronic device (such as a mobile phone, smart speaker, smart TV) can play the interaction audio (such as, "The weather forecast for the next week has been found for you") processed by TTS (Text To Speech). At the same time, the microphone is in the on state and collects the user's voice command (such as, "Book me a flight to Shanghai tomorrow"). In this case, the collected audio collected by the microphone is the mixture of the user's voice command, the interaction audio played by the speaker, and the ambient noise audio, and the reference audio is the aforementioned interaction audio. The noise audio can include but is not limited to: the sound of the user's normal activities, the sound of the wind in the natural environment, etc., which are not limited here; or, taking the voice call scenario as an example, the speaker of an electronic device (such as a mobile phone, tablet computer) can play the call audio of the remote speaker (such as, "How have you been recently"), and at the same time, the microphone is in the on state and collects the voice audio of the proximal speaker (such as, "I'm fine"). In this case, the collected audio collected by the microphone is the mixture of the call audio of the remote speaker, the voice audio of the proximal speaker, and the ambient noise audio, and the reference audio is the aforementioned call audio. The noise audio can include but is not limited to: the sound of the proximal speaker's normal activities, the sound played by other electronic devices of the proximal speaker (such as speakers, TVs), etc., which are not limited here; or, taking the speech recognition scenario as an example, the speaker of an electronic device (such as a mobile phone, tablet computer) can play music audio, and at the same time, the microphone is in the on state and collects the speech to be recognized. In this case, the collected audio collected by the microphone is the mixture of the music audio, the speech to be recognized, and the ambient noise audio, and the reference audio is the music audio. The noise audio can include but is not limited to: the sound of the user's normal activities, the sound of the wind in the natural environment, etc., which are not limited here. Other cases can be inferred by analogy and will not be exemplified one by one here.
[0024] In an implementation scenario, the short-time Fourier transform can be used to process the collected audio and the reference audio respectively to obtain the amplitude spectrum of the collected audio and the amplitude spectrum of the reference audio. Thus, the amplitude spectrum of the collected audio can be directly used as the first audio representation of the collected audio, and the amplitude spectrum of the reference audio can be directly used as the second audio representation of the reference audio.
[0025] In an implementation scenario, the collected audio and the reference audio can also be respectively subjected to Fourier transform to obtain their complex forms (i.e., the concatenation of the real part and the imaginary part). In this case, the complex form of the collected audio can be used as the first audio representation of the collected audio, and the complex form of the reference audio can be used as the second audio representation of the reference audio. It should be noted that compared with the aforementioned method of using the amplitude spectrum as the audio representation, using the complex form as the audio representation can not only introduce amplitude information but also introduce phase information.
[0026] In an implementation scenario, in order to improve the audio optimization quality as much as possible, before extracting the audio representation, echo cancellation can also be performed based on the collected audio and the reference audio. Specifically, the linear echo in the collected audio can be cancelled through adaptive filtering. It should be noted that in a real scenario, due to the existence of noise audio in the collected audio, the linear echo cannot be completely cancelled. Therefore, the residual echo and noise in the collected audio can be cancelled as much as possible through the embodiments of the present disclosure and the following disclosed embodiments to improve the audio optimization quality as much as possible. On this basis, any of the aforementioned methods can be used to extract the first audio representation of the collected audio and the second audio representation of the reference audio. For specific details, reference can be made to the relevant descriptions above, and details will not be elaborated here.
[0027] Step S12: Based on the first audio representation and the second audio representation, the first echo representation, the first speech representation, and the first noise representation are respectively extracted.
[0028] In an implementation scenario, the first audio representation and the second audio representation can be first fused to obtain a fused audio representation, and then based on the fused audio representation, echo extraction, speech extraction, and noise extraction are respectively performed to obtain the first echo representation, the first speech representation, and the first noise representation. It should be noted that the first echo representation contains echo feature information, the first speech representation contains speech feature information, and the first noise representation contains noise feature information. In addition, the first echo representation and the first noise representation may contain residual speech feature information, and the first speech representation may contain at least one of residual echo feature information and residual noise feature information. This is not limited here.
[0029] In a specific implementation scenario, the first audio representation and the second audio representation can be concatenated to obtain a fused audio representation.
[0030] In a specific implementation scenario, to improve the audio optimization effect, an audio optimization model can be pre-trained, and the audio optimization model can include an echo extraction network, a speech extraction network, and a noise extraction network. On this basis, the echo extraction network can be used to extract features from the aforementioned fused audio representation to obtain a first echo representation, the speech extraction network can be used to extract features from the aforementioned fused audio representation to obtain a first speech representation, and the noise extraction network can be used to extract features from the aforementioned fused audio representation to obtain a first noise representation. It should be noted that the echo extraction network, the speech extraction network, and the noise extraction network can respectively include, but are not limited to: LSTM (Long-Short Term Memory), CNN (Convolutional Neural Network), CRN (Convolutional Recurrent Network), etc., which are not limited herein.
[0031] Step S13: Perform interactive processing on the first speech representation with the first echo representation and the first noise representation respectively to obtain a second speech representation, a second echo representation, and a second noise representation.
[0032] In the embodiments of the present disclosure, the interactive processing includes: echo suppression, noise suppression, and speech enhancement. It should be noted that when performing interactive processing on the first speech representation with the first echo representation, the interactive processing can include echo suppression and speech enhancement, and when performing interactive processing on the first speech representation with the first noise representation, the interactive processing can include noise suppression and speech enhancement.
[0033] In an implementation scenario, echo suppression and speech enhancement can be respectively performed based on the first echo representation and the first speech representation to obtain a second echo representation and a first enhanced speech representation, and noise suppression and speech enhancement can be performed based on the first noise representation and the first speech representation to obtain a second noise representation and a second enhanced speech representation. On this basis, a second speech representation can be obtained based on the first enhanced speech representation and the second enhanced speech representation. In the above manner, performing echo suppression and speech enhancement respectively based on the first echo representation and the first speech representation can facilitate the interaction between the residual speech feature information in the first echo representation and the residual echo feature information in the first speech representation, and performing noise suppression and speech enhancement based on the first noise representation and the first speech representation can facilitate the interaction between the residual speech feature information in the first noise representation and the residual noise feature information in the first speech representation. Therefore, it is beneficial to improve the quality of speech optimization.
[0034] In a specific implementation scenario, for the sake of convenience in description, the steps of the following interaction processing do not currently reflect echo representation, voice representation, and noise representation, but are instead illustrated with the target audio representation and reference audio representation to be interacted with. Specifically, based on the correlation between the target audio representation and the reference audio representation, the reference weight of the reference audio representation can be obtained, and based on the reference weight, the target audio representation and the reference audio representation are weighted to obtain a weighted audio representation.
[0035] Taking echo suppression as an example of the interaction processing, the above-mentioned target audio representation is the first echo representation, which represents that the first echo representation is the target of echo suppression, and the above-mentioned reference audio representation is the first voice representation, which represents that the first voice representation is the reference for echo suppression. In this case, the above-mentioned weighted audio representation is the second echo representation. For the sake of convenience in description, the first echo representation can be denoted as F e and the first voice representation can be denoted as F s , and the reference weight of the first voice representation F s can be expressed as Mask(F e , F s ). Then, the second echo representation F eout can be expressed as:
[0036] F eout = F e + F s * Mask(F e , F s )......(1)
[0037] As can be seen from the above formula (1), since the reference weight Mask(F e , F s ) represents the correlation between the first echo representation F e and the first voice representation F s , by multiplying the first voice representation F s with the reference weight Mask(F e , F s ), the part of the first voice representation F s that is related to the first echo representation F e can be obtained, that is, the residual echo feature information in the first voice representation F s . Thus, the proportion of echo feature information in the second echo representation F eout can be increased, improving the accuracy of the echo representation.
[0038] Taking interactive processing for noise suppression as an example, the above-mentioned target audio representation is the first noise representation, which is used to characterize the first noise representation as the target of noise suppression, and the above-mentioned reference audio representation is the first speech representation, which is used to characterize the first speech representation as the reference of noise suppression. In this case, the above-mentioned weighted audio representation is the second noise representation. For the convenience of description, the first noise representation can be denoted as F n , and the first speech representation can be denoted as F s . The reference weight of the first speech representation F s can be expressed as Mask(F n , F s ). Then the second noise representation F nout can be expressed as:
[0039] F nout = F n + F s * Mask(F n , F s )...(2)
[0040] As can be seen from the above formula (2), since the reference weight Mask(F n , F s ) characterizes the correlation between the first noise representation F n and the first speech representation F s , multiplying the first speech representation F s by the reference weight Mask(F n , F s ) can obtain the part of the first speech representation F s that is related to the first noise representation F n , that is, the residual noise feature information in the first speech representation F s , so as to improve the proportion of noise feature information in the second noise representation F nout and improve the accuracy of the noise representation.
[0041] Taking interactive processing for speech enhancement as an example, the above-mentioned target audio representation is the first speech representation, which is used to characterize the first speech representation as the target of speech enhancement, and the above-mentioned reference audio representation can be the first echo representation, which is used to characterize the first echo representation as the reference of speech enhancement. In this case, the above-mentioned weighted audio representation is the first enhanced speech representation. For the convenience of description, the first speech representation can be denoted as F s , and the first echo representation can be denoted as F e . The reference weight of the first echo representation F e can be denoted as Mask(F s , F e ). Then the second speech representation F sout can be expressed as:
[0042] F sout = F s + F e * Mask(F s , F e )……(3)
[0043] As can be seen from the above formula (3), since the reference weight Mask(F s , F e ) represents the correlation between the first speech representation F s and the first echo representation F e , by multiplying the first echo representation F e with the reference weight Mask(F s , F e ), the part in the first echo representation F e that is related to the first speech representation F s can be obtained, that is, the residual speech feature information in the first echo representation F e . Thus, the proportion of the speech feature information in the second speech representation F sout can be increased, and the accuracy of the speech representation can be improved.
[0044] Taking the interactive processing of speech enhancement as an example, the above target audio representation is the first speech representation, which represents the first speech representation as the target of speech enhancement, and the above reference audio representation can be the first noise representation, which represents the first noise representation as the reference of speech enhancement. In this case, the above weighted audio representation is the second enhanced speech representation. For the convenience of description, the first speech representation can be denoted as F s , the first noise representation can be denoted as F n , and the reference weight of the first noise representation F n can be denoted as Mask(F s , F n ). Then the second speech representation F sout can be expressed as:
[0045] F sout = F s + F n * Mask(F s , F n )……(4)
[0046] As can be seen from the above formula (3), since the reference weight Mask(F s , F n ) represents the correlation between the first speech representation F s and the first noise representation F n , by multiplying the first noise representation F n with the reference weight Mask(F s , Fn ) Multiply to obtain the first noise representation F n In it, the part related to the first speech representation F s That is, the first noise representation F n The remaining speech feature information in it, so as to be able to improve the second speech representation F sout The proportion of speech feature information in it, and improve the accuracy of speech representation.
[0047] In a specific implementation scenario, after obtaining the first enhanced speech representation and the second enhanced speech representation, the first enhanced speech representation and the second enhanced speech representation can be concatenated to obtain the second speech representation, or alternatively, the first enhanced speech representation and the second enhanced speech representation can be averaged to obtain the second speech representation. By the above methods, by concatenating or averaging the first enhanced speech representation and the second enhanced speech representation to obtain the second speech representation, the complexity of speech processing can be reduced, which is beneficial to improving the audio optimization efficiency.
[0048] In a specific implementation scenario, as mentioned above, in order to improve the efficiency of audio optimization, an audio optimization model can be pre-trained to facilitate the audio optimization model to implement the above processes such as representation extraction and interactive processing. Please refer to Figure 2 , Figure 2 It is a schematic framework diagram of an embodiment of the audio optimization model. As Figure 2 shown, the audio optimization model can include an echo branch network, a speech branch network, a noise branch network, a first interaction network located between the echo branch network and the speech branch network, and a second interaction network located between the noise branch network and the speech branch network. Therefore, by extracting the echo representation through the echo branch network, extracting the speech representation through the speech branch network, extracting the noise representation through the noise branch network, and implementing the interactive processing between the first echo representation and the first speech representation through the first interaction network, and implementing the interactive processing between the first noise representation and the first speech representation through the second interaction network, it is thus possible to be beneficial to improving the audio optimization efficiency. It should be noted that the echo branch network, the speech branch network, and the noise branch network can be respectively implemented based on at least one of LSTM (Long-Short Term Memory), CNN (Convolutional Neural Network), and CRN (Convolutional Recurrent Network), which is not limited here. In addition, please refer to Figure 3 , Figure 3 It is a schematic framework diagram of an embodiment of the interaction network. As Figure 3As shown, still taking the aforementioned target audio representation to be interacted with and the reference audio representation to be interacted with as an example, the target audio representation and the reference audio representation can be first concatenated to obtain a concatenated audio representation. On this basis, through convolutional processing of several convolutional layers, the reference weight of the reference audio representation can be obtained. After that, the sum of the product of the reference weight and the reference audio representation and the target audio representation can be used as the weighted audio representation. In addition, for the specific training process of the audio optimization model, reference can be made to the following related disclosed embodiments, which will not be elaborated here for the time being.
[0049] In an implementation scenario, echo suppression and speech enhancement can be respectively performed based on the first echo representation and the first speech representation to obtain a second echo representation and an enhanced first speech representation, and noise suppression and speech enhancement can be performed based on the first noise representation and the enhanced first speech representation to obtain a second noise representation and a second speech representation. In the above manner, performing echo suppression and speech enhancement respectively based on the first echo representation and the first speech representation can facilitate the interaction between the residual speech feature information in the first echo representation and the residual echo feature information in the first speech representation, and performing noise suppression and speech enhancement based on the first noise representation and the enhanced first speech representation can facilitate the interaction between the residual speech feature information in the first noise representation and the residual noise feature information in the enhanced first speech representation. Therefore, it is beneficial to improve the quality of speech optimization.
[0050] In a specific implementation scenario, similar to the foregoing description, for the convenience of description, the target audio representation to be interacted with and the reference audio representation can be used to assist in explaining the interaction processing process. Specifically, reference can be made to the foregoing related description. This implementation scenario only describes the differences between the two, and the same parts can be referred to the foregoing related description, which will not be elaborated here. The processes of performing echo suppression and speech enhancement respectively based on the first echo representation and the first speech representation can be respectively referred to formula (1) and its related description and formula (3) and its related description, which will not be elaborated here. Different from the foregoing description, after obtaining the enhanced first speech representation F sout through formula (3), based on the first noise representation F n and the enhanced first speech representation F sout , the reference weight Mask(F sout , F n , F sout ) of the enhanced first speech representation F n can be obtained. Then, based on the reference weight Mask(F sout ), the first noise representation F n and the enhanced first speech representation F sout can be weighted to obtain a second noise representation F nout , which can be specifically expressed as:
[0051] Fnout = F n + F sout *Mask(F n , F sout )……(5)
[0052] As can be seen from formula (5), since the reference weight Mask(F n , F sout ) represents the correlation between the first noise representation F n and the enhanced first speech representation F sout , by multiplying the enhanced first speech representation F sout with the reference weight Mask(F n , F sout ), the part of the enhanced first speech representation F sout related to the first noise representation Fn can be obtained, that is, the residual noise feature information in the enhanced first speech representation F sout . Thus, the proportion of the noise feature information in the second noise representation F nout can be increased, and the accuracy of the noise representation can be improved. At the same time, after obtaining the enhanced first speech representation F sout through formula (3), the reference weight Mask(F n , F sout ) of the first noise representation F n can also be obtained based on the first noise representation F sout and the enhanced first speech representation F n . Then, based on the reference weight Mask(F sout , F n ), the first noise representation F n and the enhanced first speech representation F sout can be weighted to obtain the second speech representation F' sout , which can be specifically expressed as:
[0053] F' sout = F sout + F n *Mask(F sout , F n )......(6)
[0054] As can be seen from formula (6), since the reference weight Mask(F sout , F n ) represents the correlation between the enhanced first speech representation F sout and the first noise representation F n , by multiplying the first noise representation F n with the reference weight Mask(F sout , F nBy multiplying, the first noise representation F can be obtained. n The part in n that is related to the enhanced first speech representation F sout is the first noise representation F n The residual speech feature information in n , so that the proportion of speech feature information in the second speech representation F' sout can be improved, and the accuracy of the speech representation can be enhanced.
[0055] In a specific implementation scenario, as described above, to improve the efficiency of audio optimization, an audio optimization model can be pre-trained to facilitate the audio optimization model to implement the above processes such as representation extraction and interactive processing. For the specific framework of the audio processing model, reference can be made to Figure 2 and the foregoing related descriptions, which will not be elaborated here.
[0056] In an implementation scenario, noise suppression and speech enhancement can be respectively performed based on the first noise representation and the first speech representation to obtain the second noise representation and the enhanced first speech representation, and echo suppression and speech enhancement can be performed based on the first echo representation and the enhanced first speech representation to obtain the second echo representation and the second speech representation. In the above manner, performing noise suppression and speech enhancement respectively based on the first noise representation and the first speech representation can facilitate the interaction between the residual speech feature information in the first noise representation and the residual noise feature information in the first speech representation, and performing echo suppression and speech enhancement based on the first echo representation and the enhanced first speech representation can facilitate the interaction between the residual speech feature information in the first echo representation and the residual echo feature information in the enhanced first speech representation. Therefore, it is beneficial to improve the quality of speech optimization.
[0057] In a specific implementation scenario, similar to the foregoing description, for the convenience of description, the target audio representation and the reference audio representation to be interacted can be used to assist in explaining the interactive processing process. Specifically, reference can be made to the foregoing related descriptions. Only the differences between the two in this implementation scenario will be described, and the same parts can be referred to the foregoing related descriptions, which will not be elaborated here. The processes of performing noise suppression and speech enhancement respectively based on the first noise representation and the first speech representation can be respectively referred to formula (2) and its related descriptions and formula (4) and its related descriptions, which will not be elaborated here. Different from the foregoing description, after obtaining the enhanced first speech representation F sout through formula (4), based on the first echo representation F e and the enhanced first speech representation F sout , the reference weight Mask(F sout , F e , F sout ) of the enhanced first speech representation F e can be obtained. Then, based on the reference weight Mask(F sout),weight the first echo representation F e and the enhanced first speech representation F sout to obtain a second echo representation F eout , which can be specifically expressed as:
[0058] F eout = F e + F sout * Mask(F e , F sout )...(7)
[0059] As can be seen from formula (5), since the reference weight Mask(F e , F sout ) characterizes the correlation between the first echo representation F e and the enhanced first speech representation F sout , therefore, by multiplying the enhanced first speech representation F sout with the reference weight Mask(F e , F sout ), the part of the enhanced first speech representation F sout that is related to the first echo representation F e can be obtained, that is, the residual echo feature information in the enhanced first speech representation F sout , thereby improving the proportion of echo feature information in the second echo representation F eout and enhancing the accuracy of the echo representation. At the same time, after obtaining the enhanced first speech representation F sout through formula (4), the reference weight Mask(F e and the enhanced first speech representation F sout can also be used to obtain the reference weight Mask(F n ) of the first echo representation F sout , F e ), then based on the reference weight Mask(F sout , F e ), weight the first echo representation F e and the enhanced first speech representation F sout to obtain a second speech representation F' sout , which can be specifically expressed as:
[0060] F' sout = F sout + F e * Mask(F sout , F e )...(8)
[0061] As can be seen from formula (6), since the reference weight Mask(F sout , Fe ) Characterize the enhanced first speech representation F sout and the first echo representation F e between the correlations. Therefore, through the first echo representation F e and the reference weight Mask(F sout , F e ) multiplied, the part of the first echo representation F n related to the enhanced first speech representation F sout can be obtained, that is, the residual speech feature information in the first echo representation F e , so as to improve the proportion of speech feature information in the second speech representation F' sout and improve the accuracy of the speech representation.
[0062] In a specific implementation scenario, as described above, in order to improve the efficiency of audio optimization, an audio optimization model can be pre-trained to facilitate the audio optimization model to implement the above-mentioned representation extraction, interactive processing and other processes. For the specific framework of the audio processing model, reference can be made to Figure 2 and the foregoing related descriptions, which will not be elaborated here.
[0063] It should be noted that in the actual application process, any of the above three interaction methods can be selected to obtain the second speech representation, the second echo representation and the second noise representation, which is not limited here.
[0064] Step S14: Based on at least one of the second speech representation, the second echo representation and the second noise representation, obtain the optimized target audio.
[0065] Specifically, according to actual application requirements, the target audio can be restored based on at least one of the second speech representation, the second echo representation, and the second noise representation. For example, in a speech recognition scenario, it is often necessary to perform speech recognition on the audio after relevant optimizations such as echo suppression and noise suppression to improve the accuracy of speech recognition as much as possible. Therefore, the target audio can be restored based on the second speech representation. It should be noted that in this case, the target audio is speech audio; or, in a scenario of studying an echo system, the target audio can be restored based on the second echo representation. It should be noted that in this case, the target audio is echo audio; or, in a scenario of studying a noise system, the target audio can be restored based on the second noise representation. It should be noted that in this case, the target audio is noise audio. Other scenarios can be inferred by analogy and will not be exemplified one by one here. In addition, as mentioned above, the audio representation can be an amplitude spectrum or a complex form, so the second speech representation, the second echo representation, and the second noise representation can also correspondingly be an amplitude spectrum or a complex form. That is, the target speech can ultimately be restored based on the amplitude spectrum or the complex form. The specific restoration process can refer to restoration algorithms such as the Griffin Lim algorithm and will not be elaborated here.
[0066] In the above solution, the first audio representation of the collected audio is extracted, and the second audio representation of the reference audio is extracted. Based on the first audio representation and the second audio representation, the first echo representation, the first speech representation, and the first noise representation are respectively extracted. On this basis, the first speech representation is respectively interacted with the first echo representation and the first noise representation to obtain the second speech representation, the second echo representation, and the second noise representation. The interactive processing includes: echo suppression, noise suppression, and speech enhancement. And based on at least one of the second speech representation, the second echo representation, and the second noise representation, the optimized target audio is obtained. Since the interactive processing of the first speech representation and the first echo representation is beneficial to suppressing echoes and enhancing speech, and the interactive processing of the first speech representation and the first noise representation is beneficial to suppressing noise and enhancing speech, in the audio optimization process, the statistical characteristics of different signals can be considered, and the parallel method is used to interactively process the first speech representation and the first echo representation, and the first speech representation and the first noise representation, which is beneficial to improving the audio optimization effect.
[0067] Please refer to Figure 4 , Figure 4 is a schematic flowchart of another embodiment of the audio optimization method of the present application. In the embodiments of the present disclosure, the target audio is obtained through a preset numerical round optimization stage. Specifically, it can include the following steps:
[0068] Step S41: Extract the first audio representation of the collected audio and extract the second audio representation of the reference audio.
[0069] For details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0070] Step S42: Based on the first audio representation and the second audio representation, respectively extract a first echo representation, a first speech representation, and a first noise representation.
[0071] For details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0072] Step S43: Perform interactive processing on the first speech representation with the first echo representation and the first noise representation respectively to obtain a second speech representation, a second echo representation, and a second noise representation.
[0073] In the embodiments of the present disclosure, the interactive processing includes: echo suppression, noise suppression, and speech enhancement. For details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0074] Step S44: Detect whether the number of optimization rounds in the current optimization stage is lower than a preset value. If so, execute step S45; otherwise, execute step S47.
[0075] In one implementation scenario, the preset value can be set according to actual application requirements. For example, in the case of high quality requirements for audio optimization, the preset value can be set appropriately larger, such as 7, 8, 9, etc. Or, in the case of high efficiency requirements for audio optimization and relatively loose quality requirements for audio optimization, the preset value can be set appropriately smaller, such as 2, 3, 4, etc. It is not limited herein.
[0076] In one implementation scenario, as shown in the foregoing disclosed embodiments and Figure 2 shown, in order to improve the audio optimization efficiency, an audio optimization model can be pre-trained. The audio optimization model includes an echo branch network, a speech branch network, a noise branch network, a first interaction network between the echo branch network and the speech branch network, and a second interaction network between the noise branch network and the speech branch network. Please continue to refer to Figure 2, the echo branch network includes a preset number of sequentially connected echo subnetworks, the voice branch network includes a preset number of sequentially connected voice subnetworks, the noise branch network includes a preset number of sequentially connected noise subnetworks, the first interactive network includes a preset number of first interactive subnetworks, the second interactive network includes a preset number of second interactive subnetworks, and the i-th first interactive subnetwork is located between the i-th echo subnetwork and the i-th voice subnetwork, and the i-th second interactive subnetwork is located between the i-th noise subnetwork and the i-th voice subnetwork. It should be noted that the echo subnetwork, the voice subnetwork, and the noise subnetwork can be implemented based on at least one of LSTM, CNN, and CRN, respectively, and are not limited here. In addition, the network structures of the first interactive subnetwork and the second interactive subnetwork can be referred to. Figure 3 The related descriptions in the aforementioned disclosed embodiments will not be repeated here.
[0077] Step S45: extracting features from the second echo representation, the second speech representation and the second noise representation respectively to obtain a new first echo representation, a new first speech representation and a new first noise representation.
[0078] When the number of optimization rounds in the current optimization stage is less than the preset value, the next stage optimization can be continued. Specifically, the features of the second echo representation, the second speech representation, and the second noise representation can be extracted to obtain a new first echo representation, a new first speech representation, and a new first noise representation. Figure 2 , when the current optimization stage is the first stage, the first echo subnetwork extracts the first echo representation, the first speech subnetwork extracts the first speech representation, and the first noise subnetwork extracts the first noise representation. On this basis, the first first interactive subnetwork interactively processes the first echo representation and the first speech representation to obtain the second echo representation and the first enhanced speech representation. The second interactive subnetwork interactively processes the first noise representation and the first speech representation to obtain the second noise representation and the second enhanced speech representation. The second speech representation can be obtained based on the first enhanced speech representation and the second enhanced speech representation. Since the number of optimization rounds in the current optimization stage is lower than the preset value, the second echo subnetwork can be used to extract features from the second echo representation to obtain a new first echo representation, and the second speech subnetwork can be used to extract features from the second speech representation to obtain a new first speech representation, and the second noise subnetwork can be used to extract features from the second noise representation to obtain a new first noise representation. Other optimization stages can be deduced in this way, and examples are not given one by one here.
[0079] Step S46: re-execute the above step S43 and subsequent steps.
[0080] Specifically, when the new first echo representation, new first speech representation, and new first noise representation are obtained, the foregoing interactive processing and related steps can be performed. For specific details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0081] Step S47: Based on at least one of the second speech representation, second echo representation, and second noise representation, recover the target audio.
[0082] When the number of optimization rounds in the current optimization stage is not less than a preset value, the target audio can be recovered based on at least one of the second speech representation, second echo representation, and second noise representation. For specific details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0083] In the above solution, the first audio representation of the collected audio is extracted, and the second audio representation of the reference audio is extracted. Based on the first audio representation and the second audio representation, the first echo representation, first speech representation, and first noise representation are respectively extracted. On this basis, the first speech representation is respectively interacted with the first echo representation and the first noise representation to obtain the second speech representation, second echo representation, and second noise representation. The interactive processing includes: echo suppression, noise suppression, and speech enhancement. And based on at least one of the second speech representation, second echo representation, and second noise representation, the optimized target audio is obtained. Since the first speech representation and the first echo representation are interacted, it is beneficial to suppress echoes and enhance speech, and the first speech representation and the first noise representation are interacted, which is beneficial to suppress noise and enhance speech. On this basis, by further detecting whether the number of optimization rounds in the current optimization stage is lower than the preset value, and when it is lower than the preset value, feature extraction is respectively performed on the second echo representation, second speech representation, and second noise representation to obtain the new first echo representation, new first speech representation, and new first noise representation, and the above interactive processing and subsequent processes are re-executed. Therefore, echoes and noise can be further suppressed and speech can be enhanced. And when it is not less than the preset value, the target audio is recovered based on at least one of the second speech representation, second echo representation, and second noise representation. Therefore, during the audio optimization process, continuous multi-round optimization can be carried out, which is beneficial to gradually improving the audio optimization quality.
[0084] Please refer to Figure 5 , Figure 5It is a schematic flowchart of an embodiment for training an audio optimization model. As described in the foregoing disclosed embodiments, the target audio can be obtained by processing with the audio optimization model, and the audio optimization model can be trained using several groups of sample audio data. Each group of sample audio data may include sample noise audio, sample reference audio, sample speech audio, and sample acquisition audio generated from the sample noise audio, sample reference audio, and sample speech audio. Specifically, the audio optimization model can be trained through the following steps:
[0085] Step S51: Respectively extract features from the sample noise audio, sample reference audio, sample speech audio, and sample acquisition audio to obtain an initial sample noise representation, an initial sample reference representation, an initial sample speech representation, and an initial sample audio representation.
[0086] In an implementation scenario, the first room impulse response of the noise, the second room impulse response of the echo, and the third room impulse response of the speech can be obtained respectively, and the sample noise audio, sample reference audio, and sample speech audio are respectively convolved with the first room impulse response, the second room impulse response, and the third room impulse response to obtain the sample acquisition audio. In the above manner, through the simulation generation method, as rich sample data as possible can be obtained, which can overcome the problem of difficult sample acquisition in real-time scenarios and greatly improve the model training performance.
[0087] In a specific implementation scenario, a room model can be pre-constructed. The room model may include, but is not limited to: room dimensions (i.e., length, width, height), room materials (such as, hollow brick wall, solid brick wall, concrete wall, surface latex, surface wallpaper, etc.), which are not limited herein. On this basis, based on this room model, the first room impulse response, the second room impulse response, and the third room impulse response can be generated through algorithms such as the Image Method. The specific generation process can refer to the technical details of the Image Method and will not be elaborated herein.
[0088] In a specific implementation scenario, in order to enrich the sample data as much as possible, the room impulse responses at different positions of the room model can be generated. Specifically, the room impulse responses at different angles and distances deviating from the sound source in the room model can be generated. For example, the room impulse response at 1 meter in the direction 10 degrees to the left of the sound source, the impulse response at 2 meters in the direction 30 degrees to the right of the sound source, which are not limited herein.
[0089] In a specific implementation scenario, the sample voice audio is the audio of a person speaking. And to further enrich the sample data, different sample noise audios and sample reference audios can also be used according to different scenarios. Taking the voice interaction scenario as an example, white noise audio can be used as the sample noise audio, and TTS audio can be used as the sample reference audio; or music audio can be used as the sample noise audio, and TTS audio can be used as the sample reference audio. Or, taking the voice call scenario as an example, either white noise audio or music audio can be used as the sample noise audio, and either TTS audio or the speaker's audio can be used as the sample reference audio. Or, taking the speech recognition scenario as an example, white noise audio can be used as the sample noise audio, and music audio can be used as the sample reference audio. Other situations can be deduced by analogy and will not be elaborated one by one here. It should be noted that in order to make the audio optimization model applicable to different scenarios as much as possible, the sample audio data used in the training process may cover as many usage scenarios as possible.
[0090] In a specific implementation scenario, for the convenience of description, the first room impulse response can be denoted as I n , the second room impulse response can be denoted as I e , the third room impulse response can be denoted as I s , then the sample acquisition audio y can be expressed as:
[0091] y = s * I s + ref * I e + n * I n ……(9)
[0092] In the above formula (9), s represents the sample voice audio, ref represents the sample reference audio, n represents the sample noise audio, and * represents the convolution operation.
[0093] In another implementation scenario, different from the generation process of the aforementioned sample acquisition audio, in order to improve the audio optimization quality as much as possible, the sample audio after convolution can be used as the initial acquisition audio, and linear echo cancellation is performed on the initial acquisition audio to obtain the sample acquisition audio. Specifically, adaptive filtering can be used to cancel the linear echo in the initial acquisition audio based on the sample reference audio to obtain the sample acquisition audio y:
[0094] y = s * I s + ref n * I e + n * I n ……(10)
[0095] In the above formula (9), ref n represents the non-linear component in the sample reference audio.
[0096] In an implementation scenario, as described in the foregoing disclosed embodiments, any one of the amplitude spectrum and the complex form can be used for feature extraction. That is, the initial sample noise representation, the initial sample reference representation, the initial sample speech representation, and the initial sample audio representation can all be represented by the amplitude spectrum, or they can all be represented by the complex form, which is not limited herein.
[0097] Step S52: Based on the initial sample audio representation and the initial sample reference representation, the first sample echo representation, the first sample speech representation, and the first sample noise representation are respectively extracted.
[0098] Specifically, with reference to Figure 2 , the initial sample audio representation and the initial sample reference representation can be fused by means of splicing or the like to obtain the initial sample fused representation. On this basis, the first sample echo representation is extracted by using the echo branch network, the first sample speech representation is extracted by using the speech branch network, and the first sample noise representation is extracted by using the noise branch network.
[0099] Step S53: The first sample speech representation is respectively interacted with the first sample echo representation and the first sample noise representation to obtain the second sample speech representation, the second sample echo representation, and the second sample noise representation.
[0100] In an implementation scenario, similar to an interaction method described in the foregoing disclosed embodiments, the first interaction network can be used to perform echo suppression and speech enhancement on the first sample speech representation and the first sample echo representation to obtain the second sample echo representation and the first sample enhanced speech representation, and the second interaction network can be used to perform noise suppression and speech enhancement on the first sample speech representation and the first sample noise representation to obtain the second sample noise representation and the second sample enhanced speech representation. On this basis, the first sample enhanced speech representation and the second sample enhanced speech representation can be processed by means of splicing, averaging, etc. to obtain the second sample speech representation. The specific process can refer to the relevant description in the foregoing disclosed embodiments and will not be elaborated herein.
[0101] In another implementation scenario, similar to another interaction method described in the foregoing disclosed embodiments, the first interaction network can be used to perform echo suppression and speech enhancement on the first sample speech representation and the first sample echo representation to obtain the second sample echo representation and the first sample speech representation after enhancement, and then the second interaction network can be used to perform noise suppression and speech enhancement on the first sample speech representation after enhancement and the first sample noise representation to obtain the second sample noise representation and the second sample speech representation. The specific process can refer to the relevant description in the foregoing disclosed embodiments and will not be elaborated herein.
[0102] In yet another implementation scenario, similar to another interaction method described in the foregoing disclosed embodiments, a second interaction network can be used to perform noise suppression and speech enhancement on the first sample speech representation and the first sample noise representation to obtain a second sample noise representation and an enhanced first sample speech representation, and then a first interaction network is used to perform echo suppression and speech enhancement on the enhanced first sample speech representation and the first sample echo representation to obtain a second sample echo representation and a second sample speech representation. The specific process can refer to the relevant descriptions in the foregoing disclosed embodiments and will not be elaborated here.
[0103] In addition, as Figure 2 and described in the foregoing disclosed embodiments, the echo branch network includes a preset number of sequentially connected echo sub-networks, the speech branch network includes a preset number of sequentially connected speech sub-networks, the noise branch network includes a preset number of sequentially connected noise sub-networks, the first interaction network includes a preset number of first interaction sub-networks, the second interaction network includes a preset number of second interaction sub-networks, and the i-th first interaction sub-network is located between the i-th echo sub-network and the second i-th speech sub-network, and the i-th second interaction sub-network is located between the i-th noise sub-network and the i-th speech sub-network. In this case, the initial sample audio representation and the initial sample reference representation can be fused by splicing or other means to obtain an initial sample fusion representation. The first sample echo representation is extracted using the first echo sub-network, the first sample speech representation is extracted using the first speech sub-network, and the first sample noise representation is extracted using the first noise sub-network. Then, the latest extracted first sample echo representation, first sample speech representation, and first sample noise representation are respectively processed through the first first interaction sub-network and the first second interaction sub-network to obtain a second sample echo representation, a second sample speech representation, and a second sample noise representation, and it is detected whether all the preset number of sub-networks have been processed. If so, the following step S54 can be executed; otherwise, the above steps of feature extraction and interaction processing can be re-executed using the second sub-network (i.e., the second echo sub-network, the second speech sub-network, the second noise sub-network, and the second first interaction sub-network and the second second interaction sub-network) until all the preset number of sub-networks are processed.
[0104] Step S54: Based on the difference between the initial sample speech representation and the second sample speech representation, a speech optimization sub-loss is obtained, based on the difference between the initial sample noise representation and the second sample noise representation, a noise optimization sub-loss is obtained, and based on the difference between the initial sample reference representation and the second sample echo representation, an echo optimization sub-loss is obtained.
[0105] In an implementation scenario, when the relevant representation is described by the amplitude spectrum, the MSE (Mean Squared Error) can be used to calculate each sub-loss. For the specific calculation process, reference can be made to the technical details related to MSE, which will not be elaborated here.
[0106] In an implementation scenario, when the relevant description is in the complex form, the SSNR (Segmental SNR) can be used to calculate each sub-loss. For the specific process, reference can be made to the technical details related to SSNR, which will not be elaborated here.
[0107] Step S55: Based on the iteration number of this round of iteration, obtain the voice optimization weight, the noise optimization weight, and the echo optimization weight.
[0108] It should be noted that the audio optimization model is obtained through multiple rounds of iterative training. For example, the audio optimization model can be obtained through 500 rounds, 1000 rounds, etc. of iterative training. The specific number of iterations is not limited here. In this case, in order to improve the training effect, each round of iteration can focus on training different branch networks respectively. Specifically, when the iteration number meets the preset condition, the echo optimization weight can be higher than the noise optimization weight, while when the iteration number does not meet the preset condition, the echo optimization weight can be lower than the noise optimization weight.
[0109] In an implementation scenario, the preset condition can include that the iteration number can be divided evenly by 2, that is, the iteration number is an even number; or, the preset condition can also include that the iteration number cannot be divided evenly by 2, that is, the iteration number is an odd number, which is not limited here. By setting the preset condition as the iteration number being an even number or an odd number in the above manner, it is possible to alternately focus on the echo branch network and the noise branch network during multiple rounds of training, which is beneficial to improving the training effect.
[0110] In an implementation scenario, during the training process, it is also possible to slightly focus on one of the branch networks. For example, the preset condition can include that the iteration number cannot be divided evenly by 3, that is, the iteration number is 1, 2, 4, 5, 7, 8, etc. In this case, when the iteration number is 1, 2, 4, 5, 7, 8, etc., the echo optimization weight is higher than the noise optimization weight, that is, it focuses on training the echo branch network; or, the preset condition can also include that the iteration number can be divided evenly by 3. That is, when the iteration number cannot be divided evenly by 3, the noise optimization weight is higher than the echo optimization weight. That is to say, when the iteration number is 1, 2, 4, 5, 7, 8, etc., the noise optimization weight is higher than the echo optimization weight, that is, it focuses on training the noise branch network. Other situations can be deduced by analogy, and no further examples will be given here.
[0111] In one implementation scenario, the voice optimization weight may not vary with the number of iterations, that is, the voice optimization weight can be set to a fixed constant, such as 1, which is not limited herein.
[0112] Step S56: Based on the weighted results of the voice optimization sub-loss, noise optimization sub-loss, and echo optimization sub-loss with the voice optimization weight, noise optimization weight, and echo optimization weight, adjust the network parameters of the audio optimization model.
[0113] In one implementation scenario, the weighted result is the weighted sum (or weighted average) of each sub-loss. For ease of description, the voice optimization sub-loss can be denoted as L s , the noise optimization sub-loss can be denoted as L 竹 , the echo optimization sub-loss can be denoted as L e , and the voice optimization weight can be denoted as λ s , the noise optimization weight can be denoted as λ n , the echo optimization weight can be denoted as λ e , then the weighted loss L can be expressed as:
[0114] L = λ s ×L s +λ 竹 ×L 竹 +λ e ×L e ......(11)
[0115] Taking the preset condition that the number of iterations can be divided evenly by 2 as an example, in the case where the number of iterations is even, the echo optimization weight λ e can be higher than the noise optimization weight λ n , for example, the echo optimization weight λ e can be set to 0.9, and the noise optimization weight λ 竹 can be set to 0.1; or, in the case where the number of iterations is odd, the noise optimization weight λ n can be higher than the echo optimization weight λ e , for example, the noise optimization weight λ 竹 can be set to 0.9, and the echo optimization weight λ e can be set to 0.1. Other cases can be deduced by analogy and will not be elaborated herein one by one.
[0116] In one implementation scenario, the optimization method of gradient descent can be adopted to adjust the network parameters of the audio optimization model. For the specific adjustment process, refer to the technical details of optimization methods such as gradient descent, which will not be elaborated herein.
[0117] In the above solution, based on the number of iterations in the current round, the voice optimization weight, noise optimization weight, and echo optimization weight are obtained, and the network parameters of the audio optimization model are adjusted based on the weighted results of the voice optimization sub-loss, noise optimization sub-loss, and echo optimization sub-loss. During each round of iteration, it is possible to focus on training different branch networks respectively, which is beneficial to improving the training effect of the audio optimization model.
[0118] Please refer to Figure 6 , Figure 6 which is a schematic framework diagram of an embodiment of the audio optimization device 60 of the present application. The audio optimization device 60 includes: an audio feature extraction module 61, a sound representation extraction module 62, a sound representation interaction module 63, and a target audio acquisition module 64. The audio feature extraction module 61 is configured to extract a first audio representation of the collected audio and a second audio representation of the reference audio; the sound representation extraction module 62 is configured to respectively extract a first echo representation, a first voice representation, and a first noise representation based on the first audio representation and the second audio representation; the sound representation interaction module 63 is configured to perform interaction processing on the first voice representation with the first echo representation and the first noise representation respectively to obtain a second voice representation, a second echo representation, and a second noise representation; wherein, the interaction processing includes: echo suppression, noise suppression, and voice enhancement; the target audio acquisition module 64 is configured to obtain the optimized target audio based on at least one of the second voice representation, the second echo representation, and the second noise representation.
[0119] In the above solution, a first audio representation of the collected audio is extracted, and a second audio representation of the reference audio is extracted. Based on the first audio representation and the second audio representation, a first echo representation, a first voice representation, and a first noise representation are respectively extracted. On this basis, the first voice representation is further subjected to interaction processing with the first echo representation and the first noise representation respectively to obtain a second voice representation, a second echo representation, and a second noise representation, and the interaction processing includes: echo suppression, noise suppression, and voice enhancement. And based on at least one of the second voice representation, the second echo representation, and the second noise representation, the optimized target audio is obtained. Since the interaction processing between the first voice representation and the first echo representation is beneficial to suppressing echo and enhancing voice, and the interaction processing between the first voice representation and the first noise representation is beneficial to suppressing noise and enhancing voice, in the audio optimization process, it is possible to take into account the statistical characteristics of different signals, and adopting a parallel method to interactively process the first voice representation and the first echo representation, as well as the first voice representation and the first noise representation, is beneficial to improving the audio optimization effect.
[0120] In some disclosed embodiments, the voice representation interaction module 63 includes an interaction processing sub-module, which is configured to perform echo cancellation and voice enhancement respectively based on a first echo representation and a first voice representation to obtain a second echo representation and a first enhanced voice representation, and perform noise suppression and voice enhancement respectively based on a first noise representation and a first voice representation to obtain a second noise representation and a second enhanced voice representation. The voice representation interaction module 63 includes a representation fusion sub-module, which is configured to obtain a second voice representation based on the first enhanced voice representation and the second enhanced voice representation.
[0121] Therefore, performing echo cancellation and voice enhancement respectively based on the first echo representation and the first voice representation is conducive to the interaction between the residual voice feature information in the first echo representation and the residual echo feature information in the first voice representation, and performing noise suppression and voice enhancement based on the first noise representation and the first voice representation is conducive to the interaction between the residual voice feature information in the first noise representation and the residual noise feature information in the first voice representation. Therefore, it is beneficial to improve the quality of voice optimization.
[0122] In some disclosed embodiments, the interaction processing sub-module includes a reference weight acquisition unit, which is configured to obtain a reference weight of the reference audio representation based on the correlation between the target audio representation to be interacted and the reference audio representation to be interacted; the interaction processing sub-module includes an audio representation weighting unit, which is configured to weight the target audio representation and the reference audio representation based on the reference weight to obtain a weighted audio representation. Wherein, when the interaction processing is echo cancellation, the target audio representation is the first echo representation, the reference audio representation is the first voice representation, and the weighted audio representation is the second echo representation; when the interaction processing is noise suppression, the target audio representation is the first noise representation, the reference audio representation is the first voice representation, and the weighted audio representation is the second noise representation; when the interaction processing is voice enhancement, the target audio representation is the first voice representation, the reference audio representation is the first echo representation, the target audio representation is the first enhanced voice representation, or the reference audio representation is the first noise representation, and the target audio representation is the second enhanced voice representation.
[0123] Therefore, by obtaining the reference weight of the reference audio representation based on the correlation between the target audio representation to be interacted and the reference audio representation to be interacted, and weighting the target audio representation and the reference audio representation based on the reference weight to obtain a weighted audio representation, through the weighting process, the part of the reference audio representation related to the target audio representation, that is, the residual target audio feature information in the reference audio representation, can be obtained, so that the proportion of the target audio feature information in the weighted audio representation can be increased, and the accuracy of the weighted audio representation can be improved.
[0124] In some disclosed embodiments, the representation fusion sub-module includes a first fusion unit for concatenating the first enhanced speech representation and the second enhanced speech representation to obtain a second speech representation; or the representation fusion sub-module includes a second fusion unit for averaging the first enhanced speech representation and the second enhanced speech representation to obtain a second speech representation.
[0125] Therefore, by concatenating or averaging the first enhanced speech representation and the second enhanced speech representation to obtain a second speech representation, the complexity of speech processing can be reduced, which is beneficial to improving the audio optimization efficiency.
[0126] In some disclosed embodiments, the target audio is obtained through a preset number of optimization rounds; the audio optimization device 60 further includes an optimization stage detection module for detecting whether the number of optimization rounds in the current optimization stage is lower than the preset value, and the sound representation extraction module 62 is further configured to, when the optimization stage detection module detects that the number of optimization rounds in the current optimization stage is lower than the preset value, respectively extract features from the second echo representation, the second speech representation, and the second noise representation to obtain a new first echo representation, a new first speech representation, and a new first noise representation, and re-execute the steps of interactively processing the first speech representation with the first echo representation and the first noise representation respectively to obtain the second speech representation, the second echo representation, and the second noise representation and subsequent steps in combination with the sound representation interaction module 63.
[0127] Therefore, by further detecting whether the number of optimization rounds in the current optimization stage is lower than the preset value, and when it is lower than the preset value, respectively extracting features from the second echo representation, the second speech representation, and the second noise representation to obtain a new first echo representation, a new first speech representation, and a new first noise representation, and re-executing the above interactive processing and subsequent processes, echoes and noises can be further suppressed and the speech can be enhanced, which is beneficial to improving the audio optimization effect.
[0128] In some disclosed embodiments, the target audio acquisition module 64 is further configured to, when the optimization stage detection module detects that the number of optimization rounds in the current optimization stage is not lower than the preset value, detect that the number of optimization rounds in the current optimization stage is not lower than the preset value, and recover the target audio based on at least one of the second speech representation, the second echo representation, and the second noise representation.
[0129] Therefore, when the number of optimization rounds in the current optimization stage is not lower than the preset value, the target audio is recovered based on at least one of the second speech representation, the second echo representation, and the second noise representation. Therefore, during the audio optimization process, continuous multi-round optimization can be performed, which is beneficial to gradually improving the audio optimization quality.
[0130] In some disclosed embodiments, the target audio is obtained by processing with an audio optimization model, and the audio optimization model includes an echo branch network, a voice branch network, a noise branch network, a first interaction network located between the echo branch network and the voice branch network, and a second interaction network located between the noise branch network and the voice branch network.
[0131] Therefore, the target audio is obtained by processing with the audio optimization model, and the audio optimization model includes an echo branch network, a voice branch network, a noise branch network, a first interaction network located between the echo branch network and the voice branch network, and a second interaction network located between the noise branch network and the voice branch network. That is, in the audio optimization process, both the sound representation extraction and the sound representation interaction can be executed by the network model, so it is beneficial to improve the audio optimization efficiency.
[0132] In some disclosed embodiments, the echo branch network includes a preset number of sequentially connected echo sub-networks, the voice branch network includes a preset number of sequentially connected voice sub-networks, the noise branch network includes a preset number of sequentially connected noise sub-networks, the first interaction network includes a preset number of first interaction sub-networks, and the second interaction network includes a preset number of second interaction sub-networks; wherein, the i-th first interaction sub-network is located between the i-th echo sub-network and the second i-th voice sub-network, and the i-th second interaction sub-network is located between the i-th noise sub-network and the i-th voice sub-network.
[0133] Therefore, by correspondingly setting a preset number of sub-networks for each branch network, multiple rounds of optimization can be performed in the audio optimization process, which is beneficial to further improve the audio optimization quality while improving the audio optimization efficiency.
[0134] In some disclosed embodiments, the target audio is obtained by processing with an audio optimization model, and the audio optimization model is trained with several groups of sample audio data. Each group of sample audio data includes sample noise audio, sample reference audio, sample voice audio, and sample acquisition audio generated from the sample noise audio, sample reference audio, and sample voice audio.
[0135] Therefore, through the simulation generation method, as rich sample data as possible can be obtained, which can overcome the problem of difficult sample acquisition in real-time scenarios and greatly improve the model training performance.
[0136] In some disclosed embodiments, the audio optimization model is obtained through multiple rounds of iterative training. The audio optimization device 60 further includes an initial representation extraction module, which is configured to extract features from the sample noise audio, sample reference audio, sample speech audio, and sample acquisition audio respectively, to obtain an initial sample noise representation, an initial sample reference representation, an initial sample speech representation, and an initial sample audio representation; the audio optimization device 60 further includes a sample representation extraction module, which is configured to respectively extract a first sample echo representation, a first sample speech representation, and a first sample noise representation based on the initial sample audio representation and the initial sample reference representation; the audio optimization device 60 further includes a sample representation interaction module, which is configured to perform interaction processing on the first sample speech representation with the first sample echo representation and the first sample noise representation respectively, to obtain a second sample speech representation, a second sample echo representation, and a second sample noise representation; the audio optimization device 60 further includes an optimization loss calculation module, which is configured to obtain a speech optimization sub-loss based on the difference between the initial sample speech representation and the second sample speech representation, obtain a noise optimization sub-loss based on the difference between the initial sample noise representation and the second sample noise representation, and obtain an echo optimization sub-loss based on the difference between the initial sample reference representation and the second sample echo representation; the audio optimization device 60 further includes an optimization weight acquisition module, which is configured to obtain a speech optimization weight, a noise optimization weight, and an echo optimization weight based on the number of iterations of the current round of iteration; the audio optimization device 60 further includes a network parameter adjustment module, which is configured to adjust the network parameters of the audio optimization model based on the weighted results of the speech optimization sub-loss, the noise optimization sub-loss, and the echo optimization sub-loss by the speech optimization weight, the noise optimization weight, and the echo optimization weight.
[0137] Therefore, by obtaining the speech optimization weight, the noise optimization weight, and the echo optimization weight based on the number of iterations of the current round of iteration, and adjusting the network parameters of the audio optimization model based on the weighted results of the speech optimization sub-loss, the noise optimization sub-loss, and the echo optimization sub-loss by the speech optimization weight, the noise optimization weight, and the echo optimization weight, it is possible to focus on training different branch networks respectively during each round of iteration, which is beneficial to improving the training effect of the audio optimization model.
[0138] In some disclosed embodiments, when the number of iterations meets a preset condition, the echo optimization weight is higher than the noise optimization weight; and / or, when the number of iterations does not meet the preset condition, the noise optimization weight is higher than the echo optimization weight.
[0139] Therefore, each round of iteration can focus on training different branch networks respectively, which is beneficial to improving the training effect.
[0140] In some disclosed embodiments, the audio optimization device 60 further includes a sample audio generation module, which includes an impulse response acquisition sub-module for respectively acquiring a first room impulse response of noise, a second room impulse response of echo, and a third room impulse response of speech; the sample audio generation module includes an initial audio generation sub-module for convolving a sample noise audio, a sample reference audio, and a sample speech audio with the first room impulse response, the second room impulse response, and the third room impulse response respectively to obtain an initial acquired audio; the sample audio generation module includes a linear echo cancellation sub-module for performing linear echo cancellation on the initial acquired audio to obtain a sample acquired audio.
[0141] Therefore, through the simulation generation method, as rich sample data as possible can be obtained, which can overcome the problem of difficult sample acquisition in real-time scenarios and greatly improve the model training performance. Obtaining the sample acquired audio through linear echo cancellation can also greatly improve the quality of the sample audio data, which is beneficial to improving the model training performance.
[0142] Please refer to Figure 7 , Figure 7 which is a schematic framework diagram of an embodiment of the electronic device 70 of the present application. The electronic device 70 includes a mutually coupled memory 71 and a processor 72. Program instructions are stored in the memory 71, and the processor 72 is configured to execute the program instructions to implement the steps in any of the above audio optimization method embodiments. Specifically, the electronic device 70 may include, but is not limited to: a desktop computer, a laptop computer, a server, a mobile phone, a tablet computer, etc., which are not limited herein.
[0143] Specifically, the processor 72 is configured to control itself and the memory 71 to implement the steps in any of the above audio optimization method embodiments. The processor 72 may also be referred to as a CPU (Central Processing Unit). The processor 72 may be an integrated circuit chip with signal processing capabilities. The processor 72 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 72 may be implemented jointly by integrated circuit chips.
[0144] In the above solution, since the first voice representation and the first echo representation are interactively processed, it is beneficial to suppress echo and enhance voice, and since the first voice representation and the first noise representation are interactively processed, it is beneficial to suppress noise and enhance voice. Therefore, during the audio optimization process, the statistical characteristics of different signals can be taken into account, and the first voice representation and the first echo representation, as well as the first voice representation and the first noise representation, are interactively processed in parallel, which is beneficial to improving the audio optimization effect.
[0145] Please refer to Figure 8 , Figure 8 which is a schematic framework diagram of an embodiment of the computer-readable storage medium 80 of the present application. The computer-readable storage medium 80 stores program instructions 81 that can be run by a processor, and the program instructions 81 are used to implement the steps in any of the above-described audio optimization method embodiments.
[0146] In the above solution, since the first voice representation and the first echo representation are interactively processed, it is beneficial to suppress echo and enhance voice, and since the first voice representation and the first noise representation are interactively processed, it is beneficial to suppress noise and enhance voice. Therefore, during the audio optimization process, the statistical characteristics of different signals can be taken into account, and the first voice representation and the first echo representation, as well as the first voice representation and the first noise representation, are interactively processed in parallel, which is beneficial to improving the audio optimization effect.
[0147] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0148] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.
[0149] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation manners described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0150] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0151] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0152] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
Claims
1. An audio optimization method, characterized in that, Including: Extracting a first audio representation of the collected audio and extracting a second audio representation of the reference audio; Based on the first audio representation and the second audio representation, respectively extracting a first echo representation, a first speech representation, and a first noise representation; Performing interaction processing on the first speech representation with the first echo representation and the first noise representation respectively to obtain a second speech representation, a second echo representation, and a second noise representation; wherein, the interaction processing includes: echo suppression, noise suppression, and speech enhancement; Detecting whether the number of optimization rounds in the current optimization stage is lower than a preset value; If so, respectively performing feature extraction on the second echo representation, the second speech representation, and the second noise representation to obtain a new first echo representation, a new first speech representation, and a new first noise representation, and returning to the step of performing interaction processing on the first speech representation with the first echo representation and the first noise representation respectively to obtain a second speech representation, a second echo representation, and a second noise representation; If not, obtaining an optimized target audio based on at least one of the second speech representation, the second echo representation, and the second noise representation.
2. The method according to claim 1, characterized in that, The step of performing interaction processing on the first speech representation with the first echo representation and the first noise representation respectively to obtain a second speech representation, a second echo representation, and a second noise representation includes: Performing echo suppression and speech enhancement on the first echo representation and the first speech representation respectively to obtain the second echo representation and a first enhanced speech representation, and performing noise suppression and speech enhancement on the first noise representation and the first speech representation respectively to obtain the second noise representation and a second enhanced speech representation; Based on the first enhanced speech representation and the second enhanced speech representation, obtaining the second speech representation.
3. The method according to claim 2, characterized in that, The steps of the interaction processing include: Based on the correlation between the target audio representation to be interacted and the reference audio representation to be interacted, obtaining a reference weight of the reference audio representation; Based on the reference weight, weighting the target audio representation and the reference audio representation to obtain a weighted audio representation; Wherein, in the case where the interaction processing is echo suppression, the target audio representation is the first echo representation, the reference audio representation is the first speech representation, and the weighted audio representation is the second echo representation; in the case where the interaction processing is noise suppression, the target audio representation is the first noise representation, the reference audio representation is the first speech representation, and the weighted audio representation is the second noise representation; in the case where the interaction processing is speech enhancement, the target audio representation is the first speech representation, the reference audio representation is the first echo representation, the target audio representation is the first enhanced speech representation, or the reference audio representation is the first noise representation, and the target audio representation is the second enhanced speech representation.
4. The method according to claim 2, characterized in that, The step of obtaining the second speech representation based on the first enhanced speech representation and the second enhanced speech representation includes: Concatenate the first enhanced speech representation and the second enhanced speech representation to obtain the second speech representation; Alternatively, average the first enhanced speech representation and the second enhanced speech representation to obtain the second speech representation.
5. The method according to claim 1, characterized in that, The target audio is processed by an audio optimization model, which includes an echo branch network, a speech branch network, a noise branch network, a first interaction network between the echo branch network and the speech branch network, and a second interaction network between the noise branch network and the speech branch network.
6. The method according to claim 5, characterized in that, The echo branch network includes a preset number of sequentially connected echo sub-networks, the speech branch network includes the preset number of sequentially connected speech sub-networks, the noise branch network includes the preset number of sequentially connected noise sub-networks, the first interaction network includes the preset number of first interaction sub-networks, and the second interaction network includes the preset number of second interaction sub-networks; Among them, the i-th first interaction sub-network is located between the i-th echo sub-network and the second i-th speech sub-network, and the i-th second interaction sub-network is located between the i-th noise sub-network and the i-th speech sub-network.
7. The method according to claim 1, characterized in that, The target audio is processed by an audio optimization model, which is trained using several groups of sample audio data. Each group of sample audio data includes sample noisy audio, sample reference audio, sample speech audio, and sample acquisition audio generated from the sample noisy audio, the sample reference audio, and the sample speech audio.
8. The method according to claim 7, characterized in that, The audio optimization model is obtained through multiple rounds of iterative training. The training steps of the audio optimization model include: Extract features from the sample noisy audio, the sample reference audio, the sample speech audio, and the sample acquisition audio respectively to obtain an initial sample noise representation, an initial sample reference representation, an initial sample speech representation, and an initial sample audio representation; Based on the initial sample audio representation and the initial sample reference representation, a first sample echo representation, a first sample speech representation, and a first sample noise representation are respectively extracted; Perform interaction processing on the first sample speech representation with the first sample echo representation and the first sample noise representation respectively to obtain a second sample speech representation, a second sample echo representation, and a second sample noise representation; Based on the difference between the initial sample speech representation and the second sample speech representation, obtain a speech optimization sub-loss, and based on the difference between the initial sample noise representation and the second sample noise representation, obtain a noise optimization sub-loss, and based on the difference between the initial sample reference representation and the second sample echo representation, obtain an echo optimization sub-loss; Based on the iteration number of the current round of iteration, obtain a speech optimization weight, a noise optimization weight, and an echo optimization weight; Adjust the network parameters of the audio optimization model based on the weighted results of the voice optimization sub-loss, the noise optimization sub-loss, and the echo optimization sub-loss with respect to the voice optimization weight, the noise optimization weight, and the echo optimization weight.
9. The method according to claim 8, characterized in that, Obtaining the voice optimization weight, the noise optimization weight, and the echo optimization weight based on the number of iterations of the current round of iteration includes: When the number of iterations meets a preset condition, the echo optimization weight is higher than the noise optimization weight; And / or, when the number of iterations does not meet the preset condition, the noise optimization weight is higher than the echo optimization weight.
10. The method according to claim 7, wherein, The generation step of the sample acquisition audio includes: Obtain the first room impulse response of the noise, the second room impulse response of the echo, and the third room impulse response of the voice respectively; Convolve the sample noise audio, the sample reference audio, and the sample voice audio with the first room impulse response, the second room impulse response, and the third room impulse response respectively to obtain the initial acquisition audio; Perform linear echo cancellation on the initial acquisition audio to obtain the sample acquisition audio.
11. An audio optimization device, wherein, Including: An audio feature extraction module for extracting a first audio representation of the acquired audio and a second audio representation of the reference audio; A sound representation extraction module for respectively extracting a first echo representation, a first voice representation, and a first noise representation based on the first audio representation and the second audio representation; A sound representation interaction module for performing interaction processing on the first voice representation with the first echo representation and the first noise representation respectively to obtain a second voice representation, a second echo representation, and a second noise representation; wherein, the interaction processing includes: echo suppression, noise suppression, and voice enhancement; An optimization stage detection module for detecting whether the number of optimization rounds in the current optimization stage is lower than a preset value; wherein, when the number of optimization rounds in the current optimization stage is lower than the preset value, feature extraction is respectively performed on the second echo representation, the second voice representation, and the second noise representation to obtain a new first echo representation, a new first voice representation, and a new first noise representation, and return to the step of performing interaction processing on the first voice representation with the first echo representation and the first noise representation respectively to obtain a second voice representation, a second echo representation, and a second noise representation; A target audio acquisition module for obtaining the optimized target audio based on at least one of the second voice representation, the second echo representation, and the second noise representation when the number of optimization rounds in the current optimization stage is not lower than the preset value.
12. An electronic device, wherein, Including a mutually coupled memory and a processor, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the audio optimization method according to any one of claims 1 to 10.
13. A computer-readable storage medium, wherein, Store program instructions that can be run by a processor, and the program instructions are used to implement the audio optimization method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Echo cancellation and noise suppression calibration in telephony devices
US20080181392A1
System and method for performing speech enhancement using a deep neural network-based signal
US20180040333A1