Speech Enhancement Method, Apparatus and Device
Through the combination of the acoustic feature enhancement model and the vocoder, the noise classification loss adversarial multi-task learning is used to solve the problems of difficulty in speech enhancement, many distortions and poor generalization in the prior art, and high-quality speech enhancement and good generalization are achieved.
Patent Information
- Application Number
- CN202011187857.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-29
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-10-29
Smart Images

Figure CN114512140B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing technologies, and particularly to a speech enhancement method and apparatus, a speech recognition method, apparatus and system, a speech recognition text editing system, a speech enhancement model processing method and apparatus, an acoustic feature enhancement model processing method and apparatus, a user recognition method and apparatus, and an electronic device. Background Art
[0002] In the fields of machine recognition such as speech recognition and speaker recognition, noise will greatly affect the recognition accuracy. To improve the accuracy of speech recognition, speaker recognition, etc., the speech can be first separated from the background noise through single-channel speech enhancement technology, and then speech recognition, speaker recognition, etc. can be processed based on the enhanced speech data.
[0003] Currently, a typical speech enhancement scheme is to perform single-channel speech enhancement according to the energy spectrum and phase spectrum of the noisy speech, and this method will directly or indirectly enhance the phase spectrum of the noisy speech. Among them, the phase spectrum is a feature representation obtained after the speech signal undergoes short-time Fourier transform, and can restore the complete speech signal together with the amplitude spectrum. Usually, it is as random as noise and does not have a structure.
[0004] However, in the process of implementing the present invention, the inventors found that the above scheme has at least the following problems: Since the phase spectrum itself lacks a structure, if the phase spectrum of the noisy speech is directly or indirectly enhanced, it will make speech enhancement difficult, and will also cause distortion of the enhanced speech, as well as problems such as poor noise generalization and poor speaker generalization. Summary of the Invention
[0005] The present application provides a speech enhancement method to solve the problems of distortion of enhanced speech and poor noise generalization existing in the prior art. The present application also provides a speech enhancement apparatus, a speech recognition method, apparatus and system, a speech recognition text editing system, a speech enhancement model processing method and apparatus, an acoustic feature enhancement model processing method and apparatus, a user recognition method and apparatus, and an electronic device.
[0006] The present application provides a speech enhancement method, including:
[0007] Determine the acoustic feature data of the first noisy speech data to be processed;
[0008] Through an acoustic feature enhancement model, determine the enhanced acoustic features of the first noisy speech data according to the acoustic feature data; wherein, the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method;
[0009] Generate enhanced speech data of the first noisy speech data according to the enhanced acoustic features through a vocoder.
[0010] Optionally, there are different types of environmental noises between the first noisy speech data and the training data of the model; the enhanced acoustic features include enhanced acoustic features that suppress environmental noises not present in the training data.
[0011] Optionally, the determining of the enhanced acoustic features of the first noisy speech data according to the acoustic feature data through the acoustic feature enhancement model includes:
[0012] Determine the noise-independent acoustic feature coding data of the first noisy speech data according to the acoustic feature data through the encoder included in the model;
[0013] Determine the enhanced acoustic features of the first noisy speech data according to the acoustic feature coding data through the decoder included in the model.
[0014] Optionally, the acoustic feature enhancement model is obtained through adversarial multi-task learning with noise classification loss, including:
[0015] During the process of training the acoustic feature enhancement model, the training data includes the noise types of the second noisy speech data; the acoustic feature enhancement model further includes: a noise classifier; the noise classifier is used to determine the noise types of the second noisy speech data according to the acoustic feature coding data of the second noisy speech data; the training objective of the acoustic feature enhancement model includes: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss.
[0016] Optionally, it further includes:
[0017] Generate second noisy speech data according to clean speech data and noise data.
[0018] Optionally, the vocoder is learned from the correspondence set between the acoustic features and the clean speech data of multiple users' clean speech data.
[0019] Optionally, the vocoder includes: a vocoder based on a waveform recurrent neural network.
[0020] Optionally, the acoustic feature data includes: complex spectrum;
[0021] The enhanced acoustic features include: Mel spectrum.
[0022] This application also provides a speech enhancement method, including:
[0023] A vocoder is learned from the correspondence sets between the acoustic features and the clean speech data of the clean speech data of multiple users;
[0024] Through an acoustic feature enhancement model, according to the acoustic feature data of the noisy speech data, the enhanced acoustic features of the noisy speech data are determined;
[0025] Through the vocoder, according to the enhanced acoustic features, enhanced speech data of the noisy speech data is generated.
[0026] The present application also provides a method for processing a voice enhancement model, including:
[0027] Determine a first training data set and a second training data set. The first training data includes the acoustic feature data of the noisy speech data, the noise type, and the correspondence between the acoustic features of the clean speech data; the second training data includes the correspondence set between the acoustic features of the clean speech data and the clean speech data;
[0028] Construct the network structure of the voice denoising model; the voice denoising model includes an acoustic feature enhancement model and a vocoder; the acoustic feature enhancement model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the acoustic feature coding data independent of noise of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature coding data; the noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature coding data; the vocoder is used to generate enhanced speech data of the noisy speech data according to the enhanced acoustic features;
[0029] Through the noise classification loss adversarial multi-task learning method, according to the first training data set, train the network parameters of the acoustic feature enhancement model; and, according to the second training data set, train the network parameters of the vocoder.
[0030] The present application also provides a method for processing an acoustic feature enhancement model, including:
[0031] Determine a training data set, and the training data includes the acoustic feature data of the noisy speech data, the noise type, and the correspondence between the acoustic features of the clean speech data;
[0032] Construct the network structure of the acoustic feature enhancement model; the model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the acoustic feature coding data independent of noise of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature coding data; the noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature coding data;
[0033] Train the network parameters of the model according to the training data set by means of adversarial multi-task learning with noise classification loss.
[0034] This application also provides a speech recognition system, including:
[0035] A client for collecting speech data and sending the speech data to a server;
[0036] A server for determining acoustic feature data of the speech data; determining enhanced acoustic features of the speech data according to the acoustic feature data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained by means of adversarial multi-task learning with noise classification loss; generating enhanced speech data of the speech data through a vocoder; and converting the enhanced speech data into text through a speech recognition model.
[0037] This application also provides a speech recognition method, including:
[0038] Determine acoustic feature data of noisy speech data to be processed;
[0039] Determine enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained by means of adversarial multi-task learning with noise classification loss;
[0040] Generate enhanced speech data of the noisy speech data through a vocoder;
[0041] Convert the enhanced speech data into text through a speech recognition model.
[0042] This application also provides a speech recognition method, including:
[0043] Collect speech data;
[0044] Send the speech data to a server so that the server determines acoustic feature data of the speech data; determine enhanced acoustic features of the speech data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained by means of adversarial multi-task learning with noise classification loss; generate enhanced speech data of the speech data through a vocoder; and convert the enhanced speech data into text through a speech recognition model.
[0045] This application also provides a speech recognition text editing system, including:
[0046] A client for collecting speech data, sending the speech data to a server; and editing the text of the speech data recognized by the server.
[0047] A server, configured to determine the acoustic feature data of the speech data; determine the enhanced acoustic features of the speech data through an acoustic feature enhancement model, where the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; generate enhanced speech data of the speech data through a vocoder according to the enhanced acoustic features; and convert the enhanced speech data into text through a speech recognition model.
[0048] This application also provides a user identification method, including:
[0049] Determine the acoustic feature data of the noisy speech data to be processed;
[0050] Determine the enhanced acoustic features of the noisy speech data through an acoustic feature enhancement model according to the acoustic feature data, where the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method;
[0051] Generate enhanced speech data of the noisy speech data through a vocoder according to the enhanced acoustic features;
[0052] Determine the user information of the enhanced speech data through a user identification model.
[0053] This application also provides a speech enhancement device, including:
[0054] An acoustic feature extraction unit, configured to determine the acoustic feature data of the first noisy speech data to be processed;
[0055] An acoustic feature enhancement unit, configured to determine the enhanced acoustic features of the first noisy speech data through an acoustic feature enhancement model according to the acoustic feature data, where the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method;
[0056] A speech synthesis unit, configured to generate enhanced speech data of the first noisy speech data through a vocoder according to the enhanced acoustic features.
[0057] This application also provides an electronic device, including:
[0058] A processor and a memory;
[0059] A memory for storing a program for implementing a voice enhancement method. After the device is powered on and runs the program of the method through the processor, the following steps are executed: determining acoustic feature data of first noisy voice data to be processed; determining enhanced acoustic features of the first noisy voice data according to the acoustic feature data through an acoustic feature enhancement model, wherein the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss; generating enhanced voice data of the first noisy voice data through a vocoder according to the enhanced acoustic features.
[0060] This application also provides a voice enhancement device, including:
[0061] A vocoder construction unit for learning a vocoder from a correspondence set between acoustic features of clean voice data of multiple users and the clean voice data;
[0062] An acoustic feature enhancement unit for determining enhanced acoustic features of noisy voice data according to the acoustic feature data of the noisy voice data through an acoustic feature enhancement model;
[0063] A voice synthesis unit for generating enhanced voice data of the noisy voice data through a vocoder according to the enhanced acoustic features.
[0064] This application also provides an electronic device, including:
[0065] A processor and a memory;
[0066] A memory for storing a program for implementing a voice enhancement method. After the device is powered on and runs the program of the method through the processor, the following steps are executed: learning a vocoder from a correspondence set between acoustic features of clean voice data of multiple users and the clean voice data; determining enhanced acoustic features of noisy voice data according to the acoustic feature data of the noisy voice data through an acoustic feature enhancement model; generating enhanced voice data of the noisy voice data through a vocoder according to the enhanced acoustic features.
[0067] This application also provides a voice enhancement model processing device, including:
[0068] A training data determination unit for determining a first training data set and a second training data set, where the first training data includes the correspondence between the acoustic feature data and the noise type of the noisy voice data and the acoustic features of the clean voice data; the second training data includes a correspondence set between the acoustic features of the clean voice data and the clean voice data;
[0069] A model structure construction unit for constructing the network structure of a voice noise reduction model; the voice noise reduction model includes an acoustic feature enhancement model and a vocoder; the acoustic feature enhancement model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the noise-independent acoustic feature encoding data of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature encoding data; the vocoder is used to generate enhanced speech data of the noisy speech data according to the enhanced acoustic features;
[0070] A model parameter training unit for training the network parameters of the acoustic feature enhancement model according to the first training data set by means of adversarial multi-task learning with noise classification loss; and training the network parameters of the vocoder according to the second training data set.
[0071] This application also provides an electronic device, including:
[0072] A processor and a memory;
[0073] The memory is used to store a program for implementing the voice enhancement model construction method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining a first training data set and a second training data set, where the first training data includes the acoustic feature data and noise type of the noisy speech data, and the corresponding relationship between the acoustic features of the clean speech data; the second training data includes the set of corresponding relationships between the acoustic features of the clean speech data and the clean speech data; constructing the network structure of the voice noise reduction model; the voice noise reduction model includes an acoustic feature enhancement model and a vocoder; the acoustic feature enhancement model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the noise-independent acoustic feature encoding data of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature encoding data; the vocoder is used to generate enhanced speech data of the noisy speech data according to the enhanced acoustic features; training the network parameters of the acoustic feature enhancement model according to the first training data set by means of adversarial multi-task learning with noise classification loss; and training the network parameters of the vocoder according to the second training data set.
[0074] This application also provides an acoustic feature enhancement model processing device, including:
[0075] A training data determination unit for determining a training data set, where the training data includes acoustic feature data of noisy speech data and the noise type, and the corresponding relationship with the acoustic features of clean speech data;
[0076] A model structure construction unit for constructing the network structure of an acoustic feature enhancement model; the model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the noise-independent acoustic feature encoding data of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature encoding data;
[0077] A model parameter training unit for training the network parameters of the model according to the training data set by means of adversarial multi-task learning with noise classification loss.
[0078] This application also provides an electronic device, including:
[0079] A processor and a memory;
[0080] The memory is used to store a program for implementing the method for constructing an acoustic feature enhancement model. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining a training data set, where the training data includes acoustic feature data of noisy speech data and the noise type, and the corresponding relationship with the acoustic features of clean speech data; constructing the network structure of an acoustic feature enhancement model; the model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the noise-independent acoustic feature encoding data of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature encoding data; training the network parameters of the model according to the training data set by means of adversarial multi-task learning with noise classification loss.
[0081] This application also provides a speech recognition device, including:
[0082] An acoustic feature extraction unit for determining the acoustic feature data of the noisy speech data to be processed;
[0083] An acoustic feature enhancement unit for determining the enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained by means of adversarial multi-task learning with noise classification loss;
[0084] A speech synthesis unit, configured to generate enhanced speech data of noisy speech data according to enhanced acoustic features through a vocoder;
[0085] A speech conversion unit, configured to convert the enhanced speech data into text through a speech recognition model.
[0086] This application also provides an electronic device, including:
[0087] A processor and a memory;
[0088] The memory is configured to store a program for implementing a speech recognition method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining acoustic feature data of to-be-processed noisy speech data; determining enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained through an adversarial multi-task learning manner of noise classification loss; generating enhanced speech data of the noisy speech data according to the enhanced acoustic features through a vocoder; converting the enhanced speech data into text through a speech recognition model.
[0089] This application also provides a speech recognition device, including:
[0090] A speech data acquisition unit, configured to acquire speech data;
[0091] A speech data sending unit, configured to send the speech data to a server, so that the server determines acoustic feature data of the speech data; determining enhanced acoustic features of the speech data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained through an adversarial multi-task learning manner of noise classification loss; generating enhanced speech data of the speech data according to the enhanced acoustic features through a vocoder; converting the enhanced speech data into text through a speech recognition model.
[0092] This application also provides an electronic device, including:
[0093] A processor and a memory;
[0094] The memory is configured to store a program for implementing a speech recognition method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: acquiring speech data; sending the speech data to a server, so that the server determines acoustic feature data of the speech data; determining enhanced acoustic features of the speech data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained through an adversarial multi-task learning manner of noise classification loss; generating enhanced speech data of the speech data according to the enhanced acoustic features through a vocoder; converting the enhanced speech data into text through a speech recognition model.
[0095] The present application also provides a user identification device, including:
[0096] An acoustic feature determination unit, configured to determine acoustic feature data of noisy speech data to be processed;
[0097] An acoustic feature enhancement unit, configured to determine enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss;
[0098] A speech synthesis unit, configured to generate enhanced speech data of the noisy speech data through a vocoder according to the enhanced acoustic features;
[0099] A user determination unit, configured to determine user information of the enhanced speech data through a user identification model.
[0100] The present application also provides an electronic device, including:
[0101] A processor and a memory;
[0102] The memory is configured to store a program for implementing the user identification method. After the device is powered on and runs the program of the method through the processor, the following steps are executed: determining acoustic feature data of noisy speech data to be processed; determining enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss; generating enhanced speech data of the noisy speech data through a vocoder according to the enhanced acoustic features; determining user information of the enhanced speech data through a user identification model.
[0103] The present application also provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a computer, the computer is enabled to execute the above various methods.
[0104] The present application also provides a computer program product including instructions. When the computer program product runs on a computer, the computer is enabled to execute the above various methods.
[0105] Compared with the prior art, the present application has the following advantages:
[0106] The voice enhancement method provided by the embodiment of the present application determines the acoustic feature data of the first noisy voice data to be processed; through an acoustic feature enhancement model, determines the enhanced acoustic features of the first noisy voice data according to the acoustic feature data, wherein the acoustic feature enhancement model is obtained through an adversarial multi-task learning method of noise classification loss; through a vocoder, generates enhanced voice data of the first noisy voice data according to the enhanced acoustic features; this processing method enables the acoustic feature enhancement model to be obtained through an adversarial multi-task learning method based on self-supervised noise classification loss, and determines the enhanced acoustic features of the noisy voice through this model, which can avoid being sensitive to environmental noise when extracting enhanced acoustic features; therefore, it can effectively narrow the difference in voice enhancement performance among various environmental noises and improve the generalization of noises outside the training set. In addition, since this processing method performs voice synthesis on the enhanced acoustic features through a vocoder to obtain the enhanced voice of the noisy voice, avoiding directly or indirectly enhancing the phase spectrum of the noisy voice; therefore, it can effectively reduce voice distortion and improve the auditory quality of the voice.
[0107] The voice enhancement method provided by the embodiment of the present application learns a vocoder from the correspondence set between the acoustic features and the clean voice data of the clean voice data of multiple users; through an acoustic feature enhancement model, determines the enhanced acoustic features of the noisy voice data according to the acoustic feature data of the noisy voice data; through a vocoder, generates enhanced voice data of the noisy voice data according to the enhanced acoustic features; this processing method enables the use of the temporal characteristics of the voice waveform and combines a large amount of speaker training data to construct a vocoder, and performs voice enhancement on the noisy voice based on this vocoder, which can overcome the speaker generalization problem; therefore, it can effectively improve the speaker generalization degree. At the same time, this processing method can avoid directly or indirectly enhancing the phase spectrum of the noisy voice; therefore, it can effectively reduce voice distortion and improve the auditory quality of the voice.
[0108] The voice enhancement model construction method provided by the embodiment of the present application obtains an acoustic feature enhancement model through an adversarial multi-task learning method of self-supervised noise classification loss, and determines the enhanced acoustic features of the noisy voice through this model, which can avoid being sensitive to environmental noise when extracting enhanced acoustic features; therefore, it can effectively narrow the difference in voice enhancement performance among various environmental noises and improve the generalization of noises outside the training set. In addition, since this processing method performs voice synthesis on the enhanced acoustic features through a vocoder to obtain the enhanced voice of the noisy voice, avoiding directly or indirectly enhancing the phase spectrum of the noisy voice; therefore, it can effectively reduce voice distortion and improve the auditory quality of the voice.
[0109] The acoustic feature enhancement model processing method provided by the embodiments of this application obtains an acoustic feature enhancement model through an adversarial multi-task learning method with a self-supervised noise classification loss. By using this model to determine the enhanced acoustic features of noisy speech, it is possible to avoid being sensitive to environmental noise when extracting enhanced acoustic features. Therefore, it can effectively reduce the difference in speech enhancement performance among various environmental noises and improve the generalization ability of noises outside the training set.
[0110] The speech recognition system provided by the embodiments of this application collects speech data through a client and sends the speech data to a server. The server determines the acoustic feature data of the speech data. Through an acoustic feature enhancement model, based on the acoustic feature data, it determines the enhanced acoustic features of the speech data. The acoustic feature enhancement model is obtained through an adversarial multi-task learning method with a noise classification loss. Through a vocoder, based on the enhanced acoustic features, it generates the enhanced speech data of the speech data. Through a speech recognition model, it converts the enhanced speech data into text. This processing method enables an acoustic feature enhancement model to be obtained through an adversarial multi-task learning method with a self-supervised noise classification loss. By using this model to determine the enhanced acoustic features of noisy speech, it is possible to avoid being sensitive to environmental noise when extracting enhanced acoustic features, and then perform speech recognition processing on the enhanced speech. Therefore, it can effectively reduce the difference in speech recognition performance among various environmental noises and improve the generalization ability of noises outside the training set. In addition, since this processing method performs speech synthesis on the enhanced acoustic features through a vocoder to obtain the enhanced speech of the noisy speech, avoiding directly or indirectly enhancing the phase spectrum of the noisy speech. Therefore, it can effectively reduce speech distortion, improve the listening quality of speech, and thus improve the accuracy of speech recognition.
[0111] The voice recognition text editing system provided by the embodiment of the present application collects voice data through a client and sends the voice data to a server; the server determines the acoustic feature data of the voice data; through an acoustic feature enhancement model, according to the acoustic feature data, determines the enhanced acoustic features of the voice data; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, according to the enhanced acoustic features, generates the enhanced voice data of the voice data; through a voice recognition model, converts the enhanced voice data into text; the client edits the text; this processing method enables an acoustic feature enhancement model to be obtained through a self-supervised noise classification loss adversarial multi-task learning method, and determines the enhanced acoustic features of the noisy voice through this model, which can avoid being sensitive to environmental noise when extracting the enhanced acoustic features, and then performs voice recognition processing on the enhanced voice; therefore, it can effectively narrow the difference in voice recognition performance among various environmental noises and improve the generalization ability of noises outside the training set. In addition, since this processing method performs voice synthesis on the enhanced acoustic features through a vocoder to obtain the enhanced voice of the noisy voice, avoiding directly or indirectly enhancing the phase spectrum of the noisy voice; therefore, it can effectively reduce voice distortion, improve the listening quality of the voice, thereby improving the accuracy of voice recognition, and further improving the efficiency of voice recognition text editing.
[0112] The user recognition method provided by the embodiment of the present application determines the acoustic feature data of the to-be-processed noisy voice data; through an acoustic feature enhancement model, according to the acoustic feature data, determines the enhanced acoustic features of the noisy voice data; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, according to the enhanced acoustic features, generates the enhanced voice data of the noisy voice data; through a user recognition model, determines the user information of the enhanced voice data; this processing method enables an acoustic feature enhancement model to be obtained through a self-supervised noise classification loss adversarial multi-task learning method, and determines the enhanced acoustic features of the noisy voice through this model, which can avoid being sensitive to environmental noise when extracting the enhanced acoustic features, and then performs user recognition processing according to the enhanced voice; therefore, it can effectively narrow the difference in speaker recognition performance of voices under various environmental noises and improve the generalization ability of noises outside the training set. In addition, since this processing method performs voice synthesis on the enhanced acoustic features through a vocoder to obtain the enhanced voice of the noisy voice, avoiding directly or indirectly enhancing the phase spectrum of the noisy voice; therefore, it can effectively reduce voice distortion, improve the listening quality of the voice, thereby improving the accuracy of speaker recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] Figure 1 The flowchart of an embodiment of a voice enhancement method provided by the present application;
[0114] Figure 2 Schematic diagram of an application scenario of an embodiment of a voice enhancement method provided by this application;
[0115] Figure 3 Schematic diagram of a model of an embodiment of a voice enhancement method provided by this application. Detailed implementation manners
[0116] Many specific details are set forth in the following description in order to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this application. Therefore, this application is not limited by the specific implementations disclosed below.
[0117] In this application, there are provided a voice enhancement method and apparatus, a voice recognition method, apparatus and system, a voice recognition text editing system, a voice enhancement model processing method and apparatus, an acoustic feature enhancement model processing method and apparatus, a user recognition method and apparatus, and an electronic device. Each of the various solutions will be described in detail in the following embodiments.
[0118] First Embodiment
[0119] Please refer to Figure 1 , which is a flowchart of an embodiment of the voice enhancement method of this application. The execution subject of this method is a voice enhancement apparatus, which is usually deployed on a server, but is not limited to the server, and can also be any device capable of implementing the voice enhancement method. In this embodiment, the method may include the following steps:
[0120] Step S101: Determine the acoustic feature data of the first noisy voice data to be processed.
[0121] The noisy voice data may be single-channel voice data, which can be collected by a microphone. This method separates the voice from background noise (environmental noise) and can be applied in various voice processing systems, such as a voice recognition system, a speaker recognition system, a voice recognition text editing system, etc.
[0122] Please refer to Figure 2, which is a schematic diagram of the usage scenario of the embodiment of the voice enhancement method of the present application. In this embodiment, the method is applied to a speech recognition text editing system. The system includes a server and a client. The server deploys a voice enhancement device, and the user voice data is collected through the client. Due to the presence of environmental noise, the voice data is noisy voice data; the client sends the noisy voice data to the server, and the server converts the voice data into text; correspondingly, the server executes the method, performs voice enhancement processing on the noisy voice data through a voice enhancement model, that is, suppresses environmental noise, and then performs voice recognition processing on the enhanced voice through a voice recognition model, and sends the recognized text back to the client for the user to view and edit the text.
[0123] Since the first noisy voice data contains environmental noise, the acoustic feature data of this voice data is noisy acoustic feature data. The acoustic feature data can be time-frequency feature data, that is, the representation obtained by decomposing the voice signal in time and frequency. The time-frequency feature data includes but is not limited to: complex spectrum, energy spectrum, phase spectrum, Mel spectrum, and so on. The complex spectrum can be the time-frequency feature data obtained by performing short-time Fourier transform on the voice waveform. The energy spectrum, also known as the power spectrum, is the time-frequency feature data obtained after taking the modulus and squaring of the complex spectrum. The phase spectrum is the time-frequency feature data formed by taking the angle of each complex number in the complex spectrum. The Mel spectrum (Mel spectrum), also known as the Mel energy spectrum, is the time-frequency feature data obtained by filtering the energy spectrum through a Mel filter bank. Since the complex spectrum includes richer voice features, the acoustic feature data adopted in this embodiment is the complex spectrum, which will make the enhanced acoustic features more accurate and thus make the enhanced voice purer.
[0124] In specific implementation, an acoustic feature extraction algorithm can be used to extract the acoustic feature data of the first noisy voice data. Since the acoustic feature extraction algorithm is a relatively mature existing technology, it will not be elaborated here.
[0125] After determining the acoustic feature data of the first noisy voice data, the next step can be entered. Through the acoustic feature enhancement model, according to the acoustic feature data, the enhanced acoustic features of the first noisy voice data are determined.
[0126] Step S103: Through the acoustic feature enhancement model, according to the acoustic feature data, determine the enhanced acoustic features of the first noisy voice data.
[0127] The acoustic feature enhancement model is a model that can reconstruct the acoustic feature data of noisy speech data into pure speech features (i.e., enhanced acoustic features). In this embodiment, the acoustic feature enhancement model can be obtained through adversarial multi-task learning with noise classification loss, which can reduce the influence of different types of environmental noise on the acoustic feature enhancement effect. It can be seen that this model is insensitive to the type of environmental noise.
[0128] The types of the environmental noise include but are not limited to: automobile environmental noise, dining environmental noise, subway environmental noise, and so on. By adopting this acoustic feature enhancement method that is insensitive to the type of environmental noise, the difference in speech enhancement performance among various environmental noises is small. That is to say, no matter which type of environmental noise the first noisy speech data includes, a relatively high speech enhancement performance can be achieved. For example, the suppression effects of the model on automobile environmental noise, dining environmental noise, and subway environmental noise are roughly the same.
[0129] In an example, there are different types of environmental noises between the first noisy speech data and the training data of the model; the enhanced acoustic features include the enhanced acoustic features that suppress the environmental noise of the type not appearing in the training data. For example, the noisy speech in the model training data is usually automobile environmental noise and subway environmental noise, while the first noisy speech data contains dining environmental noise. In this case, the dining environmental noise can still be well suppressed through the model. Thus, it can be seen that even if the first noisy speech data includes the environmental noise not appearing in the model training data, a relatively high speech enhancement performance can still be obtained. That is to say, this acoustic feature enhancement method that is insensitive to the type of environmental noise can also effectively improve the generalization ability of the noise outside the training set.
[0130] To clearly illustrate the technical effects that can be achieved by adopting the model, the training method and training process of the model will be described below.
[0131] In this embodiment, the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss, and can be implemented as follows: During the process of training the acoustic feature enhancement model, the training data includes the acoustic feature data and noise type of the second noisy speech data, and the acoustic features of the clean speech data (without noise); the acoustic feature enhancement model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the noise-independent acoustic feature encoding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used to determine the enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss.
[0132] As Figure 3 shown, in this embodiment, the acoustic feature data of the second noisy speech data for training is used as the input data of the encoder, the acoustic feature hidden layer encoding data of the second noisy speech data output by the encoder is used as the input data of the noise classifier and the decoder, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder. Thus, it can be seen that the acoustic feature enhancement model includes two learning tasks, one is the noise classification task, and the other is the acoustic feature enhancement task. Among them, the optimization objective of the noise classification task is to minimize the noise classification loss of the noise classifier and maximize the noise classification loss of the encoder, and the optimization objective of the acoustic feature enhancement task is to minimize the enhanced acoustic feature loss. According to the training data set, the parameters in the encoder, decoder, and noise classifier are trained until the above three optimization objectives are achieved.
[0133] As Figure 3 can be seen, the minimization of the noise classification loss (L CE ) of the noise classifier means that when adjusting the parameters of the noise classifier through the gradient descent algorithm, the parameter adjustment methods that can be adopted include where ψ represents the parameters of the noise classifier; the maximization of the noise classification loss of the encoder means that when adjusting the parameters of the encoder through the gradient descent algorithm, the parameter adjustment methods that can be adopted include Among them, 0 represents the parameters of the encoder. In the method provided in this embodiment, by minimizing the noise classification loss of the noise classifier and maximizing the noise classification loss of the encoder, an adversarial multi-task learning method for the noise classification loss is realized. In this way, the influence of different types of environmental noise on the intermediate representation (acoustic feature hidden layer encoded data) output by the encoder can be reduced, so that the acoustic feature hidden layer encoded data output by the encoder is independent of the noise. The minimization of the enhanced acoustic feature loss (L MAE ) means that when adjusting the parameters of the decoder by the gradient descent algorithm, the parameter adjustment methods that can be adopted include Among them represents the parameters of the decoder; at the same time, when adjusting the parameters of the encoder, the parameter adjustment methods that can be adopted include In the method provided in this embodiment, when realizing the adversarial multi-task learning method for the noise classification loss, the enhanced acoustic features are also output by the decoder by minimizing the enhanced acoustic feature loss.
[0134] Among them, the noise classification loss (L CE ) may include: during a single training process, the cumulative value of the differences between the predicted values and the annotation information of the noise types of all noisy speech data. The model training objective includes making this cumulative value less than the noise classification loss threshold. If this loss value is greater than the noise classification loss threshold, further training is required until it is less than the noise classification loss threshold. The enhanced acoustic feature loss (L MAE ) may include: during a single training process, the cumulative value of the differences between the enhanced acoustic features of all noisy speech data output by the decoder and the acoustic features of clean speech data. The model training objective includes making this cumulative value less than the enhanced acoustic feature loss threshold. If this loss value is greater than the enhanced acoustic feature loss threshold, further training is required until it is less than the enhanced acoustic feature loss threshold.
[0135] Specifically, the acoustic feature enhancement model can be a deep neural network model or a model with other structures, as long as it can reconstruct the acoustic feature data of noisy speech data into enhanced acoustic features (clean speech features).
[0136] The noise type can be low-frequency noise, high-frequency noise or other noises. In one example, second noisy speech data is generated according to clean speech data and noise data. Specifically, second noisy speech data can be generated according to the noise data of multiple noise types in combination with clean speech data.
[0137] In one example, the enhanced acoustic feature may be a Mel spectrum. The Mel spectrum is time-frequency feature data obtained by filtering an energy spectrum through a Mel filter bank, which only includes energy information and has a lower data dimension. Therefore, it can not only obtain higher speech enhancement performance but also effectively improve the speech enhancement efficiency.
[0138] To train the acoustic feature enhancement model, training data needs to be prepared first. Specifically, during implementation, the Mel spectrum of clean speech (speech without noise) (denoted as S) and the Mel spectrum of noisy speech (denoted as X) can be extracted separately, and the Mel spectrum of clean speech can be appropriately scaled to transform the Mel spectrum into data between 0 and 1 to achieve Mel spectrum normalization, which can better train the model. Among them, S is used as the acoustic feature of clean speech, and X is used as the acoustic feature data of the second noisy speech data. In addition, the noise type of the noise signal in the second noisy speech needs to be determined. Specifically, during implementation, according to the energy distribution of the noise signal in the second noisy speech in different frequency bands, it can be divided into three categories: low-frequency noise, high-frequency noise, and other noise. In this way, the training data is prepared, including: the acoustic feature data X of the second noisy speech data, the noise type, and the acoustic feature S of the clean speech data.
[0139] After the training data is prepared, the model parameters can be trained according to the training data. Specifically, during implementation, X can be input into the encoder, and the encoder can extract the acoustic feature encoding data independent of noise (denoted as R); then R is input into the decoder, and the decoder predicts the Mel spectrum of clean speech, and the predicted value is Y; the parameters of the encoder and decoder are trained to minimize the mean square error (MSE) between S and Y. The method provided in the embodiments of the present application adopts a self-supervised adversarial multi-task training method. The model training also includes: inputting the intermediate representation R obtained by the encoder into the noise classifier to predict the noise type; training the noise classifier so that it can correctly classify the intermediate representation R, that is, minimizing the noise classification loss of the noise classifier; at the same time, training the encoder so that the noise classifier cannot correctly classify the noise type according to the intermediate representation R, that is, maximizing the noise classification loss of the encoder, so that the encoder outputs the acoustic feature encoding data independent of noise; the mean square error MSE and the noise classification loss can be combined and used together to optimize the encoder and decoder included in the acoustic feature enhancement model.
[0140] So far, the training method and process of the model have been described in detail.
[0141] Correspondingly, step S103 may include the following sub-steps: 1) Through the encoder included in the model, according to the acoustic feature data, determine the noise-independent acoustic feature coding data of the first noisy speech data; 2) Through the decoder included in the model, according to the acoustic feature coding data, determine the enhanced acoustic features of the first noisy speech data. The processing methods of the encoder and the decoder are the same in the model usage stage and the model training stage, so they will not be elaborated here.
[0142] It should be noted that in the model training stage, it not only includes the processing of the encoder and the decoder, but also includes the processing of the noise classifier, that is, the noise classifier also needs to be trained; while in the model usage stage, it does not include the processing of the noise classifier. As long as the processing of the encoder and the decoder is performed, enhanced acoustic features that are insensitive to the type of environmental noise can be obtained.
[0143] Step S105: Through a vocoder, according to the enhanced acoustic features, generate enhanced speech data of the first noisy speech data.
[0144] This method uses a vocoder to synthesize a speech waveform from the predicted acoustic features. By inputting the enhanced acoustic features output by the acoustic feature enhancement model into the vocoder, a synthesized speech waveform can be obtained.
[0145] The vocoder is used to synthesize a speech waveform from speech features and can be learned from the correspondence set between the acoustic features of clean speech data and the clean speech data. Since the vocoder belongs to a relatively mature existing technology, its learning method will not be elaborated here.
[0146] Specifically, the vocoder can adopt various vocoders in the prior art, such as the vocoder based on the waveform recurrent neural network (WaveRNN), or the vocoder based on the linear prediction method (LPCNet), the vocoder based on the streaming waveform method (Flowavenet), and so on. The Flowavenet vocoder is suitable for the enhancement of specific speakers and performs poorly in the enhancement tasks of non-specific speakers, with poor speaker generalization. The WaveRNN vocoder is a vocoder based on the recurrent neural network structure and is suitable for the enhancement tasks of non-specific speakers, with significantly improved speaker generalization.
[0147] In one example, the enhanced acoustic features may be Mel spectrograms. The Mel spectrogram is the time-frequency feature data obtained by filtering the energy spectrum through a Mel filter bank, which only includes energy information and has a low data dimension. Therefore, it can not only obtain high speech enhancement performance but also effectively improve the speech enhancement efficiency.
[0148] In this embodiment, the method may further include the following steps: learning a vocoder from a correspondence set between the acoustic features of the clean speech data of multiple users and the clean speech data. The vocoder may be a WaveRNN vocoder. Specifically in implementation, the Mel spectrum of the clean speech (denoted as S) may be input into WaveRNN to predict the clean speech waveform corresponding to S, and WaveRNN may be trained to minimize the error between the predicted value and the true value. In this way, the temporal characteristics of the speech waveform can be utilized by using WaveRNN, and at the same time, the speaker generalization problem can be overcome by combining large-scale speaker training data.
[0149] As can be seen from the above embodiments, the speech recognition system provided by the embodiments of the present application collects speech data through a client and sends the speech data to a server; the server determines the acoustic feature data of the speech data; through an acoustic feature enhancement model, based on the acoustic feature data, determines the enhanced acoustic features of the speech data; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, based on the enhanced acoustic features, generates the enhanced speech data of the speech data; through a speech recognition model, converts the enhanced speech data into text; this processing method enables an acoustic feature enhancement model to be obtained through a self-supervised noise classification loss adversarial multi-task learning method, and determines the enhanced acoustic features of the noisy speech through this model, so that it is possible to avoid being sensitive to environmental noise when extracting enhanced acoustic features, and then perform speech recognition processing on the enhanced speech; therefore, the difference in speech recognition performance among various environmental noises can be effectively reduced, and the generalization ability to noises outside the training set can be improved. In addition, since this processing method performs speech synthesis on the enhanced acoustic features through a vocoder to obtain the enhanced speech of the noisy speech, and avoids directly or indirectly enhancing the phase spectrum of the noisy speech; therefore, speech distortion can be effectively reduced, the auditory quality of the speech can be improved, and thus the speech recognition accuracy can be improved.
[0150] Second Embodiment
[0151] In the above embodiment, a speech enhancement method is provided. Correspondingly, the present application also provides a speech enhancement device. This device corresponds to the embodiment of the above method. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and for the related parts, refer to the partial description of the method embodiment. The device embodiments described below are only illustrative.
[0152] The present application further provides a speech enhancement device, including:
[0153] An acoustic feature extraction unit, configured to determine the acoustic feature data of the first noisy speech data to be processed;
[0154] An acoustic feature enhancement unit, configured to determine enhanced acoustic features of the first noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; wherein, the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with a noise classification loss;
[0155] A speech synthesis unit, configured to generate enhanced speech data of the first noisy speech data according to the enhanced acoustic features through a vocoder.
[0156] The third embodiment
[0157] In the above embodiments, a speech enhancement method is provided. Correspondingly, the present application also provides an electronic device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For related parts, please refer to the partial description of the method embodiments. The device embodiments described below are only illustrative.
[0158] An electronic device according to this embodiment, the electronic device includes: a processor and a memory; the memory is configured to store a program for implementing the speech enhancement method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining acoustic feature data of the first noisy speech data to be processed; determining enhanced acoustic features of the first noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; wherein, the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with a noise classification loss; generating enhanced speech data of the first noisy speech data according to the enhanced acoustic features through a vocoder.
[0159] The fourth embodiment
[0160] Corresponding to the above speech enhancement method, the present application also provides a speech enhancement method. The execution subject of this method includes but is not limited to: a server. For parts of this embodiment that are the same as those in the first embodiment, no further description will be given. Please refer to the corresponding parts in the first embodiment. A speech enhancement method provided by the present application includes:
[0161] Step S401: Learn a vocoder from a correspondence set between acoustic features of clean speech data of multiple users and the clean speech data.
[0162] The vocoder includes a vocoder trained by utilizing the temporal sequence of speech waveforms and simultaneously combining a large amount of speaker training data, such as a WaveRNN vocoder. This step corresponds to step S105 in the first embodiment. Please refer to the corresponding parts in the first embodiment and will not be elaborated here.
[0163] Step S403: Determine enhanced acoustic features of the noisy speech data according to the acoustic feature data of the noisy speech data through an acoustic feature enhancement model.
[0164] This step can use the acoustic feature enhancement model obtained by the adversarial multi-task learning method with noise classification loss provided in the first embodiment, or can use the single-task learning method to obtain the acoustic feature enhancement model, that is, in the model training stage, the noise classifier is not included, and there is no need to minimize the noise classification loss of the noise classifier and maximize the noise classification loss of the encoder. Only the enhancement acoustic feature loss needs to be minimized.
[0165] Step S405: Through the vocoder, according to the enhanced acoustic features, generate enhanced speech data of the noisy speech data.
[0166] This step corresponds to step S105 in the first embodiment. Please refer to the corresponding part in the first embodiment and will not be elaborated here.
[0167] As can be seen from the above embodiments, the speech enhancement method provided by the embodiments of the present application learns to obtain a vocoder from the correspondence set between the acoustic features of the clean speech data of multiple users and the clean speech data; through the acoustic feature enhancement model, according to the acoustic feature data of the noisy speech data, determine the enhanced acoustic features of the noisy speech data; through the vocoder, according to the enhanced acoustic features, generate enhanced speech data of the noisy speech data; this processing method makes use of the temporality of the speech waveform and combines a large amount of speaker training data to construct a vocoder, and performs speech enhancement on the noisy speech based on the vocoder, which can overcome the speaker generalization problem; therefore, the speaker generalization degree can be effectively improved. At the same time, this processing method can avoid directly or indirectly enhancing the phase spectrum of the noisy speech; therefore, the speech distortion can be effectively reduced and the listening quality of the speech can be improved.
[0168] Fifth Embodiment
[0169] In the above embodiment, a speech enhancement method is provided. Correspondingly, the present application also provides a speech enhancement device. This device corresponds to the embodiment of the above method. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0170] The present application further provides a speech enhancement device, including:
[0171] A vocoder construction unit, configured to learn to obtain a vocoder from the correspondence set between the acoustic features of the clean speech data of multiple users and the clean speech data;
[0172] An acoustic feature enhancement unit, configured to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature data of the noisy speech data through an acoustic feature enhancement model;
[0173] A voice synthesis unit, configured to generate enhanced speech data of noisy speech data according to enhanced acoustic features through a vocoder.
[0174] The sixth embodiment
[0175] In the above embodiments, a voice enhancement method is provided. Correspondingly, the present application further provides an electronic device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For related parts, please refer to the partial description of the method embodiments. The device embodiments described below are merely illustrative.
[0176] An electronic device according to this embodiment, the device includes: a processor and a memory; the memory is configured to store a program for implementing the voice enhancement method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: learning a vocoder from the correspondence set between the acoustic features of clean speech data of multiple users and the clean speech data; determining enhanced acoustic features of the noisy speech data according to the acoustic feature data of the noisy speech data through an acoustic feature enhancement model; generating enhanced speech data of the noisy speech data according to the enhanced acoustic features through the vocoder.
[0177] The seventh embodiment
[0178] Corresponding to the above voice enhancement method, the present application further provides a method for constructing a voice enhancement model. The execution subject of this method includes but is not limited to: a server. For parts that are the same as those in the first embodiment in this embodiment, they will not be described again. Please refer to the corresponding parts in the first embodiment. A method for constructing a voice enhancement model provided by the present application includes:
[0179] Step S701: Determine a first training data set and a second training data set. The first training data includes the acoustic feature data of noisy speech data, the noise type, and the correspondence between the acoustic features of clean speech data; the second training data includes the correspondence set between the acoustic features of clean speech data and the clean speech data;
[0180] Step S703: Construct a network structure of a voice noise reduction model; the voice noise reduction model includes an acoustic feature enhancement model and a vocoder; the acoustic feature enhancement model includes an encoder, a decoder, and a noise classifier; the encoder is configured to determine acoustic feature coding data independent of noise of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is configured to determine enhanced acoustic features of the noisy speech data according to the acoustic feature coding data; the noise classifier is configured to determine the noise type of the noisy speech data according to the acoustic feature coding data; the vocoder is configured to generate enhanced speech data of the noisy speech data according to the enhanced acoustic features;
[0181] Step S305: Through the adversarial multi-task learning method of noise classification loss, train the network parameters of the acoustic feature enhancement model according to the first training data set; and train the network parameters of the vocoder according to the second training data set.
[0182] As can be seen from the above embodiments, the method for constructing a voice enhancement model provided by the embodiments of the present application obtains an acoustic feature enhancement model through the self-supervised adversarial multi-task learning method of noise classification loss. By using this model to determine the enhanced acoustic features of noisy speech, it is possible to avoid being sensitive to environmental noise when extracting enhanced acoustic features. Therefore, it is possible to effectively reduce the difference in voice enhancement performance among various environmental noises and improve the generalization ability of noises outside the training set. In addition, since this processing method synthesizes speech from the enhanced acoustic features through a vocoder to obtain enhanced speech of noisy speech, avoiding directly or indirectly enhancing the phase spectrum of noisy speech. Therefore, it is possible to effectively reduce speech distortion and improve the listening quality of speech.
[0183] The Eighth Embodiment
[0184] In the above embodiments, a method for constructing a voice enhancement model is provided. Correspondingly, the present application also provides a device for constructing a voice enhancement model. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For related parts, refer to the partial description of the method embodiments. The device embodiments described below are only illustrative.
[0185] The present application further provides a device for constructing a voice enhancement model, including:
[0186] A training data determination unit, configured to determine a first training data set and a second training data set. The first training data includes the acoustic feature data and noise types of noisy speech data and the corresponding relationship between the acoustic features of clean speech data; the second training data includes the set of corresponding relationships between the acoustic features of clean speech data and clean speech data;
[0187] A model structure construction unit, configured to construct the network structure of a voice noise reduction model; the voice noise reduction model includes an acoustic feature enhancement model and a vocoder; the acoustic feature enhancement model includes an encoder, a decoder, and a noise classifier; the encoder is configured to determine the noise-independent acoustic feature coding data of noisy speech data according to the acoustic feature data of noisy speech data; the decoder is configured to determine the enhanced acoustic features of noisy speech data according to the acoustic feature coding data; the noise classifier is configured to determine the noise type of noisy speech data according to the acoustic feature coding data; the vocoder is configured to generate enhanced speech data of noisy speech data according to the enhanced acoustic features;
[0188] A model parameter training unit, configured to train network parameters of the acoustic feature enhancement model according to the first training dataset by means of adversarial multi-task learning with noise classification loss; and to train network parameters of the vocoder according to the second training dataset.
[0189] The ninth embodiment
[0190] In the above embodiment, a method for constructing a voice enhancement model is provided. Correspondingly, the present application also provides an electronic device. This device corresponds to the embodiment of the above method. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For related parts, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0191] An electronic device according to this embodiment, the device includes: a processor and a memory; the memory is used to store a program for implementing the method for constructing a voice enhancement model. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining a first training dataset and a second training dataset, where the first training data includes acoustic feature data of noisy speech data and the corresponding relationship between the noise type and the acoustic features of clean speech data; the second training data includes a set of corresponding relationships between the acoustic features of clean speech data and clean speech data; constructing a network structure of a voice denoising model; the voice denoising model includes an acoustic feature enhancement model and a vocoder; the acoustic feature enhancement model includes an encoder, a decoder, and a noise classifier; the encoder is configured to determine acoustic feature encoding data of noisy speech data that is independent of noise according to the acoustic feature data of the noisy speech data; the decoder is configured to determine enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; the noise classifier is configured to determine the noise type of the noisy speech data according to the acoustic feature encoding data; the vocoder is configured to generate enhanced speech data of the noisy speech data according to the enhanced acoustic features; training network parameters of the acoustic feature enhancement model according to the first training dataset by means of adversarial multi-task learning with noise classification loss; and training network parameters of the vocoder according to the second training dataset.
[0192] The tenth embodiment
[0193] Corresponding to the above voice enhancement method, the present application also provides a method for processing an acoustic feature enhancement model. The execution subject of this method includes but is not limited to: a server. Parts of this embodiment that are the same as those in the first embodiment will not be described again. Please refer to the corresponding parts in the first embodiment. A method for processing an acoustic feature enhancement model provided by the present application includes:
[0194] Step S1001: Determine a training data set, where the training data includes acoustic feature data of noisy speech data, noise types, and the corresponding relationships between the acoustic features of clean speech data;
[0195] Step S1003: Construct the network structure of the acoustic feature enhancement model; the model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the acoustic feature encoding data of the noisy speech data that is independent of noise according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature encoding data;
[0196] Step S1005: Train the network parameters of the model according to the training data set through the adversarial multi-task learning method of noise classification loss.
[0197] As can be seen from the above embodiments, the acoustic feature enhancement model processing method provided by the embodiments of the present application obtains an acoustic feature enhancement model through the self-supervised adversarial multi-task learning method of noise classification loss. The adversarial multi-task learning method of noise classification loss includes: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss. By using this model to determine the enhanced acoustic features of noisy speech, it is possible to avoid being sensitive to environmental noise when extracting enhanced acoustic features; therefore, it is possible to effectively reduce the difference in speech enhancement performance between various environmental noises and improve the generalization ability of noises outside the training set.
[0198] The eleventh embodiment
[0199] In the above embodiments, an acoustic feature enhancement model processing method is provided. Correspondingly, the present application also provides an acoustic feature enhancement model processing device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. The device embodiments described below are only illustrative.
[0200] The present application further provides an acoustic feature enhancement model processing device, including:
[0201] A training data determination unit, configured to determine a training data set, where the training data includes acoustic feature data of noisy speech data, noise types, and the corresponding relationships between the acoustic features of clean speech data;
[0202] A model structure construction unit for constructing the network structure of an acoustic feature enhancement model; the model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the acoustic feature encoding data of the noisy speech data that is independent of noise according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature encoding data.
[0203] A model parameter training unit for training the network parameters of the model according to the training data set by means of adversarial multi-task learning with noise classification loss.
[0204] The twelfth embodiment
[0205] In the above embodiment, an acoustic feature enhancement model processing method is provided. Correspondingly, the present application also provides an electronic device. This device corresponds to the embodiment of the above method. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For the related parts, please refer to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0206] An electronic device according to this embodiment, the device includes: a processor and a memory; the memory is used to store a program for implementing the acoustic feature enhancement model processing method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining a training data set, the training data including the acoustic feature data and noise type of the noisy speech data, and the corresponding relationship between the acoustic features of the clean speech data; constructing the network structure of the acoustic feature enhancement model; the model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the acoustic feature encoding data of the noisy speech data that is independent of noise according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature encoding data; training the network parameters of the model according to the training data set by means of adversarial multi-task learning with noise classification loss.
[0207] The thirteenth embodiment
[0208] Corresponding to the above speech enhancement method, the present application also provides a speech recognition system. The parts of this embodiment that are the same as those in the first embodiment will not be described in detail again. Please refer to the corresponding parts in the first embodiment. The speech recognition system provided by the present application includes: a client and a server.
[0209] Among them, the client is used to collect voice data and send the voice data to the server; the server is used to determine the acoustic feature data of the voice data; through an acoustic feature enhancement model, according to the acoustic feature data, determine the enhanced acoustic features of the voice data; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, according to the enhanced acoustic features, generate the enhanced voice data of the voice data; through a speech recognition model, convert the enhanced voice data into text.
[0210] Since the speech recognition model belongs to a relatively mature existing technology, it will not be elaborated here.
[0211] As can be seen from the above embodiments, the speech recognition system provided by the embodiments of the present application collects voice data through the client and sends the voice data to the server; the server determines the acoustic feature data of the voice data; through the acoustic feature enhancement model, according to the acoustic feature data, determine the enhanced acoustic features of the voice data; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, according to the enhanced acoustic features, generate the enhanced voice data of the voice data; through a speech recognition model, convert the enhanced voice data into text; this processing method enables an acoustic feature enhancement model to be obtained through a self-supervised noise classification loss adversarial multi-task learning method, and determine the enhanced acoustic features of the noisy speech through this model, so as to avoid being sensitive to environmental noise when extracting the enhanced acoustic features, and then perform speech recognition processing on the enhanced voice; therefore, it can effectively narrow the difference in speech recognition performance among various environmental noises and improve the generalization ability of noises outside the training set. In addition, since this processing method performs speech synthesis on the enhanced acoustic features through a vocoder to obtain the enhanced voice of the noisy speech, avoiding directly or indirectly enhancing the phase spectrum of the noisy speech; therefore, it can effectively reduce speech distortion and improve the listening quality of the voice, thereby improving the accuracy of speech recognition.
[0212] The Fourteenth Embodiment
[0213] Corresponding to the above speech recognition system, the present application also provides a speech recognition method, and the execution subject of this method includes but is not limited to: the server. The parts of this embodiment that are the same as those in the system embodiment will not be elaborated, please refer to the corresponding parts in the system embodiment. A speech recognition method provided by the present application includes:
[0214] Step S1401: Determine the acoustic feature data of the noisy speech data to be processed;
[0215] Step S1403: Determine the enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss;
[0216] Step S1405: Generate enhanced speech data of the noisy speech data through a vocoder according to the enhanced acoustic features;
[0217] Step S1407: Convert the enhanced speech data into text through a speech recognition model.
[0218] Since the speech recognition model belongs to a relatively mature existing technology, it will not be elaborated here.
[0219] The fifteenth embodiment
[0220] In the above embodiments, a speech recognition method is provided. Correspondingly, the present application also provides a speech recognition device. This device corresponds to the method embodiments above. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For related parts, refer to the partial description of the method embodiments. The device embodiments described below are only illustrative.
[0221] The present application further provides a speech recognition device, including:
[0222] An acoustic feature extraction unit, configured to determine the acoustic feature data of the noisy speech data to be processed;
[0223] An acoustic feature enhancement unit, configured to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss;
[0224] A speech synthesis unit, configured to generate enhanced speech data of the noisy speech data through a vocoder according to the enhanced acoustic features;
[0225] A speech conversion unit, configured to convert the enhanced speech data into text through a speech recognition model.
[0226] The sixteenth embodiment
[0227] In the above embodiments, a speech recognition method is provided. Correspondingly, the present application also provides an electronic device. This device corresponds to the method embodiments above. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For related parts, refer to the partial description of the method embodiments. The device embodiments described below are only illustrative.
[0228] An electronic device according to this embodiment, the device includes: a processor and a memory; the memory is used to store a program for implementing a speech recognition method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining acoustic feature data of noisy speech data to be processed; through an acoustic feature enhancement model, determining enhanced acoustic features of the noisy speech data according to the acoustic feature data; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, generating enhanced speech data of the noisy speech data according to the enhanced acoustic features; through a speech recognition model, converting the enhanced speech data into text.
[0229] The seventeenth embodiment
[0230] Corresponding to the above speech recognition system, the present application also provides a speech recognition method, and the execution subject of this method includes but is not limited to: a client. The parts of this embodiment that are the same as those of the system embodiment will not be described in detail again. Please refer to the corresponding parts in the system embodiment. A speech recognition method provided by the present application includes:
[0231] Step S1701: Collect speech data;
[0232] Step S1703: Send the speech data to the server so that the server determines the acoustic feature data of the speech data; through an acoustic feature enhancement model, determining enhanced acoustic features of the speech data; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, generating enhanced speech data of the speech data according to the enhanced acoustic features; through a speech recognition model, converting the enhanced speech data into text.
[0233] The eighteenth embodiment
[0234] In the above embodiment, a speech recognition method is provided. Correspondingly, the present application also provides a speech recognition device. This device corresponds to the embodiment of the above method. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For related parts, please refer to the partial description of the method embodiment. The device embodiments described below are only illustrative.
[0235] The present application further provides a speech recognition device, including:
[0236] A speech data acquisition unit, used to collect speech data;
[0237] A voice data sending unit, configured to send the voice data to a server so that the server determines acoustic feature data of the voice data; determine enhanced acoustic features of the voice data through an acoustic feature enhancement model, where the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; generate enhanced voice data of the voice data according to the enhanced acoustic features through a vocoder; and convert the enhanced voice data into text through a speech recognition model.
[0238] The nineteenth embodiment
[0239] In the above embodiments, a speech recognition method is provided. Correspondingly, the present application also provides an electronic device. This device corresponds to the method embodiments. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For related parts, please refer to the partial description of the method embodiments. The device embodiments described below are only illustrative.
[0240] An electronic device according to this embodiment, the device includes: a processor and a memory; the memory is configured to store a program for implementing the speech recognition method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: collect voice data; send the voice data to a server so that the server determines acoustic feature data of the voice data; determine enhanced acoustic features of the voice data through an acoustic feature enhancement model, where the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; generate enhanced voice data of the voice data according to the enhanced acoustic features through a vocoder; and convert the enhanced voice data into text through a speech recognition model.
[0241] The twentieth embodiment
[0242] Corresponding to the above speech enhancement method, the present application also provides a speech recognition text editing system. Parts of this embodiment that are the same as those in the first embodiment will not be described again. Please refer to the corresponding parts in the first embodiment. A speech recognition text editing system provided by the present application includes: a client and a server.
[0243] Among them, the client is configured to collect voice data, send the voice data to the server; and edit the text of the voice data recognized by the server; the server is configured to determine acoustic feature data of the voice data; determine enhanced acoustic features of the voice data through an acoustic feature enhancement model, where the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; generate enhanced voice data of the voice data according to the enhanced acoustic features through a vocoder; and convert the enhanced voice data into text through a speech recognition model.
[0244] Since speech recognition and the editing of speech-to-written text are both relatively mature existing technologies, they will not be elaborated here.
[0245] As can be seen from the above embodiments, the speech recognition text editing system provided by the embodiments of the present application collects speech data through a client and sends the speech data to a server; the server determines the acoustic feature data of the speech data; through an acoustic feature enhancement model, according to the acoustic feature data, determines the enhanced acoustic features of the speech data; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, according to the enhanced acoustic features, generates enhanced speech data of the speech data; through a speech recognition model, converts the enhanced speech data into text; the client edits the text; this processing method enables an acoustic feature enhancement model to be obtained through a self-supervised noise classification loss adversarial multi-task learning method, and determines the enhanced acoustic features of the noisy speech through this model, so that it is possible to avoid being sensitive to environmental noise when extracting enhanced acoustic features, and then perform speech recognition processing on the enhanced speech; therefore, it is possible to effectively reduce the difference in speech recognition performance among various environmental noises and improve the generalization ability of noises outside the training set. In addition, since this processing method performs speech synthesis on the enhanced acoustic features through a vocoder to obtain enhanced speech of the noisy speech, avoiding directly or indirectly enhancing the phase spectrum of the noisy speech; therefore, it is possible to effectively reduce speech distortion, improve the listening quality of speech, thereby improving speech recognition accuracy, and further improving the efficiency of speech recognition text editing.
[0246] Twenty-first Embodiment
[0247] Corresponding to the above speech enhancement method, the present application also provides a user recognition method, and the execution subject of this method includes but is not limited to: a server. The parts of this embodiment that are the same as those in the first embodiment will not be elaborated, please refer to the corresponding parts in the first embodiment. A user recognition method provided by the present application includes:
[0248] Step S2101: Determine the acoustic feature data of the noisy speech data to be processed;
[0249] Step S2103: Through an acoustic feature enhancement model, according to the acoustic feature data, determine the enhanced acoustic features of the noisy speech data; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method;
[0250] Step S2105: Through a vocoder, according to the enhanced acoustic features, generate enhanced speech data of the noisy speech data;
[0251] Step S2107: Through a user recognition model, determine the user information of the enhanced speech data.
[0252] Since the user identification model belongs to relatively mature existing technologies, it will not be elaborated here.
[0253] As can be seen from the above embodiments, the user identification method provided by the embodiments of the present application includes: determining acoustic feature data of noisy speech data to be processed; determining enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model, where the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with a noise classification loss; generating enhanced speech data of the noisy speech data through a vocoder according to the enhanced acoustic features; and determining user information of the enhanced speech data through a user identification model. This processing method enables the acoustic feature enhancement model to be obtained through an adversarial multi-task learning method with a self-supervised noise classification loss, and the enhanced acoustic features of the noisy speech are determined through this model, which can avoid being sensitive to environmental noise when extracting the enhanced acoustic features, and then user identification processing is performed based on the enhanced speech. Therefore, it can effectively reduce the difference in speaker identification performance of speech under various environmental noises and improve the generalization ability of noises outside the training set. In addition, since this processing method synthesizes speech from the enhanced acoustic features through a vocoder to obtain the enhanced speech of the noisy speech, and avoids directly or indirectly enhancing the phase spectrum of the noisy speech, it can effectively reduce speech distortion and improve the listening quality of the speech, thereby improving the accuracy of speaker identification.
[0254] Twenty-second Embodiment
[0255] In the above embodiments, a user identification method is provided. Correspondingly, the present application also provides a user identification device. This device corresponds to the method embodiments above. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For related parts, refer to the partial description of the method embodiments. The device embodiments described below are merely illustrative.
[0256] The present application further provides a user identification device, including:
[0257] An acoustic feature determination unit, configured to determine acoustic feature data of noisy speech data to be processed;
[0258] An acoustic feature enhancement unit, configured to determine enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model, where the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with a noise classification loss;
[0259] A speech synthesis unit, configured to generate enhanced speech data of the noisy speech data through a vocoder according to the enhanced acoustic features;
[0260] A user determination unit, configured to determine user information of the enhanced speech data through a user identification model.
[0261] Twenty-third Embodiment
[0262] In the above embodiments, a user identification method is provided. Correspondingly, the present application also provides an electronic device. This device corresponds to the embodiment of the above method. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment. The device embodiments described below are merely illustrative.
[0263] An electronic device according to this embodiment, the device includes: a processor and a memory; the memory is used to store a program for implementing the user identification method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining the acoustic feature data of the noisy speech data to be processed; through an acoustic feature enhancement model, determining the enhanced acoustic features of the noisy speech data according to the acoustic feature data; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, generating enhanced speech data of the noisy speech data according to the enhanced acoustic features; determining the user information of the enhanced speech data through a user identification model.
[0264] Although the present application is disclosed above with preferred embodiments, it is not used to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims of the present application.
[0265] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0266] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0267] 1. A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory media such as modulated data signals and carrier waves.
[0268] 2. Those skilled in the art will understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
Claims
1. A voice enhancement method, characterized in that, comprising: determining acoustic feature data of first noisy voice data to be processed; through an acoustic feature enhancement model, determining enhanced acoustic features of the first noisy voice data according to the acoustic feature data; wherein, the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with a noise classification loss; the model includes: an encoder, a decoder, and a noise classifier; in the process of training the model, taking the noise type of second noisy voice data as the output data of the noise classifier, and taking the acoustic features of clean voice data as the output data of the decoder; the encoder is used for determining acoustic feature encoding data independent of noise of the second noisy voice data according to the acoustic feature data of the second noisy voice data; the decoder is used for determining enhanced acoustic features of the second noisy voice data according to the acoustic feature encoding data; the noise classifier is used for determining the noise type of the second noisy voice data according to the acoustic feature encoding data; the training objectives of the model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; through a vocoder, generating enhanced voice data of the first noisy voice data according to the enhanced acoustic features.
2. The method according to claim 1, characterized in that, there are different types of environmental noises between the first noisy voice data and the training data of the model; the enhanced acoustic features include enhanced acoustic features that suppress environmental noises not present in the training data.
3. The method according to claim 1, characterized in that, the step of determining enhanced acoustic features of the first noisy voice data through the acoustic feature enhancement model according to the acoustic feature data includes: through the encoder included in the model, determining acoustic feature encoding data independent of noise of the first noisy voice data according to the acoustic feature data; through the decoder included in the model, determining enhanced acoustic features of the first noisy voice data according to the acoustic feature encoding data.
4. The method according to claim 1, characterized in that, further comprising: generating second noisy voice data according to clean voice data and noise data.
5. The method according to claim 1, characterized in that, the vocoder is learned from a correspondence set between acoustic features of clean voice data of multiple users and the clean voice data.
6. The method according to claim 5, characterized in that, the vocoder includes: a vocoder based on a waveform recurrent neural network.
7. The method according to claim 1, characterized in that, the acoustic feature data includes: complex spectrum; the enhanced acoustic features include: Mel spectrum.
8. A voice enhancement method, characterized in that, comprising: learning a vocoder from a correspondence set between acoustic features of clean voice data of multiple users and the clean voice data; Through an acoustic feature enhancement model, based on the acoustic feature data of noisy speech data, the enhanced acoustic features of the noisy speech data are determined; wherein, the model is obtained through an adversarial multi-task learning method with noise classification loss; the model includes: an encoder, a decoder, and a noise classifier; during the process of training the model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is used to determine the noise-independent acoustic feature encoding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used to determine the enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss. Through a vocoder, based on the enhanced acoustic features, enhanced speech data of the noisy speech data is generated.
9. A method for processing a speech enhancement model, characterized in that, it includes: determining a first training data set and a second training data set, the first training data including the correspondence between the acoustic feature data and the noise type of the noisy speech data and the acoustic features of the clean speech data; the second training data including the set of correspondences between the acoustic features of the clean speech data and the clean speech data; constructing the network structure of a speech denoising model; the speech denoising model includes an acoustic feature enhancement model and a vocoder; the acoustic feature enhancement model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the noise-independent acoustic feature encoding data of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature encoding data; the vocoder is used to generate the enhanced speech data of the noisy speech data according to the enhanced acoustic features; through an adversarial multi-task learning method with noise classification loss, according to the first training data set, training the network parameters of the acoustic feature enhancement model, and the training objectives include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; and, according to the second training data set, training the network parameters of the vocoder.
10. A method for processing an acoustic feature enhancement model, characterized in that, it includes: determining a training data set, the training data including the correspondence between the acoustic feature data and the noise type of the noisy speech data and the acoustic features of the clean speech data; Construct the network structure of the acoustic feature enhancement model; the model includes an encoder, a decoder, and a noise classifier; the encoder is used to determine the noise-independent acoustic feature encoding data of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is used to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; The noise classifier is used to determine the noise type of the noisy speech data according to the acoustic feature encoding data; Through the noise classification loss adversarial multi-task learning method, according to the training data set, train the network parameters of the model, and the training objectives include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss.
11. A speech recognition system, Characterized in that, It includes: A client for collecting speech data and sending the speech data to the server; A server for determining the acoustic feature data of the speech data; Through an acoustic feature enhancement model, according to the acoustic feature data, determine the enhanced acoustic features of the speech data; the acoustic feature enhancement model is obtained through the noise classification loss adversarial multi-task learning method; through a vocoder, according to the enhanced acoustic features, generate the enhanced speech data of the speech data; Convert the enhanced speech data into text through a speech recognition model; wherein, the acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the training of the acoustic feature enhancement model, use the noise type of the second noisy speech data as the output data of the noise classifier, and use the acoustic features of the clean speech data as the output data of the decoder; the encoder is used to determine the noise-independent acoustic feature encoding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used to determine the enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss.
12. A speech recognition method, Characterized in that, It includes: Determine the acoustic feature data of the noisy speech data to be processed; Through an acoustic feature enhancement model, according to the acoustic feature data, determine the enhanced acoustic features of the noisy speech data; The acoustic feature enhancement model is obtained through the noise classification loss adversarial multi-task learning method; The acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the training of the acoustic feature enhancement model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is used to determine the noise-independent acoustic feature encoding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used to determine the enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; Through a vocoder, enhanced speech data of the noisy speech data is generated according to the enhanced acoustic features; The enhanced speech data is converted into text through a speech recognition model.
13. A speech recognition method, characterized in that, it includes: Collect speech data; Send the speech data to the server so that the server determines the acoustic feature data of the speech data; determine the enhanced acoustic features of the speech data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, enhanced speech data of the speech data is generated according to the enhanced acoustic features; The enhanced speech data is converted into text through a speech recognition model; wherein, the acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the training of the acoustic feature enhancement model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is used to determine the noise-independent acoustic feature encoding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used to determine the enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss.
14. A speech recognition text editing system, characterized in that, it includes: A client, which is used to collect speech data, send the speech data to the server; and edit the text of the speech data recognized by the server; A server, which is used to determine the acoustic feature data of the speech data; determine the enhanced acoustic features of the speech data through an acoustic feature enhancement model; the acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; through a vocoder, enhanced speech data of the speech data is generated according to the enhanced acoustic features; Convert the enhanced speech data into text through a speech recognition model; wherein, the acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the training of the acoustic feature enhancement model, use the noise type of the second noisy speech data as the output data of the noise classifier, and use the acoustic features of the clean speech data as the output data of the decoder; the encoder is used to determine the noise-independent acoustic feature encoding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used to determine the enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss.
15. A user identification method, Characterized in that, It includes: Determine the acoustic feature data of the noisy speech data to be processed; Through an acoustic feature enhancement model, determine the enhanced acoustic features of the noisy speech data according to the acoustic feature data; The acoustic feature enhancement model is obtained through a noise classification loss adversarial multi-task learning method; The acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the training of the acoustic feature enhancement model, use the noise type of the second noisy speech data as the output data of the noise classifier, and use the acoustic features of the clean speech data as the output data of the decoder; the encoder is used to determine the noise-independent acoustic feature encoding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used to determine the enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; Through a vocoder, generate enhanced speech data of the noisy speech data according to the enhanced acoustic features; Through a user identification model, determine the user information of the enhanced speech data.
16. A speech enhancement device, Characterized in that, It includes: An acoustic feature extraction unit for determining the acoustic feature data of the first noisy speech data to be processed; An acoustic feature enhancement unit, configured to determine enhanced acoustic features of the first noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; wherein, the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss; the model includes: an encoder, a decoder, and a noise classifier; during the process of training the model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is configured to determine acoustic feature encoding data irrelevant to noise of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is configured to determine enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is configured to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; A speech synthesis unit, configured to generate enhanced speech data of the first noisy speech data according to the enhanced acoustic features through a vocoder.
17. An electronic device characterized in that it includes: a processor and a memory; The memory is configured to store a program for implementing the speech enhancement method. After the device is powered on and runs the program of the method through the processor, the following steps are executed: determining acoustic feature data of the first noisy speech data to be processed; determining enhanced acoustic features of the first noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; wherein, the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss; the model includes: an encoder, a decoder, and a noise classifier; during the process of training the model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is configured to determine acoustic feature encoding data irrelevant to noise of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is configured to determine enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is configured to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; generating enhanced speech data of the first noisy speech data according to the enhanced acoustic features through a vocoder.
18. A speech enhancement device characterized in that it includes: a vocoder construction unit, configured to learn a vocoder from a correspondence set between the acoustic features of the clean speech data of multiple users and the clean speech data; An acoustic feature enhancement unit is configured to determine enhanced acoustic features of noisy speech data based on the acoustic feature data of the noisy speech data through an acoustic feature enhancement model; wherein, the model is obtained through an adversarial multi-task learning method with noise classification loss; the model includes: an encoder, a decoder, and a noise classifier; during the process of training the model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is configured to determine acoustic feature encoding data irrelevant to noise of the second noisy speech data based on the acoustic feature data of the second noisy speech data; the decoder is configured to determine enhanced acoustic features of the second noisy speech data based on the acoustic feature encoding data; the noise classifier is configured to determine the noise type of the second noisy speech data based on the acoustic feature encoding data; the training objectives of the model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; A speech synthesis unit is configured to generate enhanced speech data of the noisy speech data based on the enhanced acoustic features through a vocoder.
19. An electronic device, characterized in that, it includes: a processor and a memory; The memory is configured to store a program for implementing the speech enhancement method. After the device is powered on and runs the program of the method through the processor, the following steps are executed: learning a vocoder from the correspondence set between the acoustic features and the clean speech data of multiple users; determining enhanced acoustic features of the noisy speech data based on the acoustic feature data of the noisy speech data through an acoustic feature enhancement model; wherein, the model is obtained through an adversarial multi-task learning method with noise classification loss; the model includes: an encoder, a decoder, and a noise classifier; during the process of training the model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is configured to determine acoustic feature encoding data irrelevant to noise of the second noisy speech data based on the acoustic feature data of the second noisy speech data; the decoder is configured to determine enhanced acoustic features of the second noisy speech data based on the acoustic feature encoding data; the noise classifier is configured to determine the noise type of the second noisy speech data based on the acoustic feature encoding data; the training objectives of the model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; generating enhanced speech data of the noisy speech data based on the enhanced acoustic features through a vocoder.
20. A speech enhancement model processing device, characterized in that, it includes: A training data determination unit for determining a first training data set and a second training data set, where the first training data includes acoustic feature data of noisy speech data and the corresponding relationship between the noise type and the acoustic features of clean speech data; the second training data includes a set of corresponding relationships between the acoustic features of clean speech data and clean speech data; A model structure construction unit for constructing the network structure of a speech denoising model; the speech denoising model includes an acoustic feature enhancement model and a vocoder; the acoustic feature enhancement model includes an encoder, a decoder, and a noise classifier; the encoder is configured to determine the noise-independent acoustic feature encoded data of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is configured to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoded data; The noise classifier is configured to determine the noise type of the noisy speech data according to the acoustic feature encoded data; the vocoder is configured to generate enhanced speech data of the noisy speech data according to the enhanced acoustic features; A model parameter training unit for training the network parameters of the acoustic feature enhancement model according to the first training data set by means of adversarial multi-task learning of noise classification loss, and the training objectives include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; And training the network parameters of the vocoder according to the second training data set.
21. An electronic device, characterized in that, it includes: a processor and a memory; The memory is used to store a program for implementing the speech enhancement model construction method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining a first training data set and a second training data set, where the first training data includes acoustic feature data of noisy speech data and the corresponding relationship between the noise type and the acoustic features of clean speech data; the second training data includes a set of corresponding relationships between the acoustic features of clean speech data and clean speech data; constructing the network structure of a speech denoising model; the speech denoising model includes an acoustic feature enhancement model and a vocoder; the acoustic feature enhancement model includes an encoder, a decoder, and a noise classifier; the encoder is configured to determine the noise-independent acoustic feature encoded data of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is configured to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoded data; The noise classifier is configured to determine the noise type of the noisy speech data according to the acoustic feature encoded data; the vocoder is configured to generate enhanced speech data of the noisy speech data according to the enhanced acoustic features; Training the network parameters of the acoustic feature enhancement model according to the first training data set by means of adversarial multi-task learning of noise classification loss, and the training objectives include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; And, according to the second training data set, train the network parameters of the vocoder.
22. An acoustic feature enhancement model processing device, characterized in that, it includes: A training data determination unit, configured to determine a training data set, where the training data includes the acoustic feature data of the noisy speech data and the noise type, and the correspondence relationship between the acoustic features of the clean speech data; A model structure construction unit, configured to construct the network structure of the acoustic feature enhancement model; the model includes an encoder, a decoder, and a noise classifier; the encoder is configured to determine the noise-independent acoustic feature encoding data of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is configured to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; The noise classifier is configured to determine the noise type of the noisy speech data according to the acoustic feature encoding data; A model parameter training unit, configured to train the network parameters of the model according to the training data set by means of adversarial multi-task learning with noise classification loss, and the training objectives include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss.
23. An electronic device, characterized in that, it includes: A processor and a memory; The memory is used to store a program for implementing the acoustic feature enhancement model construction method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining a training data set, where the training data includes the acoustic feature data of the noisy speech data and the noise type, and the correspondence relationship between the acoustic features of the clean speech data; constructing the network structure of the acoustic feature enhancement model; the model includes an encoder, a decoder, and a noise classifier; the encoder is configured to determine the noise-independent acoustic feature encoding data of the noisy speech data according to the acoustic feature data of the noisy speech data; the decoder is configured to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature encoding data; The noise classifier is configured to determine the noise type of the noisy speech data according to the acoustic feature encoding data; By means of adversarial multi-task learning with noise classification loss, train the network parameters of the model according to the training data set, and the training objectives include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss.
24. A speech recognition device, characterized in that, it includes: An acoustic feature extraction unit, configured to determine the acoustic feature data of the noisy speech data to be processed; An acoustic feature enhancement unit, configured to determine the enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; The acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss. The acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the training of the acoustic feature enhancement model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is used to determine the noise-independent acoustic feature encoding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used to determine the enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; A speech synthesis unit, configured to generate enhanced speech data of the noisy speech data through a vocoder according to the enhanced acoustic features; A speech conversion unit, configured to convert the enhanced speech data into text through a speech recognition model.
25. An electronic device, characterized in that it includes: a processor and a memory; The memory is used to store a program for implementing the speech recognition method. After the device is powered on and the program of this method is run through the processor, the following steps are executed: determining the acoustic feature data of the noisy speech data to be processed; through the acoustic feature enhancement model, determining the enhanced acoustic features of the noisy speech data according to the acoustic feature data; The acoustic feature enhancement model is obtained through an adversarial multi-task learning method of noise classification loss; The acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the training of the acoustic feature enhancement model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is used to determine the noise-independent acoustic feature encoding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used to determine the enhanced acoustic features of the second noisy speech data according to the acoustic feature encoding data; the noise classifier is used to determine the noise type of the second noisy speech data according to the acoustic feature encoding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; generating enhanced speech data of the noisy speech data through a vocoder according to the enhanced acoustic features; converting the enhanced speech data into text through a speech recognition model.
26. A speech recognition device, characterized in that it includes: A speech data acquisition unit, configured to acquire speech data; A voice data sending unit, configured to send the voice data to a server so that the server determines acoustic feature data of the voice data; determine enhanced acoustic features of the voice data through an acoustic feature enhancement model, where the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss; generate enhanced voice data of the voice data according to the enhanced acoustic features through a vocoder; convert the enhanced voice data into text through a speech recognition model; where the acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the process of training the acoustic feature enhancement model, the noise type of second noisy voice data is used as output data of the noise classifier, and the acoustic features of clean voice data are used as output data of the decoder; the encoder is configured to determine acoustic feature encoding data irrelevant to noise of the second noisy voice data according to the acoustic feature data of the second noisy voice data; the decoder is configured to determine enhanced acoustic features of the second noisy voice data according to the acoustic feature encoding data; the noise classifier is configured to determine the noise type of the second noisy voice data according to the acoustic feature encoding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss.
27. An electronic device, characterized in that, it includes: a processor and a memory; The memory is configured to store a program for implementing a speech recognition method. After the device is powered on and runs the program of the method through the processor, the following steps are executed: collect voice data; send the voice data to a server so that the server determines acoustic feature data of the voice data; determine enhanced acoustic features of the voice data through an acoustic feature enhancement model, where the acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss; generate enhanced voice data of the voice data according to the enhanced acoustic features through a vocoder; convert the enhanced voice data into text through a speech recognition model; where the acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the process of training the acoustic feature enhancement model, the noise type of second noisy voice data is used as output data of the noise classifier, and the acoustic features of clean voice data are used as output data of the decoder; the encoder is configured to determine acoustic feature encoding data irrelevant to noise of the second noisy voice data according to the acoustic feature data of the second noisy voice data; the decoder is configured to determine enhanced acoustic features of the second noisy voice data according to the acoustic feature encoding data; the noise classifier is configured to determine the noise type of the second noisy voice data according to the acoustic feature encoding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss.
28. A user identification device, characterized in that, it includes: An acoustic feature determination unit for determining acoustic feature data of noisy speech data to be processed; An acoustic feature enhancement unit for determining enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; The acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss; The acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the training of the acoustic feature enhancement model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is used for determining the noise-independent acoustic feature coding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used for determining the enhanced acoustic features of the second noisy speech data according to the acoustic feature coding data; the noise classifier is used for determining the noise type of the second noisy speech data according to the acoustic feature coding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; A speech synthesis unit for generating enhanced speech data of the noisy speech data according to the enhanced acoustic features through a vocoder; A user determination unit for determining user information of the enhanced speech data through a user recognition model.
29. An electronic device, characterized in that, it includes: a processor and a memory; The memory is used for storing a program for implementing the user recognition method. After the device is powered on and runs the program of this method through the processor, the following steps are executed: determining acoustic feature data of noisy speech data to be processed; determining enhanced acoustic features of the noisy speech data according to the acoustic feature data through an acoustic feature enhancement model; The acoustic feature enhancement model is obtained through an adversarial multi-task learning method with noise classification loss; The acoustic feature enhancement model includes: an encoder, a decoder, and a noise classifier; during the training of the acoustic feature enhancement model, the noise type of the second noisy speech data is used as the output data of the noise classifier, and the acoustic features of the clean speech data are used as the output data of the decoder; the encoder is used for determining the noise-independent acoustic feature coding data of the second noisy speech data according to the acoustic feature data of the second noisy speech data; the decoder is used for determining the enhanced acoustic features of the second noisy speech data according to the acoustic feature coding data; the noise classifier is used for determining the noise type of the second noisy speech data according to the acoustic feature coding data; the training objectives of the acoustic feature enhancement model include: minimizing the noise classification loss of the noise classifier, maximizing the noise classification loss of the encoder, and minimizing the enhanced acoustic feature loss; generating enhanced speech data of the noisy speech data according to the enhanced acoustic features through a vocoder; determining user information of the enhanced speech data through a user recognition model.
Citation Information
Patent Citations
End-to-end voice enhancement method based on generation of countermeasure network
CN110390950A
Voice enhancing method based on generative adversarial network
CN110428849A