Method and device for generating mixed language speech recognition model
By using a self-supervised learning model for feature extraction and model training in mixed language speech recognition, the problem of insufficient features caused by data sparseness is solved, and the recognition accuracy of mixed language speech recognition is improved.
Patent Information
- Application Number
- CN202210600930.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-30
AI Technical Summary
The recognition accuracy of the mixed language speech recognition model is not high, mainly due to sparse data, resulting in insufficient features.
The self-supervised learning model is used as a feature extractor to preprocess the audio samples, and a mixed language speech recognition model is generated through joint training of the language recognition network and the speech recognition network.
The accuracy of the feature vector corresponding to each frame of audio data is improved, and the recognition accuracy of the mixed language speech generation model is improved.
Smart Images

Figure CN115064154B_ABST
Abstract
Claims
1. A method for generating a mixed language speech recognition model, It is characterized in that include: Acquire a training data set, wherein the training data set includes audio samples and annotated text corresponding to the audio samples; Using a self-supervised learning model to extract features from each frame of audio data in the audio sample to obtain a feature vector corresponding to each frame of audio data, wherein the self-supervised learning model is obtained through self-supervised training using audio data in multiple languages; Inputting the feature vectors into the language recognition network and the speech recognition network in the initial mixed language speech recognition model respectively to obtain the language probability distribution and word probability distribution corresponding to each frame of audio data; Obtaining the product of the probability corresponding to each word in the word probability distribution and the probability corresponding to the language to which the word belongs in the language probability distribution, and using the product as the first probability corresponding to each word; Calculate the sum of the first probabilities corresponding to each word to obtain a second probability, calculate the ratio between the first probability of each word and the second probability, and use the ratio as the updated word probability distribution; Determining a loss value corresponding to each frame of audio data according to the language probability distribution, the updated word probability distribution and the annotated text; Based on the loss values corresponding to the frames of audio data, the language recognition network and the speech recognition network are modified respectively to obtain a mixed language speech recognition model.
2. The method according to claim 1, It is characterized in that The determining, according to the language probability distribution, the updated word probability distribution and the annotated text, a loss value corresponding to each frame of audio data includes: Determining, according to each character in the annotated text, an annotated character corresponding to each frame of audio data in the audio sample; Determining the language to which each frame of audio data belongs according to the probability corresponding to each language in the language probability distribution; Determine a language loss value according to the difference between the language of each frame of audio data and the language of the corresponding annotated character; Determining a text recognition result corresponding to each frame of audio data according to the updated word probability distribution; Determining a speech recognition loss value according to a difference between the text recognition result and the annotated character; A loss value corresponding to each frame of audio data is determined according to the language loss value and the speech recognition loss value.
3. The method according to claim 1, It is characterized in that The method of modifying the language recognition network and the speech recognition network based on the loss value corresponding to each frame of audio data to obtain a mixed language speech recognition model includes: Determining a loss value corresponding to the audio sample according to the sum of the loss values corresponding to the audio data of each frame; According to the loss value corresponding to the audio sample, the language recognition network and the speech recognition network are respectively modified.
4. The method according to claim 1, It is characterized in that The using of the self-supervised learning model to extract features from each frame of audio data in the audio sample to obtain a feature vector corresponding to each frame of audio data includes: Using the self-supervised learning model to perform feature extraction on each frame of audio data to obtain sub-feature vectors output by each hidden layer in the self-supervised learning model; The sub-feature vectors output by each hidden layer are fused to obtain a feature vector corresponding to each frame of audio data.
5. A mixed language speech recognition method, It is characterized in that include: Obtain audio data to be recognized; Using a self-supervised learning model to extract features from the audio data to be identified, so as to obtain a feature vector corresponding to each frame of audio data in the audio data to be identified, wherein the self-supervised learning model is obtained through self-supervised training using audio data in multiple languages; Inputting the feature vector corresponding to each frame of audio data into a mixed-language speech recognition model to obtain a recognition result corresponding to each frame of audio data; wherein the mixed-language speech recognition model is generated by using the method according to any one of claims 1 to 4; According to the recognition results corresponding to the audio data of each frame, the recognition results corresponding to the audio data to be recognized are determined.
6. A device for generating a mixed language speech recognition model, It is characterized in that include: A first acquisition module is used to acquire a training data set, wherein the training data set includes an audio sample and an annotated text corresponding to the audio sample; A second acquisition module is used to extract features from each frame of audio data in the audio sample using a self-supervised learning model to obtain a feature vector corresponding to each frame of audio data, wherein the self-supervised learning model is obtained through self-supervised training using audio data in multiple languages; A third acquisition module is used to input the feature vectors into the language recognition network and the speech recognition network in the initial mixed language speech recognition model respectively, so as to obtain the language probability distribution and word probability distribution corresponding to each frame of audio data; A determination module, configured to determine a loss value corresponding to each frame of audio data according to the language probability distribution, the word probability distribution and the annotated text; A training module, configured to modify the language recognition network and the speech recognition network respectively based on the loss values corresponding to the respective frames of audio data, so as to obtain a mixed language speech recognition model; The determining module comprises: An updating unit, configured to obtain the product of the probability corresponding to each word in the word probability distribution and the probability corresponding to the language to which the word belongs in the language probability distribution, and use the product as the first probability corresponding to each word; calculate the sum of the first probabilities corresponding to each word to obtain a second probability, and calculate the ratio between the first probability and the second probability of each word, and use the ratio as the updated word probability distribution; A determination unit is used to determine the loss value corresponding to each frame of audio data according to the language probability distribution, the updated word probability distribution and the annotated text.
7. The device according to claim 6, It is characterized in that The determining unit is used for: Determining, according to each character in the annotated text, an annotated character corresponding to each frame of audio data in the audio sample; Determining the language to which each frame of audio data belongs according to the probability corresponding to each language in the language probability distribution; Determine a language loss value according to the difference between the language of each frame of audio data and the language of the corresponding annotated character; Determining a text recognition result corresponding to each frame of audio data according to the updated word probability distribution; Determining a speech recognition loss value according to a difference between the text recognition result and the annotated character; A loss value corresponding to each frame of audio data is determined according to the language loss value and the speech recognition loss value.
8. The device according to claim 6, It is characterized in that The training module is used to: Determining a loss value corresponding to the audio sample according to the sum of the loss values corresponding to the audio data of each frame; According to the loss value corresponding to the audio sample, the language recognition network and the speech recognition network are respectively modified.
9. The device according to claim 6, It is characterized in that The second acquisition module is used to: Using the self-supervised learning model to perform feature extraction on each frame of audio data to obtain sub-feature vectors output by each hidden layer in the self-supervised learning model; The sub-feature vectors output by each hidden layer are fused to obtain a feature vector corresponding to each frame of audio data.
10. A mixed language speech recognition device, It is characterized in that include: A first acquisition module, used to acquire audio data to be recognized; A second acquisition module is used to extract features from the audio data to be identified using a self-supervised learning model to obtain a feature vector corresponding to each frame of audio data in the audio data to be identified, wherein the self-supervised learning model is obtained through self-supervised training using audio data in multiple languages; A third acquisition module is used to input the feature vector corresponding to each frame of audio data into a mixed language speech recognition model to obtain a recognition result corresponding to each frame of audio data; wherein the mixed language speech recognition model is generated by using the method according to any one of claims 1 to 4; The determination module is used to determine the recognition result corresponding to the audio data to be recognized according to the recognition result corresponding to each frame of audio data.
11. A computer device, It is characterized in that including a processor and a memory; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to implement the method according to any one of claims 1 to 4 or the method according to claim 5.
12. A non-transitory computer-readable storage medium having stored thereon a computer program, It is characterized in that When the program is executed by a processor, the method according to any one of claims 1 to 4 or the method according to claim 5 is implemented.
13. A computer program product, It is characterized in that The invention comprises a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 4 or implements the steps of the method according to claim 5.
Citation Information
Patent Citations
Multi-language speech recognition method based on language type and speech content collaborative classification
CN110895932A
Speech feature extraction method and device thereof, terminal and storage medium
CN112767927A
Model training method and device, voice recognition method and device, electronic equipment and storage medium
CN112951240A