Speech recognition method and related device, electronic device, and storage medium

During the training of the speech recognition model, network parameters are adjusted based on the confusion degree of sample speech, and the problem of insufficient high-quality training data is solved, which improves recognition accuracy and reduces the training cost.

CN114566153BActive Publication Date: 2025-05-23UNIV OF SCI & TECH OF CHINA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210089871.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-05-23
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

During the training process, existing speech recognition models rely on high-quality training data, but the high-quality data is insufficient and the training cost is high, resulting in low recognition accuracy.

Method used

During the training of the speech recognition model, network parameters are adjusted based on the first score (confusion degree) of several sample recognition texts of sample speech, and the confusion degree is introduced as a scoring penalty for differentiation training to strengthen model learning.

Benefits of technology

It improves the recognition accuracy of the speech recognition model and reduces the training cost of the speech recognition model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114566153B_ABST
    Figure CN114566153B_ABST
Patent Text Reader

Abstract

The present application discloses a speech recognition method and related devices, electronic devices, and storage media, wherein the speech recognition method includes: obtaining a speech to be recognized; using a speech recognition model to recognize the speech to be recognized, and obtaining a recognition text of the speech to be recognized; wherein, during the training process, the speech recognition model adjusts network parameters based on the first scores of several sample recognition texts of the sample speech, the first scores represent the confusion of the sample recognition text, and the several sample recognition texts are all obtained by the speech recognition model recognizing the sample speech. The above scheme can improve the recognition accuracy of the speech recognition model and reduce the training cost of the speech recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and in particular to a speech recognition method and related devices, electronic equipment, and storage media. Background Art

[0002] With the rapid development of electronic information technology, machine recognition of speaker voice has been widely used in many scenarios such as conferences, sampling, teaching, human-computer interaction, etc.

[0003] At present, the performance of speech recognition models depends on high-quality training data. However, in real-world scenarios, high-quality training data is often not abundant. Although the existing model training methods can make up for the lack of high-quality training data to a certain extent, the training cost is high and the training effect is poor. In view of this, how to improve the recognition accuracy of speech recognition models and reduce the training cost of speech recognition models has become an urgent problem to be solved. Summary of the invention

[0004] The main technical problem solved by the present application is to provide a speech recognition method and related devices, electronic devices, and storage media, which can improve the recognition accuracy of the speech recognition model and reduce the training cost of the speech recognition model.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a speech recognition method, including: obtaining a speech to be recognized; using a speech recognition model to recognize the speech to be recognized, and obtaining a recognition text of the speech to be recognized; wherein, during the training process of the speech recognition model, the network parameters are adjusted based on the first scores of several sample recognition texts of the sample speech, the first score represents the confusion degree of the sample recognition text, and the several sample recognition texts are all obtained by the speech recognition model recognizing the sample speech.

[0006] In order to solve the above technical problems, the second aspect of the present application provides a speech recognition device, including: an acquisition module and a recognition module, the acquisition module is used to acquire the speech to be recognized; the recognition module is used to use a speech recognition model to recognize the speech to be recognized, and obtain the recognition text of the speech to be recognized; wherein, during the training process of the speech recognition model, the network parameters are adjusted based on the first scores of several sample recognition texts of the sample speech, the first score represents the degree of confusion of the sample recognition text, and the several sample recognition texts are all obtained by the speech recognition model recognizing the sample speech.

[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other, the memory storing program instructions, and the processor being used to execute the program instructions to implement the speech recognition method of the first aspect.

[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be executed by a processor, and the program instructions are used to implement the speech recognition method of the first aspect.

[0009] The above scheme obtains the speech to be recognized, and uses the speech recognition model to recognize the speech to be recognized to obtain the recognition text of the speech to be recognized. During the training process of the speech recognition model, the network parameters are adjusted based on the first scores of several sample recognition texts of the sample speech. The first score represents the confusion degree of the sample recognition text, and the several sample recognition texts are all obtained by the speech recognition model from recognizing the sample speech. That is, during the training process of the speech recognition model, the network parameters can be adjusted according to the confusion degree of each sample recognition text of the sample speech, so that the confusion degree can be additionally introduced as a scoring penalty to perform discriminative training, so as to strengthen the model learning, which is beneficial to improve the recognition accuracy of the speech recognition model and reduce the training cost of the speech recognition model. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a flowchart of an embodiment of the speech recognition method of the present application;

[0011] Figure 2 is a flow chart of an embodiment of training a speech recognition model;

[0012] Figure 3 It is a schematic diagram of the framework of an embodiment of the speech recognition device of the present application;

[0013] Figure 4 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0014] Figure 5 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0015] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.

[0016] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0017] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.

[0018] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the speech recognition method of the present application. Specifically, it may include the following steps:

[0019] Step S11: Acquire the speech to be recognized.

[0020] In one implementation scenario, the speech to be recognized can be collected in real time. For example, the speech data of the speaker can be collected in real time as the speech to be recognized, so as to perform speech recognition in real time. That is to say, in scenarios where the real-time requirements for speech recognition are high, real-time collection and recognition can be adopted, which can specifically include but are not limited to the following scenarios: meetings, human-computer interaction, etc. For example, in a meeting scenario, while the speaker is speaking, the speaker's speech data can be recognized to obtain a recognized text, so that when the participants miss the speaker's speech during the meeting due to unstable data transmission and other reasons, they can still learn the speaker's speech content in real time through the recognized text. Other scenarios can be deduced by analogy, and examples are not given one by one here.

[0021] In an implementation scenario, the speech to be recognized can also be collected in non-real time. For example, after the speaker's entire speech is collected, all of the above speech data can be used as the speech to be recognized for recognition. That is to say, in scenarios where real-time recognition is not required, the speech can be recognized uniformly after all the speech is collected. The specific scenarios may include but are not limited to the following scenarios: interviews, lectures, etc. For example, in an interview scenario, all the speech of the interviewer and the interviewee can be stored, and after the interview is completed, all the speech data can be recognized as the speech to be recognized to obtain the recognized text, so that a text version of the interview record can be automatically formed. Other scenarios can be deduced by analogy, and examples are not given one by one here.

[0022] Step S12: using the speech recognition model to recognize the speech to be recognized, and obtaining a recognition text of the speech to be recognized.

[0023] In one implementation scenario, the speech recognition model may include, but is not limited to: LSTM (Long Short Term Memory), RNN (Recurrent Neural Network), etc. Of course, the speech recognition model may specifically adopt an end-to-end model based on the Encoder-Decoder architecture, and the network structure of the speech recognition model is not limited here.

[0024] In the embodiment of the present disclosure, during the training process, the speech recognition model can adjust the network parameters based on the first scores of several sample recognition texts of the sample speech, the first scores representing the confusion of the sample recognition texts, and the several sample recognition texts are all obtained by the speech recognition model recognizing the sample speech. Exemplarily, the speech recognition model can recognize the sample speech and obtain M sample recognition texts, and then N of the sample recognition texts can be selected as the aforementioned several recognition texts. Of course, N is less than or equal to M, that is, only some of the sample recognition texts can be selected, or all of the sample recognition texts can be selected, which is not limited here.

[0025] In an implementation scenario, a pre-trained language model such as BERT (Bidirectional Encoder Representation from Transformers, i.e., Encoder of bidirectional Transformer) can be used to process several sample recognition texts respectively, and obtain the confusion degree (perplexity, PPL) of the sample recognition text as the first score of the sample recognition text. It should be noted that the confusion degree represents the confidence of the pre-trained language model in the sample recognition text. The higher the confusion degree, the lower the confidence of the pre-trained language model in the sample recognition text, that is, the pre-trained language model believes that the sample recognition text is more unreasonable. Conversely, the lower the confusion degree, the higher the confidence of the pre-trained language model in the sample recognition text, that is, the pre-trained language model believes that the sample recognition text is more reasonable. For details, please refer to the relevant technical details about confusion degree in the field of NLP (Natural Language Processing, i.e., natural language processing), which will not be repeated here.

[0026] In an implementation scenario, the aforementioned several sample texts (i.e., the aforementioned N sample recognition texts) can be specifically selected based on the second score of each sample recognition text (i.e., the aforementioned M sample recognition texts), and the second score can specifically represent the word error rate (Word Error Rate, WER) of the sample recognition text, and the word error rate can be obtained based on the difference between the sample recognition text and the sample annotated text of the sample speech. The specific calculation method of the word error rate can refer to the technical details of the word error rate, which will not be repeated here. In the above method, several sample recognition texts are selected based on the second score of each sample recognition text, and the second score represents the word error rate of the sample recognition text, and the word error rate is obtained based on the difference between the sample recognition text and the sample annotated text of the sample speech. Therefore, in the model training process, the recognition loss of some sample recognition texts can be focused on based on the word error rate, which is conducive to reducing the computational load and improving the efficiency of model training.

[0027] In a specific implementation scenario, each sample recognition text (i.e., the aforementioned M sample recognition texts) can be sorted in order of word error rate from low to high, and the sample recognition texts ranked in the top N positions are selected as the aforementioned sample recognition texts to participate in the subsequent loss calculation. The specific value of N can be set according to the actual situation, such as 50% of M, 40% of M, etc., which is not limited here.

[0028] In a specific implementation scenario, for the convenience of description, the sample speech can be recorded as x, and the sample annotation text of the sample speech can be recorded as y. * , the jth sample recognition text is recorded as y j , the second score (i.e., word error rate) of the jth sample recognition text is recorded as W(y j ,y * ).

[0029] In an implementation scenario, after selecting and obtaining several sample recognition texts, the first weights of several sample recognition texts can be obtained based on the first scores of the several sample recognition texts, and the recognition probability values ​​of the sample recognition texts can be weighted based on the first weights of the sample recognition texts to obtain the first loss, and the recognition probability value indicates the possibility that the speech recognition model predicts that the sample speech corresponding text is the sample recognition text. Exemplarily, the higher the recognition probability value, the greater the possibility that the speech recognition model predicts that the sample speech corresponding text is the sample recognition text, and conversely, the lower the recognition probability value, the smaller the possibility that the speech recognition model predicts that the sample speech corresponding text is the sample recognition text. On this basis, the network parameters of the speech recognition model can be adjusted based on the first loss. In the above method, the first weight of the sample recognition text is obtained by the first score of the sample recognition text, and the recognition probability value of the sample recognition text is weighted by the first weight to obtain the first loss, and then the network parameters of the speech recognition model are adjusted based on the first loss, which is equivalent to introducing confusion as the scaling factor of the recognition loss corresponding to each sample recognition text, thereby being able to enhance the distinguishability of the coverage and rationality of the language model, which is conducive to improving the overall effect of the speech recognition model.

[0030] In a specific implementation scenario, the first weight of the sample recognition text can be obtained based on the difference between the first score of the sample recognition text and the first reference value, and the first reference value represents the average of the first scores of several sample recognition texts. In the above method, the first weight of the sample recognition text is obtained based on the difference between the first score of the sample recognition text and the first reference value, and the first reference value represents the average of the first scores of several sample recognition texts, which can reduce the calculation complexity of the first weight and is conducive to improving the efficiency of model training.

[0031] In a specific implementation scenario, for ease of description, the first reference value can be recorded as The first score of the jth sample recognition text can be recorded as M(y j ,y * ). On this basis, the first weight of the jth sample recognition text can be expressed as

[0032] In a specific implementation scenario, as mentioned above, the sample recognition text can be processed using a pre-trained language model to obtain a first score of the sample recognition text. Specifically, the sample recognition model can be processed using a pre-trained language model to obtain a model score representing the degree of confusion, and the model scores of several sample recognition texts can be normalized to obtain the first score of the sample recognition text. For ease of description, the jth sample recognition text is still taken as an example, and its first score M(y j ,y * ) can be expressed as:

[0033]

[0034] In the above formula (1), N represents the total number of sample recognition texts.

[0035] In an implementation scenario, as mentioned above, several sample recognition texts can be selected based on the second scores of each sample recognition text, and the second scores represent the word error rate of the sample recognition text. On this basis, before adjusting the network parameters, the second weights of several sample recognition texts can be obtained based on the second scores of the several sample recognition texts, and the recognition probability values ​​of the sample recognition texts can be weighted based on the second weights of the sample recognition texts to obtain the second loss, so that the network parameters of the speech recognition model can be adjusted based on the first loss and the second loss. In the above method, during the model training process, not only confusion is introduced as a scaling factor for the recognition loss corresponding to each sample recognition text, but also the word error rate is further introduced as a scaling factor for the recognition loss corresponding to each sample recognition text, which can further enhance the distinguishability of the coverage and rationality of the language model, which is beneficial to improving the overall effect of the speech recognition model.

[0036] In a specific implementation scenario, similar to the aforementioned first weight, the second weight of the sample recognition text can be obtained based on the difference between the second score of the sample recognition text and the second reference value, and the second reference value represents the average of the second scores of several sample recognition texts. As mentioned above, the second score (i.e., word error rate) of the jth sample recognition text can be recorded as W(y j ,y * ), then the second reference value can be recorded as Therefore, the second weight of the jth sample recognition text can be expressed as The above method obtains the second weight of the sample recognition text based on the difference between the second score of the sample recognition text and the second reference value, and the second reference value represents the average value of the second scores of several sample recognition texts. This can reduce the calculation complexity of the second weight and is conducive to improving the model training efficiency.

[0037] In a specific implementation scenario, an optimization method such as gradient descent can be used to adjust the network parameters of the speech recognition model based on the first loss and the second loss. The specific adjustment process of the network parameters can refer to the technical details of the optimization method such as gradient descent, which will not be repeated here.

[0038] In a specific implementation scenario, in order to balance the first loss and the second loss, a balance coefficient can be pre-set. On this basis, the product of the first loss and the balance coefficient and the sum of the second loss can be used as the total loss of the speech recognition model, so that the network parameters of the speech recognition model can be adjusted based on the total loss. For the convenience of description, the balance coefficient can be recorded as β, then the total loss L(x,y * ) can be expressed as:

[0039]

[0040] In the above formula (2), represents the first loss, represents the second loss, Beam(x,N) represents decoding the input sample speech x and selecting N samples with the best WER (i.e., the aforementioned second score) for recognition text, y j ∈Beam(x,N) represents the jth sample recognition text among the N sample recognition texts with the best WER. It should be noted that although there are two However, in actual processing, decoding is performed only once.

[0041] In an implementation scenario, the sample speech may include at least a first speech and a second speech, the first speech is synthesized from a first annotated text, the first annotated text is screened from a number of first candidate texts, the second speech is marked as a second annotated text, the second speech is obtained from real collection, the speech recognition model is pre-trained based on the second language before training, the first candidate texts are texts other than the second candidate texts among the candidate texts, the domain category of the second candidate texts is the target domain, and the number of second annotated texts belonging to the target domain is less than a preset threshold. In the above method, since the number of second annotated texts belonging to the target domain is less than the preset threshold, the first candidate text is regarded as non-domain data or domain data that does not belong to the key reinforcement, the first annotated text is further selected from the first candidate text, and the first speech is synthesized from the first annotated text to retrain the speech recognition model, which is conducive to further improving the deficiencies of the speech recognition model after training with the second speech and its second annotated text through the first speech and its first annotated text, and improving the recognition effect of the speech recognition model.

[0042] In a specific implementation scenario, several candidate texts can be collected in advance from many fields such as e-commerce retail, education, finance, military, technology, enterprise, sports, medical care, entertainment, games, history and humanities, etc. For example, several candidate texts can be extracted by web crawling, book extraction, etc., which is not limited here.

[0043] In a specific implementation scenario, a domain classification model can be used to classify the candidate text in the field to obtain the field category of the candidate text. At the same time, the domain classification model can be used to classify the second annotated text in the field to obtain the field category of the second annotated text. On this basis, the field category can be used as a statistical dimension to count the number (or proportion) of the second annotated text belonging to each field category. Exemplarily, the following can be statistically obtained: the proportions of e-commerce retail, education, finance, military, science and technology, enterprise, sports, medical, entertainment, games, history and humanities and other field categories are: 20%, 10%, 20%, 10%, 10%, 10%, 10%, 5%, 2%, 2%, 1%, that is, the second annotated text covers less in the above-mentioned medical, entertainment, games, history and humanities four field categories, so the above-mentioned four field categories can be used as the target field, and the candidate text belonging to any of the above-mentioned four field categories can be used as the second candidate text, and the candidate text other than the second candidate text can be used as the first candidate text. It should be noted that the field classification model can include but is not limited to convolutional neural networks, BERT, etc., which are not limited here. Specifically, in order to improve the accuracy of domain classification, sample texts annotated with sample domain categories can be collected in advance, and the sample texts can be classified using a domain classification model to obtain the predicted domain category of the sample text, and the network parameters of the domain classification model can be adjusted based on the difference between the sample domain category and the predicted domain category. For the specific measurement method of the above difference, please refer to the technical details of loss functions such as cross entropy, and for the specific adjustment process of the above parameters, please refer to the technical details of optimization methods such as gradient descent, which will not be repeated here.

[0044] In a specific implementation scenario, after selecting the first candidate text, the first candidate text can be further selected as the first annotated text based on the difference between the first sample score of the first candidate text and the first sample score of the speech recognition text. It should be noted that the speech recognition text is obtained by the speech recognition model for the synthetic speech recognition of the first candidate text, and the first sample score represents the confusion. In the above manner, the first candidate text is selected as the first annotated text by the difference between the confusion between the first candidate text and the speech recognition text, which can reduce the participation of synthetic data in the training process and is conducive to improving the model training effect. In addition, as mentioned above, the first score can be obtained by the pre-trained language model, and similarly, the first sample score can also be obtained by the pre-trained language model, that is, the first score and the first sample score can be obtained by the pre-trained language model, so the pre-trained language model participates in sample selection and model training at the same time, so that the model training can be guided by the pre-trained language model, which is conducive to improving the model training effect. Specifically, the first candidate text can be selected as the rough selection text based on the second sample score of the first candidate text, and the second sample numerator represents the word error rate of the speech recognition text corresponding to the first candidate text. Exemplarily, if the word error rate of the speech recognition text corresponding to the first candidate text is higher than a preset threshold value (e.g., 5%), the first candidate text can be selected as the rough selection text, that is, if the speech recognition model has a good recognition coverage of the first candidate text, there is no need to use the first candidate text as a training supplement for the speech recognition model, and if the speech recognition model has a poor recognition coverage of the first candidate text, the first candidate text can be used as a rough selection text to further examine whether the first candidate text is used as a training supplement for the speech recognition model. Of course, the above preset threshold can be adjusted according to actual conditions, such as it can also be set to 10%, 15%, etc., which is not limited here. Further, in response to the first sample score of the rough selection text being lower than the first sample score of the speech recognition text, the rough selection text can be selected as the first annotated text, otherwise if the first sample score of the rough selection text is not lower than the first sample score of the speech recognition text, there is no need to select the rough selection text as the first annotated text. That is to say, for the roughly selected text, if its first sample score is lower than the first sample score of the speech recognition text, it can be considered that the pre-trained language model has a higher degree of confidence in the roughly selected text, while the rationality of the speech recognition text is relatively low, that is, the recognition coverage of the speech recognition model is relatively poor, so the roughly selected text can be selected as the first annotated text; conversely, if its first sample score is not lower than the first sample score of the speech recognition text, it can be considered that the pre-trained language model has a higher degree of confidence in the speech recognition text, that is, the recognition coverage of the speech recognition model is relatively good, so there is no need to select the roughly selected text as a training supplement for the speech recognition model.In the above method, the first candidate text is selected as the rough selected text based on the second sample score of the first candidate text, and the second sample score represents the word error rate of the speech recognition text corresponding to the first candidate text. On this basis, in response to the first sample score of the rough selected text being lower than the first sample score of the speech recognition text, the rough selected text is selected as the first annotated text. Therefore, the first candidate text with poor coverage by the speech recognition model can be selected from the aspects of word error rate and confusion as a subsequent training supplement for the speech recognition model, which is conducive to strengthening the improvement of unacceptable errors in decoding results during the discriminative training process.

[0045] In a specific implementation scenario, the sample speech may further include a third speech, which is synthesized from a third annotated text, and the third annotated text source domain is the aforementioned several second candidate texts. That is to say, for the field categories that the speech recognition model does not cover well in the pre-training stage, the candidate texts belonging to these field categories can be selected as the third annotated text, and speech synthesis is performed based on the third annotated text to obtain the third speech, which is then incorporated into the sample speech, thereby serving as a subsequent training supplement for the speech recognition model. In the above manner, the sample speech also includes a third speech, which is synthesized from a third annotated text, and the third annotated text is derived from several second candidate texts, which can further strengthen the training of the speech recognition model in key areas and is conducive to improving the recognition effect of the speech recognition model.

[0046] The above scheme obtains the speech to be recognized, and uses the speech recognition model to recognize the speech to be recognized to obtain the recognition text of the speech to be recognized. During the training process of the speech recognition model, the network parameters are adjusted based on the first scores of several sample recognition texts of the sample speech. The first score represents the confusion degree of the sample recognition text, and the several sample recognition texts are all obtained by the speech recognition model from recognizing the sample speech. That is, during the training process of the speech recognition model, the network parameters can be adjusted according to the confusion degree of each sample recognition text of the sample speech, so that the confusion degree can be additionally introduced as a scoring penalty to perform discriminative training, so as to strengthen the model learning, which is beneficial to improve the recognition accuracy of the speech recognition model and reduce the training cost of the speech recognition model.

[0047] Please refer to Figure 2 , Figure 2 1 is a schematic diagram of a process of training a speech recognition model according to an embodiment. Specifically, the process may include the following steps:

[0048] Step S201: pre-training a speech recognition model based on a second speech marked with a second annotation text in the sample speech.

[0049] For details, please refer to the relevant description in the aforementioned disclosed embodiments.

[0050] Step S202: Obtain several candidate texts.

[0051] For details, please refer to the relevant description in the aforementioned disclosed embodiments.

[0052] Step S203: Based on the domain category of each candidate text and the domain category of each second annotated text, select the candidate text belonging to the target domain as the second candidate text, and select the candidate text other than the second candidate text as the first candidate text.

[0053] Specifically, the number of second annotated texts belonging to the target field may be less than a preset threshold (or the proportion is lower than a preset threshold). For details, please refer to the relevant description in the aforementioned disclosed embodiment, which will not be repeated here.

[0054] Step S204: obtaining a third annotated text based on the second candidate text.

[0055] Specifically, at least part of the second candidate text can be selected as the third annotated text. Exemplarily, only part of the second candidate text can be selected as the third annotated text. Of course, all of the second candidate text can also be selected as the third annotated text, which is not limited here.

[0056] Step S205: using a speech recognition model to recognize the synthesized speech of the first candidate text to obtain a speech recognition text.

[0057] For details, please refer to the relevant description in the aforementioned disclosed embodiments, which will not be repeated here.

[0058] Step S206: based on the second sample score of the first candidate text, select the first candidate text as the roughly selected text.

[0059] Specifically, the second sample score represents the word error rate of the speech recognition text corresponding to the first candidate text. For details, please refer to the relevant description in the aforementioned public embodiment, which will not be repeated here.

[0060] Step S207: In response to the first sample score of the roughly selected text being lower than the first sample score of the speech recognition text, the roughly selected text is selected as the first annotated text.

[0061] Specifically, the first sample score represents the degree of confusion, and details may be found in the related description in the aforementioned disclosed embodiment, which will not be repeated here.

[0062] Step S208: synthesize a first speech based on the first annotated text, synthesize a third speech based on the third annotated text, and obtain a sample speech marked with the sample annotated text based on the first speech marked with the first annotated text, the second speech marked with the second annotated text, and the third speech marked with the third annotated text.

[0063] Specifically, in a real scene, the ratio between the second voice collected in the sample voice, the synthesized first voice and the third voice can be controlled. Exemplarily, the second voice collected in the sample voice, the synthesized first voice and the third voice can be controlled to a quantitative ratio of 1:1. Of course, other quantitative ratios (such as 1:1.1, 1:1.2, etc.) can also be controlled, which is not limited here. For details, please refer to the relevant description in the aforementioned disclosed embodiment, which will not be repeated here.

[0064] Step S209: Perform speech recognition on the sample speech using the speech recognition model to obtain a number of sample recognition texts.

[0065] Specifically, several sample recognition texts are selected based on the second scores of the sample recognition texts, the second scores represent the word error rate of the sample recognition texts, and the word error rate is obtained based on the difference between the sample recognition texts and the sample annotation texts of the sample speech. For details, please refer to the relevant description in the aforementioned disclosed embodiment, which will not be repeated here.

[0066] Step S210: based on the first scores of the sample recognition texts, obtaining the first weights of the sample recognition texts, and weighting the recognition probability values ​​of the sample recognition texts based on the first weights of the sample recognition texts to obtain a first loss.

[0067] Specifically, the recognition probability value indicates the possibility that the speech recognition model predicts that the text corresponding to the sample speech is the sample recognition text. In addition, the first weight of the sample recognition text is obtained based on the difference between the first score of the sample recognition text and the first reference value, and the first reference value represents the average of the first scores of several sample recognition texts. For details, please refer to the relevant description in the aforementioned disclosed embodiment, which will not be repeated here.

[0068] Step S211: based on the second scores of the sample recognition texts, obtaining the second weights of the sample recognition texts, and weighting the recognition probability values ​​of the sample recognition texts based on the second weights of the sample recognition texts to obtain a second loss.

[0069] Specifically, the second weight of the sample recognition text is obtained based on the difference between the second score of the sample recognition text and the second reference value, and the second reference value represents the average value of the second scores of several sample recognition texts. For details, please refer to the relevant description in the aforementioned disclosed embodiment, which will not be repeated here.

[0070] Step S212: Adjust the network parameters of the speech recognition model based on the first loss and the second loss.

[0071] For details, please refer to the relevant description in the aforementioned disclosed embodiments, which will not be repeated here.

[0072] The above scheme screens text corpora based on the speech recognition model and the pre-trained language model, that is, selects text corpora with poor coverage of the current speech recognition model for speech synthesis and participates in subsequent discriminative training, which can minimize the frequency of using synthetic data during training. At the same time, the discriminative training introduces the model score (that is, the first score) of the pre-trained language model, so that the pre-trained language model can participate in the training in the form of the first score, which can not only enhance the contrast between language models and improve the coverage of language models, but also because the above process is only used during training, and after the training is completed, only the speech recognition model is used to recognize the speech to be recognized to obtain the recognized text, that is, no computational cost is added in the reasoning stage, which can be more in line with the efficiency requirements of actual usage scenarios.

[0073] See also Figure 3 , Figure 3 It is a schematic diagram of a framework of an embodiment of a speech recognition device 30 of the present application. The speech recognition device 30 may include: an acquisition module 31 and a recognition module 32, wherein the acquisition module 31 is used to acquire the speech to be recognized; the recognition module 32 is used to recognize the speech to be recognized using a speech recognition model to obtain a recognition text of the speech to be recognized; wherein, during the training process, the speech recognition model adjusts the network parameters based on the first scores of several sample recognition texts of the sample speech, the first score represents the degree of confusion of the sample recognition text, and the several sample recognition texts are all obtained by the speech recognition model recognizing the sample speech.

[0074] In the above scheme, during the training process of the speech recognition model, the network parameters can be adjusted according to the confusion degree of each sample recognition text of the sample speech, so that the confusion degree can be additionally introduced as a scoring penalty to perform discriminative training to strengthen model learning, which is beneficial to improve the recognition accuracy of the speech recognition model and reduce the training cost of the speech recognition model.

[0075] In some disclosed embodiments, the speech recognition device 30 also includes a first weight acquisition module, which is used to obtain the first weights of several sample recognition texts based on the first scores of the several sample recognition texts; the speech recognition device 30 also includes a first loss calculation module, which is used to weight the recognition probability values ​​of the sample recognition texts based on the first weights of the sample recognition texts to obtain a first loss; wherein the recognition probability value represents the possibility that the speech recognition model predicts that the text corresponding to the sample speech is the sample recognition text; the speech recognition device 30 also includes a network parameter adjustment module, which is used to adjust the network parameters of the speech recognition model based on the first loss.

[0076] Therefore, the first weight of the sample recognition text is obtained through the first score of the sample recognition text, and the recognition probability value of the sample recognition text is weighted by the first weight to obtain the first loss, and then the network parameters of the speech recognition model are adjusted based on the first loss, which is equivalent to introducing confusion as a scaling factor for the recognition loss corresponding to each sample recognition text, thereby enhancing the distinguishability of the coverage and rationality of the language model, which is conducive to improving the overall effect of the speech recognition model.

[0077] In some disclosed embodiments, the first weight of the sample recognition text is obtained based on the difference between the first score of the sample recognition text and a first reference value, and the first reference value represents an average of the first scores of several sample recognition texts.

[0078] Therefore, based on the difference between the first score of the sample recognition text and the first reference value, the first weight of the sample recognition text is obtained, and the first reference value represents the average value of the first scores of several sample recognition texts, which can reduce the calculation complexity of the first weight and help improve the model training efficiency.

[0079] In some disclosed embodiments, several sample recognition texts are selected based on a second score of each sample recognition text, the second score represents a word error rate of the sample recognition text, and the word error rate is obtained based on the difference between the sample recognition text and a sample annotated text of the sample speech.

[0080] Therefore, several sample recognition texts are selected based on the second scores of each sample recognition text. The second score represents the word error rate of the sample recognition text, and the word error rate is obtained based on the difference between the sample recognition text and the sample annotated text of the sample speech. Therefore, in the model training process, the recognition loss of some sample recognition texts can be focused on based on the word error rate, which is conducive to reducing the computational load and improving the model training efficiency.

[0081] In some disclosed embodiments, the speech recognition device 30 also includes a second weight acquisition module, which is used to obtain second weights of several sample recognition texts based on the second scores of the several sample recognition texts; the speech recognition device 30 also includes a second loss calculation module, which is used to weight the recognition probability values ​​of the sample recognition texts based on the second weights of the sample recognition texts to obtain a second loss; the network parameter adjustment module is specifically used to adjust the network parameters of the speech recognition model based on the first loss and the second loss.

[0082] Therefore, during the model training process, not only the confusion degree is introduced as the scaling factor of the recognition loss corresponding to each sample recognition text, but the word error rate is further introduced as the scaling factor of the recognition loss corresponding to each sample recognition text, which can further enhance the distinguishability of the coverage and rationality of the language model, and is conducive to improving the overall effect of the speech recognition model.

[0083] In some disclosed embodiments, the sample speech includes at least a first speech and a second speech, the first speech is synthesized from a first annotated text, the first annotated text is screened from a number of first candidate texts, the second speech is marked with the second annotated text, the second speech is obtained by real collection, the speech recognition model is pre-trained based on the second speech before training, the several first candidate texts are texts other than several second candidate texts among the several candidate texts, the domain category of the second candidate texts is the target domain, and the number of second annotated texts belonging to the target domain is less than a preset threshold.

[0084] Therefore, since the number of second annotated texts belonging to the target domain is less than the preset threshold, the first candidate text is regarded as non-domain data or domain data that does not belong to the key reinforcement, and the first annotated text is further selected from the first candidate text, and the first speech is synthesized by the first annotated text to retrain the speech recognition model. This is conducive to further improving the deficiencies of the speech recognition model after training with the second speech and its second annotated text through the first speech and its first annotated text, thereby improving the recognition effect of the speech recognition model.

[0085] In some disclosed embodiments, the first annotated text is selected based on the difference between a first sample score of a first candidate text and a first sample score of a speech recognition text, the speech recognition text is obtained by a speech recognition model performing synthetic speech recognition on the first candidate text, and the first sample score represents confusion.

[0086] Therefore, by selecting the first candidate text as the first annotated text based on the difference in confusion between the first candidate text and the speech recognition text, the participation of synthetic data in the training process can be reduced, which is beneficial to improving the model training effect.

[0087] In some disclosed embodiments, the speech recognition device 30 also includes a text coarse selection module for selecting the first candidate text as the coarse selected text based on the second sample score of the first candidate text; wherein the second sample score represents the word error rate of the speech recognition text corresponding to the first candidate text; the speech recognition device 30 also includes a text fine selection module for selecting the coarse selected text as the first annotated text in response to the first sample score of the coarse selected text being lower than the first sample score of the speech recognition text.

[0088] Therefore, based on the second sample score of the first candidate text, the first candidate text is selected as the rough selected text, and the second sample score represents the word error rate of the speech recognition text corresponding to the first candidate text. On this basis, in response to the first sample score of the rough selected text being lower than the first sample score of the speech recognition text, the rough selected text is selected as the first annotated text. Therefore, the first candidate text with poor coverage by the speech recognition model can be selected from the aspects of word error rate and confusion as a subsequent training supplement for the speech recognition model, which is conducive to improving the unacceptable errors in the decoding results during the enhanced discriminative training process.

[0089] In some disclosed embodiments, the first score and the first sample score are both obtained by a pre-trained language model.

[0090] Therefore, the pre-trained language model is involved in both sample selection and model training, so that model training can be guided by the pre-trained language model, which is conducive to improving the model training effect.

[0091] In some disclosed embodiments, the sample speech also includes a third speech, where the third speech is synthesized from a third annotated text, and the third annotated text is derived from a number of second candidate texts.

[0092] Therefore, the sample speech also includes a third speech, and the third speech is synthesized from a third annotated text. The third annotated text is derived from several second candidate texts, which can further strengthen the training of the speech recognition model on key areas and is conducive to improving the recognition effect of the speech recognition model.

[0093] See also Figure 4 , Figure 4 1 is a schematic diagram of a framework of an embodiment of an electronic device 40 of the present application. The electronic device 40 includes a memory 41 and a processor 42 coupled to each other, the memory 41 stores program instructions, and the processor 42 is used to execute the program instructions to implement the steps in any of the above-mentioned speech recognition method embodiments. Specifically, the electronic device 40 may include but is not limited to: a desktop computer, a laptop computer, a server, a mobile phone, a tablet computer, etc., which are not limited here.

[0094] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned speech recognition method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.

[0095] In the above scheme, during the training process of the speech recognition model, the network parameters can be adjusted according to the confusion degree of each sample recognition text of the sample speech, so that the confusion degree can be additionally introduced as a scoring penalty to perform discriminative training to strengthen model learning, which is beneficial to improve the recognition accuracy of the speech recognition model and reduce the training cost of the speech recognition model.

[0096] See also Figure 5 , Figure 5 1 is a schematic diagram of a framework of an embodiment of a computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps in any of the above-mentioned speech recognition method embodiments.

[0097] In the above scheme, during the training process of the speech recognition model, the network parameters can be adjusted according to the confusion degree of each sample recognition text of the sample speech, so that the confusion degree can be additionally introduced as a scoring penalty to perform discriminative training to strengthen model learning, which is beneficial to improve the recognition accuracy of the speech recognition model and reduce the training cost of the speech recognition model.

[0098] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0099] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0100] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0101] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0102] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0103] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

Claims

1. A speech recognition method, It is characterized in that include: Get the speech to be recognized; Recognize the speech to be recognized by using a speech recognition model to obtain a recognition text of the speech to be recognized; During the training process, the speech recognition model adjusts network parameters based on the first scores of several sample recognition texts of the sample speech, the first scores representing the confusion degree of the sample recognition texts, the several sample recognition texts are all obtained by the speech recognition model recognizing the sample speech, and the method for selecting the several sample recognition texts includes: obtaining the second score of the sample recognition text based on the difference between the sample recognition text and the sample annotated text of the sample speech; wherein the second score represents the word error rate of the sample recognition text; and selecting the several sample recognition texts based on the second scores of each of the sample recognition texts.

2. The method according to claim 1, It is characterized in that The training steps of the speech recognition model include: Based on the first scores of the plurality of sample recognition texts, obtaining first weights of the plurality of sample recognition texts; The recognition probability value of the sample recognition text is weighted based on the first weight of the sample recognition text to obtain a first loss; wherein the recognition probability value represents the possibility that the speech recognition model predicts that the text corresponding to the sample speech is the sample recognition text; Based on the first loss, adjust the network parameters of the speech recognition model.

3. The method according to claim 2, It is characterized in that The first weight of the sample recognition text is obtained based on the difference between the first score of the sample recognition text and a first reference value, and the first reference value represents an average value of the first scores of the plurality of sample recognition texts.

4. The method according to claim 2, It is characterized in that Before adjusting the network parameters of the speech recognition model based on the first loss, the method further includes: Based on the second scores of the plurality of sample recognition texts, obtaining second weights of the plurality of sample recognition texts; Weighting the recognition probability value of the sample recognition text based on the second weight of the sample recognition text to obtain a second loss; The adjusting the network parameters of the speech recognition model based on the first loss includes: Based on the first loss and the second loss, the network parameters of the speech recognition model are adjusted.

5. The method according to claim 4, It is characterized in that The second weight of the sample recognition text is obtained based on the difference between the second score of the sample recognition text and a second reference value, and the second reference value represents an average value of the second scores of the sample recognition texts.

6. The method according to claim 1, It is characterized in that The sample speech includes at least a first speech and a second speech, the first speech is synthesized from a first annotated text, the first annotated text is screened from a number of first candidate texts, the second speech is marked with a second annotated text, the second speech is obtained by real collection, the speech recognition model is pre-trained based on the second speech before training, the several first candidate texts are texts other than several second candidate texts among several candidate texts, the domain category of the second candidate text is a target domain, and the number of second annotated texts belonging to the target domain is less than a preset threshold.

7. The method according to claim 6, It is characterized in that The first annotated text is selected based on the difference between a first sample score of the first candidate text and a first sample score of a speech recognition text, the speech recognition text is obtained by the speech recognition model performing synthetic speech recognition on the first candidate text, and the first sample score represents confusion.

8. The method according to claim 7, It is characterized in that The first marked text screening step includes: Based on the second sample score of the first candidate text, selecting the first candidate text as the rough selection text; wherein the second sample score represents the word error rate of the speech recognition text corresponding to the first candidate text; In response to the first sample score of the roughly selected text being lower than the first sample score of the speech recognition text, the roughly selected text is selected as the first annotated text.

9. The method according to claim 7, It is characterized in that The first score and the first sample score are both obtained by a pre-trained language model.

10. The method according to claim 6, It is characterized in that The sample speech also includes a third speech, and the third speech is synthesized from a third annotated text, and the third annotated text is derived from the plurality of second candidate texts.

11. A speech recognition device, It is characterized in that include: An acquisition module, used to acquire the speech to be recognized; A recognition module, used to recognize the speech to be recognized using a speech recognition model to obtain a recognition text of the speech to be recognized; In which, during the training process, the speech recognition model adjusts network parameters based on the first scores of several sample recognition texts of the sample speech, the first scores represent the degree of confusion of the sample recognition texts, the several sample recognition texts are all obtained by the speech recognition model recognizing the sample speech, and the method for selecting the several sample recognition texts includes: obtaining the second score of the sample recognition text based on the difference between the sample recognition text and the sample annotated text of the sample speech; wherein the second score represents the word error rate of the sample recognition text; and selecting the several sample recognition texts based on the second scores of each of the sample recognition texts.

12. An electronic device, It is characterized in that It comprises a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the speech recognition method according to any one of claims 1 to 10.

13. A computer-readable storage medium, It is characterized in that Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the speech recognition method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Audio recognition method and system, machine equipment and computer readable medium

    CN110517666A