Speech recognition model training method, speech recognition method and related device
By including speech-text pairs of real speech and mixed speech in the training data and training the speech recognition model, the problem of inaccurate real speech recognition in mixed speech scenarios is solved, and the accuracy of voice interaction and user experience are improved.
Patent Information
- Application Number
- CN202410339384.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-09-23
AI Technical Summary
Existing end-to-end speech recognition models cannot accurately recognize users' real voices in mixed voice scenarios, resulting in a poor voice interaction experience.
By using speech-text pairs containing real speech and mixed speech in the training data, the speech recognition model is trained to enable it to recognize real speech in mixed speech and reduce the impact of synthetic speech on real speech recognition.
The recognition accuracy of real speech in mixed voice scenarios is improved, ensuring the accuracy of voice interaction and user experience.
Smart Images

Figure CN120690180A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a speech recognition model training method, a speech recognition method and related devices. Background Art
[0002] Automatic Speech Recognition (ASR), a branch of speech technology, refers to speech technology that converts audio data into text. With the advancement of machine learning, ASR is now commonly implemented using end-to-end speech recognition models. These models offer excellent recognition performance, require no additional language model, have low model complexity, and are easy to deploy. Text-to-Speech (TTS), another branch of speech technology, refers to technology that converts computer-generated or externally input text into fluent speech. Synthesized speech is commonly used in human-computer interaction scenarios, such as intelligent customer service and navigation.
[0003] In some scenarios, synthesized speech poses a challenge to speech recognition accuracy. For example, in navigation, when an application with integrated navigation functionality uses synthesized speech to broadcast navigation information, if a user uses voice interaction, the voice data collected by the microphone may be a mixture of the user's real speech and the synthesized speech emitted by the application. The inventors of this application have discovered that the synthesized speech in this mixture can affect the end-to-end speech recognition model's ability to recognize real speech, resulting in an inability to understand the user's true intentions during voice interaction, leading to a poor user experience. Therefore, there is an urgent need to provide a speech recognition model that can improve the accuracy of recognizing the user's real speech from this mixture. Summary of the Invention
[0004] The present application provides a training method for a speech recognition model, a speech recognition method and related devices. Through the provided training method, the trained speech recognition model can only output the recognition text of the real speech in the mixed speech in a mixed speech recognition scenario, effectively avoiding the influence of the synthesized speech in the mixed speech on the real speech recognition in the mixed speech, and improving the accuracy of real speech recognition in the mixed speech scenario.
[0005] In a first aspect, the present application provides a method for training a speech recognition model, comprising:
[0006] Acquire model training data, the model training data comprising: speech-text pairs of real speech and speech-text pairs of mixed speech, wherein a speech-text pair comprises a speech sample and its corresponding text, the speech sample in the speech-text pair of mixed speech is a mixed speech sample, the mixed speech sample is a speech sample of mixed real speech and synthesized speech, the text corresponding to the mixed speech sample is the text corresponding to the real speech in the mixed speech sample, and the speech sample in the speech sample pair of real speech is real speech;
[0007] The model training data composed of the speech-text pairs is used as input to the speech recognition model to train the speech recognition model.
[0008] In a second aspect, the present application provides a speech recognition method, comprising:
[0009] Acquire speech data to be recognized; input the speech data to be recognized into a trained speech recognition model, obtain and output the recognition text of the speech data to be recognized; wherein, the speech recognition model is trained based on the method provided in the first aspect.
[0010] In a third aspect, the present application provides a training device for a speech recognition model, comprising:
[0011] A training data acquisition module is used to acquire model training data, wherein the model training data includes speech-text pairs of real speech and speech-text pairs of mixed speech, wherein a speech-text pair includes a speech sample and its corresponding text, the speech sample in the speech-text pair of mixed speech is a mixed speech sample, the mixed speech sample is a speech sample of mixed real speech and synthesized speech, the text corresponding to the mixed speech sample is the text corresponding to the real speech in the mixed speech sample, and the speech sample in the speech sample pair of real speech is real speech;
[0012] The model training module is used to use the model training data composed of the speech-text pairs as input to the speech recognition model to train the speech recognition model.
[0013] In a fourth aspect, the present application provides a speech recognition device, comprising:
[0014] A data acquisition module, used to acquire speech data to be recognized;
[0015] A speech recognition module is used to input the speech data to be recognized into a trained speech recognition model to obtain and output the recognition text of the speech data to be recognized; wherein, the speech recognition model is trained based on the method provided in the first aspect.
[0016] In a fifth aspect, the present application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to perform the method provided in the first or second aspect of the present application.
[0017] In a sixth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method provided in the first aspect or the second aspect of the present application is implemented.
[0018] In a seventh aspect, the present application provides a program product, comprising a computer program, which, when executed by a processor, implements the method provided in the first or second aspect of the present application.
[0019] The present application provides a speech recognition model training method, a speech recognition method and related devices. In the model training stage, the speech-text pairs in the model training data include speech-text pairs of mixed speech and speech-text pairs of real speech. Since the text corresponding to the mixed speech sample in the speech-text pair of mixed speech is only the text corresponding to the real speech in the mixed speech sample, the speech recognition model trained by the above-mentioned model training data can have the ability to recognize the real speech in the mixed speech, so that the recognition result of the mixed speech output by the model is close to the recognition result of the real speech in the mixed speech, reducing the influence of synthetic speech on real speech recognition, and improving the accuracy of real speech recognition in the scenario where synthetic speech and real speech are mixed. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0021] Figure 1 A schematic diagram of a scenario for performing voice recognition on a smartphone provided in an embodiment of the present application;
[0022] Figure 2 A flowchart of a method for training a speech recognition model provided in an embodiment of the present application;
[0023] Figure 3 For this application Figure 2 A schematic diagram of a speech recognition model provided by the illustrated embodiment;
[0024] Figure 4 A flowchart of another method for training a speech recognition model provided in an embodiment of the present application;
[0025] Figure 5 For this application Figure 4 A schematic diagram of the structure of the speech recognition model provided in the illustrated embodiment;
[0026] Figure 6 A flowchart of a speech recognition method provided in an embodiment of the present application;
[0027] Figure 7 A schematic diagram of the structure of a speech recognition model training device provided in an embodiment of the present application;
[0028] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0029] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0030] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0031] First, some of the terms involved in this application are explained:
[0032] Synthetic Speech: Speech synthesized using TTS technology or other speech synthesis technologies, such as the voice of a robot, the voice of a voice assistant, the voice of a navigation announcement, etc.
[0033] Real speech: non-synthetic speech, the sound of a natural person.
[0034] Hybrid speech: A speech that is a mixture of synthesized speech and real speech.
[0035] End-to-End Speech Recognition Model: Unlike traditional speech recognition models, this model does not distinguish between acoustic and language models. Instead, it is a holistic processing architecture that converts input speech signals into text, simplifying the model's complexity and reducing the workload of model training.
[0036] Prompt: Also known as a prompt template, it is a specific instruction or prompt message. Prompt tags assist model training by providing additional information. In speech recognition scenarios, by adding prompt tags to the text of speech sample pairs, you can indicate the type of speech input to the speech recognition model.
[0037] Figure 1 A schematic diagram of a scenario for performing voice recognition on a smartphone is provided for an embodiment of the present application. If an APP (Application) installed on the smartphone supports the voice interaction function, when the user turns on the voice interaction function of the APP, the APP will collect the sound (real voice) emitted by the user through the microphone of the smartphone, that is, the APP will collect audio data through the microphone of the smartphone (not shown in the figure). When there are no other sound-generating devices around the user or the smartphone does not make any sound (such as playing songs, playing navigation voice, etc.), the audio data collected by the microphone is the real voice emitted by the user (natural person). If there are other sound-generating devices around the user or the smartphone is making a sound, the audio data collected by the microphone before the user makes a sound may be the synthesized voice output by the speech synthesis technology. There is also a situation where there are other devices around the user or the smartphone is also broadcasting a synthesized voice while the user is making a real voice. At this time, the audio data collected by the microphone is a mixed voice of real voice and synthesized voice.
[0038] When the voice collected by a smartphone's microphone is a mixed voice, the app or its server uses voice recognition technology, such as an end-to-end voice recognition model, to recognize the mixed voice. For example, if the user's actual voice is "What's the weather like today?" and the synthesized voice from other devices around the user or the smartphone is "It's 4 p.m. now," the recognition result of the mixed voice might be "What's the weather like today at 4 p.m.?" This can result in the app being unable to effectively recognize the user's actual voice, and the voice interaction function being unable to understand the user's true intent.
[0039] Exemplarily, the APP may be an application with a map navigation function, a voice assistant application, or other applications with a voice interaction function.
[0040] Taking human-computer interaction as an example, the actual user voice can be a question asked by the user, while the synthesized voice can be the response from an app, such as a voice assistant, through the smartphone's speaker. When the voice collected by the app through the smartphone's microphone is a mixed voice, such as a mixture of the user's current question and the app's previous response, the app cannot correctly understand the user's true intent, resulting in inaccurate responses or even an inability to respond, affecting the human-computer interaction experience.
[0041] Taking navigation as an example, a smartphone installed with a navigation application (an example of an app) uses a synthesized voice to announce navigation. While the navigation is being announced, the user may also interact with the app's voice assistant. In this case, the audio data collected by the smartphone's microphone and transmitted to the app will include both the navigation voice and the user's actual speech. This audio data is a mixture of both the navigation voice and the user's actual speech. The navigation voice will affect the app's voice assistant's ability to recognize the user's actual speech, and thus its understanding of the user's true intent, resulting in a failure to deliver the desired interaction results.
[0042] The above uses smartphones as an example to illustrate the application scenarios of this application. However, whether it is a smartphone, a car, a robot, an IOT (Internet of Things) device, or other devices with integrated microphones and speakers, as long as they are installed with software / applications that support voice interaction functions, they will face the aforementioned problems. Therefore, the above scenario examples should not be regarded as limitations on this application.
[0043] For the above scenarios, the existing end-to-end speech recognition model cannot accurately output the content corresponding to the real speech in the mixed speech when recognizing mixed speech. Based on this, in order to improve the accuracy of the model in recognizing the real speech in the mixed speech scenario, the present application provides a training method for a speech recognition model. In the training data used for training, the text corresponding to the mixed speech sample is only the text corresponding to the real speech sample in the mixed speech sample, such as the recognition text of the real speech, so that the trained speech recognition model has the ability to accurately recognize the text corresponding to the real speech from the mixed speech, thereby improving the accuracy of real speech recognition in the mixed speech scenario.
[0044] Figure 2 A flowchart of a method for training a speech recognition model provided in an embodiment of the present application is provided. The method can be executed by a device with corresponding data computing capabilities, and the device can be in the form of a terminal, computer, server, etc.
[0045] like Figure 2 As shown, the training method of the speech recognition model includes the following steps:
[0046] Step S201, obtaining model training data, the model training data including: speech-text pairs of real speech and speech-text pairs of mixed speech, the text corresponding to the speech sample in the speech-text pair of mixed speech is the text corresponding to the real speech in the speech sample.
[0047] A speech-to-text pair includes a speech sample and its corresponding text. This speech-to-text pair can be a speech-to-text pair for mixed speech or a speech-to-text pair for real speech. The speech sample in a speech-to-text pair for mixed speech is a mixed speech sample, which is a mixture of real speech and synthesized speech. The text corresponding to the mixed speech sample is the text corresponding to the real speech in the mixed speech sample, such as the text obtained by real speech recognition. The speech sample in a speech-to-text pair for real speech is real speech.
[0048] In some embodiments, the model training data further includes a speech-text pair of synthesized speech. The speech samples in the speech-text pair of synthesized speech are synthesized speech.
[0049] For the convenience of description, this application will refer to the speech-text pairs of mixed speech as mixed speech-text pairs, the speech-text pairs of real speech as real speech-text pairs, and the speech-text pairs of synthesized speech as synthesized speech-text pairs.
[0050] In order to train the speech recognition model, the collected speech samples need to be annotated with text to obtain the text corresponding to the speech samples, thereby obtaining multiple speech-text pairs, which are the model training data.
[0051] The text corresponding to the speech sample can be obtained by manual annotation, or by annotation or recognition based on a speech recognition module, followed by manual adjustment of the recognition results output by the speech recognition module. This application does not impose any restrictions on this. The speech recognition module here refers to a software module implemented using existing technology that has the function of recognizing real speech or synthesized speech, for example.
[0052] For mixed speech samples, manual annotation can be used to obtain the text corresponding to the mixed speech samples, thereby obtaining mixed speech sample pairs. For real speech samples or synthesized speech samples, the corresponding text can be recognized by the aforementioned speech recognition module, thereby improving the efficiency of generating training samples.
[0053] The training samples in the model training data used to train speech recognition models are speech-text pairs. These can be categorized into real speech-text pairs, synthesized speech-text pairs, and mixed speech-text pairs. The speech samples in real speech-text pairs are real speech, also known as real speech samples; the speech samples in synthesized speech-text pairs are synthesized speech, also known as synthesized speech samples; and the speech samples in mixed speech sample pairs are mixed speech, also known as mixed speech samples.
[0054] In some embodiments, the text corresponding to the real speech sample is the recognized text of the real speech sample, and the text corresponding to the synthesized speech sample is the recognized text of the synthesized speech sample.
[0055] In this embodiment, for the mixed speech-text pair, the text therein is the real speech in the mixed speech sample or the text annotated by the real speech sample.
[0056] Taking a mixed speech sample of overlapping real speech "Play the music you often listen to recently" and synthesized speech "Turn right 100 meters ahead" as an example, the text corresponding to the mixed speech sample is "Play the music you often listen to recently".
[0057] For the synthesized speech-text pairs in the model training data, you can first provide the text, and then generate the synthesized speech samples corresponding to the text based on the TTS system to obtain the synthesized speech-text pairs.
[0058] For the real speech-text pairs in the model training data, the speech input by the user during the human-computer dialogue collected at historical time can be recorded as the real speech sample, and the recognition result of the speech or the result after manual fine-tuning of the recognition result can be used as the text corresponding to the real speech sample.
[0059] Synthetic speech-text pairs and real text pairs can also be collected through other means, such as through open source datasets, which is not limited in this embodiment of the present application.
[0060] For mixed speech-text pairs in the model training data, after obtaining multiple real speech-text pairs and synthetic speech-text pairs, a mixed speech sample can be obtained by mixing the synthetic speech sample with the real speech sample, and the text corresponding to the real speech sample is the text corresponding to the mixed speech sample.
[0061] For example, in the model training data, the ratio of mixed speech-text pairs, synthesized speech-text pairs, and real speech-text pairs can be 5:2:3, 1:1:1, or other ratios.
[0062] Step S202: Using the model training data composed of the speech-text pairs as input to a speech recognition model to train the speech recognition model.
[0063] Speech recognition models are used to recognize text from input speech, such as voice samples.
[0064] After the model training data is collected, the speech recognition model is trained based on the training samples in the model training data, namely the speech-text pairs.
[0065] Specifically, multiple speech-text pairs can be divided into multiple batches. For each batch, the speech samples in the speech-text pairs of the batch are input into the speech recognition model in turn to obtain the recognized text of the speech samples recognized by the speech recognition model. The parameters of the speech recognition model are adjusted by the difference between the recognized text and the text (true value) corresponding to the speech sample, that is, the loss value, until the training end conditions are met, such as the training time reaches the upper limit time, the number of training times reaches the upper limit number of times, or the loss value meets certain conditions, such as the loss values corresponding to multiple adjacent batches are all less than the set threshold.
[0066] Each batch can include a speech sample, and the speech samples in three consecutive batches can be real speech samples, synthetic speech samples and mixed speech samples, respectively, so as to realize the alternating training of the speech recognition model using different types of speech samples.
[0067] The speech recognition model may be an end-to-end speech recognition model, such as a seq2seq model with an editor-decoder structure based on a self-attention mechanism, an end-to-end speech recognition model based on a Transformer model, or other end-to-end speech recognition models.
[0068] Exemplarily, the speech recognition model may be an end-to-end speech recognition model obtained by combining a Transformer model and an RNN-T (Recurrent Neural Network Transducer) model.
[0069] The training method of the speech recognition model provided in the present application, in the model training stage, the speech-text pairs in the model training data include speech-text pairs of mixed speech and speech-text pairs of real speech. Since the text corresponding to the mixed speech sample in the speech-text pair of mixed speech is the text corresponding to the real speech in the mixed speech sample, the speech recognition model trained by the above-mentioned model training data can have the ability to recognize the real speech in the mixed speech, so that the recognition result of the mixed speech output by the model is close to the recognition result of the real speech in the mixed speech, reducing the influence of synthetic speech on real speech recognition, and improving the accuracy of real speech recognition in the mixed speech scenario where synthetic speech and real speech are mixed, that is, in the mixed speech scenario.
[0070] Optionally, the model training data includes, in addition to the speech-text pairs of real speech and the speech-text pairs of mixed speech, text pairs of synthetic speech, where the speech sample in the speech-text pairs of synthetic speech is synthetic speech and the corresponding text is empty text.
[0071] The text pair of synthesized speech is referred to as a synthesized speech-text pair, and the speech sample therein is referred to as a synthesized speech sample. In the synthesized speech-text pair, the text corresponding to the synthesized speech sample is empty text. This enables a speech recognition model trained using the synthesized speech-text pair to learn that when the input speech is synthesized speech, the output text of the speech recognition model is empty text.
[0072] For example, Figure 3 For this application Figure 2 The schematic diagram of the speech recognition model provided by the illustrated embodiment, Figure 3 Take the encoder-decoder architecture of the speech recognition model as an example, Figure 3 As shown, the input of the speech recognition model is audio and the output is text. When the input audio X is the real speech X T When the expected value Y of the output of the speech recognition model is T For real voice X T The corresponding text Y, that is, Y T =[Y]; When the input audio X is mixed speech X M When the expected value Y of the output of the speech recognition model is M For mixed voice X M Real Voice X T The corresponding text Y, that is, Y M =[Y]; when the input audio X is synthesized speech X R When the expected value Y of the output of the speech recognition model is R If it is empty text, it is Y R =[].
[0073] In some embodiments, the text corresponding to the synthesized speech sample may not be empty text, but the text obtained by recognizing the synthesized speech sample, that is, the recognized text. In this case, in order to enable the speech recognition model to distinguish the type of speech samples in the input speech sample pair, a prompt label (Prompt Label) can be added to the text of the speech-text pair used to train the speech recognition model to indicate the type of the input speech, such as the speech sample, through the prompt label.
[0074] Optionally, the model training data includes, in addition to speech-text pairs of real speech and speech-text pairs of mixed speech, speech-text pairs of synthetic speech, in which the speech sample in the speech-text pair of synthetic speech is synthetic speech and the corresponding text is speech recognition text of the synthetic speech; the text in the speech-text pair includes a prompt label, and the prompt label is used to indicate the type of speech sample in the speech-text pair.
[0075] The prompt label indicates the type of speech sample input to the speech recognition model. Speech sample types can be divided into two or three categories. If there are two categories, the classification can be based on whether the speech sample contains real speech. According to this classification, real speech samples and mixed speech samples are both speech samples containing real speech, and their prompt label can be S. Synthetic speech samples are speech samples that do not contain real speech, and their prompt label can be T. If the speech samples are divided into three categories, namely real speech samples, synthetic speech samples, and synthetic speech samples, the corresponding prompt labels can be R, T, and M, respectively.
[0076] The prompt label added to the text corresponding to the voice sample can be added by manual annotation or automatic annotation, and can be added at a specified position in the text.
[0077] By adding prompt labels to the text of the model training data, that is, the speech-text pair, on the one hand, the speech recognition model can learn the classification of speech samples during training. On the other hand, when using the speech recognition model, for the input speech, the model can add corresponding prompt labels in the recognition text output of the speech, so that the downstream links (such as semantic understanding, etc.) know the type of the current speech, thereby assisting the downstream links to accurately understand the user's intentions, such as assisting the voice assistant to understand the user's questions and accurately respond to the user's questions, ensuring the accuracy of the response content.
[0078] The prompt labels can be manually annotated. Optionally, in order to improve the efficiency of prompt label annotating, automatic annotating can also be used. If automatic annotating is used, the training method of the speech recognition model provided in this application also includes:
[0079] Based on a speech classifier, the type of the speech sample is identified to obtain a prompt label of the speech sample; the prompt label of the speech sample is added to the text corresponding to the speech sample, and a speech-text pair including the speech sample and its corresponding text is generated.
[0080] The speech classifier may be any classifier used to distinguish the type of input audio data or speech, such as the type of speech sample.
[0081] Exemplarily, the speech classifier may be a GMM-HMM (Gaussian Mixture Model-Hidden Markov Model), a speech classification model based on a neural network, such as an LSTM (Long Short Term Memory), a Transformer, or other models.
[0082] The types of speech samples recognized by the speech classifier can be two or three categories. If it is two categories, it can be divided based on whether the speech sample contains real speech; if it is three categories, it can be divided into three categories: real speech samples, synthetic speech samples and mixed speech samples.
[0083] Different types of speech samples correspond to different prompt labels, such as the two prompt labels T and S mentioned above.
[0084] Automatic labeling of prompt labels is achieved based on a speech classifier outside the model, which improves the efficiency of prompt labeling, reduces the time spent on collecting model training data during model training, and improves the efficiency of model training.
[0085] Optionally, the designated position for adding the prompt tag in the text corresponding to the voice sample may be: before the start position or after the end position of the text content of the text record in the voice-text pair.
[0086] Taking the prompt tag as an example, where the prompt tag is located before the text content, that is, added before the beginning of the text content of the text record, the recognition text output by the speech recognition model can be expressed as [ Y], represents the prompt label, and Y is the recognition result of the input speech, such as the text obtained by real speech or synthetic speech recognition, or the text obtained by real speech recognition in mixed speech.
[0087] If the prompt tag is located after the text content, that is, added after the end of the text content of the text record, the recognition text output by the speech recognition model can be expressed as [Y ].
[0088] In some embodiments, the prompt tag can be placed before the text content in a speech-to-text pair, that is, before the start position. Experiments on open-source datasets such as Aishell-1 have shown that placing the prompt tag before the text content improves speech recognition model performance and recognition accuracy for various speech types compared to placing it after the text content.
[0089] The performance of speech recognition models under different designs can be evaluated through parameters such as character error rate (CER) and sentence accuracy (S.Corr).
[0090] Figure 4 A flow chart of another method for training a speech recognition model provided in an embodiment of the present application. Figure 2 Step S202 is further refined based on the illustrated embodiment. In this embodiment, the speech recognition module is also used to identify the type of speech, and the speech sample pairs in the model training data include three types: speech-text pairs of real speech, speech-text pairs of synthesized speech, and speech-text pairs of mixed speech.
[0091] like Figure 4 As shown, the training method of the speech recognition model provided in this embodiment may specifically include the following steps:
[0092] Step S401: Acquire model training data, where the model training data includes: speech-text pairs of real speech, speech-text pairs of synthesized speech, and speech-text pairs of mixed speech.
[0093] In some embodiments, the types of speech samples include real types and synthetic types. The real type of speech samples are real speech or mixed speech mixed with real speech. The real type of speech samples include the aforementioned real speech samples and mixed speech samples; the synthetic type of speech samples are synthetic speech, including the aforementioned synthetic speech samples.
[0094] In some other embodiments, the types of speech samples include three categories, namely, real speech samples, synthesized speech samples and mixed speech samples.
[0095] Step S402: For each batch of speech-text pairs in the model training data, the batch of speech-text pairs is input into the speech recognition model to obtain the predicted type and recognized text of each speech sample in the batch of speech-text pairs output by the speech recognition model.
[0096] A batch of speech-text pairs includes multiple speech-text pairs.
[0097] In some embodiments, the types of speech samples in each speech-text pair in a batch of speech-text pairs are the same, and the types of speech samples in adjacent batches of speech-text pairs are different.
[0098] In other embodiments, a plurality of speech-text pairs may be extracted in proportion from real speech-text pairs, synthesized speech-text pairs, and mixed speech-text pairs to form a batch of speech-text pairs.
[0099] For each batch of speech-text pairs, a speech sample in each speech-text pair in the batch of speech-text pairs is input into a speech recognition model to obtain the recognition text of the speech sample and the predicted type of the speech sample output by the speech recognition model.
[0100] The speech recognition model may include a classifier for identifying speech types, and the type of an input speech sample is predicted based on the classifier to obtain a predicted type of the speech sample.
[0101] Step S403 : Calculate the loss value of the batch of speech-text pairs based on the predicted type and recognized text of each speech sample in the batch of speech-text pairs, as well as the type and corresponding text of each speech sample in the batch of speech-text pairs.
[0102] The loss value of a batch of speech-text pairs can be calculated based on a preset loss function. The loss function includes terms related to the type of speech sample and the predicted type of the model output, as well as terms related to the text corresponding to the speech sample and the recognized text output by the model.
[0103] The loss value may include a type loss value and a text loss value. The type loss value is the loss value between the preset type output by the model and the true type value, i.e., the type in the speech-text pair. The text loss value is the loss value between the recognized text output by the model and the true text value, i.e., the text in the speech-text pair. The type loss value can be obtained based on the type of each speech sample in a batch of speech-text pairs and the predicted type of each speech sample in the speech-text pairs output by the speech recognition model; and the text loss value can be obtained based on each text in a batch of speech-text pairs and the text in the speech-text pairs output by the speech recognition model.
[0104] Exemplarily, the loss function may be a cross entropy loss function, an absolute value loss function, a logarithmic loss function, or the like.
[0105] Step S404: Adjust the parameters of the speech recognition model based on the loss value until the training end condition is met.
[0106] The training end condition can be a condition related to the loss value, such as the loss value or multiple consecutive loss values are lower than a set threshold.
[0107] The training end condition may also be a condition regarding the training time, number of iterations, etc., such as the training time is less than or equal to the upper limit time, and the number of training iterations is less than or equal to the upper limit number.
[0108] Based on the calculated loss value, backpropagation is performed to adjust the parameters of the speech recognition model, completing one round of training. Afterward, the next batch of speech-text pairs is traversed and the next round of training is performed until the training termination conditions are met, such as the number of training iterations reaching the upper limit, the training time reaching the upper limit, or the loss value falling below the set threshold.
[0109] After the training end conditions are met, the training of the speech recognition model is completed, and the model verification and testing phase can be entered. The model is verified and tested using the verification set and test set. After both verification and testing are passed, the model deployment phase is entered, and the speech recognition model is deployed to actual applications to perform speech recognition based on the trained speech recognition model.
[0110] The validation set, test set, and training set can be extracted from multiple collected speech-text pairs, i.e., model training data, in a certain proportion.
[0111] For example, the ratio of the training set, the validation set, and the test set (including the multi-speech text pairs used to train the speech recognition model) can be 8:1:1, 7:2:1, or other ratios.
[0112] The speech recognition model provided in this embodiment is capable of identifying speech types. This allows the model to identify whether the input speech contains real speech or a specific speech type. This allows downstream processes to better understand the context of the input speech based on the recognized speech type, thereby better understanding the user's true intent and making responses to the speech more accurate. For example, in a human-computer interaction scenario, the output speech type allows the robot to adopt different response strategies for different speech types, such as not responding to synthesized speech.
[0113] Optional, Figure 5 A schematic diagram of the structure of the speech recognition model provided in the embodiment of the present application is shown in FIG. Figure 5 As shown, the speech recognition model includes an encoder, a decoder, and a classifier. The encoder encodes the input audio speech, such as a speech sample, to generate a speech code M. The classifier derives and outputs a predicted type L of the input audio speech, such as a true type or a synthesized type, based on the speech code M output by the encoder. The decoder derives and outputs the recognized text Text of the input audio speech based on the speech code M output by the encoder. The outputs of the speech recognition model include the recognized text Text output by the decoder and the predicted type L output by the classifier.
[0114] The prediction type L may include a real type and a synthetic type. The real type indicates that the input audio Speech includes real speech, while the synthetic type indicates that the input audio Speech does not include real speech, that is, synthetic speech.
[0115] The speech code M is the feature representation of the input audio Speech, such as the above-mentioned speech sample. The encoder and classifier take the speech code M as input, and through analysis and processing of the speech code M, output the recognition text and prediction type of the input audio Speech respectively.
[0116] Through the structure of the above-mentioned speech recognition model, the classifier and decoder share the same encoder. Compared with the structure of adding a speech classifier outside the speech recognition model, the overall structure is simplified, the number of models is reduced, the computational cost and time of model training are greatly reduced, and the efficiency of speech recognition model training is improved.
[0117] Figure 6 A flow chart of a speech recognition method provided in an embodiment of the present application is shown as follows: Figure 6 As shown, the speech recognition method includes the following steps:
[0118] Step S601: Acquire speech data to be recognized.
[0119] The voice data to be recognized can be any one or more types of collected audio data, such as the sound of the navigation voice, the sound made by the user, the sound of the voice assistant, the sound in the surrounding environment, etc.
[0120] The voice data to be recognized can be collected through the microphone of the user terminal or the voice recognition end.
[0121] In some embodiments, the original voice data collected by the microphone may be pre-processed, such as segmentation and noise reduction, to obtain voice data to be recognized.
[0122] For original speech data with a long duration, it can be divided into multiple speech data with a shorter duration, such as speech data with a duration of 3 seconds, 5 seconds or other durations.
[0123] Step S602: input the speech data to be recognized into a trained speech recognition model to obtain and output the recognition text of the speech data to be recognized.
[0124] The speech recognition model is trained based on the speech recognition model training method provided in any of the aforementioned embodiments.
[0125] Furthermore, the method further includes generating reply content to the voice data to be recognized based on the recognized text of the voice data to be recognized, generating a reply voice of the reply content, and playing the reply voice. The reply voice of the reply content can be generated based on TTS technology to improve the naturalness and fluency of the voice.
[0126] In some embodiments, a prompt label of the speech data to be recognized is added to the recognition text output by the speech recognition model, and the prompt label is used to indicate the type of the speech data to be recognized.
[0127] Furthermore, the method also includes: extracting a prompt tag of the voice data to be recognized from the recognition text, and when the prompt tag indicates that the type of the voice data to be recognized is a synthetic type, adjusting the recognition text to empty text, so that when the voice data to be recognized is synthetic voice, its recognition result is empty text, so that the downstream link of the voice recognition end (a device that deploys a voice recognition model and is used to recognize voice) does not need to respond to the synthetic voice, such as there is no need to generate reply content for the synthetic voice.
[0128] In some embodiments, after the speech data to be recognized is input into a trained speech recognition model, the speech recognition model may further output a predicted type of the speech data to be recognized.
[0129] Optionally, inputting the speech data to be recognized into a trained speech recognition model to obtain and output a recognized text of the speech data to be recognized includes:
[0130] The speech data to be recognized is input into a trained speech recognition model to obtain the type and recognition text of the speech data to be recognized; when the predicted type of the speech data to be recognized is a synthetic type, the recognition text of the speech data to be recognized is adjusted to empty text and output.
[0131] By adjusting the recognition text of the synthesized speech to empty text, the downstream links of the speech recognition end do not need to respond to the synthesized speech, such as not needing to generate reply content for the synthesized speech, which reduces the number of responses of the downstream links and the amount of calculation of the downstream links; at the same time, by deleting the recognition text of the synthesized speech, the downstream links can more accurately understand the context of the real speech, thereby better understanding the user's true intentions and improving the accuracy of the responses of the downstream links.
[0132] In-vehicle software refers to software installed on a vehicle's onboard computer. For example, when using in-vehicle software with map navigation capabilities, the software generates and broadcasts navigation voice (a type of synthesized speech), such as "Please keep going straight." During the navigation voice broadcast, if a user interacts with an in-vehicle voice assistant (which can be integrated into the in-vehicle software with map navigation capabilities or a separate in-vehicle software), such as issuing the command "Adjust navigation destination to address 1," the original speech captured by the vehicle's microphone is a mixed speech, i.e., a mixture of the synthesized speech "Please keep going straight" and the actual speech "Adjust navigation destination to address 1." For example, "Please keep navigation straight and adjust the destination to address 1." Assume that the original speech captured by the microphone, i.e., this mixed speech, is segmented into three speech data sets, namely, speech data 1 through speech data 3. The text corresponding to speech data 1 through speech data 3 is "Please keep going straight," "Adjust navigation destination to address 1," and "Adjust to address 1," respectively.
[0133] If the in-vehicle voice assistant recognizes the aforementioned voice data using a voice recognition model deployed in the vehicle, and this voice recognition model is trained using the training method provided in the aforementioned embodiments of this application, then when recognizing voice data 1, since voice data 1 is synthesized speech (type is synthesis), the recognized text of voice data 1, i.e., recognized text 1, is empty text. Since voice data 2 is mixed speech, the recognized text of voice data 2, i.e., recognized text 2, is the recognized text "navigation destination" of the real speech in voice data 2. The recognized text of voice data 3, i.e., recognized text 3, is "adjust to address 1." By concatenating recognized text 1 and recognized text 3, the complete content of the user's voice command, i.e., "adjust navigation destination to address 1," is obtained. This avoids the influence of the synthesized speech in the mixed speech on the understanding of the real speech in the mixed speech, and enables the in-vehicle voice assistant to correctly understand the user's true intention and make the correct response, i.e., adjust the destination of the in-vehicle navigation software with map navigation function to address 1.
[0134] Figure 7 A structural diagram of a speech recognition model training device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the training device of the speech recognition model includes: a training data acquisition module 710 and a model training module 720.
[0135] The training data acquisition module 710 is used to obtain model training data, which includes speech-text pairs of real speech and speech-text pairs of mixed speech, wherein a speech-text pair includes a speech sample and its corresponding text, the speech sample in the speech-text pair of mixed speech is a mixed speech sample, the mixed speech sample is a speech sample of mixed real speech and synthesized speech, the text corresponding to the mixed speech sample is the text corresponding to the real speech in the mixed speech sample, and the speech sample in the speech sample pair of real speech is real speech; the model training module 720 is used to use the model training data composed of the speech-text pairs as the input of the speech recognition model to train the speech recognition model.
[0136] Optionally, the model training data also includes: a speech-text pair of synthesized speech, wherein the speech sample in the speech-text pair of synthesized speech is synthesized speech and the corresponding text is empty text.
[0137] Optionally, the model training data also includes: a speech-text pair of synthesized speech, wherein the speech sample in the speech-text pair of synthesized speech is synthesized speech, and the corresponding text is the speech recognition text of the synthesized speech.
[0138] Optionally, the text in the speech-text pair includes a prompt tag, and the prompt tag is used to indicate the type of the speech sample in the speech-text pair.
[0139] Optionally, the prompt tag is located before the start position or after the end position of the text content of the text record.
[0140] Optionally, the device further includes a prompt label adding module, which is used to:
[0141] Based on the speech classifier, the type of the speech sample is identified and the prompt label of the speech sample is obtained; a prompt label adding module is used to add the prompt label of the speech sample to the text corresponding to the speech sample; and a speech-text pair including the speech sample and its corresponding text is generated.
[0142] Optionally, the model training module 720 is specifically configured to:
[0143] For each batch of speech-text pairs in the multiple speech-text pairs, the batch of speech-text pairs is input into the speech recognition model to obtain the predicted type and recognized text of each speech sample in the batch of speech-text pairs output by the speech recognition model; based on the predicted type and recognized text of each speech sample in the batch of speech-text pairs, as well as the type and corresponding text of each speech sample in the batch of speech-text pairs, the loss value of the batch of speech-text pairs is calculated; and based on the loss value, the parameters of the speech recognition model are adjusted until the training end condition is met.
[0144] Optionally, the speech recognition model includes an encoder, a decoder and a classifier; the encoder is used to encode the input speech sample to obtain a speech code; the classifier is used to output the predicted type of the speech sample based on the speech code; and the decoder is used to output the recognition text corresponding to the speech sample based on the speech code.
[0145] Optionally, the type of speech sample or the predicted type includes: a real type and a synthetic type. The real type speech sample is real speech or a mixed speech of real speech, and the synthetic type speech sample is a synthetic speech. The real type speech sample includes real speech samples and mixed speech samples, and the synthetic type speech sample includes a synthetic speech sample.
[0146] The training device for the speech recognition model provided in the embodiment of the present application can be used to execute the technical solution of the training method for the speech recognition model provided in any of the above embodiments of the present application. Its implementation principle and technical effects are similar, and will not be repeated here in this embodiment.
[0147] The present invention also provides a speech recognition device, including:
[0148] A data acquisition module is used to acquire speech data to be recognized; a speech recognition module is used to input the speech data to be recognized into a trained speech recognition model to obtain and output the recognition text of the speech data to be recognized; wherein, the speech recognition model is trained based on the training method of the speech recognition model provided in any embodiment of the present application.
[0149] Optional voice recognition module, specifically used for:
[0150] The speech data to be recognized is input into a trained speech recognition model to obtain a predicted type and a recognized text of the speech data to be recognized; when the predicted type of the speech data to be recognized is a synthetic type, the recognized text of the speech data to be recognized is adjusted to an empty text and output.
[0151] The speech recognition device provided in the embodiment of the present application can be used to execute the technical solution of the sound recognition method provided in any of the above embodiments of the present application. Its implementation principle and technical effects are similar, and will not be repeated here in this embodiment.
[0152] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 8 As shown, the electronic device of this embodiment may include: at least one processor 801; and a memory 802 communicatively connected to the at least one processor; wherein the memory 802 stores instructions that can be executed by the at least one processor 801, and the instructions are executed by the at least one processor 801 to enable the electronic device to execute the method described in any of the above embodiments.
[0153] Optionally, the memory 802 may be independent or integrated with the processor 801 .
[0154] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the aforementioned embodiments and will not be described in detail here.
[0155] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in any of the above embodiments is implemented.
[0156] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method described in any of the aforementioned embodiments when executed by a processor.
[0157] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is merely a logical function division. In actual implementation, other division methods may be used. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not implemented.
[0158] The above-mentioned integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application.
[0159] It should be understood that the above-mentioned processor can be a processing unit (Central Processing Unit, CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The memory may include RAM (Random Access Memory), and may also include NVM (Non-Volatile Memory), such as at least one disk storage, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.
[0160] The storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0161] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an application-specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0162] It should be noted that, in this document, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0163] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0164] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0165] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for training a speech recognition model, characterized in that: include: Acquire model training data, the model training data comprising: speech-text pairs of real speech and speech-text pairs of mixed speech, wherein a speech-text pair comprises a speech sample and its corresponding text, the speech sample in the speech-text pair of mixed speech is a mixed speech sample, the mixed speech sample is a speech sample of mixed real speech and synthesized speech, the text corresponding to the mixed speech sample is the text corresponding to the real speech in the mixed speech sample, and the speech sample in the speech sample pair of real speech is real speech; The model training data composed of the speech-text pairs is used as input to the speech recognition model to train the speech recognition model.
2. The method according to claim 1, characterized in that The model training data also includes: a speech-text pair of synthesized speech, wherein the speech sample in the speech-text pair of synthesized speech is synthesized speech and the corresponding text is an empty text.
3. The method according to claim 1, characterized in that The model training data also includes: a speech-text pair of synthesized speech, wherein the speech sample in the speech-text pair of synthesized speech is synthesized speech and the corresponding text is the speech recognition text of the synthesized speech; The text in the speech-text pair includes a prompt tag, and the prompt tag is used to indicate the type of the speech sample in the speech-text pair.
4. The method according to claim 3, characterized in that The prompt tag is located before the start position or after the end position of the text content of the text record.
5. The method according to claim 3, characterized in that The method further comprises: Based on the speech classifier, identifying the type of the speech sample and obtaining a prompt label of the speech sample; Adding the prompt tag of the voice sample to the text corresponding to the voice sample; Generate speech-text pairs including speech samples and their corresponding texts.
6. The method according to claim 1, wherein The model training data also includes: a speech-text pair of synthesized speech, wherein the speech sample in the speech-text pair of synthesized speech is synthesized speech and the corresponding text is the speech recognition text of the synthesized speech; The speech recognition model includes an encoder, a decoder and a classifier; The encoder is used to encode the speech sample in the input speech-text pair to obtain speech code; The classifier is used to output a predicted type of the speech sample based on the speech coding; The decoder is configured to output a recognized text of the speech sample based on the speech encoding.
7. The method according to claim 3 or 6, characterized in that The types or predicted types of the speech samples include: real type and synthetic type. The real type speech sample is real speech or mixed speech of real speech, and the synthetic type speech sample is synthetic speech.
8. A speech recognition method, characterized in that: include: Obtaining voice data to be recognized; Inputting the speech data to be recognized into a trained speech recognition model to obtain and output a recognition text of the speech data to be recognized; Wherein, the speech recognition model is trained based on the method provided in any one of claims 1-7.
9. The method according to claim 8, characterized in that Inputting the speech data to be recognized into a trained speech recognition model, obtaining and outputting the recognition text of the speech data to be recognized, including: Inputting the speech data to be recognized into a trained speech recognition model to obtain a predicted type and a recognized text of the speech data to be recognized; When the predicted type of the speech data to be recognized is a synthesis type, the recognition text of the speech data to be recognized is adjusted to an empty text and output.
10. A training device for a speech recognition model, characterized in that: include: A training data acquisition module is used to acquire model training data, wherein the model training data includes speech-text pairs of real speech and speech-text pairs of mixed speech, wherein a speech-text pair includes a speech sample and its corresponding text, the speech sample in the speech-text pair of mixed speech is a mixed speech sample, the mixed speech sample is a speech sample of mixed real speech and synthesized speech, the text corresponding to the mixed speech sample is the text corresponding to the real speech in the mixed speech sample, and the speech sample in the speech sample pair of real speech is real speech; The model training module is used to use the model training data composed of the speech-text pairs as input to the speech recognition model to train the speech recognition model.