Streaming speech recognition model training method and device, electronic equipment and storage medium
By performing cropping and comparative learning training on the streaming speech recognition model, the tail word deletion and transmission delay problems in streaming speech recognition are solved, and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202510330134.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-29
AI Technical Summary
There are problems in the streaming speech recognition model with low transmission delay and recognition accuracy, especially in short sentences, the tail words cannot be predicted and the future information in the same sentence cannot be used.
By taking the first-time sample speech sequence from the initial sample speech sequence into the initial model for iterative training, the tail of the sample speech sequence is cut, and the non-stream speech recognition model is combined as the teacher model for comparison learning training, and the stream speech recognition model is obtained.
This significantly reduces the tail word deletion problem in streaming speech recognition and improves recognition accuracy while reducing transmission delay.
Smart Images

Figure CN120388560A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet technologies, and in particular, to a method, an apparatus, an electronic device, and a storage medium for training a streaming speech recognition model. Background Art
[0002] Streaming speech recognition technology aims to return the recognition result of speech content in real time during the process of processing an audio stream. Compared with traditional non-streaming speech recognition technology, streaming speech recognition can perform speech processing and text conversion in real time while the user is speaking, greatly improving the immediacy and efficiency of interaction. Streaming speech recognition technology has been widely used in many scenarios that require real-time response, including but not limited to: e-commerce live streaming, meeting recording.
[0003] An important problem of current streaming speech recognition models is emission delay. That is, the time difference between when the user speaks and when the model makes a prediction. Due to the emission delay, the last word of a short sentence often cannot be predicted due to the lack of a corresponding input block, resulting in the problem of last-word deletion.
[0004] Secondly, due to its causal constraint, streaming speech recognition cannot utilize future information in the same sentence, resulting in its recognition accuracy being often significantly lower than that of non-streaming speech recognition under the same model configuration. Summary of the Invention
[0005] The present disclosure provides a method, an apparatus, an electronic device, and a storage medium for training a streaming speech recognition model. The technical solution of the present disclosure is as follows:
[0006] According to a first aspect of an embodiment of the present disclosure, there is provided a method for training a streaming speech recognition model, including:
[0007] Taking a first sample speech sequence with a first duration from an initial sample speech sequence and inputting it into an initial model, performing corresponding speech recognition processing after each input of the first sample speech sequence is completed, and iteratively training the initial model according to the output speech recognition result to obtain a non-streaming speech recognition model; the first duration is a preset standard duration of a truncated speech block for training the non-streaming speech recognition model;
[0008] According to the duration of the initial sample speech sequence, performing a cropping process on the tail of the initial sample speech sequence to obtain a cropped sample speech sequence; combining the cropped sample speech sequence and the uncropped sample speech sequence in the initial sample speech sequence to obtain a second sample speech sequence;
[0009] Take a third sample speech sequence of a second duration from the second sample speech sequence, input the third sample speech sequence into the non-streaming speech recognition model, perform real-time speech recognition processing during the input process of the third sample speech sequence, and iteratively train the non-streaming speech recognition model according to the output speech recognition result to obtain an initial streaming speech recognition model; the second duration is a dynamically changing duration, the second duration is greater than or equal to a preset third duration and less than or equal to the first duration, and the third duration is the preset standard duration of the truncated speech block for training the streaming speech recognition model;
[0010] Take the non-streaming speech recognition model as the teacher model and the initial streaming speech recognition model as the student model, and perform contrastive learning training on the initial streaming speech recognition model according to the second sample speech sequence to obtain a first streaming speech recognition model.
[0011] Optionally, the method for cropping the tail of the initial sample speech sequence according to the duration of the initial sample speech sequence to obtain a second sample speech sequence includes:
[0012] Determine the size relationship between the duration of the initial sample speech sequence and a preset fourth duration, take the initial sample speech sequence with a duration less than the fourth duration as a short sample speech sequence, and take the initial sample speech sequence with a duration greater than or equal to the fourth duration as a long sample speech sequence; the fourth duration is greater than the first duration;
[0013] Crop the speech segment of a preset fifth duration at the tail of a preset first proportion of the short sample speech sequences to obtain the cropped short sample speech sequences; the fifth duration is a random value within a preset duration range;
[0014] Crop the speech segment of the fourth duration at the tail of a preset second proportion of the long sample speech sequences to obtain the cropped long sample speech sequences; the second proportion is less than the first proportion;
[0015] Combine the uncropped short sample speech sequences in the short sample speech sequences, the cropped short sample speech sequences, the uncropped long sample speech sequences in the long sample speech sequences, and the cropped long sample speech sequences to obtain a second sample speech sequence.
[0016] Optionally, before taking a third sample speech sequence of a second duration from the second sample speech sequence each time, it further includes:
[0017] Set a target array, where the elements in the target array are N, and the values of the elements are evenly distributed between a first value and a second value. The first value is the value of the third duration, and the second value is the value of the first duration; N is a natural number;
[0018] The number of training times for the iterative training is M times. Each time, a third sample speech sequence with a second duration is taken from the second sample speech sequence, including:
[0019] During the first-stage training, each time a first element is randomly taken from the target array, and a third sample speech sequence with a duration corresponding to the first element is taken from the second sample speech sequence;
[0020] During the second-stage training, each time a second element is randomly taken between the first element and the N / 2-th element of the target array, and a third sample speech sequence with a duration corresponding to the second element is taken from the second sample speech sequence;
[0021] During the third-stage training, each time a third sample speech sequence with a duration corresponding to the first value is taken.
[0022] Optionally, using the non-streaming speech recognition model as the teacher model and the initial streaming speech recognition model as the student model, and performing contrastive learning training on the initial streaming speech recognition model according to the second sample speech sequence to obtain a first streaming speech recognition model, including:
[0023] Take a fourth sample speech sequence with a duration corresponding to the first value from the second sample speech sequence, input the fourth sample speech sequence into the initial streaming speech recognition model, and control the initial streaming speech recognition model to perform real-time speech recognition processing during the input process of the fourth sample speech sequence; the first value is the preset standard duration of the truncated speech block for training the non-streaming speech recognition model;
[0024] Obtain the features generated by the initial streaming speech recognition model in the last neural network layer to obtain a first deep feature;
[0025] Take a fourth sample speech sequence with a duration corresponding to the first value from the second sample speech sequence, input the fourth sample speech sequence into the non-streaming speech recognition model, and control the non-streaming speech recognition model to perform corresponding speech recognition processing after each input of the fourth sample speech sequence;
[0026] Obtain the features generated by the non-streaming speech recognition model in the last neural network layer to obtain a second deep feature;
[0027] Calculate a first loss value between the first deep feature and the second deep feature according to the target loss function;
[0028] The model parameters of the initial streaming speech recognition model are updated according to the first loss value to obtain a first streaming speech recognition model.
[0029] Optionally, calculating a first loss value between the first depth feature and the second depth feature according to a target loss function includes:
[0030] Obtaining a value of the second depth feature after applying a stop gradient operation to the non-streaming speech recognition model to obtain a third depth feature;
[0031] determining a first similarity between the first depth feature and the third depth feature;
[0032] Decomposing the first depth feature from the time dimension to obtain multiple one-dimensional frequency dimension features with different time dimensions, randomly selecting a time dimension from the multiple time dimensions and combining it with the frequency dimension feature to obtain a fourth depth feature corresponding to the time dimension;
[0033] determining a second similarity between the first depth feature and the fourth depth feature;
[0034] The value of the target loss function is determined based on the first similarity, the second similarity and a preset temperature coefficient to obtain a first loss value; the temperature coefficient is used to adjust the result of the similarity calculation.
[0035] Optionally, after obtaining the first streaming speech recognition model, the method further includes:
[0036] Filling the tail of each initial test voice sequence used for performing the test with a silence segment of a preset fifth duration to obtain a plurality of first test voice sequences;
[0037] Taking a second test speech sequence of the fifth duration from the first test speech sequence, inputting the second test speech sequence into the first streaming speech recognition model for speech recognition processing, and obtaining a first text sequence;
[0038] Calculate a second loss value between the first text sequence and the real text sequence corresponding to the initial test speech sequence, and update the model parameters of the first streaming speech recognition model according to the second loss value to obtain a second streaming speech recognition model.
[0039] According to a second aspect of an embodiment of the present disclosure, a training device for a streaming speech recognition model is provided, comprising:
[0040] A non-streaming training module, configured to take a first-sample speech sequence of a first duration from an initial sample speech sequence and input it into an initial model, perform corresponding speech recognition processing after each input of the first-sample speech sequence is completed, and iteratively train the initial model according to the output speech recognition results to obtain a non-streaming speech recognition model; the first duration is a preset standard duration of a truncated speech block for training the non-streaming speech recognition model;
[0041] A cropping module, configured to perform cropping processing on the tail of the initial sample speech sequence according to the duration of the initial sample speech sequence to obtain a cropped sample speech sequence; combine the cropped sample speech sequence and the uncropped sample speech sequence in the initial sample speech sequence to obtain a second sample speech sequence;
[0042] A streaming training module, configured to take a third-sample speech sequence of a second duration from the second sample speech sequence, input the third-sample speech sequence into the non-streaming speech recognition model, perform real-time speech recognition processing during the input of the third-sample speech sequence, and iteratively train the non-streaming speech recognition model according to the output speech recognition results to obtain an initial streaming speech recognition model; the second duration is a dynamically changing duration, the second duration is greater than or equal to a preset third duration and less than or equal to the first duration, and the third duration is a preset standard duration of a truncated speech block for training the streaming speech recognition model;
[0043] A contrast training module, configured to use the non-streaming speech recognition model as a teacher model and the initial streaming speech recognition model as a student model, and perform contrast learning training on the initial streaming speech recognition model according to the second sample speech sequence to obtain a first streaming speech recognition model.
[0044] Optionally, the cropping module is specifically configured to perform:
[0045] Determine the size relationship between the duration of the initial sample speech sequence and a preset fourth duration, use the initial sample speech sequence with a duration less than the fourth duration as a short sample speech sequence, and use the initial sample speech sequence with a duration greater than or equal to the fourth duration as a long sample speech sequence; the fourth duration is greater than the first duration;
[0046] Crop a speech segment of a preset fifth duration at the tail of a preset first ratio of the short sample speech sequences to obtain a cropped short sample speech sequence; the fifth duration is a random value within a preset duration range;
[0047] Crop the speech segment of the fourth duration at the tail of the long sample speech sequence of the preset second ratio to obtain the cropped long sample speech sequence; the second ratio is smaller than the first ratio;
[0048] Combine the uncropped short sample speech sequence in the short sample speech sequence, the cropped short sample speech sequence, the uncropped long sample speech sequence in the long sample speech sequence, and the cropped long sample speech sequence to obtain the second sample speech sequence.
[0049] Optionally, the device further includes:
[0050] A setting module, configured to execute setting a target array, where the number of elements in the target array is N, the values of the elements are evenly distributed between a first value and a second value, the first value is the value of the third duration, and the second value is the value of the first duration; N is a natural number;
[0051] The streaming training module is specifically configured to execute:
[0052] During the first-stage training, randomly take a first element from the target array each time, and take a third sample speech sequence of the duration corresponding to the first element from the second sample speech sequence;
[0053] During the second-stage training, randomly take a second element between the first element and the N / 2-th element of the target array each time, and take a third sample speech sequence of the duration corresponding to the second element from the second sample speech sequence;
[0054] During the third-stage training, take a third sample speech sequence of the duration corresponding to the first value each time.
[0055] Optionally, the contrast training module is specifically configured to execute:
[0056] Take a fourth sample speech sequence of the first value duration from the second sample speech sequence, input the fourth sample speech sequence into the initial streaming speech recognition model, and control the initial streaming speech recognition model to perform real-time speech recognition processing during the input process of the fourth sample speech sequence; the first value is the preset standard duration of the truncated speech block for non-streaming speech recognition model training;
[0057] Obtain the features generated by the initial streaming speech recognition model in the last neural network layer to obtain the first depth feature;
[0058] Taking a fourth sample speech sequence having a duration corresponding to the first value from the second sample speech sequence, inputting the fourth sample speech sequence into the non-streaming speech recognition model, and controlling the non-streaming speech recognition model to perform corresponding speech recognition processing after each fourth sample speech sequence is input;
[0059] Obtaining features generated by the non-streaming speech recognition model at the last neural network layer to obtain a second deep feature;
[0060] Calculating a first loss value between the first depth feature and the second depth feature according to a target loss function;
[0061] The model parameters of the initial streaming speech recognition model are updated according to the first loss value to obtain a first streaming speech recognition model.
[0062] Optionally, the comparative training module is specifically configured to execute:
[0063] Obtaining a value of the second depth feature after applying a stop gradient operation to the non-streaming speech recognition model to obtain a third depth feature;
[0064] determining a first similarity between the first depth feature and the third depth feature;
[0065] Decomposing the first depth feature from the time dimension to obtain multiple one-dimensional frequency dimension features with different time dimensions, randomly selecting a time dimension from the multiple time dimensions and combining it with the frequency dimension feature to obtain a fourth depth feature corresponding to the time dimension;
[0066] determining a second similarity between the first depth feature and the fourth depth feature;
[0067] The value of the target loss function is determined based on the first similarity, the second similarity and a preset temperature coefficient to obtain a first loss value; the temperature coefficient is used to adjust the result of the similarity calculation.
[0068] Optionally, the device further comprises:
[0069] a filling module configured to fill the tail of each initial test voice sequence used for performing the test with a silence segment of a preset fifth duration, thereby obtaining a plurality of first test voice sequences;
[0070] an input module configured to extract a second test speech sequence of the fifth duration from the first test speech sequence, input the second test speech sequence into the first streaming speech recognition model for speech recognition processing, and obtain a first text sequence;
[0071] A calculation module, configured to execute a calculation of a second loss value between the first text sequence and the true text sequence corresponding to the initial test voice sequence, and update model parameters of the first streaming speech recognition model according to the second loss value to obtain a second streaming speech recognition model.
[0072] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0073] A processor;
[0074] A memory for storing instructions executable by the processor;
[0075] Wherein, the processor is configured to execute the instructions to implement the training method of the streaming speech recognition model as described in the first aspect, and / or the recommendation method as described in the second aspect.
[0076] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the training method of the streaming speech recognition model as described in the first aspect.
[0077] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, the computer program product includes a computer program, the computer program is stored in a readable storage medium, and at least one processor of a computer device reads and executes the computer program, enabling the computer device to execute the training method of the streaming speech recognition model as described in the first aspect.
[0078] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0079] In an embodiment of the present disclosure, a first sample speech sequence with a first duration is taken from an initial sample speech sequence and input into an initial model. After each first sample speech sequence is input, corresponding speech recognition processing is performed, and the initial model is iteratively trained according to the output speech recognition result to obtain a non-streaming speech recognition model; the first duration is a preset standard duration of a truncated speech block for training the non-streaming speech recognition model; according to the duration of the initial sample speech sequence, the tail of the initial sample speech sequence is trimmed to obtain a second sample speech sequence; a third sample speech sequence with a second duration is taken from the second sample speech sequence, the third sample speech sequence is input into the non-streaming speech recognition model, and speech recognition processing is performed in real time during the input process of the third sample speech sequence, and the non-streaming speech recognition model is iteratively trained according to the output speech recognition result to obtain an initial streaming speech recognition model; the second duration is a dynamically changing duration, the second duration is greater than or equal to a preset third duration and less than or equal to the first duration, and the third duration is a preset standard duration of a truncated speech block for training the streaming speech recognition model; the non-streaming speech recognition model is used as a teacher model, and the initial streaming speech recognition model is used as a student model, and the initial streaming speech recognition model is trained by contrastive learning according to the second sample speech sequence to obtain a first streaming speech recognition model. This solution significantly reduces the problem of tail word deletion in streaming speech recognition without increasing model parameters and reducing prediction accuracy by trimming the tail of the initial sample speech sequence. Moreover, based on the knowledge transfer method, the streaming model is trained with the non-streaming model as the teacher model and the streaming model as the student model. The obtained streaming speech recognition model not only has the recognition accuracy of the non-streaming speech recognition model but also can reduce the transmission delay.
[0080] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.
[0082] Figure 1 is a flowchart of steps of a method for training a streaming speech recognition model according to an exemplary embodiment;
[0083] Figure 2 is a block diagram of the structure of a device for training a streaming speech recognition model according to an exemplary embodiment.
[0084] Figure 3It is a block diagram of an electronic device for a training method of a streaming speech recognition model shown according to an exemplary embodiment. Detailed implementation manners
[0085] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0086] It should be noted that the terms "first", "second", etc. in the description and claims of the present disclosure and the above drawings are used to distinguish similar first objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0087] The inventors found during the research on related technologies that, compared with non-streaming speech recognition, the key to streaming speech recognition is that it can only use the cached historical blocks as historical dependencies, and the current prediction can only rely on the information before the current moment in the time series.
[0088] Currently, traditional streaming speech recognition systems generally include voice activity detection at the front end and a streaming speech recognition model at the back end. The voice activity detection model at the front end breaks sentences by detecting silent segments and clears the cached historical blocks in a timely manner. The streaming speech recognition model at the back end predicts the text corresponding to the current block based on the cached historical blocks and the current block. If the current sentence has not ended, the historical blocks and the current block are cached together as the historical blocks for the next speech block.
[0089] Therefore, traditional streaming speech recognition faces the following two problems:
[0090] First, an important problem of the current streaming speech recognition model is emission delay. In the case of few historical information in short sentences, the model has insufficient confidence in prediction, and conventional streaming speech recognition models often use more (future) context information to produce better prediction results, which leads to significant emission delay, that is, the time difference between the user's speech and the model's prediction. Due to the emission delay, the last word of a short sentence often cannot be predicted due to the lack of corresponding input blocks.
[0091] Second, due to its causal limitation, streaming speech recognition cannot utilize future information in the same sentence, resulting in its recognition accuracy being significantly lower than that of non-streaming speech recognition under the same model configuration.
[0092] Exemplarily, when using a conventional streaming speech recognition model to output the recognition result of a speech sequence, for the first few seconds when the speech starts to be output, the streaming speech recognition model does not output the corresponding recognition result in time. Moreover, when the speech has ended, the streaming speech recognition model has not output the recognition result corresponding to the last few seconds of the speech.
[0093] Therefore, there are problems of first-word delay and last-word delay in the recognition results of traditional streaming speech recognition models.
[0094] The inventors also found that in order to solve the problems faced by traditional streaming speech recognition in the prior art, improvements were made in the training loss function, that is, a regularized loss function was used to make the inference of the streaming speech recognition model tend to emit prediction results earlier.
[0095] However, the purpose of the regularized loss function is to find a balance between emission delay and prediction accuracy. Reducing the emission delay at the cost of prediction accuracy, although the situation of last-word deletion is improved, the overall accuracy will decrease.
[0096] Therefore, the embodiments of the present disclosure focus on solving the problems of low emission delay and low accuracy of traditional streaming speech recognition models. The method proposed by the present disclosure includes two aspects. From the perspective of data, without increasing model parameters and reducing prediction accuracy, the problem of last-word deletion in streaming speech recognition is significantly reduced. From the perspective of the model, a streaming model is trained based on the method of knowledge transfer. The non-streaming model is used as the target and the teacher, and the non-streaming model is used to improve the performance of the streaming model, ensuring the upper limit of the model's modeling of the data distribution.
[0097] Figure 1 It is a flowchart of the steps of a method for training a streaming speech recognition model shown according to an exemplary embodiment. As Figure 1 shown, it includes the following steps:
[0098] In step S11, a first sample speech sequence with a first duration is taken from the initial sample speech sequence and input into the initial model. After each first sample speech sequence is input, corresponding speech recognition processing is performed, and the initial model is iteratively trained according to the output speech recognition result to obtain a non-streaming speech recognition model; the first duration is the preset standard duration of the truncated speech block for training the non-streaming speech recognition model.
[0099] In the process of training a streaming speech recognition model, the training of a non-streaming speech model is used as pre-training, and then the trained non-streaming model is migrated to the streaming model. A streaming speech recognition model (Streaming ASR Model) refers to a type of model that can support real-time return of recognition results during the process of processing an audio stream. In contrast, a non-streaming model must process the entire sentence audio before returning the results. Step S11 is the training process of the non-streaming speech recognition model.
[0100] Specifically, first, an initial sample speech sequence is obtained. This sample can be a segmented speech sequence with a fixed duration. For example, one sample is 10 s (seconds). When inputting training data into the initial model, each time a chunk of the speech sequence is taken from the initial sample speech sequence and input into the initial model. Among them, one chunk corresponds to a first sample speech sequence, and the size of one chunk is the first duration. The first duration is the standard duration of the training data of the non-streaming speech recognition model. The standard duration can be set according to requirements.
[0101] At the same time, control the initial model to perform corresponding speech recognition processing after each first sample speech sequence is input. In this way, the model can use the context information of the entire first sample speech sequence for learning, so as to ensure that the model reaches the optimal effect in modeling data distribution and learning speech features.
[0102] Iteratively train the initial model according to the above method to obtain a non-streaming speech recognition model.
[0103] In step S12, according to the duration of the initial sample speech sequence, the tail of the initial sample speech sequence is cropped to obtain a cropped sample speech sequence; the cropped sample speech sequence and the uncropped sample speech sequence in the initial sample speech sequence are combined to obtain a second sample speech sequence.
[0104] In the training of the streaming speech recognition model, in order to solve the problem that the tail of the sentence cannot be predicted due to emission delay, the present disclosure reduces the emission delay of the model by cropping the tail of the sample speech sequence with a certain probability, and advances the prediction result as a whole in time.
[0105] Specifically, according to the duration of the initial sample speech sequence, different cropping probabilities are set, and the tail of the initial sample speech sequence is cropped according to the above cropping probability. After cropping, a second sample speech sequence is obtained.
[0106] The second sample speech sequence includes the uncropped initial sample speech sequence and the cropped initial sample speech sequence. The second sample speech sequence is used as the input data during the training of the streaming speech model.
[0107] Crop the sample speech sequence and retain the remaining part for training, with the aim of compressing the possible time of the model prediction delay.
[0108] From the perspective of the optimization objective of the loss function of the streaming speech recognition model, it is considered reasonable to predict characters in the last few frames of the speech. Cropping the tail is equivalent to the model being unable to predict characters in the last few frames, so the overall prediction is advanced.
[0109] In step S13, a third sample speech sequence with a second duration is taken from the second sample speech sequence, and the third sample speech sequence is input into the non-streaming speech recognition model. During the input process of the third sample speech sequence, real-time speech recognition processing is performed, and the non-streaming speech recognition model is iteratively trained according to the output speech recognition result to obtain an initial streaming speech recognition model; the second duration is a dynamically changing duration, the second duration is greater than or equal to a preset third duration and less than or equal to the first duration, and the third duration is the preset standard duration of the truncated speech block for training the streaming speech recognition model.
[0110] The present disclosure uses the non-streaming speech recognition model trained in step S11 as the initial model for streaming model training, and adopts the method of dynamic chunks to gradually migrate the non-streaming speech recognition model to the streaming speech recognition model. This method helps the model find a balance between non-streaming and streaming, enabling the non-streaming model to retain more non-streaming knowledge during the migration to the streaming model.
[0111] Specifically, each time a speech sequence of one chunk is taken from the second sample speech sequence and input into the initial model. Among them, one chunk corresponds to a third sample speech sequence, the size of one chunk is the second duration, the second duration is less than or equal to the first duration, and it is a dynamically changing duration.
[0112] Chunks of different lengths mean different lengths of delay. The larger the chunk, the richer the context information the model sees, and the better the prediction effect. The smaller the chunk, the less context but the faster the response.
[0113] Set the value range of the second duration. Exemplarily, the second duration corresponds to N values. During the training process, each time a value is randomly selected from the N values as the size of the chunk input this time. The larger the value of the chunk, the more the model training approaches non-streaming, and the smaller the value of the chunk, the more the model training approaches streaming.
[0114] Exemplarily, in the initial training, a value can be randomly selected as the size of the chunk. As the number of training iterations progresses, large chunks are gradually removed from the array until, in the final iteration, the value of the chunk is only the standard duration during the training of the streaming speech recognition model.
[0115] It can be understood that the larger the chunk, the more historical features the streaming model models for causal dependencies, and the better the model performance. The smaller the value of the chunk, the closer it is to the training of the streaming model, and the larger the value, the closer it is to the duration of the entire sentence, so the closer it is to the training of the non-streaming model.
[0116] At the same time, control the non-streaming speech recognition model to perform real-time speech recognition processing during the input of each third sample speech sequence, so that the model gradually adapts to the limitations of the streaming features.
[0117] Iteratively train the non-streaming speech recognition model according to the above method to obtain an initial streaming speech recognition model.
[0118] In step S14, use the non-streaming speech recognition model as the teacher model and the initial streaming speech recognition model as the student model, and perform contrastive learning training on the initial streaming speech recognition model according to the second sample speech sequence to obtain a first streaming speech recognition model.
[0119] Use the non-streaming speech recognition model as the teacher model and the initial streaming speech recognition model as the student model, and further improve the upper limit of the initial streaming speech recognition model by using the method of contrastive learning distillation.
[0120] Through contrastive learning distillation, use the strong model (non-streaming speech recognition model) to guide the streaming model and further improve the performance of the streaming model. During the training process, use the output of the non-streaming speech recognition model as a reference, and through contrastive learning, enable the initial streaming speech recognition model to better learn the feature representation and modeling ability of the non-streaming model, and finally train to obtain a first streaming speech recognition model.
[0121] Since the non-streaming speech recognition model does not have the problem of emission delay, the streaming speech recognition model obtained by using the method of contrastive learning training not only has the recognition accuracy of the non-streaming speech recognition model but also can reduce the emission delay.
[0122] In summary, in the embodiments of the present disclosure, a first sample voice sequence with a first duration is taken from the initial sample voice sequence and input into the initial model. After each input of the first sample voice sequence is completed, corresponding speech recognition processing is performed, and the initial model is iteratively trained according to the output speech recognition results to obtain a non-streaming speech recognition model; the first duration is the preset standard duration of the truncated speech block for training the non-streaming speech recognition model; according to the duration of the initial sample voice sequence, the tail of the initial sample voice sequence is trimmed to obtain a trimmed sample voice sequence; the trimmed sample voice sequence and the untrimmed sample voice sequence in the initial sample voice sequence are combined to obtain a second sample voice sequence; a third sample voice sequence with a second duration is taken from the second sample voice sequence, the third sample voice sequence is input into the non-streaming speech recognition model, and speech recognition processing is performed in real time during the input process of the third sample voice sequence, and the non-streaming speech recognition model is iteratively trained according to the output speech recognition results to obtain an initial streaming speech recognition model; the second duration is a dynamically changing duration, the second duration is greater than or equal to a preset third duration and less than or equal to the first duration, and the third duration is the preset standard duration of the truncated speech block for training the streaming speech recognition model; the non-streaming speech recognition model is used as the teacher model, and the initial streaming speech recognition model is used as the student model, and the initial streaming speech recognition model is trained by contrastive learning according to the second sample voice sequence to obtain a first streaming speech recognition model. In this solution, by trimming the tail of the initial sample voice sequence, the problem of tail word deletion in streaming speech recognition is significantly reduced without increasing the model parameters and reducing the prediction accuracy. Moreover, the streaming model is trained based on the knowledge transfer method, with the non-streaming model as the teacher model and the streaming model as the student model. The obtained streaming speech recognition model by using the contrastive learning training method not only has the recognition accuracy of the non-streaming speech recognition model but also can reduce the transmission delay.
[0123] In a possible implementation manner, step S12 includes:
[0124] In step S121, determine the size relationship between the duration of the initial sample voice sequence and a preset fourth duration, and use the initial sample voice sequence with a duration less than the fourth duration as a short sample voice sequence, and use the initial sample voice sequence with a duration greater than or equal to the fourth duration as a long sample voice sequence; the fourth duration is greater than the first duration.
[0125] The present disclosure processes the training data by trimming. Specifically, according to the size relationship between the duration of the initial sample voice sequence and the fourth duration, the initial sample voice sequence is divided into a short sample voice sequence and a long sample voice sequence. The fourth duration is greater than the first duration.
[0126] Exemplarily, the fourth duration is set to 3.5 seconds. Sequences shorter than 3.5 seconds are regarded as short sample speech sequences, and sequences greater than or equal to 3.5 seconds are regarded as long sample speech sequences.
[0127] In step S122, a speech segment of a preset fifth duration at the tail of a preset first proportion of the short sample speech sequences is trimmed to obtain a trimmed short sample speech sequence; the fifth duration is a random value within a preset duration range.
[0128] When trimming the tail, not all tails of the speech sequences will be trimmed, but trimming is performed according to a trimming probability of the first proportion, which can ensure the integrity of some speech sequences and enable the model to learn diverse semantic expressions as much as possible.
[0129] In addition, the trimming length is the fifth duration, and the fifth duration is a random value within a preset duration range. Exemplarily, the fifth duration can be a random value between 0.15 seconds and 0.4 seconds, such as 0.2 seconds.
[0130] Exemplarily, the first proportion is set to 70%, and the fifth duration is 0.2 seconds. In this way, the tails of 70% of the short sample speech sequences are trimmed, and the trimming length is 0.2 seconds.
[0131] In a possible implementation manner, the fifth duration is less than or equal to 75% of the duration of the initial sample speech sequence.
[0132] To ensure that the integrity of the sentence is not overly damaged, the fifth duration is randomly selected according to the duration of each initial sample speech sequence, and the fifth duration should be less than or equal to 75% of the duration of the initial sample speech sequence. In this way, sufficient context information is given to the model, which can ensure the accuracy of the model for speech recognition.
[0133] In step S123, a speech segment of the fourth duration at the tail of a preset second proportion of the long sample speech sequences is trimmed to obtain a trimmed long sample speech sequence; the second proportion is less than the first proportion.
[0134] Because the phenomenon of deleting the last word mainly occurs in short sentences. Long sentences have more context information, and the prediction delay is not so serious. Therefore, the trimming proportion of long sentences is less than that of short sentences.
[0135] Exemplarily, the second proportion is set to 50%. In this way, the tails of 50% of the long sample speech sequences are trimmed, and the trimming length is randomly selected between 0.1 seconds and 0.4 seconds.
[0136] In step S124, the uncropped short sample speech sequences, the cropped short sample speech sequences, the uncropped long sample speech sequences, and the cropped long sample speech sequences in the short sample speech sequences are combined to obtain a second sample speech sequence.
[0137] After cropping, both the uncropped speech sequences and the cropped speech sequences are used as training data.
[0138] In the embodiments of the present disclosure, a speech segment with a fifth duration is preset at the tail of the short sample speech sequences with a first ratio, and a speech segment with the fourth duration is preset at the tail of the long sample speech sequences with a second ratio. In this way, not all speech sequences need to be cropped, ensuring the integrity of some speech sequences and enabling the model to learn diverse semantic expressions as much as possible.
[0139] Moreover, since the prediction latency of long sentences is not as serious as that of short sentences, the cropping probability of long sample speech sequences is less than that of short sample speech sequences, improving the accuracy of long sentence prediction and reducing the problem of deleting the last characters of short sentences.
[0140] In a possible implementation manner, before step S13, it further includes:
[0141] In step S131, a target array is set. The number of elements in the target array is N, and the values of the elements are uniformly distributed between a first value and a second value. The first value is the value of the third duration, and the second value is the value of the first duration; N is a natural number;
[0142] The total number of training times for the iterative training is M times. The iterative training is divided into three stages. The first stage of the iterative training is from the 1st to the Pth time, the second stage is from the (P + 1)th to the Qth time, and the third stage is from the (Q + 1)th to the Mth time, where P, Q, and M are all natural numbers, 1 < P < M, and P < Q < M; Taking the third sample speech sequence with a second duration from the second sample speech sequence, the step of taking the third sample speech sequence with a second duration from the second sample speech sequence in step S13 includes:
[0143] In step S132, during the first stage of training, each time a first element is randomly selected from the target array, and a third sample speech sequence with a duration corresponding to the first element is taken from the second sample speech sequence;
[0144] In step S133, during the second stage of training, each time a second element is randomly selected between the first element and the N / 2th element of the target array, and a third sample speech sequence with a duration corresponding to the second element is taken from the second sample speech sequence;
[0145] In step S134, during the third-stage training, each time a third sample speech sequence with a duration corresponding to the first value is taken.
[0146] In steps S131 - S134, a target array [a 1, a2, a3... a N is set, and the values of N array elements are distributed between a first value and a second value. The first value is the standard duration of the training data for the streaming speech recognition model, and the second value is the standard duration of the training data for the non-streaming speech recognition model.
[0147] In the initial training, for each batch of data, an element is randomly selected from the array as the chunk of this batch of data for training. As the number of training iterations progresses, larger chunks are gradually removed from the array until, in the final iteration, only the first value remains in the array, which is consistent with the configuration of the actual streaming model. This method helps to find a balance between non-streaming and streaming, enabling the non-streaming model to retain more non-streaming knowledge during the process of migrating to the streaming model.
[0148] Let the number of training times for iterative training be M times. The M - time training is divided into three stages. The first stage is from the 1st to the Pth time, the second stage is from the (P + 1)th to the Qth time, and the third stage is from the (Q + 1)th to the Mth time.
[0149] During the first-stage training, for each batch of data, a first element is randomly selected from the target array, and the duration of the first element is used as the chunk of this batch of data. That is, each time a third sample speech sequence with a duration corresponding to the first element is taken from the second sample speech sequence for training. For example, for a batch of data, a random element is selected from the target array [a 1, a2, a3... a N , such as a3. Each time a third sample speech sequence with a duration corresponding to a3 is taken from this batch of data for training.
[0150] During the second-stage training, for each batch of data, a second element is randomly selected between the first element and the N / 2 - th element of the target array. For example, the first element and the N / 2 - th element of the target array [a 1, a2, a3... a N are a1 and a2 respectively. A random element is selected from a1 and a2, such as a2. Each time a third sample speech sequence with a duration corresponding to a2 is taken from this batch of data for training.
[0151] During the third-stage training, each time a third sample speech sequence with a duration corresponding to the first value is taken. For example, each time a third sample speech sequence with a duration corresponding to a1 is taken from this batch of data for training.
[0152] In this way, as the number of training iterations progresses, large chunks are gradually removed from the array until, in the final iteration, the array only contains the standard duration of the training data for the streaming speech recognition model, which is consistent with the configuration of the actual streaming model. This training method helps to find a balance between non-streaming model training and streaming model training, enabling the non-streaming speech recognition model to retain more non-streaming knowledge during the process of migrating to the streaming speech recognition model.
[0153] In a possible implementation manner, step S14 includes:
[0154] In step S141, a fourth sample speech sequence with a duration corresponding to the first value is taken from the second sample speech sequence, and the fourth sample speech sequence is input into the initial streaming speech recognition model, and the initial streaming speech recognition model is controlled to perform real-time speech recognition processing during the input process of the fourth sample speech sequence; the first value is the preset standard duration of the truncated speech chunks used for non-streaming speech recognition model training;
[0155] In step S142, the features generated by the initial streaming speech recognition model in the last neural network layer are obtained to get the first deep features;
[0156] In step S143, a fourth sample speech sequence with a duration corresponding to the first value is taken from the second sample speech sequence, and the fourth sample speech sequence is input into the non-streaming speech recognition model, and the non-streaming speech recognition model is controlled to perform corresponding speech recognition processing after each input of the fourth sample speech sequence;
[0157] In step S144, the features generated by the non-streaming speech recognition model in the last neural network layer are obtained to get the second deep features;
[0158] In step S145, a first loss value between the first deep features and the second deep features is calculated according to the target loss function;
[0159] In step S146, the model parameters of the initial streaming speech recognition model are updated according to the first loss value to obtain the first streaming speech recognition model.
[0160] In steps S141 - S146, each time a fourth sample speech sequence with a duration corresponding to the first value is taken from the second sample speech sequence and input into the initial streaming speech recognition model for real-time speech recognition processing. When the sample speech sequence input into the model passes through the neural network layers of the model, each layer will generate corresponding features, and the features of the last neural network layer are taken to obtain the first deep features.
[0161] Each time, a fourth sample voice sequence with a duration corresponding to the first value is also taken from the second sample voice sequence and input into the non-streaming speech recognition model, and the model is controlled to perform speech recognition processing after the input of the fourth sample voice sequence is completed. When the sample voice sequence input into the model passes through the neural network layer of the model, each layer generates corresponding features, and the features of the last neural network layer are taken to obtain the second depth feature.
[0162] In contrastive learning, the output of the non-streaming speech recognition model is used as a reference, that is, the second depth feature is used as a reference, and the first loss value between the first depth feature and the second depth feature is calculated. The parameters of the streaming speech recognition model are updated using the first loss value, and the streaming speech model is iteratively trained until the training end condition is met, and the first streaming speech recognition model is obtained.
[0163] The present disclosure uses the output of the non-streaming speech recognition model as a reference, and through contrastive learning, enables the streaming speech recognition model to better learn the feature representation and modeling ability of the non-streaming speech recognition model.
[0164] In a possible implementation manner, step S145 includes:
[0165] In step S1451, the value of the second depth feature is obtained after stopping the gradient operation on the non-streaming speech recognition model, and the third depth feature is obtained;
[0166] In step S1452, the first similarity between the first depth feature and the third depth feature is determined;
[0167] In step S1453, the first depth feature is disassembled from the time dimension to obtain a plurality of one-dimensional frequency dimension features with different time dimensions, and a time dimension is randomly selected from the plurality of time dimensions and combined with the frequency dimension feature to obtain the fourth depth feature corresponding to the time dimension;
[0168] In step S1454, the second similarity between the first depth feature and the fourth depth feature is determined;
[0169] In step S1455, the value of the target loss function is determined based on the first similarity, the second similarity, and a preset temperature coefficient to obtain the first loss value; the temperature coefficient is used to adjust the result of the similarity calculation.
[0170] In steps S1451 - S1455, the symbol for stopping the gradient calculation is sg, which is usually used to represent "stop gradient" or "gradient separation". This concept is used to prevent the gradient from propagating through certain nodes during the backpropagation process.
[0171] After obtaining the first similarity and the second similarity, respectively determine the first quotient of the first similarity and a preset temperature coefficient, and the second quotient of the second similarity and the temperature coefficient. Here, the temperature coefficient is used to adjust the result of similarity calculation. If τ is less than 1, the calculation result of the similarity changes non-linearly faster; if τ is greater than 1, the calculation result of the similarity changes non-linearly slower. Generally, τ is set greater than 1 to form a slow change process.
[0172] Then perform an exponential operation on the first quotient to obtain a first value, perform an exponential operation on the second quotient to obtain a second value. Calculate the quotient of the first value and the second value, and perform a logarithmic operation on the quotient, and take the opposite of the result of the logarithmic operation to obtain a first loss value.
[0173] In this solution, the streaming model representation is a weak representation, and the non-streaming model representation is a strong representation. The calculation of the first loss value is to make the weak representation approach the strong representation, and the strong representation is fixed without update, so there is no need to perform gradient operations on the model corresponding to the strong representation (non-streaming model). Therefore, stop the gradient operation on the strong representation model (non-streaming model).
[0174] The objective loss function proposed in this disclosure can make the streaming features approach the non-streaming features without affecting the non-streaming features, and at the same time, pull the distance between the first depth feature and the fourth depth feature of the negative sample farther to increase the diversity of the features.
[0175] In a possible implementation manner, after step S14, it further includes:
[0176] In step S15, pad the tails of each initial test speech sequence for performing the test with a preset fifth duration of silent segments to obtain a plurality of first test speech sequences;
[0177] In step S16, take a second test speech sequence of the fifth duration from the first test speech sequences, input the second test speech sequence into the first streaming speech recognition model for speech recognition processing to obtain a first text sequence;
[0178] In step S17, calculate a second loss value between the first text sequence and the true text sequence corresponding to the initial test speech sequence, and update the model parameters of the first streaming speech recognition model according to the second loss value to obtain a second streaming speech recognition model.
[0179] In steps S15 - S17, after obtaining the first streaming speech recognition model through training, it enters the stage of performance testing on the first streaming speech recognition model. In this stage, in order to further eliminate the tail emission delay, tail padding is performed on each initial test speech sequence used for prediction. The padding method is to append a silent segment with a fifth duration to the tail of the speech waveform, and use the initial test speech sequence after appending the silent segment as the first test speech sequence.
[0180] The fifth duration is specifically equal to the duration of one chunk of the first test speech sequence input to the model. For example, if one chunk is set to 0.5 seconds, then the fifth duration is also 0.5 seconds.
[0181] The first streaming speech recognition model performs speech recognition on the first test speech sequence to obtain the first text sequence. Calculate the second loss value between the first text sequence and the true text sequence, and update the model parameters according to the second loss value, thereby obtaining the second streaming speech model.
[0182] Since the second streaming speech model further eliminates the tail emission delay phenomenon, its speech recognition effect is better than that of the first streaming speech model.
[0183] After padding with the silent segment, when the model performs speech recognition, it will also regard the silent segment as the speech to be recognized. In this way, it will extend the recognition time of the model for the speech sequence, so that the model also has sufficient time to recognize and output the words at the end of the sentence. Therefore, the method of silent padding can basically eliminate the phenomenon of tail word deletion. And since what is added is a silent segment, only a small amount of computational effort is increased, and the computational burden on the model will not be increased.
[0184] Figure 2 It is a structural block diagram of a training device for a streaming speech recognition model shown according to an exemplary embodiment.
[0185] As Figure 2 shown, the device includes:
[0186] A non - streaming training module 21, configured to execute taking a first - duration first sample speech sequence from the initial sample speech sequence and inputting it into the initial model, performing corresponding speech recognition processing after each input of the first sample speech sequence is completed, and iteratively training the initial model according to the output speech recognition result to obtain a non - streaming speech recognition model; the first duration is a preset standard duration of the truncated speech block for non - streaming speech recognition model training;
[0187] A cropping module 22, configured to perform cropping processing on the tail of the initial sample voice sequence according to the duration of the initial sample voice sequence to obtain a cropped sample voice sequence; combining the cropped sample voice sequence and the uncropped sample voice sequence in the initial sample voice sequence to obtain a second sample voice sequence;
[0188] A streaming training module 23, configured to perform taking a third sample voice sequence with a second duration from the second sample voice sequence, inputting the third sample voice sequence into the non-streaming speech recognition model, performing real-time speech recognition processing during the input process of the third sample voice sequence, and performing iterative training on the non-streaming speech recognition model according to the output speech recognition result to obtain an initial streaming speech recognition model; the second duration is a dynamically changing duration, the second duration is greater than or equal to a preset third duration and less than or equal to the first duration, and the third duration is a preset standard duration of a truncated speech block for training a streaming speech recognition model;
[0189] A contrast training module 24, configured to perform using the non-streaming speech recognition model as a teacher model and the initial streaming speech recognition model as a student model, and performing contrast learning training on the initial streaming speech recognition model according to the second sample voice sequence to obtain a first streaming speech recognition model.
[0190] Optionally, the cropping module 22 is specifically configured to perform:
[0191] Determine the magnitude relationship between the duration of the initial sample voice sequence and a preset fourth duration, use the initial sample voice sequence with a duration less than the fourth duration as a short sample voice sequence, and use the initial sample voice sequence with a duration greater than or equal to the fourth duration as a long sample voice sequence; the fourth duration is greater than the first duration;
[0192] Crop a voice segment with a preset fifth duration at the tail of a preset first ratio of the short sample voice sequences to obtain a cropped short sample voice sequence; the fifth duration is a random value within a preset duration range;
[0193] Crop a voice segment with the fourth duration at the tail of a preset second ratio of the long sample voice sequences to obtain a cropped long sample voice sequence; the second ratio is less than the first ratio;
[0194] Combine the uncropped short sample voice sequences in the short sample voice sequences, the cropped short sample voice sequences, the uncropped long sample voice sequences in the long sample voice sequences, and the cropped long sample voice sequences to obtain a second sample voice sequence.
[0195] Optionally, the device 20 further includes:
[0196] A setting module, configured to execute setting a target array, where the elements in the target array are N, and the values of the elements are evenly distributed between a first value and a second value, the first value being the value of the third duration, and the second value being the value of the first duration; N is a natural number;
[0197] The total number of training times of the iterative training is M times. The iterative training is divided into three stages. The first stage of the iterative training is from the 1st time to the Pth time, the second stage is from the (P + 1)th time to the Qth time, and the third stage is from the (Q + 1)th time to the Mth time, where P, Q, and M are all natural numbers, 1 < P < M, and P < Q < M; The streaming training module 23 is specifically configured to execute:
[0198] During the first-stage training, each time a first element is randomly selected from the target array, and a third sample speech sequence with a duration corresponding to the first element is selected from the second sample speech sequence;
[0199] During the second-stage training, each time a second element is randomly selected between the first element and the N / 2th element of the target array, and a third sample speech sequence with a duration corresponding to the second element is selected from the second sample speech sequence;
[0200] During the third-stage training, each time a third sample speech sequence with a duration corresponding to the first value is selected.
[0201] Optionally, the contrast training module 24 is specifically configured to execute:
[0202] A fourth sample speech sequence with a duration corresponding to the first value is selected from the second sample speech sequence, and the fourth sample speech sequence is input into the initial streaming speech recognition model, and the initial streaming speech recognition model is controlled to perform real-time speech recognition processing during the input process of the fourth sample speech sequence; the first value is the preset standard duration of the truncated speech block for training the non-streaming speech recognition model;
[0203] The features generated by the initial streaming speech recognition model in the last neural network layer are obtained to obtain the first deep features;
[0204] A fourth sample speech sequence with a duration corresponding to the first value is selected from the second sample speech sequence, and the fourth sample speech sequence is input into the non-streaming speech recognition model, and the non-streaming speech recognition model is controlled to perform corresponding speech recognition processing after each input of the fourth sample speech sequence;
[0205] The features generated by the non-streaming speech recognition model in the last neural network layer are obtained to obtain the second deep features;
[0206] Calculate a first loss value between the first depth feature and the second depth feature according to the target loss function;
[0207] Update the model parameters of the initial streaming speech recognition model according to the first loss value to obtain a first streaming speech recognition model.
[0208] Optionally, the contrast training module 24 is specifically configured to execute:
[0209] Obtain the value of the second depth feature after applying the stop gradient operation to the non-streaming speech recognition model to obtain a third depth feature;
[0210] Determine a first similarity between the first depth feature and the third depth feature;
[0211] Break up the first depth feature from the time dimension to obtain a plurality of one-dimensional frequency dimension features with different time dimensions, randomly select one time dimension from the plurality of time dimensions and combine it with the frequency dimension feature to obtain a fourth depth feature corresponding to the time dimension;
[0212] Determine a second similarity between the first depth feature and the fourth depth feature;
[0213] Determine the value of the target loss function based on the first similarity, the second similarity and a preset temperature coefficient to obtain a first loss value; the temperature coefficient is used to adjust the result of the similarity calculation.
[0214] Optionally, the device 20 further includes:
[0215] Padding module 25, configured to execute padding a silent segment with a preset fifth duration at the tail of each initial test speech sequence for performing tests to obtain a plurality of first test speech sequences;
[0216] Input module 26, configured to execute extracting a second test speech sequence with the fifth duration from the first test speech sequences, and inputting the second test speech sequence into the first streaming speech recognition model for speech recognition processing to obtain a first text sequence;
[0217] Calculation module 27, configured to execute calculating a second loss value between the first text sequence and the true text sequence corresponding to the initial test speech sequence, and updating the model parameters of the first streaming speech recognition model according to the second loss value to obtain a second streaming speech recognition model.
[0218] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0219] Figure 3 is a block diagram of an electronic device for a training method of streaming speech recognition shown according to an exemplary embodiment. Its internal structural diagram can be as Figure 3 shown. The server or electronic device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the server or electronic device is used to provide computing and control capabilities. The memory of the server or electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the server or electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a training method for a streaming speech recognition model.
[0220] Those skilled in the art can understand that Figure 3 the structure shown in
[0221] is only a block diagram of some structures related to the solution of the present disclosure, and does not constitute a limitation on the server or electronic device to which the solution of the present disclosure is applied. The specific server or electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0222] In an exemplary embodiment, there is also provided a server or electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the training method for a streaming speech recognition model as in the embodiments of the present disclosure.
[0223] In an exemplary embodiment, there is also provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by the processor of the server or electronic device, enabling the server or electronic device to execute the training method for a streaming speech recognition model in the embodiments of the present disclosure. The computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0224] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0225] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0226] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A training method for a streaming speech recognition model, characterized in that, Including: Taking a first-sample speech sequence of a first duration from the initial sample speech sequence and inputting it into the initial model. After each input of the first-sample speech sequence is completed, corresponding speech recognition processing is performed, and the initial model is iteratively trained according to the output speech recognition result to obtain a non-streaming speech recognition model; the first duration is the preset standard duration of the truncated speech block for training the non-streaming speech recognition model; According to the duration of the initial sample speech sequence, trimming the tail of the initial sample speech sequence to obtain a trimmed sample speech sequence; combining the trimmed sample speech sequence and the untrimmed sample speech sequence in the initial sample speech sequence to obtain a second sample speech sequence; Taking a third-sample speech sequence of a second duration from the second sample speech sequence, inputting the third-sample speech sequence into the non-streaming speech recognition model, performing real-time speech recognition processing during the input of the third-sample speech sequence, and iteratively training the non-streaming speech recognition model according to the output speech recognition result to obtain an initial streaming speech recognition model; the second duration is a dynamically changing duration, the second duration is greater than or equal to a preset third duration and less than or equal to the first duration, and the third duration is the preset standard duration of the truncated speech block for training the streaming speech recognition model; Using the non-streaming speech recognition model as the teacher model and the initial streaming speech recognition model as the student model, and performing contrastive learning training on the initial streaming speech recognition model according to the second sample speech sequence to obtain a first streaming speech recognition model.
2. The method according to claim 1, wherein The step of, according to the duration of the initial sample speech sequence, trimming the tail of the initial sample speech sequence to obtain a trimmed sample speech sequence; combining the trimmed sample speech sequence and the untrimmed sample speech sequence in the initial sample speech sequence to obtain a second sample speech sequence, includes: Determining the magnitude relationship between the duration of the initial sample speech sequence and a preset fourth duration, taking the initial sample speech sequence with a duration less than the fourth duration as a short sample speech sequence, and taking the initial sample speech sequence with a duration greater than or equal to the fourth duration as a long sample speech sequence; the fourth duration is greater than the first duration; Trimming a speech segment of a preset fifth duration at the tail of a preset first proportion of the short sample speech sequences to obtain trimmed short sample speech sequences; the fifth duration is a random value within a preset duration range; Trimming a speech segment of the fourth duration at the tail of a preset second proportion of the long sample speech sequences to obtain trimmed long sample speech sequences; the second proportion is less than the first proportion; Combining the untrimmed short sample speech sequences, the trimmed short sample speech sequences, the untrimmed long sample speech sequences, and the trimmed long sample speech sequences in the short sample speech sequences to obtain a second sample speech sequence.
3. The method according to claim 1, characterized in that, Before taking a third-sample speech sequence of a second duration from the second sample speech sequence, it further includes: Set a target array, where the elements in the target array are N, and the values of the elements are evenly distributed between a first value and a second value. The first value is the value of the third duration, and the second value is the value of the first duration; N is a natural number; The total number of training times of the iterative training is M times. The iterative training is divided into three stages. The first stage of the iterative training is from the 1st to the Pth time, the second stage is from the (P + 1)th to the Qth time, and the third stage is from the (Q + 1)th to the Mth time. Among them, P, Q, and M are all natural numbers, 1 < P < M, P < Q < M; The taking of the third sample speech sequence with the second duration from the second sample speech sequence includes: During the first-stage training, each time a first element is randomly selected from the target array, and a third sample speech sequence with the duration corresponding to the first element is taken from the second sample speech sequence; During the second-stage training, each time a second element is randomly selected between the first element and the N / 2th element of the target array, and a third sample speech sequence with the duration corresponding to the second element is taken from the second sample speech sequence; During the third-stage training, each time a third sample speech sequence with the duration corresponding to the first value is taken.
4. The method according to claim 1, wherein Taking the non-streaming speech recognition model as the teacher model and the initial streaming speech recognition model as the student model, and performing contrastive learning training on the initial streaming speech recognition model according to the second sample speech sequence to obtain the first streaming speech recognition model, includes: Taking a fourth sample speech sequence with the first value duration from the second sample speech sequence, inputting the fourth sample speech sequence into the initial streaming speech recognition model, and controlling the initial streaming speech recognition model to perform real-time speech recognition processing during the input process of the fourth sample speech sequence; The first value is the preset standard duration of the truncated speech block for training the non-streaming speech recognition model; Obtaining the features generated by the initial streaming speech recognition model in the last neural network layer to obtain the first deep features; Taking a fourth sample speech sequence with the first value duration from the second sample speech sequence, inputting the fourth sample speech sequence into the non-streaming speech recognition model, and controlling the non-streaming speech recognition model to perform corresponding speech recognition processing after each input of the fourth sample speech sequence; Obtaining the features generated by the non-streaming speech recognition model in the last neural network layer to obtain the second deep features; Calculating a first loss value between the first deep features and the second deep features according to the target loss function; Updating the model parameters of the initial streaming speech recognition model according to the first loss value to obtain the first streaming speech recognition model.
5. The method according to claim 4, characterized in that, The calculating of the first loss value between the first deep features and the second deep features according to the target loss function includes: Obtaining the value of the second deep features after stopping the gradient operation on the non-streaming speech recognition model to obtain the third deep features; Determining a first similarity between the first deep features and the third deep features; Split the first depth feature in the time dimension to obtain multiple one-dimensional frequency dimension features with different time dimensions, randomly select one time dimension from the multiple time dimensions and combine it with the frequency dimension feature to obtain the fourth depth feature corresponding to the time dimension; Determine the second similarity between the first depth feature and the fourth depth feature; Based on the first similarity, the second similarity and a preset temperature coefficient, determine the value of the target loss function to obtain the first loss value; the temperature coefficient is used to adjust the result of similarity calculation.
6. The method according to claim 1, characterized in that, After obtaining the first streaming speech recognition model, it further includes: Pad the tails of each initial test speech sequence for testing with a silent segment of a preset fifth duration to obtain multiple first test speech sequences; Extract a second test speech sequence of the fifth duration from the first test speech sequences, input the second test speech sequence into the first streaming speech recognition model for speech recognition processing to obtain a first text sequence; Calculate the second loss value between the first text sequence and the true text sequence corresponding to the initial test speech sequence, and update the model parameters of the first streaming speech recognition model according to the second loss value to obtain a second streaming speech recognition model.
7. A training device for a streaming speech recognition model, characterized in that, It includes: A non-streaming training module configured to extract a first sample speech sequence of a first duration from an initial sample speech sequence and input it into an initial model, perform corresponding speech recognition processing after each input of the first sample speech sequence is completed, and iteratively train the initial model according to the output speech recognition result to obtain a non-streaming speech recognition model; the first duration is the preset standard duration of the truncated speech block for training the non-streaming speech recognition model; A cropping module configured to perform cropping processing on the tail of the initial sample speech sequence according to the duration of the initial sample speech sequence to obtain a cropped sample speech sequence; combine the cropped sample speech sequence and the uncropped sample speech sequence in the initial sample speech sequence to obtain a second sample speech sequence; A streaming training module configured to extract a third sample speech sequence of a second duration from the second sample speech sequence, input the third sample speech sequence into the non-streaming speech recognition model, perform real-time speech recognition processing during the input of the third sample speech sequence, and iteratively train the non-streaming speech recognition model according to the output speech recognition result to obtain an initial streaming speech recognition model; the second duration is a dynamically changing duration, the second duration is greater than or equal to a preset third duration and less than or equal to the first duration, and the third duration is the preset standard duration of the truncated speech block for training the streaming speech recognition model; A contrast training module configured to use the non-streaming speech recognition model as the teacher model and the initial streaming speech recognition model as the student model, and perform contrast learning training on the initial streaming speech recognition model according to the second sample speech sequence to obtain a first streaming speech recognition model.
8. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the training method of the streaming voice recognition model according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the training method of the streaming voice recognition model according to any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program, the computer program is stored in a readable storage medium, and at least one processor of the computer device reads and executes the computer program, so that the computer device executes the training method of the streaming voice recognition model according to any one of claims 1 to 6.
Citation Information
Cited By
Speech recognition large model training method and device, storage medium and equipment
CN121034291A