Speech recognition model training method and device and electronic equipment
By performing multiple speech enhancements on the speech data and combining iterative distillation training of the teacher model and student model, multiple loss values are calculated to adjust the student model parameters, solving the problem of overfitting the student model caused by knowledge distillation, and improving the accuracy and generalization of speech recognition.
Patent Information
- Application Number
- CN202510117569.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
Knowledge distillation causes student models to overfit the training data in the speech recognition model, which has poor generalization, affecting the accuracy of end-to-end speech recognition.
By performing multiple voice enhancements on the audio in the pre-noted speech dataset, multiple spectral diagrams are generated, and iterative distillation training is used to calculate hard loss values, soft loss values and consistency regularization loss values, and adjust the network parameters of the student model.
The accuracy and generalization of the speech recognition model are improved. Through the multiple speech enhancement and loss value calculation methods, the student model can better learn the speech feature representation, reduce overfitting, and output a smoother probability distribution.
Smart Images

Figure CN120048251A_ABST
Abstract
Description
Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, a variety of smart devices have been launched on the market. AI voice has made voice assistants an indispensable software for major smart devices. Automatic speech recognition (ASR) is a technology that converts speech into text and is one of the key technologies for human-smart device interaction.
[0003] In order to ensure the accuracy of speech recognition, the trained models are usually more complex and require a lot of computing resources and data sets to support them. However, in practical applications, complex speech recognition models cannot be deployed for smart devices with weak chip computing power or embedded devices. Therefore, in order to create a lightweight model that can match the accuracy of the initial model and is easy to use in practice, a proven model compression method - knowledge distillation - is proposed.
[0004] In knowledge distillation, the student model is trained by the teacher model. The same data set is used for training. The feature representation "knowledge" learned by the complex and strong learning teacher network is distilled and passed to the student network with small parameters and weak learning ability. Through knowledge distillation, the student model obtains the feature extraction ability of the teacher model while ensuring the characteristics of lightweight and easy deployment.
[0005] However, knowledge distillation focuses on narrowing the gap between the outputs of the teacher and student models, which may cause the student model to overfit the training data and have poor generalization, affecting the accuracy of end-to-end speech recognition. Summary of the invention
[0006] The embodiments of the present application provide a method, device and electronic device for training a speech recognition model, which are used to improve the accuracy of a lightweight speech recognition model.
[0007] In a first aspect, an embodiment of the present application provides a method for training a speech recognition model, comprising:
[0008] The audio in the pre-annotated first speech data set is input into the trained teacher model and the student model to be trained in batches for iterative distillation training until the preset stop condition is met, wherein each iterative distillation training includes:
[0009] Perform multiple speech enhancements on a piece of audio to obtain the spectrogram after each enhancement;
[0010] Using the teacher model to perform label prediction on multiple spectrograms of the audio segment, respectively, to obtain a first probability distribution of each spectrogram;
[0011] Using the student model to perform label prediction on multiple spectrograms of the audio segment respectively, and obtaining the second probability distribution of each spectrogram;
[0012] Calculating a hard loss value according to the predicted label output by the student model and the corresponding true label of the annotation, and calculating a soft loss value according to the first probability distribution and the second probability distribution corresponding to the same spectrogram, and calculating a consistency regularization loss value according to the second probability distributions of the multiple spectrograms;
[0013] Weighting the hard loss value, the soft loss value and the consistency regularization loss value to generate a target loss value, and adjusting the network parameters of the student model according to the target loss value.
[0014] The beneficial effects of the above technical solution are as follows: In the field of speech recognition, for the scenario of distillation training of the teacher model and the student model, each audio segment in the first speech dataset is subjected to multiple spectrum enhancement processes. In this way, when the teacher model and the student model perform label prediction based on the multiple spectrograms after spectrum enhancement, they can fully learn the speech feature representation from different data levels, thereby improving the accuracy of speech recognition. On the other hand, the parameters of the teacher model are generally larger than those of the student model, and the learning ability is strong. Therefore, during the distillation training process, in addition to calculating the hard loss value based on the predicted label corresponding to the second probability distribution output by the student model and the corresponding true label of the annotation, the first probability distribution result output by the teacher model is also used as a soft label, and the soft loss value is calculated according to the first probability distribution and the second probability distribution corresponding to the same spectrogram, realizing the training of the student model guided by the knowledge learned by the teacher model with strong learning ability, and further improving the accuracy of the student model's speech recognition. In addition, the consistency regularization loss value is calculated through the second probability distributions of multiple spectrograms after spectrum enhancement, increasing the learning of the context relationship between different spectrograms, making the speech recognition result smoother, thereby reducing the overfitting of the hard loss value to the label and improving the generalization of the model.
[0015] Optionally, the performing multiple speech enhancements on an audio segment and obtaining the spectrogram after each enhancement includes:
[0016] Performing a time warping operation on the audio segment to obtain a target audio;
[0017] Extracting the acoustic features of the target audio and creating multiple copies of the acoustic features;
[0018] Performing frequency-domain and time-domain masking on each copy to generate a spectrogram after speech enhancement.
[0019] The beneficial effects of the above technical solution are as follows: perform a warping operation on the time axis before creating a copy to prevent obvious timestamp mismatches between the branch results of multiple spectrograms output by the student model, improve the accuracy of speech recognition, and increase the data diversity by performing frequency-domain and time-domain masking processing on each copy to obtain different spectrograms. In this way, during the training process, different spectrograms can enable the model to learn knowledge of different speech representations, enhance the diversity of predictions, promote richer knowledge transfer, and complementary representation learning, and further improve the accuracy of speech recognition.
[0020] Optionally, the piece of audio includes multiple frames of speech. The calculating the consistency regularization loss value according to the second probability distributions of the multiple spectrograms includes:
[0021] Perform frame splitting on each spectrogram according to the number of speech frames to obtain sub-spectrograms of each frame of speech;
[0022] Minimize the bidirectional distance between the second probability distributions of two sub-spectrograms corresponding to each frame of speech to generate a consistency regularization loss value.
[0023] The beneficial effects of the above technical solution are as follows: the student model performs self-distillation between two spectrograms. During prediction, the prediction result of each frame of speech in one spectrogram is used as the supervision signal for the corresponding frame in the other spectrogram, enforce consistency regularization between the two spectrograms, and generate a consistency regularization loss value by minimizing the bidirectional distance between the second probability distributions of two sub-spectrograms corresponding to each frame of speech, forcing the student model to learn the context relationship between different spectrograms, reducing the overfitting of the hard loss value to the label, outputting a smoother probability distribution, thus suppressing the spike effect, avoiding overfitting on the training data set, and improving the generalization of the model.
[0024] Optionally, the formula of the target loss value is expressed as:
[0025]
[0026] Where represents the hard loss value, represents the soft loss value, CR a,b∈n (z a , z b ) represents the consistency regularization loss value, loss represents the target loss value, n is the total number of spectrograms, x i represents the i-th spectrogram, y represents the vocabulary, z a and z b represent the second probability distributions of any two spectrograms among the n spectrograms, Represents the first probability of each word in the first probability distribution of the t-th frame in the i-th spectrogram, Represents the second probability of each word in the second probability distribution of the t-th frame in the i-th spectrogram, T is the number of frames of the segment of audio, sg() represents applying a stop-gradient operation on a probability distribution, α is a hyperparameter controlling consistency regularization, and β is a hyperparameter for the teacher model to guide the student model to learn.
[0027] The beneficial effects of the above technical solution are as follows: In the target loss value, in addition to the hard loss value calculated based on the predicted label corresponding to the second probability distribution and the corresponding annotated true label, it also includes the soft loss value calculated based on the first probability distribution and the second probability distribution corresponding to the same spectrogram, and the consistency regularization loss value calculated based on the second probability distributions of multiple spectrograms. Among them, in the soft loss value, the result of the first probability distribution output by the teacher model is used as a soft label to guide the training of the student model, enabling the student model to fully learn the knowledge represented by the teacher model, reducing the difference between the output of the student model and the teacher model, thereby improving the accuracy of the student model in speech recognition. The consistency regularization loss value enables the student model to learn the context information between different spectrograms, strengthening the consistency constraint between different second probability distributions, making the prediction result of the student model smoother, thereby reducing the overfitting of the hard loss value to the label and improving the generalization of the model.
[0028] Optionally, before iterative distillation training, the method further includes:
[0029] Performing iterative training on the teacher model and the student model according to the labeled second speech dataset to obtain a pre-trained teacher model and student model; wherein the labels of the audio in the second speech dataset are pinyin without tones, the speech in the second speech dataset does not contain a wake-up word, the data volume of the second speech dataset is larger than that of the first speech dataset, the labels of the audio in the first speech dataset are pinyin without tones, and some of the audio in the first speech dataset contains a wake-up word;
[0030] Fine-tuning the teacher model according to the first speech dataset to obtain a trained teacher model.
[0031] The beneficial effects of the above technical solution are as follows: Since the second voice dataset does not contain a wake-up word, the dependence of the model on labeled data is reduced, the flexibility of data acquisition methods is increased, and the quantity of the second voice dataset is large, enabling the teacher model and the student model to have strong learning ability for the representation of general voices after pre-training. In this way, by fine-tuning the teacher model and the student model with the first voice dataset with a smaller amount of data, the teacher model and the student model can learn the specific features of the voice wake-up task, with strong transferability and low annotation cost. Moreover, using the non-tonal pinyin as the label reduces the influence of various tone changes of the wake-up word caused by dialects in different regions, improving the robustness of the model.
[0032] Optionally, some of the audio in the first voice dataset contains a wake-up word. After obtaining the trained student model, the method further includes:
[0033] Obtain voice wake-up data;
[0034] Input the voice wake-up data into the trained student model for label prediction to obtain the fused pinyin features of each character included in each frame of the voice wake-up data;
[0035] If the length of the smoothing window is greater than 1, determine the initial posterior probability of the corresponding character according to the fused pinyin features of each character;
[0036] Smooth the initial posterior probabilities of each character according to the smoothing window to obtain the target posterior probability;
[0037] Calculate the confidence of the words composed of each character according to the target posterior probabilities of each character included in each frame within the smoothing window, and use the word with the highest confidence as the wake-up word in the corresponding frame.
[0038] The beneficial effects of the above technical solution are as follows: When the lightweight student model is used to identify the wake-up word in the voice data to be recognized, the initial posterior probabilities of each character in each frame are smoothed by a smoothing window with a length greater than 1, thereby reducing the influence of noise. In this way, when calculating the confidence of the wake-up word based on the smoothed target posterior probabilities of each character, the accuracy of voice wake-up can be improved, and the smoothing method of the smoothing window has a small computational amount, is simple to implement, and has reliable output.
[0039] Optionally, if the length of the smoothing window is not greater than 1, the method further includes:
[0040] Divide the fused pinyin features of each character by a preset temperature scaling coefficient to obtain the target pinyin features; where the value of the temperature scaling coefficient is greater than 1;
[0041] Determine the target posterior probability of the corresponding character according to the respective target pinyin features;
[0042] Calculate the confidence of the words formed by each character based on the target posterior probability of each character included in each frame, and use the word with the highest confidence as the wake-up word in the corresponding frame.
[0043] The beneficial effect of the above technical solution is that when the length of the smoothing window is equal to 1, the smoothing effect cannot be achieved. Therefore, before the probability output, the temperature scaling coefficient can be used to smooth the fused pinyin features of each character in each frame, thereby reducing the influence of noise and improving the accuracy of the confidence calculation of the wake-up word in the speech data to be recognized.
[0044] In a second aspect, an embodiment of the present application provides a training device for a speech recognition model, including:
[0045] A training module for batch inputting the audio in the pre-annotated first speech dataset into the trained teacher model and the student model to be trained for iterative distillation training until a preset stop condition is met;
[0046] Wherein, the training module includes:
[0047] A speech enhancement unit for performing speech enhancement on a piece of audio multiple times to obtain the spectrogram after each enhancement;
[0048] A teacher model prediction unit for using the teacher model to perform label prediction on multiple spectrograms of the piece of audio respectively to obtain the first probability distribution of each spectrogram;
[0049] A student model prediction unit for using the student model to perform label prediction on multiple spectrograms of the piece of audio respectively to obtain the second probability distribution of each spectrogram;
[0050] A loss value calculation unit for calculating the hard loss value according to the predicted label output by the student model and the true label of the corresponding annotation, and calculating the soft loss value according to the first probability distribution and the second probability distribution corresponding to the same spectrogram, and calculating the consistency regularization loss value according to the second probability distribution of the multiple spectrograms; weighting the hard loss value, the soft loss value and the consistency regularization loss value to generate a target loss value
[0051] A parameter adjustment unit for adjusting the network parameters of the student model according to the target loss value.
[0052] In a third aspect, an embodiment of the present application provides a speech wake-up device, including:
[0053] A speech acquisition module for acquiring speech wake-up data;
[0054] A feature processing module, configured to input the voice wake-up data into a trained student model for label prediction, so as to obtain the fused pinyin features of each character included in each frame of the voice wake-up data;
[0055] A probability prediction module, configured to, if the length of the smoothing window is greater than 1, determine the initial posterior probability of the corresponding character according to the fused pinyin features of each character;
[0056] A probability smoothing module, configured to smooth the initial posterior probabilities of each character according to the smoothing window to obtain the target posterior probability;
[0057] A wake-up word determination module, configured to calculate the confidence of the words formed by each character according to the target posterior probabilities of each character included in each frame within the smoothing window, and use the word with the highest confidence as the wake-up word in the corresponding frame.
[0058] In a fourth aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, and a communication interface, where the communication interface, the memory, and the processor are connected through a bus;
[0059] The communication interface is configured to send and receive data;
[0060] The memory stores a computer program, and the processor executes the steps of any voice recognition model training method according to the computer program.
[0061] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of any voice recognition model training method are implemented.
[0062] The technical effects brought by any implementation manner in the second aspect to the fifth aspect can refer to the technical effects brought by the corresponding implementation manner in the first aspect, and will not be elaborated here. Description of the Drawings
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0064] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0065] Figure 2 It is a flowchart of a voice recognition model training method provided by an embodiment of the present application;
[0066] Figure 3 Flow chart of a voice enhancement method provided by an embodiment of the present application;
[0067] Figure 4 Schematic diagram of a traditional knowledge distillation training process provided by an embodiment of the present application;
[0068] Figure 5 Schematic diagram of a spike effect provided by an embodiment of the present application;
[0069] Figure 6 Schematic diagram of a consistency regularization process provided by an embodiment of the present application;
[0070] Figure 7 Schematic diagram of a knowledge distillation training process introducing consistency regularization provided by an embodiment of the present application;
[0071] Figure 8 Flow chart of a training method for a voice wake-up model provided by an embodiment of the present application;
[0072] Figure 9A Schematic diagram of the network structure of a voice wake-up model provided by an embodiment of the present application;
[0073] Figure 9B Schematic diagram of the structure of a residual block in a voice wake-up model provided by an embodiment of the present application;
[0074] Figure 10 Flow chart of a voice wake-up method provided by an embodiment of the present application;
[0075] Figure 11 Structure diagram of a training device for a voice recognition model provided by an embodiment of the present application;
[0076] Figure 12 Structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0077] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments recorded in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the technical solutions of the present application.
[0078] Based on the exemplary embodiments shown in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application. In addition, although the disclosure in the present application is introduced according to one or several exemplary instances, it should be understood that each aspect of these disclosures can also constitute a complete technical solution alone.
[0079] In addition, the terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to those components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0080] The term "module" used in the present application refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware and / or software code that can perform functions related to that element.
[0081] Before introducing the data processing method provided by the embodiments of the present application, some terms involved in the data processing method provided by the embodiments of the present application will be explained first.
[0082] Knowledge distillation: A method in which a student network learns knowledge from a teacher network to obtain the essence of the teacher network. After distillation, the model can effectively learn the rich knowledge of the teacher model and improve the recognition accuracy.
[0083] Teacher model: A concept in knowledge distillation, corresponding to the student model. In knowledge distillation, the teacher model is used to train the student model to achieve effects such as improving the accuracy of the student model.
[0084] Connectionist Temporal Classification (CTC): Can be understood as a neural network-based temporal classification. CTC is a method for calculating the loss value, and its advantage is that it can automatically align unaligned data, mainly used for training serialized data without prior alignment, such as speech recognition, text recognition, etc.
[0085] In traditional knowledge distillation, when training the student model with the teacher model, it focuses on narrowing the distance between the outputs of the teacher and student models, resulting in the student model may overfit the training data, with poor generalization, which affects the accuracy of end-to-end speech recognition.
[0086] In view of this, the embodiments of the present application provide a method for training a speech recognition model. When the teacher model trains the student model, frequency enhancement is performed multiple times on each audio segment in the same first speech dataset used by both the teacher and student models to obtain multiple spectrograms for each audio segment. The voice frequency bands in different spectrograms are different. In this way, when the teacher model and the student model perform label prediction on the multiple spectrograms of each audio segment, they can fully learn the speech feature representation from different data levels, thereby improving the accuracy of speech recognition. Considering that the parameters of the teacher model in knowledge distillation are generally larger than those of the student model and the speech recognition accuracy is relatively high, during the training process of the student model, in addition to calculating the hard loss value based on the predicted label and the true label, the first probability distribution result output by the teacher model can also be used as a soft label, and the soft loss value is calculated according to the first probability distribution and the second probability distribution corresponding to the same spectrogram to guide the training of the student model, further improving the accuracy of the student model's speech recognition. Moreover, the student model calculates the consistency regularization loss value through the second probability distribution of multiple spectrograms, increasing the consistency constraint on the output of the student model, making the speech recognition result smoother, thereby reducing the overfitting of the hard loss value to the label and improving the generalization of the model.
[0087] The following will combine the accompanying drawings and explain in detail a method for training a speech recognition model provided by the embodiments of the present application through specific embodiments and their application scenarios.
[0088] See Figure 1 , which is a schematic diagram of the application scenario provided by the embodiments of the present application. The user can operate the intelligent device 200 through the mobile terminal 300 and the control device 100.
[0089] In some embodiments, the control device 100 may be a remote control. The communication between the remote control and the intelligent device includes infrared protocol communication or Bluetooth protocol communication, as well as other short-distance communication methods, etc., to control the intelligent device 200 wirelessly or through other wired methods.
[0090] In some embodiments, a mobile terminal 300 such as a mobile phone, a tablet computer, a computer, a laptop computer, etc. can also be used to control the intelligent device 200. For example, an application program running on the mobile terminal is used to control the intelligent device 200. The application program can provide various controls for the user in an intuitive user interface (UI) on the screen associated with the mobile terminal through configuration.
[0091] In some embodiments, software applications can be installed on the mobile terminal 300 and the intelligent device 200, and connection communication is achieved through a network communication protocol to achieve the purpose of one-to-one control operation and data communication.
[0092] In some embodiments, voice assistants are installed on the control device 100, the mobile terminal 300, and the controlled intelligent device 200. In this way, when the user uses the control device 100 or the mobile terminal 300 to control the intelligent device 200, the user instruction can be input in a voice manner.
[0093] In some embodiments, voice recognition can be performed by the control device 100 or the mobile terminal 300, or by the intelligent device 200. Taking the intelligent device 200 as an example, the intelligent device 200 receives the voice signal sent by the user through the control device 100 or the mobile terminal 300, parses the voice signal through the deployed voice recognition model, obtains the user's voice instruction, and responds to the voice instruction.
[0094] As Figure 1 also shown in, the intelligent device 200 also communicates with the server 400 through various communication methods for data communication.
[0095] It should be noted that Figure 1 is only an example of an application scenario. In addition to being applied to the above control device 100, intelligent device 200, and mobile terminal 300, the voice recognition model can also be applied to in-vehicle terminals, smart speakers, smart refrigerators, wearable devices, etc.
[0096] In knowledge distillation, the network structures of the teacher model and the student model are the same. Among them, the backbone network in the network structure can be flexibly selected according to actual needs, including but not limited to ResNet network, Bert network, or Transform network, etc.
[0097] In some embodiments, in order to reduce the alignment operation of voice data in the time dimension during the training data preparation process, the teacher model and the student model can adopt the CTC algorithm.
[0098] Based on knowledge distillation, the training method flow of a voice recognition model provided by an embodiment of the present application is as Figure 2 shown. The audio in the pre-annotated first voice dataset is input into the trained teacher model and the student model to be trained in batches for iterative distillation training until a preset stop condition is met. Among them, the preset stop condition includes but is not limited to the convergence of the student model, reaching the preset number of iterations, etc. Each iterative distillation training includes the following steps:
[0099] S201: Perform voice enhancement on a piece of audio multiple times to obtain the spectrogram after each enhancement.
[0100] In the field of speech recognition, data diversity is increased by performing frequency-domain enhancement on audio. In one embodiment, SpecAugment, as an open-source project, enhances spectrograms through frequency-domain, time-domain, and region masking, improving the anti-interference ability and adaptability of the model.
[0101] As Figure 3 shown, the process of applying SpecAugment to perform speech enhancement on speech data mainly includes the following steps:
[0102] S2011: Perform a time warping operation on a segment of audio to obtain the target audio.
[0103] Since audio has temporal characteristics, SpecAugment involves warping along the time axis, masking frequency channels, and masking time steps, etc. The warping of the time axis will change the temporal characteristics of the audio, thus changing the timestamps output by different spectrograms in the model. Therefore, before creating a copy of the audio, perform a time warping operation on the currently input audio to prevent obvious timestamp mismatches between the outputs corresponding to different spectrograms and improve the accuracy of speech recognition.
[0104] S2012: Extract the acoustic features of the target audio and create multiple copies of the acoustic features.
[0105] In some embodiments, Fbank features are used as the acoustic features because Fbank features are more in line with the essence of speech signals and fit the acceptance characteristics of the human ear. Compared with MFCC features, Fbank features do not perform Discrete Cosine Transform (DCT). Since DCT is a linear transform, some originally highly non-linear components in the speech signal will be lost, and in deep learning, neural networks are not sensitive to highly correlated information. Therefore, Fbank features perform better than MFCC features in neural networks.
[0106] Since the frequency-domain and frequency-domain masking processing in SpecAugment is random and the spectrograms generated each time are different, to improve data diversity, multiple copies can be generated by copying the Fbank features for frequency-domain and frequency-domain masking processing.
[0107] S2013: Perform frequency-domain and frequency-domain masking on each copy to generate spectrograms after speech enhancement.
[0108] For each copy, randomly select a frequency band and set it completely to zero to simulate some obstacles that the auditory system may encounter (such as noise, data loss, etc.), thereby improving the anti-interference ability of the model. Also, randomly select a time window and set the spectrogram within this time window to zero to enhance the complexity of the sound, thereby improving the adaptability of the model.
[0109] In the embodiments of the present application, through frequency-domain and time-domain masking processing on multiple copies of the time-warped acoustic features, multiple spectrograms are obtained. Since the voice frequency bands and signal amplitudes in different spectrograms are different, the data diversity is increased. In this way, during the training process, different spectrograms can enable the model to learn the knowledge of different speech representations, enhance the diversity of predictions, promote richer knowledge transfer and complementary representation learning, and further improve the accuracy of speech recognition.
[0110] S202: Use the teacher model to perform label prediction on multiple spectrograms of an audio segment respectively to obtain the first probability distribution of each spectrogram.
[0111] In some embodiments, when the teacher model trains the student model, the target knowledge distillation method is adopted. Since the teacher model has a large number of parameters, deep layers, and large receptive fields, it can learn complex and abstract acoustic features and has good performance in scenarios with multiple noise sources or speech interruptions. The first probability distribution output by it can accurately represent the sound features. Therefore, in target knowledge distillation, the knowledge provided by the teacher model to the student is the first probability distribution of its output.
[0112] S203: Use the student model to perform label prediction on multiple spectrograms of an audio segment respectively to obtain the second probability distribution of each spectrogram.
[0113] Since the network structure of the student model is the same as that of the teacher model, but it has fewer parameters and is lighter, it aims to learn and imitate the prediction results from the more complex and better-performing teacher model to obtain a second probability distribution closer to that of the teacher model.
[0114] S204: Calculate the hard loss value according to the predicted label output by the student model and the corresponding annotated true label, and calculate the soft loss value according to the first probability distribution and the second probability distribution corresponding to the same spectrogram, and calculate the consistency regularization loss value according to the second probability distributions of multiple spectrograms.
[0115] As Figure 4 shown, in target knowledge distillation, the teacher model provides the first probability distribution output for each frame as the soft label to the student model. The student model obtains the soft loss value by calculating the distance between the soft label and the second probability distribution of each frame output by itself. At the same time, the student model uses the second probability distribution output by itself as the hard label. The hard label consists of a 1 and multiple 0s. Therefore, the category corresponding to the probability value of 1 in the hard label is the predicted label of the student model. By calculating the difference between the predicted label and the true label, the hard loss value is obtained. Finally, the total loss of knowledge distillation training is obtained by weighting the soft loss and the hard loss for backpropagation.
[0116] In some embodiments, the soft loss is the distance between the probability distributions P output by the teacher model and the student model, which can be measured by KL divergence, JS divergence, Mean Square Error (MSE), etc.
[0117] In some embodiments, the hard loss is the cross-entropy loss. In a speech recognition task, a speech sequence frame of length T is usually converted into a text sequence of length U where γ is the vocabulary, i.e., the label. The CTC algorithm expands the vocabulary γ into γ′ = γ ∪ {∈} by introducing blank frames ∈, and its goal is to maximize the total posterior probability of all valid alignments between x and y The calculation formula for the cross-entropy loss in the CTC algorithm is:
[0118]
[0119] where B(π) represents a many-to-one mapping used to merge duplicate labels and remove all blanks p(π|x) in π represents the posterior probability of the alignment.
[0120] However, when training with the CTC loss, the student model tends to give relatively high posterior probability prediction results only for a few frames, and the posterior probabilities given for other frames are very small (close to 0) to obtain a smaller loss value. In this way, the probability distribution output by the student model will show spikes one by one, and almost all non-blank labels only occupy one frame, while the remaining frames are dominated by blank labels, as Figure 5 shown, there are four spikes in an audio segment. This spike effect indicates that the student model may be overfitting to the data in the training set and is inaccurate in predicting new data outside the training set, with poor generalization.
[0121] In some embodiments, to address the spike effect in the CTC algorithm, the Consistency Regularization Connectionist Temporal Classification (CR-CTC) algorithm is adopted in knowledge distillation. Based on the CTC algorithm, the CR-CTC algorithm adds a consistency regularization loss value, thereby providing stronger constraints and better context knowledge to the student model during distillation training.
[0122] As Figure 6 shown, for the consistency regularization process in the CR-CTC algorithm, two spectrograms after speech enhancement are used as the input to the student model. These two spectrograms are respectively speech-encoded by two encoders f with weight sharing to obtain the second probability distribution z of each speech sequence frame in each spectrogram a= f{x a}, z b = f{x b}, so that during the training process, in addition to calculating the CTC loss (including CTC(x a , y) and CTC(x b , y)) based on the predicted labels corresponding to each second probability distribution and the true labels, the student model also performs self-distillation between the two second probability distributions to introduce a consistency regularization loss CR(z a , z b ), thereby strengthening the consistency between the two second probability distributions, guiding the student model to learn the average value of the two second probability distributions, finally outputting a smoother probability distribution, effectively suppressing the spike effect of the CTC loss, reducing the overconfidence of the model on the training dataset, and improving the generalization ability of the model.
[0123] As Figure 7 shown, for the target knowledge distillation process of the CR-CTC algorithm, when the student model is trained, according to the predicted labels corresponding to the second probability distribution of each spectrogram output by itself, combined with the annotated true labels, the hard loss value is calculated, and, taking the first probability distribution of each spectrogram output by the teacher model as the soft label, combined with the second probability distribution of the corresponding spectrogram output by itself, the soft loss value is calculated, and, according to the second probability distributions of multiple spectrograms output by itself, the consistency regularization loss value is calculated.
[0124] Each time during training, an input audio contains multiple frames of speech. When performing self-distillation on the student model based on multiple spectrograms of an audio, the consistency regularization is applied to each frame of speech in the spectrogram. At this time, the calculation process of the consistency regularization loss value includes: performing frame splitting on each spectrogram according to the number of speech frames of an audio to obtain sub-spectrograms of each frame of speech, and minimizing the bidirectional distance between the second probability distributions of the two sub-spectrograms corresponding to each frame of speech to generate the consistency regularization loss value.
[0125] In some embodiments, the distance between the second probability distributions can be measured by KL divergence, JS divergence, Mean Square Error (MSE), etc.
[0126] In an embodiment of the present application, the student model performs self-distillation between two spectrograms. During prediction, the prediction result of each frame of speech in one spectrogram is used as the supervision signal for the corresponding frame in the other spectrogram, enforcing consistency regularization between the two spectrograms. By minimizing the bidirectional distance between the second probability distributions of two sub-spectrograms corresponding to each frame of speech, a consistency regularization loss value is generated, forcing the student model to learn the context relationship between different spectrograms, reducing the overfitting of the hard loss value to the labels, outputting a smoother probability distribution, thereby suppressing the spike effect of the CTC algorithm, avoiding overfitting on the training dataset, and improving the generalization of the model.
[0127] Similarly, when using the teacher model to perform target knowledge distillation on the student model, the soft loss is also applied to each frame of speech in the spectrogram. Specifically, each spectrogram is framed according to the number of speech frames in an audio segment to obtain the sub-spectrogram of each frame of speech, and the distance between the first probability distribution of the sub-spectrogram of each frame of speech and the second probability distribution of the sub-spectrogram of the corresponding speech frame is calculated to obtain the soft loss value. Since the teacher model has a large number of parameters, deep layers, a large receptive field, strong feature representation ability, and high prediction accuracy, the first probability distribution output by it is used as the soft label to guide the learning of the student model, so that the second probability distribution output by the student model can better approach the first probability distribution, thereby obtaining a more accurate prediction result.
[0128] S205: Weight the hard loss value, the soft loss value, and the consistency regularization loss value to generate a target loss value, and adjust the network parameters of the student model according to the target loss value.
[0129] Specifically, the formula for the target loss value is expressed as follows:
[0130]
[0131]
[0132] Among them, represents the hard loss value, represents the soft loss value, CR a,b∈n (z a ,z b ) represents the consistency regularization loss value, loss represents the target loss value, n is the total number of spectrograms, x i represents the i-th spectrogram, y represents the vocabulary, z a and z b represent the second probability distributions of any two spectrograms among the n spectrograms, represents the first probability of each word in the first probability distribution of the t-th frame in the i-th spectrogram, Represents the second probability of each word in the second probability distribution of the t-th frame in the i-th spectrogram. T is the number of frames of an audio segment. sg() represents applying a stop-gradient operation on a probability distribution. α is a hyperparameter controlling consistency regularization, and β is a hyperparameter for the teacher model to guide the student model's learning.
[0133] In the above target loss value, in addition to the hard loss value calculated based on the predicted label corresponding to the second probability distribution and the corresponding annotated true label, it also includes the soft loss value calculated based on the first probability distribution and the second probability distribution corresponding to the same spectrogram, and the consistency regularization loss value calculated based on the second probability distributions of multiple spectrograms. Among them, in the soft loss value, the first probability distribution result output by the teacher model is used as a soft label to guide the training of the student model, enabling the student model to fully learn the knowledge represented by the teacher model, reducing the difference between the outputs of the student model and the teacher model, thereby improving the accuracy of the student model's speech recognition. The consistency regularization loss value can enable the student model to learn the context information between different spectrograms, strengthen the consistency constraint between different second probability distributions, make the prediction results of the student model smoother, thereby reducing the overfitting of the hard loss value to the label, and improving the generalization ability of the model.
[0134] Taking the example of generating two different spectrograms for each audio segment, the soft loss value and the consistency regularization loss value are represented by KL divergence. At this time, the calculation formula of the target loss value is as follows:
[0135]
[0136]
[0137] Among them, the consistency regularization loss is used to strengthen the consistency between the two second probability distributions, thereby guiding the student model to learn the average value of the predicted second probability distribution, output a smoother probability distribution, thereby suppressing the peak value of the spike effect in the CTC algorithm, reducing the confidence of the student model on the training dataset, and being able to better learn the knowledge in new data, thereby improving the generalization ability of the model.
[0138] In the embodiments of the present application, in the field of speech recognition, for the scenario of distillation training of a teacher model and a student model, each audio segment in the first speech dataset is subjected to multiple spectral enhancement processes. In this way, when the teacher model and the student model perform label prediction based on multiple spectrograms after spectral enhancement, they can fully learn the speech feature representations from different data levels, thereby improving the accuracy of speech recognition. On the other hand, the parameters of the teacher model are generally larger than those of the student model, and the speech recognition effect is better. Therefore, during the distillation training process, in addition to calculating the hard loss value based on the predicted label corresponding to the second probability distribution output by the student model and the corresponding annotated true label, a soft loss value is also calculated according to the first probability distribution and the second probability distribution corresponding to the same spectrogram. Among them, the first probability distribution result output by the teacher model in the soft loss value is used as a soft label to guide the training of the student model, further improving the accuracy of the student model's speech recognition. In addition, by calculating the consistency regularization loss value through the second probability distributions of multiple spectrograms after spectral enhancement, the learning of the context relationship between different spectrograms is increased, making the speech recognition result smoother, thereby reducing the overfitting of the hard loss value to the label and improving the generalization of the model.
[0139] It should be noted that the embodiments of the present application do not impose restrictive constraints on the application scenarios of the above speech recognition model. For example, it can be applied in speech search scenarios, intelligent question-and-answer scenarios, speech control scenarios, speech wake-up scenarios, etc.
[0140] Taking the speech wake-up scenario as an example, speech wake-up is the entry point of speech recognition. Compared with other scenarios, as a speech wake-up model, the student model with small parameters is more likely to be deployed on terminals or embedded devices with relatively weak chip computing power. At this time, the student model needs to have characteristics such as being lightweight, having a small amount of computation, and less inference time to meet the high-quality interaction experience.
[0141] As is well known, models built based on deep neural networks require a large amount of data for training. To achieve a fast and accurate speech wake-up function, a large amount of speech data of different wake-up words or custom wake-up words needs to be fed to the model for learning, increasing the costs of data collection and data annotation.
[0142] In some embodiments, in order to reduce the costs of data collection and data annotation, a pre-training + fine-tuning method is used to train the student model in the speech wake-up scenario.
[0143] See Figure 8 , which is the training method flow of the speech wake-up model provided by the embodiments of the present application, mainly including the following steps:
[0144] S801: Iteratively train the teacher model and the student model according to the labeled second speech dataset to obtain the pre-trained teacher model and student model.
[0145] Among them, the audio in the second speech dataset used for pre-training does not contain a wake word, and a part of the audio in the first speech dataset used for fine-tuning contains a wake word. The data volume of the second speech dataset is larger than that of the first dataset.
[0146] In the embodiments of the present application, since the second speech dataset does not contain a wake word, the dependence of the model on labeled data is reduced, the flexibility of data acquisition methods is increased, and the large quantity of the second speech dataset enables the teacher model and the student model to have strong learning capabilities for the representation of general speech after pre-training.
[0147] In some embodiments, some open-source speech datasets can be used for the second speech dataset. For example, the high-quality wenet dataset is selected. This dataset has been verified by multiple high-precision ASR models and has relatively high text transcription quality, with a total of 10,000 hours of audio samples.
[0148] It should be noted that the embodiments of the present application do not make restrictive requirements on the acquisition method of the second speech dataset. However, to ensure the accuracy of model training, the text transcription quality of the second speech dataset used must be accurate.
[0149] In some embodiments, the first speech dataset can be a dataset customized based on at least one wake word.
[0150] In some embodiments, the labels of the audio in the second speech dataset are full pinyin without tones, and the labels of the audio in the first speech dataset are also full pinyin without tones.
[0151] For example, for an audio of "Hisense Xiaoju", the corresponding label is ['hai', 'xin', 'xiao', 'ju'].
[0152] In the embodiments of the present application, using full pinyin without tones as the label of the training data reduces the influence of various tone changes of the wake word caused by dialects in different regions and improves the robustness of the model.
[0153] In some embodiments, the audio with labels in the first speech dataset is used as positive samples, and the audio without labels is used as negative samples. The ratio of negative samples to positive samples should be maintained at a relatively large value, that is, the number of negative samples is higher than that of positive samples. Otherwise, it will cause relatively serious false awakenings.
[0154] It should be noted that to ensure the accuracy of voice wake-up, the positive and negative samples in the first speech dataset require relatively high text transcription quality.
[0155] S802: Fine-tune the teacher model according to the first speech dataset to obtain a trained teacher model.
[0156] After pre-training, the teacher model, as a speech model with a large number of parameters, has a strong learning ability for the representation of general speech. At this time, by fine-tuning the teacher model with a small amount of the first speech dataset containing the wake word, the teacher model can learn new knowledge, and the transferability is relatively strong.
[0157] S803: According to the first speech dataset, use the trained teacher model to perform iterative distillation training on the pre-trained student model.
[0158] After the student model is pre-trained with the second speech dataset, it has initial network parameters (including weights and biases). Based on the first speech dataset containing the wake word, knowledge distillation is used for fine-tuning to obtain a small-parameter speech wake-up model.
[0159] Among them, the knowledge distillation training process based on the first speech dataset containing the wake word is the same as Figure 2 the knowledge distillation training process shown, and will not be elaborated here.
[0160] In the speech wake-up scenario, the backbone networks of the teacher model and the student model can adopt the spatio-temporal convolutional model (TC-ResNet) for real-time keyword recognition on mobile devices. As Figure 9A shown, it is the network structure diagram of TC-ResNet, which mainly includes a feature extraction layer, a convolutional layer, multiple residual blocks, a fully connected layer, and a Softmax layer. The feature extraction layer is used to extract acoustic features from the speech signal. Before entering the residual block, the extracted acoustic features are first abstracted through a convolutional layer. Among them, the stride of the convolutional layer is 1 without downsampling, the kernel size of the convolutional kernel is (1, 3), and the change of unaligned data in the time dimension is padded on both sides. The fully connected layer is used to map the model output to the dimension of the number of classifications, and the probability corresponding to each class is output by the softmax layer.
[0161] In some embodiments, the structure of each residual block is as Figure 9B shown, using a 9×1 convolutional kernel, without setting biases, and using an additional conv-BN-ReLU structure to match the size. This residual structure enables the model to have a large depth while better avoiding the problem of being unable to jump out of the local optimal solution, effectively preventing overfitting.
[0162] It should be noted that Figure 9A and Figure 9B are only examples, and the embodiments of the present application do not make restrictive requirements on the network structures of the teacher model and the student model.
[0163] In the voice wake-up scenario, considering the cost of data adoption and annotation, a large amount of second voice data sets without wake-up words are used to preselect the teacher model and the student model, enabling the teacher model and the student model to have the learning ability of general speech. Then, a small amount of first voice data sets containing wake-up words are used to fine-tune the teacher model, enabling the teacher model to have the learning ability of wake-up words. In this way, during distillation training, the knowledge about wake-up words learned by the teacher model is provided to the student model, enabling the student model to also have good learning ability for wake-up words, thus achieving lightweight deployment.
[0164] During the entire training process, only a small number of training samples containing one or several keywords need to be provided to the model, reducing the amount of data for keyword collection and annotation and the cost of data engineering. At the same time, full-pinyin without tones is used for annotation, achieving strong transferability of the model, and the voice wake-up function of different keywords or custom keywords can be realized through fine-tuning.
[0165] In the embodiments of the present application, the student model obtained through the pre-training + fine-tuning method not only has a small number of parameters and can meet the requirements of end-side deployment, but also has accurate and reliable recognition results and excellent performance in voice wake-up tasks.
[0166] In some embodiments, when using the trained student model for voice wake-up, for the noise in the voice wake-up data, a simple posterior processing process is: combining the label posterior probabilities generated for each frame into a confidence level for keyword detection.
[0167] See Figure 10 , for the voice wake-up method flow provided by the embodiments of the present application, which mainly includes the following steps:
[0168] S1001: Obtain voice wake-up data.
[0169] S1002: Input the voice wake-up data into the trained student model for label prediction to obtain the fused pinyin features of each character contained in each frame of the voice wake-up data.
[0170] S1003: Determine whether the length of the smoothing window is greater than 1. If so, execute S1004; otherwise, execute S1007.
[0171] When smoothing the posterior probability based on the smoothing window, the length of the smoothing window cannot be too long, otherwise most of the finally output confidence levels will be distributed at a relatively small value. Generally, the length of the smoothing window is set to 1 or 2.
[0172] S1004: Determine the initial posterior probability of the corresponding character according to the fused pinyin features of each character.
[0173] Among them, the student model outputs the posterior probability of each label based on the fused features, and these labels can correspond to the pinyin of keywords.
[0174] S1005: Smooth the initial posterior probability of each character according to the smoothing window to obtain the target posterior probability.
[0175] Generally, the initial posterior probability output by the deep neural network is often noisy. Therefore, the posterior probability smoothing can be completed within the smoothing window with the window length of w smooth to obtain the target posterior probability. The formula is as follows:
[0176]
[0177] Among them, h smooth = max{1, j - w smooth + 1} is the index of the first frame within the smoothing window, p ik is the initial posterior probability, and p i ′ j is the target posterior probability, i represents the index, and j represents the number of frames.
[0178] S1006: Calculate the confidence of the words composed of each character according to the target posterior probability of each character included in each frame within the smoothing window, and use the word with the highest confidence as the wake-up word in the corresponding frame.
[0179] In some embodiments, the confidence of the wake-up word in the j-th frame can be calculated within the sliding window with the window length of w max The formula is as follows:
[0180]
[0181] Among them, h max = max{1, j - w max + 1} is the index of the first frame in the sliding window, p i ′ k is the smoothed target posterior probability, n represents the confidence window length, and the confidence confidence is the value within the confidence window length for a length of time.
[0182] S1007: Divide the respective fused pinyin features of each character by the preset temperature scaling coefficient to obtain the target pinyin features.
[0183] When the length of the smoothing window is 1, it is equivalent to not performing smoothing. At this time, the method of temperature scaling can be selected to smooth the model output, where the value of the temperature scaling coefficient is greater than 1.
[0184] With Figure 9ATaking the network structure shown as an example, for the fused pinyin features output by the fully connected layer of the student model, before passing through the softmax layer, divide by a temperature scaling coefficient to obtain the target pinyin features.
[0185] It should be noted that the temperature scaling coefficient can be selected according to actual needs. When the temperature scaling coefficient is less than 1, it can make larger values larger and smaller values smaller. When the temperature scaling coefficient is greater than 1, it can play a role in smoothing the probability distribution.
[0186] S1008: Determine the target posterior probability of each corresponding character according to the respective target pinyin features.
[0187] S1009: Calculate the confidence of the words composed of each character according to the target posterior probability of each character included in each frame, and use the word with the highest confidence as the wake-up word in the corresponding frame.
[0188] Among them, the calculation method of the confidence is shown in Formula 7 and will not be elaborated here.
[0189] In the example of the present application, when using the student model with small parameters obtained by knowledge distillation to identify the wake-up word in speech, the initial posterior probability of the full pinyin of each character in each frame output by the model is smoothed through a smoothing window with a length greater than 1, thereby reducing the influence of noise. In this way, when calculating the confidence of the wake-up word based on the smoothed target posterior probability, the accuracy of speech wake-up can be improved, and the smoothing method of the smoothing window has a small computational amount, is simple to implement and the output is reliable. At the same time, for the case where the smoothing window with a length of 1 cannot play a smoothing role, before outputting the probability of the full pinyin, the temperature scaling coefficient is used to smooth the fused pinyin features of each character in each frame, thereby reducing the influence of noise and improving the accuracy of calculating the confidence of the wake-up word in the speech data to be recognized, and further improving the accuracy of speech wake-up.
[0190] Based on the same technical concept, the embodiment of the present application provides a training device for a speech recognition model, which can implement the steps of the above speech recognition model training method and achieve the same technical effect.
[0191] See Figure 11 , the device includes a training module 110, which is used to batch input the audio in the pre-annotated first speech dataset into the trained teacher model and the student model to be trained for iterative distillation training until the preset stop condition is met;
[0192] The training module 110 includes a speech enhancement unit 1101, a teacher model prediction unit 1102, a student model prediction unit 1103, a loss value calculation unit 1104, and a parameter adjustment unit 1105, where:
[0193] The voice enhancement unit 1101 performs voice enhancement on a piece of audio multiple times to obtain spectrograms after each enhancement;
[0194] The teacher model prediction unit 1102 is used to perform label prediction on multiple spectrograms of the piece of audio by using the teacher model to obtain the first probability distribution of each spectrogram;
[0195] The student model prediction unit 1103 is used to perform label prediction on multiple spectrograms of the piece of audio by using the student model to obtain the second probability distribution of each spectrogram;
[0196] The loss value calculation unit 1104 is used to calculate the hard loss value according to the predicted label output by the student model and the corresponding annotated true label, and to calculate the soft loss value according to the first probability distribution and the second probability distribution corresponding to the same spectrogram, and to calculate the consistency regularization loss value according to the second probability distribution of the multiple spectrograms; weight the hard loss value, the soft loss value and the consistency regularization loss value to generate the target loss value
[0197] The parameter adjustment unit 1105 is used to adjust the network parameters of the student model according to the target loss value.
[0198] Optionally, the voice enhancement unit 1101 is specifically used for:
[0199] Perform a time warping operation on the piece of audio to obtain the target audio;
[0200] Extract the acoustic features of the target audio and create multiple copies of the acoustic features;
[0201] Perform frequency domain and time domain masking on each copy to generate spectrograms after voice enhancement.
[0202] Optionally, the piece of audio contains multiple frames of speech, and the loss value calculation unit 1104 is specifically used for:
[0203] Perform frame splitting on each spectrogram according to the number of speech frames to obtain sub-spectrograms of each frame of speech;
[0204] Minimize the bidirectional distance between the second probability distributions of two sub-spectrograms corresponding to each frame of speech to generate the consistency regularization loss value.
[0205] Optionally, the formula of the target loss value is expressed as:
[0206]
[0207] Among them, represents the hard loss value, represents the soft loss value, CRa,b∈n (z a ,z b ) represents the consistency regularization loss value, loss represents the target loss value, n is the total number of spectrograms, x i represents the i-th spectrogram, y represents the vocabulary, z a and z b represent the second probability distributions of any two spectrograms among the n spectrograms, represents the first probability of each word in the first probability distribution of the t-th frame in the i-th spectrogram, represents the second probability of each word in the second probability distribution of the t-th frame in the i-th spectrogram, T is the number of frames of the segment of audio, sg() represents applying the stop gradient operation on a probability distribution, α is a hyperparameter controlling consistency regularization, and β is a hyperparameter for the teacher model to guide the student model to learn.
[0208] Optionally, the device further includes a pre-training module 120 and a fine-tuning module 130, where:
[0209] The pre-training module 120 is configured to iteratively train the teacher model and the student model according to the labeled second speech data set, to obtain a pre-trained teacher model and student model; where the labels of the audio in the second speech data set are non-tonal pinyin, the speech in the second speech data set does not contain a wake-up word, the data volume of the second speech data set is greater than that of the first speech data set, the labels of the audio in the first speech data set are non-tonal pinyin, and some of the audio in the first speech data set contains a wake-up word;
[0210] The fine-tuning module 130 is configured to fine-tune the teacher model according to the first speech data set, to obtain a trained teacher model.
[0211] Optionally, some of the audio in the first speech data set contains a wake-up word, and the device further includes a voice wake-up module 140, and the voice wake-up module 140 is configured to:
[0212] Obtain voice wake-up data;
[0213] Input the voice wake-up data into the trained student model for label prediction, to obtain the fused pinyin features of each word included in each frame of the voice wake-up data;
[0214] If the length of the smoothing window is greater than 1, determine the initial posterior probability of the corresponding word according to the fused pinyin features of each word;
[0215] Smooth the initial posterior probabilities of each word according to the smoothing window, to obtain the target posterior probability;
[0216] Calculate the confidence of the words formed by each character according to the target posterior probability of each character included in each frame within the smoothing window, and use the word with the highest confidence as the wake-up word in the corresponding frame.
[0217] Optionally, the voice wake-up module 140 is further configured to:
[0218] If the length of the smoothing window is not greater than 1, the method further includes:
[0219] Divide the respective fused pinyin features of each character by a preset temperature scaling coefficient to obtain target pinyin features; wherein the value of the temperature scaling coefficient is greater than 1;
[0220] Determine the target posterior probability of the corresponding character according to the respective target pinyin features;
[0221] Calculate the confidence of the words formed by each character according to the target posterior probability of each character included in each frame, and use the word with the highest confidence as the wake-up word in the corresponding frame.
[0222] For the convenience of description, the above device is divided into various modules (or units) according to functions and described separately. Of course, when implementing the present application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.
[0223] Those skilled in the art of the technical field to which the present application pertains can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0224] After introducing the training method and device of the speech recognition model according to the exemplary embodiment of the present application, next, an electronic device according to another exemplary embodiment of the present application will be introduced.
[0225] In some embodiments, the electronic device may be a server, including but not limited to a single server, a server cluster, and a cloud server, etc. The electronic device may also be a terminal device, such as a smart TV, a smart phone, a remote control device, a wearable device, a smart speaker, a vehicle-mounted terminal, a tablet, a computer, etc.
[0226] See Figure 12 , the electronic device includes a processor 1201, a memory 1202, and a communication interface 1203. The communication interface 1203, the memory 1202, and the processor 1201 are connected through a bus 1204;
[0227] The communication interface 1203 is used for sending and receiving data;
[0228] The memory 1202 stores a computer program, and the processor 1201 executes a training method for any voice recognition model according to the computer program.
[0229] In some embodiments, the memory 1202 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, programs required to run the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc. The memory 1202 may be a volatile memory, such as a random-access memory (RAM); the memory 1202 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1202 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1202 may be a combination of the above memories.
[0230] The processor 1201 may include one or more central processing units (CPUs), GPUs, or a digital processing unit, etc. The processor 1201 is used to implement the steps of the training method for any voice recognition model when calling the computer program stored in the memory 1202.
[0231] It should be noted that Figure 12 is only an example, presenting the necessary hardware for the electronic device to execute the steps of the training method for any voice recognition model provided in the embodiments of the present application. Those not shown, the electronic device may also include the hardware of conventional intelligent devices such as a pickup, a microphone, a display screen, etc.
[0232] In the embodiments of the present application, the specific connection medium between the communication interface 1203, the memory 1202, and the processor 1201 is not limited. In the embodiments of the present application, the communication interface 1203 is connected to the bus 1204 between the memory 1202 and the processor 1201, and is Figure 12 described in thick lines. The connection manners between other components are only for illustrative purposes and are not to be taken as limiting. The bus 1204 may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of description, Figure 12 only one thick line is used to describe it in, but it does not describe that there is only one bus or one type of bus.
[0233] The embodiments of the present application also provide a computer-readable storage medium for storing some instructions, which, when executed, can complete the steps of any one of the voice recognition model training methods in the foregoing embodiments.
[0234] The embodiments of the present application also provide a computer program product for storing a computer program, and the computer program is used to execute the steps of any one of the voice recognition model training methods in the foregoing embodiments.
[0235] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0236] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0237] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0238] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0239] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.
Claims
1. A method for training a speech recognition model, characterized in that: include: The audio in the pre-annotated first speech data set is input into the trained teacher model and the student model to be trained in batches for iterative distillation training until the preset stop condition is met, wherein each iterative distillation training includes: Perform multiple speech enhancements on a piece of audio to obtain the spectrogram after each enhancement; Using the teacher model to perform label prediction on multiple spectrograms of the audio segment, respectively, to obtain a first probability distribution of each spectrogram; Using the student model to perform label prediction on multiple spectrograms of the audio segment, respectively, to obtain a second probability distribution of each spectrogram; Calculating a hard loss value according to the predicted label output by the student model and the corresponding annotated true label, and calculating a soft loss value according to a first probability distribution and a second probability distribution corresponding to the same spectrogram, and calculating a consistency regularization loss value according to the second probability distribution of the plurality of spectrograms; The hard loss value, the soft loss value, and the consistency regularization loss value are weighted to generate a target loss value, and a network parameter of the student model is adjusted according to the target loss value.
2. The method according to claim 1, characterized in that The method of performing multiple speech enhancements on a segment of audio to obtain a spectrogram after each enhancement includes: Performing a time warp operation on the audio segment to obtain a target audio segment; Extracting acoustic features of the target audio and creating multiple copies of the acoustic features; Each copy is masked in frequency and time domains to generate a speech-enhanced spectrogram.
3. The method according to claim 1, characterized in that The audio segment includes multiple frames of speech, and the calculation of the consistency regularization loss value according to the second probability distribution of the multiple spectrograms includes: Perform frame processing on each spectrogram according to the number of speech frames to obtain a sub-spectrogram of each frame of speech; Minimize the bidirectional distance between the second probability distributions of the two sub-spectrograms corresponding to each frame of speech to generate a consistency regularization loss value.
4. The method according to claim 1, characterized in that The formula of the target loss value is expressed as: in, represents the hard loss value, Indicates the soft loss value, CR a,b∈n (z a ,z b ) represents the consistency regularization loss value, loss represents the target loss value, n is the total number of spectrograms, x i represents the i-th spectrogram, y represents the vocabulary, z a and z b represents the second probability distribution of any two spectrograms among n spectrograms, represents the first probability of each word in the first probability distribution of the tth frame in the i-th spectrogram, represents the second probability of each word in the second probability distribution of the tth frame in the i-th spectrogram, T is the number of frames of the audio segment, sg() represents the application of the stop gradient operation on a probability distribution, α is the hyperparameter controlling the consistency regularization, and β is the hyperparameter for the teacher model to guide the student model to learn.
5. The method according to any one of claims 1 to 4, characterized in that Before iterative distillation training, the method further includes: Iteratively training the teacher model and the student model according to the labeled second speech data set to obtain a pre-trained teacher model and a student model; wherein the label of the audio in the second speech data set is pinyin without tone, the speech in the second speech data set does not contain a wake-up word, the data volume of the second speech data set is greater than the data volume of the first speech data set, the label of the audio in the first speech data set is pinyin without tone, and part of the audio in the first speech data set contains a wake-up word; The teacher model is fine-tuned according to the first speech data set to obtain a trained teacher model.
6. The method according to any one of claims 1 to 4, characterized in that Part of the audio in the first voice data set contains a wake-up word, and after obtaining the trained student model, the method further includes: Get voice wake-up data; Input the speech wake-up data into the trained student model for label prediction, and obtain the fused pinyin features of each word contained in each frame of the speech wake-up data; If the length of the smoothing window is greater than 1, the initial posterior probability of the corresponding word is determined according to the fused phonetic features of each word; Smoothing the initial posterior probability of each word according to the smoothing window to obtain a target posterior probability; The confidence of the words composed of each word is calculated according to the target posterior probability of each word contained in each frame in the smoothing window, and the word with the highest confidence is used as the wake-up word in the corresponding frame.
7. The method according to claim 6, characterized in that If the length of the smoothing window is not greater than 1, the method further includes: Dividing the fused pinyin feature of each character by a preset temperature scaling coefficient to obtain a target pinyin feature; wherein the value of the temperature scaling coefficient is greater than 1; Determine the target posterior probability of the corresponding character according to each target phonetic feature; The confidence of the words composed of each word is calculated according to the target posterior probability of each word contained in each frame, and the word with the highest confidence is used as the wake-up word in the corresponding frame.
8. A training device for a speech recognition model, characterized in that: include: A training module, used for inputting the audio in the pre-annotated first speech data set into the trained teacher model and the student model to be trained in batches for iterative distillation training until a preset stop condition is met; Wherein, the training module includes: The speech enhancement unit performs multiple speech enhancements on a piece of audio to obtain the spectrogram after each enhancement; A teacher model prediction unit, configured to use the teacher model to perform label prediction on the multiple spectrograms of the audio segment respectively, to obtain a first probability distribution of each spectrogram; A student model prediction unit, configured to use the student model to perform label prediction on the multiple spectrograms of the audio segment respectively, to obtain a second probability distribution of each spectrogram; A loss value calculation unit, for calculating a hard loss value according to the predicted label output by the student model and the corresponding annotated true label, and calculating a soft loss value according to a first probability distribution and a second probability distribution corresponding to the same spectrogram, and calculating a consistency regularization loss value according to the second probability distribution of the plurality of spectrograms; weighting the hard loss value, the soft loss value and the consistency regularization loss value to generate a target loss value A parameter adjustment unit is used to adjust the network parameters of the student model according to the target loss value.
9. A voice wake-up device, characterized in that: include: Voice acquisition module, used to obtain voice wake-up data; A feature processing module is used to input the voice wake-up data into the trained student model for label prediction, and obtain the fused pinyin features of each word contained in each frame of the voice wake-up data; A probability prediction module, for determining the initial posterior probability of the corresponding word according to the fused phonetic features of each word if the length of the smoothing window is greater than 1; A probability smoothing module, used for smoothing the initial posterior probability of each word according to the smoothing window to obtain a target posterior probability; The wake-up word determination module is used to calculate the confidence of the words composed of each character according to the target posterior probability of each character contained in each frame in the smoothing window, and use the word with the highest confidence as the wake-up word in the corresponding frame.
10. An electronic device, characterized in that: It includes a processor, a memory and a communication interface, wherein the communication interface, the memory and the processor are connected via a bus; The communication interface is used to send and receive data; The memory stores a computer program, and the processor executes the method according to any one of claims 1 to 7.