Self-Supervised Training Method, System and Storage Medium for Anti-Noise Speech Recognition Model
By performing self-supervised training on the HuBERT model and calculating losses layer by layer, the noise resistance of the anti-noise speech recognition model is improved, the problem of limited anti-noise ability in the existing technology is solved, and the accuracy of automatic speech recognition is improved.
Patent Information
- Application Number
- CN202211710809.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-12-29
AI Technical Summary
In the prior art, the self-supervised training model has limited anti-noise capability, which affects the effect of speech recognition.
By pre-trained HuBERT model of the original speech input, a speech embedding is generated, and the noise-added speech input anti-noise speech recognition model is calculated layer by layer, and self-supervised training is performed until the masked noise speech embedding approaches the speech embedding of the pre-trained model.
The anti-noise pre-training method of the speech recognition model is improved on the HuBERT architecture, the anti-noise ability of the self-supervised training model is improved, and the accuracy of automatic speech recognition is improved.
Smart Images

Figure CN116013271B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent voice, and in particular to a self-supervised training method, system and storage medium for a noise-resistant speech recognition model. Background Art
[0002] In order to further improve the user's speech interaction experience, self-supervised learning is used to improve the performance of ASR (Automatic Speech Recognition). For example, self-supervised learning is performed by using a large amount of unlabeled speech to learn context-aware speech representations that are beneficial to ASR (or other downstream tasks). In the framework of self-supervised training, a reconstruction module from noisy speech to original speech is added, and an objective function for speech reconstruction is added to the self-supervised training, so as to improve the noise resistance of the self-supervised speech embedding, and further improve the speech recognition performance.
[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the related art:
[0004] The existing technology usually integrates a reconstruction module (for example, SE (speech enhancement module)) as a preprocessing front end for automatic speech recognition to suppress noise from noisy speech. However, due to the less interaction between the reconstruction module and the self-supervised training framework, the structure of the model is not optimized, and the noise resistance to background noise is limited, which affects the speech recognition effect. Summary of the Invention
[0005] In order to at least solve the problem of limited noise resistance of the self-supervised training model in the prior art.
[0006] In a first aspect, an embodiment of the present invention provides a self-supervised training method for a noise-resistant speech recognition model, including:
[0007] Input the original speech into a pre-trained HuBERT model, determine L speech embeddings of the original speech in the 1st to Lth layers of the encoder of the pre-trained HuBERT model, input the L speech embeddings into the aggregator of the pre-trained HuBERT model, and determine the aggregated representation generated by the aggregator as the training target of the Lth layer of the encoder of the noise-resistant speech recognition model;
[0008] Input the noisy speech generated by adding noise to the original speech into the noise-resistant speech recognition model, and determine L masked noise speech embeddings of the noisy speech in the 1st to Lth layers of the encoder through the encoder of the noise-resistant speech recognition model;
[0009] Determine the first loss of the speech embeddings of the pre-trained HuBERT model in the first to the (L-1)-th layers of the encoder corresponding to the masked noise speech embeddings of the noise-resistant speech recognition model in the first to the (L-1)-th layers of the encoder, and determine the second loss of the masked noise speech embeddings of the L-th layer of the encoder of the noise-resistant speech recognition model corresponding to the training target;
[0010] Perform self-supervised training on the noise-resistant speech recognition model based on the comprehensive loss determined by the first loss and the second loss until the masked noise speech embeddings determined by the noise-resistant speech recognition model approach the speech embeddings determined by the pre-trained HuBERT model.
[0011] In a second aspect, an embodiment of the present invention provides a self-supervised training system for a noise-resistant speech recognition model, including:
[0012] A training target determination program module, configured to input the original speech into the pre-trained HuBERT model, determine L speech embeddings of the original speech in the first to the L-th layers of the encoder through the encoder of the pre-trained HuBERT model, input the L speech embeddings into the aggregator of the pre-trained HuBERT model, and determine the aggregated representation generated by the aggregator as the training target of the L-th layer of the encoder of the noise-resistant speech recognition model;
[0013] A speech embedding determination program module, configured to input the noisy speech obtained by adding noise to the original speech into the noise-resistant speech recognition model, and determine L masked noise speech embeddings of the noisy speech in the first to the L-th layers of the encoder through the encoder of the noise-resistant speech recognition model;
[0014] A loss determination program module, configured to determine the first loss of the speech embeddings of the pre-trained HuBERT model in the first to the (L-1)-th layers of the encoder corresponding to the masked noise speech embeddings of the noise-resistant speech recognition model in the first to the (L-1)-th layers of the encoder layer by layer, and determine the second loss of the masked noise speech embeddings of the L-th layer of the encoder of the noise-resistant speech recognition model corresponding to the training target;
[0015] A self-supervised training program module, configured to perform self-supervised training on the noise-resistant speech recognition model based on the comprehensive loss determined by the first loss and the second loss until the masked noise speech embeddings determined by the noise-resistant speech recognition model approach the speech embeddings determined by the pre-trained HuBERT model.
[0016] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the self-supervised training method of the noise-resistant speech recognition model according to any embodiment of the present invention.
[0017] In a fourth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the self-supervised training method of the noise-resistant speech recognition model according to any embodiment of the present invention are implemented.
[0018] The beneficial effects of the embodiments of the present invention are as follows: A noise-resistant pre-training method for speech recognition is implemented on the HuBERT architecture, which improves the noise resistance of the self-supervised training model and further improves the accuracy of automatic speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 is a flowchart of a self-supervised training method of a noise-resistant speech recognition model provided by an embodiment of the present invention;
[0021] Figure 2 is a schematic diagram of the combination of HuBERT-NIT and HuBERT in a self-supervised training method of a noise-resistant speech recognition model provided by an embodiment of the present invention;
[0022] Figure 3 is a schematic diagram of the combination of HuBERT-AGG and HuBERT in a self-supervised training method of a noise-resistant speech recognition model provided by an embodiment of the present invention;
[0023] Figure 4 is a schematic diagram of the comparison of the word error rates of different pre-trained models on the LIBRISPEECH original and artificial test sets in a self-supervised training method of a noise-resistant speech recognition model provided by an embodiment of the present invention;
[0024] Figure 5 is a schematic diagram of the comparison of the word error rates of HuBERTAGG on the artificial noise test set in a self-supervised training method of a noise-resistant speech recognition model provided by an embodiment of the present invention;
[0025] Figure 6 It is a schematic diagram comparing word error rates of different systems on a CHiME-4 real test set according to a self-supervised training method for a noise-resistant speech recognition model provided by an embodiment of the present invention;
[0026] Figure 7 It is a structural schematic diagram of a self-supervised training system for a noise-resistant speech recognition model provided by one embodiment of the present invention;
[0027] Figure 8 A schematic diagram of the structure of an embodiment of an electronic device for self-supervised training of a noise-resistant speech recognition model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] like Figure 1 FIG. 1 is a flow chart of a self-supervised training method for a noise-resistant speech recognition model provided by an embodiment of the present invention, comprising the following steps:
[0030] S11: inputting the original speech into the pre-trained HuBERT model, determining L speech embeddings of the original speech in the 1st to Lth layers of the encoder through the encoder of the pre-trained HuBERT model, inputting the L speech embeddings into the aggregator of the pre-trained HuBERT model, and determining the aggregate representation generated by the aggregator as the training target of the Lth layer of the encoder of the noise-resistant speech recognition model;
[0031] S12: inputting the noisy speech generated by adding noise to the original speech into the anti-noise speech recognition model, and determining, through the encoder of the anti-noise speech recognition model, the L masked noise speech embeddings of the noisy speech in the 1st layer to the Lth layer of the encoder;
[0032] S13: Determine layer by layer the first loss of the speech embedding of the pre-trained HuBERT model at the 1st to the L-1th layers of the encoder corresponding to the masked noise speech embedding of the 1st to the L-1th layers of the encoder of the noise-resistant speech recognition model, and determine the second loss of the masked noise speech embedding of the Lth layer of the encoder of the noise-resistant speech recognition model corresponding to the training target;
[0033] S14: Self-supervised training of the noise-robust speech recognition model is performed based on the comprehensive loss determined from the first loss and the second loss until the masked noise speech embedding determined by the noise-robust speech recognition model approaches the speech embedding determined by the pre-trained HuBERT model.
[0034] In this embodiment, the noise-robust speech recognition model trained by this method is based on the HuBERT (Hidden-unit Bidirectional Encoder Representations from Transformers) model, where HuBERT is a self-supervised learning (SSL) method. For example, some speech features in the training speech are masked, and HuBERT only applies the prediction loss on the masked areas, forcing the model to learn the combined acoustic and language models on the continuous input. The backbone of HuBERT is a convolutional waveform encoder and a BERT mask predictor, which consists of many identical Transformer blocks. Given a sequence of speech embeddings X = [x1,..., x T , the mask predictor of HuBERT takes its masked part as input and predicts the distribution of the target encoding at each time step t:
[0035]
[0036] where C is the total number of encodings, e c is the embedding of encoding C, represents the output feature sequence at step t. The cosine similarity is calculated using sim(·,·), and τ is the degree of logistic regression.
[0037] The discrete target sequence of X is represented as Z = [z1, z2,..., z T , where z t ∈ [C] is a categorical variable of C, and the prediction loss formula for the masked area in HuBERT is:
[0038]
[0039] where, represents the masked time steps of X. It should be noted that this prediction loss is also used in the self-supervised training of the noise-robust speech recognition model of this method. HuBERT iteratively refines the assignment of aggregating hidden layer states at each level. That is, in the first iteration, the hidden layer states are assigned by clustering the MFCC features of the training data; in subsequent iterations, new hidden layer states are assigned by clustering the hidden layer representations generated by the model trained in the previous iteration.
[0040] For step S11, the pre-training of the noise-free dataset is merged with the original noisy speech into HuBERT, called HuBERT-NIT (noise-invariant training into HuBERT, bidirectional encoded representation of hidden units for noise-invariant training). The pre-trained HuBERT NIT serves as the target for the self-supervised learning of the anti-noise speech recognition model of this method, which can be understood as playing the role of the "teacher" in a similar "teacher-student" structure. The structure of HuBERT-NIT is as Figure 2 shown. The encoded original speech is represented as X, and the masked encoding of the noisy speech obtained by adding noise to the original speech representation is A regularization term is adopted. The L2 (Euclidean distance) is used to determine the cosine distance between the encoding φ l (X) of the original speech used by the pre-trained vanilla HuBERT and the encoding generated by HuBERT-NIT using the masked noisy speech encoding at the l-th layer . The regularization term is applied in a hierarchical manner:
[0041]
[0042]
[0043] where L is the number of encoder layers. When training HuBERT NIT , the parameters of the pre-trained vanilla HuBERT are frozen. The output distribution NIT of HuBERT is also supervised by the vanilla HuBERT O preprocessed by KLD (Kullback–Leibler Divergence, also known as relative entropy):
[0044]
[0045] The final loss of the pre-trained HuBERT NIT is the weighted sum of L d-NIT and L kld :
[0046] L NIT = λ1L d-NIT + λ z L KLD
[0047] The anti-noise speech recognition model of this method compared with the above pre-trained HuBERT NITBy extracting aggregated representations specifically optimized for the ASR task, the learned noise-invariant representations are improved. The noise-robust speech recognition model of this method can be called HuBERT-AGG (Hidden-unit Bidirectional Encoder Representations from Transformers - AGGREGATED, aggregated hidden-unit bidirectional encoder representations). The HuBERT-AGG training process of this method is based on the above-mentioned pre-trained HuBERT NIT and uses a labeled dataset (clean speech, noisy speech generated from clean speech) to train with the CTC (Connectionist Temporal Classification, neural network-based temporal classification) criterion for the labeled dataset. In this way, the trained aggregator can extract aggregated representations beneficial to the ASR task, and the overall structure is as Figure 3 shown.
[0048] The original speech is input into the above-mentioned pre-trained HuBERT-NIT model. The original speech representation X determined by the convolutional feature extraction layer is used, and the speech encodings of each layer of the encoder are determined φ l(X) (speech embedding). The speech encodings of each layer are input into the aggregator. The speech encodings φ l (X) (1 ≤ l ≤ L) of each layer of the pre-trained HuBERT-NIT have corresponding representations H = [h1, h2,..., h L in the aggregator. The aggregator is used to determine the weighted sum of the speech embeddings from the first layer to the L-th layer of the encoder, which is represented as a vector a = [a1, a2,..., a L in the aggregator, where a l is the weight corresponding to h l in the weighted sum. The aggregated representation can be expressed as Modify L d-NIT , and use the aggregated representation generated by the aggregator as the training target for the L-th layer (i.e., the last layer of the encoder) of the encoder of the noise-robust speech recognition model HuBERT AGG of this method.
[0049] For step S12, the noisy speech generated by adding noise to the original speech is input into the HuBERT AGG model of this method. Similarly, the noisy speech representation determined by the convolutional feature extraction layer is used, and the masked noise speech embeddings of each layer of the encoder are determined (1 ≤ l ≤ L).
[0050] For step S13, it should be noted that the training objective of the L-th layer of the encoder of the anti-noise speech recognition model of this method is determined separately by the aggregator. For the remaining losses, they can be determined layer by layer according to the speech embedding and the masked noise speech embedding. Specifically, the regularization term and the cosine regularization term 5 determined layer by layer from the speech embedding of the first layer to the (L-1)-th layer and the masked noise speech embedding of the first layer to the (L-1)-th layer are used as the first loss, and the corresponding loss can be determined through the regularization term and the cosine regularization term.
[0051] For step S14, based on the loss determined in step S13, the loss L of the anti-noise speech recognition model of this method d-AGG is determined as:
[0052]
[0053] The determined L can be used d-AGG to train the anti-noise speech recognition model until the masked noise speech embedding determined by the anti-noise speech recognition model approaches the speech embedding determined by the pre-trained HuBERT model. The trained anti-noise speech recognition model can determine a masked noise speech embedding with stronger anti-noise ability, and the final distribution is obtained through the linear layer & SoftMax of the model.
[0054] As an implementation, after determining the second loss between the masked noise speech embedding of the L-th layer of the encoder of the anti-noise speech recognition model and the training objective, the method further includes:
[0055] Using the mask predictor of the HuBERT model to determine the feature prediction sequence of the noisy speech;
[0056] Determining the feature target sequence of the original speech through the convolutional waveform encoder of the HuBERT model;
[0057] Determining the third loss based on the feature prediction sequence and the feature target sequence;
[0058] Performing self-supervised training on the anti-noise speech recognition model based on the comprehensive loss determined by the first loss, the second loss, and the third loss.
[0059] In this implementation, this method uses the prediction loss L in the masked area of HuBERT m . Since the last layer of HuBERT-AGG of this method is from the aggregated representation and comes from the aggregator, rather than being supervised by the encoding of the last layer of the pre-trained HuBERT, the final loss in training HuBERT-AGG of this method is the weighted sum of L d-AGG and L m :
[0060] L AGG= λ1L d-AGG + λ2L m
[0061] Use this loss to further improve the training effect of the HuBERT-AGG denoising speech recognition model of this method.
[0062] It can be seen from this implementation that a denoising pre-training method for speech recognition is implemented on the HuBERT architecture, which improves the denoising ability of the self-supervised training model and further improves the accuracy of automatic speech recognition.
[0063] Specific experiments were conducted on the denoising speech recognition model trained by self-supervision of this method, and the effectiveness of this method was verified using simulated and real-world noise data. Data preparation involves single-channel tracks of 4 well-known corpora in ASR: LIBRILIGHT, LibriSpeech, MUSAN, and CHiME-4. The training process consists of 3 stages, and the data usage in each stage is introduced: (1) The labeled data used to train the representation aggregator is denoted as DAGG (only for HuBERT AGG). (2) The unlabeled speech data used to preprocess HuBERT-NIT and HuBERT-AGG is denoted as DP. DP is synthesized by mixing 960 hours of LIBRISPEECH with noise randomly sampled from the music (MUSIC) and noise (NOISE) categories of MUSAN and SNR uniformly sampled between 5 and 10 dB. Among them, the MUSAN noise set has three types of noise: "music", "noise", and "human voice". Take the two categories of "noise" and "music" and mix them with the original speech. The speech category in MUSAN is not used during training to avoid confusion with actual speech when learning speech representations. The model is fine-tuned on different DFs to test simulated noise data and real noise data.
[0064] For the test of simulated noise data, 100 clean speeches from the original LIBRISPEECH training are used for DF. Prepare the original clean speech and test other tests. By mixing the original test set with different categories in MUSSAN, this method synthesizes several simulated noise test sets. The SNR of all simulated test sets is uniformly sampled from 5 to 10 dB. For the test of real noise data, all CHiME-4 data is used for direction finding. The CHiME-4 single-channel real development and evaluation sets are used for testing.
[0065] The aggregator is trained by fine-tuning HuBERT BASE, which is pre-trained on the complete 960-hour data in LIBRISPEECH. DAGG considers 5 different partitions: Libri light (LL) 10 minutes, 1 hour, 10 hours, LibriSpeech (LS) 100 hours, and CHiME-4 clean speech. For CHiME-4 clean speech, the same settings as the LS 100-hour split are used because they are of comparable size. During fine-tuning, all encoder layers are frozen to ensure that all layers can generate speech representations for aggregation.
[0066] This method uses the FAIRSEQ toolkit for model pre-training. The same architecture is adopted: 12 transformer blocks, 768 hidden states, and 8 attention heads. To converge faster, all models are initialized with the same HuBERT used in aggregator training. K-means clustering with 500 clusters is further applied to the latent features extracted from the HuBERT BASE model because it has the highest normalized phoneme purity. The same configuration is used to train the second iteration of HuBERT BASE. This method disables the hierarchical dropout and masking of the pre-trained vanilla HuBERT to produce raw speech representations of better quality for knowledge distillation.
[0067] Since the size of the CHiME-4 training set (92.28 hours) is comparable to the LS 100-hour split, the basic 100-hour settings in wav2vec 2.0 are followed in the experiments.
[0068] For decoding on the CHiME-4 test set, an LSTM-based word-level language model with a vocabulary size of 65000 is trained using Espresso on the text part of the WSJ corpus, and then the same decoding strategy is executed. The hyperparameters of decoding are adjusted for the LIBRISPEECH and CHiME-4 validation sets respectively.
[0069] Figure 4Shows the Word Error Rate (WER, %) results for the original and simulated test sets based on LIBRISPEECH. The music + noise test sets were synthesized by mixing the original test sets with music and noise categories from MUSSAN, and these test sets are in a matching state with the pre-training phase. The results of the speech test sets show a mismatch in the test conditions. In the first row, HuBERT preprocessed on the original LIBRISPEECH data shows a severe degradation on the noisy test sets. Additionally, the degradation of the speech test sets is significantly more severe than that of the music + noise tests. This is attributed to the fact that background speech is more likely to be confused with the target speech than music and noise. In the second row, if HuBERT is preprocessed using augmented data DP. At the cost of reducing the performance of the original test set, the performance on the noisy test sets is improved. In the third row, by simply penalizing the distance between the representations produced by the pre-trained vanilla HuBERT and HuBERT NIT, a consistent improvement can be seen for all test sets. The relative WER reduction (13.1% / 14.7%) under the speech mismatch test conditions is slightly less than that under the music + noise matching test conditions (i.e., music + noise: 14.8% / 17.0%). In the last row, the aggregator is trained on the Librilight 10-hour partition, and the aggregated speech representations are more beneficial for the ASR task and result in further improvements on all the noisy test sets.
[0070] Figure 5 Shows different selections of DAGG, considering 5 different partitions. The LS and LL partitions are in the same domain as the test sets, and the CHiME-4 partition is in a different domain. The amount of data used to train the aggregator has no significant impact on the final ASR performance. Training the aggregator with data from different domains (last row) results in a slight decrease in performance for most test sets. It can be argued that the benefit of introducing the aggregator lies in a better correlation between the aggregated speech representations and the downstream ASR task.
[0071] Evaluations were conducted on the channel tracks of the CHiME-4 real-world noise test set. This included fully supervised learning with and without an enhanced front-end, as well as other SSL methods. These results used the same data as in the experiments of this method (i.e., the unlabeled LIBRISPEECH 960-hour data and the labeled CHiME-4 1-channel data). In the experiments, the aggregator of HuBERT AGG was trained using the CHiME-4 clean data in the simulated partition. By training on simulated noise data synthesized from LIBRISPEECH and MUSSAN, the results on the CHiME-4 test set exceeded those obtained by HuBERT pre-trained on the original data, with the WER reduced by 18.2% / 23.2%. These results indicate that HuBERT preprocessed with simulated noise data can be well generalized to real-world noise data. As Figure 6 shown, increasing the training steps from 25k to 150k results in continuous improvement. With 150k training steps, the HuBERT-AGG of this method achieved significantly better results than all fully supervised models and other SSL models with the same data usage rate.
[0072] Overall, this method introduced an aggregator in the pre-trained HuBERT, which calculates weighted sum representations optimized for the ASR task and uses them to teach the last layer of the HuBERT-AGG encoder. By doing so, the HuBERT-AGG model learned noise-invariant representations that are also suitable for the ASR task. Experiments on the simulated noise LIBRISPEECH test set showed that compared with direct data augmentation SSL methods, the relative WER was reduced by 13.1%-17.0%, and similar performance was still maintained on the original test set. On the CHiME-4 real-world noise test set, the HuBERT AGG of this method exceeded the best results obtained by previous fully supervised learning methods and other SSL methods.
[0073] As Figure 7 shown is a schematic structural diagram of a self-supervised training system for a noise-resistant speech recognition model provided by an embodiment of the present invention. This system can execute the self-supervised training method of the noise-resistant speech recognition model described in any of the above embodiments and is configured in a terminal.
[0074] A self-supervised training system 10 for a noise-resistant speech recognition model provided in this embodiment includes: a training objective determination program module 11, a speech embedding determination program module 12, a loss determination program module 13, and a self-supervised training program module 14.
[0075] Among them, the training objective determination program module 11 is used to input the original speech into the pre-trained HuBERT model, determine L speech embeddings of the original speech in the 1st to Lth layers of the encoder through the encoder of the pre-trained HuBERT model, input the L speech embeddings into the aggregator of the pre-trained HuBERT model, and determine the aggregated representation generated by the aggregator as the training objective of the Lth layer of the encoder of the noise-resistant speech recognition model; the speech embedding determination program module 12 is used to input the noisy speech generated by adding noise to the original speech into the noise-resistant speech recognition model, and determine L masked noise speech embeddings of the noisy speech in the 1st to Lth layers of the encoder through the encoder of the noise-resistant speech recognition model; the loss determination program module 13 is used to layer by layer determine the first loss of the speech embeddings of the pre-trained HuBERT model in the 1st to L-1th layers corresponding to the masked noise speech embeddings of the noise-resistant speech recognition model in the 1st to L-1th layers, and determine the second loss of the masked noise speech embeddings of the Lth layer of the encoder of the noise-resistant speech recognition model corresponding to the training objective; the self-supervised training program module 14 is used to perform self-supervised training on the noise-resistant speech recognition model based on the comprehensive loss determined by the first loss and the second loss until the masked noise speech embeddings determined by the noise-resistant speech recognition model approach the speech embeddings determined by the pre-trained HuBERT model.
[0076] Further, the loss determination program module is further used for:
[0077] Using the mask predictor of the HuBERT model to determine the feature prediction sequence of the noisy speech;
[0078] Determining the feature target sequence of the original speech through the convolutional waveform encoder of the HuBERT model;
[0079] Determining the third loss based on the feature prediction sequence and the feature target sequence;
[0080] The self-supervised training program module is further used for:
[0081] Performing self-supervised training on the noise-resistant speech recognition model based on the comprehensive loss determined by the first loss, the second loss, and the third loss.
[0082] An embodiment of the present invention further provides a non-volatile computer storage medium, and the computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the self-supervised training method of the noise-resistant speech recognition model in any of the above method embodiments;
[0083] As an implementation manner, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as:
[0084] Input the original speech into the pre-trained HuBERT model, determine L speech embeddings of the original speech at the 1st to Lth layers of the encoder of the pre-trained HuBERT model, input the L speech embeddings into the aggregator of the pre-trained HuBERT model, and determine the aggregated representation generated by the aggregator as the training target for the Lth layer of the encoder of the noise-resistant speech recognition model;
[0085] Input the noisy speech generated by adding noise to the original speech into the noise-resistant speech recognition model, and determine L masked noise speech embeddings of the noisy speech at the 1st to Lth layers of the encoder through the encoder of the noise-resistant speech recognition model;
[0086] Determine the first loss of the speech embeddings of the pre-trained HuBERT model at the 1st to L-1th layers of the encoder corresponding to the masked noise speech embeddings of the noise-resistant speech recognition model at the 1st to L-1th layers of the encoder layer by layer, and determine the second loss of the masked noise speech embeddings of the Lth layer of the encoder of the noise-resistant speech recognition model corresponding to the training target;
[0087] Perform self-supervised training on the noise-resistant speech recognition model based on the comprehensive loss determined by the first loss and the second loss until the masked noise speech embeddings determined by the noise-resistant speech recognition model approach the speech embeddings determined by the pre-trained HuBERT model.
[0088] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium and, when executed by a processor, execute the self-supervised training method of the noise-resistant speech recognition model in any of the above method embodiments.
[0089] Figure 8 It is a schematic diagram of the hardware structure of an electronic device for the self-supervised training method of a noise-resistant speech recognition model provided in another embodiment of the present application, as Figure 8 shown. The device includes:
[0090] One or more processors 810 and a memory 820, Figure 8 Taking one processor 810 as an example. The device for the self-supervised training method of the noise-resistant speech recognition model may further include: an input device 830 and an output device 840.
[0091] The processor 810, the memory 820, the input device 830, and the output device 840 can be connected through a bus or other means, Figure 8 Taking connection through a bus as an example.
[0092] The memory 820, being a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the self-supervised training method of the noise-resistant speech recognition model in the embodiments of the present application. The processor 810 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 820, that is, implements the self-supervised training method of the noise-resistant speech recognition model in the above method embodiments.
[0093] The memory 820 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data, etc. In addition, the memory 820 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 820 may optionally include a memory remotely disposed relative to the processor 810, and these remote memories can be connected to the mobile device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0094] The input device 830 can receive input digital or character information. The output device 840 may include a display device such as a display screen.
[0095] The one or more modules are stored in the memory 820 and, when executed by the one or more processors 810, execute the self-supervised training method of the noise-resistant speech recognition model in any of the above method embodiments.
[0096] The above product can execute the method provided in the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.
[0097] A non-volatile computer-readable storage medium may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0098] An embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the self-supervised training method of the anti-noise speech recognition model according to any embodiment of the present invention.
[0099] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to:
[0100] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.
[0101] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc., such as tablet computers.
[0102] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.
[0103] (4) Other electronic devices with data processing functions.
[0104] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also include other elements not explicitly listed, or also include elements inherent in such a process, method, article, or device. Without more limitations, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article, or device including the said elements.
[0105] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0106] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A self-supervised training method for a noise-resistant speech recognition model, comprising: Input the original speech into the pre-trained HuBERT model, determine L speech embeddings of the original speech at the 1st to Lth layers of the encoder through the encoder of the pre-trained HuBERT model, input the L speech embeddings into the aggregator of the pre-trained HuBERT model, and determine the aggregated representation generated by the aggregator as the training target for the Lth layer of the encoder of the noise-robust speech recognition model; Input the noisy speech generated by adding noise to the original speech into the noise-robust speech recognition model, and determine L masked noise speech embeddings of the noisy speech at the 1st to Lth layers of the encoder through the encoder of the noise-robust speech recognition model; Determine the first loss of the speech embeddings of the pre-trained HuBERT model at the 1st to L-1th layers corresponding to the masked noise speech embeddings of the noise-robust speech recognition model at the 1st to L-1th layers layer by layer, and determine the second loss of the masked noise speech embeddings of the Lth layer of the encoder of the noise-robust speech recognition model corresponding to the training target; Perform self-supervised training on the noise-robust speech recognition model based on the comprehensive loss determined by the first loss and the second loss until the masked noise speech embeddings determined by the noise-robust speech recognition model approach the speech embeddings determined by the pre-trained HuBERT model.
2. The method according to claim 1, wherein, After determining the second loss of the masked noise speech embeddings of the Lth layer of the encoder of the noise-robust speech recognition model and the training target, the method further includes: Use the mask predictor of the HuBERT model to determine the feature prediction sequence of the noisy speech; Determine the feature target sequence of the original speech through the convolutional waveform encoder of the HuBERT model; Determine the third loss based on the feature prediction sequence and the feature target sequence; Perform self-supervised training on the noise-robust speech recognition model based on the comprehensive loss determined by the first loss, the second loss, and the third loss.
3. The method according to claim 1, wherein, The step of determining the first loss of the speech embeddings of the pre-trained HuBERT model at the 1st to L-1th layers corresponding to the masked noise speech embeddings of the noise-robust speech recognition model at the 1st to L-1th layers layer by layer includes: Use the regularization term and cosine regularization term determined layer by layer between the speech embeddings of the 1st to L-1th layers and the masked noise speech embeddings of the 1st to L-1th layers as the first loss.
4. The method according to claim 1, wherein, The aggregator of the pre-trained HuBERT model is used to determine the weighted sum of the speech embeddings at the 1st to Lth layers of the encoder.
5. A self-supervised training system for a noise-resistant speech recognition model, comprising: A training target determination program module, configured to input the original speech into the pre-trained HuBERT model, determine L speech embeddings of the original speech at the 1st to Lth layers of the encoder through the encoder of the pre-trained HuBERT model, input the L speech embeddings into the aggregator of the pre-trained HuBERT model, and determine the aggregated representation generated by the aggregator as the training target for the Lth layer of the encoder of the noise-robust speech recognition model; A voice embedding determination program module for inputting the noisy speech generated by adding noise to the original speech into the noise-resistant speech recognition model, and determining L masked noise speech embeddings of the noisy speech in the 1st to Lth layers of the encoder of the noise-resistant speech recognition model through the encoder of the noise-resistant speech recognition model; A loss determination program module for layer-by-layer determining a first loss of the speech embeddings of the pre-trained HuBERT model in the 1st to L-1th layers corresponding to the masked noise speech embeddings of the noise-resistant speech recognition model in the 1st to L-1th layers, and determining a second loss of the masked noise speech embedding in the Lth layer of the encoder of the noise-resistant speech recognition model corresponding to the training target; A self-supervised training program module for performing self-supervised training on the noise-resistant speech recognition model based on the comprehensive loss determined by the first loss and the second loss until the masked noise speech embedding determined by the noise-resistant speech recognition model approaches the speech embedding determined by the pre-trained HuBERT model.
6. The system according to claim 5, wherein, The loss determination program module is further configured to: Use the mask predictor of the HuBERT model to determine the feature prediction sequence of the noisy speech; Determine the feature target sequence of the original speech through the convolutional waveform encoder of the HuBERT model; Determine a third loss based on the feature prediction sequence and the feature target sequence; The self-supervised training program module is further configured to: Perform self-supervised training on the noise-resistant speech recognition model based on the comprehensive loss determined by the first loss, the second loss, and the third loss.
7. The system according to claim 5, wherein, The loss determination program module is configured to: Take the regularization term and the cosine regularization term determined layer by layer between the speech embeddings in the 1st to L-1th layers and the masked noise speech embeddings in the 1st to L-1th layers as the first loss.
8. The system according to claim 5, wherein, The aggregator of the pre-trained HuBERT model is used to determine the weighted sum of the speech embeddings in the 1st to Lth layers of the encoder.
9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method according to any one of claims 1-4.
10. A storage medium, on which a computer program is stored, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-4.
Citation Information
Patent Citations
End-to-end online voice detection and recognition method and system, and equipment
CN112951213A
Training method and device of dialect type prediction model and storage medium
CN114743545A