Voice content-centered self-supervised comparative representation learning method and system, electronic equipment and readable storage medium
By self-supervised fine-tuning pre-trained language model, using tone and speaker perturbation data, combined with Sinkhorn-Knopp algorithm and comparison loss function, the problem of insufficient semantic consistency of self-supervised pre-trained language model in content tasks is solved, and more efficient decoupling of voice content and speaker information is achieved, improving the performance of the model in content-related tasks.
Patent Information
- Application Number
- CN202510108344.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Self-supervised pre-trained language models have limited performance in similar content tasks requiring higher semantic consistency and are difficult to efficiently decouple voice content from speaker information under limited data.
By self-supervised fine-tuning pre-trained language model, using tone perturbations and speaker perturbations, the representations were extracted and normalized by the Sinkhorn-Knopp algorithm, and the comparison loss function was designed to optimize the semantic consistency of the representation.
Effectively decoupling voice content and speaker information, improving the performance of pre-trained models in content-related recognition tasks, especially in the case of limited data, significantly improving the effectiveness of the model.
Smart Images

Figure CN119943033A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a speech content-centered self-supervised contrast representation learning method, system, electronic device, and readable storage medium, and belongs to the field of speech recognition. Background Art
[0002] Self-supervised speech pre-trained language models, such as HuBERT and WavLM, have advanced the field of speech processing by leveraging large-scale unlabeled data to generate versatile representations. These pre-trained language models provide effective representations and initializations for downstream tasks, greatly facilitating applications such as automatic speech recognition (ASR), phoneme recognition (PR), and speaker identification (SID). Compared with pre-trained language models trained from scratch, self-supervised pre-trained language models significantly reduce training time and computational cost while providing better performance. Despite their wide applicability, these pre-trained language models also face some challenges in content-centric tasks. In particular, the performance of self-supervised pre-trained language models may be limited in content-similar tasks that require higher semantic consistency. This is because the learned representations often confuse speech content with speaker characteristics, thereby weakening the semantic association between pronunciations of the same meaning in the feature space. For example, semantically identical speeches may appear significantly different in the feature space due to variations in pitch, background noise, and speaker identity. This misalignment not only affects the accuracy of downstream tasks, but also limits the generalization ability of the pre-trained language model when the distribution of pre-training data is inconsistent with the target domain data distribution.
[0003] At present, the main research method for decoupling speech content information is to train with a large amount of data, such as the ContentVec pre-trained language model, to achieve the decoupling of speech content and speaker information. In addition, using self-supervised fine-tuning methods such as the Spin method to fine-tune the pre-trained language model can also decouple speaker information. However, existing methods still involve a trade-off between cost and performance, and efficient fine-tuning strategies with limited data remain to be explored. Summary of the invention
[0004] The technical problem solved by the present invention is: the present invention provides a speech content-centered self-supervised contrastive representation learning method, system, electronic device, and readable storage medium. The method of the present invention aims at the problem of mixing speech content and speaker information representation in a self-supervised speech model. The present invention effectively solves the problem of decoupling speech content representation and speaker representation by utilizing self-supervised fine-tuning of a pre-trained language model, thereby improving the performance of the pre-trained model on content-related recognition tasks.
[0005] The technical solution of the present invention is: a speech content-centered self-supervised contrastive representation learning method (Self-Supervised Contrastive Representation Learning, SSCLR), the method comprising:
[0006] Step 1, obtain multi-task speech recognition related datasets; the present invention uses LibriSpeech, QUESST14, Speech Commands, Fluent Commands, SNIPS and VoxCeleb1 datasets. These datasets are open source speech recognition related datasets, which are used to verify the effectiveness of the present invention on various content-related tasks.
[0007] Step 2: Preprocessing of data sets related to multi-task speech recognition: When fine-tuning the pre-trained model in self-supervision, it is necessary to pre-process the dev-clean and dev-other data sets in Librispeech and the corresponding speaker information, so as to generate speaker perturbation speech when fine-tuning the model in self-supervision.
[0008] Step 3: Use the pitch-perturbed and speaker-perturbed speech data to train the WavLM-Base or HuBERT-Base pre-trained language model, and optimize the speech representation by fine-tuning the last two layers of the pre-trained language model.
[0009] Step 4: After extracting the representation of the disturbed speech, the representation matrix is normalized using the Sinkhorn-Knopp algorithm to ensure the consistency of the representation distribution;
[0010] Step 5: Optimize the semantic consistency of representation and improve the content aggregation ability of the pre-trained language model by designing a contrast loss function.
[0011] Furthermore, the Step 3 includes:
[0012] Use a pre-trained language model based on WavLM-Base or HuBERT-Base to perform self-supervised fine-tuning on the Librispeech train-clean-100-hour dataset, and perform self-supervised comparative learning on the generated speaker-perturbed speech and pitch-perturbed speech, so that the pre-trained language model can learn content-related speech features and decouple speaker-related information.
[0013] Furthermore, the Step 3 also includes:
[0014] The pre-trained language model is used to extract the features of the perturbed audio, and then a high-dimensional feature vector is obtained through linear projection and regularization. The feature extraction process is expressed as:
[0015] Z1=Linear(Encoder(z1))
[0016] Z2=Linear(Encoder(z2))
[0017] Among them, the speaker perturbation feature z1 and the pitch perturbation feature z2 are extracted through the pre-trained language model, and then the high-dimensional speaker perturbation feature Z1 and the high-dimensional pitch perturbation feature Z2 are generated respectively through linear projection and regularization. Encoder represents the HuBERT-base or WavLM-base pre-trained language model, Linear represents the linear projection and regularization of the features extracted from the pre-trained language model, the dimension of Z1, Z2 is (batch_size*seq_length, 256), batch_size is the audio training batch size, and seq_length is the audio sequence length.
[0018] Furthermore, in order to obtain a speech representation with more speaker information perturbation and robustness, the fine-tuned pre-trained language model uses a speaker speech perturbation algorithm and a pitch perturbation algorithm, and is a pre-trained language model fine-tuned on the train-clean-100-hour English dataset of Librispeech. This part of the model uses a pre-trained language model based on the HuBERT-base and WavLM-base pre-trained by Librispeech.
[0019] Furthermore, in the Step 4, the Sinkhorn-Knopp algorithm is used, which is an iterative algorithm for solving the matrix scaling problem, and is mainly used to convert a non-negative matrix into a double random matrix by independent scaling of rows and columns. The algorithm normalizes the representation matrix to ensure the consistency of the representation distribution. In the present invention, the process of normalizing the representation matrix by the Sinkhorn-Knopp algorithm is expressed as:
[0020] Z'=Diag(u (k) )Z Diag(v (k) );
[0021] Among them, Z includes Z1 and Z2, Z1 and Z2 represent high-dimensional speaker perturbation features and pitch perturbation features respectively, Z' includes Z1' and Z2', Z1' is the high-dimensional speaker perturbation feature after algorithm normalization, Z2' is the high-dimensional pitch perturbation feature after algorithm normalization, Diag represents the operation of the diagonal matrix, u (k) and v (k) are hyperparameters, representing the scaling factors of the rows and columns of the kth iteration, initialized to u (0) =1m , v (0) =1 n ,u (k) and v (k) The iteration formula is as follows:
[0022]
[0023] Among them, Q (k) represents the matrix at the kth iteration, m, n represent the number of rows and columns of the speaker perturbation feature Z1;
[0024] At the 0th iteration, Q needs to be initialized, and the initialization of Q is expressed as:
[0025]
[0026] Among them, ∈ is a hyperparameter used to control the smoothness of the matrix, and Z' is the final result after normalization of the algorithm, that is, Z1', Z2'.
[0027] Furthermore, the Step 5 includes:
[0028] Using the self-supervised contrast loss method, we cross-compare the speaker perturbation features and the pitch perturbation features, so that the pre-trained language model can learn content-related semantic information; the total self-supervised contrast loss function is expressed as:
[0029]
[0030] in, represents the contrast loss between Z1' and Z2, represents the contrast loss between Z2' and Z1;
[0031] The formula is:
[0032]
[0033] Where τ is the temperature parameter used to control the smoothness, N is the first dimension of Z1', and Z i ' and Z j ' is the speaker perturbation audio feature and pitch perturbation feature after normalization by the Sinkhorn algorithm, Z i ' is the normalized version of Z1, Z j ' is the normalized version of Z2, Z i is the i-th row of the eigenvector Z';
[0034] Similarly, the comparison loss between Z2' and Z1 is The formula is:
[0035]
[0036] Among them, Z j is the jth row of the eigenvector Z'.
[0037] Furthermore, the Step 5 also includes:
[0038] After self-supervised fine-tuning of the upstream task, the obtained model is tested in the S3PRL (Self-Supervised Speech Pre-Training Library) framework of SUPERB (Speech Processing Universal Evaluation Benchmark); SUPERB is a multi-task benchmark for evaluating speech processing models, covering various speech understanding tasks such as automatic speech recognition (ASR), phoneme recognition, keyword spotting, speaker recognition, sentiment analysis, etc. S3PRL is a self-supervised learning method used in SUPERB, and it can be used as a framework for training and evaluating speech models; downstream tasks include: Automatic Speech Recognition (ASR) and Phoneme Recognition (PR) tasks, Query-by-Example Spoken Term Detection (QBE), Keyword Spotting (KS), Intent Classification (IC), Slot Filling (SF) and Speaker Identification (SID) tasks, all of which are trained according to the default method in the S3PRL framework.
[0039] The present invention also provides a speech content-centric self-supervised contrastive representation learning system, the system comprising: a module for executing the above-mentioned speech content-centric self-supervised contrastive representation learning method.
[0040] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned speech content-centered self-supervised contrast representation learning method when executing the program.
[0041] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the above-mentioned self-supervised contrastive representation learning method centered on speech content is implemented.
[0042] The beneficial effects of the present invention are:
[0043] 1. The present invention proposes a self-supervised contrastive representation learning method centered on speech content. The self-supervised contrastive representation learning method is used to fine-tune the WavLM-Base or HuBERT-Base pre-trained language model. The method can separate speaker information from speech content, improve the semantic consistency of content-centric tasks, decouple speaker-related information of speech, and enable the pre-trained language model to learn content-related high-dimensional representations;
[0044] 2. In the downstream test tasks of the S3PRL framework, the experimental results show that the method achieves competitive performance in content-related experiments, achieves low word error rate and other indicators in tasks such as ASR, and also shows good word error rate and phoneme error rate in cross-data tests;
[0045] 3. This method effectively improves the effectiveness of the pre-trained language model in content-related tasks and improves the performance of the model in content-related recognition tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a general design block diagram of a speech content-centered self-supervised contrastive representation learning method proposed by the present invention;
[0047] Figure 2 is the phoneme feature graph of the present invention;
[0048] Figure 3 This is the speaker recognition accuracy analysis based on the HuBERT-base pre-trained language model in the present invention. DETAILED DESCRIPTION
[0049] Example 1: Figure 1-Figure 3 As shown, a self-supervised contrastive representation learning method centered on speech content is presented. This method aims to solve the problem of the mixing of speech content and speaker information representation in self-supervised speech models. By fine-tuning the pre-trained models WavLM-Base and HuBERT-Base, the last two layers are optimized to improve the semantic consistency of speech representation. In the data processing stage, the present invention selects the train-clean-100 subset in the LibriSpeech dataset as training data, generates pitch-perturbed speech through SpeedPerturbation and PitchShift functions, and uses a speaker forgery algorithm to generate speaker-perturbed speech to enhance the robustness of the model to different disturbance conditions. In model training, this method extracts representations from the perturbed speech, and introduces the Sinkhorn-Knopp algorithm to normalize the representation matrix to ensure the balance and consistency of data distribution. At the same time, a self-supervised contrastive loss function is designed to optimize the semantic aggregation of content representation;
[0050] The method comprises:
[0051] Step 1, obtain multi-task speech recognition related datasets; the present invention uses LibriSpeech, QUESST14, Speech Commands, Fluent Commands, SNIPS and VoxCeleb1 datasets. These datasets are all open source speech recognition related datasets, which are used for automatic speech recognition (ASR) and phoneme recognition (PR), query-by-example speech word detection (QBE), keyword recognition (KS), intent classification (IC) and slot filling (SF) tasks respectively. In order to prove the effectiveness of the present invention in content decoupling, the VoxCeleb1 dataset is also used to verify the performance of the present invention in the speaker identification (SID) task.
[0052] Step 2: Preprocessing of datasets related to multi-task speech recognition: Preprocessing the dev-clean and dev-other datasets in Librispeech and the corresponding speaker information, so as to generate speaker perturbation speech when fine-tuning the model in self-supervision.
[0053] In order to generate pitch perturbation audio, the present invention uses the SpeedPerturbation and PitchShift functions of the torchaudio.transforms library in Torchaudio to implement pitch perturbation. For speaker perturbation data, the speaker perturbation algorithm in contentVec is used to perturb the audio. In addition, the present invention uses the speaker information text in the contentvec model method, as well as the phoneme files corresponding to the dev-clean and dev-other data, to calculate the discretization representation quality.
[0054] Step 3: Use the pitch-perturbed and speaker-perturbed speech data to train the WavLM-Base or HuBERT-Base pre-trained language model, and optimize the speech representation by fine-tuning the last two layers of the pre-trained language model.
[0055] The Step 3 includes:
[0056] Use a pre-trained language model based on WavLM-Base or HuBERT-Base to perform self-supervised fine-tuning on the Librispeech train-clean-100-hour dataset, and perform self-supervised comparative learning on the generated speaker-perturbed speech and pitch-perturbed speech, so that the pre-trained language model can learn content-related speech features and decouple speaker-related information.
[0057] The Step 3 also includes:
[0058] The pre-trained language model is used to extract the features of the perturbed audio, and then a high-dimensional feature vector is obtained through linear projection and regularization. The feature extraction process is expressed as:
[0059] Z1=Linear(Encoder(z1))
[0060] Z2=Linear(Encoder(z2))
[0061] Among them, the speaker perturbation feature z1 and the pitch perturbation feature z2 are extracted through the pre-trained language model, and then the high-dimensional speaker perturbation feature Z1 and the high-dimensional pitch perturbation feature Z2 are generated respectively through linear projection and regularization. Encoder represents the HuBERT-base or WavLM-base pre-trained language model, Linear represents the linear projection and regularization of the features extracted from the pre-trained language model, the dimension of Z1, Z2 is (batch_size*seq_length, 256), batch_size is the audio training batch size, and seq_length is the audio sequence length.
[0062] Furthermore, in order to obtain a speech representation with more speaker information perturbation and robustness, the fine-tuned pre-trained language model uses a speaker speech perturbation algorithm and a pitch perturbation algorithm, and is a pre-trained language model fine-tuned on the train-clean-100-hour English dataset of Librispeech. This part of the model uses a pre-trained language model based on the HuBERT-base and WavLM-base pre-trained by Librispeech.
[0063] Step 4: After extracting the representation of the disturbed speech, the representation matrix is normalized using the Sinkhorn-Knopp algorithm to ensure the consistency of the representation distribution;
[0064] Furthermore, in the Step 4, the Sinkhorn-Knopp algorithm is used. The algorithm is an iterative algorithm for solving the matrix scaling problem. It is mainly used to convert a non-negative matrix into a double random matrix by independent scaling of rows and columns. The algorithm normalizes the representation matrix to ensure the consistency of the representation distribution. Sinkhorn-Knopp prevents the gradient explosion phenomenon that occurs during the training process. In the present invention, the process of normalizing the representation matrix by the Sinkhorn-Knopp algorithm is expressed as:
[0065] Z'=Diag(u (k) )ZDiag(v (k) );
[0066] Among them, Z includes Z1 and Z2, Z1 and Z2 represent high-dimensional speaker perturbation features and pitch perturbation features respectively, Z' includes Z1' and Z2', Z1' is the high-dimensional speaker perturbation feature after algorithm normalization, Z2' is the high-dimensional pitch perturbation feature after algorithm normalization, Diag represents the operation of the diagonal matrix, u (k) and v (k) are hyperparameters, representing the scaling factors of the rows and columns of the kth iteration, initialized to u (0) =1 m , v (0) =1 n , where the number of iterations k = 3, u (k) and v (k) The iteration formula is as follows:
[0067]
[0068] Among them, Q (k) represents the matrix at the kth iteration, m, n represent the number of rows and columns of the speaker perturbation feature Z1;
[0069] At the 0th iteration, Q needs to be initialized, and the initialization of Q is expressed as:
[0070]
[0071] Among them, ∈ is a hyperparameter used to control the smoothness of the matrix. ∈ is set to 0.02, and Z' is the final result after normalization of the algorithm, that is, Z1', Z2'.
[0072] Step 5: Optimize the semantic consistency of representation and improve the content aggregation ability of the pre-trained language model by designing a contrast loss function.
[0073] Furthermore, the Step 5 includes:
[0074] Using the self-supervised contrast loss method, we cross-compare the speaker perturbation features and the pitch perturbation features, so that the pre-trained language model can learn content-related semantic information; the total self-supervised contrast loss function is expressed as:
[0075]
[0076] in, represents the contrast loss between Z1' and Z2, represents the contrast loss between Z2' and Z1;
[0077] The formula is:
[0078]
[0079] Where τ is a temperature parameter used to control the smoothness, τ is set to 0.1, N is the first dimension of Z1', Z i ' and Z j ' is the speaker perturbation audio feature and pitch perturbation feature after normalization by the Sinkhorn algorithm, Z i ' is the normalized version of Z1, Z j ' is the normalized version of Z2, Z i is the i-th row of the eigenvector Z';
[0080] Similarly, the comparison loss between Z2' and Z1 is The formula is:
[0081]
[0082] Among them, Z j is the jth row of the eigenvector Z'.
[0083] Furthermore, the Step 5 also includes:
[0084] After self-supervised fine-tuning of the upstream task, the obtained model is tested in the S3PRL (Self-Supervised Speech Pre-Training Library) framework of SUPERB (Speech Processing Universal Evaluation Benchmark); SUPERB is a multi-task benchmark for evaluating speech processing models, covering various speech understanding tasks such as automatic speech recognition (ASR), phoneme recognition, keyword spotting, speaker recognition, sentiment analysis, etc. S3PRL is a self-supervised learning method used in SUPERB, and it can be used as a framework for training and evaluating speech models; downstream tasks include: Automatic Speech Recognition (ASR) and Phoneme Recognition (PR) tasks, Query-by-Example Spoken Term Detection (QBE), Keyword Spotting (KS), Intent Classification (IC), Slot Filling (SF) and Speaker Identification (SID) tasks, all of which are trained according to the default method in the S3PRL framework.
[0085] The present invention also provides a speech content-centered self-supervised contrast representation learning system, the system comprising:
[0086] A data set acquisition module is used to acquire data sets related to multi-task speech recognition;
[0087] Dataset preprocessing module, used for preprocessing of data sets related to multi-task speech recognition;
[0088] The fine-tuning module is used to train the WavLM-Base or HuBERT-Base pre-trained language model using pitch-perturbed and speaker-perturbed speech data, and optimize the speech representation by fine-tuning the last two layers of the pre-trained language model;
[0089] A normalization module is used to extract the representation of the disturbed speech and then normalize the representation matrix using the Sinkhorn-Knopp algorithm to ensure the consistency of the representation distribution;
[0090] The optimization module is used to optimize the semantic consistency of representation and improve the content aggregation ability of the pre-trained language model by designing a contrast loss function.
[0091] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned speech content-centered self-supervised contrast representation learning method when executing the program.
[0092] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the above-mentioned self-supervised contrastive representation learning method centered on speech content is implemented.
[0093] The present invention designs comparative experiments between the present invention and current mainstream methods such as Spin, ContentVec, SCORE, and LASER on content-related tasks. Performance indicators of automatic speech recognition (ASR), phoneme recognition (PR), query-by-example speech word detection (QbE), keyword spotting (KS), intent classification (IC), and slot filling (SF) based on the SUPERB benchmark. The indicators include accuracy (Acc%), phoneme error rate (PER%), word error rate (WER%), maximum term weighted value (MTWV%), F1 score, and concept error rate (CER%). PT means pre-training, SSFT means self-supervised fine-tuning, and the results are shown in Table 1.
[0094] Table 1 shows the comparative experiments of current mainstream methods on content-related tasks.
[0095]
[0096]
[0097] From the analysis of Table 1, it can be seen that the method of the present invention performs well in the three indicators of intent classification (IC) and slot filling (SF) tasks on the WavLM and HuBERT pre-trained language models, and has obvious advantages over other methods. In the automatic speech recognition (ASR) task, the performance of the method of the present invention on the WavLM pre-trained language model is second only to ContentVec500, reaching a word error rate of 5.78%, and a word error rate of 6.16% on the HuBERT pre-trained language model, which is better than other methods. In the query-by-case speech word detection (QbE) task, the method of the present invention surpassed all other pre-trained language models on the WavLM pre-trained language model, reaching 9.67%. In the keyword detection (KS) and phoneme recognition (PR) tasks, the method of the present invention also achieved relatively ideal results compared with other methods.
[0098] In order to verify and evaluate the generalization ability of SSCRL on datasets with different domain characteristics, the present invention uses the pre-trained language model fine-tuned with LibriSpeech data to directly conduct experiments on automatic speech recognition (ASR) and phoneme recognition (PR) tasks on the TIMIT dataset. The experimental results are shown in Table 2:
[0099] Table 2 shows the automatic speech recognition and phoneme recognition performance of different methods on the TIMIT dataset
[0100]
[0101]
[0102] According to Table 2, the proposed method is second only to ContentVec in the ASR task of HuBERT and WavLM pre-trained language models. 500 Pre-trained language models, but both surpassed other methods, reaching word error rates of 27.38% and 24.66% respectively. In the PR task, the performance of this invention on the HuBERT pre-trained language model is second only to HuBERT+SPIN 256 On WavLM, the effect of the present invention is second only to the WavLM+SCORE method.
[0103] In order to verify the quality of the discretized units of the pre-trained language model, the present invention evaluates the performance of different methods on three indicators: clustering purity (ClsPur), phoneme purity (PhnPur), and phoneme normalized mutual information (PNMI). The experimental results are shown in Table 3:
[0104] Table 3 shows the performance of the discrete element quality indicators
[0105]
[0106] From the results in Table 3, it can be seen that the PNMI index of the present invention on the WavLM and HuBERT pre-trained language models is higher than that of other pre-trained language models, reaching high scores of 0.674 and 0.685 respectively. In terms of the Phn Pur index, the present invention method is also better than other pre-trained language models on the WavLM pre-trained language model, reaching a score of 0.659, and is second only to HuBERT+Spin on the HuBERT pre-trained language model. 2048 The pre-trained language model has been achieved. It can be clearly seen that the method of the present invention is effective in the discrete unit method.
[0107] In order to prove the effectiveness of the method in the feature space, the present invention randomly selects several audios and corresponding phoneme lists in the transcribed text from the original WavLM-Base pre-trained language model and the WavLM-Base pre-trained language model fine-tuned by the method of the present invention in the test-clean data set, and then uses the pre-trained language model to extract audio features, obtain the features of the last layer, and then use the PCA dimensionality reduction method to map it to the two-dimensional space. The results are shown in Figure 2. Figure 2 shown.
[0108] from Figure 2 It can be seen from the present invention that Figure 2 (a) is the feature extracted by the WavLM pre-trained language model after fine-tuning by the method of the present invention, Figure 2 (b) is the feature extracted by the WavLM pre-trained language model without fine-tuning. It can be clearly seen that after fine-tuning by the method of the present invention, the phonemes are more clustered in the feature space, and the phonemes of the same type are closer in the feature space, while the WavLM pre-trained language model without fine-tuning by the method of the present invention has the mapped phonemes scattered in the feature space and the distribution is more chaotic.
[0109] In order to verify the effectiveness of the present invention in speaker decoupling, the present invention selects the speaker recognition performance of the last six layers of the HuBERT pre-trained language model for analysis. Figure 3 shown.
[0110] from Figure 3 It can be clearly seen that after fine-tuning the SSCLR method of the present invention, the recognition accuracy of the last two layers of the HuBERT pre-trained language model is reduced to 11%, which is slightly lower than ContentVec 500Compared with the Spin method, the method of the present invention achieves similar results. This result shows that the SSCLR method can effectively reduce the interference of speaker information in speech features, proving the effectiveness of the method of the present invention in decoupling the content features of speech, and the data of the present invention only uses less than one-third of Spin.
[0111] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A speech content-centric self-supervised contrastive representation learning method, characterized by: The method comprises: Step 1. Obtain data sets related to multi-task speech recognition; Step 2: Preprocessing of data sets related to multi-task speech recognition; Step 3: Use the pitch-perturbed and speaker-perturbed speech data to train the WavLM-Base or HuBERT-Base pre-trained language model, and optimize the speech representation by fine-tuning the last two layers of the pre-trained language model. Step 4: After extracting the representation of the disturbed speech, the representation matrix is normalized using the Sinkhorn-Knopp algorithm to ensure the consistency of the representation distribution; Step 5: Optimize the semantic consistency of representation and improve the content aggregation ability of the pre-trained language model by designing a contrast loss function.
2. The method for self-supervised contrastive representation learning based on speech content according to claim 1, characterized in that: The Step 3 includes: Use a pre-trained language model based on WavLM-Base or HuBERT-Base to perform self-supervised fine-tuning on the Librispeech train-clean-100-hour dataset, and perform self-supervised comparative learning on the generated speaker-perturbed speech and pitch-perturbed speech, so that the pre-trained language model can learn content-related speech features and decouple speaker-related information.
3. The method for self-supervised contrastive representation learning based on speech content according to claim 1, characterized in that: The Step 3 also includes: The pre-trained language model is used to extract the features of the perturbed audio, and then a high-dimensional feature vector is obtained through linear projection and regularization. The feature extraction process is expressed as: Z1=Linear(Encoder(z1)) Z2=Linear(Encoder(z2)) Among them, the speaker perturbation feature z1 and the pitch perturbation feature z2 are extracted through the pre-trained language model, and then linear projection and regularization are performed to generate high-dimensional speaker perturbation feature Z1 and high-dimensional pitch perturbation feature Z2 respectively. Encoder represents the HuBERT-base or WavLM-base pre-trained language model, Linear represents linear projection and regularization of the features extracted from the pre-trained language model, the dimension of Z1, Z2 is (batch_size*seq_length, 256), batch_size is the audio training batch size, and seq_length is the audio sequence length.
4. The method for self-supervised contrastive representation learning based on speech content according to claim 1, characterized in that: The fine-tuned pre-trained language model is a pre-trained language model that uses a speaker voice perturbation algorithm and a pitch perturbation algorithm and is fine-tuned on a train-clean-100-hour English dataset of Librispeech.
5. The method for self-supervised contrastive representation learning based on speech content according to claim 1, characterized in that: In the Step 4, the process of normalizing the representation matrix by the Sinkhorn-Knopp algorithm is expressed as: Z' = D iag (u (k) )D iag (v (k) ); Among them, Z includes Z1 and Z2, Z1 and Z2 represent high-dimensional speaker perturbation features and pitch perturbation features respectively, Z' includes Z1' and Z2', Z1' is the high-dimensional speaker perturbation feature after algorithm normalization, Z2' is the high-dimensional pitch perturbation feature after algorithm normalization, Diag represents the operation of the diagonal matrix, u (k) and v (k) are hyperparameters, representing the scaling factors of the rows and columns of the kth iteration, initialized to u (0) =1 m , v (0) =1 n ,u (k) and v (k) The iteration formula is as follows: Among them, Q (k) represents the matrix at the kth iteration, m, n represent the number of rows and columns of the speaker perturbation feature Z1; At the 0th iteration, Q needs to be initialized, and the initialization of Q is expressed as: Among them, ∈ is a hyperparameter used to control the smoothness of the matrix, and Z' is the final result after normalization of the algorithm, that is, Z1', Z2'.
6. The method for self-supervised contrastive representation learning based on speech content according to claim 1, characterized in that: The Step 5 includes: Using the self-supervised contrast loss method, we cross-compare the speaker perturbation features and the pitch perturbation features, so that the pre-trained language model can learn content-related semantic information; the total self-supervised contrast loss function is expressed as: in, represents the contrast loss between Z1' and Z2, represents the contrast loss between Z2' and Z1; The formula is: Where τ is the temperature parameter used to control the smoothness, N is the first dimension of Z1', and Z i ' and Z j ' is the speaker perturbation audio feature and pitch perturbation feature after normalization by the Sinkhorn algorithm, Z i ' is the normalized version of Z1, Z j ' is the normalized version of Z2, Z i is the i-th row of the eigenvector Z'; Similarly, the comparison loss between Z2' and Z1 is The formula is: Among them, Z j is the jth row of the eigenvector Z'.
7. The method for self-supervised contrastive representation learning based on speech content according to claim 1, characterized in that: The Step 5 also includes: After self-supervised fine-tuning of the upstream tasks, the obtained models are tested in S3PRL of SUPERB. The downstream tasks include: automatic speech recognition and phoneme recognition tasks, case-by-case query voice word detection, keyword recognition, intent classification, slot filling and speaker recognition tasks. These tasks are trained according to the default method in the S3PRL framework.
8. A speech content-centric self-supervised contrastive representation learning system, characterized in that: The system comprises: a module for executing a speech content-centric self-supervised contrastive representation learning method as described in any one of claims 1 to 7.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements a speech content-centric self-supervised contrastive representation learning method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements a speech content-centric self-supervised contrastive representation learning method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Speech recognition method and model training method
CN115954001A
Improved pre-training method, electronic equipment and storage medium
CN116758904A
Decoupling type voice self-supervision pre-training method
CN118841029A