An Uncertainty-Based Integrated Self-Supervised Speaker Recognition Method
By adopting uncertainty integrated self-supervision method in speaker recognition, combining multiple self-supervision models and decision-making fusion technologies, the problems of high data annotation cost and insufficient model stability are solved, and high performance and high credibility speaker recognition are achieved.
Patent Information
- Application Number
- CN202310476907.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-04-28
AI Technical Summary
The prior art relies on a large amount of labeled voice data in speaker recognition, resulting in high data labeling costs, insufficient training data, and insufficient self-supervised models in terms of stability and credibility.
An integrated self-supervised speaker recognition method based on uncertainty is proposed. By masking the combination of self-supervised models, comparing the self-supervised models and self-regressive predictive self-supervised models, using Dirichlet distribution and Dempster rules to make decision fusion, improving the stability and credibility of the model.
This method can improve the performance and stability of speaker recognition, enhance the credibility of the algorithm, and improve the utilization rate of voice data when a small amount of labeled voice data.
Smart Images

Figure CN116386646B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of speaker recognition, and in particular to an integrated self-supervisory speaker recognition method based on uncertainty, and belongs to the application of machine learning in the field of speaker recognition. Background Art
[0002] Speaker recognition is a biometric recognition technology that identifies the speaker's identity based on the speaker's personality information in the voice signal. Existing mainstream technologies often use supervised speaker recognition methods, but they require a large amount of labeled voice data to train the model. In practical applications, the cost of manually labeling massive amounts of speaker voice data is high, which often results in insufficient training data.
[0003] The learning mechanism of transferring learning from global general features to local specific tasks can greatly reduce the model's dependence on data annotation and help improve the utilization of speech data. The self-supervised learning method can directly learn the underlying structure representation of the data from large-scale unlabeled data, thereby helping to improve the performance and convergence speed of downstream tasks. Since most existing self-supervised models tend to focus on the accuracy of learning results, in the actual training process, the stability and credibility of the algorithm are also crucial.
[0004] Therefore, how to design a reliable and stable self-supervised speaker recognition method for a small amount of labeled speech data is a problem of great value at present. Summary of the invention
[0005] In order to solve the problems of massive unlabeled data, stability and credibility in speaker recognition, the present invention proposes an integrated self-supervised speaker recognition method based on uncertainty, which greatly improves the performance of speaker recognition and improves the stability and credibility of the recognition method.
[0006] In order to achieve the above object, the present invention is achieved through the following technical solutions:
[0007] The present invention is an uncertainty-based integrated self-supervised speaker recognition method, comprising the following steps:
[0008] Step 1: Collect the speaker's voice data and annotate the voice data;
[0009] Step 2: Preprocess the original speech data, extract the Mel-level spectrogram, obtain the feature vector of the speech data, and construct the speaker recognition data set. Specifically include:
[0010] Step 2-1: Convert the preprocessed speech data into a spectrogram, and pass it through a Mel filter bank to obtain the Mel spectrogram features of the speech;
[0011] Step 2-2: For a Mel-spectrogram set X containing N speech samples = {x1, x2, ..., x N}, according to the label information, form a sample pair (x n ,z n ), where x n represents the Mel spectrogram of the nth speech sample (n=1,2,...,N), z n is the label of the nth sample.
[0012] Step 3: Input the Mel spectrogram features obtained in step 2 into several pre-trained self-supervised models in sequence. The self-supervised models of the present invention are masked self-supervised models, contrastive self-supervised models, and autoregressive prediction self-supervised models, and extract the output of the last layer of each model respectively. Specifically, it includes:
[0013] Step 3-1: Initialize the network parameters of the masked self-supervised model, contrastive self-supervised model, and autoregressive prediction self-supervised model, using the model parameters pre-trained on a large unlabeled public speech dataset as initialization parameters. The masked self-supervised and contrastive self-supervised models are both based on a multi-layer Transformer model structure, while the autoregressive prediction self-supervised model structure is a 3-layer LSTM network structure.
[0014] Step 3-2: Use different loss functions to calculate the loss value between the model prediction results and the query samples. The InfoNCE loss is used for the comparison self-supervised model, as shown below:
[0015]
[0016] In order to predict the future value x of the speech sequence t+k , and the speech potential representation w after model prediction t Together we construct the probability density function f k (x t+k ,w t ), used to retain x t+k and w t The mutual information between them, N represents X={x1,x2,...,x N}, one of which is from the distribution p(x t+k |w t ), and the rest are positive samples from distribution p(x t+k ) negative samples.
[0017] The masked self-supervised model and the autoregressive prediction self-supervised model use the L1 loss function as shown below:
[0018]
[0019] Among them, x t (t=1,2,...,n) is the input sequence, y t (t=1,2,...,n) is the output sequence.
[0020] Step 4: For downstream classification tasks, the output of the last layer of each self-supervised model in step 3 is used as the input of the fully connected layer, and the output of the fully connected layer is calculated through the ReLU activation function to obtain the evidence of the input speech data under each model. Specifically, it includes:
[0021] Step 4-1: The self-supervised model includes an L-layer neural network, and the outputs of the L-layer network are [D1, D2, ..., D L ], extract the last layer of network D L For K classification problems, the fully connected layer converts D L The output of is mapped to K dimensions.
[0022] Step 4-2: The output of the fully connected layer is calculated as evidence through the ReLU activation function. The ReLU activation function is as follows:
[0023] f(x)=max(x,0)
[0024] Step 5: Calculate the Dirichlet distribution parameters under the subjective logic framework, and then calculate the confidence quality and uncertainty of each self-supervisory model output. Specifically include:
[0025] Step 5-1: For the K-classification problem, subjective logic assigns a confidence quality to each class label based on the evidence and an uncertainty to the entire framework. For the self-supervised model q, the K+1 quality values are all non-negative and sum to 1:
[0026]
[0027] Among them, u q is the uncertainty, is the confidence quality of the kth class. In the self-supervised model, under the K classification task, the model output is used as evidence, and subjective logic uses the evidence and the parameters of the Dirichlet distribution Contact us. Evidence q That is, the decision result after the output of the self-supervised model is calculated by the ReLU activation function, and the parameters of the Dirichlet distribution are can be Export, that is
[0028] Step 5-2: Calculate confidence mass and uncertainty u q , specifically expressed as follows:
[0029]
[0030] in, is the Dirichlet intensity. The above formula actually describes the phenomenon that the more evidence of the kth class is observed, the greater the probability that the sample is classified as the kth class. Correspondingly, the less total evidence is observed, the greater the uncertainty. Confidence allocation can be regarded as a subjective opinion, and the corresponding Dirichlet distribution p cpc The formula for calculating the class probability mean is:
[0031] Step 6: Use the Dempster rule to fuse the output decision results of the three self-supervised models to obtain the final probability and overall uncertainty of each class and output the final classification result. Specifically include:
[0032] Step 6-1: Dempster-Shafer evidence theory allows combining evidence from different sources to obtain the overall uncertainty of the model. Here, we need to combine the quality sets of three different self-supervised models. Among them, it includes the confidence quality of each self-supervised model and uncertainty u q Combining the evidence of the contrastive self-supervised model, the masked self-supervised model, and the autoregressive prediction self-supervised model requires combining M cpc 、M mpc and M apc :
[0033]
[0034] The specific calculation rules are:
[0035]
[0036]
[0037] in, is a measure of the amount of conflict between the three mass sets, is the normalization factor. After the above calculation, the final classification result is obtained.
[0038] The present invention also provides an uncertainty-based integrated self-supervised speaker recognition system, the system comprising:
[0039] 1) Speech feature extraction module: pre-process the speech data, extract the Mel spectrum features of the speech, and annotate the data set;
[0040] 2) Self-supervised model training module: A large amount of unlabeled data is used to pre-train the masked self-supervised model, the contrastive self-supervised model, and the autoregressive prediction self-supervised model. The Mel-spectrogram feature results of the speech data are input into the three self-supervised models respectively, and the output of the last layer of the model is extracted;
[0041] 3) Classification decision module: The last layer output of each self-supervised model is used as the input of the fully connected layer, and the output of the fully connected layer is calculated through the ReLU activation function to obtain the evidence of the input speech data under each model;
[0042] 4) Uncertainty estimation module: Calculate the confidence quality and uncertainty of each self-supervised model output through the obtained evidence and Dirichlet distribution parameters;
[0043] 5) Decision fusion module: Use the Dempster rule to fuse the output decisions of the three self-supervised models to obtain the final probability and overall uncertainty of each class, and output the final classification result.
[0044] The beneficial effects of the present invention are:
[0045] The present invention utilizes a self-supervised model to learn useful feature information related to the current speaker recognition classification task when there is little labeled training data or no labeled training data, thereby greatly improving the performance of speaker recognition. At the same time, uncertainty estimation is used to integrate the output decision results of multiple self-supervised models, thereby improving the stability and credibility of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic diagram of the process of the present invention.
[0047] Figure 2 It is a schematic diagram of the uncertainty-based decision fusion framework of the present invention taking the contrastive self-supervision and masked self-supervision models as examples. DETAILED DESCRIPTION
[0048] The following will disclose the embodiments of the present invention with drawings. For the purpose of clear description, many practical details will be described together in the following description. However, it should be understood that these practical details should not be used to limit the present invention. That is, in some embodiments of the present invention, these practical details are not necessary. In addition, for the purpose of simplifying the drawings, some conventional structures and components will be depicted in a simple schematic manner in the drawings.
[0049] like Figure 1 As shown, the present invention is an uncertainty-based integrated self-supervision speaker recognition method, which specifically includes the following steps:
[0050] Step 1: Collect the speaker's voice data and annotate the voice data;
[0051] Step 2: Preprocess the original speech data, extract the Mel-level spectrogram, obtain the feature vector of the speech data, and construct the speaker recognition data set. Specifically include:
[0052] Step 2-1: Convert the preprocessed speech data into a spectrogram, and pass it through a Mel filter bank to obtain the Mel spectrogram features of the speech;
[0053] Step 2-2: For a Mel-spectrogram set X containing N speech samples = {x1, x2, ..., x N}, according to the label information, form a sample pair (x n ,z n ), where x n represents the Mel spectrogram of the nth speech sample (n=1,2,...,N), z n is the label of the nth sample.
[0054] Step 3: Input the Mel spectrogram features obtained in step 2 into several pre-trained self-supervised models in sequence. The present invention takes the masked self-supervised model, the contrast self-supervised model and the autoregressive prediction self-supervised model as examples, and extracts the output of the last layer of each model respectively. Specifically, it includes:
[0055] Step 3-1: Initialize the network parameters of the masked self-supervised model, the contrastive self-supervised model, and the autoregressive prediction self-supervised model: The self-supervised model was pre-trained on the train-clean-100 subset of the LibriSpeech corpus. For pre-training, the model was trained on 4 GPUs with a total batch size of 256. The present invention uses the Adam optimizer to change the learning rate, where the learning rate rises to a peak value of 4e-4 in the first 7% of the total training steps and then decays linearly. Among them, the masked self-supervised and contrastive self-supervised models are based on a multi-layer Transformer model structure, and the autoregressive prediction self-supervised model structure is a 3-layer LSTM network structure.
[0056] Step 3-2: Use different loss functions to calculate the loss value between the model prediction results and the query samples. The InfoNCE loss is used for the comparison self-supervised model, as shown below:
[0057]
[0058] In order to predict the future value x of the speech sequence t+k , and the speech potential representation w after model prediction t Together we construct the probability density function f k (x t+k ,wt ), used to retain x t+k and w t The mutual information between them, N represents X={x1,x2,...,x N}, one of which is from the distribution p(x t+k |w t ), and the rest are positive samples from distribution p(x t+k ) negative samples.
[0059] The masked self-supervised model and the autoregressive prediction self-supervised model use the L1 loss function as shown below:
[0060]
[0061] Among them, x t (t=1,2,...,n) is the input sequence, y t (t=1,2,...,n) is the output sequence.
[0062] Step 4: For downstream classification tasks, the output of the last layer of each self-supervised model in step 3 is used as the input of the fully connected layer, and the output of the fully connected layer is calculated through the ReLU activation function to obtain the evidence of the input speech data under each model. Specifically, it includes:
[0063] Step 4-1: The self-supervised model includes an L-layer neural network, and the outputs of the L-layer network are [D1, D2, ..., D L ], extract the last layer of network D L For K classification problems, the fully connected layer converts D L The output of is mapped to K dimensions.
[0064] Step 4-2: The output of the fully connected layer is calculated as evidence through the ReLU activation function. The ReLU activation function is as follows:
[0065] f(x)=max(x,0)
[0066] Step 5: Calculate the Dirichlet distribution parameters under the subjective logic framework, and then calculate the confidence quality and uncertainty of each self-supervisory model output. Specifically include:
[0067] Step 5-1: For K-classification problems, subjective logic assigns a confidence quality to each class label based on the evidence and an uncertainty to the entire framework. For example, for the self-supervised model q, the K+1 quality values are all non-negative and sum to 1:
[0068]
[0069] Among them, u qis the uncertainty, is the confidence quality of the kth class. In the self-supervised model, under the K classification task, its model output is used as evidence, and subjective logic uses the evidence and the parameters of the Dirichlet distribution Contact us. Evidence q That is, the decision result after the output of the self-supervised model is calculated by the ReLU activation function, and the parameters of the Dirichlet distribution are can be Export, that is
[0070] Step 5-2: Calculate confidence mass and uncertainty u q , specifically expressed as follows:
[0071]
[0072] in, is the Dirichlet intensity. The above formula actually describes the phenomenon that the more evidence of the kth class is observed, the greater the probability that the sample is classified as the kth class. Correspondingly, the less total evidence is observed, the greater the uncertainty. Confidence allocation can be regarded as a subjective opinion, and the corresponding Dirichlet distribution p cpc The formula for calculating the class probability mean is:
[0073] Step 6: Use the Dempster rule to fuse the output decision results of the three self-supervised models to obtain the final probability and overall uncertainty of each class and output the final classification result. Specifically include:
[0074] Step 6-1: Dempster-Shafer evidence theory allows combining evidence from different sources to obtain the overall uncertainty of the model. Here, we need to combine the quality sets of three different self-supervised models. Among them, it includes the confidence quality of each self-supervised model and uncertainty u q Combining the evidence of the contrastive self-supervised model, the masked self-supervised model, and the autoregressive prediction self-supervised model requires combining M cpc 、M mpc and M apc :
[0075]
[0076] The specific calculation rules are:
[0077]
[0078]
[0079] in, is a measure of the amount of conflict between the three mass sets, is the normalization factor. After the above calculation, the final classification result is obtained.
[0080] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An uncertainty-based integrated self-supervised speaker recognition method, characterized by: The integrated self-supervised speaker recognition method comprises the following steps: Step 1: Collect the speaker's voice data and annotate the voice data; Step 2: Preprocess the original speech data, extract the Mel-level spectrogram, obtain the feature vector of the speech data, and construct a speaker recognition data set; Step 3: Input the Mel-spectrogram features obtained in step 2 into several pre-trained self-supervised models in sequence, and extract the output of the last layer of the self-supervised model respectively; Step 4: Regarding the downstream classification task, the last layer output of each self-supervised model in step 3 is used as the input of the fully connected layer, and the output of the fully connected layer is calculated through the ReLU activation function to obtain the evidence of the classification decision of the input speech data under each model; Step 5: Calculate the Dirichlet distribution parameters under the subjective logic framework, and then calculate the confidence quality and uncertainty of each self-supervised model output; Step 6: Use the Dempster rule to fuse the output decision results of several self-supervised models to obtain the final probability and overall uncertainty of each class, and output the final classification result.
2. The method for speaker recognition based on uncertainty according to claim 1, characterized in that: The self-supervised model in step 3 is a masked self-supervised model, a contrastive self-supervised model, and an autoregressive prediction self-supervised model.
3. The method for speaker recognition based on uncertainty integrated self-supervision according to claim 2, characterized in that: The step 3 specifically includes the following steps: Step 3-1: Initialize the network parameters of the masked self-supervised model, the contrastive self-supervised model, and the autoregressive prediction self-supervised model, using the model parameters pre-trained on a large unlabeled public speech dataset as the initialization parameters. The masked self-supervised and contrastive self-supervised models are both based on a multi-layer Transformer model structure, and the autoregressive prediction self-supervised model structure is a 3-layer LSTM network structure. Step 3-2: Use different loss functions to calculate the loss values between the prediction results of the masked self-supervised model, the contrastive self-supervised model, and the autoregressive prediction self-supervised model and the query sample. The contrastive self-supervised model uses the InfoNCE loss, as shown below: In order to predict the future value x of the speech sequence t+k , and the speech potential representation w after model prediction t Together we construct the probability density function f k (x t+k ,w t ), used to retain x t+k and w t The mutual information between them, N represents X={x1,x2,...,x N }, one of which is from the distribution p(x t+k |w t ), and the rest are positive samples from distribution p(x t+k ) negative samples; The masked self-supervised model and the autoregressive prediction self-supervised model use the L1 loss function as shown below: Among them, x t is the input sequence, y t is the output sequence, t=1,2,…,n.
4. The method for speaker recognition based on uncertainty integrated self-supervision according to claim 3, characterized in that: Step 5 specifically includes the following steps: Step 5-1: For the K-classification problem, subjective logic assigns a confidence quality to each class label based on the evidence and assigns an uncertainty to the entire framework. For the self-supervised model q, the K+1 quality values are all non-negative and sum to 1: Among them, u q is the uncertainty, is the confidence quality of the kth class. In the self-supervised model, under the K classification task, its model output is used as evidence, and subjective logic uses the evidence and the parameters of the Dirichlet distribution Connect, Evidence q That is, the decision result after the output of the self-supervised model is calculated by the ReLU activation function, and the parameters of the Dirichlet distribution are can be Export, that is Step 5-2: Calculate confidence mass and uncertainty u q , expressed as: in, is the Dirichlet intensity. From the above formula, we can know that the more evidence of the kth category is observed, the greater the probability of the sample being classified as the kth category. Correspondingly, the less total evidence is observed, the greater the uncertainty. Confidence allocation is regarded as a subjective opinion, and the corresponding Dirichlet distribution p cpc The formula for calculating the class probability mean is:
5. The method for speaker recognition based on uncertainty integrated self-supervision according to claim 3, characterized in that: Step 6 uses the Dempster rule to fuse the output decision results of the three self-supervised models to obtain the final probability and overall uncertainty of each class and output the final classification result, which specifically includes the following steps: Step 6-1: Dempster-Shafer evidence theory allows combining evidence from different sources to obtain the overall uncertainty of the model, combining the quality sets of the three different self-supervised models Among them, it includes the confidence quality of each self-supervised model and uncertainty u q , in the speech data of the contrast self-supervised model and the masked self-supervised model, the combination M cpc and M mpc : The specific formula is: in, is a measure of the amount of conflict between two mass sets, is the normalization factor. After the above calculation, the final classification result is obtained.
6. The method for speaker recognition based on uncertainty integrated self-supervision according to claim 1, characterized in that: Step 4 specifically includes the following steps: Step 4-1: The self-supervised model includes an L-layer neural network, and the outputs of the L-layer network are [D1, D2, ..., D L ], extract the last layer of network D L The output of D L The output of is mapped to K dimensions; Step 4-2: The output of the fully connected layer is calculated as evidence through the ReLU activation function. The ReLU activation function is as follows: f(x)=max(x,0).
7. The method for speaker recognition based on uncertainty integrated self-supervision according to claim 1, characterized in that: In step 2, the original speech data is preprocessed, the Mel-level spectrogram is extracted, the feature vector of the speech data is obtained, and a speaker recognition data set is constructed, which specifically includes the following steps: Step 2-1: Convert the preprocessed speech data into a spectrogram, and pass it through a Mel filter bank to obtain the Mel spectrogram features of the speech; Step 2-2: For a Mel-spectrogram set X containing N speech samples = {x1, x2, ..., x N }, according to the label information, form a sample pair (x n ,z n ), where x n Represents the Mel spectrogram of the nth speech sample, z n is the label of the nth sample, n=1,2,...,N.
8. The method for speaker recognition based on uncertainty integrated self-supervision according to any one of claims 1 to 7, characterized in that: The integrated self-supervised speaker recognition method is implemented by an integrated self-supervised speaker recognition system, and the integrated self-supervised speaker recognition system includes: Speech feature extraction module: pre-processes speech data, extracts the Mel spectrum features of speech, and annotates the data set; Self-supervised model training module: Use a large amount of unlabeled data to pre-train the masked self-supervised model, contrastive self-supervised model, and autoregressive prediction self-supervised model, and input the Mel spectrogram feature results of the speech data into the three self-supervised models respectively, and extract the output of the last layer of the model; Classification decision module: The last layer output of each self-supervised model is used as the input of the fully connected layer, and the output of the fully connected layer is calculated through the ReLU activation function to obtain the evidence of the input speech data under each model; Uncertainty estimation module: Calculate the confidence quality and uncertainty of each self-supervised model output through the obtained evidence and Dirichlet distribution parameters; Decision fusion module: Use the Dempster rule to fuse the output decisions of the three self-supervised models to obtain the final probability and overall uncertainty of each class, and output the final classification result.
Citation Information
Patent Citations
Semi-supervised audio event labeling method based on self-supervised contrast learning
CN112820322A
Speaker irrelevant speech emotion recognition method and system based on unsupervised domain adversarial learning
CN113555038A