Self-Supervised Speaker Model Training Method, Electronic Device, and Storage Medium
Through the dynamic loss gate and label correction method, the Gaussian hybrid model automatically filters and corrects training data, and solves the flexibility of fixed threshold loss gates, improving the performance and efficiency of self-supervised speaker verification.
Patent Information
- Application Number
- CN202210476322.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-04-29
AI Technical Summary
In the existing self-supervised speaker verification method, when using fixed threshold loss gates to screen pseudo-labels, the threshold dependence on manual experience is not flexible enough, and the unreliable data is not fully utilized, affecting the model performance.
The Gaussian hybrid model is used to dynamically simulate the loss distribution, automatically distinguish reliable and unreliable data through dynamic thresholds, and use model prediction to label correction, filter and correct training data.
It improves the performance and convergence speed of the model, effectively utilizes unreliable data, and improves the recognition accuracy of the speaker model.
Smart Images

Figure CN114707668B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of model training, and particularly relates to a self-supervised speaker model training method, an electronic device, and a storage medium. Background Art
[0002] Currently, self-supervised speaker verification is mainly divided into two stages. In the first stage, contrastive learning is used to train a speaker feature extractor. In the second stage, the speaker features in the data are extracted using this extractor, pseudo-labels are generated by clustering, and then a new speaker feature extractor is trained using the pseudo-labels. Repeating the second stage several times can achieve very good results.
[0003] In related technologies, there are self-supervised speaker verifications that do not use a loss gate and those that use a fixed-threshold loss gate.
[0004] On the one hand, in the training process of the second stage of self-supervised speaker verification without using a loss gate, all data pseudo-labels are used for training. The inventor found that for self-supervised speaker verification without using a loss gate, a large number of low-quality pseudo-labels will damage the performance of the model, and the non-use of the loss gate cannot filter out low-quality data.
[0005] On the other hand, using a fixed-threshold loss gate believes that some of the data pseudo-labels are incorrect, and the pseudo-labels of data with small losses are more reliable. By designing a fixed-threshold loss gate to screen out those with relatively small losses, the label quality of the training data is improved, thereby enhancing the final performance. The inventor found that using a loss gate with a fixed threshold can improve performance, but this threshold depends on manual experience for design, is not flexible enough, and the low-quality data that is screened out is not fully utilized. Summary of the Invention
[0006] Embodiments of the present invention provide a self-supervised speaker model training method, an electronic device, and a storage medium, which are used to solve at least one of the above technical problems.
[0007] In a first aspect, an embodiment of the present invention provides a self-supervised speaker model training method, including: for each round of training, calculating the loss values of all training data in the current round, modeling the losses of all the training data using a Gaussian mixture model to obtain a dynamic threshold; screening the current training data based on the relationship between the loss values of the current training data and the dynamic threshold; and training the speaker model using the screened training data.
[0008] In a second aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the self-supervised speaker model training method according to any embodiment of the present invention.
[0009] In a third aspect, an embodiment of the present invention further provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute the steps of the self-supervised speaker model training method according to any embodiment of the present invention.
[0010] The method of this application uses a Gaussian mixture model to dynamically simulate the loss distribution, and uses the estimated dynamic threshold to automatically distinguish reliable data and unreliable data. In a further embodiment, in order to better utilize the unreliable data instead of directly discarding them, the embodiment of this application uses model prediction to perform label correction to correct the unreliable labels. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a flowchart of a self-supervised speaker model training method provided by an embodiment of the present invention;
[0013] Figure 2 It is a loss distribution diagram provided by an embodiment of the present invention;
[0014] Figure 3 It is a flowchart of a specific example of a self-supervised speaker model training method provided by an embodiment of the present invention;
[0015] Figure 4 It is an unlabeled distillation framework for self-supervised speaker representation learning provided by an embodiment of the present invention;
[0016] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0018] Please refer to Figure 1 , which shows a flowchart of an embodiment of the self-supervised speaker model training method of the present application. The self-supervised speaker model training method of this embodiment can be applied to training a speaker model.
[0019] As Figure 1 shown, in step 101, for each round of training, calculate the loss value of all training data in the current round, and use a Gaussian mixture model to model the loss values of all the training data to obtain a dynamic threshold;
[0020] In step 102, screen the current training data based on the relationship between the loss value of the current training data and the dynamic threshold;
[0021] In step 103, use the screened training data to train the speaker model.
[0022] In this embodiment, for step 101, during each round of model training, the loss value of the training data in the current round will be calculated first, and a Gaussian mixture model will be used to model all the loss values based on the loss values of all the training data to obtain a dynamic threshold. Since the training data corresponding to each round of training is different, the dynamic threshold also changes dynamically. Subsequently, for step 102, the training data in the current round can be screened according to the relationship between the loss value of each training data and the dynamic threshold, and the unreliable data can be screened out. Finally, the model is trained using the screened training data.
[0023] The method of this embodiment dynamically simulates the loss distribution by using a Gaussian mixture model and automatically distinguishes reliable data and unreliable data using the estimated dynamic threshold.
[0024] In some alternative embodiments, screening the training data based on the relationship between the loss value of the training data and the dynamic threshold includes: determining whether the loss value of the current training data is greater than the dynamic threshold; if the loss value of the current training data is greater than the dynamic threshold, using the output of the speaker model to correct the label of the current training data, and using the training data with the corrected label for training the speaker model; if the loss value of the current training data is not greater than the dynamic threshold, directly using the current training data for training the speaker model. Thus, for unreliable data, by correcting its label according to the prediction of the model, the unreliable data can be utilized instead of being directly discarded. For example, using the output of the model as the label of the data to correct the wrong label.
[0025] In a further alternative embodiment, using the output of the speaker model to correct the label of the current training data includes: determining whether the predicted posterior probability of the output of the speaker model corresponding to the current training data is greater than a preset probability threshold; if the predicted posterior probability of the output of the speaker model corresponding to the current training data is greater than the preset probability threshold, using the output of the speaker model to replace the label of the current training data for training; if the predicted posterior probability of the output of the speaker model corresponding to the current training data is not greater than the preset probability threshold, discarding the current training data. Thus, for unreliable data, a secondary screening is performed using the predicted posterior probability of the output of the model. For data with a relatively high preset posterior probability in the unreliable data, it indicates that the output of the model is relatively accurate. At this time, using the output of the model to replace the label of the unreliable data makes it reliable and can be used for subsequent model training.
[0026] In some alternative embodiments, modeling the loss values of all the training data using a Gaussian mixture model to obtain a dynamic threshold includes: using a Gaussian mixture model with two Gaussian components to dynamically simulate the distribution of the loss values of all the training data; determining the dynamic threshold based on the loss value at which the probabilities of the Gaussian mixture model belonging to the two Gaussian components are equal. Thus, by using a Gaussian mixture model with two Gaussian components to simulate the distribution of the loss values of the training data, the dynamic threshold can be determined according to the loss value at which the probabilities of the two Gaussian components are equal.
[0027] In some alternative embodiments, the training data includes the data to be trained and the label corresponding to the data to be trained. Thus, for data with a wrong label, only the label needs to be corrected for training.
[0028] In some alternative embodiments, before calculating the loss value of all training data in the current round for each round of training, the method further includes: using a speaker encoder pre-trained by unlabeled distillation to extract the speaker embedding of each utterance as the data to be trained; applying the k-means clustering algorithm to assign the same pseudo-label to the utterances belonging to the same cluster, where the pseudo-label is used for subsequent self-supervised training. Thus, by using a speaker encoder pre-trained by unlabeled distillation, only based on maximizing the similarity between the feature distributions of different augmented segments from the same utterance, so it does not have the problem of negative example pairs, and there will be no situation where, when training the network, because they are regarded as negative example pairs, they are pushed apart in the embedding space, resulting in more errors.
[0029] In some alternative embodiments, the speaker model includes a speaker verification model, a speaker identification model, and a speaker classification model.
[0030] It should be noted that the above method steps do not limit the execution order of each step. In fact, some steps may be executed simultaneously or in the reverse order of the step definition. This application has no restrictions here.
[0031] Next, some problems encountered by the inventor in the process of implementing the present invention and a specific embodiment of the finally determined solution will be described to enable those skilled in the art to better understand the solution of this application.
[0032] The inventor found that in the prior art, the main method to solve the above problems is to use a loss gate with a fixed threshold. Since the threshold is manually designed, it has the defect of lack of flexibility.
[0033] The solution of this application mainly starts from the following aspects for design and optimization: Dynamic loss gate: Analyze the distribution of losses. As Figure 2 shown, it can be seen that it presents an obvious bimodal distribution. The reliable data is the left peak, and the unreliable data is the right peak. In the embodiments of this application, the Gaussian mixture model can be used to model the reliable and unreliable data respectively to obtain the curve in the figure. The loss value corresponding to the intersection point of the curves is the threshold obtained in the embodiments of this application. This threshold is dynamic and changes with the training process. Label correction: The unreliable data is the data with incorrect labels. In order to utilize this unreliable data, the embodiments of this application assume that the result output by the model will be more accurate and reliable than the given label. Therefore, the embodiments of this application use the output of the model as the label of the data to correct the incorrect label.
[0034] Continue to refer to Figure 3 which shows the flowchart of the implementation of the solution of this application.
[0035] As Figure 3As shown, first, for each round of training, the embodiments of the present application first calculate the losses of all data, and then use a Gaussian mixture model to model them to obtain a dynamic threshold.
[0036] Second, the data is segmented. If the loss is greater than the threshold, it means the label is unreliable, and then the output of the model is used to correct the label.
[0037] Then, if the loss is less than the threshold, it means the label is reliable and no processing is required.
[0038] Finally, training is performed using the labels.
[0039] The inventors conducted a large number of experiments. The experimental results show that both the dynamic threshold loss gate (DLG) and label correction (LC) proposed in the embodiments of the present application can improve performance. Referring to Table 2 in the subsequent experimental data, the first row is the result without using the loss gate, the second row is the result using the fixed threshold loss gate (LGL). After adding the dynamic threshold loss gate DLG of the embodiments of the present application, the effect is significantly improved. After adding label correction LC, the best result is obtained.
[0040] Continuing to refer to Table 2, which shows the results of several iterations, it can be seen that the dynamic loss gate and label correction proposed in the embodiments of the present application can achieve the best results in each round and converge faster.
[0041] The following describes the process by which the inventors implemented the present application, the experiments conducted, and the relevant experimental data, so that those skilled in the art can better understand the solution of the present application.
[0042] In self-supervised speaker verification, due to the existence of a large number of unreliable labels, the quality of the pseudo-labels determines the upper bound of the system. In this work, the embodiments of the present application propose a dynamic threshold loss gate and label correction (DLG-LC) to mitigate the performance degradation caused by unreliable labels. In the dynamic threshold loss gate (DLG), the embodiments of the present application use a Gaussian mixture model (GMM) to dynamically simulate the loss distribution and use the estimated dynamic threshold to automatically distinguish reliable and unreliable labels. In addition, to better utilize unreliable data instead of directly discarding them, the embodiments of the present application use model predictions to perform label correction to correct unreliable labels. Compared with the current best self-supervised learning speaker verification system, the DLG-LC proposed in the embodiments of the present application converges faster and achieves relative improvements of 11.45%, 18.35%, and 15.16% in the Vox-O, Vox-E, and Vox-H trials of the Voxceleb1 evaluation dataset, respectively.
[0043] 1. Introduction
[0044] Recently, deep learning-based methods have been widely applied to the speaker verification (SV) task and achieved excellent performance. However, all these methods require a large amount of training data with speaker labels, and the collection of data with accurate speaker labels is usually very difficult and expensive.
[0045] To make full use of a large amount of unlabeled data, many researchers have proposed obtaining good speaker representations in a self-supervised learning manner. According to the iterative framework proposed in the related technology, the training of the current most commonly used self-supervised speaker verification system usually includes two stages. In the first stage, a speaker encoder is trained by the method of contrastive learning. In the second stage, pseudo-labels are estimated from the model pre-trained in the previous stage, and then a new speaker model is trained according to the estimated pseudo-labels. This step is iteratively executed multiple times to continuously improve the performance.
[0046] Currently, this phased training method has achieved excellent performance, but there are still some unreasonable assumptions that limit the further improvement of the system performance. In the first stage, the contrastive learning-based method assumes that speech segments from different utterances belong to different speakers. When training the network, they will be regarded as negative example pairs and be far away from each other in the embedding space. Undoubtedly, this inaccurate assumption will bring some errors because different utterances may come from the same speaker. In the second stage, a clustering algorithm is used to generate pseudo-labels for each utterance according to the previously learned representations, and the estimated pseudo-labels will be used for subsequent supervised training. However, related research shows that many of the estimated pseudo-labels are unreliable. Such low-quality and unreliable labels will confuse and reduce the performance of the model during the training process, which also indicates that finding a method to select high-quality labels is the key to improving the performance. Some people have applied a method based on clustering confidence to purify the pseudo-labels, but the improvement is very small. In addition, it has been observed that the labels of data with lower losses during the training process are relatively more reliable. Based on this assumption, they set a fixed loss threshold in each iteration to distinguish reliable labels from unreliable labels, and only use the samples of reliable labels to update the network. Although this approach has brought further improvement, there is still much room for improvement. For example, the threshold for each iteration is manually set, which is not flexible, and the data with unreliable labels is not fully utilized.
[0047] To address these issues and improve existing methods, the embodiments of this application introduce DINO (distillation with no labels) in the first stage. It is only based on maximizing the similarity between the feature distributions of different augmented segments from the same utterance, thus avoiding the problem of having no negative examples. Experiments in the embodiments of this application show that DINO outperforms all previous contrastive learning-based methods in the Voxceleb dataset. In the second stage, the embodiments of this application use a Gaussian mixture model (GMM) to model the loss distribution of all data. More specifically, a GMM with two components, where each Gaussian component can represent the loss distribution label of data with reliable labels or data with unreliable labels. Through the estimated GMM, the embodiments of this application can automatically obtain a loss threshold for distinguishing the two types of labels, which is more flexible than manually adjusted thresholds. Additionally, instead of directly discarding the data with unreliable labels, the embodiments of this application use the model prediction as an alternative label and use it to correct the unreliable labels. The embodiments of this application name the above two strategies proposed in the second stage as dynamic loss gating and label correction (DLG-LC). Using DLG-LC, the embodiments of this application construct a more powerful self-supervised speaker verification system and achieve an equal error rate of 1.468% on the Voxceleb 1 test set, which is the state-of-the-art today.
[0048] 2. Method
[0049] 2.1 DINO-based Self-Supervised Learning
[0050] For the contrastive learning-based methods in previous work, they share the same assumption that the segments in a batch belong to different speakers. However, this assumption does not always hold because repeated speakers may appear in the same batch. The embodiments of this application can calculate the probability of repeated speakers on Voxceleb 2 through formula (1):
[0051]
[0052] where S is the speaker number and N is the batch size. Obviously, a larger batch size will result in a higher repetition probability, which will have an adverse effect on the model.
[0053] Figure 4Shows an unlabeled distillation framework (DINO) for self-supervised speaker representation learning. Among them, the Chinese-English comparison is as follows: Short short segment, Long long segment, Encoder embedder, Student student network, Teacher teacher network, EMA exponential moving average, Cosine similiary cosine similarity, Projection head projection head, Stopgradient gradient truncation, centering softmax centered softmax.
[0054] The embodiment of the present application introduces DINO which does not require negative example pairs to solve this problem. The entire framework is as Figure 4 shown. First, the embodiment of the present application uses a multi-crop strategy to sample 4 short segments and 2 long segments from the utterance, where the long segments can be used to extract more stable speaker representations. The embodiment of the present application assumes that the segments cropped from the same utterance share the same speaker identity, and then enhances them in different ways by adding noise or room impulse response to obtain robust performance. In subsequent embodiments, all short segments pass through the student network, while only the long segments pass through the teacher network. Therefore, the short-to-long correspondence is encouraged by minimizing the cross-entropy H(·) between the two distributions:
[0055]
[0056] where P t and P s represent the output probability distributions of the teacher network g θt and the student network g θs respectively. The embodiment of the present application can calculate P by normalizing the output using the softmax function:
[0057] P s = Softmax(g θs (x) / ∈ s ) (3)
[0058] ∈ s >0 is the temperature parameter that controls the sharpness of the output distribution, and a similar formula applies to P t with temperature ∈ t . In addition, the output of the teacher network is centered on the mean calculated over the batch. Both sharpening and centering are used to avoid model collapse.
[0059] Both the teacher network and the student network have the same architecture but different parameters. The student network is updated via gradient descent, while the teacher network is updated via the exponential moving average (EMA) of the student network's parameters. The update rule is θt←λθt+(1-λ)θs, where λ follows a cosine schedule from 0.996 to 1 during training. In the embodiments of this application, the speaker embeddings are extracted by the encoder and fed into the projection head, which consists of a 3-layer MLP with a hidden dimension of 2048 and a normalization and a weight-normalized fully connected layer with K dimensions.
[0060] To make DINO more suitable for speaker verification, the embodiments of this application also add a consistency loss to maximize the cosine similarity between embeddings from the same utterance. This loss is to ensure that the speaker embeddings are in the cosine space and are more suitable for subsequent scoring and clustering. Then, the final loss is connected by the coefficient α:
[0061]
[0062] where e represents the speaker embeddings extracted by the encoder.
[0063] 2.2. Dynamic Loss Gate
[0064] The embodiments of this application use a speaker encoder pre-trained by DINO to extract the speaker embeddings of each utterance. Then, the embodiments of this application apply the k-means clustering algorithm to assign the same speaker pseudo-labels to the audios belonging to the same cluster. Based on these pseudo-labels, a new speaker encoder can be trained to generate more robust speaker embeddings. This process will be repeated several iterations until the model converges. In this process, the data with unreliable labels will undoubtedly reduce the performance of the model. Those skilled in the art observed an obvious distribution difference between the data losses with reliable and unreliable labels in the final convergence stage and proposed a loss gating learning (LGL) method that effectively selects reliable data using a fixed threshold. Using LGL achieved amazing results, but the threshold setting of this method is highly dependent on human experience, and the unreliable data is directly discarded without being fully utilized.
[0065] To obtain a suitable threshold, the embodiments of this application implemented LGL and conducted experiments to analyze the loss distribution on Voxceleb 2. The histogram can be continued to refer to Figure 2 as shown. From the loss distribution, it is obvious that there are two spikes in the distribution. Similar experiments show that these two peaks can respectively represent the data with reliable and unreliable labels. If the embodiments of this application can find a way to model this distribution, then the threshold can be directly calculated and dynamically change with the loss distribution, which can also avoid laborious manual tuning.
[0066] The Gaussian distribution is an important type of continuous probability distribution for real-valued random variables. The general form of its probability density function is shown in Equation (5).
[0067]
[0068] where μ is the location parameter and σ is the scale parameter. It is bell-shaped, low on both sides and high in the middle, and its shape is very similar to the "peak" of the loss in Figure 2 In this case, a Gaussian mixture model (GMM) with two components can be used to fit the distributions of reliable and unreliable samples:
[0069]
[0070] where λ is the weight of each Gaussian component. In the embodiments of this application, the fitting curve is plotted in Figure 2 It is obvious that the two "peaks" can be approximated by two weighted Gaussian distributions. Then, the embodiments of this application can easily obtain a threshold τ to distinguish reliable and unreliable data by finding the loss value at which the probabilities belonging to the two components are equal:
[0071] τ: p1(τ) = p2(τ) (7)
[0072] where and
[0073] For each round of training), the embodiments of this application will re-estimate the GMM according to the loss value recorded during this round of training. Therefore, the threshold τ can be dynamically adjusted according to the current training conditions.
[0074] Similar to LGL, the embodiments of this application introduce this threshold τ into the speaker classification loss function ArcMarginSoftmax (AAM) to filter out the data with smaller losses and only use the remaining data to update the parameters of the network.
[0075]
[0076] where θ j,i is the angle between the column vector W j and the embedding x i s is the scale factor and m is the hyperparameter controlling the margin. AAM can enforce a larger gap between the nearest speakers and is widely used in speaker recognition tasks.
[0077] 2.3. Label Correction
[0078] To effectively utilize unreliable data with large losses instead of directly discarding them, the embodiments of this application also propose Label Correction (LC) to dynamically correct pseudo-labels. As the model is continuously optimized, the model outputs of unreliable samples can reflect their true labels to a certain extent. To utilize this ability, the embodiments of this application assume that the output predictions of the model are more reliable than the pseudo-labels generated by clustering. Therefore, the embodiments of this application add the predicted posterior probability as the target label to the objective function to prevent the model from fitting incorrect labels. However, not all predicted labels are suitable for training. The embodiments of this application assume that if the model assigns a high probability to one of the possible classes, the predicted label has a high confidence. Therefore, the embodiments of this application introduce another threshold τ2 to retain the predicted labels whose maximum class probability is higher than τ2. And the label correction loss is defined as follows:
[0079]
[0080] where pi and p^i represent the output probabilities of the augmented segment and its corresponding clean version respectively. H(·) represents the cross-entropy between two probability distributions. The embodiments of this application also apply a sharpening operation on p^i to encourage a peak distribution with ε c ε c defined in Equation (3).
[0081] Finally, these two losses are combined to optimize the speaker model.
[0082] L = LDLG + LLC (10)
[0083] 3. Experiments
[0084] 3.1 Datasets and Configurations
[0085] In the experiments of the embodiments of this application, the development set of Voxceleb2 is used to train the network, and speaker labels are not used during this process. It contains 1,092,009 utterances of 5,994 speakers. For evaluation, the embodiments of this application report the experimental results of 3 defined test sets: Voxceleb O, E, and H test sets. The main metrics used in this paper are Equal Error Rate (EER) and Normalized Minimum Detection Cost Function (minDCF).
[0086] For the two stages of the experiments of the embodiments of this application, an 80-dimensional filter bank with a 25ms window and a 10ms shift is used as the acoustic feature. In addition, the noise MUSAN and RIRs datasets are also used for online data augmentation.
[0087] 3.1.1 First Stage: DINO
[0088] For DINO, considering the training time and memory limitations, embodiments of this application use a simplified version of ResNet to learn speaker representations from long (3 seconds) and short (2 seconds) segments, and then encode them into 256-dimensional speaker embeddings through the network. Similar to the configuration in the related art, K in the DINO projection head is set to 65536. The temperatures of the teacher network ∈t and the student network ∈s are 0.04 and 0.1 respectively. In addition, embodiments of this application set α to 0.001 to balance the two losses during training.
[0089] 3.1.2 Second stage: Iterative pseudo-label learning
[0090] In this stage, to better compare with the results in the related art, embodiments of this application use ECAPA-TDNN as the encoder in embodiments of this application. For clustering, embodiments of this application use the k-means algorithm to assign pseudo-labels to the training set. To verify the robustness of the method in embodiments of this application, embodiments of this application choose 7500 instead of 6000 as the number of clusters. When using DLG to train the network on the pseudo-labels, embodiments of this application set the margin to 0.2 and scale it to 32 in AAM. For LC, embodiments of this application set the sharpening parameter εc to 0.1 and the threshold τ2 to 0.5. Using the SGD optimizer, the learning rate decays exponentially from 0.1 to 5e-5. The training process of embodiments of this application is repeated 5 times iteratively. In the last iteration, embodiments of this application expand the channel size of the encoder (ECAPA-TDNN) to 1024.
[0091] 3.2. Comparison of DINO with previous works
[0092] To comprehensively evaluate DINO, embodiments of this application conducted experiments, and the corresponding results are listed in Table 1. All methods were trained on Voxceleb 2 and evaluated on the Vox-O test set. From the table, embodiments of this application can find that DINO without negative example pairs is better than all previous contrast-based methods, which indicates that negative example pairs are indeed the bottleneck for performance improvement. In addition, embodiments of this application also provide ablation experiments of DINO at the bottom of Table 1. Embodiments of this application can see that DINO without EMA obtained very poor results, which indicates that EMA is the key to preventing the model from crashing. When embodiments of this application add multi-crop and cosine strategies respectively during the training process, they can all improve the performance of DINO to varying degrees and prove their necessity.
[0093] Table 1: Comparison of DINO with other self-supervised speaker verification works. EER (%) and minDCF are evaluated on the Vox-O test set.
[0094]
[0095] 3.3. Dynamic Loss Gate with Label Correction
[0096] Table 2: Comparison of EER(%) of Vox-O, E, and H of the proposed DLG-LC in Iteration 1. In this experiment, the pseudo-labels were estimated by the DINO system pre-trained according to the embodiments of this application. DINO means that in the system training of the embodiments of this application, all data with estimated pseudo-labels were simply used as supervision signals without any loss gate.
[0097]
[0098] Based on the pseudo-labels generated by DINO pre-trained in the first stage, the embodiments of this application conducted some experiments to illustrate the effectiveness of the method proposed in the embodiments of this application. First, the embodiments of this application conducted an ablation study on DLG-LC in Iteration 1. As can be seen from Table 2, compared with the baseline trained without any data selection, LGL can bring significant improvement. This means that the loss gate can effectively filter out reliable labels beneficial to the model. However, the choice of threshold also has a non-negligible impact on the model performance. Based on the estimated GMM, the DLG proposed by the embodiments of this application can dynamically adjust the threshold according to the current training situation and obtain better performance than LGL that uses a fixed threshold throughout the training process. In addition, the embodiments of this application added the LC strategy to make full use of the data with unreliable labels, and the result was further improved.
[0099] Then, the embodiments of this application summarized the performance of using DLG-LC on the Vox-O set in each iteration, and the results are shown in Table 3. The embodiments of this application first compared the iterative results of generating pseudo-labels using DINO and SimCLR respectively. They were both trained with all the training data without any loss gate strategy.
[0100] Table 3: Comparison of EER(%) of the Vox-O test set for different iterations of the proposed DLG-LC and other strategies. SimCLR and DINO mean that the embodiments of this application only used all the estimated pseudo-labels in the training process without any loss gate.
[0101]
[0102] Obviously, DINO has consistent advantages over SimCLR in each iteration, indicating that DINO can generate more discriminative speaker embeddings, thus generating more reliable pseudo-labels. After applying the loss gate strategy, the results have been significantly improved. Compared with LGL, DLG-LC with a dynamic threshold proposed in the embodiments of the present application shows great advantages. It not only has superior performance but also a faster convergence speed. Only through 3 iterations, DLG-LC achieved results comparable to the final iteration of LGL, which benefits from the dynamic threshold and label correction.
[0103] Finally, Table 4 presents the performance comparison between DLG-LC proposed in the embodiments of the present application and other self-supervised speaker verification systems. Most of them are from the latest Voxceleb Speaker Recognition Challenge (VoxSRC), representing the current state-of-the-art performance. It can be clearly seen from the table that DLG-LC exceeded all existing methods with only 5 iterations on the Vox-O, Vox-E, and Vox-H sets, and outperformed the best method by relative advantages of 11.45%, 18.35%, and 15.16% respectively. In addition, the clustering algorithm adopted in the embodiments of the present application is k-means, which is simpler than the agglomerative hierarchical clustering (AHC) used in other systems. Moreover, the embodiments of the present application use 7,500 instead of 6,000 as the number of clusters because 5,994 is the actual number of speakers in the training set. Nevertheless, the system of the embodiments of the present application still achieved the state-of-the-art performance, which also shows the robustness of the system of the embodiments of the present application.
[0104] Table 4: Comparison of EER (%) of different self-supervised speaker verification methods on Vox-O, E, and H.
[0105]
[0106] 4. Conclusion
[0107] In this work, the embodiments of the present application proposed dynamic loss gate and label correction (DLG-LC) to mitigate the performance degradation of self-supervised speaker verification caused by inaccurate assumptions and labels. By modeling the loss histogram using a Gaussian distribution, the embodiments of the present application can obtain a dynamic loss gate to select reliable data during pseudo-label training. Instead of discarding unreliable data, the embodiments of the present application use the predicted posterior probability as the target distribution to prevent fitting to incorrect samples. Experiments on Voxceleb show that DLG-LC proposed in the embodiments of the present application is more robust and achieves the state-of-the-art performance.
[0108] In some other embodiments, the embodiments of the present invention further provide a non-volatile computer storage medium, which stores computer-executable instructions that can execute the self-supervised speaker model training method in any of the above method embodiments;
[0109] As an implementation manner, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as follows:
[0110] For each round of training, calculate the loss value of all training data in the current round, and use a Gaussian mixture model to model the loss values of all the training data to obtain a dynamic threshold;
[0111] Based on the relationship between the loss value of the current training data and the dynamic threshold, screen the current training data;
[0112] Use the screened training data to perform the speaker model training.
[0113] The non-volatile computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area can store an operating system and application programs required for at least one function; the storage data area can store data created according to the use of the self-supervised speaker model training device, etc. In addition, the non-volatile computer-readable storage medium may include a high-speed random access memory, and may also include non-volatile memories, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the self-supervised speaker model training device through a network. Examples of the above network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof.
[0114] The embodiments of the present invention further provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to execute any of the above self-supervised speaker model training methods.
[0115] Figure 5 is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. As shown in Figure 5 shown, the device includes: one or more processors 510 and a memory 520. Figure 5 Here, one processor 510 is taken as an example. The device for the self-supervised speaker model training method may further include: an input device 530 and an output device 540. The processor 510, the memory 520, the input device 530, and the output device 540 can be connected through a bus or other means. Figure 5Take the bus connection as an example. The memory 520 is the above-mentioned non-volatile computer-readable storage medium. The processor 510 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 520, that is, implements the self-supervised speaker model training method in the above method embodiment. The input device 530 can receive input digital or character information, and generate key signal inputs related to the user settings and function control of the communication compensation device. The output device 540 can include display devices such as a display screen.
[0116] The above product can execute the method provided by the embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiment of the present invention.
[0117] As an implementation manner, the above electronic device is applied to a self-supervised speaker model training device and is used for a client, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:
[0118] For each round of training, calculate the loss value of all training data in the current round, and use a Gaussian mixture model to model the loss value of all the training data to obtain a dynamic threshold;
[0119] Based on the relationship between the loss value of the current training data and the dynamic threshold, screen the current training data;
[0120] Use the screened training data for the speaker model training.
[0121] The electronic device in the embodiment of the present application exists in various forms, including but not limited to:
[0122] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones (such as iPhone), multimedia phones, functional phones, and low-end phones, etc.
[0123] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc., such as iPad.
[0124] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players (such as iPod), handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.
[0125] (4) Server: A device that provides computing services. The server consists of a processor, hard disk, memory, system bus, etc. The server is similar to a general computer architecture, but due to the need to provide highly reliable services, it has higher requirements in terms of processing power, stability, reliability, security, scalability, manageability, etc.
[0126] (5) Other electronic devices with data interaction functions.
[0127] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.
[0128] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A self-supervised speaker model training method, comprising: Using a speaker encoder pre-trained by unlabeled distillation to extract the speaker embedding of each utterance as the data to be trained; Applying the k-means clustering algorithm to assign the same pseudo-label to the utterances belonging to the same cluster, wherein the pseudo-label is used for subsequent self-supervised training; For each round of training, calculating the loss value of all training data in the current round, and using a Gaussian mixture model to model the loss values of all the training data to obtain a dynamic threshold, including using a Gaussian mixture model with 2 Gaussian components to dynamically simulate the loss value distribution of all the training data; determining the dynamic threshold based on the loss value at which the probabilities of the Gaussian mixture model belonging to the 2 Gaussian components are equal; Screening the current training data based on the relationship between the loss value of the current training data and the dynamic threshold, including determining whether the loss value of the current training data is greater than the dynamic threshold; if the loss value of the current training data is greater than the dynamic threshold, using the output of the speaker model to correct the label of the current training data, and using the training data with the corrected label to train the speaker model; if the loss value of the current training data is not greater than the dynamic threshold, directly using the current training data for training the speaker model; Using the screened training data to train the speaker model.
2. The method according to claim 1, wherein The using the output of the speaker model to correct the label of the current training data includes: Determining whether the predicted posterior probability of the output of the speaker model corresponding to the current training data is greater than a preset probability threshold; If the predicted posterior probability of the output of the speaker model corresponding to the current training data is greater than the preset probability threshold, using the output of the speaker model to replace the label of the current training data for training; If the predicted posterior probability of the output of the speaker model corresponding to the current training data is not greater than the preset probability threshold, discarding the current training data.
3. The method according to claim 1, wherein, The training data includes the data to be trained and the labels corresponding to the data to be trained.
4. The method according to any one of claims 1 to 3, wherein The speaker model includes a speaker verification model, a speaker recognition model, and a speaker classification model.
5. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method according to any one of claims 1 to 4.
6. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 4.