A Full-Life-Cycle Speech Emotion Recognition Method Focusing on the Feature Spacing of Samples
By employing a pre-training, fine-tuning, and inference method with supervised contrastive learning and K-nearest neighbor retrieval, the method addresses data limitations in speech emotion recognition, improving classification accuracy and clarifying feature space boundaries.
Patent Information
- Application Number
- CN202310794609.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-06-30
AI Technical Summary
The field of speech emotion recognition faces the problems of data limits and blurred boundaries of emotional categories, resulting in a decrease in recognition accuracy.
The full life cycle method focusing on the sample feature spacing is adopted, including introducing large-scale pre-training models in the pre-training stage, guiding the model fine-tuning through cross-entropy loss and supervised comparative learning loss weighted summation in the fine-tuning stage, and improving sample spacing and feature space division using K nearest neighbor retrieval enhancement technology in the inference stage.
Effective utilization of limited data improves the accuracy of speech emotion recognition, especially the Weighted Accuracy and Unweighted Accuracy indicators on the IEMOCAP dataset perform better than existing algorithms.
Smart Images

Figure CN116645980B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer processing, and more particularly, to a full-life-cycle speech emotion recognition method focusing on the feature spacing of samples. Background Art
[0002] Emotion recognition is an important aspect in the field of human-computer interaction. Speech conveys rich emotional information through different attributes such as pitch, frequency, speed, and stress. With the development of artificial intelligence technology, speech emotion recognition (SER) has been widely applied in fields such as online education, artificial customer service, and mental health.
[0003] Currently, with the development of deep learning technology, model structures based on neural networks such as recurrent neural networks, time-delay neural networks, and convolutional neural networks have become the main methods for speech emotion recognition. Compared with traditional methods, these methods rely less on manually extracted audio features, and by learning deeper speech feature representations, the accuracy of speech emotion recognition has been improved.
[0004] However, data-driven deep learning technology also poses new challenges to speech emotion recognition. In order to use models with larger scale and higher robustness to extract more accurate features, the model paradigm of "pre-training + fine-tuning" has been applied in various fields of artificial intelligence. Compared with other related fields, the dataset scale in the field of speech emotion recognition is small, and the limitation of the data volume makes it so that there is currently no general pre-trained model that can be directly applied in speech emotion recognition. This results in inaccurate speech emotion feature representations, which will directly affect the accuracy of speech emotion recognition.
[0005] In addition, due to the similarity in prosody of some emotions (such as anger and excitement), in the field of unimodal speech recognition without referring to text information, it is difficult for the model to distinguish the acoustic features of the above emotions. In the feature space, there are problems with blurred classification boundaries for some emotion features, which reduces the accuracy of speech emotion recognition. Summary of the Invention
[0006] To alleviate the limitation of the data volume in the field of speech emotion recognition on application technologies and effectively solve the problem of blurred classification boundaries of different emotion categories, the present invention provides a method focusing on sample spacing throughout the entire life cycle of speech emotion recognition. This method involves improvements in three stages of speech emotion recognition: pre-training, fine-tuning, and inference. By extracting more accurate feature representations in the pre-training stage, improving sample spacing in the fine-tuning stage, and reusing the improved sample data in the inference stage, the limited data volume can be fully utilized, and the speech emotion representations between different categories in the feature space are more clearly divided, effectively improving the accuracy of speech emotion recognition.
[0007] The present invention mainly relates to three stages of the entire life cycle of speech emotion recognition: pre-training, fine-tuning, and inference stages.
[0008] In the pre-training stage, the present invention introduces a large-scale pre-trained model to extract more accurate speech representations; in the fine-tuning stage, guided by the weighted sum of the cross-entropy loss and the supervised contrastive learning loss, the model is fine-tuned so that the sample representation spacing learned by the model is improved. Specifically, the spacing between samples of the same class is reduced, and the spacing between samples of different classes is increased; in the inference stage, first, a data storage set is constructed to store the sample representations and sample labels of the training set and the validation set. To further utilize the improved sample spacing, through the method enhanced by K-nearest neighbor retrieval, the K samples most similar to the test sample in the data storage set are retrieved, and the weighted sum of the retrieved label distribution and the inference distribution result of the model for the test sample is obtained to get the final predicted label of the test sample.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] A full-life-cycle speech emotion recognition method focusing on sample feature spacing, characterized by including the following steps:
[0011] Step S101, randomly augment the input training samples;
[0012] Step S102, introduce a model trained on a large-scale data set as a pre-trained model;
[0013] Step S103, use the pre-trained model introduced in Step S102 to extract features from the sample instances obtained in Step S101, define positive and negative samples, and calculate the supervised contrastive learning loss;
[0014] Step S104, calculate the cross-entropy loss, and perform weighted summation with the supervised contrastive learning loss calculated in Step S103 to fine-tune the pre-training of the model;
[0015] Step S105: Use the model fine-tuned in step S104 to obtain the representation-label key-value pairs of the training samples, and construct a data storage set;
[0016] Step S106: Given a test sample, retrieve the K nearest neighbor samples to the test sample in the data storage set obtained in step S105, and record their label distribution;
[0017] Step S107: For the test sample given in step S106, use the model in step S104 to predict its output distribution;
[0018] Step S108: Weightedly sum the distributions obtained in steps S106 and S107 to obtain the final predicted label of the test sample.
[0019] For further optimization of this technical solution, the supervised contrastive learning loss L is calculated in step 103 scl as follows:
[0020]
[0021] where i ∈ I = {1, ……, 2N} represents the index of an instance, N is the number of samples, A(i) represents all indices except i, P(i) represents the indices of all positive samples with the same label as sample i, a ∈ A(i) represents a specific index of a sample other than i, p ∈ P(i) represents a specific index of a positive sample with the same label as sample i; τ is a hyperparameter for calculating the supervised contrastive learning loss; x i , x p , x a respectively represent the feature vectors of the audio samples corresponding to the subscripts.
[0022] For further optimization of this technical solution, the cross-entropy loss L is calculated in step 104 ce as follows:
[0023]
[0024] where N represents the number of samples, C represents the number of categories, y i represents the audio sample label, is the probability result that the i-th sample predicted by the model belongs to the c-th category.
[0025] For further optimization of this technical solution, step 104 weightedly sums the supervised contrastive learning loss L scl and the cross-entropy loss L ce to obtain the final loss L of the model as follows:
[0026] L = (1 - μ)L ce + μLscl
[0027] Among them, μ is a hyperparameter that balances the cross-entropy loss and the contrastive learning loss.
[0028] For further optimization of this technical solution, step 105 includes: using the model fine-tuned in step S104, performing a forward pass on all training set sample data, and creating a data storage set containing all training set sample data and validation set sample data according to the feature vectors and labels of the samples. The storage format is as follows:
[0029] (K, V) = {(x i , y i ), i ∈ D}
[0030] Among them, D is the set of all sample indices of the training set and the validation set, x i represents the feature vector obtained by calculating the i-th audio sample through the model in step S104, and y i is the label corresponding to the i-th audio sample.
[0031] For further optimization of this technical solution, step 108 includes: comprehensively retrieving the results from the data storage set in step S106 and the model inference results in step S107, and performing a weighted sum on them to obtain the final predicted distribution p(y|x) of the test sample as follows:
[0032] p(y|x) = αp knn (y|x) + (1 - α)p model (y|x)
[0033] Among them, α is a hyperparameter that adjusts the proportion of p knn (y|x) and p model (y|x), p knn (y|x) is the distribution of the labels of each category recorded by retrieving the K samples closest to the test sample in step S106, and p model (y|x) is the predicted output distribution obtained by performing inference on it using the model fine-tuned in step S104 in step S107.
[0034] For further optimization of this technical solution, the pre-trained model is the Wav2vec2.0 model.
[0035] Different from the prior art, the beneficial effects of the above technical solution are as follows:
[0036] A voice emotion recognition method that focuses on the sample feature distance throughout the entire model life cycle. By introducing a large-scale pre-trained model for feature extraction, it effectively solves the problem of inaccurate voice emotion representation under the condition of data volume limitation. By constructing a new loss function to guide fine-tuning, it improves the sample feature distance, making the distribution of different types of voice emotion representations in the feature space clearer and alleviating the problem of emotion boundary confusion that existed in the past. In the inference stage, through the idea of enhancing K-nearest neighbor retrieval, the improved sample distance is reused, and without any additional training, the recognition accuracy of the model is further improved, saving the computational cost and time cost required to improve the model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic flow diagram of a full-life cycle voice emotion recognition method that focuses on the sample feature distance. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] To describe in detail the technical content, structural features, achieved objectives and effects of the technical solution, the following is described in detail in conjunction with specific embodiments and with reference to the accompanying drawings.
[0039] As Figure 1 shown, it is a schematic flow diagram of a full-life cycle voice emotion recognition method that focuses on the sample feature distance. The voice emotion recognition method of this embodiment specifically includes the following steps:
[0040] Step S101, randomly augment the input training samples.
[0041] Randomly augment a group of N input sample instances. The augmentation methods include adding noise, changing the volume, adding reverberation, changing the pitch, and mixed augmentation. The audio label after augmentation is the same as the original audio. After augmentation, a total of 2N sample instances including the original training samples and the randomly augmented samples are obtained.
[0042] Step S102, introduce a model trained on a large-scale dataset as a pre-trained model.
[0043] Data-driven deep learning technologies require a large amount of data for training to obtain a large-scale model with stronger generalization ability and better robustness. Wav2vec2.0 is a self-supervised pre-trained model trained on a large-scale speech dataset with a total duration of 960 hours, which can construct relatively accurate speech representations. In the pre-training stage, adopting the idea of transfer learning, introducing wav2vec2.0 as a feature extractor to make up for the defects brought by the scarcity of voice emotion data and extract general and accurate voice feature representations.
[0044] Step S103, define positive and negative samples and calculate the supervised contrastive learning loss.
[0045] Use the pre-trained model introduced in step S102 to extract features from the sample instances obtained in step S101. For a set of N input sample instances {x k , y k}, k = 1, ……, N, x k represents the feature vector of an input audio, and y k is the label of this audio represented by one-hot encoding. A training batch size consists of 2N sample instances, denoted as {x l , y l}, l = 1, ……, 2N, where x 2t(t=1,...,N) represents the original audio vector x k , x 2t-1 represents the randomly augmented version of x k(k=1,...,N) . After augmentation, the audio label is the same as the original audio, which can be expressed as y 2t = y 2t-1 = y k . Sample instances with the same label y are called positive samples, while sample instances with different labels are called negative samples. Calculate the supervised contrastive learning loss L scl as follows:
[0046]
[0047] where i ∈ I = {1, ……, 2N} represents the index of an instance, N is the number of samples, A(i) represents all indices except i, P(i) represents the indices of all positive samples with the same label as sample i, a ∈ A(i) represents a specific sample index except i, p ∈ P(i) represents a specific positive sample index with the same label as sample i; τ is the hyperparameter for calculating the supervised contrastive learning loss; x i , x p , x a represent the feature vectors of the audio samples corresponding to the subscripts respectively.
[0048] Step S104, calculate the cross-entropy loss, and sum it with the supervised contrastive learning loss calculated in step S103 to guide the model fine-tuning.
[0049] Calculate the cross-entropy loss L ce as follows through the N non-augmented original audio feature vectors extracted in step S103:
[0050]
[0051] where, N represents the number of samples, C represents the number of categories, y i represents the audio sample label, The probability result predicted by the model that the i-th sample belongs to the c-th category.
[0052] The supervised contrastive learning loss L scl and the cross entropy loss L ce Perform weighted summation to obtain the final loss L of the model as follows:
[0053] L=(1-μ)L ce +μL scl
[0054] Among them, μ is a hyperparameter that balances the cross entropy loss and contrastive learning loss.
[0055] By designing, calculating and minimizing the above loss function, the cross entropy loss based on supervised learning loss is used to fine-tune the model, which achieves the effect of shortening the distance between samples of the same type and increasing the distance between samples of different types, helping to alleviate the problem of fuzzy boundaries between samples of different categories.
[0056] Step S105, using the model fine-tuned in step S104, obtain the representation-label key-value pairs of the training samples and construct a data storage set.
[0057] Using the model fine-tuned in step S104, forward propagation is performed on all training set sample data. Based on the representation vectors and labels of the samples, a data storage set containing all training set sample data and validation set sample data is created. The storage format is as follows:
[0058] (K, V) = {(x i ,y i ), i∈D}
[0059] Where D is the set of all sample indexes of the training set and the validation set, x i represents the feature vector of the i-th audio sample calculated by the model in step S104, y i is the label corresponding to the i-th audio sample.
[0060] Step S106: given a test sample, retrieve K samples that are the nearest neighbors to the test sample from the data storage set obtained in step S105, and record their label distribution.
[0061] When a test sample is given, the Euclidean distance between all samples in the data storage set in step S105 and the test sample is calculated according to the feature vector of the sample, the K samples closest to the test sample are retrieved and the distribution of each category label in them is recorded, which is recorded as p knn (y|x).
[0062] Step S107: For the test samples given in step S106, use the model in step S104 to predict their output distribution.
[0063] For the test samples given in step S106, use the model fine-tuned in step S104 to perform inference on them and predict the output distribution, denoted as p model (y|x).
[0064] Step S108: Weightedly sum the distributions obtained in step S106 and step S107 to obtain the final predicted label of the test sample.
[0065] Integrate the retrieval results from the data storage set in step S106 and the model inference results in step S107, and perform a weighted sum on them to obtain the final predicted distribution p(y|x) of the test sample as follows:
[0066] p(y|x) = αp knn (y|x) + (1 - α)p model (y|x)
[0067] where α is a hyperparameter for adjusting the proportion of p knn (y|x) and p model (y|x).
[0068] A full-life-cycle speech emotion recognition method focusing on sample feature spacing improves and utilizes the sample spacing throughout the full cycle of speech emotion recognition through the interaction of supervised contrast learning and retrieval enhancement.
[0069] Supervised contrast learning can effectively improve the intra-class and inter-class sample spacings, widen the sample spacings between different classes, and narrow the sample spacings within the same class, making the distributions of speech emotion features of each class in the sample space clearer. In the improved feature space, in the inference stage, the KNN algorithm based on sample spacing calculation is further used to implement the retrieval enhancement strategy, which can improve the recognition performance of the model without any additional training. In addition, in the feature space improved by supervised contrast learning, the ideas of supervised contrast learning and retrieval enhancement based on the KNN algorithm can have a significant effect on the improvement and utilization of sample spacing and the improvement of model performance. Compared with previous speech emotion recognition algorithms, on the IEMOCAP dataset, the algorithm proposed in the present invention has achieved better results in two evaluation metrics, Weighted Accuracy (WA) and Unweighted Accuracy (UA), as shown in the following table:
[0070]
[0071]
[0072] Among the currently known speech emotion recognition algorithms, the present invention first introduces the idea of retrieval enhancement, and together with the pre-trained model and supervised contrastive learning, constitutes a full-life-cycle speech emotion recognition method focusing on the feature distance of samples.
[0073] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, the elements defined by the statement "comprising..." or "including..." do not exclude the presence of additional elements in the process, method, article or terminal device comprising the said elements. In addition, in this article, "greater than", "less than", "more than" are understood not to include the present number; "above", "below", "within" are understood to include the present number.
[0074] Although the above embodiments have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the above are only the embodiments of the present invention, and do not limit the patent protection scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.
Claims
1. A full-life-cycle voice emotion recognition method focusing on the pitch interval of sample features, characterized in that It includes the following steps: Step S101: Randomly augment the input training samples; Step S102: Introduce the model trained on a large-scale dataset as the pre-trained model; Step S103: Use the pre-trained model introduced in Step S102 to extract features from the sample instances obtained in Step S101, define positive and negative samples, and calculate the supervised contrastive learning loss; Step S104: Calculate the cross-entropy loss, perform weighted summation with the supervised contrastive learning loss calculated in Step S103, and fine-tune the pre-training of the model; Step S105: Use the model fine-tuned in Step S104 to obtain the representation-label key-value pairs of the training samples, and construct a data storage set; Step S106: Given a test sample, retrieve the K nearest neighbor samples to the test sample in the data storage set obtained in Step S105, and record their label distribution; Step S107: For the test sample given in Step S106, use the model in Step S104 to predict its output distribution; Step S108: Perform weighted summation on the distributions obtained in Step S106 and Step S107 to obtain the final predicted label of the test sample.
2. The full-life cycle voice emotion recognition method for focusing on the feature spacing of samples according to claim 1, characterized in that Calculate the supervised contrastive learning loss L in step 103 scl as follows: where \(i\in I = \{1,\ldots, 2N\}\) represents the index of an instance, \(N\) is the number of samples, \(A(i)\) represents all indices except \(i\), \(P(i)\) represents the indices of all positive samples with the same label as sample \(i\), \(a\in A(i)\) represents a specific sample index other than \(i\), \(p\in P(i)\) represents a specific index of a positive sample with the same label as sample \(i\); \(\tau\) is a hyperparameter for calculating the supervised contrastive learning loss; \(x\) i , \(x\) p , \(x\) a represent the feature vectors of the audio samples corresponding to the subscripts respectively.
3. The full-life-cycle voice emotion recognition method for focusing on the feature spacing of samples according to claim 2, wherein, The cross-entropy loss L calculated in step 104 is as follows: ce as follows: Among them, N represents the number of samples, C represents the number of categories, and y i represents the audio sample label, is the probability result that the i-th sample predicted by the model belongs to the c-th category.
4. The full-life cycle voice emotion recognition method for focusing on the feature spacing of samples according to claim 3, characterized in that, The said step 104 performs a weighted sum of the supervised contrastive learning loss L scl and the cross-entropy loss L ce to obtain the final loss L of the model as follows: L = (1 - μ)L ce + μL scl Among them, μ is the hyperparameter that balances the cross-entropy loss and the contrastive learning loss.
5. The full-life-cycle voice emotion recognition method for focusing on the feature spacing of a sample as described in claim 1, wherein The said Step 105 includes: Using the model fine-tuned in Step S104 to perform a forward propagation on all training set sample data, and creating a data storage set containing all training set sample data and validation set sample data according to the representation vectors and labels of the samples. The storage format is as follows: (K, V) = {(x i , y i ), i ∈ D} Among them, D is the set of all sample indices of the training set and the validation set, and x i represents the feature vector obtained by calculating the model for the i-th audio sample in step S104, and y i is the label corresponding to the i-th audio sample.
6. The full life cycle voice emotion recognition method for focusing on the feature spacing of samples according to claim 1, wherein The said Step 108 includes: Synthesize the retrieval results from the data storage set in Step S106 and the model inference results in Step S107, perform weighted summation on them, and obtain the final predicted distribution p(y|x) of the test sample as follows: p(y|x) = αp knn (y|x) + (1 - α)p model (y|x) where α is a hyperparameter for adjusting p knn (y|x) and p model (y|x) ratio, p knn (y|x) is to retrieve the K samples nearest to the test sample in step S106 and record the distribution of each class label among them, p model (y|x) is to perform inference on it using the model fine-tuned in step S104 in step S107 and predict the output distribution.
7. The full-life-cycle voice emotion recognition method for focusing on the pitch interval of sample features according to claim 1, wherein The pre-trained model is the Wav2vec2.0 model.
Citation Information
Patent Citations
Speech classification network training method and device, computing equipment and storage medium
CN113593611A
Rolling bearing unknown fault detection method based on cross-domain relevance representation
CN116150635A