A Speech Emotion Recognition Method and Computer Device Based on Self-Supervised Learning
Through a self-supervised learning method, the speech self-supervised learning model is trained using the labelless speech data set, and on this basis, a small amount of emotional labeled data is used for fine-tuning, which solves the dependence problem of large-scale and high-quality labeled data sets in the existing technology, and achieves high accuracy and generalized speech emotion recognition effects.
Patent Information
- Application Number
- CN202210538988.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-18
AI Technical Summary
Existing speech emotion recognition methods rely highly on large-scale, high-quality annotated data sets, resulting in low recognition accuracy and poor generalization.
A self-supervised learning method is adopted to train the voice self-supervised learning model through an unlabeled speech sample set to obtain common speech features, and on this basis, a small amount of emotional label data is used for fine-tuning to build a speech emotion recognition model.
It improves the accuracy and generalization performance of speech emotion recognition, reduces dependence on large-scale, high-quality annotated data sets, and reduces labor and time costs.
Smart Images

Figure CN114937465B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and particularly to a speech emotion recognition method and a computer device based on self-supervised learning. Background Art
[0002] Human speech contains rich content. In addition to language content, it also contains its own emotional information. To accurately understand the intention of the speaker, deeply analyzing the emotional information in speech is an effective means. Speech Emotion Recognition (SER) helps to deeply understand the true intention of the user by recognizing the emotional information in speech. The speech emotion recognition technology has been widely applied in the fields of security, education, finance, etc.
[0003] In the prior art, the training method of the speech emotion recognition method determines that it is necessary to learn speech features and emotional rules from a large number of labeled samples. If the number of labeled samples is insufficient or the quality is low, only incomplete or incorrect speech features and emotional categories can be learned, resulting in unsatisfactory recognition effects, low accuracy, and poor generalization. However, it is difficult to obtain a large-scale and high-quality speech emotion recognition labeled dataset, and the manual annotation cost is very high. Therefore, the current speech emotion recognition method has unsatisfactory effects. Summary of the Invention
[0004] In view of the above analysis, the present invention aims to provide a speech emotion recognition method and a computer device based on self-supervised learning, which solves the problems that the speech emotion recognition method in the prior art highly depends on the scale and quality of the labeled training dataset, and has low recognition accuracy and poor generalization.
[0005] The object of the present invention is mainly achieved by the following technical solutions:
[0006] On the one hand, the present invention discloses a speech emotion recognition method based on self-supervised learning, including the following steps:
[0007] Training a speech self-supervised learning model based on an unlabeled speech sample set; the speech self-supervised learning model is used to output general speech features corresponding to the unlabeled speech samples;
[0008] Based on a training sample set containing speech emotion labels, constructing and training a speech emotion recognition model including the speech self-supervised learning model;
[0009] Inputting the speech to be recognized for emotion into the speech emotion recognition model, and using the speech emotion recognition model to recognize the corresponding emotion category.
[0010] Further, the speech self-supervised learning model includes a feature encoder, a quantization module, a masking module, and a context network;
[0011] The feature encoder is used to obtain the hidden-layer speech representation of the unlabeled speech sample according to the input unlabeled speech sample;
[0012] The quantization module is used to obtain the quantized hidden-layer speech representation through product quantization according to the hidden-layer speech representation;
[0013] The masking module is used to perform random time-step masking on the hidden-layer speech representation obtained by the feature encoder to obtain a masking result;
[0014] The context network is used to obtain the overall sequence representation of the unlabeled speech sample including the sequence representation of each time step according to the masking result by using the self-attention mechanism;
[0015] When training the speech self-supervised learning model, loss iterative update is performed based on the quantized hidden-layer speech representation and the overall sequence representation.
[0016] Further, the loss iterative update based on the quantized hidden-layer speech representation and the overall sequence representation includes:
[0017] Construct a quantized candidate representation set including interference terms and the quantized hidden-layer speech representation;
[0018] According to the sequence representation c at the masking time step t t , based on the quantized candidate representation set, predict the quantized hidden-layer speech representation q corresponding to the time step t t ;
[0019] Based on the comparison error between the sequence representation c t and the quantized candidate representations in the quantized candidate representation set, perform loss iterative update to obtain the speech self-supervised learning model.
[0020] Further, there are k interference terms, and the k interference terms are the sequence representations corresponding to k time steps uniformly sampled from the time steps other than the time step t in the currently input unlabeled speech; where k is an integer greater than 1.
[0021] Further, the comparison error is expressed as:
[0022]
[0023] where sim(a, b)=a T b / ‖a‖‖b‖ represents the cosine similarity between the context representation and the quantized hidden-layer speech representation, a represents c t , b represents q t ; is the quantized candidate representation set; ct The sequence representation corresponding to time step t output for the context network; q t is the quantized hidden layer speech representation at time step t.
[0024] Further, obtaining the quantized hidden layer speech representation through product quantization includes: dividing the hidden layer speech representation of each time step output by the feature encoder into n groups of sub-vectors, where n is an integer greater than 1, clustering each group of sub-vectors to obtain n codebooks; randomly selecting a center point from each of the n codebooks at time step t and concatenating them to obtain the quantized hidden layer speech representation q of time step t t .
[0025] Further, randomly masking the hidden layer speech representation includes: randomly selecting time step t as the starting index, replacing the starting index and the subsequent M consecutive time steps with silence, where M is an integer greater than 1, to obtain the masking result of time step t; randomly masking the hidden layer speech representation according to a preset ratio to obtain the masking result of the unlabeled speech input.
[0026] Further, the speech emotion recognition model further includes a softmax layer, which is used to receive the general speech features output by the speech self-supervised learning model, perform an emotion multi-classification task, and output the emotion category corresponding to the speech to be recognized.
[0027] Further, the training sample set containing labeled tags uses the RAVDESS dataset, which includes seven emotion types: calm, happy, sad, angry, fearful, disgusted, and surprised.
[0028] On the other hand, the present invention also provides a computer device, including at least one processor and at least one memory communicatively connected to the processor;
[0029] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the foregoing speech emotion recognition method.
[0030] The present invention can at least achieve one of the following beneficial effects:
[0031] 1. The present invention adopts a two-stage training method. The model obtained by the first stage of training based on the unlabeled speech sample set is used as the initial model of the speech emotion recognition model to obtain general speech features. When performing the second stage of training based on the general speech features, only a small amount of data with emotion labels is used for fine-tuning, so that the emotion recognition effect of this method is better and the training time is short.
[0032] 2. The present invention uses a small amount of emotion-labeled data as training samples, reducing the labor cost and time cost required to obtain a large amount of labeled data;
[0033] 3. The present invention introduces self-supervised learning technology. First, it trains and learns general speech features through a large-scale unlabeled dataset, and then uses a small amount of speech emotion annotation data for fine-tuning to implement a speech emotion recognition method, improving the generalization ability of speech recognition and solving the problem that traditional speech emotion recognition methods highly rely on a large-scale and high-quality speech emotion recognition labeled dataset. The speech emotion recognition method of the present invention has better emotion recognition accuracy and generalization than traditional speech emotion recognition methods when using one percent of the data volume of the original speech emotion recognition method.
[0034] Other features and advantages of the present invention will be described in the following specification, and some will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings
[0035] The drawings are only for the purpose of showing specific embodiments and are not considered as limiting the present invention. Throughout the drawings, the same reference signs represent the same components.
[0036] Figure 1 It is a flowchart of the speech emotion recognition method based on self-supervised learning according to an embodiment of the present invention.
[0037] Figure 2 It is a structural diagram of the speech self-supervised learning model according to an embodiment of the present invention. Detailed Embodiments
[0038] The following will specifically describe the preferred embodiments of the present invention with reference to the drawings. The drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, not to limit the scope of the present invention.
[0039] A speech emotion recognition method based on self-supervised learning in this embodiment, as Figure 1 shown, includes the following steps:
[0040] Step S1: Train a speech self-supervised learning model based on an unlabeled speech sample set; the speech self-supervised learning model is used to output the general speech features corresponding to the unlabeled speech samples; among them, the unlabeled speech sample set includes pure human voice speeches in different forms such as broadcasts and conversations of different genders and different languages; the general speech features are speech features automatically obtained by the neural network from the input speech according to its own structure, which can include features such as sound intensity, loudness, pitch, short-time zero-crossing rate, fundamental frequency, and energy, and can be used for speech processing tasks such as speech emotion recognition, language recognition, and speech transcription.
[0041] Specifically, as Figure 2As shown in the figure, the voice self-supervised learning model includes a feature encoder, a quantization module, a masking module, and a context network; among them,
[0042] The feature encoder is used to obtain the hidden-layer voice representation of the unlabeled voice sample according to the input unlabeled voice sample; the hidden-layer voice representation is the voice feature obtained by the feature encoder from the input voice.
[0043] Preferably, the feature encoder adopts a 2-layer CNN structure, the convolutional kernel is 5*3, and the stride is 2; the unlabeled voices in the sample set are divided into multiple voice segments in units of a preset time interval to obtain an unlabeled voice sequence X, and X is input into the feature encoder, and the feature encoder automatically obtains the hidden-layer voice representation Z={z 1 ,z 2 ,…,z t ,…,z T} corresponding to the unlabeled voice sample from X, where z t is the hidden-layer voice representation at time step t, which is a 512-dimensional vector; t = 1, …… T; T = unlabeled voice length / preset time interval. Exemplarily, the preset time interval can be 20ms.
[0044] The quantization module is used to obtain the quantized hidden-layer voice representation through product quantization according to the hidden-layer voice representation output by the feature encoder; specifically, after the hidden-layer voice representation Z corresponding to the unlabeled voice sample output by the feature encoder is input into the quantization module, the quantization module divides the hidden-layer voice representation of each time step output by the feature encoder into n groups of sub-vectors, n is an integer greater than 1, clusters each group of sub-vectors to obtain n codebooks; randomly select a center point from each of the n codebooks at time step t and splice them to obtain the quantized hidden-layer voice representation q t ; preferably, in this embodiment, first, the voice representation of each time step in the hidden-layer voice representation Z is evenly divided into 4 groups of sub-vectors, each group of sub-vectors is 128-dimensional, and each group of sub-vectors is clustered into 256 classes using the kmeans method, that is, 256 center points are obtained, then each group of sub-vectors constitutes a codebook, and 4 codebooks are obtained. During the training process, randomly select a center point from the 4 codebooks corresponding to the hidden-layer voice representation z t at time step t and splice them to obtain the quantized voice representation q t .
[0045] The masking module is used to perform random time-step masking on the hidden-layer speech representation obtained by the feature encoder to obtain a masking result. In this embodiment, a time step t is randomly selected as the starting index, and the starting index and the subsequent M consecutive time steps are replaced with silence to obtain the masking result at time step t, where M is an integer greater than 1; the hidden-layer speech representation is randomly masked according to a preset ratio p to obtain the masking result of the unlabeled speech input. Preferably, the masked parts of the two maskings can overlap, and the value of p is between 0.06 and 0.07.
[0046] The context network is used to obtain the overall sequence representation of the unlabeled speech sample including the sequence representation of each time step by using the self-attention mechanism according to the masking result; specifically, the masking result output by the masking module is used as the input of the context network; the context network adopts the native Transformer structure and uses the self-attention mechanism to obtain the overall sequence representation C = {c 1 , c 2 , … c t , …, c T} of the unlabeled speech, where c t is the sequence representation of the unlabeled speech at time step t.
[0047] When training the speech self-supervised learning model, loss iteration update is performed based on the quantized hidden-layer speech representation and the overall sequence representation. Specifically, for the sequence representation c t at time step t output by the context network, the speech self-supervised learning model needs to predict the true quantized hidden-layer speech representation q t in a set of quantized candidate representation sets including q t and interference terms. First, a quantized candidate representation set including interference terms and the quantized hidden-layer speech representation q t is constructed. Preferably, the number of interference terms is k, and the k interference terms are the sequence representations corresponding to k time steps uniformly sampled from the time steps other than time step t in the currently input unlabeled speech; where k is an integer greater than 1. The loss of the model is the contrast error L, as shown in the following formula:
[0048]
[0049] where, sim(a, b) = a T b / ‖a‖‖b‖ represents the cosine similarity between the context representation and the quantized hidden-layer speech representation, a represents c t , b represents q t ; is the set of quantized candidate representations; c t is the sequence representation corresponding to time step t output by the context network; qt It is the quantized hidden layer speech representation corresponding to the time step t of the model output.
[0050] During the training process, the set of unlabeled speech samples is input into the model, and the contrast error L is gradually reduced using the Adam optimization method to obtain a converged speech self-supervised learning model, that is, the general speech features are obtained.
[0051] It should be noted that self-supervised learning is a method that automatically generates labels for data and learns domain-general features on the labels. The self-supervised learning method automatically generates labels for data through specific auxiliary tasks, and generates domain-general features through the automatic annotation and training of a large amount of data; the present invention introduces self-supervised learning technology, which greatly reduces the dependence of the speech emotion recognition method on labeled data. By using self-supervised learning technology to learn general speech features on a large-scale unlabeled dataset, and then using a small amount of data with emotion labels for training, a speech emotion recognition model with high accuracy and high generalization ability can be obtained.
[0052] Step S2: Based on the training sample set containing speech emotion labels, construct and train a speech emotion recognition model including the speech self-supervised learning model.
[0053] Specifically, the speech emotion recognition task can be regarded as a multi-classification task. The softmax layer can be used to receive the output of the speech self-supervised learning model and generate an N-dimensional vector. Each emotion category in the multi-classification task corresponds to a vector. At the same time, the softmax layer normalizes the value of each vector and converts it into probabilities for N emotion categories. The category with the highest probability is the emotion type corresponding to the current input speech.
[0054] Preferably, in this embodiment, the RAVDESS dataset containing annotation labels is used as the training sample. The RAVDESS dataset contains 7356 speeches, including seven emotion types: calm, happy, sad, angry, fearful, disgusted, and surprised. On the basis of the speech self-supervised learning model, a softmax layer is added to receive the general speech features output by the speech self-supervised learning model and perform emotion multi-classification tasks. The CTC loss is used as the model loss function for gradient update to obtain a converged speech emotion recognition model.
[0055] Based on the pre-trained speech self-supervised learning model, the present invention uses a small amount of data with emotion labels for simple fitting, that is, the speech emotion recognition method is realized, reducing the human cost and time cost consumed in obtaining a large amount of labeled data.
[0056] Step S3: Input the speech to be sentiment-recognized into the speech sentiment recognition model, and use the speech sentiment recognition model to recognize the corresponding sentiment category. Specifically, input the unannotated speech to be recognized into the trained speech sentiment recognition model, and the model automatically generates the sentiment type to which the input speech belongs according to the features of the input speech.
[0057] In summary, a speech sentiment recognition method based on self-supervised learning proposed by the present invention introduces self-supervised learning technology, uses an unannotated speech dataset for training, and learns general speech features on a large-scale unannotated speech dataset; then fine-tunes on the general speech features based on a small-scale sentiment-annotated data to achieve a speech sentiment recognition method with high accuracy and generalization; experiments show that the technical solution of the present invention is superior to traditional speech sentiment recognition methods.
[0058] The main process of speech sentiment recognition in the prior art is as follows: regard speech sentiment recognition as a classification problem, prepare a large number of speeches with sentiment type labels as the training dataset; construct a specific neural network structure according to the domain characteristics, use the training dataset with sentiment type labels for training to obtain a speech sentiment recognition model; input the speech without sentiment type labels into the speech sentiment recognition model to obtain the sentiment type corresponding to the speech, and achieve speech sentiment recognition. The present invention optimizes the existing speech sentiment recognition method using self-supervised learning, so that the speech sentiment recognition method can also achieve an ideal recognition effect in the case of the lack of current large-scale and high-quality annotated datasets. The present invention adopts a two-stage training method. First, use the model trained based on the unannotated speech sample set as the initial model to obtain general speech features; secondly, perform a second training on the general speech features using a small amount of datasets with sentiment labels, which solves the dependence problem of traditional speech sentiment recognition methods on large-scale and high-quality sentiment-annotated datasets, uses a small amount of training samples, and improves the accuracy and generalization performance of sentiment recognition.
[0059] Another embodiment of the present invention provides a computer device, including at least one processor and at least one memory communicatively connected to the processor; the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the speech sentiment recognition method based on self-supervised learning in the foregoing embodiment.
[0060] Those skilled in the art can understand that all or part of the processes of implementing the method in the above embodiment can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory or a random access memory, etc.
[0061] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for speech emotion recognition based on self-supervised learning, characterized in that, it includes the following steps: Training a speech self-supervised learning model based on an unlabeled speech sample set; the speech self-supervised learning model is used to output general speech features corresponding to the unlabeled speech samples; the speech self-supervised learning model includes a feature encoder, a quantization module, a masking module, and a context network; the feature encoder is used to obtain the hidden layer speech representation of the unlabeled speech sample according to the input unlabeled speech sample; the quantization module is used to obtain a quantized hidden layer speech representation by product quantization according to the hidden layer speech representation; the masking module is used to perform random time step masking on the hidden layer speech representation obtained by the feature encoder to obtain a masking result; the context network is used to obtain the overall sequence representation of the unlabeled speech sample including the sequence representation of each time step according to the masking result by using the self-attention mechanism; When training the speech self-supervised learning model, loss iterative update is performed based on the quantized hidden layer speech representation and the overall sequence representation, including: constructing a set of quantized candidate representations including interference terms and the quantized hidden layer speech representation; according to the sequence representation c of the masking time step t t , predicting the quantized hidden layer speech representation q corresponding to the time step t based on the set of quantized candidate representations t ; performing loss iterative update based on the comparison error between the sequence representation c t and the quantized candidate representations in the set of quantized candidate representations to obtain the speech self-supervised learning model; Based on a training sample set containing speech emotion labels, constructing and training a speech emotion recognition model including the speech self-supervised learning model; Inputting the speech to be emotion-recognized into the speech emotion recognition model, and using the speech emotion recognition model to recognize the corresponding emotion category.
2. The speech emotion recognition method according to claim 1, characterized in that, There are k interference items, and the k interference items are the sequence representations corresponding to k time steps uniformly sampled from the time steps other than time step t in the currently input unlabeled speech; where k is an integer greater than 1.
3. The speech emotion recognition method according to claim 2, characterized in that, The contrast error is expressed as: where sim(a, b) = a T b / ‖a‖‖b‖ represents the cosine similarity between the context representation and the quantized hidden-layer speech representation, where a represents c t and b represents q t ; is the set of quantization candidate representations; c t is the sequence representation corresponding to time step t output by the context network; q t is the quantized hidden-layer speech representation at time step t.
4. The speech emotion recognition method according to claim 1, characterized in that, The quantization of the hidden layer speech representation obtained by product quantization includes: dividing the hidden layer speech representation of each time step output by the feature encoder into n groups of sub-vectors, where n is an integer greater than 1, clustering each group of sub-vectors to obtain n codebooks; randomly selecting a center point from each of the n codebooks at time step t and concatenating them to obtain the quantization hidden layer speech representation q of time step t t .
5. The speech emotion recognition method according to claim 1, characterized in that, The random masking of the hidden layer speech representation includes: randomly selecting time step t as the starting index, and replacing the starting index and the subsequent M consecutive time steps with silence, where M is an integer greater than 1, to obtain the masking result of time step t; randomly masking the hidden layer speech representation according to a preset ratio to obtain the masking result of the input unlabeled speech.
6. The speech emotion recognition method according to claim 1, characterized in that, The speech emotion recognition model further includes a softmax layer, which is used to receive the general speech features output by the speech self-supervised learning model, perform an emotion multi-classification task, and output the emotion category corresponding to the speech to be recognized.
7. The speech emotion recognition method according to claim 1, characterized in that, The training sample set containing labeled labels uses the RAVDESS dataset, which includes seven emotion types: calm, happy, sad, angry, fearful, disgusted, and surprised.
8. A computer device, characterized in that, It includes at least one processor and at least one memory communicatively connected to the processor; The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the voice emotion recognition method according to any one of claims 1-7.
Citation Information
Patent Citations
Speech emotion recognition method and system based on semi-supervised adversarial variation self-coding
CN112863494A
Speech classification network training method and device, computing equipment and storage medium
CN113593611A