A Speech Emotion Recognition Method Based on Attention Mechanism and Multi-Task Learning
By adopting the LSTM_att-MTL model based on attention mechanism and multi-task learning in speech emotion recognition, the problems of high complexity of feature extraction calculation and poor training effect in the prior art are solved, and more efficient speech emotion recognition performance is achieved.
Patent Information
- Application Number
- CN202210546156.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-05-19
AI Technical Summary
The existing speech emotion recognition methods have high computational complexity during feature extraction, poor training results, resulting in low recognition performance.
The LSTM_att-MTL model based on attention mechanism and multi-task learning is adopted to improve the recognition performance of speech emotion recognition by preprocessing and feature extraction of speech signals, combining attention mechanism and multi-task learning.
It improves the recognition performance of speech emotion recognition, enhances the ability to extract features, improves the effect of the training process, and improves the recognition accuracy of multiple emotions.
Smart Images

Figure CN114927144B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning, and particularly relates to a speech emotion recognition method based on an attention mechanism and multi-task learning. Background Art
[0002] Speech emotion recognition is an advanced computer technology involving multiple disciplines, and its purpose is to extract emotion feature parameters from speech signals and identify speech emotions through these parameters. Speech emotion recognition has developed from a small-scale development stage into a key technology for human-computer interaction. Speech emotion recognition has been widely applied in various industries. For example, it is introduced into vehicle driving systems to record the mental state changes of drivers, increasing the interaction between people and cars; medical devices are added with speech emotion recognition systems, which can make better treatments according to the emotional changes of patients; it is introduced into the education system, and the emotional changes of students can be detected in remote teaching, thereby improving the teaching quality. The initial speech emotion recognition methods were mainly machine learning algorithms and recognition models such as HMM, which could generally only recognize a few types of emotions. The emergence of deep learning technologies such as Convolutional Neural Network (CNN) has enabled the rapid development of speech emotion recognition based on deep learning. For example, a speech emotion recognition method combining spectrograms and deep convolutional neural networks uses a network composed of convolutional layers and fully connected layers to extract features from spectrograms and effectively recognize multiple emotions.
[0003] After preprocessing the original speech for speech emotion recognition, a feature extraction method is generally used to extract features from the original speech, with the aim of identifying the emotion type to which the feature belongs. One solution is to extract prosodic features and voice quality features. Usually, energy, fundamental frequency curves, and logarithmic curves are extracted from emotional sentences, the corresponding first-order differences and second-order difference curves are calculated, and finally parameters such as skewness, kurtosis, and variance are statistically calculated. The emotion discrimination ability of the feature region of this solution is limited. For example, the recognition accuracy for emotions such as anger, fear, and happiness is not ideal. Extracting features from spectrograms is one of the important technologies for speech emotion recognition at present. It mainly normalizes the spectrogram to grayscale from both the time domain and the frequency domain to extract higher-level features, and can also be used as a visual representation of speech signals.
[0004] Most of the existing methods for extracting features from spectrograms use Long-Short Term Memory (LSTM) to solve the problem of long-term dependence in time series, but often result in low recognition rates of existing speech emotion classifiers due to high computational complexity and insufficient training data. Therefore, this patent proposes a speech emotion recognition method based on an attention mechanism and multi-task learning to improve the recognition performance of the algorithm. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a speech emotion recognition method based on an attention mechanism and multi-task learning in view of the deficiencies of the background technology. That is, an LSTM_att-MTL model is constructed, which solves the problems of high computational complexity of traditional feature extraction methods, poor training process effects, and resulting reduction in recognition performance.
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] Compared with the prior art, the present invention adopting the above technical solutions has the following technical effects:
[0008] The present invention designs a speech emotion recognition method based on an attention mechanism and multi-task learning, namely the LSTM_att-MTL model. This model solves the problems of high computational complexity of traditional feature extraction methods, poor training process effects, and resulting reduction in recognition performance. In order to extract more advanced features, first, the original speech at each moment is preprocessed to obtain a grayscale map representation of the speech, which is used as the input of a two-layer CNN network, and then speech feature extraction is performed in the CNN layer. Secondly, since the obtained speech features are correlated in the time series, LSTM_att is used to perform temporal modeling on the grayscale map features learned by the CNN network, and the LSTM_att layer is a parameter-sharing layer. Finally, since the memory of LSTM_att gradually weakens with the length of the speech, and as the speech sequence grows, the information of the starting time node has less and less influence on the current moment, an attention layer is added in the multi-task layer. The results of simulation experiments on the CASIA dataset show that this model has good speech emotion recognition performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is the overall framework diagram of a speech emotion recognition method based on an attention mechanism and multi-task learning according to the present invention;
[0010] Figure 2 is the schematic diagram of the LSTM structure with an attention gate in the present invention;
[0011] Figure 3 is the schematic diagram of the value of the weight α of LSTM_att-MTL and LSTM-MTL;
[0012] Figure 4 is the schematic diagram of the comparison of experimental results of different numbers of LSTM_att layers. DETAILED DESCRIPTION OF THE INVENTION
[0013] The following further elaborates on the technical solutions of the present invention with reference to the accompanying drawings:
[0014] A speech emotion recognition method based on an attention mechanism and multi-task learning, asFigure 1 and 2 As shown in 2 , to solve the problems such as high computational complexity of traditional feature extraction methods, poor training process effect, and resulting reduction in recognition performance, by adding attention gates in the LSTM, the purpose of reducing training parameters is achieved, and an LSTM_att-MTL speech emotion recognition model is constructed together with multi-task learning to improve the recognition performance of the algorithm. The method includes the following steps:
[0015] Step 1: Obtain the speech emotion dataset: Obtain the CASIA Chinese emotion dataset for speech emotion recognition;
[0016] Step 2: Construct the LSTM_att-MTL speech emotion recognition model: The LSTM_att-MTL speech emotion recognition model consists of a feature extraction module, a sequence modeling module, and a multi-task learning module. Input the speech emotion data in Step 1 into the recognition model for collaborative training;
[0017] Step 3: Calculate the loss function of the model: Obtain the recognition result through the softmax classifier in Step 2, and calculate the loss function between the recognition result and the training set label to adjust the loss size;
[0018] Step 4: Train the overall model and obtain the recognition result: Input the speech emotion data of the test set into the network trained in Step 3 to achieve the recognition of the speech emotion data of the test set.
[0019] The obtaining of the speech emotion dataset described in Step 1 includes the following steps:
[0020] Step 11: First, the speech emotion dataset comes from the CASIA Chinese emotion corpus developed by the Institute of Automation, Chinese Academy of Sciences. The dataset contains two pairs of actors performing 500 sentences under six emotions of anger, fear, happy, neutral, sad, and surprise in a pure recording environment, with a sampling rate of 16 kHz, 16-bit quantization, a signal-to-noise ratio of about 35 dB, stored in pcm format, and finally 9,600 sentences are selected;
[0021] Step 12: Second, divide the dataset into a training set, a validation set, and a test set according to the ratio of 6:2:2.
[0022] The constructing of the LSTM_att-MTL speech emotion recognition model described in Step 2 includes the following steps:
[0023] Step 21: First, preprocess the original speech, including the following operations: frame and window the speech with a frame length of 25 ms and a frame shift of 10 ms; perform short-time Fourier transform to obtain the spectrogram of the speech signal; perform max-min normalization on the spectrogram and quantize the spectrogram into a grayscale image, then enter Step 2;
[0024] Step 22: Input the spectrogram into the CNN network. The convolutional layer learns speech features from the spectrogram through convolutional calculations and selects the ReLU function as the activation function, then enter Step 3. Specific parameters:
[0025]
[0026] Step 23: Refer to Figure 2 As shown, use two recurrent neural networks as the shared layer, use LSTM_att as the basic unit, and the number of hidden layer units is 128. To prevent overfitting, dropout is introduced during the training process, and the parameter is set to 0.5. The first layer outputs all time series to the next layer, and the second layer outputs the result of the last time step. Output of the attention gate of LSTM_att:
[0027] att t = σ(V att × tanh(W att × c t-1 )) (1)
[0028] Among them, V att , W att ∈ R n×n are the parameters to be trained and learned from the training data; c t-1 is the node state at the previous moment; σ(·) and tanh(·) are the logistic sigmoid and hyperbolic tangent activation functions respectively; specific outputs of each gate unit:
[0029] i t = σ(W i × [c t-1 , h t-1 , x t + b i ) (2)
[0030]
[0031] o t = σ(W o × [h t-1 , x t + b o ) (5)
[0032] h t = o t·tanh(c t ) (6)
[0033] where i t represents the input gate at time t, W i and b i represent the weight matrix and bias term in the input gate unit, C t-1 and j t-1 are the cell state and hidden layer output at the previous time step respectively, and x t represents the input at the current time step; represents the candidate value for updating the cell state at time t, W c and b c represent the weight matrix and bias term when updating the state; c t represents the cell state at time t, · represents the Hadamard product; o t represents the output gate at time t, W o and b o represent the weight matrix and bias term in the output gate unit; h t represents the output of the hidden layer at time t;
[0034] Step 24: Adopt a multi-task learning method with parameter hard sharing, and use speaker gender as an auxiliary task. Since the differences between male and female voices will affect the system performance of voice-related tasks, the performance of a gender-specific emotion recognition model is better than that of other non-gender-specific emotion recognition models.
[0035] The attention layer of multi-task learning performs attention weighting on the output of LSTM_att:
[0036]
[0037] ν = ∑ i α i h i (8)
[0038] In Equation (7), α i represents the attention weight, the vector μ is the attention parameter, μ = (θ 1 , θ 2 , …, θ T ), T is the frame length, {h 1 , h 2 , …, h T} is the output of the last layer of LSTM_att. Calculate the inner product of the attention parameter vector μ and h i as the score of the importance of each time frame, and normalize it. The normalized score is the weight of each frame containing key information. Equation (8) performs a dot product of the obtained weight and the output of LSTM_att, and the resulting weighted sum v is used as the feature vector for global update weights.
[0039] The loss function of the calculation model described in step 3 includes the following steps:
[0040] Step: Input the feature vector obtained from equation (8) into the fully connected layer. The fully connected layer can achieve independent optimization while learning shared features, and perform classification through softmax to obtain the predicted value:
[0041]
[0042] where W and b represent the weights and bias terms in the fully connected layer, v represents the feature vector obtained from equation (8), and softmax(·) is the softmax function.
[0043] Step 31: During the training process, compare the recognition result with the label of the training set and calculate the loss function, and optimize it using the ADAM optimizer. In this paper, the cross-entropy function is used as the loss function for each task:
[0044]
[0045] where represents the loss function for each task, task is the type of task: emotion recognition (emotion) and gender classification (gender), y i and are the true value of the training set label and the predicted value output by the model respectively, and N task is the total number of categories of task task.
[0046] Step 32: The final total loss function:
[0047] L total = αL emo + (1 - α)L gen (10)
[0048] where L emo and L gen are the loss functions of speech emotion recognition and gender recognition respectively, and α is the weight coefficient of the emotion recognition loss function. Through the joint L total perform backpropagation and gradient update on the entire network.
[0049] The implementation of recognizing the speech emotion data of the test set described in step 4 includes the following steps:
[0050] Step 41: To demonstrate the processing ability of LSTM_att for time-series data, an LSTM-MTL comparative experiment was designed. The number of LSTM layers and nodes was the same as that of the proposed model in this paper, and emotion and gender were recognized. The values of the weight coefficients of the loss functions for the two tasks were determined. Since speech emotion recognition was the main task, its corresponding weight was greater than that of the gender recognition task. The initial value of α was set to 0.9 in the experiment. By adjusting the weights and testing on the validation set, the weights with the highest recognition accuracy were found. As Figure 3 shown, the values of the weights α of LSTM_att-MTL and LSTM-MTL were 0.8 and 0.6, respectively.
[0051] Step 42: To demonstrate the performance improvement of multi-task learning compared to single-task learning, a single-task comparative experiment LSTM_att-STL was designed. Under the premise of keeping other parameters the same, the emotion task was recognized.
[0052] Tables 1 and 2 show the recognition results of each task and emotion on the CASIA dataset, respectively. It can be seen from Table 1 that there was a significant improvement in emotion recognition for LSTM_att-MTL. Since there were few speakers in the CASIA dataset, the gender recognition rate was very high. Therefore, using gender recognition as an auxiliary task had little impact on emotion recognition. It can be seen from Table 2 that the recognition accuracy of the proposed method in this paper reached 93.1%. Except for the accuracy of neutral and surprised emotions being lower than that of LSTM-MTL, the recognition accuracy of other emotions was higher than that of the comparative experiment. The confusion matrix of the recognition results of LSTM_att-MTL is shown in Table 3. In the table, the abscissa represents the predicted results of speech emotions, and the ordinate represents the corresponding true categories of emotions. It can be seen from the table that the recognition accuracy of the happy emotion was relatively low and was easily recognized as neutral and surprised. The recognition result of neutral was relatively accurate, reaching 96.9%.
[0053] Table 1 Recognition results of each task in the comparative experiment
[0054] Tab.2 The recognition results of each task in the experiment
[0055]
[0056] Table 2 Recognition results of each emotion in the comparative experiment
[0057] Tab.3 The results of each emotion recognition in the experiment
[0058]
[0059]
[0060] Table 3 Results of each emotion recognition in LSTM_att-MTL
[0061] Tab.4Each emotion recognition result in LSTM_att-MTL
[0062]
[0063] Step 43: To select the most suitable number of LSTM_att layers in LSTM_att-MTL, three experiments were set up at the same time, with LSTM layers of 2, 3, and 4 respectively. The experimental results are as follows Figure 4 The experimental results show that the increase in the number of LSTM layers will improve the efficiency of time domain feature extraction. As shown in Table 5, the recognition accuracy of 2-layer LSTM_att is 93.1%, the recognition accuracy of 3-layer LSTM_att is 92.8%, and the recognition accuracy of 4-layer LSTM_att is 92.1%. Therefore, 2-layer LSTM_att is selected for the experiment.
[0064] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in the relevant technical field, are also included in the patent protection scope of the present invention.
Claims
1. A speech emotion recognition method based on attention mechanism and multi-task learning, characterized in that: It includes the following steps: Step 1: Obtain a speech emotion dataset: Obtain the CASIA Chinese emotion dataset for speech emotion recognition; Step 2: Construct an LSTM_att-MTL speech emotion recognition model: The LSTM_att-MTL speech emotion recognition model consists of a feature extraction module, a sequence modeling module, and a multi-task learning module. Input the speech emotion data in Step 1 into the recognition model for collaborative training; Step 3: Calculate the loss function of the model: Obtain the recognition result through the softmax classifier in Step 2, and calculate the loss function between the recognition result and the training set label to adjust the loss size; Step 4: Train the model to obtain the recognition result: Input the speech emotion data of the test set into the model trained in Step 3 to achieve the recognition of the speech emotion data of the test set; The construction of the LSTM_att-MTL speech emotion recognition model described in Step 2 includes the following steps: Step 21: First, preprocess the original speech, including the following operations: Frame and window the speech, with a frame length of 25 ms and a frame shift of 10 ms; Perform short-time Fourier transform to obtain the spectrogram of the speech signal; Perform max-min normalization on the spectrogram, and quantize the spectrogram into a grayscale image, and enter Step 2; Step 22: Input the spectrogram into the CNN network. The convolutional layer learns speech features from the spectrogram through convolutional calculations, and selects the ReLU function as the activation function, and enters Step 3; Step 23: Use two recurrent neural networks as the shared layer, use LSTM_att as the basic unit, and the number of hidden layer units is 128; to prevent overfitting, dropout is introduced during the training process, and the parameter is set to 0.5; the first layer outputs all time series to the next layer, and the second layer outputs the result of the last time step; the output att of the attention gate of LSTM_att t : att t = σ(V att × tanh(W att × c t-1 )) (1) Among them, V att , W att ∈R n×n are parameters to be trained and learned from the training data; c t-1 is the node state at the previous moment; σ(·) and tanh(·) are the logistic sigmoid and hyperbolic tangent activation functions respectively; the specific outputs of each gate unit are as follows: i t = σ(W i × [c t-1 , h t-1 , x t + b i ) (2) o t = σ(W o × [h t-1 , x t + b o ) (5) h t = o t ·tanh(c t ) (6) where i t represents the input gate at time t, W i and b i represent the weight matrix and bias term in the input gate unit, C t-1 and h t-1 are the cell state and hidden layer output at the previous time step respectively, x t represents the input at the current time step; represents the candidate value for updating the cell state at time t, W c and b c represent the weight matrix and bias term when updating the state; c t represents the cell state at time t, · represents the Hadamard product; o t represents the output gate at time t, W o and b o represent the weight matrix and bias term in the output gate unit; h t represents the output of the hidden layer at time t; Step 24: Adopt a multi-task learning method with parameter hard sharing, and use the speaker's gender as an auxiliary task. Because the differences between male and female voices will affect the system performance of speech-related tasks, the performance of a gender-specific emotion recognition model is better than that of other non-gender-specific emotion recognition models; The attention layer in the multi-task learning module performs attention weighting on the output of LSTM_att: ν = Σ i α i h i (8) In formula (7), α i represents the attention weight, the vector μ represents the attention parameter, μ = (θ 1 , θ 2 , …, θ T ), T is the frame length, {h 1 , h 2 , …, h T} is the output of the last layer LSTM_att. Calculate the inner product of the attention parameter vector μ and h i as the score of the importance of each time frame, and normalize it. The normalized score is the weight of each frame containing key information; formula (8) multiplies the obtained weight α i by the LSTM_att output h i for dot product, and the obtained weighted sum v is used as the feature vector of the global update weight.
2. The speech emotion recognition method based on attention mechanism and multi-task learning according to claim 1, characterized in that: The obtaining of the speech emotion dataset described in Step 1 includes the following steps: Step 11: First, the speech emotion dataset is from the CASIA Chinese emotion corpus developed by the Institute of Automation, Chinese Academy of Sciences. The dataset contains two pairs of actors performing 500 sentences of text in six emotions: anger, fear, happy, neutral, sad, and surprise in a pure recording environment. The sampling rate is 16 kHz, the quantization is 16 bits, the signal-to-noise ratio is about 35 db, and it is stored in pcm format. Finally, 9600 sentences are selected from them; Step 12: Secondly, divide the dataset into a training set, a validation set, and a test set according to the ratio of 6:2:
2.
3. The speech emotion recognition method based on attention mechanism and multi-task learning according to claim 1, characterized in that: The calculation of the loss function of the model described in Step 3 includes the following steps: Step 31: Input the feature vector obtained from Equation (8) into the fully connected layer. The fully connected layer can achieve independent optimization while learning shared features, and perform classification through softmax to obtain the predicted value: where W and b represent the weights and bias terms in the fully connected layer, v represents the feature vector obtained from Equation (8), and softmax(·) is the softmax function; Step 32: During the training process, compare the recognition result with the label of the training set and calculate the loss function, and optimize it using the ADAM optimizer; in this paper, the cross-entropy function is used as the loss function for each task: Among them represents the loss function for each task, where task is the type of task: emotion recognition (emotion) and gender classification (gender), and y i and are the true value of the training set label and the predicted value output by the model respectively, and N task is the total number of categories of task task; Step 33: Finally, the total loss function: L total = αL emo + (1 - α)L gen (10) where L emo and L gen are the loss functions for speech emotion recognition and gender recognition respectively, and α is the weight coefficient of the emotion recognition loss function; by jointly using L total backpropagation and gradient update are performed on the entire network.
4. A speech emotion recognition method based on attention mechanism and multi-task learning according to claim 1, characterized in that: The implementation of the recognition of the speech emotion data of the test set in step 4 includes the following steps: Step 41: To reflect the processing ability of LSTM_att for time-series data, design an LSTM-MTL comparative experiment. The number of LSTM layers and nodes is the same as that of the model in this paper, and emotion and gender are recognized; determine the value range of the weight coefficients of the two task loss functions. Speech emotion recognition is the main task, so its corresponding weight is greater than that of the gender recognition task. The initial value of α in the experiment is set to 0.9; by adjusting the weight and testing on the validation set, find the weight with the highest recognition accuracy; the values of the weights α of LSTM_att-MTL and LSTM-MTL are 0.8 and 0.6 respectively; Step 42: To reflect the improvement in performance of multi-task learning compared to single-task learning, design a single-task comparative experiment LSTM_att-STL, and recognize the emotion task on the premise of keeping other parameters the same; Step 43: To select the most suitable number of LSTM_att layers in LSTM_att-MTL, three experiments are set on the validation set at the same time, and the number of LSTM layers is 2, 3, and 4 respectively; the increase in the number of LSTM layers will improve the efficiency of time-domain feature extraction. The recognition accuracy of 2-layer LSTM_att is 93.1%, the recognition accuracy of 3-layer LSTM_att is 92.8%, and the recognition accuracy of 4-layer LSTM_att is 92.1%. Therefore, 2-layer LSTM_att is selected for the experiment.
Citation Information
Patent Citations
Domain invariance-based small sample speech emotion recognition method
CN111402929A
Cross-corpus emotion recognition method based on transfer learning and attention mechanism
CN113065344A