A Cross-Corpus Sentiment Recognition Method Based on Transfer Learning and Attention Mechanism
By using a method based on transfer learning and attention mechanism in cross-corpus emotion recognition, the problem of poor knowledge transfer effect in cross-corpus emotion recognition is solved, and effective emotion recognition under small sample training is achieved.
Patent Information
- Application Number
- CN202110330443.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-03-24
AI Technical Summary
The prior art is difficult to effectively extract appropriate emotional features in cross-corpus emotion recognition, resulting in poor knowledge transfer effects, especially in small sample training.
Using a method based on transfer learning and attention mechanism, the emotional dependence and transmission situation in the context are extracted through a recurrent neural network, and the attention transfer module is used to transfer feature parameters from the source corpus to the target corpus, controlling the migration loss within a certain range to complete knowledge transfer.
The judgment of the emotional state of the speaker is effectively completed on the target corpus with a small amount of data, solving the problem of insufficient training in small samples and improving the accuracy of emotional recognition.
Smart Images

Figure CN113065344B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of transfer learning, sentiment computing, etc., and relates to a cross-corpus sentiment recognition method based on transfer learning and attention mechanism, which is used to solve the problem of insufficient training of small samples. Background Technique
[0002] Sentiment computing aims to build a harmonious human-computer environment by endowing computers with the ability to recognize, understand, express, and adapt to human emotions, and to make computers have more efficient and comprehensive intelligence. As an important branch of artificial intelligence, sentiment computing and analysis are not only indispensable in realizing machine intelligence, but also very important in fields such as public opinion monitoring, clinical psychological dynamics detection, and human-computer interaction.
[0003] In recent years, deep learning has performed excellently in speech processing, image classification, and other machine learning-related fields, including human emotion recognition and cognitive understanding. Many works are carried out on convolutional neural networks (CNNs), recurrent neural networks (RNNs), and their variant models, and certain progress has been made. Initial research mostly identified the emotional states of target persons through single modalities such as expressions or texts on a single corpus. However, with the continuous complexity of neural network structures, a large amount of labeled data is required for network training, and the high cost of data annotation restricts the actual operation of training to a certain extent. To solve this problem, some scholars have proposed the idea of transfer learning in recent years, that is, transferring labeled data or knowledge structures from related fields to implement or improve the target field or task. In addition, in practice, due to differences in acquisition environments and devices, dialogue scenarios and topics, etc., the emotional data in the training set and the test set often vary greatly. Therefore, cross-corpus sentiment recognition is closer to real life and application scenarios. However, the difficulty of cross-corpus sentiment recognition lies in how to extract appropriate emotional features and complete knowledge transfer by continuously reducing the feature differences between the source task and the target task.
[0004] "Multimodal Sentiment Recognition Method and System Based on Neural Network and Transfer Learning" (Patent No.: CN201710698379.1). This method trains a deep neural network based on large-scale data and obtains an audio feature extractor and a video feature extractor through transfer learning, and then extracts audio features and video features from multimodal emotional data to identify the probabilities of various speech emotional categories and various video emotional categories, and judges the final emotional category through the probability values.
[0005] "A Multimodal Speech Emotion Recognition Method Based on Enhanced Deep Residual Neural Network" (Patent No.: CN201811346114.6). This method extracts the feature expressions of videos (sequence data) and speech, including converting speech data into corresponding spectrogram expressions and encoding time-series data: uses a convolutional neural network to extract the emotion feature expressions of the original data for classification. The model accepts multiple inputs with unequal input dimensions, and a cross-convolution layer is proposed to fuse the data features of different modalities. The overall network structure used by the model is an enhanced deep residual neural network: after the model is initialized, a multi-classification model is trained using speech spectrograms, sequence video information, and corresponding emotion labels. After training, the unlabeled speech and videos are predicted to obtain the probability values of emotion prediction, and the maximum probability value is selected as the emotion category of the multimodal data.
[0006] "A Multimodal Depression Detection Method and System Based on Situation Awareness" (Patent No.: 201911198356.X). This method includes: constructing a training sample set, which includes topic information, spectrograms, and corresponding text information; using a convolutional neural network and combining multi-task learning to extract acoustic features from the spectrograms of the training sample set to obtain situation-aware acoustic features; using the training sample set and a Transformer model to process word embeddings to extract situation-aware text features; establishing an acoustic channel subsystem for depression detection for the situation-aware acoustic features, establishing a text channel subsystem for depression detection for the situation-aware text features, and fusing the outputs of the acoustic channel subsystem and the text channel subsystem to obtain depression classification information.
[0007] Considering that in an actual conversation scenario, the emotional state of the speaker's target statement is often affected by the context statements. When this invention selects features for transfer, in addition to traditional emotional features, it also extracts and transfers the dynamic changes related to emotions in the context. During the transfer process, an attention transfer mechanism is used to make the feature maps of the target task as similar as possible to those of the source task, thereby completing knowledge transfer. Summary of the Invention
[0008] Based on the above difficulties in cross-corpus sentiment recognition, the present invention proposes a cross-corpus sentiment recognition method based on transfer learning and attention mechanism. Through the method of the present invention, first, each single sentence in the entire conversation is encoded on the source corpus, and the encoded vector of the single sentence is fed into a recurrent neural network (RNN). The RNN is used to extract the sentiment dependence and transmission in the context, and the feature parameters such as encoding and context sentiment dependence are transferred to the training of the target corpus. By controlling the transfer loss within a certain range through training, knowledge transfer is completed. On the target corpus, operations of encoding-context feature parameter extraction-classification are performed with the help of the knowledge of transfer learning, and finally, the task of determining the speaker's sentiment state on the target corpus is completed.
[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0010] A cross-corpus sentiment recognition method based on transfer learning and attention mechanism, the specific steps are as follows:
[0011] S1: Divide the dialogue part in the source corpus into t sentences X = [x 1 , x 2 , …, x i …, x t , and select the text data of all speakers in the dialogue.
[0012] S2: Use an encoder-decoder architecture for modeling. The encoder-decoder uses three sequential components to hierarchically construct a recurrent neural network model for the conversation: the encoder recurrent neural network is used for sentence encoding, the context recurrent neural network is used for modeling the utterance-level dialogue context, and the decoder recurrent neural network is used for generating response sentences. Each sentence divided in step S1 is fed into the recurrent neural network model for encoding-context modeling-decoding operations:
[0013] Encoding operation: First, each sentence divided in step S1 is fed into the encoder recurrent neural network for encoding, and a sentiment-related hidden layer vector during the encoding process is obtained through the attention mechanism. At a certain moment t, the following formula is calculated:
[0014]
[0015] Among them, represents the state output of the encoder at the i-th moment, f es represents the source task encoder recurrent neural network function, and Attention represents the calculation of the attention mechanism.
[0016] Context modeling: The sent to the context recurrent neural network for dialogue context modeling (where i = 1, 2, …, t), and the hidden layer state at time point t will be obtained
[0017]
[0018] where f cs represents the source task context recurrent neural network function.
[0019] Decoding operation: Use the decoder recurrent neural network to generate the response sentence x t+1 :
[0020]
[0021] where f ds represents the source task decoder recurrent neural network function. The encoder-decoder architecture is trained on the entire corpus of dialogues through the maximum likelihood estimation objective arg max θ ∑ i log p(X i ).
[0022] S3: Similarly, for each sentence in the target corpus of the target task, send it into the recurrent neural network model for encoding-context modeling operations:
[0023] Encoding operation: First, send each sentence into the encoder for encoding, and obtain the sentiment-related hidden layer vector during the encoding process through the attention mechanism. At a certain moment t, perform the following calculation:
[0024]
[0025] where represents the state output of the encoder at time i, and f et represents the target task encoder recurrent neural network function, and Attention represents the calculation of the attention mechanism.
[0026] Context modeling: Send the obtained in the encoding operation (where i = 1, 2, …, t) into the context recurrent neural network for dialogue context modeling, and the hidden layer state at time point t will be obtained
[0027]
[0028] where f ct represents the target task context recurrent neural network function.
[0029] S4: Transfer the attention information from the source corpus to the training network of the target corpus by defining a spatial attention map to complete knowledge transfer. Define the activation tensor of the recurrent neural network It consists of C channels, with a spatial dimension of H×W. The mapping function F takes A as the input and output, and the spatial attention map is calculated as follows:
[0030]
[0031] For the spatial attention map, since the absolute value of the hidden neuron activation can represent the importance of the neuron relative to a specific input, calculate the statistical information of the absolute value of the hidden neuron activation across the channel dimension and construct the following spatial attention mapping:
[0032]
[0033] where i ∈ {1, 2, …, H} and j ∈ {1, 2, …, W}, and p represents the lp-norm pooling calculation on all convolutional response channels of the activation maps in the source domain and target domain at a specific convolutional layer. In the attention transfer module, given the spatial attention map of the source task, the goal is to train the target task to not only make correct predictions but also have an attention map similar to that of the source task. The transfer loss between the source task and the target task is calculated by the following formula:
[0034]
[0035] where, and represent the losses of the source task and the target task respectively, W AT represents the weight of the transfer loss, represents the transfer loss.
[0036] The specific calculation is as follows:
[0037]
[0038] where Θ represents the spatial attention map, and represent the jth pair of spatial attention maps in the target task and the source task respectively. The l1-norm pooling calculation is used for the calculation of the selection.
[0039] S5: After completing the knowledge transfer in step S4 and conducting the encoding and modeling training on the target task corpus, use a softmax classifier to perform sentiment classification on the target statement and obtain the recognition rates of various sentiments. The final result outputs the sentiment classification matrix of the target statement, so as to be able to judge the emotional state of the speaker of each sentence.
[0040] The classification calculation of the softmax classifier and the expression of the loss function Loss during the training process are as follows:
[0041]
[0042]
[0043]
[0044] Among them, y is all the true sentiment labels, represents the hidden layer state of the context recurrent neural network at time point t in the target task, W o is the weight matrix, b o is the bias term, is the predicted probability, c is the number of sentiment classes, N represents the number of samples, y i,j represents that the i-th sentence is the true label of the j-th class of sentiment, represents the predicted probability that the i-th sentence is the j-th class of sentiment.
[0045] Advantages of the present invention: The present invention proposes a cross-corpus sentiment recognition method based on transfer learning and attention mechanism. In this method, a recurrent neural network (RNN) is used to extract the sentiment dependence and transmission in the context, and the feature parameters such as encoding and context sentiment dependence are transferred to the training of the target corpus through the attention transfer module. During the training process, the transfer loss is constrained within a certain range to complete the knowledge transfer. This method can complete the task of determining the speaker's sentiment state on the target corpus by leveraging the knowledge of transfer learning on the target corpus with less data volume, and can effectively solve the problem of insufficient training of small samples. Brief Description of the Drawings
[0046] Figure 1 is the framework flow chart of the present invention.
[0047] Figure 2 is the network structure diagram of the source task and the target task. Detailed Embodiments
[0048] The following further illustrates the detailed embodiments of the present invention in combination with the drawings and technical solutions.
[0049] The present invention can be used for cross-corpus sentiment recognition tasks based on transfer learning and attention mechanism. The process of the present invention is as Figure 1 shown, and the network structure adopted is as Figure 2 shown. This embodiment is applied to the sentiment classification task of the speaker in the dialogue. The following mainly details the implementation manner of the present invention for the problem of speaker sentiment recognition in the dialogue, including the following steps:
[0050] S1: Divide the dialogue part in the source corpus into t sentences X = [x 1 , x 2 , …, x i …, x t , and select the text data of all speakers in the dialogue.
[0051] S2: Use an encoder-decoder architecture for modeling. The encoder-decoder uses three sequential components to model the conversation in a hierarchical manner: an encoder recurrent neural network for sentence encoding, a context recurrent neural network for modeling the utterance-level dialogue context, and a decoder recurrent neural network for generating response sentences. Feed each sentence divided in step S1 into the recurrent neural network model for encoding-context modeling-decoding operations. See Figure 2 , and select a bidirectional long short-term memory network (BLSTM) model for the encoder and context modeling, and a bidirectional long short-term memory network (LSTM) model for the decoder:
[0052] Encoding operation: First, feed each sentence divided in step S1 into the encoder recurrent neural network for encoding, and obtain the hidden layer vector related to emotion during the encoding process through the attention mechanism. At a certain moment t, perform the following calculation as shown in the formula:
[0053]
[0054] where, represents the state output of the encoder at the i-th moment, f es represents the source task encoder recurrent neural network function, and Attention represents the calculation of the attention mechanism.
[0055] Context modeling: Feed the obtained in the previous step (where i = 1, 2, …, t) into the context recurrent neural network for dialogue context modeling, and obtain the hidden layer state at the t-th time point
[0056]
[0057] where, f cs represents the source task context recurrent neural network function.
[0058] Decoding operation: Use the decoder recurrent neural network to generate the response sentence x t+1 .
[0059]
[0060] where, f ds represents the source task decoder recurrent neural network function. The encoder-decoder architecture uses the maximum likelihood estimation objective arg maxθ ∑ i log p(X i ) to perform overall training on the conversations in the corpus.
[0061] S3: Similarly, each statement of the target task is fed into a recurrent neural network model for encoding-context modeling operations:
[0062] Encoding operation: First, each statement is fed into the encoder for encoding, and a sentiment-related hidden layer vector during the encoding process is obtained through the attention mechanism. At a certain moment t, the following formula is calculated:
[0063]
[0064] Among them, represents the state output of the encoder at the i-th moment, f et represents the recurrent neural network function of the target task encoder, and Attention represents the calculation of the attention mechanism.
[0065] Context modeling: Feed the obtained in the previous step (where i = 1, 2,..., t) into the context recurrent neural network for dialogue context modeling, and the hidden layer state at the t-th time point
[0066]
[0067] Among them, f ct represents the recurrent neural network function of the target task context.
[0068] S4: Attention transfer module. This module transfers attention information from the source corpus to the training network of the target corpus by defining a spatial attention map. Define the activation tensor of the bidirectional LSTM network which consists of C (for bidirectional LSTM, C = 1) channels, with a spatial dimension of H×W. The mapping function F takes A as the input and output, then the spatial attention map is calculated as follows:
[0069]
[0070] For the spatial attention map, since the absolute value of the hidden neuron activation can represent the importance of the neuron relative to a specific input, calculate the statistical information of these absolute values in the cross-channel dimension and construct the following spatial attention mapping:
[0071]
[0072] where \(i\in\{1,2,\ldots,H\}\) and \(j\in\{1,2,\ldots,W\}\), \(p\) represents the \(l_p\)-norm pooling calculation over all convolutional response channels of the activation maps in the source domain and the target domain at a specific convolutional layer. In the attention transfer module, given the spatial attention map of the source task, the goal is to train the target task to not only make correct predictions but also have an attention map similar to that of the source task. The transfer loss between the source task and the target task is calculated by the following formula:
[0073]
[0074] where and represent the losses of the source task and the target task respectively, \(W\) AT represents the weight of the transfer loss, represents the transfer loss.
[0075] Specifically, the specific calculation is as follows:
[0076]
[0077] where \(\Theta\) represents the spatial attention map, and represent the \(j\)-th pair of spatial attention maps in the target task and the source task respectively. Here the \(l_1\)-norm pooling calculation is selected.
[0078] In addition, the specific calculation is as follows:
[0079]
[0080] where \(\sigma\) is the softmax function, \(f\) s represents the source task model, which performs a classification task on the dialogue statements of \(N\) classes of labels: that is, classifying the statement \(X\) with the s label \(Y\) s and belonging to the \(n\)-th class.
[0081] Similarly, the specific calculation is as follows:
[0082]
[0083] where the first term is the conventional softmax cross-entropy loss function, and the second term is the transfer loss, and represent the \(j\)-th pair of spatial attention maps of the target task model \(f\) t and the source task model \(f\) s respectively, and \(\beta\) is the weight of the attention transfer loss.
[0084] To achieve attention transfer, pre-training is performed on the source task corpus to obtain the spatial attention map. For the training of the source task model, an encoder-context modeling-decoder model is used, where each of the forward and backward hidden layers of the BLSTM network has 128 units, and the learning rate is set to 0.001. The Movie Dialog Corpus dataset (with a large amount of data) is used as the source task database.
[0085] S5: Use a softmax classifier to perform sentiment classification on the target statement and obtain the recognition rates of various sentiments. The final result outputs the sentiment classification matrix of the target statement, enabling the determination of the emotional state of the speaker of each sentence.
[0086] The classification calculation of the softmax classifier and the calculation expression of the loss function Loss during the training process are as follows:
[0087]
[0088]
[0089]
[0090] Among them, y is all the true sentiment labels, W o is the weight matrix, b o is the bias term, is the predicted probability, c is the number of sentiment classes, N represents the number of samples, y i,j represents the true label that the i-th sentence is the j-th type of sentiment, represents the predicted probability that the i-th sentence is the j-th type of sentiment.
[0091] In this embodiment, the Adam optimizer is used to optimize the training network learning parameters, and Dropout is used to prevent overfitting. The initial learning rate is set to 0.001. In this embodiment, the Movie Dialog Corpus is selected as the source task corpus, and classification experiments on 6 types of sentiments (happy, sad, neutral, angry, excited, annoyed) are respectively carried out on the IEMOCAP and DailyDialog as the target task corpora, and the following experimental results are obtained:
[0092] Source task corpus Target task corpus Average recognition rate (%) Movie Dialog Corpus IEMOCAP 61.4 Movie Dialog Corpus DailyDialog 52.8
[0093] The above table shows that the method of the present invention can perform effective sentiment recognition on the IEMOCAP and DailyDialog as the target task corpora by leveraging the knowledge learned on the source task corpus Movie Dialog Corpus.
[0094] Although this embodiment introduces the inventive method in the training process, in practical applications, the trained network model can be used to perform classification tests on different data sets. In addition, in addition to the LSTM and bidirectional LSTM used in the examples, other models containing time series information can also be used.
Claims
1. A cross-corpus sentiment recognition method based on transfer learning and attention mechanism, characterized in that, the specific steps are as follows: S1: Divide the dialogue part in the source corpus into T sentences X = [x 1 , x 2 , …, x i …, x T , and select the text data of all speakers in the dialogue; S2: Use an encoder-decoder architecture for modeling; The encoder-decoder uses three sequential components to hierarchically construct a recurrent neural network model for conversations: the encoder recurrent neural network is used for sentence encoding, the context recurrent neural network is used for modeling the utterance-level dialogue context, and the decoder recurrent neural network is used for generating response sentences; Each statement divided in step S1 is sent into the recurrent neural network model for encoding-context modeling-decoding operations: Encoding operation: First, each statement divided in step S1 is sent into the encoder recurrent neural network for encoding, and a sentiment-related hidden layer vector during the encoding process is obtained through the attention mechanism. The following formula is calculated at a certain moment t: Among them, represents the state output of the encoder at time t, and f es represents the source task encoder recurrent neural network function, and Attention represents the calculation of the attention mechanism; Context modeling: The obtained in the encoding operation is fed into a context recurrent neural network for dialogue context modeling, and the hidden layer state at time t is obtained Among them, f cs represents the source task context recurrent neural network function; Decoding operation: Use a decoder recurrent neural network to generate a response sentence x m : Among them, f ds represents the source task decoder recurrent neural network function; the encoder-decoder architecture is trained on the entire corpus of conversations through the maximum likelihood estimation objective arg max θ ∑ i logp(X i ). S3: Each statement in the target corpus of the target task is sent into the recurrent neural network model for encoding-context modeling operations; S4: Transfer the attention information from the source corpus to the training network of the target corpus by defining a spatial attention map to complete knowledge transfer; define the activation tensor of the recurrent neural network It consists of C channels and has a spatial dimension of H×W. For the mapping function F with A as the input and output, the spatial attention map is calculated as follows: For the spatial attention map, since the absolute value of the hidden neuron activation represents the importance of the neuron relative to a specific input, calculate the statistical information of the absolute value of the hidden neuron activation in the cross-channel dimension and construct the following spatial attention mapping: where m ∈ {1, 2, …, H} and j ∈ {1, 2, …, W}, p represents the lp-norm pooling calculation on all convolutional response channels of the activation maps in the source domain and target domain of a specific convolutional layer; Given the spatial attention map of the source task, the goal is to train the target task to not only make correct predictions but also have an attention map similar to the source task; S5: After completing the knowledge transfer in step S4 and performing encoding modeling training on the target task corpus, use a softmax classifier to perform sentiment classification on the target statement and obtain the recognition rates of various sentiments; The final result outputs the sentiment classification matrix of the target statement, so as to be able to judge the emotional state of the speaker of each sentence.
2. The cross-corpus sentiment recognition method based on transfer learning and attention mechanism according to claim 1, characterized in that, the specific operations in S3 are as follows: Encoding operation: First, each statement is sent into the encoder for encoding, and a sentiment-related hidden layer vector during the encoding process is obtained through the attention mechanism. The following formula is calculated at a certain moment t: Among them, represents the state output of the encoder at the i-th moment, and f et represents the target task encoder recurrent neural network function, and Attention represents the calculation of the attention mechanism; Context modeling: The obtained in the encoding operation is fed into a context recurrent neural network for dialogue context modeling, and the hidden layer state at time t, Among them, f ct represents the target task context recurrent neural network function.
3. The cross-corpus sentiment recognition method based on transfer learning and attention mechanism according to claim 1, characterized in that, the transfer loss between the source task and the target task in S4 is calculated by the following formula: Among them, and represent the losses of the source task and the target task respectively, and W AT represents the weight of the transfer loss, represents the transfer loss; The specific calculation is as follows: where Θ represents the spatial attention map, and represent the j-th pair of spatial attention maps in the target task and the source task, respectively; Calculate and select l1-norm pooling calculation.
4. The cross-corpus sentiment recognition method based on transfer learning and attention mechanism according to claim 1, characterized in that, the classification calculation of the softmax classifier and the calculation expression of the loss function Loss during the training process are as follows: Among them, y is all the true sentiment labels, represents the hidden layer state of the context recurrent neural network at time point t in the target task, W o is the weight matrix, b o is the bias term, is the predicted probability, c is the number of sentiment classes, N represents the number of samples, y i,j indicates that the i-th sentence is the true label of the j-th class of sentiment, indicates that the predicted probability of the i-th sentence being the j-th class of sentiment.
Citation Information
Patent Citations
Multimodal emotion recognition methods and systems based on neural networks and transfer learning
CN107609572B
A multimodal speech emotion recognition method based on enhanced residual neural network
CN109460737A
Multi-modal depression detection method and system based on context awareness
CN110728997A
Comment emotion classification method and system based on deep hybrid model transfer learning
CN109271522A
An aspect-level emotion classification model and method based on dual-memory attention
CN109472031A