Dialogue emotion recognition method based on double-graph attention neural network and comparative learning
By using the dual-graph attention neural network and contrast learning method in dialogue emotion recognition, the problem of insufficient classification ability caused by noise and data imbalance in the iteration process is solved, and more accurate and efficient emotional feature extraction and classification are achieved.
Patent Information
- Application Number
- CN202510028268.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-08
AI Technical Summary
In the prior art, graph neural networks with complex model structures are prone to generate a lot of noise during the iteration process, and due to the problem of data imbalance in the dialogue emotion recognition data set, the model has poor ability to classify emotional categories with few samples.
The dialogue emotion recognition method based on dual-graph attention neural network and contrast learning is adopted, and the speech emotional characteristics related to the dialogue context are obtained through the dual-graph attention neural network, and an asymmetric graph update strategy is introduced. Combined with the contrast learning module, the similarity between emotional characteristics on speeches in the same and different emotional categories is calculated to calculate the loss.
It effectively solves the problem that graph neural network has weak ability to extract global information and complex graph structures are prone to noise, improves the model's ability to classify emotional categories for few samples, and generates more similar emotional characteristics for speeches for the same emotional categories and more different emotional characteristics for speeches for different emotional categories.
Smart Images

Figure CN119938928A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing technology and dialogue emotion recognition, and specifically relates to a dialogue emotion recognition method based on dual-graph attention neural network and contrastive learning. Background Art
[0002] Conversation emotion recognition is a very important research hotspot in the field of natural language processing and has great application value. Early conversation emotion recognition adopted a dictionary-based method, that is, the emotion value of the emotion word was obtained through the emotion dictionary, and then the sum or average of the emotion value was used to calculate the emotion of the entire sentence or document, thereby determining the emotional tendency in the text conversation. For example, Li et al. merged a sentiment dictionary called NRC-Valence arousal dominance to obtain word-level emotion distribution and assigned emotions with intensity to emotion words. However, these classification methods based only on keywords still have some problems. First, words have different meanings depending on usage and context, which will cause ambiguity. Second, these dictionaries cannot recognize text emotions that do not contain emotion keywords. Finally, the accuracy of emotion recognition is often affected by the flexibility of conversation text data, the professionalism of text data, and the lack of language-related information when performing text conversation emotion recognition based on emotion dictionaries.
[0003] In recent years, with the rapid development of deep learning models, researchers have gradually applied deep learning models to the field of natural language processing. Gupta et al. first clearly defined conversation emotion recognition and used LSTM to model the context of sentences, achieving good results. In order to distinguish different parties in a meaningful way, Majumder et al. proposed the DialogueRNN model, which uses three GRUs to capture the speaker information, context, and emotional information of previous utterances. However, these LSTM or GRU-based sequence models have the problem of difficulty in characterizing long-distance or global dependencies between clauses, thereby ignoring the deep dependencies between clauses. To solve this problem, researchers have tried to apply graph neural networks (GNNs) to conversation emotion recognition methods. Unlike LSTM and GRU, which only perform cyclic modeling on sequence data, GNN regards the problem of conversation emotion recognition as a problem of connection and message propagation between nodes in a graph, models the dependencies between nodes, and two nodes connected by edges can transmit information regardless of temporal distance. Ghosal et al. proposed a dialogue graph convolutional network (DialogueGCN), which uses dependencies within and between speakers to construct a graph neural network structure to simulate conversation context for emotion recognition. Through graph neural networks, DialogueGCN solves the context propagation problem existing in current DialogueRNN-based methods, but ignores the sequential information between discourses and information about different interlocutors. Based on this, Liang et al. proposed a speaker and position-aware graph neural network model (S+PAGE) for conversation emotion recognition, combining the advantages of relational graph convolutional networks and Transformers to achieve better context modeling. Furthermore, Zhang et al. proposed a dual graph attention neural network (DualGAT), which makes up for the shortcomings of the lack of fine-grained information in the minimalist graph structure by combining the information extracted from the directed acyclic graph that simulates the contextual sequence and the complete graph perceived by the speaker.
[0004] However, in the above existing technologies, GNNs with complex model structures often generate a lot of noise during the iteration process, and because dialogue emotion recognition datasets often have data imbalance problems, the models trained based on these datasets have poor classification capabilities for certain emotion categories with few samples. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a conversation emotion recognition method based on dual-image attention neural network and contrastive learning, comprising:
[0006] Preprocess the conversation data set; input the preprocessed data into the trained conversation emotion recognition model to obtain the probability that each sentence in the conversation belongs to a certain emotion category, and output the emotion category to which each sentence belongs based on the probability;
[0007] The dialogue emotion recognition model includes: RoBERTa model, dual-image attention neural network, cross-attention mechanism module, softmax classification layer and contrastive learning module;
[0008] The process of training a conversation emotion recognition model includes:
[0009] S1: Obtain the original conversation dataset, divide the dataset into a training set and a test set, and preprocess the data in the training set;
[0010] S2: Input the preprocessed data into the RoBERTa model to generate the initial semantic representation of each sentence;
[0011] S3: Input the initial semantic representation into the dual-graph attention neural network to obtain local context semantic representation and global context semantic representation respectively;
[0012] S4: Input the local and global context semantic representations into the cross-attention mechanism module for interaction to obtain the fused conversation semantic representation;
[0013] S5: Obtain conversation sentiment representation for sentiment classification based on the fused conversation semantic representation;
[0014] S6: Perform softmax classification on the dialogue emotion representation to obtain the emotion category of each sentence; set the initial value of the number of iterations;
[0015] S7: Input the conversation sentiment representation into the contrastive learning module assisted by the auxiliary set and calculate the contrastive learning loss;
[0016] S8: The cross entropy loss function is used to calculate the loss of sentiment classification and is used together with the contrastive learning loss to train the model;
[0017] S9: Use Adam algorithm to optimize the model;
[0018] S10: Determine whether the number of iterations reaches the maximum number of iterations. If so, the training of the dialogue emotion recognition model is completed. Otherwise, return to step S8 and increase the number of iterations by 1.
[0019] Beneficial effects of the present invention:
[0020] The present invention utilizes a dual-graph attention neural network model comprising two graph neural networks with different structures to obtain speech emotion features related to the conversation context, and introduces an asymmetric graph update strategy, which solves the problems in previous methods that graph neural networks with simple graph structures have weak ability to extract global information and graph neural networks with complex graph structures are prone to generate a large amount of noise. A comparative learning module is then used to compare the similarities between the emotion features of speeches of the same and different emotion categories, and the loss is calculated accordingly, so that the graph neural network generates more similar emotion features for speeches of the same emotion category; and generates more differentiated emotion features for speeches of different emotion categories, which solves the problem in previous conversation emotion recognition models that the model has a high probability of misclassifying small sample categories due to the imbalance of the training data set. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of the training process of the present invention;
[0022] Figure 2 This is a structural diagram of the dual-graph neural network model of the present invention. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] A conversation emotion recognition method based on dual-graph attention neural network and contrastive learning, such as Figure 1 The method comprises: preprocessing a dialogue data set; inputting the preprocessed data into a trained dialogue emotion recognition model to obtain the probability that each speech in the dialogue belongs to a certain emotion category, and outputting the emotion category to which each speech belongs according to the probability; the dialogue emotion recognition model comprises a dual-image attention neural network model and a contrastive learning module.
[0025] like Figure 1 As described above, the process of training the conversation emotion recognition model includes:
[0026] S1: Obtain the original conversation dataset, divide the dataset into a training set and a test set, and preprocess the data in the training set;
[0027] S2: Input the preprocessed data into the RoBERTa model to generate the initial semantic representation of each sentence;
[0028] S3: Input the initial semantic representation of the conversation into the dual-graph attention neural network to obtain local context semantic representation and global context semantic representation respectively;
[0029] S4: Local and global context semantic representations are fed into the cross-attention mechanism module for interaction.
[0030] Get the fused semantic representation of the conversation;
[0031] S5: Obtain conversation sentiment representation for sentiment classification based on the fused conversation semantic representation;
[0032] S6: Input the conversation sentiment representation into the contrastive learning module assisted by the auxiliary set and calculate the contrastive learning loss;
[0033] S7: Perform softmax classification on the dialogue emotion representation to obtain the emotion category of each sentence; set the initial value of the number of iterations;
[0034] S8: The cross entropy loss function is used to calculate the loss of sentiment classification and is used together with the contrastive learning loss to train the model;
[0035] S9: Use Adam algorithm to optimize the model;
[0036] S10: Determine whether the number of iterations reaches the maximum number of iterations. If so, the training of the dialogue emotion recognition model is completed. Otherwise, return to step S8 and increase the number of iterations by 1.
[0037] The MELD public dialogue dataset is used as the basic dataset. The dataset is divided into three parts. The training set contains 1039 groups of dialogues and 9989 speeches; the test set contains 280 groups of dialogues and 2610 speeches; and the validation set contains 114 groups of dialogues and 1109 speeches.
[0038] The process of preprocessing the above data includes: removing stop words in the original data set, setting the length of each speech to the length of the longest speech in each conversation, truncating speeches that are too long from the end, and padding speeches that are not long enough; the processed single-sentence speeches are input into the RoBERTa model to extract the initial representation:
[0039] Input(u i )=[CLS],w1,…,w i ,…,w n , [SEP]
[0040] h i =RoBERTa(Input(u i ))
[0041] Among them, u iIndicates a statement in a conversation, w i (1≤i≤n) represents a word contained in the speech, [CLS] is the classification mark, [SEP] is the end mark, h i For speech i After inputting the RoBERTa model, the last hidden layer output corresponding to the [CLS] position is used as the speech u i The overall input of a batch of data to the model is H = (h1||h2||h3...h N ), where || represents vector concatenation and N represents the total number of speeches in a batch.
[0042] After obtaining the initial semantic representation, the dual graph neural network model consisting of two graph attention neural networks with different structures is input for feature extraction, such as Figure 2 As shown; the structure definition and feature extraction process of the two graph neural networks are as follows:
[0043] 1) Global Context Graph Neural Network (GLGAT):
[0044] Firstly, a graph neural network with a directed complete graph structure is used as the source of global context information (Global Graph Attention Network, GLGAT). The definition of the GLGAT graph structure is as follows: G gl = {V gl , E gl}, where V gl = {v1, v2, v3…v N} is the point set of the graph, N represents the total number of speeches in a batch, and each node v i (1≤i≤N) are initialized to each speech u in the conversation i The hidden layer representation h i ; E gl ={e1, e2, e3…e M} is the edge set of the graph, M represents the total number of edges in the graph. i , v j , r) represents the node v i To node v j The relationship type is a directed edge of r. p(v i ) is the speaker mapping function, representing node v i Corresponding speech i For the design of the global context graph neural network structure, five edge types are defined as follows:
[0045] Intra-Past: If there is a node v i , v j , when i>j, p(vi )=p(v j ), then (v i , v j ,0)∈E gl ,
[0046] Intra-Future: If there is a node v i , v j , when i <j,p(v i )=p(v j ), then (v i , v j , 1)∈E gl ,
[0047] Inter-Past: If there is a node v i , v j , when i>j, p(v i )≠p(v j ), then (v i , v j , 2)∈E gl ,
[0048] Inter-Future: If there is a node v i , v j , when i <j,p(v i )≠p(v j ), then (v i , v j , 3)∈E gl ,
[0049] Intra-Loop: If there is a node v i , then (v i , v i , 4)∈E gl .
[0050] By the Intra-Past type edge (edge type is 0) and the node v i All connected nodes have speakers who are related to node v i The corresponding speech u i The speakers are the same and the timing is at speech u i Before; by the Intra-Future type edge (edge type is 1) and the node v i All connected nodes have speakers who are related to node v i The corresponding speech u i The speakers are the same and the timing is at speech u i After that, the Inter-Past type edge (edge type 2) and node v i All connected nodes have speakers who are related to node vi The corresponding speech u i The speakers are different and the timing is located at speech u i Before; Inter-Future type edge (edge type is 3) and node v i All connected nodes have speakers who are related to node v i The corresponding speech u i The speakers are different and the timing is located at speech u i After that, the Self-Loop type edge (edge type 4) and node v i The connected node is node v i The global context information of the current node is extracted based on this. The representation of each layer of nodes is updated according to the neighboring nodes connected to it. The calculation method of the graph neural network node representation update is as follows.
[0051] α ij =softmax(LRL(a T [Wh i || Wh j ||r ij ]))
[0052]
[0053] where α ij For node v i For neighbor node v j The attention coefficient is also called edge weight. Softmax(.) is the normalization function, and LRL(.) refers to the LeakyReLU activation function. || represents concatenation, T represents matrix transposition, and α and W are trainable parameter matrices. i ,h j Respectively represent the node v i , v j The corresponding speech representation vector, r ij Represents the connected node v i , v j The one-hot vector corresponding to the edge type between . Represents node v i After GLGAT updates the node representation, It represents node v i In the global context graph G gl The point set of all neighbor nodes connected by edges. After a complete update of GLGAT, the representation of each node in its output is Where || represents vector concatenation, and N represents the total number of speeches in a batch. As mentioned above, the GLGAT node representation update formula is summarized as follows:
[0054] H gl =GLGAT(H,E gl )
[0055] Where GLGAT represents the global context graph neural network, H = (h1||h2||h3…h N ) is the input of GLGAT, h i For speech i Corresponding representation, || represents vector concatenation, and N represents the total number of speeches in a batch. gl represents the edge set in the graph structure of GLGAT, H gl It is the output representation obtained after the input H is updated by GLGAT.
[0056] 2) Local Context Graph Neural Network (LOGAT):
[0057] Secondly, a graph neural network with a directed acyclic graph structure is used as the source of local context information (Local Graph Attention Network, LOGAT). The definition of the LOGAT graph structure is as follows: lo = {V lo , E lo}, where V gl = {v1, v2, v3…v N} is the point set of the graph, N represents the total number of speeches in a batch, and each node v i (1≤i≤N) are initialized to each speech u in the conversation i The hidden layer representation h i ; E gl ={e1, e2, e3…e M} is the edge set of the graph, M represents the total number of edges in the graph. j , v i , r) represents the node v i To node v j The relationship type is a directed edge of r. p(v i ) is the speaker mapping function, representing node v i Corresponding speech i For the design of the local context graph neural network structure, two edge types are defined as follows:
[0058] Intra-Speaker: If there is a node v i , v j , when i>j, p(v i )=p(v j ), and there is no node v k , p(v i )=p(v k), i>k>j, then (v j , v i , 1)∈E lo ,
[0059] Inter-Speaker: If there is a node v i , v j , when i>j, p(v i )≠p(v j ), and there is no node v k , p(v i )=p(v k ), i>k>j, then (v j , v i ,0)∈E lo ,
[0060] By the Intra-Speaker type edge (edge type is 1) and node v i The connected nodes satisfy the following conditions: i The corresponding speech u i The speakers are the same and the time sequence is in speech u i The closest node before. This type of relationship is used to obtain the contextual information of distant speakers; it is connected by an Inter-Speaker type edge (edge type 0) and node v i The connected nodes satisfy the following conditions: i The corresponding speech u i The speakers are different and the time sequence is in speech u i The closest node before. This type of relationship is used to capture the close context information that has the greatest impact on the current node. These two types of context information are captured as local context information of the current node through the update calculation of LOGAT. The specific update method of the graph neural network node representation is the same as GLGAT. The summary of the LOGAT node representation update formula is as follows:
[0061] H lo =LOGAT(H,E lo )
[0062] Where LOGAT represents the local context graph neural network, H = (h1||h2||h3…h N ) is the input of LOGAT, h i For speech i Corresponding representation, || represents vector concatenation, and N represents the total number of speeches in a batch. lo represents the edge set in the graph structure of LOGAT, H lo It is the output representation of the input H after LOGAT update.
[0063] After the dual graph neural network extracts features, the cross-attention mechanism is used to fuse the contextual semantic representations extracted by the two graph neural networks and perform asymmetric updates:
[0064] S1: Output the node representations of the two graphs H gl , H lo Through the calculation of the cross attention mechanism function, the attention coefficient matrix of the two representation matrices with respect to each other is obtained;
[0065]
[0066] Among them, softmax(.) is a normalization function, W is a trainable parameter matrix, T represents matrix transposition, and H represents the node representation output by the graph neural network after the update. n is a differential hyperparameter. Past experiments have shown that graph neural networks with too complex graph structures should not be updated with too many layers to avoid redundant noise features. Therefore, the hyperparameter n is specially set so that before the two graph neural networks interact with each other, LOGAT will perform n internal node updates, while GLGAT will only perform one update. gtl A represents the attention weight matrix of GLGAT for LOGAT representation, ltg Represents the attention weight matrix of LOGAT for GLGAT representation.
[0067] S2: representation matrix H gl , H lo And the corresponding attention coefficient matrix A gtl , A ltg Multiply them to get the graph node representation vectors of the two graphs after interaction;
[0068]
[0069] H′={H gl′ , H lo}
[0070] Among them, H l represents the node representation output by the graph neural network after the lth update, H gl′ , H lo′ are the representations obtained after cross-attention calculation of GLGAT and LOGAT, and H′ is the total representation output by the graph neural network after one update, which will be used as the overall input for the next dual-graph neural network update, where Hg l′ As the input of the next GLGAT, H lo′ As the input for the next LOGAT.
[0071] The representations extracted by the dual-graph neural network after multiple updates are concatenated and input into the classification layer to calculate the final emotion category probability:
[0072] S1: Get the speech u after L graph neural network updates i The final hidden layer representation of:
[0073]
[0074] in, Representatives speak i After all the updates of GLGAT and LOGAT, the output representations are obtained. After concatenating the output representations of these two graph neural networks, the speech u is obtained. i Final hidden layer representation for classification After concatenating all the final hidden layer representations of the conversation data in a batch, we get the overall representation for classification. Where || represents vector concatenation and N represents the total number of speeches in a batch.
[0075] S2: For speech u i The final hidden layer representation is subjected to softmax classification to obtain the emotion category with the highest corresponding probability:
[0076]
[0077] p i =softmax(W o o i +b o )
[0078]
[0079] Among them, ReLU(.) is the activation function, softmax(.) is the probability normalization function, argmax(.) is the classification function for finding the maximum value, and W h , b h , W o , b o is a trainable parameter, o i For transition characterization, p i For speech i The probability distribution of each emotion category, For speech i One-hot encoding of the sentiment category predicted by the model.
[0080] Finally, for the features extracted by the model, contrastive learning is used to calculate the similarity between the representations of various emotion categories, and the loss is calculated accordingly, so that the model generates more similar emotion features for speeches of the same emotion category, and more different emotion features for speeches of different emotion categories. label The calculation is as follows:
[0081]
[0082] Among them, N b It represents the total number of speeches in a conversation (batch), and I(.) is an indicator function, which is 1 when the condition in the function is met, otherwise it is 0. Speech u generated by the dual graph neural network model i The final emotional representation, y i For speech i The real emotional label, is a similarity function used to calculate the similarity between vectors x and y.
[0083] In the implementation process of contrastive learning, an additional auxiliary set is constructed to assist the contrastive learning algorithm, which solves the problem in the traditional contrastive learning algorithm that when there is only one sentence in a certain emotion category in a batch, the lack of positive examples will make L label Problems that cannot be calculated correctly. The construction and application of auxiliary sets are as follows:
[0084] center c =mean(h i ), y i =c
[0085] h c =NEAREST(h i , center c ), i=1,2,3…N
[0086] support c =Relu(Linear(h c ))
[0087] CInput=(H final ||support c )
[0088] Among them, mean(.) is the average function, NEAREST(x, y) is the nearest neighbor function, and its return value is the vector with the highest similarity to y in the vector set x, Relu(.) is the activation function, and Linear(.) is the linear function. center c is the average representation center of all speeches with sentiment category c in the dataset, hi For speech i The RoBERTa initial sentiment representation, y i For speech i The real emotional label, h c Indicates the initial emotional representation of all speeches with emotional category c, and center c The speech representation with the highest similarity is is the similarity function. c h c After linear transformation, it represents the auxiliary representation of emotion category c. CInput is the input of the contrastive learning module for the data in a batch. H final It is the final representation set of all conversation data in a batch after being processed by the dual-graph neural network model, and || represents vector concatenation.
[0089] After contrastive learning is completed, the loss is calculated by combining the cross entropy sentiment loss function and the contrastive learning loss function. The loss function of the overall model is calculated as:
[0090]
[0091] L=L emo +λL label
[0092] Among them, L emo is the loss function of the model sentiment classification, using cross entropy loss, N b Indicates the total number of speeches in a batch, p i For speech i The probability distribution of each emotion category, y i For speech i One-hot encoding of the true sentiment label, λ is a weight hyperparameter less than 1. L is the overall loss function of the model.
[0093] The model uses the Adam algorithm to optimize parameters to minimize the loss function. The Adam algorithm is an optimization algorithm that can replace the traditional stochastic gradient descent process. It can iteratively update the weights of the neural network based on the training data, and can design independent adaptive learning rates for different parameters by calculating the first-order moment estimate and second-order moment estimate of the gradient. The main calculation formula is as follows:
[0094]
[0095] Among them, Δθ t Represents the updated value of the parameter during training, represents the corrected first-order moment estimate, represents the corrected second-order moment estimate, ∈, η are the parameters that need to be adjusted during the training process.
[0096] In the experiment, macro-F1 is used as the evaluation index to evaluate the model, which represents the accuracy and stability of the model in classifying various emotional speeches;
[0097] The specific process includes:
[0098] Calculate the F1-score for each sentiment category c :
[0099]
[0100] Calculate macro-F1:
[0101]
[0102] Among them, F1-score c Represents the F1-score, Recall of sentiment category c c and Precision c They represent the accuracy and recall of the emotion category c respectively, Y represents the number of emotion categories in the dataset, and macro-F1 represents the final evaluation index. The higher the value, the better the model is at recognizing conversation emotions.
[0103] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for dialogue emotion recognition based on dual-image attention neural network and contrastive learning, characterized in that: include: Preprocess the dialogue dataset; Input the preprocessed data into the trained conversation emotion recognition model to obtain the probability that each sentence in the conversation belongs to a certain emotion category, and output the emotion category to which each sentence belongs based on the probability; The dialogue emotion recognition model includes: RoBERTa model, dual-image attention neural network, cross-attention mechanism module, softmax classification layer and contrastive learning module; The process of training a conversation emotion recognition model includes: S1: Obtain the original conversation dataset, divide the dataset into a training set and a test set, and preprocess the data in the training set; S2: Input the preprocessed data into the RoBERTa model to generate the initial semantic representation of each sentence; S3: Input the initial semantic representation into the dual-graph attention neural network to obtain local context semantic representation and global context semantic representation respectively; S4: Input the local and global context semantic representations into the cross-attention mechanism module for interaction to obtain the fused conversation semantic representation; S5: Obtain conversation sentiment representation for sentiment classification based on the fused conversation semantic representation; S6: Perform softmax classification on the dialogue emotion representation to obtain the emotion category of each sentence; set the initial value of the number of iterations; S7: Input the conversation sentiment representation into the contrastive learning module assisted by the auxiliary set and calculate the contrastive learning loss; S8: The cross entropy loss function is used to calculate the loss of sentiment classification and is used together with the contrastive learning loss to train the model; S9: Use Adam algorithm to optimize the model; S10: Determine whether the number of iterations reaches the maximum number of iterations. If so, the training of the dialogue emotion recognition model is completed. Otherwise, return to step S8 and increase the number of iterations by 1.
2. According to claim 1, a method for dialogue emotion recognition based on dual-graph attention neural network and contrastive learning is characterized in that: The preprocessing of the dialogue dataset includes: Stop words in the original conversation dataset are removed. For each group of conversations, the length of each speech is set to the length of the longest speech in the conversation. Speeches that are too long are truncated from the end, and speeches that are not long enough are padded.
3. According to claim 1, a method for dialogue emotion recognition based on dual-graph attention neural network and contrastive learning is characterized in that: The preprocessed data is input into the RoBERTa model to generate the initial semantic representation of each sentence, including: Input(u i )=[CLS],w1,…,w i ,…,w n ,[SEP] h i =RoBERTa(Input(u i )) Among them, Input(u i ) represents the input of the RoBERTa model, u i Indicates a statement in a conversation. w i (1≤i≤n) represents a word contained in the speech, [CLS] is the classification mark, [SEP] is the end mark, h i Indicates speech i After inputting the RoBERTa model, the last hidden layer output corresponding to the [CLS] position is used as the speech u i The initial semantic representation of .
4. According to claim 1, a method for dialogue emotion recognition based on dual-graph attention neural network and contrastive learning is characterized in that: The initial semantic representation is input into the dual-graph attention neural network to obtain local context semantic representation and global context semantic representation, including: The dual-graph attention neural network includes: a global context graph neural network (GLGAT) and a local context graph neural network (LOGAT); Local context semantic representation: H gl =GLGAT(H,E gl ) Among them, GLGAT represents the global context graph neural network, H = (h1||h2||h3…h N ) is the input of GLGAT. The overall input of a batch of data to the model is H = (h1||h2||h3…h N ), || represents vector concatenation, N represents the total number of speeches in a batch; h i Indicates speech i The initial semantic representation of E gl represents the edge set in the graph structure of GLGAT, H gl Represents local context semantic representation; Global context semantic representation: A lo =LOGAT(H,E lo ) Among them, LOGAT represents the local context graph neural network, E lo represents the edge set in the graph structure of LOGAT, H lo Represents the global context semantic representation.
5. The method for dialogue emotion recognition based on dual-graph attention neural network and contrastive learning according to claim 1, characterized in that: The local and global context semantic representations are input into the cross attention mechanism module for interaction to obtain the fused dialogue semantic representation, including: S41: Representing local and global context semantics H gl ,H lo Through the calculation of the cross attention mechanism function, the attention coefficient matrix of the two representation matrices with respect to each other is obtained; Among them, softmax(.) represents the normalization function, W represents the trainable parameter matrix, T represents the matrix transpose, and H l represents the node representation output by the graph neural network after the lth update, n is the differential hyperparameter, and A gtl A represents the attention weight matrix of GLGAT for LOGAT representation, ltg Represents the attention weight matrix of LOGAT for GLGAT representation; S42: Representing local and global context semantics H gl ,H lo And the corresponding attention coefficient matrix A gtl , A ltg Multiply them together to get the fused semantic representation of the conversation; H′={H gl′ ,H lo ′} Among them, H gl′ ,H lo They are the representations obtained after cross-attention calculation of GLGAT and LOGAT, and H′ represents the semantic representation of the dialogue output by the graph neural network after one update.
6. The method for dialogue emotion recognition based on dual-graph attention neural network and contrastive learning according to claim 1, characterized in that: The conversation sentiment representation for sentiment classification is obtained based on the fused conversation semantic representation, including: The updated conversation semantic representation H′ is used as the overall input for the next dual-graph neural network update, where H gl′ As the input of the next GLGAT, H lo As the input of the next LOGAT, get the speech u after L graph neural network updates i Conversation sentiment representation for classification: in, Representatives speak i The output representation obtained after all updates of GLGAT and LOGAT, || represents vector concatenation, Indicates speech i Conversation sentiment representation for classification; after concatenating all conversation sentiment representations for classification in a batch of conversation data, the overall conversation sentiment representation for classification is obtained. N represents the total number of speeches in a batch.
7. The method for dialogue emotion recognition based on dual-graph attention neural network and contrastive learning according to claim 1, characterized in that: Perform softmax classification on the dialogue sentiment representation to obtain the sentiment category of each sentence, including: p i =softmax(W o o i +b o ) Among them, ReLU(.) represents the activation function, softmax(.) represents the probability normalization function, argmax(.) represents the maximum classification function, W h ,b h ,W o ,b o represents the first, second, third, and fourth trainable parameters of the classification layer, o i represents transition characterization, p i Indicates speech i The probability distribution of each emotion category, Indicates speech i One-hot encoding of the sentiment category predicted by the model.
8. The method for dialogue emotion recognition based on dual-graph attention neural network and contrastive learning according to claim 1, characterized in that: The conversation sentiment representation is input into the contrastive learning module assisted by the auxiliary set, and the contrastive learning loss is calculated, including: Among them, L label represents contrastive learning loss, N b Represents the total number of speeches in a batch conversation. I(.) represents an indicator function. Its value is 1 when the condition in the function is met, otherwise it is 0. Respectively represent speech u i 、u k 、u j Conversation sentiment representation for classification, y i ,y j Respectively represent speech u i 、u j The true sentiment label, F(.) represents the similarity function; To avoid the lack of positive examples when there is only one sentence of a certain sentiment category in a batch, L label For problems that cannot be calculated correctly, the following auxiliary sets are constructed to assist contrastive learning: center c =mean(h i ),y i =c h c =NEAREST(h i ,center c ),i=1,2,3…N support c =Relu(Linear(h c )) CInput=(H final ||support c ) Among them, mean(.) is the average function, NEAREST(x,y) is the nearest neighbor function, Relu(.) is the activation function, Linear(.) is the linear function, center c is the average representation center of all speeches with sentiment category c in the dataset, h i For speech i The RoBERTa initial sentiment representation, y i For speech i The real emotional label, h c Indicates the initial emotional representation of all speeches with emotional category c and center c The speech representation with the highest similarity, support c Represents the auxiliary representation of emotion category c, CInput represents the replacement of H final As the input of a batch of data to the contrastive learning module, H final is the overall conversation sentiment representation used for classification, and N is the total number of speeches in a batch.
9. The method for dialogue emotion recognition based on dual-graph attention neural network and contrastive learning according to claim 1, characterized in that: The cross entropy loss function is used to calculate the loss of sentiment classification, including: Among them, L emo represents the sentiment classification loss, N b Indicates the total number of speeches in a conversation, p i Indicates speech i The probability distribution of each emotion category, y i Indicates speech i Real emotional labels.
10. The method for dialogue emotion recognition based on dual-graph attention neural network and contrastive learning according to claim 1, characterized in that: The Adam algorithm is used to optimize the model, including: Among them, Δθ t Represents the updated value of the parameter during training, represents the corrected first-order moment estimate, represents the corrected second-order moment estimate, ∈, η are the first and second parameters that need to be adjusted during the training process.
Citation Information
Patent Citations
Session sentiment analysis method based on multi-granularity fusion and graph convolutional network
CN115374281A
Microblog comment dialogue sentiment analysis method and system based on graph neural network
CN118228733A
Dialogue emotion correction method based on graph neural network
US20220270636A1
Method for automatically generating concrete dam defect image description on basis of graph attention network
WO2023241272A1