Method and System for Joint Recognition of Dialogue Emotion and Behavior Based on Hierarchical Interactive Learning
By constructing interactive graphs and utilizing graph neural networks, mutual guidance of dialogue emotions and behaviors is achieved, the problem of independent processing of emotion and behavior recognition in the existing technology is solved, and the accuracy and performance of recognition are improved.
Patent Information
- Application Number
- CN202510414705.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-03
AI Technical Summary
In the prior art, dialogue emotion and behavior recognition is usually regarded as an independent task, and the failure to effectively explore the mutual guidance between emotions and behaviors, resulting in insufficient recognition accuracy.
Using a hierarchical interactive learning method, by constructing speaker interaction graphs, semantic interaction graphs and label interaction graphs, a graph neural network is used to achieve mutual guidance of emotions and behaviors, and multi-perspective supervision and comparison learning is used to improve recognition performance.
Through the hierarchical interactive learning method, the accuracy and performance of joint recognition of dialogue emotions and behaviors is improved, and effective mutual guidance between emotions and behaviors is achieved.
Smart Images

Figure CN119939522B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language processing and deep learning, and in particular to a method and system for jointly recognizing dialogue emotion and behavior based on hierarchical interaction learning. Background Art
[0002] Dialogue systems aim to communicate with humans through natural language and are an important research direction in the field of artificial intelligence. With the rapid development of technologies such as big data, natural language processing, and deep learning, dialogue systems have received extensive attention and are applied in multiple fields such as intelligent customer service, online education, and healthcare. Dialogue emotion and behavior recognition are key components of dialogue systems, which help generate empathetic responses, improve the user experience, and achieve deeper human-machine interaction. Dialogue emotion recognition aims to recognize the emotion (such as happiness, sadness, etc.) expressed by each utterance in a dialogue, while dialogue behavior recognition aims to recognize the behavior label (such as informing, questioning, etc.) of each utterance in a dialogue.
[0003] In recent years, dialogue emotion and behavior recognition have been widely studied. Currently, most studies usually regard emotion recognition and behavior recognition as independent tasks and process them separately, ignoring the correlation between the two tasks. Although some studies use a unified framework to jointly recognize emotions and behaviors in a dialogue, they mainly focus on the interaction at the discourse semantic level, ignoring the close association between emotion and behavior labels and not effectively exploring the mutual guidance between emotions and behaviors. For example, the emotion "happiness" is usually accompanied by the behavior "informing", while the emotion "sadness" is usually related to the behavior "questioning". Therefore, how to design an effective method for jointly recognizing dialogue emotion and behavior to achieve the mutual guidance of emotion and behavior, thereby improving the accuracy of dialogue emotion and behavior recognition, is a key issue in current dialogue system research. Summary of the Invention
[0004] To overcome the above defects in the prior art, the present invention provides a method for jointly recognizing dialogue emotion and behavior based on hierarchical interaction learning to improve the performance of jointly recognizing dialogue emotion and behavior.
[0005] To achieve the above object, the present invention adopts the following technical solutions, including:
[0006] A method for jointly recognizing dialogue emotion and behavior based on hierarchical interaction learning, including:
[0007] Obtain a dialogue dataset for model training; the dialogue consists of consecutive N utterances, and each utterance has an emotion label and a behavior label;
[0008] Construct a speaker interaction graph according to the speakers of the utterances in the dialogue and obtain the emotional utterance representation and behavioral utterance representation of each utterance;
[0009] Construct a semantic interaction graph based on the emotional utterance representation and behavioral utterance representation of each utterance, and obtain the emotional utterance representation and behavioral utterance representation after semantic interaction;
[0010] Construct a label interaction graph based on the emotional label and behavioral label of each utterance, and obtain the emotional label representation and behavioral label representation after label interaction;
[0011] Map the emotional utterance representation and behavioral utterance representation after semantic interaction to the label space to obtain the predicted emotional label and predicted behavioral label of the utterance;
[0012] Calculate the cross-entropy loss for emotion recognition according to the predicted emotional label and the true emotional label of the utterance; calculate the cross-entropy loss for behavior recognition according to the predicted behavioral label and the true behavioral label of the utterance; calculate the emotion supervision contrast loss according to the emotional utterance representation after semantic interaction; calculate the behavior supervision contrast loss according to the behavioral utterance representation after semantic interaction; calculate the emotion-behavior supervision contrast loss according to the emotional label representation and behavioral label representation after label interaction; construct the total loss of the model according to each loss, and thus perform model training;
[0013] Use the trained model to recognize the conversation, and recognize the behavioral label and emotional label of each utterance in the conversation.
[0014] Preferably, in the speaker interaction graph, utterances are used as nodes, and there are N nodes. The connection relationship between any two nodes is self-past, self-current, self-future, other-past, or other-future; among them, self and other respectively indicate that the two nodes are the same speaker and different speakers; past, future, and current indicate the relative positions of the two nodes in the conversation;
[0015] For node v i , by aggregating the neighbor node information of node v i , obtain the utterance representation h v i of node i :
[0016] ;
[0017] Among them, α ij represents the importance of node v i to neighbor node v j ; u j is the utterance vector representation corresponding to neighbor node v j ;r ij is the vector representation of the connection relationship between nodes v i and neighbor nodes v j ; u i is the corresponding discourse vector representation of the node v i ; u k is the corresponding discourse vector representation of the node v k ; r ik is the vector representation of the connection relationship between the node v i and the node v k ; N i is the set of neighbor nodes of the node v i ; || represents the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function; a and W are learnable parameters respectively; i = 1, 2,..., N ;
[0018] Process all nodes in the speaker interaction graph to obtain the discourse representation H = [h1, h2,..., h N ;
[0019] Pass the discourse representation H = [h1, h2,..., h N through the bidirectional long short-term memory network for emotion recognition and the bidirectional long short-term memory network for behavior recognition respectively, and obtain the emotion discourse representation H e = [h1 e , h2 e ,..., h N e and the behavior discourse representation H a = [h1 a , h2 a ,..., h N a ; where, h N e and h N a are the emotion discourse representation and the behavior discourse representation of the N -th discourse respectively.
[0020] Preferably, in the semantic interaction graph, the emotion discourse representation and the behavior discourse representation of each discourse are used as nodes, and there are 2 NNodes; the connection relationship between any two nodes is same-past, same-future, same-current, different-past, different-future, or different-current; where, same or different indicates that the two nodes are the same recognition task or different recognition tasks; past, future, and current indicate the relative positions of the two nodes in the conversation;
[0021] For node w i’ , by aggregating the neighbor node information of node w i’ , the discourse representation after semantic interaction of node w i’ is obtained: :
[0022] ;
[0023] Wherein, α i’j’ represents the importance of node w i’ to neighbor node w j’ ; h j’ is the emotional discourse representation or behavioral discourse representation corresponding to neighbor node w j’ ; r i’j’ is the vector representation of the connection relationship between node w i’ and neighbor node w j’ ; a’ and W’ are learnable parameters respectively; h i’ is the emotional discourse representation or behavioral discourse representation corresponding to node w i’ ; h k’ is the emotional discourse representation or behavioral discourse representation corresponding to node w k’ ; r i’k’ is the vector representation of the connection relationship between node w i’ and node w k’ ; N i’ represents the set of neighbor nodes of node w i’ ; || is the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function, i’ = 1,2,...,2 N ;
[0024] Process all nodes in the semantic interaction graph to obtain the discourse representation after semantic interaction ;
[0025] According to whether the node is an emotional discourse representation or a behavioral discourse representation, separately extract the emotional discourse representation after semantic interaction and the behavioral discourse representation ; Among them, 、 are respectively the emotional discourse representation and the behavioral discourse representation after semantic interaction of the N th discourse.
[0026] Preferably, in the label interaction graph, U emotional labels and V behavioral labels are used as nodes, and there are a total of U + V nodes; vector representations are made for each label to obtain the initial label embeddings [e1, e2,..., e U and [a1, a2,..., a V ; Among them, e U , a V are respectively the vector representation of the U th emotional label and the vector representation of the V th behavioral label;
[0027] Use label correlation to construct the correlation matrix A of the label interaction graph; the element A xy in the correlation matrix A represents the correlation between node L x and node L y , x, y = 1, 2,..., U + V ;
[0028] Use the graph convolutional network to aggregate the nodes to obtain the label representation after label interaction :
[0029] ;
[0030] Among them, represents the regularized adjacency matrix; L is the initial label embedding, L = [e1, e2,..., e U , a1, a2,..., a V ; W l is a learnable parameter; LeakReLU(·) represents the activation function;
[0031] According to whether the node is an emotional label or a behavioral label, from Extract the emotional label representation and the behavioral label representation after label interaction respectively from it; among them, ; among them, 、 are respectively the U th emotional label representation and the V th behavioral label representation after label interaction.
[0032] Preferably, the correlation matrix A is constructed as follows:
[0033] Construct a co-occurrence matrix M, and the element M xy in the co-occurrence matrix M L x is the number of times that node L y and node
[0034] co-occur; P xy Calculate the conditional probability matrix P, and the element L x in the conditional probability matrix P L y is the probability that node P xy = M xy / N x ; where, N x is the number of times that node L x appears;
[0035] Construct a binary correlation matrix A', and the value of the element A’ xy in the binary correlation matrix A' is as follows: if P xy ≥ ω , then A’ xy = 1; if P xy < ω , then A’ xy = 0; ω is the threshold;
[0036] Re-weight the binary correlation matrix A' to obtain the correlation matrix A, and the element A xy in the correlation matrix A L x is the correlation between node L yThe correlation between them is obtained in the following way:
[0037] ;
[0038] wherein, p is a hyperparameter.
[0039] Preferably, the sentiment discourse representation and the behavior discourse representation after semantic interaction are mapped to the label space by a support vector machine to obtain the mapped sentiment discourse representation and behavior discourse representation; the dot product is taken between the mapped sentiment discourse representation and the sentiment label representation after label interaction to obtain the sentiment correlation score matrix ; the dot product is taken between the mapped behavior discourse representation and the behavior label representation after label interaction to obtain the behavior correlation score matrix ;
[0040] The softmax function is used to calculate the predicted sentiment probability distribution matrix and the predicted behavior probability distribution matrix :
[0041] ;
[0042] wherein, , , and respectively represent the predicted sentiment probability distribution and the predicted behavior probability distribution of the N -th discourse, so as to obtain the predicted sentiment label and the predicted behavior label of the N -th discourse.
[0043] Preferably, the sentiment recognition cross-entropy loss and the behavior recognition cross-entropy loss are respectively:
[0044] ;
[0045] wherein, N , U and V respectively represent the number of discourses, sentiment categories and behavior categories in the dialogue; is the sign function. If the true sentiment label of the discourse u i is the x’ -th type of sentiment, then takes the value of 1, otherwise takes the value of 0; is the sign function. If the true behavior label of the discourse u i is the y’ -th type of behavior, then Take the value of 1, otherwise Take the value of 0; Is the discourse u i The predicted sentiment label of is the probability of the x’ th type of sentiment; Is the discourse u i The predicted behavior label of is the probability of the y’ th type of behavior.
[0046] Preferably, the sentiment supervised contrastive learning loss And the behavior supervised contrastive loss Are respectively:
[0047] ;
[0048] Among them, I Is the batch sample set, which contains multiple discourses, i.e., samples; Is the sample i The sentiment discourse representation after semantic interaction of;
[0049] Is the sample p The sentiment discourse representation after semantic interaction of; P e Is the positive sentiment sample set, which contains samples in the batch that have the same sentiment as ; N e Is the negative sentiment sample set, which contains samples in the batch that have different sentiments from ; Is the sample n The sentiment discourse representation after semantic interaction of; τ Is a hyperparameter;
[0050] ;
[0051] Among them, Is the behavior discourse representation after semantic interaction of the sample i ;
[0052] Is the sample p The behavior discourse representation after semantic interaction of;
[0053] P a Is the positive behavior sample set, which contains samples in the batch that have the same behavior as ;
[0054] N a Is the negative behavior sample set, which contains samples in the batch that have different behaviors from ; For the sample n The behavioral discourse representation after semantic interaction.
[0055] Preferably, the emotion-behavior supervision contrast loss Is:[[]]
[0056] ;
[0057] Wherein, I Is the batch sample set, containing multiple discourses, i.e., samples; s i For the sample i The label representation after label interaction; s p For the sample p The label representation after label interaction; P s Is the positive label sample set, containing the samples that co-occur with s i In the batch; N s Is the negative label sample set, containing the samples that do not co-occur with s i In the batch; s n For the sample n The label representation after label interaction.
[0058] The present invention also proposes a dialogue emotion and behavior joint recognition system based on hierarchical interaction learning, which is applied to the above-mentioned dialogue emotion and behavior joint recognition method based on hierarchical interaction learning. The system includes: a discourse encoding module, a semantic interaction module, a label decoding module, and a label interaction module;
[0059] The discourse encoding module is used to encode each discourse in the dialogue to obtain the vector representation of each discourse, and construct a speaker interaction graph based on the vector representation of each discourse to obtain the emotional discourse representation and behavioral discourse representation of each discourse;
[0060] The semantic interaction module is used to construct a semantic interaction graph based on the emotional discourse representation and behavioral discourse representation of each discourse, and obtain the emotional discourse representation and behavioral discourse representation after semantic interaction of each discourse;
[0061] The label decoding module is used to identify the predicted emotional label and predicted behavioral label of each discourse according to the emotional discourse representation and behavioral discourse representation after semantic interaction of each discourse;
[0062] The label interaction module is used to construct a label interaction graph according to the emotional label and behavioral label to obtain the emotional label representation and behavioral label representation after label interaction; the label interaction module participates in the calculation during the training process and does not participate in the calculation during the prediction process.
[0063] The advantages of the present invention are:
[0064] (1) To address the problem that dialogue emotion and behavior recognition have not effectively explored the mutual guidance between the two, this paper proposes a dialogue emotion and behavior joint recognition method and system based on hierarchical interactive learning.
[0065] (2) The present invention constructs a semantic interaction graph and a label interaction graph, and uses a graph neural network to achieve mutual guidance of emotions and behaviors from the two levels of semantics and label association, thereby improving the performance of joint recognition of dialogue emotions and behaviors.
[0066] (3) The present invention designs single-task and multi-task multi-view supervised contrastive learning, which improves the ability to distinguish different labels within a task from a single-task perspective and enhances the association between labels between tasks from a multi-task perspective, thereby further improving the ability to jointly recognize emotions and behaviors from different perspectives. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 This is a flow chart of the method for jointly identifying dialogue emotions and behaviors based on hierarchical interactive learning of the present invention.
[0068] Figure 2 This is a structural diagram of the conversation emotion and behavior joint recognition system based on hierarchical interactive learning of the present invention. DETAILED DESCRIPTION
[0069] The technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all of the embodiments. Based on the embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0070] The method of joint recognition of dialogue emotion and behavior based on hierarchical interactive learning is as follows: Figure 1 As shown. First, a dataset for joint recognition of dialogue emotion and behavior is obtained; then, a speaker interaction graph is constructed to model the dependency relationship between speakers; then, a semantic interaction and label interaction graph is designed to achieve mutual guidance between the two subtasks from the two levels of semantic association and label association; in addition, multi-perspective supervised contrastive learning including single-task and multi-task is proposed to further improve the ability of emotion and behavior recognition from different perspectives. Specifically, the following steps are included:
[0071] S1, obtain a dialogue dataset for joint recognition of dialogue emotions and behaviors, and train a model based on the dialogue dataset.
[0072] The publicly available dialogue dataset DailyDialog is adopted. This dataset contains 7 emotion categories, namely Neutral, Anger, Disgust, Fear, Happiness, Sadness, and Surprise, as well as 4 act categories, namely Inform, Question, Directive, and Commissive. There are 11,118 dialogues in the training set, 1,000 dialogues in the validation set, and 1,000 dialogues in the test set of the DailyDialog dataset.
[0073] S2, Utterance encoding, modeling the dependencies between speakers.
[0074] S21, Assume that the dialogue contains N consecutive utterances { u 1, u 2,..., u N} and M speakers { s 1, s 2,..., s M}, where the utterance u i is spoken by speaker . Among them, Ф is the mapping between the utterance and the speaker index. First, use the pre-trained language model to generate the vector representation of each utterance. Specifically, add the special character [CLS] in front of the utterance u i , and pass it through the RoBERTa model (an improved BERT model), and use the vector representation corresponding to [CLS] in the last layer as the vector representation u u i of the utterance i .
[0075] S22, To capture the dependencies and context relationships between speakers, construct a speaker interaction graph according to the speaker identity and the relative position of the utterances. In the speaker interaction graph, the utterances are used as nodes, and there are a total of N nodes. The node v i is initialized with u i ; the edges represent the association between the current node r ij and the node v i under the connection relationship v j , and there is an edge between any two nodes; the connection relationshipr ij Represents the current node v i and the node v j The type of relationship between, i, j = 1, 2,... N .
[0076] There are a total of 5 connection relationships in the speaker interaction graph: self - past, self - current, self - future, other - past, other - future. Self and other respectively represent the current node v i and the node v j are the same speaker and different speakers; past, future, and current respectively refer to the relative positions of the current node v i and the node v j in the conversation.
[0077] For example, the speaker interaction graph includes 3 nodes { v 1, v 2, v 3}, as Figure 2 shown, the 3 nodes { v 1, v 2, v 3} are specifically represented as the vector representations {u1, u2, u3} of the corresponding utterances. Among them, u1 and u3 are the same speaker, and u2 and u1, u3 are different speakers. If the current node is u2, the connection relationship between the current node u2 and the node u1 is other - future; if the current node is u1, the connection relationship between u1 and u3 is self - past; if the current node is u3, the connection relationship between u3 and u1 is self - future; if the current node is u2, the connection relationship between u2 and u3 is other - past; if the current node is u2, the connection relationship between u2 and u2 is self - current.
[0078] S23. To model the dependencies between speakers and learn the context relationship, a graph attention network is used to update the nodes. For the node v i , the neighbor node information of the node v i is aggregated through the following formula:
[0079] ;
[0080] where h i is the aggregated information of the node v i , serving as the information of the node vi utterance representation; α ij representing node v i importance to neighbor nodes v j ; r ij for node v i and neighbor nodes v j vector representation of the connection relationship between; a and W are learnable parameters; u r ij ; k for node v k corresponding utterance vector representation; r ik for node v i and node v k vector representation of the connection relationship between; N i for node v i set of neighbor nodes, where neighbor nodes refer to nodes that have an edge with node v i . In the present invention, the speaker interaction graph is a fully connected graph, that is, there is an edge between any two nodes. Therefore, the set of neighbor nodes of node v i contains all nodes; || is the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function.
[0081] After updating all nodes in the speaker interaction graph and obtaining the aggregated information set of each node, the utterance representation H = [h1, h2,..., h N is obtained.
[0082] S24. To be more relevant to the task, the utterance representation H is respectively passed through a bidirectional long short-term memory network for emotion recognition and a bidirectional long short-term memory network for behavior recognition, and the emotion utterance representation H e = [h1 e , h2 e ,..., h N e and the behavior utterance representation H a = [h1 a , h2 a ,..., h N a are obtained; where h N e , hN a They are respectively the emotional utterance representation and the behavioral utterance representation of the N th utterance.
[0083] S3. Semantic interaction, modeling the dependency relationship between the semantics of the two tasks.
[0084] S31. To establish the relevance between the utterance semantics under the two recognition tasks, a semantic interaction graph is constructed. There are 2 N nodes in the semantic interaction graph, including N emotional utterance nodes and N behavioral utterance nodes, which are initialized with H e and H a obtained in step S2 respectively; the edge represents the connection relationship r i’j’ under which, the current node w i’ and the node w j’ are associated; the connection relationship r i’j’ represents the relationship type between the current node w i’ and the node w j’ , i’, j’ = 1, 2,..., 2 N .
[0085] There are 6 connection relationships in the semantic interaction graph: same - past, same - future, same - current, different - past, different - future, different - current. Same and different respectively indicate that the current node w i’ and the node w j’ are for the same recognition task and different recognition tasks; past, future, and current respectively refer to the relative positions of the current node w i’ and the node w j’ in the conversation.
[0086] For example, the semantic interaction graph contains 6 nodes { w 1, w 2, w 3, w 4, w 5, w 6}, among which there are 3 emotional utterance nodes and 3 behavioral utterance nodes. As Figure 2 shown, the 6 nodes { w 1, w 2, w 3, w 4,w 5, w 6} are specifically represented as utterances u 1, u 2, u 3's emotional utterance representations [h1 e , h2 e , h3 e and behavioral utterance representations [h1 a , h2 a , h3 a . If the current node is h2 e , then the connection relationship between h2 e and h2 a is different - current; if the current node is h2 e , then the connection relationship between h2 e and h1 a is different - future; if the current node is h2 e , then the connection relationship between h2 e and h3 a is different - past; if the current node is h1 e , then the connection relationship between h1 e and h2 e is same - past; if the current node is h3 e , then the connection relationship between h3 e and h2 e is same - future.
[0087] S32, similar to the speaker interaction graph, uses a graph attention network to update the nodes in the semantic interaction graph and establish the relevance between utterance semantics under two recognition tasks. For node w i’ , the neighbor node information of node w i’ is aggregated through the following formula to obtain the utterance representation after semantic interaction of node w i’ :
[0088] ;
[0089] where, w represents the aggregated information of node w i’ and serves as the utterance representation after semantic interaction of node w i’ ; h i’ is the corresponding emotional utterance representation or behavioral utterance representation of node w i’ ; h j’ is the neighbor node α j’The corresponding emotional discourse representation or behavioral discourse representation; α i’j’ Represents the current node w i’ For neighbor nodes w j’ Importance; r i’j’ For the current node w i’ And neighbor nodes w j’ The relationship type between r i’j’ The vector representation; a' and W' are learnable parameters respectively; h k’ Is the node w k’ The corresponding emotional discourse representation or behavioral discourse representation; r i’k’ Is the node w i’ And the node w k’ The vector representation of the connection relationship between; N i’ Represents the current node w i’ The set of neighbor nodes; || is the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function.
[0090] Update all nodes in the semantic interaction graph. After obtaining the aggregated information of each node, the discourse representation after semantic interaction is obtained .
[0091] S33. According to whether the node is an emotional discourse representation or a behavioral discourse representation, extract the emotional discourse representation after semantic interaction And the behavioral discourse representation after semantic interaction ; respectively, where ; among them, , Are respectively the emotional discourse representation and behavioral discourse representation after semantic interaction of the N th discourse.
[0092] S4. Label interaction, modeling the dependencies and correlations between labels.
[0093] S41. Assume that the emotional labels and behavioral labels contain U And V Respectively, that is, { e 1, e 2,..., e U} and { a 1, a2,..., a V}. First, use RoBERTa to represent each label as a vector, obtaining the initial label embeddings [e1, e2,..., e U and [a1, a2,..., a V ; where e U , a V are the vector representations of the U th sentiment label and the V th behavior label respectively;
[0094] S42. Construct a label interaction graph to model the dependencies between labels. In the label interaction graph, labels are used as nodes, and there are a total of U + V nodes. The correlation matrix A of the label interaction graph is constructed using label co-occurrence.
[0095] Construct a co-occurrence matrix M based on the total number of times sentiment labels and behavior labels co-occur in the training set. The element M xy in the co-occurrence matrix M is the number of times node L x and node L y co-occur (co-occur). x, y = 1, 2,..., U + V . If both node L x and node L y are sentiment labels or both are behavior labels, then the number of times node L x and node L y co-occur is 0. The element P xy in the conditional probability matrix P = M xy / N x . N x is the total number of times node L x appears in the training set, and P xy represents the probability that node L x appears when node L y appears.
[0096] Since node L x and node Ly The co-occurrence between them may follow a long-tail distribution, and some rare co-occurrences may introduce noise. Therefore, using a threshold ω to filter out noisy label pairs. In other words, there may be many terms in the co-occurrence matrix M with very small or insignificant co-occurrence frequencies. Thus, those terms with higher co-occurrence frequencies can be retained through threshold filtering. By threshold filtering, unimportant or insignificant label pairs are removed, thereby reducing noise and improving the effectiveness of the data. Specifically, using the threshold ω to filter out noisy label pairs, thus obtaining the binary correlation matrix A':
[0097] ;
[0098] To prevent the node features from being overly smoothed, a reweighting method is designed, and finally the correlation matrix A is obtained:
[0099] ;
[0100] where the element in the correlation matrix A A xy represents the correlation between node L x and node L y ; p is a hyperparameter that determines the weights assigned to the node itself and other related nodes.
[0101] S43. Using a graph convolutional network to aggregate the label nodes to obtain the label representation (matrix) after label interaction :
[0102] ;
[0103] where represents the regularized adjacency matrix; L is the initial label embedding, L = [e1, e2,..., e U , a1, a2,..., a V ; W l are learnable parameters; LeakReLU(·) represents the activation function.
[0104] S44. According to whether the node is an emotion label or a behavior label, respectively extract the emotion label representation after label interaction and the behavior label representation after label interaction from ; where and are respectively the U th emotion label representation and the V th behavior label representation after label interaction.
[0105] S5, Label decoding.
[0106] Map the emotional discourse representation and behavioral discourse representations after semantic interaction to the label space respectively, and then calculate the correlation score between the discourse and the label using the dot product. The higher the score, the closer the distance between the discourse and the label representation, and the more likely the discourse belongs to this category.
[0107] S51, Since the emotional discourse representation and behavioral discourse representation after semantic interaction are not in the same vector space as the emotional label representation and behavioral label representation after label interaction, the emotional discourse representation and behavioral discourse representation after semantic interaction are mapped to the joint label space where the emotional label representation and behavioral label representation after label interaction are located through a support vector machine, respectively, so as to obtain the mapped emotional discourse representation and behavioral discourse representation. Subsequently, the dot product is performed between the mapped emotional discourse representation and the emotional label representation after label interaction to obtain the emotional correlation score matrix ; perform the dot product between the mapped behavioral discourse representation and the behavioral label representation after label interaction to obtain the behavioral correlation score matrix .
[0108] S52, Apply the softmax function to the emotional correlation score matrix and the behavioral correlation score matrix respectively to obtain the predicted emotional probability distribution matrix and the predicted behavioral probability distribution matrix :
[0109] ;
[0110] where , , and respectively represent the predicted emotional probability distribution and predicted behavioral probability distribution of the N th discourse.
[0111] S53, Use the argmax function to calculate the predicted emotional probability distribution i of the u i th discourse and the predicted behavioral probability distribution to obtain the predicted emotional label and predicted behavioral label of the discourse u i .
[0112] S6, Calculate the total loss of the model and perform model training.
[0113] S61, Emotional recognition cross-entropy loss and behavioral recognition cross-entropy loss They are respectively:
[0114] ;
[0115] Among them, N 、 U and V respectively represent the number of utterances, emotion categories, and behavior categories in the dialogue; is the sign function. If the true emotion label of the utterance u i is the x’ -th emotion category, then takes the value of 1, otherwise takes the value of 0; is the sign function. If the true behavior label of the utterance u i is the y’ -th behavior category, then takes the value of 1, otherwise takes the value of 0; is the probability that the predicted emotion label of the utterance u i is the x’ -th emotion category; is the probability that the predicted behavior label of the utterance u i is the y’ -th behavior category.
[0116] S62. Based on the multi-level interaction method, multi-view supervised contrast learning is proposed to improve the ability to distinguish different labels within a task from a single-task perspective and further enhance the association of labels between tasks from a multi-task perspective.
[0117] Emotion supervised contrast learning loss is:
[0118] ;
[0119] Among them, I is the batch sample set, which contains multiple utterances, i.e., samples; is the emotion utterance representation after semantic interaction of the sample i ;
[0120] is the emotion utterance representation after semantic interaction of the sample p ; P e is the positive emotion sample set, which contains samples in the batch that have the same emotion as ; N e is the negative emotion sample set, which contains samples in the batch that have different emotions from ; is the emotional discourse representation after semantic interaction of the sample n ; τ is a hyperparameter.
[0121] Similarly, the behavior supervision contrast loss is obtained as:
[0122] ;
[0123] where is the behavior discourse representation after semantic interaction of the sample i ;
[0124] is the behavior discourse representation after semantic interaction of the sample p ;
[0125] P a is the set of positive behavior samples, containing the samples in the batch that have the same behavior as ;
[0126] N a is the set of negative behavior samples, containing the samples in the batch that have different behaviors from ;
[0127] is the behavior discourse representation after semantic interaction of the sample n ; τ is a hyperparameter.
[0128] The emotion - behavior supervision contrast loss from the multi - task perspective is:
[0129] ;
[0130] where I is the batch sample set, containing multiple discourses, i.e., samples; s i is the (emotional or behavioral) label representation after label interaction of the sample i ; s p is the label representation after label interaction of the sample p ; P s is the set of positive label samples, containing the samples in the batch that co - occur with s i ; N s is the set of negative label samples, containing the samples in the batch that do not co - occur with s i ; s n is the label representation after label interaction of the sample n ;
[0131] S63, the total loss of the model during training is:
[0132] ;
[0133] S7, Joint recognition of dialogue sentiment and behavior
[0134] After training is completed, the trained model is used to jointly recognize dialogue sentiment and behavior. The input is a dialogue containing N consecutive utterances, and the output is the predicted sentiment label and predicted behavior label for each utterance in the dialogue.
[0135] The method of the present invention uses a graph neural network to achieve mutual guidance between sentiment and behavior from two levels of semantics and label association, improving the performance of joint recognition of dialogue sentiment and behavior; and proposes multi-perspective supervised contrast learning for single-task and multi-task, enhancing the ability to distinguish different labels within a task from the single-task perspective and strengthening the label association between tasks from the multi-task perspective, thereby further improving the ability of joint recognition of sentiment and behavior from different angles.
[0136] The present invention also provides a system for joint recognition of dialogue sentiment and behavior based on hierarchical interaction learning. The system includes: an utterance encoding module, a semantic interaction module, a label decoding module, and a label interaction module.
[0137] The utterance encoding module is used to encode each utterance in the dialogue to obtain a vector representation of each utterance, and construct a speaker interaction graph based on the vector representation of each utterance to obtain an emotional utterance representation and a behavioral utterance representation of each utterance.
[0138] The semantic interaction module is used to construct a semantic interaction graph based on the emotional utterance representation and behavioral utterance representation of each utterance, and obtain the emotional utterance representation and behavioral utterance representation after semantic interaction of each utterance.
[0139] The label decoding module is used to identify the predicted sentiment label and predicted behavior label of each utterance based on the emotional utterance representation and behavioral utterance representation after semantic interaction of each utterance.
[0140] The label interaction module is used to construct a label interaction graph based on the sentiment label and behavior label to obtain the sentiment label representation and behavior label representation after label interaction; the label interaction module participates in the calculation during the training process and does not participate in the calculation during the prediction process.
[0141] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for joint recognition of dialogue emotions and behaviors based on hierarchical interactive learning, characterized in that: include: Obtain a conversation dataset for model training; The dialogue consists of continuous N Each utterance has a sentiment label and a behavior label. According to the speakers of the utterances in the dialogue, a speaker interaction graph is constructed and the emotional utterance representation and behavioral utterance representation of each utterance are obtained; According to the emotional discourse representation and behavioral discourse representation of each discourse, a semantic interaction graph is constructed and the emotional discourse representation and behavioral discourse representation after semantic interaction are obtained; According to the emotion labels and behavior labels of each utterance, a label interaction graph is constructed to obtain the emotion label representation and behavior label representation after label interaction; Map the emotional discourse representation and behavioral discourse representation after semantic interaction to the label space to obtain the predicted emotional label and predicted behavioral label of the discourse; Calculate the emotion recognition cross entropy loss based on the predicted emotion label and the real emotion label of the utterance; calculate the behavior recognition cross entropy loss based on the predicted behavior label and the real behavior label of the utterance; calculate the emotion supervision contrast loss based on the emotion utterance representation after semantic interaction; According to the behavioral discourse representation after semantic interaction, the behavioral supervision contrast loss is calculated; according to the emotional label representation and behavioral label representation after label interaction, the emotion-behavior supervision contrast loss is calculated; Construct the total loss of the model based on each loss to train the model; The trained model is used to recognize the conversation and identify the behavior label and emotion label of each utterance in the conversation.
2. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: The speaker interaction graph uses utterances as nodes, with a total of N nodes, and the connection relationship between any two nodes is self-past, self-present, self-future, others-past or others-future; among them, self and others respectively indicate that the two nodes are the same speaker and different speakers; past, future and present indicate the relative positions of the two nodes in the conversation; For nodes v i , through the aggregation node v i The neighbor node information of node v i The words expressed h i : ; in, α ij Representation Node v i For neighbor nodes v j The importance of j Neighbor node v j The corresponding utterance vector representation; r ij For Node v i and neighbor nodes v j The vector representation of the connection relationship between i For Node v i The corresponding discourse vector representation; u k For Node v k The corresponding utterance vector representation; r ik For Node v i and nodes v k The vector representation of the connection relationship between them; N i For Node v i The set of neighbor nodes; || represents the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function; a and W are learnable parameters respectively; i =1,2,..., N ; Process all nodes in the speaker interaction graph to obtain the discourse representation H=[h1,h2,...,h N ]; Let the utterance be represented by H=[h1,h2,...,h N ] The emotional speech representation H is obtained by using a bidirectional long short-term memory network for emotion recognition and a bidirectional long short-term memory network for behavior recognition. e =[h1 e ,h2 e ,...,h N e ] and behavioral discourse representation H a =[h1 a ,h2 a ,...,h N a ]; where h N e 、h N a Respectively N The emotional discourse representation and behavioral discourse representation of each utterance.
3. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: The semantic interaction graph uses the emotional discourse representation and behavioral discourse representation of each discourse as nodes, with a total of 2 N nodes; the connection relationship between any two nodes is same-past, same-future, same-current, different-past, different-future or different-current; where same or different means that the two nodes are the same recognition task or different recognition tasks; past, future and current represent the relative positions of the two nodes in the conversation; For nodes w i’ , through the aggregation node w i’ The neighbor node information of node w i’ The discourse representation after semantic interaction : ; in, α i’j’ Representation Node w i’ For neighbor nodes w j’ The importance of j’ Neighbor node w j’ The corresponding emotional discourse representation or behavioral discourse representation; r i’j’ For Node w i’ and neighbor nodes w j’ The vector representation of the connection relationship between them; a' and W' are learnable parameters; h i’ For Node w i’ The corresponding emotional discourse expression or behavioral discourse expression; h k’ For Node w k’ The corresponding emotional discourse representation or behavioral discourse representation; r i’k’ For Node w i’ and nodes w k’ The vector representation of the connection relationship between them; N i’ Representation Node w i’ The set of neighbor nodes; || is the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function, i’ =1,2,...,2 N ; Process all nodes in the semantic interaction graph to obtain the discourse representation after semantic interaction ; According to whether the node is a sentiment discourse representation or a behavior discourse representation, Extract the emotional discourse representation after semantic interaction respectively and behavioral discourse representation ;in, , Respectively N The emotional discourse representation and behavioral discourse representation after the semantic interaction of the discourses.
4. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: The label interaction diagram is used U emotional labels and V behavior labels as nodes, a total of U + V nodes; each label is represented by a vector, and the initial label embedding [e1,e2,...,e U ] and [a1,a2,...,a V ]; where e U 、a V Respectively U The sentiment label vector representation and V Behavior label vector representation; The correlation matrix A of the label interaction graph is constructed using label correlation; the elements in the correlation matrix A A xy Representation Node L x and nodes L y The correlation between x,y =1,2,..., U + V ; Use graph convolutional networks to aggregate nodes and obtain label representations after label interaction : ; in, represents the regularized adjacency matrix; L is the initialization label embedding, L=[e1,e2,...,e U ,a1,a2,...,a V ]; W l is a learnable parameter; LeakReLU(·) represents the activation function; According to whether the node is a sentiment label or a behavior label, Extract the sentiment label representation after label interaction and behavior labels ;in, , They are the first U The sentiment label represents V Behavior labels are represented.
5. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 4 is characterized in that: The correlation matrix A is constructed as follows: Construct the co-occurrence matrix M. The elements in the co-occurrence matrix M M xy For Node L x and nodes L y the number of co-occurrences; Calculate the conditional probability matrix P, the elements in the conditional probability matrix P P xy For Node L x When the node appears L y The probability of occurrence, P xy = M xy / N x ;in, N x For Node L x Number of occurrences; Construct the binary correlation matrix A'. The elements in the binary correlation matrix A' A’ xy The value of is: if P xy ≥ ω ,but A’ xy =1; if P xy < ω ,but A’ xy =0; ω is the threshold value; The binary correlation matrix A' is reweighted to obtain the correlation matrix A. The elements in the correlation matrix A are A xy For Node L x and nodes L y The correlation between them is obtained as follows: ; in, p is a hyperparameter.
6. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: The emotional discourse representation and behavioral discourse representation after semantic interaction are mapped to the label space through support vector machine to obtain the mapped emotional discourse representation and behavioral discourse representation; the mapped emotional discourse representation is dot-producted with the emotional label representation after label interaction to obtain the emotional correlation score matrix ; Perform dot product between the mapped behavior discourse representation and the behavior label representation after label interaction to obtain the behavior-related score matrix ; Using the softmax function, calculate the predicted emotion probability distribution matrix And the predicted behavior probability distribution matrix : ; in, , , and Respectively represent N The predicted emotion probability distribution and predicted behavior probability distribution of each utterance are obtained. N The predicted sentiment label and predicted behavior label of each utterance.
7. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: Cross Entropy Loss for Emotion Recognition and cross entropy loss for action recognition They are: ; in, N , U and V Respectively represent the number of utterances, emotion categories, and behavior categories in the conversation; is a symbolic function, if the utterance u i The true emotional label is x’ Emotional The value is 1, otherwise The value is 0; is a symbolic function, if the utterance u i The actual behavior label is y’ Class behavior, The value is 1, otherwise The value is 0; For discourse u i The predicted sentiment label is x’ Probability of class sentiment; For discourse u i The predicted behavior label is y’ The probability of class behavior.
8. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: Sentiment Supervised Contrastive Learning Loss Comparative loss with behavior supervision They are: ; in, I It is a batch sample set, which contains multiple utterances, i.e., samples; For sample i The emotional discourse representation after semantic interaction; For sample p The emotional discourse representation after semantic interaction; P e is the emotional positive sample set, including Samples with the same emotions; N e is the emotional negative sample set, including There are samples with different emotions; For sample n The emotional discourse representation after semantic interaction; τ is a hyperparameter; ; in, For sample i The behavioral discourse representation after the semantic interaction; For sample p The behavioral discourse representation after the semantic interaction; P a is the behavior positive sample set, including Samples with the same behavior; N a is the behavior negative sample set, including There are samples of different behaviors; For sample n The behavioral discourse representation after semantic interaction.
9. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: Sentiment-behavior supervised contrastive loss for: ; in, I is a batch sample set, which contains multiple utterances, i.e., samples; i For sample i The label representation after the label interaction; s p For sample p Label representation after label interaction; P s is the label positive sample set, including the batch with s i co-occurring samples; N s is a set of labeled negative samples, including i Samples that do not co-occur; n For sample n The label representation after label interaction.
10. A dialogue emotion and behavior joint recognition system based on hierarchical interactive learning, characterized by: The method for joint recognition of dialogue emotions and behaviors based on hierarchical interactive learning applied to any one of claims 1 to 9 above, the system comprising: a discourse encoding module, a semantic interaction module, a label decoding module, and a label interaction module; The speech encoding module is used to encode each speech in the conversation to obtain a vector representation of each speech, and to construct a speaker interaction graph based on the vector representation of each speech to obtain an emotional speech representation and a behavioral speech representation of each speech; The semantic interaction module is used to construct a semantic interaction graph according to the emotional discourse representation and behavioral discourse representation of each discourse, and obtain the emotional discourse representation and behavioral discourse representation after the semantic interaction of each discourse; The label decoding module is used to identify the predicted emotion label and predicted behavior label of each utterance according to the emotion utterance representation and behavior utterance representation after the semantic interaction of each utterance; The label interaction module is used to construct a label interaction graph according to the emotion label and the behavior label to obtain the emotion label representation and the behavior label representation after the label interaction; the label interaction module participates in the calculation during the training process and does not participate in the calculation during the prediction process.
Citation Information
Patent Citations
Emotion guiding method and system based on emotion semantic transfer graph
CN111914556A
Conversation emotion recognition method based on common sense perception and hierarchical multi-task learning
CN114722838A