Conversation emotion and behavior combined recognition method and system based on hierarchical interactive learning

By constructing interactive graphs and utilizing graph neural networks, mutual guidance between dialogue emotions and behaviors is achieved, and the problem of insufficient accuracy of dialogue emotions and behavior recognition in the prior art is solved, and the recognition performance is improved.

CN119939522AActive Publication Date: 2025-05-06HEFEI UNIV OF TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510414705.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing technology has failed to effectively explore the mutual guidance between dialogue emotions and behaviors, resulting in insufficient accuracy in dialogue emotions and behavior recognition.

Method used

Using a hierarchical interactive learning method, by constructing speaker interaction graphs, semantic interaction graphs and label interaction graphs, a graph neural network is used to achieve mutual guidance of emotions and behaviors, and multi-perspective supervision and comparison learning is used to improve recognition performance.

Benefits of technology

The accuracy of joint recognition of dialogue emotions and behaviors is improved, and the correlation between tags between tasks is enhanced through multi-perspective supervision and comparison learning, and the recognition effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939522A_ABST
    Figure CN119939522A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue emotion and behavior combined recognition method and system based on hierarchical interactive learning, and relates to the technical field of natural language processing and deep learning, and the method comprises the steps: obtaining a dialogue data set; constructing a speaker interaction graph to obtain an emotion utterance representation and a behavior utterance representation; constructing a speaker interaction graph and obtaining emotion utterance representation and behavior utterance representation of each utterance; constructing a semantic interaction graph, obtaining an emotion utterance representation and a behavior utterance representation after semantic interaction, and mapping the emotion utterance representation and the behavior utterance representation to a tag space to obtain a predicted emotion tag and a predicted behavior tag of the utterance; constructing a label interaction graph and obtaining emotion label representation and behavior label representation after label interaction; based on single-task and multi-task multi-view supervised contrast learning, calculating the total loss of the model to carry out model training; and recognizing the dialogue by using the trained model, and recognizing a behavior tag and an emotion tag of each utterance in the dialogue. According to the invention, the performance of joint recognition of dialogue emotion and behaviors is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing and deep learning technology, and in particular to a method and system for jointly identifying dialogue emotions and behaviors based on hierarchical interactive learning. Background Art

[0002] The dialogue system aims to communicate with humans through natural language and is an important research direction in the field of artificial intelligence. With the rapid development of technologies such as big data, natural language processing, and deep learning, dialogue systems have received widespread attention and have been applied to many fields such as intelligent customer service, online education, and healthcare. Dialogue emotion and behavior recognition are key components of dialogue systems, which help generate empathetic responses, improve user experience, and achieve deeper human-computer interaction. Dialogue emotion recognition aims to identify the emotion expressed by each utterance in the dialogue (such as happiness, sadness, etc.), while dialogue behavior recognition aims to identify the behavior label of each utterance in the dialogue (such as telling, asking, etc.).

[0003] In recent years, conversation emotion and behavior recognition has been widely studied. At present, most studies usually regard emotion recognition and behavior recognition as independent tasks and handle them separately, ignoring the correlation between the two tasks. Although some studies use a unified framework to jointly identify emotions and behaviors in conversations, they mainly focus on the interaction at the semantic level of discourse, ignore the close relationship between emotion and behavior labels, and do not effectively explore the mutual guidance between emotions and behaviors. For example, the emotion "happy" is usually accompanied by the behavior "telling", while the emotion "sad" is usually related to the behavior "asking questions". Therefore, how to design an effective method for joint recognition of conversation emotions and behaviors, realize the mutual guidance of emotions and behaviors, and thus improve the accuracy of conversation emotion and behavior recognition is a key issue in the current research on dialogue systems. Summary of the invention

[0004] In order to overcome the defects in the above-mentioned prior art, the present invention provides a method for joint recognition of dialogue emotions and behaviors based on hierarchical interactive learning to improve the performance of joint recognition of dialogue emotions and behaviors.

[0005] To achieve the above object, the present invention adopts the following technical solutions, including: The method of joint recognition of dialogue emotion and behavior based on hierarchical interactive learning includes: Get the conversation dataset for model training; the conversation consists of continuous N Each utterance has a sentiment label and a behavior label. According to the speakers of the utterances in the dialogue, a speaker interaction graph is constructed and the emotional utterance representation and behavioral utterance representation of each utterance are obtained; According to the emotional discourse representation and behavioral discourse representation of each discourse, a semantic interaction graph is constructed and the emotional discourse representation and behavioral discourse representation after semantic interaction are obtained; According to the emotion labels and behavior labels of each utterance, a label interaction graph is constructed to obtain the emotion label representation and behavior label representation after label interaction; Map the emotional discourse representation and behavioral discourse representation after semantic interaction to the label space to obtain the predicted emotional label and predicted behavioral label of the discourse; According to the predicted emotion label and the real emotion label of the discourse, the emotion recognition cross entropy loss is calculated; according to the predicted behavior label and the real behavior label of the discourse, the behavior recognition cross entropy loss is calculated; according to the emotion discourse representation after semantic interaction, the emotion supervision contrast loss is calculated; according to the behavior discourse representation after semantic interaction, the behavior supervision contrast loss is calculated; according to the emotion label representation and the behavior label representation after label interaction, the emotion-behavior supervision contrast loss is calculated; according to each loss, the total loss of the model is constructed to train the model; The trained model is used to recognize the conversation and identify the behavior label and emotion label of each utterance in the conversation.

[0006] Preferably, the speaker interaction graph uses utterances as nodes, with a total of N nodes, and the connection relationship between any two nodes is self-past, self-present, self-future, others-past or others-future; among them, self and others respectively indicate that the two nodes are the same speaker and different speakers; past, future and present indicate the relative positions of the two nodes in the conversation; For nodes v i , through the aggregation node v i The neighbor node information of node v i The words h i : ; in, α ij Representation Node v i For neighbor nodes v j The importance of j Neighbor node v j The corresponding utterance vector representation; r ij For Node v i and neighbor nodes v j The vector representation of the connection relationship betweeni For Node v i The corresponding discourse vector representation; u k For Node v k The corresponding utterance vector representation; r ik For Node v i and nodes v k The vector representation of the connection relationship between them; N i For Node v i The set of neighbor nodes; || represents the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function; a and W are learnable parameters respectively; i = 1,2,..., N ; Process all nodes in the speaker interaction graph to obtain the discourse representation H=[h1,h2,...,h N ]; Let the utterance be represented by H=[h1,h2,...,h N ] The emotional speech representation H is obtained by using a bidirectional long short-term memory network for emotion recognition and a bidirectional long short-term memory network for behavior recognition. e =[h1 e ,h2 e ,...,h N e ] and behavioral discourse representation H a =[h1 a ,h2 a ,...,h N a ]; where h N e 、h N a Respectively N The emotional discourse representation and behavioral discourse representation of each utterance.

[0007] Preferably, the semantic interaction graph uses the emotional discourse representation and behavioral discourse representation of each discourse as nodes, with a total of 2 N nodes; the connection relationship between any two nodes is same-past, same-future, same-current, different-past, different-future or different-current; where same or different means that the two nodes are the same recognition task or different recognition tasks; past, future and current represent the relative positions of the two nodes in the conversation; For nodes w i’, through the aggregation node w i’ The neighbor node information of node w i’ The discourse representation after semantic interaction : ; in, α i’j’ Representation Node w i’ For neighbor nodes w j’ The importance of j’ Neighbor node w j’ The corresponding emotional discourse representation or behavioral discourse representation; r i’j’ For Node w i’ and neighbor nodes w j’ The vector representation of the connection relationship between them; a' and W' are learnable parameters; h i’ For Node w i’ The corresponding emotional discourse expression or behavioral discourse expression; h k’ For Node w k’ The corresponding emotional discourse representation or behavioral discourse representation; r i’k’ For Node w i’ and nodes w k’ The vector representation of the connection relationship between them; N i’ Representation Node w i’ The set of neighbor nodes; || is the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function, i’ = 1,2,...,2 N ; Process all nodes in the semantic interaction graph to obtain the discourse representation after semantic interaction ; According to whether the node is a sentiment utterance or a behavior utterance, Extract the emotional discourse representation after semantic interaction respectively and behavioral discourse representation ;in, , Respectively N The emotional discourse representation and behavioral discourse representation after the semantic interaction of the discourses.

[0008] Preferably, the label interaction diagram is U emotional labels and V behavior labels as nodes, a total of U + V nodes; each label is represented by a vector, and the initial label embedding [e1,e2,...,e U ] and [a1,a2,...,a V ]; where e U 、a V Respectively U The sentiment label vector representation and V Behavior label vector representation; The correlation matrix A of the label interaction graph is constructed using label correlation; the elements in the correlation matrix A A xy Representation Node L x and nodes L y The correlation between x,y = 1,2,..., U + V ; Use graph convolutional networks to aggregate nodes and obtain label representations after label interaction : ; in, represents the regularized adjacency matrix; L is the initialization label embedding, L=[e1,e2,...,e U ,a1,a2,...,a V ]; W l is a learnable parameter; LeakReLU(·) represents the activation function; According to whether the node is a sentiment label or a behavior label, Extract the sentiment label representation after label interaction and behavior labels ;in, , They are the first U The sentiment label represents V Behavior labels.

[0009] Preferably, the correlation matrix A is constructed as follows: Construct the co-occurrence matrix M. The elements in the co-occurrence matrix M M xy For Node L x and nodes L ythe number of co-occurrences; Calculate the conditional probability matrix P, the elements in the conditional probability matrix P P xy For Node L x When the node appears L y The probability of occurrence, P xy = M xy / N x ;in, N x For Node L x Number of occurrences; Construct the binary correlation matrix A'. The elements in the binary correlation matrix A' A’ xy The value of is: if P xy ≥ oh ,but A’ xy =1; if P xy < oh ,but A’ xy =0; oh is the threshold value; The binary correlation matrix A' is reweighted to obtain the correlation matrix A. The elements in the correlation matrix A are A xy For Node L x and nodes L y The correlation between them is obtained as follows: ; in, p is a hyperparameter.

[0010] Preferably, the emotional discourse representation and the behavioral discourse representation after semantic interaction are mapped to the label space through a support vector machine to obtain the mapped emotional discourse representation and the behavioral discourse representation; the mapped emotional discourse representation and the emotional label representation after label interaction are dot-producted to obtain the emotional correlation score matrix ; Perform dot product between the mapped behavior discourse representation and the behavior label representation after label interaction to obtain the behavior-related score matrix ; Using the softmax function, calculate the predicted emotion probability distribution matrix And the predicted behavior probability distribution matrix : ; in, , , and Respectively represent N The predicted emotion probability distribution and predicted behavior probability distribution of each utterance are obtained. N The predicted sentiment label and predicted behavior label of each utterance.

[0011] Preferably, sentiment recognition cross entropy loss and cross entropy loss for action recognition They are: ; in, N , U and V Respectively represent the number of utterances, emotion categories, and behavior categories in the conversation; is a symbolic function, if the utterance u i The true emotional label is x’ Emotional The value is 1, otherwise The value is 0; is a symbolic function, if the utterance u i The actual behavior label is y’ Class behavior, The value is 1, otherwise The value is 0; For discourse u i The predicted sentiment label is x’ Probability of class sentiment; For discourse u i The predicted behavior label is y’ The probability of class behavior.

[0012] Preferably, sentiment supervised contrastive learning loss Comparative loss with behavior supervision They are: ; in, I It is a batch sample set, which contains multiple utterances, i.e., samples; For sample i The emotional discourse representation after semantic interaction; For sample p The emotional discourse representation after semantic interaction; P eis the emotional positive sample set, including Samples with the same emotions; N e is the emotional negative sample set, including There are samples with different emotions; For sample n The emotional discourse representation after semantic interaction; t is a hyperparameter; ; in, For sample i The behavioral discourse representation after the semantic interaction; For sample p The behavioral discourse representation after the semantic interaction; P a is the behavior positive sample set, including Samples with the same behavior; N a is the behavior negative sample set, including There are samples of different behaviors; For sample n The behavioral discourse representation after semantic interaction.

[0013] Preferably, sentiment-behavior supervised contrastive loss for: ; in, I is a batch sample set, which contains multiple utterances, i.e., samples; i For sample i The label representation after the label interaction; s p For sample p Label representation after label interaction; P s is the label positive sample set, including the batch with s i co-occurring samples; N s is a set of labeled negative samples, including those in the batch i Samples that do not co-occur; n For sample n The label representation after label interaction.

[0014] The present invention also proposes a dialogue emotion and behavior joint recognition system based on hierarchical interactive learning, which is applied to the above-mentioned dialogue emotion and behavior joint recognition method based on hierarchical interactive learning. The system includes: a discourse encoding module, a semantic interaction module, a label decoding module, and a label interaction module; The speech encoding module is used to encode each speech in the conversation to obtain a vector representation of each speech, and to construct a speaker interaction graph based on the vector representation of each speech to obtain an emotional speech representation and a behavioral speech representation of each speech; The semantic interaction module is used to construct a semantic interaction graph according to the emotional discourse representation and behavioral discourse representation of each discourse, and obtain the emotional discourse representation and behavioral discourse representation after the semantic interaction of each discourse; The label decoding module is used to identify the predicted emotion label and predicted behavior label of each utterance according to the emotion utterance representation and behavior utterance representation after the semantic interaction of each utterance; The label interaction module is used to construct a label interaction graph according to the emotion label and the behavior label to obtain the emotion label representation and the behavior label representation after the label interaction; the label interaction module participates in the calculation during the training process and does not participate in the calculation during the prediction process.

[0015] The advantages of the present invention are: (1) To address the problem that dialogue emotion and behavior recognition have not effectively explored the mutual guidance between the two, this paper proposes a dialogue emotion and behavior joint recognition method and system based on hierarchical interactive learning.

[0016] (2) The present invention constructs a semantic interaction graph and a label interaction graph, and uses a graph neural network to achieve mutual guidance of emotions and behaviors from the two levels of semantics and label association, thereby improving the performance of joint recognition of dialogue emotions and behaviors.

[0017] (3) The present invention designs single-task and multi-task multi-view supervised contrastive learning, which improves the ability to distinguish different labels within a task from a single-task perspective and enhances the association between labels between tasks from a multi-task perspective, thereby further improving the ability to jointly recognize emotions and behaviors from different perspectives. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of the method for jointly identifying dialogue emotions and behaviors based on hierarchical interactive learning of the present invention.

[0019] Figure 2 This is a structural diagram of the dialogue emotion and behavior joint recognition system based on hierarchical interactive learning of the present invention. DETAILED DESCRIPTION

[0020] The technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all of the embodiments. Based on the embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0021] The method of joint recognition of dialogue emotion and behavior based on hierarchical interactive learning is as follows: Figure 1 As shown. First, a dataset for joint recognition of dialogue emotion and behavior is obtained; then, a speaker interaction graph is constructed to model the dependency relationship between speakers; then, a semantic interaction and label interaction graph is designed to achieve mutual guidance between the two subtasks from the two levels of semantic association and label association; in addition, multi-perspective supervised contrastive learning including single-task and multi-task is proposed to further improve the ability of emotion and behavior recognition from different perspectives. Specifically, the following steps are included: S1, obtain a dialogue dataset for joint recognition of dialogue emotions and behaviors, and train a model based on the dialogue dataset.

[0022] The public dialogue dataset DailyDialog is used. This dataset contains 7 emotion categories: Neutral, Anger, Disgust, Fear, Happiness, Sadness, and Surprise, and 4 behavior categories: Inform, Question, Directive, and Commissive. The training set of the DailyDialog dataset contains 11,118 dialogues, the validation set contains 1,000 dialogues, and the test set contains 1,000 dialogues.

[0023] S2, discourse encoding, models dependencies between speakers.

[0024] S21, assuming the conversation contains N A continuous utterance u 1, u 2,..., u N}and M Speaker s 1, s 2,..., s M}, words u i Is the speaker Said, among them, F is a mapping between utterances and speaker indices. First, a pre-trained language model is used to generate a vector representation of each utterance. Specifically, u i Add the special character [CLS] in front and use the vector representation corresponding to [CLS] in the last layer as the discourse through the RoBERTa model (an improved BERT model) u i The vector representation ui .

[0025] S22, in order to capture the dependencies and contextual relationships between speakers, a speaker interaction graph is constructed based on the speaker identities and the relative positions of the utterances. The speaker interaction graph uses utterances as nodes, with a total of N nodes, nodes v i Use u i Initialize the representation; the edge represents the connection relationship r ij Next, the current node v i and nodes v j There is an edge between any two nodes; connection relationship r ij Indicates the current node v i and nodes v j The type of relationship between i, j = 1,2,..., N .

[0026] There are five types of connection relationships in the speaker interaction graph: self-past, self-present, self-future, others-past, and others-future. Self and others represent the current node v i and nodes v j for the same speaker and different speakers; past, future and present refer to the current node respectively v i and nodes v j The relative position in the conversation.

[0027] For example, the speaker interaction graph consists of three nodes { v 1, v 2, v 3}, such as Figure 2 As shown, 3 nodes { v 1, v 2, v3} is specifically represented as the vector representation of the corresponding speech {u1,u2,u3}, where u1 and u3 are the same speaker, and u2 and u1, u3 are different speakers. If the current node is u2, the connection relationship between the current node u2 and node u1 is others-future; if the current node is u1, the connection relationship between u1 and u3 is self-past; if the current node is u3, the connection relationship between u3 and u1 is self-future; if the current node is u2, the connection relationship between u2 and u3 is others-past; if the current node is u2, the connection relationship between u2 and u2 is self-current.

[0028] S23, in order to model the dependencies between speakers and learn contextual relationships, the graph attention network is used to update the nodes. v i , aggregate nodes through the following formula v i Neighbor node information: ; Among them, h i For Node v i The aggregated information of v i The words expressed; α ij Representation Node v i For neighbor nodes v j The importance of r ij For Node v i and neighbor nodes v j Connection relationship r ij vector representation; a and W are learnable parameters; u k For Node v k The corresponding utterance vector representation; r ik For Node v i and nodes v k The vector representation of the connection relationship between them; N i For Node v i The neighbor node set of node v i In the present invention, the speaker interaction graph is a fully connected graph, that is, there is an edge between any two nodes. Therefore, the node vi The neighbor node set of includes all nodes; || is the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function.

[0029] After updating all nodes in the speaker interaction graph and obtaining the aggregated information set of each node, the discourse representation H=[h1,h2,...,h N ].

[0030] S24, in order to be more relevant to the task, the discourse representation H is passed through the bidirectional long short-term memory network for emotion recognition and the bidirectional long short-term memory network for behavior recognition to obtain the emotional discourse representation H e =[h1 e ,h2 e ,...,h N e ] and behavioral discourse representation H a =[h1 a ,h2 a ,...,h N a ]; where h N e 、h N a Respectively N The emotional discourse representation and behavioral discourse representation of each utterance.

[0031] S3, semantic interaction, models the dependency between the semantics of two tasks.

[0032] S31, in order to establish the relevance between the semantics of the discourse under the two recognition tasks, a semantic interaction graph is constructed. N nodes, including N emotional discourse nodes and N behavior discourse nodes, respectively using the H obtained in step S2 e and H a Initialize; the edge represents the connection relationship r i’j’ Next, the current node w i’ and nodes w j’ The relationship between r i’j’ Indicates the current node w i’ and nodes w j’ The type of relationship between i', j' = 1,2,...,2 N .

[0033] There are 6 types of connection relationships in the semantic interaction graph: same-past, same-future, same-current, different-past, different-future, and different-current. Same and different represent the current node w i’ With Node w j’ for the same recognition task and different recognition tasks; past, future and current refer to the current node respectively w i’ and nodes w j’ The relative position in the conversation.

[0034] For example, the semantic interaction graph contains 6 nodes { w 1, w 2, w 3, w 4, w 5, w 6}, which contains 3 emotional discourse nodes and 3 behavioral discourse nodes, such as Figure 2 As shown, 6 nodes { w 1, w 2, w 3, w 4, w 5, w 6} Specifically expressed as discourse u 1. u 2. u 3. Emotional discourse representation [h1 e ,h2 e ,h3 e ] and behavioral discourse representation [h1 a ,h2 a ,h3 a ]. If the current node is h2 e , then h2 e and h2 a The connection relationship is different-current; if the current node is h2 e , then h2 e and h1 a The connection relationship is different-future; if the current node is h2 e , then h2 e and h3 a The connection relationship is different-past; if the current node is h1 e , then h1 e and h2 e The connection relationship is the same-past; if the current node is h3 e , then h3 e and h2 e The connection relationship is same-future.

[0035] S32, similar to the speaker interaction graph, uses the graph attention network to update the nodes in the semantic interaction graph and establish the correlation between the semantics of the discourse under the two recognition tasks. w i’ , aggregate nodes through the following formula w i’ The neighbor node information of node w i’ The discourse representation after semantic interaction : ; in, Representation Node w i’ The aggregated information of w i’ The discourse representation after semantic interaction; h i’ For Node w i’ The corresponding emotional discourse expression or behavioral discourse expression; h j’ Neighbor node w j’ The corresponding emotional discourse representation or behavioral discourse representation; α i’j’ Indicates the current node w i’ For neighbor nodes w j’ importance; r i’j’ For the current node w i’ and neighbor nodes w j’ Type of relationship r i’j’ vector representation of; a' and W' are learnable parameters; h k’ For Node w k’ The corresponding emotional discourse representation or behavioral discourse representation; r i’k’ For Node w i’ and nodes w k’ The vector representation of the connection relationship between them; N i’ Indicates the current node w i’ is the set of neighbor nodes; || is the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function.

[0036] Update all nodes in the semantic interaction graph, obtain the aggregate information of each node, and obtain the discourse representation after semantic interaction .

[0037] S33, according to whether the node is an emotional discourse representation or a behavioral discourse representation, Extract the emotional discourse representation after semantic interaction respectively Behavior discourse representation after interaction with semantics ;in, , Respectively N The emotional discourse representation and behavioral discourse representation after the semantic interaction of the discourses.

[0038] S4, label interaction, models the dependencies and associations between labels.

[0039] S41, assuming that the emotion label and behavior label contain U and V , that is, { e 1, e 2,..., e U}and{ a 1, a 2,..., a V First, use RoBERTa to represent each label as a vector and get the initial label embedding [e1,e2,...,e U ] and [a1,a2,...,a V ]; where e U 、a V Respectively U The sentiment label vector representation and V Behavior label vector representation; S42, construct a label interaction graph to model the dependencies between labels. Labels are used as nodes in the label interaction graph. U + V The correlation matrix A of the label interaction graph is constructed using label co-occurrence.

[0040] The co-occurrence matrix M is constructed based on the total number of times the emotion labels and behavior labels co-occur in the training set. The elements in the co-occurrence matrix M M xy For Node L x and nodes L y The number of co-occurrences (co-occurrences), x,y = 1,2,..., U + V , if the node Lx and nodes L y If all are emotion labels or all are behavior labels, then the node L x and nodes L y The number of co-occurrences is 0. The elements in the conditional probability matrix P P xy = M xy / N x , N x For Node L x The total number of occurrences in the training set, P xy Representation Node L x When the node appears L y Probability of occurrence.

[0041] Since the node L x and nodes L y The co-occurrence between may be a long-tail distribution, and some rare co-occurrences will bring noise. Therefore, using the threshold oh Filter out noisy label pairs. In other words, there may be many items in the co-occurrence matrix M with very small or insignificant co-occurrence times, so the items with high co-occurrence times can be retained by threshold filtering. By threshold filtering, those unimportant or insignificant label pairs are deleted, thereby reducing noise and improving the validity of the data. Specifically, using the threshold oh Filter out the noise label pairs to obtain the binary correlation matrix A': ; In order to prevent the node features from being over-smoothed, a re-weighting method is designed, and finally the correlation matrix A is obtained: ; Among them, the elements in the correlation matrix A A xy Representation Node L x and nodes L y The correlation between p is a hyperparameter that determines the weights assigned to the node itself and other related nodes.

[0042] S43, use graph convolutional network to aggregate label nodes and obtain label representation (matrix) after label interaction : ; in, represents the regularized adjacency matrix; L is the initialization label embedding, L=[e1,e2,...,e U ,a1,a2,...,a V ]; W l is a learnable parameter; LeakReLU(·) represents the activation function.

[0043] S44, according to whether the node is an emotion label or a behavior label, Extract the sentiment label representation after label interaction Behavior label representation after interaction with label ;in, , They are the first U The sentiment label represents V Behavior labels.

[0044] S5, label decoding.

[0045] The emotional discourse after semantic interaction is expressed as and behavioral discourse The representations are mapped to the label space respectively, and then the dot product is used to calculate the correlation score between the utterance and the label. The higher the score, the closer the distance between the utterance and the label representation is, and the more likely the utterance belongs to this category.

[0046] S51, since the emotional discourse representation and behavioral discourse representation after semantic interaction are not in the same vector space as the emotional label representation and behavioral label representation after label interaction, the emotional discourse representation and behavioral discourse representation after semantic interaction are respectively mapped to the joint label space where the emotional label representation and behavioral label representation after label interaction are located through a support vector machine, thereby obtaining the mapped emotional discourse representation and behavioral discourse representation. Subsequently, the mapped emotional discourse representation is dot-producted with the emotional label representation after label interaction to obtain the emotional correlation score matrix ; Perform dot product between the mapped behavior discourse representation and the behavior label representation after label interaction to obtain the behavior-related score matrix .

[0047] S52, apply the softmax function to the emotion correlation score matrix and the behavior correlation score matrix respectively to obtain the predicted emotion probability distribution matrix And the predicted behavior probability distribution matrix : ; in, , , and Respectively represent N The predicted emotion probability distribution and predicted behavior probability distribution of each utterance.

[0048] S53, using the argmax function i Discourse u i The predicted emotion probability distribution and predicting the probability distribution of behavior Calculate and get the words u i Predicted sentiment labels and predicted behavior labels.

[0049] S6, calculate the total loss of the model and perform model training.

[0050] S61, cross entropy loss for emotion recognition and cross entropy loss for action recognition They are: ; in, N , U and V Respectively represent the number of utterances, emotion categories, and behavior categories in the conversation; is a symbolic function, if the utterance u i The true emotional label is x’ Emotional The value is 1, otherwise The value is 0; is a symbolic function, if the utterance u i The actual behavior label is y’ Class behavior, The value is 1, otherwise The value is 0; For discourse u i The predicted sentiment label is x’ Probability of class sentiment; For discourse u i The predicted behavior label is y’ The probability of class behavior.

[0051] S62, based on the multi-level interaction method, proposes multi-perspective supervised contrastive learning to improve the ability to distinguish different labels within a task from a single-task perspective, and further enhance the association between labels between tasks from a multi-task perspective.

[0052] Sentiment Supervised Contrastive Learning Loss for: ; in, I It is a batch sample set, which contains multiple utterances, i.e., samples; For sample i The emotional discourse representation after semantic interaction; For sample p The emotional discourse representation after semantic interaction; P e is the emotional positive sample set, including Samples with the same emotions; N e is the emotional negative sample set, including There are samples with different emotions; For sample n The emotional discourse representation after semantic interaction; t is a hyperparameter.

[0053] Similarly, we obtain the behavior supervision contrast loss for: ; in, For sample i The behavioral discourse representation after the semantic interaction; For sample p The behavioral discourse representation after the semantic interaction; P a is the behavior positive sample set, including Samples with the same behavior; N a is the behavior negative sample set, including There are samples of different behaviors; For sample n The behavioral discourse representation after the semantic interaction; t is a hyperparameter.

[0054] Emotion-behavior supervised contrastive loss from a multi-task perspective for: ; in, I is a batch sample set, which contains multiple utterances, i.e., samples; i For sample i The (emotional or behavioral) label representation after label interaction; p For sample p Label representation after label interaction; P sis the label positive sample set, including the batch with s i co-occurring samples; N s is a set of labeled negative samples, including those in the batch i Samples that do not co-occur; n For sample n The label representation after label interaction.

[0055] S63, total model loss during training for: ; S7, Joint recognition of conversational sentiment and behavior.

[0056] After the training is completed, the trained model is used to jointly identify the conversation emotions and behaviors. N The output is a dialogue of consecutive utterances, and the predicted emotion label and predicted behavior label of each utterance in the dialogue.

[0057] The method of the present invention utilizes graph neural networks from two levels, namely semantics and label association, to achieve mutual guidance of emotions and behaviors, thereby improving the performance of joint recognition of dialogue emotions and behaviors. It also proposes single-task and multi-task multi-perspective supervised comparative learning, which improves the ability to distinguish different labels within a task from a single-task perspective, and enhances the association between labels between tasks from a multi-task perspective, thereby further improving the ability of joint recognition of emotions and behaviors from different perspectives.

[0058] The present invention also provides a dialogue emotion and behavior joint recognition system based on hierarchical interactive learning, the system comprising: a discourse encoding module, a semantic interaction module, a label decoding module, and a label interaction module.

[0059] The speech encoding module is used to encode each speech in the conversation to obtain a vector representation of each speech, and to construct a speaker interaction graph based on the vector representation of each speech to obtain an emotional speech representation and a behavioral speech representation of each speech.

[0060] The semantic interaction module is used to construct a semantic interaction graph according to the emotional discourse representation and behavioral discourse representation of each discourse, and obtain the emotional discourse representation and behavioral discourse representation of each discourse after the semantic interaction.

[0061] The label decoding module is used to identify the predicted emotion label and predicted behavior label of each utterance according to the emotion utterance representation and behavior utterance representation after the semantic interaction of each utterance.

[0062] The label interaction module is used to construct a label interaction graph according to the emotion label and the behavior label to obtain the emotion label representation and the behavior label representation after the label interaction; the label interaction module participates in the calculation during the training process and does not participate in the calculation during the prediction process.

[0063] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for joint recognition of dialogue emotions and behaviors based on hierarchical interactive learning, characterized in that: include: Obtain a conversation dataset for model training; The dialogue consists of continuous N Each utterance has a sentiment label and a behavior label. According to the speakers of the utterances in the dialogue, a speaker interaction graph is constructed and the emotional utterance representation and behavioral utterance representation of each utterance are obtained; According to the emotional discourse representation and behavioral discourse representation of each discourse, a semantic interaction graph is constructed and the emotional discourse representation and behavioral discourse representation after semantic interaction are obtained; According to the emotion labels and behavior labels of each utterance, a label interaction graph is constructed to obtain the emotion label representation and behavior label representation after label interaction; Map the emotional discourse representation and behavioral discourse representation after semantic interaction to the label space to obtain the predicted emotional label and predicted behavioral label of the discourse; Calculate the emotion recognition cross entropy loss based on the predicted emotion label and the real emotion label of the utterance; calculate the behavior recognition cross entropy loss based on the predicted behavior label and the real behavior label of the utterance; calculate the emotion supervision contrast loss based on the emotion utterance representation after semantic interaction; According to the behavioral discourse representation after semantic interaction, the behavioral supervision contrast loss is calculated; according to the emotional label representation and behavioral label representation after label interaction, the emotion-behavior supervision contrast loss is calculated; Construct the total loss of the model based on each loss to train the model; The trained model is used to recognize the conversation and identify the behavior label and emotion label of each utterance in the conversation.

2. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: The speaker interaction graph uses utterances as nodes, with a total of N nodes, and the connection relationship between any two nodes is self-past, self-present, self-future, others-past or others-future; among them, self and others respectively indicate that the two nodes are the same speaker and different speakers; past, future and present indicate the relative positions of the two nodes in the conversation; For nodes v i , through the aggregation node v i The neighbor node information of node v i The words h i : ; in, α ij Representation Node v i For neighbor nodes v j The importance of j Neighbor node v j The corresponding utterance vector representation; r ij For Node v i and neighbor nodes v j The vector representation of the connection relationship between i For Node v i The corresponding discourse vector representation; u k For Node v k The corresponding utterance vector representation; r ik For Node v i and nodes v k The vector representation of the connection relationship between them; N i For Node v i The set of neighbor nodes; || represents the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function; a and W are learnable parameters respectively; i =1,2,..., N ; Process all nodes in the speaker interaction graph to obtain the discourse representation H=[h1,h2,...,h N ]; Let the utterance be represented by H=[h1,h2,...,h N ] The emotional speech representation H is obtained by using a bidirectional long short-term memory network for emotion recognition and a bidirectional long short-term memory network for behavior recognition. e =[h1 e ,h2 e ,...,h N e ] and behavioral discourse representation H a =[h1 a ,h2 a ,...,h N a ]; where h N e 、h N a Respectively N The emotional discourse representation and behavioral discourse representation of each utterance.

3. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: The semantic interaction graph uses the emotional discourse representation and behavioral discourse representation of each discourse as nodes, with a total of 2 N nodes; the connection relationship between any two nodes is same-past, same-future, same-current, different-past, different-future or different-current; where same or different means that the two nodes are the same recognition task or different recognition tasks; past, future and current represent the relative positions of the two nodes in the conversation; For nodes w i’ , through the aggregation node w i’ The neighbor node information of node w i’ The discourse representation after semantic interaction : ; in, α i’j’ Representation Node w i’ For neighbor nodes w j’ The importance of j’ Neighbor node w j’ The corresponding emotional discourse representation or behavioral discourse representation; r i’j’ For Node w i’ and neighbor nodes w j’ The vector representation of the connection relationship between them; a' and W' are learnable parameters; h i’ For Node w i’ The corresponding emotional discourse expression or behavioral discourse expression; h k’ For Node w k’ The corresponding emotional discourse representation or behavioral discourse representation; r i’k’ For Node w i’ and nodes w k’ The vector representation of the connection relationship between them; N i’ Representation Node w i’ The set of neighbor nodes; || is the concatenation operation; LeakReLU(·) represents the activation function; exp(·) represents the exponential function, i’ =1,2,...,2 N ; Process all nodes in the semantic interaction graph to obtain the discourse representation after semantic interaction ; According to whether the node is a sentiment utterance or a behavior utterance, Extract the emotional discourse representation after semantic interaction respectively and behavioral discourse representation ;in, , Respectively N The emotional discourse representation and behavioral discourse representation after the semantic interaction of the discourses.

4. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: The label interaction diagram is used U emotional labels and V behavior labels as nodes, a total of U + V nodes; each label is represented by a vector, and the initial label embedding [e1,e2,...,e U ] and [a1,a2,...,a V ]; where e U 、a V Respectively U The sentiment label vector representation and V Behavior label vector representation; The correlation matrix A of the label interaction graph is constructed using label correlation; the elements in the correlation matrix A A xy Representation Node L x and nodes L y The correlation between x,y =1,2,..., U + V ; Use graph convolutional networks to aggregate nodes and obtain label representations after label interaction : ; in, represents the regularized adjacency matrix; L is the initialization label embedding, L=[e1,e2,...,e U ,a1,a2,...,a V ]; W l is a learnable parameter; LeakReLU(·) represents the activation function; According to whether the node is a sentiment label or a behavior label, Extract the sentiment label representation after label interaction and behavior labels ;in, , They are the first U The sentiment label represents V Behavior labels are represented.

5. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 4 is characterized in that: The correlation matrix A is constructed as follows: Construct the co-occurrence matrix M. The elements in the co-occurrence matrix M M xy For Node L x and nodes L y the number of co-occurrences; Calculate the conditional probability matrix P, the elements in the conditional probability matrix P P xy For Node L x When the node appears L y The probability of occurrence, P xy = M xy / N x ;in, N x For Node L x Number of occurrences; Construct the binary correlation matrix A'. The elements in the binary correlation matrix A' A’ xy The value of is: if P xy ≥ ω ,but A’ xy =1; if P xy < ω ,but A’ xy =0; ω is the threshold value; The binary correlation matrix A' is reweighted to obtain the correlation matrix A. The elements in the correlation matrix A are A xy For Node L x and nodes L y The correlation between them is obtained as follows: ; in, p is a hyperparameter.

6. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: The emotional discourse representation and behavioral discourse representation after semantic interaction are mapped to the label space through support vector machine to obtain the mapped emotional discourse representation and behavioral discourse representation; the mapped emotional discourse representation is dot-producted with the emotional label representation after label interaction to obtain the emotional correlation score matrix ; Perform dot product between the mapped behavior discourse representation and the behavior label representation after label interaction to obtain the behavior-related score matrix ; Using the softmax function, calculate the predicted emotion probability distribution matrix And the predicted behavior probability distribution matrix : ; in, , , and Respectively represent N The predicted emotion probability distribution and predicted behavior probability distribution of each utterance are obtained. N The predicted sentiment label and predicted behavior label of each utterance.

7. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: Cross Entropy Loss for Emotion Recognition and cross entropy loss for action recognition They are: ; in, N , U and V Respectively represent the number of utterances, emotion categories, and behavior categories in the conversation; is a symbolic function, if the utterance u i The true emotional label is x’ Emotional The value is 1, otherwise The value is 0; is a symbolic function, if the utterance u i The actual behavior label is y’ Class behavior, The value is 1, otherwise The value is 0; For discourse u i The predicted sentiment label is x’ Probability of class sentiment; For discourse u i The predicted behavior label is y’ The probability of class behavior.

8. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: Sentiment Supervised Contrastive Learning Loss Comparative loss with behavior supervision They are: ; in, I It is a batch sample set, which contains multiple utterances, i.e., samples; For sample i The emotional discourse representation after semantic interaction; For sample p The emotional discourse representation after semantic interaction; P e is the emotional positive sample set, including Samples with the same emotions; N e is the emotional negative sample set, including There are samples with different emotions; For sample n The emotional discourse representation after semantic interaction; τ is a hyperparameter; ; in, For sample i The behavioral discourse representation after the semantic interaction; For sample p The behavioral discourse representation after the semantic interaction; P a is the behavior positive sample set, including Samples with the same behavior; N a is the behavior negative sample set, including There are samples of different behaviors; For sample n The behavioral discourse representation after semantic interaction.

9. The method for joint recognition of dialogue emotion and behavior based on hierarchical interactive learning according to claim 1 is characterized in that: Sentiment-behavior supervised contrastive loss for: ; in, I is a batch sample set, which contains multiple utterances, i.e., samples; i For sample i The label representation after the label interaction; s p For sample p Label representation after label interaction; P s is the label positive sample set, including the batch with s i co-occurring samples; N s is a set of labeled negative samples, including i Samples that do not co-occur; n For sample n The label representation after label interaction.

10. A dialogue emotion and behavior joint recognition system based on hierarchical interactive learning, characterized by: The method for joint recognition of dialogue emotions and behaviors based on hierarchical interactive learning applied to any one of claims 1 to 9 above, the system comprising: a discourse encoding module, a semantic interaction module, a label decoding module, and a label interaction module; The speech encoding module is used to encode each speech in the conversation to obtain a vector representation of each speech, and to construct a speaker interaction graph based on the vector representation of each speech to obtain an emotional speech representation and a behavioral speech representation of each speech; The semantic interaction module is used to construct a semantic interaction graph according to the emotional discourse representation and behavioral discourse representation of each discourse, and obtain the emotional discourse representation and behavioral discourse representation after the semantic interaction of each discourse; The label decoding module is used to identify the predicted emotion label and predicted behavior label of each utterance according to the emotion utterance representation and behavior utterance representation after the semantic interaction of each utterance; The label interaction module is used to construct a label interaction graph according to the emotion label and the behavior label to obtain the emotion label representation and the behavior label representation after the label interaction; the label interaction module participates in the calculation during the training process and does not participate in the calculation during the prediction process.

Citation Information

Patent Citations

  • Emotion guiding method and system based on emotion semantic transfer graph

    CN111914556A

  • Conversation emotion recognition method based on common sense perception and hierarchical multi-task learning

    CN114722838A

  • Emotion recognition method based on human-object space-time interaction behavior

    CN114926837A

  • Method for assisting in enhancing emotion recognition in dialogue based on generative common sense

    CN116595985A

  • Emotion recognition method and system based on dialogue situation knowledge base and comparative learning

    CN119202854A