A multimodal dialogue emotion recognition method and system based on graph neural network
By constructing hypergraphs and adopting multi-frequency propagation modules in multi-modal dialogue emotion recognition, the shortcomings of the existing technology in modeling multi-variable relationships and utilizing high-frequency information are solved, and more efficient and accurate emotion recognition effects are achieved.
Patent Information
- Application Number
- CN202310437725.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-04-21
AI Technical Summary
The existing multimodal dialogue emotion recognition method based on graph neural networks has shortcomings in modeling multivariate relationships and utilizing high-frequency information, resulting in low recognition accuracy and efficiency.
By constructing a hypergraph to capture the multivariate relationship between multimodal and context, and using a multi-frequency propagation module to extract the importance of different frequency components from node features, data fusion is performed to improve the accuracy of emotion recognition.
It improves the accuracy and efficiency of dialogue emotion recognition, can more fully model complex multivariate relationships and utilize high-frequency information, and improves the performance of the model.
Smart Images

Figure CN116467416B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of emotion computing technology, and in particular to a multimodal dialogue emotion recognition method and system based on graph neural network. Background Art
[0002] The task of Emotion Recognition in Conversation (ERC) aims to enable machines to detect human emotions in conversations using multi-sensory data, including text, visual, and auditory information. Unlike traditional emotion computing tasks performed in a single modality (such as text, speech, or facial images) or in non-conversational scenarios, there is a unique and critical challenge in the ERC task, namely, the complex multivariate relationships across modalities and contextual dimensions.
[0003] Researchers have been exploring how to capture these complex relationships more effectively. In existing ERC models, a major approach is to use context-aware modules (such as recurrent units or Transformers) to capture contextual relationships while modeling multimodal relationships through various fusion methods. Despite some progress, this approach often underestimates the multivariate relationships between modalities and contexts because it encodes multimodal and contextual relationships in a loosely coupled manner, limiting the natural interactions between them and resulting in insufficient learning of relationships.
[0004] Recently, Graph Neural Network (GNN) has shown some advantages in capturing data relationships and provides a new solution for conversation emotion recognition. A common solution is to construct a heterogeneous graph in which each modality of the discourse is regarded as a node and connected to other modalities of the same discourse and to discourses of the same modality in the same conversation. A carefully tuned edge weighting strategy is usually followed. On this basis, multimodal and contextual dependencies between discourses can be modeled simultaneously through message passing, providing tighter entanglement and richer interactions. Although these GNN-based methods are very powerful, they still have two limitations:
[0005] i) Inadequate modeling of multivariate relationships. Traditional GNNs assume that objects of interest have pairwise relationships and can only provide approximations of high-order and multivariate relationships through multiple pairwise combinations. However, degenerating these multivariate relationships into pairwise combinations may impair the expressive power. Therefore, existing GNN-based methods may not be able to adequately model the complex multivariate relationships in ERC.
[0006] ii) Underestimation of high-frequency information. Studies have shown that the propagation rule of GNN (i.e., aggregating and smoothing messages from neighboring nodes) is similar to a fixed low-pass filter, and that low-frequency messages mainly flow in the graph, while the effect of high-frequency information is greatly weakened. In addition, studies have shown that low-frequency messages can preserve the commonality of node features and perform better on homogeneous graphs (homogeneous graphs are graphs in which linked nodes tend to have similar features and share the same labels). In contrast, high-frequency information that reflects differences and inconsistencies is more important in heterogeneous graphs. For ERC, the constructed graphs are usually highly heterogeneous, where inconsistent sentiment information may exist between modalities or short-distance contexts. Therefore, high-frequency information may provide critical guidance, but previous GNN-based ERC models have seriously ignored this, resulting in a bottleneck in performance improvement. Summary of the invention
[0007] The purpose of the present invention is to provide a multimodal conversation emotion recognition method and system based on graph neural network, which can improve the accuracy and efficiency of conversation emotion recognition by studying the multivariate relationship between modalities and contexts and making full use of different frequency information reflecting emotional differences and emotional commonalities.
[0008] The technical solution of the present invention to solve the above technical problems is as follows:
[0009] The present invention provides a multimodal conversation emotion recognition method based on a graph neural network, and the multimodal conversation emotion recognition method based on a graph neural network comprises:
[0010] S1: Obtain speaker- and context-aware unimodal representation, wherein the unimodal representation includes text, vision, and hearing;
[0011] S2: extracting multivariate and high-order information between each modality and conversation context according to the speaker- and context-aware unimodal representation to obtain multivariate representation data;
[0012] S3: Extract the different importance of different frequency components between each modality and conversation context to obtain multi-frequency representation data;
[0013] S4: performing data fusion on the multivariate representation data and the multi-frequency representation data to obtain an emotional representation of the input dialogue;
[0014] S5: Obtain a predicted label of the input dialogue based on the emotion representation, and output the predicted label as a multimodal dialogue emotion recognition result.
[0015] Optionally, the S1 includes:
[0016] S11: Encode the text features of the input dialogue using a bidirectional gated recurrent unit to obtain text encoding data;
[0017] S12: Encode the auditory features and visual features of the input dialogue using the first fully connected network and the second fully connected network respectively to obtain visual encoding data and auditory encoding data;
[0018] S13: Calculate the speaker's embedded representation;
[0019] S14: Obtain a text unimodal representation, a visual unimodal representation and an auditory unimodal representation according to the text encoding data, the visual encoding data and the auditory encoding data, and the embedded representation, respectively.
[0020] Optionally, the S11 includes:
[0021]
[0022] The S12 includes:
[0023]
[0024]
[0025] in, Represents text-encoded data, represents auditory coded data, represents visual encoding data, represents a bidirectional gated recurrent unit function, Text features representing the input dialogue, express or That is, input bidirectional gated recurrent unit The text below or above, W1 represents the first fully connected network, represents the auditory features of the input dialogue, represents the auditory bias, W2 represents the second fully connected network, Visual features representing the input dialogue, Indicates visual bias;
[0026] The S13 includes:
[0027] S i =W s s i
[0028] Among them, S i is the embedding feature of the speaker in the i-th round of dialogue, W s is the trainable weight, s i A one-hot vector represents each speaker;
[0029] The S14 includes:
[0030]
[0031] in, represents the unimodal representation of the speaker and context awareness of the i-th round of dialogue. When x = t, Indicates text encoding data; when x = a, represents auditory coding data; when x=v, represents visual encoding data, S i Represents speaker embedding representation.
[0032] Optionally, the S2 includes:
[0033] S21: determining a plurality of first nodes according to the speaker- and context-aware unimodal representation;
[0034] S22: constructing a multimodal hyperedge and a contextual hyperedge of each first node;
[0035] S23: assigning weights to each hyperedge and each first node respectively;
[0036] S24: generating a hypergraph according to each of the first nodes, each of the hyperedges, each of the hyperedge allocation weights, and each of the first node allocation weights;
[0037] S25: performing a first node convolution on the hypergraph, updating a hyperedge embedding by aggregating node features, and performing a hyperedge convolution to propagate a hyperedge message to the first node;
[0038] S26: Repeat S25 until the last iteration, and use the output of the last iteration as the multivariate characterization data.
[0039] Optionally, the S25 includes:
[0040]
[0041] Among them, V (l) represents the input of layer l and represents a node in the lth layer of the hypergraph neural network, v H represents the set of nodes in the hypergraph, D h represents the feature dimension of the network hidden layer nodes, σ() is a nonlinear activation function, W e is the hyperedge weight matrix and diag() represents the diagonal matrix, w() represents the weight, e1 represents the first hyperedge, Represents |ε H |Hyperedge, ε H represents the set of hyperedges in the hypergraph, and They are the node degree matrix and the hyperedge degree matrix respectively, H represents the association matrix between the hypergraph nodes and the edges. represents the weighted incidence matrix and T represents the transpose operation.
[0042] Optionally, the S3 includes:
[0043] S31: determining a plurality of second nodes according to the speaker- and context-aware unimodal representation;
[0044] S32: construct an undirected graph according to all second nodes;
[0045] S33: extracting high-frequency messages and low-frequency messages of node features of the current node in the undirected graph by using the high-pass filter and the low-pass filter respectively;
[0046] S34: combining the high-frequency message and the low-frequency message by weighted sum;
[0047] S35: according to the weighted contribution of the high-frequency signal of the neighboring node and the low-frequency message to the current node, considering the correlation between the current node and the neighboring node, determining the dominant information of the current node and whether to receive the difference information between the current node and the neighboring node;
[0048] S36: propagating the dominant information and the difference information to the entire undirected graph, by stacking K layers, so that each second node receives a multi-frequency signal from a K-hop neighboring node;
[0049] S37: Use the output of the last layer as multi-frequency representation data.
[0050] Optionally, in S32, the adjacency matrix of the undirected graph is:
[0051] The normalized graph Laplacian matrix of the undirected graph is: Among them, D g is a diagonal degree matrix, I is the identity matrix, A is the adjacency matrix of the undirected graph, v g is a node;
[0052] The S34 includes:
[0053]
[0054] Among them, F (k) is the input of the kth layer and R l , are the weight matrices of low-frequency information and high-frequency information, respectively, so the above formula can be rewritten as:
[0055]
[0056] N i is the neighbor node of node i, N j is the neighbor node of node j, and are the contributions of low-frequency information and high-frequency information of node j to node i, respectively, satisfying the constraint
[0057] Optionally, the S4 includes:
[0058]
[0059] The S5 includes:
[0060]
[0061] in, is the predicted label of the input dialogue, P i Indicates and W4 is a trainable weight matrix, is the normalized sentiment representation and e i represents the emotional features of the input dialogue, P i [τ] represents the probability value of the τth category, τ represents the τth category, b4 represents the bias of the trainable weight matrix, represents multivariate representation data, f i x , i∈[1,N], x∈{t,a,v} represents multi-frequency representation data.
[0062] Optionally, the loss function L of the multimodal dialogue emotion recognition method based on graph neural network is:
[0063]
[0064] Where Num is the number of dialogues, c(i) is the number of sentences in dialogue i, and p ij and i,j are the predicted label probability distribution and true label of sentence j in dialogue i, λ is the regularization weight of L2, θ represents all trainable parameters in the model, and c(s) represents the number of sentences in dialogue s.
[0065] The present invention also provides a multimodal conversation emotion recognition system based on the above-mentioned multimodal conversation emotion recognition method based on graph neural network, and the multimodal conversation recognition system comprises:
[0066] A modality encoding module, the modality encoding module is used to obtain speaker and context-aware unimodal representation;
[0067] A multivariate propagation module, the multivariate propagation module is used to extract multivariate and high-order information between each modality and conversation context according to the speaker and context-aware unimodal representation to obtain multivariate representation data;
[0068] A multi-frequency propagation module, which is used to extract the different importance of different frequency components between each modality and conversation context to obtain multi-frequency representation data;
[0069] The emotion classification module is used to perform data fusion on the multivariate representation data and the multi-frequency representation data to obtain the emotion representation of the input dialogue; and obtain the predicted label of the input dialogue based on the emotion representation, and output the predicted label as the multimodal dialogue emotion recognition result.
[0070] The present invention has the following beneficial effects:
[0071] 1) The present invention studies the multivariate relationship between modality and context, and makes full use of different frequency information reflecting emotional differences and emotional commonalities, so as to improve the accuracy and efficiency of dialogue emotion recognition;
[0072] 2) The hyperedges in the hypergraph of the present invention can connect any number of nodes, and thus can naturally encode more diverse relationships; at the same time, by adopting a set of frequency filters to extract different frequency components from node features, multi-frequency information is modeled on an undirected graph network, so that different frequency signals can be adaptively integrated to capture the different importance of emotional differences and emotional commonalities in local neighborhoods, thereby realizing an adaptive information sharing mode. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 This is a flow chart of a multimodal conversation emotion recognition method based on graph neural network of the present invention;
[0074] Figure 2 Schematic diagram of a multimodal conversation emotion recognition system based on graph neural network of the present invention;
[0075] Figure 3 It is a schematic diagram of the results of the multimodal conversation emotion recognition system based on graph neural network of the present invention based on different graph network layers;
[0076] Figure 4 This is a schematic diagram comparing the effects of the multimodal conversation emotion recognition system based on graph neural network of the present invention and FAGCN. DETAILED DESCRIPTION
[0077] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0078] The present invention provides a multimodal dialogue emotion recognition method based on graph neural network, referring to Figure 1 As shown, the multimodal dialogue emotion recognition method based on graph neural network includes:
[0079] S1: Obtain speaker- and context-aware unimodal representation, wherein the unimodal representation includes text, vision, and hearing;
[0080] A conversation consists of N sentences {(u1, p1), (u2, p2), ..., (u N , p N )}, where each statement u i By speaker p i Spoken, including multisensory data, i.e. text Vision and hearing
[0081] Conversations are essentially sequential and consist of multiple speakers. Therefore, the present invention first uses speaker and context information to process unimodal sentences to obtain speaker and context-aware unimodal representations.
[0082] Specifically include:
[0083] S11: Encode the text features of the input dialogue using a bidirectional gated recurrent unit to obtain text encoding data;
[0084] S12: Encode the auditory features and visual features of the input dialogue using the first fully connected network and the second fully connected network respectively to obtain visual encoding data and auditory encoding data;
[0085] The S11 includes:
[0086]
[0087] The S12 includes:
[0088]
[0089]
[0090] in, Represents text-encoded data, represents auditory coded data, represents visual encoding data, represents a bidirectional gated recurrent unit function, Text features representing the input dialogue, express or That is, input bidirectional gated recurrent unit The text below or above, W1 represents the first fully connected network, represents the auditory features of the input dialogue, represents the auditory bias, W2 represents the second fully connected network, Visual features representing the input dialogue, Indicates visual bias.
[0091] S13: Calculate the speaker's embedded representation;
[0092] S i =W s s i
[0093] Among them, S i is the embedding feature of the speaker in the i-th round of dialogue, W s is the trainable weight, s i is a one-hot vector representing each speaker.
[0094] S14: Obtain a text unimodal representation, a visual unimodal representation and an auditory unimodal representation according to the text encoding data, the visual encoding data and the auditory encoding data, and the embedded representation, respectively.
[0095]
[0096] in, represents the unimodal representation of the speaker and context awareness of the i-th round of dialogue. When x = t, Indicates text encoding data; when x = a, represents auditory coding data; when x=v, represents visual encoding data, S i Represents the embedded representation.
[0097] S2: extracting multivariate and high-order information between each modality and conversation context according to the speaker- and context-aware unimodal representation to obtain multivariate representation data;
[0098] The main idea of this step is to explore the multivariate and high-level information between multiple modalities and conversation contexts. The present invention first constructs a hypergraph H from the above sequence-encoded sentences. In general, given a sentence sequence containing N conversation turns, a hypergraph H = (V H ,ε H ,ω,γ), where each first node v∈V H (|V H |=3N) corresponds to a unimodal sentence, and each hyperedge e∈ε H (|ε H|=3+N) encodes multimodal or contextual dependencies. For each hyperedge e∈ε H Assign a weight ω(e), and assign a weight γ to each node v connected to each hyperedge e e (v) Use represents the incidence matrix, where the non-zero entries H ve = 1 means that node v is connected to hyperedge e; otherwise H ve =0.
[0099] Based on this, the S2 includes:
[0100] S21: determining a plurality of first nodes according to the speaker- and context-aware unimodal representation;
[0101] Each mode of each sentence is represented as a node in the hypergraph, i.e. Indicates text mode. represents the auditory modality, Representing visual modalities, using sequence encoding representations Initialize node embedding
[0102] S22: constructing a multimodal hyperedge and a contextual hyperedge of each first node;
[0103] refer to Figure 2 As shown, each first node First, connect to all other statements in the same dialogue and the same modality through a context hyperedge In addition, each first node are connected to other modes of the same sentence through a multimodal hyperedge In this way, the constructed hypergraph is able to capture high-order and multivariate information beyond pairwise combinations.
[0104] S23: assigning weights to each hyperedge and each first node respectively;
[0105] Unlike existing GNN-based ERC models, which use complex relational learning or similarity metrics to manually adjust edge weighting strategies, the present invention uses randomly initialized weight values to avoid making the model unnecessarily complicated. Specifically, the present invention defines two types of weights in the hypergraph:
[0106] i) The edge weight ω(e) of each hyperedge e;
[0107] ii) The node weight γ of each node v connected to the hyperedge e e (v) (also known as the node weight of edge dependency).
[0108] Intuitively, γ e(v) Measure the contribution of node v to the hyperedge e, thereby enforcing fine-grained multimodal and contextual dependencies. Therefore, the node weights of edge dependencies can be expressed as the weighted association matrix express:
[0109]
[0110] S24: generating a hypergraph according to each of the first nodes, each of the hyperedges, each of the hyperedge allocation weights, and each of the first node allocation weights;
[0111] S25: performing a first node convolution on the hypergraph, updating a hyperedge embedding by aggregating node features, and performing a hyperedge convolution to propagate a hyperedge message to the first node;
[0112]
[0113] Among them, V (l) represents the input of layer l and represents a node in the lth layer of the hypergraph neural network, v H represents the set of nodes in the hypergraph, D h represents the feature dimension of the network hidden layer nodes, σ() is a nonlinear activation function, W e is the hyperedge weight matrix and diag() represents the diagonal matrix, w() represents the weight, e1 represents the first hyperedge, Represents |ε H |Hyperedge, ε H represents the set of hyperedges in the hypergraph, and They are the node degree matrix and the hyperedge degree matrix respectively, H represents the association matrix between the hypergraph nodes and the edges. represents the weighted incidence matrix and T represents the transpose operation.
[0114] S26: Repeat S25 until the last iteration, and use the output of the last iteration as the multivariate characterization data.
[0115] After L iterations, the output of the last iteration is Characterize data as multivariate.
[0116]
[0117] S3: Extract the different importance of different frequency components between each modality and conversation context to obtain multi-frequency representation data;
[0118] The above-mentioned multivariate propagation module is able to capture high-order dependencies beyond pairwise combinations, but it still follows the general graph learning protocol, namely aggregating and smoothing signals from the neighborhood. This can be interpreted as a form of low-pass filter, that is, the smoothness of the message basically propagates low-frequency information while erasing high-frequency information. However, as mentioned earlier, high-frequency information that reflects the difference in node sentiment is crucial, and combining the effects of messages of different frequencies is worth exploring. Therefore, the present invention proposes a multi-frequency propagation module to extract different frequency components with different importance. To this end, the present invention further constructs an undirected graph g=(v g , ε g ), in parallel with the multivariate module.
[0119] Specifically, the present invention constructs an undirected graph g=(v g , ε g ), whose node v g Same as the nodes in H, denoted as {f i t ,f i a ,f i v}. Node embedding also uses sequence encoding representation Different from H, the present invention constructs a set of edges ε with pairwise connections g Similarly, the present invention also sets each node f i x Connected to all other sentences of the same modality in the same conversation {f i x |j∈[1,N],j≠i}, and other modal forms of the same sentence {f i z |z∈{t,a,v},z≠x}. Includes:
[0120] S31: determining a plurality of second nodes according to the speaker- and context-aware unimodal representation;
[0121] S32: construct an undirected graph according to all second nodes;
[0122] The undirected graph constructed is Figure 2 As shown, the adjacency matrix of the undirected graph is: The normalized graph Laplacian matrix of the undirected graph is: Among them, D g is a diagonal degree matrix, I is the identity matrix, A is the adjacency matrix of the undirected graph, v g For the node.
[0123] S33: extracting high-frequency messages and low-frequency messages of node features of the current node in the undirected graph by using the high-pass filter and the low-pass filter respectively;
[0124] The present invention first designs a low-pass filter F l and a high-pass filter F h Extract signals from node features:
[0125]
[0126]
[0127] It can be noted that the high-pass filter is equivalent to the normalized graph Laplacian matrix, which is consistent with the theory that the Laplacian kernel can be used to highlight high-frequency edge information in image signal processing. According to the graph Fourier transform theory, given a signal F l and F h The filtering operation can be regarded as And the convolution between the corresponding convolution kernels*c:
[0128]
[0129] Specifically, F l and F h They are all filters. Given a signal When this signal The filtering operation using these two filters can be seen as Convolution operation between convolution kernels corresponding to filters.
[0130] S34: combining the high-frequency message and the low-frequency message by weighted sum;
[0131]
[0132] Among them, F (k) is the input of the kth layer and are the weight matrices of low-frequency information and high-frequency information, respectively, so the above formula can be rewritten as:
[0133]
[0134] N i is the neighbor node of node i, N j is the neighbor node of node j, and are the contributions of low-frequency information and high-frequency information of node j to node i, respectively, satisfying the constraint
[0135] S35: according to the weighted contribution of the high-frequency signal of the neighboring node and the low-frequency message to the current node, considering the correlation between the current node and the neighboring node, determining the dominant information of the current node and whether to receive the difference information between the current node and the neighboring node;
[0136] Consider the correlation between the central node and its neighbors:
[0137]
[0138] in represents tensor concatenation operation, is a trainable weight matrix, and tanh() is the hyperbolic tangent function, which limits the value to [-1,1]. In this way, the coefficient The different importance of different frequency components can be easily simulated. For example, if Then high-frequency messages dominate, and node i receives the difference information between node i and neighbor j (i.e., f i,(k) -f j,(k) );vice versa.
[0139] S36: propagating the dominant information and the difference information to the entire undirected graph, by stacking K layers, so that each second node receives a multi-frequency signal from a K-hop neighboring node;
[0140] S37: Use the output of the last layer as multi-frequency representation data.
[0141]
[0142] S4: performing data fusion on the multivariate representation data and the multi-frequency representation data to obtain an emotional representation of the input dialogue;
[0143]
[0144] S5: Obtain a predicted label of the input dialogue based on the emotion representation, and output the predicted label as a multimodal dialogue emotion recognition result.
[0145]
[0146] in, is the predicted label of the input dialogue, P i Indicates and W4 is a trainable weight matrix, is the normalized sentiment representation and e i represents the emotional features of the input dialogue, P i [τ] represents the probability value of the τth category, τ represents the τth category, b4 represents the bias of the trainable weight matrix, represents multivariate representation data, f i x , i∈[1,N], x∈{t,a,v} represents multi-frequency representation data.
[0147] The present invention follows the conventional setting and uses categorical cross entropy with L2 regularization as the loss function:
[0148]
[0149] Where Num is the number of dialogues, c(i) is the number of sentences in dialogue i, and p ij and i,j are the predicted label probability distribution and true label of sentence j in dialogue i, λ is the regularization weight of L2, θ represents all trainable parameters in the model, and c(s) represents the number of sentences in dialogue s.
[0150] The present invention also provides a multimodal conversation emotion recognition system based on the above-mentioned multimodal conversation emotion recognition method based on graph neural network, and the multimodal conversation recognition system comprises:
[0151] A modality encoding module, the modality encoding module is used to obtain speaker and context-aware unimodal representation;
[0152] A multivariate propagation module, the multivariate propagation module is used to extract multivariate and high-order information between each modality and conversation context according to the speaker and context-aware unimodal representation to obtain multivariate representation data;
[0153] A multi-frequency propagation module, which is used to extract the different importance of different frequency components between each modality and conversation context to obtain multi-frequency representation data;
[0154] The emotion classification module is used to perform data fusion on the multivariate representation data and the multi-frequency representation data to obtain the emotion representation of the input dialogue; and obtain the predicted label of the input dialogue based on the emotion representation, and output the predicted label as the multimodal dialogue emotion recognition result.
[0155] Example 2
[0156] The present invention proposes a multimodal dialogue recognition system (M 3 Net) is validated on two popular multimodal datasets IEMOCAP and MELD. Three modalities are adopted, namely text, video and audio. The present invention uses pre-extracted unimodal features, following the same extraction procedure in previous work.
[0157] The proposed model is implemented using PyTorch and torch-geometric toolkit. The model is trained on a machine equipped with 1 NVIDIA GeForce RTX 3090. Accuracy and F1 score are used as performance metrics. The model is trained using Adam optimizer with a batch size of 16 on both datasets. L and K are tested in the range of 1 to 7 and present the best performance results. The complete details of the hyperparameters for both datasets are shown in Table 1.
[0158] Dataset Batch Optimizer <![CDATA[D h ]]> L K Dropout IEMOCAP 16 Adam (learning rate = 1e-4) 512 3 4 0.5 MELD 16 Adam (learning rate = 1e-4) 512 3 3 0.4
[0159] Table 1 Detailed information of hyperparameters
[0160] The proposed model is compared with the existing state-of-the-art methods, and the results are shown in Table 2. It can be seen that on these two datasets, M 3 Net surpasses previous methods and achieves new state-of-the-art records in terms of accuracy and F1 score metrics. In particular, M 3 Net outperforms previous GNN-based methods, including DialogueGCN, MMGCN, and MM-DFN, which use complex relational learning or similarity metrics to manually adjust edge weighting strategies to capture multimodal and contextual relationships. It can be seen that the advantages of the present invention are due to the study of multivariate and multi-frequency information between modalities and contexts, which is ignored by previous methods.
[0161]
[0162]
[0163] Table 2. Comparison with previous state-of-the-art methods on IEMOCAP and MELD.
[0164] (Bold indicates best performance. Indicates that the data comes from CMN; *Indicates that the data comes from ICON; Indicates that the data comes from MetaDrop; Indicates that the data comes from the reproduction using open source code)
[0165] The present invention is to 3 We conduct an ablation study on the key components of Net to explore the effectiveness of each component module, and the results are presented in Table 3.
[0166]
[0167] Table 3
[0168] The role of multivariate information
[0169] The present invention first explores the role of multivariate information in modality and context. To achieve this, the present invention removes the multivariate propagation module (i.e., hypergraph H) and performs classification based only on multi-frequency representation, which is shown as variant 1 in Table 3. Under this setting, it can be observed that the accuracy on IEMOCAP drops by 2.40% and the F1 score drops by 2.44%. The accuracy on MELD drops by 0.54% and the F1 score drops by 0.69%. This demonstrates the effectiveness of introducing multivariate propagation, which can effectively encode more diverse relationships.
[0170] The role of multi-frequency information
[0171] M 3 Another core component of Net is the multi-frequency propagation module. Again, the present invention tests the importance of this module by removing it and performing predictions using only multivariate representations. Variant 2 shows the results of this configuration, from which a sharp drop in performance can be observed. This demonstrates the effectiveness of introducing different frequency information in ERC, which can guide the model to capture the different importance of sentiment differences and sentiment commonalities in local neighborhoods.
[0172] The role of weights in hypergraphs
[0173] Two types of weights are defined in the hypergraph H to capture fine-grained multivariate relationships. Therefore, the present invention conducts experiments to verify the effects of these two weights. From variants 3 to 5, it can be seen that deleting one or two of the weights (i.e., the weight values ω(e) or / and γ e Setting (v) to 1) hurts the performance on both datasets. This suggests that the formulated weights are beneficial for training.
[0174] The role of parallel modeling
[0175] In M 3 Net, the present invention propagates multivariate and multifrequency information in parallel. Parallel modeling is compared with two-step serial modeling, and the results are shown as variants 6 and 7. Serial modeling slightly degrades the performance on MELD, but leads to a significant drop in performance on IEMOCAP, which means that parallel modeling is effective.
[0176] M 3 Net contains two parallel graphs, and graph propagation plays a key role. In order to study the impact of stacking different numbers of graph network layers, this paper conducts a grid search on the number of layers. Specifically, the number of layers for multivariate propagation (L) and multifrequency propagation (K) is searched in the range of 1 to 7, and the results are summarized in Figure 3On IEMOCAP, the effects of L and K are similar. At first, the results steadily improve as more layers are stacked, with peaks occurring at L=3 and K=4, respectively. Further stacking more layers has little positive impact on the performance. On the other hand, it can be noticed that the results on MELD are not very sensitive to the number of graphic layers, with no particular pattern as either shallow or deep layers can produce decent performance.
[0177] As mentioned above, the graph propagation rules of the multi-frequency module of the present invention are closely related to FAGCN, but there are important differences. To further demonstrate the effectiveness of the method of the present invention, additional experimental comparisons with FAGCN are shown. Specifically, the multivariate module is retained, and the multi-frequency modeling strategy of the present invention (step S3) is replaced by the strategy proposed in FAGCN. Since FAGCN introduces a hyperparameter ∈∈[0,1] when defining the filter, ∈ is tested in the range of [0,1] with a step size of 0.1. The comparison is summarized in Figure 4 Obviously, ∈ is a crucial factor that greatly affects the performance, especially on IEMOCAP. However, in any case, these variants with FAGCN cannot outperform the original M 3 Net. This shows the superiority of the multi-frequency modeling mechanism of the present invention.
[0178] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A multimodal dialogue emotion recognition method based on graph neural network, characterized in that: The multimodal dialogue emotion recognition method based on graph neural network includes: S1: Obtain speaker- and context-aware unimodal representation, wherein the unimodal representation includes text, vision, and hearing; S2: extracting multivariate and high-order information between each modality and conversation context according to the speaker- and context-aware unimodal representation to obtain multivariate representation data; S3: Extract the different importance of different frequency components between each modality and conversation context to obtain multi-frequency representation data; S4: performing data fusion on the multivariate representation data and the multi-frequency representation data to obtain an emotional representation of the input dialogue; S5: Obtain a predicted label of the input dialogue according to the emotion representation, and output the predicted label as a multimodal dialogue emotion recognition result; The S3 includes: S31: determining a plurality of second nodes according to the speaker- and context-aware unimodal representation; S32: construct an undirected graph according to all second nodes; S33: extracting high-frequency messages and low-frequency messages of the current node feature in the undirected graph by using a high-pass filter and a low-pass filter respectively; S34: combining the high-frequency message and the low-frequency message by weighted sum; S35: according to the weighted contribution of the high-frequency signal of the neighboring node and the low-frequency message to the current node, considering the correlation between the current node and the neighboring node, determining the dominant information of the current node and whether to receive the difference information between the current node and the neighboring node; S36: propagating the dominant information and the difference information to the entire undirected graph, by stacking K layers, so that each second node receives a multi-frequency signal from a K-hop neighboring node; S37: Use the output of the last layer as multi-frequency representation data.
2. According to claim 1, the multimodal dialogue emotion recognition method based on graph neural network is characterized in that: The S1 includes: S11: Encode the text features of the input dialogue using a bidirectional gated recurrent unit to obtain text encoding data; S12: Encode the auditory features and visual features of the input dialogue using the first fully connected network and the second fully connected network respectively to obtain visual encoding data and auditory encoding data; S13: Calculate the speaker's embedded representation; S14: Obtain a text unimodal representation, a visual unimodal representation and an auditory unimodal representation according to the text encoding data, the visual encoding data and the auditory encoding data, and the embedded representation, respectively.
3. According to claim 2, the multimodal dialogue emotion recognition method based on graph neural network is characterized in that: The S11 includes: The S12 includes: in, Represents text-encoded data, represents auditory coded data, represents visual encoding data, represents a bidirectional gated recurrent unit function, Text features representing the input dialogue, express or , that is, input bidirectional gated recurrent unit The text below or above, represents the first fully connected network, represents the auditory features of the input dialogue, Indicates auditory bias, represents the second fully connected network, Visual features representing the input dialogue, Indicates visual bias; The S13 includes: in, For the i The embedding features of the speaker in the turn-based dialogue, are trainable weights, A one-hot vector represents each speaker; The S14 includes: in, Indicates i Turn-based dialogue speakers and context-aware unimodal representations, when x = t hour, Represents text-encoded data; when x = a hour, Represents auditory coding data; when x = v hour, represents visual encoding data, Represents speaker embedding representation.
4. According to claim 1, the multimodal dialogue emotion recognition method based on graph neural network is characterized in that: The S2 includes: S21: determining a plurality of first nodes according to the speaker- and context-aware unimodal representation; S22: constructing a multimodal hyperedge and a contextual hyperedge of each first node; S23: assigning weights to each hyperedge and each first node respectively; S24: generating a hypergraph according to each of the first nodes, each of the hyperedges, each of the hyperedge allocation weights, and each of the first node allocation weights; S25: performing a first node convolution on the hypergraph, updating a hyperedge embedding by aggregating node features, and performing a hyperedge convolution to propagate a hyperedge message to the first node; S26: Repeat S25 until the last iteration, and use the output of the last iteration as the multivariate characterization data.
5. According to claim 4, the multimodal dialogue emotion recognition method based on graph neural network is characterized in that: The S25 includes: in, represents the input of layer l and , represents a node in the lth layer of the hypergraph neural network, represents the set of nodes in the hypergraph, Represents the feature dimension of the network hidden layer nodes, is a non-linear activation function, is the hyperedge weight matrix and , represents a diagonal matrix, represents the weight, represents the first hyperedge, express Super edge, represents the set of hyperedges in the hypergraph, and are the node degree matrix and the hyperedge degree matrix, respectively. represents the association matrix between the hypergraph nodes and edges and , represents the weighted incidence matrix and , T Represents a transpose operation.
6. The multimodal dialogue emotion recognition method based on graph neural network according to claim 1 is characterized in that: In S32, the adjacency matrix of the undirected graph is: The normalized graph Laplacian matrix of the undirected graph is: ,in, is a diagonal degree matrix, is the identity matrix, is the adjacency matrix of the undirected graph, is a node; The S34 includes: in, For the k The input of the layer and , represents a low-pass filter, represents a high-pass filter, are the weight matrices of low-frequency information and high-frequency information, respectively, so the above formula can be rewritten as: Is a node i The neighbor nodes of Is a node j The neighbor nodes of and The nodes are j The low-frequency information and high-frequency information of the node i The contribution of .
7. The multimodal dialogue emotion recognition method based on graph neural network according to claim 1 is characterized in that: The S4 includes: The S5 includes: in, is the predicted label for the input dialogue, Indicates and , is the trainable weight matrix, is the normalized sentiment representation and , represents the emotional features of the input dialogue, Indicates The probability value of each category, Indicates categories, represents the bias of the trainable weight matrix, represents multivariate representation data, Represents multi-frequency characterization data.
8. The multimodal dialogue emotion recognition method based on graph neural network according to any one of claims 1 to 7, characterized in that: Loss function of the multimodal conversation emotion recognition method based on graph neural network for: in, Num is the number of conversations, It's a dialogue i The number of statements in and The dialogue i Chinese sentence j The predicted label probability distribution and the true label, yes L The regularization weight is 2, represents all trainable parameters in the model, Indicates dialogue s The number of statements in .
9. A multimodal conversation emotion recognition system based on the multimodal conversation emotion recognition method based on graph neural network according to any one of claims 1 to 8, characterized in that: The multimodal dialogue emotion recognition system comprises: A modality encoding module, the modality encoding module is used to obtain speaker and context-aware unimodal representation; A multivariate propagation module, the multivariate propagation module is used to extract multivariate and high-order information between each modality and conversation context according to the speaker and context-aware unimodal representation to obtain multivariate representation data; A multi-frequency propagation module, which is used to extract the different importance of different frequency components between each modality and conversation context to obtain multi-frequency representation data; The emotion classification module is used to perform data fusion on the multivariate representation data and the multi-frequency representation data to obtain the emotion representation of the input dialogue; and obtain the predicted label of the input dialogue based on the emotion representation, and output the predicted label as the multimodal dialogue emotion recognition result.
Citation Information
Patent Citations
A method and an apparatus for tracking microblog messages for relevancy to an entity identifiable by an associated text and an image
CN105593851A
Emotion recognition method, intelligent device and computer readable storage medium
CN111164601A