Session emotion recognition model based on session directed graph convolution and modal biased fusion

Through the combination of a conversation directed graph convolution network and a modal biased Transformer, the mining problem of long-distance context information and modal interaction information in multimodal conversation emotion recognition is solved, and more efficient information fusion and accurate emotional classification are achieved.

CN120448479APending Publication Date: 2025-08-08CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411763012.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively explore long-distance context information and modal interaction information in multimodal conversation emotion recognition, and modal private information is easily lost during the information fusion process.

Method used

Session directed graph convolution network (CDAGN) is used to capture long-distance context information, and the fusion of shared information between modals is achieved through modal biased Transformer (MBTransformer), while retaining the private information of the modality using orthogonal constraints.

Benefits of technology

It improves the accuracy and F1-score of multimodal emotion recognition, effectively reduces information redundancy, retains modal private information, and improves the classification effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448479A_ABST
    Figure CN120448479A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of emotion recognition, and discloses a session emotion recognition model based on session directed graph convolution and modal biased fusion, which comprises a context information extraction module, a modal biased information aggregation module and a multi-modal emotion classification module. The model connects the utterances for each mode according to the sequence of the sessions instead of directly connecting all the utterances in a full connection mode, so that redundant information is reduced during information transmission. In the aspect of information fusion, the model can enhance the representation of each mode through a biased fusion mode and is matched with an orthogonal constraint to ensure that the specific information of the mode is not lost when the mode information is fused. Model evaluation is carried out on a disclosed data set, and the performance of the model exceeds the result of a previous optimal model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of emotion recognition technology, and in particular to a conversation emotion recognition model that combines conversation directed graph convolution with modality bias. Background Art

[0002] Emotions are an integral part of everyday human communication. With the development of artificial intelligence, there is growing interest in building dialogue systems that can understand user emotions and intent and conduct effective conversations. Understanding the diverse information generated during a conversation, such as text, sound, and facial expressions, is key to accurately understanding user emotions. The goal of multimodal emotion recognition in conversation (MERC) is to analyze and model this diverse data to identify the emotion of each utterance in a conversation. Developing empathetic dialogue systems and advancing human-computer interaction are essential for research on MERC. Mining contextual information in conversations and conducting effective multimodal fusion are key challenges in MERC research.

[0003] Early work on context mining used recursive structures to mine context within conversations. However, recursive models are ineffective at mining long-range context within conversations. Therefore, some research has explored graph-based approaches to mining context. While graph-based approaches can address the inability to mine long-range context, these approaches all rely on fully connected conversation graphs, which can lead to redundant information during information transfer. Furthermore, fully connected conversation graphs can exhibit oversmoothing when multiple messages are passed.

[0004] Regarding multimodal information fusion, some studies have explored intermodal interaction fusion by mining intermodal interaction information to enhance modal information. While these methods have achieved encouraging results, simple feature concatenation or modal weighted fusion cannot fully exploit intermodal interaction information. To address this issue, some researchers have leveraged the Transformer Encoder architecture to mine intermodal interaction information for modal information fusion. Others have used an intra-modal Transformer Encoder to mine contextual information within the modalities and then employed an inter-modal Transformer Encoder to achieve intermodal interaction fusion. However, due to the attention mechanism within the Transformer Encoder, these studies only focus on shared information (similarities) between modalities when fusion is performed, thereby losing private information (dissimilarities). Private information between modalities is also crucial for emotion recognition. Only by considering both shared and private information between modalities can effective integration of modal information be achieved.

[0005] In order to solve the above-mentioned problems, the present invention proposes a conversation emotion recognition model that combines conversation directed graph convolution and modality biased fusion. Summary of the Invention

[0006] To address the above technical problems, the present invention provides a conversational emotion recognition model that combines conversational directed graph convolution and modality-biased fusion. For short-range contextual information, the present invention uses BiGRU to obtain it. For long-range contextual information, the present invention constructs a conversational directed graph network (CDAGN) for each modality by simulating real-world conversation sequences. This reduces redundant information during information transfer, thereby better mining long-range contextual information within each modality. For multimodal information fusion, the present invention first uses the CDAGN in a shared manner across different modalities, implicitly learning the shared information between modalities. The modality-biased transformer (MBTransformer), proposed in the modality-biased information aggregation module, then explicitly mines the shared information between modalities while also reasonably preserving the private information within each modality. The MBTransformer is also used in a shared manner across different modalities to implicitly fuse shared information. Furthermore, the present invention incorporates orthogonal constraints to further ensure that modality-specific information is not overwritten after modality information fusion. Finally, a large number of experiments were conducted on two datasets, MELD and IEMOCAP. The experimental results show that the proposed model is effective compared with the existing best methods.

[0007] The conversation emotion recognition model based on conversation directed graph convolution and modality biased fusion provided by the present invention includes:

[0008] Context information extraction module, used to extract context information in the conversation;

[0009] The modality-biased information aggregation module uses MBTransformer to achieve biased mutual information fusion between modalities while retaining the modality's own private information. Orthogonal constraints are also used to further ensure that the modality's private information is not lost.

[0010] Multimodal sentiment classification module, used to classify sentiment based on contextual information and modality information;

[0011] The context information extraction module obtains the context information within a single modality through the bidirectional gated recurrent unit (BiGRU) and the conversational directed graph convolutional network (CDAGN);

[0012] The modality-biased information aggregation module aggregates modality information through modality bias while retaining the private information of the modality;

[0013] The multimodal sentiment classification module concatenates the enhanced unimodal features and obtains the final sentiment classification result through the fully connected layer and normalization layer.

[0014] Preferably, the bidirectional gated recurrent unit (BiGRU) is used to capture short-range context information in a conversation; and the conversational directed graph convolutional network (CDAGN) is used to capture long-range context information in a conversation.

[0015] Preferably, the modality biased information aggregation module includes:

[0016] A hybrid contrastive learning unit, which is used to achieve inter-modal discourse-level alignment and intra-modal sentiment-level alignment through unsupervised contrastive learning and supervised contrastive learning;

[0017] MBTransformer unit, which takes each modality as the central modality for fusion in turn, mines the shared information between modalities through the attention mechanism, and retains the private information;

[0018] The orthogonal constraint unit is used to ensure that the private information of the modalities is not lost when fusing modal information through orthogonal constraints.

[0019] Preferably, the multimodal sentiment classification module includes:

[0020] A feature splicing unit, used to splice the three enhanced single-modal features;

[0021] The multi-layer perceptron unit is used to pass the concatenated features through the fully connected layer and normalization to obtain the final sentiment classification.

[0022] Preferably, the model uses standard cross entropy as the classification loss function and is optimized in combination with unsupervised contrast loss, supervised contrast loss and orthogonal constraint loss.

[0023] Preferably, the model is used to identify the emotion of each utterance in a conversation, including but not limited to happy, neutral, angry, excited, frustrated, surprised, scared, sad, disgusted and angry.

[0024] Compared with related technologies, the conversation emotion recognition model based on conversation directed graph convolution and modality biased fusion provided by the present invention has the following beneficial effects:

[0025] (1) The present invention constructs a conversation directed graph by simulating the actual conversation state to reduce redundant information in the information transmission process.

[0026] (2) The present invention innovatively proposes a modal biased fusion method, which can not only fuse the shared information between modalities but also alleviate the loss of modal private information, and further retain the modal private information by combining orthogonal constraints.

[0027] (3) The present invention uses modules in a shared manner on each modality to achieve implicit fusion of modal common information.

[0028] (4) Experiments on the MELD and IEMOCAP datasets show that the model proposed in this invention has significant improvements in accuracy and F1-score, and has better sentiment classification effect than other models. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 The overall architecture diagram of a conversation emotion recognition model that combines conversation directed graph convolution and modality biased fusion provided by the present invention;

[0030] Figure 2 This is the MBTransformer architecture diagram in the present invention;

[0031] Figure 3 This is the ablation experiment diagram on the IEMOCAP dataset in the present invention;

[0032] Figure 4 This is the ablation experiment diagram on the MELD dataset in the present invention;

[0033] Figure 5 This is a parameter sensitivity experiment diagram of γ1 in the present invention;

[0034] Figure 6 This is a parameter sensitivity experiment diagram of γ2 in the present invention;

[0035] Figure 7 This is a parameter sensitivity experiment diagram of γ3 in the present invention;

[0036] Figure 8 This is the F1-score graph of the model with different numbers of CDAGN layers in the present invention;

[0037] Figure 9 This is the F1-score graph of the MBTransformer model with different numbers of layers in the present invention. DETAILED DESCRIPTION

[0038] First of all, it should be noted that, assuming there are N utterances from L interlocutors in a conversation, the conversation can be defined as a series of utterances U = {u1,u2,…,u n} and speaker information X={x1,x2,…,x l}, where u i is the i-th utterance in the conversation, x iis the i-th speaker in the conversation. Each sentence contains sentence alignment data from three modalities: text, audio, and visual. Multimodal conversational sentiment analysis is to predict the sentiment labels of utterances based on speaker information and available multimodal information.

[0039] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0040] 1. Conversational emotion recognition model based on directed graph convolution and modality biased fusion

[0041] The overall architecture of the conversational emotion recognition model based on the conversational directed graph convolution and modality biased fusion proposed in this paper is shown in the figure below: Figure 1 As shown, it consists of the following parts:

[0042] (1) Context information extraction module: The context information within a single modality is obtained through a bidirectional gated recurrent unit (BiGRU) and CDAGN.

[0043] (2) Modality biased information aggregation module: MBTransformer is used to achieve biased common information fusion between modalities while retaining the modality’s own private information, and orthogonal constraints are used to further ensure that the modality’s private information is not lost.

[0044] (3) Multimodal sentiment classification module: The three enhanced unimodal features are concatenated and passed through a fully connected layer and normalized to obtain the final sentiment classification.

[0045] 1.1 Unimodal Feature Extraction

[0046] In order to better extract the discourse text representation with interpretability, the present invention uses the large-scale general pre-trained language model RoBERTa-Large as the text vector encoding extraction. Specifically, the original text of each discourse sentence is constructed into a general format: {[CLS],w i1 ,w i2 ,…,w in}, where i represents the i-th utterance and n represents the n-th word in the i-th utterance. This format is input into RoBERTa-Large and the output [CLS] feature is used as the feature of the utterance. The feature dimension is 1024.

[0047] OpenSmile and DenseNet extract features from raw audio and images, respectively. For each raw emotional audio data set, OpenSmile first extracts 6,373 features. The audio features in the IEMOCAP dataset are then reduced to 1,582 dimensions, and the audio features in the MELD dataset are reduced to 300 dimensions. Image features are 342 dimensions.

[0048] 1.2 Context Information Extraction

[0049] For each modality there are utterances with the same sequence length:

[0050]

[0051] where l represents the number of utterances, and Since the feature dimensions of the modalities are not uniform, in order to facilitate the processing of subsequent models, a fully connected layer will be used to project the dimensions of the three modalities into the same dimension:

[0052] U {t,a,v} =Linear {t,a,v} (U {t,a,v} ) (2)

[0053] Among them U {t,a,v} ∈R l×d .

[0054] The contextual information in a conversation is crucial for identifying the sentiment of an utterance. According to the size of the distance, the contextual information is divided into close contextual information and distant contextual information.

[0055] Short-range context information: The emotion expressed by the current utterance in a conversation is easily influenced by the most recent utterance, so capturing short-range context information can help identify the emotion of the utterance. Since BiGRU uses update gates and reset gates to learn how to adjust the importance of input information and the hidden state information of the previous time step at each time step, it effectively captures short-range dependencies. Therefore, in this invention, BiGRU is used to capture short-range context information in the conversation:

[0056] F x =BiGRU x (U x ),x∈{t,a,v} (3)

[0057] in

[0058] Long-range context information: Capturing long-range context information is very helpful for emotional speech caused by many factors. Previous works have captured long-range context information through composition, but these works are based on a fully connected approach. This approach can easily generate redundant information when transmitting information. Therefore, in order to alleviate the generation of redundant information when capturing long-range context information in conversations, the present invention constructs CDAGN by simulating the actual conversation sequence (that is, the current speech can only receive information from the previous speech), transmitting information in a natural and reasonable way.

[0059] The present invention represents a directed graph as G = {V, E, R}. The nodes in the graph are the sentences in the conversation, that is, V = {u1, u2, ... u l The edges in the graph represent the direction of information transmission and the type of information, for example (i, j, r i,j )∈E represents node u i To node u j Information transmission, r i,j ∈R represents the edge relationship type. The edge type set is R = {1,0}, where 1 indicates that two connected utterances in the graph come from the same speaker, and 0 indicates that they come from different speakers.

[0060] For the transmission of information in the graph, the present invention adds three constraints:

[0061] (1) Directionality: Each statement in the graph can only receive information from previous statements.

[0062] (2) Long-range information: Except for the first sentence, each sentence u in the figure i The long-range information is defined as u τ , where i>τ, u i and u τ Utterances spoken by the same speaker. τ contains the information of the previous speech, so in the process of information dissemination, u τ As the remote information is propagated to u i .

[0063] (3) Short-distance information: The present invention defines τ and u i All words of u l As short-range information.

[0064] Due to the flow of temporal information, the hidden states of utterances should be computed sequentially from the first utterance to the last utterance.

[0065] Taking the text mode as an example, for each utterance By using The (l-1)-th characteristics and its predecessor The l-th feature calculates the attention weight α between them ij .

[0066]

[0067] in is a learnable parameter matrix, and || represents the concatenation operation.

[0068] The speaker's personality traits will affect the emotions of their own speech as well as the emotions of other speakers' speech. If the speaker (intra-speaker) has a cheerful personality, then the speech they speak will tend to be happy. At the same time, other speakers (inter-speakers) will also be influenced by their cheerful personality and speak with happy emotions. Inspired by R-GCN, CDAGN adds intra-speaker and inter-speaker type edges during information transmission to enable the model to perceive which information comes from itself and which comes from other speakers, thereby modeling the influence of the speaker. The specific formula is defined as follows:

[0069]

[0070] in is a learnable parameter used to learn the speaker type information.

[0071] Got Required aggregate information Afterwards, GRU is used to aggregate the information and In the (l-1)-th hidden representation Merge to obtain the l-th hidden representation.

[0072]

[0073] Taking into account Contains the information of the neighbor nodes of the current layer, and uses guide Similarly, in order to retain more historical context information, the present invention uses As a guiding signal spread.

[0074]

[0075] final The l-th is expressed as the sum of the two.

[0076]

[0077] in Similarly, the speech representation of the audio modality and the visual modality is obtained.

[0078]

[0079] in Finally, the feature of each modality is the concatenation of the output results of all layers.

[0080]

[0081] 1.3 Biased Modality Fusion

[0082] Aligned modal information facilitates information fusion between modalities. Therefore, before modal information fusion, the present invention uses hybrid contrastive learning, i.e., unsupervised contrastive learning and supervised contrastive learning, to achieve alignment at the discourse level between modalities and alignment at the sentiment level within modalities.

[0083] Specifically, for unsupervised contrastive learning, the positive sample pair is: Where w∈{t,a,v},o∈{t,a,v},w≠o. That is, the formula for unsupervised contrastive learning is defined as:

[0084]

[0085]

[0086] where B represents the batch size, l represents the number of utterances, and τ represents the temperature hyperparameter.

[0087] For supervised contrastive learning, utterances with the same label within a single modality are considered positive sample pairs, and utterances with different labels are considered negative sample pairs. The specific supervised contrastive learning formula is defined as follows:

[0088]

[0089] Previous work only focused on the shared information between modalities when performing multimodal information fusion, thereby maximizing the shared information between modalities. However, this will lose the private information of the modalities. Therefore, we proposed MBTransformer and orthogonal constraints to ensure that the private information of the modalities is not lost when fusing the shared information between modalities. MBTransformer will take each modality as the central modality of fusion in turn, and the remaining modalities will assist in improving the shared information of the central modality. At the same time, the central modality will pay attention to its own private information to ensure that the private information will not be covered by the shared information. Taking the text modality as an example, the architecture of MBTransformer is as follows: Figure 2As shown in the figure, it first takes the result of the context information extraction modality as the initial input:

[0090]

[0091] Then calculate the attention weight matrix of the text modality, audio modality, visual modality and the attention weight matrix of the text modality itself respectively:

[0092]

[0093] Where d is the feature dimension, is a learnable parameter. ta(0) ,att tv(0) as a guiding signal of common information to enhance the common information of text modality, and at the same time t(0) It serves as a guiding signal for private information to enhance the private information of text modality.

[0094] Then use a sum function to add the attention weight matrices of the three modalities to obtain the final multimodal weight matrix:

[0095] att tav(0) =Softmax(sum(att t(0) ,att a(0) ,att v(0) )) (twenty one)

[0096] In this way, att tav(0) It contains both the guidance signal of common information and the guidance signal of private information, so when using it to enhance the text modality, both common information and private information will be taken into account. The specific formula is as follows:

[0097]

[0098] after Replace Z t(0) The new text modality input is used, while the audio and visual modality inputs remain unchanged. After passing through N layers of MBTransformers, the final enhanced text modality features are obtained. The present invention replaces Formulas 19 to 22 with the following formulas:

[0099] Z t′ =MBTransformer-t N (Z t(0) ,Z a(0) ,Z v(0) ) (twenty three)

[0100] Where N means N layers are stacked.

[0101] Similarly, enhanced audio and visual modes can be obtained:

[0102] Z a′ =MBTransformer-a N (Z a(0) ,Z t(0) ,Z v(0) ) (twenty four)

[0103] Z v′ =MBTransformer-v N (Z v(0) ,Z t(0) ,Z a(0) )

[0104] In order to further ensure that the private information of the modality is not lost, the present invention proposes an orthogonal constraint to achieve this goal. The specific formula is defined as follows:

[0105]

[0106] in Represents the squared Frobenius norm. This loss further preserves the private information of the modalities by minimizing the similarity between the modalities.

[0107] 1.4 Multimodal Sentiment Classification

[0108] After obtaining the enhanced representations of the three modalities through MBTransformer, they are concatenated to obtain the final sentiment features, which are then input into a multi-layer perceptron to obtain the final sentiment classification.

[0109]

[0110] 1.5 Loss Function

[0111] The present invention uses standard cross entropy as the classification loss function.

[0112]

[0113] Where D represents the number of sessions, c i represents the number of utterances in the ith session, p i,j represents the probability distribution of the predicted sentiment label of the jth utterance in the i-th session, y i,j represents the predicted category of the jth utterance in the i-th session. λ represents the L2 regularization weight, and Θ represents the set of trainable parameters. Thus, the overall loss can be expressed as:

[0114] L all =L ce +γ1L ucl +γ2L scl+γ3L or (29)

[0115] Among them, γ1, γ2, and γ3 are hyperparameters, which are used to control the size of unsupervised contrast loss, supervised contrast loss, and orthogonal constraint loss, respectively.

[0116] 2 Experiments

[0117] The above model will be analyzed in detail through experiments below.

[0118] 2.1 Dataset and Evaluation Metrics

[0119] The present invention uses two datasets, IEMOCAP and MELD, to evaluate the effectiveness of the model. These datasets contain information from text, audio, and visual modalities. Table 1 shows the data distribution of the two datasets.

[0120] Table 1 Statistics of the dataset

[0121]

[0122] IEMOCAP: Each dialogue in this dataset is performed by two actors according to a script. IEMOCAP contains 7,433 utterances and 151 dialogues. Each sentence in the dialogue is labeled with six emotions: happy, neutral, angry, excited, and frustrated.

[0123] MELD: This dataset is based on the television show "Friends," consisting of 13,708 utterances and 1,433 dialogues. Unlike the IEMOCAP dataset, MELD has three or more speakers in each dialogue. Each utterance is labeled with seven emotions: neutral, surprise, fear, sadness, happy, disgust, and anger.

[0124] Evaluation indicators: This paper uses weighted average accuracy and F1-score to evaluate the overall performance of the model.

[0125] 2.2 Experimental Setup

[0126] This experiment used an RTX 2080Ti graphics card with 12GB of video memory for model training. The experiment used the PyTorch deep learning framework. On the IEMOCAP dataset, the CDAGN was set to 4 layers, the MBTransformer to 1 layer, the learning rate and batch size to 0.0001 and 16, respectively, and the values of γ1, γ2, and γ3 to 0.05, 0.007, and 1.0.

[0127] On the MELD dataset, CDAGN is set to 1 layer, MBTransformer is set to 3 layers, the learning rate and batch size are set to 0.0002 and 32 respectively, and the values of γ1, γ2 and γ3 are set to 0.03, 0.003 and 1.4. The Adam optimizer is used to train the model.

[0128] 2.3 Comparative Experiment

[0129] In this section, the performance comparison of our proposed model and previous baseline models is presented.

[0130] ICON: Two GRUs are used to model the speaker's information, an additional global GRU is used to track changes in emotional state throughout the conversation, and a multi-layer memory network is used to model the global emotional state. However, ICON is only designed for two-person conversations.

[0131] DialogueRNN: Models the speaker and sequential information in a conversation by using three GRUs with different functions (global GRU, speaker GRU, and emotion GRU).

[0132] DialogueGCN: Applying GCN to the MERC task, the generated features can integrate rich information, but this also introduces redundant information.

[0133] DialogueCRN: Introduces a cognitive stage that extracts and integrates emotional cues from the context retrieved in the perceptual stage.

[0134] MMGCN: Using the GCN network, we not only consider the global contextual information within the modality, but also try to model other modal information.

[0135] MM-DFN: By designing a new graph-based dynamic fusion module to fully integrate multimodal context information, it can better interpret the emotions in the discourse.

[0136] SCMM: We design a modality interaction module, which consists of three sub-modules, including full, partial, and biased interactions, to fully study the correlation between different modalities.

[0137] CTNet: A Transformer-based model-level fusion strategy is proposed to capture context dependency and speaker dependency.

[0138] DialogueTRM: A new understanding of intra-modal and inter-modal emotional dynamics is provided to better capture intra-modal and inter-modal interactions.

[0139] DER-GCN: A weighted multi-relation graph is constructed to simultaneously capture the dependencies between speakers and event relations in a conversation, and a self-supervised masked graph autoencoder (SMGAE) is introduced to improve the fusion representation ability of features and structures.

[0140] Table 2 Overall experimental results

[0141]

[0142]

[0143] The bolded results in Table 2 are the best, and the underlined results are the second best. '-' indicates that the corresponding data is not given in the original text. '*' indicates that the improvement in the experimental results is statistically significant (p value is less than 0.05) under the one-sided T-test. The accuracy and F1-score of the model proposed in the present invention are 1.32% and 0.79% higher than those of DER-GCN in the overall emotion category on the MELD dataset. On the IEMOCAP dataset, the accuracy is 2.78% higher than that of DER-GCN, and the F1-score is 2.84% higher than that of DialogueTRM. Because DER-GCN simply performs a weighted sum when fusing modal information, it cannot capture the interactive information between modalities well. DialogueTRM only considers the private information of the modalities when performing modal information fusion. The model of the present invention takes into account both the common information and private information of the modalities when fusing modal information, thereby achieving effective modal information fusion, so the performance will be better than DER-GCN and DialogueTRM.

[0144] 2.4 Ablation Experiment

[0145] (1) Modal ablation experiment

[0146] In order to verify whether the multimodal combination can make up for the problem of insufficient emotional expression ability of a single modality, the present invention conducts sentiment analysis experiments on a single modality, a text-audio modality combination, a text-visual modality combination, and a three-modality combination.

[0147] Table 3 Modal combination test results

[0148]

[0149] The results in Table 3 show that, in single-modality experiments using the same dataset, the text modality performs better than the other two modalities. This is because the text modality contains the richest semantic information, resulting in higher emotion recognition accuracy than the other two modalities. The dual-modality experimental results show that the dual-modality results are superior to any single-modality results, and the combination of the three modalities achieves the best results. These experimental results demonstrate that the modalities can complement each other's emotional capabilities. A greater number of modalities results in stronger emotion recognition capabilities.

[0150] (2) Module ablation experiment

[0151] In order to further verify the effectiveness of the method proposed in the present invention, the present invention conducted relevant ablation experiments. Figure 3 and Figure 4 From the results, it can be concluded that removing any method will cause the performance of the model to decline, indicating that all methods are crucial to the present invention.

[0152] For hybrid contrastive learning loss and orthogonality constraint loss, modal semantic alignment before modal information fusion helps improve model performance. Using orthogonal constraints to properly preserve modal private information after fusion also helps improve model performance.

[0153] For BiGRU, the model metrics all decrease after removal, which indicates that the recursive architecture is helpful in modeling short-range contextual information in the session.

[0154] Regarding speaker information, the IEMOCAP dataset shows the most significant drop after removing the speaker-aware method. This is because the IEMOCAP dataset only has two speakers in a conversation, which corresponds to the two speaker-type edges set in this paper. Therefore, the convolution process can effectively model the dependency between the two speakers. However, in the MELD dataset, there are more than two speakers in a conversation, and simply classifying speaker types into two categories may not be able to effectively model the dependencies between speakers.

[0155] For the BMTransformer, the model's performance dropped significantly after the removal. This is because the BMTransformer, when fusing modal information, can absorb both shared information between modalities while also taking into account private information within the modalities, thus preventing private information from being overwritten. This results in excellent results in the modal information fusion process.

[0156] 2.5 Parameter sensitivity experiments

[0157] The present invention conducts sensitivity experiments on the main hyperparameters γ1, γ2 and γ3 in the model, wherein the present invention uses F1-score to measure the impact of hyperparameters on the experimental results.

[0158] γ1 represents the alignment strength of the inter-modal discourse level. The present invention sets the step size to 0.2 and increases it from 0.01 to 0.09 to observe the change of model performance. Figure 5 The results show that when γ1 is 0.05, the model achieves the best effect on the IEMOCAP dataset. When γ1 is 0.03, the model achieves the best effect on the MELD dataset. This situation shows that appropriately increasing γ1, that is, increasing the alignment strength between modalities, helps to improve the effect of the model on the two datasets. As for γ2, the higher γ2 means the higher the alignment strength of the emotion level within the modality. The present invention sets the step size to 0.002 and increases it from 0.001 to 0.009 to observe the changes in model performance. Figure 6 The results show that when γ2 is 0.007 and 0.003, the model achieves the best results on the IEMOCAP and MELD datasets. This indicates that appropriately increasing the strength of intra-modal sentiment alignment can help improve the performance of the model.

[0159] γ3 represents the orthogonal constraint strength. The larger the value, the stronger the orthogonal constraint. The present invention sets the step size to 0.2 and increases from 1.0 to 2.0 to observe the change in model performance. Figure 7 The results show that when γ3 is 1.0 and 1.4, the model achieves optimal results on both datasets. This is due to the longer conversations in the IEMOCAP dataset, which results in more private information within the modal sequence than in the MELD dataset. Therefore, the private information of each modality can be retained without overly strong orthogonal constraints. In contrast, the MELD dataset requires larger constraints to ensure that the private information of the modalities is not lost during fusion. This also indirectly reflects that retaining the private information of the modalities can improve the accuracy of modal emotion recognition, which is consistent with the results of the ablation experiment.

[0160] 2.6 Experiment on the effectiveness of contextual information

[0161] In order to verify the impact of contextual information on emotion recognition, the present invention conducts experiments on the number of convolutional layers of CDAGN. Figure 8The F1-score of the model with different numbers of layers is shown. From the experiment, it can be seen that on the IEMCOAP dataset, the 4-layer CDAGN achieved the optimal effect of the model. On the MELD dataset, the 1-layer CDAGN achieved the optimal effect. The reason for this may be that the average length of each conversation in the IEMOCAP dataset is about 48, while the average length of each conversation in the MELD dataset is about 9. The number of CDAGN layers determines the distance of the context information mined. The larger the number of layers, the more distant the context information can be mined. For IEMOCAP, since the conversation length is relatively long, only multiple convolution layers can fully mine the long-distance context information in the conversation so that there will be no missing information. On the contrary, the conversation length of MELD is relatively short, and CDAGN only needs 1 layer to model the context information of the discourse well. Based on this situation, the present invention draws the following conclusions:

[0162] (1) For the IEMOCAP dataset, due to the long length of the conversation, long-distance context information is more important than short-distance context information in emotion recognition.

[0163] (2) For the MELD dataset, due to the short length of the conversation, close context information is more important than long context information in emotion recognition. This is consistent with the conclusion obtained in the ablation experiment.

[0164] 2.7 Multimodal Information Fusion Effectiveness Experiment

[0165] In order to verify the effectiveness of modal fusion, the present invention guides the experiment on the number of MBTransformer layers. Figure 9 The F1-score of the model with different numbers of layers is shown. From the experiment, it can be seen that when the number of layers is 1 and 3, the model achieves the best results on IEMOCAP and MELD. Through the experiment, the author draws the following conclusions:

[0166] (1) The average conversation length in the IEMOCAP dataset is 48, indicating that conversations across different modalities already contain rich contextual information. At this point, simply fusing the information between modalities appropriately can improve the accuracy of emotion recognition. Therefore, a single-layer MBTransformer can achieve optimal results on IEMOCAP. Further increasing the intensity of multimodal information fusion will, in turn, lead to lower emotion recognition results due to redundant information.

[0167] (2) The average conversation length in the MELD dataset is 8. Therefore, insufficient contextual information may affect emotion recognition in conversations between different modalities. In this case, the integration of information from different modalities will be more effective to compensate for the lack of contextual information within a single modality. This is why the MBTransformer model achieves optimal results in MELD at the third layer.

[0168] Compared with related technologies, the conversational emotion recognition model based on conversational directed graph convolution and modality biased fusion proposed in this paper has the following beneficial effects:

[0169] This paper builds a novel multimodal emotion recognition model called Conversational Directed Graph Convolution and Modality Biased Fusion. This paper constructs a conversational directed graph for each modality and performs convolution, thereby extracting contextual information between modalities at a deeper level. Secondly, this paper uses MBTransformer to simultaneously consider both the shared and private information of the modalities to achieve effective modal information fusion. Orthogonal constraints are then used to further ensure that private information of the modalities is not lost when fusing the modal information. Finally, experimental results demonstrate that the proposed model has certain advantages in the performance of multimodal emotion recognition.

[0170] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A conversational emotion recognition model based on conversational directed graph convolution and modality bias fusion, characterized by: include: Context information extraction module, used to extract context information in the conversation; The modality-biased information aggregation module uses MBTransformer to achieve biased mutual information fusion between modalities while retaining the modality's own private information. Orthogonal constraints are also used to further ensure that the modality's private information is not lost. Multimodal sentiment classification module, used to classify sentiment based on contextual information and modality information; The context information extraction module obtains the context information within a single modality through a bidirectional gated recurrent unit and a session-directed graph convolutional network. The modality-biased information aggregation module aggregates modality information through modality bias while retaining the private information of the modality; The multimodal sentiment classification module concatenates the enhanced unimodal features and obtains the final sentiment classification result through the fully connected layer and normalization layer.

2. The conversation emotion recognition model based on conversation directed graph convolution and modality biased fusion according to claim 1 is characterized in that: The bidirectional gated recurrent unit is used to capture short-range context information in the conversation; the conversation directed graph convolutional network is used to capture long-range context information in the conversation.

3. The conversation emotion recognition model based on conversation directed graph convolution and modality biased fusion according to claim 1 is characterized in that: The modal biased information aggregation module includes: A hybrid contrastive learning unit, which is used to achieve inter-modal discourse-level alignment and intra-modal sentiment-level alignment through unsupervised contrastive learning and supervised contrastive learning; MBTransformer unit, which takes each modality as the central modality for fusion in turn, mines the shared information between modalities through the attention mechanism, and retains the private information; The orthogonal constraint unit is used to ensure that the private information of the modalities is not lost when fusing modal information through orthogonal constraints.

4. The conversation emotion recognition model based on conversation directed graph convolution and modality bias fusion according to claim 1 is characterized in that The multimodal sentiment classification module includes: A feature splicing unit, used to splice the three enhanced single-modal features; The multi-layer perceptron unit is used to pass the concatenated features through the fully connected layer and normalization to obtain the final sentiment classification.

5. The conversation emotion recognition model based on conversation directed graph convolution and modality biased fusion according to claim 1 is characterized in that: The model uses standard cross entropy as the classification loss function and is optimized in combination with unsupervised contrastive loss, supervised contrastive loss, and orthogonality constraint loss.

6. The conversation emotion recognition model based on conversation directed graph convolution and modality bias fusion according to claim 1 is characterized in that: The model is used to identify the sentiment of each utterance in a conversation, including but not limited to happy, neutral, angry, excited, frustrated, surprised, scared, sad, disgusted, and angry.