A Method and Device for Dialogue Emotion Recognition Based on Multimodal Large Models

By using multimodal large model and multi-layer residual graph convolution network based on semantic graphs in dialogue emotion recognition, the problem of insufficient understanding of multimodal fine-grained emotional semantic information and difficulty in modeling dialogue context for long-term complex learners is solved, and higher accuracy and robustness of emotion recognition are achieved.

CN119416035BActive Publication Date: 2025-05-27ZHEJIANG NORMAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510018288.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-27
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

In the prior art, in dialogue emotion recognition, there are problems such as insufficient understanding of multimodal fine-grained emotional semantic information and difficulty in long-term complex learners' dialogue context modeling, resulting in insufficient accurate emotion recognition.

Method used

The dialogue emotion recognition method based on multimodal large model is adopted, and the feature extraction and emotion recognition of multimodal data are carried out by constructing a model including feature extraction layer, bidirectional gating unit, multimodal large model, BERT language model, modal information complementary module and semantic graph-based multi-layer residual graph convolution network.

Benefits of technology

Through the detailed-grained emotional cues understanding and graphical representation learning of multimodal large models, the accuracy and robustness of dialogue emotion recognition are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416035B_ABST
    Figure CN119416035B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for dialogue emotion recognition based on a multimodal large model, relating to the field of emotion recognition. The method includes obtaining all sets of sentences in the dialogue in the current scenario; each sentence includes three modalities: audio, video, and text; constructing a dialogue emotion recognition model; the dialogue emotion recognition model includes: a feature extraction layer, a bidirectional gated unit, a multimodal large model, a BERT language model, a modality information complementary module, a multi-layer residual graph convolutional network based on a semantic graph, and a fully connected layer; according to all sets of sentences in the dialogue, using the trained dialogue emotion recognition model to obtain an emotion recognition result. The present application can improve the accuracy and robustness of dialogue emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of emotion recognition, and particularly to a method and device for dialogue emotion recognition based on a multimodal large model. Background Art

[0002] In recent years, the development of artificial intelligence (AI) technology has turned many seemingly science fiction concepts into reality. For example, more and more household robots are capable of providing comfort according to user requests. The time to process such requests depends on the intelligence level of the machine and the efficiency of human-machine interaction. Therefore, efficiency has become a key metric for enhancing machine intelligence. Making machines more human, that is, enabling them to quickly adjust their behavior when the user's mood changes, such as providing comfort when the user is in a low mood and timely reminding distracted learners in a smart classroom, is particularly important in the field of intelligent robots. In addition, accurately capturing the user's emotional transition can enhance human-machine interaction, thereby providing users with a higher level of AI services. Accurately identifying the user's emotional state is not only a necessary condition for enhancing machine functions but also an important research focus in the field of artificial intelligence. Therefore, the emotion recognition task plays a key role in many fields and has received extensive attention from scholars.

[0003] Dialogue emotion recognition (ERC), also known as dialogue emotion detection, is the task of discerning the emotional state of a speaker by analyzing the signals (such as text, audio, or facial expressions) conveyed by the speaker in the context of a dialogue. ERC has significant potential in multiple fields, including: (i) Customer service: ERC promotes the interaction between users and customer service by identifying and responding to the emotional state of users, thereby improving overall satisfaction. (ii) Social media analysis: In the field of social media, ERC facilitates the capture and analysis of user emotions, enabling companies to collect user satisfaction with their products or services. (iii) Educational technology: ERC can timely detect changes in the emotions of each student in the classroom and provide feedback to teachers, providing a valuable emotional basis for adjusting teaching methods. (iv) Healthcare: In the field of healthcare, it can also evaluate the emotional well-being of patients in telemedicine consultations, providing insights for medical professionals for further examinations. (v) Human-machine interaction: ERC makes human-machine interaction more intuitive and responsive, enabling virtual assistants and chatbots to effectively identify and adapt to the emotions of users. Overall, the ability to recognize emotions in a dialogue helps AI achieve stronger empathy and more effective communication in a wide range of corpus backgrounds, and this ability is particularly important in the field of educational technology.

[0004] In the field of educational technology, ERC technology is widely used to improve teaching effectiveness and the student experience. It can be roughly divided into two categories. One is sentiment analysis based on natural language processing (NLP), which identifies the emotional state of students by analyzing their words and tones in the classroom. This enables educators to better understand students' emotional feedback, adjust teaching strategies in a timely manner, and provide personalized learning support. The other is audio- and video-based emotion recognition, which identifies emotional changes during the learning process by analyzing students' speech features and facial expressions. This technology can be applied in remote teaching to capture learners' emotions in real time through speech recognition technology, providing a basis for educators to adjust teaching methods. However, the above technologies have some inherent disadvantages. First, both of the above methods rely on temporal modeling. When the conversation is too long, temporal modeling cannot obtain good feature representations due to the long-range dependence problem, resulting in inaccurate emotion recognition. Second, with the wide application of multi-modal data, multi-modal dialogue emotion recognition has become a research hotspot. Traditional emotion recognition mainly focuses on a single modality, but in actual conversations, people communicate information through multiple means such as text, speech, and video. Multi-modal dialogue emotion recognition aims to use multi-modal data to provide richer and more comprehensive emotion representations. In a conversation, emotions are often expressed in multiple ways and are affected by context and inter-modal relationships. By integrating multi-modal data, the conversation context can be understood more comprehensively, improving the accurate understanding of emotions. Traditional multi-modal emotion recognition fuses data from different modalities through different feature fusion methods, ignoring the natural shared and unique information within each modality. There may be some correlations between this information that affect emotion judgment. Second, most traditional graph-based methods only serve as an extension of temporal modeling methods. Simply constructing a graph of multiple modality information cannot fully explore the key factors affecting emotion transitions.

[0005] Regarding the text, speech, and video modality data of learners, the emotional state of the learner in the current conversation is implicitly contained. How to use the data of the three modalities for high-quality representation learning directly affects the accuracy of the final emotional judgment. Among them, the two major challenges that need to be addressed urgently are the understanding of multi-modal fine-grained emotional semantic information and the modeling of long-term complex learner conversation contexts. Therefore, compared with traditional time-series modeling methods, integrating multi-modal large language models (MMLLMs) and graph representation learning for conversation modeling has natural advantages. Specifically, multi-modal large language models integrate key components such as modality feature encoders, attention feature fusion mechanisms, large language models, and external sub-task models. Under the premise of cross-modal semantic alignment such as vision-language and speech-language, they can accurately process complex cross-modal instructions and have high-order fine-grained multi-modal semantic understanding capabilities. This ability can serve many cross-modal downstream tasks, such as visual understanding and question answering, pixel-level image or video editing, and speech human-computer interaction. By introducing multi-modal large language models, the fine-grained understanding of potential emotional clues in learners' multi-modal conversation data can be utilized to obtain emotional semantic information directly related to the judgment of emotional states, realizing the gain of multi-modal fine-grained emotional clue semantic information in the model reasoning process, and thus effectively improving the representation quality of conversation data in complex teaching scenarios. Graph representation learning performs excellently in dealing with complex non-linear relationships and structured data. Many practical problems can be better represented through graph-based modeling, such as social networks, molecular structures in bioinformatics, and knowledge graphs. Traditional representation learning may not be able to effectively capture these complex relationships. At the same time, graph representation learning can consider the context information between a node and its neighboring nodes. This makes the learned representation better reflect the relationship between an entity or event and its surrounding environment, helping to express the context more accurately. In addition, the methods of graph representation learning are usually more general and can adapt to a variety of different tasks. This flexibility makes graph models more applicable when dealing with problems in various fields. Therefore, graph representation learning has always been one of the popular machine learning techniques, and a large number of researchers are conducting research work in this area. In a conversation graph, an object is usually regarded as a node, and the relationships between objects are represented by the edges between the nodes. Specifically, the three modalities of a single sentence in a conversation are regarded as three different nodes, and different types of edges are formed based on different semantic backgrounds. For multi-modal data, graph representation learning transforms the features of each modality into the relationship between edges and nodes through the above different types of edges, thus taking the three modalities into consideration.

[0006] At present, multi-modal large language models and algorithms based on graph representation learning have made great progress. However, common modal information interaction methods are mainly divided into early fusion, late fusion, and cross-learning fusion, etc. Among them, the early fusion method first extracts feature expressions from each independent modality, and mixes each feature through methods such as multiplying and adding corresponding position elements, alleviating the problem of inconsistent representation of raw data in different modalities. The late fusion trains different models for different modalities, and then outputs the learning results of multiple models through methods such as maximum combination, mean combination, or ensemble learning. This method solves the asynchrony of modal data processing and improves the scalability of the modality, but ignores the mutual correlation between each modality and lacks the representation of shared information. The cross-learning fusion strengthens the transmission of shared information between modalities by mapping the feature representations of different modalities to a common semantic representation space, but cannot distinguish redundant information in the shared information. The deep fusion of multi-modal features is one of the key links of the ERC technology. The design of this part directly affects the inference accuracy of the final emotional state category. Therefore, traditional multi-modal emotion recognition methods based on deep neural networks often align the multi-modal feature semantics through deep representation and cross-modal shared semantic space learning, aiming to effectively fuse cross-modal features and infer potential emotional categories on the basis of cross-modal information consistency learning and information complementarity. However, such methods generally use the design of loss functions and single-stage simple fusion of features (such as early and late fusion) to achieve cross-modal feature semantic alignment and complementary utilization, which easily leads to the loss and misunderstanding of emotional semantic information (modality-specific semantics) contained in each modality. In addition, existing ERC methods still have the problem of insufficient perception granularity of emotional semantic information in each modality. These two limitations make it difficult for existing ERC methods to accurately and robustly identify the emotional state of learners in complex teaching and cognitive interaction scenarios.

[0007] Although the traditional graph neural network GCN has been very successful, variants of convolutional networks based on GCN such as GCN2 and GraphSage have also improved in algorithm complexity. Generally speaking, these models have deficiencies. For example, GCN has the problem of over-smoothing. When the number of network layers is stacked, the model performance decreases significantly. This phenomenon limits the development of traditional graph convolutional networks and cannot obtain the maximum utilization of raw data. At the same time, in many large dialogue scenarios, these traditional network models cannot distinguish which information has a long-term impact on the emotions of the interlocutors, that is, the long-term topic factors in the dialogue, and which information is the reason for the sudden change of the interlocutor's emotions, that is, the short-term sensitive information in the dialogue.

[0008] Generally speaking, however, the accuracy and robustness of using existing technologies for the task of dialogue emotion recognition of learners still need to be improved.

[0009] Therefore, based on the above problems, there is an urgent need to provide a new method for dialogue emotion recognition to improve the accuracy and robustness of dialogue emotion recognition. Summary of the Invention

[0010] The purpose of this application is to provide a method and device for dialogue emotion recognition based on a multimodal large model, which can improve the accuracy and robustness of dialogue emotion recognition.

[0011] To achieve the above purpose, this application provides the following solutions:

[0012] First Invention This application provides a method for dialogue emotion recognition based on a multimodal large model. The method for dialogue emotion recognition based on a multimodal large model includes:

[0013] Obtain all the sentence sets in the dialogue in the current scenario; each sentence includes three modalities: audio, video, and text;

[0014] Build a dialogue sentiment recognition model; the dialogue sentiment recognition model includes: a feature extraction layer, a bidirectional gated unit, a multimodal large model, a BERT language model, a modal information complementary module, a multi-layer residual graph convolutional network based on a semantic graph, and a fully connected layer; the feature extraction layer is used to extract the initial features of each modality respectively; the bidirectional gated unit is used to extract the dialogue person information embedding; the multimodal large model is used to extract the fine-grained sentiment clue features of each modality; and fine-tune the BERT language model; the fine-tuned BERT language model is used to encode the fine-grained sentiment clue features of each modality, and use the multimodal large model to splice the encoded features with the initial features of each modality to obtain the final features of the corresponding modality; the modal information complementary module is used to obtain the initial feature vector representation of complementary learning of the corresponding modality according to the final features of each modality and the dialogue person information embedding; the complementary learning determines the attention weight score matrix between different modalities according to the initial feature vector representations corresponding to different modalities; and according to the attention weight score matrix between different modalities, adopt attention operation to obtain the modality-shared information feature and the modality-unique information feature, and then adopt the differential regularization operator loss to ensure the information discrimination of the modality-shared information feature and the modality-unique information feature; the multi-layer residual graph convolutional network based on the semantic graph is used to splice the modality-shared information feature and the modality-unique information feature to obtain the modality-completed feature; and use the modality-completed feature as a node to construct a semantic relationship graph; furthermore, use the multi-layer residual graph convolutional network to perform graph representation learning on the semantic relationship graph to obtain the multi-layer residual graph convolutional network; and extract the long-term factors and emotional transient factors that affect the emotional transformation of the learner according to the multi-layer residual graph convolutional network; the long-term factor is the global and long-term dialogue information of the semantic graph node features obtained by using shallow one-dimensional convolution in the long-term channel; the emotional transient factor is the local and short-term dialogue information after deep one-dimensional convolution and non-linear mapping through the short-term channel; the fully connected layer is used to determine the sentiment recognition result according to the long-term factors and emotional transient factors that affect the emotional transformation of the learner;

[0015] According to all the statement sets in the dialogue, use the trained dialogue sentiment recognition model to obtain the sentiment recognition result.

[0016] Optionally, the feature extraction layer specifically includes: a multi-layer perceptron, a long short-term memory network, and a text-based multi-layer perceptron;

[0017] Use the multi-layer perceptron to perform corresponding initial feature extraction on the original unprocessed features of the audio modality;

[0018] Use the long short-term memory network to perform corresponding initial feature extraction on the original unprocessed features of the video modality;

[0019] Use the text-based multi-layer perceptron to perform corresponding initial feature extraction on the original unprocessed features of the text modality.

[0020] Optionally, use the formula to determine the final feature of each modality ;

[0021] where () is the splicing function, is the initial feature of each modality, () is the encoding function of the BERT language model, is the fine-grained sentiment cue feature of each modality, and the parameter , is the modal video, is the audio modality, is the text modality.

[0022] Optionally, the modal information complementary module is used to obtain the initial feature vector representation of complementary learning of the corresponding modality according to the final feature of each modality and the interlocutor information embedding; the complementary learning determines the attention weight score matrix between different modalities according to the initial feature vector representation corresponding to different modalities; and according to the attention weight score matrix between different modalities, the attention operation is used to obtain the modality-shared information feature and the modality-unique information feature, and then the differential regularization operator loss is used to ensure the information discrimination between the modality-shared information feature and the modality-unique information feature, which specifically includes:

[0023] Use the formula to determine the initial feature expression for complementary learning ;

[0024] Use the formula to determine the attention weight score matrix between the m th modality and the n th modality ;

[0025] Use the formula to determine the modality-shared information feature l of the text modality ;

[0026] Use the formula to determine the modality-unique information feature;

[0027] Use the formula to ensure the discrimination between the modality-shared information features of all modalities and the modality-unique information features ;

[0028] where is the interlocutor information embedding, , represents the speaker group At the layer hidden state, , is the set of words spoken by all speaker groups . Indicates that for the speaker group At the -1 layer hidden state, is the original input statement, () is a mathematical function, represents the mapping matrix, is the cross-modal semantic information, , represents the identity matrix, is a hyperparameter that controls the degree of retention of the original modal information, is the attention weight score matrix, with the superscript T being the transpose, is the value matrix, () represents the vectorization operation, is a function to obtain the modality-specific information, is the differential regularization operator loss, F represents the Frobenius norm, is the query matrix for the m modality, is the key matrix for the n modality, is the scaling factor used to control the tightness of the attention calculation.

[0029] Optionally, in the semantic relationship graph, the formula is used to determine the edge weights ;

[0030] Among them, () is the cosine similarity calculation function.

[0031] Optionally, the formula is used to determine the loss function for training the dialogue sentiment recognition model ;

[0032] Among them, is the total classification loss, is the preliminary classification loss calculated from the modality complementary feature information and the true label, and are hyperparameters that balance the two parts of the loss, is the set of modality complementary features, is the set of true labels.

[0033] In a second aspect, the present application provides a dialogue sentiment recognition device based on a multi-modal large model. The dialogue sentiment recognition device based on a multi-modal large model includes:

[0034] A data acquisition module, configured to acquire all sets of statements in a conversation in the current scenario; each statement includes three modalities: audio, video, and text.

[0035] A model construction module, configured to construct a dialogue sentiment recognition model; the dialogue sentiment recognition model includes: a feature extraction layer, a bidirectional gated unit, a multi-modal large model, a BERT language model, a modality information complementary module, a multi-layer residual graph convolutional network based on a semantic graph, and a fully connected layer; the feature extraction layer is configured to extract initial features of each modality respectively; the bidirectional gated unit is configured to extract dialogue participant information embeddings; the multi-modal large model is configured to extract fine-grained sentiment clue features of each modality; and fine-tune the BERT language model; the fine-tuned BERT language model is configured to encode the fine-grained sentiment clue features of each modality, and use the multi-modal large model to splice the encoded features with the initial features of each modality to obtain the final features of the corresponding modality; the modality information complementary module is configured to obtain an initial feature vector representation of complementary learning of the corresponding modality according to the final features of each modality and the dialogue participant information embeddings; the complementary learning determines an attention weight score matrix between different modalities according to the initial feature vector representations corresponding to different modalities; and according to the attention weight score matrix between different modalities, perform an attention operation to obtain modality-shared information features and modality-unique information features, and then use a differential regularization operator loss to ensure the information distinguishability of the modality-shared information features and the modality-unique information features; the multi-layer residual graph convolutional network based on the semantic graph is configured to splice the modality-shared information features and the modality-unique information features to obtain modality-completed features; and construct a semantic relationship graph with the modality-completed features as nodes; furthermore, use the multi-layer residual graph convolutional network to perform graph representation learning on the semantic relationship graph to obtain the multi-layer residual graph convolutional network; and extract long-term factors and emotional transient factors that affect the emotional transformation of the learner according to the multi-layer residual graph convolutional network; the long-term factors are the global and long-term dialogue information of the semantic graph nodes obtained by using a shallow one-dimensional convolution in the long-term channel; the emotional transient factors are the local and short-term dialogue information after non-linear mapping through a deep one-dimensional convolution in the short-term channel; the fully connected layer is configured to determine the sentiment recognition result according to the long-term factors and the emotional transient factors that affect the emotional transformation of the learner.

[0036] A recognition result output module, configured to obtain a sentiment recognition result by using the trained dialogue sentiment recognition model according to all sets of statements in the conversation.

[0037] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the dialogue sentiment recognition method based on the multi-modal large model.

[0038] According to the specific embodiments provided in the present application, the present application has the following technical effects:

[0039] The present application provides a method and device for dialogue emotion recognition based on a multi-modal large model. The constructed dialogue emotion recognition model includes: a feature extraction layer, a bidirectional gated unit, a multi-modal large model, a BERT language model, a modal information complementary module, a multi-layer residual graph convolutional network based on a semantic graph, and a fully connected layer. Among them, through the modal information complementary module (AMC module), the attention weight scores of different modalities are calculated to extract features between target modalities, which maximally ensures information sharing between modalities. At the same time, different from the cross-learning fusion in the prior art, the proposed modal information complementary module can independently extract modal shared and independent information, so as to judge and attribute emotion information from multiple perspectives and improve the robustness of the algorithm. The modal information complementary module performs cross-modal information extraction, which can respectively take the text, video, and audio modalities as target modalities (in parallel) and extract valuable information between it and the other two complementary modalities. By introducing a multi-modal large model and utilizing its ability to understand fine-grained emotion clues in each modality's information, the implicit fine-grained emotion information in different modalities is shared and mapped to the text semantic space. With this kind of multi-modal fine-grained emotion clue semantic information gain, the accuracy and robustness of the learner's conversation emotion recognition in complex teaching scenarios are improved. In order to capture this special information to the greatest extent and reduce the interference of redundant information, the finally obtained features are more reliable and richer. The present application proposes a multi-layer residual graph convolutional network, which solves the problem that traditional graph neural networks (GCNs) cannot be stacked in multiple layers while capturing dialogue emotion-sensitive information, improves the learning effect of the model, and also provides a new tool for the efficient extraction of multi-modal dialogue-sensitive data. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0041] Figure 1 It is a schematic flowchart of a method for dialogue emotion recognition based on a multi-modal large model in an embodiment of the present application;

[0042] Figure 2 It is a schematic diagram of a multi-modal learner dialogue scenario;

[0043] Figure 3 It is a schematic diagram of the structure of a dialogue emotion recognition model;

[0044] Figure 4Schematic diagram of cross-modal instruction examples in the video modality;

[0045] Figure 5 Schematic diagram of cross-modal instruction examples in the audio modality;

[0046] Figure 6 Schematic diagram of instruction examples in the text modality;

[0047] Figure 7 Schematic diagram of a multi-layer residual graph convolutional network based on a semantic graph. Detailed implementation manners

[0048] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0049] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0050] In an exemplary embodiment, as Figure 1 shown, a method for dialogue emotion recognition based on a multi-modal large model is provided, and this method includes the following S101 to S103. Among them:

[0051] S101, obtaining all the statement sets in the dialogue in the current scenario; each statement includes three modalities: audio, video, and text;

[0052] As Figure 2 shown, define a dialogue in a scenario , , where represents the total number of all statements in this dialogue, respectively represent the original inputs of the three modalities (audio, video, text) corresponding to the rd sentence. Define the speaker set in this dialogue, where is the total number of speakers. For the th sentence, which is stated by the st speaker, then this statement-speaker mapping relationship can be defined as . For the set of words spoken by all speaker groups it can be defined as , and the goal of the present application is to construct a large semantic relationship graph And learn and obtain the current utterance on the graph through the proposed multi-layer residual graph convolutional network model corresponding sentiment label .

[0053] S102. Construct a dialogue sentiment recognition model. As Figure 3 shown, the dialogue sentiment recognition model includes: a feature extraction layer, a bidirectional gated unit, a multi-modal large model, a BERT language model, a modal information complementary module, a multi-layer residual graph convolutional network based on a semantic graph, and a fully connected layer. The feature extraction layer is used to extract the initial features of each modality respectively. The bidirectional gated unit is used to extract the dialogue person information embedding. The multi-modal large model is used to extract the fine-grained sentiment clue features of each modality, and fine-tune the BERT language model. The fine-tuned BERT language model is used to encode the fine-grained sentiment clue features of each modality, and use the multi-modal large model to splice the encoded features with the initial features of each modality to obtain the final features of the corresponding modality. The modal information complementary module is used to obtain the initial feature vector representation of complementary learning of the corresponding modality according to the final features of each modality and the dialogue person information embedding. The complementary learning determines the attention weight score matrix between different modalities according to the initial feature representations corresponding to different modalities, and according to the attention weight score matrix between different modalities, uses attention operation to obtain the modality-shared information feature and the modality-unique information feature, and then uses the differential regularization operator loss to ensure the information discrimination of the modality-shared information feature and the modality-unique information feature, avoiding model collapse caused by non-discrimination of the two pieces of information. The multi-layer residual graph convolutional network based on the semantic graph is used to splice the modality-shared information feature and the modality-unique information feature to obtain the modality-completed feature, and construct a semantic relationship graph with the modality-completed feature as the node. Furthermore, use the multi-layer residual graph convolutional network to perform graph representation learning on the semantic relationship graph to obtain the multi-layer residual graph convolutional network, and extract the long-term factors and emotional transient factors that affect the emotional transformation of the learner according to the multi-layer residual graph convolutional network. The long-term factor is the global and long-term dialogue information of the semantic graph node features obtained by using shallow one-dimensional convolution in the long-term channel. The emotional transient factor is the local and short-term dialogue information after deep one-dimensional convolution and non-linear mapping through the short-term channel. The fully connected layer is used to determine the sentiment recognition result according to the long-term factors and emotional transient factors that affect the emotional transformation of the learner

[0054] The feature extraction layer specifically includes: a multi-layer perceptron, a long short-term memory network, and a text-based multi-layer perceptron

[0055] For the audio modality, perform feature extraction on the original feature input through a multi-layer perceptron (MLP), and the calculation is as follows .

[0056] ;

[0057] For the video modality, due to its strong context relevance, the long short-term memory network (LSTM) is used to extract the original video features and is expressed as:

[0058] ;

[0059] Specifically, for the audio and video modalities, the feature extraction is expressed as:

[0060] ;

[0061] For the original text modality features, the text-based multi-layer perceptron (TMLP) is used to extract the features and is expressed as:

[0062] ;

[0063] where represents the output after the first feature encoding, that is, the initial input statement; , represents the learnable parameters. At the same time, in order to simulate the scenario in real conversations, the bidirectional gated unit is used here to extract the interlocutor information embedding so as to improve the accuracy of emotion recognition, which is expressed as:

[0064] ;

[0065] where represents the hidden state of the speaker group at the th layer, , is the set of all the words spoken by the speaker group , represents the hidden state of the speaker group at the -1th layer, is the original input statement;

[0066] Specifically, when the human brain analyzes multi-modal learner data such as video, audio, and text and judges the target emotional state, there is an implicit mining process for the multi-dimensional fine-grained emotional clues of each modality. This process provides rich granular-level valuable information for the human brain to comprehensively analyze and complementarily utilize multi-modal emotional semantics. Therefore, a multi-modal large model is introduced in the ERC task. Using its powerful cross-modal semantic understanding ability, it fully and effectively mines the fine-grained emotional clue information in the learner data of each modality to improve the sensitivity and accuracy of the ERC model's perception of the learner's emotional state.

[0067] The multi-modal large model "iFlytek Spark" developed by iFlytek is used as the backbone network of the multi-modal fine-grained emotion clue mining module. It takes video, audio, and text three-modal rich semantic data as input and outputs the semantic text information of the emotion clues at the granularity level of each modality. This process explicitly completes the shared mapping of the fine-grained emotion clue information of the three modalities of video, audio, and text to the text feature semantic space, reducing the natural semantic gap between the emotion information of each modality.

[0068] Cross-modal instructions for fine-grained emotion clue semantic mining in video, audio, and text three-modal data, and valuable emotion clue text information of each modality is obtained in JSON format. Among them, the cross-modal instructions take emotion state perception as an implicit goal and the output of fine-grained emotion clues in the three modalities as an explicit goal.

[0069] Specifically, for the video modality, cross-modal instructions are used to guide the multi-modal large model to perceive, understand, and output the fine-grained visual emotion clues related to the learner's emotion state in the video , including the muscle change information related to the learner's facial expression in the video modality (muscle changes in parts such as the corners of the eyes and eyebrows), body movements, etc. The instruction design and examples are as Figure 4 shown.

[0070] For the audio modality, the cross-modal instructions mainly aim to obtain the fine-grained auditory emotion clues explicitly related to the learner's emotion state, such as the tone, intonation, and speech rate of the audio as the goal. The instruction design and examples are as Figure 5 shown.

[0071] For the text modality, it focuses on obtaining the fine-grained language emotion clues that reflect the learner's emotion state, such as the context emotion semantics, sentence structure, and special punctuation in the text sequence, through the multi-modal large model , and the instruction design and examples are as Figure 6 shown.

[0072] After obtaining the semantic text information of the fine-grained emotion clues in the three modalities of the learner's dialogue video, audio, and text sentences , and , the BERT language model pre-trained on a large-scale Chinese corpus and fine-tuned on an emotion information corpus is used to encode the semantic text of the fine-grained emotion clues to obtain the fine-grained emotion clue semantic embeddings, and they are respectively concatenated with the initial features of the corresponding modalities to obtain the final features of the video, audio, and text three modalities before complementary learning of modality information are expressed as:

[0073] ;

[0074] Among them, () is a splicing function, is the initial feature of each modality, () is the encoding function of the BERT language model, is the fine-grained sentiment cue feature of each modality, parameter , is the modal video, is the audio modality, is the text modality.

[0075] Since single-modal data rarely has the ability to provide enough information for accurate evaluation, a mature and reliable sentiment recognition model should be able to utilize the commonalities and complementarities of multi-modal data. Generally speaking, the commonalities between multi-modal data are considered to represent the consistency of the dialogue or sentiment polarity. On the contrary, complementary information means specific emotions or emotional transitions. Intuitively, combining the commonalities and unique information between modalities can improve the accuracy of sentiment recognition. Therefore, the initial feature expression for complementary learning is obtained by splicing the interlocutor embedding and the initial feature , expressed as:

[0076] ;

[0077] Then for all statement features , there is a mapping relationship , where is used to extract modality-shared information from the statement features, while is used to extract modality-unique information from the statement features, expressed as:

[0078] ;

[0079] where ( ), ( ) represents the modality-shared information and modality-unique information corresponding to each modality. Based on this, for a given statement 's modality complementary representation can be expressed as:

[0080] ;

[0081] Regarding the feature embeddings of the three modalities as word vectors, the modality-shared and modality-unique semantic information can be defined as cross-modal attention. Taking the text modality as an example, for a given text feature , the query matrix can be defined as , similarly, the key matrix is , and the value matrix is , where and If it is a conversion matrix, then use the formula to determine the attention weight score matrix m between the n th mode and the th mode; based on , the cross-modal semantic information extraction can be expressed as:

[0082] ;

[0083] Among them, represents the identity matrix, = 1, which is a hyperparameter for controlling the retention degree of the original modal information. Adding the identity matrix is to prevent gradient dissipation. Therefore, for the text modality, its modal shared information feature can be expressed as:

[0084] ;

[0085] Among them, represents the mapping matrix, () represents the vectorization operation, then the mapping relationship can be expressed as:

[0086] ;

[0087] Among them, is the function definition symbol, The m-th row of represents the degree of association of the m-th mode with other modes. Regarding the m-th mode as the recognition target, the unique semantic information in it, including sudden emotional changes in the discourse, can be reflected by the degree of difference in the attention of the m-th mode to each other mode. Therefore, the attention matrix of each utterance in the dialogue can be regarded as a representation in the modal unique embedding subspace, used to represent the difference in unique information between modes. Then the mapping relationship can be expressed as:

[0088] ;

[0089] To enhance the difference between modal shared information and modal unique information and reduce the repeated extraction of redundant information, a differential regularization operator is introduced, expressed as:

[0090] ;

[0091] Among them, is the modal unique information feature, the superscript T is the transpose, is the value matrix, () represents the vectorization operation, is the differential regularization operator loss, Fdenotes the Frobenius norm, is the query matrix of the m-th modality, is the key matrix of the n-th modality, is the scale factor, which is used to control the tightness of attention calculation.

[0092] Graph models are naturally suitable for processing heterogeneous relationships, can more flexibly model the relationships between different types of edges and nodes, and can better capture the changing process of emotions through the dynamic evolution of graphs at different time points or contexts. Considering that in many emotion recognition scenarios, different modalities may affect each other, in order to capture such interactions between modalities, a semantic graph based on modality-completed features is constructed to improve the model recognition accuracy. Specifically, in a conversation, each sentence can be represented as three types of nodes , representing audio nodes, video nodes, and text nodes respectively, , then for N sentences, there are a total of nodes, , in order to establish the connection between utterances and different modalities, whether there is an edge connection between two nodes depends on the following rules: 1: Connect any two nodes of the same modality in the same utterance. 2: Each node maintains a connection with nodes from different modalities but corresponding to the same utterance. 3: Any utterance should be connected to its three modality nodes and . To distinguish the importance of different adjacent nodes, different weights are assigned to the edges according to their importance, that is, a larger weight is assigned to nodes that are more important than other nodes. The weight of the edge can be expressed as:

[0093] ;

[0094] where, () is the cosine similarity calculation function.

[0095] After obtaining the semantic graph , in order to obtain an accurate feature representation for the sentence , for the feature of each node , it is initialized to . Specifically, since the relationship between utterances in a conversation is changing, it is necessary to retain the historical interaction information of the interlocutors, and at the same time, retaining the current interaction relationship information also plays a crucial role in the ERC task. In order to balance the interactive part that should be retained or discarded, this application sets up a short-term storage unit to retain the current interaction state or emotional mutation information. is initialized to 0, indicating that there is no current interaction to retain. At the same time, a long-term memory unit is used to store historical interaction information or topic information of a conversation, such as Figure 7 shown.

[0096] such as Figure 7 shown, an improved graph convolution operation is selected to aggregate complementary semantics in a unique subspace. Here, considering the input semantic graph as a sequence, the input graph stream can be represented as , where represents the semantic graphs of different layers, and the message passing process can be represented as:

[0097] ;

[0098] where represents the regularized graph Laplacian matrix, is the regularized diagonal matrix, is the regularized adjacency matrix, and are hyperparameters, , is also a hyperparameter, is the weight matrix, is the identity matrix, is the node feature matrix of the semantic graph at the -th layer. By using an improved graph convolutional network instead of a common graph convolutional network, strict control of the filter is provided by avoiding explicit use of the graph Fourier basis, and the computational complexity is reduced. , is the node feature of the semantic graph at the -th layer, is the -th short-term storage unit, indicating that the current interaction information is embedded into the graph matrix . Meanwhile, in order to accurately retain and clear the interaction information, there is:

[0099] ;

[0100] where represents the message passing gate, , are learnable parameters, represents the sigmoid function. For the parameter , there is controlling the information to be written into the storage unit, controlling the redundant interaction information to be erased from the memory unit, controlling the final output feature representation. Meanwhile, the hidden state update of the long-term memory unit can be represented as:

[0101] ;

[0102] Among them, represents the activation function, , represents the learnable parameter, then the update of the long-term memory unit can be expressed as:

[0103] ;

[0104] Among them, represents the dot product calculation operation. This memory unit stores the information that should be forgotten from the previous layer and also stores the updated information from the current layer. Stacking layers of the above structure to enrich the final feature representation, the final feature calculation process can be expressed as:

[0105] ;

[0106] Among them, MREGCN is the proposed multi-layer residual graph convolutional network, are the set of semantic graph nodes, the set of semantic graph edges, and the node feature representation of the th layer and the th layer in sequence, The feature is initialized as the modality complementary feature ;

[0107] After encoding through layers, the final output feature matrix is obtained. In addition, in order to retain the original interaction information, the final feature is expressed as:

[0108] ;

[0109] Among them, represents the concatenation operation, represents the learnable parameter. The finally obtained comprehensively considers the information of 6 different levels ( ), and is used for the final sentiment classification task. The prediction process can be expressed as:

[0110] ;

[0111] ;

[0112] ;

[0113] Among them, represents the feature vector of the sentence ​ Represents the intermediate hidden layer feature representation of the fully connected layer during prediction, is the output of the model, indicating for the statement the sentiment probability distribution, represents for the statement the final sentiment prediction result (taking the sentiment category corresponding to the index value of the maximum probability in ), and represent the activation function and the normalization function, , , , are the learnable parameters of the fully connected layer adopted in the prediction process, and the loss of the classification prediction result can be expressed as:

[0114] ;

[0115] Among them, represents all conversation samples, represents all statements in the conversation, represents for the sample the class label of the i-th sentence, represents the predicted label, and respectively represent the learnable parameter and the regularization coefficient. Based on this, the loss function of the prediction result finally used to train the model is:

[0116] ;

[0117] Among them, is the total classification loss, is the preliminary classification loss calculated from the modal complementary feature information and the true label, is the difference regularization operator loss, and are the hyperparameters for balancing the losses of the two parts, is the set of modal complementary features, is the set of true labels.

[0118] S103. According to all the statement sets in the conversation, use the trained conversation sentiment recognition model to obtain the sentiment recognition result.

[0119] Based on the same inventive concept, an embodiment of the present application further provides a dialogue emotion recognition device for implementing the dialogue emotion recognition method involved above. The implementation solutions provided by this device to solve problems are similar to those recorded in the above method. Therefore, the specific limitations in one or more embodiments of the dialogue emotion recognition device provided below can refer to the limitations on the dialogue emotion recognition method in the above text and will not be elaborated here.

[0120] In an exemplary embodiment, a dialogue emotion recognition device based on a multimodal large model is provided, including:

[0121] A data acquisition module, configured to acquire all statement sets in the dialogue in the current scenario; each statement includes three modalities: audio, video, and text.

[0122] A model construction module, configured to construct a dialogue emotion recognition model; the dialogue emotion recognition model includes: a feature extraction layer, a bidirectional gated unit, a multimodal large model, a BERT language model, a modality information complementary module, a multi-layer residual graph convolutional network based on a semantic graph, and a fully connected layer; the feature extraction layer is configured to extract initial features of each modality respectively; the bidirectional gated unit is configured to extract dialogue person information embeddings; the multimodal large model is configured to extract fine-grained emotion clue features of each modality; and fine-tune the BERT language model; the fine-tuned BERT language model is configured to encode the fine-grained emotion clue features of each modality, and use the multimodal large model to splice the encoded features with the initial features of each modality to obtain the final features of the corresponding modality; the modality information complementary module is configured to obtain an initial feature vector representation of complementary learning for the corresponding modality according to the final features of each modality and the dialogue person information embeddings; the complementary learning determines an attention weight score matrix between different modalities according to the initial feature vector representations corresponding to different modalities; and according to the attention weight score matrix between different modalities, an attention operation is used to obtain modality shared information features and modality unique information features, and then a differential regularization operator loss is used to ensure the information discrimination between the modality shared information features and the modality unique information features; the multi-layer residual graph convolutional network based on the semantic graph is configured to splice the modality shared information features and the modality unique information features to obtain modality completion features; and construct a semantic relationship graph with the modality completion features as nodes; furthermore, the multi-layer residual graph convolutional network is used to perform graph representation learning on the semantic relationship graph to obtain the multi-layer residual graph convolutional network; and according to the multi-layer residual graph convolutional network, long-term factors and emotion transient factors that affect the emotion change of the learner are extracted; the long-term factors are the global and long-term dialogue information of the semantic graph node features obtained by the long-term channel using shallow one-dimensional convolution; the emotion transient factors are the local and short-term dialogue information after deep one-dimensional convolution and non-linear mapping through the short-term channel; the fully connected layer is configured to determine the emotion recognition result according to the long-term factors and the emotion transient factors that affect the emotion change of the learner.

[0123] The recognition result output module is used to obtain the sentiment recognition result by using the trained dialogue sentiment recognition model based on all the statement sets in the dialogue.

[0124] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for dialogue sentiment recognition based on a multimodal large model.

[0125] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor implements the steps in the above method embodiments.

[0126] In an exemplary embodiment, a computer program product is provided, including a computer program, which when executed by a processor implements the steps in the above method embodiments.

[0127] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0128] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0129] In the embodiments provided in the present application, the databases involved can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. In the embodiments provided in the present application, the processors involved can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0130] In the present application, all actions of obtaining signals, information, or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining authorization from the owner of the corresponding device.

[0131] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0132] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. To sum up, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A conversation emotion recognition method based on a multimodal large model, characterized in that: The conversation emotion recognition method based on the multimodal large model includes: Get all the sentences in the conversation in the current scene; each sentence includes three modes: audio, video and text; Construct a conversation emotion recognition model; the conversation emotion recognition model includes: a feature extraction layer, a bidirectional gating unit, a multimodal large model, a BERT language model, a modal information complementary module, a multi-layer residual graph convolutional network based on a semantic graph, and a fully connected layer; the feature extraction layer is used to extract the initial features of each modality respectively; the bidirectional gating unit is used to extract the interlocutor information embedding; the multimodal large model is used to extract the fine-grained emotional clue features of each modality; and the BERT language model is fine-tuned; the fine-tuned BERT language model is used to encode the fine-grained emotional clue features of each modality, and the encoded features are spliced ​​with the initial features of each modality using the multimodal large model to obtain the final features of the corresponding modality; the modal information complementary module is used to obtain the initial feature vector representation of the complementary learning of the corresponding modality according to the final features of each modality and the interlocutor information embedding; the complementary learning determines the attention weight score matrix between different modalities according to the initial feature vector representations corresponding to different modalities; and according to different The attention weight score matrix between the same modalities uses attention operation to obtain modal shared information features and modal unique information features, and then uses differential regularization operator loss to ensure the information discrimination of modal shared information features and modal unique information features; the multi-layer residual graph convolution network based on semantic graph is used to splice modal shared information features and modal unique information features to obtain modal completion features; and the modal completion features are used as nodes to construct a semantic relationship graph; and then the multi-layer residual graph convolution network is used to perform graph representation learning on the semantic relationship graph to obtain a multi-layer residual graph convolution network; and the long-term factors and emotional transient factors that affect the learner's emotional transformation are extracted based on the multi-layer residual graph convolution network; the long-term factor is a long-term channel that uses shallow one-dimensional convolution to obtain the global and long-term dialogue information of the semantic graph node features; the emotional transient factor is a short-term channel that uses deep one-dimensional convolution to obtain the local and short-term dialogue information after nonlinear mapping; the fully connected layer is used to determine the emotion recognition result based on the long-term factors and emotional transient factors that affect the learner's emotional transformation; According to the set of all sentences in the conversation, the emotion recognition result is obtained by using the trained conversation emotion recognition model.

2. The method for dialogue emotion recognition based on a multimodal large model according to claim 1 is characterized in that: The feature extraction layer specifically includes: a multi-layer perceptron, a long short-term memory network, and a text-based multi-layer perceptron; Use a multi-layer perceptron to extract the corresponding initial features of the original unprocessed features of the audio modality; The long short-term memory network is used to extract the corresponding initial features of the original unprocessed features of the video modality; A text-based multi-layer perceptron is used to extract the corresponding initial features of the original unprocessed features of the text modality.

3. The method for dialogue emotion recognition based on a multimodal large model according to claim 1 is characterized in that: Using the formula Determine the final characteristics of each mode ; in, () is the concatenation function, is the initial feature of each mode, () is the encoding function of the BERT language model, is the fine-grained emotional cue feature of each modality, parameter , For modal video, For audio mode, For text mode.

4. The method for dialogue emotion recognition based on a multimodal large model according to claim 3 is characterized in that: The modal information complementation module is used to obtain the initial feature vector representation of the complementary learning of the corresponding modality according to the final features of each modality and the information embedding of the interlocutor; the complementary learning determines the attention weight score matrix between different modalities according to the initial feature vector representation corresponding to different modalities; and according to the attention weight score matrix between different modalities, the attention operation is used to obtain the modal shared information features and the modal unique information features, and then the differential regularization operator loss is used to ensure the information discrimination of the modal shared information features and the modal unique information features, specifically including: Using the formula Determining initial feature representations for complementary learning ; Using the formula Determine m The mode and n The attention weight score matrix between modalities ; Using the formula Determine text mode l The modality sharing information feature ; Using the formula Determine the unique information characteristics of the modality; Using the formula Ensure modality sharing of information features across all modalities Unique information features of the modality The degree of distinction between in, To embed the conversation person information, , For speaker groups In the The hidden state of the layer, , For all speaker groups A collection of words spoken, For speaker groups In the -1 layer of hidden state, is the original input statement, () is a mathematical function, represents the mapping matrix, is the cross-modal semantic information, , represents the identity matrix, is a hyperparameter that controls the degree of preservation of the original modal information. is the attention weight score matrix, the superscript T is the transpose, is the value matrix, () indicates vectorized operation, A function to obtain modality-specific information. is the difference regularization operator loss, F represents the Frobenius norm, is the query matrix of m modes, is the bond matrix of n modes, is a scale factor used to control the tightness of attention calculation.

5. The method for dialogue emotion recognition based on a multimodal large model according to claim 4 is characterized in that: Using formula in semantic relationship graph Determine edge weights ; in, ()Cosine similarity calculation function.

6. The method for dialogue emotion recognition based on a multimodal large model according to claim 5 is characterized in that: Using the formula Determine the loss function for training the conversation emotion recognition model ; in, is the total classification loss, The preliminary classification loss calculated for the modality complementary feature information and the true label, and To balance the hyperparameters of the two losses, is the set of modal complementary features, is the true label set.

7. A conversation emotion recognition device based on a multimodal large model, characterized in that: The conversation emotion recognition device based on the multimodal large model includes: The data acquisition module is used to obtain the set of all sentences in the conversation in the current scene; each sentence includes three modes: audio, video and text; A model building module is used to build a conversation emotion recognition model; the conversation emotion recognition model includes: a feature extraction layer, a bidirectional gating unit, a multimodal large model, a BERT language model, a modal information complementation module, a multi-layer residual graph convolutional network based on a semantic graph, and a fully connected layer; the feature extraction layer is used to extract the initial features of each modality respectively; the bidirectional gating unit is used to extract the interlocutor information embedding; the multimodal large model is used to extract the fine-grained emotional clue features of each modality; and the BERT language model is fine-tuned; the fine-tuned BERT language model is used to encode the fine-grained emotional clue features of each modality, and the encoded features are spliced ​​with the initial features of each modality using the multimodal large model to obtain the final features of the corresponding modality; the modal information complementation module is used to obtain the initial feature vector representation of the complementary learning of the corresponding modality according to the final features of each modality and the interlocutor information embedding; the complementary learning determines the attention weight score matrix between different modalities according to the initial feature vector representations corresponding to different modalities; According to the attention weight score matrix between different modalities, the attention operation is used to obtain the modal shared information features and the modal unique information features, and then the differential regularization operator loss is used to ensure the information discrimination of the modal shared information features and the modal unique information features; the multi-layer residual graph convolution network based on the semantic graph is used to splice the modal shared information features and the modal unique information features to obtain the modal completion features; and the modal completion features are used as nodes to construct a semantic relationship graph; and then the multi-layer residual graph convolution network is used to perform graph representation learning on the semantic relationship graph to obtain a multi-layer residual graph convolution network; and according to the multi-layer residual graph convolution network, the long-term factors and emotional transient factors that affect the learner's emotional transformation are extracted; the long-term factor is a long-term channel using a shallow one-dimensional convolution to obtain the global and long-term dialogue information of the semantic graph node features; the emotional transient factor is a short-term channel using a deep one-dimensional convolution to obtain the local and short-term dialogue information after nonlinear mapping; the fully connected layer is used to determine the emotion recognition result according to the long-term factors and emotional transient factors that affect the learner's emotional transformation; The recognition result output module is used to obtain the emotion recognition result based on the set of all sentences in the conversation using the trained conversation emotion recognition model.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the conversation emotion recognition method based on a multimodal large model as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method and device

    CN118260711A

  • Private domain live broadcast room hotspot prediction and adaptive routing method based on swarm intelligence

    CN119254688A