Multi-modal fine-grained emotion recognition method oriented to human-computer interaction and based on large model
Through the multimodal fine-grained emotion recognition method of large models, the bilinear attention network and enhanced dependency graph attention network are used to solve the coarse-grained problem of cross-modal emotion understanding in the robot system, and the fine-grained alignment and noise suppression are achieved, which improves the accuracy of emotion recognition and multimodal understanding ability.
Patent Information
- Application Number
- CN202510420762.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has coarse-grained problems in multimodal emotion understanding in robot systems, making it difficult to achieve fine-grained alignment of cross-modal emotion cues, and visual noise interference and multimodal fusion bottlenecks of large models lead to a decrease in interaction accuracy.
A multimodal fine-grained emotion recognition method based on large models is adopted, through fine-grained visual-text alignment and emotional information fusion, a bilinear attention network and an enhanced dependency graph attention network are used to achieve hierarchical alignment of text and visual features and precise capture of emotional information.
It improves the robot system's accurate recognition of user emotions, suppresses environmental noise interference, improves the understanding of multimodal emotional semantics, and supports the application of home service robots and medical accompanying assistants.
Smart Images

Figure CN120296668A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a multi-modal fine-grained emotion recognition method based on a large model for human-computer interaction. Background Art
[0002] With the rapid development of artificial intelligence technology, the application of robot systems and large multi-modal models (such as GPT-4, CLIP, etc.) in the fields of intelligent interaction, service robots, emotional companionship, etc. has become increasingly widespread. Robots need to perceive the user's intention and emotional state in real time through multi-modal data (such as voice, vision, text) to provide natural and personalized responses; while large models need to break through the limitations of a single modality and achieve deep alignment and semantic understanding of cross-modal information. However, the existing technologies still face significant challenges:
[0003] 1. The coarse-grained problem of multi-modal emotion understanding: Traditional robot systems rely on independently processing each modality data (such as analyzing facial expressions or speech separately), lacking fine-grained alignment of cross-modal emotional cues, resulting in misjudgment of complex emotions of users (such as sarcasm, ambivalent emotions).
[0004] 2. Visual noise interference: In dynamic environments (such as homes, public places), the images collected by robots often contain irrelevant backgrounds or occlusions. Existing models are difficult to distinguish visual features related to emotions (such as facial expression details) from noise, affecting the accuracy of interaction.
[0005] 3. The multi-modal fusion bottleneck of large models: Although large models perform excellently in single-modal tasks, their multi-modal interaction mechanisms (such as simple splicing or single attention) are difficult to achieve precise semantic matching between text and images. Especially in aspect-level emotion analysis, they are easily interfered by cross-modal redundant information.
[0006] For example, service robots need to combine the user's language ("This function is too slow") with a frowning expression to judge their dissatisfaction and optimize the response strategy. However, existing methods may fail due to visual noise (such as a cluttered background) or text-image misalignment (such as a contradiction between the description and the expression). In addition, when generating multi-modal content, large models often lack a fine-grained alignment mechanism, resulting in inconsistent semantics in the generated results (such as describing a "happy scene" but with a conflicting picture). Summary of the Invention
[0007] In view of this, the present invention proposes a multi-modal fine-grained emotion recognition method based on a large model for human-computer interaction, which effectively alleviates the above problems through fine-grained visual-text alignment and emotion information fusion.
[0008] A multi-modal emotion recognition method includes the following steps:
[0009] Step 1: Perform multi-modal feature extraction on the input text-image pair: Generate an image description C and a facial expression description D based on the visual features V of the image; then convert the text features including the text, image description, and facial expression description into text embeddings respectively.
[0010] Step 2: Based on the bilinear attention network to capture the pairwise local interactions between text features and image features, including using the bilinear interaction graph to capture pairwise attention weights, and the utilization of the interaction graph by the variant multi-modal residual network to perform selective fusion to achieve hierarchical alignment of text and visual features, specifically including:
[0011] Input the multi-modal features obtained in Step 1 into the BART encoder to obtain multi-modal hidden states, which are respectively: the hidden state H corresponding to the text embedding T T , the hidden state H corresponding to the visual feature V V , the hidden state H corresponding to the image description C , the hidden state H corresponding to the facial expression description D , and H M = {H V , H T};
[0012] Based on the hidden state H T , the hidden state H C , use the bilinear interaction graph to obtain the single-head pairwise interaction matrix of the two; introduce a bilinear pooling layer to obtain the attention map F; then introduce a residual layer to integrate multiple bilinear attention maps to obtain the enhanced text feature H TC ;
[0013] Based on the hidden state H T , the hidden state H D , use the bilinear interaction graph to obtain the single-head pairwise interaction matrix of the two; introduce a bilinear pooling layer to obtain the attention map F; then introduce a residual layer to integrate multiple bilinear attention maps, gradually fuse the fine-grained visual semantics of the facial expression, and finally generate the locally fine-grained aligned joint representation H TD ;
[0014] Finally, learn the parameter β t to selectively fuse the features H TC and H TD to obtain the final text representation: Concatenate H V with , and finally generate the image-text feature
[0015] Step 3: Weightedly fuse the sentiment scores of SenticNet with the multi-modal features obtained by the attention pairing interaction module, construct a multi-modal dependency graph, and aggregate sentiment information through a graph attention network, specifically including:
[0016] For a word t in the input text i , retrieve its sentiment score from SenticNet and project it into the same dimensional space as the multi-modal features to obtain the common sense sentiment feature s i ; Then, integrate this sentiment feature s i into the output of APIM :
[0017]
[0018] where, W s and b s are learnable parameters, represents the feature containing sentiment knowledge; represents the feature representation of the i-th node of the image-text feature ; After the above formula, the hidden state H S rich in common sense sentiment features is obtained, where is the feature representation of the i-th node in H S ;
[0019] Adopt the global attention mechanism to obtain the dependency matrix A TT between the input text words, and use the multi-head attention mechanism to obtain the dependency relationship A VV between the visual features, as well as the dependency relationships A VT and A TV between the visual features and the text features, to obtain the final dependency matrix D:
[0020]
[0021] Adopt the graph attention network graph structure data H S to learn the importance between different nodes and obtain a richer sentiment information output
[0022] Step 4: The BART decoder receives and the previous decoder output Y <t , and predicts the token probability distribution as follows:
[0023]
[0024] where, λ 1 and λ 2 are two hyperparameters, representation For the text part, W represents the text embedding representation of Step 1; Decoder() represents the processing of the BART decoder; softmax() represents the processing of the softmax function on the data in the parentheses C d The embedding representation corresponding to the sentiment label; represents the hidden state of the decoder at time step t; P(y t ) represents the probability distribution of the predicted sentiment label generated at time step t.
[0025] Preferably, in the said Step 2, the single-head pairwise interaction matrix I ∈ R N×Z :
[0026]
[0027] wherein, and are learnable weight matrices of the text and subtitle representations, q ∈ R K is a learnable weight vector, 1 ∈ R K is a fixed all-ones vector, ° represents the Hadamard product; I represents the interaction strength of each word or subphrase in the text with the image description pair.
[0028] Preferably, the specific method for obtaining the enhanced text feature H TC includes:
[0029] Introduce a bilinear pooling layer to obtain a bilinear joint representation F′. Specifically, the k-th element of F′ is calculated as:
[0030]
[0031] I i,j represents the element in I; Concatenate F′ k along the feature dimension to obtain F′, and perform a linear transformation on F′ to obtain the attention map F;
[0032] wherein, U k and V k respectively represent the k-th columns in the text representation weight matrix U and the image description weight matrix V;
[0033] Introduce a residual layer to integrate multiple bilinear attention maps:
[0034] F i+1 = BNA i (F i , H C ; I i )·1 T + F i
[0035] F i is the input feature of the i-th layer in the residual module, and its initial value F0 is the original text feature H T , I i represents the i-th attention map, BNA i represents the i-th bilinear attention module, 1 T represents the transposed all-1 vector, F i+1 represents the output feature map of the (i + 1)-th layer, that is, by I i calculate the weighted interaction between the text and the visual features to generate joint features, and after passing the joint features through batch normalization and linear transformation, add them to the original feature F i to obtain the updated F i+1 .
[0036] Preferably, the learning parameter β t is expressed as:
[0037] β t = sigmoid(W β [W1H TC ; W2H TD +b β )
[0038] W β is a weight matrix used to perform a linear transformation on the input features to help calculate β t ; W1 is a weight matrix used to perform a linear transformation on the input feature H TC ; W2 is a weight matrix used to perform a linear transformation on the input feature H TD ; b β is a bias vector used to add an offset to the result after the linear transformation to increase the fitting ability of the model.
[0039] Preferably, the method for obtaining the dependency matrix A TT includes: using the Stanza parser to parse the text into a dependency tree G, and converting the dependency tree G into a fixed induced syntactic graph G dis , this graph is fully connected, and using a symmetric adjacency matrix A dis to represent the distance between the i-th and j-th words;
[0040] A dis = tanh -1 (1 - A dis / max_tree_ dis )
[0041] max_tree_dis represents the maximum distance value that may exist between two nodes in a tree; the range of distance attenuation is controlled by adjusting the parameter max_tree_dis;
[0042] Assign numerical identifiers to the types and convert G into the induced syntactic graph D represented by the adjacency matrix A dep ; then initialize the type feature matrix H ∈ R dep , where U represents the total number of dependency types, D represents the dimension of the feature matrix set by oneself, and q ∈ R U×D serves as the transposed query vector; subsequently, use the softmax function to derive the attention weight matrix and apply the gather method; extract the weights corresponding to A D×1 ; A′ dep is the initially obtained initial adjacency matrix, and finally, obtain the type dependency matrix A dep : dep
[0043]
[0044] Finally, obtain the final sub-matrix A by adding the two, TT and its expression is:
[0045] A TT = A dis + A dep .
[0046] Preferably, the method for obtaining the dependency matrix A VV includes: adopting the multi-head attention mechanism to obtain the hidden feature h of the visual feature from H S , and the final sub-matrix A sv is expressed as: VV
[0047]
[0048] h sv is the column element in H V , W i Q is a weight matrix used to project the feature h sv into the query space and is a key parameter for calculating the query vector in the multi-head attention mechanism; is also a weight matrix responsible for projecting the feature h sv into the key space for calculating the key vector and is also a parameter optimized during training. Their role is to help the model extract information from different angles of the input features for attention calculation.
[0049] Preferably, obtain the dependency matrix A VT and A TVThe method includes: obtaining the hidden features h S and h sv of visual and text features from H st , and the final sub-matrices A VT and A TV are expressed as:
[0050]
[0051] Preferably, the graph attention network structure data H S is used to learn the importance between different nodes, and a richer emotional information output is obtained. The method includes:
[0052] Taking H S as the initial node representation in the graph, first calculate the attention coefficient between nodes i and j:
[0053]
[0054] where and are the hidden states of node i and node j respectively, a represents the learnable parameter, || represents the vector concatenation operation, and e ij represents the correlation between nodes i and j;
[0055]
[0056] where M i,j is the weight matrix after masking the attention coefficient using A i,j , and further operations are performed to obtain α ij ;
[0057]
[0058] where K represents the number of attention heads, W k , W out are the learnable weight matrices, and the outputs of multiple attention heads are concatenated and then passed through a linear transformation to obtain the final node feature representation
[0059] Furthermore, it also includes optimizing the model using the cross-entropy loss function, generating the joint prediction results of aspect terms and their sentiment polarities by weighting the contributions of multiple modules with hyperparameters.
[0060] A multi-modal sentiment recognition system implemented using the AVTAF model includes the following modules: a feature extraction module, an attention pairing interaction module APIM, a reinforced dependency graph attention network RD-GAT, and a prediction and optimization module;
[0061] The feature extraction module is used to implement the method in step 1;
[0062] The attention pairing interaction module APIM is used to implement the method in step 2;
[0063] The enhanced dependency graph attention network RD-GAT is used to implement the method in step 3;
[0064] The prediction and optimization module is used to implement the method in step 4.
[0065] The present invention has the following beneficial effects:
[0066] The multi-modal fine-grained emotion recognition method based on a large model for human-computer interaction proposed by the present invention, based on the aspect-driven visual-text alignment and fusion network (AVTAF), realizes cross-modal alignment from coarse-grained to fine-grained through the attention pairing interaction module (APIM), can accurately capture emotion-related visual features (such as facial expressions, gestures) in the robot scenario, and suppress environmental noise; at the same time, the enhanced dependency graph attention network (RD-GAT) improves the reasoning ability of the large model for multi-modal emotion semantics by integrating external emotion knowledge (such as SenticNet). This technology provides a new paradigm for the intelligent upgrade of robot emotion interaction and the multi-modal understanding of large models, and is expected to promote breakthrough applications in fields such as domestic service robots, medical escort assistants, and multi-modal content generation. Description of the Drawings
[0067] Figure 1 It is a schematic diagram of the AVTAF framework of the present invention.
[0068] Figure 2 It is a structural diagram of the attention pairing interaction module (APIM).
[0069] Figure 3 It is the visualization result of the AVTAF model in the multi-modal emotion analysis task. Detailed Embodiments
[0070] 1. System Architecture
[0071] As Figure 1 shown, the AVTAF model constitutes a multi-modal emotion analysis framework, and its core goal is to improve the accuracy of emotion analysis through the fine-grained alignment and fusion of visual and text information. The system architecture of this model mainly includes the following modules: a feature extraction module, an attention pairing interaction module (APIM), an enhanced dependency graph attention network (RD-GAT), and a prediction and optimization module.
[0072] The operation mechanism of the AVTAF model can be divided into several steps:
[0073] First, perform feature extraction. Select to interact with the frozen image encoder for visual feature extraction. To establish a coarse-grained alignment between the global image and text, apply the image captioning tool CATR, which generates high-quality captions for the scene. At the same time, to extract fine-grained emotional visual information, use the facial expression description template to generate facial descriptions. At the same time, for the text input, generate context-aware text features through the BART model.
[0074] Subsequently, design and utilize the APIM module. First, capture the pairwise local interactions between text features and image features through a bilinear attention network, introduce residual layers to integrate multiple bilinear attentions, selectively fuse the coarse-grained alignment and fine-grained alignment established between the global image and text, obtain the final image-text features after visual-text alignment from coarse-grained to fine-grained, and generate a joint representation.
[0075] Furthermore, perform the aggregation of emotional information. Adopt the RD-GAT module to integrate SenticNet emotional knowledge, construct a multi-modal dependency matrix, and aggregate node information through a graph attention network to output a feature representation rich in emotional information.
[0076] Finally, input the outputs of APIM and RD-GAT into the BART decoder to generate an aspect-sentiment pair sequence, and use the cross-entropy loss function to optimize the model, adjust the hyperparameters to balance the contributions of each module, and complete prediction and optimization.
[0077] The AVTAF model can analyze the multi-modal information posted by users, identify the emotional tendency, and can be widely applied in the fields of market research, e-commerce review analysis, intelligent customer service, and human-computer interaction to improve the service personalization and interaction intelligence level.
[0078] 2. Feature Extraction Module
[0079] 2.1 Image Representation
[0080] 2.1.1 Generate Descriptive Captions
[0081] The present invention uses an image encoder to extract visual features. It takes an image as input, extracts general and stable visual features as output after pre-training, and does not need to be retrained after freezing the parameters, saving computing resources and time, avoiding overfitting on small data sets. In addition, the extracted features have strong transferability. The present invention uses V = {v1, v2,..., v 32} to represent visual features. In addition, the present invention uses an image caption model to establish a close connection between the visual and text modalities and convey the semantic essence of visual content at a broader level.
[0082] CATR (Cross-Attention Transformer) is a cross-attention model based on Transformer. It fuses visual and language modalities through cross-attention layers to accurately locate and segment the image regions corresponding to natural language descriptions. To achieve coarse-grained alignment between the global image and text, the present invention uses this tool to generate descriptive captions for the scene, expressed as:
[0083] C = Caption(image) (1)
[0084] 2.1.2 Generating Facial Descriptions
[0085] Facial expressions, as a direct and intuitive way for humans to convey emotions, play a crucial role in accurately identifying object-level emotions in images. Analyses of the TWITTER-2015 and TWITTER-2017 datasets show that a significant portion of images contain facial expressions.
[0086] Yang et al. proposed a facial expression description framework, which is a deep learning-based multimodal sentiment analysis system aiming to more accurately identify and understand facial expressions by combining visual information and text descriptions. This framework first uses a convolutional neural network (CNN) to extract high-level visual features from facial images to capture detailed information about expressions; meanwhile, through natural language processing (NLP) techniques, it generates text descriptions related to the expressions, such as "smiling" or "frowning". Then, a cross-modal attention mechanism or Transformer architecture is adopted to align and fuse the visual features and text descriptions to capture the semantic associations between the two. Finally, based on the fused multimodal features, emotion classification (such as happy, sad, angry, etc.) or emotion intensity regression is performed. The core advantage of this framework is that through the collaborative analysis of vision and text, it can more comprehensively understand the semantics and emotional connotations of facial expressions, significantly improving the accuracy and robustness of sentiment analysis.
[0087] To capture subtle emotional visual cues, the present invention applies the above facial expression description framework to generate detailed facial descriptions.
[0088] D = Face_Description(image) (2)
[0089] 2.2 Text Representation
[0090] For text inputs (such as sentence text, Caption and Face_Description generated from images, etc.), through BART Embedding, high-dimensional and discrete data can be converted into low-dimensional and continuous vectors to obtain initial word embeddings. The initial word embeddings of sentence text are represented as T = {t1, t2,..., t j,...,t n}, where t j refers to the feature of the j-th word in the text. The initial word embedding of the Caption is represented as C’ = {c1, c2,..., c j ,..., c z}, where c j refers to the feature of the j-th word in the Caption. The initial word embedding of the Face Description is represented as D’ = {d1, d2,..., d j ,..., d x}, where d j refers to the feature of the j-th word in the Face Description.
[0091] 3.2.1 BART-based Generation Framework
[0092] 3.2.1.1 Encoder
[0093] The encoder of the model of the present invention uses a multi-layer bidirectional transformer and can process context information simultaneously. To distinguish the inputs from different modalities, special tokens are added: and are placed before and after the visual features to mark their boundaries, while <bos>and <eos>respectively used to represent the start and end of text, Caption, and Face Description features. In the present invention, multimodal features are concatenated as the input X to the BART encoder, and the encoder outputs a multimodal hidden state H M ={H V ,H T} and H C and H D , where
[0094] 3.2.1.2 Decoder
[0095] The decoder of the model of the present invention uses multiple layers of transformers. Different from this, the decoder generates outputs unidirectionally, while the encoder generates outputs bidirectionally. The present invention introduces a special marker <bos>To indicate the start of generation, the Decoder interacts with the output of the Encoder through self-attention mechanism and cross-attention mechanism to generate the predicted word at the current time step. The generated word is used as the input for the next step, and this process is repeated until the end token is generated, finally generating the complete target sequence.
[0096] 3. Attention Pairing Interaction Module (APIM)
[0097] 3.1 Module Composition
[0098] The bilinear attention network (BAN) effectively extends the single attention network using bilinear attention maps for multimodal learning. It evaluates each pair of multimodal input channels, such as image regions and text words, to learn interaction representations. Compared with applying a single attention mechanism to multimodal data, BAN provides richer joint information while maintaining a similar computational cost.
[0099] The present invention designs an Attention Pairing Interaction Module (APIM), as Figure 2 shown. The present invention uses BAN to capture local pairwise interactions between text and image features. This module consists of two parts: (I) A bilinear interaction graph for capturing pairwise attention weights. (II) A variant of the multimodal residual network to effectively utilize the interaction graph.
[0100] 3.2 Bilinear Interaction Graph
[0101] 3.2.2 Bilinear Interaction Capturing Features
[0102] Given the text hidden state representation H T encoded by the BART Encoder C and the caption hidden state representation H N×Z : Let N and Z represent the number of features in the text and caption respectively, and R represent the set of extended real numbers. Then the bilinear interaction graph can obtain a single-head pairwise interaction matrix I ∈ R
[0103]
[0104] where and are learnable weight matrices for text and caption representations, q ∈ R K is a learnable weight vector, 1 ∈ R K is a fixed all-ones vector, and ° represents the Hadamard (element-wise) product. I represents the interaction strength between each word (or subphrase) in the text and the caption pair.
[0105] To better understand the bilinear interaction, the element I in formula (3) can also be expressed as:
[0106]
[0107] Among them, represents the i-th column of H T , and represents the j-th column of H C , respectively representing the i-th and j-th feature representations of the text and the subtitle. Therefore, the present invention can regard the bilinear interaction as first mapping the text representation and the subtitle feature to a common feature space with weight matrices U and V, and then learning the interaction on the Hadamard (element-wise) product and the vector q weights.
[0108] 3.2.3 Bilinear Pooling Sharing
[0109] Introduce a bilinear pooling layer to obtain a bilinear joint representation F' ∈ R K . Specifically, the k-th element of F' is calculated as:
[0110]
[0111] This formula calculates the interaction representation of the text feature H T and the image subtitle feature H C through bilinear pooling, generating the bilinear joint feature F' k of the k-th channel. Its core is to capture the fine-grained association between modalities by weighting the element-wise product of the text and image features with the attention weight I i,j . F' k is concatenated along the feature dimension to obtain F', and a linear transformation is performed on F' to obtain the attention map F.
[0112] Among them, U k and V k respectively represent the k-th columns in the weight matrices U and V. To reduce the model parameter quantity and prevent overfitting, the present invention shares these two weight matrices between the current interaction layer and the previous interaction layer. This sharing strategy helps to improve the performance and stability of the model.
[0113] 3.3 Variants of the Multimodal Residual Network
[0114] The present invention introduces a residual layer to integrate multiple bilinear attention maps:
[0115] F i+1 = BNA i (F i , H C ; I i ) · 1 T + F i (6)
[0116] F i is the input feature of the i-th layer in the residual module, and its initial value F0 is the original text feature H T , I i represents the i-th attention map. The residual layer adds this attention-enhanced feature to the original input feature through formula (6), BNA i represents the i-th bilinear attention module, 1 T represents the transposed all-1 vector, F i+1 represents the output feature map of the (i + 1)-th layer, that is, through I i calculates the weighted interaction between text and visual features to generate a joint feature. After passing the joint feature through batch normalization and linear transformation, it is added to the original feature F i to obtain the updated F i+1 .
[0117] The input text feature H T and the image caption feature H C , calculate the attention map I of text and image caption through formula (3), then generate the joint feature F′ by bilinear pooling according to formula (5), and use the residual network of formula (6) to iteratively update to obtain the enhanced text feature H TC .
[0118] Calculate the fine-grained interaction weight matrix I of the text feature H T and the facial description feature H D through the bilinear attention mechanism, map the interacted features to a local joint representation F′ using bilinear pooling, and iteratively update the text feature through a multi-layer residual network to gradually fuse the fine-grained visual semantics of facial expressions, and finally generate a locally fine-grained aligned joint representation H TD , the core of which is to accurately capture the emotional cues (such as facial expression details) related to text-related words through face-description-driven cross-modal alignment, thereby reducing the interference of irrelevant visual noise.
[0119] Finally, learn the parameter β t to selectively fuse the features H TC and H TD to obtain the final text representation
[0120]
[0121] W β is a weight matrix used to perform a linear transformation on the input features to help calculate β t ; W1 is a weight matrix used to perform a linear transformation on the input feature H TC ; W2 is a weight matrix used to perform a linear transformation on the input feature H TD ; b β is a bias vector used to add an offset to the result after linear transformation, enhancing the fitting ability of the model.
[0122] Under the action of the attention pairing interaction module, through the visual-text alignment process from coarse-grained to fine-grained between the image and the text, accurate semantic matching is gradually achieved. Through this multi-level alignment method, the H V and are concatenated, and finally an image-text feature containing rich information is generated
[0123] As shown in Table 1, compared with the complete model, after disabling the attention pairing interaction module (APIM), the accuracy rates on the two datasets decreased by approximately 1.9% and 0.7% respectively. This prominently shows that the APIM module can effectively coordinate different modalities, and this coarse-to-fine image-text alignment method can better achieve multi-modal feature fusion.
[0124] Table 1
[0125]
[0126] 4. Enhanced Dependency Graph Attention Network (RD-GAT)
[0127] 4.1 Extracting Emotional Features
[0128] SenticNet is an affective computing knowledge base designed to support sentiment analysis in natural language processing (NLP). It assigns sentiment polarities (such as positive, negative, neutral) and sentiment intensities to a large number of concepts by combining semantics and psychological methods, and considers the multi-dimensional representation of emotions (such as pleasantness, attention, etc.).
[0129] To effectively capture emotional information from multi-modal data, the present invention develops an enhanced dependency graph attention network (RD-GAT). First, the present invention integrates common sense knowledge related to emotional concepts into the multi-modal features Specifically, for a word t i in the input text, the present invention retrieves its emotional score from SenticNet and projects it into the same dimensional space as the multi-modal features to obtain the common sense emotional feature s i . Then, the present invention integrates this emotional feature s i into the output of the APIM:
[0130]
[0131] where, W s and b s is a learnable parameter, represents the feature containing sentiment knowledge; represents the feature representation of the i-th node of the image text feature . Through the above formula, the hidden state H rich in commonsense sentiment features can be obtained S , where is the feature representation of the i-th node in H S .
[0132] 4.2 Calculate the dependency matrix
[0133] 4.2.1 Distance importance calculation
[0134] The present invention designs a dependency matrix A to establish the connection between the image and the text. The sub-matrix A of this matrix TT describes the relationship between words, mainly considering two aspects: the importance of the distance between words, and the importance of the dependency type between them. To generate this matrix, the present invention uses the Stanza parser (a powerful tool that can perform word segmentation, part-of-speech tagging, syntactic analysis, etc. on multiple languages). It can parse the text into a dependency tree G, which is a special graph structure that shows the directed relationship between words in a sentence (such as who depends on whom). To simplify the modeling, the present invention changes the directed dependency relationship into an undirected edge to ensure that any two words in the sentence are connected. Next, the present invention calculates the minimum tree distance between words, that is, the number of edges on the shortest path connecting them. Finally, the present invention converts the dependency tree G into a fixed induced syntactic graph G dis , this graph is fully connected, and is represented by a symmetric adjacency matrix A dis , where A dis [i,j] reflects the distance between the i-th and j-th words. Simply put, the present invention parses the text, constructs a dependency tree, and finally generates a matrix to quantify the distance and relationship between words, thereby helping to establish the connection between the image and the text.
[0135] Considering that the weight of the edge is usually inversely proportional to the distance, the present invention innovatively introduces a distance importance calculation method based on the inverse hyperbolic tangent function to more intuitively map the influence of the distance on the weight. Finally, an enhanced distance dependency matrix A dis is generated.
[0136] A dis = tanh -1 (1 - A dis / max_tree_dis) (12)
[0137] max_tree_dis represents the maximum distance value that may exist between two nodes in a tree. The range of distance attenuation is controlled by adjusting the parameter max_tree_dis. The inverse hyperbolic tangent function can smoothly adjust the weights of distant nodes and provide high weights when the distance approaches zero, which is suitable for hierarchical structured data.
[0138] 4.2.2 Type Edge Importance Calculation
[0139] To calculate the importance weight of type edges, the present invention adopts a global attention mechanism. Specifically, the present invention first assigns numerical identifiers to types and converts G into an induced syntactic graph D represented by an adjacency matrix A dep representing dep . Then initialize the type feature matrix H ∈ R U×D , where U represents the total number of dependency types, D represents the dimension of the feature matrix set by oneself, and q ∈ R D×1 as the transposed query vector. Subsequently, use the softmax function to derive the attention weight matrix and apply the gather method (an operation to collect elements from a tensor or array according to specified indices) to extract the weights corresponding to A dep . A′ dep is the initially obtained initial adjacency matrix. Finally, obtain the type dependency matrix A dep :
[0140]
[0141] Finally, the final sub-matrix A TT is obtained by adding the two, and its expression is:
[0142] A TT = A dis + A dep (14)
[0143] 4.2.3 Obtaining Dependencies
[0144] The present invention adopts a multi-head attention mechanism to obtain the dependencies A VV between visual features, as well as the dependencies A VT and A TV between visual features and text features. In the multi-head attention mechanism, the attention weight is a key output. It determines the "attention" degree of each query vector to different key vectors. The attention weight can be understood as a probability distribution, indicating from which key vectors each query vector should obtain more information. First, the present invention obtains the attention weight A VV , which represents the degree of mutual attention between different positions in the same modality. The hidden feature h of the visual feature is obtained from H S sv , the final sub - matrix A VV is expressed as:
[0145]
[0146] h sv is the column element in H V , and W i Q is a weight matrix used to project the feature h sv into the query space and is a key parameter for calculating the query vector in the multi - head attention mechanism; Similarly, it is a weight matrix responsible for projecting the feature h sv into the key space for calculating the key vector and is also a parameter optimized during training. Their role is to help the model extract information from input features from different perspectives for attention calculation.
[0147] 4.2.4 Calculating the Dependency Matrix
[0148] To establish element - level (word - level / region - level) interaction between visual and text features, a cross - attention mechanism is adopted to obtain attention weights as the visual - text / text - visual dependency matrix. The hidden features h S of visual and text features are obtained from H sv and h st , and the final sub - matrix A VT and A TV are expressed as:
[0149]
[0150] The attention weight A VT represents the degree of attention of a certain position in the visual modality (such as image features) to a certain query vector in the text modality (such as the representation of a word), and the attention weight A TV represents the degree of attention of a query vector in the text modality (such as the representation of a word) to a certain position in the visual modality (such as image features). Combining A TT , A VV , A VT and A TV , the present invention obtains the final dependency matrix D.
[0151]
[0152] Let the dependency matrix D be the adjacency matrix A of the present invention.
[0153] 4.3 Principle of RD - GAT
[0154] The Graph Attention Network (GAT) is a model that combines graph neural networks and attention mechanisms, aiming to improve the quality of node representations by adaptively aggregating information from neighboring nodes. This method can learn the importance between different nodes in graph-structured data, thereby effectively capturing the features of the graph.
[0155] H S is a multi-modal feature rich in commonsense sentiment information. The Reinforced Dependent Graph Attention Network (RD-GAT) uses H S as the initial node representation in the graph and first calculates the attention coefficient between nodes i and j:
[0156]
[0157] where and are the hidden states of nodes i and j respectively, a represents learnable parameters, || represents the vector concatenation operation, and e ij represents the correlation between nodes i and j.
[0158]
[0159] where, M i,j is the weight matrix obtained by masking the attention coefficients using A i,j , and further operations yield α ij .
[0160]
[0161] where, K represents the number of attention heads, W k , W out are learnable weight matrices, and the outputs of multiple attention heads are concatenated and then passed through a linear transformation to obtain the final node feature representation The final output is obtained through the above formula This output fuses information from the image-text pair and is rich in abundant sentiment information.
[0162] 4.4 Ablation Experiments
[0163] RD-GAT Ablation Experiment: As shown in the following table, after deactivating RD-GAT, the overall performance of the model drops significantly. This is because the present invention incorporates common sense knowledge related to emotion-related concepts in RD-GAT, that is, using external emotion information to assist emotion prediction to improve the prediction accuracy. After removing the external emotion information, the F1 scores on the two datasets decreased by approximately 1.8% and 0.9% respectively, which highlights the role of this module in enhancing the emotion understanding ability of the model. Similarly, after deactivating the RD-GAT module, the F1 scores on the two datasets decreased by approximately 1.3% and 2.3% respectively, indicating that this module is crucial for maintaining the model performance. This highlights the effectiveness of RD-GAT in aggregating emotion information in multi-modal features.
[0164]
[0165] 5. Prediction and Optimization
[0166] The present invention proposes a multi-task prediction and optimization framework based on the BART decoder, the core of which is to uniformly model complex tasks as index generation tasks to achieve collaborative reasoning and end-to-end optimization of cross-modal information. Specifically, the BART decoder dynamically receives three parts of input during the generation process: (1) the output of the APIM multi-modal features obtained from coarse-to-fine image-text alignment; (2) the output of the RD-GAT multi-modal features aggregating multi-modal emotion information; (3) the previous decoder output Y <t , maintaining the causal consistency of the task sequence. The predicted token probability distribution is as follows:
[0167]
[0168] where λ 1 and λ 2 are hyperparameters controlling the contributions of the two modules, denotes the text part of, W represents the embedding representation of the input token, and C d corresponds to the emotion label (such as [positive, neutral, negative, <eos>)'s embedded representation.
[0169] The loss function is defined as follows:
[0170]
[0171] Where O = 2M + 2N + 2 is the length of Y, and X represents the multimodal input.
[0172] Joint ablation experiment of APIM and RD-GAT: As shown in the following table, compared with the complete model, after simultaneously deactivating the Attention Pairing Interaction Module (APIM) and the Reinforced Dependence Graph Attention Network (RD-GAT), the recall rates on the two datasets decreased by approximately 1.4% and 2.3% respectively. This prominently shows that the model of the present invention is effective in alleviating the problems of visual noise and emotional interference between different aspects.
[0173] The experimental results of AVTAF on the TWITTER-2015 and TWITTER-2017 datasets show that it is significantly superior to the existing methods in the multimodal aspect sentiment analysis task, especially in terms of visual noise and the problem of emotional interference between aspects.
[0174]
[0175] The above specific embodiments only describe the design principle of the present invention. The shapes and names of the components in this description can be different and are not limited. Therefore, those skilled in the art of the present invention can modify or equivalently replace the technical solutions recorded in the foregoing embodiments; and these modifications and replacements do not depart from the purpose and technical solutions of the present invention, and shall all fall within the protection scope of the present invention.< / eos> < / bos> < / eos> < / bos>
Claims
1. An emotion recognition method, characterized in that, It includes the following steps: Step 1: Perform multi-modal feature extraction on the input text-image pair: Generate an image description C and a facial expression description D based on the visual features V of the image; then convert the text features including text, image description, and facial expression description into text embeddings respectively; Step 2: Based on the bilinear attention network to capture the pairwise local interactions between text features and image features, including using the bilinear interaction graph to capture pairwise attention weights, and the variant multi-modal residual network to utilize the interaction graph for selective fusion to achieve hierarchical alignment of text and visual features, specifically including: Input the multimodal features obtained in step 1 into the BART encoder to obtain multimodal hidden states, which are respectively: the hidden state H corresponding to the text embedding T T , the hidden state H corresponding to the visual feature V V , the hidden state H corresponding to the image description C , the hidden state H corresponding to the facial expression description D , and H M = {H V , H T}; Based on the hidden state H T and the hidden state H C , a single-head pairwise interaction matrix of the two is obtained by using a bilinear interaction graph; a bilinear pooling layer is introduced to obtain the attention map F; and then a residual layer is introduced to integrate multiple bilinear attention maps to obtain the enhanced text feature H TC ; Based on the hidden state H T and the hidden state H D , a single-headed pairwise interaction matrix of the two is obtained by using a bilinear interaction graph; a bilinear pooling layer is introduced to obtain the attention map F; and then a residual layer is introduced to integrate multiple bilinear attention maps, gradually fusing the fine-grained visual semantics of facial expressions, and finally generating a locally fine-grained aligned joint representation H TD ; Finally, the learning parameter β t is used to selectively fuse the feature H TC and H TD to obtain the final text representation: Concatenate H V with to finally generate the image text feature Step 3: Weightedly fuse the sentiment scores of SenticNet with the multi-modal features obtained by the attention pairing interaction module, construct a multi-modal dependency graph, and aggregate sentiment information through the graph attention network, specifically including: For a word t in the input text i , retrieve its sentiment score from SenticNet and project it into the same dimensional space as the multi-modal features to obtain the commonsense sentiment feature s i ; then, integrate this sentiment feature s i into the output of APIM : Among them, W s and b s are learnable parameters, representing features containing sentiment knowledge; represents the feature representation of the i-th node of the image text feature ; the hidden state H rich in commonsense sentiment features is obtained through the above formula S , where is the feature representation of the i-th node in H S ; The global attention mechanism is adopted to obtain the dependency matrix A between the words of the input text TT The multi-head attention mechanism is used to obtain the dependency relationship A between visual features VV as well as the dependency relationship A between visual features and text features VT and A TV to obtain the final dependency matrix D: Adopt the graph attention network structure data H S to learn the importance between different nodes and obtain a richer emotional information output Step 4. The BART decoder receives and the previous decoder output Y <t , and predicts the token probability distribution as follows: Among them, λ 1 and λ 2 are two hyperparameters, denotes the text part of, W denotes the text embedding representation of Step 1; Decoder() denotes the processing of the BART decoder; softmax() denotes the processing of the softmax function on the data in the parentheses C d the embedding representation corresponding to the sentiment label; denotes the hidden state of the decoder at time step t; P(y t ) denotes the probability distribution of the predicted sentiment label generated at time step t.
2. The emotional recognition method according to claim 1, wherein In step 2, the single-headed pairwise interaction matrix I ∈ R N×Z : Among them, and are learnable weight matrices for text and caption representations, \(q\in\mathbb{R}^{ K}\) K is a learnable weight vector, \(1\in\mathbb{R}^{ K}\) K is a fixed all-ones vector, denotes the Hadamard product; \(I\) represents the interaction strength between each word or sub-phrase in the text and the image description pair.
3. The emotional recognition method according to claim 1, wherein, Obtain the enhanced text feature H TC The specific methods include: Introduce a bilinear pooling layer to obtain a bilinear joint representation F′. Specifically, the k-th element of F′ is calculated as: I i,j represent the elements in I; concatenate F' k along the feature dimension to obtain F', and perform a linear transformation on F' to obtain the attention map F; Among them, U k and V k respectively represent the k-th column in the text representation weight matrix U and the image description weight matrix V; Introduce a residual layer to integrate multiple bilinear attention graphs: F i+1 = BNA i (F i , H C ; I i )·1 T +F i F i is the input feature of the i-th layer in the residual module, and its initial value F0 is the original text feature H T , I i represents the i-th attention map, BNA i represents the i-th bilinear attention module, 1 T represents the transposed all-1 vector, F i+1 represents the output feature map of the (i + 1)-th layer, that is, by I i calculates the weighted interaction between the text and visual features to generate joint features. After passing the joint features through batch normalization and linear transformation, they are added to the original feature F i to obtain the updated F i+1 .
4. The emotional recognition method according to claim 1, wherein Learning parameter β t Expressed as: β t = sigmoid(W β [W1H TC ; W2H TD + b β ) W β is a weight matrix used to linearly transform the input features to assist in calculating β t ; W1 is a weight matrix used to perform a linear transformation on the input feature H TC ; W2 is a weight matrix used to perform a linear transformation on the input feature H TD ; b β is a bias vector used to add an offset to the result after the linear transformation, enhancing the model's fitting ability.
5. The emotional recognition method according to claim 1, characterized in that Method for obtaining dependency matrix A TT The method includes: using a Stanza parser to parse the text into a dependency tree G, and converting the dependency tree G into a fixed induced syntactic graph G dis , which is fully connected, and using a symmetric adjacency matrix A dis to represent the distance between the i-th and j-th words; A dis = tanh -1 (1 - A dis / max_tree_dis) max_tree_dis represents the maximum possible distance value between two nodes in the tree; control the range of distance attenuation by adjusting the parameter max_tree_dis; Assign numeric identifiers to the types and transform G into a matrix consisting of the adjacency matrix A. dep The induced syntactic graph D dep ; Then initialize the type feature matrix H∈R U×D , where U represents the total number of dependency types, D represents the dimension of the feature matrix set by oneself, and q∈R D×1 as the transposed query vector; then use the softmax function to derive the attention weight matrix and apply the gather method; extract dep The corresponding weight; A′ dep is the initial adjacency matrix obtained initially, and finally, the type dependency matrix A is obtained dep : Finally, the final sub-matrix A is obtained by adding the two together TT , and its expression is: A TT = A dis + A dep 。 6. The emotional recognition method according to claim 1, characterized in that Method for obtaining dependence matrix A VV includes: using a multi-head attention mechanism to obtain the hidden feature h of visual features from H S , and the final sub-matrix A sv is expressed as: VV h sv is the column element in H V and is a weight matrix used to project the feature h sv into the query space and is a key parameter for calculating the query vector in the multi-head attention mechanism; similarly, it is also a weight matrix responsible for projecting the feature h sv into the key space for calculating the key vector and is also a parameter optimized during training. Their role is to help the model extract information from input features from different angles for attention calculation.
7. The emotional recognition method according to claim 1, characterized in that, Obtain the dependency matrix A VT and A TV The method includes: obtaining the hidden features h S of visual and text features from H sv and h st The final sub-matrices A VT and A TV are expressed as:
8. The emotional recognition method according to claim 1, characterized in that Adopt the graph attention network structure data H S to learn the importance between different nodes and obtain a richer emotional information output The method includes: Taking H S as the initial node representation in the graph, first calculate the attention coefficient between nodes i and j: where and are the hidden states of node i and node j respectively, a represents a learnable parameter, || represents the vector concatenation operation, and e ij represents the correlation between node i and node j; Among them, M i,j is the weight matrix after masking the attention coefficients using A i,j , and further operations are performed to obtain α ij ; Among them, K represents the number of attention heads, and W k , W out are learnable weight matrices. After concatenating the outputs of multiple attention heads, the final node feature representation is obtained through a linear transformation 9. The emotional recognition method according to claim 1, characterized in that It also includes optimizing the model using the cross-entropy loss function, and generating a joint prediction result of aspect terms and their sentiment polarities by weighting the contributions of multiple modules with hyperparameters.
10. A system for implementing the emotion recognition method according to any one of claims 1-9, characterized in that, Implemented using the AVTAF model, which includes the following modules: a feature extraction module, an attention pairing interaction module APIM, an enhanced dependency graph attention network RD-GAT, and a prediction and optimization module; The feature extraction module is used to implement the method of Step 1; The attention pairing interaction module APIM is used to implement the method of Step 2; The enhanced dependency graph attention network RD-GAT is used to implement the method of Step 3; The prediction and optimization module is used to implement the method of Step 4.