Multi-modal Sentiment Analysis Method and System Based on Attention and Graph-Enhanced Text

Through adaptive cross-modal interaction module and hierarchical multimodal graph fusion network, the local and global dependencies between modes are explicitly modeled, and the problem of context aberration between modes in multimodal sentiment analysis is solved, and more accurate emotional prediction is achieved.

CN119622559BActive Publication Date: 2025-07-11SHANDONG JIAOTONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510167640.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-07-11
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The existing multimodal emotion analysis methods lack the ability to capture context alignment and complex dependency between modes, resulting in inaccurate emotion prediction.

Method used

A multimodal sentiment analysis model (MAG-TE) based on attention and graph enhancement text is adopted, and the local and global dependencies between modes are explicitly modeled through adaptive cross-modal interaction modules and hierarchical multimodal graph fusion networks, and a multi-level interaction of multimodal emotional features is captured using a jump-connection graph convolution network.

Benefits of technology

It significantly improves the accuracy and comprehensiveness of multimodal emotional prediction, effectively solves the problem of context aberration between modals, and enhances the understanding and prediction ability of emotional expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119622559B_ABST
    Figure CN119622559B_ABST
Patent Text Reader

Abstract

The present invention proposes a multi-modal sentiment analysis method and system based on attention and graph-enhanced text, belonging to the technical field of multi-modal sentiment analysis. The method includes: obtaining text features, image features, and speech features in video data and performing preprocessing; using an adaptive cross-modal interaction module to calculate attention weights between text features and image features and speech features to obtain enhanced text features; inputting the enhanced text features into a hierarchical multi-modal graph fusion network, and using a self-attention mechanism to construct an adjacency matrix; inputting the adjacency matrix and the enhanced text features into a skip connection graph convolutional network to obtain a final feature matrix; combining the feature matrix and the adjacency matrix, and using an encoder and a classifier to obtain a prediction result of sentiment analysis. It solves the problem of insufficient context alignment between different modalities and performs sentiment polarity prediction more comprehensively and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimodal sentiment analysis, and particularly relates to a multimodal sentiment analysis method and system based on attention and graph-enhanced text. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Humans understand and experience the world through the perception of multimodal signals, and emotional expressions also show multi-level and three-dimensional characteristics. In addition to language content, humans also convey emotions and attitudes through intonation, volume, facial expressions, gestures, etc. In recent years, the rapid rise of short video platforms has led to an explosive growth in the number of short videos based on human activities, accompanied by the generation of a large amount of multimodal data. This phenomenon highlights the importance of research in the field of multimodal sentiment analysis and has attracted extensive attention to the recognition of emotional states using information from multiple modalities.

[0004] Most existing multimodal sentiment analysis methods are mainly dominated by text information, and supplemented by visual and auditory information to make up for the deficiencies of text information, so as to improve the comprehensiveness and accuracy of emotion prediction. However, the existing methods still face the following problems: First, there are obvious deficiencies in the context alignment between different modalities. Due to the locality and inconsistency in semantic expression among modalities such as text, audio, and vision, for example, the intonation changes in audio or facial expressions in vision may only be related to part of the semantics in the text. Most existing methods adopt simple fusion strategies such as weighting or splicing, such as LMF, TFN, etc., but fail to effectively capture the fine-grained associations between modalities, resulting in the problem of context misalignment. Second, the existing methods have insufficient ability to capture the complex dependence relationships between different modalities. Many methods fail to explicitly model the global and local dependencies between modalities, resulting in the neglect of potential semantic interactions between modalities during the fusion process of modal information, thus affecting the accuracy of emotion prediction. Summary of the Invention

[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a multimodal sentiment analysis model method and system based on attention and graph-enhanced text. By using a multimodal sentiment analysis model based on attention and graph-enhanced text (MAG-TE), through an adaptive cross-modal interaction module, a discrete sequence and a cross-modal attention mechanism are introduced to dynamically align and integrate image and speech features. A hierarchical multimodal graph fusion network is adopted to explicitly model the local and global dependence relationships between modalities through a graph structure, and a skip-connected graph convolutional network is used to capture the multi-level interactions of multimodal sentiment features, significantly improving the performance of multimodal emotion prediction.

[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0007] The first aspect of the present invention provides a multi-modal sentiment analysis method based on attention and graph-enhanced text;

[0008] The multi-modal sentiment analysis method based on attention and graph-enhanced text includes:

[0009] Obtain text features, image features, and speech features in video data and perform preprocessing;

[0010] Input the preprocessed text features, image features, and speech features into a trained multi-modal sentiment analysis model to obtain the prediction result of sentiment analysis;

[0011] Among them, the trained multi-modal sentiment analysis model uses an adaptive cross-modal interaction module to calculate the attention weights between text features and image features and speech features, and obtain enhanced text features;

[0012] Input the enhanced text features into a hierarchical multi-modal graph fusion network, and use the self-attention mechanism to construct an adjacency matrix; input the adjacency matrix and the enhanced text features into a skip connection graph convolutional network to obtain the final feature matrix;

[0013] Combine the feature matrix and the adjacency matrix, and use an encoder and a classifier to obtain the prediction result of sentiment analysis.

[0014] As a further technical solution, the preprocessing process includes:

[0015] Process the text features, use a pre-trained language model to tokenize and embed the text, and generate word embeddings and position embeddings;

[0016] Extract image information of facial expressions, action units, and postures from video frames; perform a dimensionality reduction operation on the image information to obtain image features;

[0017] Extract speech emotion information from speech frames, and perform a dimensionality reduction operation on the speech emotion information to obtain speech features.

[0018] As a further technical solution, the process of using an adaptive cross-modal interaction module to calculate the attention weights between text features and image features and speech features and obtain enhanced text features is as follows:

[0019] Use an adaptive cross-modal interaction module to map the image features and speech features into a unified index sequence;

[0020] Process the index vector using a cross-modal attention mechanism, calculate the attention weights between text features and image and speech features, and obtain enhanced text features.

[0021] As a further technical solution, the process of processing the index sequence using a cross-modal attention mechanism, calculating the attention weights between text features and image and speech features, and obtaining enhanced text features is as follows:

[0022] Map the discrete index sequence to a continuous high-dimensional vector representation through an embedding layer;

[0023] Use the text features as query vectors, and the image and speech features as key vectors and value vectors to obtain the attention weight matrix between text features and image and speech features;

[0024] Based on the attention weight matrix, extract enhanced non-verbal information from the image and speech features to generate non-verbal embedding information corresponding to the text features;

[0025] Fuse the non-verbal embedding information and the text features to obtain enhanced text features.

[0026] As a further technical solution, the adjacency matrix constructed using the self-attention mechanism is:

[0027]

[0028] In the formula, is the adjacency matrix representing the correlation between nodes; is the matrix transpose of, representing the similarity between word pairs in text features; is a scaling factor used to stabilize the calculation and prevent the dot product from being too large; is the operation applied to each row of the matrix, representing the relationship strength between each pair of words.

[0029] As a further technical solution, the process of inputting the adjacency matrix and the enhanced text features into a skip connection graph convolutional network to obtain the final feature matrix is as follows:

[0030] Use the adjacency matrix and the enhanced text features as the initial inputs of the skip connection graph convolutional network, where the enhanced text features serve as the initial feature matrix;

[0031] Use the adjacency matrix to perform weighted aggregation on the initial feature matrix. The aggregated feature matrix represents the weighted average features of each node and its neighbors;

[0032] The input feature matrix is added to the transformed feature matrix using skip connections, and the ReLU activation function is used to obtain the final feature matrix.

[0033] As a further technical solution, the process of combining the feature matrix and the adjacency matrix and using the encoder and classifier to obtain the prediction result of sentiment analysis is as follows:

[0034] Use The function concatenates the final feature matrix and the adjacency matrix to obtain the fused features;

[0035] The fused features are added to the initial text embedding to obtain the text enhanced feature representation;

[0036] The text enhanced feature representation is input into the trained Transformer encoder and classifier to obtain the prediction result of sentiment analysis.

[0037] The second aspect of the present invention provides a multi-modal sentiment analysis system based on attention and graph-enhanced text.

[0038] A multi-modal sentiment analysis system based on attention and graph-enhanced text includes:

[0039] A multi-modal feature acquisition module, configured to: acquire text features, image features, and speech features in video data and perform preprocessing;

[0040] A text feature enhancement module, configured to: calculate the attention weights between the text features and the image features and speech features using an adaptive cross-modal interaction module to obtain the enhanced text features;

[0041] A final feature matrix acquisition module, configured to: input the enhanced text features into a hierarchical multi-modal graph fusion network, and use the self-attention mechanism to construct an adjacency matrix; input the adjacency matrix and the enhanced text features into a skip connection graph convolutional network to obtain the final feature matrix;

[0042] A sentiment analysis prediction module, configured to: combine the final feature matrix and the adjacency matrix, and use the encoder and classifier to obtain the prediction result of sentiment analysis.

[0043] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the multi-modal sentiment analysis method based on attention and graph-enhanced text as described in the first aspect of the present invention.

[0044] A fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the multi-modal sentiment analysis method based on attention and graph-enhanced text as described in the first aspect of the present invention.

[0045] The above one or more technical solutions have the following beneficial effects:

[0046] The present invention uses an adaptive cross-modal interaction module to convert the features of visual and audio modalities into discrete index sequence representations, alleviating the formal differences between different modalities. Subsequently, based on the cross-modal attention mechanism, it can dynamically extract sentiment signals related to text semantics, thereby enhancing the understanding of text sentiment expression. It can effectively ignore irrelevant interference information, such as meaningless smiles in vision or background noise in audio.

[0047] The hierarchical multi-modal graph fusion network explicitly models the dependencies between modalities by constructing a multi-modal adjacency matrix. It captures the fine-grained dependencies between modalities through local convolution operations and ensures the efficient propagation of cross-layer information through the skip connection mechanism, avoiding information loss in traditional deep graph convolution networks. Finally, the hierarchical multi-modal graph fusion network generates a unified multi-modal sentiment representation by hierarchically fusing local and global features, thereby achieving more comprehensive and accurate sentiment polarity prediction.

[0048] Advantages of additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0050] Figure 1 It is a flowchart of the method for the first embodiment.

[0051] Figure 2 It is a system structure diagram for the second embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0053] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0054] In the case of no conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0055] The present invention maps the original features of vision and audition into a unified index sequence through an adaptive cross-modal interaction module to reduce the heterogeneity differences between modalities, and then uses a cross-modal attention mechanism to calculate the attention weights between the vision-audition embedding and the text embedding; uses a hierarchical multi-modal graph fusion network to construct an adjacency matrix between multi-modal features, and uses a skip-connected graph convolutional neural network to capture the complex dependencies between different modalities to achieve deep fusion of multi-modal features. Finally, the fused multi-modal features are combined with the enhanced text features and input into a pre-trained Transformer encoder and classifier for sentiment polarity prediction.

[0056] Embodiment 1

[0057] This embodiment discloses a multi-modal sentiment analysis method based on attention and graph-enhanced text;

[0058] As Figure 1 shown, the multi-modal sentiment analysis method based on attention and graph-enhanced text includes:

[0059] Step S1, obtaining text features, image features and speech features in video data and performing preprocessing;

[0060] Multi-modal sentiment analysis uses multi-modal signals to detect the sentiment of video clips. For a video clip X, it consists of three parts: a text sequence (t), an acoustic sequence (a), and a visual sequence (v), expressed as , n ∈ {t, a, v}. Among them, represents the sequence length of modality n, respectively represent the i-th word, audio frame, and visual frame.

[0061] For the text sequence , we use a tokenizer of a pre-trained language model (such as SentiLARE) to extract sentiment-related features such as sentiment polarity and part of speech to enhance the sentiment expression ability of text representation. By segmenting the text into the corresponding word index sequence . Among them, is the length of the text sequence.

[0062] Next, word embeddings, position embeddings, and sentiment-related feature embeddings (such as sentiment polarity and part of speech) are generated through a pre-trained language model:

[0063]

[0064] In the formula, represents the word embedding, Represents positional embedding, represents sentiment polarity embedding.

[0065] For acoustic sequences and visual sequences, Facet and COVAREP are respectively used to extract raw features. The raw acoustic and visual feature sequences are denoted as and , where are respectively the number of frames of the acoustic and visual sequences, and are respectively the acoustic and visual feature dimensions. Specifically, for visual sequences, Facet is used to extract image information such as facial expressions, action units, and postures from video frames. The image information is dimensionally reduced to obtain image features, facilitating subsequent modeling and fusion.

[0066] For speech features, the COVAREP tool is used to extract relevant sentiment information from audio signals, such as speech sentiment information like Mel-frequency cepstral coefficients (MFCCs), pitch, and speech segment segmentation. The speech sentiment information is dimensionally reduced to obtain speech features.

[0067] Through the extraction and dimensionality reduction of text, image, and speech features, the generated multi-modal features lay the foundation for subsequent modal alignment and fusion.

[0068] Step S2, the multi-modal sentiment analysis model based on attention and graph-enhanced text uses an adaptive cross-modal interaction module to calculate the attention weights between text features and image features and speech features, obtaining enhanced text features;

[0069] In multi-modal sentiment analysis tasks, information from different modalities always has problems of misalignment in time and content, which leads to biases in the model's understanding of the overall sentiment. Therefore, to solve the problem of misalignment between asynchronous modalities, the visual and auditory feature representations are unified into index sequences through an adaptive cross-modal interaction module (ACIM), and the information interaction between text features and image features and speech features is explicitly modeled through a cross-modal attention mechanism, effectively capturing the complementary information between text and non-text modalities and enhancing text representation.

[0070] Specifically, text features are usually represented as index sequences obtained from a vocabulary, while the representations of image features and speech features are real vector sequences.

[0071] Step S21, first obtain the image frames and speech frames in the video data and construct a feature set where represents image features or auditory features, represents the number of frames, represents the feature dimension per frame, Represent image features; Represent speech features; Use k-means clustering to divide image frames and speech frames into clusters, as follows:

[0072]

[0073] where represents the cluster center of modality n. Through these cluster centers, visual vocabulary and auditory vocabulary are constructed, and the image features and speech features are lexicalized through the cluster centers, so that the index sequence matches the discrete lexical form of the text sequence, providing a consistent representation for subsequent modality fusion and attention mechanism calculation.

[0074] Step S22, convert the feature sequence into an index. Specifically, given a feature sequence , it needs to be converted into the corresponding index sequence . For the of the i-th frame, its index is calculated as follows:

[0075]

[0076] In the formula, is the index of the cluster center corresponding to the i-th frame of modality n; is to find the index j of the cluster center nearest to ; is to calculate the Euclidean distance between the feature and the cluster center; finally, the index sequence is obtained as the representation of modality n.

[0077] Through the above means, the modality difference can be better reduced, and the video and speech features can be mapped into discrete indexes, which can effectively narrow the difference between modalities and improve the fusion efficiency. This mapping enables non-verbal features to better fuse with text features and improves the effect of multi-modal sentiment analysis.

[0078] Step S23, input the index vector output in the above steps into the embedding layer of the model:

[0079]

[0080] where is the output of the embedding layer, is the embedding dimension, is the embedding layer function. The main role of the embedding layer is to map discrete indexes into continuous high-dimensional vector representations. Using the high-dimensional vector representation as the input of the model can provide a richer and more learnable feature representation.

[0081] Step S24: Use the text features as the query vector (Q), and the image features and speech features as the key vector (K) and value vector (V) to obtain the attention weight matrix between the text features and the image features and speech features, as shown in the following formula:

[0082]

[0083]

[0084]

[0085] where is a learnable parameter matrix, are the dimensions of the query, key, and value vectors respectively.

[0086] Step S25: By obtaining the attention weights and correlations between the text features and the image features and speech features, generate the attention weight matrix of each word in the text features for each video frame or speech frame:

[0087]

[0088] where is the attention weight matrix, representing the attention distributions between the text features and the image features and speech features respectively.

[0089] Step S26: Based on the obtained attention weight matrix, extract the enhanced non-verbal information from the image features and speech features to generate the non-verbal embedding information corresponding to the text:

[0090]

[0091] These enhanced non-verbal embeddings can be regarded as the emotional information of the images and speech selected by the text features. After obtaining the enhanced embeddings from the image features and speech features, they are fused through a gating mechanism (Gate) to combine the text, image, and speech information:

[0092]

[0093] where ';' represents the concatenation operation, and Gate(;) is a gating mechanism composed of fully connected layers, is the enhanced information extracted from the image features and aligned with the text features; is the enhanced information extracted from the speech features and aligned with the text features; used to fuse the information of the two modalities.

[0094] Step S3: Input the enhanced text features into the hierarchical multi-modal graph fusion network, and construct the adjacency matrix using the self-attention mechanism; input the adjacency matrix and the enhanced text features into the skip-connected graph convolutional network to obtain the final feature matrix;

[0095] Step S31: Take the text features enhanced by ACIM as the input of the hierarchical multi-modal graph fusion network (HMGFN), and use the self-attention mechanism to construct the adjacency matrix A:

[0096]

[0097] where is the adjacency matrix representing the correlation between nodes; is the matrix transpose of, and the dot product calculates the similarity between word pairs in the text features, that is, the similarity between each word pair. is the scaling factor used to stabilize the calculation and prevent the dot product from being too large. is to apply the operation to each row of the matrix to ensure that the sum of the elements in each row is 1, that is, a probability distribution is obtained, representing the relationship strength between each pair of words.

[0098] Step S32: Use the skip-connected graph convolutional network (SGCN) for multi-level information aggregation. Each layer of the graph convolutional network combines the node features with the adjacency matrix to capture the dependencies between modalities. Take the adjacency matrix A and the enhanced text features obtained by ACIM as the initial input of SGCN, and let , then for the th layer of SGCN:

[0099]

[0100] where is the feature matrix of the i-th layer; is the learnable weight matrix of the th layer; ReLU is the activation function; the skip connection term is used to retain the original feature information, is the length of the text sequence; d is the feature dimension of each text word embedding.

[0101] The key step of SGCN is to aggregate the neighbor information of each node. The adjacency matrix A is used to perform weighted aggregation on the input feature matrix , and the aggregated feature matrix represents the weighted average feature of each node and its neighbors. By adding the input feature matrix to the transformed feature matrix through the skip connection, the input features can be retained and information loss can be avoided.

[0102] Step S33: Introduce a non - linear transformation by applying the ReLU activation function to obtain the final feature matrix S, ensuring that the model can learn more complex non - linear relationships, thereby increasing the model's representation ability.

[0103] Step S4: Combine the final feature matrix and the adjacency matrix, and use the encoder and classifier to obtain the prediction result of sentiment analysis.

[0104] The final feature matrix S can capture the complex dependency relationships between different modalities and has structured global information; while the adjacency matrix A obtained through ACIM aligns the information in terms of time and content of different modalities and captures the local correlations between text and non - text modalities.

[0105] Step S41: Use the function to concatenate the final feature matrix and the adjacency matrix to obtain the fused features;

[0106]

[0107] where concat(·) is the concatenation function.

[0108] Then, use the Linear layer for linear transformation to transform the dimension of H:

[0109]

[0110] where represents a linear transformation operation.

[0111] Step S42: Add the fused features to the initial text embedding T to obtain the text - enhanced feature representation;

[0112]

[0113] In the formula, is the text - enhanced feature representation. By adding the text embeddings, it is ensured that the comprehensive information of all modalities can achieve the best performance in downstream tasks.

[0114] Step S43: Input the text - enhanced feature representation into the trained Transformer encoder and classifier to obtain the prediction result of sentiment analysis.

[0115] Furthermore, in this embodiment, experiments were also conducted on two public datasets CMU - MOSI and CMU - MOSEI with the method proposed in the present invention.

[0116] The CMU-MOSI dataset is a publicly available dataset commonly used in the fields of multimodal sentiment analysis and emotion recognition. It contains information in three modalities: text, vision, and audio. It consists of 93 personal videos from YouTube, which are subjectively annotated to form a total of 2,199 annotated segments, each of which is a part of a video. Each segment has an emotional intensity annotation ranging from [-3, 3]. Among them, -3 represents strong negative emotion, 0 represents neutral, and 3 represents strong positive emotion. We divide the dataset into three parts: a training set (1,284 segments), a validation set (229 segments), and a test set (686 segments).

[0117] The CMU-MOSEI dataset is an extended version of MOSI, making it applicable to larger-scale multimodal sentiment analysis tasks. It also comes from YouTube videos, covering various topics such as movie reviews and speeches. It contains 3,229 videos, which are subjectively annotated and decomposed into 23,454 annotated segments according to emotional intensity. Each annotated segment in it not only has an emotional intensity label in the range of [-3, 3] but also adds binary classification labels for 6 basic emotions (happiness, sadness, anger, fear, disgust, and surprise).

[0118] To evaluate the performance of the model proposed in the present invention, two evaluation tasks of binary classification and regression are constructed. The binary classification accuracy (Acc-2) and weighted F1-score (F1-score) are used to evaluate the binary classification task. And there are two classification methods: negative / non-negative and negative / positive. For the regression task, the model performance is evaluated by the mean absolute error (MAE) and Pearson correlation (Corr), as shown in Table 1.

[0119] Table 1 Performance comparison and evaluation metrics of MAG-TE with 19 baselines on the CMU-MOSI and CMU-MOSEI datasets. An upward arrow indicates that the higher the metric, the better, and a downward arrow indicates the opposite.

[0120]

[0121] The experimental results show that the performance of the MAG-TE model on the four benchmark datasets is better than the compared models, which verifies the effectiveness of the proposed model in the MSA task.

[0122] As can be seen from the table, Transformer-based models (such as MAG, Self-MM, MISA, CENet, etc.) generally outperform traditional multimodal methods (such as TFN, LMF, MFN, MulT) in various evaluation metrics on the two datasets. This trend indicates that the Transformer architecture can better capture the interaction relationships of multimodal information and fully exploit the dependencies between modalities through its global attention mechanism. Among all Transformer-based models, the MAG-TE model proposed in this application shows significant advantages in all metrics, which further proves the effectiveness of the multimodal enhancement module (MTE) in aligning and fusing multimodal information.

[0123] For the current mainstream multimodal alignment-based methods (such as MAG, Self-MM, CENet), the present invention leads in the two core metrics of MAE and Corr. On the CMU-MOSI dataset, the MAE of MAG-TE reaches 0.581, which is 2.5% and 18.5% higher than the best-performing CENet (0.596) and Self-MM (0.713), respectively. At the same time, on the CMU-MOSEI dataset, the MAE of MAG-TE is further reduced to 0.516, significantly outperforming other models. This shows that MAG-TE effectively solves the problems of modality asynchrony and context misalignment through the adaptive cross-modal interaction module (ACIM), greatly improving the prediction accuracy. In addition, in terms of the Corr metric reflecting emotional correlation, MAG-TE reaches the highest value of 0.868 on the CMU-MOSI dataset, and the Corr on the CMU-MOSEI dataset also reaches 0.850, comprehensively verifying the significant ability of MAG-TE in modeling multimodal correlation.

[0124] In the comparison of multimodal fusion methods, the MAG-TE method also shows better performance in classification metrics such as Acc-2 and F1. Specifically, on the CMU-MOSI dataset, the Acc-2 of MAG-TE is 87.76% and the F1 is 87.69%, respectively exceeding CENet (Acc-2 is 86.74% and F1 is 86.69%) and MAG (Acc-2 is 84.20% and F1 is 84.14%). Similarly, on the CMU-MOSEI dataset, the Acc-2 and F1 metrics of MAG-TE also achieve comprehensive superiority. This shows that MAG-TE effectively improves the accuracy and robustness of emotion classification by explicitly modeling the global and local dependencies between modalities through the hierarchical multimodal graph fusion network (HMGFN).

[0125] Example 2

[0126] This embodiment discloses a multi-modal sentiment analysis system based on attention and graph-enhanced text;

[0127] As Figure 2 shown, the multi-modal sentiment analysis system based on attention and graph-enhanced text includes:

[0128] A multi-modal feature acquisition module, configured to: acquire text features, image features, and speech features in video data and perform preprocessing;

[0129] A text feature enhancement module, configured to: calculate attention weights between text features and image features and speech features using an adaptive cross-modal interaction module to obtain enhanced text features;

[0130] A final feature matrix acquisition module, configured to: input the enhanced text features into a hierarchical multi-modal graph fusion network, and use a self-attention mechanism to construct an adjacency matrix; input the adjacency matrix and the enhanced text features into a skip connection graph convolutional network to obtain a final feature matrix;

[0131] A sentiment analysis prediction module, configured to: combine the final feature matrix and the adjacency matrix, and use an encoder and a classifier to obtain a prediction result of sentiment analysis.

[0132] Embodiment Three

[0133] The purpose of this embodiment is to provide a computer-readable storage medium.

[0134] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the multi-modal sentiment analysis method based on attention and graph-enhanced text as described in Embodiment 1.

[0135] Embodiment Four

[0136] The purpose of this embodiment is to provide an electronic device.

[0137] An electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in the multi-modal sentiment analysis method based on attention and graph-enhanced text as described in Embodiment 1.

[0138] The steps involved in the devices in Embodiments Two, Three, and Four above correspond to those in Method Embodiment One, and the specific implementation manners can refer to the relevant description part of Embodiment One. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0139] Those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0140] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. A multimodal sentiment analysis method based on attention and graph-enhanced text, characterized in that, Including: Obtain the text features, image features, and speech features in the video data and perform preprocessing; where, for video segment X, it consists of a text sequence t, an acoustic sequence a, and a visual sequence v; Input the preprocessed text features, image features, and speech features into the trained multi-modal sentiment analysis model to obtain the prediction results of sentiment analysis; Among them, the trained multi-modal sentiment analysis model uses an adaptive cross-modal interaction module to calculate the attention weights between the text features and the image features and speech features, and obtain the enhanced text features, including: Use the adaptive cross-modal interaction module to map the image features and speech features into a unified index sequence; where, the process of using the adaptive cross-modal interaction module to map the image features and speech features into a unified index sequence includes: Obtain the image frames and speech frames in the video data and construct a feature set; Use the k-means method to cluster the image frames and speech frames, construct visual vocabulary and auditory vocabulary through the cluster centers, and lexicalize the image features and speech features through the cluster centers, so that the index sequence matches the discrete lexical form of the text sequence; Convert the feature sequence into an index. Specifically, given a feature sequence , it is necessary to convert it into the corresponding index sequence ; for the of the i-th frame, its index is calculated by the following method: In the formula, is the i-th frame of mode n corresponding clustering center index; is to find the distance to the nearest clustering center index j; is to calculate the Euclidean distance between the feature and the clustering center; finally, the index sequence is obtained as the representation of mode n; Use the cross-modal attention mechanism to process the index vectors, calculate the attention weights between the text features and the image features and speech features, and obtain the enhanced text features; Among them, using the cross-modal attention mechanism to process the index vectors, calculate the attention weights between the text features and the image features and speech features, and obtain the enhanced text features, including: Map the discrete index sequence into a continuous high-dimensional vector representation through an embedding layer; wherein is the output of the embedding layer, is the embedding dimension; is the embedding layer function; represents the sequence length of modality n; Use the text features as query vectors, and the image features and speech features as key vectors and value vectors; Among them, is a learnable parameter matrix, d is the feature dimension of each text word embedding, are the dimensions of the key and value vectors respectively; is the length of the text sequence; Generate an attention weight matrix of each word in the text features for each video frame or speech frame by obtaining the attention weights and correlations between the text features and the image features and speech features; Among them is the attention weight matrix, representing the attention distributions between text features and image features, and speech features respectively; Based on the attention weight matrix, extract enhanced non-verbal information from the image features and speech features to generate non-verbal embedding information corresponding to the text features; Among them, is non-verbal embedding information; Fuse the non-verbal embedding information and the text features to obtain the enhanced text features; the enhanced text features are: In the formula, is the enhanced text feature; ';' represents the concatenation operation, and Gate(;) is a gating mechanism composed of fully connected layers, is the enhanced information extracted from the image feature and aligned with the text feature; is the enhanced information extracted from the speech feature and aligned with the text feature, which is used to fuse the information of the two modalities; Input the enhanced text features into a hierarchical multi-modal graph fusion network, and use the self-attention mechanism to construct an adjacency matrix; Among them, is the adjacency matrix representing the association between nodes; is the matrix transpose of, and the dot product calculates the similarity between word pairs in the text features, that is, the similarity between each word pair; is the scaling factor, used to stabilize the calculation and prevent the dot product from being too large; is to apply operation to each row of the matrix to ensure that the sum of the elements in each row is 1, that is, to obtain a probability distribution representing the relationship strength between each pair of words; Input the adjacency matrix and the enhanced text features into a skip-connection graph convolutional network to obtain the final feature matrix; where, the process of inputting the adjacency matrix and the enhanced text features into the skip-connection graph convolutional network to obtain the final feature matrix is: Use the adjacency matrix and the enhanced text features as the initial input of the skip-connection graph convolutional network, where the enhanced text features are used as the initial feature matrix; Use the adjacency matrix to perform weighted aggregation on the initial feature matrix, and the aggregated feature matrix represents the weighted average features of each node and its neighbors; The input feature matrix is added to the transformed feature matrix using skip connections, and the ReLU activation function is used to obtain the final feature matrix; specifically, a skip connection graph convolutional network is used for multi-level information aggregation, and the graph convolutional network of each layer combines node features with the adjacency matrix; the adjacency matrix and the enhanced text features are used as the initial inputs of the skip connection graph convolutional network, and let , then for the -th layer of the skip connection graph convolutional network: Where A is the adjacency matrix; is the feature matrix of the i-th layer; is the learnable weight matrix of the -th layer; ReLU is the activation function; the skip connection term is used to retain the original feature information; Combine the feature matrix and the adjacency matrix, and use an encoder and a classifier to obtain the prediction results of sentiment analysis. Specifically, the process of combining the feature matrix and the adjacency matrix and using an encoder and a classifier to obtain the prediction results of sentiment analysis is: Adopt The function is used to splice the final feature matrix and the adjacency matrix to obtain the fused features; Add the fused features to the initial text embedding to obtain a text-enhanced feature representation; Input the text-enhanced feature representation into the trained Transformer encoder and classifier to obtain the prediction result of sentiment analysis.

2. The multimodal sentiment analysis method based on attention and graph-enhanced text according to claim 1, wherein The preprocessing process includes: Process the text features, tokenize and embed the text using a pre-trained language model to generate word embeddings and position embeddings; Extract the image information of facial expressions, action units, and postures from the video frames; perform a dimensionality reduction operation on the image information to obtain image features; Extract the speech emotion information from the speech frames, and perform a dimensionality reduction operation on the speech emotion information to obtain speech features.

3. A multimodal sentiment analysis system based on attention and graph-enhanced text, characterized in that: It includes: A multimodal feature acquisition module configured to: acquire and preprocess text features, image features, and speech features in video data; where, for a video segment X, it consists of a text sequence t, an acoustic sequence a, and a visual sequence v; A text feature enhancement module configured to: calculate the attention weights between text features and image features and speech features using an adaptive cross-modal interaction module to obtain enhanced text features, including: Use the adaptive cross-modal interaction module to map the image features and speech features into a unified index sequence; where the process of using the adaptive cross-modal interaction module to map the image features and speech features into a unified index sequence includes: Obtain the image frames and speech frames in the video data and construct a feature set; Use the k-means method to cluster the image frames and speech frames, construct visual vocabulary and auditory vocabulary through the cluster centers, and lexicalize the image features and speech features through the cluster centers so that the index sequence matches the discrete vocabulary form of the text sequence; Convert the feature sequence into an index. Specifically, given a feature sequence , it is necessary to convert it into the corresponding index sequence ; for the of the i-th frame, its index is calculated as follows: wherein, is the index of the i-th frame of the n-th mode corresponding to the cluster center; is to find the distance to the nearest cluster center index j; is to calculate the Euclidean distance between the feature and the cluster center; finally obtain the index sequence as the representation of the n-th mode; Use the cross-modal attention mechanism to process the index vectors, calculate the attention weights between text features and image features and speech features, and obtain enhanced text features; Among them, using the cross-modal attention mechanism to process the index vectors, calculate the attention weights between text features and image features and speech features, and obtain enhanced text features, including: Map the discrete index sequence into a continuous high-dimensional vector representation through an embedding layer; where is the output of the embedding layer, is the embedding dimension, is the embedding layer function; represents the sequence length of modality n; Use the text features as query vectors, and the image features and speech features as key vectors and value vectors; Among them, is a learnable parameter matrix, d is the feature dimension of each text word embedding, are the dimensions of the key and value vectors respectively; is the length of the text sequence; Generate an attention weight matrix for each word in the text features for each video frame or speech frame by obtaining the attention weights and correlations between text features and image features and speech features; Among them is the attention weight matrix, representing the attention distributions between text features and image features, speech features respectively; Based on the attention weight matrix, extract enhanced non-verbal information from the image features and speech features to generate non-verbal embedding information corresponding to the text features; Among them, is non-verbal embedding information; Fuse the non-verbal embedding information and text features to obtain enhanced text features; the enhanced text features are: In the formula, is the enhanced text feature; ';' represents the concatenation operation, and Gate(;) is a gating mechanism composed of fully connected layers. is the enhanced information extracted from the image feature and aligned with the text feature; is the enhanced information extracted from the speech feature and aligned with the text feature, which is used to fuse the information of the two modalities. A final feature matrix acquisition module configured to: input the enhanced text features into a hierarchical multimodal graph fusion network and use the self-attention mechanism to construct an adjacency matrix; Among them, is the adjacency matrix representing the association between nodes; is the matrix transpose of, and the dot product calculates the similarity between word pairs in the text features, that is, the similarity between each word pair; is the scaling factor, used to stabilize the calculation and prevent the dot product from being too large; is to apply operation to each row of the matrix to ensure that the sum of the elements in each row is 1, that is, to obtain a probability distribution representing the relationship strength between each pair of words; Input the adjacency matrix and the enhanced text features into the skip connection graph convolutional network to obtain the final feature matrix. Among them, the process of inputting the adjacency matrix and the enhanced text features into the skip connection graph convolutional network to obtain the final feature matrix is as follows: Use the adjacency matrix and the enhanced text features as the initial input of the skip connection graph convolutional network, where the enhanced text features serve as the initial feature matrix. Use the adjacency matrix to perform weighted aggregation on the initial feature matrix. The aggregated feature matrix represents the weighted average features of each node and its neighbors. The input feature matrix is added to the transformed feature matrix using skip connections, and the ReLU activation function is used to obtain the final feature matrix; specifically, a skip connection graph convolutional network is used for multi-level information aggregation, and the graph convolutional network of each layer combines node features with the adjacency matrix; the adjacency matrix and the enhanced text features are used as the initial inputs of the skip connection graph convolutional network, and let , then for the -th layer of the skip connection graph convolutional network: Where A is the adjacency matrix; is the feature matrix of the i-th layer; is the learnable weight matrix of the -th layer; ReLU is the activation function; the skip connection term is used to retain the original feature information; The sentiment analysis prediction module is configured to: combine the final feature matrix and the adjacency matrix, and use the encoder and the classifier to obtain the prediction result of sentiment analysis. Specifically, the process of combining the feature matrix and the adjacency matrix and using the encoder and the classifier to obtain the prediction result of sentiment analysis is as follows: Adopt The function is used to splice the final feature matrix and the adjacency matrix to obtain the fused features; Add the fused features to the original text embedding to obtain the text enhanced feature representation. Input the text enhanced feature representation into the trained Transformer encoder and classifier to obtain the prediction result of sentiment analysis.

4. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the multi-modal sentiment analysis method based on attention and graph-enhanced text according to any one of claims 1-2.

5. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the multi-modal sentiment analysis method based on attention and graph-enhanced text according to any one of claims 1-2.

Citation Information

Patent Citations

  • Video question and answer method based on combination of layered and overall multi-modal features

    CN117764086A

  • Mongolian multi-modal sentiment analysis method based on cross-modal transformer

    CN118364427A

  • Method for predicting combustion efficiency and NOx emission of pulverized coal fired boiler in power plant

    CN119025885A

  • Emotion analysis method and device for promoting multi-modal information fusion

    CN119248924A