A Multimodal Sentiment Analysis Method Based on Cross-Modal Attention

By adopting the cross-modal attention mechanism and the multi-headed attention mechanism in multimodal sentiment analysis, the features of text, visual and speech modalities are extracted and fused, and the problem of insufficient feature fusion and interactivity between modals in the prior art is solved, and more accurate sentiment analysis is achieved.

CN119295994BActive Publication Date: 2025-06-24SOUTHWEST PETROLEUM UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411349902.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-06-24
Estimated Expiration
2044-09-26

AI Technical Summary

Technical Problem

The existing multimodal sentiment analysis methods have shortcomings in the characteristics fusion and interactivity between modals, resulting in the limitation of the accuracy and effectiveness of sentiment analysis.

Method used

The multimodal emotion analysis method based on cross-modal attention is adopted, and the timing characteristics of text, visual and speech modalities are extracted through bidirectional gating cyclic units and multi-headed attention mechanisms, and a cross-modal interaction layer and multi-modal fusion layer are designed to enhance information interaction and feature fusion between modals.

Benefits of technology

Through effective intermodal feature fusion and interactive enhancement, the accuracy and depth of sentiment analysis are improved, and the speaker's emotional tendencies can be more accurately reflected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119295994B_ABST
    Figure CN119295994B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent emotion recognition technology, and particularly to a multi-modal sentiment analysis method based on cross-modal attention, including: obtaining a video segment to be analyzed; inputting the video segment to be analyzed into a preset multi-modal sentiment analysis model, and outputting a sentiment analysis result, wherein the multi-modal sentiment analysis model includes an input layer for extracting features of the video segment to be analyzed; a single-modal feature extraction layer for extracting single-modal context information from the features to obtain single-modal context features; a cross-modal interaction layer for performing cross-modal feature interaction on the single-modal context features to obtain a plurality of bimodal joint features; a multi-modal fusion layer for performing fusion processing on the bimodal joint features to obtain cross-modal interaction features; and an output layer for splicing the single-modal context features and the cross-modal interaction features to obtain multi-modal fusion features and output a sentiment analysis result. The present invention can more accurately reflect the emotional tendency of the speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent emotion recognition, and particularly to a multi-modal sentiment analysis method based on cross-modal attention. Background Art

[0002] With the rapid development of industry and the Internet, people's ways of expressing their opinions and views have gradually tended to use a combination of text, images and videos. These users share and display their daily lives on various social media such as Twitter, YouTube, Sina Weibo, etc., and express their views and discussions on various social events. These social information contains rich emotional information. As an important data source for sentiment analysis, if these social information can be collected and analyzed, the research of sentiment analysis can be widely applied to various fields, such as product recommendation, restaurant review, public opinion monitoring and control, mental health guidance, etc.

[0003] Previous sentiment analysis research mainly focused on the text modality. For the existing sources of sentiment information, the emotional expressions presented by a single modality have defects and certain limitations. Ignoring the different inherent sampling rates of temporal information and cross-modal context dependencies between different modalities, multi-modal sentiment analysis not only needs to consider the single-modal feature information extraction technologies of text, vision and speech, but also needs to consider effective fusion strategies between different modalities. Due to the different sampling methods of different modal features, there are large differences in time series and semantic information, and a large amount of low-order signals and redundant information are contained in visual and speech feature information, making the feature information between modalities unable to be effectively fused.

[0004] Therefore, there is an urgent need for a multi-modal sentiment analysis method based on cross-modal attention. Summary of the Invention

[0005] The purpose of the present invention is to provide a multi-modal sentiment analysis method based on cross-modal attention for the defects existing in the current multi-modal sentiment analysis method. With the supplement of visual and speech modalities, sentiment analysis can more accurately reflect the emotional tendency of the speaker at that time.

[0006] To achieve the above purpose, the present invention provides the following solution:

[0007] A multi-modal sentiment analysis method based on cross-modal attention, comprising:

[0008] Obtain the video segment to be analyzed;

[0009] Input the video segment to be analyzed into a preset multi-modal sentiment analysis model, and output the sentiment analysis result. Among them, the multi-modal sentiment analysis model includes an input layer, a single-modal feature extraction layer, a cross-modal interaction layer, a multi-modal fusion layer, and an output layer. The input layer is used to extract the features of the video segment to be analyzed; the single-modal feature extraction layer is used to extract single-modal context information from the features to obtain single-modal context features; the cross-modal interaction layer is used to perform cross-modal feature interaction on the single-modal context features to obtain a number of bimodal joint features; the multi-modal fusion layer is used to perform fusion processing on the bimodal joint features to obtain cross-modal interaction features; the output layer is used to splice the single-modal context features and the cross-modal interaction features to obtain multi-modal fusion features, and output the sentiment analysis result based on the multi-modal fusion features.

[0010] Optionally, the input layer extracting the features of the video segment to be analyzed includes:

[0011] Obtain the text sequence, speech sequence, and visual sequence in the video segment to be analyzed;

[0012] Extract text feature vectors, speech feature vectors, and visual feature vectors respectively based on the text sequence, speech sequence, and visual sequence.

[0013] Optionally, the single-modal feature extraction layer extracting single-modal context information from the features to obtain single-modal context features includes:

[0014] Map the text feature vector, speech feature vector, and visual feature vector to the same dimension through one-dimensional convolution of a fixed dimension to obtain local temporal features;

[0015] Input the local temporal features into a BiGRU to obtain the bidirectional hidden state and context features at each moment, then use the multi-head attention mechanism to strengthen the context features, and calculate the query index similarity of the context features to obtain a query matrix;

[0016] Splice the output of the multi-head attention mechanism to obtain an output result, and use layer normalization to process the output result and the query matrix, and then integrate through a fully connected layer to obtain a single-modal feature vector with context temporal information.

[0017] Optionally, the cross-modal interaction layer performing cross-modal feature interaction on the single-modal context features to obtain a number of bimodal joint features includes:

[0018] Perform pairwise interaction on the single-modal context features;

[0019] Reconstruct the two single-modal context features of the interaction with each other using low-order signals, and then obtain two one-way cross-modal features through several cross-modal transformers;

[0020] Concatenate the two one-way cross-modal features through an activation function to obtain a concatenated feature;

[0021] Calculate the attention distribution of the concatenated feature and multiply the attention distribution by the concatenated feature to obtain the bimodal joint feature.

[0022] Optionally, the working method of the cross-modal transformer is as follows:

[0023]

[0024] Among them, Y α represents the cross-modal attention implementation operator, and the enhanced query (Query) comes from the target modality α, and the key and value come from the auxiliary modality β, are the corresponding weight matrices respectively, represent the sequence features of the target modality α and the auxiliary modality β respectively, and L and d represent the sequence length and the feature dimension respectively.

[0025] Optionally, the multi-modal fusion layer performs fusion processing on the bimodal joint feature to obtain the cross-modal interaction feature, including:

[0026] Concatenate the pairwise interaction features with the corresponding bimodal joint features to obtain a joint feature;

[0027] Based on the joint feature, use the self-attention mechanism to obtain the intra-modal correlation features of the pairwise interaction features;

[0028] After filtering the intra-modal correlation features, obtain the intra-modal attention features of the pairwise interaction features, that is, the cross-modal interaction features.

[0029] Optionally, the method for the output layer to concatenate the single-modal context feature and the cross-modal interaction feature to obtain the multi-modal fusion feature and output the sentiment analysis result based on the multi-modal fusion feature is:

[0030]

[0031] F1 = ReLU(W R Z2 + b R )

[0032] ICA-MSA = Softmax(F1)

[0033] Among them, represents the concatenation operation of vectors. Z2 is the feature vector obtained by concatenating the unimodal context feature and the modality attention feature. Self VT is the text-visual modality attention feature. Self AV is the visual-audio modality attention feature. Self TA is the audio-text modality attention feature. F1 is the multi-modal fusion feature obtained by integrating the concatenated feature vector through a fully connected layer. W R and b R are the initial weights and bias terms of the activation function ReLU. ICA-MSA is the final classification result.

[0034] The beneficial effects of the present invention are as follows:

[0035] Aiming at the problems of insufficient modality fusion and weak interactivity caused by the semantic feature differences between different modalities in multi-modal sentiment analysis, the present invention builds a multi-modal sentiment analysis model by studying and analyzing the potential correlations between different modalities. First, the model uses a bidirectional gated recurrent unit and a multi-head attention mechanism to extract the sequential features of text, visual, and audio modalities with context semantic information. Then, a cross-modal interaction layer is designed to continuously strengthen the target modality with the low-order signals of the auxiliary modality, enabling the target modality to learn the information of the auxiliary modality and capture the potential adaptability between modalities. Next, the enhanced features are input into the multi-modal fusion layer, and the conditional vector is used to further capture the similarity between different modalities, enhance the correlation degree of important features, and explore the deeper interactivity between modalities. Finally, the multi-head attention mechanism is used to concatenate and fuse the features after compound cross-modal reinforcement with the low-order signals, improve the weights of important features within the modality, retain the unique feature information of the initial modality, and perform the final sentiment classification task on the obtained multi-modal fusion features. The model can more effectively explore the correlations between different modalities and make more accurate sentiment judgments. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0037] Figure 1 is a flowchart of a multi-modal sentiment analysis method based on cross-modal attention according to an embodiment of the present invention;

[0038] Figure 2Schematic diagram of the working process and network structure of the cross-modal converter according to an embodiment of the present invention;

[0039] Figure 3 Schematic diagram of the working process and network structure of the cross-modal interaction layer according to an embodiment of the present invention;

[0040] Figure 4 Schematic diagram of the working process and network structure of the multi-modal sentiment analysis model according to an embodiment of the present invention;

[0041] Figure 5 Influence results of different cross-modal attention layers on the model according to an embodiment of the present invention, where (a) is the experimental result of the Acc-2 index, and (b) is the experimental result of the F1-Score index. Detailed implementation manners

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0044] This embodiment provides a multi-modal sentiment analysis method based on cross-modal attention, including:

[0045] Obtain the video segment to be analyzed;

[0046] Input the video segment to be analyzed into a preset multi-modal sentiment analysis model, and output a sentiment analysis result. The multi-modal sentiment analysis model includes an input layer, a single-modal feature extraction layer, a cross-modal interaction layer, a multi-modal fusion layer, and an output layer. The input layer is used to extract the features of the video segment to be analyzed; the single-modal feature extraction layer is used to extract single-modal context information from the features to obtain single-modal context features; the cross-modal interaction layer is used to perform cross-modal feature interaction on the single-modal context features to obtain a plurality of bimodal joint features; the multi-modal fusion layer is used to perform fusion processing on the bimodal joint features to obtain cross-modal interaction features; the output layer is used to splice the single-modal context features and the cross-modal interaction features to obtain multi-modal fusion features, and output the sentiment analysis result based on the multi-modal fusion features.

[0047] Specifically, as Figure 4 shown, the multi-modal sentiment analysis model of this embodiment specifically includes:

[0048] (1) Input layer: Extract primary features of modalities, including text features, data features, and audio features.

[0049] (2) Single-modal feature extraction layer: Use a bidirectional gated recurrent neural network to extract single-modal context feature vectors, mine unique features within the modality, and extract context information.

[0050] (3) Cross-modal interaction layer: Design a multi-modal fusion method based on an improved cross-modal attention mechanism, and continuously strengthen the target modality using low-order signals of the auxiliary modality, enabling the target modality to learn information from the auxiliary modality, capture the potential adaptability between modalities, and achieve information enhancement between different modalities.

[0051] (4) Multi-modal fusion layer: Construct a multi-modal feature fusion module through a self-attention mechanism and a fully connected layer to fuse the enhanced modalities and obtain cross-modal interaction features.

[0052] (5) Output layer: Concatenate the context features and cross-modal joint features, and send them to Softmax through a fully connected layer for the final sentiment analysis task.

[0053] In this embodiment, aiming at the problems of insufficient modality fusion and weak interactivity caused by the semantic feature differences between different modalities in multi-modal sentiment analysis, by studying and analyzing the potential correlations between different modalities, a multi-modal sentiment analysis model is built. The model first uses a bidirectional gated recurrent unit and a multi-head attention mechanism to extract the temporal features of text, visual, and speech modalities with context semantic information; then designs a cross-modal interaction layer to continuously strengthen the target modality using low-order signals of the auxiliary modality, enabling the target modality to learn information from the auxiliary modality and capture the potential adaptability between modalities; then inputs the enhanced features into the multi-modal fusion layer to further capture the similarity between different modalities through a conditional vector, enhance the correlation degree of important features, and mine the deeper interactivity between modalities; finally, use the multi-head attention mechanism to splice and fuse the features strengthened by the composite cross-modal with the low-order signals, improve the weight of important features within the modality, retain the unique feature information of the initial modality, and perform the final sentiment classification task on the obtained multi-modal fusion features. The model can more effectively mine the correlations between different modalities and make more accurate sentiment judgments.

[0054] Further, the features extracted by the input layer for the video segment to be analyzed include:

[0055] Obtain the text sequence, speech sequence, and visual sequence in the video segment to be analyzed;

[0056] Extract text feature vectors, speech feature vectors, and visual feature vectors based on the text sequence, speech sequence, and visual sequence respectively.

[0057] Specifically, the input layer extracts modal primary features, including text data, using 300-dimensional Glove word embeddings with global vectors as text features; for each sentiment polarity audio extraction, the COVAREP speech analysis framework is used to extract 74-dimensional speech high-order statistical features from the speech signal; the Openface facial expression analysis framework is used to obtain visual features, detect the facial action units of the speaker in each frame, extract 35 muscle movement units, and the text, audio, and visual feature dimensions of each video sequence are d T = 300, d A = 74, d V = 35.

[0058] Furthermore, the single-modal feature extraction layer extracts single-modal context information from the features to obtain single-modal context features, including:

[0059] The text feature vector, speech feature vector, and visual feature vector are mapped to the same dimension through one-dimensional convolution with a fixed dimension to obtain local temporal features;

[0060] The local temporal features are input into the BiGRU to obtain the bidirectional hidden states and context features at each moment, and then the multi-head attention mechanism is used to strengthen the context features and calculate the query index similarity of the context features to obtain the query matrix;

[0061] The outputs of the multi-head attention mechanism are concatenated to obtain the output result, and after processing the output result and the query matrix using layer normalization, and then integrating through a fully connected layer, a single-modal feature vector with context temporal information is obtained.

[0062] Specifically, the feature vector is input into the single-modal feature extraction layer, and the temporal information of the text, speech, and visual feature sequences is obtained through the bidirectional gated recurrent unit and the self-attention method respectively respectively represent the text, visual, and audio feature vectors with context temporal information in the final output.

[0063] Furthermore, the cross-modal interaction layer performs cross-modal feature interaction on the single-modal context features to obtain several bimodal joint features, including:

[0064] The single-modal context features are interacted pairwise;

[0065] The two single-modal context features after interaction are reconstructed with each other using low-order signals, and then several cross-modal transformers are passed through to obtain two unidirectional cross-modal features;

[0066] The two unidirectional cross-modal features are concatenated through an activation function to obtain the concatenated feature;

[0067] Calculate the attention distribution of the spliced features, multiply the attention distribution by the spliced features, and obtain the bimodal joint features.

[0068] Specifically, input the three modalities Z T , Z V , Z A into the cross-modal interaction layer. By pairwise combination of different modalities, perform feature fusion between cross-modalities. Taking text (Z T ) and vision (Z V ) as examples, continuously strengthen the target modality through the low-order signals of the auxiliary modality to enhance the semantic correlation between the two modalities, and obtain the bimodal features with enhanced modality information Then, use the feature fusion method of the multi-modal fusion layer to obtain the bimodal joint feature ICA VT .

[0069] Furthermore, the working method of the cross-modal transformer is as follows:

[0070]

[0071] Among them, Y α represents the cross-modal attention implementation operator, and the enhanced query (Query) comes from the target modality α, the key and the value come from the auxiliary modality β. are the corresponding weight matrices respectively, represent the sequence features of the target modality α and the auxiliary modality β respectively, and L and d represent the sequence length and the feature dimension respectively.

[0072] Furthermore, the multi-modal fusion layer performs fusion processing on the bimodal joint features to obtain cross-modal interaction features including:

[0073] Splice the pairwise interaction features with the corresponding bimodal joint features to obtain joint features;

[0074] Based on the joint features, use the self-attention mechanism to obtain the intra-modal correlation features of the pairwise interaction features;

[0075] After filtering the intra-modal correlation features, obtain the intra-modal attention features of the pairwise interaction features, that is, cross-modal interaction features.

[0076] Specifically, input the text-vision features vision-speech features speech-text features into the multi-modal fusion layer to obtain the text-vision modality attention feature Self VT, the visual - speech modality attention feature Self AV , the speech - text modality attention feature Self TA .

[0077] Furthermore, the method for the output layer to splice the unimodal context feature and the cross - modal interaction feature, obtain the multimodal fusion feature, and output the sentiment analysis result based on the multimodal fusion feature is as follows:

[0078]

[0079] F1 = ReLU(W R Z2 + b R )

[0080] ICA - MSA = Softmax(F1)

[0081] where Z2 is the feature vector obtained by splicing the unimodal context feature and the modality attention feature, Self VT is the text - visual modality attention feature, Self AV is the visual - speech modality attention feature, Self TA is the speech - text modality attention feature, F1 is the multimodal fusion feature obtained by integrating the feature vector spliced through the fully - connected layer, W R and b R are the initial weights and bias terms of the activation function ReLU, and ICA - MSA is the final classification result.

[0082] Next, in combination with Figures 1 - 3 a multimodal sentiment analysis method based on cross - modal attention proposed by the present invention will be described in detail, which specifically includes the following steps:

[0083] Step 1: Obtain the text, speech, and visual sequences of a video segment. Each sample data {x1, x2, ···, x L} is divided into a sequence of length L, and each sample sequence is divided into three sequence modalities: text (T), visual (V), and audio (A), which is represented by the feature vector X = [X T , X V , X A .[[]END]]

[0084] Step 2: In Step 1, the text, audio, and visual feature dimensions of each video sequence obtained are d T ==300, d A ==74, d V ==35. Input the feature vector into the unimodal feature extraction layer, and respectively obtain the temporal information of the text, speech, and visual feature sequences through the bidirectional gated recurrent unit and the self - attention method respectively represent the text, visual, and audio feature vectors with context temporal information in the final output.

[0085] Step 3: Input the three modalities Z T , Z V , Z A obtained in Step 2 into the cross-modal interaction layer. By pairwise combining different modalities, perform feature fusion between cross-modalities. Taking text (Z T ) and vision (Z V ) as an example, continuously strengthen the target modality through the low-order signals of the auxiliary modality to enhance the semantic correlation between the two modalities, and obtain the bimodal features with enhanced modality information Then, use the feature fusion method in the composite fusion module to obtain the bimodal joint feature ICA VT .

[0086] Step 4: Repeat Step 3 to obtain the text-vision features with enhanced information vision-speech features speech-text features Then, input the output features of the improved cross-modal interaction into the multi-modal fusion layer to obtain the intra-modal related features Self VT . Finally, perform pairwise cross-modal interactions on the text, speech, and vision modalities respectively to obtain the text-vision modality attention feature Self VT , the vision-speech modality attention feature Self AV , and the speech-text modality attention feature Self TA .

[0087] Step 5: Concatenate the obtained attention feature vectors, and use the fully connected layer to integrate the cross-modal interaction features and intra-modal features obtained, and input them into the Softmax function to perform sentiment classification on the final multi-modal fusion features.

[0088] Furthermore, Step 1 specifically includes using the =Multimodal= SDK toolkit to obtain the unimodal features in the video sequence:

[0089] Step 1.1: For text data, use 300-dimensional Glove word embeddings as text features;

[0090] Step 1.2: For each audio with emotional polarity, extract 74-dimensional speech high-order statistical features from the speech signal using the COVAREP speech analysis framework;

[0091] Step 1.3: Use the FACET facial expression analysis framework to obtain visual features, detect the facial action units of the speaker in each frame, and extract 35 muscle movement units;

[0092] The text, audio, and visual feature dimensions obtained for each video sequence are d T == 300, d A == 74, d V == 35.

[0093] Furthermore, step 2 specifically includes:

[0094] Step 2.1: Use a set of convolutional kernels with the same fixed feature dimension d k (k ∈ {T, V, A}) to extract local temporal features, and map the features of different modalities to the same dimension d:

[0095]

[0096] where L represents the sequence length.

[0097] Step 2.2: After processing by the one-dimensional convolutional layer, input the obtained features into the BiGRU to obtain the bidirectional hidden state at each moment, H = BiGRU(X {T,V,A} ), k ∈ {T, V, A}, capture the long-term dependencies of the context, and then use the multi-head attention mechanism to strengthen the weights of the intra-modal context features and calculate the similarity of the query indices in different subspaces:

[0098]

[0099] where M is the number of attention heads, d h is the hidden state dimension at a certain moment, W Q , W K , W V are the corresponding Querys (Q), Keys (K), and Values (V) mapping matrices respectively.

[0100] Step 2.3: Concatenate all attention heads to obtain the complete output result, and use layer normalization (Layer = Normalization, = LN) to process the accumulated output result and the query matrix elements, M LN= LN(H + M a (H)), and then integrate through the fully connected layer to obtain the single-modal feature vector Z:

[0101]

[0102] where Z = [Z T , Z V , Z A , respectively representing the text, visual, and audio feature vectors with context temporal information in the final output.

[0103] Further, step 3 specifically includes:

[0104] Step 3.1: The cross-modal attention network is improved in this embodiment as Figure 2 shown. Since the sequence sampling rates of different modalities are different, most of the data is unaligned in the time series. The Cross-Modal Transformer (CM) considers the interaction between different time series and solves the problem of data misalignment in the multi-modal fusion process. For the target modality α and the auxiliary modality β, the sequence features of each modality (unaligned) are respectively represented as and where L and d represent the sequence length and the feature dimension. The potential adaptation from modality β to modality α is expressed as follows:

[0105]

[0106] where, Y α represents the cross-modal attention implementation operator, and the enhanced query comes from the target modality α, the key and the value come from modality β, are the corresponding weight matrices respectively. The source modality β reconstructs the feature information of the target modality α with low-order signals, interacts through the key / value, and captures important correlation information across modalities.

[0107] Step 3.2: The cross-modal attention network is improved in this embodiment as Figure 3 shown. Taking the text modality (Z T ) and the visual modality (Z V ) as examples, the single-modal feature sequences with context time series information are input into the cross-modal attention interaction layer. Let Z V→T , Z T→V respectively represent the reconstruction of text features by visual low-order signals and the reconstruction of visual features by text low-order signals. To avoid gradient explosion caused by excessive numerical values, layer normalization LN and residual connections are added. Then, the similarity between the auxiliary modality and the target modality is calculated bidirectionally using two CM modules, and the target modality is continuously strengthened. The feed-forward calculation is performed according to the i = 1, 2,..., n layers. Taking Z V→T as an example, the calculation process is as follows:

[0108]

[0109] Among them, P1 is used as an intermediate variable to obtain the feed-forward value of cross-modal attention. f is a position feed-forward sub-layer parameterized by θ, which gives sufficient attention to the (i-1)-th time step of the visual modality through the i-th time step of the text modality. The i-th time step of P1 is the weighted sum of V β , and the weights are determined by the i-th row in Softmax.

[0110] Step 3.3 specifically includes: continuously strengthening the text modality by using low-level visual modality information, so that the text modality can learn information from the visual modality. After n layers of cross-modal attention stacking, unidirectional cross-modal text features are obtained Meanwhile, the visual features after the text modality strengthens the visual modality are also obtained The enhanced features of the two modalities are concatenated through the activation function tanh to capture the hidden correlation between modalities and obtain the feature vector The attention distribution matrix att1 is calculated using the Softmax function, and then the obtained attention distribution matrix is multiplied by the matrix of multi-modal feature fusion to obtain the final weighted cross-modal interaction features ICA of text and vision VT , and the specific calculation is as follows.

[0111]

[0112] Among them, represents the concatenation operation of vectors, W VT and b VT represent the mapping matrix and bias term after feature concatenation respectively, and · represents matrix multiplication.

[0113] Furthermore, Step 4 specifically includes: inputting the output features of the improved cross-modal interaction into the self-attention layer to obtain the intra-modal relevant features Self VT , and the specific calculation is as follows:

[0114]

[0115]

[0116] Self VT = att2·Z1

[0117] Among them, Z1 represents the feature vector obtained by concatenating text features, speech features and cross-modal interaction features, and Self VT represents the attention feature vector after information filtering.

[0118] Furthermore, Step 5 specifically includes: performing pairwise cross-modal interactions on the text, speech, and visual modalities respectively to obtain the text-visual modality attention feature Self VT, the visual-audio modality attention feature Self AV , the audio-text modality attention feature Self TA , concatenate the obtained attention feature vectors, and use a fully connected layer to integrate the cross-modal interaction features and intra-modal features obtained, and input them into Sofmax for sentiment classification. The calculation process is as follows:

[0119]

[0120] F1 = ReLU(W R Z2 + b R )

[0121] ICA-MSA = Softmax(F1)

[0122] Among them, Z2 is the feature vector obtained by concatenating the attention feature vector and the context feature vector, F1 is the multi-modal fusion feature obtained by integrating the concatenated feature vector through a fully connected layer, W R and b R are the initial weights and bias terms of the activation function ReLU, and ICA-MSA is the final classification result.

[0123] Experimental verification:

[0124] (1) Comparative experiment:

[0125] Table 1

[0126]

[0127] The experimental results are shown in Table 1. In the binary sentiment classification task, the ICA-MSA model proposed in this embodiment shows better performance in all indicators compared with other models. Under two different data sets, compared with the simple feature-level fusion EF-LSTM and decision-level fusion LF-LSTM models, the ICA-MSA model has significant improvements in all indicators. Among them, the binary classification accuracy has increased by about 5%. Compared with the models based on non-attention methods for multi-modal fusion such as TFN, LMF, and MFM, the binary classification accuracy has increased by about 3.5%. This fully shows that the multi-modal fusion method based on the attention fusion strategy can better capture the interaction between different modalities. Compared with the MTCN and MulT models, the improved cross-modal attention method used in this embodiment also shows certain advantages, further indicating that through the improved cross-modal fusion strategy, considering the different effects of the modality's own temporal sequence and low-high order modality information on the sentiment classification result, it can more accurately focus on the deeper correlation within the modality, enabling the model to learn a more complete emotional expression and thus make a more accurate emotional judgment.

[0128] (2) Model ablation experiment comparison:

[0129] Table 2

[0130]

[0131] The experimental results are shown in Table 2. The sentiment analysis results of the fusion of text, visual, and audio modal features are the best, fully demonstrating the importance of multimodal fusion. From different datasets, it shows that larger-scale training data can improve the performance of sentiment analysis to a certain extent. In the experiments of single-modal and dual-modal on the same dataset, the text single-modal and the dual-modal with text modality have significantly better results than the other two, indicating that in general, the sentiment polarity of text modal features is the most significant. Compared with single-modal, the two index results of the model are significantly better in dual-modal, but the accuracy of the V+A modality is slightly lower than that of the visual modality. This is due to the relatively weak sentiment polarity of visual and audio modalities and the interference of redundant information. Therefore, effectively fusing text, visual, and audio three-modal features helps to improve the performance of sentiment classification and makes up for the problem of incomplete single-modal sentiment expression.

[0132] (3) Influence of hyperparameters:

[0133] The number of stacked cross-modal attention layers n in the cross-modal interaction layer is also a main hyperparameter affecting the final performance of the model, which specifies the number of cross-modal Transformers. Feedforward calculations are performed from layer i = 1 to layer n, and the auxiliary modality continuously updates its sequence information. To study the influence of the number of cross-modal attention layers n on the overall performance, n is set from 1 to 9, and the experimental indicators include Acc-2 and F1-Score. Figure 5 (a) represents the experimental results of the Acc-2 index. Figure 5 (b) represents the experimental results of the F1-Score index. According to the analysis of the experimental results, the two indicators have a similar trend of change. The accuracy and F1 value first increase and then decrease. When the number of stacked cross-modal modules is 4, the model performance is the best, indicating that the number of layers and performance are not positively correlated. When the number of layers is higher than 4, the model effect begins to decline significantly, and the model training is more difficult and uncontrollable.

[0134] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A multimodal sentiment analysis method based on cross-modal attention, characterized in that: include: Get the video segment to be analyzed; Input the video segment to be analyzed into a preset multimodal sentiment analysis model, and output the sentiment analysis result, wherein the multimodal sentiment analysis model includes an input layer, a unimodal feature extraction layer, a cross-modal interaction layer, a multimodal fusion layer, and an output layer, wherein the input layer is used to extract the features of the video segment to be analyzed; the unimodal feature extraction layer is used to extract unimodal context information from the features to obtain unimodal context features; the cross-modal interaction layer is used to perform cross-modal feature interaction on the unimodal context features to obtain a number of bimodal joint features; the multimodal fusion layer is used to fuse the bimodal joint features to obtain cross-modal interaction features; the output layer is used to splice the unimodal context features and the cross-modal interaction features to obtain multimodal fusion features, and output the sentiment analysis result based on the multimodal fusion features; The cross-modal interaction layer performs cross-modal feature interaction on the unimodal context features to obtain a plurality of bimodal joint features, including: Interacting the unimodal context features in pairs; The two interacting unimodal context features are reconstructed using low-order signals, and then passed through several cross-modal converters to obtain two unidirectional cross-modal features; The two unidirectional cross-modal features are concatenated through an activation function to obtain concatenated features; Calculating the attention distribution of the splicing feature, and multiplying the attention distribution by the splicing feature to obtain the bimodal joint feature; Among them, the two unidirectional cross-modal features are spliced ​​through the activation function to obtain the spliced ​​features. The calculation method is: in, is the concatenation feature of unidirectional cross-modal text features and unidirectional cross-modal visual features. is a unidirectional cross-modal text feature, is a unidirectional cross-modal visual feature, n represents the number of cross-modal converters, Represents the concatenation operation of vectors, W VT and b VT Respectively represent the mapping matrix and bias term after feature concatenation; The attention distribution of the splicing feature is calculated, and the attention distribution is multiplied by the splicing feature to obtain the bimodal joint feature. The calculation method is: Among them, att1 is the attention distribution matrix calculated using the Softmax function, ICA VT It is a bimodal joint feature of text features and visual features; The multimodal fusion layer performs fusion processing on the bimodal joint features to obtain cross-modal interaction features, including: The pairwise interaction features are concatenated with the corresponding bimodal joint features to obtain the joint features; Based on the joint features, a self-attention mechanism is used to obtain the modal internal correlation features of the pairwise interactive features; After filtering the relevant features within the modality, modal attention features of the features interacting with each other are obtained, that is, the cross-modal interaction features; The calculation method of the multimodal fusion layer is: Self VT =att2·Z1 Among them, Z1 represents the text feature Z T , speech feature Z V And the corresponding bimodal joint feature ICA VT The concatenated feature vector, att2 represents the attention distribution matrix calculated by the self-attention mechanism, Self VT Represents text-visual modality attention features.

2. The multimodal sentiment analysis method based on cross-modal attention according to claim 1, characterized in that: The input layer extracts the features of the video segment to be analyzed including: Obtaining a text sequence, a voice sequence, and a visual sequence in the video segment to be analyzed; A text feature vector, a speech feature vector, and a visual feature vector are extracted based on the text sequence, the speech sequence, and the visual sequence, respectively.

3. The multimodal sentiment analysis method based on cross-modal attention according to claim 2 is characterized in that: The unimodal feature extraction layer extracts unimodal context information from the feature, and obtaining the unimodal context feature includes: Mapping the text feature vector, the speech feature vector, and the visual feature vector to the same dimension through a fixed-dimensional one-dimensional convolution to obtain local temporal features; Input the local time series features into BiGRU, obtain the bidirectional hidden state and context features at each moment, then use the multi-head attention mechanism to enhance the context features, calculate the query index similarity of the context features, and obtain the query matrix; The outputs of the multi-head attention mechanism are spliced ​​to obtain the output result, and the output result and the query matrix are processed by layer normalization, and then integrated through a fully connected layer to obtain a unimodal feature vector with contextual temporal information.

4. The multimodal sentiment analysis method based on cross-modal attention according to claim 1, characterized in that: The working method of the cross-modal converter is: Among them, Y α Represents the cross-modal attention implementation operator, the enhanced query (Query) From target modality α, key and value From the auxiliary modal β, are the corresponding weight matrices, They represent the sequence features of the target modality α and the auxiliary modality β respectively, L and d represent the sequence length and feature dimension respectively.

5. The multimodal sentiment analysis method based on cross-modal attention according to claim 1, characterized in that: The output layer splices the unimodal context features and the cross-modal interaction features to obtain multimodal fusion features, and the method for outputting the sentiment analysis result based on the multimodal fusion features is: F1=ReLU(W R Z2+b R ) ICA-MSA=Softmax(F1) in, Represents the concatenation operation of the vector, Z2 is the feature vector obtained by concatenating the unimodal context feature and the modal attention feature, Self VT is the text-visual modality attention feature, Self AV is the visual-speech modality attention feature, Self TA is the speech-text modality attention feature, F1 is the multimodal fusion feature obtained by integrating the concatenated feature vectors through the fully connected layer, and W R and b R are the initialization weights and bias terms of the activation function ReLU, and ICA-MSA is the final classification result.