Sentiment Analysis Method for Image-Text Fusion Based on Multimodal Cross-Attention Mechanism
Through the multimodal cross-attention mechanism of ALBert and DenseNet121 combined with the CBAM attention mechanism, the problem of poor coordination of information redundancy and modal correlation in multimodal sentiment analysis is solved, and more accurate sentiment analysis is achieved.
Patent Information
- Application Number
- CN202310848751.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-07-11
AI Technical Summary
The existing multimodal sentiment analysis methods have problems of poor coordination of information redundancy and modal correlation in the process of feature extraction and fusion, resulting in insufficient accuracy of sentiment analysis.
Using a multimodal cross attention mechanism, text features are extracted through the ALBert model and emotional vocabulary and context features are processed using BiLSTM, image features are extracted in combination with DenseNet121 network and CBAM attention mechanism, feature fusion is used for feature fusion, and finally sentiment analysis is performed through softmax classifier.
The accuracy and robustness of multimodal sentiment analysis are improved, and the effect of sentiment analysis is improved by fully exploring the complementarity and correlation between images and text.
Smart Images

Figure CN116844179B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sentiment analysis method based on image-text fusion with a multimodal cross-attention mechanism. Background Art
[0002] Research on single-modal sentiment analysis has achieved great success, but emotions in real life are mostly multimodal. Not only text, but also pictures, audio, video and other forms, each modality promotes each other. If the connection between each modality can be explored, the accuracy of sentiment analysis will be further improved.
[0003] With the deepening of research, multimodal sentiment analysis has received more and more attention in recent years. It is also a very challenging research topic. Compared with unimodal sentiment analysis, multimodality is an interdisciplinary research topic that combines natural language processing, computer vision and other fields.
[0004] Based on the research on different fusion methods, it is mainly divided into three categories, namely early fusion, intermediate fusion and late fusion.
[0005] Based on the early fusion method, the features of multiple modal information are first extracted, and then fused by splicing, weighting, etc. For example: by fusing text and images in a unified bag of words, feature splicing is performed, and the final representation is output; the text features are extracted using a binary representation, the underlying image features are extracted using the mutual information method, and the results are classified into two categories based on the similarity neighborhood classifier. Late fusion will train the data of each modality separately, select the most appropriate classifier, and fuse the final results for output. For example: first use CNN to extract image features and text features respectively, then use logistic regression to predict and analyze different emotions, and finally use average and weighted fusion strategies to perform the final emotion prediction analysis; a deep multimodal attention fusion model, first use attention to extract text and image features respectively, then obtain image-text fusion features through early fusion, and finally use late fusion to integrate the classifier results; a hybrid deep learning model, first perform fine-grained analysis of multimodal data, and then use decision-level multimodal combination to classify and output multimodal data; a new bidirectional multi-level attention model to study the complementarity and correlation between images and text.
[0006] Early fusion mainly extracts features from multiple modalities of information and then fuses them through concatenation, weighting, etc. However, the fusion result may contain a large number of redundant vectors, causing information redundancy and information dependence, which in turn leads to poor fusion effects. Intermediate fusion mainly uses neural networks to share the intermediate layers of the network during the fusion process. Late fusion allows each modality to be trained in the most suitable way for itself, predicts using different classifiers, and then performs decision fusion. However, late fusion cannot well coordinate the correlation between modalities. Summary of the Invention
[0007] To solve the above technical problems, the present invention provides a sentiment analysis method based on multimodal cross-attention mechanism for text-image fusion.
[0008] A sentiment analysis method based on multimodal cross-attention mechanism for text-image fusion, including:
[0009] Obtain the text and image to be processed;
[0010] Vectorize the text to obtain text features; process the text features through BiLSTM to obtain context features containing sentiment words;
[0011] Obtain the image features of the image, and use CBAM attention to obtain the sentiment feature region features in the image features from both spatial and channel aspects;
[0012] Fuse the extracted context features and sentiment feature region features through a cross-attention mechanism to obtain cross features after cross-attention fusion, and perform sentiment classification based on the obtained cross features and a classifier to obtain the sentiment analysis result.
[0013] Further, the vectorizing the text to obtain text features includes: using the ALBert pre-trained model to vectorize the text to obtain text features.
[0014] Further, the ALBert pre-trained model includes an input layer, an ALBert pre-training layer, and a Transformer encoder;
[0015] The using the ALBert pre-trained model to vectorize the text to obtain text features includes:
[0016] The text passes through the input layer to the ALBert pre-training layer, and the ALBert pre-training layer is used to obtain the numbers of each character in the text and vectorize them to obtain an information sequence;
[0017] The information sequence is input into the Transformer encoder to mine deep semantic feature information, and text features are obtained after transformation.
[0018] Further, the obtaining of the image features of the image includes:
[0019] Using the pre-trained DenseNet121 network model to obtain the image features of the image.
[0020] Further, it is set that the feature map obtained after the i-th image passes through the j-th layer of convolution in the DenseNet121 network model is where C is the number of channels, H is the length of the feature map, and W is the width of the feature map;
[0021] The using of CBAM attention to obtain the emotional feature region features in the image features from both spatial and channel aspects includes:
[0022] Using CBAM attention to obtain the attention weights of the j-th feature map of the i-th image The calculation method is as follows:
[0023]
[0024] where, M c (F ij ) is the channel attention weight of the j-th feature map of the i-th picture; M s (F ij ) is the spatial attention weight of the j-th feature map of the i-th picture;
[0025] The channel attention weight M c (F ij ) is calculated as follows:
[0026] M c (F ij ) = σ(MLP(AvgPool(F ij )) + MLP(MaxPool(F ij )))
[0027] = σ(W1(W0(AvgPool(F ij )))+W1(W0(MaxPool(F ij ))))
[0028] Among them, σ is the Sigmoid activation function; the weights W0 and W1 of the MLP are shared by two inputs, and the ReLu activation function is followed by W0; AvgPool(·) is the global average pooling function, which calculates the average value of all feature regions of each feature map; MaxPool(·) is the global maximum pooling function, which calculates the maximum eigenvalue of each feature map;
[0029] Spatial attention weight M s (F ij ) is calculated as follows:
[0030] M s (F ij ) = σ(f 7×7 ([AvgPool(F ij ), MaxPool(F ij )]))
[0031] Among them, [·] is the concatenation operation; f 7×7 (·) represents the convolution operation, and the convolution kernel size is 7×7; AvgPool(·) is the average pooling function, and MaxPool(·) is the maximum pooling function;
[0032] The calculation formula of the CBAM attention feature map is as follows:
[0033]
[0034] Among them, represents element-wise multiplication of corresponding elements;
[0035] After the convolution operation, the feature mapping vector I = [I1, I2,..., I n of the key region of each image is obtained.
[0036] Furthermore, for the extracted context features and sentiment feature regions, they are fused through the cross-attention mechanism to obtain the cross features after cross-attention fusion. According to the obtained cross features and the classifier, sentiment classification is performed to obtain the sentiment analysis results, including:
[0037] The cross-attention mechanism is composed of scaled dot-product attention. The context features are used as the query matrix in the scaled dot-product attention, and the sentiment feature regions are used as the key-value matrix; through the scaled dot-product attention, the cross-attention features are obtained as follows:
[0038]
[0039]
[0040]
[0041]
[0042] Among them, W Q , W K , W V are all parameter matrices; d k is the number of columns of Q and K; I i , T i are the emotional feature region and the context feature respectively; Att(·) is the cross-attention mechanism module network;
[0043] Obtain the cross feature after cross-attention fusion: Y c =[Y c1 , Y c2 , …, Y cn ;
[0044] For the output Y c after cross-attention fusion, as the input of the linear function softmax, perform the final sentiment classification, as shown in the following formula:
[0045] y = softmax(W c Y c + b c )
[0046] Among them, W c represents the weight matrix, and b c represents the bias term.
[0047] The present invention has the following beneficial effects: The obtained text and image are processed differently. Among them, text features are obtained, and the text features are processed by BiLSTM to obtain context features containing emotional words; image features of the image are obtained, and CBAM attention is used to obtain the emotional feature region features in the image features from both spatial and channel aspects. Finally, the extracted context features and emotional feature region features are fused through the cross-attention mechanism to obtain the cross feature after cross-attention fusion. According to the obtained cross feature and the classifier, sentiment classification is performed to obtain the sentiment analysis result. Combining the image features and text features, the multi-modal sentiment analysis process is more accurate and has better effects compared with the existing single-modal sentiment analysis methods and other multi-modal sentiment analysis processes. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is the overall flowchart of a sentiment analysis method based on multi-modal cross-attention mechanism for text and image fusion provided by the present invention;
[0049] Figure 2 is the flowchart of the text feature extraction network;
[0050] Figure 3 is the flowchart of the image feature extraction network;
[0051] Figure 4 is the flowchart of the network of the sentiment analysis method provided by the invention, which is based on multi-modal cross-attention mechanism for text and image fusion;
[0052] Figure 5 is the statistical chart of the text length frequency of the dataset;
[0053] Figure 6 is the comparison chart of the performance indicators of different models on the MVSA and TumEmo datasets;
[0054] Figure 7 is the schematic diagram of the change of the training set and accuracy of different proportions of different models on the MVSA and TumEmo datasets;
[0055] Figure 8 is the PR curve graph of different classifier models on the MVSA and TumEmo datasets;
[0056] Figure 9 is the ablation experiment result graph of different datasets. Detailed implementation mode
[0057] This embodiment provides a sentiment analysis method based on multi-modal cross-attention mechanism for text and image fusion, and proposes a multi-modal sentiment analysis model based on cross-attention mechanism for text and image fusion, which utilizes the complementarity and relevance between images and texts. The main idea is to use the attention mechanism to extract the emotional regions and emotional words from images and texts respectively; then use the cross-attention mechanism to fuse the extracted emotional features; finally, pass the fused features through a classifier to output the prediction results. The main idea of the sentiment analysis method based on multi-modal cross-attention mechanism for text and image fusion is as follows: for the text, first use the ALBert pre-trained model to transform the text into vectorization, and then use BiLSTM to obtain the text context features; for the image, first use the DenseNet121 network to extract features, and then use the CBAM mechanism to obtain the features of the key regions from two aspects of channels and space, and obtain the emotional regions of the corresponding regions; for the obtained emotional words and emotional region features respectively, perform feature fusion through the cross-attention mechanism, and perform the final result output through the softmax classifier.
[0058] Before specifically describing the sentiment analysis method based on multi-modal cross-attention mechanism for text and image fusion, the related technical means used will be introduced first.
[0059] DenseNet network:
[0060] DenseNet (Densely Connected Convolutional Network) is a deep convolutional neural network model proposed by Gao Huang et al. in 2017. It uses the idea of dense connection, enabling the network to have stronger feature transfer and reuse capabilities. It can train very deep neural networks with high accuracy.
[0061] In traditional convolutional neural networks, the output of each layer is obtained by applying a convolution and a non-linear activation function to the input of the previous layer.
[0062] Assume the input is an image X0, passing through an L-layer neural network, where the non-linear transformation of the i-th layer is denoted as H i (*), H i (*) can be the accumulation of various functional operations such as BN, ReLU, Pooling, or Conv, etc. The feature output of the i-th layer is denoted as X i .
[0063] Traditional convolutional feedforward neural networks take the output X i of the i-th layer as the input of the (i + 1)-th layer, which can be written as X i = H i (X i-1 ). ResNet adds a bypass connection and can be written as follows:
[0064] X l = H l (X l-1 ) + X l-1
[0065] One of the main advantages of ResNet is that the gradient can flow through the identity function to reach the earlier layers. However, the way of superimposing the identity mapping and the output of the non-linear transformation is addition, which to some extent disrupts the information flow in the network. DenseNet proposes the concept of dense blocks. Each dense block contains several convolutional layers and a concatenation, enabling the features of all previous layers to be directly passed to the subsequent layers, thus forming a dense network. A dense block consists of several convolutional layers and a concatenation.
[0066] The input of each layer is the concatenation of all the feature maps generated by all previous layers in the same dense block of DenseNet. The output of the (l - 1)-th layer is recorded as X l-1 . The input of the i-th layer is related not only to the output of the (i - 1)-th layer but also to the outputs of all previous layers. The output of the l-th layer is as follows:
[0067] X l =H l ([X0,X1,…,X l-1 )
[0068] where [] represents concatenation, that is, all the output feature maps of layers from X0 to X l-1 are combined by channel. The non-linear transformation function H l (·) here is a combination of BN + ReLU + Conv(3×3).
[0069] Among them, the input of each convolutional layer contains the outputs of all the previous layers, and the output will be directly passed to all the subsequent layers. If the dimensions of the input and output are inconsistent, an additional convolutional layer is required for transformation so that the input and output can be connected.
[0070] The advantage of using a dense block is that it can make the network have stronger feature transfer and reuse capabilities, thereby improving the accuracy. In addition, skip connections can also make the network have stronger feature reuse and generalization capabilities, thus further improving the accuracy.
[0071] ALBert pre-trained model:
[0072] ALBert (A Lite BERT) is a natural language processing model based on the BERT (Bidirectional Encoder Representations from Transformers) model. It is a self-supervised pre-trained model that can be used for natural language processing tasks such as text classification, named entity recognition, question answering, and text generation.
[0073] The main feature of ALBert is that it adopts two pre-training tasks: Masked Language Model (MLM) and Next Sentence Prediction (NSP). The MLM task randomly selects some tokens in the input sequence and replaces them with [MASK], and then lets the model predict the masked tokens; the NSP task is to let the model predict whether two sentences are adjacent sentences. Through the pre-training of these two tasks, ALBert can learn richer language representations, thereby improving its performance in natural language processing tasks.
[0074] The structure of ALBert is similar to BERT. It consists of multiple Transformer modules, and each Transformer module contains a multi-head self-attention layer and a feed-forward neural network layer. Different from BERT, ALBert adopts cross-layer connections and global pathways to enhance the ability of feature transfer and information flow, thereby improving the efficiency and performance of the model.
[0075] The advantages of ALBERT compared to BERT are mainly as follows:
[0076] (1) More efficient training: ALBERT uses two new training strategies: cross-layer parameter sharing and sentence order prediction. Cross-layer parameter sharing can significantly reduce the number of model parameters and the complexity of model training, enabling larger models to be trained. Sentence order prediction can increase the amount of training data and improve the generalization ability of the model. These improvements enable ALBERT to train larger and more efficient models with the same computational resources.
[0077] (2) Better performance: ALBERT achieved better results than BERT in the GLUE (General Language Understanding Evaluation) benchmark evaluation task, indicating that ALBERT has better performance in natural language understanding tasks compared to BERT.
[0078] (3) Better generalization ability: ALBERT uses more comprehensive pre-training and can better learn general language representations, thus having better generalization ability in downstream tasks. In addition, ALBERT also uses more data augmentation techniques, can better adapt to the data distribution of different tasks, and improves the robustness of the model.
[0079] (4) More flexible model structure: ALBERT provides a variety of different model structures and hyperparameter options, which can be flexibly adjusted according to specific application scenarios. For example, different numbers of layers, hidden layer sizes, cross-layer sharing strategies, etc. can be selected, making ALBERT more suitable for different scenarios and task requirements.
[0080] In summary, compared to BERT, ALBERT has improvements in terms of training efficiency, performance, generalization ability, and model flexibility. Therefore, in this embodiment, ALBert is selected as the pre-trained model.
[0081] CBAM attention mechanism:
[0082] CBAM (Convolutional Block Attention Module) is an attention mechanism module for image recognition, proposed by Sanghyun Woo et al. from the KAIST Machine Learning Research Center in 2018. The CBAM module can introduce spatial and channel attention mechanisms into convolutional neural networks, thereby improving the accuracy and robustness of image recognition.
[0083] Specifically, the CBAM module includes two attention sub-modules: the channel attention module and the spatial attention module. The channel attention module is used to adaptively adjust the weights of different channels to increase the model's attention to important features. The spatial attention module is used to adaptively adjust the weights of different spatial positions to increase the model's attention to important positions. Through the combination of these two sub-modules, the CBAM module can adaptively adjust the weights of each channel and position in the feature map, thereby reducing the dependence on global features and improving the robustness and generalization ability of the model.
[0084] Given an intermediate feature map F ∈ R C×H×W as the input, the operation process of CBAM is generally divided into two parts. First, perform global max-pooling and average pooling on the input by channel, send the two one-dimensional vectors after pooling into a fully connected layer for operation and then add them to generate a one-dimensional channel attention M C ∈ R C×1×1 , then multiply the channel attention with the input elements to obtain the feature map F' adjusted by the channel attention; secondly, perform global max-pooling and average pooling on F' by space, splice the two two-dimensional vectors generated by pooling and then perform a convolution operation to finally generate a two-dimensional spatial attention M S ∈ R 1×H×W , and then multiply the spatial attention with F' element-wise. The process of CBAM generating attention is as follows:
[0085]
[0086]
[0087] where denotes element-wise multiplication. Before the multiplication operation, the channel attention and the spatial attention need to be broadcasted accordingly according to the spatial dimension and the channel dimension respectively.
[0088] As Figure 1 shown, a sentiment analysis method based on multi-modal cross-attention mechanism for text and image fusion provided in this embodiment includes the following steps:
[0089] Step 1: Obtain the text and image to be processed:
[0090] Obtain the text and image that need sentiment analysis.
[0091] Step 2: Vectorize the text to obtain text features; process the text features through BiLSTM to obtain context features containing sentiment words:
[0092] It should be understood that for an input sentence, its sentiment information is often related to certain words in the text. Therefore, in the process of text sentiment analysis, first, the text content is vectorized through the pre-trained model ALBert, and then the sentiment degree of the text content is further obtained through BiLSTM by combining the information in the forward and backward directions of the input sequence. The extraction process of text features is as Figure 2 shown.
[0093] First, use the ALBert pre-trained model to vectorize the text to obtain text features. In this embodiment, the ALBert pre-trained model includes an input layer, an ALBert pre-training layer, and a Transformer encoder.
[0094] Use the ALBert pre-trained model to vectorize the text sequence. The ALBert pre-trained model performs a sub-word operation and can vectorize the corpus text at the character level. For each input text m composed of n characters, it can be expressed as: m = [m1, m2, …, m n , where m i represents the i-th character in the text.
[0095] The text m passes through the input layer to the ALBert pre-training layer. First, for each character in the text information, mark its position number in the dictionary to obtain the corresponding number, and perform vectorization of the text content to obtain the information sequence e. e i represents the vectorization corresponding to the i-th character, as shown in the following formula:
[0096] e = [e1, e2, …, e n
[0097] For the information sequence e, then input it into the Transformer encoder in the ALBert pre-trained model to mine deep semantic feature information, and finally obtain the feature vector z of the text sequence after transformation. z i represents the feature vector corresponding to the i-th character, as shown in the following formula:
[0098] z = [z1, z2, …, z n
[0099] The feature vector z corresponding to each character i i As an input, the BiLSTM network can simultaneously combine the information of the input sequence in both the forward and backward directions, better mine the deep semantic information, and further strengthen the sentiment information in the text. For the input z at time t it , its forward output and backward output are calculated as follows:
[0100]
[0101]
[0102] For the feature vector z i of each character i, by concatenating its forward and backward context information, a feature representation with sentiment information is obtained, as follows:
[0103]
[0104] where [·] is the vector concatenation operation.
[0105] After calculation, the output feature vector of each character i in the input text at time t is obtained, as follows:
[0106]
[0107] Finally, the feature vector of each input text containing sentiment information, that is, the context feature containing sentiment words is: T = [T1, T2, …, T n .
[0108] Step 3: Obtain the image features of the image, and use CBAM attention to obtain the feature of the sentiment feature region in the image features from both the spatial and channel aspects:
[0109] In an image, usually certain regions can better reflect the sentiment tendency of the whole image. If these regions with the most sentiment features can be mined, the result of sentiment analysis will be more accurate. First, use the pre-trained DenseNet121 network model to extract the image features, and then use CBAM attention to extract the regions with the most sentiment features in the picture from both the spatial and channel aspects, which improves the expression ability of the overall and local features. The process of extracting the image features is as Figure 3 shown. Let X = {X1, X2, …, X n} represent a dataset of n images. For each image X i , use the DenseNet network to preprocess the image in the input layer, and then through the CBAM attention mechanism, finally extract the feature vector I i of each image, that is, the feature of the sentiment feature region in the image features.
[0110] Specifically:
[0111] It is set that the feature map obtained after the i-th image passes through the j-th convolution layer in the DenseNet121 network model is where C is the number of channels, H is the length of the feature map, and W is the width of the feature map. F′ ij is the attention feature map obtained after attention weighting.
[0112] Then the attention weight of the j-th feature map of the i-th image is calculated as follows:
[0113]
[0114] where M c (F ij ) is the channel attention weight of the j-th feature map of the i-th picture; M s (F ij ) is the spatial attention weight of the j-th feature map of the i-th picture.
[0115] Channel attention focuses on which features on which channels are meaningful, specifically manifested in the contribution of each feature in the feature map after convolution to the key information. The channel attention weight M c (F ij ) is calculated as follows:
[0116] M c (F ij ) = σ(MLP(AvgPool(F ij )) + MLP(MaxPool(F ij )))
[0117] = σ(W1(W0(AvgPool(F ij )))+W1(W0(MaxPool(F ij ))))
[0118] where σ is the Sigmoid activation function; the weights W0 and W1 of MLP are shared by two inputs, followed by W0 with the ReLu activation function; AvgPool(·) is the global average pooling function, calculating the average value of all feature regions of each feature map; MaxPool(·) is the global maximum pooling function, calculating the maximum eigenvalue of each feature map.
[0119] The input feature map is First, a global average pooling and a global maximum pooling are respectively performed to obtain two results as The feature maps are then fed into a two-layer fully-connected neural network separately. For these two feature maps, the two-layer fully-connected neural network shares parameters. Then, the two obtained feature maps are added together, and then a weight coefficient between 0 and 1 is obtained through the Sigmoid function. Then, the weight coefficient is multiplied by the input feature map to obtain the final output feature map.
[0120] Spatial attention focuses on which parts of the features in the space are meaningful, specifically manifested in the contribution of local regions of the picture to key information, and can identify the regions that need to be focused on in the picture information. The spatial attention weight M s (F ij ) is calculated as follows:
[0121] M s (F ij ) = σ(f 7×7 ([AvgPool(F ij ), MaxPool(F ij )]))
[0122] where σ is the Sigmoid activation function; [·] is the concatenation operation; f 7×7 (·) represents the convolution operation, which can obtain the influence of different local regions of the feature map on key information, and the convolution kernel size is 7×7; AvgPool(·) is the average pooling function, and MaxPool(·) is the maximum pooling function.
[0123] The input feature map is Max-pooling and average pooling are performed separately in the channel dimension to obtain two feature maps, and then these two feature maps are concatenated in the channel dimension, and the obtained feature map is Then, it passes through a convolutional layer and is reduced to 1 channel. The convolution kernel uses 7×7, and at the same time, H×W remains unchanged. The output feature map is Then, the spatial weight coefficient is generated through the Sigmoid function, and then it is multiplied by the input feature map to obtain the final feature map.
[0124] The calculation formula of the CBAM attention feature map is as follows:
[0125]
[0126] where represents element-wise multiplication;
[0127] After the convolution operation, the feature mapping vector I = [I1, I2, …, I n of the key region of each image is obtained.
[0128] Step 4: Fuse the extracted context features and sentiment feature region features through a cross-attention mechanism to obtain cross features after cross-attention fusion. Based on the obtained cross features and the classifier, perform sentiment classification to obtain the sentiment analysis result:
[0129] Based on the above analysis, this embodiment proposes a multi-modal cross-attention model (Multi-Cross Attentive Model, MCAM), which is a multi-modal model that can process multiple types of data (such as text, images, audio, etc.). Traditional sentiment analysis models can only use one or several of these data types for sentiment analysis and cannot fully utilize the interaction between different data types. However, MCAM can process multiple data types simultaneously and use the cross-attention mechanism to learn the interaction between different data types, thereby improving the accuracy and robustness of sentiment analysis.
[0130] The specific architecture of the model is as Figure 4 shown. First, use ALBert and DenseNet121 to obtain the vectorized features of text and images respectively; then, the text features are processed by BiLSTM to obtain context features containing sentiment words, and CBAM attention obtains the regions with the most sentiment features in the images from both spatial and channel aspects; the single-modal attention mechanism considers the relationships within the modality, and the cross-attention mechanism can consider the relationships between two modalities; then, the sentiment features of the extracted text and images are fused through the cross-attention mechanism, fully considering the complementarity and relevance between different modalities; finally, through fusion, the final sentiment analysis result is obtained through the softmax classifier.
[0131] Use the cross-attention module to model the inter-modal relationships between image regions and text words. The cross-attention mechanism is composed of scaled dot-product attention, which can fully consider the complementarity and relevance between images and texts and can improve the ability to recognize the sentiment features of image and word segments.
[0132] The cross-attention mechanism is composed of scaled dot-product attention. Take the context features as the query matrix in the scaled dot-product attention and the sentiment feature region as the key-value matrix; through the scaled dot-product attention, obtain the cross-attention features, as shown in the following formula:
[0133]
[0134]
[0135]
[0136]
[0137] Among them, W Q , W K , W V are all parameter matrices; d k is the number of columns of Q and K; I i , T i are the emotional feature region and the context feature respectively; Att(·) is the cross-attention mechanism module network.
[0138] Obtain the cross feature Y after cross-attention fusion: c = [Y c1 , Y c2 , …, Y cn .
[0139] For the output Y c after cross-attention fusion, as the input of the linear function softmax, perform the final sentiment classification, as shown in the following formula:
[0140] y = softmax(W c Y c + b c )
[0141] Among them, W c represents the weight matrix, and b c represents the bias term.
[0142] Based on the above technical solutions, the experimental process is given below. By designing comparative experiments, the performance of the MCAM model is evaluated. The model is tested on the MVSA and TumEmo datasets, and its performance evaluation is given qualitatively.
[0143] (1) Dataset introduction:
[0144] Two publicly available datasets for multimodal sentiment analysis of images and text, MVSA and TumEmo, are used. The MVSA dataset is messages containing text, pictures, etc. crawled from Twitter, also known as tweets. TumEmo is image-text sentiment data crawled from Tumblr. Tumblr, also known as Taobler in Chinese, is the world's largest microblogging website. The multimedia content posted by users on it usually contains content in the form of pictures, text, etc. These datasets are publicly available datasets in the field of multimodal sentiment analysis of images and text. Table 1 shows the information of the MVSA dataset, and Table 2 shows the information of the TumEmo dataset.
[0145] Table 1
[0146] Dataset Positive Neutral Negative Total MVSA 2683 470 1358 4511
[0147] Table 2
[0148] Dataset Angry Bored Calm Fear Happy Love Sad Total TumEmo 2167 1667 3577 1136 2476 1975 854 13852
[0149] The original data of MVSA was obtained from http: / / mcrlab.net / research / mvsa-sentiment-analysis-on-multi-view-social-data / ; the original data of TumEmo was obtained from https: / / github.com / YangXiaocui1215 / MVAN.
[0150] The MVSA dataset contains 4,869 text-image pairs. Each sample contains a set of text and images, and the sentiment labels are unimodal. There are only three modalities for the sentiment labels in the dataset: positive, neutral, and negative. The final screening results are shown in Table 1. Different from the MVSA dataset, the TumEmo dataset further subdivides the sentiment labels, specifically including seven emotion types: Angry, Bored, Calm, Happy, Love, and Sad. The TumEmo dataset contains 195,265 samples. In this embodiment, 13,852 data are obtained from it according to the proportion of each sentiment label. The TumEmo data information and content are shown in Table 2.
[0151] All image-text pairs in each dataset are divided into a training set, a test set, and a validation set, with a ratio of 6:2:2. The experimental environment of the model in this embodiment has a CPU of 12 vCPU Intel(R)Xeon(R)Platinum 8255C CPU@2.50GHz, a memory of 40GB, a GPU of RTX 3080, a Windows 10 64-bit operating system, a programming language of Python, and a version of Python 3.8 (ubuntu20.04). The deep learning-based architecture is TensorFlow 2.9.0.
[0152] (1) Model parameter settings:
[0153] For the input image part, the shape of the input image is specified as (224, 224, 3), representing the height, width, and number of channels of the image, and the batch input size is 32. In the CNN layer, the pre-trained DenseNet-121 network that achieved good results in the ImageNet2017 dataset
[33] classification challenge was used. The pooling size used in the AveragePooling2D layer is set to (7, 7), and the output dimension of the Dense layer is set to 1024 / / 8 and 1024, indicating that the dimension of the output vector is reduced to 1 / 8 of the original, that is, 128, while 1024 means keeping the dimension of the output vector unchanged. The purpose of this is to reduce the computational amount by dimensionality reduction while retaining the original feature information.
[0154] For the input text part, the ALBert pre-trained model is adopted to initialize the word embedding layer of the text. The dimension of its hidden layer is set to 128, that is, each word is represented by a 128-dimensional vector. The maximum input length of the text is obtained based on the analysis of the dataset, as Figure 5 shown. It is found that most of the text lengths are below 150, and the number of texts with a length below 150 is also the largest. Therefore, 150 is selected as the maximum length of the input text. For texts with an input length greater than 150, the text will be truncated, and for texts with an input length less than 150, zero-value padding will be performed.
[0155] For the multi-modal fusion part, the cross-attention mechanism is used. During the model training process, the learning rate is set to 0.01, and the Adam optimizer is used to optimize the model parameters; to prevent overfitting, a random dropout rate (Dropout Value) of 0.1 is set in the model, and early stopping technology and L2 regularization are adopted. Ten-fold cross-validation and grid search are used to verify different parameter combinations. The hyperparameter combination of the final model is shown in Table 3.
[0156] Table 3
[0157]
[0158]
[0159] (3) Evaluation metrics:
[0160] To evaluate the performance of the model in this embodiment, the commonly used ones in the classification task: accuracy, precision, recall, and F1-score are selected as evaluation metrics. The classification results are: true positive (TP), false positive (FP), false negative (FN), and true negative (TN).
[0161] Accuracy: It reflects the judgment ability of the model for the entire dataset, that is, the proportion of the correct prediction quantity in the positive and negative examples to the total quantity. The formula is as follows:
[0162]
[0163] Precision: Based on the prediction result as the judgment basis, the proportion of the correctly predicted samples among the samples predicted as positive examples. The formula is as follows:
[0164]
[0165] Recall: Based on the actual samples as the judgment basis, among the samples that are actually positive examples, the proportion of the correctly predicted positive examples to the total actual positive example samples. The formula is as follows:
[0166]
[0167] F1 Score: The F1 score takes into account both the precision and recall of a classification model and can be regarded as a weighted average of the model's precision and recall, with a value range of [0, 1]. The formula is as follows:
[0168]
[0169] (4) Baseline Method:
[0170] To evaluate the robustness and generalization ability of the model in this embodiment, the following comparative experiments are set up. The comparison methods are introduced as follows:
[0171] 1) Single-modal text model: Use the contextual word representations from the pre-trained language model BERT and the fine-tuning method with additional generated text for sentiment analysis; propose a sentiment analysis method based on BiGRU information enhancement.
[0172] 2) Single-modal image model: Propose a method based on the ResNet50 network for image recognition and classification; propose a method using the VGG19 model for classifying crop leaf diseases.
[0173] 3) Multi-modal text-image fusion model: Propose the Deep Multi-modal Attention Fusion (DMAF) model to perform joint sentiment classification using the internal correlation between visual and text features; propose a new Image-Text Interaction Network (ITIN) model to study the relationship between sentiment image regions and text for multi-modal sentiment analysis; propose a new multi-modal sentiment analysis model based on the Multi-View Attention Network (MVAN), which uses an updated memory network to obtain the deep semantic features of image text; propose the Attention-based Modal Gating Network (AMGN) model, which utilizes the correlation between image and text modalities and extracts discriminative features for multi-modal sentiment analysis for sentiment analysis.
[0174] (5) Experimental Results:
[0175] To evaluate the performance of the MCAM model, the following comparative experiments are set up. Comparisons are made respectively from: single-modal text models, single-modal image models, and multi-modal text-image fusion models, as shown in Tables 4 and 5. Table 4 is the performance comparison table of different models on the MVSA dataset. Table 5 is the performance comparison table of different models on the TumEmo dataset.
[0176] Table 4
[0177]
[0178] Table 5
[0179]
[0180] (5-1) Experimental results of the MVSA dataset:
[0181] As can be seen from Table 5 and Figure 6 analysis, the MCAM model proposed in this embodiment achieved the best results on the MVSA dataset. For single-modal image and text data, their performance in sentiment classification is relatively low, and the classification effect is average. For multi-modal sentiment analysis, by learning the correlation between the two modalities, the classification effect has been greatly improved immediately. In particular, the MCAM model proposed in this embodiment, by introducing the cross-attention mechanism to learn the correlation and complementarity between different modalities, has increased the prediction accuracy by 11.7% and the F1 score by 9.8% compared with the MVAN model, and increased the accuracy by 2.4% and the F1 score by 1.7% compared with the DMAF model, obtaining the best prediction effect, indicating that the introduction of cross-attention plays a great role in improving the model performance.
[0182] In addition, the single-modal picture and single-modal text models are also better than the baseline methods. In the text model, not only the more effective pre-trained model ALBert is used, but also the bidirectional long short-term memory network (BiLSTM) is introduced, which can further improve the prediction effect; in the picture model, the DenseNet121 network is used. The dense connection structure of DenseNet121 enables better feature transmission, smoother gradient flow, and lower overfitting risk. At the same time, with fewer parameters, it can achieve performance comparable to ResNet50 and VGG19, and by introducing the CBAM attention mechanism to extract attention from both the channel and spatial aspects, the classification effect is further improved.
[0183] To further prove the effectiveness of the MCAM model, different proportions of data from 20% to 100% are randomly selected from the MVSA dataset and the TumEmo dataset, and then the changes in the accuracy of the four baseline models and the MCAM model are observed. From Figure 7 it can be seen that no matter what proportion of training data is extracted, the accuracy of the model is always better than the baseline models, showing an absolute competitive advantage, and further indicating that the model can also achieve good results in the case of insufficient training data.
[0184] (5-2) Experimental results of the TumEmo dataset:
[0185] As can be seen from Table 5 and Figure 6Analysis shows that the performance of the TumEmo dataset has decreased significantly compared to that of the MVSA dataset. The main reason for the analysis is that the emotion categories in the TumEmo dataset are more diverse, reaching 7, resulting in a decrease in model performance. In multi-class classification problems, as the number of categories increases, the classifier needs to distinguish more categories, thus increasing the classification difficulty, and the corresponding classification effect may decline. However, the model still has obvious competitive advantages and has achieved the best results in all evaluation metrics. Although the improvement effect is not very obvious compared with AMGN and DMAF, there is still a slight improvement effect. Compared with AMGN, the accuracy has increased by 1.5% and the F1 score has increased by 1.9%. Compared with DMAF, the accuracy has increased by 1.4% and the F1 score has increased by 1.6%. This also proves once again that the introduction of the cross-attention mechanism plays a certain role in learning the complementarity and correlation between different modalities.
[0186] (5-3) Analysis of the PR curve results of different models on different datasets:
[0187] As Figure 8 shown, it can be seen from the left subgraph that the PR curves of the TumEmo dataset show various different shapes, indicating that different models have different performance when classifying the TumEmo dataset. Among them, the PR curve of the model MCAM shows the optimal shape, indicating that this model has the best performance when classifying the TumEmo dataset. While the PR curves of other models such as ITIN and MVAN show relatively poor shapes, indicating that their performance is relatively poor.
[0188] It can be seen from the right subgraph that the PR curves of the MVSA dataset show different shapes from those of the TumEmo dataset. The PR curve of the model MCAM still shows the optimal shape, while the PR curves of other models such as ITIN and MVAN show relatively poor shapes, indicating that their performance is relatively poor. It is worth noting that the performance of the model DMAF on the MVSA dataset is relatively poor, which is the opposite of its performance on the TumEmo dataset.
[0189] By placing the PR curves of the two datasets on the same graph, the performance of different models in the two datasets can be more intuitively compared. It can be seen that the model MCAM shows the best performance in both datasets, while the performance of other models varies. In addition, it can also be found that there are obvious differences in the shapes of the PR curves of the TumEmo dataset and the MVSA dataset, indicating that the differences in model performance between different datasets may be relatively large, and appropriate datasets need to be selected for specific tasks.
[0190] (6) Ablation experiment:
[0191] To further verify the performance of the MCAM model, the following ablation experiments were set up: Table 6 is the comparison table of ablation experiments.
[0192] Table 6
[0193]
[0194] From the performance analysis of Table 6 and Figure 9 the ablation experiments, it can be seen that when there is only an image or text, the performance of the model on different datasets is average. This is mainly because unimodal sentiment analysis only learns within the modality. Sometimes the image may play a dominant role in classifying the sentiment category, and sometimes the text may play a dominant role. Therefore, unimodal sentiment analysis cannot fully consider the complementarity and correlation between different modalities, ultimately resulting in poor classification performance.
[0195] However, when the cross-attention mechanism is introduced, the performance of the model is greatly improved. Specifically, in the cross-attention mechanism, the model first uses independent networks to learn the feature representations of each modality, and then uses the cross-modal attention mechanism to capture the relationships between different modalities. This is achieved by calculating the similarity between different modalities, and then these similarities are used as the weights of the cross-modal attention mechanism to weighted fuse the feature representations of different modalities. Finally, the model uses these fused feature representations to perform sentiment analysis predictions, improving the performance and robustness of the model.
[0196] (7) Visual analysis of attention weights:
[0197] In this section, a qualitative analysis of the sentiment analysis of the fusion of images and text is carried out, focusing on the attention score before and after the introduction of cross-attention. The sentiment score of attention is any value between 0 and 1. Through visual analysis, it can be clearly seen the changes in the heat map effect of the image area before and after the introduction of attention.
[0198] (7-1) Visual attention processing process:
[0199] For image data, the CBAM enhanced convolutional neural network is used to improve the ability to extract image features. First, the input image tensor (224, 224, 3) is subjected to average pooling and max pooling operations to calculate the average value and maximum value of the channel dimension; then these two results are compressed into vectors with dimensions of 1024 / 8 = 128 and 1024 through two fully connected layers (Dense), and then restored to the original dimension through another fully connected layer. These two results are added together, and then a weight between 0 and 1 is obtained through the sigmoid function.
[0200] Next, for the tensor weighted by the channel attention mechanism, perform the attention mechanism in the spatial dimension. First, calculate the average and maximum values in the spatial dimension of the tensor weighted in the channel dimension, and then concatenate them in the channel dimension. Input the concatenated tensor into a 1×1 convolutional layer to obtain a 1×1×1 feature map, and then obtain a weight from 0 to 1 through the sigmoid function.
[0201] Finally, multiply the weights obtained from the channel attention mechanism and the spatial attention mechanism to obtain the weighted attention weight, which is the final attention score.
[0202] For the image processed by the CBAM attention mechanism, a heat map is drawn to make its color more distinct. The transparency of the original image is set to 0.5, and the transparency of the processed image is set to 0.8 to highlight the areas that the attention focuses on more. The higher the regional attention score, the redder the color of that area.
[0203] (7-2) Text attention processing process:
[0204] For text data, first tokenize the input text data to obtain a token sequence, convert the token sequence into an input tensor that can be processed by the ALBert model, and use the attention mask tensor to represent the position of each token.
[0205] Next, input the input tensor and the attention mask into the ALBert model. The model will encode the input tokens to obtain an encoded tensor. For each self-attention layer, the model will divide the encoded tensor into multiple heads. Each head calculates the query, key, and value respectively, and uses this information to calculate the attention score between each token and other tokens.
[0206] Finally, through weighted averaging of the attention scores obtained from each attention head, a weight vector is obtained, which represents the importance of each token in the entire text data, that is, the attention weight.
[0207] For the text processed by self-attention, color annotation is performed on the emotional words that the attention focuses on. To distinguish the contribution of different words to the attention score, annotations of the same color but different depths are made. The darker the color, the greater the attention weight of the word.
[0208] (7-3) Cross-attention fusion multimodal processing process:
[0209] For multi-modal data that combines images and text, first, image features and text features are respectively passed into two fully connected layers, and their outputs are passed as inputs into a dot product layer to calculate the attention weights.
[0210] Next, the attention weights are respectively multiplied by the image features and text features to obtain the attention tensors of the image and text.
[0211] Finally, the Concatenate layer is used to concatenate the attention tensors of the image and text along the last dimension to obtain the fused feature tensor, and the attention score of the fused feature tensor is calculated, which is the cross-attention score of the fused multi-modal.
[0212] For the image and text after cross-attention fusion, it is obvious that the area of the image being attended to is further enlarged, the red color is further deepened, and the color of the words being attended to by the attention is also deepened. The attention scores have all increased to varying degrees, and the attention score after fusion has also been further improved compared to the score of simple vector concatenation. Therefore, introducing the cross-attention mechanism has a good fusion effect on learning the complementarity and correlation between different modalities.
[0213] In this embodiment, a multi-modal sentiment analysis method based on the cross-attention mechanism (MCAM) is proposed. This model can adaptively calculate the correlation between images and text, thereby better fusing their features. First, two different processing methods are proposed for single modalities respectively to learn text and image features; then, the cross-attention mechanism is used to adaptively fuse the features of different modalities, thereby improving the performance of sentiment analysis. Finally, the sentiment result is output through a classifier. Through experimental analysis, the results show that the method provided in this embodiment obtains the best results on two publicly available multi-modal sentiment analysis datasets, outperforming the four proposed baseline methods.
Claims
1. A sentiment analysis method based on text-image fusion with a multi-modal cross-attention mechanism, characterized in that Including: Obtain the text and image to be processed; Vectorize the text to obtain text features; Process the text features through BiLSTM to obtain context features containing sentiment words; Obtain the image features of the image, and use CBAM attention to obtain the feature region features of the sentiment feature in the image features from two aspects of space and channel respectively; For the extracted context features and sentiment feature region features, fuse them through a cross-attention mechanism to obtain cross features after cross-attention fusion. According to the obtained cross features and the classifier, perform sentiment classification to obtain the sentiment analysis result; The obtaining of the image features of the image includes: Use the pre-trained DenseNet121 network model to obtain the image features of the image; Setting: The feature map obtained after the i-th image passes through the j-th convolutional layer in the DenseNet121 network model is where C is the number of channels, H is the length of the feature map, and W is the width of the feature map; The using of CBAM attention to obtain the feature region features of the sentiment feature in the image features from two aspects of space and channel respectively includes: Obtain the attention weight of the j-th feature map of the i-th image using CBAM attention The calculation method is as follows: Among them, M c (F ij ) is the channel attention weight of the j-th feature map of the i-th picture; M s (F ij ) is the spatial attention weight of the j-th feature map of the i-th picture; Channel attention weight M c (F ij ) is calculated as follows: Where σ is the Sigmoid activation function; the weights W0 and W1 of the MLP are shared by two inputs, and ReLu activation function is followed by W0; AvgPool(·) is the global average pooling function, which calculates the average value of all feature regions of each feature map; MaxPool(·) is the global maximum pooling function, which calculates the maximum eigenvalue of each feature map; Spatial attention weight M s (F ij ) is calculated as follows: M s (F ij ) = σ(f 7×7 ([AvgPool(F ij ), MaxPool(F ij )])) where [·] is the concatenation operation; f 7×7 (·) represents the convolution operation with a convolution kernel size of 7×7; AvgPool(·) is the average pooling function, and MaxPool(·) is the max pooling function; The calculation formula of the CBAM attention feature map is as follows: Among them, means bitwise multiplication of corresponding elements; After the convolution operation, the feature mapping vector I = [I1, I2, …, I n of the key region of each image is obtained.
2. The method for sentiment analysis based on multimodal cross-attention mechanism graphic-text fusion according to claim 1, wherein The vectorizing of the text to obtain text features includes: using the ALBert pre-trained model to vectorize the text to obtain text features.
3. The sentiment analysis method based on multi-modal cross-attention mechanism for image-text fusion according to claim 2, wherein The ALBert pre-trained model includes an input layer, an ALBert pre-training layer, and a Transformer encoder; The using of the ALBert pre-trained model to vectorize the text to obtain text features includes: The text passes through the input layer to the ALBert pre-training layer, and the ALBert pre-training layer is used to obtain the numbers of each character in the text and vectorize them to obtain an information sequence; The information sequence is input into the Transformer encoder to mine deep semantic feature information, and is transformed to obtain text features.
4. The method for sentiment analysis based on multi-modal cross-attention mechanism for image-text fusion according to claim 1, wherein The fusing of the extracted context features and sentiment feature regions through a cross-attention mechanism to obtain cross features after cross-attention fusion. According to the obtained cross features and the classifier, perform sentiment classification to obtain the sentiment analysis result includes: The cross-attention mechanism is composed of scaled dot-product attention. The context features are used as the query matrix in the scaled dot-product attention, and the sentiment feature region is used as the key-value matrix; through the scaled dot-product attention, the cross-attention feature is obtained, as shown in the following formula: Among them, W Q , W K , W V are all parameter matrices; d k is the number of columns of Q and K; I i , T i are the emotional feature region and the context feature respectively; Att(·) is the cross-attention mechanism module network; Obtain the cross feature Y after cross-attention fusion: Y c = [Y c1 , Y c2 , …, Y cn ; For the output Y after cross-attention fusion c , as the input of the linear function softmax, perform the final sentiment classification as follows: y = softmax(W c Y c + b c ) Among them, W c represents the weight matrix, and b c represents the bias term.
Citation Information
Patent Citations
Dual-mode sentiment analysis method based on attention mechanism
CN112860888A
Postoperative cataract vision prediction system based on multi-modal fusion network
CN114782394A