Multi-modal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement
By extracting multi-scale text visual features in multimodal aspect-level sentiment analysis and adopting a dynamic attention pooling mechanism, the problem of insufficient image representation perception in the prior art is solved, and more efficient multimodal information integration and sentiment analysis accuracy is achieved.
Patent Information
- Application Number
- CN202510197147.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The existing multimodal aspect-level sentiment analysis methods fail to effectively perceive image representation from the semantic level of the image, resulting in visual attention being unable to cover the corresponding visual representation of the target and being unable to effectively associate the target with image information.
A multimodal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement is proposed. By extracting multi-scale visual features and text and aspects, filtering noise with dynamic attention pooling mechanism, and integrating multimodal information to improve model performance.
Through the multi-scale text visual feature enhancement and dynamic attention pooling mechanism, the accuracy and performance of multimodal aspect-level sentiment analysis models are effectively improved, ensuring the quality of multimodal fusion features and the generalization ability of the model.
Smart Images

Figure CN120123978A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing technology, computer vision and multimodal sentiment analysis, and in particular to a multimodal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement. Background Art
[0002] With the development of intelligent new media technology, the number of Internet users has increased dramatically, and social networks generate massive amounts of social graph and text data every day. By analyzing these social texts containing images, the application of multimodal aspect-level sentiment analysis methods can help decision-making, public opinion response, etc. The aspect-based multimodal sentiment analysis task, as a multimodal sentiment classification task, aims to calculate the sentiment polarity of specific aspects in the social graphs and texts posted by users. Compared with text aspect-level sentiment analysis, data with image content can help improve the accuracy of sentiment discrimination. In the multimodal aspect-level sentiment analysis method, how to align the image areas involving aspect items in the text and effectively use image modal data to make the image and text content complementary is crucial to improving the performance of the multimodal aspect-level sentiment analysis model. For example, the user mentioned three aspect entities in the text but there are only two salient target entities in the image. When the number of aspect items in the text is greater than the number of salient target entities in the image, it is necessary to extract fine-grained visual semantic information to align the image target-level entities with the aspect items.
[0003] In the scenario of fine-grained emotion recognition on multimodal data, the semantics of the input image-text pair are complex, the text modality contains limited emotional information, but the visual content is associated with one or more targets in the text, so different targets and image emotional features are easily confused when aligned. The mainstream models mainly use memory neural networks, graph convolutional neural networks, attention mechanisms and pre-trained language models to perform multimodal aspect-level sentiment analysis.
[0004] In the paper "Exploiting BERT for MultimodalTarget Sentiment Classification through Input Space Translation" in the proceedings of the 29th ACM International Conference on Multimedia, Yu et al. designed an auxiliary reconstruction module based on the idea of automatic encoding, reconstructed the intermediate layer input of text and image through a shared decoder, and narrowed the data difference between text and image.
[0005] The literature "Sentiment-aware multimodal pre-training for multimodal sentiment analysis" in Knowledge-Based Systems, Volume 258 in 2022, proposed a sentiment-aware multimodal pre-training (SMP) framework for multimodal sentiment analysis. Ye et al. adopted a cross-modal contrastive learning method, introduced fine-grained sentiment labels, and learned rich sentiment information from multimodal data.
[0006] The literature "AMIFN: Aspect-guided multi-view interactions and fusion network for multimodal aspect-based sentiment analysis" in Neurocomputing, Volume 573 in 2024, proposed an aspect-guided multi-view interaction and fusion network (AMIFN) for multimodal aspect-based sentiment analysis. Yang et al. used the RoBerta pre-training model to extract hidden feature representations from sentences and targets, and adopted a multi-head attention mechanism to perform different interactions of aspect images to extract aspect-related visual information.
[0007] However, none of the above methods can perceive the image representation from the local target semantic level of the image. They only implicitly align the image regions with the target aspect words, and the granularity of the aspect terms in the two modalities is inconsistent, which results in the visual attention sometimes unable to cover the corresponding visual representation of the target aspect and cannot effectively associate the target and image information. Summary of the Invention
[0008] The purpose of the present invention is to provide a multimodal aspect-level sentiment analysis method with enhanced multi-scale text-visual features, which expands the feature information to enable effective interaction between text and images, and reduces the noise generated by text and picture data to ensure the quality of multimodal fusion features. First, multi-scale visual feature information is extracted and interacted with text and aspects in multiple dimensions, making full use of the fine-grained information in the visual modality to further improve the accuracy of sentiment prediction. Second, a dynamic attention pooling mechanism is proposed in the multimodal data fusion stage to filter noise and efficiently integrate multimodal information to improve the performance of the model.
[0009] The above technical purpose of the present invention is achieved through the following technical solutions: A multimodal aspect-level sentiment analysis method based on enhanced multi-scale text-visual features, comprising the following steps:
[0010] Step 1: Obtain a multi-modal aspect-level sentiment dataset, which contains a piece of text information, a sequence of aspect words, and an associated image;
[0011] Step 2: Extract the scene semantics of the associated image, fuse the image scene semantics with the text information, and input it into a pre-trained language model to obtain global context features;
[0012] Step 3: Extract the object semantics of the associated image, divide the image object semantics into object noun semantics and object emotion semantics according to different part-of-speech features, align the aspect word sequence with the object noun semantics based on a dynamic gating mechanism, introduce a multi-head cross-attention layer to obtain a visual information representation related to the aspect, and combine the object emotion semantics to obtain a visual emotion feature related to the aspect;
[0013] Step 4: Obtain the dependency relationship between words in the text information according to dependency parsing, construct an adjacency matrix by combining the word association weights, introduce a graph convolutional neural network to capture context syntactic information and long-distance word dependencies, and obtain syntactic-enhanced text features;
[0014] Step 5: Fuse the global context features, syntactic-enhanced text features, and visual emotion features related to the aspect, and perform noise reduction on the multi-scale text visual features based on a pooling mechanism to obtain multi-modal fusion features;
[0015] Step 6: Output the multi-modal fusion features to an emotion prediction layer to predict the emotion polarity of the aspect words.
[0016] Preferably, the specific content of Step 2 is as follows:
[0017] Step 2-1: Input the image into the Caption Transformer encoding framework to generate an image caption as the image scene semantics D i ;
[0018] Step 2-2: Tokenize the text information S i , the aspect word sequence T i and the scene description D obtained in Step 2-1 i respectively, and use [SEP] to separate the internal text sequences, and [CLS] as the global context start marker to obtain the integrated global context sequence C : i :
[0019]
[0020] Step 2-3: Calculate the global context embedding, and use the Bert-based-uncased pre-trained language model to process Ci 映Projected into a continuous vector space to obtain word vector embedding SD i and segmentation embedding SE i and position embedding PE i . Add SD i and SE i to obtain the global context embedding sequence
[0021]
[0022] Step 2-4: Calculate the attention mask. Through tensor expansion of the position embedding PE described in Step 2-3 i , obtain the attention mask MA i through masked_softmax:
[0023] MA i = Masked_Softmax(expand(PE i ));
[0024] Step 2-5: Input the attention mask MA described in Step 2-4 i and the global context embedding sequence described in Step 2-3 into the multi-layer self-attention encoder of the Transformer to obtain the global context features where k is the maximum length of the global context and d is the dimension of the hidden layer:
[0025]
[0026] Preferably, the specific content of Step 3 is as follows:
[0027] Step 3-1: Use SentiBank to extract the adjective pairs ANP = (A k , N k ) contained in the image. Select the top 5 ANPs with the highest confidence for each image to reduce the introduction of irrelevant noise as the image target semantics, and divide the image target semantics into target emotion semantics A k and target noun semantics N k ;
[0028] Step 3-2: Input the aspect word sequence described in Step 1, the target emotion semantics A k and the target noun semantics N k described in Step 3-1 into the Bert pre-trained language model to obtain the aspect word encoding where k t is the maximum length of the aspect word, d t is the dimension of the aspect word hidden layer, the target emotion representation the target noun representation s represents the length of the maximum adjective or noun sequence;
[0029] Step 3-3: Construct a dynamic gating adjustment vector, calculate the semantic similarity between the aspect term representation and the noun representation in the high-dimensional vector space, and construct the gating adjustment vector G A :
[0030] G A = expand(cos<H T , H N >);
[0031] The gating adjustment vector G A is multiplied by the target noun representation H described in Step 3-2 to obtain the aggregated result after constraining the target noun representation and the aspect representation H N to obtain T the gated adjusted aspect representation:
[0032]
[0033] Step 3-4: Obtain the aspect-related visual features. Input the associated image described in Step 1 into the ResNet-152 image pre-training model to obtain the image encoding P. Use the gated adjusted aspect representation described in Step 3-3 as the Query, and the image encoding representation P as the Value and Key and input them into the multi-head cross-attention to obtain the aspect-related visual attention A T→V :
[0034]
[0035] CATT i is the cross-attention of the i-th head, σ is the Gaussian error linear activation function Gelu, is the learnable parameter matrix, and m is the number of attention heads of the multi-head attention;
[0036] Considering the difference in the numerical range of different pixel points in the image data, add the Gaussian error linear activation unit Gelu to ensure the generalization of the neural network in the random regularization process, reduce the variable shift, and obtain the aspect-related visual representation H T→V :
[0037]
[0038] Step 3-5: Incorporate the target emotion information to obtain the fine-grained visual emotion representation. By analogy with the embedding method of the target noun representation in the aspect term described in Step 3-3, use the gating vector G described in Step 3-3 A to multiply with the adjective representation H AMultiply to obtain the visual emotional representation
[0039]
[0040] Perform attention pooling on the visual representation H related to the aspects described in steps 3 - 4 T→V and the visual emotional representation respectively, so that the model can more effectively process and understand the visual information related to the aspects, and then perform tensor dimension expansion:
[0041]
[0042] Steps 3 - 6: Integrate the and described in steps 3 - 5 to obtain the visual emotional feature H related to the aspect T→VA :
[0043]
[0044] Preferably, the specific steps of step 4 are as follows:
[0045] Step 4 - 1: Construct a dependency matrix, and obtain the dependency matrix between words by means of spacy dependency parsing
[0046] Step 4 - 2: Calculate the word lemma representation, and input the text information described in step 1 into the BERT encoder to obtain the initial text representation n t represents the total number of word lemmas in the sentence after word segmentation, and integrate the word lemmas belonging to the same word as the word representation:
[0047]
[0048] represents the k - th word representation obtained after integrating the corresponding word lemmas, n represents the number of words in the sentence;
[0049] Step 4 - 3: Construct an adjacency matrix, consider the matching between the syntactic parsing tree and the word - to - word dependency relationship, and calculate the relevance between each pair of words in the text as the relationship weight matrix
[0050]
[0051] The relationship weight matrix Q relevance is combined with the D described in step 4 - 1 ij to establish the adjacency matrix A ij ,
[0052] A ij = Dij Q relevance ;
[0053] Step 4-4: Obtain a syntax-enhanced text representation based on the graph convolutional neural network layer, and input the adjacency matrix A described in Step 4-3 ij and the word token representation described in Step 4-2 into the graph convolutional network layer H S , to capture context syntax information and long-distance word dependencies:
[0054]
[0055] represents the hidden layer representation of the l-th layer of the i-th node, is the adjacent node of node i, W l , b l are learnable parameters, and σ is the activation function ReLU;
[0056] Obtain the output of the last layer of the graph convolutional layer as the syntax-enhanced text representation;
[0057] Step 4-5: Integrate the global context representation H described in Step 2-5 C and the syntax-enhanced text representation H described in Step 4-4 S as the text feature H CS :
[0058] H CS = Concat(H C , H S ).
[0059] Preferably, Step 5 is specifically:
[0060] Step 5-1: Input the text feature H described in Step 4-5 CS and the visual emotional feature H related to the aspect described in Step 3-4 T→VA into the dynamic attention pooling layer to obtain a multi-modal fusion feature encoding H m :
[0061] H m(l) = ComATT(concat(H T→AV , H CS ))
[0062]
[0063] ComATT represents joint self-attention, and the multi-layer feature representation is where H m(l) is the feature of the l-th layer, and L is the number of layers of the dynamic pooling attention mechanism of the multi-modal fusion module.
[0064] Preferably, step 6 is specifically as follows:
[0065] Step 6-1: Put the multi-modal fusion features described in step 5-1 into the softmax layer for sentiment classification:
[0066]
[0067] Step 6-2: Calculate the mean square error to calculate the reconstruction loss. The multi-modal aspect-level sentiment dataset is the training sample set D. For a training sample set D, the first part of the loss comes from the mean square error loss L generated when the aspect words interact with the vision in step 3-4 r :
[0068]
[0069] Step 6-3: Calculate the classification loss of the multi-modal fusion features. For a training sample set D, the second part of the loss comes from the multi-modal aspect-level sentiment analysis method described in claim 1. The cross entropy is used to quantify and calculate the classification loss L of the multi-modal fusion features m :
[0070]
[0071] where |y| is the number of sentiment classification categories, which is 3, and t ij is the true label;
[0072] Step 6-4: Calculate the gradient loss of the model backpropagation. The proportion of the mean square error loss L described in step 6-2 in the model training loss is controlled by the hyperparameter λ, and the loss value L of the model backpropagation training in the gradient descent process is obtained: r In the model training loss, the proportion is obtained, and the loss value L of the model backpropagation training in the gradient descent process is obtained:
[0073] L = L m + λL r .
[0074] In summary, the present invention has the following beneficial effects:
[0075] The present invention proposes a multi-modal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement. The image semantics are mined from two levels of scene and target. The scene semantic information is combined with the text content to complement and enrich the global context information representation, and the fine-grained target semantic information is incorporated into the visual information representation related to the aspect.
[0076] The present invention constructs a syntactic graph according to the text information, and combines the adjacent nodes in the graph convolutional neural network to capture the long-distance word dependencies, enhancing the perception of the context syntactic information related to the aspect.
[0077] To capture the correlations between different modalities, in the multi-modal feature fusion stage, a dynamic attention pooling strategy is adopted to fuse multi-image and text features. Considering the inconsistent learning efficiencies of different modalities, a dynamic gating reconstruction loss function is designed to improve the loss calculation in real time and optimize the learning path to reduce the risk of model overfitting. Description of the Drawings
[0078] Figure 1 It is a flowchart of the multi-modal aspect-level sentiment analysis method based on multi-scale text-visual feature enhancement according to the present invention;
[0079] Figure 2 It is an overall framework diagram of the multi-modal aspect-level sentiment analysis method based on multi-scale text-visual feature enhancement according to the present invention;
[0080] Figure 3 It is a schematic diagram of the framework of the visual feature extraction module related to the present invention;
[0081] Figure 4 They are three randomly selected cases of the present invention;
[0082] Figure 5 It is an aspect-guided visual perception attention heat map in the visual information extraction module related to the present invention;
[0083] Figure 6 It is a syntactic dependency relationship information diagram of the present invention;
[0084] Figure 7 It is a syntactic feature weight attention matrix diagram of text graph convolution of the present invention. Detailed Embodiment
[0085] The present invention will be further described in detail below with reference to the drawings.
[0086] This specific embodiment is only an interpretation of the present invention and is not a limitation thereof. Those skilled in the art can make modifications without creative contributions to this embodiment according to needs after reading this specification, but as long as they are within the scope of the claims of the present invention, they are protected by the patent law.
[0087] As Figure 1 shown, the multi-scale text-visual feature enhancement network provided by the present invention is specifically as follows:
[0088] First, to solve the problem that text and image data cannot effectively interact under inconsistent feature spaces. Based on the CaptionTransformerBERT image captioning task, the image scene semantics and text content are complementary. The global image representation is perceived from the semantic level of the image, and the image scene information is fused with multi-level text information to enrich the context information perception ability of the text.
[0089] Secondly, to accurately extract visual features related to aspect terms, the present invention proposes a dynamic cross-attention gating mechanism to extract visual features under the guidance of aspects. The aspect words and the semantics of image targets are shown to be aligned, and fine-grained visual target information is incorporated to achieve cross-modal interaction between the image and the aspect words.
[0090] In addition, to capture long-distance word dependencies, an adjacency matrix is constructed, and the input of word tokens is integrated into a graph convolutional layer to enhance the representation of context syntactic information and obtain syntactic-enhanced text information features.
[0091] Finally, to strengthen the interaction effect of cross-modal information and enable the model to better understand and fuse complex multi-modal data, self-attention pooling is used to filter out the noise generated in the fusion process, efficiently integrate multi-modal information to capture the correlation between different modalities, and reduce the risk of model overfitting.
[0092] As Figure 2 shown, the multi-scale text-visual feature enhancement network provided by the present invention is specifically as follows:
[0093] S1. Multi-modal data acquisition and preprocessing module. The public datasets of Twitter-2015 and Twitter-2017 are acquired. After preprocessing, the text information, aspect word sequences, and associated images are input into a pre-trained model to obtain initial encodings.
[0094] S11. Given a set of multi-modal samples M, each sample m i ∈M, each multi-modal data contains a piece of text, an associated image, and aspect words, expressed as: m i =(S i , P i , T i ). Each sentence contains n words S i =(w 1 , w 2 ,..., w n ), and l aspect terms T i =(t 1 , t 2 ,..., t l ).
[0095] S12. Preprocess the text information, mask the aspect word part in the original text content to obtain the text information S i , and use the Bert-based-cased pre-trained model to encode the text information to obtain the initial text information representation n t represents the total number of word tokens in the sentence after word segmentation.
[0096] S13. Preprocess the associated images, standardize the three-channel feature maps of the images, map the values of each channel to the interval [-1, 1], and mark it as P.
[0097] S2. Global context information extraction module. Concatenate the image scene semantic information, text information, and aspect word sequence, input them into the Bert pre-trained model, fuse the image scene semantic information with the multi-level text information, enrich the context information perception ability of the text, and obtain the global context features.
[0098] S21. Input the image into the Caption Transformer encoding framework to generate an image caption as the image scene semantics D i ;
[0099] S22. Tokenize the text information S i , aspect word sequence T i and the scene description D i respectively to obtain Use [SEP] to separate the internal text sequences, and [CLS] as the global text start marker to obtain the integrated global text sequence. Among them, C i represents the global text sequence:
[0100]
[0101] S23. Calculate the global text embedding. Use the Bert-based-uncased pre-trained language model to map C i , TD i , PD i into a continuous vector space to obtain the word vector embedding SD i , segment embedding SE i , position embedding PE i . Add SD i and SE i to obtain the global text embedding sequence
[0102]
[0103] S24. Calculate the attention mask. Obtain the attention mask MA i through tensor expansion masked_softmax on the step position embedding PE i .
[0104] S25. Input the attention mask MAi and the global text embedding sequence into the multi-layer self-attention encoder of the Transformer to obtain the global text representation k is the maximum length of the global text, and d is the dimension of the hidden layer:
[0105]
[0106] S3. Aspect-related visual information extraction module, which aligns the aspect word sequence with the semantic meaning of the target noun based on the dynamic gating mechanism, extracts aspect-related visual features through aspect-guided use of multi-head cross-attention, further integrates the target emotion semantics into the aspect-related visual features, enhances the perception of aspect-related visual features, and obtains aspect-related visual emotion features. Obtain aspect-related visual features (as Figure 3 shown).
[0107] S31. Use SentiBank to extract the adjective pairs ANP=(A k , N k ) contained in the image. Select the top 5 ANPs with the highest confidence in each image to reduce the introduction of irrelevant noise, and use them as the image target semantics. Divide the image target semantics into target emotion semantics A k , target noun semantics N k .
[0108] S32. Input the target emotion semantics A k and the target noun semantics N k into the Bert pre-trained language model to obtain the target emotion representation
[0109] target noun representation where s represents the length of the maximum adjective (or noun) sequence.
[0110] S33. Construct a dynamic gating adjustment vector. Calculate the semantic similarity between the aspect word representation and the noun representation in the high-dimensional vector space, and construct the gating adjustment vector G A :
[0111] G A = expand(cos<H T , H N >).
[0112] The gating adjustment vector G A is multiplied by the target noun representation H N to obtain the aggregation of the constrained target noun representation and the aspect representation H T to obtain the gated adjusted aspect representation
[0113]
[0114] S34. Obtain visual features related to the acquisition aspect. Characterize the gating adjustment aspect As Query, represent the image encoding as P as Value and Key and input them into the multi-head cross-attention to obtain the acquisition aspect-related visual attention A T→V :
[0115]
[0116] CATT i Is the cross-attention of the i-th head, σ is the Gaussian error linear activation function Gelu, Is the learnable parameter matrix, m is the number of attention heads of the multi-head attention.
[0117] Considering the differences in the numerical ranges of different pixel points in the image data, add the Gaussian error linear activation unit Gelu to ensure the generalization of the neural network in the random regularization process, reduce variable offset, and obtain the acquisition aspect-related visual representation H T→V :
[0118]
[0119] S35. Incorporate target emotion information to obtain fine-grained visual emotion representation. Analogy the embedding method of the target noun representation in the aspect words. Use the gating vector G A Multiply with the adjective representation H A To obtain the visual emotion representation
[0120]
[0121] For the acquisition aspect-related visual representation H T→V And the visual emotion representation Perform attention pooling respectively so that the model can more effectively process and understand the acquisition aspect-related visual information, and then perform tensor dimension expansion:
[0122]
[0123] S36. The And Obtain the acquisition aspect-related visual emotion feature H T→VA :
[0124]
[0125] S4. Text syntactic feature enhancement module. It uses dependency parsing to obtain the dependency relationships between information words in the text and constructs a dependency relationship matrix. It constructs a word association degree matrix according to semantic similarity. Multiply the word association degree matrix by the dependency relationship matrix to obtain an adjacency matrix. At the same time, integrate the tokens belonging to the same word in the text representation as the word representation. Input the word representation and the adjacency matrix into a graph convolutional network to enhance the context syntactic information representation and capture long-distance word dependencies, and obtain the syntactic enhanced text information features.
[0126] S41. Construct a dependency relationship matrix. Use spacy dependency parsing to obtain the dependency relationship matrix between words
[0127] S42. Calculate the word token representation. Input the text information into the BERT encoder to obtain the initial text representation nt represents the total number of tokens in the sentence after word segmentation. Integrate the tokens belonging to the same word as the word representation:
[0128]
[0129] h k s represents the k-th word representation obtained after integrating the corresponding tokens, n represents the number of words in the sentence.
[0130] S43. Construct an adjacency matrix. Consider the matching between the syntactic parse tree and the word dependencies, and calculate the relevance between each pair of words in the text as the relationship weight matrix
[0131]
[0132] Relationship weight matrix Q relevance Combine with D described in step 4-1) ij to establish an adjacency matrix A ij ,
[0133] A ij = D ij Q relevance .
[0134] S44. Obtain the syntactic enhanced text representation based on the graph convolutional neural network layer. Input the adjacency matrix A ij and the word token representation into the graph convolutional network layer H S , to capture context syntactic information and long-distance word dependencies:
[0135]
[0136] represents the l-th layer hidden layer representation of the i-th node, is the adjacent node of node i. W l , b l are learnable parameters, and σ is the activation function ReLU.
[0137] Obtain the output of the last layer of the graph convolutional layer as the text representation enhanced by syntax.
[0138] S45. Integrate the global context text representation H C with the text representation H S enhanced by syntax as the text feature H CS :
[0139] H CS = Concat(H C , H S ).
[0140] S5. Multimodal sentiment classification module. Use the pooling mechanism to denoise the global context features, aspect-related visual sentiment features, and syntax-enhanced text information features, and input them into the multimodal dynamic attention encoding layer to perform cross-modal interactions to obtain multimodal fusion features.
[0141] S51. Input the text feature H CS and the aspect-related visual sentiment feature H T→VA into the dynamic attention pooling layer to obtain the multimodal fusion feature encoding H m :
[0142] H m(l) = ComATT(concat(H T→AV , H CS ))
[0143]
[0144] ComATT represents joint self-attention, and the multi-layer feature representation is where H m(l) is the feature of the l-th layer, and L is the number of layers of the dynamic pooling attention mechanism of the multimodal fusion module.
[0145] S6. Input the multimodal fusion feature into the Softmax layer to obtain the multimodal sentiment category probability distribution. Calculate the cross-entropy loss for the multimodal fusion feature. Calculate the mean square error loss between the aspect-related nouns and the aspect-related visual sentiment to reduce the representation difference and ensure more accurate perception of the visual features related to the aspect term. Minimize the joint loss of the two parts, optimize the model training parameters, improve the prediction ability of the multimodal aspect-level sentiment analysis model, and obtain the sentiment prediction result.
[0146] S61. Put the multi-modal fusion features into the softmax layer for sentiment classification:
[0147]
[0148] S62. Calculate the mean squared error to calculate the reconstruction loss. For a training sample set D, the first part of the loss comes from the mean squared error loss L generated when the aspect words in step 3-4) interact with vision. r :
[0149]
[0150] S63. Calculate the classification loss of the multi-modal fusion features. For a training sample set D, the second part of the loss comes from the multi-modal aspect-level sentiment analysis method described in claim 1 based on multi-scale text-visual feature enhancement. The cross-entropy quantization is used to calculate the classification loss L of the multi-modal fusion features:
[0151]
[0152] where |y| is the number of sentiment classification categories, which is 3, and t ij is the true label. Predict the sentiment polarity of the given aspect word in each text-image pair. In the prediction results, 0, 1, and 2 correspond to negative, neutral, and positive sentiments respectively.
[0153] S64. Calculate the gradient loss of the model's backpropagation. Control the proportion of the mean squared error loss L r in the model training loss through the hyperparameter λ, and obtain the loss value L of the model's backpropagation training during the gradient descent process:
[0154] L = L m + λL r .
[0155] To study the role of each module in the present invention, ablation experiments are conducted on three main modules: the image scene semantic extraction module (ISS), the text syntactic feature enhancement module (TSE), and the image target semantic extraction module (ITS).
[0156] Table 1 Ablation study of the main modules of MSPAF
[0157]
[0158] Table 1 shows the results of the ablation experiment of the model.
[0159] First, to prove the effectiveness of the Image Scene Semantic Extraction module (ISS) for this model, in this example, the Image Scene Semantic Extraction module (ISS) is removed, and the original text is directly used for encoding and visual modality fusion. As shown in Table 1, the accuracy of MSPAF decreases by 3.15% and 4.18% respectively, indicating that fusing image scene semantic information with multi-level text information can complement text information, thus improving the accuracy of multi-modal sentiment prediction.
[0160] In addition, to prove that adding the Syntactic Feature Enhancement module (TSE) can improve the extraction of aspect-related effective information in text by capturing long-distance dependencies, in this example, the Text Syntactic Feature Enhancement module (TSE) is removed. It can be found from Table 1 that the performance of MSPAF w / o TSE decreases by 2.03% and 2.65% respectively in terms of classification accuracy compared to MSPAF on the two Twitter datasets, indicating that constructing a text graph convolutional network by analyzing the dependencies of each sentence component can enhance the extraction of text syntactic information and provide support for the sentiment classification task of specific targets.
[0161] Finally, to prove that the Image Target Semantic Extraction module (ITS) can provide fine-grained visual information to compensate for general visual features, as shown in Table 1, in this example, the Image Target Semantic Extraction module (ITS) is removed, and general image encoding features are used to interact with aspect words. The accuracy of the model performance decreases by 1.64 and 2.34 percentage points respectively. This shows that adding the Image Target Semantic Extraction module can promote cross-modal interaction between images and aspect words, thus extracting more effective aspect-related visual features and further improving the sentiment classification effect of the model.
[0162] The above ablation experiments evaluate the impact of each main component or factor in the MSPAF model on the model performance. By removing or changing a certain part or factor in the network, the change in the model performance can be observed, and then the importance of this part can be determined.
[0163] Experimental Results
[0164] Table 2 Performance Comparison between MSPAF and Baseline Models
[0165]
[0166] The MSPAF model proposed in this chapter performs best on the public datasets Twitter-2015 and Twitter-2017. Compared with the best text modality model Dual-GCN, the Acc increases by 2.82% and 7.71%, and the Mac-F1 increases by 3.44% and 8.69%. This shows that the initial graph nodes of Dual-GCN use LSTM encoding, resulting in insufficient word semantic content and increasing the difficulty of extracting information by the graph neural network. While MSPAF uses BERT to represent the initial graph nodes with word embeddings, further enhancing the word semantics. At the same time, compared with the best multimodal baseline model TomBERT, the Acc of the model in this chapter increases by 1.15% and 1.45%, and the Mac-F1 increases by 2.04% and 1.14%. This indicates that through multi-scale semantic perception and semantic association modeling, multi-scale graphic and text semantic information can be effectively mined, promoting the effective interaction between text and image in a unified feature space and performing more accurate sentiment analysis on multimodal data.
[0167] Case Analysis
[0168] To better analyze the roles of the image scene semantic extraction module ISS and the text syntactic feature enhancement module TSE, three representative samples were randomly selected for case studies. At the same time, the two optimal models Dual-GCN and TomBERT in single-modal and multimodal were selected for comparison with MSPAF. The sample information and prediction results are as Figure 4 shown.
[0169] Case (1) contains two types of aspect words with positive and neutral sentiment polarities. Dual-GCN only analyzes based on the text content, while MSPAF extracts image scene information to supplement the text content and obtains the correct prediction result. In Case (2), Dual-GCN fails to judge the sentiment of the user towards the aspect word "Netanyahu" through the text. While Tom-BERT and MSPAF utilize cross-modal attention to associate the semantics of the image region and the aspect word, and correctly infer the sentiment polarity of the aspect word "Netanyahu" through the unhappy facial expression of the man in the image. In Case (3), TomBERT predicts incorrectly due to confusing the sentiments of different aspect words. MSPAF correctly predicts the sentiment polarity of "Blackhawks" by parsing the sentence component relationship and makes correct judgments on the sentiment tendencies of different aspect entities.
[0170] Attention Visualization
[0171] To better demonstrate the actual effect of the present invention, this section selects Figure 4 Case (3) in it to visualize the details of the aspect-guided visual attention mapping and the aspect-guided visual perception attention heat map is as Figure 5 shown.
[0172] Figure 5 The display model pays more attention to the red thermal imaging area in the attention heat map, which is highly correlated with the aspect term "BrandonSaad". This indicates that visual attention reduces the noise brought by irrelevant information and maximally retains the visual area related to "Brandon Saad". It is proved that the dynamic gating attention mechanism helps the alignment of images and aspect terms, guiding the model to capture the visual area information related to aspect terms. In the text graph convolutional feature enhancement module, after adding Figure 6 the syntactic dependency information of adjacent nodes, "Blackhawks" is affected by the noun modifier (nummod) - "3-" and the adverbial modifier (advmod) - "early.", and the context semantics related to the aspect term in the text is enhanced. From Figure 7 the syntactic feature weight attention matrix of the text graph convolution, it can be seen that the attention block at the intersection of two words with sentence dependency has a darker color. The emotion of "Blackhawks" is not affected by the irrelevant word "1", and it is also observed that the attention block at this place has a lighter color. This shows that the injection of syntactic information enhances the semantic understanding of the context related to the aspect entity.
[0173] In summary, the present invention proposes a multi-modal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement, which promotes the effective interaction between text and images in a unified feature space, and fuses image scene semantic information with multi-level text information; uses gated cross-attention to model the semantic association between aspects and images, integrates fine-grained visual sentiment features into the image feature representation related to aspect terms, and reduces the impact of the quality difference of heterogeneous modal information on cross-modal interaction. An adjacency matrix is constructed according to syntactic parsing and semantic similarity, and a text graph convolutional network is used to obtain the deep dependency relationship between adjacent nodes, and a syntactic and semantic enhanced context representation is obtained. In the feature fusion stage, a multi-layer attention pooling mechanism is used to denoise the multi-modal fusion features, effectively reducing the feature dimension while learning the correlation of different modal features, and improving the accuracy of the multi-modal aspect-level sentiment analysis model.
Claims
1. A multimodal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement, characterized in that: The following steps are involved: Step 1: Obtain a multimodal aspect-level sentiment dataset, wherein the multimodal aspect-level sentiment dataset includes a piece of text information, an aspect word sequence, and a related image; Step 2: extract the scene semantics of the associated image, fuse the image scene semantics with the text information, and input the pre-trained language model to obtain global context features; Step 3: Extract the target semantics of the associated image, divide the image target semantics into target noun semantics and target emotional semantics according to different part-of-speech features, align the aspect word sequence with the target noun semantics based on the dynamic gating mechanism, introduce a multi-head cross-attention layer to obtain the aspect-related visual information representation, and combine the target emotional semantics to integrate the acquisition of the aspect-related visual emotional features; Step 4: Obtain the dependency relationship between words in the text information based on dependency parsing, build an adjacency matrix based on the weights of the associations between words, introduce a graph convolutional neural network to capture contextual syntactic information and long-distance word dependencies, and obtain syntactically enhanced text features; Step 5: Fusing the global context features, syntactically enhanced text features and aspect-related visual sentiment features, denoising the multi-scale text visual features based on a pooling mechanism to obtain multimodal fusion features; Step 6: The multimodal fusion features are output to the sentiment prediction layer to predict the sentiment polarity of the aspect words.
2. According to claim 1, a multimodal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement is characterized in that: The step 2 is specifically as follows: Step 2-1: Input the image into the Caption Transformer encoding framework to generate image captions as image scene semantics D i ; Step 2-2: The text message S described in step 1 i , aspect word sequence T i and the scene description D obtained in step 2-1 i Separate the words to get Use [SEP] to separate the internal text sequence, [CLS] as the global context start marker, and obtain the integrated global context sequence C i : Step 2-3: Calculate the global context embedding and use the Bert-based-uncased pre-trained language model to embed C i映 Projected into a continuous vector space, the word vector embedding SD is obtained i , segmentation embedding SE i , position embedded PE i , SD i with SE i Add together to get the global context embedding sequence Step 2-4: Calculate the attention mask by embedding the PE at the position described in step 2-3 i After tensor expansion masked_softmax, we get the attention mask MA i : MA i =Masked_Softmax(expand(PE i )); Step 2-5: Add the attention mask MA described in step 2-4 i and the global context embedding sequence described in step 2-3 Input into the multi-layer self-attention encoder of Transformer to obtain global context features k is the maximum length of the global context, and d is the dimension of the hidden layer:
3. The multimodal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement according to claim 2 is characterized in that: The step 3 is specifically as follows: Step 3-1: Use SentiBank to extract the adjective pairs contained in the image ANP = (A k ,N k ), in order to reduce the introduction of irrelevant noise, the top 5 ANPs with the highest confidence in each image are selected as the image target semantics, and the image target semantics are divided into target emotional semantics A k , target noun semantics N k ; Step 3-2: The aspect word sequence described in step 1 and the target emotional semantics A described in step 3-1 are combined. k and the target noun semantics N k , input the Bert pre-trained language model to obtain the aspect word encoding where k t is the maximum length of aspect words, d t is the dimension of the hidden layer of aspect words, target emotion representation Target noun representation s represents the length of the maximum adjective or noun sequence; Step 3-3: Construct a dynamic gated adjustment vector, calculate the semantic similarity between the aspect word representation and the noun representation in the high-dimensional vector space, and construct the gated adjustment vector G A : G A =expand(cos<H T ,H N >); Gating adjustment vector G A The target noun representation H described in step 3-2 N Multiply to obtain the target noun representation constraint and the aspect representation H T Aggregation, acquisition Characterization for gating adjustment: Step 3-4: Obtain the visual features related to the aspect, input the associated image described in step 1 into the ResNet-152 image pre-training model to obtain the image code P, Characterize the gating adjustment aspects described in step 3-3 As Query, the image encoding representation P is input into the multi-head cross attention as Value and Key to obtain aspect-related visual attention A. T→V : CATT i is the cross attention of the i-th head, σ is the Gaussian error linear activation function Gelu, is a learnable parameter matrix, m is the number of attention heads of multi-head attention; Considering the differences in the numerical ranges of different pixels in the image data, a Gaussian error linear activation unit Gelu is added to ensure the generalization of the neural network in the random regularization process, reduce variable offset, and obtain aspect-related visual representation H. T→V : Step 3-5: Incorporate the target emotion information to obtain a fine-grained visual emotion representation. Analogously, the target noun representation is embedded in the aspect word in step 3-3, and the gated vector G in step 3-3 is used. A With adjective characterization H A Multiply to get visual emotion representation The visual representation H related to the aspects described in step 3-4 T→V , Visual Emotion Representation Attention pooling is performed separately so that the model can more effectively process and understand the relevant visual information, and then the tensor dimension is expanded: Step 3-6: Integrate the steps described in step 3-5 and Get aspect-related visual emotion features H T→VA :
4. The multimodal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement according to claim 3 is characterized in that: The step 4 is specifically as follows: Step 4-1: Build a dependency matrix and use spacy dependency parsing to obtain the dependency matrix between words Step 4-2: Calculate word lemma representation and input the text information in step 1 into the BERT encoder to obtain the initial text representation n t Represents the total number of lemmas in the sentence after word segmentation, and integrates lemmas belonging to the same word as word representation: represents the k-th word representation obtained after integrating the corresponding word unit, n represents the number of words in the sentence; Step 4-3: Construct an adjacency matrix, consider the dependency matching between the syntactic parse tree and the words, and calculate the correlation between each word in the text as a relationship weight matrix Relation weight matrix Q relevance Same as step 4-1 ij Combine to build the adjacency matrix A ij , A ij =D ij Q relevance ; Step 4-4: Obtain syntactically enhanced text representation based on the graph convolutional neural network layer, and transform the adjacency matrix A described in step 4-3 ij The word unit representation described in step 4-2 is input into the graph convolutional network layer H S , capturing contextual syntactic information and long-range word dependencies: f i l represents the lth hidden layer representation of the ith node, is the adjacent node of node i, W l , b l is a learnable parameter, σ is the activation function ReLU; Get the last layer output of the graph convolution layer Textual representation as syntactic enhancement; Step 4-5: Represent the global context H described in step 2-5 C The text representation H enhanced with the syntax described in step 4-4 S Integration of text features H CS : H CS =Concat(H C ,H S )。 5. The multimodal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement according to claim 4 is characterized in that: Step 5 is as follows: Step 5-1: Convert the text feature H described in step 4-5 CS Visual emotion features H related to the aspects described in step 3-4 T→VA , input the dynamic attention pooling layer to obtain the multimodal fusion feature encoding H m : H m(l) =ComATT(concat(H T→AV ,H CS )), ComATT stands for joint self-attention, and the multi-layer features are represented as Among them, H m(l) is the feature of the lth layer, and L is the number of layers of the dynamic pooling attention mechanism in the multimodal fusion module.
6. The multimodal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement according to claim 5 is characterized in that: Step 6 is as follows: Step 6-1: Put the multimodal fusion features described in step 5-1 into the softmax layer for sentiment classification: p(y|H m )=softmax(W ° H m ): Step 6-2: Calculate the mean square error to calculate the reconstruction loss. The multimodal aspect-level sentiment dataset is a training sample set D. For a training sample set D, the first part of the loss comes from the mean square error loss L generated when the aspect words interact with the visual in step 3-4. r : Step 6-3: Calculate the classification loss of the multimodal fusion feature. For a training sample set D, the second part of the loss comes from the multimodal aspect-level sentiment analysis method based on multi-scale text visual feature enhancement described in claim 1. The classification loss L of the multimodal fusion feature is calculated using cross entropy quantization. m : Among them, |y| is the number of sentiment classification categories, which is 3, t ij is the true label; Step 6-4: Calculate the gradient loss of the model back propagation, and control the mean square error loss L described in step 6-2 through the hyperparameter λ r The proportion of model training loss in the gradient descent process to obtain the loss value L of the model back propagation training: L=L m +λL r 。
Citation Information
Cited By
Data visualization method and system based on artificial intelligence
CN120852561A
Trusted decoding method for multi-mode large model sequence generation
CN121074159A
A trusted decoding method for multi-modal large model sequence generation
CN121074159B
Multi-modal aspect-level emotion analysis method and system based on interactive feedback
CN121958871A
Risk behavior detection method based on large model and multi-field data fusion
CN122432889A