Multi-modal aspect-level sentiment analysis method for enhancing visual features of target

Through a multimodal aspect-level emotion analysis method that enhances the visual characteristics of the target, the problems of difficulty in modal fusion, inaccurate emotional prediction and large resource consumption in the prior art are solved, and more efficient multimodal emotion analysis is achieved, improving the accuracy and resource utilization efficiency of emotional prediction.

CN120014320APending Publication Date: 2025-05-16JIANGSU OCEAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510001985.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing multimodal emotion analysis technology has problems in the difficulty of modal fusion, inaccurate emotional prediction and large resource consumption.

Method used

A multimodal aspect-level sentiment analysis method that enhances the visual characteristics of the target is proposed. Through multimodal data acquisition, text and image preprocessing, graph convolution network processing, multimodal interactive attention and emotion prediction modules, deep fusion and emotion prediction of text and image information are realized.

Benefits of technology

It improves the accuracy and efficiency of emotion prediction, reduces the consumption of computing resources, and can more accurately identify and utilize key visual and text features related to emotions, significantly improving the comprehensiveness and depth of multimodal emotion analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014320A_ABST
    Figure CN120014320A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal aspect-level sentiment analysis method for enhancing target visual features, and aims to improve sentiment analysis accuracy and detail analysis capability by integrating texts, aspect words and image data. According to the method, syntactic dependency relationship analysis and a KNN algorithm are combined, fine-grained information in a modal is fully mined, and deep semantic understanding of text content is enhanced. By using a CLIP model, similarity calculation and Faster R-CNN, a key visual area related to aspect words in an image is accurately positioned, and processing of visual information is optimized. In addition, an interactive attention mechanism is adopted to deeply mine associated features between modals, and texts, aspect words and image information are effectively fused. Experimental results show that the expression of the model on a public data set exceeds that of a plurality of baseline models, and the application potential and effect of the model in the field of multi-modal sentiment analysis are proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and natural language processing, and in particular relates to a multimodal aspect-level sentiment analysis method for enhancing target visual features. Background Art

[0002] With the rapid development of Internet technology, social media and e-commerce platforms have become important platforms for people's daily communication and business activities. User interactions on these platforms are mostly expressed through text comments, and more and more users tend to attach relevant pictures to enhance the expressiveness of information. The popularity of this multimodal (text and image) data provides new research directions and application scenarios for sentiment analysis technology, that is, to more comprehensively understand users' emotions and opinions by analyzing the comprehensive information of text and images.

[0003] Multimodal sentiment analysis refers to combining data from different information sources (such as text, images, audio, etc.) to predict users' emotional attitudes. In particular, the combination of text and images provides richer context and background information for sentiment analysis. For example, a picture with rich expressions and negative comments may express sarcastic emotions, which may be difficult to accurately capture in a single text analysis. Aspect-level sentiment analysis goes a step further and not only analyzes the overall emotional tendency, but also refines the sentiment analysis of specific aspects in the text, such as the specific attributes or characteristics of the evaluation object. This analysis helps companies accurately understand consumers' specific feelings about various aspects of the product. At present, the widely used public technologies include the following: Convolutional Neural Network (CNN): widely used in image processing, capable of capturing the spatial hierarchy in images. Long Short-Term Memory Network (LSTM): superior to traditional recurrent neural network (RNN), it can more effectively process long sequence text data. Attention mechanism: improves the model's ability to focus on important parts of information, especially widely used in aspect-level sentiment analysis. BERT (Bidirectional Encoder Representations from Transformers): as a pre-trained model, it is trained through a large-scale corpus and can effectively extract text features. Faster R-CNN: An efficient object detection technique commonly used to quickly and accurately identify and locate objects in images.

[0004] Although the current multimodal sentiment analysis technology has made some progress, it still has the following defects: Difficulty in modal fusion: There are large representation differences between different modalities, and how to effectively fuse text and image information is a challenge. Existing methods often fail to fully utilize the detailed information in the image, especially when the image content is directly related to the text information. Inaccurate sentiment prediction: Existing models often have low accuracy when processing text with sarcastic or ambiguous emotional expressions. This is because the model has difficulty capturing subtle emotional changes, especially when the literal meaning of the text is inconsistent with the emotion expressed in the image. High resource consumption: Multimodal analysis usually requires large computing resources, especially when processing high-resolution images and large-scale text data, the computational cost and time cost are high. Summary of the invention

[0005] In view of the problems of difficulty in modal fusion, inaccurate sentiment prediction and high resource consumption in existing multimodal sentiment analysis technology, the present invention proposes a multimodal aspect-level sentiment analysis method that enhances the visual features of the target to improve the accuracy and efficiency of sentiment prediction while reducing the consumption of computing resources. To achieve the above purpose, the technical solution of the present invention includes the following steps:

[0006] S1: Multimodal data acquisition: obtain a set of multimodal data samples D, where each sample d∈D includes text comments T={t1,t2,…,t n}, one or more aspect words A={a1,a2,…,a r} and an image I corresponding to the text comment;

[0007] S2: Text preprocessing and encoding: perform word segmentation, invalid character removal and case normalization preprocessing on the text comment T; use the BERT model to encode the preprocessed text comment T and the feature extraction encoder of the aspect word A, and the calculation formula is:

[0008] T text =BERT(T) (1)

[0009] T aspect =BERT(A) (2)

[0010] Among them, T text ∈R s×d , T aspect ∈R q×d , s and q represent the lengths of the encoded text and aspect word sequences, respectively, and d represents the hidden dimension of each word vector;

[0011] S3: Image preprocessing and encoding, including the following steps:

[0012] S3-1: resize and normalize the image I, and input the processed image into the pre-trained ResNet50 model to extract the image feature map, the calculation formula of which is:

[0013]

[0014] I ij =I[i×14:(i+1)×14,j×14:(j+1)×14] (4)

[0015] I ij image =ResNet50 (I ij ) (5)

[0016] Among them, I ij image ∈R 196×2048 , W and H are the width and height of the original image, N h and N W Indicates the number of image blocks, I ij represents the segmented image block;

[0017] S3-2: The feature dimension of the image modality I ij image Flattened into a two-dimensional matrix representation and mapped to the same dimension as the text embedding through a linear transformation layer, the calculation formula is:

[0018] I ij v =W0I ij image +b0 (6)

[0019] Among them, I ij v ∈R 196×d , W0 and b0 are learnable parameters of the linear transformation layer;

[0020] S4: Text graph construction and graph convolution network processing, using dependency syntax analysis tools to perform dependency syntax analysis on the words of the text comment T to obtain the adjacency matrix A of the text dependency relationship t ∈R n×n , providing adjacency matrix graph convolution network GCN processing, the calculation formula is:

[0021]

[0022] F t =H t (l+1) (8)

[0023] Among them, H t(l) represents the input feature of the GCN at layer l, H t (l+1) represents the output feature of the (l+1)th layer GCN, is the adjacency matrix A t The degree matrix of , W(l) represents the trainable parameters;

[0024] S5: Image graph construction and graph convolution network processing, for each image block I ij v The feature vector uses the K nearest neighbor algorithm to find the nearest neighbors of each feature vector, finds the K nearest neighbors of each feature vector, and records the indexes of these neighbors. And constructs the adjacency matrix A of the graph based on the neighbor information v ∈R m×m The specific calculation formula is as follows:

[0025]

[0026] F v =H v (s+1) (10)

[0027] Among them, H v (s) represents the input feature of the GCN at layer s, H v (s+1) represents the output feature of the (s+1)th layer GCN, is the degree matrix of the adjacency matrix Av, s represents the trainable parameter, and K is set to 8 by default;

[0028] S6: Encode the text and image simultaneously through the CLIP model, calculate the cosine similarity between the text and the image, and between the aspect words and the image, and obtain the comprehensive similarity through the weighted coefficient. The calculation formula is:

[0029] F image ,F text ,F aspect =CLIPModel(I,T,A) (11)

[0030]

[0031] Among them, F text ,F aspect ,F image Feature vectors representing text, aspect words, and images, sim Text,I and sim Aspect,I Represents the cosine similarity between text and image and between aspect words and image, sim combined It represents the comprehensive cosine similarity between text and aspect words and images. Its value range is [-1,1]. The closer the value is to 1, the more similar the vectors are. μ represents the proportional value of the similarity weight parameter.

[0032] S7: Multimodal interactive attention, using aspect words as queries and text as keys and values ​​to perform a cross-attention calculation; using the target visual features as queries and the text description generated by the BLIP model as keys and values ​​to perform another cross-attention calculation, dynamically adjusting the attention distribution. The calculation formula is:

[0033] Q i =T aspect W i Q , K i =T text W i K , V i =T text W i V (15)

[0034]

[0035] head i =Attention i (Q i, K i, V i ) (17)

[0036] MultiHead(Q , K , V)=Concat (head1,head2,…head h )W O (18)

[0037] Among them, W i Q , W i K , W i V As a weight matrix, h represents the number of attention heads, R A-T =MultiHead(Q , K , V) represents the aspect words as the text feature information of the query, where R A-T ∈R g×d ;

[0038] S8: Multimodal interactive attention, by splicing the feature vectors output by multiple modules, realizes the comprehensive integration of different modal information. The calculation formula is:

[0039] E ATV =concat (F t ; F v ; R A-T ; RV-T ) (19)

[0040] Z ATV =ReLU(W ATV E ATV +b ATV ) (20)

[0041] Among them, W ATV, b ATV is a trainable weight parameter;

[0042] S9: Perform sentiment prediction through sentiment prediction module.

[0043] As a technical preferred solution of the present invention, step S1 includes obtaining a set of multimodal image and text datasets D, each sample d∈D includes a text comment T, an image I and an aspect sequence A; the aspect sequence A is a subsequence of the text comment T; the length of the text comment is n, the length of the aspect word is r, and the image and text information (T, I) is comprehensively used to predict the sentiment polarity of the aspect sequence A, wherein the sentiment polarity includes three sentiment categories: positive, neutral and negative.

[0044] As a technical preferred solution of the present invention, the image preprocessing in step S3 includes the following steps: the input image is first divided into image blocks of size 14×14, and then each image block is adjusted to a standard size of 224×224 pixels; then, the resized image is converted into a tensor form, and its pixel value is normalized to the range of [0,1]; finally, through a standardization operation, the pixel value of each color channel is normalized, and its mean and standard deviation are adjusted to [0.485, 0.456, 0.406] and [0.229, 0.224, 0.225] respectively, to ensure that the image data is highly consistent with the input format of the pre-trained network.

[0045] As a technical preferred solution of the present invention, step S6 further uses the Faster R-CNN model to generate candidate regions in the image; by combining the Faster R-CNN with the CLIP model, the candidate regions related to the text description are accurately generated, wherein the determination of the candidate regions includes the bounding box of the object and the corresponding confidence score, and the extracted best target visual feature vector information, and the specific calculation formula is:

[0046] D = Faster R-CNN (I) (21)

[0047]

[0048] Among them, D contains the bounding box of the object and the corresponding confidence score, and F R is the optimal target visual feature vector information.

[0049] As a technical preferred solution of the present invention, step S7 further uses the best target visual feature as a query, and uses the text description T generated by the BLIP model image As keys and values, multi-head cross attention calculation is performed. This method uses the text feature information R guided by the target visual features. V-T ∈R 196×d ; This process uses the multi-head cross-attention mechanism to deeply fuse cross-modal information and capture the rich context of images and texts in multiple dimensions.

[0050] As a technical preferred solution of the present invention, the emotion prediction module in step S9 performs emotion prediction through the following steps:

[0051] S9-1: Input the multimodal fusion vector F into a multilayer perceptron, passing through several fully connected layers and using an activation function;

[0052] S9-2: Softmax function is used in the output layer to obtain the sentiment prediction result y of the aspect word, where the sentiment polarity includes positive, neutral and negative;

[0053] S9-3: During the training process, the model is trained with standard gradient descent using the standard cross entropy loss function plus the L2 regularization term as the loss function. The calculation formula is:

[0054] y=softmax(W s MLP(Z ATV )+b s ) (twenty three)

[0055]

[0056] Among them, softmax is the activation function, W s and b s is a trainable weight matrix, i represents a sample in the data set, l represents the set containing all samples, and y i is the true value of the sample label, is the predicted label value of the sample; λ is the regularization coefficient, and Θ is all trainable parameters.

[0057] Compared with the related prior art, the beneficial effects of the present invention are:

[0058] Effective extraction of fine-grained feature information: By combining syntactic dependencies and the KNN algorithm, this model can more deeply mine fine-grained information in text and images. Compared with traditional multimodal sentiment analysis techniques, this method can more accurately identify and utilize key visual and text features related to sentiment.

[0059] Accurate visual feature localization: By using the CLIP model combined with Faster R-CNN to accurately locate the target visual features, the present invention can effectively identify the image areas most relevant to the textual aspect words. This method improves the accuracy of sentiment analysis, especially when dealing with complex visual situations, and can better parse the association between image content and textual meaning compared to existing technologies.

[0060] Mining deep inter-modal associations: Through the interactive attention mechanism, the present invention strengthens the information interaction between different modalities, so that the model not only processes the surface modal data, but can deeply understand the intrinsic relationship between modalities. This deep information fusion method is a major innovation in multimodal sentiment analysis, which can significantly improve the comprehensiveness and depth of analysis.

[0061] Optimization of performance and resource efficiency: This paper is committed to designing a more lightweight model structure, reducing computational complexity and resource consumption, making the model more suitable for running in a resource-constrained environment. This is particularly important for real-time or large-scale data processing scenarios, and helps to improve the practicality and scalability of the model.

[0062] Adaptability and generalization: By introducing more modal information and improving the target visual module, the present invention is able to handle a wider range of data types and more complex sentiment analysis tasks. This flexibility and strong generalization ability make the model not only suitable for current social media and e-commerce platforms, but also adapt to new platforms and new types of multimodal data that may appear in the future. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 is a flow chart of a multimodal aspect-level sentiment analysis method for enhancing target visual features of the present invention;

[0064] Figure 2 The present invention provides case a (left picture) and case b (right picture) of the embodiment;

[0065] Figure 3 It is an overall structural diagram of the multimodal aspect-level sentiment analysis model of the embodiment provided by the present invention;

[0066] Figure 4 is a structural diagram of a target visual feature extraction module according to an embodiment of the present invention;

[0067] Figure 5 is a performance graph of the number of GCN layers of an embodiment provided by the present invention on the Twitter2015 dataset;

[0068] Figure 6 is a performance graph of the number of GCN layers of an embodiment provided by the present invention on the Twitter2017 dataset;

[0069] Figure 7 is a performance graph of the u value of the embodiment provided by the present invention on the Twitter2015 data set;

[0070] Figure 8 is the performance of the u value of the embodiment provided by the present invention on the Twitter2017 dataset;

[0071] Fig. 9 It is a case study effect diagram of an embodiment provided by the present invention. DETAILED DESCRIPTION

[0072] The present invention is further described below in conjunction with the accompanying drawings and examples. However, the present invention can be implemented in many different ways and should not be construed as being limited to the embodiments shown; on the contrary, these embodiments provide those skilled in the art with implementation methods that meet applicable legal requirements.

[0073] Example 1: According to Figure 1 As shown, this embodiment provides a specific implementation process of a multimodal aspect-level sentiment analysis method for enhancing target visual features. The steps are as follows:

[0074] S1: Figure 2 The cases a and b shown in the figure show two typical multimodal sentiment analysis scenarios. The multimodal data acquisition is specifically as follows: a set of multimodal data samples D is acquired, where each sample d∈D includes text comments T={t1,t2,…,t n}, one or more aspect words A={a1,a2,…,a r} and an image I corresponding to the text comment;

[0075] like Figure 3 As shown, the overall structure of the multimodal aspect-level sentiment analysis model described in the present invention includes a picture-text feature extraction module, a target visual feature extraction module, a picture-text neural network module, a multimodal interactive attention module, a multimodal feature fusion module and a sentiment prediction module.

[0076] S2: Text preprocessing and encoding: perform word segmentation, invalid character removal and case normalization preprocessing on the text comment T; use the BERT model to encode the preprocessed text comment T and the feature extraction encoder of the aspect word A, and the calculation formula is:

[0077] T text =BERT(T) (1)

[0078] T aspect =BERT(A) (2)

[0079] Among them, T text ∈R s×d, T aspect ∈R q×d , s and q represent the lengths of the encoded text and aspect word sequences, respectively, and d represents the hidden dimension of each word vector;

[0080] S3: Image preprocessing and encoding, including the following steps:

[0081] S3-1: resize and normalize the image I, and input the processed image into the pre-trained ResNet50 model to extract the image feature map, the calculation formula of which is:

[0082]

[0083] I ij =I[i×14:(i+1)×14,j×14:(j+1)×14] (4)

[0084] I ij image =ResNet50 (I ij ) (5)

[0085] Among them, I ij image ∈R 196×2048 , W and H are the width and height of the original image, N h , N W Indicates the number of image blocks, I ij represents the segmented image block;

[0086] S3-2: The feature dimension of the image modality I ij image Flattened into a two-dimensional matrix representation and mapped to the same dimension as the text embedding through a linear transformation layer, the calculation formula is:

[0087] I ij v =W0I ij image +b0 (6)

[0088] Among them, I ij v ∈R 196×d , W0 and b0 are learnable parameters of the linear transformation layer;

[0089] S4: Text graph construction and graph convolution network processing, using dependency syntax analysis tools to perform dependency syntax analysis on the words of the text comment T to obtain the adjacency matrix A of the text dependency relationship t ∈R n×n , providing adjacency matrix graph convolution network GCN processing, the calculation formula is:

[0090]

[0091] F t =H t (l+1) (8)

[0092] Among them, H t (l) represents the input feature of the GCN at layer l, H t (l+1) represents the output feature of the (l+1)th layer GCN, is the adjacency matrix A t The degree matrix of , W(l) represents the trainable parameters;

[0093] S5: Image graph construction and graph convolution network processing, for each image block I ij v The feature vector uses the K nearest neighbor algorithm to find the nearest neighbors of each feature vector, finds the K nearest neighbors of each feature vector, and records the indexes of these neighbors. And constructs the adjacency matrix A of the graph based on the neighbor information v ∈R m×m The specific calculation formula is as follows:

[0094]

[0095] F v =H v (s+1) (10)

[0096] Among them, H v (s) represents the input feature of the GCN at layer s, H v (s+1) represents the output feature of the (s+1)th layer GCN, is the degree matrix of the adjacency matrix Av, s represents the trainable parameter, and K is set to 8 by default;

[0097] like Figure 4 As shown in the figure, a target visual feature extraction module is designed, which uses the CLIP model to process text and image information simultaneously, so as to effectively capture the semantic connection between the two. CLIP (Contrastive Language-Image Pre-training) is an advanced multimodal learning model developed by OpenAI, which aims to map text and images to the same semantic space through contrastive learning. Through training, the representation between the image and the corresponding text is made as close as possible, while the representation with irrelevant text or image is kept at a large distance. Next, by calculating the similarity, the text can quantitatively evaluate the degree of match between the text and the image. This process not only helps to identify the subject content in the image, but also deeply analyzes the relevant features in the image, and then extracts key information that supports the text intent.

[0098] S6: Encode the text and image simultaneously through the CLIP model, calculate the cosine similarity between the text and the image, and between the aspect words and the image, and obtain the comprehensive similarity through the weighted coefficient. The calculation formula is:

[0099] F image ,F text ,F aspect =CLIPModel(I,T,A) (11)

[0100]

[0101]

[0102] Among them, F text ,F aspect ,F image Feature vectors representing text, aspect words, and images, sim Text,I and sim Aspect,I Represents the cosine similarity between text and image and between aspect words and image, sim combined It represents the comprehensive cosine similarity between text and aspect words and images. Its value range is [-1,1]. The closer the value is to 1, the more similar the vectors are. μ represents the proportional value of the similarity weight parameter.

[0103] S7: Multimodal interactive attention, using aspect words as queries and text as keys and values ​​to perform a cross-attention calculation; using the target visual features as queries and the text description generated by the BLIP model as keys and values ​​to perform another cross-attention calculation, dynamically adjusting the attention distribution. The calculation formula is:

[0104] Q i =T aspect W i Q , K i =T text W i K , V i =T text W i V (15)

[0105]

[0106] head i =Attention i (Q i, K i, V i ) (17)

[0107] MultiHead(Q , K , V)=Concat (head1,head2,…head h )W O (18)

[0108] Among them, W i Q , W i K , W i V As a weight matrix, h represents the number of attention heads, R A-T =MultiHead(Q , K , V) represents the aspect words as the text feature information of the query, where R A-T ∈R g×d ;

[0109] S8: Multimodal interactive attention, by splicing the feature vectors output by multiple modules, realizes the comprehensive integration of different modal information. The calculation formula is:

[0110] E ATV =concat (F t ; F v ; R A-T ; R V-T ) (19)

[0111] Z ATV =ReLU(W ATV E ATV +b ATV ) (20)

[0112] Among them, W ATV, b ATV is a trainable weight parameter;

[0113] S9: Use the emotion prediction module to perform emotion prediction and transform the feature vector Z ATV Through MLP, we further process and organize the complex relationships between features and improve the model’s expressiveness. Finally, we use softmax to get the predicted probability distribution of sentiment polarity, and use the standard cross entropy loss function plus the L2 regularization term as the loss function to perform standard gradient descent training on the model.

[0114] y=softmax(W s MLP(Z ATV )+b s ) (twenty three)

[0115]

[0116] Among them, softmax is the activation function, W s and b s is a trainable weight matrix, i represents a sample in the data set, l represents the set containing all samples, and y i is the true value of the sample label, is the predicted label value of the sample; λ is the regularization coefficient, and Θ is all trainable parameters.

[0117] Example 2: In order to verify the effectiveness of the proposed multimodal aspect-level sentiment analysis method (GCN-TOVF) for enhancing target visual features, a specific experimental study was conducted to compare the model with the prior art. The experiments used the multimodal aspect-level public datasets Twitter2015 and Twitter2017, which cover a large number of user tweets, including images associated with text and clearly annotated aspect word sentiment polarity.

[0118] Table 1: Statistics of Twitter2015 and Twitter2017 datasets

[0119]

[0120] The Twitter2015 and Twitter2017 datasets contain rich text and image data. Each sample has one or more aspect words, and each aspect word has a clear sentiment label (positive, neutral, negative). The dataset is divided into training set, validation set, and test set in a ratio of 3:1:1. Table 1 shows the detailed information statistics of the two datasets.

[0121] In order to measure the model performance of the multimodal aspect-level sentiment analysis task, the experiment uses Accuracy (ACC), Macro-F1 (F1), Precision (P) and Recall (R) as the final evaluation indicators of the model.

[0122] To verify the performance of the model, the text model is compared with the following representative multimodal aspect-level sentiment analysis baseline models on the Twitter2015 and Twitter2017 datasets.

[0123] Table 2: Comparative experimental results

[0124]

[0125] The experimental results listed in Table 2 show that the performance comparison between the proposed model and the mainstream baseline models on the Twitter2015 and Twitter2017 datasets shows that the GCN-TOVF model outperforms most of the baseline models in terms of accuracy (Acc) and F1 value.

[0126] Specifically, the EF-Net and VLP-MABSA models improve model performance by introducing attention mechanisms to capture the correlation information between modalities, but they ignore the fine-grained feature information within each modality. To address this problem, the GCN-TOVF model uses the syntactic dependencies in the text and the KNN algorithm to associate the fine-grained feature information within each modality, thereby significantly enhancing the performance of the model. Compared with VLP-MABSA, the accuracy (Acc) and F1 value of GCN-TOVF on the Twitter2015 dataset increased by 0.56% and 1.71% respectively.

[0127] Secondly, the HIMT, ITM, KEF-TomBERT, M2DF, and MGFN-SD models mine deep semantic features in images through aspect words and text modalities to reduce the impact of visual noise on the model. In contrast, the GCN-TOVF model designs a target visual extraction module to accurately locate the visual features associated with the target aspect words, further enhance the visual feature information and reduce the impact of noise information. Compared with MGFN-SD, the accuracy (Acc) and F1 value of GCN-TOVF on the Twitter2017 dataset increased by 0.04% and 0.14%, respectively.

[0128] Finally, although the accuracy (Acc) of the text model GCN-TOVF on the Twitter2015 dataset is slightly lower than that of the MGFN-SD model, it still performs well in terms of overall performance. This fully proves that GCN-TOVF has significant advantages in fine-grained feature extraction and target visual feature localization.

[0129] Example 3: In order to explore the impact of the number of GCN layers on model performance, the present invention conducted an experiment in which the number of GCN layers in both text and image modalities was set to 1 to 6 layers, and the model performance under different layer settings was evaluated on two datasets. The experimental results are shown in Figure 3. Figure 5 and Figure 6 shown.

[0130] From the experimental results, we can see that when the number of GCN layers is set to 2, the model achieves the best performance on the Twitter2015 and Twitter2017 datasets. However, as the number of GCN layers increases, the model performance gradually decreases. This phenomenon can be attributed to the following points:

[0131] First, increasing the number of GCN layers will increase the complexity of the model, which may lead to overfitting. Secondly, the increase in the number of GCN layers leads to the problem of gradient vanishing or exploding, especially when training deep networks. This interferes with the training process of the model and affects its convergence and performance. In addition, as the number of GCN layers increases, the consumption of computing resources will also increase significantly. Finally, excessive layer settings may lead to the dilution of information transfer. In a graph convolutional network, each layer aggregates the information of neighboring nodes. When there are too many layers, the information is transmitted too far, and the information of distant nodes is over-smoothed, resulting in the loss of local feature details.

[0132] In summary, setting the number of GCN layers to 2 is the best choice in the text model. This not only achieves the best performance of the model, but also takes into account the computational efficiency and generalization ability of the model.

[0133] In order to analyze the weight ratio of the similarity between text and image and between aspect words and image, the present invention conducted experiments on two data sets and evaluated the impact of different similarity weight parameter values ​​on model performance. The experimental results are shown in Figure 2. Figure 7 and Figure 8 As shown. The experimental results show that when the similarity weight parameter value is set to 0.4, the performance of the model reaches the best on both datasets. Specifically, the weight ratio of aspect words to image similarity is 0.4, and the weight ratio of text to image similarity is 0.6. This shows that text information plays a more prominent role in locating the best target visual area. And text information can provide direct and clear semantic information, which helps to understand the image content more accurately and establish associations with aspect words. However, as the similarity weight parameter value gradually increases, the model performance becomes unstable and eventually decreases. This shows that too high a proportion of association weights may cause the model to over-rely on a single source of information in multimodality, thereby ignoring other important information. Therefore, the optimal weight distribution needs to balance the contributions of multiple information sources to ensure the stability of the model and the optimization of performance.

[0134] In order to study the impact of each component in the GCN-TOVF model on the model performance, the present invention conducts ablation experiments on the text graph convolutional network layer (TextGCN), image graph convolutional network layer (ImageGCN), multi-modal interactive attention module (Multi-Head Cross-Attention) and target visual feature extraction module respectively. Table 3 shows the ablation experiment results on the Twitter2015 and Twitter2017 datasets.

[0135] Experimental results show that removing the text graph convolutional network layer and the image graph convolutional network layer will lead to a significant decrease in model performance. This shows that TextGCN and ImageGCN are crucial for fully mining the fine-grained feature information deep inside the modality and realizing accurate association of information within the modality. In addition, removing the multimodal interactive attention module will result in the inability to fully interact and associate the feature information between the text, aspect words and image modalities, making it easy for the model to ignore important associated information when capturing feature information, affecting its overall performance. This highlights the important role of the multimodal interactive attention module in strengthening the interaction between modalities and fully understanding the complex information between modalities.

[0136] Finally, removing the target visual feature extraction module will cause the model's accuracy (Acc) and F1 value to drop by 3.97% and 2.3%, and 0.95% and 0.94% on the two datasets, respectively. This shows that the target visual module plays a significant role in helping the model accurately locate the image areas associated with aspect words. Through this module, the model can not only enhance the ability to extract image feature information, but also effectively reduce the interference of image noise information, thereby improving the overall performance. Therefore, the target visual feature extraction module plays a key role in optimizing the model's utilization of visual information and improving the model's robustness.

[0137] Table 3: Ablation experiment results

[0138]

[0139]

[0140] In order to verify the performance advantages of the text model, three representative data samples were selected from the two datasets for experiments. Fig. 9The comparison of the prediction results of the text model GCN-TOVF and the baseline model HIMT in these samples is shown. First, in the first sample, the sentiment polarity of the aspect word "Harry Gulliver" is positive. Both GCN-TOVF and HIMT can correctly judge the result, indicating that both models can fully extract the word "Happy" expressing emotions in the sample. Secondly, in the second sample, because there are no emotional words directly given in the text, HIMT makes mistakes in predicting the aspect word "Meghan Trainor", which shows that analyzing the sentiment of the text not only requires capturing significant emotional words, but also further associating the semantic structure within the text. In this regard, GCN-TOVF combines syntactic dependencies to enhance the extraction of fine-grained feature information in the text, thereby obtaining correct sentiment prediction results when predicting the aspect word "Meghan Trainor". Finally, in the third sample, the text discovery HIMT makes an error in predicting the aspect word "nba" because it does not fully combine image information to extract information related to the aspect word. In contrast, GCN-TOVF designs a target visual feature extraction module to focus on image regions related to aspect words, which enables GCN-TOVF to correctly predict the sentiment polarity of the aspect word "nba".

[0141] In order to further verify the effect of the target visual feature extraction module designed by the model, the present invention conducts in-depth visual analysis on three sample images. The experimental results show that when the aspect words in the sample images are "Harry Gulliver", "Meghan Trainor" and "Kyrie Irving", GCN-TOVF can accurately detect the parts of the image related to these aspect words. This shows that the target visual feature extraction module not only significantly improves the performance of the model, but also effectively reduces the noise interference in the image.

[0142] This embodiment demonstrates the advanced multimodal aspect-level sentiment analysis method proposed in the text, namely a multimodal aspect-level sentiment analysis method that enhances the target visual features. This method effectively mines fine-grained features in text and images by fusing syntactic dependencies and the KNN algorithm, significantly improving the deep parsing capabilities of intra-modal information. By combining the CLIP model and Faster R-CNN, this model can accurately locate key visual areas related to aspect words, thereby more accurately predicting sentiment polarity. Experimental results show that the GCN-TOVF model outperforms existing technologies on public datasets, verifying its efficiency and reliability in practical applications.

[0143] The above embodiments only express several implementation methods of the present invention, and the description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.

Claims

1. A multimodal aspect-level sentiment analysis method for enhancing target visual features, characterized by: The following steps are involved: S1: Multimodal data acquisition: obtain a set of multimodal data samples D, where each sample d∈D includes text comments T={t1,t2,…,t n }, where one or more aspect words A={a1,a2,…,a r } and an image I corresponding to the text comment; S2: Text preprocessing and encoding: perform word segmentation, invalid character removal and case normalization preprocessing on the text comment T; use the BERT model to encode the preprocessed text comment T and the feature extraction encoder of the aspect word A, and the calculation formula is: T text =BERT(T) (1) T aspect =BERT(A) (2) Among them, T text ∈R s×d , T aspect ∈R q×d , s and q represent the lengths of the encoded text and aspect word sequences, respectively, and d represents the hidden dimension of each word vector; S3: Image preprocessing and encoding, including the following steps: S3-1: resize and normalize the image I, and input the processed image into the pre-trained ResNet50 model to extract the image feature map, the calculation formula of which is: I ij =I[i×14:(i+1)×14,j×14:(j+1)×14] (4) I ij image =ResNet50 (I ij ) (5) Among them, I ij image ∈R 196×2048 , W and H are the width and height of the original image, N h , N W Indicates the number of image blocks, I ij Represents the segmented image block, i, j represents the pixel position; S3-2: The feature dimension of the image modality I ij image Flattened into a two-dimensional matrix representation and mapped to the same dimension as the text embedding through a linear transformation layer, the calculation formula is: I ij v =W0I ij image +b0 (6) Among them, I ij v ∈R 196×d , W0 and b0 are learnable parameters of the linear transformation layer; S4: Text graph construction and graph convolution network processing, using dependency syntax analysis tools to perform dependency syntax analysis on the words of the text comment T to obtain the adjacency matrix A of the text dependency relationship t ∈R n×n , providing adjacency matrix graph convolution network GCN processing, the calculation formula is: F t =H t (l+1) (8) Among them, H t (l) represents the input feature of the GCN at layer l, H t (l+1) represents the output feature of the (l+1)th layer GCN, is the adjacency matrix A t The degree matrix of , W(l) represents the trainable parameters; S5: Image graph construction and graph convolution network processing, for each image block I ij v The feature vector uses the K nearest neighbor algorithm to find the nearest neighbors of each feature vector, finds the K nearest neighbors of each feature vector, records the indexes of these neighbors, and constructs the adjacency matrix A of the graph based on the neighbor information v ∈R m×m , the specific calculation formula is as follows: F v =H v (s+1) (10) Among them, H v (s) represents the input feature of the GCN at layer s, H v (s+1) represents the output feature of the (s+1)th layer GCN, is the degree matrix of the adjacency matrix Av, s represents the trainable parameter, and K is set to 8 by default; S6: Encode the text and image simultaneously through the CLIP model, calculate the cosine similarity between the text and the image, and between the aspect words and the image, and obtain the comprehensive similarity through the weighted coefficient. The calculation formula is: F image ,F text ,F aspect =CLIPModel(I,T,A) (11) Among them, F text ,F aspect ,F image Feature vectors representing text, aspect words, and images, sim Text,I and sim Aspect,I Represents the cosine similarity between text and image, aspect word and image, sim combined It represents the comprehensive cosine similarity between text and aspect words and images. Its value range is [-1,1]. The closer the value is to 1, the more similar the vectors are. μ represents the proportional value of the similarity weight parameter. S7: Multimodal interactive attention, using aspect words as queries and text as keys and values ​​to perform a cross-attention calculation; using the target visual features as queries and the text description generated by the BLIP model as keys and values ​​to perform another cross-attention calculation, dynamically adjusting the attention distribution. The calculation formula is: Q i =T aspect W i Q , K i =T text W i K , V i =T text W i V (15) head i =Attention i (Q i, K i, V i ) (17) MultiHead(Q , K , V)=Concat (head1,head2,…head h )W O (18) Among them, W i Q , W i K and W i V As a weight matrix, h represents the number of attention heads, R A-T =MultiHead(Q , K , V) represents the aspect words as the text feature information of the query, where R A-T ∈R g×d ; S8: Multimodal interactive attention, by splicing the feature vectors output by multiple modules, realizes the comprehensive integration of different modal information. The calculation formula is: E ATV =concat (F t ;F v ;R A-T ;R V-T ) (19) WITH ATV =ReLU(W ATV E ATV +b ATV ) (20) Among them, W ATV, b ATV is a trainable weight parameter; S9: Perform sentiment prediction through sentiment prediction module.

2. The multimodal aspect-level sentiment analysis method for enhancing target visual features according to claim 1, characterized in that: The step S1 includes obtaining a set of multimodal image-text datasets D, each sample d∈D includes a text comment T, an image I and an aspect sequence A; the aspect sequence A is a subsequence of the text comment T; the length of the text comment is n, the length of the aspect word is r, and the image-text information (T, I) is comprehensively used to predict the sentiment polarity of the aspect sequence A, wherein the sentiment polarity includes three sentiment categories: positive, neutral and negative.

3. The multimodal aspect-level sentiment analysis method for enhancing target visual features according to claim 1, characterized in that: In step S3, the image preprocessing includes the following steps: the input image is first divided into image blocks of size 14×14, and then each image block is adjusted to a standard size of 224×224 pixels; then, the resized image is converted into a tensor form, and its pixel value is normalized to the range of [0,1]; finally, through a standardization operation, the pixel value of each color channel is normalized, and its mean and standard deviation are adjusted to [0.485, 0.456, 0.406] and [0.229, 0.224, 0.225] respectively.

4. The multimodal aspect-level sentiment analysis method for enhancing target visual features according to claim 1, characterized in that: In step S6, the FasterR-CNN model is further used to generate candidate regions in the image; by combining FasterR-CNN with the CLIP model, the candidate regions related to the text description are accurately generated, wherein the determination of the candidate regions includes the bounding box of the object and the corresponding confidence score, as well as the extracted best target visual feature vector information, and the specific calculation formula is: D = Faster R-CNN (I) (21) Among them, D contains the bounding box of the object and the corresponding confidence score, and F R is the optimal target visual feature vector information.

5. The multimodal aspect-level sentiment analysis method for enhancing target visual features according to claim 1, characterized in that: In step S7, the best target visual feature is further used as a query, and the text description T generated by the BLIP model is used image As the key and value, multi-head cross attention calculation is performed. This method uses the text feature information R guided by the target visual feature V-T ∈R 196×d ; This process uses the multi-head cross-attention mechanism to deeply fuse cross-modal information and capture the rich context of images and texts in multiple dimensions.

6. The multimodal aspect-level sentiment analysis method for enhancing target visual features according to claim 1, characterized in that: In step S9, the emotion prediction module performs emotion prediction by the following steps: S9-1: Input the multimodal fusion vector F into a multilayer perceptron, passing through several fully connected layers and using an activation function; S9-2: Softmax function is used in the output layer to obtain the sentiment prediction result y of the aspect word, where the sentiment polarity includes positive, neutral and negative; S9-3: During the training process, the model is trained with standard gradient descent using the standard cross entropy loss function plus the L2 regularization term as the loss function. The calculation formula is: y =softmax(W s MLP(Z ATV )+b s ) (23) Among them, softmax is the activation function, W s and b s is a trainable weight matrix, i represents a sample in the data set, l represents the set containing all samples, and y i is the true value of the sample label, is the predicted label value of the sample; λ is the regularization coefficient, and Θ is all trainable parameters.

Citation Information

Cited By

  • Multi-modal picture understanding method and device based on cross-modal mark fusion

    CN120611154A

  • Aspect-level multi-modal sentiment analysis method, system, equipment and medium

    CN120873962A

  • News broadcasting method based on artificial intelligence and related device

    CN121075311A