GCN-based Aspect-level Multimodal Sentiment Analysis Method

By integrating grammatical information and image block feature fusion into the graph convolution neural network, the problem of low accuracy of sentiment analysis in the prior art is solved, and efficient fine-grained emotional information extraction for aspect-level multimodal sentiment analysis is achieved.

CN116756314BActive Publication Date: 2025-07-29XIDIAN UNIV +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310719473.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-16
Publication Date
2025-07-29
Estimated Expiration
2043-06-16

AI Technical Summary

Technical Problem

The existing aspect-level multimodal emotion analysis methods cannot accurately extract emotional information related to aspect-words, and there is noisy information in the image mode, resulting in low accuracy in emotional polarity classification.

Method used

Graph convolution neural network is used to integrate syntax information into text features, and the image is divided into image blocks, feature fusion is performed through attention mechanism, and finally the graph convolution neural network is used to aggregate the associated fusion features.

Benefits of technology

It improves the accuracy of sentiment analysis, reduces the noise influence in the image mode, and can accurately extract emotional characteristics related to the aspect words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116756314B_ABST
    Figure CN116756314B_ABST
Patent Text Reader

Abstract

The present invention discloses an aspect-level multi-modal sentiment analysis method based on GCN, which includes the following steps: generating a training set; constructing an aspect-level multi-modal sentiment analysis network; training the aspect-level multi-modal sentiment analysis network; and classifying the aspect-level multi-modal sentiment polarity. The present invention constructs an aspect-level multi-modal sentiment analysis network based on GCN, and uses a graph convolutional neural network to integrate syntactic information into text features, enabling aspect words to pay attention to the sentiment information expressed by correct opinion words and reducing the influence of noise in the text. The present invention divides an image into image patches and fuses them with the text modality, constructs a graph convolutional neural network for fusing modalities, fuses relevant image information, obtains fine-grained image information related to aspect words in the image, reduces the influence of noise information in the image, and effectively improves the accuracy of sentiment analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of electronic digital data processing technology, and further relates to a method for aspect-level multimodal sentiment analysis based on graph convolutional neural networks (GCNs) in the field of multimodal content understanding and data analysis. This invention can be used to analyze and understand the sentiment of aspect entities in multimodal data such as images and text. Background Art

[0002] The goal of aspect-level multimodal sentiment analysis is to analyze the sentiment polarity of specified aspect terms within a given multimodal dataset (such as text and images). With the development of the internet, the amount of multimodal data generated is increasing. Using computers to analyze and determine the sentiment polarity of specified aspect terms within this multimodal data has become a critical approach. Because multiple aspect terms exist within a piece of multimodal data, and these terms may exhibit different sentiment tendencies, considering only coarse-grained, overall sentiment information can lead to erroneous sentiment judgments. Unlike multimodal sentiment analysis, this task is a fine-grained sentiment analysis task. Therefore, how to accurately and efficiently mine the potential fine-grained sentiment information within multimodal data is a pressing challenge in the field of aspect-level multimodal sentiment analysis. Textual modality dominates sentiment analysis, making it particularly important to ensure that aspect terms accurately focus on the information expressed by the correct opinion terms. While the proper utilization of image information can improve sentiment analysis performance, the noise contained in images can lead to erroneous sentiment judgments. Therefore, it is necessary to extract information related to aspect terms within images to mitigate the impact of noise on sentiment analysis and improve performance.

[0003] Many methods have been proposed to address aspect-level multimodal sentiment analysis. For example, Guilin University of Electronic Technology, in its patent application, "A Method for Aspect-Level Multimodal Sentiment Analysis Based on Collaborative Attention Fusion" (Application No.: 2022109650599, Publication No.: CN 115293170 A), discloses a collaborative attention-based global-local feature fusion network for aspect-level multimodal sentiment analysis. This method first acquires aspect-guided text and image features, then utilizes a cross-modal feature interaction mechanism to obtain fused features from the image and text modalities. Finally, a gated multimodal fusion mechanism is used to obtain the final sentiment features for sentiment classification. However, this method has drawbacks: when processing the text modality, it uses a long short-term memory network to obtain text representations related to aspect words, which can cause aspect words to focus on incorrect opinion words. When processing the image modality, the use of a ResNet to acquire visual features can introduce noise information unrelated to the aspect words, affecting the sentiment classification results and reducing the accuracy of sentiment polarity classification.

[0004] In their published paper "Hierarchical Interactive Multimodal Transformer for Aspect-Based Multimodal Sentiment Analysis" (IEEE Transactions on Affective Computing, 2021), Yu et al. disclosed an aspect-level multimodal sentiment analysis method based on a hierarchical interactive multimodal Transformer. This method uses Faster R-CNN to extract prominent features with semantic concepts from images to reduce the influence of noise in the images. The hierarchical interaction module in this method first models the interaction between aspect-text and aspect-image, and then captures the text-image interaction. To eliminate the semantic gap between the text and image modalities, an auxiliary reconstruction module is designed. This method uses the self-attention mechanism to obtain text features. Since the self-attention mechanism has superior long-term and short-term feature capture capabilities, it is also vulnerable to the influence of other irrelevant opinion words in the sentence, resulting in incorrect sentiment judgments. At the same time, this method uses Faster R-CNN to extract image features related to aspect words. However, due to the complexity of image information, it is impossible to accurately extract image information related to aspect words using this method, and even no image information can be extracted in some images.

[0005] In summary, for the fine-grained multimodal sentiment analysis task, existing methods fail to fully extract the sentiment information related to aspect words from multimodal data. In the present invention, syntactic information is introduced into the text features to construct the relationship between aspect terms and opinion words, and the fine-grained information in the images is extracted to construct complete scene information to obtain image information related to aspect words. Summary of the Invention

[0006] The purpose of the present invention is to propose an aspect-level multimodal sentiment analysis method based on GCN in view of the above deficiencies of the existing technologies. It is used to solve the technical problem that in the existing technologies, the sentiment information related to aspect words in different modalities cannot be fully extracted, resulting in a large amount of noise information in the extracted features, and the accuracy of aspect-level multimodal sentiment polarity classification is relatively low.

[0007] The idea to achieve the object of the present invention is that the present invention uses a graph convolutional neural network to incorporate syntactic information into text features, enabling aspect words to correctly focus on the sentiment information expressed by correct opinion words, thereby solving the problem that the existing technology cannot accurately extract text features related to aspect words and introduces a large amount of noise information, resulting in poor aspect-level multi-modal sentiment polarity classification effect. The present invention divides an image into image patches and extracts the fine-grained features of each image patch, and uses an attention mechanism to fuse the features of the text modality and the image modality. Finally, a graph convolutional neural network is used to aggregate the associated fused features, thereby solving the problem that the existing technology cannot accurately extract image features related to aspect words in an image.

[0008] The technical solution adopted by the present invention includes the following steps:

[0009] Step 1, generating a training set:

[0010] Step 1.1, selecting at least 3000 pieces of multi-modal data, extracting the syntactic parse tree of the text modality in each piece of multi-modal data, and constructing a text adjacency matrix through the syntactic parse tree; evenly dividing the images in the multi-modal data into 16×16 image patches; forming a training set with all the multi-modal data, the text adjacency matrix, the image patches, and the true sentiment labels.

[0011] Step 2, constructing an aspect-level multi-modal sentiment analysis network:

[0012] Step 2.1, constructing a text modality processing sub-network composed of a text feature extraction module, a text graph convolutional module, and an aspect text feature aggregation module connected in series, and the output vector dimension of this sub-network is 768; the text feature extraction module adopts a BERT network structure, the text graph convolutional module is composed of a graph convolutional neural network, and this graph convolutional neural network has 3 layers, and the aspect text feature aggregation module is implemented by 1 layer of average pooling.

[0013] Step 2.2, constructing a fused modality processing sub-network composed of an image feature extraction module, a text-image modality fusion module, a fused graph convolutional module, and an aspect fused feature aggregation module, and the output vector dimension of this sub-network is 768; the image feature extraction module adopts a ViT network structure, the text-image modality fusion module is implemented by 1 layer of cross-modal Transformer layer, the fused graph convolutional module is composed of a fused adjacency matrix calculation layer and a graph convolutional neural network connected in series, and this graph convolutional neural network has 3 layers, and the aspect fused feature aggregation module is implemented by 1 layer of average pooling layer.

[0014] Step 2.3, constructing a sentiment prediction sub-network composed of a feature aggregation layer and 1 layer of fully connected layer, and the output vector dimension of this sub-network is 3.

[0015] Step 2.4, cascade the text modality processing sub-network and the fusion modality processing sub-network into an aspect-level sentiment feature extraction group, and then connect it in series with the sentiment prediction sub-network to form an aspect-level multi-modal sentiment analysis network;

[0016] Step 3, train the aspect-level multi-modal sentiment analysis network:

[0017] Input the training set into the aspect-level multi-modal sentiment analysis network. The text feature extraction module and the image feature extraction module perform forward propagation to extract text features and image features respectively; the text-image modality fusion module performs forward propagation to obtain fusion features; the text graph convolution module and the fusion graph convolution module perform forward propagation to obtain graph-based text features and graph-based fusion features respectively; the aspect text aggregation module and the aspect fusion feature aggregation module perform forward propagation to obtain aspect-based text features and aspect-based fusion features respectively; the sentiment prediction sub-network performs forward propagation to obtain predicted sentiment labels; use the cross-entropy loss function to calculate the loss value between the predicted sentiment labels and the true sentiment labels, and use the gradient descent method to train the randomly generated initial network weights, iteratively update the network parameters until the network loss function converges, and obtain the trained aspect-level multi-modal sentiment analysis network;

[0018] Step 4, classify the aspect-level multi-modal sentiment polarity:

[0019] For the multi-modal data to be analyzed, respectively obtain the corresponding text adjacency matrix and image patches according to the method in Step 1, and input them into the trained aspect-level multi-modal sentiment analysis network respectively, and output the sentiment polarity corresponding to the aspect words.

[0020] Compared with the prior art, the present invention has the following advantages:

[0021] First, the present invention uses a graph convolutional neural network to integrate syntactic information into text features, enabling aspect words to correctly focus on the sentiment information expressed by correct opinion words, overcoming the problem that existing aspect-level multi-modal sentiment analysis methods cannot accurately extract text features related to aspect words and introduce a large amount of noise information, resulting in poor aspect-level multi-modal sentiment polarity classification effect. Therefore, the present invention has the advantage of being able to accurately extract the sentiment information related to aspect words in the text.

[0022] Second, the present invention divides the image into image patches and extracts the fine-grained features of each image patch, and uses the attention mechanism to fuse the text modality and the image modality features, and finally aggregates the relevant fusion features, overcoming the problem that the prior art cannot accurately extract the sentiment information related to aspect words in the image. Therefore, the present invention reduces the influence of noise in the image modality, and the aspect-level multi-modal sentiment analysis network constructed by the present invention can accurately extract the sentiment features related to aspect words in the image modality. Description of the Drawings

[0023] Figure 1 is the implementation flow chart of the present invention.

[0024] Figure 2 is the schematic diagram of the network structure of the present invention. Specific implementation manners

[0025] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments.

[0026] Referring to Figure 1 further describes the implementation steps of the embodiments of the present invention.

[0027] Step 1, generate a training set.

[0028] Step 1.1, select at least 3000 pieces of multimodal data, extract the syntactic parse tree of the text modality in each piece of multimodal data, and construct a text adjacency matrix through the syntactic parse tree; evenly divide the images in the multimodal data into 16×16 image patches; form a training set with all the multimodal data, the text adjacency matrix, the image patches, and the true sentiment labels.

[0029] The construction of the text adjacency matrix through the syntactic parse tree means that a probability matrix of dependency arcs is obtained by using a dependency analyzer, and related words are connected to facilitate the aspect word to pay attention to the information expressed by the correct opinion word, so as to capture rich sentiment information. Compared with the discrete output as the adjacency matrix, the probability matrix can alleviate the problem of dependency parsing errors, and thus can capture rich sentiment information.

[0030] The multimodal data consists of text, images, and aspect words, where the aspect words are a subset of the corresponding text.

[0031] Step 2, construct an aspect-level multimodal sentiment analysis network.

[0032] Step 2.1, construct a text modality processing sub-network composed of a text feature extraction module, a text graph convolution module, and an aspect text feature aggregation module connected in series. The output vector dimension of this sub-network is 768; the text feature extraction module adopts the BERT network structure, the text graph convolution module is composed of a graph convolutional neural network, which has 3 layers in total, and the aspect text feature aggregation module is implemented by 1 layer of average pooling.

[0033] Step 2.2, construct a fused modality processing sub-network composed of an image feature extraction module, a text-image modality fusion module, a fused graph convolution module, and an aspect fusion feature aggregation module. The output vector dimension of this sub-network is 768; the image feature extraction module adopts a ViT network structure, the text-image modality fusion module is implemented by 1 cross-modal Transformer layer, the fused graph convolution module is composed of a fused adjacency matrix calculation layer in series with a graph convolutional neural network, and this graph convolutional neural network has 3 layers in total. The aspect fusion feature aggregation module is implemented by 1 average pooling layer.

[0034] The cross-modal Transformer layer is composed of two sub-layers in series. The structure of the first sub-layer is composed of a multi-head attention sub-layer, a feature aggregation layer, and layer normalization in series. The number of heads of the multi-head attention sub-layer is 4. The structure of the second sub-layer is composed of a feed-forward neural network sub-layer, a feature aggregation layer, and layer normalization in series.

[0035] The fused adjacency matrix calculation layer is implemented by the following formula:

[0036]

[0037]

[0038] where s i,j represents and 's cosine similarity, and represent the i-th and j-th feature vectors in the fused features respectively. S represents the similarity matrix formed by combining the similarities between every two feature vectors. A t represents the text adjacency matrix, represents the Hadamard product calculation, and || ||2 represents the L2 normalization operation. By calculating the cosine similarity between every two fused features, the similarity between image patches can be indirectly obtained. Calculating the Hadamard product between the feature vector similarity matrix and the text adjacency matrix gives the fused feature adjacency matrix. Then, using the graph convolutional neural network to aggregate the related features can extract the complete image scene information related to the aspect word.

[0039] Step 2.3, construct an emotion prediction sub-network composed of a feature aggregation layer and 1 fully connected layer. The output vector dimension of this sub-network is 3.

[0040] Step 2.4, cascade the text modality processing sub-network and the fused modality processing sub-network into an aspect-level emotion feature extraction group, and then connect it in series with the emotion prediction sub-network to form an aspect-level multi-modal emotion analysis network.

[0041] Step 3, train the aspect-level multi-modal emotion analysis network.

[0042] The training set is input into the aspect-level multi-modal sentiment analysis network. The text feature extraction module and the image feature extraction module perform forward propagation to extract text features and image features respectively. The text-image modality fusion module performs forward propagation to obtain fused features. The text graph convolution module and the fused graph convolution module perform forward propagation respectively to obtain graph-based text features and graph-based fused features. The aspect text aggregation module and the aspect fused feature aggregation module perform forward propagation to obtain aspect-based text features and aspect-based fused features. The sentiment prediction sub-network performs forward propagation to obtain predicted sentiment labels. The cross-entropy loss function is used to calculate the loss value between the predicted sentiment labels and the true sentiment labels, and the gradient descent method is used to train the randomly generated initial weights of the network, iteratively updating the network parameters until the network loss function converges, and a trained aspect-level multi-modal sentiment analysis network is obtained.

[0043] The text-image fusion features are obtained through the following calculation process:

[0044]

[0045]

[0046] H m = LN(H l + MHCA(H l , H p ))

[0047] H f = LN(H m + MLP(H m ))

[0048] where represents the feature vector calculated by the i-th attention head, softmax represents the softmax activation function, H l and H p represent text features and image features respectively, W Q , W K , W V and W C represent randomly selected different weight parameters, d represents the dimension of the feature vector, T represents the transpose operation, MHCA(H l , H p ) represents the feature vector output by the multi-head attention sub-layer for H l and H p , concat represents the concatenation operation, H m represents the feature vector obtained by performing layer normalization on the feature vector calculated by the multi-head attention sub-layer and the text feature vector, LN represents the layer normalization operation, Hf denotes the fused feature, and MLP denotes the feedforward neural network.

[0049] The forward propagation processes of the text graph convolution module and the fused graph convolution module to obtain the graph-based text feature and the graph-based fused feature are as follows:

[0050]

[0051]

[0052] Among them, and respectively denote the graph-based text feature and the graph-based fused feature of the i-th word, and denote the accumulation of graph convolution operations, K t and K f respectively denote the number of graph convolution times of the text graph convolution module and the fused graph convolution module, σ denotes the sigmoid activation function, N i denotes the set of subscripts of all words associated with the i-th word, and respectively denote the association degree between the i-th and the j-th in the text adjacency matrix and the fused adjacency matrix, W t and W f denote randomly generated different weight parameters, and respectively denote the text feature and the fused feature of the j-th word, b t and b f denote the bias parameters.

[0053] The forward propagation processes of the aspect text aggregation module and the aspect fused feature aggregation module to obtain the aspect-based text feature and the aspect-based fused feature are to aggregate the graph-based text feature and the graph-based fused feature corresponding to the aspect words.

[0054] The cross-entropy loss function is as follows:

[0055]

[0056] Among them, denotes the loss value between the predicted sentiment label and the true sentiment label, D denotes the number of samples in the training set, p denotes the probability distribution, y denotes the true sentiment label, denotes the predicted sentiment label.

[0057] Step 4, classify the aspect-level multi-modal sentiment polarity:

[0058] The multi-modal data to be analyzed are respectively used to obtain the corresponding text adjacency matrix and image patches according to the method in step 1, and are respectively input into the trained aspect-level multi-modal sentiment analysis network to output the sentiment polarity corresponding to the aspect words.

Claims

1. An aspect-level multimodal sentiment analysis method based on GCN, characterized in that, Construct and train an aspect-level multi-modal sentiment analysis model; the steps of the analysis method are as follows: Step 1, generate a training set: Select at least 3000 multi-modal data, extract the syntactic parse trees of the text modality in each multi-modal data, and construct a text adjacency matrix through the syntactic parse trees; evenly divide the images in the multi-modal data into 16×16 image patches; combine all the multi-modal data, text adjacency matrix, image patches and real sentiment labels to form a training set; Step 2, construct an aspect-level multi-modal sentiment analysis network: Step 2.1, construct a text modality processing sub-network composed of a text feature extraction module, a text graph convolution module, and an aspect text feature aggregation module in series, and the output vector dimension of this sub-network is 768; the text feature extraction module adopts a BERT network structure, the text graph convolution module is composed of a graph convolutional neural network, and this graph convolutional neural network has 3 layers, and the aspect text feature aggregation module is implemented by 1 layer of average pooling; Step 2.2, construct a fusion modality processing sub-network composed of an image feature extraction module, a text-image modality fusion module, a fusion graph convolution module, and an aspect fusion feature aggregation module, and the output vector dimension of this sub-network is 768; the image feature extraction module adopts a ViT network structure, the text-image modality fusion module is implemented by 1 layer of cross-modal Transformer layer, the fusion graph convolution module is composed of a fusion adjacency matrix calculation layer and a graph convolutional neural network in series, and this graph convolutional neural network has 3 layers, and the aspect fusion feature aggregation module is implemented by 1 layer of average pooling layer; Step 2.3, construct a sentiment prediction sub-network composed of a feature aggregation layer and 1 layer of fully connected layer, and the output vector dimension of this sub-network is 3; Step 2.4, cascade the text modality processing sub-network and the fusion modality processing sub-network into an aspect-level sentiment feature extraction group, and then connect it in series with the sentiment prediction sub-network to form an aspect-level multi-modal sentiment analysis network; Step 3, train the aspect-level multi-modal sentiment analysis network: Input the training set into the aspect-level multi-modal sentiment analysis network, the text feature extraction module and the image feature extraction module perform forward propagation to extract text features and image features respectively; the text-image modality fusion module performs forward propagation to obtain fusion features; the text graph convolution module and the fusion graph convolution module perform forward propagation to obtain graph-based text features and graph-based fusion features respectively; the aspect text aggregation module and the aspect fusion feature aggregation module perform forward propagation to obtain aspect-based text features and aspect-based fusion features respectively; the sentiment prediction sub-network performs forward propagation to obtain predicted sentiment labels; use the cross-entropy loss function to calculate the loss value between the predicted sentiment labels and the real sentiment labels, and use the gradient descent method to train the randomly generated initial network weights, iterate and update the network parameters until the network loss function converges, and obtain the trained aspect-level multi-modal sentiment analysis network; Step 4, classify the aspect-level multi-modal sentiment polarity: The multi-modal data to be analyzed is respectively used to obtain the corresponding text adjacency matrix and image patches according to the method in step 1, and are respectively input into the trained aspect-level multi-modal sentiment analysis network to output the sentiment polarity corresponding to the aspect words.

2. The aspect-level multi-modal sentiment analysis method based on GCN according to claim 1, wherein The construction of the text adjacency matrix through the syntactic parse tree in step 1 refers to using a dependency analyzer to obtain the probability matrix of dependency arcs, connecting related words, facilitating the aspect words to focus on the information expressed by the correct opinion words, and capturing rich sentiment information.

3. The aspect-level multi-modal sentiment analysis method based on GCN according to claim 1, characterized in that, The multi-modal data described in step 1 consists of text, images, and aspect words.

4. The aspect-level multi-modal sentiment analysis method based on GCN according to claim 1, characterized in that The cross-modal Transformer layer described in step 2.2 is composed of two sub-layers in series. The structure of the first sub-layer is composed of a multi-head attention sub-layer, a feature aggregation layer, and layer normalization in series. The number of heads of the multi-head attention sub-layer is 4. The structure of the second sub-layer is composed of a feed-forward neural network sub-layer, a feature aggregation layer, and layer normalization in series.

5. The aspect-level multi-modal sentiment analysis method based on GCN according to claim 1, characterized in that The fused adjacency matrix calculation layer described in step 2.2 is implemented by the following formula: Among them, s i,j represents and 's cosine similarity. and respectively represent the i-th and j-th feature vectors in the fused features. S represents the similarity matrix formed by combining the similarities between every two feature vectors. A t represents the text adjacency matrix. represents the Hadamard product calculation, and || ||2 represents the two-norm normalization operation.

6. The aspect-level multi-modal sentiment analysis method based on GCN according to claim 1, characterized in that The text-image fusion feature described in step 3 is obtained through the following calculation process: H m = LN(H l + MHCA(H l ,H p )) H f = LN(H m + MLP(H m )) Among them, represents the feature vector calculated by the i-th attention head, softmax represents the softmax activation function, and H l and H p represent the text feature and the image feature respectively, and W Q , W K , W V and W C represent different randomly selected weight parameters respectively, d represents the dimension of the feature vector, represents the transpose operation, and MHCA(H l , H p ) represents the feature vector output by the multi-head attention sub-layer for H l and H p , concat represents the concatenation operation, and H m represents the feature vector obtained by performing layer normalization on the feature vector calculated by the multi-head attention sub-layer and the text feature vector, LN represents the layer normalization operation, and H f represents the fused feature, and MLP represents the feed-forward neural network.

7. The aspect-level multi-modal sentiment analysis method based on GCN according to claim 1, characterized in that, The forward propagation of the text graph convolution module and the fused graph convolution module described in step 3 to obtain the graph-based text feature and the graph-based fused feature respectively is as follows: Among them, and respectively represent the graph-based text feature and graph-based fusion feature of the i-th word, and represent the accumulation of graph convolution operations, K t and K f respectively represent the number of graph convolution times of the text graph convolution module and the fusion graph convolution module, σ represents the sigmoid activation function, N i represents the set of subscripts of all words associated with the i-th word, and respectively represent the degree of association between the i-th and j-th in the text adjacency matrix and the fusion adjacency matrix, W t and W f represent randomly generated different weight parameters, and respectively represent the text feature and fusion feature of the j-th word, b t and b f represent bias parameters.

8. The aspect-level multimodal sentiment analysis method based on GCN according to claim 1, characterized in that, The forward propagation of the aspect text aggregation module and the aspect fused feature aggregation module described in step 3 to obtain the aspect-based text feature and the aspect-based fused feature is achieved by aggregating the graph-based text feature and the graph-based fused feature corresponding to the aspect words.

9. The aspect-level multi-modal sentiment analysis method based on GCN according to claim 1, characterized in that The cross-entropy loss function described in step 3 is as follows: Among them, represents the loss value between the predicted sentiment label and the true sentiment label, D represents the number of samples in the training set, p represents the probability distribution, y represents the true sentiment label, represents the predicted sentiment label.

Citation Information

Patent Citations

  • Aspect-level multi-modal sentiment analysis method based on collaborative attention fusion

    CN115293170A

  • Multi-dimensional fine-grained dynamic sentiment analysis method and system

    CN114154077A

  • Multi-modal emotion recognition method based on multi-head attention and graph neural network

    CN116245102A