A method of sentiment analysis that deeply integrates social content and relationships

By constructing a weighted undirected graph and a multi-layer neural network to integrate social content and relationships, the problem of unutilized social network links in multimodal sentiment analysis is solved, and more accurate sentiment prediction is achieved.

CN118861856BActive Publication Date: 2026-08-25MOUTAI INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410893762.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-04
Publication Date
2026-08-25
Estimated Expiration
2044-07-04

AI Technical Summary

Technical Problem

Existing multimodal sentiment analysis methods fail to effectively utilize the link relationships in social networks, resulting in unsatisfactory sentiment analysis results.

Method used

By constructing a weighted undirected graph, extracting image and text features using ResNet and GloVe word embeddings, combining a Transformer encoder and a graph convolutional neural network, integrating multimodal features of social content and relationships, and employing a memory attention mechanism and a multilayer perceptual neural network for sentiment prediction.

Benefits of technology

It achieves deep fusion of image and text features, improves the accuracy of multimodal sentiment analysis, captures finer-grained feature associations, and significantly improves the accuracy of sentiment prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861856B_ABST
    Figure CN118861856B_ABST
Patent Text Reader

Abstract

The application discloses a kind of deep integration social content and relationship sentiment analysis method of multimodal data analysis technical field, including feature preprocessing, picture-text attention based on Transformer, node-guided picture-text joint attention and four modules of multimodal graph reasoning.Wherein, feature preprocessing module is mainly used to extract and process picture, text and social relationship network features;Picture-text attention based on Transformer module is used to capture fine-grained emotional feature association and complementary information between picture and text;Node-guided picture-text joint attention module and multimodal graph reasoning module explore the influence of social relationship on picture-text sentiment polarity from two aspects of social network topology and neighbor content information respectively.The application comprehensively utilizes the data of three modalities of picture, text and relationship network, and carries out deep integration and reasoning of social content and relationship through cross-modal attention mechanism, graph neural network and other technologies, effectively improves the accuracy of multimodal sentiment analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of modal data analysis technology, specifically to a sentiment analysis method that deeply integrates social content and relationships. Background Technology

[0002] The rise of multimodal sentiment analysis technology is rooted in a profound understanding of the diversity and multidimensionality of human emotional expression. Especially with the rise and widespread adoption of social media, people tend to express their emotional tendencies through various forms such as text, images, and audio. However, traditional sentiment analysis techniques primarily focus on text data, failing to fully explore the rich emotional information hidden in perceptual modalities such as images and audio, thus limiting the comprehensiveness and accuracy of sentiment analysis. Against this backdrop, multimodal sentiment analysis technology has emerged, representing a major innovation in the field of sentiment analysis and a significant expansion of the depth and breadth of our understanding of human emotions. With the rapid development of computer vision and natural language processing technologies, the accurate and efficient acquisition and processing of data from different modalities has become a reality, providing a solid technical foundation for multimodal sentiment analysis and opening up vast possibilities for its performance improvement. The core objective of multimodal sentiment analysis technology is to integrate data from different modalities and deeply explore the emotional connections within them to gain a more comprehensive understanding of the human emotional world. In practical applications, multimodal sentiment analysis technology is gradually demonstrating its unique value. Whether in areas such as brand sentiment insight and election analysis, or in various other application scenarios, it can help us more accurately grasp people's emotional needs and provide more intelligent and personalized services.

[0003] While the explosive growth of multimodal social content on social media has greatly propelled the development of multimodal sentiment analysis (MMA) technology, research in this field remains in its early stages. Typically, information uploaded by users to social networks exists in multiple modalities and may convey rich emotional tendencies. Current MAA methods primarily focus on capturing complementary features between different modalities within social content to effectively and comprehensively understand cross-modal information. Although current methods have achieved significant success in modeling cross-modal sentiment feature associations, they primarily focus on user-uploaded multimodal social content, neglecting the impact of the links and interactions between nodes in the social network on sentiment tendencies. In reality, multimodal social content with linkages often shares a certain degree of emotional commonality, which can be utilized as supplementary information for sentiment analysis. However, due to the lack of social link networks, current methods cannot establish effective emotional connections between social content, resulting in less than ideal sentiment analysis results. Therefore, integrating social relationship networks into MAA to achieve more accurate sentiment prediction has become an urgent and challenging technology. Summary of the Invention

[0004] This invention aims to provide a sentiment analysis method that deeply integrates social content and relationships, in order to solve the problem that existing multimodal sentiment analysis lacks a social link network and cannot establish effective emotional connections between social content.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a sentiment analysis method that deeply integrates social content and relationships, comprising the following steps: S1. Feature Preprocessing: Image information is extracted into an image feature matrix using a ResNet residual neural network; the text words corresponding to the image are encoded using GloVe word embeddings, and the embedding matrix corresponding to the word sequence is input into a bidirectional long short-term memory network Bi-LSTM to output a text feature matrix; a weighted undirected graph is constructed using social attributes, and the graph node embedding features of the fused topology are calculated using the improved DeepWalk algorithm; S2. Transformer-based image and text attention: Multiple Transformer encoders are used to encode the image feature matrix and text feature matrix obtained by the feature preprocessing module respectively. The output matrices of multiple Transformer encoders are concatenated and then input into a new Transformer encoder. The new image feature matrix and text feature matrix output at the corresponding positions are fused to generate a joint image and text feature matrix. S3, Node-guided graph-text joint attention: The graph node features obtained in S1 and the image and text joint features obtained in S2 are used as inputs. The heterogeneous relationship between the two features is calculated using the memory attention mechanism. The topological structure of the relationship graph and the graph-text joint features are fused through this attention mechanism to generate multimodal graph node features. S4. Multimodal Graph Reasoning: Based on the weighted undirected graph constructed in S1, the multimodal graph node features calculated in S3 are used to represent each node on the undirected graph. Multimodal heterogeneous graph reasoning is performed on the undirected graph through a multilayer graph convolutional neural network. The graph node embeddings obtained from the reasoning are used as input to the multimodal sentiment classifier to achieve sentiment polarity prediction.

[0006] Furthermore, the weighted undirected graph described in step S1 is constructed using the social attribute information of the images, and its graph nodes correspond to multimodal image and text information; step S1 creates different links for different attribute co-occurrence relationships between images. Co-occurrence relationships indicate that they have the same tag, were uploaded by the same user, were taken at the same location, etc. Each link relationship is assigned a weight of 1, that is, if there are w link relationships between two images, the corresponding link weight is w.

[0007] Furthermore, the graph node embedding feature vector described in step S1 is calculated by improving the DeepWalk algorithm to make it applicable to the weighted undirected graph constructed in step S1. Using the idea of ​​DeepWalk, the graph node embedding in step S1 is mainly used to represent the topological structure information of nodes in the social relationship network graph.

[0008] Furthermore, step S2 employs two Transformer encoders to encode the image feature matrix and the text feature matrix respectively.

[0009] Furthermore, in step S2, after splitting the output of the last Transformer encoder into new image and text feature matrices, the neural network is used to perform feature concatenation and mapping, and the two types of features are deeply fused to generate joint image and text features.

[0010] Furthermore, the memory attention mechanism used in step S3 generates multimodal node features by iteratively calculating the attention of graph node features to the joint features of images and text multiple times. This is to model the heterogeneous relationship between the local topological structure of graph nodes and the joint information of graph and text, and to capture the influence of the graph's topological structure on sentiment.

[0011] Furthermore, the multimodal sentiment classifier described in step S4 is a multilayer perceptual neural network structure. It takes multimodal graph node embeddings that integrate image, text, network topology and neighbor content information as input and outputs the corresponding multimodal sentiment polarity, i.e., positive, negative or neutral.

[0012] The beneficial effects of this invention are as follows: 1. Based on image and text modalities, this invention introduces the link relationships between images in social networks, deeply exploring the topological structure of social network graphs and the potential impact of neighbor content information on the sentiment tendencies of images and text. In this way, this method can achieve deep fusion of features from three modalities: image, text, and network, maximizing the use of social data information for comprehensive sentiment reasoning and analysis; 2. This invention employs multiple Transformer encoders to deeply learn the feature dependencies within and between image and text modalities. This method can capture more fine-grained feature associations, effectively mine and utilize cross-modal sentiment complementarity information, thereby significantly improving the accuracy of multimodal sentiment analysis. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating a sentiment analysis method that deeply integrates social content and relationships according to the present invention. Detailed Implementation

[0014] The following detailed description illustrates the specific implementation methods: The basic implementation examples are as follows: Figure 1 The following describes a sentiment analysis method that deeply integrates social content and relationships. Given an image, its associated text description, and a relationship network constructed from image attributes, the method predicts the sentiment tendency shared by three modalities: image, text, and network data. The specific implementation steps are as follows: The specific implementation process is as follows: Step 1: Feature Preprocessing For a given image, this invention uses a ResNet network to encode image information into an image region feature matrix. ,in Indicates the number of image regions. Indicates the first The feature vectors corresponding to each image region. For image-related text descriptions, this invention first uses 300-dimensional GloVe word vectors to encode each word in the text, and then uses the word sequence features as input to a Bi-LSTM bidirectional long short memory network, the output of which is used as the text feature matrix. , where represents the number of text segments. Indicates the first The feature vectors of each text segmentation.

[0015] Regarding the social attributes of images, this invention creates different links for different attribute co-occurrence relationships between images (having the same tag, uploaded by the same user, taken at the same location, etc.). Each type of link is assigned a uniform weight of 1. If there are w types of link information between two images, then the edge weight between them is set to w. Thus, this invention constructs a weighted undirected graph. Here, V represents the set of nodes in the graph, and E represents the set of edges. Each node in the graph contains a specific image and its corresponding descriptive text. To incorporate the topological structure of the network graph G into the embedded representation of graph nodes, this invention improves the DeepWalk random walk algorithm, making it applicable to weighted networks. First, a point is selected from graph G. The method begins a random walk by selecting a neighboring node as the next node to move to, and this process continues until the walk reaches a maximum length of m. In the context of a weighted graph, the probability of selecting a node is proportional to the weight of the edge connecting it. Thus, this walk method produces a sequence of nodes. For nodes in the node sequence Generate a context sequence Where l represents the window size. This invention uses Skip-Gram as the objective function to maximize the number of target nodes within the prediction window. This process can be described as follows: .in, Defined by the softmax objective equation, i.e.

[0016] In the formula, This represents the embedding vector of a graph node and is also a product of model learning. Using the improved DeepWalk algorithm, a node embedding matrix Z of a graph G that incorporates the topological structure can be trained. ).

[0017] Step 2: Transformer-based image and text attention In social media content, images often appear alongside their corresponding descriptive text, complementing each other to convey the user's emotions. To fully capture the multimodal sentiment tendencies contained in images and text, this invention utilizes a Transformer encoder to capture the feature dependencies within and between image and text modalities, thereby learning deep-level cross-modal sentiment complementarity information. First, a fully connected neural network is used to map the image feature matrix R and the text feature matrix T to a shared space, i.e. , .in, and The weight matrix is ​​a learnable matrix. and For learnable bias terms, It is the ReLU non-linear activation function. Next, the image feature matrices are... and text feature matrix The input is fed into two different Transformer encoders to encode fine-grained dependencies between features within each modality. The core of the Transformer encoder is processing the input features through a multi-head self-attention mechanism. This self-attention mechanism uses a fully connected neural network to map the input feature matrix X to a set of query matrices. Key matrix Sum matrix The input matrix X is the image feature matrix. or text feature matrix Self-attention distribution The calculation method is as follows: Where s is and Corresponding shared space dimension. Self-attention mechanism output. In this invention, the multi-head self-attention fusion and subsequent feature processing are consistent with the original Transformer model. Thus, two Transformer encoders process image and text features respectively, outputting corresponding feature matrices. and Next, in order to model the cross-modal sentiment feature associations between images and text, we will... and Concatenate into a new image and text feature matrix ,in This represents feature concatenation. The concatenated feature matrix. The input is fed into a new Transformer encoder to model the sentiment association between the two modalities through a self-attention mechanism, capture cross-modal sentiment complementarity information, and output a feature matrix. To obtain image and text features associated with heterogeneous modalities, this invention will... Matrix according to The concatenation is reversed and decomposed into two matrices. and It is worth noting that from the input matrix and To output matrix and Throughout the entire calculation process, this invention iteratively executes the calculation P times, with P ultimately taking the value of 3. Through this method, the model can deeply capture the feature correlations within and between image and text modalities, thereby fully exploring the complementary emotional information between images and text. Finally, this invention uses the iteratively calculated feature matrix... and Merged into a feature matrix .in, and These are the learnable weight matrix and bias terms. This represents the image-text joint features learned from a Transformer-based image-text attention network, which contain cross-modal sentiment information.

[0018] Step 3: Node-guided joint attention of text and graph In relational networks, social content nodes with similar topologies typically exhibit consistent sentiment tendencies. To incorporate the influence of network topology into multimodal sentiment analysis, it is crucial to thoroughly and effectively integrate the topological information of social nodes into the analysis of multimodal sentiment tendencies. Based on this, this invention designs a node-guided graph-text joint attention network, aiming to mine the heterogeneous relationship between the topological structure of graph nodes and image-text joint features, thereby integrating the graph's topological information into multimodal sentiment analysis. For a social relational graph node, step one learns a node embedding vector containing the corresponding topological structure. Step two learned the joint image-text features at the nodes. Based on this, the present invention employs a memory network to model the heterogeneous relationship between the two. First, a fully connected neural network is used to connect the matrices... Convert to a memory matrix of the same size Then, by calculating the node vector and memory matrix The similarity between them is used to obtain node-guided joint attention weights for the graph and text. Finally, The attention weights are multiplied, and the memory matrix is ​​updated through another fully connected layer. To generate a clearer attention distribution focused on important emotional information, this process is iterated H times, with H ultimately set to 3. The entire process can be expressed by the formula: superscript Indicates the first Number of iterations, Represents the attention weight vector. It is an update operation; ⊙ indicates element-wise multiplication after broadcasting on the appropriate axis. and These are the learnable parameters in the fully connected layer. After rounds of iterative computation, the output matrix of the memory network This can be viewed as the result of a deep interaction between image-text joint features and node topology information. This invention further utilizes neural networks to... Perform the following calculations: .in, ReLU represents the nonlinear activation function. and These represent the learnable weight matrix and bias term, respectively. This indicates that the features embedding images, text, graph node topology, and the close relationships between them are multimodal node features.

[0019] Step 4: Multimodal Graph Reasoning In social networks, the content of two linked graph nodes often exhibits similar sentiment tendencies. For example, in the context of the same wedding event, multimodal graph and text information uploaded by different users is likely to convey positive emotions. Therefore, when analyzing the multimodal sentiment tendencies of network nodes, it is essential to use the multimodal content of neighboring nodes as supplementary information. To this end, this invention designs and implements a multimodal graph inference network, aiming to fully integrate the multimodal content information of neighboring nodes in social networks, thereby improving the accuracy of sentiment analysis. Since step three represents each multimodal graph node as a vector... Then the embedding representations of all N nodes in the graph together form the node embedding matrix. Additionally, step one represents the social relationship network as a weighted undirected graph. ,in This represents the weight matrix of the edges in the network graph. Building upon this, the present invention further utilizes a graph convolutional neural network to perform cross-modal sentiment reasoning on graph G. The graph convolutional neural network consists of... Layer graph convolution operations consist of... The final value is 2. The graph convolution operation of a layer can be represented as: .in, Indicates the first The hidden feature matrix of the nodes in the layer, Represents the original node embedding matrix . Representation diagram The augmented adjacency matrix is ​​obtained by adding the identity matrix to the original adjacency matrix. This represents the corresponding degree matrix. Then it means the first The trainable convolution weight matrix in the graph convolution operation. The result of the last graph convolution, i.e., the... Graph node embedding matrix obtained by convolution This is used for subsequent sentiment classification. By utilizing multi-layer graph convolution operations, this invention incorporates multimodal content information from neighbors into the embedding learning of graph nodes, enabling the node features in the graph to cover a wider range of elements influencing sentiment tendencies. Thus, the node embedding matrix... This invention integrates images, text, and the topology and neighbor content of social relationship networks from social networks. Finally, it employs a multi-layer fully connected neural network to construct a multimodal sentiment classifier, with the final layer using softmax as the activation function and embedding each node... The data is used as input to the classifier. The three candidate classification options correspond to the three sentiments (positive, neutral, and negative) contained in the graph nodes. During training, cross-entropy is used as the loss function for model training. In actual sentiment prediction, the candidate with the highest classification probability is used as the result of the multimodal sentiment prediction.

[0020] The above descriptions are merely embodiments of the present invention, and common knowledge regarding specific structures and characteristics is not elaborated upon here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the structure of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications will not affect the effectiveness of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A sentiment analysis method that deeply integrates social content and relationships, characterized in that, Includes the following steps: S1. Feature Preprocessing: Image information is extracted into an image feature matrix using a ResNet residual neural network; GloVe word embeddings are used to encode the corresponding text words in the image, and the embedding matrix corresponding to the word sequence is input into a Bi-LSTM network to output the text feature matrix; a weighted undirected graph is constructed using social attributes, and an improved DeepWalk algorithm is used to calculate the graph node embedding features that fuse the topological structure. Specifically, a node vi is selected from the graph G as the starting point of the random walk, and a neighboring node is selected as the next node in the walk. This process continues until the walk reaches the maximum length m. In the context of a weighted graph, the probability of selecting a node is proportional to the weight of the edge connecting it. Thus, this walk algorithm generates a node sequence {n1, ..., n}. i ,...,n m }, for node n in the node sequence i Generate a context sequence {n i-l ,…,n i+l }, where 1 represents the window size. Then, the Skip-Gram model is used to maximize the predicted probability of the target node within the window as the loss function, and the node embedding matrix Z of the fused topology structure is obtained by training. The embedding vector of each node is the product of model learning. S2. Transformer-based image and text attention: Multiple Transformer encoders are used to encode the image feature matrix and text feature matrix obtained by the feature preprocessing module respectively. The output matrices of multiple Transformer encoders are concatenated and then input into a new Transformer encoder. The new image feature matrix and text feature matrix output at the corresponding positions are fused to generate a joint image and text feature matrix. S3, Node-guided graph-text joint attention: The graph node embedding features obtained in S1 and the image and text joint features obtained in S2 are used as inputs. The heterogeneous relationship between the two features is calculated using the memory attention mechanism. The topological structure of the relationship graph and the graph-text joint features are fused through this attention mechanism to generate multimodal graph node features. S4. Multimodal Graph Reasoning: Based on the weighted undirected graph constructed in S1, the multimodal graph node features calculated in S3 are used to represent each node on the undirected graph. Multimodal heterogeneous graph reasoning is performed on the undirected graph through a multilayer graph convolutional neural network. The graph node embeddings obtained from the reasoning are used as input to the multimodal sentiment classifier to achieve sentiment polarity prediction.

2. The sentiment analysis method for deep integration of social content and relationships according to claim 1, characterized in that: The weighted undirected graph described in step S1 is constructed using the social attribute information of images, and its graph nodes correspond to multimodal image and text information. Step S1 creates different links for different attribute co-occurrence relationships between images. Co-occurrence relationship means that they have the same tag, are uploaded by the same user, and are taken at the same location. Each link relationship is assigned a weight of 1, that is, if there are w link relationships between two images, the corresponding link weight is w.

3. The sentiment analysis method for deep integration of social content and relationships according to claim 2, characterized in that: The graph node embedding feature vector described in step S1 is calculated by improving the DeepWalk algorithm to make it applicable to the weighted undirected graph constructed in step S1. Using the idea of ​​DeepWalk, the graph node embedding in step S1 is used to represent the topological structure information of nodes in the social relationship network graph.

4. A sentiment analysis method for deep integration of social content and relationships according to claim 3, characterized in that: Step S2 uses two Transformer encoders to encode the image feature matrix and the text feature matrix respectively.

5. A sentiment analysis method for deep integration of social content and relationships according to claim 4, characterized in that: In step S2, after splitting the output of the last Transformer encoder into new image and text feature matrices, the neural network is used to perform feature concatenation and mapping, and the two types of features are deeply fused to generate joint image and text features.

6. The sentiment analysis method for deep integration of social content and relationships according to claim 5, characterized in that: The memory attention mechanism used in step S3 generates multimodal node features by iteratively calculating the attention of graph node features to the joint features of images and text. This model models the heterogeneous relationship between the local topological structure of graph nodes and the joint information of graph and text, capturing the influence of the graph's topological structure on sentiment.

7. The sentiment analysis method for deep integration of social content and relationships according to claim 6, characterized in that: The multimodal sentiment classifier described in step S4 is a multilayer perceptual neural network structure. It takes multimodal graph node embeddings that integrate image, text, network topology and neighbor content information as input and outputs the corresponding multimodal sentiment polarity, i.e., positive, negative or neutral.

Citation Information

Patent Citations

  • Multi-modal sentiment classification method based on progressive neural network

    CN115795020A

  • Aspect-level multi-modal sentiment analysis method based on double channels and attention mechanism

    CN116662924A