Multimodal Sentiment Analysis Method and System Integrating Multi-Granularity Visual and Textual Features

By adopting a method of fusion of multi-grained vision and text features in multi-modal sentiment analysis, combined with graph attention network and ANP parser, the lack of noise problems and aspects of word emotional information enhancement in multi-modal data is solved, and the accuracy of emotional polarity prediction is improved.

CN116776287BActive Publication Date: 2025-06-13FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310794745.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-06-13
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

The existing multimodal sentiment analysis technology has noise problems caused by information irrelevance or redundant information when processing multimodal data, and lacks a strategy to enhance multimodal information for the aspect words, which affects the prediction accuracy of emotional polarity.

Method used

A multimodal emotion analysis method that integrates multi-grained vision and text features is adopted to extract image features through pre-training models and ResNet, and a graph attention network is constructed by combining syntactic dependencies and component tree structures. Multi-layer graph attention network and ANP parser are used to fusion of multi-grained feature and expression of aspect words to reduce noise and enhance the emotional information of aspect words.

Benefits of technology

The accuracy of emotional polarity prediction is improved, the noise problem in multimodal information extraction is reduced, and the accuracy of sentiment analysis is improved through multimodal information enhancement for aspect words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116776287B_ABST
    Figure CN116776287B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-modal sentiment analysis method and system that integrates multi-granularity visual and text features. The method includes: A. Initializing the text representation of graph nodes and extracting the feature representation of pictures; B. Determining the edge relation adjacency matrix in the graph attention network according to the syntactic dependency relationship and the constituent tree structure; using the multi-modal attention mechanism to obtain the joint text-visual feature representations at the word level, phrase level, and sentence level respectively, and finally obtaining the multi-granularity text-visual fusion feature representation; predicting the position of the aspect word according to the text-visual joint feature representation output by the multi-layer graph attention network and forming the aspect word representation; parsing out the most relevant ANP pair according to the aspect word position using the ANP parser, and constructing an aspect graph according to the adjacency relationship of the aspect word in the syntactic dependency relationship; C. Using the output of the aspect graph as the sentiment representation of the aspect word to predict the sentiment polarity corresponding to the aspect word. This method and system are beneficial to improving the accuracy of sentiment polarity prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing, and particularly relates to a multi-modal sentiment analysis method and system that integrates multi-granularity visual and text features. Background Art

[0002] Sentiment analysis aims to effectively analyze and mine the sentiment information in data, and is an important task in the field of natural language processing. Different from text-level sentiment analysis, aspect-level sentiment analysis does not simply analyze the entire document or sentence and identify its sentiment polarity, but predicts the position of the aspect word and its corresponding different sentiment polarities respectively according to multiple aspects of the entity described in the sentence. For example, given the comment "The price is reasonable, but the service is very poor", the aspect words are "price" and "service" respectively, and the corresponding sentiment polarities are positive and negative. Multi-modal sentiment analysis refers to finding key sentiment information by switching between different modalities based on multi-modal data including text, pictures, videos, and audio, and correlating them with each other, so as to predict more accurate aspect words and their corresponding sentiment polarities. Compared with the sentiment analysis problem limited to text, multi-modal data provides a more diverse information source for the sentiment analysis task. In addition, when humans process and analyze sentiment, they mainly rely on their own sentiment reasoning ability to switch between different modalities, find key sentiment information, and correlate them with each other. Therefore, sentiment analysis in a multi-modal context is closer to the human reasoning mode. However, extracting sentiment information only through text will ignore the information of other modalities, and if the multi-modal information is not properly fused, it will introduce noise. Therefore, the extraction and fusion methods of sentiment features for different modalities are crucial for the final sentiment word extraction and sentiment polarity prediction tasks.

[0003] In recent years, the key to processing fine-grained sentiment expressions in multimodal data is to find important aspect-related information from different modalities (such as text and images), and then utilize the connections between different modalities to interact with this information to help the model further identify sentiment information. Therefore, in order to better explore the connection between the text modality and the image modality, researchers have adopted deep neural networks (such as CNN, RNN, GNN) combined with attention mechanisms to solve this problem, and have also achieved remarkable results. However, there are still certain limitations in the above work. For example, the noise problem caused by irrelevant or redundant information in different modalities. Most studies directly extract the information of pictures using pre-trained models and then interact with the original text information, which inevitably leads to the introduction of picture information noise. One of the effective ways to reduce noise is to reduce the extraction granularity of multimodal information. However, most studies have not considered this point. Secondly, most existing multimodal models for image-text data ignore the role of aspect words. Existing studies capture the global features of text and image information in the entire dataset by introducing multi-channel graph attention networks and other methods, or apply multi-head attention mechanisms to deeply fuse the information between text and image. However, for the multimodal aspect sentiment analysis task, there is a lack of strategies and methods for enhancing the use of multimodal information for aspect words, ignoring the special role of aspect words in multimodal tasks. Different aspect words have different sentiment polarities, and different, aspect-specific sentiment information should also be extracted from the picture information for different aspect words. Thirdly, since syntactic dependency information represents the syntactic dependency relationship between words in a sentence, it has been proven effective by predecessors and is often used in conjunction with graph attention networks, taking the syntactic dependency relationship as the connection relationship between different nodes in the graph. However, this method is prone to introducing sentiment noise, resulting in the spread of incorrect sentiment information and affecting the final sentiment prediction accuracy.

[0004] In summary, the graph attention network has achieved certain achievements in fusing text and picture information, but there are still deficiencies in controlling picture information noise and enhancing the use of multimodal information for aspect words, and it is easy to introduce and generate noise due to syntactic dependency information during the composition process. By analyzing the human emotion judgment process, it can be seen that people will first look at the text information when reading, and then extract aspect information in a targeted manner, and combine the information related to the key information in the picture to obtain the final correct answer. Summary of the Invention

[0005] The purpose of the present invention is to provide a multimodal sentiment analysis method and system that fuse multi-granularity visual and text features, which is beneficial to improving the accuracy of sentiment polarity prediction.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is: a multimodal sentiment analysis method that fuses multi-granularity visual and text features, including the following steps:

[0007] Step A: Initialize the text representation of graph nodes using a pre-trained model, and preliminarily extract the image feature representation using ResNet;

[0008] Step B: Determine the edge relation adjacency matrix in the graph attention network according to the syntactic dependency relationship and the constituent tree structure; use the multi-modal attention mechanism to obtain the joint text-visual feature representations at the word level, phrase level, and sentence level respectively; fuse the joint text-visual feature representations at the word level, phrase level, and sentence level with the help of a multi-layer graph attention network to finally obtain the multi-granularity text-visual fusion feature representation; predict the aspect word position according to the text-visual joint feature representation output by the multi-layer graph attention network and form the aspect word representation; parse out the most relevant ANP pair using the ANP parser according to the aspect word position, and construct the aspect graph according to the adjacency relationship of the aspect word in the syntactic dependency relationship;

[0009] Step C: Predict the sentiment polarity corresponding to the aspect word according to the output of the aspect graph as the aspect word sentiment representation.

[0010] Further, step B specifically includes the following steps:

[0011] Step B1: Perform syntactic dependency parsing on the dataset text to obtain the syntactic dependency relationship A between different nodes 1 ; construct the constituent tree structure A based on the text with the help of constituent tree parsing 2 , obtain the text division of different clauses, specifically manifested in word-level, intermediate-level phrase-level, and sentence-level texts Use the syntactic dependency relationship A 1 and the constituent tree structure A 2 to fuse and obtain the edge relation adjacency matrices of different layers of the graph attention network

[0012] Step B2: Use the edge relation adjacency matrix of the multi-layer graph attention network obtained in step B1 to obtain text representations of different granularities Input into the multi-modal attention network layer by layer to obtain the joint text-visual feature representation specific to this layer Complete the interaction between the multi-granularity text features and the image features; finally output the multi-granularity text-visual fusion feature representation

[0013] Step B3: Use the multi-granularity text-visual fusion feature representation obtained in step B2 Calculate the probabilities p start and p end of the start and end positions of each position being the aspect word respectively, and select the most likely start and end positions of the aspect word according to the probabilities (index start , indexend );

[0014] Step B4: According to the aspect word positions (index start , index end ) marked in Step B3 and the multi-granularity text-visual fusion feature representation obtained in Step B2 obtain the final aspect word representation a i ;

[0015] Step B5: Use the word-level adjacency matrix obtained in Step B1 centered on the aspect word positions (index start , index end ) obtained in Step B3, take the directly connected edges as the edge relationships of the aspect graph G a and construct the adjacency matrix of the aspect graph G a based on this

[0016] Step B6: According to the picture feature representation obtained in Step A, the ANP parser parses out the top-K (noun N, adjective adj a ) pairs most relevant to the picture. According to the aspect word representation a i obtained in Step B4 and its directly connected word representations, calculate the cosine similarity with the nouns in the ANP respectively; and select the K (noun N, adjective adj a ) pairs with the highest cosine similarity as the supplementary nodes of the aspect graph G a to enhance the adjacency matrix in Step B4

[0017] Step B7: Use the graph attention network to obtain the fusion feature representation senti a of the aspect graph G a ; establish a Mask matrix, apply the fusion feature representation senti a of the aspect graph G a to retain the sentiment representation aspect of the aspect word w ; based on the obtained sentiment representation aspect of the aspect word w , calculate the sentiment polarity of the aspect word.

[0018] Furthermore, the specific steps of Step B1 are as follows:

[0019] Step B11: Perform syntactic dependency parsing on the dataset text to obtain the syntactic dependency relationship A between different nodes 1 ; construct a text-based constituent tree structure A by means of constituent tree parsing 2 to obtain the text partitioning of different clauses, specifically manifested at the word level, intermediate-level phrase level, and sentence level

[0020]

[0021]

[0022] Among them, is the numerical value at position (i, j) in syntactic dependency relationship A 1 in the component tree structure A is the numerical value at position (i, j); Dep.Tree represents the syntactic dependency relationship; Con.Tree represents the component tree structure; 2 in the component tree structure A

[0023] Step B12: Taking the component tree clause partitioning strategy obtained in Step B11 as a criterion, eliminate the dependencies across clauses in the syntactic dependency relationship to achieve dependency relationship noise reduction;

[0024] Step B13: Fuse the noise-reduced syntactic dependency relationship A 1 obtained in Step B12 and the adjacency relationship A 2 of each layer of the component tree as the edge relationship adjacency matrix of the graph attention network

[0025]

[0026] Among them, represents the numerical value at position (i, j) in the edge relationship adjacency matrix of the graph attention network.

[0027] Furthermore, the specific steps of Step B2 are as follows:

[0028] Step B21: Use the word representation T obtained by the pre-trained model to interact with the picture, and use the multi-modal attention network to complete the interaction between the text features and the picture features, so as to obtain the joint text-visual feature representation at the word level

[0029]

[0030]

[0031]

[0032] Among them, Q is the word representation T of the currently calculated node i, K and V are the picture representations output by ResNet, and MH(·) is the multi-head attention;

[0033]

[0034]

[0035] Among them, Q is the word representation T of the currently calculated node i, M image is the image representation corresponding to the text, LN(·) is layer normalization, and FFN(·) is a feed-forward neural network;

[0036] Step B22: With the help of the edge relation adjacency matrix of the graph attention network in Step B1 and the joint text-visual feature representation at the word granularity extracted in Step B21 initially construct the graph attention network of the first layer, and output the word granularity text-visual fusion representation

[0037]

[0038]

[0039]

[0040] Among them, is the edge relation adjacency matrix at the word granularity in Step B1 the adjacent nodes in, is the final output of the first layer of the graph attention network FC is a fully connected layer, is the result after the word node passes through the masked self-attention mechanism, || represents vector concatenation, Z is the number of attention heads, and б is the activation function; is the trainable parameter of the l-th layer of the Z-th attention head, and f(·) is a function to measure the correlation between two words; by stacking multiple GAT layers, the previous layer is used as the input of the next layer, and image information is fused between layers;

[0041] Step B23: Group the phrases divided in Step B1 and the word granularity text-visual fusion representation output in Step B22 perform average pooling, and the result is used as the text input for phrase-level multimodal interaction, and use the multimodal attention mechanism to complete the interaction between phrase-level text features and image features

[0042]

[0043]

[0044] Among them, Q is the word granularity text-visual fusion representation output in Step B22 According to the phrase grouping divided in Step B1 the result of performing average pooling, M imageIt is the picture representation corresponding to the text, LN(·) is layer normalization, and FFN(·) is a feed-forward neural network;

[0045] Step B24: Use the interaction between the phrase-level text features and the picture features output in Step B23 and the word-grained text-visual fusion representation in Step B22 to perform fusion. Based on the edge relation adjacency matrix of the graph attention network obtained in Step B1 to generate the phrase-grained text-visual fusion representation as the graph attention output

[0046] Step B25: Repeat Step B24 until reaching the sentence-level text division level;

[0047] Step B26: Average pool the top-level phrase-grained text-visual fusion representation obtained in Step B25 as the text features at the sentence-grained level, and input them into the multi-modal attention mechanism module to complete the interaction between the sentence-level text features and the picture features and obtain their fusion representation

[0048]

[0049]

[0050] where Q is the phrase-grained text-visual fusion representation output in Step B22 Perform average pooling according to the sentence grouping in Step B1 to obtain the result M image It is the picture representation corresponding to the text, LN(·) is layer normalization, and FFN(·) is a feed-forward neural network;

[0051] Step B27: Average pool the sentence-level text-picture feature fusion representation obtained in Step B26 and the output word vectors at each position of the graph attention network to obtain the multi-grained text-visual fusion feature representation

[0052] Furthermore, Step B3 specifically includes the following steps:

[0053] Step B31: Use the multi-grained text-visual fusion feature representation obtained in Step B2 to calculate the probabilities p start and p end respectively for the start and end positions of the aspect words at each position:

[0054]

[0055]

[0056] Among them, w start , b start , w end , b end are all trainable parameters;

[0057] Step B32: Select the most likely starting and ending positions of the aspect word according to the probability (index start , index end ):

[0058]

[0059]

[0060] Among them, l represents the largest word subscript in the sentence.

[0061] Furthermore, in the said step B4, the calculation formula of the aspect representation a i is as follows:

[0062]

[0063] Among them, l represents the largest word subscript in the sentence.

[0064] Furthermore, the said step B5 specifically includes the following steps:

[0065] Step B51: Let the graph be the aspect graph G a =(V, E), where V is the set of graph nodes and E is the set of edge relations in the graph; use the word-level adjacency matrix obtained in step B1

[0066]

[0067] Step B52: Take the edges directly connected to the aspect word in as the edge relations of the aspect graph G a , and construct the adjacency matrix of the aspect graph G a based on this

[0068]

[0069] Among them, represents the value of the adjacency matrix a of the aspect graph G at the position (i, j).

[0070] Furthermore, the said step B6 specifically includes the following steps:

[0071] Step B61: According to the image feature representation obtained in Step A, the ANP parser parses out the TOP-K most relevant (noun N, adjective adj a ) pairs;

[0072]

[0073] where Anps is the set of the TOP-K most relevant (noun N, adjective adj a ) pairs;

[0074] Step B62: Based on the aspect word representation a i and the directly connected word representation w a in the aspect graph G i obtained in Step B4, calculate the cosine similarity between the nouns of the ANP pairs or

[0075]

[0076]

[0077] where, represents the cosine similarity between the aspect word a i and the noun in the ANP pair, represents the cosine similarity between the directly connected word w i and the noun in the ANP pair; i

[0078] Step B63: Based on the cosine similarity obtained in Step B62, use the Top-K algorithm to screen out the K most relevant (noun N, adjective adj a ) pairs for this word:

[0079]

[0080] where score i represents the cosine similarity calculated in Step B61; is the set of nouns of the most relevant ANP pairs;

[0081] Step B64: Obtain the corresponding adjective adj a in the (noun N, adjective adj a ) pairs screened out in Step B63:

[0082]

[0083] where, is the most relevant adjective corresponding in the ANP pair;​

[0084] Step B65: Establish an edge relationship between the nodes selected in Step B64 as the nodes in the aspect graph G a and their corresponding words, and enhance the adjacency matrix in Step B4

[0085]

[0086] where i and j respectively represent the abscissa and ordinate in the adjacency matrix; wi represents the word with subscript i, and wj represents the word with subscript j.

[0087] Furthermore, Step B7 specifically includes the following steps:

[0088] Step B71: Use a graph attention network to obtain the fused feature representation senti a of the aspect graph G a ;

[0089]

[0090] where senti a l represents the feature representation of the l-th layer. When l = 1, senti a l-1 is the text representation of the multi-granularity image information obtained in Step B2 w l and b l are the weight and bias values of the l-th layer, ReLu is the activation function, is the adjacency matrix of the aspect graph G a obtained in Step B6;

[0091] Step B72: Establish a Mask matrix and apply it to the fused feature representation senti a of the aspect graph G a , and retain the sentiment representation aspect w of the aspect word;

[0092]

[0093] Based on the obtained sentiment representation aspect w of the aspect word, calculate the sentiment polarity of the aspect word.

[0094] The present invention also provides a multi-modal sentiment analysis system that fuses multi-granularity visual and text features, including a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the above method steps can be implemented.

[0095] Compared with the prior art, the present invention has the following beneficial effects: First, the present invention addresses the noise problem caused by irrelevant or redundant information in different modalities. Instead of directly adding the information of the entire image, it uses a multi-granularity joint visual-text feature representation to reduce the noise problem caused by redundant information. By leveraging the natural phrase grouping in the constituency tree, multiple words form a higher semantic representation, making phrases more perceptible than individual words (such as adjectives). Using a multi-layer GAT, an adjacency matrix is constructed from the constituency tree and the dependency tree to remove cross-clause dependencies, fuse text information, and during the fusion process at different layers, different granularity visual feature representations are added. Initially, the fused word-level joint visual feature representation is added, the phrase-level joint visual feature representation is added in the middle layer, and the sentence-level joint visual feature representation is added at the top layer. Different levels of interaction between text and image are explored. Additionally, the present invention uses ANP to adopt a strategy and method of enhancing multi-modal information for aspect words. By using the sentiment word corresponding to the ANP most relevant to the aspect word as a bridge, the multi-modal sentiment information corresponding to the aspect word is introduced to help the model improve the accuracy of sentiment polarity prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] Figure 1 is a flowchart of the method implementation of an embodiment of the present invention;

[0097] Figure 2 is a model architecture diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0098] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0099] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0100] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0101] As Figure 1-2 shown, this embodiment provides a multi-modal sentiment analysis method that fuses multi-granularity visual and text features, including the following steps:

[0102] Step A: Initialize the text representation of graph nodes using a pre-trained model, and preliminarily extract the image feature representation using ResNet.

[0103] Step B: Determine the edge relation adjacency matrix in the graph attention network according to the syntactic dependency relationship and the constituent tree structure; use the multi-modal attention mechanism to obtain the joint text-visual feature representations at the word level, phrase level, and sentence level respectively; fuse the joint text-visual feature representations at the word level, phrase level, and sentence level with the help of a multi-layer graph attention network to finally obtain the multi-granularity text-visual fusion feature representation; predict the aspect word position based on the text-visual joint feature representation output by the multi-layer graph attention network and form the aspect word representation; parse the most relevant ANP pairs using the ANP parser according to the aspect word position, and construct the aspect graph based on the adjacency relationship of the aspect word in the syntactic dependency relationship.

[0104] Step C: Predict the sentiment polarity corresponding to the aspect word based on the output of the aspect graph as the aspect word sentiment representation.

[0105] In this embodiment, step B specifically includes the following steps:

[0106] Step B1: Perform syntactic dependency parsing on the dataset text to obtain the syntactic dependency relationship A between different nodes; construct the constituent tree structure A based on the text with the help of constituent tree parsing, obtain the text partitioning of different clauses, specifically manifested as word-level, intermediate-layer phrase-level, and sentence-level texts 1 ; fuse the syntactic dependency relationship A 2 and the constituent tree structure A to obtain the edge relation adjacency matrix of different layers of the graph attention network 1 and the constituent tree structure A 2 to obtain the edge relation adjacency matrix of different layers of the graph attention network

[0107] In this embodiment, step B1 specifically includes the following steps:

[0108] Step B11: Perform syntactic dependency parsing on the dataset text to obtain the syntactic dependency relationship A between different nodes; construct the constituent tree structure A based on the text with the help of constituent tree parsing, obtain the text partitioning of different clauses, specifically manifested as word-level, intermediate-layer phrase-level, and sentence-level 1 ; construct the constituent tree structure A based on the text with the help of constituent tree parsing 2 to obtain the text partitioning of different clauses, specifically manifested as word-level, intermediate-layer phrase-level, and sentence-level

[0109]

[0110]

[0111] where is the value at position (i, j) in the syntactic dependency relationship A 1 in the syntactic dependency relationship A is the value at position (i, j) in the constituent tree structure A 2 ; Dep.Tree represents the syntactic dependency relationship; Con.Tree represents the constituent tree structure.

[0112] Step B12: Taking the constituent tree clause partitioning strategy obtained in Step B11 as the standard, eliminate the cross-clause dependency relationships in the syntactic dependency relationships to achieve dependency relationship noise reduction.

[0113] Step B13: Combine the noise-reduced syntactic dependency relationship A 1 obtained in Step B12 and the adjacency relationship of each layer of the constituent tree 2 as the edge relationship adjacency matrix of the graph attention network

[0114]

[0115] wherein, represents the value at position (i, j) in the edge relationship adjacency matrix of the graph attention network.

[0116] Step B2: Using the edge relationship adjacency matrix of the multi-layer graph attention network obtained in Step B1, obtain text representations at different granularities Input into the multi-modal attention network layer by layer to obtain the joint text-visual feature representation specific to that layer Complete the interaction between the multi-granularity text features and the picture features; finally output the multi-granularity text-visual fusion feature representation

[0117] In this embodiment, the specific steps of Step B2 include the following steps:

[0118] Step B21: Use the word representation T obtained by the pre-trained model to interact with the picture, and use the multi-modal attention network to complete the interaction between the text features and the picture features, so as to obtain the joint text-visual feature representation at the word granularity

[0119]

[0120]

[0121]

[0122] wherein, Q is the word representation T of the currently calculated node i, K and V are the picture representations output by ResNet, MH(·) is the multi-head attention, and d k = 768.

[0123]

[0124]

[0125] Among them, Q is the word representation T of the currently calculated node i, and M image is the image representation corresponding to the text, LN(·) is layer normalization, and FFN(·) is a feed-forward neural network.

[0126] Step B22: With the help of the edge relation adjacency matrix of the graph attention network in Step B1 and the joint text-visual feature representation at the word granularity extracted in Step B21 initially construct the graph attention network of the first layer, and output the word granularity text-visual fusion representation

[0127]

[0128]

[0129]

[0130] Among them, is the edge relation adjacency matrix at the word granularity in Step B1 is the adjacent node in is the final output of the first layer of the graph attention network FC is a fully connected layer, is the result after the word node passes through the masked self-attention mechanism, || represents vector concatenation, Z is the number of attention heads, and б is the activation function; is the trainable parameter of the l-th layer of the Z-th attention head, and f(·) is a function to measure the correlation between two words; by stacking multiple GAT layers, the previous layer is used as the input of the next layer, and picture information is fused between layers.

[0131] Step B23: Group the phrases divided in Step B1 and the word granularity text-visual fusion representation output in Step B22 perform average pooling, and the result is used as the text input for phrase-level multimodal interaction. Use the multimodal attention mechanism to complete the interaction between phrase-level text features and picture features

[0132]

[0133]

[0134] Among them, Q is the word granularity text-visual fusion representation output in Step B22 according to the phrase grouping divided in Step B1 the result of performing average pooling, M imageIt is the picture representation corresponding to the text, LN(·) is layer normalization, and FFN(·) is the feed-forward neural network.

[0135] Step B24: Use the interaction between the phrase-level text features and the picture features output in Step B23 and the word-grained text-visual fusion representation in Step B22 to perform fusion. Based on the edge relation adjacency matrix of the graph attention network obtained in Step B1 to generate the phrase-grained text-visual fusion representation through graph attention output

[0136] Step B25: Repeat Step B24 until reaching the sentence-level text division level.

[0137] Step B26: Average pool the top-level phrase-grained text-visual fusion representation obtained in Step B25 as the text features at the sentence-grained level, and input them into the multi-modal attention mechanism module to complete the interaction between the sentence-level text features and the picture features and obtain their fusion representation

[0138]

[0139]

[0140] where Q is the phrase-grained text-visual fusion representation output in Step B22 Perform average pooling according to the sentence grouping in Step B1 to obtain the result M image It is the picture representation corresponding to the text, LN(·) is layer normalization, and FFN(·) is the feed-forward neural network.

[0141] Step B27: Average pool the sentence-level text-picture feature fusion representation obtained in Step B26 and the output word vectors at each position of the graph attention network to obtain the multi-grained text-visual fusion feature representation

[0142] Step B3: Use the multi-grained text-visual fusion feature representation obtained in Step B2 to calculate the probabilities p start and p end of the start and end positions of the aspect words at each position respectively, and select the most likely start and end subscript positions (index start , index end ) according to the probabilities.

[0143] In this embodiment, Step B3 specifically includes the following steps:

[0144] Step B31: Use the multi-granularity text-visual fusion feature representation obtained in Step B2 Calculate the probabilities p start and p end respectively for the start and end positions of the aspect word at each position:

[0145]

[0146]

[0147] where w start , b start , w end , and b end are all trainable parameters;

[0148] Step B32: Select the most likely start and end subscript positions (index start , index end ) of the aspect word according to the probabilities:

[0149]

[0150]

[0151] where l represents the largest word subscript in the sentence.

[0152] Step B4: Obtain the final aspect word representation a start according to the aspect word positions (index end ) marked in Step B3 and the multi-granularity text-visual fusion feature representation obtained in Step B2 i .

[0153] where the calculation formula for the aspect representation a i is as follows:

[0154]

[0155]

[0155] where l represents the largest word subscript in the sentence.

[0156] Step B5: Use the word-level adjacency matrix obtained in Step B1 Take the directly connected edges with the aspect word positions (index start , index end ) obtained in Step B3 as the edge relationships of the aspect graph G a , and construct the adjacency matrix of the aspect graph G a based on this

[0157] In this embodiment, step B5 specifically includes the following steps:

[0158] Step B51: Set the graph as the aspect graph G a =(V, E), where V is the set of graph nodes and E is the set of edge relationships in the graph; use the word-level adjacency matrix obtained in step B1

[0159]

[0160] Step B52: Take the edges directly connected to the aspect words as the edge relationships of the aspect graph G a and construct the adjacency matrix of the aspect graph G a based on this

[0161]

[0162] where represents the value of the adjacency matrix of the aspect graph G a at position (i, j).

[0163] Step B6: According to the image feature representation obtained in step A, the ANP parser parses out the top-K most relevant (noun N, adjective adj a ) pairs. According to the aspect word representation a i obtained in step B4 and its directly connected word representations, calculate the cosine similarity with the nouns in the ANP respectively; and select the K most relevant (noun N, adjective adj a ) pairs as the supplementary nodes of the aspect graph G a to enhance the adjacency matrix in step B4

[0164] In this embodiment, step B6 specifically includes the following steps:

[0165] Step B61: According to the image feature representation obtained in step A, the ANP parser parses out the TOP-K most relevant (noun N, adjective adj a ) pairs;

[0166]

[0167] where Anps is the set of the TOP-K most relevant (noun N, adjective adj a ) pairs;

[0168] Step B62: According to the aspect word representation a i obtained in step B4 and the aspect graph G a ​The directly connected words in it represent w i , calculate the cosine similarity between the nouns of the ANP pair respectively or

[0169]

[0170]

[0171] Among them, represents the aspect word a i and the cosine similarity between the noun in the ANP pair, represents the word w i directly connected to the aspect word a i and the cosine similarity between the noun in the ANP pair;

[0172] Step B63: Based on the cosine similarity obtained in Step B62, use the Top-K algorithm to screen out the K (noun N, adjective adj a ) pairs most relevant to this word:

[0173]

[0174] Among them, score i represents the cosine similarity calculated in Step B61; is the set of nouns of the most relevant ANP pairs;

[0175] Step B64: Obtain the corresponding adjective adj a in the (noun N, adjective adj a ) pairs screened out in Step B63:

[0176]

[0177] Among them, is the most relevant adjective corresponding in the ANP pair;

[0178] Step B65: Take the screened out in Step B64 as the nodes in the aspect graph G a , and establish an edge relationship between it and its corresponding word to enhance the adjacency matrix in Step B4

[0179]

[0180] Among them, i and j represent the abscissa and ordinate in the adjacency matrix respectively; wi represents the word with subscript i, and wj represents the word with subscript j.

[0181] Step B7: Use the graph attention network to obtain the aspect graph G a 's fused feature representation senti a ; Establish a Mask matrix and apply the fused feature representation senti a of the aspect graph G a to retain the sentiment representation aspect w of the aspect term; Based on the obtained sentiment representation aspect w of the aspect term, calculate the sentiment polarity of this aspect term.

[0182] In this embodiment, the specific steps of step B7 include the following steps:

[0183] Step B71: Use the graph attention network to obtain the fused feature representation senti a of the aspect graph G a ;

[0184]

[0185] Among them, senti a l represents the feature representation of the l-th layer. When l = 1, senti a l-1 is the text representation of the multi-granularity image information obtained in step B2 w l and b l are the weight and bias values of the l-th layer, ReLu is the activation function, is the adjacency matrix of the aspect graph G a obtained in step B6;

[0186] Step B72: Establish a Mask matrix and apply it to the fused feature representation senti a of the aspect graph G a to retain the sentiment representation aspect w ;

[0187]

[0188] Based on the obtained sentiment representation aspect w of the aspect term, calculate the sentiment polarity of this aspect term.

[0189] This embodiment also provides a multi-modal sentiment analysis system that fuses multi-granularity visual and text features, which is characterized in that it includes a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the above method steps can be implemented.

[0190] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0191] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0192] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0194] As mentioned above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A multimodal sentiment analysis method that integrates multi-granularity visual and text features, It is characterized in that The following steps are involved: Step A: Use the pre-trained model to initialize the graph node text representation, and use ResNet to preliminarily extract the image feature representation; Step B: Determine the edge relationship adjacency matrix in the graph attention network based on the syntactic dependency and component tree structure; Use multimodal attention mechanism to obtain joint text visual feature representation at word level, phrase level, and sentence level respectively; The multi-layer graph attention network is used to fuse the joint text-visual feature representations at the word level, phrase level, and sentence level, and finally a multi-granular text-visual fusion feature representation is obtained; the position of aspect words is predicted and the aspect word representation is formed according to the text-visual joint feature representation output by the multi-layer graph attention network; the most relevant ANP pairs are parsed according to the position of aspect words using the ANP parser, and an aspect graph is constructed according to the adjacency relationship of aspect words in syntactic dependencies; Step C: Based on the output of the aspect graph as the sentiment representation of the aspect word, predict the sentiment polarity corresponding to the aspect word; The step B specifically comprises the following steps: Step B1: Perform syntactic dependency parsing on the dataset text to obtain the syntactic dependency relationship A between different nodes 1 ; construct a text-based constituency tree structure A by means of constituency tree parsing 2 , obtain the text partitioning of different clauses, specifically manifested in word-level, intermediate-level phrase-level, and sentence-level texts Use the syntactic dependency relationship A 1 and the constituency tree structure A 2 to fuse and obtain the edge relationship adjacency matrix of different layers of the graph attention network Step B2: Obtain text representations at different granularities by using the edge relation adjacency matrix of the multi-layer graph attention network obtained in Step B1 Input into the multi-modal attention network layer by layer to obtain the joint text-visual feature representation specific to that layer Complete the interaction between text features and image features at multiple granularities; finally output the multi-granularity text-visual fusion feature representation Step B3: Use the multi-granularity text-visual fusion feature representation obtained in Step B2 Calculate the probabilities p start , p end of the start and end positions of the aspect word at each position respectively, and select the start and end positions (index start , index end ) of the aspect word with the maximum probability according to the probabilities; Step B4: According to the aspect word positions (index start , index end ) marked in Step B3 and the multi-granularity text-visual fusion feature representation obtained in Step B2 obtain the final aspect word representation a i ; Step B5: Use the word-level adjacency matrix obtained in Step B1 With the aspect word positions (index start , index end ) obtained in Step B3 as the center, take the directly connected edges as the edge relationships of the aspect graph G a Based on this, construct the adjacency matrix of the aspect graph G a ​ Step B6: According to the image feature representation obtained in Step A, the ANP parser parses out the top-K (noun N, adjective adj a ) pairs that are most relevant to the image. Based on the aspect word representation a i obtained in Step B4 and its directly connected word representations, calculate the cosine similarity with the nouns in ANP respectively; and select the K (noun N, adjective adj a ) pairs with the highest cosine similarity as the supplementary nodes of the aspect graph G a to enhance the adjacency matrix in Step B4 Step B7: Use the graph attention network to obtain the aspect graph G a 's fused feature representation senti a ; Establish a Mask matrix and apply the fused feature representation senti a of the aspect graph G a to retain the sentiment representation aspect w of the aspect term; Based on the obtained sentiment representation aspect w of the aspect term, calculate the sentiment polarity of this aspect term; The step B1 specifically comprises the following steps: Step B11: Perform syntactic dependency parsing on the dataset text to obtain the syntactic dependency relationship A between different nodes 1 ; Construct a text-based constituent tree structure A by means of constituent tree parsing 2 , and obtain the text partitioning of different clauses, specifically manifested at the word level, intermediate phrase level, and sentence level Among them, is the numerical value at position (i, j) in syntactic dependency relationship A 1 is the numerical value at position (i, j) in constituent tree structure A Dep.Tree represents the syntactic dependency relationship; Con.Tree represents the constituent tree structure; 2 is the numerical value at position (i, j) in the constituent tree structure A; Dep.Tree represents the syntactic dependency relationship; Con.Tree represents the constituent tree structure; Step B12: Using the component tree clause partitioning strategy obtained in step B11 as a standard, remove cross-clause dependencies in the syntactic dependency relationship to achieve dependency noise reduction; Step B13: Take the noise-reduced syntactic dependency relationship A obtained in step B12 1 and the adjacency relationship A of each layer of the constituent tree 2 and fuse them as the edge relationship adjacency matrix of the graph attention network Among them, represents the value at position (i, j) in the edge relation adjacency matrix of the graph attention network; The step B2 specifically comprises the following steps: Step B21: Interact the word representation T obtained using the pre-trained model with the picture, and use the multi-modal attention network to complete the interaction between the text features and the picture features, so as to obtain the joint text-visual feature representation at the word level Where Q is the word representation of the currently calculated node i, K and V are the image representations output by ResNet, and MH(·) is the multi-head attention; Among them, Q is the word representation T of the currently calculated node i, M image is the image representation corresponding to the text, LN(·) is layer normalization, and FFN(·) is a feed-forward neural network; Step B22: With the help of the edge relation adjacency matrix of the graph attention network in Step B1 and the joint text visual feature representation at the word granularity extracted in Step B21 initially construct the graph attention network of the first layer and output the word granularity text-visual fusion representation Among them, is the edge relation adjacency matrix at the word granularity in step B1 The adjacent nodes in is the final output of the first layer of the graph attention network, FC is the fully connected layer, is the result after the word node passes through the masked self-attention mechanism, || represents vector concatenation, Z is the number of attention heads, and б is the activation function; is the trainable parameter of the l-th layer of the Z-th attention head, f(·) is a function that measures the correlation between two words; by stacking multiple GAT layers, the previous layer is used as the input of the next layer, and picture information is fused between layers; Step B23: Group the phrases divided in Step B1 and the word-grained text-visual fusion representation output in Step B22 Perform average pooling, and use the result as the text input for phrase-level multimodal interaction. Use the multimodal attention mechanism to complete the interaction between phrase-level text features and image features Among them, Q is the word-grained text-visual fusion representation output in step B22 According to the phrase grouping divided in step B1 The result of average pooling, M image is the image representation corresponding to the text, LN(·) is layer normalization, and FFN(·) is a feed-forward neural network; Step B24: Interaction between the phrase-level text features and image features output in Step B23 and the word-grained text-visual fusion representation in Step B22 are fused, based on the edge relation adjacency matrix of the graph attention network obtained in Step B1 to generate a phrase-grained text-visual fusion representation as the graph attention output Step B25: Repeat step B24 until reaching the sentence-level text segmentation level; Step B26: Use the top-phrase granularity text-visual fusion representation obtained in Step B25 Perform average pooling to obtain text features at the sentence granularity level, and input them into the multi-modal attention mechanism module to complete the interaction between sentence-level text features and image features, and obtain their fusion representation Among them, Q is the phrase-level text-visual fusion representation output in step B22 Group sentences according to step B1 The result of average pooling, M image is the image representation corresponding to the text, LN(·) is layer normalization, and FFN(·) is a feed-forward neural network; Step B27: Obtain the sentence-level text-image feature fusion representation obtained in step B26 And perform average pooling on the output word vectors at each position of the graph attention network to obtain a multi-granularity text-visual fusion feature representation The step B3 specifically comprises the following steps: Step B31: Use the multi-granularity text-visual fusion feature representation obtained in Step B2 Calculate the probabilities p start , p end : Among them, w start , b start , w end , b end are all trainable parameters; Step B32: Select the start and end positions (index start , index end ) of the aspect word with the highest probability according to the probability start ,index end ) Among them, l represents the largest word subscript in the sentence; In step B4, the aspect represents a i The calculation formula is as follows: Among them, l represents the largest word subscript in the sentence; The step B5 specifically comprises the following steps: Step B51: Set the graph as the aspect graph G a =(V, E), where V is the set of graph nodes and E is the set of edge relationships in the graph; use the word-level adjacency matrix obtained in Step B1 Step B52: Take The edges directly connected to the neutral aspect words as the aspect graph G a Edge relationship of, and construct the aspect graph G based on this a Adjacency matrix of Among them, represents the adjacency matrix of the aspect graph G a at the position (i, j); the value at position (i, j); The step B6 specifically comprises the following steps: Step B61: Based on the image feature representation obtained in Step A, the ANP parser parses out the TOP-K most relevant (noun N, adjective adj a ) pairs; Among them, Anps is the set of the TOP-K most relevant (noun N, adjective adj a ) pairs; Step B62: Obtain the aspect word representation a according to the aspect obtained in step B4 i and the aspect graph G a The directly connected words in it are represented as w i , and calculate the cosine similarity between the nouns of the ANP pair respectively Or Among them, represents the cosine similarity between aspect word a i and the nouns in the ANP pair, represents the cosine similarity between the word w i directly connected to aspect word a i and the nouns in the ANP pair; Step B63: Based on the cosine similarity obtained in Step B62, use the Top-K algorithm to screen out the K (noun N, adjective adj a ) pairs that are most relevant to this word: Among them, score i represents the cosine similarity calculated in step B61; is the noun set of the most relevant ANP pairs; Step B64: Obtain the corresponding adjective adj for the pair (noun N, adjective adj) selected in step B63 a ) a : Among them, is the most relevant adjective corresponding to the ANP pair; Step B65: For the ones selected in step B64 serve as nodes in aspect graph G a and establish edge relationships between them and their corresponding words, enhancing the adjacency matrix in step B4 Where i and j represent the horizontal and vertical coordinates in the adjacency matrix respectively; wi represents the word with subscript i, and wj represents the word with subscript j; The step B7 specifically includes the following steps: Step B71: Use the graph attention network to obtain the aspect graph G a 's fused feature representation senti a ; Among them, senti a l represents the feature representation of the l-th layer. When l = 1, senti a l-1 is the text representation of the multi-granularity image information obtained in step B2 w l and b l are the weight and bias values of the l-th layer, and ReLu is the activation function is the adjacency matrix of the aspect graph G a obtained in step B6; Step B72: Establish a Mask matrix and apply it to the aspect graph G a of the fused feature representation senti a , and retain the sentiment representation aspect of the aspect term w ; Sentiment representation of the aspect word based on the obtained aspect words w , calculate the sentiment polarity of the aspect word.

2. A multimodal sentiment analysis system that integrates multi-granularity visual and textual features, It is characterized in that The method comprises a memory, a processor and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method steps as claimed in claim 1 can be implemented.