An Aspect-Level Sentiment Analysis Method Based on Dual Graph Convolutional Network

By using a dual-graph convolutional network in aspect-level sentiment analysis, combining GloVe word vectors and bidirectional long and short-term memory networks, syntactic and semantic features are obtained, the shortcomings of context and aspect interaction processing are solved, and the accuracy of sentiment analysis is improved.

CN115759116BActive Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211445352.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2025-05-27
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

Existing aspect-level sentiment analysis techniques are insufficient in dealing with the interaction between context and aspects. A single attention mechanism cannot effectively capture syntactic dependencies, which may lead to incorrect associations.

Method used

Using a method based on a double-graph convolutional network, the context and aspects are transformed into word vector embedding representations through the GloVe word vector model, feature extraction is performed by combining bidirectional long and short-term memory networks, and syntactic graph convolutional networks and semantic graph convolutional networks are used to obtain syntactic features and semantic features, and finally the features predicting emotional polarity are generated through interaction matrix and average pooling operations.

Benefits of technology

The accuracy of aspect-level sentiment analysis is improved, and by combining syntactic and semantic information, the ability to capture the mutual influence between context and aspects is enhanced, reducing the possibility of wrong associations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115759116B_ABST
    Figure CN115759116B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of sentiment analysis, and specifically relates to an aspect-level sentiment analysis method based on a dual graph convolutional network, including: obtaining the context and aspects of the text to be analyzed, converting them into word vector embedding representations as inputs; using a bidirectional long short-term memory network to extract features from the word vector embedding representations of the context and aspects to obtain the hidden state representations of the context and aspects; after obtaining the hidden state of the context, performing position encoding on it and then using it as the initial input of the syntactic graph convolutional network to obtain syntactic features; at the same time, combining guiding vectors to encode the hidden states of the context and aspects to obtain an interaction matrix; using it as the initial input of the semantic graph convolutional network to obtain semantic features; connecting the syntactic features and semantic features as the final features to predict the sentiment polarity. The aspect-level sentiment analysis method provided by the present invention can improve the accuracy of aspect-level sentiment analysis from both semantic and syntactic aspects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of sentiment analysis, and in particular relates to an aspect-level sentiment analysis method based on a dual-graph convolutional network. Background Art

[0002] With the rapid development of information technology, the Internet has become an inseparable part of people's daily life. Obviously, posting comments on the Internet has become an important way for people to express their opinions and convey their experiences. More and more people are willing to express their attitudes and emotions on the Internet, and the number of online comments has also begun to explode. Internet comment texts have gradually become an important source of reference information for people to find decision-making, but how to accurately and quickly extract valuable information from massive information has become a difficult problem that people need to solve urgently. Text sentiment analysis, also known as opinion mining, is a study on the calculation of people's opinions, comments and emotions expressed on entities (including products, services, organizations, etc.). Text sentiment analysis can be divided into three different levels according to the granularity of analysis: paragraph-level sentiment analysis, sentence-level sentiment analysis and aspect-level sentiment analysis. In the early stage, paragraph-level and sentence-level sentiment analysis tasks were the focus of research. However, overall sentiment analysis of text will obscure its details, and overall sentiment cannot reflect people's fine-grained emotional expression of opinion targets. Users also pay more attention to fine-grained information when browsing comments, such as price, quality, size, taste, etc. Therefore, aspect-level sentiment analysis of comments can help users make better decisions. For businesses or other organizations, there is no need to collect public opinions on certain aspects through time-consuming questionnaires, because such information is already very abundant. In summary, in order to conduct a more comprehensive sentiment analysis, the system needs to determine the emotional information expressed by the comment text on each aspect, which is the aspect-level sentiment analysis technology.

[0003] In the early work, traditional methods mainly focused on artificial feature engineering. However, a lot of work is usually required for part-of-speech polarity tagging and rule formulation, and manual feature screening of data is also required, so the labor cost is extremely high. In addition, since the polarity and rules of part-of-speech may be different in different fields, the model generalization ability is poor. In recent years, deep learning has become a powerful technology and has produced state-of-the-art results in many application fields. This method shows strong feature extraction and text representation capabilities, so it has good scalability. In the task of aspect-level sentiment analysis, deep learning methods have gradually become a research hotspot, and researchers have proposed various models based on deep learning to improve task performance. Deep learning models can be divided into methods based on convolutional neural networks (CNN), methods based on recurrent neural networks (RNN), methods based on memory neural networks (Memory Networks) and other methods. The above methods can only improve the accuracy of aspect-level sentiment analysis from a semantic perspective, but the accuracy of aspect-level sentiment analysis from a syntactic perspective has not been solved. By combining the sentence dependency tree to establish a graph convolutional network (GCN), syntactic information and word dependencies can be used to improve the efficiency of aspect-level sentiment analysis.

[0004] The current aspect-level sentiment analysis technology still has the following problems:

[0005] (1) There is a lack of interaction between context and aspects. Some studies have recognized the importance of aspects in sentiment classification and have accurately modeled their context by generating specific aspect representations. However, they have ignored the mutual influence between the two. In a sentence, there may be more than one aspect word and context word, and each word has a different degree of influence on the final classification. Only by coordinating the context and aspects can the performance of semantic analysis be truly improved.

[0006] (2) A single attention mechanism (semantically) is not enough. Although many attention-based models have improved the effect of aspect-level sentiment analysis to a certain extent, attention-based models are not sufficient to capture the syntactic dependencies between aspects and contexts. Current attention mechanisms may cause a given aspect to mistakenly use syntactically irrelevant context words as descriptive information.

[0007] (3) It may lead to incorrect associations. Building a graph convolutional network on the syntactic parse tree, although it does contain useful syntactic information, it may still incorrectly associate irrelevant words to the target aspect through the iterations of graph convolution propagation. Summary of the invention

[0008] In order to solve the above technical problems, the present invention proposes an aspect-level sentiment analysis method based on a dual-graph convolutional network, comprising:

[0009] S1. Obtain the context and aspects of the text to be analyzed, and convert the context and aspects of the text to be analyzed into word vector embedding representations through the GloVe word vector model;

[0010] S2, using a bidirectional long short-term memory network to extract features from the context and aspect word vector embedding representations to obtain the hidden state representations of the context and aspect;

[0011] S3, position encoding the hidden state representation of the context, and inputting the encoded hidden state representation of the context into the syntactic graph convolutional network to obtain syntactic features;

[0012] S4, averaging the hidden states of the context and the aspect to obtain a guide vector, using the guide vector to encode the hidden state representation of the context and the aspect, obtaining an interaction matrix according to the encoded hidden state representation of the context and the aspect, and inputting the interaction matrix into a semantic graph convolutional network to obtain semantic features;

[0013] S5. Through the average pooling operation, the hidden state representation of non-aspect words in the syntactic features output by the syntactic graph convolutional network and the semantic features output by the semantic graph convolutional network are masked to obtain the final syntactic features and semantic features. The final syntactic features and semantic features are connected as the final features, and the sentiment polarity is predicted based on the final features.

[0014] Beneficial effects of the present invention:

[0015] The present invention connects syntactic features and semantic features as the final features to predict sentiment polarity, thereby improving the accuracy of aspect-level sentiment analysis from both semantic and syntactic aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flow chart of the present invention;

[0017] Figure 2 is the syntactic dependency graph of the present invention;

[0018] Figure 3 It is a schematic diagram of the adjacency matrix of the present invention;

[0019] Figure 4 It is a single-layer graph convolutional network graph. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] A dual-graph convolutional network-based aspect-level sentiment analysis method, such as Figure 1 As shown, including:

[0022] S1. Obtain the context and aspects of the text to be analyzed, and convert the context and aspects of the text to be analyzed into word vector embedding representations through the GloVe word vector model;

[0023] S2, using a bidirectional long short-term memory network to extract features from the context and aspect word vector embedding representations to obtain the hidden state representations of the context and aspect;

[0024] S3, position encoding is performed on the hidden state representation of the context, and the encoded hidden state representation of the context is input into the syntactic graph convolutional network to obtain syntactic features;

[0025] S4, averaging the hidden states of the context and the aspect to obtain a guide vector, using the guide vector to encode the hidden state representations of the context and the aspect respectively, to obtain a context representation integrated with the aspect and an aspect representation integrated with the context, obtaining an interaction matrix based on the context representation integrated with the aspect and the aspect representation integrated with the context, and inputting the interaction matrix into a semantic graph convolutional network to obtain a semantic feature;

[0026] S5. Shield the hidden state representation of non-aspect words in the syntactic features output by the syntactic graph convolutional network and the semantic features output by the semantic graph convolutional network, obtain the final syntactic features and semantic features through average pooling operation, connect the final syntactic features and semantic features as the final features, and predict the sentiment polarity based on the final features.

[0027] The obtained text to be analyzed is preprocessed: (1) the text is segmented; (2) noise characters are removed, numbers, letters, hyphens, and punctuation marks are retained, and other characters are regarded as noise characters; (3) all letters are converted to lowercase.

[0028] The word set of the text to be analyzed is represented as For a given text, it can be regarded as a sequence of words in the Ws set, expressed as Then the length of the text is n, and For a given aspect subsequence of text Wc, it is represented as Where m≤n. Convert the input text into word vector embedding, here we use GloVe embedding to get the embedding matrix E∈R |V|×de , where |V| is the size of the vocabulary and de is the embedding dimension of the word vector.

[0029] Positional encoding of the hidden state representation of the context, including:

[0030]

[0031] Among them, p i represents the result of position encoding of the hidden state representation of the context; F() represents the function of assigning position weights to enhance the importance of context words close to aspect words, in order to reduce the noise and bias naturally generated during dependency parsing q i represents the position weight of the i-th token, τ represents the position of the aspect word, i represents the i-th word in the sentence, m represents the length of the aspect word, n represents the length of the context, Represents the hidden state of the context.

[0032] Use the spaCy syntactic analyzer to build a dependency tree for each sentence, and then Figure 2 The dependency relationship shown in Figure 2 obtains the adjacency matrix A of each sentence; Figure 3 As shown in the figure, if two words have a dependency relationship in syntactic analysis, an edge is established between the two words. If there is an edge between the two words, the weight of the edge is 1, otherwise it is 0; different parts of speech are given different weights according to the part of speech standard.

[0033] The hidden state representation of the encoded context is input into the syntactic graph convolutional network. The single-layer graph convolutional network is as follows Figure 4 As shown, the syntactic features are obtained, including:

[0034]

[0035] The syntactic features obtained from the syntactic graph convolutional network are represented as:

[0036] Among them, H syn represents the syntactic features obtained from the syntactic graph convolutional network, Represents the result of the current graph convolutional network, n represents the length of the context, l represents the current graph convolution layer, and ReLU represents the nonlinear function; A ij represents the adjacency matrix constructed according to the syntactic parse tree; p i represents the result of position encoding of the hidden state of the context; W represents the weight matrix, b represents the bias term, di represents the degree of the i-th token, and j represents the sequence number of the token.

[0037] The hidden states of context and aspect are averaged to obtain the guidance vector, which consists of:

[0038] The guidance vector of the context is expressed as:

[0039]

[0040] The guidance vector of the aspect is expressed as:

[0041]

[0042] Among them, v c The guidance vector representing the context, v a represents the guidance vector of the aspect, n and m represent the length of the context and aspect words respectively; and Represent the hidden states of context and aspect words respectively.

[0043] The hidden state representations of the context and the aspect are encoded using the guidance vectors respectively, so as to obtain the context representation integrated with the aspect and the aspect representation integrated with the context, including:

[0044] The hidden state representation of the aspect is encoded using the guidance vector to obtain the contextual representation that incorporates the aspect:

[0045]

[0046] Among them, s c Represents the contextual representation of the integration aspect; α i represents the attention weight, v a Represents the guidance vector of the aspect, n represents the length of the context, exp() represents the exponential function, and score represents the hidden state of the calculated context A scoring function for importance in context, W represents the weight matrix, represents the hidden state of the jth word.

[0047] Similarly, the hidden state representation of the context is encoded using the guidance vector to obtain the aspect representation s that is integrated into the context a .

[0048] The interaction matrix is ​​obtained based on the hidden state representation of the encoded context and aspects, including:

[0049] I ij =s c s aT

[0050] Among them, I ij represents the interaction matrix, s c The context representation of the integration aspect, s a To incorporate the contextual representation, T represents the transpose operation.

[0051] Input the interaction matrix into the semantic graph convolutional network to obtain semantic features;

[0052]

[0053] The semantic features obtained from the semantic graph convolutional network are expressed as:

[0054] Among them, H sem represents the semantic features obtained from the semantic graph convolutional network, Represents the result of the current graph convolutional network, n represents the length of the context, l represents the current graph convolution layer, W represents the weight matrix, b represents the bias term, ReLU represents the nonlinear function, and h represents the j represents the hidden state of the jth token; I ij represents the interaction matrix.

[0055] The hidden state representation of non-aspect words in the syntactic features output by the syntactic graph convolutional network and the semantic features output by the semantic graph convolutional network are shielded, and the final syntactic features and semantic features are obtained through average pooling operations.

[0056] include:

[0057]

[0058]

[0059] in, and Respectively, they represent filtering out non-aspect words and retaining only high-level aspect-specific syntactic and semantic features, MeanPool represents the average pooling operation, represents the grammatical features of the i-th aspect word, represents the semantic feature of the i-th aspect word, and m is the length of the aspect word.

[0060] The final syntactic features and semantic features are connected as the final features, including:

[0061]

[0062] Among them, u represents the final feature of the connection between syntactic features and semantic features, and Respectively, they represent filtering out non-aspect words and retaining only high-level aspect-specific syntactic and semantic features.

[0063] Predict sentiment polarity based on the final features, including:

[0064] P = softmax(W p u+b p )

[0065] Where P∈R dp represents the predicted sentiment polarity, W p and b p They represent the trainable weight matrix and bias respectively; R represents a set of real numbers, and dp represents the number of categories of sentiment polarity.

[0066] The training strategy of graph convolutional network is: choose to have L 2 The cross entropy function of the regularization term is used as the loss function in the neural network. The data is batch normalized so that the data satisfies the normal distribution with a mean of 0 and a variance of 1, avoiding the gradient vanishing problem and the gradient exploding problem generated during the training process, speeding up the model training speed and improving the model generalization ability. The optimization method of the neural network adopts the standard gradient descent algorithm, and the deep learning framework adopted is Pytorch. The parameters of the deep neural network are learned and determined by training to continuously reduce the function value of the objective function.

[0067]

[0068] Among them, C represents the training data set, And represents the true sentiment polarity label, Indicates the first elements, Θ represents all training parameters, λ represents L 2 Regularization coefficient.

[0069] The trained model is used to classify the sentiment polarity of the aspect words in the text to be analyzed and the output results are evaluated, using accuracy and F1-Measure as evaluation indicators.

[0070]

[0071] Among them, TP+TN represents the number of samples predicted correctly, and C represents the total number of samples.

[0072]

[0073] Among them, P is the precision rate, which means the real true samples in the samples predicted to be true, and it refers to the prediction results; R is the recall rate, which means how many true samples in the samples are predicted to be true, and it refers to the original samples.

[0074] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A dual-graph convolutional network-based aspect-level sentiment analysis method. It is characterized in that include: S1. Obtain the context and aspects of the text to be analyzed, and convert the context and aspects of the text to be analyzed into word vector embedding representations through the GloVe word vector model; S2, using a bidirectional long short-term memory network to extract features from the context and aspect word vector embedding representations to obtain the hidden state representations of the context and aspect; S3, position encoding the hidden state representation of the context, and inputting the encoded hidden state representation of the context into the syntactic graph convolutional network to obtain syntactic features; S4, averaging the hidden states of the context and the aspect to obtain a guide vector, using the guide vector to encode the hidden state representations of the context and the aspect respectively, to obtain a context representation integrated with the aspect and an aspect representation integrated with the context, obtaining an interaction matrix based on the context representation integrated with the aspect and the aspect representation integrated with the context, and inputting the interaction matrix into a semantic graph convolutional network to obtain a semantic feature; The hidden state representations of the context and the aspect are encoded using the guidance vectors respectively, so as to obtain the context representation integrated with the aspect and the aspect representation integrated with the context, including: The hidden state representation of the aspect is encoded using the guidance vector to obtain the contextual representation that incorporates the aspect: Among them, s c Represents the contextual representation of the integration aspect; α i represents the attention weight, v a Represents the guidance vector of the aspect, n represents the length of the context, exp() represents the exponential function, and score represents the hidden state of the calculated context A scoring function for importance in context, W represents the weight matrix, represents the hidden state of the jth word; Similarly, the hidden state representation of the context is encoded using the guidance vector to obtain the aspect representation s that is integrated into the context a ; The interaction matrix is ​​obtained based on the hidden state representation of the encoded context and aspects, including: I ij =s c s aT Among them, I ij represents the interaction matrix, s a To incorporate the aspect representation into the context, T represents the transpose operation; Input the interaction matrix into the semantic graph convolutional network to obtain semantic features; The semantic features obtained from the semantic graph convolutional network are expressed as: Among them, H sem represents the semantic features obtained from the semantic graph convolutional network, Represents the result of the current graph convolutional network, n represents the length of the context, l represents the current graph convolution layer, b represents the bias term, ReLU represents the nonlinear function, and h represents the j represents the hidden state of the jth token; S5, shielding the hidden state representation of non-aspect words in the syntactic features output by the syntactic graph convolution network and the semantic features output by the semantic graph convolution network, obtaining the final syntactic features and semantic features through an average pooling operation, connecting the final syntactic features and semantic features as the final features, and predicting the sentiment polarity according to the final features; The hidden state representation of non-aspect words in the syntactic features output by the syntactic graph convolutional network and the semantic features output by the semantic graph convolutional network are shielded, and the final syntactic features and semantic features are obtained through average pooling operations, including: in, and Respectively, they represent filtering out non-aspect words and retaining only high-level aspect-specific syntactic and semantic features, MeanPool represents the average pooling operation, represents the grammatical features of the i-th aspect word, represents the semantic feature of the i-th aspect word, and m is the length of the aspect word.

2. According to claim 1, an aspect-level sentiment analysis method based on a dual graph convolutional network, It is characterized in that Positional encoding of the hidden state representation of the context, including: Among them, p i represents the result of position encoding of the hidden state representation of the context; F() represents the function of assigning position weights to enhance the importance of context words close to aspect words, in order to reduce the noise and bias naturally generated during dependency parsing w i represents the position weight of the i-th token, τ represents the position of the aspect word, i represents the i-th word in the sentence, m represents the length of the aspect word, n represents the length of the context, Represents the hidden state of the context.

3. According to claim 1, the aspect-level sentiment analysis method based on dual graph convolutional network, It is characterized in that The hidden state representation of the encoded context is input into the syntactic graph convolutional network to obtain syntactic features, including: The syntactic features obtained from the syntactic graph convolutional network are represented as: Among them, H syn represents the syntactic features obtained from the syntactic graph convolutional network, Represents the result of the current graph convolutional network, n represents the length of the context, l represents the current graph convolution layer, and ReLU represents the nonlinear function; A ij represents the adjacency matrix constructed according to the syntactic parse tree, p i represents the result of position encoding of the hidden state of the context, W represents the weight matrix, b represents the bias term, di represents the degree of the i-th token, and j represents the sequence number of the token.

4. The aspect-level sentiment analysis method based on dual graph convolutional network according to claim 1, It is characterized in that The hidden states of context and aspect are averaged to obtain the guidance vector, which consists of: The guidance vector of the context is expressed as: The guidance vector of the aspect is expressed as: Among them, v c The guidance vector representing the context, v a represents the guidance vector of the aspect, n and m represent the length of the context and aspect words respectively, and Represent the hidden states of context and aspect words respectively.

5. The aspect-level sentiment analysis method based on dual graph convolutional network according to claim 1, It is characterized in that The final syntactic features and semantic features are connected as the final features, including: Among them, u represents the final feature of the connection between syntactic features and semantic features, and Respectively, they represent filtering out non-aspect words and retaining only high-level aspect-specific syntactic and semantic features.

6. The aspect-level sentiment analysis method based on dual graph convolutional network according to claim 1, It is characterized in that Predict sentiment polarity based on the final features, including: P=softmax(W p u+b p ) Where P∈R dp represents the predicted sentiment polarity, W p and b p They represent the trainable weight matrix and bias respectively; R represents a set of real numbers, and dp represents the number of categories of sentiment polarity.

Citation Information

Patent Citations

  • Emotion classification method utilizing graph convolutional neural network and Chinese syntax

    CN112001186A

  • Aspect-level sentiment analysis method and device based on graph convolutional neural network

    CN112528672A