An Aspect-Level Sentiment Analysis Method Based on Hybrid Weights and Dual-Channel Graph Convolution
The integration of word nature, syntactic relations, positional features, and emotional polarity knowledge through dual-channel graph convolutional networks enhances aspect-level sentiment analysis, addressing the limitations of existing methods and improving classification accuracy.
Patent Information
- Application Number
- CN202310053263.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-02-03
AI Technical Summary
The existing aspect-level emotion analysis methods fail to fully combine knowledge of part-of-speech, grammatical relationship, location, physical distance and emotional polarity, resulting in a low accuracy of emotion classification.
Using a method based on mixed weights and dual-channel graph convolution, we incorporate part of speech, grammatical relationships and position features into word vector representation, and construct a grammatical distance weight enhancement graph and emotional polarity structure graph, and use the Bi-GRU model to obtain context information, and combine attention mechanisms and classification functions to predict emotional polarity.
The accuracy and performance of aspect-level emotion classification are improved. By comprehensively considering multiple characteristic information, noise interference is reduced and the ability to judge the emotional polarity of the other aspect-level words is enhanced.
Smart Images

Figure CN116340507B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to an aspect-level sentiment analysis method based on hybrid weights and dual-channel graph convolution. Background Art
[0002] Aspect-level sentiment analysis (ALSA) is a fine-grained sentiment analysis task, whose purpose is to automatically classify the sentiment related to a specific aspect of the text. The research work of ALSA is carried out based on machine learning methods. Wu et al. proposed a sentiment analysis method based on a probabilistic graphical model, combined with a support vector machine for classification, and had good performance in experiments. However, traditional machine learning-based methods rely on the quality of constructing feature engineering, which limits the development of ALSA. Deep learning is not subject to these limitations and has achieved great success in various sentiment analysis tasks. Currently, most research on aspect-level sentiment classification mainly focuses on improving grammar or obtaining semantic information through neural networks. Among them, syntactic information is mostly used to obtain the structure of sentences through dependency trees. A dependency tree refers to the syntactic relationship between words in a sentence, which appears in the form of a triplet to form a tree structure of a sentence. On this basis, Hou XC et al. generated dependency trees with different syntactic structures through multiple parsings and combined them into a directed graph network for training. Wang K et al. constructed a dependency tree with the aspect word as the root node and combined it with a relational graph attention network to achieve sentiment prediction. Su Jintian et al. defined the syntactic distance and syntactic distance weight based on the syntactic dependency tree and finally achieved good classification results. Although these methods improve the syntactic structure or prune the syntactic dependency tree, they ignore the research on sentiment polarity knowledge, resulting in the problem of low accuracy of aspect-level sentiment classification.
[0003] Sentiment knowledge is usually used to enhance the sentiment feature representation in sentiment analysis tasks. Ma et al. incorporated the common sense knowledge of sentiment-related concepts into a long short-term memory network for aspect-oriented sentiment classification. Bin Liang et al. introduced the SenticNet dictionary to polish the syntactic dependency tree. Existing research has not effectively incorporated the sentiment polarity label and sentiment polarity value in sentiment knowledge into the sentiment classification model, lacking in-depth mining of sentiment information.
[0004] To improve the classification performance, P. Chen et al. integrated the linear position information of a given aspect into the aspect-level sentiment classification model. Li et al. added a masking mechanism and sentence linear position encoding on the basis of GCN and used the attention mechanism for classification, achieving better classification results compared with previous methods. However, they lacked comprehensive consideration of part-of-speech features and physical distance features, resulting in the model's inability to focus on context words that have a greater impact on aspect words. Therefore, combining part-of-speech features and physical distance features is also a research direction worthy of consideration. In addition, Fu Chaoyan et al. incorporated part-of-speech and syntactic relationship into the vector representation of each word to enrich the vector representation of words, but lacked consideration of the position features of words. Therefore, integrating position features into the vector representation of words can be more conducive to semantic learning.
[0005] The above methods have all been well verified in practice, but they all study sentiment classification from a certain perspective. There is still a lack of a method that comprehensively considers part-of-speech, syntactic relationship, position, physical distance, syntactic distance, and sentiment polarity knowledge, etc., which are important features for accurately identifying the sentiment of aspect words. The mining of sentiment polarity knowledge is also insufficient, and there is a lack of consideration of the mixed weights of part-of-speech and physical distance. Finally, the present invention proposes an aspect-level sentiment analysis method based on mixed weights and dual-channel graph convolution. Summary of the Invention
[0006] The present invention provides an aspect-level sentiment analysis method based on mixed weights and dual-channel graph convolution to solve the problem of inaccurate aspect-level sentiment analysis caused by insufficient extraction of semantic features, syntactic dependency relationship features, and external sentiment polarity knowledge features in the prior art.
[0007] The present invention provides an aspect-level sentiment analysis method based on mixed weights and dual-channel graph convolution, including the following steps:
[0008] Step 1: Generate a vector representation for each word in the sentence;
[0009] Step 2: Incorporate the part-of-speech, syntactic relationship, and position of each word in the sentence into the vector representation of each word;
[0010] Step 3: Based on the vector representation of each word obtained in Step 2, use the Bi-GRU model to obtain the context information of each word, where the context information includes: the context information of the aspect word;
[0011] Step 4: Obtain the part-of-speech-distance mixed weight of the context words of each aspect word in the sentence relative to the aspect word;
[0012] Step 5: Construct a dual-channel graph convolutional network. The dual-channel graph convolutional network performs convolutional operations on the syntax distance weight enhanced graph and the sentiment polarity structure graph of the sentence respectively to obtain the output feature vector of the graph convolutional network based on syntax distance and the output feature vector of the graph convolutional network based on sentiment polarity;
[0013] Among them, the syntax distance weight enhanced graph adds syntax distance on the basis of the syntactic dependency tree;
[0014] The sentiment polarity structure graph adds sentiment polarity labels and sentiment polarity values on the basis of the syntactic dependency tree;
[0015] Step 6: Respectively perform aspect masking on the two feature vectors obtained in Step 5 to obtain a feature vector containing only the hidden features of the aspect words;
[0016] Step 7: Perform attention weight allocation on the output feature vector of the graph convolutional network based on syntax distance containing only the hidden features of the aspect words obtained in Step 6 and the context information of the aspect words in Step 3 through the attention mechanism to obtain the feature vector processed by the attention mechanism; among them, the context information of the aspect words in Step 3 serves as the key matrix and value matrix of the attention mechanism, and the output feature vector of the graph convolutional network based on syntax distance containing only the hidden features of the aspect words obtained in Step 6 serves as the query matrix of the attention mechanism;
[0017] Step 8: Concatenate and fuse the feature vector processed by the attention mechanism obtained in Step 7 and the output feature vector of the graph convolutional network based on sentiment polarity containing only the hidden features of the aspect words obtained in Step 6, and then input it into the classification function, and use the output result of the classification function as the sentiment polarity prediction result of the target word.
[0018] Furthermore, the specific process of Step 2 is as follows:
[0019] First, map the part-of-speech, syntactic relationship, and position of each word in the sentence to a low-dimensional, continuous, and dense space to obtain the part-of-speech embedding, syntactic relationship embedding, and position embedding, and then integrate the part-of-speech embedding, syntactic relationship embedding, and position embedding into the vector representation of each word to complete the integration process.
[0020] Furthermore, the specific process of Step 4 is as follows:
[0021] Obtain the physical distance between the current aspect word and each context word corresponding to the current aspect word, and assign physical distance weights from high to low to each context word according to the physical distance from near to far;
[0022] Obtain the adjectives in each context word corresponding to the current aspect word, assign high-weight part-of-speech weights to the adjectives, and assign zero part-of-speech weights to other words;
[0023] Add the physical distance weight of each context word corresponding to the current aspect word to the part-of-speech weight to obtain the part-of-speech-distance hybrid weight of each context word corresponding to the current aspect word;
[0024] Further, in step 4, it also includes obtaining the articles in the current aspect word and each context word corresponding to the current aspect word, assigning a high weight part-of-speech weight to adjectives, assigning a low weight part-of-speech weight to articles, and assigning a zero part-of-speech weight to other words.
[0025] Further, when the dual-channel graph convolutional network performs a convolution operation on the grammar distance weight enhanced graph of the sentence, it includes:
[0026] Weight the output of the current convolution of the graph convolutional network based on grammar distance in the dual-channel graph convolutional network with the part-of-speech-distance hybrid weight of the vector representation as the input for the next convolution.
[0027] Advantages of the present invention:
[0028] In the present invention, since part-of-speech, grammatical relations, and position are incorporated into the vector representation of each word to enrich the vector representation of each word, and a hybrid weight is constructed using part-of-speech weights and physical distance weights to exclude the interference of context words that are not important for the aspect word, and finally the grammatical distance feature and the sentiment polarity knowledge feature are fused so that the method captures information about the aspect word from multiple perspectives. Therefore, the aspect-level sentiment classification performance of the present invention is superior to other methods;
[0029] In the present invention, since part-of-speech embedding, grammatical relation embedding, and position embedding are incorporated into the vector representation of each word, the vector representation of the word can be enriched. Therefore, the present invention enables the model to learn more information and helps to improve the classification effect;
[0030] In the present invention, since the hybrid weight based on part-of-speech features and physical distance features can give greater weight to key information in the sentence linear structure, the present invention can exclude the noise and deviation brought by unimportant context words;
[0031] In the present invention, since the grammatical distance feature and the sentiment polarity knowledge (sentiment polarity label, sentiment polarity value) are fused on the sentence tree structure, the present invention can capture information about the aspect word from different angles and helps to improve the accuracy of sentiment classification. Description of the Drawings
[0032] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings. The drawings are schematic and should not be construed as imposing any limitation on the present invention. In the drawings:
[0033] Figure 1 It is a schematic flow chart of a specific embodiment of the present invention;
[0034] Figure 2 The part-of-speech-distance hybrid coding diagram of the specific embodiment of the present invention;
[0035] Figure 3 The grammar distance weight enhancement diagram of the specific embodiment of the present invention;
[0036] Figure 4 The emotional polarity structure diagram of the specific embodiment of the present invention;
[0037] Figure 5 The matrix representation of the emotional polarity structure diagram of the specific embodiment of the present invention;
[0038] Figure 6 The relationship diagram between the number of graph convolutional network layers and the accuracy rate under the Lap14, Rest14, and Twitter data sets in the specific embodiment of the present invention;
[0039] Figure 7 The relationship diagram between the number of graph convolutional network layers and the macro F1 value under the Lap14, Rest14, and Twitter data sets in the specific embodiment of the present invention. Detailed implementation manners
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] As Figure 1 shown, the embodiment of the present invention provides an aspect-level sentiment analysis method based on hybrid weights and dual-channel graph convolution, including the following steps:
[0042] Step 1: Generate vector representations for each word in the sentence;
[0043] The word embedding layer can generate vector representations for each word in the sentence. In this paper, the GloVe model is used to generate vector representations for each word. GloVe is based on LSA and Word2vec, combines the advantages of both, complements each other's disadvantages, and realizes a non-deep learning upgraded language model. Therefore, the Glove model can be trained faster, can be extended to large-scale corpora, is also applicable to small-scale corpora and small vectors, and usually has better final results. Given an aspect-word sentence pair (a, s), where: a = {w t , w t+1 , …, w t+m} is the sentence s = {w1, w2, …, a, …, wi , …, w n ; the subsequence of aspect words of {w1, w2, …, wm}; m is the sentence length, and m is the aspect word length. The embedding matrix representation of the sentence obtained by GloVe is where: n represents the number of words in the sentence; d m represents the embedding dimension of the word. In this way, the vector representation of each word w i is obtained
[0044] Step 2: Incorporate the part-of-speech, syntactic relationship, and position of each word in the sentence into the vector representation of each word;
[0045] Furthermore, first map the part-of-speech and syntactic relationship of each word in the sentence to a low-dimensional, continuous, and dense space to obtain the part-of-speech embedding, syntactic relationship embedding, and position embedding, and then incorporate the part-of-speech embedding, syntactic relationship embedding, and position embedding into the vector representation of each word to complete the incorporation process;
[0046] Drawing on the idea of the word embedding method, map the part-of-speech, syntactic relationship, and position to a low-dimensional, continuous, and dense space, and respectively obtain the part-of-speech embedding syntactic relationship embedding and position embedding and splice them behind the embedding matrix W of the sentence to obtain the input matrix as the semantic learning layer where: d Pos is the part-of-speech embedding dimension; d Dep is the syntactic relationship embedding dimension; d Post is the position embedding dimension. In this way, the vector representation of a word contains part-of-speech features, syntactic relationship features, and position features.
[0047] Step 3: Based on the vector representation of each word obtained in Step 2, use the Bi-GRU model to obtain the context information of each word, where the context information includes: the context information of the aspect word;
[0048] The Bidirectional-Gated Recurrent Unit (Bi-GRU), as a variant of the recurrent neural network, can better learn the long-distance dependencies and context features in the sentence, and can effectively alleviate the problems of gradient disappearance and gradient explosion. Moreover, Bi-GRU has one less gate unit than Bi-LSTM, saving network training time and improving efficiency while maintaining almost the same accuracy. Therefore, the Bi-GRU model is selected in the semantic learning layer to obtain the context information of each word. The hidden states of the k-th layer learned by Bi-GRU in the forward and backward directions of I are respectively represented as and Finally, splice them to obtain the hidden state representation where d h is the dimension of the Bi-GRU hidden state. In this way, the context information of each word, that is, the semantic information, is obtained.
[0049] Step 4: Obtain the part-of-speech-distance hybrid weight of each context word relative to the aspect word in the sentence, as Figure 2 shown,
[0050] Specifically:
[0051] Obtain the physical distance between the current aspect word and each context word corresponding to the current aspect word, and assign physical distance weights from high to low to each context word according to the physical distance from near to far;
[0052] Obtain the adjectives and articles in each context word corresponding to the current aspect word, assign high-weight part-of-speech weights to adjectives, assign low-weight part-of-speech weights to articles, and assign zero part-of-speech weights to other words;
[0053] Add the physical distance weight and the part-of-speech weight of each context word corresponding to the current aspect word to obtain the part-of-speech-distance hybrid weight of each context word corresponding to the current aspect word;
[0054] For each context word w corresponding to the aspect word i The influence of the physical distance from the current aspect word is also different on the sentiment analysis of the aspect word. Therefore, it is necessary to calculate the physical distance weight for text words. The calculation method is as follows:
[0055]
[0056] In the formula: p i is the physical distance weight of the word with index i; j τ and j τ+m are the start index and end index of the aspect word.
[0057] In addition, in order to better remove noise and bias, different part-of-speech weights are assigned to context words based on their part of speech. Adjectives in the sentence have a greater impact on the sentiment analysis of the aspect word, and the influence of adjectives on the sentiment analysis of the aspect word should be increased; articles have a smaller impact on the sentiment analysis of the aspect word, and their influence on the sentiment analysis of the aspect word should be reduced; the part-of-speech weights of other context words are assigned zero. The calculation method of part-of-speech weights is as follows:
[0058]
[0059] In the formula: m iIt represents the part-of-speech weight with the word index being i; J represents adjectives; K represents articles; α = 2 is the part-of-speech range. Here, the value of α being 2 is set based on the basic structure of English sentences. A too large value will introduce noise, and a too small value will lose key sentiment information.
[0060] Then, combine the physical distance weight and the part-of-speech weight of each context word corresponding to the current aspect word to obtain the part-of-speech-distance hybrid weight q=(q1, q2, …, q i , …, q n ):
[0061] q i = p i + m i (3)
[0062] Finally, introduce the part-of-speech-distance hybrid weight q i to update the hidden state representation h s of the semantic learning layer to obtain a new hidden state representation The calculation method is as follows:
[0063] H G = F(h s ) = q i h s (4)
[0064] In the formula: F(·) represents the weight function; q i represents the part-of-speech-distance hybrid weight with the word index being i.
[0065] Step 5: Construct a two-channel graph convolutional network. The two-channel graph convolutional network performs convolutional operations on the syntactic distance weight enhanced graph and the sentiment polarity structure graph of the sentence respectively to obtain the output feature vector of the graph convolutional network based on syntactic distance and the output feature vector of the graph convolutional network based on sentiment polarity;
[0066] Among them, when the two-channel graph convolutional network performs convolutional operations on the syntactic distance weight enhanced graph of the sentence, it includes:
[0067] Weight the output of the current convolution of the graph convolutional network based on syntactic distance in the two-channel graph convolutional network with the part-of-speech-distance hybrid weight of the vector representation as the input for the next convolution
[0068] The syntactic distance weight enhanced graph adds syntactic distance on the basis of the syntactic dependency tree;
[0069] The sentiment polarity structure graph adds sentiment polarity labels and sentiment polarity values on the basis of the syntactic dependency tree;
[0070] This paper uses the Stanford CoreNLP tool to generate the dependency tree T = {V, E, A} of sentences. Where: V is the set of nodes; E is the set of node pairs with syntactic dependency relationships; |V| is the number of nodes in the sentence; A ∈ R |V|×|V| is the adjacency matrix. At the same time, self-connections of nodes are added. Finally, the weight calculation method between nodes v i ∈ V and v j ∈ V is as follows:
[0071]
[0072] Considering that the node representation of the dependency tree needs to include more comprehensive syntax, the present invention introduces syntactic distance. The syntactic distance refers to the shortest distance from the context word of the aspect word above the dependency tree to the aspect word, reflecting the syntactic association degree between the aspect word and the context word of the aspect word. For example Figure 3 the syntactic distance between "features" and "great" in is 1, and the syntactic distance between "features" and "that" is 2. The syntactic distance between each context word of the aspect word and the aspect word can be represented by a vector D = (d 1,a , d 2,a , …, d i,a , …, d n,a ), where d i,a is the syntactic distance between each context word w i of the aspect word and the aspect word. Based on the syntactic distance, the syntactic distance weight l i is calculated, and the calculation method is as follows, where d max is the maximum syntactic distance value in the sentence:
[0073]
[0074] Using the syntactic distance weight l i to update the adjacency matrix A to obtain the adjacency matrix A G enhanced by the syntactic distance weight. Finally, the syntactic distance weight enhanced graph is represented as G G = {V G , E G , A G}, where V G is the same as V; E G is the same as E.
[0075] In sentence structure, only considering the syntactic relationship will result in the loss of some key emotional information. Different nodes not only have syntactic relationships but also emotional polarity features, which are crucial for the judgment of emotional polarity. To improve the accuracy of emotion classification, emotion polarity labels and emotion polarity values are used as knowledge sources and embedded into the dependency tree. First, obtain the emotion polarity labels from SenticNet. Emotion polarity labels include positive labels, negative labels, and neutral labels, represented by "positive", "negative", and "neutral" respectively. Then, as Figure 4 shown in the example, different emotion polarity labels are added to words with emotional polarity features such as "old", "great", and "offers", which can support polarity prediction. In the constructed emotion polarity structure graph G S = {V S , E S , A S}, the edge set E S includes the set of node pairs with syntactic dependency relationships and the set of node pairs with emotional polarity relationships. V s is the node set, |V S | = n + 3, where n is the number of word nodes in sentence s, and 3 is the number of emotion polarity labels. The weight i between nodes v j and v is calculated as follows:
[0076]
[0077] Then, in B S , it is not just simply represented by 0 or 1, but there are rich emotion polarity values hidden between different nodes. For example, the emotion polarity value of "good" is 0.191, and the emotion polarity value of "old" is -0.81, which can be easily obtained from SenticNet. By calculating the emotion polarity values between nodes, the model can focus on the parts with more explicit emotional tendencies. The emotion polarity value S i between nodes v j and v ij is calculated as follows:
[0078] S ij = |Sent(v i ) + Sent(v j )| (8)
[0079] Where: Sent(v) ∈ [-1, 1] represents the mapped sentiment polarity value of node v in SenticNet. In SenticNet, a strongly positive sentiment polarity value is very close to 1, while a strongly negative sentiment polarity value is close to -1. Sent(v) = 0 indicates that v is a neutral word or not in SenticNet. As Figure 5 shown, and then the adjacency matrix A of the sentiment polarity structure graph is obtained according to the following formula s :
[0080]
[0081] In addition, the established affective - space sentiment word embedding space document maps the concepts of SenticNet to a continuous low - dimensional embedding matrix E aff , without losing the semantics and sentiment relevance of the original space. By looking up the embedding matrix E aff to calculate the vector representation of the sentiment polarity nodes where n k is the number of sentiment polarity nodes. The feature matrix serves as the node embedding representation in G S . Each row is the feature vector of the word or sentiment polarity label node.
[0082] After constructing the graph, the grammar distance graph convolution (Grammar distance - GraphConvolutional Network, G - GCN) and the sentiment polarity graph convolution (Sentiment polarity - GraphConvolutional Network, S - GCN) are used to perform convolution operations on the grammar distance weight - enhanced graph and the sentiment polarity structure graph respectively. The update formula for the nodes in the graph is as follows:
[0083]
[0084]
[0085] Where: and are the node hidden state representations of the previous layer networks of G - GCN and S - GCN respectively. It should be noted that the output of each layer of graph convolution of G - GCN has to go through the part - of - speech - distance hybrid coding weighting to reduce the noise information in the dependency tree, as shown in the formula.
[0086]
[0087] Then the text feature H output by G - GCN is obtained G={h1, h2, …, hn} and the text features output by S-GCN
[0088] Step 6: Respectively perform aspect masking on the two feature vectors obtained in Step 5 to obtain feature vectors that only contain the hidden features of aspect words;
[0089] For the two feature vectors H G and H S Performing aspect masking respectively can exclude the interference of non-aspect words and highlight the importance of aspect words. The positions corresponding to aspect words are set to 1, and the positions corresponding to non-aspect words are set to 0. The calculation method is as follows:
[0090] M mask =[0, 0, 1, 0, 1, 0, …, 0] T (13)
[0091]
[0092] In the formula: M mask represents the aspect masking matrix; represents the feature matrix that only retains aspect words after aspect masking; h K represents the output text feature matrix of the graph convolutional network layer, where h K can be H G or can also be H S . Finally, the output text feature matrices of G-GCN and S-GCN respectively obtain the graph convolutional network output feature vectors based on syntactic distance that only contain the hidden features of aspect words and the graph convolutional network output feature vectors based on sentiment polarity that only contain the hidden features of aspect words
[0093] Step 7: Perform attention weight assignment on the graph convolutional network output feature vectors based on syntactic distance that only contain the hidden features of aspect words obtained in Step 6 and the context information of aspect words in Step 3 through the attention mechanism to obtain the feature vectors processed by the attention mechanism; among them, the context information of aspect words in Step 3 serves as the key matrix and value matrix of the attention mechanism, and the graph convolutional network output feature vectors based on syntactic distance that only contain the hidden features of aspect words obtained in Step 6 serve as the query matrix of the attention mechanism;
[0094] The graph convolutional network output feature vectors based on syntactic distance that only contain the hidden features of aspect words obtained in Step 6 and the context information h of aspect words in Step 3 sAttention weight distribution is carried out through the attention mechanism to highlight the words that play an important role in the sentiment polarity judgment of aspect words, and finally the feature vector r processed by the attention mechanism is obtained. The weight calculation method is as follows:
[0095]
[0096]
[0097]
[0098] In the formula: γ represents the weight to be allocated; the context information h of the aspect word in step 3 s is used as the key matrix and value matrix of the attention mechanism; the output feature vector of the graph convolutional network based on syntactic distance that only contains the hidden features of the aspect word obtained in step 6 is used as the query matrix of the attention mechanism.
[0099] Step 8: Concatenate and fuse the feature vector processed by the attention mechanism obtained in step 7 and the output feature vector of the graph convolutional network based on sentiment polarity that only contains the hidden features of the aspect word obtained in step 6, and then input it into the classification function, and use the output result of the classification function as the sentiment polarity prediction result of the target word.
[0100] Concatenate and fuse the feature vector r processed by the attention mechanism obtained in step 7 and the output feature vector of the graph convolutional network based on sentiment polarity that only contains the hidden features of the aspect word obtained in step 6 to obtain the feature vector In this way, both the syntactic distance feature and the sentiment polarity feature are retained.
[0101] Then pass z to the fully connected softmax classification function. The output of the classification function is the probability distribution of different sentiment polarities. The sentiment polarity prediction result of the target word is obtained according to the probability distribution. End-to-end training of the model is realized by backpropagation, where the objective function Q to be minimized is the cross-entropy error, and the formula is as follows:
[0102]
[0103]
[0104] y = (a, s) (20)
[0105] In the formula: y represents the data sample; D represents the total number of samples; G represents the sentiment category; y c (y) represents the sentiment polarity; represents the predicted sentiment polarity.
[0106] The specific embodiments of the present invention provide the following experimental proofs:
[0107] The performance of the evaluation method of the present invention on SemEval 2014 includes restaurant reviews (Rest14) and laptop reviews (Laptop14). The present invention also conducts experiments on the Twitter dataset. The statistical summary of the dataset is shown in Table 1, the sample distribution divided by class labels on the public dataset:
[0108]
[0109] Table 1
[0110] A series of experiments conducted by the present invention uniformly select the PyTorch framework for implementation. The selected word vector dimension is 300, the batch size is 32, the learning rate is 0.01, the optimizer selects Adam, and the Dropout of Bi-GRU and graph convolution are 0.3 and 0.01 respectively. The word vector dimension, syntactic relation dimension, and position dimension are all set to 30.
[0111] The method proposed by the present invention is compared with other methods on the benchmark dataset. The methods considered by the present invention include: Sentic-GCN mainly uses sentiment dictionaries to complete classification tasks; R-GAT reshapes an aspect-based dependency tree by pruning the dependency tree and uses relational graph attention to encode the tree structure; ASGCN uses a graph convolutional network to process dependencies and uses the inter-sentence syntactic dependency structure to solve the long-term dependency problem; Repwalk proposes a new type of neural network that uses a multi-path syntax graph and performs a random walk strategy on the graph; CDT proposes a convolutional dependency model that identifies the sentiment of words for specific aspects in a sentence and fuses the dependency tree with graph convolution for representation learning. The comparison results are shown in Table 2, the comparison experiments of different methods:
[0112]
[0113] Table 2
[0114] It can be seen from the table that graph convolutional models using sentiment dictionaries or improving syntactic dependency trees all have good effects. However, the present invention combines syntactic distance and introduces sentiment polarity labels and sentiment polarity values at the same time. Compared with single-channel models, the dual-channel network can better obtain two kinds of information focusing on syntactic distance features and sentiment polarity features through two different graph convolution operations, which is helpful for the improvement of sentiment analysis tasks.
[0115] To study the independent factors affecting the method classification effect, the present invention sets up several ablation experiments. The factors considered include: grammar distance weight enhanced graph, sentiment polarity structure graph, part-of-speech - position mixed coding, and combinations thereof. M-G is a single-channel network with only the grammar distance weight enhanced graph removed; M-S is a single-channel graph convolutional network with only the sentiment polarity structure graph removed; M-GS is a single-channel graph convolutional network that removes the grammar distance weight enhanced graph and the sentiment polarity structure graph and only uses an ordinary syntactic dependency tree; M-N is a single-channel network with only the part-of-speech features removed; M-P is a single-channel network with only the physical distance features removed; M-PN is a single-channel network with the part-of-speech - position mixed coding removed. The ablation experiment results are shown in Table 3 of the ablation experiment:
[0116]
[0117] Table 3
[0118] It can be found from the experimental results that removing any one of them will cause the accuracy A and macro F1 value of the method to decrease, which indicates the effectiveness of the fusion of part-of-speech coding and physical distance coding and the use of a two-channel network to combine grammar distance and sentiment polarity knowledge. This is mainly because the HCDC-GCN method fuses more feature information.
[0119] To better study the effectiveness of the number of GCN layers L, the number of GCN layers is set to L = {1, 2, 3, 4, 5} respectively. The accuracy A and macro F1 value on the public dataset are shown in Appendix Figure 6 and Appendix Figure 7 respectively.
[0120] It can be found from the experimental results that first, the performance improves as L increases, and then gradually decreases. When HCDC-GCN reaches the best performance at the network layer L = 2. The L-layer GCN model can capture the information of neighbors within L steps. Nodes within three steps are sufficient to complete this task, and too many layers will introduce noise to the model.
[0121] In summary, the present invention proposes an aspect-level sentiment analysis method based on hybrid coding and two-channel GCN, which combines semantic, relationship type, part-of-speech, physical distance, grammar distance, and sentiment polarity knowledge. First, the physical distance features and part-of-speech features are combined on the sentence linear structure to remove noise, and experiments prove that the hybrid coding improves the classification effect of the method. Then, a grammar distance weight enhanced graph is constructed from the grammar level, and a sentiment polarity structure graph is constructed from the sentiment knowledge level. Finally, a two-channel graph convolutional network is used to combine the two, which enables the aspect words to match the context more accurately. The experimental results verify the effectiveness of the method proposed by the present invention in the field of aspect-level sentiment analysis.
[0122] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations fall within the scope defined by the appended claims.
Claims
1. An aspect-level sentiment analysis method based on hybrid weights and dual-channel graph convolution, characterized in that, It includes the following steps: Step 1: Generate vector representations for each word in the sentence; Step 2: Incorporate the part-of-speech, syntactic relationship, and position of each word in the sentence into the vector representation of each word; Step 3: Based on the vector representation of each word obtained in Step 2, use the Bi-GRU model to obtain the context information of each word, where the context information includes: the context information of the aspect word; Step 4: Obtain the part-of-speech-distance hybrid weight of each context word relative to the aspect word in the sentence. The specific process is as follows: Obtain the physical distance between the current aspect word and each context word corresponding to the current aspect word, and assign physical distance weights from high to low to each context word according to the physical distance from near to far; Obtain the adjectives in each context word corresponding to the current aspect word, assign high-weight part-of-speech weights to the adjectives, and assign zero part-of-speech weights to other words; Add the physical distance weights of each context word corresponding to the current aspect word to the part-of-speech weights to obtain the part-of-speech-distance hybrid weight of each context word corresponding to the current aspect word; Step 8: Construct a dual-channel graph convolutional network. The dual-channel graph convolutional network performs convolutional operations on the syntactic distance weight enhanced graph and the sentiment polarity structure graph of the sentence respectively to obtain the graph convolutional network output feature vector based on syntactic distance and the graph convolutional network output feature vector based on sentiment polarity; Among them, the syntactic distance weight enhanced graph adds syntactic distance on the basis of the syntactic dependency tree; The sentiment polarity structure graph adds sentiment polarity labels and sentiment polarity values on the basis of the syntactic dependency tree; Step 11: Perform aspect masking on the two feature vectors obtained in Step 5 to obtain a feature vector that only contains the hidden features of the aspect word; Step 12: Perform attention weight allocation on the graph convolutional network output feature vector based on syntactic distance that only contains the hidden features of the aspect word obtained in Step 6 and the context information of the aspect word in Step 3 through the attention mechanism to obtain the feature vector processed by the attention mechanism; among them, the context information of the aspect word in Step 3 serves as the key matrix and value matrix of the attention mechanism, and the graph convolutional network output feature vector based on syntactic distance that only contains the hidden features of the aspect word obtained in Step 6 serves as the query matrix of the attention mechanism; Step 13: Concatenate and fuse the feature vector processed by the attention mechanism obtained in Step 7 and the graph convolutional network output feature vector based on sentiment polarity that only contains the hidden features of the aspect word obtained in Step 6, and then input it into the classification function, and use the output result of the classification function as the sentiment polarity prediction result of the target word.
2. The aspect-level sentiment analysis method based on hybrid weights and dual-channel graph convolution as shown in claim 1, characterized in that The specific process of Step 2 is as follows: First, map the part-of-speech, syntactic relationship, and position of each word in the sentence to a low-dimensional, continuous, and dense space to obtain part-of-speech embedding, syntactic relationship embedding, and position embedding, and then incorporate the part-of-speech embedding, syntactic relationship embedding, and position embedding into the vector representation of each word to complete the incorporation process.
3. The aspect-level sentiment analysis method based on hybrid weights and dual-channel graph convolution as shown in claim 1 or 2, characterized in that In step 4, it also includes obtaining the articles in each context word corresponding to the current aspect word, assigning a high-weight part-of-speech weight to adjectives, assigning a low-weight part-of-speech weight to articles, and assigning a zero part-of-speech weight to other words.
4. The aspect-level sentiment analysis method based on hybrid weights and dual-channel graph convolution as shown in claim 1, characterized in that, When the dual-channel graph convolutional network performs a convolution operation on the sentence's grammatical distance weight enhanced graph in step 5, it includes: The output of the current convolution of the graph convolutional network based on grammatical distance in the dual-channel graph convolutional network is weighted with the part-of-speech-distance hybrid weight of the vector representation as the input for the next convolution.