An emotion - cause pair extraction method and device
The method improves emotion-cause pair extraction by using graph neural networks to integrate syntactic and semantic features, enhancing feature representation and task interaction, thereby improving extraction accuracy.
Patent Information
- Application Number
- CN202310474990.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-04-27
AI Technical Summary
The existing multi-task model cannot fully utilize the promotion effect of sub-tasks in emotional-cause extraction tasks, lacks interactive collaboration between tasks, and has poor clause feature expression ability and lacks external knowledge introduction.
The emotional causal pair extraction method is adopted, combined with the graph neural network and attention mechanism, and the semantic dependence of adjacency matrix and graph attention network are aggregated node information to improve the accuracy of emotional-cause extraction.
Through text semantic dependency analysis and pre-trained model embedding, the graph attention network is used to aggregate node information, which improves the accuracy of emotion-cause extraction, solves the problem of insufficient interaction and collaboration between tasks, and enhances the clause feature expression ability.
Smart Images

Figure CN116578671B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of deep learning and sentiment analysis, and more particularly relates to a method and device for extracting sentiment-reason pairs. Background Art
[0002] With the rapid development of the Internet, people can publish information and share feelings anytime and anywhere through the Internet. With the rapid increase in the number of text messages posted on various social media, how to analyze an individual's sentiment based on text has become an important research direction in the field of natural language processing (NLP). Currently, the research on sentiment analysis focuses on sentiment classification and sentiment expression. For example, detecting the sentiment expressed by the text author or predicting the sentiment of the text reader, with the aim of analyzing the sentiment expressed towards a certain aspect of things. In addition to obtaining surface information, researchers also desire to extract and analyze deep information about sentiment. Through the research on sentiment analysis problems, sentiment causes are considered to be one of the key elements for in-depth sentiment analysis.
[0003] Regarding the research on the potential causes behind certain sentiments in text, in order to better explore the connections between sentiment causes, researchers have proposed the task of extracting sentiment-reason pairs. For the research on the task of extracting sentiment-reason pairs, two subtasks of sentiment extraction and reason extraction have been derived. To solve these two subtasks and the extraction of sentiment-reason pairs, step-by-step methods and end-to-end methods are usually used.
[0004] Subsequently, in the method of extracting sentiment-reason pairs, linguistic knowledge, deep learning models, sequence labeling, and attention mechanisms are mostly used to solve this task. Researchers use the above methods to design various models for extracting sentiment-reason pairs and exploring the causes behind sentiment. However, most of these methods rely heavily on the quality of text features, and there is a lack of interaction between features, and the content included is relatively single.
[0005] Although all of the above technologies and models can provide relatively good results, the research trend in the task of extracting sentiment-reason pairs has been committed to solving two problems: the clause feature expression ability is poor, and there is a lack of interaction between multiple tasks.
[0006] In order to represent the natural language of a clause as a numerical matrix that can be processed by a machine, there must be a transformation from word-level features to clause-level features. For the changes in feature vectors of different granularities, researchers have tried to generate text feature vectors using pre-trained word vectors, and then use the attention mechanism to select and fuse clause features. On this basis, a classifier based on sentiment-reason pairs is trained to perform sentiment-reason pair extraction. However, word vector models such as word2vec and Golve, and the clause feature generation mechanism, generate feature vectors with single content and poor semantics. Moreover, existing feature fusion methods only focus on modeling the content relationship in the text and ignore the introduction of external knowledge.
[0007] Most current multi-task methods first predict sentiment clauses and reason clauses, and then encode distance information to obtain fused features under the multi-task model through concatenation. However, existing multi-task models cannot fully utilize the promoting effects of the two sub-tasks on the sentiment-reason pair extraction task and lack interaction and collaboration between tasks. Summary of the Invention
[0008] An embodiment of the present invention provides a method and device for extracting sentiment-reason pairs, which are used to solve the problems that existing multi-task models cannot fully utilize the promoting effects of the two sub-tasks on the sentiment-reason pair extraction task, lack interaction and collaboration between tasks, and have poor clause feature expression ability and lack the introduction of external knowledge.
[0009] An embodiment of the present invention provides a method for extracting sentiment-reason pairs, including:
[0010] Determine the first clause feature and the first clause feature vector corresponding to the ECPE text according to the word hidden state weights and the hidden states of each word included in the ECPE text extracted by sentiment causal pair extraction.
[0011] Obtain a semantic dependency adjacency matrix according to the syntactic dependency parsing of the hanlp tool; obtain the clause adjacency matrix of the graph neural network according to the first clause feature, the attention mechanism, and the semantic dependency adjacency matrix; obtain the second clause feature and the second clause feature vector corresponding to the graph neural network according to the first clause feature vector and the clause adjacency matrix of the graph neural network.
[0012] Obtain a matching possibility matrix between clause pairs and sentiment-reason pair features according to the second clause feature vector, and obtain the clause pair feature and the clause pair feature vector corresponding to the graph neural network according to the sentiment-reason pair features, the attention mechanism, and the matching possibility matrix between clause pairs.
[0013] Predict the sentiment-reason pair prediction probability by training a classifier on the clause pair feature vector corresponding to the graph neural network, and improve the accuracy of the sentiment-reason pair prediction probability based on the loss function of sentiment-reason extraction.
[0014] Preferably, obtaining the clause pair features corresponding to the graph neural network from the features, graph attention network, and matching possibility matrix between clause pairs according to emotional reasons specifically includes:
[0015] Obtaining the clause pair attention weight matrix from the features and graph attention network according to emotional reasons, and obtaining the clause pair adjacency matrix corresponding to the graph neural network based on the clause pair attention weight matrix and the matching possibility matrix between clause pairs;
[0016] According to the clause pair adjacency matrix corresponding to the graph neural network and the features of emotional reasons, the clause pair features corresponding to the graph neural network are obtained through the following formula:
[0017]
[0018] where represents the clause pair ij feature corresponding to the t-th layer of the graph neural network, represents the weight corresponding to the clause pair ij and xy of the t-th layer of the graph neural network in the adjacency matrix, W t , b t are learnable parameters, d ij represents the degree of the node ij in the clause pair adjacency matrix P t of the t-th layer of the graph neural network.
[0019] Preferably, obtaining the second clause feature and the second clause feature vector corresponding to the graph neural network from the first clause feature vector and the clause adjacency matrix of the graph neural network specifically includes:
[0020] The second clause feature corresponding to the graph neural network is as follows:
[0021]
[0022] where represents the second clause feature corresponding to the t-th layer of the graph neural network, represents the weight corresponding to the clauses i and j of the t-th layer of the graph neural network in the adjacency matrix, W t , b t are learnable parameters, d i represents the degree of the node i in the adjacency matrix A t of the t-th layer of the graph neural network.
[0023] Preferably, predicting the emotion-reason pair prediction probability by training a classifier on the clause pair feature vector corresponding to the graph neural network specifically includes:
[0024] For the clause pair feature vector corresponding to the graph neural network through the following formula:
[0025]
[0026] Among them, W p and b p represent learning parameters, represents the sentiment - reason pair prediction probability, represents the clause pair feature vector corresponding to the graph neural network.
[0027] Preferably, obtaining the clause adjacency matrix of the graph neural network according to the first clause feature, the attention mechanism, and the semantic dependency adjacency matrix specifically includes:
[0028] The first clause feature is as follows:
[0029]
[0030] The semantic dependency adjacency matrix is as follows:
[0031]
[0032] Among them, A w represents the adjacency matrix, M clause represents the semantic dependency adjacency matrix, h i represents the first clause feature, α i,j represents the word hidden state weight, represents the hidden state of the word, and R represents the relationship matrix between all words and all clauses.
[0033] Preferably, obtaining the matching possibility matrix between clause pairs and the sentiment - reason pair feature according to the second clause feature vector specifically includes:
[0034] Taking the Cartesian product of the prediction labels of a group of sentiment clauses and the prediction labels of a group of reason clauses to obtain the sentiment - reason label product;
[0035] Taking the Cartesian product of the sentiment - reason label product and the sentiment - reason label to obtain the matching possibility matrix between clause pairs;
[0036] Determining the sentiment - reason pair feature through the following formula:
[0037]
[0038] Among them, represents the feature vector of the second clause i, represents the feature vector of the second clause j, p i,j represents the distance between the second clause i and the second clause j, represents the connection between feature vectors.
[0039] Preferably, before obtaining the matching possibility matrix between clause pairs and the feature of sentiment - reason pairs based on the second clause feature vector, the following steps are further included:
[0040] Input the second clause feature vector into two prediction layers respectively, and obtain the sentiment clause prediction probability and the reason clause prediction probability respectively through the following formulas:
[0041]
[0042]
[0043] According to the sentiment clause prediction probability and the reason clause prediction probability, obtain the loss function of the sentiment clause extraction task and the loss function of the reason clause extraction task respectively through the following formulas:
[0044]
[0045]
[0046] Wherein, represents the sentiment clause prediction probability of clause i, represents the reason clause prediction probability of clause i, W e 、b e 、W c 、b c are learnable parameters, represents the feature vector of the second clause i, represents the sentiment true label of clause i, represents the reason true label of clause i, L emo represents the loss function of the sentiment clause extraction task, L cause represents the loss function of the reason clause extraction task.
[0047] An embodiment of the present invention provides a sentiment - reason pair extraction device, including:
[0048] A first determination unit, configured to determine the first clause feature and the first clause feature vector corresponding to the ECPE text according to the word hidden state weights and the hidden states of each word included in the ECPE text for sentiment - cause pair extraction;
[0049] A first obtaining unit, configured to obtain a semantic dependency adjacency matrix according to the syntactic dependency parsing of the hanlp tool; obtain a clause adjacency matrix of the graph neural network according to the first clause feature, the attention mechanism and the semantic dependency adjacency matrix; obtain the second clause feature and the second clause feature vector corresponding to the graph neural network according to the first clause feature vector and the clause adjacency matrix of the graph neural network;
[0050] A second obtaining unit, configured to obtain a matching possibility matrix between clause pairs and sentiment cause pair features according to the second clause feature vector, and obtain clause pair features and clause pair feature vectors corresponding to the graph neural network according to the sentiment cause pair features, the attention mechanism, and the matching possibility matrix between clause pairs;
[0051] A second determining unit, configured to predict the sentiment-cause pair prediction probability by training a classifier on the clause pair feature vectors corresponding to the graph neural network, and improve the accuracy of the sentiment-cause pair prediction probability based on the loss function of sentiment-cause extraction.
[0052] An embodiment of the present invention provides a computer device, where the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the above-mentioned sentiment-cause pair extraction method.
[0053] An embodiment of the present invention provides a computer-readable storage medium, characterized in that it stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the above-mentioned sentiment-cause pair extraction method.
[0054] An embodiment of the invention provides a sentiment-cause pair extraction method, the method comprising: determining first clause features and a first clause feature vector corresponding to an ECPE text according to the word hidden state weights and the hidden state of each word included in the ECPE text for sentiment causal pair extraction; obtaining a clause adjacency matrix of the graph neural network according to the first clause features, the graph attention network, and the semantic dependency adjacency matrix; obtaining second clause features and a second clause feature vector corresponding to the graph neural network according to the first clause feature vector and the clause adjacency matrix of the graph neural network; obtaining a matching possibility matrix between clause pairs and sentiment cause pair features according to the second clause feature vector, and obtaining clause pair features and clause pair feature vectors corresponding to the graph neural network according to the sentiment cause pair features, the graph attention network, and the matching possibility matrix between clause pairs; predicting the sentiment-cause pair prediction probability by training a classifier on the clause pair feature vectors corresponding to the graph neural network, and improving the accuracy of the sentiment-cause pair prediction probability based on the loss function of sentiment-cause extraction. Based on the original sentiment-cause pair extraction data set, this method converts the text into a graph structure through text semantic dependency analysis and pre-trained model embedding, and uses the graph attention network to aggregate node information under the guidance of prior knowledge to obtain a richer text sentiment cause feature representation. Then, through the aggregation of sentiment-cause pairs and the interaction mechanism between multiple tasks, the accuracy of sentiment-cause pair extraction is effectively improved. This method solves the problems that existing multi-task models cannot fully exert the promotion effect of two sub-tasks on the sentiment-cause pair extraction task, lack of interaction and cooperation between tasks, and the problem of poor clause feature expression ability and lack of introduction of external knowledge. Brief Description of the Drawings
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0056] Figure 1 It is a schematic flowchart of a method for extracting emotion - cause pairs provided by an embodiment of the present invention;
[0057] Figure 2 It is a schematic structural diagram of a device for extracting emotion - cause pairs provided by an embodiment of the present invention. Detailed Embodiments
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0059] Figure 1 It is a schematic flowchart of a method for extracting emotion - cause pairs provided by an embodiment of the present invention; As Figure 1 shown, a method for extracting emotion - cause pairs provided by an embodiment of the present invention mainly includes the following steps:
[0060] Step 101, determine the first clause feature and the first clause feature vector corresponding to the ECPE text according to the attention weights of the hidden states of the words included in the ECPE text and the hidden state of each word;
[0061] Step 102, obtain the clause adjacency matrix of the graph neural network according to the first clause feature, the graph attention network, and the semantic dependency adjacency matrix; obtain the second clause feature and the second clause feature vector corresponding to the graph neural network according to the first clause feature vector and the clause adjacency matrix of the graph neural network;
[0062] Step 103, obtain the matching possibility matrix between clause pairs and the emotion - cause pair feature according to the second clause feature vector, and obtain the clause pair feature and the clause pair feature vector corresponding to the graph neural network according to the emotion - cause pair feature, the graph attention network, and the matching possibility matrix between clause pairs;
[0063] Step 104: Predict the clause pair feature vectors corresponding to the graph neural network through a trained classifier to obtain the sentiment-cause pair prediction probability, and improve the accuracy of the sentiment-cause pair prediction probability based on the loss function extracted from the sentiment-cause.
[0064] Before processing the data, the data needs to be preprocessed. In the embodiments of the present invention, in combination with the ECPE (emotion-cause pair extraction) dataset, the method for extracting sentiment-cause pairs is introduced in detail.
[0065] In practical applications, the ECPE dataset is a Chinese dataset containing text information and labels. The text information includes: file number, number of file clauses, clause number, sentiment category of the clause, sentiment words in the clause, and text content after clause word segmentation. The file label refers to the sentiment clause number and the corresponding cause clause number of the file, that is, the sentiment-cause pair of each file confirmed by the dataset publisher.
[0066] According to the data required by the model for sentiment-cause pair extraction applications, the dataset needs to be screened. The data required by the model includes: file number, number of file clauses, clause number, and text content after clause word segmentation. Therefore, only the data with these five attributes of file number, number of file clauses, clause number, and text content after clause word segmentation in the dataset are retained, and other data are deleted.
[0067] After the above operations, a text dataset with a unified data format is obtained.
[0068] Furthermore, in order to enable the model to have good generalization ability, the dataset needs to be divided. Specifically, after the above data screening, the ECPE dataset includes a total of 1945 file data, among which 1746 files contain one sentiment-cause pair, 177 files contain two sentiment-cause pairs, and 22 files contain more than two sentiment-cause pairs, accounting for 89.77%, 9.10%, and 1.13% respectively. Randomly select 90% of the data as the training set for training, and the remaining 10% of the data as the test set for testing. Finally, the ratio of the training set to the test set is 9:1. This ensures consistency with the full dataset and better improves the generalization ability of the model.
[0069] Furthermore, after the ECPE dataset is divided into a test set and a training set, whether it is the test set data or the training set data, its text data has a "word-clause-document" structure.
[0070] In the embodiments of the present invention, a file is defined as d = {c1, c2,..., c n}, where the file contains n clauses, and c i Define the word sequence of the clause as where |c i | is the length of the clause sequence.
[0071] Furthermore, use the pre-trained BERT (Bidirectional Encoder Representations from Transformers, an autoencoding language) model to extract the features of all texts. First, insert [CLS] and [SEP] tokens at the beginning and end of each clause. Then, provide all the words and two special characters included in each clause to BERT as input. Use the pre-trained BERT model to perform sequence embedding e i,j ={e i,1 , e i,2 , …, e i,m}, where e i,j ∈R m×d , k is the dimension of the word embedding. m = |c i |+l, and l is the sum of the special [CLS] and [SEP] token characters.
[0072] The formula for obtaining the context representation of the words included in the clause can be obtained through the BERT model, as shown in formula (1):
[0073] e i,j = BERT(w i,j ) (1)
[0074] Furthermore, the text feature vectors {e i,1 , e i,2 , …, e i,m} can be obtained through the BERT model. In the embodiments of the present invention, the text feature vectors {e i,1 , e i,2 , …, e i,m} pass through a group of word-level Bi-LSTM neural networks. Each Bi-LSTM network unit corresponds to a word, and each word interacts with its surrounding words to accumulate context information. The hidden state of the j-th word in the i-th clause obtained through Bi-LSTM is h i,j . By capturing the sequence features, a sequence of word hidden states is obtained where each word hidden state h i,j is the concatenation of the forward word hidden state and the backward word hidden state , as specifically shown in formula (2):
[0075]
[0076] Among them, represents the word hidden state, represents the forward word hidden state, represents the backward word hidden state.
[0077] In step 101, after obtaining the hidden state of each word, the attention weight of the word hidden state can be calculated through the attention model. Further, according to the attention weight of each word hidden state and the hidden state of the word, they are aggregated into a clause matrix feature according to the weight, which is called the first clause feature corresponding to the ECPE text here.
[0078] Specifically, the attention weight of each word hidden state is determined by the following formula (3):
[0079]
[0080] The first clause feature corresponding to the ECPE text is determined by the following formula (4):
[0081]
[0082] Among them, represents the feature of the j-th word in the i-th clause, α i,j represents the attention weight of the hidden state of the j-th word in the i-th clause, m represents the number of words included in clause i, W a and b a are both unknowns in deep learning and are automatically calculated according to the backpropagation gradient descent algorithm.
[0083] In practical applications, when the first clause feature h i is determined, the first clause feature sequence {h1, h2,..., h |d|} can be obtained. Further, the clause feature sequence {h1, h2,..., h |d|} is input into the clause-level Bi-LSTM to model the potential context relationship between clauses in the document, and the first clause feature vector {r1, r2,..., r |d|} corresponding to the ECPE text can be obtained.
[0084] Before step 102, the clause semantic dependency relationship is introduced first:
[0085] In practical applications, since most existing methods use attention mechanisms or vector concatenation methods to obtain semantic information, they ignore the dependency relationships between clauses. However, the dependency relationships between clauses contain richer structural information, which helps to reduce information loss and thus understand the text more profoundly. Existing models for clause syntax analysis that are relatively mature can use the model to establish a dependency tree based on the dependency relationships between words, and then obtain the syntactic dependency relationships between clauses through transformation.
[0086] First, use the Hanlp tool to perform syntactic and semantic dependency analysis on the text information, and use the adjacency matrix A w to represent the semantic dependency of words. Then, use R to represent the relationship between clauses and words, as shown in formula (5) specifically:
[0087]
[0088] r i,j represents whether word j is in clause i (logical 1 or 0 in the relationship matrix). If word w j is within clause c i , then r i,j = 1, otherwise r i,j = 0.
[0089] After obtaining A w and R, the semantic dependency relationship between clauses can be as shown in formula (6):
[0090] M clause = RA w R T (6)
[0091] Among them, M clause represents the adjacency matrix of the semantic dependency between clauses obtained from the dependency relationships between words, that is, the semantic dependency adjacency matrix. R represents the relationship matrix between all words and all clauses.
[0092] In step 102, in the embodiment of the present invention, in order to promote the fusion between the first clause feature vectors, the first clause features are used as nodes, and the semantic dependency adjacency matrix M clause of clauses is used to form the edges between nodes, and then all nodes and edges form a semantic graph. Guided by the semantic dependency adjacency matrix, use the graph attention network to perform graph convolution on the clauses. Since emotion and reason are inseparable, the reason clause corresponding to some emotion clauses is the emotion clause itself, so a self-loop needs to be added to the dependency matrix of the clauses before performing graph convolution.
[0093] Specifically, the graph attention calculation formula is as shown in formulas (7) and (8):
[0094]
[0095]
[0096] Among them, since the graph neural network has a multi-layer structure and the current layer is layer t, represents the output of the previous layer. When the current layer is the first layer, its input first clause feature vectors {r1, r2, …, r |d|}, represents the concatenation of vectors, represents the clause feature and the clause feature are vectors concatenated after dimensional transformation; represents all nodes adjacent to node i, represents the attention weight after regularization between node i and neighbor node j, that is represents the attention weight between clauses; ReLU() is the activation function, w t , W t are learnable parameters.
[0097] Furthermore, after determining the attention weights between clauses, the attention weight matrix between all clauses in layer t can be obtained. Further, according to the attention weight matrix between clauses and the semantic dependency adjacency matrix, the clause adjacency matrix of the graph neural network in layer t can be obtained, as shown in the following formula (9):
[0098]
[0099] Among them, A t represents the clause adjacency matrix of the graph neural network in layer t, represents the attention weight matrix between clauses.
[0100] Furthermore, according to the result of multiplying the determined clause adjacency matrix of the graph neural network in layer t by the first clause feature vector for aggregation, the second clause feature corresponding to the graph neural network can be obtained, as shown in the specific formula (10):
[0101]
[0102] Among them, represents the second clause feature corresponding to the graph neural network in layer t, d i represents the degree of node i in the adjacency matrix A t of the graph neural network in layer t, W t , b t are learnable parameters, represents the second clause feature corresponding to the graph neural network in layer t - 1.
[0103] Furthermore, the second clause feature vector of the graph neural network can be obtained through the second clause feature corresponding to the graph neural network
[0104] In step 103, when the second clause feature vector of the graph neural network that fuses the context relationship and semantic features is obtained through the multi-layer graph attention network it can input the same second clause feature vector into two prediction layers respectively to predict whether the clause is an emotion clause or a reason clause, that is, obtain the emotion clause prediction probability and the reason clause prediction probability respectively. Specifically, the emotion clause prediction probability and the reason clause prediction probability are obtained through the following formulas (11) and (12):
[0105]
[0106]
[0107] where represents the emotion clause prediction probability of clause i, represents the reason clause prediction probability of clause i, W e and b e and W c and b c are learnable parameters.
[0108] Furthermore, according to the emotion clause prediction probability and the true label of the emotion clause, the loss function of the emotion clause extraction task can be determined, and according to the reason clause prediction probability and the true label of the reason clause, the loss function of the reason clause extraction task can be determined, as shown in formulas (13) and (14) specifically
[0109]
[0110]
[0111] where represents the emotion clause prediction probability of clause i, represents the true emotion label of clause i. represents the reason clause prediction probability of clause i, represents the true reason label of clause i, L emo represents the loss function of the emotion clause extraction task, L cause represents the loss function of the reason clause extraction task.
[0112] Furthermore, according to the extraction of the above emotion clause prediction probability and reason clause prediction probability, a set of predicted labels of emotion clauses and a set of predicted labels of reason clauses In practical applications, since the predicted label of the sentiment clause or the predicted label of the reason clause represents the magnitude of possibility, the Cartesian product of the predicted label of a sentiment clause and the predicted labels of a group of reason clauses is performed to construct a set of sentiment-reason label products {x 1,1 , x 1,2 , …, x |d|,|d|}, as specifically shown in formula (15):
[0113]
[0114] where x i,j represents the possibility of pairing between sentiment clause i and reason clause j.
[0115] Furthermore, the Cartesian product of the sentiment-reason label product x i,j = {x 1,1 , x 1,2 , …, x |d|,|d|} with itself can obtain the matching possibility matrix between clause pairs, as specifically shown in formula (16):
[0116] M pair = x i,j ×x i,j (16)
[0117] where M pair represents the matching possibility matrix between clause pairs.
[0118] In practical applications, since most of the existing methods for sentiment-reason pairing only utilize the feature vector predicted by the sentiment clause and the feature vector predicted by the reason clause and the distance encoding p i,j are concatenated to form a feature vector composed of three features Then, the pair i,j vector is used as the input, and a classifier is trained to determine whether i and j form a sentiment-reason pair. However, such a simple processing does not consider the information interaction between sentiment-reason pairs and the promotion of subtasks to the overall task.
[0119] In the embodiments of the present invention, a feature vector composed of three features is used to represent each pair of sentiment-reason pairs (sentiment-reason pair features), as shown in formula (17):
[0120]
[0121] where represents the feature vector of the second clause i, represents the feature vector of the second clause j, and p i,jDenotes the distance between the second clause i and the second clause j, p i,j = j - i, Denotes the connection between vectors.
[0122] Furthermore, the sentiment cause pair features Are used as nodes in the graph, M pair As the relational adjacency matrix of the nodes, with M pair Matrix as a guide to aggregate the information between nodes. At the same time, since the M pair Matrix utilizes the results of the subtasks, the subtasks are used to further enhance the extraction accuracy of the sentiment-cause pairs.
[0123] The graph attention calculation formulas are shown in (18) and (19):
[0124]
[0125]
[0126] Among them, since the graph neural network is a multi-layer structure, the current layer is the t-th layer, Denotes the output of the previous layer. When the current layer is the first layer, the input is Denotes all the nodes adjacent to the node pair ij, Denotes the attention weight after regularization of the node pair ij and the neighbor node pair xy, that is, the clause pair attention weight. ReLU() is the activation function, w t 、W t Are learnable parameters.
[0127] Furthermore, according to the clause pair attention weights, the clause pair attention weight matrix corresponding to the t-th layer graph neural network can be obtained. By multiplying the clause pair attention weight matrix and the possibility matrix of the sentiment cause pairs, the clause pair adjacency matrix of the t-th layer graph neural network can be obtained, as shown in formula (20) specifically:
[0128]
[0129] Among them, P t Denotes the clause pair adjacency matrix of the t-th layer graph neural network, Denotes the clause pair attention weight matrix.
[0130] Furthermore, according to the result of multiplying the determined clause pair adjacency matrix of the t-th layer graph neural network and the second clause features, aggregation can be performed to obtain the clause pair features corresponding to the graph neural network, as shown in formula (21) specifically:
[0131]
[0132] Among them, Represents the clause pair ij feature corresponding to the t - layer graph neural network, represents the clause pair ij and xy adjacency matrices of the t - layer graph neural network, represents the clause pair xy feature corresponding to the t - 1 layer graph neural network; represents the clause pair xy feature corresponding to the t - layer graph neural network, W t , b t are learnable parameters, d ij represents the degree of node ij in the clause pair adjacency matrix P t of the t - layer graph neural network.
[0133] Furthermore, the clause pair feature vector corresponding to the graph neural network can be obtained according to the clause pair feature corresponding to the graph neural network that is, the feature vector that fuses different pairs of sentiment reasons is obtained through a multi - layer graph attention network.
[0134] In step 103, after obtaining the clause pair feature vector corresponding to the graph neural network obtained through aggregation, a classifier can be trained to predict the pairs containing causal relationships, and a label as a sentiment - cause pair is predicted for each clause pair.
[0135] Specifically, the classifier is as shown in formula (22):
[0136]
[0137] where, represents the probability that the second clause pair i, j is predicted as a sentiment - cause pair, that is, it represents represents the sentiment - cause pair prediction probability of the second clause pair i, j, represents the clause pair feature vector, b p and W p represent learning parameters.
[0138] Furthermore, the accuracy of the sentiment - cause pair prediction probability is determined according to the loss function for sentiment - cause pair extraction. The loss function for sentiment - cause pair extraction is as shown in formula (23):
[0139]
[0140] where, represents the sentiment - cause pair prediction probability of the second clause pair i, j, y i,j represents the true label of the sentiment - cause pair of the second clause ij, L pair represents the loss function for the sentiment - cause pair extraction task.
[0141] It should be noted that in this embodiment, the loss function of the model is composed of the loss function of the sentiment clause extraction task, the loss function of the reason clause extraction task, and the loss function of the sentiment-reason pair extraction task, that is, the loss function of the sentiment clause extraction task + the loss function of the reason clause extraction task + the loss function of the sentiment-reason pair extraction task = the loss function of the model, which can also be expressed by the following formula (24):
[0142] Loss = L pair + L emo + L cause (24)
[0143] where Loss represents the loss function of the model, and L pair represents the loss function of the sentiment-reason pair extraction task, L emo represents the loss function of the sentiment clause extraction task, and L cause represents the loss function of the reason clause extraction task.
[0144] In the embodiment of the present invention, the cross-entropy function of the sentiment clause extraction task and the true label of the sentiment clause is used as the loss function of the sentiment clause extraction task; the cross-entropy function of the reason clause extraction task and the true label of the reason clause is used as the loss function of the reason clause extraction task; the cross-entropy function of the sentiment-reason pair extraction task and the true label of the sentiment-reason pair is used as the loss function of the sentiment-reason pair extraction task.
[0145] In summary, the embodiments of the present invention provide an emotion - cause pair extraction method, which includes: determining the first clause feature and the first clause feature vector corresponding to the ECPE text according to the word hidden state weight and the hidden state of each word included in the ECPE text extracted by the emotion - causality pair; obtaining the clause adjacency matrix of the graph neural network according to the first clause feature, the graph attention network, and the semantic dependency adjacency matrix; obtaining the second clause feature and the second clause feature vector corresponding to the graph neural network according to the first clause feature vector and the clause adjacency matrix of the graph neural network; obtaining the matching possibility matrix between clause pairs and the emotion - cause pair feature according to the second clause feature vector, and obtaining the clause pair feature and the clause pair feature vector corresponding to the graph neural network according to the emotion - cause pair feature, the graph attention network, and the matching possibility matrix between clause pairs; predicting the emotion - cause pair prediction probability by training a classifier for the clause pair feature vector corresponding to the graph neural network, and improving the accuracy of the emotion - cause pair prediction probability based on the loss function of emotion - cause extraction. Based on the original emotion - cause pair extraction data set, this method transforms the text into a graph structure through text semantic dependency analysis and pre - trained model embedding, and uses the graph attention network to guide by prior knowledge, aggregates node information to obtain a richer text emotion - cause feature representation. Then, through the aggregation of emotion - cause pairs and the interaction mechanism between multiple tasks, the accuracy of emotion - cause pair extraction is effectively improved.
[0146] Based on the same inventive concept, the embodiments of the present invention provide an emotion - cause pair extraction device. Since the principle of this device for solving technical problems is similar to that of the emotion - cause pair extraction method, the implementation of this device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0147] Figure 2 The following is a schematic structural diagram of an emotion - cause pair extraction device provided by the embodiments of the present invention, as Figure 2 shown, the device includes: a first determination unit 201, a first obtaining unit 202, a second obtaining unit 203, and a second determination unit 204.
[0148] The first determination unit 201 is configured to determine the first clause feature and the first clause feature vector corresponding to the ECPE text according to the word hidden state weight and the hidden state of each word included in the ECPE text extracted by the emotion - causality pair;
[0149] The first obtaining unit 202 is configured to obtain the clause adjacency matrix of the graph neural network according to the first clause feature, the graph attention network, and the semantic dependency adjacency matrix; obtain the second clause feature and the second clause feature vector corresponding to the graph neural network according to the first clause feature vector and the clause adjacency matrix of the graph neural network;
[0150] A second obtaining unit 203, configured to obtain a matching possibility matrix between clause pairs and sentiment reason pair features according to the second clause feature vector, and obtain clause pair features and clause pair feature vectors corresponding to the graph neural network according to the sentiment reason pair features, the graph attention network, and the matching possibility matrix between clause pairs;
[0151] A second determining unit 204, configured to predict a sentiment-reason pair prediction probability by training a classifier on the clause pair feature vectors corresponding to the graph neural network, and improve the accuracy of the sentiment-reason pair prediction probability based on a loss function for sentiment-reason extraction.
[0152] Further, the second obtaining unit 203 is specifically configured to:
[0153] Obtain a clause pair attention weight matrix according to the sentiment reason pair features and the graph attention network, and obtain a clause pair adjacency matrix corresponding to the graph neural network according to the clause pair attention weight matrix and the matching possibility matrix between clause pairs;
[0154] According to the clause pair adjacency matrix corresponding to the graph neural network and the sentiment reason pair features, obtain the clause pair features corresponding to the graph neural network through the following formula:
[0155]
[0156] where represents the clause pair ij feature corresponding to the t-th layer of the graph neural network, represents the clause pair ij and xy adjacency matrix of the t-th layer of the graph neural network, W t , b t are learnable parameters, d ij represents the degree of node ij in the clause pair adjacency matrix P t of the t-th layer of the graph neural network.
[0157] Further, the first obtaining unit 202 is specifically configured to:
[0158] The second clause features corresponding to the graph neural network are as follows:
[0159]
[0160] where represents the second clause feature corresponding to the t-th layer of the graph neural network, represents the clause adjacency matrix of the t-th layer of the graph neural network, d i represents the degree of node i in the adjacency matrix A t of the t-th layer of the graph neural network, W t , b t are learnable parameters.
[0161] Further, the second obtaining unit 203 is specifically configured to:
[0162] For the clause pair feature vector corresponding to the graph neural network through the following formula:
[0163]
[0164] where W p and b p represent learning parameters, represents the sentiment - cause pair prediction probability, represents the clause pair feature vector corresponding to the graph neural network.
[0165] Further, the first obtaining unit 202 is specifically configured to:
[0166] The first clause feature is as follows:
[0167]
[0168] The semantic dependency adjacency matrix is as follows:
[0169]
[0170] where A w represents the adjacency matrix, M clause represents the semantic dependency adjacency matrix, h i represents the first clause feature, α i,j represents the word hidden state weight, represents the hidden state of the word.
[0171] Further, the second obtaining unit 203 is specifically configured to:
[0172] Perform a Cartesian product on the prediction labels of a set of sentiment clauses and the prediction labels of a set of reason clauses to obtain a sentiment - reason label product;
[0173] Perform a Cartesian product on the sentiment - reason label product and the sentiment - reason label to obtain a matching possibility matrix between clause pairs;
[0174] Determine the sentiment - cause pair feature through the following formula:
[0175]
[0176] where, represents the feature vector of the second clause i, represents the feature vector of the second clause j, p i,j represents the distance between the second clause i and the second clause j, represents the connection between feature vectors.
[0177] Further, the second obtaining unit 203 is further configured to:
[0178] Input the second clause feature vectors into two prediction layers respectively, and obtain the sentiment clause prediction probability and the reason clause prediction probability respectively through the following formulas:
[0179]
[0180]
[0181] According to the sentiment clause prediction probability and the reason clause prediction probability, obtain the loss function of the sentiment clause extraction task and the loss function of the reason clause extraction task respectively through the following formulas:
[0182]
[0183]
[0184] Wherein, represents the sentiment clause prediction probability of clause i, represents the reason clause prediction probability of clause i, W e and b e 、W c 、b c are learnable parameters, represents the feature vector of the second clause i, represents the sentiment true label of clause i, represents the reason true label of clause i, L emo represents the loss function of the sentiment clause extraction task, L cause represents the loss function of the reason clause extraction task.
[0185] It should be understood that the units included in the above sentiment-reason pair extraction device are only logically divided according to the functions implemented by the device. In actual applications, the above units can be stacked or split. And the functions implemented by the sentiment-reason pair extraction device provided in this embodiment correspond one by one to the sentiment-reason pair extraction method provided in the above embodiment. For the more detailed processing flow implemented by this device, it has been described in detail in the first method embodiment above, and will not be described in detail here.
[0186] Another embodiment of the present invention further provides a computer device, which includes: a processor and a memory; the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device executes each step of the sentiment-reason pair extraction method in the method flow shown in the above method embodiment.
[0187] Another embodiment of the present invention further provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a computer device, the computer device is caused to execute each step of the emotion-cause pair extraction method in the method flow shown in the above method embodiment.
[0188] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0189] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for extracting emotion - cause pairs, characterized in that, Including: Determine the first clause feature and the first clause feature vector corresponding to the ECPE text according to the word hidden state weight and the hidden state of each word included in the ECPE text extracted from the emotional causality pair. Obtain the clause adjacency matrix of the graph neural network according to the first clause feature, the attention mechanism, and the semantic dependency adjacency matrix; obtain the second clause feature and the second clause feature vector corresponding to the graph neural network according to the first clause feature vector and the clause adjacency matrix of the graph neural network. Obtain the matching possibility matrix between clause pairs and the emotional cause pair feature according to the second clause feature vector, and obtain the clause pair feature and the clause pair feature vector corresponding to the graph neural network according to the emotional cause pair feature, the attention mechanism, and the matching possibility matrix between clause pairs. Predict the clause pair feature vector corresponding to the graph neural network through a trained classifier to obtain the emotional-cause pair prediction probability, and improve the accuracy of the emotional-cause pair prediction probability based on the loss function for emotional-cause extraction.
2. The method according to claim 1, characterized in that, The obtaining of the clause pair feature corresponding to the graph neural network according to the emotional cause pair feature, the graph attention network, and the matching possibility matrix between clause pairs specifically includes: Obtain the clause pair attention weight matrix according to the emotional cause pair feature and the graph attention network, and obtain the clause pair adjacency matrix corresponding to the graph neural network according to the clause pair attention weight matrix and the matching possibility matrix between clause pairs. According to the clause pair adjacency matrix corresponding to the graph neural network and the emotional cause pair feature, obtain the clause pair feature corresponding to the graph neural network through the following formula: Among them, represents the clause pair ij feature corresponding to the t-th layer graph neural network, represents the weight corresponding to the clause pairs ij and xy of the t-th layer graph neural network in the adjacency matrix, W t , b t are learnable parameters, d ij represents the degree of node ij in the clause pair adjacency matrix P t of the t-th layer graph neural network.
3. The method according to claim 1, characterized in that, The obtaining of the second clause feature and the second clause feature vector corresponding to the graph neural network according to the first clause feature vector and the clause adjacency matrix of the graph neural network specifically includes: The second clause feature corresponding to the graph neural network is as follows: Among them, represents the second clause feature corresponding to the t-th layer graph neural network, represents the weight corresponding to clauses i and j of the t-th layer graph neural network in the adjacency matrix, W t , b t are learnable parameters, d i represents the adjacency matrix A of the t-th layer graph neural network t is the degree of node i in 4. The method according to claim 1, wherein The predicting of the emotional-cause pair prediction probability by predicting the clause pair feature vector corresponding to the graph neural network through a trained classifier specifically includes: For the clause pair feature vector corresponding to the graph neural network through the following formula: Among them, W p and b p represent learning parameters, represents the prediction probability of the sentiment-cause pair, represents the clause pair feature vector corresponding to the graph neural network.
5. The method according to claim 1, characterized in that The obtaining of the clause adjacency matrix of the graph neural network according to the first clause feature, the attention mechanism, and the semantic dependency adjacency matrix specifically includes: The first clause feature is as follows: The semantic dependency adjacency matrix is as follows: Among them, A w represents the adjacency matrix, M clause represents the semantic dependency adjacency matrix, h i represents the first clause feature, α i,j represents the word hidden state weight, represents the hidden state of the word, and R represents the relationship matrix between all words and all clauses.
6. The method according to claim 1, wherein The obtaining of the matching possibility matrix between clause pairs and the emotional cause pair feature according to the second clause feature vector specifically includes: Perform the Cartesian product of the predicted labels of a group of emotional clauses and the predicted labels of a group of cause clauses to obtain the emotional-cause label product; Perform the Cartesian product of the emotional-cause label product and the emotional-cause label to obtain the matching possibility matrix between clause pairs; Determine the emotional cause pair feature through the following formula: Among them, represents the feature vector of the second clause i, represents the feature vector of the second clause j, p i,j represents the distance between the second clause i and the second clause j, represents the connection between the feature vectors.
7. The method according to claim 1, wherein Before the obtaining of the matching possibility matrix between clause pairs and the emotional cause pair feature according to the second clause feature vector, it further includes: Input the second clause feature vector into two prediction layers respectively, and obtain the emotional clause prediction probability and the cause clause prediction probability respectively through the following formula: According to the emotional clause prediction probability and the cause clause prediction probability, obtain the loss function for the emotional clause extraction task and the loss function for the cause clause extraction task respectively through the following formula: Among them, represents the sentiment clause prediction probability of clause i, represents the reason clause prediction probability of clause i, W e , b e , W c , b c are learnable parameters, represents the feature vector of the second clause i, represents the sentiment true label of clause i, represents the reason true label of clause i, L emo represents the loss function of the sentiment clause extraction task, L cause represents the loss function of the reason clause extraction task.
8. An emotion - cause pair extraction device, characterized in that, Including: A first determination unit, configured to determine a first clause feature and a first clause feature vector corresponding to the ECPE text according to the word hidden state weight and the hidden state of each word included in the ECPE text according to the emotional causality pair extraction; A first obtaining unit, configured to obtain a clause adjacency matrix of the graph neural network according to the first clause feature, the graph attention network, and the semantic dependency adjacency matrix; and obtain a second clause feature and a second clause feature vector corresponding to the graph neural network according to the first clause feature vector and the clause adjacency matrix of the graph neural network; A second obtaining unit, configured to obtain a matching possibility matrix between clause pairs and an emotional cause pair feature according to the second clause feature vector, and obtain a clause pair feature and a clause pair feature vector corresponding to the graph neural network according to the emotional cause pair feature, the graph attention network, and the matching possibility matrix between clause pairs; A second determination unit, configured to predict the clause pair feature vector corresponding to the graph neural network through a trained classifier to obtain an emotion-cause pair prediction probability, and improve the accuracy of the emotion-cause pair prediction probability based on the loss function of the emotion-cause extraction.
9. A computer device, characterized in that, The computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the emotion-cause pair extraction method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored. When the computer program is executed by a processor, the processor executes the emotion-cause pair extraction method according to any one of claims 1-7.
Citation Information
Patent Citations
Emotion cause clause pair extraction method based on semantic decision graph neural network
CN113505583A
Text emotion reason identification method, system and equipment and storage medium
CN115391534A