Aspect sentiment triple extraction method and system based on bidirectional MRC and dual span

By introducing two-way MRC and double-span methods in aspect emotion triple extraction, the interference problem of the model in analyzing multi-faceted sentences and the noise and high computational cost caused by enumeration spans are solved, and a more accurate and efficient triple extraction is achieved.

CN119783659BActive Publication Date: 2025-05-09JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510251904.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-05-09
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Existing aspect emotion triple extraction methods based on machine reading comprehension are susceptible to interference when analyzing sentences containing multiple aspects, making it difficult for the model to obtain the correct correspondence between aspect terms and opinion terms, and generate noise and high computational costs when enumerating spans.

Method used

A method of aspect emotion triple extraction based on bidirectional MRC and double span is proposed. The masked context is encoded through the BERT model, and the part-of-speech graph and syntactic dependency graph are constructed. The attention module and gate mechanism are used to generate the double span candidate set to reduce noise and reduce calculation costs.

Benefits of technology

Through a two-way inference method, a more accurate triplet is generated, noise and calculation costs are reduced, and the model extraction efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783659B_ABST
    Figure CN119783659B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and system for extracting aspect sentiment triples based on bidirectional MRC and dual spans, the method comprising: using shielded context combined with construction rules to construct a part-of-speech graph and a syntactic dependency graph, using a gate mechanism for fusion, and generating a dual span candidate set; performing prediction calculations on the dual span candidate set, selecting the highest prediction score as the opinion item and the aspect item; detecting the aspect item or the opinion item on the semantic representation of the shielded context, and obtaining a detection flag; using multi-head attention to sequentially fuse the representation of the aspect item span, the representation of the opinion item span, and the shielded context to obtain sentiment polarity; using a discriminant model to query the aspect item and the opinion item through a bidirectional reasoning method, and obtaining a triple set of aspect item traversal completed triple set of opinion item traversal completed. The present invention alleviates the dependence of the model effect on the first extracted aspect item through a bidirectional reasoning method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aspect-level sentiment analysis, and in particular to a method and system for extracting aspect sentiment triples based on bidirectional MRC and dual spans. Background Art

[0002] Aspect-based-Sentiment-Analysis (ABSA) is a fine-grained sentiment analysis that can determine people's attitude towards a specific subject by extracting aspect terms and detecting the sentiment polarity of each extracted aspect term. Different sentiment elements are analyzed by introducing different scenario tasks. The task of simultaneously extracting the three sentiment elements of aspect terms, opinion terms and sentiment polarity is called Aspect Sentiment Triple Extraction (ASTE).

[0003] In the past, many researchers have introduced the Machine Reading Comprehension (MRC) framework in the task of extracting aspect-sentiment triples and achieved good results. However, the machine reading comprehension-based method may encounter interference problems when analyzing sentences containing multiple aspects, making it difficult for the model to obtain the correct correspondence between aspect items and opinion items. In order to reduce the interference of irrelevant aspect items, some researchers proposed a mask-based MRC model, which performs context enhancement by enumerating all masked contexts of each aspect item, and then iteratively masks the aspect items through reasoning methods to extract opinion items and sentiment polarity, effectively alleviating the interference problem.

[0004] However, this type of method has two problems: the accuracy of the extracted triples is highly dependent on the accuracy of the first extracted aspect item; and enumerating all spans in a sentence can easily cause a lot of noise and high computational cost. Summary of the invention

[0005] In view of the above situation, the main purpose of the present invention is to propose a method and system for extracting aspect sentiment triples based on bidirectional MRC and dual spans to solve the above technical problems.

[0006] The present invention proposes a method for extracting aspect emotion triples based on bidirectional MRC and dual spans, and the method comprises the following steps:

[0007] Step 1: Use the BERT model to encode the fixed query and the masked context to obtain the semantic representation of the masked context;

[0008] Step 2: Use the semantic representation of the shielded context to construct a graph, obtain a part-of-speech graph and a syntactic dependency graph, characterize the part-of-speech graph and the syntactic dependency graph, and obtain part-of-speech features and syntactic dependencies;

[0009] Step 3: Use the part-of-speech graph attention network and the syntactic dependency graph attention network based on the attention module to learn the part-of-speech graph and the syntactic dependency graph respectively to obtain the part-of-speech relationship features and the syntactic dependency features;

[0010] Step 4: Use the gate mechanism to combine the part-of-speech relationship features and the syntactic dependency features to fuse the part-of-speech features and syntactic features to generate a dual-span candidate set;

[0011] Step 5: perform prediction calculation on the double span candidate set to obtain the calculated prediction score, perform span selection on the calculated prediction score to obtain the candidate span;

[0012] Based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the aspect item;

[0013] Step 6: Based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the opinion item;

[0014] Step 7: Detect aspect items or opinion items of the semantic representation of the shielded context and obtain a detection flag;

[0015] Among them, the detection mark includes the detection mark of aspect item and the detection mark of opinion item;

[0016] Step 8: Use multi-head attention to fuse the representation of aspect item span, the representation of opinion item span, and the masked context in turn to obtain sentiment polarity;

[0017] The discriminant model is used to query aspect items and opinion items through a bidirectional reasoning method to form aspect item triplets and opinion item triplets, and the aspect item triplets and opinion item triplets are added to sets respectively to obtain a set of triples traversed through aspect items and a set of triples traversed through opinion items.

[0018] The present invention also proposes an aspect sentiment triple extraction system based on bidirectional MRC and dual span, the system comprising:

[0019] Input modules for:

[0020] The BERT model is used to encode the fixed query and the masked context to obtain the semantic representation of the masked context.

[0021] The dual-span candidate set generation module is used to:

[0022] The semantic representation of the shielded context is used to construct a graph to obtain a part-of-speech graph and a syntactic dependency graph, and the part-of-speech graph and the syntactic dependency graph are represented to obtain part-of-speech features and syntactic dependencies;

[0023] The part-of-speech graph attention network and syntactic dependency graph attention network based on the attention module are used to learn the part-of-speech graph and syntactic dependency graph respectively to obtain the part-of-speech relationship features and syntactic dependency features;

[0024] The gate mechanism is used to combine the part-of-speech relationship features and the syntactic dependency features to fuse the part-of-speech features and syntactic features to generate a dual-span candidate set;

[0025] Aspect item extraction module, used for:

[0026] Perform prediction calculation on the double span candidate set to obtain the calculated prediction score, perform span selection on the calculated prediction score to obtain the candidate span;

[0027] Based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the aspect item;

[0028] Opinion item extraction module, used for:

[0029] Based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the opinion item;

[0030] Detection module for:

[0031] Detecting aspect items or opinion items on the semantic representation of the shielded context and obtaining a detection flag;

[0032] Among them, the detection mark includes the detection mark of aspect item and the detection mark of opinion item;

[0033] Sentiment classification module, used for:

[0034] Multi-head attention is used to fuse the representation of aspect term span, the representation of opinion term span and the masked context in turn to obtain sentiment polarity.

[0035] The discriminant model is used to query aspect items and opinion items through a bidirectional reasoning method to form aspect item triplets and opinion item triplets, and the aspect item triplets and opinion item triplets are added to sets respectively to obtain a set of triples traversed through aspect items and a set of triples traversed through opinion items.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1. The present invention alleviates the dependence of the model effect on the first extracted aspect item through a bidirectional reasoning method;

[0038] 2. The present invention utilizes the syntactic dependency and part-of-speech correlation between spans to generate a span candidate set that is much smaller than the enumerated span candidate set, thereby reducing the noise and high computational cost caused by the enumerated span;

[0039] 3. The present invention combines the syntactic dependency and part-of-speech features of the sentence with the span representation by obtaining a dual-span candidate set, which helps the model extract triples.

[0040] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flowchart of the steps of the aspect sentiment triple extraction method based on bidirectional MRC and dual span proposed in the present invention.

[0042] Figure 2 Schematic diagram of the syntactic dependency matrix and part-of-speech adjacency matrix in a sentence of the aspect sentiment triple extraction method based on bidirectional MRC and dual spans proposed in the present invention.

[0043] Figure 3 This is a structural diagram of the triplet extraction model and a schematic diagram of the bidirectional reasoning method of the aspect sentiment triplet extraction method based on bidirectional MRC and dual spans proposed in the present invention.

[0044] Figure 4 This is a structural diagram of the aspect sentiment triplet extraction system based on bidirectional MRC and dual spans proposed in the present invention. DETAILED DESCRIPTION

[0045] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout are the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0046] These and other aspects of the embodiments of the present invention will be apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to provide some ways to implement the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0047] See also Figure 1 The embodiment of the present invention proposes a method for extracting aspect emotion triples based on bidirectional MRC and dual spans, and the method comprises the following steps:

[0048] Step 1: Use the BERT model to encode the fixed query and masked context to obtain the semantic representation of the masked context.

[0049] In step 1, a fixed query and a masked context are used as input; the fixed query is used to prompt the model to complete tasks in different directions. In one direction, the leftmost aspect item in the sentence and its corresponding opinion item are first extracted. The first fixed query is expressed as:

[0050] ;

[0051] in, Represents the first fixed query.

[0052] In the other direction, first extract the leftmost opinion item in the sentence and its corresponding aspect item, and the second fixed query is expressed as:

[0053] ;

[0054] in, Represents the second fixed query.

[0055] Two types of operations are performed on each aspect item or opinion item: execute or not execute. Masking is achieved by setting the attention score of the token to 0 and obtaining the mask matrix. The corresponding relationship in the process is as follows:

[0056] ;

[0057] in, Indicates Node The mask matrix of columns, represents the number of columns of the matrix, Indicates the token sequence number.

[0058] The mask matrix is ​​applied to the attention module of the BERT model to obtain the semantic representation of the masked context. The corresponding relationship of the process is as follows:

[0059] ;

[0060] in, represents the attention module, Indicates a query, Indicates the key, Indicates the value, Indicates the key Transpose, represents the transpose operation, represents the normalization operation, Represents the dimension, Represents a mask matrix.

[0061] It should be noted that BERT stands for Transformer's bidirectional encoder.

[0062] Step 2: Use the semantic representation of the masked context to construct a graph to obtain a part-of-speech graph and a syntactic dependency graph, characterize the part-of-speech graph and the syntactic dependency graph, and obtain part-of-speech features and syntactic dependencies.

[0063] See also Figure 2 In step 2, the semantic representation of the shielded context is used to construct a graph to obtain a part-of-speech graph and a syntactic dependency graph, and the part-of-speech graph and the syntactic dependency graph are characterized to obtain part-of-speech features and syntactic dependencies. The corresponding relationship in the process is as follows:

[0064] ;

[0065] ;

[0066] in, Represents the constructed part-of-speech graph, represents the edge of the part-of-speech graph, Indicates the value, represents the constructed syntactic dependency graph, Represents the edges of the syntactic dependency graph.

[0067] Step 3: Use the part-of-speech graph attention network and syntactic dependency graph attention network based on the attention module to learn the part-of-speech graph and syntactic dependency graph respectively to obtain the part-of-speech relationship features and syntactic dependency features.

[0068] In step 3, the part-of-speech graph and the syntactic dependency graph are learned respectively using the part-of-speech graph attention network and the syntactic dependency graph attention network based on the attention module to obtain the part-of-speech relationship features and the syntactic dependency features. The corresponding relationship between the process is as follows:

[0069] ;

[0070] in, Indicates Tier The syntactic features of the nodes, represents the syntactic dependency graph attention network, represents the number of attention heads, and Both represent the node numbers. Representation Node The neighboring set of represents the sigmoid activation function, and Represents two different Tier The normalized attention coefficient of the attention head, , , and Represent four different parameter matrices, Indicates The syntactic features of neighboring nodes, represents the syntactic feature vector, No. Tier The part-of-speech features of the nodes, represents the part-of-speech graph attention network, Indicates The part-of-speech features of neighboring nodes, Represents the part-of-speech feature vector.

[0071] Step 4: Use the gate mechanism to combine the part-of-speech relationship features and the syntactic dependency features to fuse the part-of-speech features and syntactic features to generate a dual-span candidate set.

[0072] In step 4, the part-of-speech features and syntactic features are fused by combining the part-of-speech relationship features and the syntactic dependency features using a gate mechanism to generate a dual-span candidate set, which specifically includes the following steps:

[0073] The gate mechanism is used to fuse the node's part-of-speech relationship features and the node's syntactic relationship features to obtain the fused features;

[0074] Among them, the relationship between the fusion features and the corresponding process is as follows:

[0075] ;

[0076] in, represents the intermediate variable, represents the first trainable weight, represents the first trainable bias, Representing syntactic dependency features and part-of-speech relationship features The concatenation of represents the element-wise product operation, Indicates fusion features;

[0077] The semantic representation of the shielded context is given, and the part-of-speech tag is determined in combination with the part-of-speech relationship features of the node to obtain the part-of-speech tag determination result, and the part-of-speech tag determination result is combined with the fusion feature to form the word span;

[0078] Among them, the relationship between the word span and the corresponding process is as follows:

[0079] ;

[0080] in, represents the word span, Indicates the starting position of the word span, Indicates the end position of the word span. Representing words and words Trainable cross-length embeddings, Representing words Part of speech, Representing words , Indicates a noun, It indicates an adjective;

[0081] The semantic representation of the shielding context is given, and the dependency edge is determined in combination with the syntactic relationship features of the nodes to obtain the dependency edge determination result, and the dependency edge determination result is combined with the fusion features to form the syntactic span;

[0082] Among them, the relationship between the formation of syntactic span and the corresponding existence of the process is as follows:

[0083] ;

[0084] in, represents the syntactic span, Representing words and words The syntactic span between;

[0085] The word span and syntactic span are combined with the fusion features to generate a dual span candidate set;

[0086] Among them, the relationship between the generation of double span candidate sets and the corresponding process is as follows:

[0087] ;

[0088] in, Represents a double span candidate set.

[0089] Step 5: perform prediction calculation on the double span candidate set to obtain the calculated prediction score, perform span selection on the calculated prediction score to obtain the candidate span;

[0090] Based on the calculated prediction scores, the span that obtains the highest prediction score is selected from the candidate spans as the aspect item.

[0091] In step 5, based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the aspect item. The corresponding relationship in the process is as follows:

[0092] ;

[0093] in, represents the prediction score of the span aspect item, represents a ReLU activated feedforward neural network, Indicates that the truth value is an aspect item, Representing words and words The span between

[0094] Step 6: Based on the calculated prediction scores, select the span with the highest prediction score from the candidate spans as the opinion item.

[0095] In step 6, based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the opinion item. The corresponding relationship of the process is as follows:

[0096] ;

[0097] in, represents the prediction score of the opinion item with span, Indicates that the truth value is an opinion item;

[0098] Step 7: Detect aspect items or opinion items of the semantic representation of the shielded context and obtain a detection flag;

[0099] Among them, the detection flags include aspect item detection flags and opinion item detection flags.

[0100] In step 7, the semantic representation of the shielded context is detected for aspect items or opinion items to obtain a detection flag. The corresponding relationship between the process and the existing one is as follows:

[0101] ;

[0102] in, represents the intermediate vector, represents the maximum pooling operation, Indicates the marking symbol The expression, represents the probability of labels being true and false, The aspect item indicates that, The opinion item indicates that represents the second trainable weight, represents the second trainable bias, Indicates a serial operation;

[0103] Step 8: In step 8, the representation of the aspect item span, the representation of the opinion item span and the masked context are sequentially fused using multi-head attention to obtain the sentiment polarity. The corresponding relationship of the process is as follows:

[0104] ;

[0105] in, and represents two different intermediate variables, represents the masked context representation, Presentation layer specification processing, represents multi-head attention processing, represents the probability of sentiment polarity label, represents the third trainable weight, represents the third trainable bias;

[0106] See also Figure 3 ,The bidirectional reasoning method includes two consecutive stages, the first consecutive stage is the ,aspect reasoning stage and the aspect attachment reasoning stage, and the second consecutive stage is the ,opinion reasoning stage and the opinion attachment reasoning stage;

[0107] The discriminant model is used to query the aspect items and opinion items through a bidirectional reasoning method to form aspect item triplets and opinion item triplets, and the aspect item triplets and opinion item triplets are added to the sets respectively to obtain the triple set of aspect item traversal and the triple set of opinion item traversal, which specifically includes the following sub-steps:

[0108] In the aspect reasoning stage, the discriminant model is used to obtain the aspect item detection flag of the semantic representation of the first fixed query and the masked context and the first aspect item, and the aspect item detection flag is discriminated to obtain the aspect item detection flag result;

[0109] When the result of the aspect item detection flag is yes, the first aspect item is added to the aspect set, and the first aspect item of the semantic representation of the context is masked to obtain the context masked by the aspect item;

[0110] The discriminant model is used to obtain the next aspect item of the context masked by the aspect item in a cyclic manner, and the discriminant model is used again to discriminate the aspect item detection flag;

[0111] When the result of the aspect item detection flag is negative, a discriminant model of aspect item training of the aspect item set of aspect items added cyclically is obtained;

[0112] In the aspect attachment reasoning stage, the context masked by the aspect item and the first fixed query are input into the discriminant model trained on the aspect item to obtain information, and the sentiment polarity corresponding to the aspect item of the opinion item set of the aspect item is obtained. Based on the aspect item set and the sentiment polarity corresponding to the aspect item of the opinion item set, a combination construction is performed to obtain the triples of the aspect item, and the triples of the aspect item are added to the set in a cyclic manner to obtain the triple set of the aspect item traversal completed;

[0113] In the opinion reasoning stage, the discriminant model trained by aspect items is used to obtain the opinion item detection flags of the semantic representation of the second fixed query and the shielded context and the first opinion item, and the opinion item detection flags are discriminated to obtain the opinion item detection flag results;

[0114] When the result of the opinion item detection flag is yes, the first opinion item is added to the opinion item set, and the first opinion item of the semantic representation of the context is masked to obtain the context masked by the opinion item;

[0115] The discriminant model trained by the aspect item is used to obtain the next opinion item of the context shielded by the opinion item in a cyclic manner, and the discriminant model trained by the aspect item is used again to discriminate the opinion item detection flag;

[0116] When the result of the opinion item detection flag is negative, a discriminant model for opinion item training of the opinion item set of opinion items added cyclically is obtained;

[0117] In the opinion attachment reasoning stage, the masked context of the opinion item and the second fixed query are input into the discriminant model trained on the opinion item to obtain information, and the aspect item set and the corresponding sentiment polarity of the opinion item are obtained. Based on the aspect item set of the opinion item and the sentiment polarity of the opinion item in the opinion item set, a combination construction is performed to obtain the triples of the opinion item, and the triples of the opinion item are set-added in a cyclic manner to obtain the triple set of the aspect item traversal.

[0118] It should be noted that in Figure 2 and Figure 3 In the above figure, q represents query, x represents masked context, MP represents max pooling, e represents aspect item detection flag or opinion item detection flag, C represents concatenation, s represents sentiment polarity, a represents aspect item, and o represents opinion item.

[0119] Furthermore, the bidirectional reasoning method includes two directions, each of which includes two consecutive stages: aspect reasoning stage and aspect attachment reasoning stage, opinion reasoning stage and opinion attachment reasoning stage. In one direction, all aspect items are queried first, and then the opinion items corresponding to each aspect are queried; in the other direction, all opinion items are queried first, and then their corresponding aspect items are queried.

[0120] In the aspect reasoning stage, the trained model is first used to obtain the aspect detection flag and the first aspect item a of the fixed query q1 and the masked context x; if the aspect item detection flag is yes, the aspect item a is added to the aspect set, and then the aspect item a in the context is masked; with the masked context, the trained model is used again to obtain the next aspect item detection flag and the next aspect item; the above steps are repeated until the aspect item detection flag is no, and finally the aspect item set A is obtained. In the aspect attachment reasoning stage, the sentence with all aspect items except aspect item a masked and the fixed query q1 are sent to the trained model to obtain the opinion item set O and the corresponding sentiment polarity s of aspect item a; triples are obtained based on the sets A, O and sentiment polarity s, and the obtained triples are added to the triple set T; the above steps are repeated until all aspect items in the aspect set A are traversed.

[0121] In the opinion reasoning stage, the trained model is used to obtain the opinion item detection flag and the first opinion item of the fixed query q2 and the shielded context x. If the opinion item detection flag is yes, the opinion item o is added to the opinion set, and then the opinion item o in x is shielded; the trained model is repeatedly used to obtain the next opinion item detection flag and the next opinion until the opinion item detection flag is no, and finally the opinion item set O is obtained. In the opinion attachment reasoning stage, the sentence with all opinions except opinion item o shielded and the fixed query q2 are sent to the trained model to obtain the aspect item set A and the corresponding sentiment polarity s of the opinion item o; based on the set A, O and sentiment polarity s, a triple is obtained, and the obtained triple is added to the triple set T; the above steps are repeated until all opinion items in the opinion set O are traversed.

[0122] In executing the above steps 5 to 8, the corresponding training method includes the following training steps:

[0123] Based on the negative log-likelihood loss function, the highest score predicted in the dual-span candidate set is used as the logarithmic probability of the representation prediction of the aspect item to generate the aspect item span to construct the loss function of the aspect item span. The corresponding relationship in the process of constructing the loss function of the aspect item span is as follows:

[0124] ;

[0125] in, represents the loss of aspect item span, represents the logarithmic function, represents the conditional probability, A representation that represents the span of an aspect item.

[0126] Based on the negative log-likelihood loss function, the highest score predicted in the dual-span candidate set is used as the logarithmic probability of the opinion item to generate the representation of the opinion item span to construct the loss function of the aspect item span. The corresponding relationship in the process of constructing the loss function of the aspect item span is as follows:

[0127] ;

[0128] in, represents the loss of the opinion term span, A representation representing the span of an opinion item.

[0129] Based on the negative log-likelihood loss function, the detection loss is constructed by using the probabilities of true and false labels to generate the logarithmic probability of the label truth prediction. The corresponding relationship in the process of constructing the detection loss is as follows:

[0130] ;

[0131] in, represents the detection loss, represents the true value of the label, The probability of the label being true or false Node elements.

[0132] Based on the negative log-likelihood loss function, the sentiment polarity loss function is constructed by using the sentiment polarity label probability to generate the logarithmic probability of sentiment polarity label prediction. The process of constructing the sentiment polarity loss function corresponds to the following relationship:

[0133] ;

[0134] in, represents the sentiment classification loss, represents the sentiment polarity label, The probability of sentiment polarity label Node elements.

[0135] The total loss of the discriminant model is constructed by using the loss function in the direction of extracting the aspect item first and then the opinion item, and the loss function in the direction of extracting the opinion item first and then the aspect item. The extraction error of the discriminant model is reduced based on the total loss of the discriminant model.

[0136] Among them, the relationship between the total loss of the construction of the discriminant model and the existence of the following equation is:

[0137] ;

[0138] in, represents the loss function in the direction of extracting the aspect item first and then the opinion item, represents the loss function in the direction of extracting the opinion item first and then the aspect item, , , , , , , and Both represent hyperparameters used to adjust the corresponding loss effects. Represents the total loss of the discriminative model.

[0139] See also Figure 4 The present invention also provides an aspect sentiment triple extraction system based on bidirectional MRC and dual span, the system comprising:

[0140] Input modules for:

[0141] The BERT model is used to encode the fixed query and the masked context to obtain the semantic representation of the masked context.

[0142] The dual-span candidate set generation module is used to:

[0143] The semantic representation of the shielded context is used to construct a graph to obtain a part-of-speech graph and a syntactic dependency graph, and the part-of-speech graph and the syntactic dependency graph are represented to obtain part-of-speech features and syntactic dependencies;

[0144] The part-of-speech graph attention network and syntactic dependency graph attention network based on the attention module are used to learn the part-of-speech graph and syntactic dependency graph respectively to obtain the part-of-speech relationship features and syntactic dependency features;

[0145] The gate mechanism is used to combine the part-of-speech relationship features and the syntactic dependency features to fuse the part-of-speech features and syntactic features to generate a dual-span candidate set;

[0146] Aspect item extraction module, used for:

[0147] Perform prediction calculation on the double span candidate set to obtain the calculated prediction score, perform span selection on the calculated prediction score to obtain the candidate span;

[0148] Based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the aspect item;

[0149] Opinion item extraction module, used for:

[0150] Based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the opinion item;

[0151] Detection module for:

[0152] Detecting aspect items or opinion items on the semantic representation of the shielded context and obtaining a detection flag;

[0153] Among them, the detection mark includes the detection mark of aspect item and the detection mark of opinion item;

[0154] Sentiment classification module, used for:

[0155] Multi-head attention is used to fuse the representation of aspect term span, the representation of opinion term span and the masked context in turn to obtain sentiment polarity.

[0156] The discriminant model is used to query aspect items and opinion items through a bidirectional reasoning method to form aspect item triplets and opinion item triplets, and the aspect item triplets and opinion item triplets are added to sets respectively to obtain a set of triples traversed through aspect items and a set of triples traversed through opinion items.

[0157] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0158] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0159] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A method for extracting aspect sentiment triples based on bidirectional MRC and dual span, characterized in that: The method comprises the following steps: Step 1: Use the BERT model to encode the fixed query and the masked context to obtain the semantic representation of the masked context; Step 2: Use the semantic representation of the shielded context to construct a graph, obtain a part-of-speech graph and a syntactic dependency graph, characterize the part-of-speech graph and the syntactic dependency graph, and obtain part-of-speech features and syntactic dependencies; Step 3: Use the part-of-speech graph attention network and the syntactic dependency graph attention network based on the attention module to learn the part-of-speech graph and the syntactic dependency graph respectively to obtain the part-of-speech relationship features and the syntactic dependency features; Step 4: Use the gate mechanism to combine the part-of-speech relationship features and the syntactic dependency features to fuse the part-of-speech features and syntactic features to generate a dual-span candidate set; Step 5: perform prediction calculation on the double span candidate set to obtain the calculated prediction score, perform span selection on the calculated prediction score to obtain the candidate span; Based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the aspect item; Step 6: Based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the opinion item; Step 7: Detect aspect items or opinion items of the semantic representation of the shielded context and obtain a detection flag; Among them, the detection mark includes the detection mark of aspect item and the detection mark of opinion item; Step 8: Use multi-head attention to fuse the representation of aspect item span, the representation of opinion item span, and the masked context in turn to obtain sentiment polarity; The discriminant model is used to query aspect items and opinion items through a bidirectional reasoning method to form aspect item triplets and opinion item triplets, and the aspect item triplets and opinion item triplets are added to sets respectively to obtain a set of triples traversed through aspect items and a set of triples traversed through opinion items.

2. The aspect sentiment triple extraction method based on bidirectional MRC and dual span according to claim 1, characterized in that: In step 2, the semantic representation of the shielded context is used to construct a graph to obtain a part-of-speech graph and a syntactic dependency graph, and the part-of-speech graph and the syntactic dependency graph are characterized to obtain part-of-speech features and syntactic dependencies. The corresponding relationship in the process is as follows: ; ; in, Represents the constructed part-of-speech graph, represents the edge of the part-of-speech graph, Indicates the value, represents the constructed syntactic dependency graph, Represents the edges of the syntactic dependency graph.

3. The aspect sentiment triple extraction method based on bidirectional MRC and dual span according to claim 2 is characterized in that: In step 3, the part-of-speech graph and the syntactic dependency graph are learned respectively using the part-of-speech graph attention network and the syntactic dependency graph attention network based on the attention module to obtain the part-of-speech relationship features and the syntactic dependency features. The corresponding relationship between the process is as follows: ; in, Indicates Tier The syntactic features of the nodes, represents the syntactic dependency graph attention network, represents the number of attention heads, and Both represent the node numbers. Representation Node The neighboring set of represents the sigmoid activation function, and Represents two different Tier The normalized attention coefficient of the attention head, , , and Represent four different parameter matrices, Indicates The syntactic features of neighboring nodes, represents the syntactic feature vector, No. Tier The part-of-speech features of nodes, represents the part-of-speech graph attention network, Indicates The part-of-speech features of neighboring nodes, Represents the part-of-speech feature vector.

4. The aspect sentiment triple extraction method based on bidirectional MRC and dual span according to claim 3 is characterized in that: In step 4, the gate mechanism is used to combine the part-of-speech relationship feature and the syntactic dependency feature to fuse the part-of-speech feature and the syntactic feature to generate a dual-span candidate set, which specifically includes the following steps: The gate mechanism is used to fuse the node's part-of-speech relationship features and the node's syntactic relationship features to obtain the fused features; Among them, the relationship between the fusion features and the corresponding process is as follows: ; in, represents the intermediate variable, represents the first trainable weight, represents the first trainable bias, Representing syntactic dependency features and part-of-speech relationship features The concatenation of represents the element-wise product operation, Indicates fusion features; The semantic representation of the shielded context is given, and the part-of-speech tag is determined in combination with the part-of-speech relationship features of the node to obtain the part-of-speech tag determination result, and the part-of-speech tag determination result is combined with the fusion feature to form the word span; Among them, the relationship between the word span and the process is as follows: ; in, represents the word span, Indicates the starting position of the word span, Indicates the end position of the word span. Representing words and words Trainable cross-length embeddings, Representing words Part of speech, Representing words , Indicates a noun, It indicates an adjective; The semantic representation of the shielding context is given, and the dependency edge is determined in combination with the syntactic relationship features of the nodes to obtain the dependency edge determination result, and the dependency edge determination result is combined with the fusion features to form the syntactic span; Among them, the relationship between the formation of syntactic span and the corresponding existence of the process is as follows: ; in, represents the syntactic span, Representing words and words The syntactic span between; The word span and syntactic span are combined with the fusion features to generate a dual span candidate set; Among them, the relationship between the generation of double span candidate sets and the corresponding process is as follows: ; in, Represents a double span candidate set.

5. The aspect sentiment triple extraction method based on bidirectional MRC and dual span according to claim 4 is characterized in that: In step 5, based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the aspect item. The corresponding relationship in the process is as follows: ; in, represents the prediction score of the span aspect item, represents a ReLU activated feedforward neural network, Indicates that the truth value is an aspect item, Representing words and words The span between.

6. The aspect sentiment triple extraction method based on bidirectional MRC and dual span according to claim 5, characterized in that: In step 6, based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the opinion item. The corresponding relationship of the process is as follows: ; in, represents the predicted score of the opinion item with span, Indicates that the truth value is an opinion item.

7. The aspect emotion triple extraction method based on bidirectional MRC and dual span according to claim 6 is characterized in that: In step 7, the semantic representation of the shielded context is detected for aspect items or opinion items to obtain a detection flag. The corresponding relationship between the process and the existing one is as follows: ; in, represents the intermediate vector, represents the maximum pooling operation, Indicates the marking symbol The expression, represents the probability of labels being true and false, The aspect item indicates that, The opinion item indicates that represents the second trainable weight, represents the second trainable bias, Represents a concatenation operation.

8. The aspect sentiment triple extraction method based on bidirectional MRC and dual span according to claim 7, characterized in that: In step 8, the representation of the aspect item span, the representation of the opinion item span and the masked context are sequentially fused using multi-head attention to obtain the sentiment polarity. The corresponding relationship of the process is as follows: ; in, and represents two different intermediate variables, represents the masked context representation, Presentation layer specification processing, represents multi-head attention processing, represents the probability of sentiment polarity label, represents the third trainable weight, Represents the third trainable bias.

9. The aspect sentiment triple extraction method based on bidirectional MRC and dual span according to claim 8, characterized in that: The bidirectional reasoning method includes two consecutive stages, the first consecutive stage is the aspect reasoning stage and the aspect attachment reasoning stage, and the second consecutive stage is the opinion reasoning stage and the opinion attachment reasoning stage; The discriminant model is used to query the aspect items and opinion items through a bidirectional reasoning method to form aspect item triplets and opinion item triplets, and the aspect item triplets and opinion item triplets are added to the sets respectively to obtain the triple set of aspect item traversal and the triple set of opinion item traversal, which specifically includes the following sub-steps: In the aspect reasoning stage, the discriminant model is used to obtain the aspect item detection flag of the semantic representation of the first fixed query and the masked context and the first aspect item, and the aspect item detection flag is discriminated to obtain the aspect item detection flag result; When the result of the aspect item detection flag is yes, the first aspect item is added to the aspect set, and the first aspect item of the semantic representation of the context is masked to obtain the context masked by the aspect item; The discriminant model is used to obtain the next aspect item of the context masked by the aspect item in a cyclic manner, and the discriminant model is used again to discriminate the aspect item detection flag; When the result of the aspect item detection flag is negative, a discriminant model of aspect item training of the aspect item set of aspect items added cyclically is obtained; In the aspect attachment reasoning stage, the context masked by the aspect item and the first fixed query are input into the discriminant model trained on the aspect item to obtain information, and the sentiment polarity corresponding to the aspect item of the opinion item set of the aspect item is obtained. Based on the aspect item set and the sentiment polarity corresponding to the aspect item of the opinion item set, a combination construction is performed to obtain the triples of the aspect item, and the triples of the aspect item are added to the set in a cyclic manner to obtain the triple set of the aspect item traversal completed; In the opinion reasoning stage, the discriminant model trained by aspect items is used to obtain the opinion item detection flags of the semantic representation of the second fixed query and the shielded context and the first opinion item, and the opinion item detection flags are discriminated to obtain the opinion item detection flag results; When the result of the opinion item detection flag is yes, the first opinion item is added to the opinion item set, and the first opinion item of the semantic representation of the context is masked to obtain the context masked by the opinion item; The discriminant model trained by the aspect item is used to obtain the next opinion item of the context shielded by the opinion item in a cyclic manner, and the discriminant model trained by the aspect item is used again to discriminate the opinion item detection flag; When the result of the opinion item detection flag is negative, a discriminant model for opinion item training of the opinion item set of opinion items added cyclically is obtained; In the opinion attachment reasoning stage, the masked context of the opinion item and the second fixed query are input into the discriminant model trained on the opinion item to obtain information, and the aspect item set and the corresponding sentiment polarity of the opinion item are obtained. Based on the aspect item set of the opinion item and the sentiment polarity of the opinion item in the opinion item set, a combination construction is performed to obtain the triples of the opinion item, and the triples of the opinion item are set-added in a cyclic manner to obtain the set of triples of the opinion item traversal completed.

10. A system for extracting aspect sentiment triples based on bidirectional MRC and dual span, characterized in that: The system applies any one of the above-mentioned methods for extracting aspect emotion triples based on bidirectional MRC and dual spans as claimed in claims 1 to 9, and the system comprises: Input modules for: The BERT model is used to encode the fixed query and the masked context to obtain the semantic representation of the masked context. The dual-span candidate set generation module is used to: The semantic representation of the shielded context is used to construct a graph to obtain a part-of-speech graph and a syntactic dependency graph, and the part-of-speech graph and the syntactic dependency graph are represented to obtain part-of-speech features and syntactic dependencies; The part-of-speech graph attention network and syntactic dependency graph attention network based on the attention module are used to learn the part-of-speech graph and syntactic dependency graph respectively to obtain the part-of-speech relationship features and syntactic dependency features; The gate mechanism is used to combine the part-of-speech relationship features and the syntactic dependency features to fuse the part-of-speech features and syntactic features to generate a dual-span candidate set; Aspect item extraction module, used for: Perform prediction calculation on the double span candidate set to obtain the calculated prediction score, perform span selection on the calculated prediction score to obtain the candidate span; Based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the aspect item; Opinion item extraction module, used for: Based on the calculated prediction scores, a span with the highest prediction score is selected from the candidate spans as the opinion item; Detection module for: Detecting aspect items or opinion items on the semantic representation of the shielded context and obtaining a detection flag; Among them, the detection mark includes the detection mark of aspect item and the detection mark of opinion item; Sentiment classification module, used for: Multi-head attention is used to fuse the representation of aspect term span, the representation of opinion term span and the masked context in turn to obtain sentiment polarity. The discriminant model is used to query aspect items and opinion items through a bidirectional reasoning method to form aspect item triplets and opinion item triplets, and the aspect item triplets and opinion item triplets are added to sets respectively to obtain a set of triples traversed through aspect items and a set of triples traversed through opinion items.

Citation Information

Patent Citations

  • Emotion triple extraction method based on span sharing and grammar dependency relationship enhancement

    CN113743097A

  • Knowledge enhancement-based aspect emotion triple extraction method and system

    CN117171610A