A method and system for event extraction based on graph parsing
By constructing event graphs based on graph analysis and using Transformer generative model for event extraction, the problem of relying on trigger words and difficulty in modeling multi-event correlation in the existing technology is solved, and efficient event extraction and accurate argument sharing modeling is achieved.
Patent Information
- Application Number
- CN202210845805.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-19
AI Technical Summary
The existing event extraction method relies on trigger words, making it difficult to effectively model the correlation and argument sharing phenomenon between multiple events, and fails to fully utilize the tag semantic information of event types and argument roles.
Using a graph-based analysis method, multiple events in the input sentence are treated as a whole, an event graph is constructed, and linearized through the Seq2Seq sequence-sequence framework. Using a generative model and decoding algorithm based on Transformer, combined with the pre-trained language model BART, event extraction is realized.
Without relying on trigger words, effectively model the correlation and argument sharing phenomenon between multiple events, improving the accuracy and performance of event extraction and alleviating the problem of data sparseness.
Smart Images

Figure CN115169285B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information extraction in natural language processing, and in particular to an event extraction method based on graph parsing. Background Art
[0002] In practical applications, event extraction has been widely used in fields such as automatic question answering, information retrieval, human-computer interface, trend analysis, and financial social networking. Nowadays, there are more and more studies on the intelligent application of event extraction technology. For example, the writing robot developed by the Baidu Knowledge Graph team automatically generates some articles about major events based on the event graph. Event extraction is widely used in intelligence work in fields such as business and military, and also has many practical applications in life and social networking. We can quickly extract important events from news and obtain key information more efficiently. Regarding the mentality and state of people facing the event, as well as the responses of the government, we can monitor social public opinion through event extraction technology, which is of great significance for us to understand the domestic and international epidemic situations, as well as the scientific prevention and control and emergency management of the government.
[0003] Generally speaking, the event extraction task can be divided into two subtasks, namely event detection (identifying the specified event type) and argument extraction (identifying the arguments of each event type and labeling their roles).
[0004] Existing research work usually transforms event extraction into two subtasks: trigger word classification task and argument classification task, and solves these two tasks through classification methods. However, such methods have some inherent drawbacks. First, existing methods rely heavily on trigger words. On the one hand, trigger words are not necessary for event detection and event extraction tasks. On the other hand, trigger words restrict the accurate identification of events by the model to a certain extent. Secondly, the current model cannot explicitly model the correlation between multiple events and multiple arguments playing different roles in a sentence, nor can it well solve the problem of argument sharing. Finally, existing methods cannot utilize the label semantic information of event types and argument roles. Summary of the Invention
[0005] Object of the Invention: The object of the present invention is to provide an event extraction method based on graph parsing, which, without relying on trigger words, solves the long-tail phenomenon, and at the same time constructs a Transformer-based model for the correlation between multiple events to solve the argument sharing phenomenon; linearizes the sentences of the given input text and uses the sequence-to-sequence conversion method to solve the data sparsity problem.
[0006] Technical Solution: An event extraction method based on graph parsing provided by the present invention includes the following steps:
[0007] 1) Given the input text, determine whether the sentences in the given input text contain event types. Extract any node in the sentences of the given input text as the root node, and append the event type node as a child node after the root node. If the event type is included, the value of this root node is EVTS, and connect the events contained in the sentences of the given input text to form an event graph. If the event type is not included, the value of this root node is NA, and the event extraction ends;
[0008] 2) Based on the Seq2Seq sequence-to-sequence framework, linearize the event graph to obtain the linearized sequence of the event graph;
[0009] 3) Based on the linearized sequence of the event graph, design a Transformer-based generation model and decoding algorithm. The specific method of this generation model and decoding algorithm is as follows:
[0010] Set \(x = \langle x 1 ,\ldots,x n \rangle\) as the sentence of the given input text, where \(x i \) represents the \(i\)-th word in the sentence, \(i = 1,2,\ldots,n\). At the same time, set \(E = \langle e 1 ,\ldots,e j ,\ldots,e n \rangle\) as the entity mentions in the sentence, where \(e i \) represents the \(i\)-th entity mention in the sentence, \(i = 1,2,\ldots,j,\ldots,n\). The entity mentions in total contain \(k\) entity mentions, \(k\) is a natural number, and each entity mention contains a head entity and an entity type. Build a Transformer-based generation model in this way;
[0011] The Transformer-based generation model needs to decode the token list \(y=\langle y 1 ,\ldots,y m \rangle\) in sequence, where \(y i \) is the \(i\)-th token in the token list, \(i = 1,2,\ldots,m\). The value of the token \(y i \) is any one of the event type, event argument (i.e., entity mention), argument role, entity type, and special pointer symbol;
[0012] 4) Use the pre-trained language model BART to convert the Transformer-based generation model into an Encoder-Decoder architecture to complete the learning of event extraction results.
[0013] Furthermore, in step 2), two linearization methods, depth-first traversal and breadth-first traversal, are used for linearization processing, and entity type nodes are added at the same time. Design to swap the traversal order of nodes and edges.
[0014] Further, in step 3), when the generation model generates argument nodes for specific event type nodes, the generation model outputs the head entity of each argument as the argument. Let Y be the space of the output solution, and the goal of the generation model becomes finding the node sequence of the sentence x of the given input text:
[0015]
[0016] Then, an encoder based on Transformer is used for encoding, and the following decoding algorithm is used for decoding:
[0017]
[0018] P(y j |x,y <j )=softmax(g(s j ))
[0019] where p and P are both probabilities, y <j is the value of each node before the j-th position, h i is the context hidden vector, Encoder is the encoder of Transformer, s j is the m decoding symbols sequentially generated by the decoder, Decoder is the decoder, and y j is the j-th token.
[0020] Further, after the decoding algorithm is completed, based on the event extraction method of graph parsing, the dependency syntactic information is used to further extract the result after decoding the generation model based on Transformer; this dependency syntactic information encodes the dependency graph using a graph attention neural network, and fuses the dual attention mechanisms of the dependency graph encoding layer and the sentence encoding layer, specifically including:
[0021] Obtain the context hidden vectors {h i ,h i+1 ,...,h i+m-1} and the syntactic dependency graph hidden vectors {h′ i ,h′ i+1 ,...,h′ i+m-1}, apply the average pooling function to obtain the context representation h con and the dependency graph representation h syn ;
[0022] h con =pool(h i ,h i+1 ,...,h i+m-1 )
[0023] h syn =pool(hi ′, h i ′ +1 ,..., h i ′ +m-1 )
[0024] Among them, a gating mechanism is adopted to fuse the features of the context representation h con and the dependency graph representation h syn The dependency graph representation h syn is fused into the context representation h con as shown in the following formula:
[0025]
[0026] Where is the product operation, and the specific calculation method of the function g is shown in the following formula:
[0027] g = σ(W g [h syn ; h con + b g )
[0028] Where, [h syn ; h con is the concatenation of h con and h syn , and W g and b g are the parameters of the generation model.
[0029] The present invention correspondingly provides an event extraction system based on graph parsing, including a judgment module, a linearization module, a generation model and a decoding module, and a transfer learning module;
[0030] The judgment module is used to give an input text, judge whether the sentence of the given input text contains an event type, extract any node in the sentence of the given input text as the root node, attach the event type node after the root node as a child node. If the event type is included, the value of the root node is EVTS, and the events included in the sentence of the given input text are connected to form an event graph. If the event type is not included, the value of the root node is NA, and the extraction of the event ends;
[0031] The linearization module is used to linearly process the event graph based on the Seq2Seq sequence-to-sequence framework to obtain a linearized sequence of the event graph;
[0032] The generation model and the decoding module are used to design a generation model and a decoding algorithm based on Transformer based on the linearized sequence of the event graph. The specific method of the generation model and the decoding algorithm is as follows:
[0033] Set x = <x 1,...,x n >For the sentence of the given input text, where x i represents the i-th word in the sentence, i = 1, 2... n. At the same time, set E = <e 1 ,..e j ..,e n >is the entity mention in the sentence, where e i represents the i-th entity mention in the sentence, i = 1, 2..j..n. The entity mention contains a total of k entity mentions, k is a natural number, and each entity mention contains a head entity and an entity type, so as to build a Transformer-based generation model;
[0034] The Transformer-based generation model needs to decode the token list y = <y 1 ,...,y m >in sequence, where yi is the i-th token in the token list, i = 1, 2... m. The value of the token y i is any one of the event type, event argument (i.e., entity mention), argument role, entity type, and special pointer symbol;
[0035] The transfer learning module is used to convert the Transformer-based generation model into an Encoder-Decoder architecture by using the pre-trained language model BART to complete the learning of event extraction results.
[0036] Furthermore, in the linearization module, the linearization process adopts two linearization methods: depth-first traversal and breadth-first traversal. At the same time, entity type nodes are added, and the traversal order of nodes and edges is designed to be swapped.
[0037] Furthermore, in the generation model and decoding module, when the generation model generates an argument node for a specific event type node, the generation model outputs the head entity of each argument as the argument. Set Y as the space of the output solution, and the goal of the generation model becomes to find the node sequence of the sentence x of the given input text:
[0038]
[0039] Then, the Transformer-based encoder is used for encoding, and the following decoder is used for decoding:
[0040]
[0041] P(y j |x,y <j ) = softmax(g(s j ))
[0042] where p and P are both probabilities, and y <j is the values of each node before the j-th position, h i is the context hidden vector, Encoder is the encoder of the Transformer, s j is the m decoded symbols generated sequentially by the decoder, Decoder is the decoder, and y j is the j-th token.
[0043] Furthermore, after the decoding algorithm is completed, based on the event extraction method of graph parsing, the dependency syntactic information is used to further extract the results after decoding the generation model based on the Transformer; the dependency syntactic information encodes the dependency graph using a graph attention neural network, and fuses the dual attention mechanisms of the dependency graph encoding layer and the sentence encoding layer, specifically including:
[0044] Obtain the context hidden vectors {h i , h i+1 ,..., h i+m-1} and the syntactic dependency graph hidden vectors {h′ i , h′ i+1 ,..., h′ i+m-1}, and apply the average pooling function to obtain the context representation h con and the dependency graph representation h syn ;
[0045] h con = pool(h i , h i+1 ,..., h i+m-1 )
[0046] h syn = pool(h i ′, h i ′ +1 ,..., h i ′ +m-1 )
[0047] where pool is the pooling layer, and the gating mechanism is used to fuse the features of the context representation h con and the dependency graph representation h syn , and the dependency graph representation h syn is fused into the context representation h con as shown in the following formula:
[0048]
[0049] where is the product operation, and the specific calculation method of the function g is shown in the following formula:
[0050] g = σ(Wg [h syn ; h con +b g )
[0051] Among them, [h syn ; h con is the concatenation of h con and h syn , and W g and b g are the parameters of the generation model.
[0052] Beneficial effects: Compared with the prior art, the remarkable feature of the present invention is that multiple events included in the input sentence are regarded as a whole, multiple events are linked to form an event graph, and the method of extracting events from the input sentence is transformed into a graph parsing method for analyzing and generating an event graph for the input sentence, which no longer depends on trigger words, avoids the long-tail phenomenon, linearizes the input sentence in a sequence-to-sequence manner, so that the data sparsity problem is improved. At the same time, a generation model based on Transformer is built, and through the decoding algorithm, the performance of event extraction is improved, the correlation between multiple events is explicitly modeled, the argument sharing phenomenon is solved, the dependency syntactic information is utilized in the generation model based on the event graph, the graph attention neural network is used to encode the dependency information, and the dual attention mechanisms of the dependency graph encoding layer and the sentence encoding layer are fused to further improve the performance of event extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is the flow schematic diagram of the present invention;
[0054] Figure 2 is the schematic diagram of the event graph in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0056] Embodiment 1
[0057] Please refer to Figure 1 shown, the event extraction method based on graph parsing provided by the present invention includes the following steps:
[0058] 1) Given the input text, determine whether the sentence of the given input text contains an event type, extract any node in the sentence of the given input text as the root node, attach the event type node after the root node as a child node. If the event type is included, the value of the root node is EVTS, connect the events included in the sentence of the given input text to form an event graph. If the event type is not included, the value of the root node is NA, and the event extraction ends.
[0059] For a given input text, the purpose of event extraction is to identify and predict the event types contained therein, as well as the corresponding arguments and their roles. To model the correlation between multiple events, the multiple events contained in the sentences of the given input text are regarded as a whole and linked together to form an event graph, as Figure 2 shown.
[0060] Please refer to Figure 2 shown. For the sentences of the given input text, the underlined marks in the sentences are all entity mentions in the sentence, which are known information for the model. The event elements contained in this sentence are linked together as a whole to form an event graph.
[0061] Specifically, a special node is first introduced as the root node, and each event type node is attached after the root node as a child node; multiple arguments of each specific event type are linked as child nodes of the event type, where these edges are marked with the argument roles played by the arguments in this event.
[0062] It should be noted that the root node of the event graph is not a virtual root. It has two possible values: EVTS and NA. If the input sentence does not contain any event types, the value of the root node is NA and the event extraction ends. If it contains event types, the value of the root node is EVTS. Predicting the value of the root node is to first determine whether the input sentence contains event types.
[0063] 2) Based on the Seq2Seq sequence-to-sequence framework, linearize the event graph to obtain a linearized sequence of the event graph.
[0064] The linearization process adopts two linearization methods: depth-first traversal and breadth-first traversal. At the same time, entity type nodes are added to the event graph, and the traversal order of nodes and edges is designed to make the edge of the argument role appear after the argument node. Finally, the linearized sequence of the event graph is as Figure 2 shown.
[0065] 3) Based on the linearized sequence of the event graph, design a generation model and a decoding algorithm based on Transformer. The specific methods of this generation model and decoding algorithm are as follows:
[0066] Set \(x = \lt x 1 ,\ldots,x n \gt\) as the sentence of the given input text, where \(x i \) represents the \(i\)-th word in the sentence, \(i = 1,2,\ldots,n\). At the same time, set \(E=\lt e 1 ,\ldots,e j ,\ldots,e n \gt\) as the entity mentions in the sentence, where \(e iDenote the i-th entity mention in the sentence, where i = 1, 2..j..n. The entity mention contains a total of k entity mentions, and k is a natural number. Each entity mention contains a head entity and an entity type, and a Transformer-based generation model is constructed accordingly;
[0067] The Transformer-based generation model needs to decode the token list y = <y 1 ,...,y m > in sequence, where y i is the i-th token in the token list, i = 1, 2...m. The value of the token y i is any one of the event type, event argument (i.e., entity mention), argument role, entity type, and special pointer symbol.
[0068] When the generation model generates argument nodes for a specific event type node, the generation model outputs the head entity of each argument as the argument. Let Y be the space of output solutions, and the goal of the generation model becomes finding the node sequence of the sentence x of the given input text:
[0069]
[0070] Then, a Transformer-based encoder is used for encoding, and the following decoder is used for decoding:
[0071]
[0072]
[0073] P(y j |x,y <j ) = softmax(g(s j )) (4)
[0074] where h i is the context hidden vector, Encoder is the Transformer encoder, s j is the m decoding symbols generated sequentially by the decoder, Decoder is the decoder, and y j is the j-th token.
[0075] 4) To alleviate the data sparsity, the pre-trained language model BART is used to convert the Transformer-based generation model into an Encoder-Decoder architecture, and the pre-trained language model BART is used to learn latent knowledge such as semantic information and syntactic relationships in advance, and then complete the learning of event extraction results.
[0076] After the decoding algorithm is completed, the event extraction method based on graph parsing further extracts the results after decoding the Transformer-based generation model by using dependency syntactic information, thereby further improving the performance of the present invention. The dependency syntactic information encodes the dependency graph using a graph attention neural network, and fuses the dual attention mechanisms of the dependency graph encoding layer and the sentence encoding layer, specifically including:
[0077] Obtain the context hidden vectors {h i , h i+1 ,..., h i+m-1} and the syntactic dependency graph hidden vectors {h′ i , h′ i+1 ,..., h′ i+m-1}, and apply the average pooling function to obtain the context representation h con and the dependency graph representation h syn ;
[0078] h con = pool(h i , h i+1 ,..., h i+m-1 ) (5)
[0079] h syn = pool(h i ′, h i ′ +1 ,..., h i ′ +m-1 ) (6)
[0080] Among them, pool is the pooling layer, which uses a gating mechanism to fuse the features of the context representation h con and the dependency graph representation h syn . The dependency graph representation h syn is fused into the context representation h con as shown in the following formula:
[0081]
[0082] Among them is the product operation, and the specific calculation method of the function g is shown in the following formula:
[0083] g = σ(W g [h syn ; h con + b g ) (8)
[0084] Among them, [h syn ; h con is the concatenation of h con and h syn , and Wg and b g are parameters of the generation model.
[0085] Example 2
[0086] Corresponding to the event extraction method based on graph parsing in Example 1, this Example 2 provides an event extraction system based on graph parsing, including a judgment module, a linearization module, a generation model, a decoding module, and a transfer learning module;
[0087] The judgment module is used to give the input text, judge whether the sentence of the given input text contains an event type, extract any node in the sentence of the given input text as the root node, attach the event type node after the root node as a child node. If it contains an event type, the value of this root node is EVTS, connect the events contained in the sentence of the given input text to form an event graph. If it does not contain an event type, the value of this root node is NA, and the event extraction ends.
[0088] For the given input text, the purpose of event extraction is to identify and predict the contained event types, corresponding arguments, and the roles they play. To model the correlation between multiple events, the multiple events contained in the sentence of the given input text are regarded as a whole and linked together to form an event graph, as Figure 2 shown.
[0089] Please refer to Figure 2 shown. For the sentence of the given input text, the underlined marks in the sentence are all entity mentions in the sentence, which are known information of the model. The event elements contained in this sentence are linked together as a whole to form an event graph.
[0090] Specifically, first, a special node is introduced as the root node, and each event type node is attached after the root node as a child node; multiple arguments of each specific event type are linked as child nodes of the event type, and the edges are marked with the argument roles that the arguments play in this event.
[0091] It should be noted that the root node of the event graph is not a virtual root. It has two possible values: EVTS and NA. If the input sentence does not contain any event type, the value of the root node is NA, and the event extraction ends. If it contains an event type, the value of the root node is EVTS. The prediction of the root node value is to first judge whether the input sentence contains an event type.
[0092] The linearization module is used to linearly process the event graph based on the Seq2Seq sequence-to-sequence framework to obtain the linearized sequence of the event graph.
[0093] The linearization process adopts two linearization methods: depth-first traversal and breadth-first traversal. At the same time, entity type nodes are added to the event graph, and the traversal order of nodes and edges is designed to make the edge of the argument role appear after the argument node. Finally, the linearization sequence of the event graph is as Figure 2 shown.
[0094] The generation model and decoding module are used to design a Transformer-based generation model and decoding algorithm based on the linearization sequence of the event graph. The specific method of this generation model and decoding algorithm is as follows:
[0095] Set \(x = \langle x 1 ,\ldots,x n \rangle\) as the sentence of the given input text, where \(x i \) represents the \(i\)-th word in the sentence, \(i = 1,2,\ldots,n\). At the same time, set \(E=\langle e 1 ,\ldots,e j ,\ldots,e n \rangle\) as the entity mentions in the sentence, where \(e i \) represents the \(i\)-th entity mention in the sentence, \(i = 1,2,\ldots,j,\ldots,n\). The entity mentions in total contain \(k\) entity mentions, \(k\) being a natural number, and each entity mention contains a head entity and an entity type. Based on this, a Transformer-based generation model is built;
[0096] The Transformer-based generation model needs to decode the token list \(y=\langle y 1 ,\ldots,y m \rangle\) in sequence, where \(y i \) is the \(i\)-th token in the token list, \(i = 1,2,\ldots,m\). The value of the token \(y i \) is any one of the event type, event argument (i.e., entity mention), argument role, entity type, and special pointer symbol.
[0097] When the generation model generates argument nodes for a specific event type node, the generation model outputs the head entity of each argument as the argument. Set \(Y\) as the space of the output solution, and the goal of the generation model becomes to find the node sequence of the sentence \(x\) of the given input text:
[0098]
[0099] Then, a Transformer-based encoder is used for encoding, and the following decoder is used for decoding:
[0100]
[0101]
[0102] P(yj |x,y <j ) = softmax(g(s j )) (4)
[0103] where p and P are both probabilities, and y <j is the value of each node before the j-th position, h i is the context hidden vector, Encoder is the encoder of the Transformer, s j is the m decoded symbols sequentially generated by the decoder, Decoder is the decoder, and y j is the j-th token.
[0104] The transfer learning module is used to alleviate the data sparsity. It uses the pre-trained language model BART to convert the Transformer-based generation model into an Encoder-Decoder architecture, and uses the pre-trained language model BART to learn potential knowledge such as semantic information and syntactic relationships in advance, and then completes the learning of event extraction results.
[0105] After the decoding algorithm is completed, the event extraction method based on graph parsing further extracts the results after decoding the Transformer-based generation model by using dependency syntactic information, so as to further improve the performance of the present invention. The dependency syntactic information encodes the dependency graph by using a graph attention neural network and fuses the dual attention mechanisms of the dependency graph encoding layer and the sentence encoding layer, specifically including:
[0106] Obtain the context hidden vectors {h i , h i+1 ,..., h i+m-1} and the syntactic dependency graph hidden vectors {h′ i , h′ i+1 ,..., h′ i+m-1}, and apply the average pooling function to obtain the context representation h con and the dependency graph representation h syn ;
[0107] h con = pool(h i , h i+1 ,..., h i+m-1 ) (5)
[0108] h syn = pool(h i ′, h i ′ +1 ,..., h i ′ +m-1 ) (6)
[0109] Among them, pool is the pooling layer, which uses a gating mechanism to represent the context h con and the dependency graph representation h syn to perform feature fusion between the two. The dependency graph representation h syn is fused into the context representation h con as shown in the following formula:
[0110]
[0111] where is the product operation. The specific calculation method of the function g is shown in the following formula:
[0112] g = σ(W g [h syn ; h con + b g ) (8)
[0113] where [h syn ; h con is the concatenation of h con and h syn . W g and b g are the parameters of the generation model.
Claims
1. A method for event extraction based on graph parsing, characterized in that, it includes the following steps: 1) Given the input text, determine whether the sentence in the given input text contains an event type, extract any node in the sentence of the given input text as the root node, append the event type node after the root node as a child node. If the event type is included, the value of this root node is EVTS, connect the events included in the sentence in the given input text to form an event graph. If the event type is not included, the value of this root node is NA, and the event extraction ends; 2) Based on the Seq2Seq sequence-to-sequence framework, linearize the event graph to obtain a linearized sequence of the event graph; 3) Based on the linearized sequence of the event graph, design a generation model and a decoding algorithm based on Transformer. The specific manner of this generation model and decoding algorithm is: Set \(x = \lt x\lt i , \lt j , \lt n , \lt i , \lt 1 , \lt 1 , \lt n ,\ldots,x\lt n > as the sentences of the given input text, where \(x\lt i \) represents the \(i\)-th word in the sentence, \(i = 1, 2,\ldots,n\). At the same time, set \(E=\lt e\lt 1 ,\ldots,e\lt j ,\ldots,e\lt n > as the entity mentions in the sentence, where \(e\lt i \) represents the \(j\)-th entity mention in the sentence, \(i = 1, 2,\ldots,j,\ldots,n\). The entity mentions in total contain \(k\) entity mentions, \(k\) is a natural number, and each entity mention contains a head entity and an entity type, so as to build a Transformer-based generation model; The Transformer-based generation model needs to decode the token list y = <y 1 ,..., y m > in sequence, where yi is the i-th token of the token list, i = 1, 2... m, and the value of the token y i is any one of the event type, event argument i.e., entity mention, argument role, entity type, and special pointer symbol; 4) Use the pre-trained language model BART to convert the generation model based on Transformer into an Encoder-Decoder architecture to complete the learning of event extraction results.
2. The method for event extraction based on graph parsing according to claim 1, characterized in that, in step 2), two linearization processing methods, depth-first traversal and breadth-first traversal, are used for linearization processing. At the same time, entity type nodes are added, and the traversal order of nodes and edges is designed to be swapped.
3. The method for event extraction based on graph parsing according to claim 2, characterized in that, in step 3), when the generation model generates argument nodes for specific event type nodes, the generation model outputs the head entity of each argument as the argument. Let Y be the space of the output solution, and the goal of the generation model becomes to find the node sequence of the sentence x of the given input text: then use the encoder based on Transformer for encoding and the following decoder for decoding: P(yj|x,y<j)=softmax(g(sj)) Among them, both p and P are probabilities, and y <j is the values of each node before the j-th position, h i is the context hidden vector, Encoder is the encoder of the Transformer, s j are m decoded symbols sequentially generated by the decoder, Decoder is the decoder, and y j is the j-th token.
4. The method for event extraction based on graph parsing according to claim 3, characterized in that, after the decoding algorithm is completed, for the method for event extraction based on graph parsing, use dependency syntactic information to further extract the result after decoding the generation model based on Transformer; this dependency syntactic information uses a graph attention neural network to encode the dependency graph, and fuses the dual attention mechanisms of the dependency graph encoding layer and the sentence encoding layer, specifically including: Obtain the context hidden vectors {h i , h i+1 ,..., h i+m-1} and the syntactic dependency graph hidden vectors {h i ′, h i ′ +1 ,..., h i ′ +m-1}. Apply the average pooling function to obtain the context representation hcon and the dependency graph representation h syn ; h con = pool(h i , h i+1 ,..., h i+m-1 ) h syn = pool(h i ′, h i ′ +1 ,..., h i ′ +m-1 ) Among them, pool is the pooling layer, and a gating mechanism is used to fuse the context representation hcon and the dependency graph representation h syn to perform feature fusion between the two. The dependency graph representation hsyn is fused into the context representation h con as shown in the following formula: Among them is a multiplication operation, and the specific calculation method of function g is shown in the following formula: g=σ(Wg[hsyn;hcon]+bg) Among them, [hsyn; hcon] is the concatenation of h con and hsyn, and Wg and b g are the parameters of the generation model.
5. An event extraction system based on graph parsing, characterized in that, it includes a judgment module, a linearization module, a generation model and a decoding module, and a conversion learning module; The judgment module is used to give the input text, determine whether the sentence in the given input text contains an event type, extract any node in the sentence of the given input text as the root node, append the event type node after the root node as a child node. If the event type is included, the value of this root node is EVTS, connect the events included in the sentence in the given input text to form an event graph. If the event type is not included, the value of this root node is NA, and the event extraction ends; The linearization module is used to linearize the event graph based on the Seq2Seq sequence-to-sequence framework to obtain the linearized sequence of the event graph; The generation model and decoding module is used to design a Transformer-based generation model and decoding algorithm based on the linearized sequence of the event graph. The specific method of the generation model and decoding algorithm is as follows: Set x = <x 1 ,...,x n >For a sentence of a given input text, where x i represents the i-th word in the sentence, i=1,2...n, and at the same time, set E= <e 1 ,..e j ..,e n > is an entity mention in a sentence, where e i Represents the i-th entity mention in the sentence, i = 1, 2..j..n, the entity mention contains a total of k entity mentions, k is a natural number, each of which contains a head entity and an entity type, so as to build a Transformer-based generation model; The Transformer-based generation model needs to decode the list of tokens y = <y 1 ,..., y m > in sequence, where yi is the i-th token in the list of tokens, i = 1, 2... m, and the value of the token y i is any one of the event type, event arguments i.e., entity mentions, argument roles, entity types, and special pointer symbols; The transfer learning module is used to convert the Transformer-based generation model into an Encoder-Decoder architecture using the pre-trained language model BART to complete the learning of event extraction results.
6. The graph parsing-based event extraction system according to claim 5, wherein, In the linearization module, the linearization process adopts two linearization methods, depth-first traversal and breadth-first traversal. At the same time, entity type nodes are added, and the traversal order of nodes and edges is designed to be swapped.
7. The graph parsing-based event extraction system according to claim 6, wherein, In the generation model and decoding module, when the generation model generates argument nodes for specific event type nodes, the generation model outputs the head entity of each argument as the argument. Let Y be the space of output solutions, and the goal of the generation model becomes to find the node sequence of the sentence x of the given input text: Then, a Transformer-based encoder is used for encoding, and the following decoder is used for decoding: P(yj|x,y<j)=softmax(g(sj)) Among them, both p and P are probabilities, and y <j is the values of each node before the j-th position, h i is the context hidden vector, Encoder is the encoder of the Transformer, s j are m decoded symbols sequentially generated by the decoder, Decoder is the decoder, and y j is the j-th token.
8. The graph parsing-based event extraction system according to claim 7, wherein, After the decoding algorithm is completed, for the graph parsing-based event extraction method, dependency syntactic information is used to further extract the result after decoding the Transformer-based generation model; the dependency syntactic information encodes the dependency graph using a graph attention neural network and fuses the dual attention mechanisms of the dependency graph encoding layer and the sentence encoding layer, specifically including: Obtain the context hidden vectors {h i , h i+1 ,..., h i+m-1} and the syntactic dependency graph hidden vectors {h i ′, h i ′ +1 ,..., h i ′ +m-1}. Apply the average pooling function to obtain the context representation h con and the dependency graph representation h syn ; h con = pool(h i , h i+1 ,..., h i+m-1 ) h syn = pool(h i ′, h i ′ +1 ,..., h i ′ +m-1 ) Among them, pool is the pooling layer, which uses a gating mechanism to represent the context h con and the dependency graph representation h syn to perform feature fusion between the two. The dependency graph representation h syn is fused into the context representation h con as shown in the following formula: Among them is a product operation, and the specific calculation method of function g is shown in the following formula: g = σ(W g [h syn ; h con + b g ) Among them, [h syn ; h con is the concatenation of h con and h syn , and W g and b g are the parameters of the generation model.
Citation Information
Patent Citations
Event extraction system and method oriented to open domain
CN106951438A
Event extraction method, event extraction device and electronic equipment
CN111401033A