An event element extraction system and method
By combining Bi-GRU encoding layer and graph Transformer Encoder layer, the attention mechanism is improved, and the problem of low efficiency in syntactic dependency modeling of existing methods is solved, and the accuracy of event element extraction and complex semantic recognition capabilities are improved.
Patent Information
- Application Number
- CN202310814588.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-07-04
AI Technical Summary
The existing event element extraction method is inefficient when utilizing syntactic information, making it difficult to effectively model syntactic dependencies, especially long-distance dependencies, which affects the performance of event element extraction.
The Bi-GRU encoding layer and the graph Transformer Encoder layer are combined, and the attention mechanism of syntactic distance is modeled, and the syntactic distance information is calculated on the graph structure using the dependent syntactic structure, and the attention mechanism of Transformer Encoder is improved to identify complex semantics.
It improves the model's modeling ability when the syntax distance is zero or too long, improves the accuracy of event element extraction, especially the recognition ability of complex semantics.
Smart Images

Figure CN116932683B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the event extraction technology in the field of natural language processing, and particularly relates to an event element extraction system and method integrating bidirectional encoding and graph Transformer Encoder. Background Art
[0002] The goal of event extraction is to detect event instances in text, and if any exist, extract information such as the event type and all participants and attributes related to the event. Event extraction consists of a two-stage task. The first stage is event detection, with the aim of extracting the key information of the event; the second stage is event element extraction, with the aim of extracting the relevant attributes of the event: event elements and their corresponding element roles. Compared with event detection, event element extraction has more entities to be extracted and greater difficulty.
[0003] Event element extraction has been widely applied in multiple practical scenarios. Specifically, the structured events after event element extraction can not only be directly used to expand the knowledge base, such as constructing an event knowledge graph, but also can be used for further event logical reasoning to help people analyze the rich information contained in the event and promote the practical applications in multiple fields such as information retrieval, recommendation systems, intelligent question answering, and automatic summarization.
[0004] The existing event element extraction mainly includes: [1] methods based on sequence labeling; [2] methods based on graph neural networks.
[0005] The existing methods have the following disadvantages:
[0006] (1) Methods based on sequence labeling cannot utilize the rich syntactic information in natural language sentences, resulting in the model ignoring the syntactic structure and having low efficiency when capturing the dependency relationships of sentences, thereby affecting the performance of event elements.
[0007] (2) Methods based on graph neural networks have difficulty in modeling long-distance dependency relationships or words without direct connections in the syntactic dependency graph, affecting the performance of event element extraction; and there is a problem of ignoring the information transmitted by the syntactic distance between entities, while the syntactic distance between entities can characterize certain event element types. Summary of the Invention
[0008] The present invention is achieved through the following technical solutions:
[0009] An event element extraction system, comprising:
[0010] An input layer, used to convert a sentence into a vector sequence containing multiple features to enrich the feature expression of the sentence;
[0011] Bi-GRU encoding layer, which is used to encode information in both forward and backward directions along the words in a sentence, and capture the Context representation of the input text;
[0012] Graph Transformer Encoder layer for modeling syntactic distance, which converts a sentence into a graph structure using the dependency syntactic structure and calculates the attention with syntactic distance information added on the graph structure;
[0013] Output layer, which is used to predict event elements and their roles as classification results.
[0014] In some embodiments, the input layer is used to encode each word w in sentence S i into a real-valued vector x i , so that the node representation contains richer feature expressions. The real-valued vector x i is composed of the concatenation of the word embedding vector word i , the part-of-speech tagging embedding vector pos i and the entity type label embedding vector et i , representing information such as the semantic features, part-of-speech features, position features, and entity features of the word respectively. The formula is x i =(word i ,pos i ,et i ).
[0015] In some embodiments, the Bi-GRU encoding layer is used to adaptively accumulate context, realizing the expansion of information flow without increasing the number of graph neural network layers, and at the same time avoiding the problems of gradient explosion and gradient disappearance existing in the basic RNN. The formula is
[0016] where is the forward GRU, which processes the words between x i and x n ; is the backward GRU, which processes the words between x i and x1, and || is the concatenation operation.
[0017] In some embodiments, the Graph Transformer Encoder layer for modeling syntactic distance includes:
[0018] Adding a Transformer Encoder structure to the GNN to learn the dependency relationships between words at different distances in the syntactic structure, which helps to improve the model's ability in extreme cases such as zero or too long syntactic distances.
[0019] The Transformer Encoder structure of the L layer takes the output vector H = {h1, h2,..., h n} of the Bi-GRU bidirectional encoding layer as the input of the first layer After passing through Hl = Transformer l (H l-1 )(l ∈ [1, L]) processing, recursively generate hidden state representations of different levels To aggregate the output vectors of the previous layer, n h self-attention heads are used in each layer. For the l-th layer Transformer Encoder layer, first use the parameter matrices and to linearly project the output of the previous layer onto the Q, K, and V matrices. The calculation formula is
[0020]
[0021]
[0022]
[0023] Further improve the attention mechanism in the Transformer Encoder. When calculating the attention weights between a word and other words in a sentence, add a strategy of pairwise syntactic distance to help identify complex semantics in the event.
[0024] The key idea is to add syntactic structure when operating on the mask matrix of the sentence and weight the attention based on pairwise syntactic distances.
[0025] The specific steps are as follows: First, analyze the general dependency syntactic structure of the sentence and calculate the shortest path syntactic distance between each pair of words. Then, add the syntactic distance to the original adjacency matrix of the sentence, and set the values that were originally 0 in the adjacency matrix (i.e., the two words are not connected) to the syntactic distance (the shortest number of steps to connect), generating a distance matrix D.
[0026] Represent the distance matrix as where D ij represents the syntactic distance between the words at positions i and j in the input sequence. If allowing words to add their adjacent words (syntactic distance of 1 step) at each layer, then set the mask matrix M as
[0027] Generalize the setting idea of the mask matrix to the distance-based attention model, allowing the mask matrix to incorporate distances within the vocabulary, thereby achieving the function of controlling the scale of syntactic distances incorporated into the mask matrix. At this time, the mask matrix is set to
[0028] Incorporate the mask matrix M that reflects syntactic and semantic information into the calculation process of the attention head A l in the attention mechanism, that is
[0029] After the softmax operation, an attention matrix is generated where T i represents the attention of the i-th word to other words in the sentence, and F is a function that modifies the attention weights, learning based on syntactic distances to make the model pay more attention to words that are closer in the dependency syntactic structure and assign smaller weights to words that are farther away. Define the calculation of the (i, j)-th element of the generated attention matrix by the F function as
[0030] where D ij is the distance between the i-th and j-th words. Z i is the normalization factor, calculated as
[0031] In some embodiments, the output layer is used to predict event elements and their corresponding element roles through a linear layer and a softmax function. In the sentence mark the event trigger word of the event element candidate word The output layer formula is
[0032] where and b a ∈R r are parameters, and r is the total number of element role label types. The probability of the t-th event type is expressed as P(a t |s, e t , e a ), the t-th element in the output O a of the corresponding role label.
[0033] To train the neural network, the cross-entropy loss function is minimized to optimize the parameters in the model,
[0034] where N is the number of event elements in the input sentence s, θ represents all the parameters in the model, represents the event trigger word e t and the event element e aThe true information (Ground Truth) event type among...
[0035] An event element extraction method, comprising the following steps:
[0036] Step 100, vectorize the sentence using the input layer;
[0037] Step 200, encode the context information of words in the sentence using a Bi-GRU structure;
[0038] Step 300, model the syntactic distance using the attention mechanism of the graph Transformer Encoder layer to identify complex semantics in the event;
[0039] Step 400, predict the event elements and role classification results using the event element classification layer.
[0040] In some embodiments, the step 100 includes the following sub-steps:
[0041] Step 110, input a sentence S of length n = {w1,..., w t ..., w i ,... w n}, and the event trigger word w t annotated at position t (t ∈ [1, n]) in the sentence;
[0042] Step 120, convert each word w i in the sentence S of step 110 into a real-valued vector x i , where x i is composed of the concatenation of the word embedding vector word i of the word w i , the part-of-speech tagging embedding vector pos i and the entity type label embedding vector et i , obtaining x i = (word i , pos i , et i ).
[0043] In some embodiments, the step 200 includes the following sub-steps:
[0044] Step 210, obtain the input vector X = {x1, x2,..., x n} of the Bi-GRU encoding layer;
[0045] Step 220, through the operation of the Bi-GRU, the vector X = {x1, x2,..., x n} encoded as H = {h1, h2,..., h n}.
[0046] In some embodiments, step 300 includes the following sub-steps:
[0047] Step 310, the Transformer Encoder structure of layer L takes the output vector H = {h1, h2,..., h n} of the Bi-GRU bidirectional encoding layer as the input of the first layer of the Transformer Encoder
[0048] Step 320, process H of step 310 through H l = Transformer l (H l-1 ) (l ∈ [1, L]) to recursively generate hidden state representations at different levels
[0049] Step 330, for the l-th layer Transformer Encoder layer of step 320, use the parameter matrices and to linearly project the output of the previous layer onto the Q, K, and V matrices
[0050]
[0051]
[0052]
[0053] Step 340, analyze the general dependency syntactic structure of the sentence, calculate the shortest path syntactic distance between each pair of words, and generate a symmetric distance matrix where D ij represents the syntactic distance between the words at positions i and j in the input sequence;
[0054] Step 350, set the mask matrix M according to the distance matrix D of step 340, allowing the mask matrix to include the vocabulary within the distance to achieve the function of controlling the scale of the syntactic distance added to the mask matrix. The mask matrix M is set to
[0055] Step 360, modify the self-attention mechanism in the Transformer Encoder, add a strategy for modeling syntactic distance, and add the mask matrix M reflecting syntactic semantics obtained in step 350 to the attention head A in the attention mechanisml During the calculation process
[0056] In some embodiments, step 400 includes the following sub-steps:
[0057] Step 410, mark event trigger words in the input sentence and candidate words of event elements
[0058] Step 420, for the output of step 410, predict the role labels of event elements through a linear layer and a softmax function
[0059] Step 430, in order to train the neural network, adopt the minimization of the cross-entropy loss function to optimize the parameters in the model
[0060] Compared with the prior art, the positive effects of the present invention are as follows:
[0061] Through the above-mentioned event element extraction method that fuses multiple neural networks, the present invention combines Bi-GRU, dependency syntax, and Transformer Encoder, and can learn the dependencies between words at different distances in the syntactic structure, which helps to improve the ability of the model to model extreme cases where the syntactic distance is zero or too long. On the other hand, this method notices that the syntactic distance between entities can characterize certain event element types, further improves the attention mechanism in Transformer Encoder, and adds a strategy of pairwise syntactic distance to help identify complex semantics in events. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is the overall flowchart of the system of the present invention.
[0063] Figure 2 is the algorithm flowchart of an event element extraction method that fuses bidirectional encoding and graph Transformer Encoder according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The present invention will be further described below in conjunction with the drawings and specific embodiments.
[0065] Figure 1 As shown, an event element extraction system includes an input layer for converting a sentence into a vector sequence containing multiple features to enrich the feature expression of the sentence.
[0066] In some embodiments, the input layer is used to encode each word w in sentence S i into a real-valued vector x i , so that the node representation contains richer feature expressions. The real-valued vector x i is composed of the concatenation of the word embedding vector word i , the part-of-speech tagging embedding vector pos i and the entity type label embedding vector et i , resulting in x i =(word i ,pos i ,et i ).
[0067] The event element extraction system further includes a Bi-GRU encoding layer, which is used to encode information in both the forward and backward directions of the words in the sentence, capturing the context information of the input text.
[0068] Specifically, in some embodiments, the Bi-GRU encoding layer is used to adaptively accumulate context, and the formula is
[0069] The event element extraction system further includes a graph Transformer Encoder layer for modeling syntactic distance, which converts the sentence into a graph structure using the dependency syntactic structure and calculates the attention with syntactic distance information added on the graph structure.
[0070] Specifically, in some embodiments, the graph Transformer Encoder layer for modeling syntactic distance includes adding a Transformer Encoder structure in the GNN, and using the parameter matrices and to linearly project the output of the previous layer onto the Q, K, and V matrices, and the calculation formula is
[0071]
[0072]
[0073]
[0074] Further improve the attention mechanism in the Transformer Encoder. First, analyze the general dependency syntactic structure of the sentence, calculate the shortest path syntactic distance between each pair of words, and generate the distance matrix D. Set the mask matrix as Allow the mask matrix to add distance The words within are used to control the scale of syntactic distance in the mask matrix added, thereby realizing the function of controlling the scale of syntactic distance in the mask matrix added. The mask matrix M reflecting syntactic and semantic information is added to the attention head A in the attention mechanism l during the calculation process, that is
[0075] The event element extraction system further includes an output layer, which is used to predict event elements and their roles as classification results through a linear layer and a softmax function
[0076] Specifically, in some embodiments, in the sentence mark the event trigger word candidate words of event elements The formula of the output layer is In order to train the neural network, the cross-entropy loss function is minimized to optimize the parameters in the model
[0077] such as Figure 2 shown, an event element extraction method mainly includes the following four steps
[0078] Step 100, vectorize the sentence using the input layer
[0079] Step 200, use the Bi-GRU structure to encode the context information of words in the sentence
[0080] Step 300, use the attention mechanism of the graph Transformer Encoder layer to model syntactic distance and identify complex semantics in the event
[0081] Step 400, use the event element classification layer to predict the event element and role classification results
[0082] Through the above event element extraction method that fuses multiple neural networks, the present invention combines Bi-GRU, dependency syntax, and Transformer Encoder, and can learn the dependency relationships between words at different distances in the syntactic structure, which helps to improve the ability of the model in extreme cases where the syntactic distance is zero or too long; on the other hand, this method notices that the syntactic distance between entities can characterize certain event element types, further improves the attention mechanism in the Transformer Encoder, and adds a strategy of pairwise syntactic distance to help identify complex semantics in the event, greatly improving the accuracy of event element extraction
[0083] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. All within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for extracting event elements, characterized in that, It includes the following steps: Step 100, vectorize the sentence using the input layer; Step 110: Input a sentence S with a length of n = {w1,…,w t …,w i ,...w n }, and the event trigger word w marked at sentence t t , t∈[1,n]; Step 120, for each word w in sentence S of Step 110 i Convert it into a real-valued vector x i , x i which is composed of the concatenation of the word embedding vector word i of word w i , the part-of-speech tagging embedding vector pos i and the entity type label embedding vector et i to obtain x i =(word i , pos i , et i ); Step 200, encode the context information of words in the sentence using the Bi-GRU structure; Step 210, obtain the input vector X = {x1, x2, …, x n}; Step 220, through the operation, encode the vector X = {x1, x2, …, x n} in Step 210 into H = {h1, h2, …, h n}; Step 300, model the syntactic distance using the attention mechanism of the graph Transformer Encoder layer to identify complex semantics in the event; Step 310, the Transformer Encoder structure of layer L uses the output vector H = {h1, h2, …, h n} of the Bi-GRU bidirectional encoding layer as the input of the first layer of the Transformer Encoder Step 320, subject the H in Step 310 to H l = Transformer l (H l-1 ) process to recursively generate hidden state representations at different levels Step 330: For the l-th Transformer Encoder layer in Step 320, use the parameter matrices and to linearly project the output of the previous layer onto the Q, K, and V matrices Q l = Linear(H l-1 ) = H l-1 W l Q K l = Linear(H l-1 ) = H l-1 W l K V l = Linear(H l-1 ) = H l-1 W l V Step 340, analyze the general dependency syntactic structure of the sentence, calculate the shortest path syntactic distance between each pair of words, and generate a symmetric distance matrix Step 350, set a mask matrix M according to the distance matrix D in step 340, allowing the mask matrix to include words within the distance and set them to Step 360: Modify the self-attention mechanism in the Transformer Encoder, add a strategy for modeling syntactic distance, and incorporate the mask matrix M obtained in Step 350 into the calculation process of the attention head A in the attention mechanism. l during the calculation Step 400, predict the event elements and role classification results using the event element classification layer; Step 410, in the input sentence mark the event trigger words and the candidate words of event elements Step 420, for the output of Step 410, predict the role labels of event elements through a linear layer and a softmax function. Step 430, to train the neural network, the cross-entropy loss function is minimized to optimize the parameters in the model.
2. An event element extraction system, characterized in that, Applied to the method described in claim 1, the system includes: An input layer, which inputs a text sentence and performs word embedding, part-of-speech tagging embedding, and entity type label embedding processing on each word in the sentence, converting the sentence into a vector sequence containing multiple features to enrich the feature expression of the sentence; A Bi-GRU encoding layer, before the sentence is input into the graph Transformer Encoder network, uses the Bi-GRU structure to encode information in both the forward and backward directions along the words in the sentence, capturing the context representation of the input text and reducing the limitations brought by the graph model; A graph Transformer Encoder layer for modeling syntactic distance. First, add the Transformer Encoder structure to the GNN to learn the dependencies between words at different distances in the syntactic structure. Then, further improve the attention mechanism in the Transformer Encoder and add a pairwise syntactic distance strategy to help identify complex semantics in the event; An event element classification layer, which predicts the event elements and their roles as the classification results through a linear layer and a softmax function.
Citation Information
Patent Citations
Carbon transaction text event extraction method based on graph neural network
CN114637827A
Event extraction method and system based on graph analysis
CN115169285A