Role Representation Multi-Round Learning Method and System for Discourse-Level Event Argument Extraction
By introducing multiple rounds of learning methods and graph attention networks in chapter-level event argument extraction, the problem of neglecting argument associations and relying on pre-trained models in the existing technology is solved, and more accurate and rich role representations and argument span optimization are achieved.
Patent Information
- Application Number
- CN202510332201.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In the chapter-level event argument extraction, the prior art ignores the interrelationship between the arguments in event instances, and relies on pre-trained language models to obtain role representations. It is limited to the event mode level or the overall level of the chapter, and fails to effectively capture the direct or indirect relationships between roles.
A multi-round learning method for character representation is proposed. By obtaining the initial role representation from the chapter, an event pattern-instance graph is established, a graph attention network is used to iterate the role representation, and the initial and historical role representations are integrated in each round of learning to gradually correct the argument span boundary.
It realizes the integrated capture of role associations in event mode and argument associations in event instances, outputs more accurate and rich role representations, and optimizes the argument span boundaries.
Smart Images

Figure CN119849574B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information extraction, and particularly to a multi-round learning method and system for role representation for discourse-level event argument extraction. Background Art
[0002] Discourse-level event argument extraction is a key task in information extraction, which extracts event-related arguments from a discourse and accurately identifies their roles. The existing role-based span selection strategies mainly capture role representations and discourse representations through pre-trained language models and prompt tuning techniques, so as to construct a span selector to predict spans for each role. These methods have limitations: (1) Regarding arguments as independent units, ignoring the mutual associations between arguments in an event instance; (2) Depending on pre-trained language models to obtain role representations and discourse representations, or only capturing the semantics of event types (roles) at the event pattern level, or only capturing the semantics of roles at the overall discourse level.
[0003] Therefore, there is a need to develop a role learning model that can integrally capture the role semantics contained in various direct or indirect associations between roles in an event pattern, between arguments in an event instance, and between pattern-instance matches. Summary of the Invention
[0004] In view of the above situation, the main object of the present invention is to propose a multi-round learning method and system for role representation for discourse-level event argument extraction to solve the above technical problems.
[0005] The present invention proposes a multi-round learning method for role representation for discourse-level event argument extraction, and the method includes the following steps:
[0006] Step 1, obtain an initial role representation from a discourse, and use the initial role representation to predict argument spans;
[0007] Step 2, establish role associations at the abstract event pattern level, span associations at the specific event instance level, and role-span associations between patterns and instances according to the event pattern information, event instance information, and argument spans in the discourse, to obtain an event pattern-instance graph;
[0008] Step 3, take the edges in the event pattern-instance graph as virtual nodes, and use a graph attention network to iteratively update the nodes and edges in the event pattern-instance graph, and output the updated role representation;
[0009] Step 4: Repeat Steps 1 to 3 iteratively for multiple rounds of learning, and during each round of learning, fuse the initial role representation, historical role representation, and the updated role representation in the current round as the prediction input for the argument span in the next round to gradually correct the span boundary. After the learning is completed, output the final role representation.
[0010] The present invention also proposes a multi-round learning system for role representation in discourse-level event argument extraction. Among them, the system applies the multi-round learning method for role representation in discourse-level event argument extraction as described above. The system includes:
[0011] A graph construction module, used for:
[0012] Obtain the initial role representation from the discourse and use the initial role representation to predict the argument span;
[0013] Establish role associations at the abstract event pattern level, span associations at the specific event instance level, and role-span associations between patterns and instances based on the event pattern information, event instance information, and argument span in the discourse to obtain an event pattern-instance graph;
[0014] An iterative interaction update module, used for:
[0015] Take the edges in the event pattern-instance graph as virtual nodes, and use a graph attention network to iteratively update the nodes and edges in the event pattern-instance graph, and output the updated role representation;
[0016] A role memory fusion module, used for:
[0017] During each round of learning, fuse the initial role representation, historical role representation, and the updated role representation in the current round as the prediction input for the argument span in the next round;
[0018] A multi-round span prediction optimization module, used for:
[0019] Perform multi-round learning in an iterative manner to gradually correct the span boundary. After the learning is completed, output the final role representation.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] 1. The present invention proposes the construction of an event pattern-instance graph, establishes role associations at the abstract event pattern level, span associations at the specific event instance level, and role-span associations between patterns and instances, and realizes the integrated capture of the potential semantics of roles from the graph.
[0022] 2. The present invention proposes a multi-round learning network for node-edge interaction. Through an optimized graph attention network, it iteratively updates the representations of nodes and edges in the graph, and designs a role memory fusion module to obtain a new round of role representations, optimizing the argument boundaries predicted in the new round.
[0023] 3. The present invention develops a discourse-level event argument extraction scheme that integrates the construction of event pattern-instance graphs and the multi-round learning network for node-edge interaction.
[0024] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flowchart of the multi-round learning method for role representation for discourse-level event argument extraction proposed by the present invention;
[0026] Figure 2 is a framework diagram of the multi-round learning method for role representation for discourse-level event argument extraction proposed by the present invention;
[0027] Figure 3 is an example diagram of the construction of the event pattern-instance graph proposed by the present invention;
[0028] Figure 4 is a structural diagram of the multi-round learning system for role representation for discourse-level event argument extraction proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0030] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0031] Please refer to Figure 1 and Figure 2 , where in the figure represents a product operation, represents an attention coefficient calculation operation.
[0032] This embodiment provides a multi-round learning method for role representation for discourse-level event argument extraction, and the method includes the following steps:
[0033] Step 1: Obtain the initial role representation from the passage, and use the initial role representation to predict the argument span.
[0034] As a further preferred embodiment of the present invention, the method for obtaining the initial role representation from the passage and using the initial role representation to predict the argument span includes the following steps:
[0035] Given a prompt, and input the prompt and the passage into the pre-trained language model BART to obtain the role representation of the prompt and the passage representation.
[0036] Preliminary span selection layer: First, select the representation of the P th role k as the initial role representation according to the representation of the prompt, and then obtain the preliminary predicted span according to the initial role representation and the passage representation. The corresponding process has the following relational expressions: r k ;
[0037] ;
[0038] ;
[0039] ;
[0040] ;
[0041] ;
[0042] where and respectively represent two different learnable parameters, and respectively represent the role representations of the start position and the end position of the role r k , and respectively represent the probability distributions of the start position and the end position of the predicted span corresponding to the role r k in the whole context, represents the initial role representation, represents the role representation in the u-th round, represents the passage ; u represents the current learning round of the model, u ∈U, where U is the total number of learning rounds of the model; when u = 1, is ; when u > 1, The calculation can be seen in the last formula of the final character representation acquisition part.
[0043] As a further preferred embodiment of the present invention, a method for obtaining the character representation of a given prompt by inputting the prompt and the passage into the pre-trained language model BART includes the following steps:
[0044] Define two different special tokens <t>< / t> and ;
[0045] Locate the trigger word in the passage, wrap the trigger word with two different special tokens to generate the preprocessed passage ; where , represents the n th word in the passage, represents the trigger word, <t>< / t> and respectively represent two different special tokens;
[0046] Input the preprocessed passage into the encoder of the pre-trained language model BART to obtain the passage representation of the passage. The corresponding process has the following relationship:
[0047] ;
[0048] where represents the encoder of the pre-trained language model BART, represents the passage representation of the passage, represents the preprocessed passage;
[0049] Given the prompt corresponding to a certain event type, input the passage representation and the prompt into the decoder of the pre-trained language model BART to obtain the character representation of the entire prompt P The corresponding process has the following relationship:
[0050] ;
[0051] where represents the decoder of the pre-trained language model BART, represents the given prompt; represents the representation of the entire prompt P which contains the representations of all characters under this event type.
[0052] Step 2: Establish the role association at the abstract event pattern level, the span association at the specific event instance level, and the role-span association between the pattern and the instance based on the event pattern information, event instance information, and argument span in the passage to obtain the event pattern-instance graph;
[0053] Please refer toFigure 3 , as a further preferred embodiment of the present invention, a method for obtaining an event pattern-instance graph by establishing role associations at the abstract event pattern level, span associations at the specific event instance level, and role-span associations between patterns and instances based on the event pattern information, event instance information, and argument spans in the text includes the following steps:
[0054] Construct an event pattern subgraph: Generate corresponding nodes according to the event types involved in the event in the text. At the same time, generate corresponding nodes for each role under this event type;
[0055] Establish an edge of attribute type between the event type node and the role node to obtain the event pattern subgraph;
[0056] Construct an event instance subgraph: Generate a trigger word node according to the event trigger word given in the event in the text. At the same time, generate corresponding span nodes according to the involved event, and fill the predicted spans into the corresponding span nodes;
[0057] Establish an edge between the trigger word node and the span node to obtain the event instance subgraph, where the type of the edge is served by the role describing the corresponding span;
[0058] Construct subgraph matching associations: Construct an edge of instance type between the event type node and the trigger word node, and construct an edge of value type between the role node and its corresponding span node to obtain the event pattern-instance graph;
[0059] Among them, an event instance refers to a predicted event triggered by a given event trigger word, which is composed of the trigger word and the spans predicted for each role.
[0060] Step 3: Use the edges in the event pattern-instance graph as virtual nodes, and use the graph attention network to iteratively update the nodes and edges in the event pattern-instance graph, and output the updated role representations;
[0061] , as a further preferred embodiment of the present invention, a method for using the edges in the event pattern-instance graph as virtual nodes and using the graph attention network to iteratively update the nodes and edges in the event pattern-instance graph and output the updated role representations specifically includes the following steps:
[0062] In terms of the embedded representation of nodes, except for the predicted spans to be filled, for other nodes in the event pattern-instance graph, first use the BERT model to obtain the representations of all words contained in each node in the event pattern-instance graph, and then use the average pooling method to obtain the representation of the entire node. For the edges in the graph, since the types of edges in the event pattern-instance graph are all represented by corresponding words, the BERT model is also used for vector initialization to obtain the representations of the corresponding edges.k A kind of role r k At the u representation of the span node in the
[0063] ;
[0064] ;
[0065] ;
[0066] ;
[0067] ;
[0068] Among them, represents the BERT model, represents average fusion, represents the n th word of the trigger word, represents the n th word of the event type, represents the role k of the n th word, represents the k th representation of the role edge, represents the trigger word representation, represents the event type representation, represents the k th role representation, respectively represent the edges of the attribute type, instance type, and value type, respectively represent the representations of the edges of the attribute type, instance type, and value type, and respectively represent the start position and end position of the predicted span corresponding to the role r k ; represents and within the range of the j th word representation, represents the k th role r k at the u th round of prediction of the span node representation;
[0069] Multi-round Learning Network for Node-Edge Interaction: Since the event pattern-instance graph has some special features and aims to enrich the semantics of role representation, and there are both role nodes and role edges in the graph. Therefore, when using the Graph Attention Network (GAT) to update the node representation, its connected edges are also regarded as virtual nodes and incorporated into the scope of information transmission. The specific process is to concatenate the edge representation into the corresponding neighbor node representation, then use the softmax function to normalize the masked attention coefficients to obtain the final attention distribution, and then update the node representation according to the attention distribution; the i attention coefficient, normalization process, and the representation of the j node between the i th and
[0070] th nodes are shown as follows:
[0071] ;
[0072] ;
[0073] Among them, represents the concatenation operation, represents the attention vector, represents the activation function, represents the transpose operation, represents the representation of node i , represents the representation of node j , and , the representation of the node is determined by the node type itself, represents the i th and j th nodes, , the representation of the edge is determined by the edge type between the nodes, represents a learnable weight matrix for transforming node features, represents the i node and the j node, represents the exponential function, represents the neighbor set of node i , represents the i node and the j node, represents the m th attention head, l in the i layer, j between the th attention coefficient between them The weight matrix of one attention head in the l layer, denotes the representation of node l- in the j 1st layer, denotes the representation of node l in the i updated representation, denotes the activation function, denotes the regularization operation. To prevent overfitting, dropout needs to be applied to the attention coefficients. denotes the number of attention heads, denotes the i th neighbor node set of the node, l ∈ L , L denotes the total number of layers of the GAT.
[0074] After the node representation is updated, the representation of the role edge is updated by fusing the following: the role node, the trigger word node, the updated span node, and the edge's own representation. This method fully considers the semantic associations of various elements in the graph. The u th round, the l layer, the k th type of role edge representation is as follows:
[0075] ;
[0076] Among them, denotes the edge weight matrix, , , respectively denote the span node representation of the u th round, the l layer, the k th type of role, the trigger word node representation of the u th round, the l layer, and the role node representation of the u th round, the l layer, the k th type of role. denotes the u th round, the l- 1st layer, the k th type of role edge representation. Through average fusion, all relevant semantic features are integrated into the new edge representation, making the semantics of the edge representation richer.
[0077] During the update process from layer 1 to L layer, the node representation and the edge representation will influence each other, continuously generating the latest role node representation and role edge representation. In the U th round, the LThe layer fuses the role node representation and the role edge representation to obtain the role representation finally output by the multi-round learning network, as follows:
[0078] ;
[0079] where g represents the meaning of U rounds of learning based on the event pattern-instance graph structure, represents the U round, the L layer, and the k th role edge representation, represents the U round, the L layer, and the k th role node representation, represents the role representation finally output by the multi-round learning network.
[0080] Step 4: Repeat Steps 1 to 3 in an iterative manner for multi-round learning, and in each round of learning, fuse the initial role representation, the historical role representation, and the role representation updated in the current round as the prediction input for the argument span in the next round to gradually correct the span boundary. After the learning is completed, output the final role representation.
[0081] As a further preferred embodiment of the present invention, fusing the initial role representation, the historical role representation, and the role representation updated in the current round specifically includes the following steps:
[0082] For role r k , assuming that the representation in the memory unit is (initially empty), fuse the historical role representation u in the current ( th) round, the role representation u output by the multi-round learning network in the th round, and the initial role representation . In addition, to ensure the comparability of different role representations (historical role representation, the representation output by the multi-round learning network, initial role representation) and avoid dimensional deviation, first normalize each representation, and the obtained role representation is as follows:
[0083] ;
[0084] ;
[0085] ;
[0086] where represents the historical role representation in the u th round, Represents the role representation of the output of the multi-round learning network in the u th round, represents the normalization operation, represents the historical role representation of the u th round after normalization, represents the role representation of the output of the multi-round learning network in the u th round after normalization, represents the initial role representation after normalization;
[0087] Next, the normalized representations are averaged and fused to generate the final role representation. The role representation after fusion in the u th round is calculated as follows:
[0088] ;
[0089] where, represents the historical role learning matrix.
[0090] As a further preferred embodiment of the present invention, in performing the above steps 1 to 4, it is implemented through an argument extraction model, and the training method of the argument extraction model includes the following steps:
[0091] After U rounds of learning, using the obtained final role representation containing rich semantic information ;
[0092] Adopt the span selection strategy in the primary selection span selection layer to output the probability distribution of the start position and end position of the predicted span in the entire context;
[0093] According to the probability distribution of the start position and end position of the predicted span in the entire context, generate a candidate span set through greedy search;
[0094] Based on minimizing the matching cost, use the Hungarian algorithm to assign a unique true span label to each predicted candidate span;
[0095] Construct a cross-entropy loss function according to the probability distribution of the start position and end position of the predicted span in the entire context and the cross-entropy of the corresponding true span label. The corresponding process has the following relationship:
[0096] ;
[0097] where, D represents the number of passages in the training set, represents the optimal assignment calculated by the Hungarian algorithm, and respectively represent the gold start position and gold end position of the k th role;
[0098] The argument extraction model is trained by updating weights and learning parameters to minimize the loss.
[0099] Please refer to Figure 4 This embodiment also provides a multi-round learning system for role representation in discourse-level event argument extraction. Among them, the system applies the multi-round learning method for role representation in discourse-level event argument extraction as described above. The system includes:
[0100] A graph construction module for:
[0101] Obtaining an initial role representation from the discourse and predicting argument spans using the initial role representation;
[0102] Establishing role associations at the abstract event pattern level, span associations at the specific event instance level, and role-span associations between patterns and instances based on event pattern information, event instance information, and argument spans in the discourse to obtain an event pattern-instance graph;
[0103] An iterative interaction update module for:
[0104] Regarding the edges in the event pattern-instance graph as virtual nodes, using a graph attention network to iteratively update the nodes and edges in the event pattern-instance graph, and outputting the updated role representation;
[0105] A role memory fusion module for:
[0106] Fusing the initial role representation, historical role representation, and the role representation updated in the current round as the prediction input for the argument spans in the next round during each round of learning;
[0107] A multi-round span prediction optimization module for:
[0108] Performing multi-round learning in an iterative manner to gradually correct the span boundaries. After the learning is completed, the final role representation is output.
[0109] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0110] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0111] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0112] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A multi-round learning method for role representation for extracting event arguments at the chapter level, characterized in that: The method comprises the following steps: Step 1: Obtain the initial role representation from the text and use it to predict the argument span; Step 2: Based on the event pattern information, event instance information and argument span in the text, establish the role association at the abstract event pattern level, the span association at the specific event instance level and the role-span association between the pattern and the instance to obtain the event pattern-instance graph; Step 3: Use the edges in the event pattern-instance graph as virtual nodes, use the graph attention network to iteratively update the nodes and edges in the event pattern-instance graph, and output the updated role representation; Step 4: Repeat steps 1 to 3 in an iterative manner for multiple rounds of learning. In each round of learning, the initial role representation, the historical role representation, and the updated role representation of the current round are fused as the prediction input of the argument span of the next round to gradually correct the span boundary. After the learning is completed, the final role representation is output; In step 2, according to the event pattern information, event instance information and argument span in the text, role associations at the abstract event pattern level, span associations at the specific event instance level and role-span associations between patterns and instances are established, and the method for obtaining an event pattern-instance graph includes the following steps: Generate corresponding nodes according to the event types involved in the chapter, and generate corresponding nodes for each role under the event type; Establish an edge of attribute type between the event type node and the role node to obtain the event pattern subgraph; Generate trigger word nodes according to the event trigger words given by the events in the chapter, and at the same time, generate corresponding span nodes according to the events involved, and fill the predicted span into the corresponding span nodes; Establish edges between trigger word nodes and span nodes to obtain event instance subgraphs; the edge type is played by the role describing the corresponding span; An instance type edge is constructed between the event type node and the trigger word node, and a value type edge is constructed between the role node and its corresponding span node. An event pattern-instance graph is obtained based on the instance type edge and the value type edge.
2. The multi-round learning method for character representation for extracting event arguments at the chapter level according to claim 1, characterized in that: In step 1, the method of obtaining the initial role representation from the passage and using the initial role representation to predict the argument span includes the following steps: Given a prompt, input the prompt and the text into the pre-trained language model BART to obtain the role representation and text representation of the prompt; According to the prompt P The expression is selected k Type of role r k The representation of is used as the initial role representation, and then the preliminary prediction span is obtained based on the initial role representation and the chapter representation. The corresponding process has the following relationship: ; ; ; ; ; in, and Represent two different learnable parameters, and Represents roles r k The role representation of the start and end positions of and Represents roles r k The probability distribution of the start and end positions of the corresponding prediction span in the entire context, represents the initial role representation, represents the role representation of round u, Indicates chapter The chapter indicates; u Indicates the number of rounds the model is currently learning. u ∈U, U is the total number of rounds of model learning; when u =1, That is .
3. The multi-round learning method for character representation for extracting event arguments at the chapter level according to claim 2, characterized in that: Given a prompt, and inputting the prompt and the passage into a pre-trained language model BART to obtain a role representation of the prompt includes the following steps: Define two different special tags <t>< / t> and ; Locate the trigger words in the article, use two different special tags to wrap the trigger words, and generate the preprocessed article ;in, , Indicates the first n words, Indicates trigger words, <t>< / t> and Respectively represent two different special marks; The preprocessed passage is input into the pre-trained language model BART encoder to obtain the passage representation of the passage. The corresponding process has the following relationship: ; in, represents the pre-trained language model BART encoder, The chapter that represents the chapter indicates that Represents the chapter after preprocessing; Given a prompt corresponding to a certain event type, the chapter representation and the prompt are input into the pre-trained language model BART decoder to obtain the entire prompt P The role representation of the corresponding process has the following relationship: ; in, represents the pre-trained language model BART decoder, Indicates a given prompt; Indicates the entire prompt P The representation includes the representation of all roles under this event type.
4. The multi-round learning method for character representation for extracting event arguments at the chapter level according to claim 2, characterized in that: In step 3, the edges in the event pattern-instance graph are used as virtual nodes, and the nodes and edges in the event pattern-instance graph are iteratively updated using a graph attention network. The method for outputting the updated role representation specifically includes the following steps: The BERT model is used to obtain the representation of all words contained in each node in the event pattern-instance graph, and then the average pooling method is used to obtain the representation of the entire node. The corresponding process has the following relationship: ; ; ; in, represents the BERT model, represents average fusion, Indicates the trigger word n words, Indicates the event type n words, Representing roles k No. n words, Indicates k The representation of role edges, Indicates trigger words, Indicates the event type. Indicates k Types of role representation; The edges in the event pattern-instance graph are initialized with BERT to obtain the representation of the corresponding edges. The corresponding process has the following relationship: ; in, The edges represent the attribute type, instance type, and value type respectively. The edge representations represent the attribute type, instance type, and value type respectively; According to the obtained prediction span, several word representations contained in the prediction span are fused to obtain the representation of the predicted span node. The corresponding process has the following relationship: ; in, and Represents roles r k The corresponding start and end positions of the prediction span, express and In the range j The expression of a word, Indicates k Type of role r k In the u Representation of span nodes for round prediction; The edge representation is concatenated to the corresponding neighbor node representation, and then the softmax function is used to normalize the masked attention coefficient to obtain the attention distribution. The corresponding process has the following relationship: ; in, Represents a splicing operation, represents the attention vector, represents the activation function, represents the transpose operation, Representation Node i ' Representation Node j , and ,The representation of a node is determined by the node type itself; Indicates i Nodes and j The representation of the edges between nodes, and ,The representation of edges is determined by the edge type between nodes; Represents a learnable weight matrix used to transform node features; Representation Node i and nodes j The original attention distribution between Regularization operation is applied to the attention distribution to obtain the final attention distribution. The corresponding process has the following relationship: ; in, represents the exponential function, Representation Node i The neighbor set of represents the regularization operation, Representation Node i and nodes j The final attention distribution between According to the final attention distribution, the representation of the node is updated to obtain the updated representation of the node. The corresponding process has the following relationship: ; in, Indicates m The attention head is l Node in layer i and nodes j The attention coefficient between Indicates m The attention head is l The weight matrix in the layer, Indicates l −1st layer mid-node j The expression, Indicates l Node in layer i The updated representation is, represents the activation function, represents the number of attention heads, Indicates i The set of neighbor nodes of a node, l ∈ L , L Indicates the total number of layers of GAT; After the node representation is updated, the updated k The span node representation of the role, the updated representation of the trigger word node, the updated k The representation of the role node is averaged and fused with the representation of the unupdated role edge to obtain the updated role edge representation. The corresponding process has the following relationship: ; in, represents the edge weight matrix, , , Respectively represent u Round l Tier k The span node representation of the role u Round l The representation of the trigger word node of the first u Round l Tier k The representation of a role node. Indicates u Round l-1 Tier k The representation of role edges; After the update, the role node representation is merged with the role edge representation to obtain the role representation that is finally output by the multi-round learning network, as shown below: ; Among them, g represents the event pattern-instance graph structure. U The meaning of round learning, Indicates U Round L Tier k The representation of role edges, Indicates U Round L Tier k The representation of a role node. Represents the character representation of the final output of the multi-round learning network.
5. The multi-round learning method for character representation for extracting event arguments at the chapter level according to claim 2, characterized in that: In step 4, fusing the initial role representation, the historical role representation, and the role representation updated in the current round specifically includes the following steps: The first u The historical role representation of the round, the multi-round learning network u The role representation of the round output and the initial role representation are normalized respectively, and the corresponding process has the following relationship: ; ; ; in, Indicates u The historical role of the wheel indicates that Represents the multi-round learning network u The role representation of the round output, represents the normalization operation, Represents the normalized u The historical role of the wheel indicates that Represents the normalized multi-round learning network u The role representation of the round output, represents the normalized initial role representation; The normalized representations are averaged and fused to generate the fused character representation. The corresponding process has the following relationship: ; in, Represents the historical role learning matrix.
6. The multi-round learning method for character representation for extracting event arguments at the chapter level according to claim 2, characterized in that: In executing the above steps 1 to 4, the argument extraction model is used for implementation. The training method of the argument extraction model includes the following steps: go through U Round learning, using the final character representation containing rich semantic information ; Adopt the span selection strategy in the preliminary span selection layer and output the probability distribution of the start and end positions of the predicted span in the entire context; Generate a set of candidate spans through greedy search according to the probability distribution of the start and end positions of the predicted span in the entire context; Based on minimizing the matching cost, the Hungarian algorithm is used to assign a unique true span label to each predicted candidate span; The cross entropy loss function is constructed based on the start and end positions of the predicted span, the probability distribution in the entire context, and the cross entropy of the corresponding true span label. The corresponding process has the following relationship: ; in, D represents the number of chapters in the training set, represents the optimal allocation calculated by the Hungarian algorithm, and Respectively represent k The gold starting position and gold ending position of each character; The argument extraction model is trained by updating the weights and learning the parameters to minimize the loss.
7. A multi-round learning system for role representation for extracting event arguments at the chapter level, characterized in that: The system applies the multi-round learning method for character representation for extracting chapter-level event arguments according to any one of claims 1 to 6, and the system comprises: Graph building blocks for: Obtaining initial role representation from the text and using the initial role representation to predict argument span; According to the event pattern information, event instance information and argument span in the text, role associations at the abstract event pattern level, span associations at the specific event instance level and role-span associations between patterns and instances are established to obtain an event pattern-instance graph. Iterative interactive update module for: The edges in the event pattern-instance graph are used as virtual nodes, and the graph attention network is used to iteratively update the nodes and edges in the event pattern-instance graph, and the updated role representation is output; Role memory fusion module, used for: In each round of learning, the initial role representation, historical role representation, and updated role representation of the current round are fused as the prediction input of the argument span of the next round; Multi-round span prediction optimization module, used for: Multiple rounds of learning are performed in an iterative manner to gradually correct the span boundary. After learning is completed, the final role representation is output.
Citation Information
Patent Citations
Document-level event extraction method and system based on integrated joint learning
CN116579338A
Role-aware chapter theme event argument extraction method and device
CN117194672A