A method, device, computer equipment and storage medium for event coreference resolution
By explicitly modeling event arguments and introducing confidence scores and gated filtering mechanisms, the problems of poor generalization and poor effect of event co-referential digestion in the prior art are solved, and a more accurate and globally consistent event co-referential chain reconstruction is achieved.
Patent Information
- Application Number
- CN202211052666.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-08-30
AI Technical Summary
The prior art has problems of poor generalization and poor results in event co-reference digestion, especially in applications in unlabeled corpus, and usually only considers local consistency and cannot guarantee the global optimal inference result.
A method of event co-referential dissolution is proposed. By explicitly modeling the event argument, the argument is divided into the actuator, the recipient, the time, the place and other five roles, and a confidence score is introduced in the argument representation, a gated filtering mechanism is designed to filter noise information, and the event chain is reconstructed to ensure global consistency.
It improves the generalization of the event co-reference digestion model and the accuracy of inference results, ensures a globally consistent event co-reference chain, and significantly improves the performance of the model.
Smart Images

Figure CN115422368B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of event graphs, and in particular to a method, apparatus, computer equipment and storage medium for event co-reference resolution by modeling arguments and reconstructing event chains. Background Art
[0002] In-document event coreference resolution is a task to identify and cluster event mentions in a text that refer to the same real event. It is a challenging research topic with many applications. Events are mainly composed of trigger words and event arguments. Trigger words are the main words in a sentence that can most clearly describe the occurrence of an event, while event arguments include other important information about the event, such as the agent, patient, time, and place.
[0003] In the relevant definition of ACE2005, an event mention has one and only one trigger word. Compared with the trigger word, the argument is more complex. How to use the argument reasonably to help solve the event co-reference is a difficult problem. On the one hand, different types of events have event arguments with different roles, and a single event may have more than one argument with the same role. On the other hand, it is impossible for all arguments of event mentions in the text to exist. In other words, arguments corresponding to certain roles are missing.
[0004] Previous work on event coreference resolution usually directly uses annotated event information (Bejan and Harabagiu, 2010; Krause et al., 2016). Such methods are heavily dependent on annotated data and have poor generalization. Current research mainly conducts event coreference resolution in unannotated corpora (Peng et al., 2016; Lu and Vincent, 2017). Such methods are more challenging and have more practical significance. Peng et al. designed a pipeline model for event extraction and event coreference resolution to extract event information from unannotated text and resolve coreference, but the pipeline method has the problem of error propagation. Lu and Vincent proposed a joint learning model to jointly learn event detection tasks, event anaphora detection and event coreference resolution tasks. The subtasks of the joint model promote each other, which can effectively alleviate the error propagation problem and enable the model to achieve the best performance.
[0005] Event arguments, as key information of events, are widely used in event coreference resolution tasks (Huang et al., 2019; Zeng et al., 2020; Lu and Vincent, 2021; Chen et al., 2021; Lu et al., 2022). Using event arguments to resolve event coreference resolution is based on the fact that if two event mentions are coreference, then the arguments of their corresponding roles are also coreference. Lu et al. implicitly modeled event arguments using BERT and jointly learned event detection and event coreference resolution. Lu and Vincent do not distinguish between events and arguments, but use spans instead of them. Based on the SpanBERT model, they jointly learn event coreference resolution and entity coreference resolution, and restrict the results by adding soft and hard constraints. Chen et al. only used trigger words and some arguments to resolve event coreference, missing some important argument information.
[0006] In addition, the existing technology usually only considers local consistency when inferring event coreference chains, and the obtained inference results cannot guarantee the global optimality.
[0007] The existing technology has the problems of poor generalization and poor effect. Summary of the invention
[0008] Based on this, it is necessary to provide an event coreference resolution method, apparatus, computer equipment and storage medium that can explicitly model event arguments in response to the above technical problems.
[0009] A method for event coreference resolution, comprising:
[0010] Obtain a training data set for event coreference resolution;
[0011] Inputting the training data set into an event coreference resolution model; the event coreference resolution model comprises an event extraction component, a mention encoder component, a coreference scorer component and an event coreference chain reconstruction module;
[0012] The event extraction component is used to obtain multiple event mentions according to the document data in the training data set; each event mention includes a trigger word, an argument and an event subtype; the argument is divided into five roles: agent, patient, time, place and others;
[0013] The mention encoder component is used to obtain a trigger word representation and an argument representation of an argument role of any event mention according to the word metadata of the event mention and the corresponding document; wherein the argument representation includes an argument confidence score; further obtain a trigger word pair representation and an argument pair representation of any two event mentions, filter the argument pair representation according to the trigger word pair representation through a gated filtering mechanism to obtain a filtered argument pair representation, and then obtain a mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation;
[0014] The coreference scorer component is used to obtain the coreference score of any two event mentions according to the mention pair representation;
[0015] The event co-reference chain reconstruction module is used to obtain the initial event co-reference chain of the corresponding document in the training data set according to the co-reference score, verify the single chain by calculating the co-reference score between the single chain in the initial event co-reference chain and other event chains, verify the long chain by calculating the co-reference score between any event mention in the long chain in the initial event co-reference chain and other event mentions in the long chain, and then reconstruct the initial event co-reference chain to obtain the predicted final event co-reference chain;
[0016] The event coreference resolution model is trained by using the training data set and the predicted final event coreference chain to obtain a trained event coreference resolution model;
[0017] The document data to be subjected to event coreference resolution is input into the trained event coreference resolution model to obtain a final event coreference chain corresponding to the document data.
[0018] In one embodiment, the method further includes: obtaining k event mentions {m1, m2, ..., m k} and the n-word metadata of the corresponding document;
[0019] The transformer encoder forms a context representation for each input word as X = (X1, X2, ..., X n );in, d represents the dimension of the vector after each word is encoded;
[0020] For each event mention m i , the event mentioned m i The trigger word indicates t i is defined as the average of its token embeddings:
[0021]
[0022] Among them, s i and e iRespectively represent the start and end index of the trigger word;
[0023] The incident mentioned m i The argument corresponding to role r is expressed as:
[0024]
[0025]
[0026] Among them, r∈{agent,patient,time,place,other}, agent,patient,time,place,other are five argument roles, namely agent, patient, time, place and others. It is mentioned i The representation of the lth argument corresponding to role r, and Respectively represent the start and end indexes of the lth argument, c represents the confidence score of the lth argument, and u represents m i The number of arguments corresponding to the role r; when m i The argument corresponding to role r is default or non-existent and is represented by a d-dimensional 0 vector.
[0027] In one embodiment, the method further includes: given two events mentioning m i and m j , respectively define the trigger word pair representation and the argument pair representation of the corresponding role r as:
[0028]
[0029]
[0030] Among them, FFNN t is a A standard feed-forward neural network, Code m i and m j Element-level similarity.
[0031] In one embodiment, the method further comprises: representing the trigger word pair according to the trigger word pair; ij For the pair of arguments Perform orthogonal decomposition to obtain the argument pair representation The orthogonal components and parallel components They are:
[0032]
[0033]
[0034] Define the orthogonal components and the parallel component The weight coefficients are:
[0035]
[0036] ω p =1-ω o
[0037] Among them, ω o and ω p are the orthogonal components and the parallel component The weight coefficient of FFNN p is a A feedforward neural network, σ is the sigmoid activation function;
[0038] Through the gated filtering mechanism, the filtered argument pair representation is obtained:
[0039]
[0040] In one embodiment, the method further includes: obtaining a mention pair representation f of any two events mentioned according to the trigger word pair representation and the filtered argument pair representation ij for:
[0041]
[0042] In one embodiment, the method further comprises: referring to the representation f ij The coreference score s(i,j) between any two event mentions is obtained as:
[0043] s(i,j)=FFNN a (f ij )
[0044] Among them, FFNN a yes Feedforward neural network.
[0045] In one embodiment, the method further includes: obtaining an initial event co-reference chain C1 of the corresponding document in the training data set according to the co-reference score = {c1, c2, ..., c l};
[0046] For each single-chain c i (c i ∈C={c1,c2,…,c l}), calculate its relationship with other event chains c j (c j ∈C-{c i})'s coreference score, if c i and c j The score of c is the highest and greater than the threshold ω1, and c is merged i and c j And update the event coreference chain;
[0047] For each long chain c with a length greater than 2, calculate the score of each event mention m and c-{m} in the chain in turn. If the score is less than the threshold ω2, move m out of c and update the event co-reference chain as a single chain;
[0048] The reconstruction of the initial event coreference chain is completed to obtain the predicted final event coreference chain.
[0049] An event co-reference elimination device, comprising:
[0050] A training data set acquisition module, used to acquire a training data set for event coreference resolution;
[0051] A training data input module is used to input the training data set into an event coreference resolution model; the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component and an event coreference chain reconstruction module; the event extraction component is used to obtain multiple event mentions based on the document data in the training data set; each event mention includes a trigger word, an argument and an event subtype of the event; the argument is divided into five roles: agent, patient, time, place and others; the mention encoder component is used to obtain a trigger word representation of any event mention and an argument representation of the argument role based on the event mention and the word metadata of the corresponding document; wherein the argument representation includes an argument confidence score; further, a trigger word pair representation and an argument pair representation of any two event mentions are obtained, and a gated filtering mechanism is used to obtain the trigger word pair representation and the argument pair representation of any two event mentions. , filtering the argument pair representation according to the trigger word pair representation to obtain the filtered argument pair representation, and then obtaining the mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation; the co-reference scorer component is used to obtain the co-reference score of any two event mentions according to the mention pair representation; the event co-reference chain reconstruction module is used to obtain the initial event co-reference chain of the corresponding document in the training data set according to the co-reference score, verifying the single chain by calculating the co-reference score of the single chain in the initial event co-reference chain with other event chains, and verifying the long chain by calculating the co-reference score of any event mention in the long chain in the initial event co-reference chain with other event mentions in the long chain, and then reconstructing the initial event co-reference chain to obtain the predicted final event co-reference chain;
[0052] A model training module, used for training the event coreference resolution model through the training data set and the predicted final event coreference chain to obtain a trained event coreference resolution model;
[0053] The model application module is used to input the document data to be subjected to event coreference resolution into the trained event coreference resolution model to obtain the final event coreference chain corresponding to the document data.
[0054] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0055] Obtain a training data set for event coreference resolution;
[0056] Inputting the training data set into an event coreference resolution model; the event coreference resolution model comprises an event extraction component, a mention encoder component, a coreference scorer component and an event coreference chain reconstruction module;
[0057] The event extraction component is used to obtain multiple event mentions according to the document data in the training data set; each event mention includes a trigger word, an argument and an event subtype; the argument is divided into five roles: agent, patient, time, place and others;
[0058] The mention encoder component is used to obtain a trigger word representation and an argument representation of an argument role of any event mention according to the word metadata of the event mention and the corresponding document; wherein the argument representation includes an argument confidence score; further obtain a trigger word pair representation and an argument pair representation of any two event mentions, filter the argument pair representation according to the trigger word pair representation through a gated filtering mechanism to obtain a filtered argument pair representation, and then obtain a mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation;
[0059] The coreference scorer component is used to obtain the coreference score of any two event mentions according to the mention pair representation;
[0060] The event co-reference chain reconstruction module is used to obtain the initial event co-reference chain of the corresponding document in the training data set according to the co-reference score, verify the single chain by calculating the co-reference score between the single chain in the initial event co-reference chain and other event chains, verify the long chain by calculating the co-reference score between any event mention in the long chain in the initial event co-reference chain and other event mentions in the long chain, and then reconstruct the initial event co-reference chain to obtain the predicted final event co-reference chain;
[0061] The event coreference resolution model is trained by using the training data set and the predicted final event coreference chain to obtain a trained event coreference resolution model;
[0062] The document data to be subjected to event coreference resolution is input into the trained event coreference resolution model to obtain a final event coreference chain corresponding to the document data.
[0063] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0064] Obtain a training data set for event coreference resolution;
[0065] Inputting the training data set into an event coreference resolution model; the event coreference resolution model comprises an event extraction component, a mention encoder component, a coreference scorer component and an event coreference chain reconstruction module;
[0066] The event extraction component is used to obtain multiple event mentions according to the document data in the training data set; each event mention includes a trigger word, an argument and an event subtype; the argument is divided into five roles: agent, patient, time, place and others;
[0067] The mention encoder component is used to obtain a trigger word representation and an argument representation of an argument role of any event mention according to the word metadata of the event mention and the corresponding document; wherein the argument representation includes an argument confidence score; further obtain a trigger word pair representation and an argument pair representation of any two event mentions, filter the argument pair representation according to the trigger word pair representation through a gated filtering mechanism to obtain a filtered argument pair representation, and then obtain a mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation;
[0068] The coreference scorer component is used to obtain the coreference score of any two event mentions according to the mention pair representation;
[0069] The event co-reference chain reconstruction module is used to obtain the initial event co-reference chain of the corresponding document in the training data set according to the co-reference score, verify the single chain by calculating the co-reference score between the single chain in the initial event co-reference chain and other event chains, verify the long chain by calculating the co-reference score between any event mention in the long chain in the initial event co-reference chain and other event mentions in the long chain, and then reconstruct the initial event co-reference chain to obtain the predicted final event co-reference chain;
[0070] The event coreference resolution model is trained by using the training data set and the predicted final event coreference chain to obtain a trained event coreference resolution model;
[0071] The document data to be subjected to event coreference resolution is input into the trained event coreference resolution model to obtain a final event coreference chain corresponding to the document data.
[0072] The event coreference resolution method, device, computer equipment and storage medium above construct an event coreference resolution model, including an event extraction component, a mention encoder component, a coreference scorer component and an event coreference chain reconstruction module. By explicitly modeling argument information and dividing the arguments into five roles including agent, patient, time, place and others, it can not only meet the needs of processing the arguments of corresponding roles separately, but also ensure that all argument information is included, and will not cause the loss of some argument information; by introducing confidence scores in the argument representation, the negative impact of error propagation is alleviated; by designing a gated filtering mechanism, trigger words are used to filter the noise in the arguments, further alleviating error propagation and obtaining the most useful information in a specific context; through the algorithm for reconstructing the event chain, the legitimacy of the single chain and the long chain in the inferred initial event coreference chain is further verified, so that the event coreference chain is updated and reconstructed to ensure the global consistency of the final event coreference chain, and the performance of the event coreference resolution model is further improved. The present invention improves the event coreference resolution model from multiple perspectives, improves the generalization of the event coreference resolution model, and optimizes the accuracy of the inference results. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 1 is a flow chart of an event coreference resolution method according to an embodiment;
[0074] Figure 2 A schematic diagram of a model and workflow of an event coreference resolution method in one embodiment;
[0075] Figure 3 A schematic diagram of the structure of an encoder mentioned in an embodiment;
[0076] Figure 4 is a schematic diagram of the structure of a door control module in one embodiment;
[0077] Figure 5 is a structural block diagram of an event coreference resolution device in one embodiment;
[0078] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0079] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0080] In one embodiment, Figure 1 As shown, a method for event coreference resolution is provided, comprising the following steps:
[0081] Step 102: Obtain a training data set for event coreference resolution.
[0082] The training data set of this embodiment is collected from the ACE2005 data set.
[0083] Step 104: input the training data set into the event coreference resolution model.
[0084] like Figure 2 As shown in FIG, the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component, and an event coreference chain reconstruction module. The event coreference chain reconstruction module is the part from the initial event coreference chain to the final event coreference chain.
[0085] Event extraction:
[0086] The event extraction component is used to obtain multiple event mentions based on the document data in the training dataset; each event mention includes the event trigger word, argument and event subtype; the argument is divided into five roles: agent, patient, time, place and others.
[0087] This embodiment uses OneIE to identify event mentions and their subtypes and arguments. OneIE is the most advanced event extraction method and has achieved the best results on the ACE2005 dataset. Unlike OneIE, the present invention modifies the results of argument prediction so that the model only outputs five roles including agent, patient, time, place and others. The present invention explicitly models argument information and divides the arguments into five roles to make it more general. The first four are the basic information of the event, and "others" contain argument information that cannot be divided into the first four. This improved division method can not only meet the needs of processing the arguments of corresponding roles separately, but also ensure that all argument information is included, and will not cause the loss of some argument information. It allows different types of arguments to be treated differently and make full use of various argument information to enable the model to achieve the best performance.
[0088] Mention the encoder:
[0089] The mention encoder component is used to obtain the trigger word representation and argument representation of the argument role of any event mention based on the tokens data of the event mention and the corresponding document; the argument representation includes the argument confidence score; further obtain the trigger word pair representation and argument pair representation of any two event mentions, and filter the argument pair representation according to the trigger word pair representation through the gated filtering mechanism to obtain the filtered argument pair representation, and then obtain the mention pair representation of any two event mentions based on the trigger word pair representation and the filtered argument pair representation.
[0090] Specifically, the input of the mention encoding is a matrix containing n tokens and k event mentions {m1,m2,…,m k} document D. In English, a word can be decomposed into multiple tokens, and each character in Chinese is a token. Figure 3 shown.
[0091] The model first uses a transformer encoder to form a contextual representation for each input token, and uses X = (X1, X2, ..., Xn) to represent the output of the encoder, where d represents the dimension of the vector after encoding each token. i , use s i and e i Respectively represent the start and end index of the trigger word (argument), and its trigger word represents t i is defined as the average of its token embeddings:
[0092]
[0093] Among them, t i It is mentioned i However, as mentioned above, m i There may be more than one argument corresponding to r, and errors in information extraction may have a negative impact on event coreference resolution. Each argument obtained by OneIE has a confidence score c∈(0,1], which indicates whether the argument is a reference to m. i The probability of the argument corresponding to role r. The present invention believes that when the argument confidence score c is closer to 1, the possibility of using it to introduce errors is smaller, and vice versa. In order to alleviate the negative impact of error propagation, the present invention introduces confidence scores in the argument representation, corresponding to the representation of the argument of role r The definition is as follows:
[0094]
[0095]
[0096] Among them, r∈{agent,patient,time,place,other}, It is mentioned i The representation of the lth argument corresponding to role r, and Respectively represent the start and end indexes of the lth argument, c represents the confidence score of the lth argument, and u represents m i The number of arguments corresponding to the role r. i After all the arguments corresponding to role r are represented, the pooling strategy is used to obtain the final argument representation When m i The argument corresponding to role r is default or non-existent, represented by a d-dimensional zero vector.
[0097] Given two mentions m i and m j , the representation of the trigger word pair and the argument pair corresponding to role r are defined as:
[0098]
[0099]
[0100] Among them, FFNN t is a A standard feed-forward neural network, Code m i and m j Element-level similarity.
[0101] In order to further alleviate error propagation and obtain the most useful information in a specific context, the present invention designs a gated filtering mechanism that uses trigger words to filter noise in arguments. Figure 4 shown.
[0102] First, the arguments are orthogonally decomposed based on the trigger word representation. yes In t ij The projection in the direction, which can be regarded as containing the ij In contrast, With t ij is orthogonal, so it can be considered to contain new information. When it is very clean and has complementary information, it should be used new information in the , and vice versa.
[0103]
[0104]
[0105] Then, use the following method to get the weight coefficients of the two components:
[0106]
[0107] ω p =1-ω o
[0108] Among them, ω o and ω p are the weight coefficients on the orthogonal component and the horizontal component respectively, FFNN p is a is a feedforward neural network, σ is the sigmoid activation function.
[0109] Get the filtered argument pair representation:
[0110]
[0111] Simply concatenate the trigger-based representation and all argument-based representations to construct the final mention pair representation:
[0112]
[0113] Coreference Scorer:
[0114] The coreference scorer component is used to obtain the coreference score of any two event mentions based on the mention pair representation.
[0115] m i and m j The coreference score s(i,j) of is:
[0116] s(i,j)=FFNN a (f ij )
[0117] Among them, FFNN a yes Feedforward neural network.
[0118] Event coreference chain reconstruction:
[0119] The event co-reference chain reconstruction module is used to obtain the initial event co-reference chain of the corresponding document in the training dataset according to the co-reference score, verify the single chain by calculating the co-reference score between the single chain and other event chains in the initial event co-reference chain, verify the long chain by calculating the co-reference score between any event mention in the long chain in the initial event co-reference chain and other event mentions in the long chain, and then reconstruct the initial event co-reference chain to obtain the predicted final event co-reference chain.
[0120] Specifically, for each event mention m i , the model will be drawn from all candidate mentions y iAssign it an antecedent m j Or the virtual antecedent ε:m j ∈y i ={ε,m1,m2,…,m i-1}. The virtual antecedent represents two situations: 1) m i Not an event mention; 2)m i is an event mention, but it does not co-reference with any of the previous event mentions. Set s(i,ε) = 0. A necessary condition for two mentions to co-reference is that they have the same event subtype, so here only mention pairs with the same event subtype are considered candidate co-referencing mention pairs.
[0121] The most straightforward way to construct an event coreference chain is to find the best event mention from each candidate event mention, i.e. the one with the highest coreference score:
[0122]
[0123] in, Indicates m i The candidate co-reference pair with the highest score. However, this greedy algorithm only considers local consistency and cannot guarantee the global optimal. The present invention designs a new chaining algorithm. Since singletons account for a large proportion, the present invention believes that it is necessary to consider the legitimacy of singletons again. In addition, the present invention regards event chains with a number of mentions greater than 2 as long chains. The complexity of the event chain increases with the length of the event chain, so it is necessary to additionally verify the legitimacy of each mention in the long chain. Specifically, given the event mentions {m1, m2, …, m k}, the model first obtains the initial event coreference chain through the above greedy algorithm, and then uses Algorithm 1 to obtain the final event coreference chain.
[0124]
[0125] In Algorithm 1, for each event chain c in D i (c i ∈C={c1,c2,…,c l}), the average pooling of all mentions in the chain is used as the chain representation. Lines 2 to 8 in Algorithm 1 verify the single chain. For each single chain c in document D i , use scorer to calculate c respectively i Chain with other events j (c j ∈C-{c i}) score, if c i and c j The score of c is the highest and greater than the threshold ω1, and c is merged i and cj And update the event chain. In lines 9 to 15 of Algorithm 1, for each event chain c with a length greater than 2, the scores of each event mention m and c-{m} in the chain are calculated in turn. If the score is less than the threshold ω2, it is considered that the possibility of m and other mentions in c being co-referenced is not high, so m is removed from c and used as a single chain to update the event co-reference chain. By using Algorithm 1, all single chains and long chains in D are reconstructed once.
[0126] Step 106 , training the event coreference resolution model using the training data set and the predicted final event coreference chain to obtain a trained event coreference resolution model.
[0127] The model goal is to output all the co-referencing event chains in the document. When the predicted antecedent of an event mention is its true co-referencing event, the predicted antecedent is considered to be the correct antecedent. In order to get the best results for the model, the present invention optimizes the marginal log-likelihood of all correct antecedents:
[0128]
[0129]
[0130] Among them, GOLD(i) represents m i The real coreference event chain of m i There is no real coreference event, then GOLD(i) = {ε}, P(i,j) represents m i With m j The probability of coreference.
[0131] Step 108 : input the document data to be subjected to event coreference resolution into the trained event coreference resolution model to obtain a final event coreference chain corresponding to the document data.
[0132] In the above event coreference resolution method, an event coreference resolution model is constructed, including an event extraction component, a mention encoder component, a coreference scorer component and an event coreference chain reconstruction module. By explicitly modeling argument information and dividing the arguments into five roles including agent, patient, time, place and others, it can not only meet the needs of processing the arguments of corresponding roles separately, but also ensure that all argument information is included, and will not cause the loss of some argument information; by introducing confidence scores in the argument representation, the negative impact of error propagation is alleviated; by designing a gated filtering mechanism, trigger words are used to filter the noise in the arguments, further alleviating error propagation and obtaining the most useful information in a specific context; through the algorithm for reconstructing the event chain, the legitimacy of the single chain and the long chain in the inferred initial event coreference chain is further verified, so that the event coreference chain is updated and reconstructed to ensure the global consistency of the final event coreference chain, further improving the performance of the event coreference resolution model. The present invention improves the event coreference resolution model from multiple perspectives, improves the generalization of the event coreference resolution model, and optimizes the accuracy of the inference results.
[0133] In a specific embodiment, the event coreference resolution model was trained according to the method of the present invention, and the results were analyzed. The details are as follows:
[0134] Dataset:
[0135] All experiments are conducted on the ACE2005 English dataset, which contains 599 documents. For comparison, 30 news articles are selected as the test set for experiments, and 40 other documents of different genres are randomly selected as the validation set. The remaining 529 documents are used to train the model.
[0136] Experimental setup:
[0137] The CoNLL and AVG metrics are used to measure the F1 score of the results. By definition, these two metrics are the other standard co-reference metrics B 3 、MUC、CEAF e and the average of BLANC. SpanBERT (spanbert-base-cased) is used as the Transformer encoder. Different learning rates are set for different tasks. The learning rate of SpanBERT is 5 e-5 , the task learning rate is 5 e-4 In the encoder, we set d=786, which is the encoding dimension of SpanBERT, and set the dimension of FFNN to p=500 and the depth to 1. In Algorithm 1, we set the thresholds ω1 and ω2 to 0. We set dropout=0.5, the batch size of each training to 8, and epoch=50.
[0138] Results Overview (Predicted Mentions):
[0139]
[0140] Table 1: Overall results of the end-to-end model on the ACE2005 dataset (using extracted trigger words and arguments).
[0141] Table 1 shows the overall results of the end-to-end model on the ACE2005 dataset. OneIE is used to extract event mentions, types, and arguments. Table 1 shows that compared with previous work, the model of the present invention achieves the best results. In terms of CoNLL and AVG, the performance is improved by 1.41 and 1.64 compared with the most advanced model.
[0142] Results summary (marked with mentions):
[0143]
[0144] Table 2: Results on the ACE 2005 dataset (using the trigger words and arguments annotated in the dataset).
[0145] In order to better analyze the effectiveness of the model of the present invention, experiments are also conducted using event mentions annotated in the dataset (Table 2). For this purpose, the confidence scores c of the arguments are all set to 1. As can be seen from Table 2, compared with Lai, the experimental results of the model of the present invention on real event mentions are improved by 2.25 and 1.66, which further illustrates the effectiveness of the method of the present invention.
[0146] Effects of event arguments:
[0147]
[0148] Table 3: Results of ablation experiments on event arguments on the ACE2005 dataset
[0149] In order to further explore the effects of each component of the model, an ablation experiment was conducted on the ACE2005 dataset. Table 3 shows the results of the ablation experiment on event arguments. The results show that, first, when the argument confidence score is deleted, that is, all argument confidence scores are set to 1, the CoNLL and AVG indicators decrease by 0.74 and 1.04 respectively, which shows that the introduction of argument confidence scores has a positive effect on resisting errors in the information extraction stage. Secondly, when the event argument corresponding to the role of 'other' is deleted, that is, only the four arguments with the roles of agent, patient, time and place are used, the CoNLL and AVG indicators decrease by 1.25 and 1.58 respectively, indicating that the argument with the role of 'other' contains information that is positive for solving event coreference, which proves the rationality of the argument division method. Finally, when all argument information is deleted, that is, only the trigger word is used, it can be seen from Table 3 that the effect of the model decreases by 2.90 and 4.03. The experimental results once again verified the conjecture. Although the pre-trained language model indirectly encodes arguments into the trigger word representation, which is the method currently used by most researchers, the method of the present invention explicitly models arguments, which has a significant impact on resolving event coreference.
[0150]
[0151] Table 4: Results of ablation experiments on the heavy chain algorithm on the ACE2005 dataset
[0152] In addition, this embodiment also conducts an ablation study on the heavy chain algorithm (Table 4). The results show that when the greedy algorithm is directly used to construct the event chain and used as the final event coreference chain, the CoNLL and AVG indicators decrease by 0.85 and 1.48 respectively. The study shows that the heavy chain algorithm of the present invention has a positive effect on constructing event coreference chains and improving the performance of the coreference resolution model.
[0153] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0154] In one embodiment, Figure 5As shown, an event coreference resolution device is provided, comprising: a training data set acquisition module 502, a training data input module 504, a model training module 506 and a model application module 508, wherein:
[0155] A training data set acquisition module 502 is used to acquire a training data set for event coreference resolution;
[0156] The training data input module 504 is used to input the training data set into the event coreference resolution model; the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component and an event coreference chain reconstruction module; the event extraction component is used to obtain multiple event mentions according to the document data in the training data set; each event mention includes the trigger word, argument and event subtype of the event; the argument is divided into five roles: agent, patient, time, place and others; the mention encoder component is used to obtain the trigger word representation of any event mention and the argument representation of the argument role according to the word metadata of the event mention and the corresponding document; the argument representation includes the argument confidence score; further obtain the trigger word pair representation and argument pair representation of any two event mentions, through The gated filtering mechanism filters the argument pair representation according to the trigger word pair representation to obtain the filtered argument pair representation, and then obtains the mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation; the co-reference scorer component is used to obtain the co-reference score of any two event mentions according to the mention pair representation; the event co-reference chain reconstruction module is used to obtain the initial event co-reference chain of the corresponding document in the training dataset according to the co-reference score, and verify the single chain by calculating the co-reference score of the single chain in the initial event co-reference chain with other event chains, and verify the long chain by calculating the co-reference score of any event mention in the long chain in the initial event co-reference chain with other event mentions in the long chain, and then reconstruct the initial event co-reference chain to obtain the predicted final event co-reference chain;
[0157] A model training module 506 is used to train the event coreference resolution model using the training data set and the predicted final event coreference chain to obtain a trained event coreference resolution model;
[0158] The model application module 508 is used to input the document data to be subjected to event coreference resolution into the trained event coreference resolution model to obtain the final event coreference chain corresponding to the document data.
[0159] The training data input module 504 is also used to obtain k event mentions {m1, m2, ..., m k} and the n word metadata of the corresponding document; the transformer encoder is used to form a context representation for each input word as X = (X1, X2, ..., X n );in, d represents the dimension of the vector after encoding each word; for each event mention m i , event mentions m i The trigger word indicates t i is defined as the average of its token embeddings:
[0160]
[0161] Among them, s i and e i Respectively represent the start and end index of the trigger word;
[0162] Event mentions m i The argument corresponding to role r is expressed as:
[0163]
[0164]
[0165] Among them, r∈{agent,patient,time,place,other}, agent,patient,time,place,other are five argument roles, namely agent, patient, time, place and others. It is mentioned i The representation of the lth argument corresponding to role r, and Respectively represent the start and end indexes of the lth argument, c represents the confidence score of the lth argument, and u represents m i The number of arguments corresponding to the role r; when m i The argument corresponding to role r is default or non-existent and is represented by a d-dimensional 0 vector.
[0166] The training data input module 504 is also used to give two event mentions m i and m j , respectively define the trigger word pair representation and the argument pair representation of the corresponding role r as:
[0167]
[0168]
[0169] Among them, FFNN t is a A standard feed-forward neural network, Code m i and m j Element-level similarity.
[0170] The training data input module 504 is also used to represent t according to the trigger word pair. ijArgument pair representation Perform orthogonal decomposition to obtain the argument pair representation The orthogonal components and parallel components They are:
[0171]
[0172]
[0173] Defining the orthogonal components and parallel components The weight coefficients are:
[0174]
[0175] ω p =1-ω o
[0176] Among them, ω o and ω p They are orthogonal components and parallel components The weight coefficient of FFNN p is a A feedforward neural network, σ is the sigmoid activation function;
[0177] Through the gated filtering mechanism, the filtered argument pair representation is obtained:
[0178]
[0179] The training data input module 504 is also used to obtain the mention pair representation f of any two event mentions based on the trigger word pair representation and the filtered argument pair representation ij for:
[0180]
[0181] The training data input module 504 is also used to represent f according to the mention pair ij The coreference score s(i,j) between any two event mentions is obtained as:
[0182] s(i,j)=FFNN a (f ij )
[0183] Among them, FFNN a yes Feedforward neural network.
[0184] The training data input module 504 is also used to obtain the initial event co-reference chain C1 of the corresponding document in the training data set according to the co-reference score = {c1, c2, ..., cl}; For each single chain c i (c i ∈C={c1,c2,…,c l}), calculate its relationship with other event chains c j (c j ∈C-{c i})'s coreference score, if c i and c j The score of c is the highest and greater than the threshold ω1, and c is merged i and c j And update the event co-reference chain; for each long chain c with a length greater than 2, calculate the score of each event mention m and c-{m} in the chain in turn. If the score is less than the threshold ω2, move m out of c and update the event co-reference chain with m as a single chain; complete the reconstruction of the initial event co-reference chain and obtain the predicted final event co-reference chain.
[0185] For the specific definition of the event coreference resolution device, please refer to the definition of the event coreference resolution method above, which will not be repeated here. Each module in the above-mentioned event coreference resolution device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0186] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an event coreference resolution method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0187] Those skilled in the art will understand that Figure 6The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0188] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above method embodiment when executing the computer program.
[0189] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.
[0190] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0191] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0192] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for event coreference resolution, characterized in that: The method comprises: Obtain a training data set for event coreference resolution; Inputting the training data set into an event coreference resolution model; the event coreference resolution model comprises an event extraction component, a mention encoder component, a coreference scorer component and an event coreference chain reconstruction module; The event extraction component is used to obtain multiple event mentions according to the document data in the training data set; each event mention includes a trigger word, an argument and an event subtype; the argument is divided into five roles: agent, patient, time, place and others; The mention encoder component is used to obtain a trigger word representation and an argument representation of an argument role of any event mention according to the word metadata of the event mention and the corresponding document; wherein the argument representation includes an argument confidence score; further obtain a trigger word pair representation and an argument pair representation of any two event mentions, filter the argument pair representation according to the trigger word pair representation through a gated filtering mechanism to obtain a filtered argument pair representation, and then obtain a mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation; The coreference scorer component is used to obtain the coreference score of any two event mentions according to the mention pair representation; The event co-reference chain reconstruction module is used to obtain the initial event co-reference chain of the corresponding document in the training data set according to the co-reference score, verify the single chain by calculating the co-reference score between the single chain in the initial event co-reference chain and other event chains, verify the long chain by calculating the co-reference score between any event mention in the long chain in the initial event co-reference chain and other event mentions in the long chain, and then reconstruct the initial event co-reference chain to obtain the predicted final event co-reference chain; The event coreference resolution model is trained by using the training data set and the predicted final event coreference chain to obtain a trained event coreference resolution model; The document data to be subjected to event coreference resolution is input into the trained event coreference resolution model to obtain a final event coreference chain corresponding to the document data.
2. The method according to claim 1, characterized in that According to the event mention and the word metadata of the corresponding document, the trigger word representation and the argument representation of the argument role of any event mention are obtained, including: Get k event mentions {m1,m2,…,m k } and the n-word metadata of the corresponding document; The transformer encoder forms a context representation for each input word as X = (X1, X2, ..., Xn); where d represents the dimension of the vector after each word is encoded; For each event mention m i , the event mentioned m i The trigger word indicates t i is defined as the average of its token embeddings: Among them, s i and e i Respectively represent the start and end index of the trigger word; The incident mentioned m i The argument corresponding to role r is expressed as: Among them, r∈{agent,patient,time,place,other}, agent,patient,time,place,other are five argument roles, namely agent, patient, time, place and others. It is mentioned i The representation of the lth argument corresponding to role r, and Respectively represent the start and end indexes of the lth argument, c represents the confidence score of the lth argument, and u represents m i The number of arguments corresponding to the role r; when m i The argument corresponding to role r is default or non-existent and is represented by a d-dimensional 0 vector.
3. The method according to claim 2, characterized in that Further, we can get the trigger word pair representation and argument pair representation of any two events, including: Given two events mentioning m i and m j , respectively define the trigger word pair representation and the argument pair representation of the corresponding role r as: Among them, FFNN t is a A standard feed-forward neural network, Code m i and m j Element-level similarity.
4. The method according to claim 3, characterized in that The argument pair representation is filtered according to the trigger word pair representation through a gated filtering mechanism to obtain a filtered argument pair representation, including: According to the trigger word pair representation t ij For the pair of arguments Perform orthogonal decomposition to obtain the argument pair representation The orthogonal components and parallel components They are: Define the orthogonal components and the parallel component The weight coefficients are: oh p =1-h o Among them, ω o and ω p are the orthogonal components and the parallel component The weight coefficient of FFNN p is a A feedforward neural network, σ is the sigmoid activation function; Through the gated filtering mechanism, the filtered argument pair representation is obtained: 。 5. The method according to claim 4, characterized in that According to the trigger word pair representation and the filtered argument pair representation, a mention pair representation of any two event mentions is obtained, including: According to the trigger word pair representation and the filtered argument pair representation, the mention pair representation f of any two events mentioned is obtained. ij for:
6. The method according to claim 5, characterized in that The co-reference score of any two event mentions is obtained according to the mentioned pair representation, including: According to the reference, f ij The coreference score s(i,j) between any two event mentions is obtained as: s(i,j)=FFNN a (f ij ) Among them, FFNN a yes Feedforward neural network.
7. The method according to claim 6, characterized in that The initial event co-reference chain of the corresponding document in the training data set is obtained according to the co-reference score, and the single chain is verified by calculating the co-reference score between the single chain and other event mentions in the initial event co-reference chain, and the long chain is verified by calculating the co-reference score between any event mention in the long chain in the initial event co-reference chain and other event mentions in the long chain, and then the initial event co-reference chain is reconstructed to obtain the predicted final event co-reference chain, including: The initial event co-reference chain C1 of the corresponding document in the training data set is obtained according to the co-reference score. l }; For each single-chain c i (c i ∈C={c1,c2,…,c l }), calculate its relationship with other event chains c j (c j ∈C-{c i })'s coreference score, if c i and c j The score of c is the highest and greater than the threshold ω1, and c is merged i and c j And update the event coreference chain; For each long chain c with a length greater than 2, calculate the score of each event mention m and c-{m} in the chain in turn. If the score is less than the threshold ω2, move m out of c and update the event co-reference chain as a single chain; The reconstruction of the initial event coreference chain is completed to obtain the predicted final event coreference chain.
8. An event coreference resolution device, characterized in that: The device comprises: A training data set acquisition module, used to acquire a training data set for event coreference resolution; A training data input module is used to input the training data set into an event coreference resolution model; the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component and an event coreference chain reconstruction module; the event extraction component is used to obtain multiple event mentions based on the document data in the training data set; each event mention includes a trigger word, an argument and an event subtype of the event; the argument is divided into five roles: agent, patient, time, place and others; the mention encoder component is used to obtain a trigger word representation of any event mention and an argument representation of the argument role based on the event mention and the word metadata of the corresponding document; wherein the argument representation includes an argument confidence score; further, a trigger word pair representation and an argument pair representation of any two event mentions are obtained, and a gated filtering mechanism is used to obtain the trigger word pair representation and the argument pair representation of any two event mentions. , filtering the argument pair representation according to the trigger word pair representation to obtain the filtered argument pair representation, and then obtaining the mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation; the co-reference scorer component is used to obtain the co-reference score of any two event mentions according to the mention pair representation; the event co-reference chain reconstruction module is used to obtain the initial event co-reference chain of the corresponding document in the training data set according to the co-reference score, verifying the single chain by calculating the co-reference score of the single chain in the initial event co-reference chain with other event chains, and verifying the long chain by calculating the co-reference score of any event mention in the long chain in the initial event co-reference chain with other event mentions in the long chain, and then reconstructing the initial event co-reference chain to obtain the predicted final event co-reference chain; A model training module, used for training the event coreference resolution model through the training data set and the predicted final event coreference chain to obtain a trained event coreference resolution model; The model application module is used to input the document data to be subjected to event coreference resolution into the trained event coreference resolution model to obtain the final event coreference chain corresponding to the document data.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Event co-reference resolution method and device based on modeling argument, equipment and medium
CN115422325A