Event Coreference Resolution Method, Device, Equipment and Medium Based on Modeling Arguments

The model-based event coreference resolution addresses the limitations of existing methods by explicitly modeling event arguments and using a gating mechanism to enhance accuracy and generalization.

CN115422325BActive Publication Date: 2025-07-15NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211045398.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-07-15
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

The prior art has problems of poor generalization and poor adaptability in event co-referential digestion, especially in unlabeled corpus, the challenges and mispropagation of event co-referential digestion.

Method used

Building events refers to a digestion model. By explicitly modeling event arguments, the arguments are divided into five roles: the actuator, the recipient, the time, the place and the other five roles, and a confidence score and a gated filtering mechanism are introduced to alleviate the propagation of errors and obtain the most useful information.

Benefits of technology

It improves the generalization and adaptability of event co-referencing, reduces the impact of argument information loss and error propagation, and improves the performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115422325B_ABST
    Figure CN115422325B_ABST
Patent Text Reader

Abstract

The present application relates to an event coreference resolution method, apparatus, computer device, and storage medium based on modeling arguments. The method includes: constructing an event coreference resolution model, including an event extraction component, a mention encoder component, a coreference scorer component, and an event coreference chain determination module. By explicitly modeling argument information and dividing the arguments into five roles including agent, patient, time, location, and others, it can not only meet the needs of separately processing the corresponding role arguments but also ensure that all argument information is included, without causing the lack of some argument information; by introducing a confidence score in the argument representation, the negative impact brought by error propagation is alleviated; by designing a gating filtering mechanism to filter out the noise in the arguments using trigger words, error propagation is further alleviated, and the most useful information in a specific context is obtained. The method of the present invention has the advantages of good effect and good adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of event graphs, and particularly to an event coreference resolution method, apparatus, computer device, and storage medium based on modeling arguments. Background Art

[0002] Event coreference resolution within a document is a task of identifying and clustering event mentions in a piece of text that refer to the same real event. It is a challenging research and has many applications. An event mainly consists of a trigger word and event arguments. The trigger word is the main word in a sentence that can most clearly describe the occurrence of an event, and event arguments include other important information of the event, such as the agent, patient, time, and location, etc.

[0003] In the relevant definition of ACE2005, an event mention has exactly one trigger word. Compared with trigger words, arguments are more complex, and how to reasonably use arguments to help solve event coreference is a difficult problem. On the one hand, different types of events have event arguments with different roles, and there may be more than one argument of the same role in a single event. On the other hand, the arguments of event mentions in the text may not all exist. In other words, arguments corresponding to certain roles are missing.

[0004] Previous work on event coreference resolution usually directly uses labeled event information (Bejan and Harabagiu, 2010; Krause et al., 2016). Such methods rely heavily on labeled data and have poor generalization. Current research mainly focuses on event coreference resolution in unlabeled corpora (Peng et al., 2016; Lu and Vincent, 2017). Such methods are more challenging and practically significant. Peng et al. designed a pipeline model for event extraction and event coreference resolution, extracting event information from unlabeled text and performing coreference resolution, but the pipeline method has an error propagation problem. Lu and Vincent proposed a joint learning model, jointly learning the event detection task, event anaphora detection, and event coreference resolution task. The subtasks of the joint model promote each other, which can effectively alleviate the error propagation problem and enable the model to achieve the best performance.

[0005] Event arguments, as the key information of events, are widely used in the event coreference resolution task (Huang et al., 2019; Zeng et al., 2020; Lu and Vincent, 2021; Chen et al., 2021; Lu et al., 2022). Using event arguments to solve event coreference resolution is based on the fact that if two event mentions are coreferential, then their corresponding role arguments are also coreferential. Lu et al. implicitly model event arguments using BERT and jointly learn event detection and event coreference resolution. Lu and Vincent do not distinguish between events and arguments, use spans to replace them, jointly learn event coreference resolution and entity coreference resolution based on the SpanBERT model, and limit the results by adding soft constraints and hard constraints. Chen et al. only use trigger words and some arguments to solve event coreference, missing some important argument information.

[0006] The prior art has problems of poor generalization and poor adaptability. Summary of the Invention

[0007] Based on this, in view of the above technical problems, it is necessary to provide an argument-based event coreference resolution method, device, computer device and storage medium that can explicitly model event arguments.

[0008] An argument-based event coreference resolution method, the method includes:

[0009] Obtain a training data set to be resolved for event coreference;

[0010] Input the training data set into an event coreference resolution model; the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component, and an event coreference chain determination module;

[0011] The event extraction component is used to obtain multiple event mentions according to the document data in the training data set; each event mention includes a trigger word, an argument, and an event subtype of the event; the argument is divided into five roles: agent, patient, time, location, and others;

[0012] The mention encoder component is used to obtain the trigger word representation of any event mention and the argument representation of the argument role according to the event mention and the token data of the corresponding document; the argument representation includes an argument confidence score; further obtain the trigger word pair representation and the argument pair representation of any two event mentions, and filter the argument pair representation according to the trigger word pair representation through a gating filtering mechanism to obtain the filtered argument pair representation, and then obtain the mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation;

[0013] The co-reference scorer component is used to obtain the co-reference score for any two event mentions according to the mentions;

[0014] The event co-reference chain determination module is used to obtain the predicted event co-reference chain of the corresponding document in the training data set according to the co-reference score;

[0015] The event co-reference resolution model is trained by the training data set and the predicted event co-reference chain to obtain a trained event co-reference resolution model;

[0016] The document data to be resolved for event co-reference is input into the trained event co-reference resolution model to obtain the event co-reference chain data corresponding to the document data.

[0017] In one of the embodiments, it further includes: obtaining k event mentions {m1, m2,..., m k} output by the event extraction component and n token data of the corresponding document;

[0018] The Transformer encoder is used to form a context representation for each input token as X = (X1, X2,..., X n ); where, d represents the vector dimension after encoding each token;

[0019] For each event mention m i , the trigger word representation t i of the event mention m i is defined as the average of its token embeddings:

[0020]

[0021] where, s i and e i represent the start and end indices of the trigger word respectively;

[0022] The argument representation of the event mention m i corresponding to the role r is:

[0023]

[0024]

[0025] where, r ∈ {agent, patient, time, place, other}, and agent, patient, time, place, and other are five argument roles of agent, patient, time, place, and others respectively, is the representation of the l-th argument of the mention m i corresponding to the role r, and respectively represent the start and end indices of the l-th argument, c represents the confidence score of the l-th argument, and u represents m i the number of arguments corresponding to the role r; when m i the argument corresponding to the role r is default or does not exist, it is represented by a d-dimensional zero vector.

[0026] In one embodiment, it further includes: given two event mentions m i and m j , the trigger word pair representation and the argument pair representation corresponding to the role r are respectively defined as:

[0027]

[0028]

[0029] where FFNN t is a standard feedforward neural network, encoding the element-wise similarity of m i and m j .

[0030] In one embodiment, it further includes: performing orthogonal decomposition on the argument pair representation ij according to the trigger word pair representation t to obtain the orthogonal component and the parallel component of the argument pair representation respectively as:

[0031]

[0032]

[0033] Define the weight coefficients of the orthogonal component and the parallel component respectively as:

[0034]

[0035] ω p = 1 - ω o

[0036] where ω o and ω p are respectively the weight coefficients of the orthogonal component and the parallel component , FFNN p is a feedforward neural network, and σ is the sigmoid activation function;

[0037] Through the gating filtering mechanism, the filtered argument pair representation is obtained:

[0038]

[0039] In one embodiment, it further includes: obtaining the mention pair representation f of any two event mentions according to the trigger word pair representation and the filtered argument pair representation ij That is:

[0040]

[0041] In one embodiment, it further includes: according to the mention pair representation f ij The coreference score s(i, j) of any two event mentions is obtained as:

[0042] s(i, j) = FFNN a (f ij )

[0043] Wherein, FFNN a is a feedforward neural network.

[0044] In one embodiment, it further includes: for any event mention, obtaining event mentions with the same event subtype as it as candidate coreference mention pairs;

[0045] According to the coreference scores of the candidate coreference mention pairs, the coreference event of the event mention is determined by the greedy algorithm;

[0046] Furthermore, the predicted event coreference chain of the corresponding document in the training data set is determined.

[0047] An event coreference resolution device based on modeling arguments, the device includes:

[0048] A training data set acquisition module, configured to acquire a training data set to be resolved for event coreference;

[0049] A training data input module for inputting the training data set into an event coreference resolution model; the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component, and an event coreference chain determination module; the event extraction component is used to obtain a plurality of event mentions according to the document data in the training data set; each of the event mentions includes a trigger word, an argument, and an event subtype of the event; the argument is divided into five roles: agent, patient, time, location, and others; the mention encoder component is used to obtain a trigger word representation of any event mention and an argument representation of the argument role according to the event mention and the token data of the corresponding document; wherein the argument representation includes an argument confidence score; further obtain a trigger word pair representation and an argument pair representation of any two event mentions, and through a gating filtering mechanism, filter the argument pair representation according to the trigger word pair representation to obtain a filtered argument pair representation, and then obtain a mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation; the coreference scorer component is used to obtain a coreference score of any two event mentions according to the mention pair representation; the event coreference chain determination module is used to obtain a predicted event coreference chain of the corresponding document in the training data set according to the coreference score;

[0050] A model training module for training the event coreference resolution model through the training data set and the predicted event coreference chain to obtain a trained event coreference resolution model;

[0051] A model application module for inputting the document data to be resolved for event coreference into the trained event coreference resolution model to obtain the event coreference chain data corresponding to the document data.

[0052] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0053] Obtain a training data set to be resolved for event coreference;

[0054] Input the training data set into an event coreference resolution model; the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component, and an event coreference chain determination module;

[0055] The event extraction component is used to obtain a plurality of event mentions according to the document data in the training data set; each of the event mentions includes a trigger word, an argument, and an event subtype of the event; the argument is divided into five roles: agent, patient, time, location, and others;

[0056] The mentioned encoder component is used to obtain the trigger word representation of any event mention and the argument representation of the argument role according to the event mention and the token metadata of the corresponding document; wherein the argument representation includes an argument confidence score; further obtaining the trigger word pair representation and the argument pair representation of any two event mentions, and through a gating filtering mechanism, filtering the argument pair representation according to the trigger word pair representation to obtain a filtered argument pair representation, and then obtaining the mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation;

[0057] The coreference scorer component is used to obtain the coreference score of any two event mentions according to the mention pair representation;

[0058] The event coreference chain determination module is used to obtain the predicted event coreference chain of the corresponding document in the training data set according to the coreference score;

[0059] Training the event coreference resolution model through the training data set and the predicted event coreference chain to obtain a trained event coreference resolution model;

[0060] Inputting the document data to be resolved for event coreference into the trained event coreference resolution model to obtain the event coreference chain data corresponding to the document data.

[0061] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0062] Obtaining a training data set to be resolved for event coreference;

[0063] Inputting the training data set into an event coreference resolution model; the event coreference resolution model includes an event extraction component, an encoder component for mentions, a coreference scorer component, and an event coreference chain determination module;

[0064] The event extraction component is used to obtain a plurality of event mentions according to the document data in the training data set; each event mention includes a trigger word, an argument, and an event subtype of the event; the argument is divided into five roles: agent, patient, time, location, and others;

[0065] The mentioned encoder component is used to obtain the trigger word representation of any event mention and the argument representation of the argument role according to the event mention and the token metadata of the corresponding document; wherein the argument representation includes an argument confidence score; further obtaining the trigger word pair representation and the argument pair representation of any two event mentions, and through a gating filtering mechanism, filtering the argument pair representation according to the trigger word pair representation to obtain a filtered argument pair representation, and then obtaining the mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation;

[0066] The co-reference scorer component is used to obtain the co-reference score for any two event mentions according to the mention;

[0067] The event co-reference chain determination module is used to obtain the predicted event co-reference chain of the corresponding document in the training data set according to the co-reference score;

[0068] The event co-reference resolution model is trained by the training data set and the predicted event co-reference chain to obtain a trained event co-reference resolution model;

[0069] The document data to be resolved for event co-reference is input into the trained event co-reference resolution model to obtain the event co-reference chain data corresponding to the document data.

[0070] The above-mentioned method, device, computer device and storage medium for event co-reference resolution based on modeling arguments construct an event co-reference resolution model, including an event extraction component, a mention encoder component, a co-reference scorer component and an event co-reference chain determination module. By explicitly modeling argument information and dividing arguments into five roles including agent, patient, time, location and others, it can not only meet the needs of processing corresponding role arguments separately, but also ensure that all argument information is included, without causing the lack of some argument information; by introducing a confidence score in the argument representation, the negative impact brought by error propagation is alleviated; by designing a gating and filtering mechanism to filter out the noise in the argument using trigger words, the error propagation is further alleviated and the most useful information in a specific context is obtained. Description of the Drawings

[0071] Figure 1 It is a schematic flow chart of an event co-reference resolution method in an embodiment;

[0072] Figure 2 It is a schematic diagram of the model and working process of an event co-reference resolution method in an embodiment;

[0073] Figure 3 It is a schematic structural diagram of a mention encoder in an embodiment;

[0074] Figure 4 It is a schematic structural diagram of a gating module in an embodiment;

[0075] Figure 5 It is a block diagram of the structure of an event co-reference resolution device in an embodiment;

[0076] Figure 6 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0077] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.

[0078] In one embodiment, as Figure 1 shown, a method for event coreference resolution based on modeling arguments is provided, including the following steps:

[0079] Step 102, obtain a training data set for which event coreference resolution is to be performed.

[0080] The training data set in this embodiment is collected from the ACE2005 data set.

[0081] Step 104, input the training data set into the event coreference resolution model.

[0082] As Figure 2 shown, the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component, and an event coreference chain determination module. Among them, the event coreference chain determination module is the part from the initial event coreference chain to the final event coreference chain.

[0083] Event extraction:

[0084] The event extraction component is used to obtain multiple event mentions based on the document data in the training data set; each event mention includes a trigger word, arguments, and an event subtype of the event; the arguments are divided into five roles: agent, patient, time, location, and others.

[0085] In this embodiment, OneIE is used to identify event mentions and their subtypes and arguments. OneIE is the most advanced event extraction method and has obtained the best results on the ACE2005 data set. Different from OneIE, the present invention modifies the result of argument prediction so that the model only outputs the five roles including agent, patient, time, location, and others. The present invention explicitly models argument information and divides the arguments into five roles, making it more general. The first four are the basic information of the event, and "others" includes argument information that cannot be classified into the first four. This improved division method can not only meet the needs of separately processing the arguments of corresponding roles, but also ensure that all argument information is included, without causing the lack of some argument information, allowing different treatment of various types of arguments, making full use of various argument information, and enabling the model to achieve the best performance.

[0086] Mention encoder:

[0087] The mention encoder component is used to obtain the trigger word representation of any event mention and the argument representation of the argument role according to the event mention and the token data of the corresponding document; the argument representation includes an argument confidence score; further, the trigger word pair representation and the argument pair representation of any two event mentions are obtained, and through a gating filtering mechanism, the argument pair representation is filtered according to the trigger word pair representation to obtain the filtered argument pair representation, and then the mention pair representation of any two event mentions is obtained according to the trigger word pair representation and the filtered argument pair representation.

[0088] Specifically, the input of the mention encoding is a document D containing n tokens and k event mentions {m1, m2,..., m k}. In English, a word can be decomposed into multiple tokens, and in Chinese, each character is a token. The mention encoder is as Figure 3 shown.

[0089] The model first uses a Transformer encoder to form a context representation for each input token, and uses X = (X1, X2,..., X n ) to represent the output of the encoder, where d represents the vector dimension of each encoded token. For each m i , use s i and e i to represent the start and end indices of the trigger word (argument) respectively, and its trigger word representation t i is defined as the average value of its token embeddings:

[0090]

[0091] where, t i is the trigger word representation of the mention m i . However, as mentioned above, there may be more than one argument corresponding to r for the mention m i , and errors from information extraction may have a negative impact on event coreference resolution. Each argument obtained by OneIE has a confidence score c ∈ (0, 1], which represents the probability that the argument is an argument of the role r corresponding to the mention m i . The present invention believes that when the argument confidence score c is closer to 1, the possibility of introducing errors using it is smaller, and vice versa. To mitigate the negative impact brought by error propagation, the present invention introduces a confidence score into the argument representation, and the representation of the argument corresponding to the role r is defined as follows:

[0092]

[0093]

[0094] where r ∈ {agent, patient, time, place, other}, is the representation of the l-th argument i corresponding to the role r of m, and respectively represent the start and end indices of the l-th argument, c represents the confidence score of the l-th argument, and u represents the number of arguments i corresponding to the role r of m. After obtaining all the argument representations i corresponding to the role r of m, a pooling strategy is adopted to obtain the final argument representation When the argument i corresponding to the role r of m is default or does not exist, it is represented by a d-dimensional zero vector.

[0095] Given two mentions m i and m j , the representations of the trigger word pair and the argument pair corresponding to the role r are respectively defined as:

[0096]

[0097]

[0098] where FFNN t is a standard feedforward neural network that encodes the element-wise similarity of m i and m j .

[0099] To further mitigate error propagation and obtain the most useful information in a specific context, the present invention designs a gated filtering mechanism to filter out noise in the arguments using the trigger word. As Figure 4 shown.

[0100] First, the argument is orthogonally decomposed based on the trigger word representation. The parallel component is the projection of in the direction of t ij , which can be regarded as containing the information that is already part of t ij . In contrast, is orthogonal to t ij , so it can be regarded as containing new information. When the original vector is very clean and has complementary information, the new information in should be utilized, and vice versa.

[0101] Parallel

[0102] Orthogonal

[0103] Then, the following method is used to obtain the weight coefficients of the two components:

[0104]

[0105] ω p = 1 - ω o

[0106] where ω o and ω p are the weight coefficients on the orthogonal component and the horizontal component respectively, FFNN p is a feed - forward neural network, and σ is the sigmoid activation function.

[0107] Obtain the filtered argument pair representation:

[0108]

[0109] Simply concatenate the trigger - based representation and all argument - based representations to construct the final mention pair representation:

[0110]

[0111] Coreference scorer:

[0112] The coreference scorer component is used to obtain the coreference score between any two event mentions according to the mention pair representation.

[0113] m i and m j The coreference score s(i, j):

[0114] s(i, j)= FFNN a (f ij )

[0115] where FFNN a is a feed - forward neural network.

[0116] Event coreference chain determination:

[0117] The event coreference chain determination module is used to obtain the predicted event coreference chain of the corresponding document in the training dataset according to the coreference score.

[0118] Specifically, for each event mention m i , the model will assign an antecedent m i or a dummy antecedent ε to it from all candidate mentions y j : m j ∈ y i = {ε, m1, m2, …, mi-1}. The dummy antecedent represents two cases: 1) m i is not an event mention; 2) m i is an event mention, but it is not co-referential with all the previous event mentions. Set s(i, ε) = 0. A necessary condition for two mentions to be co-referential is to have the same event subtype. Therefore, only the mention pairs with the same event subtype are used as candidate co-referential mention pairs here.

[0119] The most direct method to construct an event co-reference chain is to find the best one from each candidate event mention, that is, the event mention with the highest co-reference score:

[0120]

[0121] where represents the candidate co-reference pair with the highest score for m i .

[0122] According to the candidate co-reference pair with the highest score for each event mention m i , the predicted event co-reference chain for the entire document is further determined.

[0123] Step 106, train the event co-reference resolution model with the training dataset and the predicted final event co-reference chain to obtain a trained event co-reference resolution model.

[0124] The model objective is to output all co-reference event chains in the document. When the predicted antecedent of an event mention is its true co-reference event, this predicted antecedent is considered the correct antecedent. To enable the model to obtain the best results, the present invention optimizes the marginal log-likelihood of all correct antecedents:

[0125]

[0126]

[0127] where GOLD(i) represents the true co-reference event chain of m i , if there is no true co-reference event for m i , then GOLD(i) = {ε}, and P(i, j) represents the probability that m i is co-referential with m j .

[0128] Step 108, input the document data to be resolved for event co-reference into the trained event co-reference resolution model to obtain the event co-reference chain data corresponding to the document data.

[0129] In the above event coreference resolution method based on modeling arguments, an event coreference resolution model is constructed, including an event extraction component, a mention encoder component, a coreference scorer component, and an event coreference chain determination module. By explicitly modeling argument information and dividing arguments into five roles including agent, patient, time, location, and others, it can not only meet the needs of separately processing corresponding role arguments but also ensure that all argument information is included, without causing the lack of some argument information. By introducing a confidence score in the argument representation, the negative impact brought by error propagation is alleviated. By designing a gating filtering mechanism and using trigger words to filter out the noise in the arguments, error propagation is further alleviated, and the most useful information in a specific context is obtained.

[0130] In a specific embodiment, the event coreference resolution model is trained according to the method of the present invention, and the results are analyzed as follows:

[0131] Dataset:

[0132] All experiments are carried out on the ACE2005 English dataset, which contains 599 documents. For comparison, 30 news articles are selected as the test set for experiments, and 40 other documents of different genres are randomly selected as the validation set. The remaining 529 documents are used to train the model.

[0133] Experimental settings:

[0134] The F1 scores of the results are measured using the CoNLL and AVG metrics. By definition, these two metrics are the average values of other standard coreference metrics B 3 , MUC, CEAF e and BLANC. SpanBERT (spanbert-base-cased) is used as the Transformer encoder. Different learning rates are set for different tasks. The learning rate of SpanBERT is 5 e-5 , and the task learning rate is 5 e-4 . In the mention encoder, d = 786 is set, which is the encoding dimension of SpanBERT, and the dimension p of the FFNN is set to 500, with a depth of 1.

[0135] Overview of results (predicted mentions):

[0136]

[0137] Table 1: Overall results of the end-to-end model on the ACE2005 dataset (using the extracted trigger words and arguments).

[0138] Table 1 shows the overall results of the end-to-end model on the ACE2005 dataset. OneIE is used to extract event mentions, types, and arguments. Table 1 shows that compared with previous work, the model of the present invention achieves the best results, with a 0.46 and 0.16 improvement in performance over the state-of-the-art model in terms of the CoNLL and AVG metrics, respectively.

[0139] It should be understood that although Figure 1 the steps in the flowchart are shown sequentially as indicated by the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction on the execution of these steps, and these steps may be executed in other orders. Moreover, Figure 1 at least a portion of the steps in

[0140] In one embodiment, as Figure 5 shown, there is provided an event coreference resolution device based on modeling arguments, including: a training dataset acquisition module 502, a training data input module 504, a model training module 506, and a model application module 508, wherein:

[0141] The training dataset acquisition module 502 is configured to acquire a training dataset for which event coreference resolution is to be performed;

[0142] A training data input module 504 for inputting the training data set into an event coreference resolution model; the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component, and an event coreference chain determination module; the event extraction component is used to obtain a plurality of event mentions according to the document data in the training data set; each event mention includes a trigger word, arguments, and an event subtype of the event; the arguments are divided into five roles: agent, patient, time, location, and others; the mention encoder component is used to obtain a trigger word representation of any event mention and an argument representation of the argument role according to the event mention and the token data of the corresponding document; wherein the argument representation includes an argument confidence score; further obtaining a trigger word pair representation and an argument pair representation of any two event mentions, and filtering the argument pair representation according to the trigger word pair representation through a gating filtering mechanism to obtain a filtered argument pair representation, and then obtaining a mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation; the coreference scorer component is used to obtain a coreference score of any two event mentions according to the mention pair representation; the event coreference chain determination module is used to obtain a predicted event coreference chain of the corresponding document in the training data set according to the coreference score;

[0143] A model training module 506 for training the event coreference resolution model through the training data set and the predicted event coreference chain to obtain a trained event coreference resolution model;

[0144] A model application module 508 for inputting document data to be resolved for event coreference into the trained event coreference resolution model to obtain event coreference chain data corresponding to the document data.

[0145] The training data input module 504 is further configured to obtain k event mentions {m1, m2,..., m k} output by the event extraction component and n token data of the corresponding document; forming a context representation X=(X1, X2,..., X n ) for each input token through a transformer encoder; wherein, d represents the vector dimension after encoding each token; for each event mention m i , the trigger word representation t i of the event mention m i is defined as the average of its token embeddings:

[0146]

[0147] wherein, s i and e i respectively represent the start and end indices of the trigger word;

[0148] The event mentions m i The argument corresponding to the role r is represented as:

[0149]

[0150]

[0151] where r ∈ {agent, patient, time, place, other}, and agent, patient, time, place, and other are five argument roles of agent, patient, time, place, and others respectively. is the mention m i the representation of the l-th argument corresponding to the role r, and respectively represent the start and end indices of the l-th argument, c represents the confidence score of the l-th argument, and u represents m i the number of arguments corresponding to the role r; when m i the argument corresponding to the role r is default or does not exist, it is represented by a d-dimensional zero vector.

[0152] The training data input module 504 is also used to given two event mentions m i and m j , respectively define the trigger word pair representation and the argument pair representation corresponding to the role r as:

[0153]

[0154]

[0155] where FFNN t is a standard feedforward neural network, encoding the element-wise similarity of m i and m j .

[0156] The training data input module 504 is also used to perform orthogonal decomposition on the argument pair representation ij according to the trigger word pair representation t to obtain the orthogonal component and the parallel component of the argument pair representation respectively as:

[0157]

[0158]

[0159] Define the orthogonal component and the parallel component The weight coefficients are as follows:

[0160]

[0161] ω p = 1 - ω o

[0162] where ω o and ω p are the weight coefficients of the orthogonal component and the parallel component respectively, and FFNN p is a feed - forward neural network, and σ is the sigmoid activation function;

[0163] Through the gated filtering mechanism, the filtered argument pair representation is obtained:

[0164]

[0165] The training data input module 504 is also used to obtain the mention pair representation f of any two event mentions based on the trigger word pair representation and the filtered argument pair representation ij as:

[0166]

[0167] The training data input module 504 is also used to obtain the coreference score s(i, j) of any two event mentions based on the mention pair representation f ij as:

[0168] s(i, j)= FFNN a (f ij )

[0169] where FFNN a is a feed - forward neural network.

[0170] The training data input module 504 is also used to, for any event mention, obtain event mentions with the same event subtype as it as candidate coreference mention pairs; determine the coreference event of the event mention through the greedy algorithm according to the coreference scores of the candidate coreference mention pairs; and further determine the predicted event coreference chain of the corresponding document in the training data set.

[0171] For the specific limitations of the event coreference resolution device based on modeled arguments, reference may be made to the limitations of the event coreference resolution method based on modeled arguments in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned event coreference resolution device based on modeled arguments can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0172] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an event coreference resolution method based on modeled arguments. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0173] Those skilled in the art can understand that Figure 6 the structure shown in

[0174] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0175] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in the above method embodiment.

[0176] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0177] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0178] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.

Claims

1. An event coreference resolution method based on modeling arguments, characterized in that, The method includes: Obtaining a training data set for event coreference resolution to be performed; Inputting the training data set into an event coreference resolution model; the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component, and an event coreference chain determination module; The event extraction component is used to obtain a plurality of event mentions according to the document data in the training data set; each event mention includes a trigger word, arguments, and an event subtype of the event; the arguments are divided into five roles: agent, patient, time, location, and others; The mention encoder component is used to obtain a trigger word representation of any event mention and an argument representation of the argument role according to the event mention and the token data of the corresponding document; the argument representation includes an argument confidence score; further obtain a trigger word pair representation and an argument pair representation of any two event mentions, and filter the argument pair representation according to the trigger word pair representation through a gating filtering mechanism to obtain a filtered argument pair representation, and then obtain a mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation; The coreference scorer component is used to obtain a coreference score of any two event mentions according to the mention pair representation; The event coreference chain determination module is used to obtain a predicted event coreference chain of the corresponding document in the training data set according to the coreference score; Training the event coreference resolution model through the training data set and the predicted event coreference chain to obtain a trained event coreference resolution model; Inputting the document data for which event coreference resolution is to be performed into the trained event coreference resolution model to obtain the event coreference chain data corresponding to the document data.

2. The method according to claim 1, wherein Obtaining a trigger word representation of any event mention and an argument representation of the argument role according to the event mention and the token data of the corresponding document includes: Obtain k event mentions {m1, m2, …, m k} output by the event extraction component and n token data of the corresponding document; Form context representations for each input token through a Transformer encoder as X = (X1, X2, …, X n ); where, d represents the vector dimension after encoding each token; For each event mention m i , the trigger word representation t i of the event mention m i is defined as the average of its token embeddings: Among them, s i and e i represent the start and end indices of the trigger word, respectively; The event mentions m i The argument corresponding to the role r is expressed as: where r ∈ {agent, patient, time, place, other}, and agent, patient, time, place, and other are five argument roles of agent, patient, time, place, and others respectively. is the mention of m i representation of the l-th argument corresponding to the role r, and respectively represent the start and end indices of the l-th argument, c represents the confidence score of the l-th argument, and u represents the number of arguments of m i corresponding to the role r; when m i the argument corresponding to the role r is default or does not exist, it is represented by a d-dimensional zero vector.

3. The method according to claim 2, wherein Further obtaining a trigger word pair representation and an argument pair representation of any two event mentions includes: Given two event mentions m i and m j , the trigger word pair representation and the argument pair representation corresponding to the role r are respectively defined as: Among them, FFNN t is a standard feed-forward neural network that encodes the element-wise similarity of m i and m j .

4. The method according to claim 3, wherein Filtering the argument pair representation according to the trigger word pair representation through a gating filtering mechanism to obtain a filtered argument pair representation includes: According to the trigger word, the representation of t ij For the argument pair representation Perform orthogonal decomposition to obtain the orthogonal component of the argument pair representation and the parallel component respectively as follows: Define the orthogonal component and the weight coefficients of the parallel component are respectively:[[]] ω p = 1 - ω o where, ω o and ω p are the weight coefficients of the orthogonal component and the parallel component respectively, FFNN p is a feedforward neural network, and σ is the sigmoid activation function; Obtaining a filtered argument pair representation through a gating filtering mechanism: 。 5. The method according to claim 4, wherein Obtaining a mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation includes: Obtain the mention pair representation f of any two event mentions based on the trigger word representation and the filtered argument pair representation ij as follows:

6. The method according to claim 5, characterized in that Obtaining a coreference score of any two event mentions according to the mention pair representation includes: According to the mentioned representation f ij The co-reference score s(i, j) obtained for any two event mentions is as follows: s(i,j) = FFNN a (f ij ) Among them, FFNN a is a feedforward neural network.

7. The method according to claim 6, characterized in that, Obtaining a predicted event coreference chain of the corresponding document in the training data set according to the coreference score includes: For any event mention, obtaining event mentions with the same event subtype as it as candidate coreference mention pairs; Determining the coreference event of the event mention through a greedy algorithm according to the coreference scores of the candidate coreference mention pairs; Furthermore, determining the predicted event coreference chain of the corresponding document in the training data set.

8. An event coreference resolution device based on modeling arguments, characterized in that The device includes: A training data set acquisition module for obtaining a training data set for event coreference resolution to be performed; A training data input module for inputting the training data set into an event coreference resolution model; the event coreference resolution model includes an event extraction component, a mention encoder component, a coreference scorer component, and an event coreference chain determination module; the event extraction component is used to obtain a plurality of event mentions according to the document data in the training data set; each of the event mentions includes a trigger word, arguments, and an event subtype of the event; the arguments are divided into five roles: agent, patient, time, location, and others; the mention encoder component is used to obtain a trigger word representation of any event mention and an argument representation of the argument role according to the event mention and the token data of the corresponding document; wherein the argument representation includes an argument confidence score; further obtaining a trigger word pair representation and an argument pair representation of any two event mentions, and through a gating filtering mechanism, filtering the argument pair representation according to the trigger word pair representation to obtain a filtered argument pair representation, and then obtaining a mention pair representation of any two event mentions according to the trigger word pair representation and the filtered argument pair representation; the coreference scorer component is used to obtain a coreference score of any two event mentions according to the mention pair representation; the event coreference chain determination module is used to obtain a predicted event coreference chain of the corresponding document in the training data set according to the coreference score; A model training module for training the event coreference resolution model through the training data set and the predicted event coreference chain to obtain a trained event coreference resolution model; A model application module for inputting document data to be resolved for event coreference into the trained event coreference resolution model to obtain event coreference chain data corresponding to the document data.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Prompt-based event argument extraction method and system

    CN114880431A

  • Event argument extraction method and apparatus and electronic device

    US20210200947A1