Question-Answering Based Multi-Task Event Extraction Method and Device

Through a multi-task event extraction method based on question and answer, special problem templates are generated for different models, trigger word information and event types are introduced, which solves the problems of error propagation, nested entity processing and trigger word length limitation in the existing technology, and achieves a more efficient event information extraction effect.

CN115563253BActive Publication Date: 2025-06-27SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211079899.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2025-06-27
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

The existing event extraction model has problems such as error propagation, lack of nested entity processing and trigger word length limitation, making it difficult to effectively process complex event information.

Method used

A multi-task event extraction method based on question and answer is adopted. By generating special problem templates for different models, introducing trigger word information and event types, gradually filtering and extracting event information, reducing error propagation, and supporting nested entity processing.

Benefits of technology

It effectively reduces the negative impact of error propagation on model performance, improves the processing ability of complex event information, especially on the ACE2005 dataset, which has significantly improved performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115563253B_ABST
    Figure CN115563253B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-task event extraction method and device based on question and answer. The method includes: obtaining a first input vector; inputting the first input vector into a trigger word extraction model to obtain the position information of the trigger word in the original text; obtaining a second input vector; inputting the second input vector into an event recognition model to screen out event samples with correct trigger words and obtain the event types to which the event samples belong; generating questions for a specified argument role type according to a question template capable of introducing trigger word information and event types, and obtaining a third input vector; inputting the third input vector into an argument role extraction model to screen out event samples with correct trigger word extraction and event recognition results and obtain the argument positions of the specified argument roles. The present invention reduces the negative impact on the model performance caused by error propagation by introducing an auxiliary screening task. The present invention can be widely applied to the field of natural language processing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a multi-task event extraction method and device based on question and answer. Background Art

[0002] Event extraction is a classic information extraction task in the field of NLP and is also the basis for practical applications such as information retrieval, intelligent question and answer, and knowledge graph construction. In the era of information explosion, humans cannot retrieve a large amount of unstructured data at a sufficiently fast speed and with stable quality, and it is impossible to let all information rely on manual processing. In the field of knowledge graphs, the Automatic Content Extraction (ACE) evaluation conference defines an event as something or a change in state that occurs at a specific time point or time period and within a specific geographical area and consists of one or more actions participated by one or more roles. The event extraction task studies how to extract event information of interest to users from text describing event information and present it in a structured form.

[0003] Event extraction technology extracts events of interest to users from unstructured information and presents them in a structured manner to users. The event extraction task can be decomposed into 4 sub-tasks: trigger word extraction, event recognition, argument extraction, and role classification task. Among them, argument extraction and role classification can be combined into an argument role extraction task. Event recognition determines the event type to which each word in a sentence belongs and is a multi-classification task based on words. The argument role extraction task is a multi-classification task based on word pairs, which determines the role relationship between any pair of trigger words and entities in a sentence.

[0004] Existing event extraction models have the following problems: 1) Existing work based on the pipeline model often has the problem of error propagation, and the output results of the previous models have a certain negative impact on the performance of the subsequent models; 2) In the event extraction task on the ACE2005 dataset, there is little research on nested entities; 3) A large amount of existing work directly processes trigger words as having a length of 1, but there are some trigger words with a length not equal to 1 in the ACE2005 dataset. Summary of the Invention

[0005] To at least to some extent solve one of the technical problems existing in the prior art, an object of the present invention is to provide a multi-task event extraction method and device based on question and answer.

[0006] The technical solution adopted by the present invention is:

[0007] A multi-task event extraction method based on question and answer, comprising the following steps:

[0008] Generate a question according to a question template for trigger word extraction and obtain a first input vector;

[0009] Input the first input vector into the trigger word extraction model to obtain the position information of the trigger word in the original text;

[0010] Generate a question according to the question template that can introduce trigger word information, and obtain the second input vector;

[0011] Input the second input vector into the event recognition model, filter out the event samples with correct trigger words, and obtain the event types to which the event samples belong;

[0012] Generate a question for the specified argument role type according to the question template that can introduce trigger word information and event types, and obtain the third input vector;

[0013] Input the third input vector into the argument role extraction model, filter out the event samples with correct trigger word extraction and event recognition results, and obtain the argument positions of the specified argument roles.

[0014] Further, the generating a question according to the question template for trigger word extraction and obtaining the first input vector includes:

[0015] Separate the original text by spaces and represent it as where m is the number of words in the original text;

[0016] Generate a question according to the question template for trigger word extraction to obtain the input text, which is represented as

[0017] Tokenize the input text by the tokenization method provided by the pre-trained model BERT, convert the tokenized input text into a real-valued vector representation to obtain word vectors

[0018] Obtain the corresponding block encoding according to the input text, which is represented as B a ={0,0,0,11,12,…,1 m ,1}; where the subscript of 1 corresponds to the subscript in the input text T a in;

[0019] Convert the block encoding into a real-valued vector through the block vector matrix provided by the pre-trained model BERT to obtain block vectors

[0020] Obtain the absolute position information of each word in the input text according to the position encoding method provided by the pre-trained model BERT to obtain position vectors

[0021] The word vectors block vectors and position vectors The addition of three vectors yields the first input vector v of the trigger word extraction model a .

[0022] Furthermore, inputting the first input vector into the trigger word extraction model to obtain the position information of the trigger word in the original text includes:

[0023] Input the first input vector v a into the pre-trained model BERT to obtain the output of the last layer of the BERT network, followed by a fully connected layer, and then use the Softmax function to obtain the type probability corresponding to each word;

[0024] Extract the trigger word according to the obtained type probability, and each predicted trigger word is represented as where k represents the number of words of the trigger word;

[0025] Among them, the extraction of the trigger word according to the obtained type probability includes:

[0026] Fix four types, namely B, I, O, and PAD. The PAD type is used to construct the labels corresponding to the problem positions, CLS, and SEP positions during the training process;

[0027] During prediction, regard the PAD type as the O type, and use the BIO triple annotation method to decode each position of the original text to obtain the position information of the trigger word in the original text, represented as where represents the start position of the predicted trigger word in the original text, represents the end position of the predicted trigger word in the original text.

[0028] Furthermore, calculate the cross-entropy loss between the predicted word type probability and the original type label to obtain the loss of the trigger word extraction model, and use it to train and optimize the trigger word extraction model.

[0029] Furthermore, generating a question according to the question template that can introduce trigger word information and obtaining the second input vector includes:

[0030] Generate a question according to the question template that can introduce trigger word information to obtain the input text T b ;

[0031] Tokenize the input text T through the tokenization method provided by the pre-trained model BERT b to convert the tokenized input text into a real-valued vector representation to obtain the word vector

[0032] According to the input text T b obtain the corresponding block encoding;

[0033] Convert the block encoding into a real-valued vector through the block vector matrix provided by the pre-trained model BERT to obtain the block vector

[0034] Obtain the absolute position information of each word in the input text T according to the position encoding method provided by the pre-trained model BERT to obtain the position vector b in it, to obtain the position vector

[0035] The word vector The block vector And the position vector Add the three vectors to obtain the second input vector v of the event recognition model b ;

[0036] Among them, the input text T b is expressed as

[0037] Furthermore, input the second input vector into the event recognition model, screen out the event samples with correct trigger words, and obtain the event types to which the event samples belong, including:

[0038] Input the second input vector v b into the pre-trained model BERT, obtain the output at the CLS position of the BERT network, followed by a classification fully connected layer, and then use the Softmax function to obtain the probability of the event type to which the sentence belongs, to obtain the event samples with correct trigger words and their event types, and the predicted event type name is expressed as where n represents the number of words of the event type name

[0039] Furthermore, calculate the cross-entropy loss between the predicted event type probability and the original type label to obtain the loss of the event recognition model, and use it to train and optimize the event recognition model

[0040] Furthermore, there are 34 fixed event types in total. The first 33 correspond to the 33 event types of the ACE2005 dataset, and the last one is the None type, indicating that the event sample is a sample with an incorrect trigger word and needs to be discarded

[0041] Furthermore, generate questions for a specified argument role type according to the question template that can introduce trigger word information and event types, and obtain the third input vector, including:

[0042] According to the obtained event type E p Obtain all the argument role types under this event, where each argument role name is expressed as

[0043] Generate questions for the specified argument role according to the question template to obtain the input text Tc ;

[0044] Tokenize the input text T using the tokenization method provided by the pre-trained model BERT c to obtain token vectors by converting the tokenized input text into a real-valued vector representation

[0045] Obtain the corresponding block encoding according to the input text T c ;

[0046] Convert the block encoding into a real-valued vector through the block vector matrix provided by the pre-trained model BERT to obtain block vectors

[0047] Obtain the absolute position information of each word in the input text T according to the position encoding method provided by the pre-trained model BERT to obtain position vectors c ;

[0048] Add the token vectors block vectors and position vectors to get the third input vector v of the event recognition model c ;

[0049] where the input text T c is represented as

[0050] Furthermore, inputting the third input vector into the argument role extraction model, screening out the event samples with correct trigger word extraction and event recognition results, and obtaining the argument positions of the specified argument roles includes:

[0051] Input the third input vector v c into the pre-trained model BERT to obtain the output O at the CLS position of the BERT network 1 and the output O of the last layer 2 , O 1 is followed by a screening fully-connected layer, O 2 is followed by a classification fully-connected layer, and then use the Softmax function to obtain the error probability of the event sample and the type probability corresponding to each word;

[0052] According to the error probability of the event sample and the type probability corresponding to each word, obtain the correct event sample and the arguments of the specified argument roles in the event.

[0053] Furthermore, the obtaining the correct event sample and the arguments of the specified argument roles in the event according to the error probability of the event sample and the type probability corresponding to each word includes:

[0054] There are five fixed types, namely B, I, O, BandI, and PAD. During the training process, the PAD type is used to construct the labels corresponding to the problem positions, CLS, and SEP positions.

[0055] During prediction, the PAD type is regarded as the O type. Based on the BIO triple annotation method, the BandI type is added to solve the nested entity problem. By decoding each position of the original text, the position information of the argument is obtained, expressed as where represents the start position of the predicted argument in the original text, represents the end position of the predicted argument in the original text.

[0056] Furthermore, the training model is optimized by obtaining the loss, including:

[0057] Calculating the cross-entropy loss between the predicted sample error probability and the original label to obtain the first loss;

[0058] Calculating the cross-entropy loss between the predicted word type probability and the original type label to obtain the second loss;

[0059] Adding the first loss and the second loss, and jointly training and optimizing the argument role extraction model according to the obtained loss.

[0060] Furthermore, based on the BIO triple annotation method, the BandI type is added. For the BandI type, it means that the word represented by the current position serves both as an independent word of length one and as part of the previously matched word.

[0061] Furthermore, the trigger word extraction model, event recognition model, and argument role extraction model are all optimized using the Adam algorithm, and an early stopping training method is used to prevent overfitting.

[0062] Another technical solution adopted by the present invention is:

[0063] A multi-task event extraction device based on question answering, including:

[0064] At least one processor;

[0065] At least one memory for storing at least one program;

[0066] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0067] The beneficial effects of the present invention are as follows: By introducing an auxiliary screening task, the present invention reduces the negative impact of error propagation on the model performance; for the selection of problem templates, dedicated solutions for different models are determined, and good results can be achieved. Description of the Drawings

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings related to the technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0069] Figure 1 is a flowchart of the steps of a multi-task event extraction method based on question and answer in an embodiment of the present invention;

[0070] Figure 2 is a structural diagram of a multi-task event extraction model based on question and answer in an embodiment of the present invention. Detailed Embodiments

[0071] The embodiments of the present invention are described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adjusted adaptively according to the understanding of those skilled in the art.

[0072] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0073] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is two or more, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0074] In the description of the present invention, unless otherwise clearly defined, terms such as "set", "install", "connect", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution.

[0075] As Figure 1 shown, this embodiment provides a multi-task event extraction method based on question and answer, including the following steps:

[0076] S1. Generate a question according to the question template extracted for the trigger word, and obtain the first input vector.

[0077] As Figure 2 shown, for the input part of the trigger word extraction model, the present invention splices the word "action" and the original text, introduces the question intention, and uses "SEP" as the separator between "action" and the original text, and marks the beginning and end with "CLS" and "SEP", expressed as Taking the event sample D as an example, the original text is "He visited all his friends", the trigger word is "visited", the event type is the Meet type, and the corresponding argument roles are "Individual" and "Group", corresponding to "He" and "all his friends" in the original text respectively. According to the original text of the event sample D, obtain T a ={CLS, action, SEP, He, visited, all, his, friends, SEP}. In this step, both the word "action" and the original text need to be tokenized using the WordPiece method provided by BERT, and then spliced with "CLS" and "SEP" to obtain the tokenized input sequence, which is converted into the corresponding one-hot vector, denoted as where N represents the length of the input sequence, and |V| represents the size of the vocabulary. Therefore, the word vector representation corresponding to the input sequence is where W t ∈R |V|×e , is the word vector matrix provided by BERT, and e represents the dimension of the word vector. Similarly, according to T a obtain the corresponding block sequence representation as: B a ={0, 0, 0, 11, 12,…, 1 m , 1}. Among them, the subscript of 1 corresponds to the subscript in the input text, m corresponds to the length of the input sequence, and m is 5 in the event sample D. Convert the block sequence into the corresponding block encoding, and through the vector matrix W s ∈R |s|×e convert the block encoding into a real-valued vector to obtain the block vector where |S| represents the number of blocks; use the position vector matrix Perform one-hot encoding on the position Convert it into a real-valued vector to obtain the position vector Word vector Block vector And the position vector The dimensions of all three vectors are (BS, N, e). Among them, BS represents the batch size of training. Corresponding to the BERT-Base model, the dimension of the input vector obtained from the event sample D is (BS, 9, 768). The three vectors are added together to obtain the input vector of the trigger word extraction model, denoted as v a .

[0078] S2. Input the first input vector into the trigger word extraction model to obtain the position information of the trigger word in the original text.

[0079] As an alternative implementation, input the first input vector v a into the trigger word extraction model, followed by a fully connected layer to convert the vector into a new vector with a dimension of (BS, N, 4), and then use the Softmax function for calculation to obtain the probabilities that each word in the original text is of type B, type I, type O, and type PAD respectively. During prediction, the PAD type is processed as the O type, and the BIO triple annotation method is used for the extraction of trigger words. For the event sample D, it is predicted that the probability that the position of "visited" is of type B is the largest, and the probabilities that the other positions in the original text are of type O are the largest. Therefore, the position information of the trigger word is extracted Among them represents the start position of the predicted trigger word in the original text represents the end position of the predicted trigger word in the original text. In the sample D, it is represented as Therefore, it is predicted that the trigger word is "visited". During the training process, PAD is used to construct the labels corresponding to the problem positions, CLS, and SEP positions. For the event sample D, the constructed label sequence is label = {PAD, PAD, PAD, O, B, O, O, PAD}. Calculate the cross-entropy loss between the predicted word type probabilities and the original type labels to obtain the loss of the trigger word extraction model, and use it to train and optimize the trigger word extraction model

[0080] S3. Generate a question according to the question template that can introduce trigger word information, and obtain the second input vector

[0081] As Figure 2 shown, for the input part of the event recognition model, the trigger word text information and trigger word position information of the previous model are introduced. Taking the event sample D as an example, construct the corresponding input text according to the question template, denoted as T b={CLS,The,trigger,word,visited,pos,1,1,pos,SEP,He,visited,all,his,friends,SEP}. In a similar manner to step S2, the input text Tb is subjected to a series of transformations to obtain a word vector block vector and a position vector The sum of these three vectors gives the input vector of the event recognition model, denoted as the second input vector v b .

[0082] S4. Input the second input vector into the event recognition model, filter out the event samples with correct trigger words, and obtain the event types to which the event samples belong.

[0083] As an alternative implementation, input the second input vector v b into the trigger word extraction model, followed by a fully connected layer to convert the vector into a new vector with a dimension of (BS, N, 34), and then use the Softmax function for calculation to obtain the probabilities of each word in the original text being the 33 event types and the None type in the ACE2005 dataset. During prediction, for the event sample D, the probability of predicting the event type as the Meet type is the highest. During the training process, the output result of the previous model will be combined for training. If the result of extracting the trigger word for a certain sample in the previous model is incorrect, the label of this sample is set to None, and the labels of the remaining correct samples are set to the types belonging to the 33 event types in the ACE2005 dataset originally. Calculate the cross-entropy loss between the predicted event type probabilities and the original type labels to obtain the loss of the event recognition model, and use it to train and optimize the event recognition model.

[0084] S5. Generate questions for a specified argument role type according to a question template that can introduce trigger word information and event types, and obtain the third input vector.

[0085] As Figure 2 shown, for the input part of the argument role extraction model, the text information and position information of the trigger word, as well as the corresponding event type, are introduced. For each type belonging to the 33 event types in the ACE2005 dataset, there is a corresponding argument role group. Taking the event sample D as an example, the event of the Meet type corresponds to two argument roles, namely the Individual role and the Group role. Construct two input texts corresponding to different argument roles according to the question template, denoted as In a similar manner to step S2, the input text and A series of transformations are performed to obtain the corresponding word vectors, block vectors, and position vectors respectively. The sum of the three vectors yields two input vectors for the argument role extraction model, denoted as and

[0086] S6. The third input vector is input into the argument role extraction model to filter out event samples where both the trigger word extraction and event recognition results are correct, and the argument positions of the specified argument roles are obtained.

[0087] As an alternative implementation, and are respectively input into the argument role extraction model. Then, a fully connected layer is connected to convert the vectors into new vectors with a dimension of (BS, N, 5). Subsequently, the Softmax function is used for calculation to obtain the probabilities that each word in the original text belongs to the B type, I type, O type, BandI type, and PAD type respectively. During prediction, the PAD type is processed as the O type, and a new annotation method is adopted for argument extraction to ensure support for nested argument extraction. Specifically, for the BandI type, it means that the word represented by the current position serves both as an independent word of length one and as part of the previously matched word. For the event sample D, the predicted result shows that the probability that the position of "He" is of the B type is the highest, and the probabilities that the other positions are of the O type are the highest. The predicted result shows that the probabilities that the three positions of "all his friends" are of the B type, I type, and I type respectively are the highest, and the probabilities that the other positions are of the O type are the highest. That is, the argument corresponding to the argument role Individual is "He", and the argument corresponding to the argument role Group is "all his friends". During the training process, PAD is used to construct the labels corresponding to the problem positions, CLS, and SEP positions. For the event sample D, the constructed label sequence is label1 = {PAD, PAD, PAD, PAD, PAD, PAD, PAD, PAD, PAD, PAD, PAD, B, O, O, O, O, PAD}, the constructed label sequence is: label2 = {PAD, PAD, PAD, PAD, PAD, PAD, PAD, PAD, PAD, PAD, PAD, O, O, B, I, I, PAD}. The cross-entropy loss is calculated between the predicted word type probabilities and the original type labels to obtain the loss of the trigger word extraction model, and the trigger word extraction model is trained and optimized.

[0088] In summary, compared with the prior art, the present embodiment has the following advantages and beneficial effects: The present invention proposes to reduce the negative impact of error propagation on model performance by introducing a specific auxiliary screening task and utilizing the dependency relationship between pipeline models. For the selection of problem templates, the present invention determines dedicated solutions for different models and achieves good results. For nested entity problems, the present invention proposes a new encoding method to achieve the extraction of argument roles of nested structures. Experiments prove that it performs very well in the ACE2005 dataset.

[0089] The present embodiment further provides a question-based multi-task event extraction device, including:

[0090] At least one processor;

[0091] At least one memory for storing at least one program;

[0092] When the at least one program is executed by the at least one processor, the at least one processor is caused to implement Figure 1 The method shown.

[0093] A question-based multi-task event extraction device according to the present embodiment can execute a question-based multi-task event extraction method provided by an embodiment of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0094] The embodiment of the present application also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 The method shown.

[0095] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutive blocks shown can actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are foreseeable, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.

[0096] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of such modules would be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0097] If the described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0098] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device). For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0099] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0100] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0101] In the foregoing description of the present specification, descriptions with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0102] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0103] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A multi-task event extraction method based on question answering, characterized in that It includes the following steps: Generate a question according to the question template for trigger word extraction, and obtain a first input vector; Input the first input vector into the trigger word extraction model to obtain the position information of the trigger word in the original text; Generate a question according to the question template that can introduce trigger word information, and obtain a second input vector; Input the second input vector into the event recognition model, filter out the event samples with correct trigger words, and obtain the event types to which the event samples belong; Generate a question for a specified argument role type according to the question template that can introduce trigger word information and event types, and obtain a third input vector; Input the third input vector into the argument role extraction model, filter out the event samples with correct trigger word extraction and event recognition results, and obtain the argument positions of the specified argument role; The step of generating a question for a specified argument role type according to the question template that can introduce trigger word information and event types, and obtaining a third input vector includes: According to the obtained event type E p Obtain all the argument role types under this event, where each argument role name is represented as Generate questions for the specified argument role according to the question template to obtain the input text T c ; Tokenize the input text T using the tokenization method provided by the pre-trained model BERT c and convert the tokenized input text into a real-valued vector representation to obtain word vectors According to the input text T c obtain the corresponding block code; Convert the block encoding into a real-valued vector through the block vector matrix provided by the pre-trained model BERT to obtain the block vector Obtain the absolute position information of each word in the input text T according to the position encoding method provided by the pre-trained model BERT, and obtain the position vector c ​ The word vector block vector and the position vector are added together to obtain the third input vector v of the event recognition model c ; where the input text T c is expressed as The step of inputting the third input vector into the argument role extraction model, filtering out the event samples with correct trigger word extraction and event recognition results, and obtaining the argument positions of the specified argument role includes: Input the third input vector v c into the pre-trained model BERT to obtain the output O at the CLS position of the BERT network 1 and the output O of the last layer 2 , O 1 is followed by a screening fully connected layer, and O 2 is followed by a classification fully connected layer, and then the Softmax function is used to obtain the error probability of the event sample and the type probability corresponding to each word; Obtain the correct event sample and the argument of the specified argument role in the event according to the error probability of the event sample and the type probability corresponding to each word.

2. The multi-task event extraction method based on question answering according to claim 1, wherein The step of generating a question according to the question template for trigger word extraction, and obtaining a first input vector includes: Separate the original text by spaces and represent it as where m is the number of words in the original text; Generate a question according to the question template for trigger word extraction to obtain the input text, denoted as Tokenize the input text using the tokenization method provided by the pre-trained model BERT, and convert the tokenized input text into a real-valued vector representation to obtain word vectors Obtain the corresponding block code according to the input text, denoted as B a ={0, 0, 0, 11, 12, …, 1 m , 1}; where the subscript of 1 corresponds to the subscript in the input text T a in the subscript; Convert the block encoding into a real-valued vector through the block vector matrix provided by the pre-trained model BERT to obtain the block vector Obtain the absolute position information of each word in the input text according to the position encoding method provided by the pre-trained model BERT to obtain the position vector The word vector block vector and the position vector Add the three vectors to obtain the first input vector v of the trigger word extraction model a .

3. A method for multi-task event extraction based on question answering according to claim 2, characterized in that The step of inputting the first input vector into the trigger word extraction model to obtain the position information of the trigger word in the original text includes: Input the first input vector v a into the pre-trained model BERT to obtain the output of the last layer of the BERT network, followed by a fully connected layer, and then use the Softmax function to obtain the type probability corresponding to each word; Extract the trigger words according to the obtained type probabilities, and each predicted trigger word is represented as where k represents the number of words in the trigger word; Among them, the step of extracting the trigger word according to the obtained type probability includes: Fix four types, namely B, I, O, and PAD. The PAD type is used to construct the labels corresponding to the question position, CLS, and SEP positions during the training process; When making predictions, the PAD type is regarded as the O type, and the BIO triple annotation method is used to decode each position of the original text to obtain the position information of the trigger word in the original text, expressed as where represents the start position of the predicted trigger word in the original text, represents the end position of the predicted trigger word in the original text.

4. A multi-task event extraction method based on question and answer according to claim 1, characterized in that, The step of generating a question according to the question template that can introduce trigger word information, and obtaining a second input vector includes: Generate a question based on a question template that can introduce trigger word information to obtain the input text T b ; Tokenize the input text T using the tokenization method provided by the pre-trained model BERT b and convert the tokenized input text into a real-valued vector representation to obtain word vectors According to the input text T b Obtain the corresponding block code; Convert the block encoding into a real-valued vector through the block vector matrix provided by the pre-trained model BERT to obtain the block vector Obtain the absolute position information of each word in the input text T according to the position encoding method provided by the pre-trained model BERT, and obtain the position vector b ​ The word vector block vector and the position vector are added together to obtain the second input vector v of the event recognition model b ; among them, the input text T b is expressed as 5. A multi-task event extraction method based on question answering according to claim 4, characterized in that The step of inputting the second input vector into the event recognition model, filtering out the event samples with correct trigger words, and obtaining the event types to which the event samples belong includes: Input the second input vector v b into the pre-trained model BERT, obtain the output at the CLS position of the BERT network, followed by a classification fully-connected layer, and then use the Softmax function to obtain the probability of the event type to which the sentence belongs, obtain the correct event sample of the trigger word and its event type, and the predicted event type name is denoted as where n represents the number of words in the event type name.

6. A multi-task event extraction method based on question and answer according to claim 1, characterized in that The step of obtaining the correct event sample and the argument of the specified argument role in the event according to the error probability of the event sample and the type probability corresponding to each word includes: Fix five types, namely B, I, O, BandI, and PAD. The PAD type is used to construct the labels corresponding to the question position, CLS, and SEP positions during the training process; When making predictions, the PAD type is regarded as the O type. On the basis of the BIO triple annotation method, the BandI type is added to solve the problem of nested entities. By decoding each position of the original text, the position information of the argument is obtained, expressed as where represents the start position of the predicted argument in the original text, represents the end position of the predicted argument in the original text.

7. A method for multi-task event extraction based on question answering according to claim 1, characterized in that, The trigger word extraction model, event recognition model, and argument role extraction model are all optimized using the Adam algorithm, and an early stopping training method is used to prevent overfitting.

8. A multi-task event extraction device based on question answering, characterized in that, It includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Intelligent automated assistant

    CN102792320A

  • Event extraction method and device and storage medium

    CN114936563A