A Construction Method for a Pipeline Multi-Event Extraction Model
By building a pipeline-based multi-event extraction model, using T5 model to train and gradually build a predicted sample set, the problem of low accuracy of multi-event extraction and overlapping multi-event extraction in the existing technology is solved, and event extraction with high accuracy and loyalty is achieved.
Patent Information
- Application Number
- CN202211733205.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In the multi-event, overlapping multi-event extraction tasks, the extraction accuracy rate of the existing event extraction model is not high due to identification missing and event elements cannot be matched.
The construction method of the pipeline multi-event extraction model is adopted. By obtaining the marked text data, an event feature data set is constructed, and a training set of positive and negative samples containing event types and event elements is constructed. The T5 model is used for training, and the prediction sample set of each step is gradually constructed, and the final extraction result is integrated.
The model's understanding and prediction ability of multiple events is improved, the extraction accuracy and loyalty is ensured, and the identification missing and event elements cannot be matched when extracting multiple events and overlapping multiple events is solved.
Smart Images

Figure CN116028812B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method for constructing a pipeline multi-event extraction model. Background Art
[0002] Event extraction (EE) is one of the important tasks in the field of natural language processing (NLP). The purpose of event extraction is to identify the event type, event trigger, event argument, and argument role contained in a given corpus. At present, the application scenarios of event extraction technology are very extensive, which can efficiently extract useful information from a large amount of text and provide strong data support for the construction of knowledge graphs.
[0003] The extraction methods of existing mainstream event extraction models include sequence labeling method, pointer discriminant method, and generative method. The sequence labeling method is essentially a multi-label multi-classification method, which predicts possible labels for each token; the pointer discriminant method extracts events by predicting the start and end positions of the text corresponding to each label; the generative method is an end-to-end method, which extracts context information through a deeper network and directly outputs event information in text format. The above three methods have good performance in single-event or non-overlapping multi-event corpora with a small number of events. However, when there are many events in the corpus, especially when one or several elements overlap, it is very easy to have problems such as missing recognition, misrecognition, and inability to match event elements, resulting in very low accuracy. Since overlapping multi-events are common in actual corpora, there is an urgent need for a more optimized method for constructing a multi-event extraction model to solve the problem of low extraction accuracy caused by missing recognition and inability to match event elements in the event extraction model of the prior art in multi-event and overlapping multi-event extraction tasks. Summary of the Invention
[0004] In view of the above analysis, an embodiment of the present invention aims to provide a method for constructing a pipeline multi-event extraction model to solve the problem of low extraction accuracy caused by missing recognition and inability to match event elements in the event extraction model of the prior art in multi-event and overlapping multi-event extraction tasks.
[0005] On the one hand, an embodiment of the present invention provides a method for constructing a pipeline multi-event extraction model, including the following steps:
[0006] Obtain the labeled text data as the original data set;
[0007] Obtain a set of event feature data based on the original data set, and further construct a positive sample data set D of event types+1 , positive sample data set D of event elements +2 , all-negative sample data set D of event types -1 and random negative sample data set D of event elements -2 , and finally obtain the model training data set D all ;
[0008] Use the training data set D all to train the T5 model and obtain the trained pipeline multi-event extraction model M trained ;
[0009] During multi-event extraction, gradually construct the prediction sample set for each step, and the trained model M trained is used to obtain the prediction result for each step based on the prediction sample set for each step, and integrate to obtain the final extraction result.
[0010] Furthermore, the acquisition of the labeled text data includes:
[0011] Acquire the original text data;
[0012] Annotate the original text data; where the annotation includes: determining the event types included in the sentences in the text data; extracting trigger words, event elements and their positions according to the event types; and assigning appropriate event role labels to the event elements.
[0013] Furthermore, the event feature data set includes:
[0014] The correspondence schema between event types and all event roles, the corresponding set S between event types and a single event role type_role , the set S of all event types type , the set S of all trigger words trigger and the set S of all event elements argument ; where the schema records all event types in the original data set and all event roles corresponding to them respectively; S type_role is obtained according to the schema, including the pairwise combinations of the event types and all event roles of each event in the schema, and which event role this event role belongs to in the schema; S type records all event types; S trigger records all trigger words that appear in the original data set; S argument records all event elements included in the original data set.
[0015] Furthermore, the model training data set D all is constructed through the following steps:
[0016] Summarize and organize the annotation information of the original dataset to obtain the correspondence schema between event types and all event roles, the corresponding set S of event types and a single event role, type_role and the set S of all event types type Three sets of event feature data;
[0017] Use the original dataset and the dataset schema to construct the positive sample dataset D of event types +1 and the positive sample dataset D of event elements +2 , as well as the set S of all trigger words that appear in the original dataset trigger and the set S of all event elements argument Two sets of event feature data;
[0018] Use the positive sample dataset D of event types +1 and the event type dataset S type to construct the all-negative sample dataset D of event types -1 ;
[0019] Use the positive sample dataset D of event elements +2 , the set S of trigger words trigger , the set S of event elements argument and the corresponding set S of event types and a single event role type_role to construct the random negative sample dataset D of event elements -2 ;
[0020] Mix and shuffle D +1 , D +2 , D -1 , D -2 to finally obtain the model training dataset D all .
[0021] Furthermore, the positive sample dataset D of event types +1 and the positive sample dataset D of event elements +2 are constructed through the following steps:
[0022] A1. Extract the event type e corresponding to a certain event contained in the text data text_p of the original dataset type , the trigger word w trigger , the event roles e role_1 ~e role_n , and the corresponding event elements w arg_1 ~w arg_n (n is the number of event roles included in this event, and is also equal to the number of event elements); the input for constructing the positive sample of the event type of this event is text_p + e type + "trigger word", and the output is w trigger; The input for constructing the positive samples of event elements for this event is text_p + prompt arg , and the output is w arg_1 ~w arg_n ; Among them, the event element prompt arg can be obtained by the following formula:
[0023]
[0024] A2. For each event in the text data text_p, use the method in (1) to construct positive samples of event types and positive samples of event elements, and obtain the positive sample dataset D +1 of event types and the positive sample dataset D +2 of event elements;
[0025] Further, the all-negative sample dataset D -1 of event types is constructed through the following steps:
[0026] B1. Replace the e type of a positive sample of a certain event type with other event types of this event in the event type dataset S type in turn, and the target output is empty, to obtain the all-negative sample of the event type of this event;
[0027] B2. For all events in the positive sample dataset D +1 of event types, use the method in (1) to construct the all-negative sample dataset D -1 of event types.
[0028] Further, the random negative sample dataset D -2 of event elements is constructed through the following steps:
[0029] (1) Find all positive samples of event elements of a certain event in D +2 , and find all event element prompts prompt arg from the positive samples of event elements, and form a set S prompt ;
[0030] (2) Randomly select a trigger word from S trigger to obtain w trigger_random ; Randomly select an element from S type_role to obtain an event type e type_random , an event role e role_random and the position p where the event role is located;
[0031] (3) Randomly select p event elements from the event element set S argument to obtain w arg_r_1 ~w arg_r_p, combine to obtain the event element random prompt in the following format arg_random ;
[0032] prompt arg_random = e type_random + w trigger_random + w arg_r_1 +…+ w arg_r_p + e role_random
[0033] (4) Judge whether prompt arg_random exists in S prompt . If it exists, repeat steps 2, 3, and 4. If it does not exist, use prompt arg_random to construct a negative sample, and add prompt arg_random to S prompt ;
[0034] (5) Repeat steps (1) to (4) until 5n event element random negative samples are obtained.
[0035] (6) For all event samples in D +2 , use the method in (1) to (5) to construct an event element random negative sample dataset D -2 .
[0036] Furthermore, the training of the T5 model includes:
[0037] Divide the model training dataset D all into a training set D train , a validation set D eval and a test set D test in a certain proportion;
[0038] Use the training set D train to fine-tune and train the T5 model for n rounds. After each round of training, use the validation set D eval for validation. Select the model of the round with the best validation set result as the final model, and use the test set D test for testing, and finally obtain the trained model M trained ;
[0039] During the training process, use the following formula to calculate the model loss and update the parameters:
[0040] Loss = CrossEntropy(x pred , x gold )
[0041] where x pred is the prediction result, and x gold is the target output.
[0042] Further, the step-by-step construction of the prediction sample set for each step includes:
[0043] Construct the first-step prediction sample set D based on the text to be extracted text and the event feature data set step_1 ;
[0044] Based on the text to be extracted text, the event feature data set and the prediction result of the previous-step model M trained Construct the prompt information prompt, and construct the prediction sample set of the next-step model in the structure of text+prompt, so as to realize the step-by-step construction of the 2nd to n+1th step prediction sample sets D step_2 ~D step_(n+1) .
[0045] Furthermore, the trained model M trained is used to obtain the prediction result of each step based on the prediction sample set of each step, and integrate to obtain the final extraction result, including:
[0046] Input D step_1 into the model M trained , and obtain all the trigger words p included in the first-step prediction result text trigger ;
[0047] Construct the 2nd to n+1th step prediction sample sets D step_2 ~D step_(n+1) in the format of text+prompt_x, input D step_x into the model M trained , and obtain the x-1th event role corresponding to each trigger word and the x-1th event element corresponding to the event type where x∈[2,n+1]; where prompt_x is expressed as:
[0048]
[0049] Combine the prompt information of the last step with the extraction result to obtain the complete event.
[0050] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:
[0051] 1. By constructing an event feature data set based on the original data set, and further constructing a training set including positive and negative samples of event types and event elements, the T5 model is trained using the training set, enabling the model to effectively learn the internal relationships among various event types, event roles, event elements, and trigger words. In particular, it improves the model's ability to understand and predict multiple events. The overall training process uses the method of prompt information, which to a certain extent ensures the extraction accuracy and loyalty, and an event extraction model with a high recognition rate for event texts is obtained.
[0052] 2. Based on the trained model, event texts are extracted. Events can be extracted in a progressive manner by using prompt information. All event types are used as prompt information to extract the corresponding trigger words, and then the trigger words and the element roles to be extracted are sequentially added step by step to prompt the extraction of event elements. After all the event elements included in the event type are extracted, the prompt information of the last step is combined with the extraction result to obtain a complete event. This pipeline-based extraction method provides a separate extraction path for each possible event, and focuses on solving the problems of missing recognition and unmatched event elements during the extraction of multiple events and overlapping multiple events, greatly improving the extraction accuracy.
[0053] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can be made obvious from the description, or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the content specifically pointed out in the description and the drawings. Brief Description of the Drawings
[0054] The drawings are only for the purpose of showing specific embodiments, and are not considered as a limitation to the present invention. Throughout the drawings, the same reference signs represent the same components.
[0055] Figure 1 It is a schematic flow chart of the construction method of the pipeline-based multi-event extraction model according to the embodiment of the present invention;
[0056] Figure 2 It is a schematic overall implementation flow chart of the construction method of the pipeline-based multi-event extraction model according to the embodiment of the present invention including actual prediction;
[0057] Figure 3 It is a schematic flow chart of the construction of training data provided by the embodiment of the present invention;
[0058] Figure 4 It is a schematic flow chart of obtaining the prediction result provided by the embodiment of the present invention. Detailed Description of the Embodiments
[0059] The preferred embodiments of the present invention will be specifically described below in conjunction with the accompanying drawings, where the accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0060] A specific embodiment of the present invention discloses a method for constructing a pipeline multi-event extraction model, as Figure 1 shown, including the following steps:
[0061] Step S110: Obtain the labeled text data as the original data set;
[0062] Step S120: Obtain the event feature data set based on the original data set, and further construct the positive sample data set D +1 of event types, the positive sample data set D +2 of event elements, the full negative sample data set D -1 of event types, and the random negative sample data set D -2 of event elements, and finally obtain the model training data set D all ;
[0063] Step S130: Use the training data set D all to train the T5 model to obtain the trained pipeline multi-event extraction model M trained ;
[0064] During multi-event extraction, gradually construct the prediction sample set for each step, and the trained model M trained is used to obtain the prediction result for each step based on the prediction sample set for each step, and the final extraction result is integrated.
[0065] The embodiment of the present invention uses a training set containing positive and negative samples of event types and event elements to train the T5 model to obtain a multi-event extraction model. By constructing an event feature data set based on the original data set, and further constructing a training set containing positive and negative samples of event types and event elements, and using the training set to train the T5 model, the model effectively learns the internal relationships between various event types, event roles, event elements, and trigger words, especially improving the model's understanding and prediction ability for multi-events. The overall training process uses the method of prompt information, which to a certain extent ensures the extraction accuracy and loyalty, and obtains an event extraction model with a high recognition rate for event texts.
[0066] Based on the above embodiments, specifically, the labeled text data in the above step S110 is obtained by the following method:
[0067] Directly use the event extraction data set of Baidu;
[0068] Annotate the original text data by itself; among them, the annotation method is as follows: determine the event types included in the sentences in the text data; extract the trigger words, event elements and their positions according to the event types; assign appropriate event role labels to the event elements.
[0069] Specifically, the above step S120 can also be optimized into the following steps:
[0070] Step S210: Summarize and organize the annotation information of the original data set to obtain the corresponding relationship schema between event types and all event roles, the corresponding set S between event types and a single event role type_role and the set S of all event types type Three event feature data sets;
[0071] Specifically, summarize and organize all event types and event roles in the original data set to construct the data set schema, S type_role and S type ; where the schema records all event types in the original data set and their corresponding all event roles; S type_role Obtained according to the schema, including the pairwise combinations of the event types and all event roles of each event in the schema, and which event role this event role belongs to in the schema; S type Records all event types; preferably, store the schema in the file json, S type_role and S type are both stored using sets.
[0072] Exemplarily, for an event with an event type of "acquisition" and event roles including "acquisition time, acquirer, acquiree", its records in the schema, S type_role and S type are shown in Table 1.
[0073] Table 1 Records of events of type "acquisition" in the schema, S type_role and S type Example records
[0074]
[0075] Step S220: Use the original data set and the data set schema to construct the positive sample data set D of event types +1 and the positive sample data set D of event elements +2 , as well as the set S of all trigger words that appear in the original data set trigger and the set S of all event elements argument Two event feature data sets;
[0076] Specifically, the constructed positive sample dataset D of event types +1 and the positive sample dataset D of event elements +2 , as well as all trigger word sets S that appear in the original dataset trigger and all event element sets S argument include:
[0077] (1) Extract the event type e corresponding to a certain event contained in the text data text_p of the original dataset type , trigger word w trigger , event role e role_1 ~e role_n , corresponding event element w arg_1 ~w arg_n (n is the number of event roles included in this event, and is also equal to the number of event elements); the input for constructing the positive sample of the event type of this event is text_p + e type + "trigger word", and the output is w trigger ; the input for constructing the positive sample of the event element of this event is text_p + prompt arg , and the output is w arg_1 ~w arg_n ; among them, the event element prompt arg can be obtained by the following formula:
[0078]
[0079] (2) For each event in the text data text_p, use the method in (1) to construct the positive sample of the event type and the positive sample of the event element, and obtain the positive sample dataset D of event types +1 and the positive sample dataset D of event elements +2 ;
[0080] (3) Save all the trigger words w trigger of all events in the text data text_p in the trigger word set S trigger , and save all event elements w arg_1 ~w arg_n in the event element set S argument , and obtain the trigger word dataset S trigger and the event element dataset S argument .
[0081] Exemplarily, for an event of the type "acquisition", the constructed positive samples of the event type and the event element, and the saving examples in S trigger and S argument are shown in Table 2.
[0082] Examples of positive samples of event types and positive samples of event elements for an event of type "acquisition" in Table 2, and examples of preservation in S trigger and S argument and examples of preservation in
[0083]
[0084]
[0085] It should be noted that in this example, the elements of the prompt information are separated by "-", and in fact, other symbols or spaces can also be used for separation. When constructing positive samples of event elements, the order of event elements in the prompt arg must be consistent with that recorded in the schema.
[0086] For events in complex situations, the output may be multiple event elements. When constructing the input, event element prompts need to be constructed separately. Exemplarily, Table 3 shows examples of positive samples of event elements for multiple events sharing a trigger word:
[0087] Table 3 Examples of positive samples of event elements for multiple events sharing a trigger word
[0088]
[0089] Step S230: Use the positive sample dataset D +1 of event types and the dataset S type of event types to construct a full negative sample dataset D -1 of event types;
[0090] Specifically, the construction of the full negative sample dataset D -1 of event types includes:
[0091] (1) Replace the e type of a positive sample of a certain event type with other event types of this event in the dataset S type of event types one by one. If the target output is empty for all replacements, a full negative sample of this event's event type is obtained;
[0092] (2) Apply the method in (1) to all events in the positive sample dataset D +1 of event types to construct the full negative sample dataset D -1 of event types.
[0093] Exemplarily, if there are m event types in the dataset S type of event types, then there are m - 1 full negative samples of the event type for each event;
[0094] For the training of the model, positive samples are samples with target output results, and negative samples are samples without output results. Adding negative samples during training can effectively improve the recognition accuracy of the model.
[0095] Step S240: Use the positive sample dataset D of event elements +2 , trigger word set S trigger , event element set S argument and the corresponding set S of event types and single event roles type_role to construct a random negative sample dataset D of event elements -2 ;
[0096] The input format of the random negative sample of event elements is the same as that of the positive sample of event elements. The difference lies in the different prompt information of the random negative sample of event elements, and the output results are all empty. For a certain event, the number of random negative samples of event elements is generally recommended to be 5 times that of the positive sample of event elements;
[0097] Specifically, the steps of constructing the random negative sample dataset D of event elements -2 are as follows:
[0098] (1) Find all positive samples of event elements of a certain event in D +2 , and find all event element prompts prompt arg from the positive samples of event elements to form a set S prompt ;
[0099] (2) Randomly select a trigger word from S trigger to get w trigger_random ; Randomly select an element from S type_role to get an event type e type_random , an event role e role_random and the position p where the event role is located;
[0100] (3) Randomly select p event elements from the event element set S argument to get w arg_r_1 ~ w arg_r_p , and combine them in the following format to get the random prompt prompt of event elements arg_random ;
[0101] prompt arg_random = e type_random + w trigger_random + w arg_r_1 + … + w arg_r_p + e role_random
[0102] (4) Judge whether prompt arg_random exists in S promptIf it exists, repeat steps 2, 3, and 4; if not, use the prompt arg_random Construct negative samples and add the prompt arg_random to S prompt ;
[0103] (5) Repeat steps (1) to (4) until 5n randomly generated negative samples of event elements are obtained.
[0104] (6) For all event samples in D +2 Use the methods in (1) to (5) to construct a dataset D of randomly generated negative samples of event elements -2 ;
[0105] Step S250: Mix and shuffle D +1 , D +2 , D -1 , D -2 to finally obtain the model training dataset D all ;
[0106] Specifically, the training of the T5 model in step S130 includes:
[0107] Divide the model training dataset D all into a training set D train , a validation set D eval and a test set D test in a certain proportion; preferably, the proportion is 8:1:1; use the training set D train to fine-tune the T5 model for n rounds. After each round of training, use the validation set D eval for validation. Select the model of the round with the best validation set result as the final model, and use the test set D test for testing to finally obtain the trained model M trained ; preferably, the number of training epochs n is 20;
[0108] Furthermore, during the training process, use the following formula to calculate the model loss and update the parameters:
[0109] Loss = CrossEntropy(x pred , x gold )
[0110] where x pred is the prediction result and x gold is the annotation.
[0111] Even further, use the trained model M trained ; the extraction of actual event texts includes the following steps:
[0112] Step S310: Obtain the text to be extracted, denoted as text. Among them, the text to be extracted, text, can be news text data crawled from websites.
[0113] Step S320: Based on the text to be extracted, text, the set of event feature data obtained from the original dataset, and the prediction result of the previous model M train Construct the prediction sample sets D step_1 ~D step_(n+1) from step 1 to step n + 1 in the structure of text+prompt, and input D step_1 ~D step_(n+1) into the model M step by step train to obtain the prediction results of the model M from step 1 to step n + 1. Here, n is the number of event roles of the event type corresponding to the prediction result of the first step. train Specifically, the construction of the prediction sample sets D
[0114] ~D step_1 ~D step_(n+1) and the obtaining of the prediction results of the model M from step 1 to step n + 1 include the following steps: train
[0115] (A) Traverse all event types e type in S type in sequence. For any event type add the sample to the prediction sample set D step1 of the first step: After the traversal ends, the number of samples in D step1 is m (m is the number of event types, k ∈ [1, m]).
[0116] (B) Input the prediction sample set D step1 of the first step into M trained . When a certain sample has an output result, its output result is the trigger word of the event type in the text to be extracted, text. Denote it as Look up the corresponding first event role in the schema and add the output result to the prediction sample set D _2 of the next step in the format of text+prompt step2 ; Among them
[0117] For samples without output results, it means that there is no trigger word of the input event type in the text text, that is, the text text does not contain an event with the event type .
[0118] (C) Take D step2 Input M trained and predict each trigger word corresponding to the first event role of the event elements, denoted as Determine the event type by checking the schema whether there are other event roles. If not, go to step S330;
[0119] If the event type has other event roles in the schema then for this event type of other event roles successively construct the next prediction sample set D step_3 ~D step_(n+1) and input D step_3 ~D step_(n+1) into the model M successively trained for event element extraction until all event elements corresponding to the included event roles are extracted by the model M trained and then go to step S330;
[0120] More specifically, the method for constructing the next prediction sample set D step_3 ~D step_(n+1) is as follows:
[0121] Construct the sample in the format text + prompt _X and add it to the next prediction sample set D step_(x) ; where is the prompt _(x-1) Based on this, replace with and add at the end, where x ∈ [3, n + 1] and n is the number of event roles included in this event type in the schema;
[0122] The event elements used in the prompt information prompt include and its determination method is as follows:
[0123]
[0124] where j ∈ [1, n - 1] and n is the number of event roles included in this event type in the schema; if contains multiple prediction results, separate the multiple results according to the format in this step to construct the prediction samples.
[0125] Step S330: Based on the n + 1th prediction sample set Dstep_(n+1) and the prediction result of the model M in the (n + 1)-th step train are integrated to obtain the final recognition result;
[0126] Specifically, the integration to obtain the final recognition result includes:
[0127] According to D step_n+1 and the n-th event element of the prediction result, the event extraction result is sorted out as:
[0128] Event type:
[0129] Trigger word:
[0130] Event role / argument:
[0131]
[0132] Exemplarily, the event extraction result can be integrated in the format of Table 4.
[0133] Table 4 Example of integrating event extraction results
[0134]
[0135] In summary, the beneficial effects of this embodiment are as follows:
[0136] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0137] 1. By constructing an event feature data set based on the original data set, and further constructing a training set including positive and negative samples of event types and event elements, and using the training set to train the T5 model, the model effectively learns the internal relationships among various event types, event roles, event elements, and trigger words. In particular, it improves the model's understanding and prediction ability for multiple events. The overall training process uses the method of prompt information, which to a certain extent ensures the extraction accuracy and loyalty, and obtains an event extraction model with a high recognition rate for event texts.
[0138] 2. Extract event texts based on the trained model. Events can be extracted in a progressive manner by using prompt information (prompt). All event types are used as prompt information to extract the corresponding trigger words, and then the trigger words and the elements to be extracted are added step by step to the prompt to extract event elements. After all event elements included in this event type are extracted, the prompt information of the last step and the extraction results are combined to obtain a complete event. This pipeline extraction method provides a separate extraction path for each possible event, focusing on solving the problems of missing recognition and unmatched event elements during the extraction of multiple events and overlapping multiple events, and greatly improving the extraction accuracy.
[0139] Those skilled in the art can understand that all or part of the processes of implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0140] As mentioned above, the above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for constructing a pipeline multi-event extraction model, characterized in that It includes the following steps: Obtain the labeled text data as the original data set; Obtain an event feature data set based on the original data set, and further construct a positive sample data set D of event types +1 , a positive sample data set D of event elements +2 , a full negative sample data set D of event types -1 and a random negative sample data set D of event elements -2 , and finally obtain a model training data set D all ; Use the training dataset D all Train the T5 model to obtain the trained pipeline multi-event extraction model M trained ; When performing multi-event extraction, gradually construct the prediction sample set for each step, and the trained model M trained is used to obtain the prediction result for each step based on the prediction sample set for each step, and the final extraction result is integrated; The model training dataset D all , is constructed through the following steps: Summarize and organize the annotation information of the original dataset to obtain the correspondence schema between event types and all event roles, the corresponding set S between event types and individual event roles type_role and the set S of all event types type Three sets of event feature data; Construct the positive sample dataset D of event types using the original dataset and the dataset schema +1 and the positive sample dataset D of event elements +2 , as well as the set S of all trigger words that appear in the original dataset trigger and the set S of all event elements argument Two sets of event feature data; Use the positive sample dataset D of event types +1 and the dataset S of event types type to construct the full negative sample dataset D of event types -1 ; Use the positive sample dataset D of event elements +2 , trigger word set S trigger , event element set S argument and the corresponding set S of event types and single event roles type_role Construct the random negative sample dataset D of event elements -2 ; Mix D +1 , D +2 , D -1 , D -2 , and shuffle them to finally obtain the model training dataset D all ; The positive sample data set D of the event type +1 and the positive sample data set D of the event elements +2 are constructed through the following steps: A1. Extract the event type e corresponding to a certain event contained in the text data text_p of the original dataset type , trigger word w trigger , event role e role_1 ~e role_n , corresponding event element w arg_1 ~w arg_n , where n is the number of event roles included in this event; the input for constructing the positive sample of the event type of this event is text_p + e type + "trigger word", and the output is w trigger ; the input for constructing the positive sample of the event element of this event is text_p + prompt arg , and the output is w arg_1 ~w arg_n ; among them, the event element prompt arg can be obtained by the following formula: For each event in the text data text_p, construct positive samples of event types and positive samples of event elements using the method in (1) to obtain the positive sample dataset D of event types +1 and the positive sample dataset D +2 ; The event type all-negative sample data set D -1 is constructed through the following steps: B1. Replace the positive sample e of a certain event type type successively with other event types of this event in the event type dataset S type and if the target output is empty for all of them, obtain the event type all-negative samples of this event; B2. For all events in the positive sample dataset D of event types +1 use the method in (1) to construct the full negative sample dataset D of event types -1 ; The random negative sample data set D of the event elements -2 is constructed through the following steps: (1) Find all positive samples of event elements for a certain event in D +2 and find all event element prompts from the positive samples of event elements arg to form a set S prompt ; (2) Randomly select a trigger word from S trigger to obtain w trigger_random ; Randomly select an element from S type_role to obtain an event type e type_random , an event role e role_random and the position p where the event role is located; (3) Randomly select p event elements from the event element set S argument to obtain w arg_r_1 ~w arg_r_p , and combine them in the following format to obtain the random prompt prompt of event elements arg_random ; prompt arg_random = e type_random + w trigger_random + w arg_r_1 + … + w arg_r_p + e role_random (4) Judge the prompt arg_random Whether it exists in S prompt If it exists, repeat steps 2, 3, and 4; if it does not exist, use the prompt arg_random Construct negative samples and add the prompt arg_random To S prompt ; (5) Repeat steps (1) to (4) until 5n random negative samples of event elements are obtained; (6) For D +2 For all event samples in, a random negative sample dataset D of event elements is constructed using the methods in (1) to (5). -2 ; The step-by-step construction of the prediction sample set for each step includes: Construct the first-step prediction sample set D based on the text to be extracted text and the event feature data set step_1 ; Based on the text to be extracted, the event feature data set, and the prediction result of the previous model M trained Construct a prompt message prompt based on the prediction result, and construct a prediction sample set for the next step model in the structure of text+prompt, so as to sequentially construct the prediction sample sets D step_2 ~D step_(n+1) ; The trained model M trained is used to obtain the prediction result of each step based on the prediction sample set of each step and integrate it to obtain the final extraction result, including: Input D step_1 into the model M trained to obtain all the trigger words p contained in the first-step prediction result text text trigger ; In the format text+prompt _X Construct the prediction sample set D for the 2nd to (n + 1)th steps step_2 ~D step_(n+1) , and input D step_x into the model M trained to obtain each trigger word corresponding event type of the (x - 1)th event role and the corresponding (x - 1)th event element where x ∈ [2, n + 1]; where prompt _X is expressed as: Combine the hint information of the last step with the extraction result to obtain a complete event.
2. The method according to claim 1, characterized in that, The obtaining of the labeled text data includes: Obtain the original text data; Label the original text data; wherein, the labeling includes: determining the event type included in the sentence in the text data; extracting the trigger word, event elements and their positions according to the event type; and assigning appropriate event role labels to the event elements.
3. The method according to claim 1, wherein The event feature data set includes: The correspondence schema between event types and all event roles, and the set S of correspondences between event types and individual event roles type_role , the set S of all event types type , the set S of all trigger words trigger and the set S of all event elements argument ; where the schema records all event types in the original dataset and all their corresponding event roles; S type_role is derived from the schema and includes the pairwise combinations of the event types and all event roles of each event in the schema, as well as which event role this event role belongs to in the schema; S type records all event types; S trigger records all trigger words that appear in the original dataset; S argument records all event elements contained in the original dataset.
4. The method according to claim 1, characterized in that, The training of the T5 model includes: Divide the model training dataset D all into a training set D train , a validation set D eval and a test set D test ; Use the training set D train Fine-tune and train the T5 model for n rounds. After each round of training, use the validation set D eval for validation. Select the model with the best results on the validation set as the final model, and use the test set D test for testing to finally obtain the trained model M trained ; During the training process, the following formula is used to calculate the model loss and update the parameters: Loss=CrossEntropy(x pred ,x gold ) where x pred is the prediction result, and x gold is the target output.
Citation Information
Patent Citations
Text event acquisition method and device, electronic equipment and storage medium
CN111597302A
Method for Summarizing Event-Related Texts To Answer Search Queries
US20130013535A1