An event detection method based on multi-task joint learning
By splitting the event detection task into two subtasks: event type judgment and trigger word recognition, and using multi-task joint learning method, the problems of low recall rate and noise information influence of existing event detection methods are solved, and better recognition and learning effects are achieved.
Patent Information
- Application Number
- CN202211743006.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-12-30
AI Technical Summary
The existing event detection methods have problems such as low recall, noise information affects the effect and difficult implementation when identifying event trigger words and event types.
Using a multi-task joint learning method, the event detection task is split into two subtasks: event type judgment and trigger word recognition. Joint modeling is carried out through deep pre-trained models such as BERT, sharing hidden layer parameters of neural networks, and learning the relationship between labels and text through attention mechanisms.
The recognition effect of event detection is improved, the relationship between labels and text is better learned through the attention mechanism, the correlation between tasks is used to improve the learning effect of the model, and the number of samples is expanded to learn more fully.
Smart Images

Figure CN116628189B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of event detection in information extraction tasks in natural language processing, and particularly relates to an event detection method based on multi-task joint learning. Background Art
[0002] The goal of event extraction is to automatically identify trigger words, event types, event arguments, etc. of the events that occur in unstructured text, which is an important research field in natural language processing. As one of the subtasks of event extraction, event detection aims to identify event trigger words from the given text and classify them into the correct event types. The trigger word refers to the core word or phrase that marks the occurrence of an event, and the event type is the type of event that needs to be detected predefined in the task.
[0003] The current mainstream event detection methods have the following several solutions:
[0004] I. Lexicon- or rule-based method
[0005] The lexicon- or rule-based method constructs a trigger word lexicon or designs a trigger word detection template for each event, and then identifies the trigger word and event type through matching.
[0006] II. Deep learning-based method
[0007] The neural network model can automatically learn high-level feature representations related to trigger words from the original text. Therefore, taking the original text as input, using deep learning models such as LSTM and Transformer to automatically learn text features, then performing character-level classification to achieve trigger word recognition, and further realizing event type judgment.
[0008] III. Multi-feature fusion method
[0009] This method usually fuses various features such as syntactic dependency features, part-of-speech features, word vector features, etc., and then inputs them into networks such as LSTM, Transformer, GCN (Graph Convolutional Network) for learning, so as to achieve trigger word recognition and classification. Syntactic dependency relations can represent the dependency relations between words. This method fuses morphological and syntactic features, etc., and can better locate the information most relevant to the trigger word, thus achieving a better recognition effect.
[0010] The lexicon- or rule-based method usually can obtain a relatively high precision rate, but its recall rate is low, and at the same time, it is too dependent on the effectiveness of the rules and usually has very poor generalization.
[0011] The deep learning-based method does not require manual feature construction. However, since the input is the entire text, most words in the text are irrelevant noise information for the judgment of trigger words and event types, which may affect the effect of the event detection task.
[0012] The multi-feature fusion method needs to use tools such as lexical analysis and syntactic analysis. However, there are a certain proportion of parsing errors in these existing analysis tools, which will still cause the retention of some noise information. And this method is relatively difficult to implement.
[0013] Therefore, it is necessary to propose an event detection method based on multi-task joint learning to solve the above problems. Summary of the Invention
[0014] The purpose of the present invention is to provide an event detection method based on multi-task joint learning to solve the problems proposed in the above background technology.
[0015] To achieve the above purpose, the present invention provides the following technical solution: An event detection method based on multi-task joint learning, including the following steps:
[0016] S1: Sample generation. For example, the predefined event types in a certain business scenario are Label = ["injured", "sentenced", "theft"]. For the text "The defendant Yuan used his fist to injure the face of the victim Guo", the trigger word contained therein is "injured", and the event type is "injured". Concatenate the event type with the text, and then label the trigger word and the event type respectively. The trigger word recognition uses the BIO annotation mode as a sequence annotation task. "B" represents the start of the trigger word, "I" represents the middle or end of the trigger word, and "O" represents not belonging to the trigger word; the event type judgment is a binary classification task, "1" represents that the text contains this event type, and "0" represents that the text does not contain this event type;
[0017] S2: Multi-task joint learning based on a deep pre-training model. Jointly model the trigger word recognition and event type judgment. The two sub-tasks share the neural network hidden layer parameters, and then construct their own classifiers for different tasks to complete their respective task objectives. The detailed model calculation steps are as follows:
[0018] a. Concatenate the event type l i and the text content text, and add the "[CLS]" and "[SEP]" flags at the beginning and end respectively, and then perform segmentation to obtain the sequence X = [[CLS], x 1 , x 2 , x 3 ,..., x n , [SEP]];
[0019] b. Input the sequence X into the BERT model to obtain the representation vector E = [e[CLS] , e 1 , e 2 , e 3 ,..., e n , e [SEP] ;
[0020] c1. Trigger word recognition
[0021] (1). Input the representation vector e of each character in the text n into a fully connected neural network, and after passing through the softmax layer, output to obtain the probability P = [p B , p I , p O that the character belongs to each type in "BIO";
[0022] (2). Calculate the cross-entropy loss loss1 between the probability P that each character belongs to each type and the true trigger word label;
[0023] c2. Event type judgment
[0024] (1). Take the vector e at the position of "[CLS]" in E [CLS] , and input e [CLS] into a fully connected neural network, and after passing through the softmax layer, output to obtain the event type probability P = [p 1 , p 2 , p 1 indicating the probability that the text contains the event type l i , p 2 indicating the probability that the text does not contain the event type l i ;
[0025] (2). Calculate the cross-entropy loss loss2 between the predicted event type probability P of the text and the true event type;
[0026] d. Perform weighted summation on loss1 and loss2 to obtain loss, and then perform backpropagation on loss to update the model parameters by the gradient descent method.
[0027] Preferably, one piece of text in S1 generates three pieces of annotation data. In the annotation data, text is formed by concatenating the event type and the text content, with "[SEP]" as the separator between the two. trigger_tag and event_type_tag respectively represent the annotation results of the trigger word and the event type. Since the event type included in this text is "injury" and the trigger word is "beaten", when the concatenated event type is "injury", event_type_tag is 1. The starting positions of "beaten" in the concatenated text are 20 and 21 respectively. Therefore, the 20th position in trigger_tag is "B", the 21st position is "I", and all other positions are "O".
[0028] Preferably, when the concatenated event types are "sentencing" and "fraud", since these two event types are not included in the text, event_type_tag is 0 for both, and all positions in trigger_tag are also "O".
[0029] The technical effects and advantages of the present invention:
[0030] 1. The present invention splits the event detection task into two subtasks: event type judgment and trigger word recognition, and then jointly models and learns the two subtasks. The input of the model includes the event type and the text content. Through the attention mechanism, the relationship between the label and the text can be better learned, and the learning effect of the model is further improved by utilizing the correlation between the tasks;
[0031] 2. The present invention concatenates the event type and the text content as the input of the model, and then respectively performs trigger word recognition and event type judgment under the limited event type. Assuming that the number of preset event types is n, each event type is concatenated with the text content respectively. In this way, each piece of data generates n samples, which not only avoids the problem that the input of the model during model inference is inconsistent with the input of the model during the training stage due to not knowing the event type of the text, but also expands the number of samples, enabling the model to be more fully learned;
[0032] 3. For each sample formed by concatenating the event type and the text content, only the trigger word recognition and event type judgment under the current event type are performed, avoiding the problem that when a text contains multiple event types and each event type corresponds to a different trigger word, the corresponding relationship between the trigger word and the event type cannot be distinguished. Description of the Drawings
[0033] Figure 1 It is a flowchart of the event detection method based on multi-task joint learning of the present invention. Detailed Embodiments
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] The present invention provides a Figure 1 method for event detection based on multi-task joint learning as shown. The event detection is split into two subtasks: event type judgment and trigger word recognition, and then the two subtasks are jointly modeled.
[0036] The type of an event and the event trigger word are often closely related. Given the event type, the recognition effect of the trigger word will be significantly improved. In order to utilize the event type information in the trigger word recognition subtask, in this proposal, the label content of the event type and the text content to be detected are concatenated as the input of the model. The event type label is no longer a symbol independent of the text. The model can better learn the relationship between the label and the text through the attention mechanism, thereby achieving a better recognition effect.
[0037] Assume that the predefined event type labels are Label = [l 1 , l 2 ,... l n . For a specific text t. Each label l i in Label is concatenated with t respectively to obtain a new text t i . t i is used as the input of the model and encoded by the deep pre-trained language model BERT to obtain the representation vector of t i . Then, event type judgment and trigger word recognition are performed under the condition that the event type is l i . Among them, the event type judgment is a binary classification task, which only needs to judge whether the text contains the event l i ; the trigger word recognition is a character-level classification task, which only needs to recognize the trigger word when the event type is l i .
[0038] The specific steps include the following aspects:
[0039] Sample generation
[0040] If the predefined event types for a certain business scenario are Label=["injured", "sentenced", "stolen"], for the text "The defendant Yuan used his fist to injure the face of the victim Guo", the trigger word contained is "injured", and the event type is "injured". Concatenate the event type with the text, then label the trigger word and the event type respectively. For trigger word recognition, the BIO annotation mode is used for the sequence annotation task. "B" indicates the start of the trigger word, "I" indicates the middle or end of the trigger word, and "O" indicates not belonging to the trigger word; for event type judgment, it is a binary classification task, "1" indicates that the text contains this event type, and "0" indicates that the text does not contain this event type.
[0041]
[0042] The generated samples are shown in the above table. There are three event types, so one text will generate three labeled data. In the labeled data, text is concatenated by the event type and the text content, and "[SEP]" is the separator between the two. trigger_tag and event_type_tag respectively represent the annotation results of the trigger word and the event type. Since the event type contained in this text is "injured" and the trigger word is "injured", when the concatenated event type is "injured", event_type_tag is 1. The starting positions of "injured" in the concatenated text are 20 and 21 respectively. Therefore, the 20th position in trigger_tag is "B", the 21st position is "I", and all other positions are "O"; when the concatenated event types are "sentenced" and "scammed", since the text does not contain these two event types, event_type_tag is 0 for both, and all positions in trigger_tag are also "O".
[0043] Multi-task joint learning based on a deep pre-trained model
[0044] Jointly model trigger word recognition and event type judgment. The two sub-tasks share the hidden layer parameters of the neural network, and then construct their own classifiers for different tasks to achieve their respective task goals. The detailed model calculation steps are as follows:
[0045] a. Concatenate the event type li and the text content text, and add the "[CLS]" and "[SEP]" flags at the beginning and end respectively, then perform segmentation to obtain the sequence X = [[CLS], x 1 , x 2 , x 3 ,..., x n , [SEP]];
[0046] b. Input the sequence X into the BERT model to obtain the representation vector E = [e [CLS] , e1 , e 2 , e 3 ,..., en, e [SEP] ;
[0047] c1. Trigger word recognition
[0048] (1). Input the representation vector e of each character in the text n into a fully connected neural network, and after passing through the softmax layer, output to obtain the probability P = [p B , p I , p O that the character belongs to each type in "BIO";
[0049] (2). Calculate the cross-entropy loss loss1 between the probability P that each character belongs to each type and the true trigger word label.
[0050] c2. Event type judgment
[0051] (1). Take the vector e at the position of "[CLS]" in E [CLS] , and input e [CLS] into a fully connected neural network, and after passing through the softmax layer, output to obtain the event type probability P = [p 1 , p 2 , p 1 indicating the probability that the text contains the event type l i , p 2 indicating the probability that the text does not contain the event type l i ;
[0052] (2). Calculate the cross-entropy loss loss2 between the event type probability P predicted by the text and the true event type.
[0053] d. Perform weighted summation on loss1 and loss2 to obtain loss, and then perform backpropagation on loss to update the model parameters by the gradient descent method.
Claims
1. An event detection method based on multi-task joint learning, characterized in that: It includes the following steps: S1: Sample generation. For example, the predefined event types in a certain business scenario are Label = ["Injury", "Sentencing", "Theft"]. For the text "The defendant Yuan used his fist to injure the face of the victim Guo", the trigger word contained therein is "injure", and the event type is "Injury". Concatenate the event type with the text body, and then label the trigger word and the event type respectively. For trigger word recognition, use the BIO annotation mode for sequence annotation tasks. "B" represents the start of the trigger word, "I" represents the middle or end of the trigger word, and "O" represents not belonging to the trigger word. Event type judgment is a binary classification task. "1" means the text contains this event type, and "0" means the text does not contain this event type. S2: Multi-task joint learning based on a deep pre-trained model. Jointly model trigger word recognition and event type judgment. The two sub-tasks share the neural network hidden layer parameters, and then construct their own classifiers for different tasks to achieve their respective task goals. The detailed model calculation steps are as follows: a. Concatenate the event type l i and the text content text, and add "[CLS]" and "[SEP]" at the beginning and end respectively, and then perform segmentation to obtain the sequence X = [[CLS], x 1 , x 2 , x 3 ,..., x n , [SEP]]; b. Input sequence X into the BERT model to obtain the representation vector E = [e [CLS] , e 1 , e 2 , e 3 ,..., e n , e [SEP] ; c1. Trigger word recognition (1) Input the representation vector \(e\) of each character in the text n into a fully-connected neural network, and after passing through the softmax layer, output the probability \(P = [p\) B , p I , p O that the character belongs to each type in "BIO"; (2). Calculate the cross-entropy loss loss1 between the probability P that each character belongs to each type and the true trigger word label. c2. Event type judgment (1). Take the vector e at the "[CLS]" position in E [CLS] , and input e [CLS] into a fully-connected neural network, and output it after passing through the softmax layer to obtain the event type probability P = [p 1 , p 2 , where p1 represents the probability that the text contains the event type l i , and p 2 represents the probability that the text does not contain the event type l i ; (2). Calculate the cross-entropy loss loss2 between the predicted event type probability P of the text and the true event type. d. Perform weighted summation on loss1 and loss2 to obtain loss, and then perform backpropagation on loss, and update the model parameters by the gradient descent method.
2. The event detection method based on multi-task joint learning according to claim 1, characterized in that: In S1, one piece of text will generate three pieces of annotated data. In the annotated data, text is concatenated by the event type and the text content, and "[SEP]" is the separator between the two. trigger_tag and event_type_tag respectively represent the annotation results of the trigger word and the event type. Since the event type contained in this text is "Injury" and the trigger word is "injure", when the concatenated event type is "Injury", event_type_tag is 1. The starting positions of "injure" in the concatenated text are 20 and 21 respectively. Therefore, the 20th position in trigger_tag is "B", the 21st position is "I", and all other positions are "O".
3. The event detection method based on multi-task joint learning according to claim 2, characterized in that: When the concatenated event types are "Sentencing" and "Fraud", since the text does not contain these two event types, event_type_tag are both 0, and all positions in trigger_tag are also "O".
Citation Information
Patent Citations
Event detection model construction method and device, electronic equipment and storage medium
CN111813931A
Literature-based cancer-related biomedical event database construction method
CN111859935A