A Few-Shot Event Detection Method and System Based on Continual Learning

By using the knowledge distillation and empirical playback mechanism methods in the event detection model, new types of embedding vectors are updated phase by phase, and the forgetting problem in the incremental learning stage in the existing technology is solved, and the efficiency and accuracy of event detection are improved.

CN116542320BActive Publication Date: 2025-05-30GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310506290.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-06
Publication Date
2025-05-30
Estimated Expiration
2043-05-06

AI Technical Summary

Technical Problem

The existing technology has the problem of forgetting in the incremental learning stage, and cannot effectively distinguish between old and new categories, and has high requirements for rapid learning ability.

Method used

By establishing an event detection model including instance encoder, type encoder and prototype network, using knowledge distillation and experience playback mechanisms, continuous learning training is carried out phase by phase, new types of embedding vectors are updated, and the model is optimized through the total loss function.

Benefits of technology

It solves the problem of forgetting in traditional models, enriches the connection between new and old types, and improves the efficiency and accuracy of event detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116542320B_ABST
    Figure CN116542320B_ABST
Patent Text Reader

Abstract

The present invention provides a few-shot event detection method and system based on continual learning. The method includes: obtaining an initial event detection data set and performing preprocessing, establishing a few-shot event detection incremental task set including several stages according to the preprocessed initial event detection data set, establishing a few-shot event detection framework based on continual learning. When facing continual learning of successive new tasks, first learning the prototype representation of the new type and saving it as new type knowledge, and then obtaining the event detection results of the new type through an experience replay mechanism, knowledge distillation, and knowledge transfer between the old type and the new type. The method of the present invention can solve the forgetting problem of traditional models, while enriching the connection between the new type and the old type, and improving the efficiency and accuracy of event detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer deep learning, and more specifically, to a small-sample event detection method and system based on continuous learning. Background Art

[0002] With the rapid development and popularization of the Internet, users are frequently faced with a large number of interactive behaviors when surfing the Internet every day. How to extract useful key information from the explosive information has become a topic of great concern.

[0003] In information extraction, event detection is an important task, which aims to identify and classify event trigger words of predefined event types in text. It includes trigger word recognition and event classification. The purpose of trigger recognition is to extract all event triggers from a sentence, while event classification requires classifying them into corresponding event types. Previous event detection methods can be roughly divided into two categories; one is a two-stage model, which first performs trigger recognition and then event classification, and the other model has only one stage, which takes all tokens in the sentence as trigger candidates and uses an additional event type NA for event classification, which means that the corresponding tokens are not triggers.

[0004] Although supervised event detection algorithms have achieved good results, when adapting to new event types and different domains, a large amount of manually annotated event data is required, which is costly. Therefore, training an event detection model with continuous learning ability and applied to small samples to continuously detect newly added event types has important research and application value.

[0005] The prior art discloses a video event detection method based on continuous learning, including an initial learning stage and an incremental learning stage; in the initial learning stage, video data with labels is prepared, and the sparse autoencoder is used to learn the video data with labels to train a prior model; in the incremental learning stage, the trained prior model is used to classify newly arrived video data, calculate the probability score and gradient distance, and use active learning to decide whether to automatically add labels or manually add labels to the newly arrived video data according to the calculation results; the method in the prior art combines deep learning and active learning, uses unsupervised learning to extract features and uses active learning to reduce the work of manual classification, but in the incremental learning stage, due to the unavailability of old classes, the distribution fitted based on new classes often overlaps with the distribution fitted based on old classes, so it will cause it to be unable to distinguish between new and old classes, and there is a catastrophic forgetting problem, that is, it tends to overfit new classes and forget the knowledge of old classes; in addition, the method in the prior art has a high requirement for the ability to learn quickly due to the limited number of labeled data from new classes. Summary of the Invention

[0006] To overcome the problems of forgetting in the above-mentioned existing technologies and the relatively high requirements for the ability of rapid learning, the present invention provides a few-shot event detection method and system based on continual learning, which can solve the forgetting problem of traditional models, enrich the connection between new types and old types, and improve the efficiency and accuracy of event detection.

[0007] To solve the above technical problems, the technical solution of the present invention is as follows:

[0008] A few-shot event detection method based on continual learning, comprising the following steps:

[0009] S1: Obtain an initial event detection data set and perform preprocessing. The initial event detection data set includes a support set and a query set. Both the support set and the query set include a number of instances and their corresponding triggers and event types;

[0010] S2: Establish a few-shot event detection incremental task set including a number of stages according to the preprocessed initial event detection data set;

[0011] S3: Establish an event detection model, which includes an instance encoder, a type encoder, and a prototype network connected in sequence;

[0012] S4: Input the tasks of the few-shot event detection incremental task set into the initial event detection model stage by stage for continual learning training;

[0013] For the tasks of each stage, perform knowledge distillation on the output data of the prototype network of the current stage and the previous stage respectively. Using the experience replay mechanism, save the output data of the prototype network of the previous stage after knowledge distillation as old type knowledge, and save the output data of the current stage after knowledge distillation as new type knowledge;

[0014] S5: For the tasks of each stage, update the new type knowledge with the old type knowledge, obtain the embedding vector of the new type, and perform event detection on the data of the current stage task again according to the embedding vector of the new type to obtain the updated event detection result of the new type;

[0015] S6: Set a total loss function. For the tasks of each stage, optimize the event detection model using the total loss function. When the value of the total loss function is the smallest, complete the optimization of the event detection model and obtain the optimal event detection model;

[0016] S7: Obtain the data to be event-detected and input it into the optimal event detection model for event detection to obtain the optimal event detection result.

[0017] Preferably, the specific method for obtaining the initial event detection data set and performing preprocessing in step S1 is:

[0018] The initial dataset for event detection is specifically the RAMS, ACE, and LR-KBP datasets, including a support set S and a query set Q. Both the support set S and the query set Q include a number of instances and their corresponding triggers and event types:

[0019]

[0020]

[0021] Among them, represents the i-th event mention in the support set S, including the instance trigger and event type N is the number of event types in the support set S, and K is the number of instances included in each event type; represents the i-th event mention in the query set Q, including the instance trigger and event type M is the number of instances in the query set Q; each instance or is represented as a word sequence L is the maximum length of the corresponding event mention; all the event types in the support set S and the query set Q are denoted as the set ε;

[0022] The specific method for preprocessing the initial dataset for event detection is: dividing the support set S and the query set Q into subsets.

[0023] Preferably, in step S2, the specific method for establishing a few-shot event detection incremental task set including several stages according to the preprocessed initial dataset for event detection is:

[0024] The few-shot event detection incremental task set T includes tasks in t stages, denoted as: T = {S, Q};

[0025] Each subset in the support set S and the query set Q is used as a task in one stage of the few-shot event detection incremental task set.

[0026] Preferably, the instance encoder in step S3 is specifically:

[0027] Encoding the input data using a preset token marking sequence, the length of the preset token marking sequence is L, denoted as The trigger of the task in the t-th stage is denoted as

[0028] Encode the trigger of the preset token sequence using a pre-trained BERT model to obtain the context representation of the trigger. The context representation of the trigger at the t-th stage is denoted as

[0029] Use the token embedding of [CLS] as the context representation of the preset token sequence. The context representation of the token sequence X i of the i-th event type is denoted as X i ;

[0030] During the event detection process, use the context representation of the trigger at the t-th stage as an event trigger candidate to calculate the probability of the corresponding event type.

[0031] Preferably, the type encoder in step S3 is specifically:

[0032] Calculate the prototype P k of the k-th event type e k according to the following formula:

[0033]

[0034] where N k is the number of instances in the event type e k .

[0035] Preferably, the prototype network in step S3 is specifically:

[0036] Calculate the probability of the corresponding k-th event type e k according to the prototype P k of the input data. The specific formula is:

[0037]

[0038] where y is the label of the context representation of the trigger at the t-th stage, ‖·‖ represents the Euclidean distance, and N e is the total number of event types.

[0039] Preferably, in step S4, for the tasks of each stage, perform knowledge distillation on the output data of the prototype network of the current stage and the previous stage respectively. Using the experience replay mechanism, save the output data of the prototype network of the previous stage after knowledge distillation as old type knowledge, and save the output data of the current stage after knowledge distillation as new type knowledge. The specific method is:

[0040] For the tasks in the t-th stage and the (t-1)-th stage, knowledge distillation is performed on the output data of the prototype network in the current stage and the previous stage respectively, and the experience replay mechanism is used to allocate k slots for each event type as storage space;

[0041] Save the output data of the prototype network in the previous stage after knowledge distillation as old type knowledge, specifically:

[0042]

[0043]

[0044] where, T 1 is the first temperature factor for knowledge distillation, satisfying T 1 > 1, is the output prediction of the previous stage after adjusting the temperature factor, P t-1 is the output prediction of the prototype network in the previous stage, is the output prediction of the current stage after adjusting the temperature factor, P t is the output prediction of the prototype network in the current stage.

[0045] Preferably, in the step S5, for the tasks in each stage, the specific method for updating the new type of knowledge with the old type of knowledge to obtain the embedding vector of the new type is:

[0046] For the tasks in the t-th stage and the (t-1)-th stage, update the prototype embedding vector of the new type of knowledge with the old type of knowledge, and the obtained embedding vector z of the new type is specifically:

[0047] z = g α,β (N)μ+(1 - g α,β (N))r

[0048] where, g α,β (N) is the gate function, satisfying g α,β (N)=αexp(-βN), α and β are the first and second hyperparameters respectively; μ is the first initialization vector, and r is the second initialization vector;

[0049] The first initialization vector μ satisfies μ = ω + ν, where,

[0050] The second initialization vector r is a randomly initialized vector, satisfying r ∼ N(0, d 2 I / dim(r)), where,

[0051] Preferably, the total loss function set in the step S6 is specifically:

[0052]

[0053] Among them, l is the total loss function value, l S is the self-training loss function value, l D is the knowledge distillation loss function value, l ED is the cross-entropy loss function value, and λ is the third hyperparameter;

[0054] In the total loss function, the self-training loss function is specifically:

[0055]

[0056] Among them, O t-1 is the event type in the task of the (t - 1)th stage, q t-1 is the pseudo-label of the old type knowledge, satisfying where τ is the second temperature factor, satisfying τ < 1;

[0057] The knowledge distillation loss function is specifically:

[0058]

[0059] The cross-entropy loss function is specifically:

[0060]

[0061] Among them, y is the context representation of the trigger in the tth stage, N e is the total number of event types.

[0062] The present invention also provides a few-shot event detection system based on continual learning, applying the above-mentioned few-shot event detection method based on continual learning, including:

[0063] Data acquisition unit: used to acquire the initial event detection data set and perform preprocessing. The initial event detection data set includes a support set and a query set, and both the support set and the query set include several instances and their corresponding triggers and event types;

[0064] Task construction unit: used to establish a few-shot event detection incremental task set including several stages according to the preprocessed initial event detection data set;

[0065] Model establishment unit: used to establish an event detection model, and the event detection model includes an instance encoder, a type encoder, and a prototype network connected in sequence;

[0066] Knowledge distillation unit: used to input the tasks of the few-shot event detection incremental task set into the initial event detection model in sequence by stages for continual learning training;

[0067] For the tasks of each stage, knowledge distillation is respectively performed on the output data of the prototype network in the current stage and the previous stage. Using the experience replay mechanism, the output data of the prototype network in the previous stage after knowledge distillation is saved as old-type knowledge, and the output data of the current stage after knowledge distillation is saved as new-type knowledge;

[0068] Data update unit: For the tasks of each stage, it is used to update the new-type knowledge with the old-type knowledge, obtain the embedding vector of the new type, and perform event detection on the data of the current stage task again according to the embedding vector of the new type to obtain the updated event detection result of the new type;

[0069] Model optimization unit: It is used to set the total loss function. For the tasks of each stage, the event detection model is optimized using the total loss function. When the value of the total loss function is the smallest, the optimization of the event detection model is completed to obtain the optimal event detection model;

[0070] Event detection unit: It is used to obtain the data to be event-detected and input it into the optimal event detection model for event detection to obtain the optimal event detection result.

[0071] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0072] The present invention provides a few-shot event detection method based on continual learning, including: obtaining an initial event detection dataset and performing preprocessing, where the initial event detection dataset includes a support set and a query set, and both the support set and the query set include a number of instances and their corresponding triggers and event types; establishing a few-shot event detection incremental task set including a number of stages according to the preprocessed initial event detection dataset; establishing an event detection model, where the event detection model includes an instance encoder, a type encoder, and a prototype network connected in sequence; inputting the tasks of the few-shot event detection incremental task set into the initial event detection model stage by stage for continual learning training; for the tasks of each stage, respectively performing knowledge distillation on the output data of the prototype network of the current stage and the previous stage, and using the experience replay mechanism, saving the output data of the prototype network of the previous stage after knowledge distillation as old type knowledge, and saving the output data of the current stage after knowledge distillation as new type knowledge; for the tasks of each stage, updating the prototype embedding vector of the new type knowledge with the old type knowledge, and performing event detection on the new type again according to the updated prototype embedding vector of the new type knowledge to obtain the event detection result of the updated new type; setting a total loss function, and for the tasks of each stage, optimizing the event detection model with the total loss function, and when the value of the total loss function is the smallest, completing the optimization of the event detection model to obtain the optimal event detection model; obtaining the data to be event-detected and inputting it into the optimal event detection model for event detection to obtain the optimal event detection result;

[0073] The present invention establishes a few-shot event detection framework based on continual learning. When facing the continual learning of successive new tasks, it first learns the prototype representation of the new type, and then obtains the event detection result of the new type through the experience replay mechanism, knowledge distillation, and knowledge transfer between the old type and the new type, which can solve the forgetting problem of traditional models, while enriching the connection between the new type and the old type and improving the efficiency and accuracy of event detection. Brief Description of the Drawings

[0074] Figure 1 It is a flowchart of a few-shot event detection method provided for Embodiment 1.

[0075] Figure 2 It is a schematic diagram of a few-shot event detection method provided for Embodiment 2.

[0076] Figure 3 It is a comparison chart of the accuracy rates of a few-shot event detection method provided for Embodiment 2 on different datasets.

[0077] Figure 4 It is a structural diagram of a few-shot event detection system provided for Embodiment 3. Detailed implementation manners

[0078] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent;

[0079] To better illustrate this embodiment, some components in the accompanying drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product;

[0080] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.

[0081] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0082] Embodiment 1

[0083] As Figure 1 shown, this embodiment provides a few-shot event detection method based on continuous learning, including the following steps:

[0084] S1: Obtain the initial event detection dataset and perform preprocessing. The initial event detection dataset includes a support set and a query set. Both the support set and the query set include a number of instances and their corresponding triggers and event types;

[0085] S2: Establish a few-shot event detection incremental task set including several stages according to the preprocessed initial event detection dataset;

[0086] S3: Establish an event detection model. The event detection model includes an instance encoder, a type encoder, and a prototype network connected in sequence;

[0087] S4: Input the tasks of the few-shot event detection incremental task set into the initial event detection model stage by stage for continuous learning training;

[0088] For the tasks of each stage, perform knowledge distillation on the output data of the prototype network of the current stage and the previous stage respectively. Using the experience replay mechanism, save the output data of the prototype network of the previous stage after knowledge distillation as old type knowledge, and save the output data of the current stage after knowledge distillation as new type knowledge;

[0089] S5: For the tasks of each stage, update the new type knowledge with the old type knowledge, obtain the embedding vector of the new type, and perform event detection on the data of the current stage task again according to the embedding vector of the new type to obtain the updated event detection result of the new type;

[0090] S6: Set the total loss function. For the tasks of each stage, optimize the event detection model using the total loss function. When the value of the total loss function is the smallest, complete the optimization of the event detection model and obtain the optimal event detection model;

[0091] S7: Obtain the data to be detected for events and input it into the optimal event detection model for event detection, and obtain the optimal event detection result.

[0092] In the specific implementation process, first obtain the initial event detection data set and perform preprocessing. The initial event detection data set includes a support set and a query set. Both the support set and the query set include a number of instances and their corresponding triggers and event types; establish a small-sample event detection incremental task set including several stages according to the preprocessed initial event detection data set; establish an event detection model, and the event detection model includes an instance encoder, a type encoder, and a prototype network connected in sequence; input the tasks of the small-sample event detection incremental task set into the initial event detection model stage by stage for continuous learning and training;

[0093] For the tasks of each stage, perform knowledge distillation on the output data of the prototype network of the current stage and the previous stage respectively. Using the experience replay mechanism, save the output data of the prototype network of the previous stage after knowledge distillation as old type knowledge, and save the output data of the current stage after knowledge distillation as new type knowledge; for the tasks of each stage, update the new type knowledge with the old type knowledge, obtain the embedding vector of the new type, and perform event detection on the data of the new task according to the updated new type to obtain the event detection result of the updated new type; set the total loss function, and for the tasks of each stage, optimize the event detection model using the total loss function. When the value of the total loss function is the smallest, complete the optimization of the event detection model and obtain the optimal event detection model; finally, obtain the data to be detected for events and input it into the optimal event detection model for event detection, and obtain the optimal event detection result;

[0094] This method establishes a small-sample event detection framework based on continuous learning. When facing continuous learning of successive new tasks, it first learns the prototype representation of the new type, and then obtains the event detection result of the new type through the experience replay mechanism, knowledge distillation, and knowledge transfer between the old type and the new type, which can solve the forgetting problem of traditional models, enrich the connection between the new type and the old type at the same time, and improve the efficiency and accuracy of event detection.

[0095] Embodiment 2

[0096] As Figure 2 shown, this embodiment provides a small-sample event detection method based on continuous learning, including the following steps:

[0097] S1: Obtain the initial event detection data set and perform preprocessing. The initial event detection data set includes a support set and a query set. Both the support set and the query set include a number of instances and their corresponding triggers and event types;

[0098] S2: Establish a small-sample event detection incremental task set including several stages based on the preprocessed initial event detection data set;

[0099] S3: Establish an event detection model, where the event detection model includes an instance encoder, a type encoder, and a prototype network connected in sequence;

[0100] S4: Input the tasks of the small-sample event detection incremental task set into the initial event detection model stage by stage for continuous learning and training;

[0101] For the tasks of each stage, perform knowledge distillation on the output data of the prototype network of the current stage and the previous stage respectively. Using the experience replay mechanism, save the output data of the prototype network of the previous stage after knowledge distillation as old type knowledge, and save the output data of the current stage after knowledge distillation as new type knowledge;

[0102] S5: For the tasks of each stage, update the new type knowledge with the old type knowledge, obtain the embedding vectors of the new type, and perform event detection on the data of the current stage task again according to the embedding vectors of the new type to obtain the updated event detection results of the new type;

[0103] S6: Set the total loss function. For the tasks of each stage, optimize the event detection model using the total loss function. When the value of the total loss function is the smallest, complete the optimization of the event detection model and obtain the optimal event detection model;

[0104] S7: Obtain the data to be event-detected and input it into the optimal event detection model for event detection to obtain the optimal event detection results;

[0105] The specific method for obtaining and preprocessing the initial event detection data set in step S1 is as follows:

[0106] The initial event detection data set is specifically the RAMS, ACE, and LR-KBP data sets, including a support set S and a query set Q. Both the support set S and the query set Q include several instances and their corresponding triggers and event types:

[0107]

[0108]

[0109] Among them, represents the i-th event mention in the support set S, including the instance trigger and event type N is the number of event types in the support set S, and K is the number of instances included in each event type; Denote the i-th event mention in the query set Q, including the instance Trigger and the event type M is the number of instances in the query set Q; each instance or is represented as a sequence of words L is the maximum length of the corresponding event mention; denote all the event types in the support set S and the query set Q as the set ε;

[0110] The specific method for preprocessing the initial event detection data set is: divide the support set S and the query set Q into subsets;

[0111] In step S2, specifically establishing a small-sample event detection incremental task set including several stages according to the preprocessed initial event detection data set is as follows:

[0112] The small-sample event detection incremental task set T includes tasks in t stages, denoted as: T = {S, Q};

[0113] Take each subset in the support set S and the query set Q as a task in one stage of the small-sample event detection incremental task set;

[0114] The instance encoder in step S3 is specifically:

[0115] Encode the input data using a preset token marker sequence, and the length of the preset token marker sequence is L, denoted as Denote the trigger of the t-th stage task as

[0116] Use the pre-trained BERT model to encode the trigger of the preset token marker sequence to obtain the context representation of the trigger. Denote the context representation of the trigger in the t-th stage as

[0117] Take the token embedding of [CLS] as the context representation of the preset token marker sequence. Denote the context representation of the token sequence X i of the i-th event type as X i ;

[0118] During the event detection process, take the context representation of the trigger in the t-th stage as the event trigger candidate for calculating the probability of the corresponding event type;

[0119] The type encoder in step S3 is specifically:

[0120] Calculate the prototype P of the k-th event type e according to the following formula k ofk :

[0121]

[0122] Among them, N k is the number of instances in event type e k ;

[0123] The prototype network in step S3 is specifically:

[0124] Calculate the probability corresponding to the k-th event type e k for the prototype P of the input data, and the specific formula is: k

[0125]

[0126] Among them, y is the context representation of the trigger in the t-th stage label, ‖·‖ represents the Euclidean distance, and N e is the total number of event types;

[0127] In step S4, for the tasks of each stage, knowledge distillation is performed on the output data of the prototype networks of the current stage and the previous stage respectively. Using the experience replay mechanism, the output data of the prototype network of the previous stage after knowledge distillation is saved as old type knowledge, and the output data of the current stage after knowledge distillation is saved as new type knowledge. The specific method is:

[0128] For the tasks of the t-th stage and the (t - 1)-th stage, knowledge distillation is performed on the output data of the prototype networks of the current stage and the previous stage respectively, and k slots are allocated as storage spaces for each event type using the experience replay mechanism;

[0129] Save the output data of the prototype network of the previous stage after knowledge distillation as old type knowledge, specifically:

[0130]

[0131]

[0132] Among them, T 1 is the first temperature factor for knowledge distillation, satisfying T 1 > 1, is the output prediction of the previous stage after temperature factor adjustment, and P t-1 is the output prediction of the prototype network of the previous stage, is the output prediction of the current stage after temperature factor adjustment, and P t is the output prediction of the prototype network of the current stage;

[0133] In step S5, for the tasks of each stage, the specific method of using the old type of knowledge to update the new type of knowledge and obtain the embedding vector of the new type is as follows:

[0134] For the tasks of the t-th stage and the (t - 1)-th stage, use the old type of knowledge to update the prototype embedding vector of the new type, and the obtained embedding vector z of the new type is specifically:

[0135] z = g α,β (N)μ+(1 - g α,β (N))r

[0136] where g α,β (N) is the gate function, satisfying g α,β (N)=αexp(-βN), where α and β are the first and second hyperparameters respectively; μ is the first initialization vector, and r is the second initialization vector;

[0137] The first initialization vector μ satisfies μ = ω + ν, where

[0138] The second initialization vector r is a randomly initialized vector, satisfying r ∼ N(0,d 2 I / dim(r)), where

[0139] The total loss function set in step S6 is specifically:

[0140]

[0141] where l is the value of the total loss function, l S is the value of the self-training loss function, l D is the value of the knowledge distillation loss function, l ED is the value of the cross-entropy loss function, and λ is the third hyperparameter;

[0142] In the total loss function, the self-training loss function is specifically:

[0143]

[0144] where O t-1 is the event type in the task of the (t - 1)-th stage, q t-1 is the pseudo-label of the old type of knowledge, satisfying where τ is the second temperature factor, satisfying τ < 1;

[0145] The knowledge distillation loss function is specifically:

[0146]

[0147] The cross-entropy loss function is specifically as follows:

[0148]

[0149] where y is the context representation of the trigger in the t-th stage of the label, and N e is the total number of event types.

[0150] In the specific implementation process, with the development of the Internet and the popularity of social networks, there is a vast amount of user data in the network. However, this data is presented in a semi-structured form. Currently, news websites generate a large amount of data every day, and the amount of information is so huge that it has exceeded people's control ability; there is a large amount of useless information in the network, which brings difficulties to people in obtaining useful information; how to quickly obtain the latest information and how to pay attention to events in a timely manner have become difficult problems faced by people;

[0151] A method for small-sample event detection based on continuous learning provided in this embodiment detects events from semi-structured news texts and classifies event types;

[0152] First, obtain the initial event detection dataset and perform preprocessing. The initial event detection dataset includes a support set and a query set, and both the support set and the query set include a number of instances and their corresponding triggers and event types;

[0153] In this embodiment, the initial event detection dataset is specifically the RAMS, ACE, and LR-KBP datasets. RAMS is a large-scale dataset that provides 9,124 manually annotated event triggers for 139 event subtypes; ACE is an event extraction benchmark dataset with 33 event subtypes; LR-KBP is a large-scale event detection dataset for FSL. It combines the ACE-2005 and TAC-KBP datasets and extends some event types by automatically collecting data from Freebase and Wikipedia;

[0154] Since the RAMS and ACE datasets are designed for supervised learning, they need to be re-split for FSL training; first, perform an exact training / development / test split on ACE-2005 and RAMS, and discard 5 event subtypes with insufficient sample numbers for sampling;

[0155] In this embodiment, the initial event detection dataset includes a support set S and a query set Q, and both the support set S and the query set Q include a number of instances and their corresponding triggers and event types:

[0156]

[0157]

[0158] Among them, represents the i-th event mention in the support set S, including the instance trigger and event type N is the number of event types in the support set S, K is the number of instances included in each event type. In this embodiment, K = 5; represents the i-th event mention in the query set Q, including the instance trigger and event type M is the number of instances in the query set Q; each instance or is represented as a word sequence L is the maximum length of the corresponding event mention; all event types in the support set S and the query set Q are denoted as the set ε;

[0159] The specific method for preprocessing the initial event detection data set is: partitioning and sorting the support set S and the query set Q, and partitioning them into 5 subsets;

[0160] Specifically, establishing a small-sample event detection incremental task set including several stages based on the preprocessed initial event detection data set is as follows:

[0161] The small-sample event detection incremental task set T includes a total of 5 stages of tasks, denoted as: T = {S, Q};

[0162] Taking each subset in the support set S and the query set Q as a task in one stage of the small-sample event detection incremental task set;

[0163] Establishing an event detection model, which includes an instance encoder, a type encoder, and a prototype network connected in sequence;

[0164] The specific instance encoder is as follows:

[0165] Encoding the input data using a preset token marking sequence, the length of the preset token marking sequence is L, denoted as The trigger of the t-th stage task is denoted as

[0166] Using the pre-trained BERT model to encode the trigger of the preset token marking sequence to obtain the context representation of the trigger, and the context representation of the trigger in the t-th stage is denoted as

[0167] Taking the token embedding of [CLS] as the context representation of the preset token marking sequence, and the token sequence X of the i-th event typei The context representation of i ;

[0168] During the event detection process, the context representation of the trigger in the t-th stage is used as an event trigger candidate to calculate the probability of the corresponding event type;

[0169] The type encoder is specifically:

[0170] Calculate the prototype P of the k-th event type e k according to the following formula: k :

[0171]

[0172] where N k is the number of instances in the event type e k ;

[0173] The prototype network is specifically:

[0174] Calculate the probability of the corresponding k-th event type e k according to the prototype P of the input data, and the specific formula is: k :

[0175]

[0176] where y is the label of the context representation of the trigger in the t-th stage , ‖·‖ represents the Euclidean distance, and N e is the total number of event types;

[0177] The cross-entropy loss function is specifically:

[0178]

[0179] where y is the label of the context representation of the trigger in the t-th stage , and N e is the total number of event types;

[0180] Input the tasks of the small-sample event detection incremental task set into the event detection initial model stage by stage for continuous learning and training;

[0181] For the tasks in each stage, perform knowledge distillation on the output data of the prototype networks of the current stage and the previous stage respectively. Using the experience replay mechanism, save the output data of the prototype network of the previous stage after knowledge distillation as old type knowledge, and save the output data of the current stage after knowledge distillation as new type knowledge;

[0182] In existing continuous learning methods, most methods either allocate k slots for each type or fix a total of K slots and evenly distribute them among all learned types with k or K being hyperparameters; although the latter setting method has certain advantages in terms of memory, this framework is limited to learning at most K types of tasks; therefore, this embodiment adopts the former as the continuous learning environment to learn as many event types as possible in continuous learning;

[0183] For the tasks in the t-th stage and the (t - 1)-th stage, knowledge distillation is respectively performed on the output data of the prototype networks in the current stage and the previous stage, and the experience replay mechanism is used to allocate k slots for each event type as storage space;

[0184] The output data of the prototype network in the previous stage after knowledge distillation is saved as old type knowledge, specifically:

[0185]

[0186]

[0187] where, T 1 is the first temperature factor for knowledge distillation. In this embodiment, T 1 = 2, is the output prediction of the previous stage after adjustment of the temperature factor, P t-1 is the output prediction of the prototype network in the previous stage, is the output prediction of the current stage after adjustment of the temperature factor, P t is the output prediction of the prototype network in the current stage;

[0188] The knowledge distillation loss function is specifically:

[0189]

[0190] For the tasks in each stage, the prototype embedding vectors of the new type knowledge are updated using the old type knowledge, and event detection is performed on the new type again according to the updated prototype embedding vectors of the new type knowledge to obtain the updated event detection results of the new type;

[0191] Traditional continuous learning methods solve the catastrophic forgetting problem by retaining the knowledge of the old model and cannot effectively update the old knowledge. If the new type is related to some old types, some instances of the new type may share similarities with the old types; therefore, this embodiment updates the learned knowledge by extending the knowledge distillation loss that only retains the old knowledge to a new self-training loss;

[0192] The self-training loss function is specifically:

[0193]

[0194] Among them, O t-1 is the event type in the task of the (t - 1)-th stage, and q t-1 is the pseudo-label of the old type of knowledge, satisfying where τ is the second temperature factor, and τ = 0.5;

[0195] For the tasks of the t-th stage and the (t - 1)-th stage, the prototype embedding vector of the new type of knowledge is updated using the old type of knowledge, so as to transfer the old knowledge to the new knowledge;

[0196] When a new type has enough training instances, random initialization is usually adopted. Therefore, for frequently occurring new types, this embodiment adopts a knowledge-aware random initialization. The norm of the learned features also contains knowledge about the feature space, and it can be used to help the learning of new types. The resulting second initialization vector r is a random initialization vector, satisfying r ∼ N(0, d 2 I / dim(r)), where is the average norm of the prototype embedding vectors of the old type of knowledge;

[0197] For event types with less training data, it may be difficult to train from random initialization. At this time, more knowledge needs to be transferred from the old type. Assume X 1:h is an instance of a new type, then the output of the prototype network is used as this relevant measure, aggregating the type embeddings of the learned types, and representing the new knowledge by aggregating the encoded exemplar instances. Specifically:

[0198]

[0199] Re-scale using the average norm d of the prototype embedding vectors of the old type of knowledge, and use to weight each sample instance, representing its irrelevance degree to the learned event types. Specifically:

[0200]

[0201] Thus, the first initialization vector μ is obtained, satisfying μ = ω + ν;

[0202] After that, the first initialization vector μ and the second initialization vector r are combined through a gate function, and the updated prototype embedding vector z of the new type of knowledge obtained is specifically:

[0203] z = g α,β (N)μ+(1 - g α,β (N))r

[0204] where g α,β (N) is the gate function, satisfying g α,β(N) = α exp(-βN), where α and β are the first and second hyperparameters respectively. In this embodiment, α = 0.5 and β = 0.05;

[0205] Set the total loss function. For the tasks in each stage, use the total loss function to optimize the event detection model. When the value of the total loss function is minimized, the optimization of the event detection model is completed, and the optimal event detection model is obtained;

[0206] The specifically set total loss function is as follows:

[0207]

[0208] where l is the value of the total loss function, l S is the value of the self-training loss function, l D is the value of the knowledge distillation loss function, l ED is the value of the cross-entropy loss function, and λ is the third hyperparameter. In this embodiment, if the probability of the task is less than 0.9, then λ = 0.5; otherwise, λ = 0;

[0209] Finally, obtain the data to be event-detected and input it into the optimal event detection model for event detection, and obtain the optimal event detection result;

[0210] As Figure 3 shown, the accuracy of the optimal event detection model on each dataset is presented. It can be seen from Figure 3 this that this method has a high accuracy on each dataset;

[0211] This method establishes a few-shot event detection framework based on continual learning. When facing the continual learning of successive new tasks, it first learns the prototype representations of new types, and then obtains the event detection results of new types through the experience replay mechanism, knowledge distillation, and knowledge transfer between old types and new types. It can solve the forgetting problem of traditional models, while enriching the connection between new types and old types, and improving the efficiency and accuracy of event detection.

[0212] Embodiment 3

[0213] As Figure 4 shown, this embodiment provides a few-shot event detection system based on continual learning, which applies the above-mentioned few-shot event detection method based on continual learning, and includes:

[0214] A data acquisition unit 301: used to acquire the initial event detection dataset and perform preprocessing. The initial event detection dataset includes a support set and a query set, and both the support set and the query set include a number of instances and their corresponding triggers and event types;

[0215] Task construction unit 302: It is used to establish a small-sample event detection incremental task set including several stages according to the preprocessed event detection initial data set;

[0216] Model establishment unit 303: It is used to establish an event detection model, and the event detection model includes an instance encoder, a type encoder, and a prototype network connected in sequence;

[0217] Knowledge distillation unit 304: It is used to input the tasks of the small-sample event detection incremental task set into the initial event detection model in sequence by stages for continuous learning and training;

[0218] For the tasks of each stage, knowledge distillation is respectively performed on the output data of the prototype networks of the current stage and the previous stage. Using the experience replay mechanism, the output data of the prototype network of the previous stage after knowledge distillation is saved as old type knowledge, and the output data of the current stage after knowledge distillation is saved as new type knowledge;

[0219] Data update unit 305: For the tasks of each stage, it is used to update the new type knowledge with the old type knowledge, obtain the embedding vectors of the new type, and perform event detection on the data of the tasks of the current stage again according to the embedding vectors of the new type to obtain the updated event detection results of the new type;

[0220] Model optimization unit 306: It is used to set the total loss function. For the tasks of each stage, the event detection model is optimized using the total loss function. When the value of the total loss function is the smallest, the optimization of the event detection model is completed, and the optimal event detection model is obtained;

[0221] Event detection unit 307: It is used to obtain the data to be event-detected and input it into the optimal event detection model for event detection to obtain the optimal event detection result.

[0222] In the specific implementation process, first, the data acquisition unit 301 acquires the initial event detection data set and performs preprocessing. The initial event detection data set includes a support set and a query set, and both the support set and the query set include several instances and their corresponding triggers and event types; the task construction unit 302 establishes a small-sample event detection incremental task set including several stages according to the preprocessed initial event detection data set; the model establishment unit 303 establishes an event detection model, and the event detection model includes an instance encoder, a type encoder, and a prototype network connected in sequence; the knowledge distillation unit 304 inputs the tasks of the small-sample event detection incremental task set into the initial event detection model in sequence by stages for continuous learning and training;

[0223] For the tasks in each stage, knowledge distillation is respectively performed on the output data of the prototype network in the current stage and the previous stage. Using the experience replay mechanism, the output data of the prototype network in the previous stage after knowledge distillation is saved as old-type knowledge, and the output data of the current stage after knowledge distillation is saved as new-type knowledge. For the tasks in each stage, the data update unit 305 updates the new-type knowledge with the old-type knowledge to obtain the embedded vector of the new type. According to the updated new type, event detection is performed on the data of the new task to obtain the updated event detection result of the new type. The model optimization unit 306 sets the total loss function. For the tasks in each stage, the event detection model is optimized using the total loss function. When the value of the total loss function is the smallest, the optimization of the event detection model is completed, and the optimal event detection model is obtained. Finally, the event detection unit 307 obtains the data to be event-detected and inputs it into the optimal event detection model for event detection to obtain the optimal event detection result.

[0224] This system establishes a few-shot event detection framework based on continual learning. When facing the continual learning of successive new tasks, it first learns the prototype representation of the new type, and then obtains the event detection result of the new type through the experience replay mechanism, knowledge distillation, and knowledge transfer between the old type and the new type. It can solve the forgetting problem of traditional models, enrich the connection between the new type and the old type, and improve the efficiency and accuracy of event detection.

[0225] The same or similar reference numerals correspond to the same or similar components.

[0226] The terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent.

[0227] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A few-shot event detection method based on continual learning, characterized in that, it includes the following steps: S1: Obtain the initial event detection dataset and perform preprocessing; the initial event detection dataset is specifically the RAMS, ACE, and LR-KBP datasets, including a support set S and a query set Q, and both the support set S and the query set Q include a number of instances and their corresponding triggers and event types: Among them, represents the i-th event mention in the support set S, including an instance , a trigger and an event type , where N is the number of event types in the support set S, and K is the number of instances included in each event type; represents the i-th event mention in the query set Q, including an instance , a trigger and an event type , where M is the number of instances in the query set Q; each instance or is represented as a word sequence , where L is the maximum length of the corresponding event mention; denote all the event types in the support set S and the query set Q as a set ; The specific method for preprocessing the initial event detection dataset is: divide the support set S and the query set Q into subsets; S2: Establish a few-shot event detection incremental task set including several stages according to the preprocessed initial event detection dataset; S3: Establish an event detection model, and the event detection model includes an instance encoder, a type encoder, and a prototype network connected in sequence; S4: Input the tasks of the few-shot event detection incremental task set into the initial event detection model stage by stage for continual learning training; For the tasks of each stage, perform knowledge distillation on the output data of the prototype network of the current stage and the previous stage respectively. Using the experience replay mechanism, save the output data of the prototype network of the previous stage after knowledge distillation as old type knowledge, and save the output data of the current stage after knowledge distillation as new type knowledge; S5: For the tasks of each stage, update the new type knowledge with the old type knowledge, obtain the embedding vector of the new type, and perform event detection on the data of the current stage task again according to the embedding vector of the new type to obtain the updated event detection result of the new type; S6: Set the total loss function, and for the tasks of each stage, optimize the event detection model using the total loss function. When the value of the total loss function is the smallest, complete the optimization of the event detection model and obtain the optimal event detection model; S7: Obtain the data to be event-detected and input it into the optimal event detection model for event detection to obtain the optimal event detection result.

2. A few-shot event detection method based on continual learning according to claim 1, characterized in that, in step S2, establishing a few-shot event detection incremental task set including several stages according to the preprocessed initial event detection dataset is specifically: The small-sample event detection incremental task set T includes tasks in t stages, expressed as: ; Take each subset in the support set S and the query set Q as the task of one stage in the few-shot event detection incremental task set.

3. A few-shot event detection method based on continual learning according to claim 2, characterized in that, the instance encoder in step S3 is specifically: Encode the input data using a preset token marking sequence, where the length of the preset token marking sequence is L, denoted as , and the trigger for the task in the t-th stage is denoted as ; Encode the trigger of the preset token tag sequence using a pre-trained BERT model to obtain the context representation of the trigger. The context representation of the trigger at the t-th stage is denoted as ; Embed the token of [CLS] as the context representation of the preset token sequence, and the token sequence of the i-th event type is denoted as ; During the event detection process, the context representation of the trigger at the t-th stage is used as an event trigger candidate to calculate the probability of the corresponding event type.

4. A few-shot event detection method based on continual learning according to claim 3, characterized in that, the type encoder in step S3 is specifically: Calculate the k-th event type according to the following formula prototype : Among them, is the number of instances in the event type.

5. A few-shot event detection method based on continual learning according to claim 4, characterized in that, the prototype network in step S3 is specifically: Based on the prototype of the input data Calculate the probability of the corresponding k-th event type The specific formula is as follows: where y is the context representation of the trigger in the t-th stage is the label of, denotes the Euclidean distance, is the total number of event types.

6. A few-shot event detection method based on continual learning according to claim 1 or 5, characterized in that, In step S4, for the tasks in each stage, knowledge distillation is performed on the output data of the prototype networks in the current stage and the previous stage respectively. Using the experience replay mechanism, the output data of the prototype network in the previous stage after knowledge distillation is saved as old-type knowledge, and the output data of the current stage after knowledge distillation is saved as new-type knowledge. The specific method is as follows: For the tasks in the t-th stage and the (t - 1)-th stage, knowledge distillation is performed on the output data of the prototype networks in the current stage and the previous stage respectively. Using the experience replay mechanism, k slots are allocated as storage spaces for each event type; The output data of the prototype network in the previous stage after knowledge distillation is saved as old-type knowledge, specifically: Among them, T 1 is the first temperature factor for knowledge distillation, satisfying T 1 > 1, is the output prediction of the previous stage after temperature factor adjustment, is the output prediction of the prototype network in the previous stage, is the output prediction of the current stage after temperature factor adjustment, is the output prediction of the prototype network in the current stage.

7. A few-shot event detection method based on continual learning according to claim 6, characterized in that, In step S5, for the tasks in each stage, the new-type knowledge is updated using the old-type knowledge. The specific method for obtaining the embedding vector of the new type is as follows: For the tasks in the t-th stage and the (t - 1)-th stage, use the old type of knowledge to update the prototype embedding vector of the new type of knowledge, and obtain the embedding vector of the new type Specifically: Among them, is a gate function that satisfies , and are the first and second hyperparameters respectively; is the first initialization vector, is the second initialization vector; The first initialization vector satisfies , where , ; Second initialization vector is a random initialization vector that satisfies , where .

8. A few-shot event detection method based on continual learning according to claim 7, characterized in that, The total loss function set in step S6 is specifically: Among them, is the total loss function value, is the self-training loss function value, is the knowledge distillation loss function value, is the cross-entropy loss function value, is the third hyperparameter; In the total loss function, the self-training loss function is specifically: Among them, is the event type in the task of the (t - 1)-th stage, is the pseudo-label of the old type of knowledge, satisfying , where is the second temperature factor, satisfying < 1; The knowledge distillation loss function is specifically: The cross-entropy loss function is specifically: where y is the context representation of the trigger in the t-th stage label of is the total number of event types 9. A few-shot event detection system based on continual learning, applying the few-shot event detection method described in any one of claims 1 to 8, characterized in that, including: Data acquisition unit: used to acquire the initial event detection dataset and perform preprocessing. The initial event detection dataset includes a support set and a query set. Both the support set and the query set include a number of instances and their corresponding triggers and event types; Task construction unit: used to establish an incremental task set for few-shot event detection including several stages according to the preprocessed initial event detection dataset; Model establishment unit: used to establish an event detection model. The event detection model includes an instance encoder, a type encoder, and a prototype network connected in sequence; Knowledge distillation unit: used to input the tasks of the few-shot event detection incremental task set into the initial event detection model stage by stage for continual learning training; For the tasks in each stage, knowledge distillation is performed on the output data of the prototype networks in the current stage and the previous stage respectively. Using the experience replay mechanism, the output data of the prototype network in the previous stage after knowledge distillation is saved as old-type knowledge, and the output data of the current stage after knowledge distillation is saved as new-type knowledge; Data update unit: used to update the new-type knowledge using the old-type knowledge for the tasks in each stage, obtain the embedding vector of the new type, and perform event detection on the data of the current stage task again according to the embedding vector of the new type to obtain the updated event detection result of the new type; Model optimization unit: used to set the total loss function, and optimize the event detection model using the total loss function for the tasks in each stage. When the value of the total loss function is the smallest, the optimization of the event detection model is completed, and the optimal event detection model is obtained; Event detection unit: used to obtain data to be detected for events and input it into the optimal event detection model for event detection, and obtain the optimal event detection result.

Citation Information

Patent Citations

  • Continuous small sample intention recognition method based on natural language prompt mechanism

    CN115688872A

  • Knowledge distillation and gradient pruning-based compression of artificial intelligence-based base caller

    WO2021168014A1