Low-resource event extraction method based on large language model enhancement
By using a large language model and event multi-level feature similarity fusion algorithm in low-resource scenarios, the problem of limited labeling data in event extraction tasks is solved, and high-quality event extraction and generalization capabilities are improved.
Patent Information
- Application Number
- CN202510310325.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-01
AI Technical Summary
In low-resource scenarios, event extraction tasks face the problems of limited labeling data and limited improvement of existing methods to extract the results.
A low-resource event extraction method based on a large language model is adopted. Through the pre-extraction mechanism, a multi-level feature similarity fusion algorithm is designed to use unlabeled text, sample samples with high similarity are selected, and a multi-dimensional constraint hint template is constructed to generate event samples maintained by dependencies, and finally filter the generated samples through AMR edge label matching.
It effectively alleviates the problem of scarce event labeling data in low-resource scenarios, improves the accuracy and generalization ability of event extraction, high quality of generated samples, and less computing resources.
Smart Images

Figure CN120234380A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence knowledge graphs, and particularly relates to a low-resource event extraction method enhanced by a large language model. Background Art
[0002] Event extraction is an important task in the field of information extraction, aiming to extract structured information from unstructured or semi-structured text. With the development of deep learning technology, event extraction tasks have achieved many important results in the fully supervised paradigm. The fully supervised paradigm relies on a large amount of labeled data. Event extraction tasks require more fine-grained annotation, and such annotation is often time-consuming and laborious. In low-resource scenarios, the labeled data is very limited. Training a model with limited labeled data to obtain better extraction results is an important challenge for this task. Currently, methods such as meta-learning and transfer learning are used to solve low-resource event extraction, but the differences in tasks, incomplete semantic understanding, and the neglect of dependencies lead to limited improvement in extraction by these methods. Data augmentation is a method of synthesizing more data from limited data, effectively alleviating the problem of insufficient labeled data. At the same time, by adding constraints of different dimensions and degrees, data augmentation can generate domain-related text, improving the generalization ability of the model in different tasks, domains, and contexts.
[0003] To solve the above problems, the current existing technologies are as follows:
[0004] Application No. CN202411025466.7, titled "Large Language Model Data Augmentation Method and Device for Event Extraction". This application obtains event patterns from the initial dataset and determines the sampling probability based on the frequency of event categories to ensure that low-frequency event categories have a higher sampling chance; uses the sampled data as context prompts and prompt templates to input into the large language model to generate new data, and screens high-quality data through preset evaluation metrics and the reasoning ability of the large language model. Finally, the screened data is stored in the initial dataset to form a target dataset. The characteristics of this method are: 1) The prompt template includes an event triple (event category, event trigger word, event argument). By constructing a trigger word dictionary and determining the event trigger word, diverse event triples are generated; 2) The generated data is screened through preset evaluation metrics (negative metric, information integrity metric, fact consistency metric) to ensure the quality of the generated data. Using the reasoning ability of the large language model, an inference chain is generated to further verify the logic and consistency of the generated data.
[0005] However, it has the following situations:
[0006] 1) Different initial sampling data: The "Data Augmentation Method and Device for Event Extraction Oriented Large Language Model" relies on the data distribution characteristics of the full dataset during initial sampling, and the sampling probability is related to the frequency of events in the dataset. This invention focuses on the event extraction problem in low-resource scenarios. The initial dataset has limited data and each event type is evenly distributed, without relying on event distribution.
[0007] 2) Limitations of evaluation metrics: Although the "Data Augmentation Method and Device for Event Extraction Oriented Large Language Model" introduces negative metrics, information integrity metrics, and fact consistency metrics to evaluate the quality of generated data, these metrics may not comprehensively cover all potential problems of the generated data. This invention uses an evaluation method that matches the shortest dependency path edges of the AMR graph consistent with the use and retrieval process, ensuring the consistency of the generation and screening processes.
[0008] 3) Complexity of inference chain generation: The "Data Augmentation Method and Device for Event Extraction Oriented Large Language Model" uses a large language model to generate inference chains and perform procedural event extraction. Although it can improve data quality, this process has a high computational complexity, which may increase time and resource consumption. Especially when applied to large-scale datasets, efficiency may become a bottleneck. This invention efficiently calculates the event similarity degree during example retrieval, can quickly locate examples, has higher generation efficiency, and consumes less computing resources.
[0009] The application number is CN202310858248.0, named "A Method and Device for Event Extraction Based on Information Enhancement and Prompt Learning". This application mainly optimizes the event extraction task in low-resource scenarios. Its core idea is to improve the effect of prompt learning through information enhancement strategies, thereby enhancing the accuracy and generalization ability of the event extraction model. The characteristics of this method are: 1) Using a lexical chain to construct historical event information to provide more context information for the model; counting high-frequency trigger words to enhance the event extraction model's ability to identify key events. 2) Combining a pre-trained language model (such as BART) for prompt learning, and identifying events through the method of target template matching to improve the model's adaptability to specific event categories. 3) Event extraction is based on the Lattice LSTM structure to improve the ability to capture event context relationships.
[0010] However, the following situations exist:
[0011] 1) Relying on manually designed enhancement strategies with limited generalization ability: Relying on lexical chains to construct historical event information and high-frequency trigger word optimization strategies requires a large amount of domain knowledge and manual statistics, and cannot automatically adapt to different tasks and domains. When the event category expands or the event data distribution changes, the model needs to re-collect statistical information, and the maintenance cost is high.
[0012] 2) Limitations in event argument recognition: Event argument extraction relies on Lattice LSTM, which can only capture local sequence relationships and performs poorly in dealing with complex event dependencies. At the same time, it lacks in-depth analysis of event structures and semantic dependencies, making it prone to errors when extracting complex events. Using BART for prompt learning, but its prompt information lacks the ability to model event reasoning and complex dependencies.
[0013] The present invention proposes a low-resource event extraction method based on a large language model. First, the large language model is used to pre-judge event types, extract trigger words and arguments in a large amount of unlabeled text and a small amount of labeled original samples to form an example search set; secondly, the labeled original samples are traversed, and the large language model is used to generate an argument word candidate set for each sample; then, the present invention designs an event multi-level feature similarity fusion algorithm, including a semantic embedding module, a structure parsing module and a dependency path parsing module, which calculate the semantic similarity, structure similarity and dependency path similarity of two event samples respectively, and select the top k samples with the highest similarity as examples for context learning after weighting the three similarities; combine the examples with the target samples to construct a prompt template that combines topic constraints, keyword constraints and dependency relationship constraints, and prompt the large model to generate diverse event samples with dependency relationships maintained; finally, screen the generated samples according to the matching degree of the dependency path edge labels, automatically label the samples and merge them into the original samples as an enhanced training set.
[0014] However, the following situations exist:
[0015] 4) Different initial sampling data: The "Data Augmentation Method and Device for Large Language Models for Event Extraction" relies on the data distribution characteristics of the full dataset during initial sampling, and the sampling probability is related to the frequency of events in the dataset. The present invention focuses on the event extraction problem in low-resource scenarios. The initial dataset has limited data and each event type is evenly distributed, without relying on event distribution.
[0016] 5) Limitations in evaluation metrics: Although the "Data Augmentation Method and Device for Large Language Models for Event Extraction" introduces negative metrics, information integrity metrics and fact consistency metrics to evaluate the quality of generated data, these metrics may not comprehensively cover all potential problems of the generated data. The present invention uses the evaluation method of matching the shortest dependency path edges of the AMR graph that is consistent with the use and retrieval process, ensuring the consistency of the generation and screening processes.
[0017] 6) Complexity of inference chain generation: The "Large Language Model Data Augmentation Method and Device for Event Extraction" uses a large language model to generate inference chains and perform process-based event extraction. Although it can improve data quality, this process has a high computational complexity, which may increase time and resource consumption. Especially when applied to large-scale datasets, efficiency may become a bottleneck. The present invention efficiently calculates the event similarity degree during example retrieval, can quickly locate examples, has higher generation efficiency, and consumes less computing resources.
[0018] The present invention proposes a low-resource event extraction method based on a large language model. First, use the large language model to pre-judge event types, extract trigger words and arguments from a large amount of unlabeled text and a small amount of labeled original samples to form an example search set; secondly, traverse the labeled original samples and use the large language model to generate an argument word candidate set for each sample; then, the present invention designs an event multi-level feature similarity fusion algorithm, including a semantic embedding module, a structure parsing module, and a dependency path parsing module, to calculate the semantic similarity, structure similarity, and dependency path similarity of two event samples respectively, and select the top k samples with the highest similarity as examples for context learning after weighting the three similarities; combine the examples with the target samples to construct a prompt template that combines topic constraints, keyword constraints, and dependency relationship constraints to prompt the large model to generate diverse event samples with dependency relationship preservation; finally, screen the generated samples according to the matching degree of the dependency path edge labels, automatically label the samples and merge them into the original samples as an enhanced training set. Summary of the Invention
[0019] To solve the above technical problems, the present invention proposes a low-resource event extraction method based on a large language model. First, the present invention uses a large language model to perform zero-shot event element pre-extraction on a large number of unlabeled samples and a small number of labeled original samples, constructing a search space for examples. Secondly, a prompt template is designed to generate an argument candidate set for each labeled original sample. Then, the present invention proposes an event multi-level feature similarity fusion algorithm for searching for pseudo-labeled events close to the target event in the example search space. Specifically, the algorithm is divided into three modules. The semantic embedding module uses Sentence-Bert to calculate the semantic similarity between sample texts. The structure parsing module calculates the event structure similarity according to the proportion of the event type and argument role labels of two samples in all labels. The dependency path similarity uses AMR to model the event samples as a graph structure, constructs a set of dependency paths from the trigger word to all arguments using the shortest path algorithm in the graph, and calculates the similarity of the dependency path sets of two samples to obtain the dependency path similarity. After weighting the three similarities, the top k samples with the highest similarity are selected as context learning examples. Subsequently, the searched example samples are combined with the target sample to construct a prompt template, and event topic constraints, keyword constraints, and dependency relationships between the trigger word and each argument are fused in the prompt. After inputting into the large model, generated samples are obtained. Finally, a scoring function based on the matching degree of AMR edge labels of the dependency path set is proposed to screen the generated samples, and samples with high scores are selected for automatic annotation and merged into the original samples, finally obtaining an enhanced training set with dependency preservation.
[0020] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0021] A low-resource event extraction method enhanced based on a large language model, the specific steps are as follows:
[0022] 1) Event sample pre-extraction based on a large language model;
[0023] For unlabeled texts and a small number of labeled event samples, a zero-shot event sample pre-extraction mechanism is proposed, and a prompt template is designed to prompt the large language model to predict the event type of the unlabeled text and pre-extract the trigger word and arguments, constructing a search space for examples;
[0024] 2) Construction of argument candidate sets based on a large language model;
[0025] For each labeled event sample, a prompt template is designed to construct an argument word candidate set for a single sample in the form of filling in the blanks;
[0026] 3) Example search based on the event multi-level feature similarity fusion algorithm;
[0027] A sample search mechanism based on the fusion of multi-level event feature similarities is proposed to retrieve D samples with the highest similarity to the annotated event samples in the constructed sample search space from three aspects: text semantics, event structure, and the dependency relationship between trigger words and arguments.
[0028] 4) Generation of dependency-preserving event samples based on large language models;
[0029] Design a prompt template that explicitly shows the dependency relationship between event trigger words and arguments. Use the d highly similar samples searched in 3) as example samples, and use the trigger words and arguments randomly sampled and combined in 2) as keyword constraints to prompt the large language model to generate m event samples with rich context and maintaining the dependency between event trigger words and arguments.
[0030] 5) Generation sample screening and automatic annotation;
[0031] Design a text generation quality scoring function based on edge label matching. Among the m event samples generated in 4), select the r samples with the highest scores as high-quality event samples, and annotate the event type, trigger words, and arguments for the high-quality event samples.
[0032] As a further improvement of the present invention, step 1) pre-extraction of event samples based on large language models is as follows;
[0033] Given an annotated event sample set containing N event types, with K samples for each event type Given a set containing several unannotated texts For each sample in;
[0034] First, use the sample text as input to prompt the large model to perform a binary classification task of whether an event is contained.
[0035] Secondly, for samples containing events, input the text and candidate event types, and prompt the large model to perform a classification task of judging the event type in the form of multiple-choice. The output result of the large model should be stored in the form of an array.
[0036] Then, traverse each event type predicted by the large model in the previous step, and prompt the large model to extract the trigger words and arguments in the text in the form of table filling, and match the arguments and the corresponding argument roles.
[0037] Finally, organize the results pre-extracted by the large model to form a sample search space
[0038] As a further improvement of the present invention, step 2) construction of argument candidate sets based on large language models is as follows;
[0039] Traverse the set of labeled event samples For each sample in the set, traverse all the arguments of the sample and prompt the large model to generate a set of synonymous candidates for the argument
[0040] As a further improvement of the present invention, step 3) example search based on the event multi-level feature similarity fusion algorithm is specifically divided into three modules;
[0041] First, the semantic embedding module, based on Sentence-BERT, encodes and to obtain the embedding representation of the text, and uses cosine similarity to calculate the text semantic similarity degree:
[0042]
[0043] Secondly, the structure parsing module determines the structural elements of the event based on the event type and its argument roles in the text; calculates the number of identical structural elements in two event samples, and calculates the ratio of this number to the total number of all structural elements; normalizes the obtained ratio to obtain the event structure similarity:
[0044]
[0045] where e stru represents the set of type labels and argument role labels of the event sample;
[0046] Then, the dependency path parsing module models the event text as a directed graph structure using the abstract semantic representation AMR; aligns the trigger word and arguments of the event with the nodes in the AMR graph; takes the trigger word as the central node and extracts the shortest path from the trigger word to each argument, and the path contains the node V; finally obtains the dependency path set of the event sample:
[0047]
[0048] P is the text form of the obtained dependency path set; uses Glove to encode the nodes in the dependency path and performs average pooling on the encoding results to obtain the argument-trigger word dependency path representation, and the dimension of the representation is 1×300:
[0049]
[0050] Iteratively calculate all the dependency paths in the sample to generate a dependency path representation matrix M for a single sample, and the dimension of the matrix is p×300, where p is the number of dependency paths;
[0051] Dependency path representation matrices M of two samples m×300 and N n×300 , and normalize them:
[0052]
[0053] Calculate the cosine similarity for the normalized matrices M' and N':
[0054] S = M · N T
[0055] where S is an m×n similarity matrix, and each element S i,j represents the similarity between the i-th vector of M and the j-th vector of N, with a value range of [-1, 1];
[0056] Take the average of all values in S as the similarity of the final dependency path sets of the two event samples:
[0057]
[0058] Finally, based on the semantic similarity structural similarity and the dependency path set similarity calculate the final similarity degree of the two event samples, and take the top k samples as the final examples:
[0059]
[0060] where α, β, γ are hyperparameters and satisfy α + β + γ = 1.
[0061] As a further improvement of the present invention, step 4) generating dependency-preserving event samples based on a large language model is specifically as follows;
[0062] Given a target event sample, construct a prompt, input it into the large model to generate r samples, with the event type of the target event sample as the topic constraint; randomly sample from the argument candidate set pre-generated in step 2) and combine it with the trigger word to form the keyword constraint of the prompt; describe the dependency relationship between event elements in the form that the "relationship between the 'argument' and the 'trigger word' is the 'argument role'" as the dependency-preserving semantic constraint; input the above-constructed prompt information together with the example samples searched in step 3) into the large language model; generate m enhanced samples that meet the above constraints through the large language model.
[0063] As a further improvement of the present invention, step 5) sample screening and automatic annotation are specifically as follows;
[0064] Model the generated samples as an AMR directed graph, where nodes represent semantic concepts and directed edges represent semantic relationships; align the trigger words and arguments of the event with the nodes in the AMR graph, and calculate the shortest path from the trigger word to each argument node; design a quality evaluation function based on edge label matching, which needs to consider the matching degree between the edge labels in the path and the target event frame and the compliance of the edge direction attribute with the event structure; calculate the comprehensive score of each candidate sample according to the quality evaluation function, and select the top r samples with the highest scores as high-quality event samples; perform automated expression on the selected high-quality samples, including: the same event type as the target event sample, locate and label the trigger words, and locate the arguments and their semantic roles. The labeled samples and the initial dataset jointly form an enhanced event dataset. Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0065] The present invention proposes a low-resource event extraction method based on a large language model, which alleviates the problem of scarce event annotation data in low-resource scenarios through the understanding ability, reasoning ability and generation ability of the large language model, and fully utilizes unlabeled text through a pre-extraction mechanism. The example search mechanism based on the event multi-level feature similarity fusion algorithm flexibly and efficiently selects examples, ensuring that the examples carry different granularity information from the text level to the token level in terms of semantics, structure and dependency relationship, and fully stimulating the emergence ability of the large language model. In terms of the construction of the prompt words, multi-dimensional constraints are added to control the generation results of the large language model, so that the generated text satisfies diversity while maintaining the dependency relationship from the trigger word to the argument. The generated sample screening mechanism based on AMR edge label matching further ensures the consistency of the dependencies between the generated samples and the target samples. Therefore, the present invention has good application prospects and a wide range of promotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is the logical flow chart of the method of the present invention;
[0067] Figure 2 is the sub-flow chart of the event pre-extraction process of the method of the present invention;
[0068] Figure 3 is the sub-flow chart of the construction process of the argument candidate set of the method of the present invention;
[0069] Figure 4 is the sub-flow chart of the event multi-level feature similarity fusion example search process of the method of the present invention;
[0070] Figure 5 is the sub-flow chart of the dependency-preserving event generation process of the method of the present invention;
[0071] Figure 6 is the sub-flow chart of the AMR edge label matching screening process of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments:
[0073] The present invention proposes a low-resource event extraction method based on a large language model, which alleviates the problem of scarce event annotation data in low-resource scenarios through the understanding ability, reasoning ability, and generation ability of the large language model, and makes full use of unannotated text through a pre-extraction mechanism. The example search mechanism based on the event multi-level feature similarity fusion algorithm flexibly and efficiently selects examples, ensuring that the examples carry different granularity information from the text level to the token level in terms of semantics, structure, and dependency relationship, and fully stimulating the emergence ability of the large language model. In the construction of the prompt, multi-dimensional constraints are added to control the generation results of the large language model, so that the generated text satisfies diversity while maintaining the dependency relationship from the trigger word to the argument. The generated sample screening mechanism based on AMR edge label matching further ensures the consistency of the dependencies between the generated samples and the target samples. Therefore, the present invention has good application prospects and a wide range of promotion.
[0074] As a specific embodiment of the present invention, the present invention provides a logic flowchart as Figure 1 shown, and the sub-flowchart of the event pre-extraction process is as Figure 2 shown. This process inputs low-resource annotated samples and a large number of unannotated samples, uses the prompt template shown in the figure to prompt the large language model to pre-extract event trigger words and arguments, and outputs annotated pseudo-event samples as the example search space; the sub-flowchart of the argument candidate set construction process is as Figure 3 shown. This process inputs low-resource annotated samples and constructs a single-sample argument candidate set using the prompt template shown in the figure; the sub-flowchart of the event multi-level feature similarity fusion example search process is as Figure 4 shown. This process inputs low-resource annotated samples and pseudo-annotated event samples, and outputs the sample closest to the original sample as an example after calculating the similarity through a semantic embedding calculation model, a structure parsing module, and a dependency path parsing module; the sub-flowchart of the dependency-preserving event generation process is as Figure 5 shown. This process inputs examples, low-resource annotated samples, and a candidate set of paths, and uses the prompt template shown in the figure to prompt the large language model to generate dependency-preserving event samples; the sub-flowchart of the AMR edge label matching screening process is as Figure 6 shown. This process inputs the generated event samples and outputs the sample with the highest matching degree after AMR edge label matching. A low-resource event extraction method based on a large language model includes the following steps:
[0075] 1) Event sample pre-extraction based on a large language model;
[0076] Given an annotated event sample set containing N event types, with K samples for each event type Given a set containing several unannotated texts For each sample in ;
[0077] First, using the sample text as input, prompt the large language model to perform a binary classification task of judging whether the text contains an event. The prompt template is:
[0078] "Please judge whether the given text contains an event. Just answer 'yes' or 'no', without any additional information.
[0079] Text: ";
[0080] Secondly, for samples containing events, input the text and candidate event types, and prompt the large language model to perform a classification task of judging the event type in the form of multiple-choice. The output result of the large language model should be stored in the form of an array. The prompt template is:
[0081] "Please judge which event type the event in the given text belongs to. The text contains at least one event type. Please organize the results in the form of an array.
[0082] Text:
[0083] The event types to be selected are: Conflict: Attack; Life: Death;...; Business: Merger of organizations";
[0084] Then, traverse each event type predicted by the large language model in the previous step, and prompt the large language model to extract the trigger word and arguments in the text in the form of filling in a table, and match the arguments and the corresponding argument roles. The prompt template is:
[0085] "The text contains a 'Conflict: Attack' event. Please extract the trigger word and the corresponding arguments from the text according to the given event type and argument roles. Complete this task in the form of filling in a table. Note that the trigger word and the arguments are both sub-fragments of the text. If no argument that undertakes the argument role can be found in the text, use 'None' to represent it.
[0086] Text:
[0087] Event type:
[0088] The trigger word is:
[0089] Location:
[0090] Attacker:
[0091] Attack target: ";
[0092] Finally, organize the events pre-extracted by the large language model into Json format to form an example search space
[0093] 2) Construction of the argument candidate set based on the large language model;
[0094] Traverse the samples in the set of annotated event samples For each sample, traverse all the arguments of the sample and prompt the large model to generate a set of synonymous candidates for the argument The prompt template is as follows:
[0095] "Given that the text contains a 'Conflict: Attack' event, generate 10 text fragments such that the generated text fragments can still act as the 'Location' role in the original sentence. Note that if the original fragment is a word, a word needs to be generated, and if the original fragment is a text fragment, fragments with the same number of words as the original fragment need to be generated.
[0096] Text:
[0097] Original argument:
[0098] Argument role: ";
[0099] The purpose of constructing the argument candidate set is to increase the diversity of event arguments, while ensuring that the arguments fit the argument characteristics of the original sample as much as possible and maintaining the argument boundaries as much as possible.
[0100] 3) Example search based on the event multi-level feature similarity fusion algorithm;
[0101] First, the semantic embedding module, based on Sentence-BERT, encodes and to obtain the embedding representation of the text, and uses cosine similarity to calculate the text semantic similarity degree:
[0102]
[0103] Secondly, the structure parsing module, based on the event type and its argument role in the text, determines the structural elements of the event; calculates the number of identical structural elements in two event samples, and calculates the ratio of this number to the total number of all structural elements; normalizes the obtained ratio to obtain the event structure similarity:
[0104]
[0105] where e stru represents the set of type labels and argument role labels of the event sample;
[0106] Then, the dependency path parsing module uses the abstract semantic representation (AMR) to model the event text as a directed graph structure; aligns the trigger word and arguments of the event with the nodes in the AMR graph; takes the trigger word as the central node and extracts the shortest path from the trigger word to each argument, and the path contains the node V; finally obtains the dependency path set of the event sample:
[0107]
[0108] P is the text form of the obtained set of dependency paths; the nodes in the dependency paths are encoded using Glove And average pooling is performed on the encoding results to obtain the dependency path representation of the argument-trigger word, and the dimension of the representation is 1×300:
[0109]
[0110] Iterative calculation is performed on all the dependency paths in the sample to generate the dependency path representation matrix M of a single sample, and the dimension of the matrix is p×300, where p is the number of dependency paths;
[0111] The dependency path representation matrices M m×300 and N n×300 of two samples are normalized:
[0112]
[0113] The cosine similarity is calculated for the normalized matrices M' and N':
[0114] S = M·N T
[0115] where S is a similarity matrix of m×n, and each element S i,j represents the similarity between the i-th vector of M and the j-th vector of N, and the value range is [-1, 1];
[0116] The average value of all the values in S is taken as the similarity of the dependency path sets of the final two event samples:
[0117]
[0118] Finally, based on the semantic similarity structural similarity and the dependency path set similarity the final similarity degree of the two event samples is calculated, and the top k samples are taken as the final examples:
[0119]
[0120] where α, β, and γ are hyperparameters and satisfy α + β + γ = 1. When searching for samples, the values of the hyperparameters can be dynamically adjusted to find the best-performing examples.
[0121] 4) Generation of dependency-preserving event samples based on large language models;
[0122] Given a target event sample, construct a prompt. Input it into a large model to generate r samples, with the event type of the target event sample as the topic constraint; randomly sample from the set of pre-generated argument candidates in step 2) and combine them with the trigger word to form the keyword constraint of the prompt; describe the dependency relationship between event elements in the form that the "relationship between the argument and the trigger word" is the "argument role", as the semantic constraint for maintaining dependencies; input the above-constructed prompt information together with the example samples retrieved in step 3) into the large language model; generate m enhanced samples that meet the above constraints through the large language model. The specific prompt is as follows:
[0123] "Example 1
[0124] Example 2
[0125] Generate 10 sentences, each containing a 'conflict: attack' event.
[0126] The requirements for the generated sentence format are:
[0127] 1. Do not modify the content of the trigger word and the arguments, but the context of the sentence should be as rich as possible;
[0128] 2. Only generate sentences and do not generate other redundant information;
[0129] 3. The number of words in the generated sentences should be close to the original sentences;
[0130] 4. Organize the results in an array form.
[0131] The requirements for the content of the generated sentences are:
[0132] 1. The keywords that need to be included in the sentence are: trigger word, argument 1, argument 2,..., argument n;
[0133] 2. The trigger word is:
[0134] 3. The arguments are:
[0135] 4. The constraint relationship is: the relationship between 'argument 1' and the 'trigger word' is 'argument role 1', the relationship between 'argument 2' and the 'trigger word' is 'argument role 2',..., the relationship between 'argument n' and the 'trigger word' is 'argument role n';
[0136] Output: "
[0137] 5) Generate sample screening and automated annotation;
[0138] Model the generated samples as an AMR directed graph, where nodes represent semantic concepts and directed edges represent semantic relationships; align the trigger words and arguments of the event with the nodes in the AMR graph, and calculate the shortest path from each argument to the trigger word; design a quality evaluation function based on edge label matching. Since the constructed AMR graph is a directed graph, if the argument points to the trigger word, the direction is positive, and if the trigger word points to the argument, the direction is negative. The function needs to consider the matching degree between the edge labels in the path and the target event frame, as well as the compliance between the direction attribute of the edge and the event structure. The scoring function is:
[0139]
[0140] where n represents the number of argument-trigger word paths for the event sample, p edge ={p edge (w arg ,w tri )|p edge (w arg ,w tri )=[e arg-o1 ,e o1-o2 ,...,e ok-tri ,e∈E}. Calculate the comprehensive score of each candidate sample according to the quality evaluation function, and select the top r samples with the highest scores as high-quality event samples; perform automatic annotation on the selected high-quality samples, including: the same event type as the target event sample, locate and annotate the trigger word, locate the argument and its semantic role. The annotated samples and the initial dataset jointly form an enhanced event dataset.
[0141]
Example 1
[0142] In the example, the low-resource event extraction method based on the large language model is used to generate, train, and predict on the general datasets ACE05-EN and ACE05-EN+, and the same data as in this example is used in all other examples. The ACE05-EN dataset has a total of 18,927 samples, of which 5,055 samples contain events, with 33 event types. Among them, the validation set has 450 event samples, and the test set has 403 event samples. To construct a low-resource scenario, an N-way-K-shot low-resource training set is constructed from 4,202 event samples in the training set, where N is 33 event types and K is an integer from 1 to 5. The ACE05-EN+ dataset has a total of 20,793 samples, of which 5,311 samples contain events, with 33 event types. Among them, the validation set has 468 event samples, and the test set has 424 event samples. An N-way-K-shot low-resource training set is constructed from 4,419 selected online samples. Since the model has good robustness and generalization, the same hyperparameter settings can be used in different general datasets.
[0143] The specific implementation is as follows: The large language model used in this invention is ChatGPT-4, and the AMR modeling tool used is the open-source tool Hanlp. Randomly sample 2000 pieces of data from the training set and combine it with the low-resource training set as the dataset for large language pre-extracted events. The samples containing events after pre-extraction constitute the search space of the examples. For each event in the N-way-K-shot low-resource dataset, generate its argument candidate set, and the number of candidate words for each argument is 10. When searching for examples, the weights of semantic similarity, structural similarity, and dependency path similarity are α = 0.2, β = 0.4, and γ = 0.4 respectively, and the samples with the top 2 similarity are selected as examples. After the examples and the target event samples are input into the large language model, the generated 10 samples are screened. Finally, the top 3 samples that meet the dependencies are selected from each event sample as enhanced event samples, which are combined with the initial low-resource dataset, and the ratio of the original samples to the enhanced samples is 1:3.
[0144] To verify the effectiveness of the enhanced samples for the event extraction task, use the extraction model based on sequence labeling in the open-source event extraction project OmniEvent for training. The extraction task is divided into event detection (ED) and event argument extraction (EAE). The pre-trained language model for both tasks is roberta-large. In the training stage of the event detection task, use adamw_torch as the optimizer, set the learning rate to 5.0e-5, set the batch size for each training to 8, and set the number of training epochs to 30. The evaluation metric for this task is the micro-F1 value of trigger word classification; in the training stage of the event argument extraction task, use adamw_torch as the optimizer, set the learning rate to 7.0e-5, set the batch size for each training to 32, and set the number of training epochs to 40. The evaluation metric for this task is the micro-F1 value of argument classification.
[0145] The trained event extraction model is applied to the test set, and the results of the event trigger words and event arguments extracted by the model are compared with the actual results. It is found that on the ACE05-EN dataset, at an augmentation ratio of 1:3 and the initial dataset size of k-shot, the micro-F1 of event trigger word classification in the event detection task is increased by 45.63%, 5.99%, 7.09%, 4.95%, and 8.01% respectively, and the micro-F1 of argument classification in the event argument extraction task is increased by 15.38%, 32.62%, 5.5%, 16.59%, and 6.41% respectively. On the ACE05-EN+ dataset, at an augmentation ratio of 1:3 and the initial dataset size of k-shot, the micro-F1 of event trigger word classification in the event detection task is increased by 54.16%, 5.75%, 5.44%, 3.39%, and 0.76% respectively, and the micro-F1 of argument classification in the event argument extraction task is increased by 16.39%, 35.28%, 13.00%, 8.25%, and 10.49% respectively.
[0146]
Example 2
[0147] The example search algorithm of the event multi-level feature similarity fusion algorithm for the low-resource event extraction method based on large language models can significantly improve the performance of event extraction. The example samples selected by this module contain rich information on semantics, event structure, and the dependency path from arguments to trigger words, and can provide multi-granularity event patterns for the large language model to learn. Taking the ACE05-EN+ dataset as an example, during generation, using this algorithm to select examples increases the micro-F1 value of trigger word classification in the event detection task by 0.76% - 5.75% compared with randomly selecting examples, and increases the micro-F1 value of argument classification in the event argument extraction task by 2.13% - 8.76%. In practical applications, since this module does not need to train a deep neural network and the similarity can be cached, the processing of the original data can be quickly realized to ensure the efficiency of this method in real scenarios.
[0148]
Example 3
[0149] The screening method based on AMR dependency path edge label matching for the generated samples in the low-resource event extraction method based on large language models plays a crucial role in identifying higher-quality generated samples. The screening process considers the significant role of edge labels in maintaining relationships in the dependency path from the argument to the trigger word, and the set scoring function can distinguish the generated samples. Taking the ACE05-EN+ dataset as an example, the generated samples selected after being screened by this module have an increase of 1.54% - 6.26% in the micro-F1 value of trigger word classification in the event detection task, and an increase of 2.14% - 4.34% in the micro-F1 value of argument classification in the event argument extraction task compared to randomly selected generated samples. It meets the requirements for screening high-quality samples in actual application scenarios.
[0150] The above are only the preferred embodiments of the present invention, and do not impose any other form of limitation on the present invention. Any modification or equivalent change made according to the technical essence of the present invention still falls within the scope claimed by the present invention.
Claims
1. A low-resource event extraction method based on large language model enhancement, the specific steps are as follows, characterized in that: 1) Event sample pre-extraction based on large language model; For unlabeled text and a small number of labeled event samples, a zero-shot event sample pre-extraction mechanism is proposed. A prompt template is designed to prompt the large language model to predict the event type for unlabeled text and pre-extract trigger words and arguments to construct an example search space. 2) Construction of argument candidate sets based on large language models; For each annotated event sample, a prompt template is designed to construct a candidate set of argument words for a single sample in the form of filling in the blanks; 3) Example search based on event multi-level feature similarity fusion algorithm; An example search mechanism based on event multi-level feature similarity fusion is proposed. From the three aspects of text semantics, event structure, and trigger word and argument dependency, D samples with the highest similarity to the annotated event samples are retrieved in the example search space constructed in 1). 4) Dependency-preserving event sample generation based on large language models; In designing a prompt template that explicitly indicates the dependency between event trigger words and arguments, we use the d high-similarity samples searched in 3) as example samples, and the trigger words and arguments randomly sampled from 2) as keyword constraints to prompt the large language model to generate m event samples with rich contexts that maintain the dependency between event trigger words and arguments. 5) Generate sample screening and automatic annotation; Design a text generation quality scoring function based on edge label matching. Among the m event samples generated in 4), select the r samples with the highest scores as high-quality event samples, and annotate the event type, trigger word, and argument for the high-quality event samples.
2. According to claim 1, a low-resource event extraction method based on large language model enhancement is characterized in that: Step 1) Pre-extraction of event samples based on a large language model. The specific steps are as follows: Given a set of labeled event samples containing N event types and K samples of each event type Given a collection of unlabeled texts right Each sample in First, the sample text is used as input to prompt the large model to perform a binary classification task of whether it contains events; Secondly, for samples containing events, input text and candidate event types, and prompt the large model to perform the classification task of determining the event type in the form of multiple choices. The output results of the large model should be stored in the form of an array; Then, traverse each event type predicted by the big model in the previous step, prompt the big model to extract the trigger words and arguments in the text in the form of a table, and match the arguments with the corresponding argument roles; Finally, the results of the large model pre-extraction are sorted to form an example search space 3. The low-resource event extraction method based on large language model enhancement according to claim 1 is characterized in that: Step 2) constructing an argument candidate set based on the large language model. The specific steps are as follows: Traverse the set of labeled event samples For each sample, traverse all the arguments of the sample and prompt the large model to generate a synonym candidate set for the argument.
4. The low-resource event extraction method based on large language model enhancement according to claim 1 is characterized in that: The step 3) is based on the example search of the event multi-level feature similarity fusion algorithm, which is specifically divided into three modules; First, the semantic embedding module, based on Sentence-BERT, encodes and Get the embedded representation of the text and use cosine similarity to calculate the semantic similarity of the text: Secondly, the structural analysis module determines the structural elements of the event based on the event type and its argument roles in the text; calculates the number of identical structural elements in two event samples and calculates the ratio of this number to the total number of all structural elements; and normalizes the obtained ratio to obtain the event structure similarity: Among them, e stru Represents the type label and argument role label set of event samples; Then, the dependency path parsing module uses the abstract semantic representation AMR to model the event text as a directed graph structure; aligns the event trigger words and arguments with the nodes in the AMR graph; uses the trigger word as the central node, extracts the shortest path from the trigger word to each argument, and the path contains the node V; finally, the dependency path set of the event sample is obtained: P is the text form of the obtained dependency path set; Glove is used to encode the nodes in the dependency path. The encoding results are average pooled to obtain the argument-trigger dependency path representation, the dimension of which is 1×300: Iteratively calculate all dependency paths in the sample to generate a dependency path representation matrix M of a single sample, where the dimension of the matrix is p×300, where p is the number of dependency paths; The dependency path representation matrix M of two samples m×300 and N n×300 , normalize it: Calculate the cosine similarity of the normalized matrices M' and N': S=M·N T Among them, S is an m×n similarity matrix, in which each element S i,j Represents the similarity between the i-th vector of M and the j-th vector of N, with a value range of [-1,1]; Take the average value of all values in S as the similarity of the final two event sample dependency path sets: Finally, based on the semantic similarity Structural similarity Similarity with dependent path set Calculate the final similarity between two event samples and take the topk samples as the final example: Among them, α, β, and γ are hyperparameters, and they satisfy α+β+γ=1.
5. The low-resource event extraction method based on large language model enhancement according to claim 1 is characterized in that: The step 4) generates event samples based on the dependency preservation of the large language model, and the specific process is as follows; Given a target event sample, construct a prompt and input it into the large model to generate r samples, with the event type of the target event sample as the topic constraint; randomly sample from the argument candidate set pre-generated in step 2) and combine it with the trigger word to form a keyword constraint for the prompt; describe the dependency relationship between event elements in the form of "argument" and "trigger word" as "argument role" as a semantic constraint for dependency preservation; input the above-constructed prompt information and the example sample searched in step 3) into the large language model together; generate m enhanced samples that meet the above constraints through the large language model.
6. The low-resource event extraction method based on large language model enhancement according to claim 1 is characterized in that: Step 5) generates sample screening and automatic annotation, and the specific process is as follows; The generated samples are modeled as an AMR directed graph, where nodes represent semantic concepts and directed edges represent semantic relationships; the trigger words and arguments of the event are aligned with the nodes in the AMR graph, and the shortest path from the trigger word to each argument node is calculated; a quality evaluation function based on edge label matching is designed, which needs to consider the matching degree between the edge label in the path and the target event framework and the conformity of the edge direction attribute with the event structure; the comprehensive score of each candidate sample is calculated according to the quality evaluation function, and the r samples with the highest scores are selected as high-quality event samples; Automated description is performed on the selected high-quality samples, including: the same event type as the target event sample, locating and labeling trigger words, locating arguments and their semantic roles. The labeled samples and the initial dataset together constitute an enhanced event dataset.
Citation Information
Patent Citations
Prompt learning event extraction method and device based on information enhancement
CN116911254A
Large language model data enhancement method and device for event extraction
CN118551194B
Cited By
Event extraction-oriented labeling guide automatic generation and iterative optimization method
CN121524707A
An Automatic Generation and Iterative Optimization Method for Annotation Guidelines Oriented to Event Extraction
CN121524707B