Semantic injection enhancement-based event extraction model construction method and event extraction method
By introducing semantic injection enhancement methods in event extraction, semantic features are extracted using event knowledge prototypes and small models, combined with the loss optimization strategy of semantic anchors, the problems of simplicity of semantic representation and limited generalization ability of event extraction in the existing technology are solved, significantly improving the accuracy and performance of event extraction.
Patent Information
- Application Number
- CN202510639331.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
AI Technical Summary
In the event extraction, the prior art there are problems such as simple semantic representation, limited generalization capability, requiring a large number of intermediate and result annotations, and inadequate task performance in complex and variable scenarios and sparse data.
By introducing semantic injection enhancement methods, the event knowledge prototype is obtained to obtain the key semantics of event types, and small models such as BERT extract semantic features from event text, construct semantic enhancement tables, and directly integrate these semantic information into the task prompt template. At the same time, the loss optimization strategy combined with semantic anchor points is designed to optimize the training of large language models.
It improves the accuracy and generalization ability of large language models in event extraction, reduces the dependence on a large number of intermediate and result annotations, and significantly improves the event extraction performance in complex scenarios and sparse data.
Smart Images

Figure CN120163231A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and particularly to the construction of an event extraction model enhanced by semantic injection and an event extraction method. Background Art
[0002] Event Extraction (EE) is a key task for identifying event types and their related elements from unstructured text, and has broad application prospects in fields such as information retrieval and knowledge graph construction. Traditional methods rely on static models and a large amount of labeled data, and perform supervised training through rule matching or feature-based machine learning algorithms, especially deep learning. Common algorithms include sequence labeling methods based on conditional random fields (CRF) and pre-trained models (such as BERT). These methods mostly use cross-entropy loss and require intermediate steps such as extracting event trigger words. Therefore, the semantic representation of traditional methods is relatively simple, the generalization ability is limited, and a large amount of intermediate and result annotations are required. Especially in complex and changeable scenarios and data-sparse situations, the task performance is insufficient.
[0003] In recent years, large language models (LLMs) have brought new ideas and methods for event extraction with their powerful generation capabilities and semantic understanding potential. On pre-trained general language large models, by directly inputting prompts, structured outputs can be generated, such as "Attack#we_attacker;". This method simplifies the event extraction process, avoids complex intermediate steps, and shows great potential for performance improvement, and has broad application prospects in fields such as news event monitoring and financial risk analysis.
[0004] However, in practical applications, after directly injecting task semantics through prompts, there are still many challenges in optimizing the performance of the event extraction task through efficient fine-tuning on large models: 1) Lack of deep semantic optimization: The commonly used cross-entropy loss in large models only focuses on token-level predictions and ignores the deep semantics of event types. For example, when extracting events related to "economic crisis", the model may simply match the keywords in the text, but cannot deeply understand the complex economic phenomena and semantic relationships contained in the event type of "economic crisis", resulting in inaccurate extraction results.
[0005] 2) Problem of sparse labeled data: The scarcity of high-quality labeled data severely restricts the model performance. Obtaining a large amount of high-quality labeled data is costly. In the case of insufficient data, it is difficult for the model to learn comprehensive and accurate event patterns and features, and the extraction ability for complex events is greatly reduced. For example, in some emerging fields or specific scenarios, due to the lack of sufficient labeled data, the model cannot effectively identify and extract relevant events.
[0006] 3) Few studies on the fusion of fine-tuning strategies: There are few studies on the fusion of large model fine-tuning strategies optimized for specific tasks. Different fine-tuning strategies have their own advantages and disadvantages for different tasks and data, but currently there is a lack of effective methods to organically combine multiple fine-tuning strategies to give full play to the advantages of large models and improve event extraction performance. Summary of the Invention
[0007] The purpose of the present invention is to provide a method for constructing an event extraction model based on semantic injection enhancement and an event extraction method, which improves the event extraction ability of large language models by introducing semantic injection enhancement.
[0008] To achieve the above purpose, the present technical solution provides a method for constructing an event extraction model based on semantic injection enhancement, including the following steps: S1: Obtain at least one event text, and determine the event type and event role table of the event text; S2: Construct a semantic enhancement table based on dataset samples of different event types; S3: Construct a prompt template based on the task description, event type, event role table, semantic enhancement table, and structured output constraints that define the event extraction task, and use the event text and the corresponding prompt template as the fine-tuning dataset, where the structured output constraints specify the output content and output format of the event extraction task; S4: Perform enhancement processing on the fine-tuning dataset; S5: Train the output content of the large language model using the fine-tuning dataset until the loss function meets the requirements, where the loss function used for training the large language model is the combined value of the prototype loss and the cross-entropy loss, and the prototype loss reflecting the prototype structure problem is calculated from the predicted event type extracted from the output content of the large language model and the semantic anchor dictionary constructed based on the semantic enhancement table.
[0009] The present technical solution provides an event extraction method based on semantic injection enhancement. The event text to be subjected to event extraction is input into the event extraction model based on semantic injection enhancement constructed by the method for constructing an event extraction model based on semantic injection enhancement to obtain the event extraction result.
[0010] The present technical solution also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method for constructing an event extraction model based on semantic injection enhancement or the event extraction method based on semantic injection enhancement.
[0011] Compared with the prior art, the present technical solution has the following characteristics and beneficial effects: 1. Propose a semantic injection enhancement method: For the event extraction task, on the one hand, obtain the key semantics of relevant event types from the event knowledge prototype, such as extracting the core semantics of events like "contract disputes" from the event knowledge prototype in the legal field; on the other hand, use small models like BERT to extract semantic features from event texts, such as vocabulary and sentence structure-related features, and directly integrate these rich semantics into the task prompt template as a semantic enhancement table to improve the model's understanding of event texts and enhance the accuracy of event extraction.
[0012] 2. Design a loss optimization strategy combined with semantic anchors: Transform event extraction into a question-and-answer pair form to fit the generation ability of large language models. Construct a semantic anchor dictionary. For example, determine semantic anchors such as "earthquake magnitude" for the "natural disaster" event type. During the training of large language models, its loss function is composed of the cross-entropy loss based on token-level prediction and the prototype loss calculated according to the output prediction event type and the semantic anchor dictionary, balancing the contributions of both to optimize model training and improve event extraction performance.
[0013] 3. Build a complete event extraction framework: Expand data through operations such as semantic injection enhancement and synonym replacement for the event types of event texts, and carefully construct a prompt template based on task descriptions, event types, event role tables, semantic enhancement tables, and structured output constraints. In this way, in complex scenarios, such as when dealing with texts with cross-domain knowledge, the model's generalization ability can be effectively improved to complete accurate event extraction. For example, in news monitoring, news manuscripts can be analyzed in real time to extract structured event information to help understand news dynamics. In the field of financial risk analysis, relevant financial texts are mined to extract structured content of risk events to provide support for financial decision-making. Brief Description of the Drawings
[0014] Figure 1 is a flowchart showing the construction method of an event extraction model based on semantic injection enhancement proposed in this solution.
[0015] Figure 2 is a schematic diagram of the data map of this solution.
[0016] Figure 3 is a logical diagram showing the construction method of an event extraction model based on semantic injection enhancement proposed in this solution.
[0017] Figure 4 is a framework diagram of an electronic device implementing this solution. Detailed Implementation Modes
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0019] Those skilled in the art should understand that in the disclosure of the present invention, the orientation or positional relationships indicated by the terms "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the above terms should not be construed as limiting the present invention.
[0020] Embodiment 1 As Figure 1 shown, this solution proposes a construction method for an event extraction model based on semantic injection enhancement, including the following steps: S1: Obtain at least one event text and determine the event type and event role table of the event text; S2: Construct a semantic enhancement table based on dataset samples of different event types; S3: Construct a prompt template based on the task description, event type, event role table, semantic enhancement table, and structured output constraints that define the event extraction task. Use the event text and the corresponding prompt template as the fine-tuning dataset, where the structured output constraints specify the output content and output format of the event extraction task; S4: Perform enhancement processing on the fine-tuning dataset; S5: Train the output content of the large language model using the fine-tuning dataset until the loss function meets the requirements to obtain an event extraction model based on semantic injection enhancement. The loss function used for training the large language model is the combined value of the prototype loss and the cross-entropy loss. The prototype loss that reflects the prototype structure problem is calculated from the predicted event type extracted from the output content of the large language model and the semantic anchor dictionary constructed based on the semantic enhancement table.
[0021] This solution constructs a semantic enhancement table for event texts that require event extraction by means of semantic injection enhancement, directly injects rich semantics into the prompt template based on the semantic enhancement table, and combines the prompt template to convert the event extraction task into the form of question-and-answer pairs for large language models, thereby leveraging the generation ability of large language models to improve the accuracy of event extraction of the event extraction model. In this way, the event extraction model optimized by this solution can better understand the semantic information of event texts, accurately identify event types and event roles, and thus improve the performance of event extraction.
[0022] Specifically, in step S1: This solution can directly obtain event texts with known event type and event role tables, or define the event type and event role table manually after obtaining the event texts.
[0023] Event text: The text description that requires event extraction. Event texts come from a wide range of sources, covering news reports, social media posts, legal documents, medical records, etc. Text data can be collected as event texts through web crawlers from major news websites and social media platforms; specific domain texts can also be obtained from professional databases as event texts. After obtaining the event texts, preprocessing is performed, including removing special characters, HTML tags, unifying case, tokenizing, and part-of-speech tagging.
[0024] Event type: Defines the type to which the current event text belongs. When the event text is related to the legal field, the event type can be any one of "declaring bankruptcy, transferring ownership, transferring funds, getting married, transporting, death, arrest and imprisonment, conviction, prosecution", etc., and the event type can be adjusted according to actual scenario requirements.
[0025] Event role table: Defines the roles involved in the current event text. Similarly, when the event text is related to the legal field, the event role table can be any one of beneficiary, prosecutor, plaintiff, seller, victim, buyer, agent, destination, organization, recipient, location.
[0026] In step S2: The semantic enhancement table records semantic enhancement words and / or semantic features corresponding to the current event type, where the semantic enhancement words are words with similar semantic meanings to the current event type, and the semantic features are features representing the semantic enhancement information of the current event type.
[0027] In some embodiments, semantic enhancement words for the current event type are artificially defined, and the corresponding semantic enhancement words are extended based on the event type. At this time, a semantic enhancement database with semantic enhancement words corresponding to different event types can be preset, and semantic enhancement words related to the current event type are screened from the semantic enhancement database and a semantic enhancement table is constructed. These semantic enhancement words can describe the semantic information of the event type more comprehensively and accurately. For example, for the event type of "fire", the semantic enhancement vocabulary may include "point of origin of the fire", "spread of the fire", "fire fighting and rescue", etc.; for example, for the event type of "attack", the semantic enhancement words may include "bombing, assault, raid, invade", etc.
[0028] In some embodiments, the event text is input into a trained event type extraction model to output semantic features related to the event type. At this time, the event type extraction model is trained to extract semantic features from the event text, and the event type extraction model is a small model similar to BERT, GRU or LSTM.
[0029] Preferably, the semantic enhancement table records the semantic enhancement words and semantic features corresponding to the current event type, that is, the semantic enhancement words and semantic features are combined to form a semantic enhancement table corresponding to the current event type.
[0030] In step S3: The prompt template is the key to guiding the large language model to perform event extraction. This solution designs a reasonable prompt template by combining the task description, event type, event role table, semantic enhancement table and structured output constraints, and uses the prompt template and the event text that needs to be extracted as a fine-tuning data set after structured processing. Subsequently, the large language model can be fine-tuned with the fine-tuning data set.
[0031] Regarding the prompt template of this solution: Task description: Used to define the task of event extraction for the current event text, such as "extracting event information from the given event text".
[0032] Structured output constraints: Used to define the output content and output format of the event extraction task. In some embodiments, the output format is a structured key-value pair, and the output content is the event extraction result.
[0033] It should be noted that the purpose of this solution is to enable the large language model to extract structured event information from unstructured event texts. Therefore, through structured output constraints, the structuring of event information can be achieved to facilitate subsequent data analysis of the extraction results.
[0034] For example, if the task description is "Extract event information from the given text", the event type is "traffic accident", the event role table contains "driver", "passenger", "vehicle", etc., and the semantic enhancement table has information such as "collision location" and "accident cause", and the structured output constraint stipulates that the output format is JSON. Then the prompt template can be designed as: "Please extract traffic accident-related information from the following text, and the output format is JSON, including event type, driver, passenger, vehicle, collision location, accident cause, etc.: [event text]".
[0035] Of course, before inputting the fine-tuning dataset into the large language model for fine-tuning training, data standardization processing is performed on the data in the fine-tuning dataset. For example, the time and date in the fine-tuning dataset can be normalized to prevent training errors of the large language model caused by non-standard data in the fine-tuning dataset.
[0036] In step S4: This solution utilizes the excellent generation ability and semantic understanding ability of the large language model itself to complete the event extraction task, and at the same time introduces semantic injection enhancement technical means to solve many problems that occur in the event extraction task of traditional large language models.
[0037] Specifically, in order to prevent abnormal samples in the fine-tuning dataset from interfering with the fine-tuning training process of the large language model, that is, to ensure that the samples in the fine-tuning dataset conform to the original data distribution characteristics, this solution further includes the steps: In the fine-tuning training, the samples in the fine-tuning dataset are input into the large language model in batches, and the loss value of each sample in the multi-round training process is recorded. Based on the loss value, a data map of each sample is constructed, and abnormal samples in the fine-tuning dataset are located and removed from the data map to construct a new fine-tuning dataset.
[0038] Furthermore, taking the standard deviation of each sample within multiple rounds of training as the abscissa of the data map and the mean value of each sample within multiple rounds of training as the ordinate of the data map, the data deviating from the normal samples on the data map is removed.
[0039] As Figure 2 shown, Figure 2 is an example of the data map given by this solution. The data deviating from the normal samples refers to the data with a relatively large abscissa value and a relatively large ordinate value in the data map. It can be seen that the Figure 2 red dots in it are abnormal samples, and the blue dots and yellow dots are normal samples.
[0040] The formula for calculating the mean value of each sample within multiple rounds of training is as follows: ; The formula for calculating the standard deviation of each sample within multiple rounds of training is as follows: ; where E represents the total number of training rounds, and S represents the number of epochs starting from which statistics are counted, represents the loss value of the i-th sample in the e-th epoch, is the mean value of the i-th sample, is the standard deviation of the i-th sample.
[0041] It should be noted that when the sample size of the fine-tuning dataset is large, the large language model may already have good performance on some samples in the first few rounds, which may cause a certain unfairness problem. Therefore, this solution does not record the loss values of samples from the beginning, thereby alleviating the impact brought by this problem.
[0042] In some embodiments, the abnormal samples include samples with incorrect annotations and samples that do not conform to the data distribution, where the semantic information of the samples that do not conform to the data distribution is quite different from that of most samples in the fine-tuning sample set.
[0043] In step S5: In some embodiments, when fine-tuning the large language model using the fine-tuning dataset, a plug-and-play event semantic injection enhancement component is inserted into the output layer of the large language model, where the event semantic injection enhancement component can disassemble the event type from the output content and optimize the loss function of the large language model using the prototype loss function.
[0044] The loss function of the large language model in this solution is the weighted sum of the cross-entropy loss and the prototype loss, where the cross-entropy loss is the prediction of the large language model at the token level, and the loss function of the large language model is expressed as follows: ; where is the cross-entropy loss, is the prototype loss, is a hyperparameter used to balance the contributions of the prototype loss and the cross-entropy loss, ensuring the coordinated expression of the event core semantics and the context.
[0045] Furthermore, the prototype loss is the cosine distance between the predicted event type extracted from the output content of the large language model and the semantic anchor corresponding to the current event type in the semantic anchor dictionary, and the calculation formula is as follows:
[0046] ; ; where h is the mean value of the predicted event type extracted from the hidden state of the large language model, p is the semantic anchor corresponding to the event type, hi refers to the predicted event type of the i-th hidden state, and n represents the number of hidden states; e i refers to the word embedding vector of the i-th semantic enhancement word or semantic feature corresponding to the current event type, and m represents the number of semantic enhancement words or semantic features corresponding to the current event type.
[0047] The method for obtaining the semantic anchor dictionary is as follows: Use a tokenizer to encode the semantic enhancement words or semantic features in the semantic enhancement table into a token sequence, and convert the token sequence into a word embedding vector through a word embedding layer. Take the mean of the word embedding vectors of the semantic enhancement words or semantic features as the semantic anchor. Use the semantic anchor as the value and the event type as the key to construct key-value pairs, and store the key-value pairs of different event types to obtain the semantic anchor dictionary.
[0048] In some embodiments, the tokenizer of ChatGLM3-6B can be used to encode the semantic enhancement words or semantic features in the semantic enhancement table into a token sequence.
[0049] In some embodiments, locate and extract the predicted event type from the output layer of the large language model. For example, invalid tokens can be skipped and the start and end positions of the event type (marked by "#") can be found.
[0050] Figure 3 is a logical schematic diagram of the construction method of the event extraction model based on semantic injection enhancement provided by this solution. First, construct a semantic enhancement table based on the event type of the event text by manually defining semantic enhancement words and extracting semantic features by the model. Then, construct an indication template based on the semantic enhancement table, combine the indication template and the event text to construct a fine-tuning dataset, and input the fine-tuning dataset into the large language model for fine-tuning; at the same time, use a tokenizer to extract word embedding vectors from the semantic enhancement table and construct a semantic anchor dictionary corresponding to different event types; construct a prototype loss based on the predicted event type and the semantic anchor dictionary in the model prediction result of the large language model, and combine the cross-entropy loss and the prototype loss of the large language model as the loss function of the large language model.
[0051] Embodiment 2 This solution proposes an event extraction model based on semantic injection enhancement, which is constructed according to the event extraction method based on semantic injection enhancement mentioned in the above Embodiment 1.
[0052] Furthermore, this solution provides an event extraction method based on semantic injection enhancement, including the following steps: Input the event text to be subjected to event extraction into the event extraction model based on semantic injection enhancement trained in Embodiment 1, and output the event extraction result.
[0053] The content of Example 2 that is the same as that of Example 1 will not be elaborated here to avoid redundancy.
[0054] Example 3 This embodiment also provides an electronic device. Refer to Figure 4 , including a memory 404 and a processor 402. A computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any of the above embodiments of the method for constructing an event extraction model enhanced by semantic injection.
[0055] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured as one or more integrated circuits implementing the embodiments of the present application. Among them, the memory 404 may include a mass memory 404 for data or instructions.
[0056] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any of the above-mentioned methods for constructing an event extraction model enhanced by semantic injection.
[0057] Optionally, the above-mentioned electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above-mentioned processor 402, and the input / output device 408 is connected to the above-mentioned processor 402.
[0058] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above-mentioned network may include wired or wireless networks provided by a communication provider of the electronic device. In one example, the transmission device includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0059] The input / output device 408 is used to input or output information. In this embodiment, the input information may be event text, etc., and the output information may be event extraction results, etc.
[0060] Optionally, in this embodiment, the above-mentioned processor 402 may be configured to execute the following steps through a computer program: S1: Obtain at least one event text, and determine the event type and event role table of the event text; S2: Construct a semantic enhancement table based on dataset samples of different event types; S3: Construct a prompt template based on the task description, event types, event role table, semantic enhancement table, and structured output constraints for defining the event extraction task. Use the event text and the corresponding prompt template as the fine-tuning dataset, where the structured output constraints specify the output content and output format of the event extraction task; S4: Perform enhancement processing on the fine-tuning dataset; S5: Train the output content of the large language model using the fine-tuning dataset until the loss function meets the requirements to obtain an event extraction model enhanced by semantic injection. The loss function used for training the large language model is the combined value of the prototype loss and the cross-entropy loss. The prototype loss reflecting the prototype structure problem is calculated from the predicted event types extracted from the output content of the large language model and the semantic anchor dictionary constructed based on the semantic enhancement table.
[0061] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0062] Generally, various embodiments can be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, microprocessor, or other computing device. However, the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, by way of non-limiting example, the blocks, devices, techniques, or methods described herein can be implemented in hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controllers, or other computing devices, or some combination thereof.
[0063] Embodiments of the present invention can be implemented by computer software, which is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components configured to perform the embodiments when the program runs. The one or more computer-executable components can be at least one software code or a part thereof. Additionally, in this regard, it should be noted that any box in the logical flow in the figure can represent a program step, or interconnected logical circuits, boxes, and functions, or a combination of program steps and logical circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media are non-transitory media.
[0064] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0065] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for constructing an event extraction model based on semantic injection enhancement, characterized in that: The following steps are involved: S1: obtaining at least one event text, and determining an event type and an event role table of the event text; S2: Construct corresponding semantic enhancement tables based on dataset samples of different event types; S3: Construct a prompt template based on the task description, event type, event role table, semantic enhancement table, and structured output constraints that define the event extraction task, and use the event text and the corresponding prompt template as a fine-tuning dataset, where the structured output constraints specify the output content and output format of the event extraction task; S4: Enhance the fine-tuning dataset; S5: Use the fine-tuning dataset to train the output content of the large language model until the loss function meets the requirements. The loss function used in the large language model training is the sum of the prototype loss and the cross entropy loss. The prototype loss that reflects the prototype structure problem is calculated by the predicted event type extracted from the output content of the large language model and the semantic anchor dictionary constructed based on the semantic enhancement table.
2. The method for constructing an event extraction model based on semantic injection enhancement according to claim 1, characterized in that: The semantic enhancement table records the semantic enhancement words and / or semantic features corresponding to the current event type, where the semantic enhancement words are words with similar semantic expression meanings to the current event type, and the semantic features are features that characterize the semantic enhancement information of the current event type.
3. The method for constructing an event extraction model based on semantic injection enhancement according to claim 1, characterized in that: Artificially define the semantic enhancers of the current event type, and expand the corresponding semantic enhancers based on the event type; Input the event text into the trained event type extraction model to output semantic features related to the event type.
4. The method for constructing an event extraction model based on semantic injection enhancement according to claim 1, characterized in that: The task description in the prompt template is used to define the task of event extraction from the current event text, and the structured output constraint is used to define the output content and output format of the event extraction task.
5. The method for constructing an event extraction model based on semantic injection enhancement according to claim 1, characterized in that: During fine-tuning training, the samples in the fine-tuning dataset are input into the large language model in batches and the loss value of each sample during multiple rounds of training is recorded. A data map of each sample is built based on the loss value, and abnormal samples in the fine-tuning dataset are located and removed from the data map to build a new fine-tuning dataset.
6. The method for constructing an event extraction model based on semantic injection enhancement according to claim 5, characterized in that: The standard deviation of each sample in multiple rounds of training is used as the horizontal coordinate of the data map, and the mean of each sample in multiple rounds of training is used as the vertical coordinate of the data map, and the data that deviates from the normal samples on the data map are eliminated.
7. The method for constructing an event extraction model based on semantic injection enhancement according to claim 1, characterized in that: The loss function of the large language model is expressed as follows: ; in is the cross entropy loss, is the prototype loss, is a hyperparameter used to balance the contribution of prototype loss and cross entropy loss.
8. The method for constructing an event extraction model based on semantic injection enhancement according to claim 1, characterized in that: The semantic enhancement words or semantic features in the semantic enhancement table are encoded into a token sequence by using a token segmenter, and the token sequence is converted into a word embedding vector through a word embedding layer. The mean of the word embedding vectors of the semantic enhancement words or semantic features is taken as the semantic anchor point, and a key-value pair is constructed with the semantic anchor point as the value and the event type as the key. The key-value pairs of different event types are stored to obtain a semantic anchor point dictionary.
9. An event extraction method based on semantic injection enhancement, characterized in that: The event text for which event extraction is required is input into the event extraction model based on semantic injection enhancement constructed according to the method for constructing the event extraction model based on semantic injection enhancement according to any one of claims 1 to 8, to obtain the event extraction result.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method for constructing an event extraction model based on semantic injection enhancement as described in any one of claims 1 to 8 or the event extraction method based on semantic injection enhancement as described in claim 9.
Citation Information
Patent Citations
Event parameter extraction method and system based on abstract semantic graph
CN117874254A
Event extraction system and method based on pre-training model and word meaning enhancement
CN119848229A
Event detection method and system capable of efficiently utilizing limited annotation data
CN119884379A
Event extraction method and device based on causal knowledge constraint
CN119886126A