An event extraction method, device and system based on multi-source annotation

CN117874232BActive Publication Date: 2026-08-28NAT UNIV OF DEFENSE TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311778708.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2026-08-28
Estimated Expiration
2043-12-22

AI Technical Summary

Technical Problem

首先,不同标注方对不同类别的数据的标注水平不平衡

Benefits of technology

[0047]与现有技术相比,本发明有以下有益效果:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117874232B_ABST
    Figure CN117874232B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of natural language processing, and discloses an event extraction method, device and system based on multi-source annotation, which comprises the following steps: creating an event set containing multiple different event types; collecting a large amount of text from various data sources; dividing the corpus dataset into two subsets; annotating the training set; fusing the labels; taking the final label obtained through label aggregation as the training label, taking the text in the training set as the input, and training the deep neural network; the trained neural network model is used for event extraction on new text; the trained model is used for predicting the event type of new text, thereby completing the event extraction task. The application effectively evaluates the distinguishing ability of the annotation party, weights the performance of the label and the annotation party on the corresponding category, obtains the quality of the appropriate difficult-to-annotate event label, and obtains high-quality event extraction labeled data under limited conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of natural language processing and artificial intelligence technology, and particularly relates to an event extraction method, apparatus and system based on multi-source annotation. Background Technology

[0002] Event extraction is a key task in Natural Language Processing (NLP), aiming to identify and extract specific types of events from text. Event extraction has significant value for many applications, including information retrieval, knowledge graph construction, news summarization, and surveillance.

[0003] The core problem of event extraction lies in understanding semantics, a challenging task. Recently, deep learning has been widely applied to event extraction. However, the success of deep learning models often relies on training with large amounts of high-quality labeled data. Acquiring such data is expensive, requiring meticulous annotation by human experts, which is both time-consuming and labor-intensive. To balance quality and cost, multiple non-professional annotators are often organized to repeatedly annotate the same data. For the same data, the label appearing most frequently among different annotators is considered high-quality and regarded as the true label. However, for event data annotation, a dataset may contain dozens of categories. Distinguishing so many categories is already a laborious task, and the existence of similar categories further increases the difficulty of annotation. When labeling data categories that are difficult to distinguish, annotators may easily make mistakes. This can lead to more incorrect duplicate labels than correct labels, thus breaking the standard for selecting accurate labels in existing methods, namely the mode voting method.

[0004] In the prior art, Chinese invention application CN202111624377.0 designed an event annotation system based on crowdsourcing technology for multi-level annotators, completing dataset construction, corpus construction, annotation mechanism, crowdsourcing allocation and aggregation mechanism, and result database export mechanism, with a relatively complete process. However, this invention application only uses simple majority voting as the crowdsourcing aggregation mechanism, without considering the differences in the level of each annotator, and even less considering the impact of the existence of confusing classes on the distribution of annotation results. Therefore, the label aggregation module of this invention application has low accuracy and is difficult to apply in complex annotation scenarios.

[0005] As mentioned above, crowdsourced annotation also faces many challenges. First, the annotation quality of different annotators varies across different categories of data. Second, the annotation quality of different annotators can differ significantly. Therefore, how to effectively extract valuable information from crowdsourced annotation and how to design algorithms to appropriately aggregate these annotations to obtain high-quality annotation results are key issues currently faced. This invention addresses this problem by proposing an event extraction method, apparatus, and system based on multi-source annotation. Summary of the Invention

[0006] In view of this, this invention proposes an event extraction method, apparatus, and system based on multi-source labeling. It conducts fundamental research on crowdsourced label selection and designs a quality assessment process that assigns higher weight to more capable labelers than less capable ones. Through multiple rounds of quality assessment algorithms, correct labels can outnumber incorrect labels, significantly improving the accuracy of the aggregated labeling and resulting in a more robust event extraction model.

[0007] To achieve the above-mentioned objectives, this invention discloses an event extraction method based on multi-source annotation, comprising the following steps: Create an event set containing multiple different event types, where each event consists of one or more roles; Text is collected from multiple data sources, containing the events to be extracted; a corpus is built based on the text, and the data in the corpus includes multiple domains, topics, and text types; The constructed corpus dataset Divided into two subsets: training set and inference set The training set The inference set is used for model learning and training. The part used by the model for event extraction and inference; The training set is labeled to obtain the labels; The tags are merged, including inferring easily confused categories, evaluating tag quality, and merging tags. The final labels obtained through label aggregation are used as training labels, and the text in the training set is used as input to train a deep neural network. Trained neural network model Used for event extraction from new text: given a set of inferences Using models For each text Make predictions and obtain the predicted event labels. , ,in is the total number of reasoning texts, where i is the text label; By using a trained model, the event type can be predicted for new text, thereby completing the task of event extraction.

[0008] Furthermore, create an event set containing multiple different event types, each event consisting of one or more roles, including: First, define the role for each event based on the event types extracted; then, obtain relevant event and role information from the domain expert's knowledge base to enrich the event set; finally, repeatedly review and revise the event set to ensure its completeness and accuracy. Completed event set Indicated as containing Collection of events Each event It is by A collection of roles r m This refers to the m-th character.

[0009] Furthermore, the dataset partitioning includes: training set This constitutes a small portion of the entire dataset, but ensures a sufficient number of samples; the remaining samples form the inference set. Furthermore, the training set Divide into known sets and unknown set known set The samples in the set are precisely labeled portions, while those in the unknown set are not. In this dataset, the correct label of a sample is unknown, thus the entire dataset... Divided into known sets Unknown set and inference set .

[0010] Furthermore, the inferred easily confused categories include: Construct a set of various common and easily confused categories, denoted as M represents the total number of easily confused categories; for each annotator, they are pre-tested to determine their accuracy in annotating classes; if the accuracy is too low, the easily confused categories corresponding to that group are added to the annotator's easily confused category set, i.e. ; λ is Represents the test set of easily confused categories Upper The label accuracy of the annotation results of each annotator, where M is the total number of easily confused categories and l is the number of the easily confused category.

[0011] Furthermore, the quality of the label follows a Gaussian distribution, the distribution function of which is given by the mathematical expectation. and variance The decision is: The expected value of annotation quality is estimated through continuous iteration. The basis for iterating the quality of a particular annotation is its consistency with other annotations. At the start of the iteration, the quality of all annotations is initialized as follows: p i For the annotator The validation set obtained from a pre-random sampling The annotation precision is expressed as: in It is the first Annotator's annotation The label accuracy of the results; Total The first label, of which the... The annotation result of each annotation is Its confusion class is Therefore, the quality of its labeling is represented as follows: Where Y is the set of annotation results, L k E k These are the confusion class and mathematical expectation of the k-th label, respectively. This is the updated tag. quality It is an S-shaped function used to smooth the output. yes The inverse functions of are expressed as follows: The term is calculated using the following equation: in It is the first The annotation results of each annotator Indicates the annotator and Consistency between them This indicates the annotation results that are categorized as "confusion-related". It is the update amount of the confidence level; The variance of label quality is calculated using the following equation: go through Through rounds of iteration, the final label quality of each annotator is finally confirmed. and its variance , V represents the variance of the label quality, the final label quality. It is a label quality with fault tolerance boundaries.

[0012] Furthermore, the quality of the obtained tags is used as the weight for the mode vote, and the final aggregated tag results are voted on to obtain the final tags with high reliability, i.e.: It is an indicator function, when hour, ,otherwise It is the first The annotation results labels of each annotator.

[0013] Furthermore, training a deep neural network includes: The training set is ,in Indicates the first The text of a sample, Indicates the first Event labels for each sample This is the total number of training samples; the goal is to train a deep neural network model. This makes the model Able to train set Minimize a certain loss function ,Right now: in, Representative model For input text The predicted output, Indicates the predicted output and real event tags The losses between This means finding the model that minimizes the subsequent expression. .

[0014] The second aspect of this invention discloses an event extraction device based on multi-source annotation, comprising: Create a unit: Create an event set containing multiple different event types, with each event consisting of one or more roles; Collection Unit: Collects text from various data sources. The text contains the events to be extracted. A corpus is built based on the text. The data in the corpus includes multiple domains, topics, and text types. Unit partitioning: The constructed corpus dataset Divided into two subsets: training set and inference set training set This is the part used for model learning and training, while the inference set... It is the part used by the model for event extraction and reasoning; Labeling unit: Labels the training set; Fusion Unit: This unit fuses tags, including inferring easily confused categories, evaluating tag quality, and merging tags. Training Unit: The final labels obtained through label aggregation are used as training labels, and the text in the training set is used as input to train a deep neural network; Extraction unit: a trained neural network model Used for event extraction from new text: given a set of inferences Using models For each text Make predictions and obtain the predicted event labels. , ,in It is the total number of reasoning texts; Prediction Unit: Using a trained model, it predicts the event type of new text, thereby completing the task of event extraction.

[0015] The third aspect of this invention discloses an event extraction system based on multi-source annotation, comprising: a data acquisition terminal and an event extraction server. Attached Figure Description

[0016] Figure 1 Framework diagram of the present invention; Figure 2 A flowchart illustrating the event extraction method based on multi-source annotation. Detailed Implementation

[0017] The present invention will be further described below with reference to the accompanying drawings, but this is not intended to limit the present invention in any way. Any modifications or substitutions made based on the teachings of the present invention shall fall within the protection scope of the present invention.

[0018] This invention provides the following technical solution, and the main technical process is as follows: Figure 1The specific technical solution process is as follows: I. Data Collection and Data Labeling Step 1.1: Event Set Construction First, create an event set containing multiple different event types. These event types include, but are not limited to, buying, selling, and collaboration. Each event consists of one or more roles; for example, a buying event might include roles such as buyer, seller, and product. This event set is created based on the knowledge and experience of domain experts to ensure the comprehensiveness and accuracy of the events.

[0019] Specifically, first, the role of each event is defined according to the types of events to be extracted; then, relevant event and role information is obtained from the knowledge base of domain experts to enrich the event set; finally, the event set is repeatedly reviewed and revised to ensure its completeness and accuracy.

[0020] Completed event set It can be represented as containing Collection of events Each event It is by A collection of roles The following explanation focuses on event classification tasks; the identification of roles follows the same principle as event identification.

[0021] Step 1.2: Corpus Construction Corpus construction is a crucial step in the entire event extraction process. This step involves collecting a large amount of text from various data sources, containing the events to be extracted. Corpus construction must ensure data diversity and representativeness, including different domains, topics, and text types.

[0022] First, text is collected from various online and offline data sources, such as news websites, social media platforms, forums, blogs, books, and research reports. Then, the collected text is preprocessed, including word segmentation, part-of-speech tagging, and named entity recognition. Finally, the preprocessed text is cleaned and filtered to remove irrelevant text and noise, and retain the text containing the target event.

[0023] The above steps have constructed a corpus containing a large amount of diverse and representative texts, providing data support for subsequent event extraction.

[0024] Step 1.3: Dataset Partitioning In this stage, the corpus dataset will be constructed. Divided into two subsets: training set and inference set Training set This is the part used for model learning and training, while the inference set... It is the part used by the model for event extraction and reasoning.

[0025] training set This can occupy a small portion of the entire dataset, but a sufficient number of samples must be ensured. The remaining samples constitute the inference set. Furthermore, the training set Divide into known sets and unknown set known set The samples in the dataset are those that have been evaluated and precisely labeled by experts. In the unknown set... In this dataset, the correct label of a sample is unknown. Thus, the entire dataset... Divided into known sets (Precisely labeled samples), unknown set (Unlabeled training samples) and inference set (Samples used for model inference).

[0026] The advantage of this partitioning method is that it allows the model to operate on a known set. To learn effectively on, and in unknown sets Learning from unlabeled samples is performed on the inference set. It provides application scenarios for event extraction and reasoning of the model, enabling the model to perform reasoning and verification in real-world environments.

[0027] Step 1.4: Add annotations Label the training set. Let there be a total of... There are 1 annotation party, and the number of text segments in the corpus is 1. The total annotation size is The annotator selects an event type suitable for the semantics and context of the text segment from the pre-defined event set in step 1.1, and uses this as the annotation for that event type in the text segment. In this invention, the annotation method can be manual or automated by machine. Machine annotation refers to annotation using a machine trained with an existing neural network. For example, the Label Studio open-source data annotation tool can be used for annotation.

[0028] II. Tag Fusion Methods Step 2.1 Inference of Easily Confused Categories First, construct a set of common and easily confused categories, denoted as [list of categories]. For each annotator, pre-testing is performed to determine its accuracy on the labeled classes. If the accuracy is too low, the corresponding easily confused classes are added to the annotator's easily confused class set. , Step 2.2 Label Quality Assessment The quality of the labels is calculated below. The quality of a label can be viewed as a Gaussian distribution, whose distribution function is given by the expected value. and variance The decision is: The expected value of annotation quality is estimated through iterative iteration. The basis for each annotation quality iteration is its consistency with other annotations. At the start of the iteration, the quality of all annotations is initialized as follows: Assume there is a total The annotation method, of which the first annotation method The annotation results of each annotation party Its confusion class is Therefore, the quality of its labeling can be expressed as follows: in It is an S-shaped function used to smooth the output. yes The inverse functions of are expressed as follows: The term is calculated using the following equation: in It is the first The annotation results of each annotator Indicates the annotator and Consistency between them This indicates the annotation results that are categorized as "confusion-related". It is the update amount of the confidence level. This is the updated tag. The quality.

[0029] The variance of crowdsourced tag quality is calculated using the following equation: go through Through rounds of iteration, the final label quality of each labeler was finally confirmed. and its variance , V represents the variance of the label quality, the final label quality. It is a label quality with fault tolerance boundaries.

[0030] Step 2.3 Tag Fusion The quality of the tags obtained in the previous step is used as the weight for the mode vote, and a vote is then cast on the final aggregated tag results to obtain the final tags with high reliability. That is: It is an indicator function, when hour, ,otherwise It is the first The annotation results labels of each annotator.

[0031] III. Model Training and Inference Step 3.1 Deep Model Training In the event extraction task of this invention, the final labels obtained through label aggregation are used as training labels, and the text in the training set is used as input to train a deep neural network. Deep neural network models include bidirectional long short-term memory network models, etc., and this invention does not limit them.

[0032] Assume the training set is ,in Indicates the first The text of a sample, Indicates the first Event labels for each sample This represents the total number of training samples. The goal of this invention is to train a deep neural network model. This makes the model Able to train set Minimize a certain loss function ,Right now: in, Representative model For input text The predicted output, Indicates the predicted output and real event tags The losses between This means finding the model that minimizes the subsequent expression. .

[0033] Step 3.2 Deep Model Inference Trained neural network model It can be used to extract events from new text. Given a set of inferences... We can use a model For each text Make predictions and obtain the predicted event labels. , ,in It represents the total number of reasoning texts.

[0034] In this way, a trained model can be used to predict the event type of new, unseen text, thereby completing the task of event extraction.

[0035] refer to Figure 2 To facilitate the explanation of the solution process, a specific example of extracting global news events from news text using text event extraction technology is given below. This example is a detailed illustration of the entire method (for ease of explanation, the example parameters are used in this invention, but the specific implementation is not limited to the specific parameters of this example): I. Data Collection and Preprocessing Step 1.1: Event Set Construction In constructing the event set, we aim to extract specific event types and their associated roles from news reports. For example, events might include "product launch," "management appointment," "management resignation," "strategic plan release," and "bankruptcy declaration." For a "product launch" event, associated roles might include "the launching organization," "the launching department," and "the host." This process may require expert assistance to ensure that all important event types and roles are considered.

[0036] Step 1.2: Corpus Construction Corpus construction primarily involves collecting and preprocessing a large volume of news report text. This text can be obtained from publicly available online sources such as various news websites, social media, and blogs. Preprocessing steps include word segmentation, part-of-speech tagging, and named entity recognition to aid in subsequent event extraction. Then, the text is filtered and cleaned, including removing short paragraphs, removing irrelevant text, and removing and replacing special characters. Articles containing the events defined in step 1.1 are retained.

[0037] Step 1.3: Dataset Partitioning The constructed corpus dataset Divided into two subsets: training set and inference set .in The number accounts for 20% of the total sample size. This number accounts for 80% of the total sample size. Only in... The trained model can then be used in... This will play a role. Furthermore, the training set... Divide into known sets and unknown set Known set The samples in the training set are precisely labeled, accounting for 30% of the total. The unknown set... The samples in the training set account for 70%.

[0038] Step 1.4: Organize annotations To train the event extraction model, events and roles in the training set are labeled using a multi-source approach. The text in the training set is then labeled. During the labeling process, the labeler needs to label all relevant events and their roles for each article based on the events and roles defined in step 1.1.

[0039] II. Tag Fusion and Model Building Step 2.1 Inference of Easily Confused Categories Build a collection containing various common and easily confused event categories. This set includes the following subsets: ["Product Launch", "Management Appointment"], ["Product Launch", "Strategic Plan Release"], ["Strategic Plan Release", "Bankruptcy Declaration"], etc. This step is conducted under the guidance of expert prior knowledge to ensure the comprehensiveness and accuracy of the event set. Each annotator needs to undergo pre-defined tests to evaluate their accuracy across all confusion categories. If an annotator's accuracy is too low, the corresponding confusion category will be added to that annotator's confusion category set. This means that for the first... Annotator: , Step 2.2 Label Quality Assessment The quality of the labeled area is calculated below.

[0040] First, initialize the quality of all annotators to: Assume there is a total The number of annotators, of which the first The annotation results of each annotator are Its confusion class is Therefore, the quality of its labeling can be expressed as follows: The term is calculated using the following equation: in It is the first The annotation results of each annotator Indicates the annotator and Consistency between them This indicates the annotation results that are categorized as "confusion-related". It is the update amount of the confidence level. This is the updated tag. The quality.

[0041] The variance of label quality is calculated using the following equation: After 20 rounds of iteration, the final label quality for each annotator was confirmed. and its variance , .Will Label quality serves as the final confidence criterion. For example, for a piece of text, five labelers assigned event types such as "product launch," "administrator appointment," "product launch," "administrator appointment," and "strategic plan release." These types... , They can calculate their .

[0042] Step 2.3 Tag Fusion According to the formula Calculate the vote values ​​for all categories to determine the weight of the "Product Launch" category. The weight of the "Administrator Appointment" class is The weight of the "Strategic Plan Release" category is Ultimately, the text was tagged with "product launch".

[0043] III. Model Training and Inference 3.1 Deep Model Training Using a bidirectional long short-term memory network model The training process takes the text of a sentence as input and outputs event predictions from that sentence. The goal of this invention is to minimize the cross-entropy loss between the model's predicted events and the actual events. in, The model is for the first Event prediction for each sentence, It is the cross-entropy loss between the predicted event and the actual event.

[0044] 3.2 Deep Model Inference After training, the trained bidirectional long short-term memory network model is used. Event extraction is performed on the new sentences. Assume there is a set of sentences to be predicted. Models can be used For each sentence Make predictions and obtain predicted events. .

[0045] For example, given the sentence "Company XX launched its latest flagship phone," the model will extract the event "product launch." This completes the event extraction task for the sentence.

[0046] This invention also discloses an event extraction system, including a data acquisition terminal and an event extraction server. The event extraction server uses the event extraction method described above. The data acquisition terminal includes computers, mobile devices, etc., and is used to collect text from various data sources in the real world. The text contains the events to be extracted. The acquisition method is existing technology, such as obtaining from the Internet or from public datasets, which will not be described in detail in this invention.

[0047] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention proposes a crowdsourced label selection framework that can infer accurate results from the results of crowdsourced labeling.

[0048] (2) In view of the occurrence of easily confused categories in event categories, the present invention designs a test program that can effectively evaluate the annotator’s ability to distinguish between different categories of events, thereby providing reliable prior quality of labels.

[0049] (3) The present invention designs an iterative quality assessment algorithm, which weights the label and the labeler’s performance in the corresponding category, thereby obtaining a more appropriate quality label for difficult-to-label events. (4) The method proposed in this invention can obtain high-quality event extraction labeled data under limited conditions, train an event extraction model with high reliability, and realize accurate analysis of natural language text content and event extraction.

[0050] As used herein, the term "preferred" is meant as an example, illustration, or illustration. Any aspect or design described herein as "preferred" need not be construed as being more advantageous than other aspects or designs. Rather, the use of the term "preferred" is intended to present the concept in a specific manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusionary "or." That is, unless otherwise specified or clear from the context, "X uses A or B" naturally includes either of the permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.

[0051] Furthermore, although this disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art based on a reading and understanding of this specification and the accompanying drawings. This disclosure includes all such modifications and variations and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the aforementioned components (e.g., elements, etc.), the terminology used to describe such components is intended to correspond to any component (unless otherwise indicated) that performs the specified function of said component (e.g., is functionally equivalent to it), even if structurally not equivalent to the disclosed structure performing the functions in the exemplary implementations of this disclosure shown herein. Moreover, although specific features of this disclosure have been disclosed with respect to only one of several implementations, such features may be combined with one or more features of other implementations that may be desirable and advantageous for a given or particular application. Furthermore, with regard to the use of the terms “comprising,” “having,” “containing,” or variations thereof in the Detailed Description or claims, such terms are intended to be included in a manner similar to the term “including.”

[0052] The functional units in this invention embodiment can be integrated into a processing module, or each unit can exist physically separately, or multiple units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. The aforementioned devices or systems can execute the storage methods in the corresponding method embodiments.

[0053] In summary, the above embodiments are one implementation of the present invention, but the implementation of the present invention is not limited to the embodiments described above. Any changes, modifications, substitutions, combinations, or simplifications made that deviate from the spirit and principle of the present invention should be considered equivalent substitutions and are included within the protection scope of the present invention.

Claims

1. An event extraction method based on multi-source annotation, characterized in that, Includes the following steps: Create an event set containing multiple different event types, where each event consists of one or more roles; Collect text from multiple data sources; the text contains events to be extracted. A corpus is built based on text, and the data in the corpus includes multiple domains, topics, and text types; The constructed corpus dataset Divided into two subsets: training set and inference set The training set The inference set is used for model learning and training. The part used by the model for event extraction and inference; Label the training set to obtain the labels; The tags are merged, including inferring easily confused categories, evaluating tag quality, and merging tags. The final labels obtained through label aggregation are used as training labels, and the text in the training set is used as input to train a deep neural network. Trained neural network model Used for event extraction from new text: given a set of inferences Using models For each text Make predictions and obtain the predicted event labels. , ,in is the total number of reasoning texts, where i is the text label; Using the trained model, the event type is predicted for new text, thus completing the task of event extraction; The quality of the label follows a Gaussian distribution, the distribution function of which is given by the mathematical expectation. and variance The decision is: The expected value of annotation quality is estimated through continuous iteration. The basis for iterating the quality of a particular annotation is its consistency with other annotations. At the start of the iteration, the quality of all annotations is initialized as follows: p i For the labeling party The validation set obtained from a pre-random sampling The annotation precision is expressed as: in It is the first Annotator's annotation The label accuracy of the results; Total The first label, of which the... The annotation result of each annotation is Its confusion class is Therefore, the quality of its labeling is represented as follows: Where Y is the set of annotation results, L k E k These are the confusion class and mathematical expectation of the k-th label, respectively. This is the updated tag. quality It is an S-shaped function used to smooth the output. yes The inverse functions of are expressed as follows: The term is calculated using the following equation: in It is the first The annotation results of each annotator Indicates the annotator and Consistency between them This indicates the annotation results that are categorized as "confusion-related". It is the update amount of the confidence level; The variance of label quality is calculated using the following equation: go through Through rounds of iteration, the final label quality of each annotator is finally confirmed. and its variance , ; Will As the final confidence level for tag quality, V is the variance of the tag quality, representing the final tag quality. It is the label quality with fault tolerance boundaries; the obtained label quality is used as the weight of the mode vote, and the final aggregated label results are voted on to obtain the final label with high reliability.

2. The event extraction method based on multi-source annotation according to claim 1, characterized in that, Create an event set containing multiple different event types, each event consisting of one or more roles, including: Define the role for each event based on the event types to be extracted; then, obtain relevant event and role information from the domain expert's knowledge base to enrich the event set; and repeatedly revise the event set to ensure its completeness and accuracy. Completed event set Indicated as containing Collection of events Each event It is by A collection of roles r m This refers to the m-th character.

3. The event extraction method based on multi-source annotation according to claim 2, characterized in that, The dataset partitioning includes: training set Ensure there is a sufficient number of samples; the remaining samples constitute the inference set. Further training set Divide into known sets and unknown set known set The samples in the set are precisely labeled portions, while those in the unknown set are not. In this dataset, the correct label of a sample is unknown; therefore, the entire dataset... Divided into known sets Unknown set and inference set .

4. The event extraction method based on multi-source annotation according to claim 1, characterized in that, Inferences about easily confused categories include: Construct a set of various common and easily confused categories, denoted as M represents the total number of easily confused categories; for each labeler, it is pre-tested to determine its accuracy in labeling categories; if the accuracy is too low, the easily confused categories corresponding to that group are added to the labeler's easily confused category set, i.e. ; λ is the accuracy threshold. Represents the test set of easily confused categories Upper The label accuracy of the annotation results of each annotator, where M is the total number of easily confused categories and l is the number of the easily confused category.

5. The event extraction method based on multi-source annotation according to claim 4, characterized in that, The high reliability label is: It is an indicator function, when hour, ,otherwise It is the first The labeling results of each annotation party.

6. The event extraction method based on multi-source annotation according to claim 5, characterized in that, Training a deep neural network includes: The training set is ,in Indicates the first The text of a sample, Indicates the first Event labels for each sample This is the total number of training samples; the goal is to train a deep neural network model. This makes the model Able to train set Minimize a certain loss function ,Right now: in, Representative model For input text The predicted output, Indicates the predicted output and real event tags The losses between This means finding the model that minimizes the subsequent expression. .

7. An event extraction apparatus based on multi-source annotation using the method described in any one of claims 1-6, characterized in that, include Create a unit: Create an event set containing multiple different event types, with each event consisting of one or more roles; Collection Unit: Collects text from various data sources. The text contains the events to be extracted. A corpus is built based on the text. The data in the corpus includes multiple domains, topics, and text types. Unit partitioning: The constructed corpus dataset Divided into two subsets: training set and inference set training set This is the part used for model learning and training, while the inference set... It is the part used by the model for event extraction and reasoning; Labeling unit: Labels the training set; Fusion Unit: This unit fuses tags, including inferring easily confused categories, evaluating tag quality, and merging tags. Training Unit: The final labels obtained through label aggregation are used as training labels, and the text in the training set is used as input to train a deep neural network; Extraction unit: a trained neural network model Used for event extraction from new text: given a set of inferences Using models For each text Make predictions and obtain the predicted event labels. , ,in It is the total number of reasoning texts; Prediction Unit: Using a trained model, it predicts the event type of new text, thereby completing the task of event extraction.

8. An event extraction system based on multi-source annotation, characterized in that, include: A data acquisition terminal, and an event extraction server that applies the method of any one of claims 1-6.

Citation Information

Patent Citations

  • A method for constructing an event annotation system for multi-level annotators based on crowdsourcing technology

    CN114281998B

  • Data processing method and system, network model and training method thereof, and electronic equipment

    CN113469205A

  • Artificial intelligence (AI) method for cleaning data to train (AI) model

    CN115699208A