Method executed by electronic equipment, electronic equipment and storage medium
Through the encoder-decoder network independently learning the event concept and building a general event element set, the problems of high cost of event extraction and poor generalization in the prior art are solved, and higher extraction accuracy and cross-domain adaptability are achieved.
Patent Information
- Application Number
- CN202311801290.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-07-01
AI Technical Summary
The existing event extraction technology relies on manual writing of description text, which is cost-effective, and has degraded performance when migrating between different fields, poor generalization, and lack of data leads to insufficient model robustness and generalization.
Reusable concept learning method based on encoder-decoder network is adopted to independently learn event concept knowledge, and by retrieving and splicing related event definition information, a general event element collection is constructed, the correlation between event concepts and instances is established, and the accuracy of extraction is improved.
Improves the accuracy of event extraction and the generalization ability of the model across different fields, surpassing the performance of existing models in the ACE 2005 dataset and WIKIEVENTS zero-sample evaluation.
Smart Images

Figure CN120234416A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of artificial intelligence, machine learning, etc. This application relates to a method executed by an electronic device, an electronic device, and a storage medium. Background Art
[0002] Event extraction refers to identifying various defined events from unstructured natural language. Event extraction is an important basis for understanding natural language. With the development of deep learning, event extraction technology based on neural networks has emerged. However, since the research on event extraction technology was carried out, there has always been room for improvement in the extraction effect. How to better use neural networks for event extraction remains the focus of current research. Summary of the Invention
[0003] This application provides a method executed by an electronic device, an electronic device, and a storage medium. The technical solution is as follows:
[0004] On the one hand, a method executed by an electronic device is provided. The method includes:
[0005] Obtain the text to be processed;
[0006] Based on the text to be processed and the first prompt information, use the trained artificial intelligence (AI) network to extract the event extraction result of the text to be processed;
[0007] Wherein, the first prompt information is used to prompt extraction based on at least one event definition information in the candidate event definition set.
[0008] In a possible implementation, the first prompt information is obtained through the following steps:
[0009] Retrieve at least one associated event definition information associated with the text to be processed from the candidate event definition set;
[0010] Based on the at least one associated event definition information, obtain the first prompt information, and the at least one event definition information includes the at least one associated event definition information.
[0011] In a possible implementation, obtaining the first prompt information based on the at least one associated event definition information includes any one of the following:
[0012] Concatenate the at least one associated event definition information and the second prompt information to obtain the first prompt information. The second prompt information refers to extraction based on specific event definition information, and the specific event definition information refers to the at least one associated event definition information being concatenated;
[0013] Concatenate at least one event type information corresponding to the at least one associated event definition information and the second prompt information to obtain the first prompt information, where the specific event definition information refers to at least one associated event definition information corresponding to the at least one event type information concatenated.
[0014] In a possible implementation, the event extraction result is obtained by the AI network by performing the following operations:
[0015] Based on the text to be processed and the first prompt information, extract the description information of the event composition units of the text to be processed;
[0016] Based on the description information of the event composition units, obtain the event extraction result of the text to be processed.
[0017] In a possible implementation, the extracting the description information of the event composition units of the text to be processed based on the text to be processed and the first prompt information includes:
[0018] Based on the text to be processed and the first prompt information, extract at least one event type information associated with the text to be processed, the description information of at least one general event element corresponding to each target event, and the description information of the event trigger element;
[0019] Wherein, the at least one general event element is an element in the general element set, and the general element set includes the general event elements corresponding to each candidate event in the candidate event set corresponding to the candidate event definition set; the event composition units of the text to be processed include the event composition units corresponding to each target event associated with the text to be processed, and each event composition unit includes an event type, each general event element, and an event trigger element.
[0020] In a possible implementation, the at least one target event includes any one of the following:
[0021] If the at least one event definition information includes at least one associated event definition information associated with the text to be processed, the at least one target event is the event associated with the text to be processed among the at least one event corresponding to the at least one associated event definition information;
[0022] If the at least one event definition information includes each event definition information in the candidate event definition set, the at least one target event is the event associated with the text to be processed in the candidate event set.
[0023] In a possible implementation, obtaining an event extraction result corresponding to each target event based on the event type information of each extracted target event, as well as the description information of each general event element associated with each target event and the description information of the event trigger element, includes:
[0024] Based on the event type information of each target event, converting the description information of each general event element corresponding to each target event into the description information of each special event element in the set of special event elements corresponding to each target event;
[0025] Based on the description information of each special event element corresponding to each target event and the description information of the event trigger element, obtaining an event extraction result corresponding to each target event.
[0026] In a possible implementation, the description information of a general event element includes at least the general event element;
[0027] The converting, based on the event type information of each target event, the description information of each general event element corresponding to each target event into the description information of each special event element corresponding to each target event includes:
[0028] Based on the event type information of each target event, converting the general event elements in the description information of each general event element corresponding to each target event to obtain the description information of each special event element corresponding to each target event.
[0029] In a possible implementation, the description information of a general event element includes the general event element and the element value of the general event element;
[0030] In a possible implementation, obtaining an event extraction result corresponding to each target event based on the event type information of each extracted target event, as well as the description information of each general event element associated with each target event and the description information of the event trigger element, includes:
[0031] Based on the event type information of each target event, determining the set of target special elements for each target event, where the set of target special elements includes the special event elements corresponding to the general event elements associated with the corresponding target event;
[0032] Using each special event element in the set of target special elements of each target event to replace the general event elements in the description information of the general event elements of the corresponding target event, to obtain the description information of the special event elements corresponding to each target event;
[0033] Based on the description information of each dedicated event element corresponding to each target event and the description information of the event trigger element, obtain the event extraction result corresponding to each target event.
[0034] In a possible implementation manner, the training method of the AI network includes:
[0035] Obtain a training data set, where the training data set includes first training data. The first training data includes a plurality of first samples, and each first sample includes a first sample input and a first label. The first sample input includes a sample text and a first hint information of the sample text, and the first label represents the true event extraction result corresponding to the sample text;
[0036] Based on the training data set, iteratively perform the following training operations on the initial AI network until a preset condition is met to obtain the event extraction network:
[0037] Based on each first sample input, use the initial AI network to extract the first event extraction result of each sample text;
[0038] Based on the first label and the first event extraction result of each sample text, determine the first training loss;
[0039] Adjust the network parameters of the initial AI network based on the total training loss, where the total loss includes the first training loss.
[0040] In a possible implementation manner, the training data set further includes second training data. The second training data includes a plurality of second samples, and each second sample includes a second sample input and a second label. The second sample input includes an initial definition of a candidate event, and the second label represents the true event type corresponding to the initial definition of the candidate event;
[0041] The total training loss further includes a second training loss, and the training operation further includes:
[0042] Based on each second sample input, use the initial AI network to extract the event type information corresponding to each candidate event;
[0043] Based on the second label corresponding to the initial definition of each candidate event and the extracted event type information, determine the second training loss.
[0044] In a possible implementation manner, the training data set further includes third training data. The third training data includes a third sample corresponding to each sample text, and each third sample includes a third sample input and a third label. The third sample input includes the sample text, a third hint information of the sample text, and a fourth hint information;
[0045] Among them, the third prompt message prompts extraction based on the event definition information of the target event associated with the sample text; the fourth prompt message prompts extraction of at least one target element corresponding to the sample text; the at least one target element includes at least one of an event type, each general event element, or an event trigger unit; the third tag represents the true description information of each target element corresponding to the sample text.
[0046] The total training loss further includes a third training loss, and the training operation further includes:
[0047] Based on each third sample input, using the initial AI network, a second event extraction result of each sample text is extracted, and the second event extraction result includes the description information of each target element corresponding to the sample text.
[0048] Based on the third tag and the second event extraction result of each sample text, the third training loss is determined.
[0049] On the other hand, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the method executed by the electronic device as described above.
[0050] On the other hand, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the method executed by the electronic device as described above is implemented.
[0051] The beneficial effects brought by the technical solution provided in the embodiments of the present application are:
[0052] The method executed by the electronic device provided in the present application, based on the text to be processed and the general prompt information, extracts the event extraction result of the text to be processed through the trained event extraction network; the general prompt information is used to prompt extraction based on at least one event definition information in the candidate event definition set, and using this first prompt information can enable the AI network to extract in combination with at least one event definition information, which can effectively improve the accuracy of event extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application.
[0054] Figure 1 It is a schematic diagram of a prompt text in a related technology provided in an embodiment of the present application;
[0055] Figure 2 It is a schematic flowchart of a method executed by an electronic device provided in an embodiment of the present application;
[0056] Figure 3 A schematic diagram of a retrieval process provided by an embodiment of the present application;
[0057] Figure 4 A schematic diagram of the structure of an event extraction network provided by an embodiment of the present application;
[0058] Figure 5 A schematic diagram of the process of a training method for an AI network provided by an embodiment of the present application;
[0059] Figure 6 A schematic diagram of a sample data structure for adding prompt text provided by an embodiment of the present application;
[0060] Figure 7 A schematic diagram for comparing prompt information between multiple training tasks provided by an embodiment of the present application;
[0061] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0062] The following description with reference to the accompanying drawings is provided to facilitate a thorough understanding of various embodiments of the present disclosure defined by the claims and their equivalents. This description includes various specific details to facilitate understanding but should only be considered exemplary. Therefore, those of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and structures may be omitted for clarity and conciseness.
[0063] The terms and phrases used in the following specification and claims are not limited to their dictionary meanings but are used solely by the inventors to enable a clear and consistent understanding of the present disclosure. Therefore, it should be apparent to those skilled in the art that the following description of the various embodiments of the present disclosure is provided only for illustrative purposes and not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.
[0064] It should be understood that the singular forms "a", "an", and "the" may also include plural referents unless the context clearly indicates otherwise. Thus, for example, a reference to "the surface of a component" includes a reference to one or more such surfaces. When we say that one element is "connected" or "coupled" to another element, that one element may be directly connected or coupled to the other element, or it may mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include a wireless connection or wireless coupling.
[0065] The term "comprises" or "may comprise" refers to the presence of the corresponding disclosed functions, operations, or components that can be used in various embodiments of the present disclosure, rather than limiting the presence of one or more additional functions, operations, or features. In addition, the term "comprises" or "has" may be interpreted to mean certain characteristics, numbers, steps, operations, components, components, or combinations thereof, but should not be construed to exclude the possibility of the presence of one or more other characteristics, numbers, steps, operations, components, components, or combinations thereof.
[0066] The term "or" used in various embodiments of the present disclosure includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, more, or all of the multiple items. For example, for the description of "parameter A includes A1, A2, A3", it may be implemented as parameter A includes A1 or A2 or A3, or it may also be implemented as parameter A includes at least two of the three items of parameter A1, A2, and A3.
[0067] Unless otherwise defined, all terms (including technical terms or scientific terms) used in the present disclosure have the same meaning as understood by those skilled in the art described in the present disclosure. Commonly used terms defined in a dictionary are interpreted to have a meaning consistent with the context in the relevant technical field, and should not be interpreted idealistically or overly formally unless explicitly defined as such in the present disclosure.
[0068] At least part of the functions in the device or electronic device provided in the embodiments of the present disclosure can be implemented by an AI model. For example, at least one of the multiple modules of the device or electronic device can be implemented by an AI model. The functions associated with AI can be executed by a non-volatile memory, a volatile memory, and a processor.
[0069] The processor may include one or more processors. At this time, the one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., or a pure graphics processing unit, such as a graphics processing unit (GPU), a vision processing unit (VPU), and / or an AI dedicated processor, such as a neural processing unit (NPU).
[0070] The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence (AI) models stored in the non-volatile memory and the volatile memory. The predefined operation rules or artificial intelligence models are provided through training or learning.
[0071] Here, providing through learning refers to obtaining a predefined operation rule or an AI model with desired characteristics by applying a learning algorithm to multiple learning data. This learning can be performed in the device or electronic device itself where the AI according to the embodiment is executed, and / or can be implemented by a separate server / system.
[0072] The AI model can include multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network calculations through the calculation between the input data of the layer (such as the calculation result of the previous layer and / or the input data of the AI model) and the multiple weight values of the current layer. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial network (GAN), and deep Q-network.
[0073] A learning algorithm is a method of using multiple learning data to train a predetermined target device (e.g., a robot) to enable, permit, or control the target device to make a determination or prediction. Examples of this learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0074] The method provided by this disclosure may relate to one or more fields in technical fields such as speech, language, image, video, or data intelligence.
[0075] Optionally, when involving the field of speech or language, according to this disclosure, in the method executed by an electronic device, a method for recognizing a user's speech and interpreting the user's intention can receive a speech signal as an analog signal via a speech signal acquisition device (e.g., a microphone), and convert the speech part into computer-readable text using an automatic speech recognition (ASR) model. The intention of the user's words can be obtained by interpreting the converted text using a natural language understanding (NLU) model. The ASR model or the NLU model can be an artificial intelligence model. The artificial intelligence model can be processed by an artificial intelligence dedicated processor designed in a hardware structure specified for artificial intelligence model processing. The artificial intelligence model can be obtained through training. Here, "obtained through training" means obtaining a predefined operation rule or an artificial intelligence model configured to perform desired features (or purposes) by training a basic artificial intelligence model with multiple training data using a training algorithm. Language understanding is a technology for recognizing and applying / processing human language / text, including, for example, natural language processing, machine translation, dialogue systems, question answering, or speech recognition / synthesis.
[0076] Optionally, when it comes to the field of images or videos, according to the present disclosure, in the method executed in an electronic device, output data for recognizing an image or in the image can be obtained by using the image data as input data of an artificial intelligence model. The artificial intelligence model can be obtained through training. Here, "obtained through training" means obtaining a predefined operation rule or artificial intelligence model configured to perform an expected feature (or purpose) by training a basic artificial intelligence model with multiple pieces of training data using a training algorithm. The method of the present disclosure may relate to the field of visual understanding of artificial intelligence technology, which is a technology for recognizing and processing things like human vision and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / positioning, or image enhancement.
[0077] Optionally, when it comes to the field of data intelligent processing, according to the present disclosure, in the method executed in an electronic device, the method for inference or prediction can be recommended / executed by using an artificial intelligence model. The processor of the electronic device can perform a preprocessing operation on the data to convert it into a form suitable for use as input to the artificial intelligence model. The artificial intelligence model can be obtained through training. Here, "obtained through training" means obtaining a predefined operation rule or artificial intelligence model configured to perform an expected feature (or purpose) by training a basic artificial intelligence model with multiple pieces of training data using a training algorithm. Inference and prediction are techniques for logical reasoning and prediction by determining information and include, for example, knowledge-based reasoning, optimization prediction, preference-based planning, or recommendation.
[0078] Event extraction (EE) is a task of understanding human-defined events from unstructured natural language data, including identifying the event trigger words and each event argument. The event trigger word is the word that most clearly expresses the occurrence of an event, and the event arguments clarify the participants and attributes of the event. Therefore, the evaluation of event extraction includes the following two parts:
[0079] (1) Event detection (ED): Identifying event trigger words and event types;
[0080] (2) Event argument extraction (EAE): Extracting each predefined argument of the event. Event extraction has a wide range of applications, such as general information extraction, finance, multimedia, law, social, etc.
[0081] In the related art, in the training stage, it is necessary to construct prompts for each type of event and each argument in the event, and splice the constructed prompt text to each instance text. The model learns the trigger words and event elements of each event on each instance text. For example, Figure 1As shown, the prompt text constructed for each event example includes four parts: example text, event type description text, event keyword description text, and event trigger word / event element description template.
[0082] The inventors of the present application found through research on related technologies that the related technologies have at least one of the following problems:
[0083] (1) It is necessary to manually write description texts and formulate templates for each event type, which highly depends on manual writing, with a fine granularity down to each sample, and the implementation cost is high.
[0084] (2) Each instance data needs to be concatenated with prompt text, making the knowledge learning of events tightly coupled with the instance text. When migrating between different fields, for example, when the instance context of an event undergoes a field switch, the model performance deteriorates and the generalization ability is poor.
[0085] (3) The common elements in different events have good domain invariance, but the existing methods destroy the semantic structural similarity of these elements, which is not conducive to the robustness and generalization of the model.
[0086] It should be noted that a major challenge in event extraction is the lack of data. Prompt learning can improve performance by obtaining additional knowledge support through manually written text or learned tags. In related technologies, although the prompt information is constructed for each sample, the model does not independently learn the conceptual knowledge of events, but learns the co-occurrence patterns of events and their instances, and when the downstream field changes, their performance will decline.
[0087] To solve the problems existing in the related technologies, the reusable concept learning method based on an encoder-decoder as the backbone network proposed in the present application independently learns the conceptual knowledge of events and applies the event knowledge to various downstream tasks. By splitting the conceptual knowledge of events and event instances, an independent event concept learning task, an event extraction task, and an association learning task are designed in this application to combine event knowledge and event extraction task instances. Experimental results show that the method of the present application has achieved state-of-the-art results in both full-training settings and zero-shot settings, making the model have better generalization ability in downstream tasks.
[0088] Next, through the description of several optional embodiments, the technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure will be described. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0089] Figure 2 This is a flowchart of a method performed by an electronic device provided by the present application. The execution subject of the method may be an electronic device, for example, the electronic device may be any device such as a server, a terminal or a cloud computing center device, and the present application does not limit this. Figure 2 As shown, the method includes the following steps.
[0090] Step 201: The electronic device obtains a text to be processed.
[0091] The text to be processed is a text that needs to be processed for event extraction, and the text to be processed can be a natural language text. For example, a sentence input by a user, or a paragraph including multiple sentences, an article, etc. In this application, a trained AI (Artificial Intelligence) network can be used to extract the event extraction result of the text to be processed; the AI network is used to extract events from event texts, and the AI network can also be called an event extraction network.
[0092] Step 202: The electronic device uses a trained artificial intelligence AI network to extract event extraction results of the text to be processed based on the text to be processed and the first prompt information.
[0093] The first prompt information is used to prompt extraction based on at least one event definition information in the candidate event definition set. The candidate event definition set includes event definition information of each candidate event in the candidate event set; the candidate event set includes at least one predefined candidate event. The event definition information of an event can be information used to describe the event or summarize the characteristics of the event.
[0094] For example, an event definition information may be a concept text of a candidate event. For example, the concept text corresponding to the Be-Born event is: "whenever a person is given birth to, there is an Be-Born event of LIFE category"; its corresponding Chinese expression is "whenever a person is given birth to, there is a Be-Born event of LIFE category".
[0095] In a possible implementation, the first prompt information is prompt information related to the text to be processed. Before step 202, the following steps are further included: The electronic device obtains the first prompt information of the text to be processed. Correspondingly, step 202 is replaced with: The electronic device, based on the text to be processed and the first prompt information of the text to be processed, uses the trained AI network to extract the event extraction result of the text to be processed.
[0096] Among them, the at least one event definition information prompted by the first prompt information may include at least one associated event definition information associated with the text to be processed in the candidate event definition set; that is, the first prompt information is used to prompt extraction based on at least one associated event definition information associated with the text to be processed in the candidate event definition set. Correspondingly, the first prompt information is obtained through the following steps A1 - step A2:
[0097] Step A1: Retrieve at least one associated event definition information associated with the text to be processed from the candidate event definition set;
[0098] Step A2: Obtain the first prompt information based on the at least one associated event definition information, and the at least one event definition information includes the at least one associated event definition information.
[0099] In this application, based on the text to be processed and the candidate event definition set, a retriever can be used to retrieve at least one associated event definition information associated with the text to be processed. For example, the retriever can be used to obtain the relevance between the text to be processed and each event definition information in the candidate event definition set, so as to screen out a specified number of associated event definition information that meets the target conditions from the candidate event definition set. The relevance indicates the degree of relevance between the text to be processed and the event definition information. The target conditions may include, but are not limited to: the first specified number among the event definition information arranged in descending order of relevance, the relevance higher than the target relevance threshold, etc.
[0100] Such as Figure 3As shown, the set of candidate event definitions can be stored in a DB (Database). Based on the event context text for which event extraction is to be performed and each event definition information (such as each concept text), the Retriever module can be used to retrieve the top 5 event definition information with the highest degree of relevance to the text to be processed as the associated event definition information by using a pre-configured retrieval method. For example, the pre-configured retrieval algorithms include, but are not limited to: BM25 (Best Matching 25) method, FAISS (Facebook AI Similarity Search) method, sentence-BERT method, SimCSE (Simple Contrastive Learning of Sentence Embeddings) method.
[0101] In step A2, the first hint information can be directly obtained according to each associated event definition information. Alternatively, the first hint information can also be obtained according to the event type information corresponding to each associated event definition information. Correspondingly, the implementation manner of step A2 can include any one of the following steps A21 and A22:
[0102] Step A21: Concatenate the at least one associated event definition information and the second hint information to obtain the first hint information;
[0103] Step A22: Concatenate the at least one event type information corresponding to the at least one associated event definition information and the second hint information to obtain the first hint information.
[0104] The second hint information refers to extraction according to specific event definition information. For example, the second hint information can be "According to the definitions of events". In step A21, the specific event definition information refers to the at least one associated event definition information concatenated. In step A22, the specific event definition information refers to the at least one associated event definition information corresponding to the at least one event type information concatenated.
[0105] In step A21, the information of each associated event definition can be sequentially concatenated after the second prompt message. Sequential concatenation can be to concatenate the information of each associated event definition after the second prompt message in the order of the degree of relevance between the information of each associated event definition and the text to be processed. For example, the first 5 event concept texts most relevant to the text to be processed can be sequentially concatenated together according to the degree of relevance, denoted as [Retrieved Event Concepts], and concatenated after the second prompt message; where Retrieved Event Concepts refers to the retrieved event concept text.
[0106] Among them, the first prompt message can be: According to the definitions of events, [Retrieved Event Concepts]. For example, if Retrieved Event Concepts are the first 5 retrieved event concept texts, sorted in descending order of relevance as Concept1, Concept2,..., Concept5; then the first prompt message can be: According to the definitions of events, Concept1+Concept2+Concept3+Concept4+Concept5.
[0107] In step A22, the information of each event type can be concatenated after the second prompt message. The event type information corresponding to the information of the associated event definition indicates the type of the event corresponding to the information of the associated event definition. For example, after retrieving the first 5 information of the associated event definitions most relevant to the text to be processed, the 5 event type information corresponding to the first 5 information of the associated event definitions can be further obtained, denoted as [Retrieved Event Subtypes], and concatenated after the second prompt message; where "Event Subtypes" are the retrieved event type information, obtained based on the information of the associated event definition.
[0108] In this application, multiple candidate events can be predefined. For example, as shown in Table 1 below, in the ACE2005 dataset, 8 major categories of events are predefined, including a total of 33 subcategories of events. The candidate event set can include 33 types of candidate events shown in Table 1, as specifically shown in Table 1 below:
[0109] Table 1
[0110]
[0111] As shown in Table 1, the Types column in Table 1 shows 8 major types of events included in the ACE2005 dataset; among them, event type information can be used to represent different types of events, and the Subtype column shows the event type information of 33 types of events included in 8 major types.
[0112] In step A22, the first prompt message can be: According to the definitions of events, [Retrieved Event Subtypes]. For example, [Retrieved Event Subtypes] are the event type information corresponding to the first 5 retrieved event definition information, sorted in descending order of relevance as Subtype1, Subtype2,..., Subtype5, and one of the Subtypes can be one type keyword shown in Table 1, such as Be - Born, or Marry, or Die...; then the first prompt message can be expressed as: According to the definitions of events, Subtype1 + Subtype2 + Subtype3 + Subtype4 + Subtype5.
[0113] It should be noted that for "According to the definitions of events, [Retrieved Event Concepts]" in step A21 and "According to the definitions of events, [Retrieved Event Subtypes]" in step A22, the "the definitions of events" therein both refer to the definitions of the retrieved events concatenated thereafter. For example, in step A21, "the definitions of events" refers to the 5 directly concatenated Event Concepts thereafter, and in step A22, "the definitions of events" refers to the definitions corresponding to the 5 Event Subtypes concatenated with it.
[0114] In yet another possible implementation, the first prompt message may also be a prompt message that is common to each event text. The at least one event definition information prompted by the first prompt message may include each event definition information in the candidate event definition set; that is, the first prompt message may prompt extraction based on each event definition information in the candidate event definition set. Among them, the first prompt message may be: According to the definitions of events; regardless of which event file, this prompt message is used for prompting.
[0115] In step 202, the electronic device may input the text to be processed and the first prompt message into the event extraction network, and use the event extraction network to perform event extraction on the text to be processed to obtain an event extraction result.
[0116] The following introduces how to perform event extraction on the text to be processed:
[0117] Event extraction may be a process of extracting description information of the event composition units of the text to be processed. In one possible way, in step 202, the event extraction result is obtained by the event extraction network by performing the following steps B1 - B2:
[0118] Step B1: Based on the text to be processed and the first prompt message, extract the description information of the event composition units of the text to be processed;
[0119] Step B2: Based on the description information of the event composition units, obtain the event extraction result of the text to be processed.
[0120] Exemplarily, the event composition unit refers to the unit that constitutes the event corresponding to the text to be processed. The event composition unit may include but is not limited to: the type of the event, the event trigger word, the time of the event, the location of the event, and other event elements. The description information of the event composition unit can be used to describe the unit value of the text to be processed corresponding to each event composition unit.
[0121] In one possible example, the JSON format may be used to structurally represent each event composition unit and its description information. For example, for the event trigger word, the main participant of the event, the time, and the location, they can be represented in sequence as follows: "event trigger": "xxx", "subject participant": "xxx", "time": "xxx", "place": "xxx".
[0122] In another possible example, complete sentences can also be used to represent each event component unit and its description information. For example, for the event trigger word, the main participant of the event, the time, and the place, they can be represented in sequence as follows: The event trigger is xxx, the subject participant is xxx, the time is xxx, the place is xxx.
[0123] Exemplarily, the text to be processed can correspondingly express one or more target events, and the description information of the corresponding event component units can be obtained for each target event. In one possible way, the implementation manner of step B1 can include the following steps C1:
[0124] Step C1: Based on the text to be processed and the first prompt information, extract the event type information of at least one target event associated with the text to be processed, the description information of at least one general event element corresponding to each target event, and the description information of the event trigger element.
[0125] Among them, the event component units of the text to be processed include each event component unit corresponding to each target event associated with the text to be processed, and each event component unit includes an event type, each general event element, and an event trigger element.
[0126] Among them, the at least one general event element is an element in the general element set, and the general element set includes the general event elements corresponding to each candidate event in the candidate event set corresponding to the candidate event definition set; that is, each candidate event commonly corresponds to a set of general element sets. For example, 33 types of events in Table 1 commonly correspond to a set of general element sets.
[0127] Among them, the event trigger element is an element that can trigger the occurrence of an event. For example, the event trigger element can be an event trigger word, and the event trigger word can be a verb that triggers the occurrence of an event.
[0128] The target event associated with the text to be processed refers to the event described by the natural language of the text to be processed. A text to be processed can describe one or more target events; for example, a sentence can describe one event, and a natural language paragraph or an article can describe or express multiple events.
[0129] In step C1, for each target event, the description information of the general event elements corresponding to the target event, the description information of the event trigger elements, and the event type information can be extracted. For example, if a passage contains both Be-Born events and Marry events, then the description information of the general event elements, the event trigger words, and the event type information of the Be-Born events can be extracted, and the description information of the general event elements, the event trigger words, and the event type information of the Marry events can be extracted. Correspondingly, in step C2, the event extraction results corresponding to each target event associated with the text to be processed can be obtained.
[0130] In a possible way, based on different first prompt information, the at least one target event includes any one of the following (1)-(2):
[0131] (1) If the at least one event definition information includes at least one associated event definition information associated with the text to be processed, the at least one target event is an event associated with the text to be processed among the at least one event corresponding to the at least one associated event definition information;
[0132] (2) If the at least one event definition information includes each event definition information in the candidate event definition set, the at least one target event is an event associated with the text to be processed in the candidate event set.
[0133] In (1), the first prompt information is used to prompt extraction based on at least one associated event definition information associated with the text to be processed; correspondingly, the at least one target event associated with the text to be processed is within the event range corresponding to the retrieval result of step A1; for example, the target event associated with the text to be processed is at least one of the 5 most relevant events retrieved.
[0134] In (2), the first prompt information is used to prompt extraction based on each event definition information in the candidate event definition set; correspondingly, the at least one target event associated with the text to be processed is within the candidate event set range corresponding to the candidate event definition set; for example, the target event associated with the text to be processed is at least one of the 33 events in the candidate event set.
[0135] In this application, event extraction can be to extract the description information of predefined event constituent units from unstructured natural language text. The candidate event set shares a set of general event element sets; and each candidate event in the candidate event set can respectively correspond to its own dedicated event element set. In step 202, the description information of the general event elements can also be converted into the description information of the dedicated event elements corresponding to each target event.
[0136] In one possible way, the implementation of step C2 may include the following steps C21 - C22:
[0137] Step C21: Based on the event type information of each target event, convert the description information of each general event element corresponding to each target event into the description information of each special event element in the set of special event elements corresponding to each target event;
[0138] Step C22: Based on the description information of each special event element corresponding to each target event and the description information of the event trigger element, obtain the event extraction result corresponding to each target event.
[0139] Each candidate event may correspond to its own set of special event elements. Exemplarily, for each target event, a corresponding set of special event elements can be obtained based on the event type information of the target event, so as to convert the description information of each general event element corresponding to the target event into the description information of the corresponding special event elements. Among them, the special event elements corresponding to different event types can be different. For example, a set of special elements corresponding to type A may include: special element A1, A2, A3; a set of special elements corresponding to type B may include: special element B1, B2, B3. Of course, the same special element may be included in different groups of special elements. For example, both type A and type B include a time element.
[0140] For example, if a passage contains both event 1 of the Be - Born type and event 2 of the Marry type; then, according to the Be - Born type of event 1, the description information of each general event element corresponding to event 1 can be converted into the description information of the special event elements corresponding to the Be - Born type; and, according to the Marry type of event 2, the description information of each general event element corresponding to event 2 can be converted into the description information of the special event elements corresponding to the Marry type.
[0141] In step C22, the event extraction result corresponding to each target event means that each target event may correspond to its own event extraction result.
[0142] In one possible way, the description information of a general event element includes at least the general event element; the conversion process in step C21 can be achieved by converting the general event element. Exemplarily, the implementation of step C21 may include the following step C21 - 1:
[0143] Step C21 - 1: Based on the event type information of each target event, convert the general event elements in the description information of each general event element corresponding to each target event to obtain the description information of each special event element corresponding to each target event.
[0144] Exemplarily, for each target event, the general event elements in the description information of the general event elements corresponding to the target event can be converted into the corresponding specific event elements in the set of specific event elements of the target event; thus, the description information of the specific event elements corresponding to each target event is obtained.
[0145] In a possible way, the description information of a general event element includes the general event element and the element value of the general event element; for example, the description information of a general event element is "subject participant": "A". Then the general event element is the subject participant, and the element value is A. In step C21, it can be to replace the general event element in the description information. Exemplarily, the implementation manner of step C21 may include the following steps D1 - D3:
[0146] Step D1: Based on the event type information of the respective target events, determine the respective target specific element sets of the respective target events, where the target specific element sets include the respective specific event elements corresponding to the general event elements associated with the corresponding target events;
[0147] Step D2: Use the respective specific event elements in the respective target specific element sets of the respective target events to replace the general event elements in the description information of the general event elements of the corresponding target events, and obtain the description information of the specific event elements corresponding to each target event;
[0148] Step D3: Based on the description information of the specific event elements corresponding to the respective target events and the description information of the event trigger elements, obtain the event extraction results corresponding to each target event.
[0149] Among them, there is a mapping relationship between the general event element set and the set of specific event elements corresponding to each event type. That is to say, each specific event element corresponding to each event type has a corresponding general event element in the general element set. In this step, for each target event, the target specific element set corresponding to the target event can be obtained based on the event type information of the target event; and based on the mapping relationship between the general element set and the target specific element set corresponding to the target event, the general event element in the description information of the general event element can be directly replaced with the corresponding specific event element, and thus the description information of the specific event elements of the target event can be obtained.
[0150] As shown in Table 2 below, Table 2 shows the specific event elements corresponding to several event types:
[0151] Table 2
[0152]
[0153] As can be seen from Table 2, different events have some common elements, such as the implementer of the event, the recipient of the event, the medium, the time, the location, etc. The attributes in this semantic structure have good domain invariance. That is to say, events in different domains have common elements in the semantic structure, and the common elements in different events have good domain invariance.
[0154] Based on this, in this application, a general event element set can be uniformly constructed based on the dedicated event elements with common characteristics among each candidate event. In one possible way, the construction method of the general event element set includes: defining the dedicated event elements with common characteristics in the dedicated element sets corresponding to each candidate event as the same general event element, and constructing the general element set. Based on this, the mapping from the dedicated element naming in the original dataset to the new general element naming can be completed.
[0155] As shown in Table 3 below, Table 3 shows the mapping relationship between each general event element and each dedicated event element:
[0156] Table 3
[0157] General event element Special event element Subject participant includes person / agent / giver... Object participant includes person / org / recipient.. Third-party affected participant includes beneficiary-arg.. Medium for subject participant includes vehicle / artifact... Place includes start place / destination place...
[0158] For example, for the Life:BE-BORN event type, the subject is person, for the Business:START-ORG event type, the subject is agent; for the Conflict:ATTACK event type, the subject is attacker; then these subjects can be uniformly renamed as Subject participant. Specifically, the mapping relationship between the general event element set and the dedicated event element sets of each candidate event is as shown in Table 2 and Table 3 above, and will not be listed and elaborated one by one here.
[0159] It should be noted that since different events have some common elements, such as the implementer of the event, the recipient of the event, the medium, the time, the location, etc., the attributes in this semantic structure have good domain invariance. In this application, by constructing the mapping relationship between the general event element set and the dedicated event sets of each candidate event, during extraction, the description information of the general event elements can be extracted first, effectively protecting the similarity of these elements in the semantic structure; then the conversion from the general event elements to the dedicated event elements is performed, which further improves the robustness and generalization of the network model without affecting the final event extraction result.
[0160] Moreover, in this application, by splicing at least one associated event definition information or its corresponding event type information, a first prompt message is obtained, so as to directly list the event definition information or event type related to the event text as part of the spliced context, and explicitly list the retrieved related event definitions or related event types in the input text, effectively establishing the association between the event instance and the related event concept, reducing the difficulty of establishing the association between the instance and the concept by the network model, and helping to improve the accuracy of event extraction.
[0161] Figure 3 Shows the retrieval process corresponding to the retrieval module. Figure 4 Shows a network structure included in an event extraction network. As Figure 4 shown, the event extraction network may include an encoder-decoder network structure. The following combines Figure 3 and Figure 4 to further introduce the overall process of this application:
[0162] As Figure 3 shown, taking the manner of step A21 as an example, the input of the retrieval module is the text to be processed; at least one associated event definition information can be retrieved from the DB by the retrieval module, and a first prompt message is generated based on the at least one associated event definition information and the second prompt message. Among them, the input (Input event) of the encoder includes an event file with prompt information, specifically including the text obtained by splicing the text to be processed and the first prompt message. The encoder is used to extract the implicit features of the text to be processed and the first prompt message, and input the implicit features of the text to be processed and the first prompt message into the Figure 4 shown decoder, and the decoder is used to output the generated prediction, that is, the event extraction result.
[0163] For example, splice the first prompt message and the text to be processed to obtain an instance, and input the instance into the encoder of the event extraction network to obtain the result output by the decoder, specifically as follows:
[0164] According to the definitions of events, [Retrieved Event Concepts], does the following context contain an event? If so what’s the event’s arguments?
[0165] Context:[passage]
[0166] [Model output]:
[0167] {"contain event": "yes", "event name": "xxx", "event trigger": "xxx",
[0168] "subject participant": "xxx", "object participant": "xxx", "third - party involved": "xxx",
[0169] "medium for subject participant": ”xxx”, "time": "xxx", "place": "xxx"}。
[0170] Among them, the description information of the general event elements output by the decoder is listed in the above [Model output]; further, a conversion module can be used to convert the description information of each general event element in [Model output] by using the mapping relationship between the set of general event elements and the set of dedicated event elements of each candidate event, so as to obtain the description information corresponding to each dedicated event element.
[0171] Among them, the above [Model output] corresponds to the case where the instance contains an event; if the instance does not contain an event, the [Model output] is: {"contain event": "no"}.
[0172] In some embodiments, the present application can pre - train the event extraction network, and the training method of the event extraction network will be introduced below.
[0173] Figure 5 It is a schematic flow chart of a training method for an event extraction network provided by the present application. As Figure 5 shown, the training method of the event extraction network includes the following steps.
[0174] Step 501, the electronic device obtains a training data set.
[0175] Among them, the training data set includes first training data, and the first training data includes a plurality of first samples. Each first sample includes a first sample input and a first label. The first sample input includes sample text and the first prompt information, and the first label represents the true event extraction result corresponding to the sample text.
[0176] It should be noted that each sample text corresponds to its own first hint information. The first hint information of the sample text is used to prompt extraction based on at least one associated event definition information associated with the sample text. The electronic device can retrieve at least one associated event definition information associated with each sample text from the candidate event definition set, and obtain the first hint information of the sample text based on at least one associated event definition information associated with each sample text. Among them, this retrieval step can be executed by a trained retrieval module, and it is default that the retrieval module is trained and can give accurate retrieval results. For example, the first hint information of the sample text can be: According to the definitions of events, [Retrieved Event Concepts]. The acquisition method and steps of the first hint information of the sample text are the same as those in step A2, and will not be elaborated here one by one.
[0177] In a possible example, the true event extraction result of the sample text may include the true description information of at least one event constituent unit of the sample text. Exemplarily, the at least one event constituent unit includes an event type, at least one general event element, and an event trigger element.
[0178] The electronic device can iteratively perform the following training operations on the initial AI network based on the training data set until a preset condition is met to obtain the event extraction network. The training operations include the following steps 502-step 504:
[0179] Step 502: The electronic device uses the initial AI network based on each first sample input to extract the first event extraction result of each sample text.
[0180] Exemplarily, as Figure 4 shown, the initial AI network includes an encoder and a decoder. The implementation manner of step 502 includes: based on each first sample input, using the encoder in the initial AI network to obtain the implicit features of each sample text and its first hint information; based on the implicit features of each sample text and its first hint information, using the decoder in the initial AI network to obtain the first event extraction result of each sample text.
[0181] Among them, the first event extraction result of each sample text includes the event type information of at least one target event associated with the sample text, as well as the description information of at least one general event element and the description information of the event trigger element of each target event. The first event extraction result can be output in JSON format. It should be noted that using the JSON format can explicitly describe each element of the event and enhance the network model's ability to describe the event structure.
[0182] Step 503: The electronic device determines a first training loss based on the first tags and first event extraction results of the various sample texts.
[0183] Step 504: The electronic device adjusts the network parameters of the initial AI network based on the total training loss, where the total loss includes the first training loss.
[0184] The electronic device can determine the first training loss based on the description information of each event component of each sample text and the true description information of the corresponding event component in the first tag of each sample text. For example, a cross-entropy loss function can be used to calculate the first training loss corresponding to each sample text.
[0185] In a possible example, the processes of the above steps 502-504 can be the training process of the first training task for an instance, where the instance refers to a sample instance constructed by splicing the first prompt information and the sample text; the first training task corresponds to the first training loss Loss1. In the first training task, the input of the initial AI network is the sample text + the first prompt information. For example, "sample text + According to the definitions of events, [Retrieved Event Concepts]". The output is the event type information of the sample text, the event trigger, and the description information of each general event element.
[0186] In a possible implementation, the event definition information of each candidate event is learned by the initial AI network based on sample data. The training data set further includes second training data, where the second training data includes multiple second samples, each second sample includes a second sample input and a second tag, the second sample input includes an initial definition of a candidate event, and the second tag represents the true event type corresponding to the initial definition of the corresponding candidate event.
[0187] Among them, the total training loss further includes a second training loss. Correspondingly, this training operation further includes steps E1-E2:
[0188] Step E1: Based on each second sample input, using the initial AI network, extract the event type information corresponding to each candidate event.
[0189] Exemplarily, as Figure 4 shown, based on each second sample input, the encoder in the initial AI network can be used to obtain the implicit features of the initial definition of each candidate event; based on the implicit features of the initial definition of each candidate event, the decoder in the initial AI network is used to output the event type information corresponding to the initial definition of each candidate event.
[0190] Step E2. Determine the second training loss based on the second tags corresponding to the initial definitions of the candidate events and the extracted event type information.
[0191] Exemplarily, the initial definition of each candidate event can be a natural language text used to define the candidate event, such as a sentence or a passage of text. Additionally, the initial definition of each candidate event may include multiple synonymous definition texts for the candidate event, where the multiple synonymous definition texts have the same meaning but differ in literal text, that is, the individual words or phrases included in the text may be different.
[0192] In a possible example, for Steps E1 - E2, it can be the training process of a second training task for the event type corresponding to the concept. This second training task corresponds to a second training loss Loss2. In the second training task, the input to the initial AI network includes the concept definitions of the candidate events, where the same event can correspond to multiple sets of synonymous description texts, such as a candidate event corresponding to 5 synonymous sentences; the output is the event type information corresponding to the concept definition of the candidate event, for example, the Life Be - born event. Exemplarily, the extracted event type information can be in JSON format.
[0193] In a possible example, the second sample input may further include associated learning prompt information; the associated learning prompt information is used to prompt associated learning between the training task corresponding to the first training data and the training task corresponding to the second training data. For example, the second sample input can be: Here is a definition of an event. whenever a person is given birth to, there is an BE_BORN event of LIFE category. Among them, "Here is a definition of an event" is the associated learning prompt information, and "a definition of an event" in this associated learning prompt information has the same meaning as "the definitions of events" in the first prompt information in the training task corresponding to the first training data. Based on this, in multi - task training, associated learning can be enabled between the two training tasks through the associated learning prompt information, so that the training task corresponding to the first training data can also learn the knowledge of the training task corresponding to the second training data.
[0194] In a possible example, the second sample input may further include element hint information of the candidate event, and the element hint information is used to hint at the general event elements corresponding to the candidate event. For example, the second sample input may be: Hereis a definition of an event; whenever a person is given birth to, there is anBE_BORN event of LIFE category; the BE_BORN event contains subject participantrefers to the person who is born, time when the birth takes place, and placewhere the birth takes place. Among them, "the BE_BORN event contains subjectparticipant refers to the person who is born, time when the birth takes place, and place where the birth takes place" is the element hint information, which hints at what general elements the BE_BORN event corresponds to. Based on this, the network model can further learn the definitions of each event and the knowledge of the event elements included, thereby improving the accuracy of event extraction.
[0195] It should be noted that in this application, the concept of an event and the instance of an event can be split first, that is, the knowledge of the event concept and the knowledge of parameter extraction of the instance are split, and independent data are constructed respectively. Taking the ACE2005 LIFE-Marry event as an example, the split is as follows:
[0196] Event concept knowledge:
[0197] Here is a definition of an event. Whenever a person is given birth to, there is an BE_BORN event of LIFE category.
[0198] Event instance:
[0199] Does the following context contain an event? If so what’s the event’sarguments?
[0200] Context:[passage]
[0201] [Model output]:Yes,the context contains a xxx event.Event trigger:xxx。
[0202] Moreover, the association between concepts and instances is created through hint messages.
[0203] For example, association learning hint messages are concatenated in the event concept, specifically as follows: Here is a definition of an event. Whenever a person is given birth to, there is an BE_BORN event of LIFE category. And, the first hint message is concatenated in the event instance, specifically as follows: According to the definitions of events, [Retrieved Event Concepts], does the following context contain an event? If so what’s the event’s arguments? Context: [passage]. Through "a definition of an event" and "the definitions of events", by establishing the association learning between the training task corresponding to the second training data and the training task corresponding to the first training data, the network model learns the association between concepts and instances while learning independent concepts.
[0204] In a possible implementation manner, the training data set further includes third training data, the third training data includes a third sample corresponding to each sample text, each third sample includes a third sample input and a third label, and the third sample input includes the sample text, the third hint message of the sample text, and a fourth hint message;
[0205] Wherein, the third hint message prompts to extract event definition information of a target event associated based on the sample text; the fourth hint message prompts to extract at least one target element corresponding to the sample text; the at least one target element includes at least one of an event type, each general event element, or an event trigger unit; the third label represents the true description information of each target element corresponding to the sample text;
[0206] Among them, the total training loss further includes a third training loss. Correspondingly, the training operation further includes steps F1 - F2:
[0207] Step F1: Based on each third - sample input, use the initial AI network to extract the second event extraction result of each sample text, where the second event extraction result includes the description information of each target element corresponding to the sample text;
[0208] Among them, the third prompt information of a sample text may include: According to definition of the event, [True Event Concept]. Among them, [True Event Concept] is the event - definition information corresponding to the true event type of this sample text, which can be considered as the precise concept of this sample text.
[0209] The fourth prompt information of a sample text can be considered as the task - prompt information of this sample text, that is, it prompts which one or several target elements this task is to extract from this sample text.
[0210] Figure 6 shows a possible sample data structure. As Figure 6 shown, the at least one target element can be selected from event types, event trigger words, at least one general element (such as event participants, locations, times). Figure 6 Among them, the right - hand entries of event type, event trigger word, event participant, location, and time are their respective corresponding prompt texts. For example, if the at least one target element includes event trigger, the fourth prompt information includes: The trigger word that best indicates the action is <trigger>Among them, " <trigger>"It means that the target elements to be extracted include trigger.
[0211] For example, when the input of an instance is Figure 6 the event text + event type + trigger + each general element in, the corresponding output contains descriptive information of elements such as event trigger, subject participant, object participant, medium, time, place, etc.; which corresponds one by one to the hint text of the target elements included in the input. Another example, when the input of an instance is Figure 6 the event text + a certain target element in, the corresponding output is the predicted descriptive information of this target element.
[0212] Exemplarily, as Figure 4 shown, the input (Input event text) of the encoder in the initial AI network can include each third sample input, obtaining implicit features corresponding to each sample text and the third hint information spliced with the real event definition information of each sample text.
[0213] As Figure 4 shown, in the training task corresponding to the third training data, the input of the decoder includes each sample text output by the encoder, the real event definition information, the implicit features corresponding to the third hint information, and the fourth hint information. Figure 4 In, this fourth hint information (prompt template text) can be called "prefix hint text for decoding". For example, if at least one target element includes event trigger and event participants, then Figure 6 the right entries corresponding to the trigger and event participants in, that is, "The trigger word that best indicates the action is <trigger>” and "In this event, the subject participant (doer of the event) is <subject>,the object participant being acted upon is <object>", as a prefix, prompts the input of the decoder with text; and uses the decoder to obtain the description information of the trigger and event participants of the sample text.
[0214] Step F2: Determine the third training loss based on the third tags and the second event extraction results of the respective sample texts.
[0215] Exemplarily, the fourth prompt information further includes the element definition information of each target element corresponding to the sample text. The element definition information is information used to define the corresponding event element, and the corresponding event element can be an event trigger element or a general event element.
[0216] In a possible example, the process of the above Step F1 - Step F2 can be a training process for the third training task of "exact concept + instance type"; the instance refers to a specific sample text instance; the concept refers to the true event definition information corresponding to each sample text.
[0217] This third training task corresponds to the third training loss Loss3. In the third training task, the input of the initial AI network is "sample text + third prompt information + fourth prompt information"; the output is the prediction result corresponding to at least one target element of the sample text.
[0218] In a possible way, the at least one target element can be an element selected from the event type, at least one general event element, and the event trigger unit. In another possible way, the at least one target element can be selected by combining the prediction performance information corresponding to each event element in at least one training process; for example, the prediction performance information corresponding to each event element in at least one training process can be obtained; for example, the accuracy rate; based on the prediction performance information corresponding to each event element in at least one training process, select the at least one target element that meets the preset conditions from at least one general event element and the event trigger element. For example, select the event element with an accuracy rate lower than a certain threshold as the target element.
[0219] In a possible way, the initial AI network can also be trained by combining the above three training tasks. For example, based on the above steps 502 - 503, steps E1 - E2, and steps F1 - F2, the total training loss can be obtained to execute step 504. Based on this, through the hybrid multi-task learning of "concept type", "exact concept + instance type", and "retrieved related concept + instance type", the model can learn independent concepts and establish the association between concepts and instances at the same time. When the domain changes, the model has better robustness in understanding concepts and the transferability is improved.
[0220] It should be noted that in the training task corresponding to the third training data, "According to definition of the event" in the third prompt information includes "definition of the event", and "definition of the event" has the same meaning as "a definition of an event" and "the definitions of events" in the previous two training tasks. Based on this, in the multi-task training including 3 training tasks, by associatively learning the prompt information, the first prompt information, and the third prompt information, the three training tasks can be associatively learned, so that the knowledge of the training task corresponding to the second training data can also be learned in the training tasks corresponding to the first training data and the third training data. Based on this, by establishing the associative learning between these three training tasks, the associative learning of the network model for concepts and instances is further promoted.
[0221] It should be noted that the generation of the target element is applied to the above 3 training tasks (the training tasks corresponding to the first training data, the second training data, and the third training data respectively). Learn event concept knowledge and event extraction task knowledge. For event concept knowledge, that is, the training process of steps E1 - E2 above, the network model generates an event name and an event parameter description. For the event detection task, the model needs to generate an event name and an event trigger word. For the event element extraction task, the network model needs to generate a description information of the general event element.
[0222] The present application can use cross-entropy loss to implement the above training process. For example, the following formula (1) can be used to obtain any of the above training losses:
[0223]
[0224] where N is the number of all labeled parameters in a mini-batch, C is the number of parameter types, x n,i is the logit of the parameter of each type; x n,c is the logits of the target parameter; where logit refers to the output vector of the decoder. y n,c is the true label of the target parameter. w c represents the weight of the parameter of each type.
[0225] In one possible way, for each sample text, the encoder in the initial AI network can also be trained by combining the description information of the general event element extracted from each sample text.
[0226] Exemplarily, the training operation further includes: determining a fourth training loss based on the similarity between the general event elements of each sample text. Among them, the implicit features output by the encoder can be used to determine the similarity between each general element. For example, the sample text and the position identifier of the sample text can be input into the encoder, and the implicit features of the sample text output by the encoder can be obtained; for each general event element of each sample text, the implicit features of the description information of the general event element can be obtained from the implicit features of the sample text based on the position marker of the general event element in the sample text; based on this, the implicit features of the description information of at least one general event element in each sample text can be obtained.
[0227] Then, determine the similarity between the implicit features of the description information of the same general event element in each sample text to obtain the fourth training loss. For example, the description information of the same general event element can be the element value of the subject element of the marry event and the element value of the subject element in the be-born event; or, the element value of the time element of the marry event and the element value of the time element in the be-born event.
[0228] Specifically, using the positions of each known event element in the sample text, extract the feature representation of the event element sequence marker from the output of the encoder through position indexing, and construct the feature representation of the event element through average pooling. In a batch, the same type of general event elements in different samples will be pulled closer, while different types of general event elements will be pulled apart due to the contrast loss. This enables the model to find syntactic and semantic similarities of event elements in different event types and different event contexts.
[0229] The following formula (2) can be used to calculate the fourth training loss:
[0230]
[0231] where x + represents a positive sample, for example, the subject participant of one of the batch inputs. x represents all negative samples, such as object participants, third-party affected participants, time, location, etc., which should be separated from the subject participant in the semantic space. c represents the context, that is, the text composition of the first sample input and the text composition of the third sample input; the specific context is as Figure 7 shown. f enc represents the feature vector output by the encoder. E represents the mathematical expectation.
[0232] like Figure 4 As shown, the contrastive loss function can be used to calculate the fourth training loss. For example, a group of triangles, a group of rectangles, and a group of pentagons can represent the description information of the general event element A, the general event element B, and the general event element C corresponding to each sample text, respectively. Contrastive loss can be used for training to shorten the distance between the same elements and increase the distance between different elements.
[0233] Correspondingly, step 504 includes: the electronic device adjusts the network parameters of the encoder and the decoder in the initial AI network based on the first training loss; and adjusts the network parameters of the encoder based on the fourth training loss.
[0234] The general event elements of the sample texts may include at least one of the following: general event elements of the sample texts corresponding to the first sample input; and general event elements of the sample texts corresponding to the third sample input.
[0235] It should be noted that the final training loss can be the weighted sum of various training losses; for example, the final training loss can be calculated by the following formula: Among them, the first item on the right side of the formula "=" can be considered as a training loss calculated based on the cross entropy loss, such as at least one of the first training loss, the second training loss, and the third training loss; the second item on the right side is the fourth training loss calculated based on the contrast loss, wherein α can be the weight of the fourth training loss.
[0236] It should be noted that by comparative learning of the classification of common event elements on the encoder side, the transferability is enhanced by learning the similarity of similar event elements. On the encoder side, several sample texts input in the same batch, the same type of event elements corresponding to each sample text, such as subject-participant, time, place, etc., in the feature space, shorten the distance between the features of the description information of the same type of event elements, and shorten the distance between the features of the description information of different types of event elements, and enhance the transferability by learning the similarity of the same type of event elements.
[0237] The method executed by an electronic device provided in this application extracts an event extraction result of the text to be processed through a trained event extraction network based on the text to be processed and the first prompt information; the first prompt information is used to prompt extraction based on at least one event definition information in a candidate event definition set. Utilizing the first prompt information enables the AI network to perform extraction in combination with at least one event definition information, effectively improving the accuracy of event extraction.
[0238] It should be noted that to solve the problems existing in the related art, this application proposes an event extraction method based on prompt-based reusable concept learning. Different from existing prompt learning, this application designs a reusable concept learning method, enabling the network model to independently learn event concepts and apply event concept knowledge to various instances. The network model proposed in this application is a network model with an encoder-decoder as the backbone network as shown in Figure 4
[0239] First, this application decouples event concept knowledge from event samples, and event concept learning and event detection / event argument detection are constructed as independent tasks. Second, by injecting the same textual pattern into each task data, the connection between event concepts and event samples is established. Third, through further research, it is found that many semantically similar elements are given different element names in different events. This application reconstructs different event elements of various event types into a unified set of general elements, and uses contrastive learning to enable the network model to learn the universal semantic characteristic of the event arguments, thereby further enhancing the generalization ability of the network model. This network model outperforms the current state-of-the-art models in the ACE 2005 dataset with complete training data and zero-shot evaluation of WIKIEVENTS, and constructs separate prompts for event concepts and event samples.
[0240] The method of this application will be introduced systematically as follows:
[0241] In the prompt method for solving the event extraction task, the uniqueness of this application mainly lies in the following two aspects.
[0242] First, the event detection task is highly human-oriented. The same semantic role is often given different parameter names in different events, which is convenient for human experts to match parameters with event instances. In fact, the scattered naming method hardly provides additional information, but instead obscures the common semantics among these parameters. These common semantics are more general in different fields. By mapping different parameters to a unified set of event parameters, the event extraction task is reformulated. How to utilize these shared role semantics to obtain more robust performance for various downstream fields is our second design focus.
[0243] Secondly, conceptual knowledge is universal across various fields. In related technologies, concepts are tightly coupled with each sample, restricting the model to solve ED / EAE tasks based only on the conceptual text of a single sample. We design an independent event concept learning task for the model to obtain complete knowledge about each event.
[0244] To establish the connection between event concept knowledge and the event extraction task, this application designs two strategies: (1) In the event concept learning task and the event detection / event argument extraction task, inject a shared textual pattern that hints at the event concept definition, and let the co-occurrence of this shared textual pattern be a connection clue between different tasks. (2) In the event detection / event argument extraction task, retrieve the top k most relevant event concepts of the event sample and splice them into the second hint information to make the first hint information more explicit.
[0245] Next, the concept and system architecture of this application are further introduced. The communication of this application mainly includes the following 1.1 - 1.4.
[0246] 1.1. Unified Event Parameters across Events
[0247] Each event has its own named parameters, and parameters of the same role have different names. In this application, according to the shared semantic role of these parameters in the event, these parameters are uniformly named. For example, as shown in Tables 2 and 3 in the above embodiments, in the ACE2005 dataset, the parameter buyer (purchaser) in the TRANSFER-OWNERSHIP event, the parameter giver in the TRANSFER-MONEY event, and the parameter agent in the START-ORG event are all unified as subject participant, which means they are the agents who perform specific actions / behaviors. Similarly, the object participant is the object affected by specific actions / behaviors. Although the details of specific events (such as buyer / giver) are removed, the composition and semantic information of event parameters are well preserved, and the general attributes between different events are explicitly retained. Through contrastive learning, the network model further learns the variable attributes of event parameters and obtains a robust understanding in different contexts and different domains.
[0248] 1.2 Decoupling Event Knowledge and Event Samples
[0249] 1.2.1 Learning Independent Event Concept Knowledge
[0250] For each event type, this application constructs a descriptive text to describe the definition of the event and each event parameter. The prompt information structure of this application includes two parts: the first part is the definition of the event, and the second part is the description of the event trigger word and each parameter in the event, and a unified parameter naming pattern is used to represent. The event definition and element description are from the original dataset, and this application also rewrites multiple synonymous expressions to cover the diversity of expressions.
[0251] Regarding the implementation method of learning independent event concept knowledge, reference can be made to the process of steps E1 - E2 in the above embodiments. It will not be elaborated here.
[0252] In one example, in the prompt information of only the event concept, the prompt scheme includes the event definition and the concepts and indications regarding the event trigger word and event elements, as shown below:
[0253] Event Definition (Event definition):
[0254] Whenever a person is born, there is a BEBORN event of LIFE category;
[0255] Event Argument Description:
[0256] The BE BORN event contains a subject participant that specifies the person who is born, a time argument indicating when the birth takes place, and a place argument denoting where the birth takes place;
[0257] Event Concept Context (including the combination of the above two):
[0258] Event Definition + Event Argument Description.
[0259] The event definition is stored in a database (DB). This application can train a retriever based on the SimCSE model. Given an event text, the retriever is trained to retrieve the event definition most relevant to the event text. In an event detection / argument extraction task, the k most relevant event definitions retrieved are used as event knowledge prompts.
[0260] 1.2.2 Event Extraction Learning
[0261] The event extraction task is reconstructed as a prompt-based generation task, including an event detection subtask and an event argument extraction subtask.
[0262] For an independent event extraction task, the input is the event instance description text and the description text of each of its arguments / triggers / event types. The goal of the detection task can be to randomly select target elements from the trigger word, event type recognition, and event arguments for detection, or to include all elements of the trigger word, event type recognition, and event arguments for detection.
[0263] Regarding the implementation method of event extraction learning, reference can be made to the process of steps F1 - F2 in the above embodiments, as well as the description text of each event and the description text of each argument. For details, see Figure 6 ; details are not described here again.
[0264] 1.3 Establish the association between event concept knowledge and event extraction tasks
[0265] Through the multi-task learning mode, the network model can learn event concept knowledge and event extraction task knowledge simultaneously. In the event extraction task, this application does not rely on black-box knowledge, but instead uses event concept knowledge. This application designs a method to associate text patterns with event concept learning tasks and event extraction learning tasks, so that there is an explicit association between the two.
[0266] As Figure 7 shown, EK, Precise EE, and Retrieved EE respectively represent three training tasks: event concept knowledge learning (event concept knowledge), event concept knowledge learning based on precise concepts, and event concept knowledge learning based on retrieved relevant event concepts. As Figure 7 shown, "the definition of an event" is the repeated text that appears in the event concept learning task, event detection task, and event argument extraction task. That is to say, in the above three training tasks, they respectively contain "a definition of an event", "the definition of the event", and "the definitions of the events", all of which contain the meaning of the definition of an event. In the event extraction task based on this, there are two types of texts that prompt event concept knowledge in the samples. One is to add precise concepts to the sample prompt text (corresponding to the training tasks of steps F1 - F2), and the other is to add the retrieved k most relevant event definitions to the sample prompt text (corresponding to the first training task of steps 502 - 504). Specifically, event arguments are usually formatted in templates, and the model needs to replace the entries in the template to achieve the argument extraction task. In this application, the event definition does not directly guide the model to extract event arguments based on detailed argument descriptions; instead, through the three training tasks and the text of "the definition of an event" they all contain, the model is guided to associate with each event concept learning task, and detailed argument knowledge is learned through association in multiple tasks.
[0267] 1.4 Design of learning objectives
[0268] As Figure 4 shown, this application uses a network model with an encoder-decoder as the backbone network. The training loss adopted in this application can use the cross-entropy loss based on the above formula (1) and the contrast loss calculated based on formula (2). Details are not elaborated here.
[0269] An embodiment of the present disclosure further provides an electronic device, which includes a processor. Optionally, it may further include a transceiver and / or a memory coupled to the processor, and the processor is configured to execute the steps of the method provided in any optional embodiment of the present disclosure.
[0270] Figure 8 FIG. shows a schematic structural diagram of an electronic device applicable to an embodiment of the present invention, as Figure 8 shown Figure 8 The electronic device 4000 shown in FIG. includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present disclosure. Optionally, the electronic device may be a first network node, a second network node or a third network node.
[0271] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure content of the present disclosure. The processor 4001 may also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0272] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 8 only a thick line is shown in FIG., but it does not mean that there is only one bus or one type of bus.
[0273] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.
[0274] The memory 4003 is used to store the computer program for implementing the embodiments of the present disclosure and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0275] The embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0276] The embodiments of the present disclosure further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0277] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and drawings of the present disclosure are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than that shown in the drawings or described in words.
[0278] It should be understood that although the flowcharts of the embodiments of the present disclosure indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present disclosure do not limit this.
[0279] The above text and drawings are provided only as examples to assist the reader in understanding the present disclosure. They are not intended and should not be construed as limiting the scope of the present disclosure in any way. Although certain embodiments and examples have been provided, it will be apparent to those skilled in the art based on the content disclosed herein that changes can be made to the illustrated embodiments and examples without departing from the scope of the present disclosure, and other similar implementation means based on the technical idea of the present disclosure can be adopted, which also fall within the protection scope of the embodiments of the present disclosure.< / object> < / subject> < / trigger> < / trigger> < / trigger>
Claims
1. A method performed by an electronic device, characterized in that, The method includes: Obtaining the text to be processed; Based on the text to be processed and the first prompt information, using the trained artificial intelligence (AI) network to extract the event extraction result of the text to be processed; Wherein, the first prompt information is used to prompt extraction based on at least one event definition information in the candidate event definition set.
2. The method according to claim 1, characterized in that, The first prompt information is obtained through the following steps: Retrieving at least one associated event definition information associated with the text to be processed from the candidate event definition set; Based on the at least one associated event definition information to obtain the first prompt information, and the at least one event definition information includes the at least one associated event definition information.
3. The method according to claim 2, wherein The obtaining the first prompt information based on the at least one associated event definition information includes any one of the following: Concatenating the at least one associated event definition information and the second prompt information to obtain the first prompt information, where the second prompt information refers to extraction according to the specific event definition information, and the specific event definition information refers to the at least one associated event definition information being concatenated; Concatenating the at least one event type information corresponding to the at least one associated event definition information and the second prompt information to obtain the first prompt information, where the specific event definition information refers to the at least one associated event definition information corresponding to the at least one event type information being concatenated.
4. The method according to any one of claims 1-3, characterized in that, The event extraction result is obtained by the AI network through performing the following operations: Based on the text to be processed and the first prompt information, extracting the description information of the event composition units of the text to be processed; Based on the description information of the event composition units, obtaining the event extraction result of the text to be processed.
5. The method according to claim 4, wherein The extracting the description information of the event composition units of the text to be processed based on the text to be processed and the first prompt information includes: Based on the text to be processed and the first prompt information, extracting the event type information of at least one target event associated with the text to be processed, the description information of at least one general event element corresponding to each target event, and the description information of the event trigger element; Wherein, the at least one general event element is an element in the general element set, and the general element set includes the general event elements corresponding to each candidate event in the candidate event set corresponding to the candidate event definition set; the event composition units of the text to be processed include each event composition unit corresponding to each target event associated with the text to be processed, and each event composition unit includes an event type, each general event element, and an event trigger element.
6. The method according to claim 5, wherein The at least one target event includes any one of the following: If the at least one event definition information includes at least one associated event definition information associated with the text to be processed, the at least one target event is the event associated with the text to be processed among the at least one event corresponding to the at least one associated event definition information; If the at least one event definition information includes each event definition information in the candidate event definition set, the at least one target event is the event associated with the text to be processed in the candidate event set.
7. The method according to claim 5, characterized in that, Obtaining an event extraction result corresponding to each target event based on the event type information of each extracted target event, as well as the description information of each general event element associated with each target event and the description information of the event trigger element, includes: Based on the event type information of each of the target events, converting the description information of each general event element corresponding to each of the target events into the description information of each dedicated event element in the dedicated event element set corresponding to each of the target events; Based on the description information of each dedicated event element corresponding to each of the target events and the description information of the event trigger element, obtaining an event extraction result corresponding to each of the target events.
8. The method according to claim 7, wherein The description information of a general event element includes at least the general event element; The converting the description information of each general event element corresponding to each of the target events into the description information of each dedicated event element corresponding to each of the target events based on the event type information of each of the target events includes: Based on the event type information of each of the target events, converting the general event elements in the description information of each general event element corresponding to each of the target events to obtain the description information of each dedicated event element corresponding to each of the target events.
9. The method according to any one of claims 5-8, characterized in that, The description information of a general event element includes the general event element and the element value of the general event element; Obtaining an event extraction result corresponding to each target event based on the event type information of each extracted target event, as well as the description information of each general event element associated with each target event and the description information of the event trigger element, includes: Based on the event type information of each of the target events, determining the target dedicated element set of each of the target events, where the target dedicated element set includes the dedicated event elements corresponding to the general event elements associated with the corresponding target event; Using each dedicated event element in the target dedicated element set of each of the target events to replace the general event elements in the description information of the general event elements of the corresponding target event, to obtain the description information of the dedicated event elements corresponding to each of the target events; Based on the description information of each dedicated event element corresponding to each of the target events and the description information of the event trigger element, obtaining an event extraction result corresponding to each of the target events.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 9.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 9.