Document-level event extraction method and device based on heterogeneous graph interactive learning and public context sensing fusion

By constructing heterogeneous graphs and performing information interaction learning, we can capture the long-distance dependence and complex interaction of entities in the document, and solve the problem of low document-level event extraction accuracy, achieving higher event detection accuracy and multi-event correlation analysis capabilities.

CN120373441APending Publication Date: 2025-07-25HENAN UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510382711.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing document-level event extraction methods are difficult to fully capture long-distance dependencies and complex interactions between entities, resulting in low event extraction accuracy, which is easy to be confused when multiple events coexist.

Method used

Using a method based on heterogeneous graph interactive learning and public context-aware fusion, a heterogeneous graph containing sentences and entity mentions is constructed, and an initial representation vector is generated using a BiLSTM encoder, and information interaction is performed through a heterogeneous graph attention network, combining global representation vectors to predict event types and entity combinations to capture cross-sentence dependencies.

Benefits of technology

It improves the accuracy of document-level event extraction, reduces information confusion between different events, enhances the ability to handle multiple event coexistence situations, and improves the accuracy and comprehensiveness of event detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373441A_ABST
    Figure CN120373441A_ABST
Patent Text Reader

Abstract

The invention provides a document level event extraction method and device based on heterogeneous graph interactive learning and public context sensing fusion. The method comprises the following steps: acquiring a to-be-detected document; inputting a to-be-detected document into the trained event extraction model to generate an extraction result; comprising the steps that sentences in an input document and entity mentions extracted from the sentences are coded, and initial representation vectors of the sentences and the entity mentions are generated; constructing a heterogeneous graph of the input document, performing information interaction on nodes in the heterogeneous graph according to the initial representation vectors of the nodes, and generating global representation vectors mentioned by sentences and entities; performing binary classification on each predefined event type according to the global representation vector of the sentence and predicting the occurrence probability of each predefined event type; extracting a public context between any two entities according to the global representation vectors mentioned by all the entities so as to predict an adjacent matrix between any two entities, and extracting an entity combination; and combining and pairing the predicted event type and the extracted entity to generate an event record.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a document-level event extraction method and device based on the fusion of heterogeneous graph interaction learning and common context awareness. Background Art

[0002] In the era of Internet big data, with the continuous expansion of online text data, the massive information generated by various social media, news websites, and forums has grown exponentially, which greatly increases the difficulty of manual information screening. Therefore, how to quickly and efficiently extract event information automatically from massive data and ensure the accuracy and comprehensiveness of the extraction results has become a key problem that needs to be solved urgently. This not only affects the progress of intelligent information processing technology but also directly relates to the effects and quality of actual application scenarios such as emergency response, public opinion analysis, and news push.

[0003] Document-level event extraction, as an important task in the field of information extraction, aims to automatically identify and extract events and their related structured information from long texts, including event types, participating entities, and their roles, etc. Different from traditional sentence-level event extraction, document-level event extraction needs to capture and integrate information in a broader context to handle cross-sentence complex dependencies and associations. Currently, most document-level event extraction methods are mostly based on traditional sequence models or ordinary graph neural networks. These methods usually have difficulty in comprehensively capturing the long-distance dependencies and complex interactions between entities in the document, and are not sufficient in dealing with the coexistence of multiple events, which easily leads to the confusion of different event information and reduces the accuracy of document-level event extraction. Summary of the Invention

[0004] To solve the problem of low accuracy of the above-mentioned document-level event extraction, the present invention provides a document-level event extraction method and device based on the fusion of heterogeneous graph interaction learning and common context awareness.

[0005] In a first aspect, the present invention provides a document-level event extraction method based on the fusion of heterogeneous graph interaction learning and common context awareness, including:

[0006] Step 1: Obtain the document to be detected;

[0007] Step 2: Input the document to be detected into the trained event extraction model to generate an extraction result; wherein, the extraction process of the event extraction model specifically includes:

[0008] Encode the sentences in the input document and the entity mentions extracted from the sentences to generate initial representation vectors for the sentences and entity mentions; for the input document, construct a heterogeneous graph containing sentence nodes and entity mention nodes, and perform information interaction on the nodes in the heterogeneous graph according to the initial representation vectors of the sentences and entity mentions to generate global representation vectors for the sentences and entity mentions; perform binary classification on each predefined event type according to the global representation vector of the sentence and predict the probability of its occurrence; extract the common context between any two entities according to the global representation vectors of all entity mentions to predict the adjacency matrix between any two entities, and extract entity combinations according to the adjacency matrix; wherein, the adjacency matrix is used to characterize whether there is a correlation between two entities; pair the predicted event types and the extracted entity combinations to generate the final event record.

[0009] Further, the encoding of the sentences in the input document and the entity mentions extracted from the sentences to generate initial representation vectors for the sentences and entity mentions specifically includes:

[0010] Segment the input document into sentences, input each sentence into the first encoder BiLSTM to generate representation vectors for each character in the sentence and the first representation vector of the sentence;

[0011] Based on the representation vectors of each character, use conditional random fields and the BIO annotation method to perform named entity recognition on each sentence, extract the entity mentions in the sentence, and perform max pooling on the representation vectors of the characters included in the entity mention to obtain the initial representation vector of the entity mention;

[0012] Input the first representation vectors of all sentences into the second encoder BiLSTM for redefinition to generate the initial representation vector of the sentence.

[0013] Further, constructing a heterogeneous graph containing sentence nodes and entity mention nodes for the input document specifically includes: connecting the entity mention nodes to their corresponding sentence nodes; connecting different entity mention nodes in the same sentence; connecting all entity mention nodes pointing to the same entity; connecting entity mention nodes of the same type; connecting all sentence nodes; adding self-loops to each individual node.

[0014] Further, using a heterogeneous graph attention network to perform information interaction on the nodes in the heterogeneous graph according to the initial representation vectors of the sentences and entity mentions to generate global representation vectors for the sentences and entity mentions specifically includes:

[0015] For a node p in the heterogeneous graph, calculate its representation vector at the l+1 layer in the heterogeneous graph attention network using the following formula

[0016]

[0017] Among them, represents a set of edges composed of different types of edges, represents the set of neighbors connected to node p through the k-th type of edge, W k (l) are trainable parameters, σ represents the ReLU activation function, is the representation vector of node q at the l-th layer, represents the attention weight of node q for node p in the (l + 1)-th layer, and its calculation method is:

[0018]

[0019] Among them, α is a single-layer feedforward neural network, [;] is the concatenation operation, W c is a learnable parameter, LeakyRelu() is the activation function, exp() is the exponential function with the natural constant as the base;

[0020] Average-pool the representation vectors of node p in each layer of the heterogeneous graph attention network, and derive the global representation vector h of node p through a linear layer and an activation function p :

[0021]

[0022] Among them, N g represents the number of layers of the heterogeneous graph attention network, W p is a trainable parameter, MeanPooling represents the average-pooling operation, is the initial representation vector of node p.

[0023] Furthermore, the binary classification of each predefined event type based on the global representation vector of the sentence and the prediction of its occurrence probability specifically includes:

[0024] Randomly initialize a trainable query vector for each event type, and use the attention mechanism to perform binary classification on each event type and predict its occurrence probability according to the following formula:

[0025]

[0026] P(t|D) = softmax(f D W t )

[0027] Among them, Q t is the trainable query vector of the t-th event type, softmax is the activation function, and S is the sentence feature matrix composed of the global representation vectors of all sentences in the input document, denotes the transpose operation, W t is a trainable parameter, d k denotes the vector dimension, and D denotes the document to be detected.

[0028] Furthermore, extracting the common context between any two entities based on the global representation vectors of all entity mentions to predict the adjacent matrix between any two entities specifically includes:

[0029] Using max pooling to aggregate the global representation vectors of all entity mentions pointing to the same entity to generate the representation vector of this entity;

[0030] Feeding the representation vectors of all entities into a BiLSTM to derive a set of entity representations where N e denotes the number of entities;

[0031] Counting the sentences in which each entity appears respectively to construct the sentence set corresponding to this entity;

[0032] Taking the intersection of the sentence sets corresponding to any two entities, and performing max pooling on the global representation vectors of all sentences in the intersection to obtain the common context vector between the two entities;

[0033] Predicting entity correlation according to the representation vectors and the common context vector of the two entities according to the following formula:

[0034]

[0035]

[0036] where, tanh and sigmoid are activation functions, W i 、W j 、 are all trainable parameters, v ij denotes entity e i and entity e j the common context vector between them, denotes the transpose operation, A i,j = 1 indicates that there is a correlation from entity e i to entity e j and vice versa, there is no; γ is a preset threshold.

[0037] Furthermore, each of the entity combinations consists of a core entity and one or more ordinary entities, and there is a correlation between the core entity and any one of the ordinary entities.

[0038] In a second aspect, the present invention provides a document-level event extraction device based on heterogeneous graph interaction learning and common context awareness fusion, including:

[0039] A document acquisition unit for acquiring a document to be detected;

[0040] An event extraction unit for inputting the document to be detected into a trained event extraction model to generate an extraction result; wherein, the event extraction model includes:

[0041] A document representation and entity extraction module for encoding sentences in the input document and entity mentions extracted from the sentences to generate initial representation vectors of the sentences and entity mentions;

[0042] A heterogeneous graph interaction module for constructing a heterogeneous graph containing sentence nodes and entity mention nodes for the input document, and performing information interaction on the nodes in the heterogeneous graph according to the initial representation vectors of the sentences and entity mentions to generate global representation vectors of the sentences and entity mentions;

[0043] An event type detection module for performing binary classification on each predefined event type based on the global representation vector of the sentence and predicting the probability of its occurrence;

[0044] A parameter combination extraction module for extracting the common context between any two entities based on the global representation vectors of all entity mentions to predict the adjacent matrix between any two entities, and extracting entity combinations according to the adjacent matrix; wherein, the adjacent matrix is used to characterize whether there is a correlation between two entities;

[0045] An event record generation module for pairing the predicted event type with the extracted entity combination to generate a final event record.

[0046] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in the first aspect is implemented.

[0047] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0048] The beneficial effects of the present invention are:

[0049] (1) The document-level event extraction method and device provided by the present invention construct a heterogeneous graph with sentences and entity mentions in the target document as nodes, and interact with the heterogeneous graph, which can comprehensively capture the long-distance dependencies and complex interactions between entities in the target document. At the same time, the common context between two entities is extracted based on the global representation vectors of all entity mentions, and whether there is a correlation between the two entities is judged based on this common context, and then entity combinations are extracted, which can better handle the situation of coexisting multiple events, as much as possible avoid information confusion between different events, and finally improve the accuracy of document-level event extraction.

[0050] (2) The present invention uses BiLSTM as an encoder to encode sentences and entity mentions in the document, which can consider both forward and backward information in the sentence, so as to more comprehensively capture the context dependencies and better capture the long-distance dependencies. Moreover, by obtaining the context information before and after each character in the sentence through BiLSTM, the model can more accurately understand the semantics of the sentence, thereby improving the performance of the model in the event extraction task.

[0051] (3) The present invention takes sentences and entity mentions as nodes in the heterogeneous graph, which can jointly model the interaction between sentences and entities from a global perspective, capture long-distance dependencies, enrich semantic information, and improve the performance of event extraction and relationship extraction. Among them, the relationship between any two sentences in the document can be modeled through sentence-sentence edges, better understanding cross-sentence events; the context information of entity mentions within a sentence can be better understood through sentence-entity mention edges; the relationship between entity pairs can be more accurately identified through three different types of entity mention-entity mention edges. In particular, by connecting entity mention nodes of the same type to each other, global dependencies can be more effectively captured; and by adding self-loops, each node can include its own features when aggregating information, thereby enhancing the feature expression ability of the node.

[0052] (4) In document-level event detection, the semantics of an entity often depends on its context information. By considering the common context of two entities, the present invention can better understand the relationship between entities, thereby reducing semantic ambiguity. Moreover, the entities of an event may be scattered in multiple sentences. By using the common context, cross-sentence entities can be effectively captured, thereby more accurately identifying events. In addition, there may be multiple related events in the document, and there may be complex interrelationships between these events (for example, two events may share some entity mentions or entities). By analyzing the common context of two entities, the association between these events can be better identified, thereby improving the accuracy of multi-event detection. In short, the present invention uses the common context of two entities for document event detection, which can enhance context understanding, capture cross-sentence entities, and improve the accuracy of multi-event detection. Description of the Drawings

[0053] Figure 1 Flowchart of the document-level event extraction method based on heterogeneous graph interaction learning and common context awareness fusion provided for the implementation of the present invention;

[0054] Figure 2 Event extraction flowchart of the event extraction model provided for the embodiments of the present invention;

[0055] Figure 3 Structure diagram of the document-level event extraction device based on heterogeneous graph interaction learning and common context awareness fusion provided for the embodiments of the present invention;

[0056] Figure 4 Structure block diagram of an electronic device provided for the embodiments of the present invention. Detailed implementation manners

[0057] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] Combined with Figure 1 and Figure 2 As shown, the embodiments of the present invention provide a document-level event extraction method based on heterogeneous graph interaction learning and common context awareness fusion, including the following steps:

[0059] S101: Obtain the document to be detected;

[0060] S102: Input the document to be detected into the trained event extraction model to generate an extraction result.

[0061] Specifically, the extraction process of the event extraction model specifically includes the following 5 stages:

[0062] Document representation and entity extraction stage: Encode the sentences in the input document and the entity mentions extracted from the sentences to generate initial representation vectors of the sentences and entity mentions;

[0063] Heterogeneous graph interaction stage: For the input document, construct a heterogeneous graph containing sentence nodes and entity mention nodes, and perform information interaction on the nodes in the heterogeneous graph according to the initial representation vectors of the sentences and entity mentions to generate global representation vectors of the sentences and entity mentions;

[0064] Event type detection stage: Perform binary classification on each predefined event type according to the global representation vector of the sentence and predict the probability of its occurrence;

[0065] Parameter combination extraction stage: Extract the common context between any two entities according to the global representation vectors of all entity mentions to predict the adjacent matrix between any two entities, and extract entity combinations according to the adjacent matrix; wherein, the adjacent matrix is used to characterize whether there is a correlation between two entities.

[0066] Event record generation stage: Pair the predicted event types and the extracted entity combinations to generate the final event records.

[0067] The document-level event extraction method provided by the embodiments of the present invention constructs a heterogeneous graph with sentences and entity mentions in the target document as nodes, and performs interactions on the heterogeneous graph, which can comprehensively capture the long-distance dependencies and complex interactions between entities in the target document; at the same time, extract the common context between two entities according to the global representation vectors of all entity mentions, and judge whether there is a correlation between the two entities based on this common context, and then extract entity combinations, which can better handle the situation of co-existing multiple events, avoid information confusion between different events as much as possible, and finally improve the accuracy of document-level event extraction.

[0068] In one embodiment, the following process is adopted to train the event extraction model:

[0069] S201: Construct a training set, a validation set, and a test set;

[0070] Specifically, obtain the ChFinAnn dataset. The ChFinAnn dataset contains 35 parameter roles, 5 event types, and 32040 documents, which are divided into a training set / validation set / test set according to the ratio of 25632 / 3204 / 3204. Preprocess the dataset. To ensure the consistency of the text data, set the maximum sentence length MaxLength, and fill in shorter sentences and truncate longer sentences according to MaxLength. In this embodiment, the maximum sentence length MaxLength is set to 128.

[0071] S202: Iteratively train the event extraction model using the training set until the number of training times reaches the preset number of training times to obtain a trained event extraction model;

[0072] Specifically, input the training set into the model in batches for iterative training, and set the number of iterations, learning rate, dropout rate, batch size, and loss function. Every time the preset number of times is trained, use the validation set to verify and evaluate the model during the training process, and save the model parameters during the training process. After the training is completed, select the model with the best performance from all the saved models as the trained event extraction model.

[0073] In this embodiment, the number of iterations is set to 100, the learning rate is set to 0.0005, the batch size is set to 64, the optimizer is set to Adam, the dropout rate is set to 0.1, and the threshold γ is set to 0.5.

[0074] In order to comprehensively capture entity and context information and generate a complete event record, in this embodiment, the loss function is composed of four loss components, and different weights are assigned to calculate the total loss:

[0075]

[0076] Among them, respectively represent the losses of entity extraction, event type detection, parameter combination extraction, and role prediction in event record generation; β1, β2, β3, and β4 are all hyperparameters representing the weights of each loss component. In this embodiment, they are respectively set to 0.05, 0.95, 0.95, and 0.95; is the total loss.

[0077] As an example, the negative log-likelihood function is used as the loss function for entity extraction and the loss function for event type detection The binary cross-entropy loss function is used as the loss function for parameter combination extraction and the loss function for role prediction

[0078] S203: Use the test set to perform performance testing on the trained event extraction model, obtain the performance test results, and optimize the event extraction model;

[0079] Specifically, input the test set into the trained event extraction model for testing. According to the test results, use the three evaluation metrics of precision (P), recall (R), and F1 score for performance evaluation. The calculation methods are as follows:

[0080]

[0081] F1 = 2 * P * R / (P + R)

[0082] Among them, TP represents the number of positive samples predicted as positive samples, FN represents the number of positive samples predicted as negative samples, and FP represents the number of negative samples predicted as positive samples.

[0083] Through repeated experiments and adjustments, according to the performance of the performance metrics, select the model with the best performance as the optimized event extraction model.

[0084] S204: Use the optimized event extraction model to perform event extraction on the document to be detected and generate extraction results.

[0085] In one embodiment, as Figure 2 shown, for the document representation and entity extraction stage, the sentences in the input document and the entity mentions extracted from the sentences are encoded to generate initial representation vectors for the sentences and entity mentions, specifically including:

[0086] S301: Segment the input document into sentences, and input each sentence into the first encoder BiLSTM to generate the representation vector for each character in the sentence and the first representation vector of the sentence;

[0087] Specifically, first perform sentence segmentation on the document, and input each sentence into the encoder BiLSTM respectively to obtain the forward hidden vector sequence and the backward hidden vector sequence. Then concatenate the forward and backward hidden vector sequences to obtain the representation vector for each character, and concatenate the last hidden states in both directions as the representation vector of the sentence. For the convenience of distinguishing from the subsequent representation vectors, the sentence representation vector obtained at this time is denoted as the first representation vector of the sentence.

[0088] S302: Based on the representation vector of each character, use the conditional random field and BIO annotation method to perform named entity recognition on each sentence, extract the entity mentions in the sentence, and perform max pooling on the representation vectors of the characters included in the entity mention to obtain the initial representation vector of the entity mention; taking the j-th entity mention in the document as an example, perform max pooling on the representation vectors of all characters inside it to obtain the representation vector of the entity mention

[0089] Specifically, the BIO annotation method is a sequence annotation method in natural language processing (NLP), mainly used for tasks such as named entity recognition (NER). It identifies the entity category to which each word or symbol in the text belongs by assigning a label to it. B (Begin): indicates that the current word is the start of a certain entity. I (Inside): indicates that the current word is inside a certain entity, that is, the word is a component of a certain entity. O (Outside): indicates that the current word does not belong to any entity, that is, the word is not a component of any entity.

[0090] S303: Input the first representation vectors of all sentences into the second encoder BiLSTM for redefinition to generate the initial representation vector of the sentence. This process can be represented by the following formula:

[0091]

[0092] where, is the first representation vector of the i-th sentence, is the initial representation vector of the i-th sentence after redefinition, N sis the number of sentences in the input document. In step S301, the first encoder BiLSTM only encodes a single sentence; in step S303, by inputting the first representation vectors of multiple sentences into the second encoder BiLSTM for re - encoding, the information between sentences can be learned from each other, making the semantic information of the sentence representation vectors more sufficient.

[0093] In the embodiments of the present invention, BiLSTM is used as the encoder, which can consider both the forward and backward information in the sentence, so as to capture the context - dependence relationship more comprehensively and better capture the long - distance dependence relationship. Moreover, by simultaneously obtaining the context information of each character in the sentence through BiLSTM, the model can more accurately understand the semantics of the sentence, thereby improving the performance of the model in the event extraction task.

[0094] In one embodiment, in the heterogeneous graph interaction stage, for the input document, a heterogeneous graph including sentence nodes and entity mention nodes is constructed, specifically including: connecting the entity mention nodes with their corresponding sentence nodes; connecting different entity mention nodes in the same sentence; connecting all entity mention nodes pointing to the same entity; connecting entity mention nodes of the same type; connecting all sentence nodes; and adding self - loops to each individual node.

[0095] In the embodiments of the present invention, taking sentences and entity mentions as nodes in the heterogeneous graph can jointly model the interaction between sentences and entities from a global perspective, capture long - distance dependence relationships, enrich semantic information, and improve the performance of event extraction and relation extraction. Among them, the relationship between any two sentences in the document can be modeled through sentence - sentence edges, better understanding cross - sentence events; the context information of entity mentions within sentences can be better understood through sentence - entity mention edges; and the relationship between entity pairs can be more accurately identified through three different types of entity mention - entity mention edges. In particular, by connecting entity mention nodes of the same type, global dependence relationships can be more effectively captured; and by adding self - loops, each node can include its own features when aggregating information, thereby enhancing the feature expression ability of the node.

[0096] In one embodiment, the heterogeneous graph attention network is used to perform information interaction on the nodes in the heterogeneous graph according to the initial representation vectors of sentences and entity mentions, generating global representation vectors of sentences and entity mentions, specifically including:

[0097] For node p in the heterogeneous graph, the following formula is used to calculate its representation vector at the (l + 1)-th layer in the heterogeneous graph attention network

[0098]

[0099] Among them, represents a set of edges composed of different types of edges, represents the set of neighbors connected to node p through the k-th type of edge, W k (l) is a trainable parameter, σ represents the ReLU activation function, is the representation vector of node q at the l-th layer, represents the attention weight of node q for node p in the (l + 1)-th layer, and its calculation method is:

[0100]

[0101] Among them, α is a single-layer feedforward neural network, [;] is the concatenation operation, W c is a learnable parameter, LeakyRelu() is the activation function, exp() is the exponential function with the natural constant as the base;

[0102] Through the iterative update of the above heterogeneous graph attention network, the representation vector of node p at each layer of the heterogeneous graph attention network can be obtained. Then, the representation vectors of node p at each layer of the heterogeneous graph attention network are subjected to average pooling, and a global representation vector h of node p is derived through a linear layer and an activation function p :

[0103]

[0104] Among them, N g represents the number of layers of the heterogeneous graph attention network, W p is a trainable parameter, MeanPooling represents the average pooling operation, is the initial representation vector of node p. If node p is a sentence node corresponding to the i-th sentence, its initial representation vector is If node p is a mention node corresponding to the j-th entity mention, its initial representation vector is

[0105] In one embodiment, in the event type detection stage, binary classification is performed on each predefined event type according to the global representation vector of the sentence and the probability of its occurrence is predicted, specifically including:

[0106] A trainable query vector is randomly initialized for each event type, and binary classification is performed on each event type by using the attention mechanism according to the following formula and the probability of its occurrence is predicted:

[0107]

[0108] P(t|D) = softmax(f D W t )

[0109] Among them, Q t is the trainable query vector for the t-th event type, softmax is the activation function, S is the sentence feature matrix composed of the global representation vectors of all sentences in the input document, represents the transpose operation, W t is the trainable parameter, d k represents the vector dimension, and D represents the document to be detected.

[0110] In one embodiment, in the parameter combination extraction stage, the common context between any two entities is extracted according to the global representation vectors of all entity mentions to predict the adjacent matrix between any two entities, specifically including:

[0111] S401: Use max pooling to aggregate the global representation vectors of all entity mentions pointing to the same entity to generate the representation vector of this entity;

[0112] S402: Feed the representation vectors of all entities into the BiLSTM to derive the entity representation set where N e represents the number of entities;

[0113] S403: Count the sentences in which each entity appears respectively to construct the sentence set corresponding to this entity;

[0114] S404: Take the intersection of the sentence sets corresponding to any two entities, and perform max pooling on the global representation vectors of all sentences in the intersection to obtain the common context vector between the two entities;

[0115] S405: Predict entity correlation according to the representation vectors of the two entities and the common context vector according to the following formula:

[0116]

[0117] where, tanh and sigmoid are activation functions, W i , W j , are all trainable parameters, v ij represents the common context vector between entity e i and entity e j , represents the transpose operation, A i,j =1 indicates that there is a correlation from entity e i to entity e j , otherwise, there is no such correlation; γ is a preset threshold (belonging to hyperparameters), when the predicted value is not less than the threshold γ, it represents that there is a correlation from e i to e j .

[0118] Further, entity combinations are extracted according to the adjacent matrix, specifically including: each entity combination consists of a core entity and one or more ordinary entities, and there is a correlation between the core entity and any ordinary entity.

[0119] Specifically, in document-level event detection, the semantics of an entity often depends on its context information. In the embodiments of the present invention, by considering the common context of two entities, the relationship between entities can be better understood, thereby reducing semantic ambiguity. Moreover, the entities of an event may be scattered in multiple sentences. By using the common context, entities across sentences can be effectively captured, thereby more accurately identifying events. In addition, there may be multiple related events in a document, and there may be complex interrelationships between these events (for example, two events may share some entity mentions or entities). By analyzing the common context of two entities, the association between these events can be better identified, thereby improving the accuracy of multi-event detection. In summary, the present embodiment uses the common context of two entities for document event detection, which can enhance context understanding, capture entities across sentences, and improve the accuracy of multi-event detection.

[0120] In one embodiment, in the event record generation stage, the predicted event type and the extracted entity combination are paired to generate the final event record, specifically including:

[0121] Perform a Cartesian product operation on the detected event type results and the extracted entity combinations to generate all type-combination pairs. For each pair, use a feedforward neural network of a specific event type and a sigmoid activation function to calculate the probability of each entity corresponding to all roles under the current event type. Select the entity with the highest probability under each role to fill the event table, where if the probability of all entities under a role is lower than 0.7, it is considered that the role is empty.

[0122] In event detection, some events may involve multiple entities or multiple types. In the present embodiment, by performing a Cartesian product operation on the detected event type results and the extracted entity combinations, all possible event instances can be generated, thereby more comprehensively covering complex scenarios; and by combining different event types with related entity combinations, the relationship between events can be better analyzed, improving the comprehensiveness and accuracy of event detection, supporting multi-event association analysis, and enhancing the generalization ability of the model.

[0123] Based on the same inventive concept, as Figure 3 shown, the embodiments of the present invention further provide a document-level event extraction device based on heterogeneous graph interaction learning and common context awareness fusion, including: a document acquisition unit and an event extraction unit.

[0124] Among them, the document acquisition unit is used to acquire the document to be detected; the event extraction unit is used to input the document to be detected into the trained event extraction model to generate an extraction result; wherein the event extraction model includes a document representation and entity extraction module, a heterogeneous graph interaction module, an event type detection module, a parameter combination extraction module and an event record generation module.

[0125] The document representation and entity extraction module is used to encode the sentences in the input document and the entity mentions extracted from the sentences, and generate the initial representation vectors of the sentences and entity mentions. The heterogeneous graph interaction module is used to construct a heterogeneous graph containing sentence nodes and entity mention nodes for the input document, and to perform information interaction on the nodes in the heterogeneous graph according to the initial representation vectors of the sentences and entity mentions, and to generate the global representation vectors of the sentences and entity mentions. The event type detection module is used to perform binary classification on each predefined event type according to the global representation vector of the sentence and predict its probability of occurrence. The parameter combination extraction module is used to extract the common context between any two entities according to the global representation vectors of all entity mentions to predict the adjacent matrix between any two entities, and extract entity combinations according to the adjacent matrix; wherein the adjacent matrix is used to characterize whether there is a correlation between the two entities. The event record generation module is used to pair the predicted event type with the extracted entity combination to generate the final event record.

[0126] The document-level event extraction device provided by the embodiment of the present invention can comprehensively capture the long-distance dependencies and complex interactions between entities in the target document by constructing a heterogeneous graph with sentences and entity mentions in the target document as nodes and interacting with the heterogeneous graph; at the same time, the common context between two entities is extracted according to the global representation vector of all entity mentions, and whether there is a correlation between the two entities is judged based on the common context, and then the entity combination is extracted, which can better handle the coexistence of multiple events, avoid information confusion between different events as much as possible, and ultimately improve the accuracy of document-level event extraction.

[0127] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4As shown in the figure, the electronic device may include: a processor 401, a communications interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communications interface 402, and the memory 403 complete communication with each other through the communication bus 404. The processor 401 may call the logical instructions in the memory 403 to execute a document-level event extraction method based on heterogeneous graph interaction learning and common context awareness fusion. The method includes: Step 1: Obtain the document to be detected; Step 2: Input the document to be detected into the trained event extraction model to generate an extraction result. Among them, the extraction process of the event extraction model specifically includes: encoding the sentences in the input document and the entity mentions extracted from the sentences to generate initial representation vectors of the sentences and entity mentions; for the input document, constructing a heterogeneous graph including sentence nodes and entity mention nodes, and performing information interaction on the nodes in the heterogeneous graph according to the initial representation vectors of the sentences and entity mentions to generate global representation vectors of the sentences and entity mentions; performing binary classification on each predefined event type according to the global representation vector of the sentence and predicting the probability of its occurrence; extracting the common context between any two entities according to the global representation vectors of all entity mentions to predict the adjacent matrix between any two entities, and extracting entity combinations according to the adjacent matrix; where the adjacent matrix is used to characterize whether there is a correlation between two entities; pairing the predicted event types and the extracted entity combinations to generate the final event record.

[0128] In addition, when the logical instructions in the above-mentioned memory 403 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0129] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the document-level event extraction method based on heterogeneous graph interaction learning and common context awareness fusion provided by each of the above method embodiments.

[0130] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the document-level event extraction method based on heterogeneous graph interaction learning and common context awareness fusion provided by each of the above method embodiments.

[0131] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in each of the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A document-level event extraction method based on the fusion of heterogeneous graph interaction learning and common context awareness, characterized in that Including: Step 1: Obtain the document to be detected; Step 2: Input the document to be detected into the trained event extraction model to generate an extraction result; wherein, the extraction process of the event extraction model specifically includes: Encode the sentences in the input document and the entity mentions extracted from the sentences to generate initial representation vectors of the sentences and entity mentions; for the input document, construct a heterogeneous graph containing sentence nodes and entity mention nodes, and perform information interaction on the nodes in the heterogeneous graph according to the initial representation vectors of the sentences and entity mentions to generate global representation vectors of the sentences and entity mentions; perform binary classification on each predefined event type according to the global representation vector of the sentence and predict the probability of its occurrence; extract the common context between any two entities according to the global representation vectors of all entity mentions to predict the adjacent matrix between any two entities, and extract entity combinations according to the adjacent matrix; wherein, the adjacent matrix is used to characterize whether there is a correlation between two entities; pair the predicted event types and the extracted entity combinations to generate the final event record.

2. The method for document-level event extraction based on heterogeneous graph interaction learning and public context awareness fusion according to claim 1, wherein The encoding of the sentences in the input document and the entity mentions extracted from the sentences to generate initial representation vectors of the sentences and entity mentions specifically includes: Segment the input document into sentences, input each sentence into the first encoder BiLSTM to generate the representation vector of each character in the sentence and the first representation vector of the sentence; Based on the representation vectors of each character, use a conditional random field and the BIO annotation method to perform named entity recognition on each sentence, extract the entity mentions in the sentence, and perform max pooling on the representation vectors of the characters included in the entity mention to obtain the initial representation vector of the entity mention; Input the first representation vectors of all sentences into the second encoder BiLSTM for redefinition to generate the initial representation vector of the sentence.

3. The method for document-level event extraction based on heterogeneous graph interaction learning and public context awareness fusion according to claim 1, wherein, The construction of a heterogeneous graph containing sentence nodes and entity mention nodes for the input document specifically includes: connecting the entity mention nodes with their corresponding sentence nodes; connecting different entity mention nodes in the same sentence; connecting all entity mention nodes pointing to the same entity; connecting entity mention nodes of the same type; connecting all sentence nodes; adding self-loops to each individual node.

4. The method for document-level event extraction based on the fusion of heterogeneous graph interaction learning and common context awareness according to claim 3, characterized in that, Use a heterogeneous graph attention network to perform information interaction on the nodes in the heterogeneous graph according to the initial representation vectors of the sentences and entity mentions to generate global representation vectors of the sentences and entity mentions, specifically including: For node p in the heterogeneous graph, the following formula is used to calculate its representation vector at the (l + 1)-th layer in the heterogeneous graph attention network Among them, represents a set of edges composed of different types of edges, represents a set of neighbors connected to node p through the k-th type of edge, are trainable parameters, and σ represents the ReLU activation function, is the representation vector of node q at the l-th layer, represents the attention weight of node q for node p in the (l + 1)-th layer, and its calculation method is: where α is a single-layer feedforward neural network, [;] is the concatenation operation, W c is a learnable parameter, LeakyRelu() is the activation function, and exp() is the exponential function with the natural constant as the base; Average pool the representation vectors of node p at each layer of the heterogeneous graph attention network, and derive the global representation vector h of node p through a linear layer and an activation function p : Among them, N g represents the number of layers of the heterogeneous graph attention network, and W p is a trainable parameter. MeanPooling represents the average pooling operation, is the initial representation vector of node p.

5. The method for document-level event extraction based on the fusion of heterogeneous graph interaction learning and common context awareness according to claim 1, characterized in that The binary classification of each predefined event type according to the global representation vector of the sentence and the prediction of its occurrence probability specifically includes: Randomly initialize a trainable query vector for each event type, and use the attention mechanism to perform binary classification on each event type and predict the probability of its occurrence according to the following formula: P(t|D) = softmax(f D W t ) Among them, Q t is the trainable query vector for the t-th event type, softmax is the activation function, S is the sentence feature matrix composed of the global representation vectors of all sentences in the input document, represents the transpose operation, W t is the trainable parameter, d k represents the vector dimension, and D represents the document to be detected.

6. The method for document-level event extraction based on the fusion of heterogeneous graph interaction learning and common context awareness according to claim 1, wherein The extraction of the common context between any two entities according to the global representation vectors of all entity mentions to predict the adjacent matrix between any two entities specifically includes: Use max pooling to aggregate the global representation vectors of all entity mentions pointing to the same entity to generate the representation vector of the entity; Feed the representation vectors of all entities into a BiLSTM to derive a set of entity representations where N e represents the number of entities; Count the sentences in which each entity appears separately to construct the sentence set corresponding to the entity; Take the intersection of the sentence sets corresponding to any two entities, and perform max pooling on the global representation vectors of all sentences in the intersection to obtain the common context vector between the two entities; Predict entity relevance according to the representation vectors of the two entities and the common context vector according to the following formula: Among them, tanh and sigmoid are activation functions, and W i , W j , are all trainable parameters. v ij represents the common context vector between entity e i and entity e j . represents the transpose operation. A i,j = 1 indicates that there is a correlation from entity e i to entity e j , otherwise, there is no such correlation; γ is a preset threshold.

7. The method for document-level event extraction based on the fusion of heterogeneous graph interaction learning and common context awareness according to claim 6, wherein Each of the entity combinations consists of a core entity and one or more ordinary entities, and there is a relevance between the core entity and any ordinary entity.

8. A document-level event extraction device based on the fusion of heterogeneous graph interaction learning and common context awareness, characterized in that Including: A document acquisition unit for acquiring a document to be detected; An event extraction unit for inputting the document to be detected into a trained event extraction model to generate an extraction result; wherein, the event extraction model includes: A document representation and entity extraction module for encoding sentences in the input document and entity mentions extracted from the sentences to generate initial representation vectors of the sentences and entity mentions; A heterogeneous graph interaction module for constructing a heterogeneous graph containing sentence nodes and entity mention nodes for the input document, and performing information interaction on the nodes in the heterogeneous graph according to the initial representation vectors of the sentences and entity mentions to generate global representation vectors of the sentences and entity mentions; An event type detection module for performing binary classification on each predefined event type according to the global representation vector of the sentence and predicting the probability of its occurrence; A parameter combination extraction module for extracting the common context between any two entities according to the global representation vectors of all entity mentions to predict the adjacency matrix between any two entities, and extracting entity combinations according to the adjacency matrix; wherein, the adjacency matrix is used to characterize whether there is a relevance between two entities; An event record generation module for pairing the predicted event type and the extracted entity combination to generate a final event record.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Document-level event extraction model training method, event extraction method and system

    CN121542744A

  • Document-level event extraction model training method, event extraction method and system

    CN121542744B