Event Causality Recognition Method Based on Hypergraph Modeling of Document-Level Causal Structure
Through a combination of hypergraph modeling and prompt learning, the interdependence of event causality and document-level context semantic information of event causality in documents is captured, and the problem of ignoring interdependence and inability to fully utilize document-level context semantics is solved in the prior art, and a more efficient document-level event causality recognition is achieved.
Patent Information
- Application Number
- CN202310595004.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-05-24
AI Technical Summary
The existing causal relationship recognition methods ignore the interdependence of event causal relationships in documents and cannot fully utilize document-level context semantic information, resulting in a low causal relationship recognition rate across sentences.
The document-level causal structure recognition method based on hypergraph modeling is adopted, and through the steps of text preprocessing, paired event semantic learning, document-level causal structure learning and causal relationship recognition, combined with prompt learning and hypergraph convolutional neural network, the interdependence relationship and document-level context semantic information of events in the document are captured.
The accuracy of document-level event causal relationship recognition is improved, and the ability to understand and recognize the interactions of multiple events is enhanced by combining paired event semantic information and document-level causal structure information.
Smart Images

Figure CN116719900B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to a method for identifying document-level event causal relationships based on prompt learning and neural networks, and particularly relates to a method for identifying event causal relationships by modeling document-level causal structures using hypergraphs, which is used in the technical field of event relationship recognition. Background Art
[0002] Event Causal Identification (denoted as ECI) aims to detect whether there is a causal relationship between two events in a document. The ECI task is crucial for many Natural Language Processing (denoted as NLP) applications, such as question answering, information extraction, etc. For causal relationship identification, a variety of techniques have been developed. The latest methods can be roughly divided into knowledge-base-based methods, methods based on the prompt learning paradigm, and methods based on graph neural networks.
[0003] The idea of knowledge-base-based methods is to utilize external causal knowledge from an external knowledge base to enhance causal relationship identification. The idea of methods based on the prompt learning paradigm is to model the probability of text in a pre-trained language model, transform the identification task into a text prediction task, so as to achieve the purpose of causal relationship identification. Currently, good results have been achieved in many NLP tasks and it has been successfully applied to the ECI task. The idea of methods based on graph neural networks is usually to model the ECI task as a node classification problem, apply graph neural networks to learn event node representation vectors from document-level context semantics, and then use common classifiers in machine learning for classification. In addition to node classification, some studies have also investigated potential causal edges in the event graph for causal relationship identification.
[0004] Knowledge-base-based methods can effectively enhance causal relationship detection and have good results for sentence-level event causal relationship identification. However, the disadvantage is that they cannot fully utilize document-level context semantics, and the recognition rate for cross-sentence causal relationships is relatively low. Methods based on the prompt learning paradigm can deeply mine the context semantic information of text, learn feature vectors, and have a relatively high recognition rate for sentence-level event causal relationships. However, they have poor interpretability and are limited by the input length limit of the pre-trained language model and cannot fully learn document-level context semantic information, resulting in a relatively low recognition rate for cross-sentence causal relationships. Methods based on graph neural networks can fully utilize document-level context semantic information and have a relatively high recognition rate for cross-sentence causal relationships. However, they do not make full use of the semantic information within sentences. Moreover, existing research has also
[0005] ignored the fact of the mutual dependence of event causal relationships in the document. Summary of the Invention
[0006] In view of the defects or improvement requirements of the prior art, the present invention provides an event causality recognition method based on hypergraph modeling of document-level causal structures, aiming to solve the problem in the existing event causality recognition methods that the fact of mutual dependence of event causal relationships in the document is ignored, and at the same time, using pairwise event semantics to assist in the recognition of causal relationships between event pairs to further improve the accuracy of document-level event causality recognition.
[0007] To achieve the above object, the present invention provides an event causality recognition method based on hypergraph modeling of document-level causal structures, which is characterized by including a text preprocessing step, a pairwise event semantics learning step, a pre-trained pairwise event semantics learning module, a document-level causal structure learning step, a document-level causal relationship recognition step, and a training and testing network step; where:
[0008] (1) Text preprocessing step: Input the original document, and preprocess it in two ways: pairwise events and events, to obtain two types of data: the sentence pairs where the pairwise events are located and the sentences where the events are located.
[0009] (2) Pairwise event semantics learning step: Adopt a PLM based on prompt learning, and combine a custom template to model the sentence pairs where the pairwise events in step (1) are located, to obtain the context semantic representation of the pairwise events and the prediction result of the causal relationship between the pairwise events.
[0010] (3) Pre-training pairwise event semantics learning module step: Construct a cross-entropy loss function based on the predicted virtual answer words and the real answer words, and pre-train the pairwise event semantics learning module by minimizing the loss function, and select the model with the highest overall F1 in the validation set as the pairwise semantics learning model.
[0011] (4) Document-level causal structure learning step: Input the sentences where the events in step (1) are located into another PLM to obtain sentence-level event representations, and combine the prediction results of the causal relationships between the pairwise events in step (2) to construct a document-level causal hypergraph, and obtain document-level event representations through a hypergraph convolutional neural network.
[0012] (5) Document-level causal relationship recognition step: Concatenate the context semantic representations of the pairwise events in step (2) and the document-level event representations in step (4), and predict the probability of the existence of causal relationships for each event pair in the document through a multi-layer perceptron network.
[0013] (6) Training and testing network step: Based on the predicted causal probability distribution and the real causal label y to construct a loss function, and then train the network to minimize the loss function. After training, input the validation set and test set documents, and select the model with the highest F1 value on the validation set documents, so as to obtain the causal relationship prediction results of the corresponding test samples.
[0014] Furthermore, step (2) includes the following sub-steps:
[0015] (2-1) First, construct each event pair in the document x k =(Evt i ; Evt j ) into a prompt template T p (x) that can describe the potential causal relationship between the two event mentions:
[0016] T p (x k ) = In this sentence, Evt i [MASK]Evt j .
[0017] Among them, Evt i and Evt j are two event mentions. Insert the specific token [MASK] of the PLM between them for relationship prediction, and then connect the original sentence T s of the event mention with the constructed prompt template T p as the input sentence T of the PLM. Use the PLM-specific tokens [CLS] and [SEP] to represent the beginning and end of the input sentence T, and another [SEP] for the separation token between T s and T p ;
[0018] (2-2) Encode the input sentence T using the PLM and obtain the hidden vectors of the two event mentions and the specific token [MASK] from the output:
[0019]
[0020] Among them and are the hidden vectors of the two event mentions, is the hidden vector of [MASK], and d is the dimension of the hidden vector;
[0021] If the event mention consists of multiple words, use the average of their hidden vectors as the event mention representation, and combine the hidden vectors of the two event mentions to obtain the context semantic representation of the paired events:
[0022]
[0023] The context semantic representation of the paired events in a document:
[0024]
[0025] Among them, k is the number of pairs of all paired events in the document.
[0026] Furthermore, in step (2), by adding two virtual answer words, namely, Casual and None, to the PLM vocabulary, the PLM-based MLM classifier is based on To estimate the probability that [MASK] is two virtual answer words, the predicted virtual answer word with higher probability is used as the prediction result of the causal relationship of paired events.
[0027] Furthermore, in step (4), a hypergraph is used to model the causal structure of each document, wherein each node represents an event and a hyperedge is a causal relationship between multiple events that are mutually dependent, specifically including:
[0028] The sentences containing the events in step (1) are input into another PLM sentence by sentence, and the sentence-level event representation is obtained through PLM encoding. For event mentions consisting of multiple words, the average value of the hidden vectors of the constituent words is used as the sentence-level event representation, which is then encoded into the initial representation of the hypergraph node:
[0029]
[0030] The initial representation of the hypergraph nodes of the entire document is:
[0031]
[0032] Connect the two nodes with causal relationship of paired events in step (2). For each event node, aggregate all its paired causal relationships to create a hyperedge. Construct a hypergraph using event nodes and hyperedges, denoted as:
[0033]
[0034] Among them, ε is the set of event nodes, It is a super edge set.
[0035] Furthermore, in step (4), a hypergraph neural network is used to obtain a document-level event representation of each event in the document through hypergraph convolution learning based on the initial representation and causal structure of the hypergraph nodes:
[0036]
[0037] Furthermore, in step (5), the paired event vector H obtained in step (2) PES and the document-level event vector E obtained in step (4) (l) , concatenate the two event representation vectors of each event pair as the final representation of causal relationship classification:
[0038] v k =[(e i -ej ) || (e i + e j ) || h j || h i
[0039] where || represents the concatenation operation,
[0040] Furthermore, in the step (5), the representation v of each event pair is transformed into a causal probability distribution through a multi-layer perceptron network, and the probability distribution is softmax-normalized to obtain the probability that there is a causal relationship between each event pair, which is expressed by the formula as follows: k where W
[0041]
[0042] where W c , b c are learnable parameters, is the probability value that the predicted event pair has a causal relationship.
[0043] Furthermore, in the step (2), the pre-trained language model uses a RoBERTa-based masked language model to predict the causal relationship between paired events in the document using the masked model specific token [MASK] in prompt learning.
[0044] Furthermore, in the step (4), the pre-trained language model uses the RoBERTa model; the hypergraph neural network consists of two hypergraph convolutional layers, and the transformation function of this hypergraph convolutional layer is:
[0045]
[0046] where σ is the ReLU activation function, represents the incidence matrix of the hypergraph If the event e n is a node of the hyperedge r on the hypergraph m , Α(e n , r m ) = 1, otherwise Α(e n , r m ) = 0, D e and D r respectively represent the diagonal matrices of the node degree and the edge degree, W is the identity matrix, is a learnable parameter, and the output after l layers of convolution is and E (0) = H DCS .
[0047] Further, in step (6), the loss function adopts the cross - entropy loss function, which is expressed by the formula as follows:
[0048]
[0049] where y (k) and are the true label and the predicted label of the k - th event pair in the document respectively, and λ and θ are regularization hyperparameters.
[0050] All in all, compared with the prior art, the above - mentioned technical solution conceived by the present invention can achieve a better event causality recognition effect: Since the prompting learning paradigm is adopted, the pairwise event semantic information is fused into the event representation vector based on the document, which is beneficial to the mining of specific semantic information; Since the hypergraph neural network is adopted to capture the interdependence of potential document - level events, the final event representation vector contains more causal structure information. The combination of pairwise event semantic information and document - level event causal structure information promotes the improvement of the event causality recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is the structural diagram of the NCHM model proposed by the present invention;
[0052] Figure 2 is the performance of the model proposed by the present invention on different hyper - edge degrees in the ECS 0.9 dataset;
[0053] Figure 3 is the performance of the model proposed by the present invention on different event - pair distances in the ECS 0.9 dataset;
[0054] Figure 4 is the visual display of the intra - sentence and inter - sentence ECI results of the model proposed by the present invention in a specific document. DETAILED DESCRIPTION OF THE INVENTION
[0055] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0056] The overall idea of the method for identifying event causal relationships based on hypergraph modeling of document-level causal structures in the present invention is that the method first preprocesses the text in two ways: paired events and events, and then uses a pre-trained language model based on prompt learning to model the paired event sentences. After paired event semantic learning, the hidden vectors of event mentions are obtained, and the causal relationship between event pairs is predicted simultaneously. Then, a pre-trained language model is used to model the sentences to obtain the hidden vectors of event mentions. Combining the predicted causal relationship between event pairs, a document-level causal relationship structure is modeled based on hypergraphs, and combined with hypergraph convolution learning to obtain a document-level event representation. Finally, the two representations are concatenated and input into a multi-layer perceptron to obtain the probability that there is a causal relationship between event pairs in the document.
[0057] As Figure 1 shown, the method for identifying event causal relationships based on hypergraph modeling of document-level causal structures in the present invention includes the following steps:
[0058] (1) Text preprocessing step: Input the original document and preprocess it in two ways: paired events and events to obtain two types of data: sentence pairs where paired events are located and sentences where events are located;
[0059] (2) Paired event semantic learning step: Use a PLM based on prompt learning, combined with a custom template, to model the sentence pairs where paired events are located in step (1) to obtain the context semantic representation of paired events and the prediction result of the causal relationship between paired events;
[0060] Specifically, it includes the following sub-steps:
[0061] (2-1) First, construct each event pair in the document x k =(Evt i ; Evt j ) into a prompt template T p (x) that can describe the potential causal relationship between two event mentions:
[0062] T p (x k ) = In this sentence, Evt i [MASK]Evt j .
[0063] Among them, Evt i and Evt j are two event mentions, and a specific token [MASK] of the PLM is inserted between them for relationship prediction.
[0064] is the complete context semantics including the sentence, and then the original sentence T s of the event mention is combined with the constructed prompt template T pConnect them as the input sentence T of the PLM. Use the PLM-specific tokens [CLS] and [SEP] to represent the beginning and end of the input sentence T, and another [SEP] for T s and T p as the separation token.
[0065] (2-2) Encode the input sentence T using the PLM, and two event mentions and the hidden vector of the specific token [MASK] can be obtained from the output:
[0066]
[0067] where and are the hidden vectors of the two event mentions, is the hidden vector of [MASK], and d is the dimension of the hidden vector.
[0068] If an event mention consists of multiple words, use the average of their hidden vectors as the event mention representation. Combining the hidden vectors of the two event mentions can obtain the context semantic representation of the paired events:
[0069]
[0070] The context semantic representation of the paired events in a document:
[0071]
[0072] where k is the number of pairs of all paired events in the document.
[0073] (2-3) Combine the two virtual answer words added by the present invention in the PLM vocabulary, namely Casual and None, and use through the MLM classifier of the PLM to estimate the probabilities that [MASK] is the two virtual answer words, and use the predicted virtual answer word with the higher probability as the prediction result of the causal relationship of the paired events.
[0074] In step (2), the pre-trained language model uses the RoBERTa-based masked language model to predict the causal relationship of the paired events in the document by using the masked model-specific token [MASK] in prompt learning.
[0075] (3) Steps for pre-training the paired event semantic learning module: Construct a cross-entropy loss function based on the predicted virtual answer word and the true answer word, pre-train the paired event semantic learning module by minimizing the loss function, and select the model with the highest overall F1 in the validation set as the paired semantic learning model;
[0076] (4) Document-level causal structure learning steps: Input the sentences where the events in step (1) are located into another PLM to obtain sentence-level event representations, and construct a document-level causal hypergraph by combining the prediction results of pairwise event causal relationships in step (2). Through the hypergraph convolutional neural network, obtain the document-level event representations;
[0077] Construct a document causal hypergraph according to the pairwise causal relationships predicted in step (2-3), and combine it with the hypergraph convolutional neural network to obtain the document-level event representations. It includes the following sub-steps:
[0078] (4-1) Use a hypergraph to model the causal structure of each document, where each node represents an event, and the hyperedges are the causal relationships of mutual dependence among multiple events.
[0079] Input the sentences where the events in step (1) are located into another PLM one by one. Through the encoding of the PLM, obtain the sentence-level event representations. For event mentions composed of multiple words, use the average value of the hidden vectors of the constituent words as the sentence-level event representation, and then encode it as the initial representation of the hypergraph node:
[0080]
[0081] Initial representation of the hypergraph nodes of the entire document:
[0082]
[0083] Connect the two nodes with pairwise causal associations of events in step (2). For each event node, aggregate all its pairwise causal relationships to create a hyperedge, and construct a hypergraph with the event node and the hyperedge, denoted as:
[0084]
[0085] where ε is the set of event nodes, is the set of hyperedges.
[0086] (4-2) Use a hypergraph neural network. Based on the initial representation of the hypergraph nodes and the causal structure, through hypergraph convolution learning, obtain the document-level event representations of each event in the document:
[0087]
[0088] In step (4), the pre-trained language model uses the RoBERTa model; the hypergraph neural network consists of two hypergraph convolutional layers, and the transformation function of this hypergraph convolutional layer is:
[0089]
[0090] where σ is the ReLU activation function, represents the hypergraph The incidence matrix, if event e n is a hypergraph on the hyperedge r m is a node, Α(e n , r m ) = 1, otherwise Α(e n , r m ) = 0, D e and D r respectively represent the diagonal matrices of node degree and edge degree, W is the identity matrix, is a learnable parameter, and the output after l - layer convolution is and E (0) = H DCS .
[0091] (5) Document - level causal relationship recognition steps: Concatenate the context semantic representations of paired events in step (2) and the document - level event representations in step (4), and predict the probability of the existence of a causal relationship for each event pair in the document through a multi - layer perceptron network;
[0092] Connect the two representations obtained in steps (2 - 2) and (4 - 2), and predict the probability of the existence of a causal relationship for each event pair in the document through a multi - layer perceptron network. It includes the following sub - steps:
[0093] (5 - 1) According to the paired event vector H PES obtained in step (2 - 2) and the document - level event vector E (l) obtained in step (4 - 2), concatenate the two event representation vectors of each event pair as the final representation for causal relationship classification:
[0094]
[0095] where || represents the concatenation operation,
[0096] (5 - 2) Transform the representation v k of each event pair into a causal probability distribution through a multi - layer perceptron network, and normalize the probability distribution by softmax to obtain the probability of the existence of a causal relationship for each event pair. It is expressed by the formula as follows:
[0097]
[0098] where W c , b c are learnable parameters, is the probability value of the existence of a causal relationship for the predicted event pair.
[0099] (6) Training and testing the network steps: Based on the predicted causal probability distribution Construct a loss function with the true causal label y, then train the network to minimize the loss function. After training is completed, input the validation set and test set documents, and select the model with the highest F1 value on the validation set documents, so as to obtain the causal relationship prediction results of the corresponding test samples.
[0100] In step (6), the loss function adopts the cross-entropy loss function, which is expressed by the formula as follows:
[0101]
[0102] where y (k) and are respectively the true label and the predicted label of the k-th event pair in the document, and λ and θ are regularization hyperparameters.
[0103] Taking the EventStoryLine 0.9Corpus (denoted as: ESCv0.9) dataset widely used in the ECI task as an example, the performance effect of the event causal relationship recognition method based on hypergraph modeling of document-level causal structures proposed in the present invention is demonstrated. The ESC dataset consists of news documents from different news websites, including 22 topics, 258 documents, and a total of 5334 event mentions. A total of 5625 pairs of event pairs are marked as having a causal relationship, among which 1770 pairs are intra-sentence causal relationships and 3855 pairs are inter-sentence causal relationships. The same as the standard data division, the last two topics are used as the validation set, and the remaining 20 topics are used for 5-fold cross-validation. The accuracy (P), recall (R), and F1 value of the average result are used as performance indicators.
[0104] Use the 768-dimensional pre-trained language model RoBERTa provided by HuggingFace transformers and run the PyTorch framework with CUDA on NVIDIA GTX 3090 GPUs. RoBERTa is a language model proposed by Facebook that is pre-trained in an unsupervised manner by performing cloze tasks on a large amount of unlabeled text. The learning rate of the experiment is set to 1e-5, the number of hyper-layers is set to 2, and all trainable parameters are randomly initialized from a normal distribution. We use the Adam optimizer with L2 regularization and combine dropout for model training.
[0105] To further explore the influence of the interaction of multiple causal relationships, Figure 2 Show the NCHM model proposed in the present invention and the method that only uses the document-level event vector E in the final causal classification step in the form of a column-line chart (l)Performance of the NCHM model (denoted as: NCHM w / o PES) at different hyper-degrees. Here, the hyper-degree refers to the number of nodes connected by this hyper-edge. As can be seen from the figure, even when the number of instances decreases, both models benefit from the increase in hyper-degree. This shows that by learning the causal structure of the document to mine the interactions between multiple cross-sentence events, it helps to improve the performance of document-level causal relation recognition.
[0106] To explore the effectiveness of the NCHM model proposed in the present invention in document-level causal recognition, Figure 3 The performance of the model at different event distances is shown in the form of a column-line chart. Here, the number of event mentions between two event mentions is used as the event distance, and NCHM w / o DCS refers to only using the pairwise event vector H in the final causal classification step PES of the NCHM model. As can be seen from the figure, the recognition of causal relationships for events that are far apart benefits from the introduction of the document-level event vector. This further shows that hypergraph modeling of the document-level causal structure helps to mine the interdependencies between multiple events and is conducive to improving the causal relation recognition effect of the document-level NCHM model.
[0107] Figure 4 Shows the causal recognition of intra-sentence and inter-sentence event pairs in a certain document in the ESC v0.9 dataset for the NCHM model and the NCHM w / o DCS model. As can be seen from the figure, for intra-sentence event pairs, the recognition effects of both models are good; for inter-sentence event pairs, the NCHM w / o DCS model is much worse than the NCHM model, and the gap is concentrated in the same event "charged". This further shows that the causal relationships of events in the document are usually interdependent, and this should be utilized to enhance document-level causal relation recognition.
[0108]
[0109]
[0110] Table 1
[0111] Table 1 shows the performance comparison of the NCHM model proposed in the present invention with existing competing models in three aspects: intra-sentence, inter-sentence, and overall. As can be seen from the table, the performance of the model proposed in the present invention is significantly better than that of existing competing models, and the improvement in inter-sentence causal relation recognition is particularly obvious, indicating that both the learned mutual relationships between multiple events and the encoding of pairwise event semantics based on hypergraph modeling of the document-level causal structure contribute to the causal relation recognition of document-level event pairs.
[0112] To compare the performance of the present invention based on hypergraph modeling of the document-level causal structure and prompting learning, the present invention tested the causal relation recognition effects of 3 schemes, which are respectively:
[0113] (1) NCHM (layer=x): Use x-layer hypergraph convolutional layers, that is, set the number of hypergraph convolutional layers in step (4) to x layers, and keep the other steps unchanged.
[0114] (2) NCHM w / o DCS: Do not adopt hypergraph modeling of document-level causal structure, and only use the pairwise event vector H in the final causal classification step PES , that is, skip step (4) and subsequent steps, and directly use the pairwise event causal relationship predicted in step (2) as the prediction result of the final model;
[0115] (3) NCHM w / o PES: In the final causal classification step proposed by the present invention, only the document-level event vector E is used (l) , that is, only input the document-level event vector E obtained in step (4-2) in step (5-1) (l) , and then enter step (5-2) and subsequent steps;
[0116]
[0117]
[0118] Table 2
[0119] Table 2 shows the causal relationship recognition performance of schemes (1)-(3). It can be seen from the table that setting the number of hypergraph convolutional layers x to 2 has higher causal relationship recognition performance compared to other numbers of layers. The results prove that when the number of hypergraph convolutional layers is too small, the aggregated information is insufficient, and when the number of hypergraph convolutional layers is too large, there will be an over-smoothing effect. Compared with the complete model, the causal relationship recognition performance of schemes (2) and (3) is inferior. Therefore, it shows that the mutual relationship between multiple events learned based on hypergraph modeling of document-level causal structure and the pairwise event semantic information mined based on prompt learning are both beneficial to improving the effect of causal relationship recognition.
[0120] It is easy for those skilled in the art to understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An event causal relationship recognition method based on hypergraph modeling of document-level causal structures, characterized in that, it includes a text preprocessing step, a pairwise event semantic learning step, a pre-trained pairwise event semantic learning module, a document-level causal structure learning step, a document-level causal relationship recognition step, and a training and testing network step; where: (1) Text preprocessing step: Input the original document, and preprocess it in two ways: pairwise events and events, to obtain two types of data: the sentence pairs where pairwise events are located and the sentences where events are located; (2) Pairwise event semantic learning step: Use a PLM based on prompt learning, combined with a custom template, to model the sentence pairs where pairwise events are located in step (1), to obtain the context semantic representation of pairwise events and the prediction result of the causal relationship between pairwise events; (3) Pre-training pairwise event semantic learning module step: Construct a cross-entropy loss function based on the predicted virtual answer words and the true answer words, and pre-train the pairwise event semantic learning module by minimizing the loss function, and select the model with the highest overall F1 in the validation set as the pairwise semantic learning model; (4) Document-level causal structure learning step: Input the sentences where events are located in step (1) into another PLM to obtain sentence-level event representations, combine the prediction results of the causal relationships between pairwise events in step (2) to construct a document-level causal hypergraph, and pass through a hypergraph convolutional neural network to obtain document-level event representations; (5) Document-level causal relationship recognition step: Concatenate the context semantic representation of pairwise events in step (2) and the document-level event representation in step (4), and predict the probability that each event pair in the document has a causal relationship through a multi-layer perceptron network; (6) Steps for training and testing the network: Based on the predicted causal probability distribution and the true causal label y, construct a loss function, and then train the network to minimize the loss function. After training, input the validation set and test set documents, and select the model with the highest F1 value on the validation set documents, so as to obtain the causal relationship prediction results of the corresponding test samples.
2. The event causal relationship recognition method according to claim 1, characterized in that, the step (2) includes the following sub-steps: (2-1) First, document x k =(Evt i ; Evt j ) each pair of events in is constructed into a hint template T p (x): T p (x k ) = In this sentence, Evt i [MASK]Evt j . Among them, Evt i and Evt j are two event mentions. A specific marker [MASK] of PLM is inserted between them for relationship prediction, and then the original sentence T s of the event mention is connected with the constructed prompt template T p to form the input sentence T of the PLM. The specific markers [CLS] and [SEP] of the PLM are used to indicate the beginning and end of the input sentence T, and another [SEP] is used for the separation marker between T s and T p ; (2-2) Use the PLM to encode the input sentence T, and obtain the hidden vectors of two event mentions and the specific token [MASK] from the output: where and are the hidden vectors mentioned in two events, is the hidden vector of [MASK], and d is the dimension of the hidden vector; If the event mention consists of multiple words, use the average of their hidden vectors as the event mention representation, and combine the hidden vectors of the two event mentions to obtain the context semantic representation of the pairwise events: The context semantic representation of pairwise events in a document: where k is the number of pairs of pairwise events in the document.
3. The event causal relationship recognition method according to claim 1 or 2, characterized in that, In the step (2), two virtual answer words, namely Casual and None, are added to the PLM vocabulary. Based on the MLM classifier of the PLM, according to to estimate the probability that [MASK] is the two virtual answer words, and the predicted virtual answer word with a higher probability is used as the prediction result of the causal relationship of the paired events. is the hidden vector of [MASK].
4. The event causal relationship recognition method according to claim 1 or 2, characterized in that, in the step (4), a hypergraph is used to model the causal structure of each document, where each node represents an event, and the hyperedge is the causal relationship of mutual dependence between multiple events, specifically including: Input the sentences where events are located in step (1) into another PLM one by one, and obtain sentence-level event representations through PLM encoding. For event mentions composed of multiple words, use the average of the hidden vectors of the constituent words as the sentence-level event representation, and then encode it into the initial representation of the hypergraph node: The initial representation of the hypergraph nodes of the entire document: Connect the two nodes with causal associations of paired events in step (2). For each event node, aggregate all its paired causal relationships to create a hyperedge. Construct a hypergraph using the event nodes and hyperedges, denoted as: Among them, ε is the set of event nodes, is the set of hyperedges.
5. The method for identifying event causal relationships according to claim 1 or 2, characterized in that, in the said step (4), a hypergraph neural network is used, and based on the initial representation of hypergraph nodes and the causal structure, the document-level event representation of each event in the document is obtained through hypergraph convolution learning:
6. The method for identifying event causal relationships according to claim 5, characterized in that, In step (5), according to the paired event vector H obtained in step (2) PES and the document-level event vector Ε obtained in step (4) (l) , concatenate the two event representation vectors of each event pair as the final representation for causal relationship classification: where ∥ represents the series operation, i and j range from 1 to n, and x and y range from 1 to k.
7. The method for identifying event causal relationships according to claim 1 or 2, characterized in that, In the step (5), the representation v of each event pair is transformed into a causal probability distribution through a multi-layer perceptron network, and the probability distribution is normalized by softmax to obtain the probability that there is a causal relationship between each event pair, which is expressed by the following formula: k and the probability distribution is normalized by softmax to obtain the probability that there is a causal relationship between each event pair, which is expressed by the following formula: Among which W c , b c are learnable parameters, is the probability value of the predicted causal relationship existing between event pairs.
8. The method for identifying event causal relationships according to claim 1 or 2, characterized in that, in the said step (2), the pre-trained language model uses a RoBERTa-based masked language model, and uses the masked model specific token [MASK] in prompt learning to predict the causal relationships of paired events in the document.
9. The method for identifying event causal relationships according to claim 1 or 2, characterized in that, in the said step (4), the pre-trained language model uses the RoBERTa model; the hypergraph neural network consists of two hypergraph convolution layers, and the transformation function of this hypergraph convolution layer is: where σ is the ReLU activation function, represents the incidence matrix of the hypergraph . If the event e n is a node of the hyperedge r on the hypergraph m , then Α(e n , r m ) = 1; otherwise, Α(e n , r m ) = 0. D e and D r represent the diagonal matrices of the node degree and the edge degree respectively. W is the identity matrix. is a learnable parameter. The output after l - layer convolution is and Ε (0) = Η DCS .
10. The method for identifying event causal relationships according to claim 1 or 2, characterized in that, in the said step (6), the loss function adopts the cross-entropy loss function, and is expressed by the formula as follows: where y (k) and are the true label and the predicted label of the k-th event pair in the document, respectively, and λ and θ are regularization hyperparameters.
Citation Information
Patent Citations
Emergency clue extraction method based on news reports
CN110737819A
Document-level event causal relationship identification method and system, medium, equipment and terminal
CN115577678A