Causal inference method fusing co-reference relationship and external knowledge enhancement

By integrating the causal inference method of coreference relationship and external knowledge enhancement, and using the improved BERT model and graph attention mechanism to construct a causal graph, the problem of insufficient causal chain construction in the existing technology is solved, and more accurate and reliable causal relationship identification is achieved.

CN120611784AActive Publication Date: 2025-09-09GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510583164.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-09
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

Existing methods fail to fully utilize external knowledge bases and text language features in document-level event causal relationship identification, resulting in the causal chain construction not fully reflecting the complex logical relationship between events and failing to maximize the use of potential information.

Method used

A causal inference method that integrates coreference relationships and external knowledge enhancement is adopted. The semantic features of event words are extracted through the improved BERT model. The contrastive learning mechanism and graph attention mechanism are combined to construct a causal graph of the external knowledge base. The coreference features and context features are integrated, and the multi-head attention mechanism is used to generate a comprehensive feature representation.

Benefits of technology

It improves the accuracy and robustness of causal relationship identification, enhances the understanding of complex logical relationships between events and the in-depth utilization of knowledge, and improves the accuracy and reliability of causal inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611784A_ABST
    Figure CN120611784A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, and provides a causal inference method fusing a co-reference relationship and external knowledge enhancement, and the method comprises the steps: extracting event words in a text as mask target words, and carrying out the mask operation of non-event words, using an improved BERT model to predict a masked event word through context reconstruction training in a pre-training stage; extracting co-referring event pairs of the reason event and the result event to obtain co-referring features of co-referring events; extracting potential indirect events associated with the cause / result events from the external knowledge base to construct an external knowledge base cause and effect graph, constructing a structured semantic graph, and performing node feature fusion; the co-reference features and the fused node features are fused and then nonlinear transformation is carried out; and fusing the context semantic features and the external knowledge base features to obtain comprehensive feature representation, and inputting the comprehensive feature representation into a prediction layer to obtain a final prediction result of the event pair relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a causal inference method that integrates coreference relationships with external knowledge enhancement. Background Art

[0002] Event Causality Identification (ECI) is a key and challenging task in the field of natural language processing (NLP). It aims to identify the causal relationship between events from text. This is of great significance for applications such as understanding the deep semantics of text, building knowledge graphs, question-answering systems, and text summarization.

[0003] Document-Level Event Causality Identification (D-ECI) is a highly challenging task in natural language processing. Constructing relationships between document-level events typically relies on external knowledge bases to enhance the coherence of causal chains. However, existing methods often only consider the target event for causal identification as the starting and ending points of the causal chain. This not only limits the in-depth utilization of external knowledge bases but also undermines the full exploration of text semantic features. This limitation results in the construction of causal chains that cannot fully reflect the complex logical relationships between events and fail to maximize the potential information of the text and knowledge base. Summary of the Invention

[0004] In response to the above-mentioned defects, the purpose of the present invention is to propose a causal inference method that integrates coreference relationships and external knowledge enhancement, aiming to achieve accurate semantic mapping of external knowledge and document events, and improve the accuracy and robustness of causal relationship identification.

[0005] To achieve this object, the present invention adopts the following technical solutions:

[0006] A causal inference method that integrates coreference relationships and external knowledge enhancement includes the following steps:

[0007] S1: Event words are extracted from document-level text as masked target words and non-event words are randomly masked. The improved BERT model is used to predict masked event words through context reconstruction training in the pre-training phase, so that event words can fully learn contextual semantic features.

[0008] S2: Randomly extract event pairs from the text and label them as cause events and result events, extract coreference event pairs between the cause events and the result events, obtain labeled coreference event positive samples from an external knowledge base, and simultaneously generate non-coreference event pairs as negative samples. Utilize a contrastive learning mechanism to design a contrastive loss function to shorten the semantic distance between the coreference event pairs, and obtain the coreference features of the coreference events.

[0009] S3: Extract potential indirect events associated with cause / result events from the external knowledge base to construct an external knowledge base causal graph. A structured semantic graph is constructed with events in the external knowledge base causal graph as nodes and causal relationships as edges. A graph attention mechanism and a graph convolutional neural network are used to fuse node features. The potential indirect events include indirect events between cause events and result events.

[0010] S4: Fuse the coreference features in step S2 and the node features fused in step S3, and perform nonlinear transformation through a fully connected layer to obtain the external knowledge base features of the cause event and the result event;

[0011] S5: A multi-head attention mechanism is used to fuse the contextual semantic features and the external knowledge base features to obtain a comprehensive feature representation, which is input into the prediction layer to obtain the final prediction result of the event pair relationship.

[0012] Preferably, in step S1, the relationship is satisfied:

[0013]

[0014] Where X=[x1,x2,…,x i ,…x N ] represents the word sequence of the input text, x i represents the i-th word, E represents the event word set, N text Indicates the length of the word sequence of the text; P mask (x i ) represents the probability of masking the i-th word, α and β are the masking probabilities, Represents the contextual semantic features of the i-th word, that is, the embedding vector of the event word.

[0015] Preferably, step S2 includes:

[0016] By designing a contrastive loss function through contrastive learning strategy, the semantic distance between the cause event and its coreference event, and between the result event and its coreference event are shortened, while the semantic distance between non-coreference event pairs is shortened.

[0017] Construct positive sample co-reference pairs (e i ,e j)∈C, randomly combined from non-coreferential events A projection layer is introduced to map events into a new semantic space, and cosine similarity is used to calculate the contextual relevance of co-referencing event pairs, satisfying the relationship:

[0018]

[0019] Among them, W and b represent the weight matrix and bias of the projection layer respectively, represents the contextual semantic features of the i-th word, z i Denotes the embedding vector h i The projection vector of is the loss function of contrast training, s(z i ,z j ) represents z i With z j The similarity between them is τ, which represents the temperature coefficient.

[0020] Preferably, in step S3, the process includes extracting potential indirect events associated with the cause / result events from an external knowledge base, and identifying and extracting the causal relationships between the events using entity linking and relationship extraction techniques;

[0021] The causal relationship is serialized and then semantically encoded using the parameter-frozen EMBert model to satisfy the relationship:

[0022] Input=[CLS]e0[indirect]e1[indirect]…[indirect]e n [effect];

[0023]

[0024] Among them, Input represents the input sequence, [indirect] represents the indirect causal relationship between events, and [effect] represents the final result event of the causal chain. Indicates event e i The embedding vector in sequence k.

[0025] Preferably, in step S3, the process includes calculating the embedding vectors of the event nodes between the event nodes of the causal graph. and Combine the attention weights to fuse the weights of the same nodes and edges in different sequences to satisfy the relationship:

[0026]

[0027]

[0028] in, Represents event node e in sequence k i To event node e j The weight, W g is a learnable matrix, Represents event node e in sequence k i With e j The attention weight between T is a learnable vector, is the weight between global event nodes.

[0029] Preferably, in step S3, mean pooling is used to fuse the features of the same event nodes in multiple different sequences into a unified global feature, satisfying the relationship:

[0030]

[0031] in Indicates event e i Embedding vectors in the sequence, Indicates nodes with the same event e i the number of sequences, To fuse the global feature representation of each sequence event node.

[0032] Preferably, in step S3, the attention coefficient between each node and its adjacent nodes is calculated according to the graph attention mechanism, and the features of the adjacent nodes are fused. At the same time, the edge weight information is fused into the node features using a graph convolutional neural network to satisfy the relationship:

[0033]

[0034] in, Represents the attention coefficient between adjacent nodes, A is the attention coefficient matrix, H (0) Represents the node features of the initial layer, W GAT represents the learnable feature transformation matrix, H (1) Represents the features after the fusion of the first-layer graph attention mechanism, H (2) Represents the features after fusion of the second layer of graph convolutional neural network, represents the degree matrix, Represents the normalized adjacency matrix of the graph, W GCN Represents the weight between global event nodes, each element is

[0035] Preferably, step S4 satisfies the relationship:

[0036]

[0037] f=[y1;y2;…;y n ];

[0038] in Represents the node features of the structured semantic graph, z i Represents the coreference feature of the event node, y i It represents the fusion result of the node features and coreference features of the structured semantic graph, and f represents the feature concatenation vector of each coreference event of the cause event or result event.

[0039] Preferably, step S5 satisfies the relationship:

[0040]

[0041] M=[head1;…;head h ]W O ;

[0042] Among them, C cause with C effect Respectively represent the concatenated cause event and result event feature vectors, W i Q , Represents the weight matrices of Q, K, and V, respectively. O is the weight matrix of the multi-head attention output mapping, M is the comprehensive feature representation of the cause event and the result event, The contextual semantic features of the cause event in the text, Represents the contextual semantic features of the result event in the text, f cause The external knowledge base features representing the cause event, f effect Represents the external knowledge base features of the result event.

[0043] Preferably, step S5 includes inputting the comprehensive feature representation into the prediction layer, determining whether there is a causal relationship between the event pairs, and calculating the loss thereof, to obtain the final prediction result of the event pair relationship, satisfying the relationship:

[0044] p i =sigmoid(WM+b);

[0045]

[0046] Among them, p i Expressed as predicted probability, y i is the true label, N is the number of samples of event pairs, Expressed as a binary cross entropy loss function, is the comprehensive loss, M is the comprehensive feature representation of the cause event and the result event, b is the bias of the projection layer, and λ is the weight coefficient of the loss function. represents the loss function for contrastive training.

[0047] One of the above technical solutions has the following advantages or beneficial effects:

[0048] The present invention optimizes the pre-training process through an event masking strategy, improves the accuracy of event semantic analysis, and helps to enhance the efficiency of event semantic feature extraction in a contextual environment; utilizes a contrastive learning mechanism to shorten the semantic distance of co-referencing event pairs, effectively captures the co-referencing relationship between events, enhances the model's ability to deeply mine text semantics, and deeply considers the relevance of co-referencing events with target event pairs in the text context. Furthermore, by integrating co-referencing event nodes in a causal graph constructed from an external knowledge base and incorporating external knowledge associated with these co-referencing events into the causal inference process, the semantic information of the causal graph can be enriched, and by strengthening the logical reasoning between events, the comprehensive analysis capability of event causal relationships can be significantly improved; by fusing co-referencing features with external knowledge base features through a fully connected layer nonlinear transformation, a more comprehensive external knowledge base feature representation is obtained; and by adopting a multi-head attention mechanism to fuse contextual semantic features and external knowledge base features, a high-quality comprehensive feature representation is generated, thereby improving the accuracy of causal relationship prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0050] Figure 1 This is a flow chart of a causal inference method that integrates coreference relationships and external knowledge enhancement, provided by an embodiment of the present invention;

[0051] Figure 2 This is a diagram illustrating an implementation framework of a causal inference method that integrates coreference relationships and external knowledge enhancement, as provided in an embodiment of the present invention;

[0052] Figure 3 This is a schematic diagram of extracting causal chains and constructing coreference sample pairs through an external knowledge base in a causal inference method that integrates coreference relationships and external knowledge enhancement, provided by an embodiment of the present invention;

[0053] Figure 4 It is a schematic diagram of the fusion of coreference features and causal graph node features of the causal inference method that integrates coreference relationships and external knowledge enhancement provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0055] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "axial", "radial", "circumferential" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0056] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "plurality" means two or more.

[0057] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to direct connections, indirect connections through an intermediary, or internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0058] Causal inference methods that integrate coreference relationships with external knowledge enhancement, such as Figure 1 As shown, in a preferred embodiment of the present invention, the causal inference method proposed in the present invention can be executed by a corresponding causal inference model or system, and the causal inference method includes the following steps:

[0059] S1: Event words are extracted from document-level text as masked target words and non-event words are randomly masked. The improved BERT model is used to predict masked event words through context reconstruction training in the pre-training phase, so that event words can fully learn contextual semantic features.

[0060] Specifically, in the field of natural language processing, event words, as the carriers of key actions or state information in the text, contain rich semantic information and contextual associations, while document-level texts often have complex semantic structures and long dependencies. Step S1 preprocesses the document-level text, first extracting event words from the text and using them as masked target words, while randomly masking non-event words. The improved BERT model is used to predict masked event words through context reconstruction training in the pre-training stage. This enables event words to fully learn contextual semantic features on large-scale unsupervised data, thereby obtaining more accurate contextual embedding representations. Its purpose is to improve the ability to understand event semantics and context perception, lay a solid foundation for subsequent causal inference tasks, and enhance the ability to capture potential causal relationships between events in the text.

[0061] The Improved BERT model is a deep language model optimized on the original BERT architecture to address issues with the traditional BERT model (insufficiently accurate semantic feature extraction of event words and difficulty capturing the logical connections between event words). It has stronger context encoding and semantic understanding capabilities, and therefore requires the use of the Improved BERT model.

[0062] S2: Randomly extract event pairs from the text and label them as cause events and result events, extract coreference event pairs between the cause events and the result events, obtain labeled coreference event positive samples from an external knowledge base, and simultaneously generate non-coreference event pairs as negative samples. Utilize a contrastive learning mechanism to design a contrastive loss function to shorten the semantic distance between the coreference event pairs, and obtain the coreference features of the coreference events.

[0063] In text data, the causal relationship between event pairs is often implicit and complex and changeable. Co-referential events, as different expressions of the same event at different positions in the text, contain rich semantic associations and contextual clues, which can provide strong support for causal inference. Step S2 randomly extracts event pairs from the text and labels them as cause events and result events, extracts co-referential event pairs between the two, and simultaneously obtains co-referential event positive samples and generates non-co-referential event negative samples by combining with an external knowledge base. The contrastive learning mechanism is used to design a contrastive loss function to shorten the semantic distance of co-referential event pairs, thereby obtaining the co-referential features of co-referential events. Its purpose is to mine the co-referential information contained in the text, strengthen the understanding and modeling of the semantic associations between co-referential events, so as to more accurately capture the potential causal logic between event pairs, improve the accuracy of causal inference, and enhance the deep mining and utilization efficiency of text semantics.

[0064] Among them, event pairs are combinations of two events in the text with potential causal relationships, and are the core analysis objects of causal inference tasks; cause events and result events represent the events that cause and are caused by the results in the event pair, respectively. Clearly labeling the two helps focus on the causal logic direction; co-referential event pairs refer to different forms of expression referring to the same event in the text. They are highly semantically consistent. By extracting co-referential event pairs, we can enrich our understanding of event semantics and provide multi-angle clues for causal inference; external knowledge bases are structured data sets that store a large amount of domain knowledge or common sense knowledge, which contain rich co-referential event annotated samples and causal relationship information, which can be used to supplement the missing semantic information in the text and enhance knowledge coverage; contrastive learning mechanism is an effective machine learning strategy that designs a contrastive loss function to make similar samples closer in the semantic space and dissimilar samples farther apart, thereby enhancing the ability to discriminate sample features; contrastive loss function is a key component in contrastive learning, which is used to measure the semantic similarity difference between sample pairs and guide parameter optimization, promote the convergence of semantic embedding of co-referential event pairs, and enable more accurate distinction between co-referential and non-co-referential event pairs.

[0065] S3: Extract potential indirect events associated with cause / result events from the external knowledge base to construct an external knowledge base causal graph. A structured semantic graph is constructed with events in the external knowledge base causal graph as nodes and causal relationships as edges. A graph attention mechanism and a graph convolutional neural network are used to fuse node features. The potential indirect events include indirect events between cause events and result events.

[0066] External knowledge bases contain vast amounts of domain knowledge and causal relationship information, which complement the events in the text and provide a more comprehensive semantic perspective and knowledge context for causal inference. The purpose of step S3 is to integrate the causal links in the external knowledge base to form a structured causal knowledge representation. By leveraging the feature fusion capabilities of graph neural networks, we can fully exploit the semantic information in the external knowledge base and effectively integrate it into the modeling process of the target event pair, thereby improving the knowledge depth and logical coherence of causal inference and enhancing the ability to reason about complex causal logic between events.

[0067] Among them, potential indirect events refer to intermediate events that have indirect causal relationships with cause events and result events. They play a key bridging role in the causal chain and can expand the depth and complexity of causal relationships; the causal graph is a semantic graph structure constructed with events as nodes and causal relationships as edges. It intuitively displays the causal logic and association paths between events and is an important basis for knowledge fusion and reasoning; the graph attention mechanism is a graph neural network technology that dynamically aggregates the features of adjacent nodes by calculating the attention coefficient between nodes, highlights the influence of important nodes on target nodes, and improves the pertinence and effectiveness of feature fusion; the graph convolutional neural network is deep learning based on graph structured data, which can effectively fuse node features and edge weight information, perform feature learning and information propagation of graph structured data, and make full use of the structured semantic information of the causal graph.

[0068] S4: Fuse the coreference features in step S2 and the node features fused in step S3, and perform nonlinear transformation through a fully connected layer to obtain the external knowledge base features of the cause event and the result event;

[0069] Among them, the co-reference feature is a semantic feature extracted from co-referential event pairs in the text through a contrastive learning mechanism. It reflects the semantic consistency of events within the text and provides semantic clues for events within the text; the node feature is a feature fused from the causal graph of the external knowledge base through a graph neural network. It contains event semantics and causal relationship information in the external knowledge base and can supplement rich background knowledge; the fully connected layer is a basic component in the neural network. It realizes nonlinear transformation and fusion of features through the fully connected structure of neurons, and can learn complex associations and interaction patterns between features, making the fused features more discriminative and expressive; the external knowledge base feature is a fused feature representation. It integrates the information of the co-reference feature and the node feature of the external knowledge base. It is an important feature basis for subsequent causal inference tasks and determines the degree of comprehensive understanding of the semantics and knowledge of event pairs by the causal inference method.

[0070] S5: A multi-head attention mechanism is used to fuse the contextual semantic features and the external knowledge base features to obtain a comprehensive feature representation, which is input into the prediction layer to obtain the final prediction result of the event pair relationship.

[0071] In causal inference models, contextual semantic features and external knowledge base features provide rich feature descriptions of event pairs from the perspectives of internal text semantic information and external knowledge, respectively. However, they focus on different semantic levels and knowledge dimensions. A multi-head attention mechanism is used to fuse contextual semantic features and external knowledge base features to produce a comprehensive feature representation. This multi-head attention mechanism can capture the multi-dimensional correlations and interactions between features, highlighting the impact of key information on causal inference. The comprehensive feature representation is input into the prediction layer, where classification or regression calculations are performed to obtain the final prediction result of the event pair relationship. This aims to integrate the advantages of internal text semantics and external knowledge to form a more discriminative and explanatory feature representation, enhance the causal inference model's ability to comprehensively judge the causal relationship between event pairs, provide accurate prediction results for the final causal inference, and meet the requirements for accurate and reliable causal relationship identification in practical application scenarios.

[0072] Among them, contextual semantic features refer to the contextual semantic information of event pairs extracted from the text through a pre-trained model (such as the improved BERT). They reflect the lexical, grammatical, and semantic environment characteristics of the events in the text and are the key basis for the causal inference model to understand the internal semantics of the text. External knowledge base features refer to the features fused in step S4, which include coreference features and event semantics and causal relationship information in the external knowledge base, providing knowledge supplement and logical support beyond the text for the causal inference model. The multi-head attention mechanism is a variant of the attention model. By calculating the association between features in parallel through multiple attention heads, it can capture different semantic levels and interaction patterns between features, improving the effect of feature fusion and the expressiveness of features. The comprehensive feature representation is the final feature form after fusing contextual semantic features and external knowledge base features. It includes comprehensive information of internal and external knowledge of the text and is the direct feature basis for the causal inference model to perform causal inference. The prediction layer is the output layer of the model, which can include components such as classifiers or regressors. It is used to calculate the probability or strength of the causal relationship between event pairs based on the comprehensive feature representation and output the final prediction result. It is the key link for the causal inference model to transform features into practical application value.

[0073] Preferably, in step S1, the relationship is satisfied:

[0074]

[0075] Where X=[x1,x2,…,x i ,…x N ] represents the word sequence of the input text, x i represents the i-th word, E represents the event word set, N text Indicates the length of the word sequence of the text; P mask (x i) represents the probability of masking the i-th word, α and β are the masking probabilities, Represents the contextual semantic features of the i-th word, that is, the embedding vector of the event word.

[0076] Specifically, event words are first extracted from the text as masked target words, while non-event words are randomly masked. During this process, an improved BERT model (EM-Bert) is used for context reconstruction training to predict masked event words. This model can fully learn contextual semantic features, thereby obtaining more accurate contextual embedding representations.

[0077] In this embodiment, if Figure 2 As shown, the document-level text obtained is: Kenneth Dorsey says the woman accused of killing two co-workers and critically injuring a third at the Kraft plant in Northeast Philly is a good person. And so were the two women she's accused of gunning down with a.357 Magnum. When processing the input text, the causal inference method strictly controls the mask word ([MASK]) ratio to 15%, extracts event keywords through the label recognition mechanism and masks them, and adopts a random masking strategy to process the remaining non-event words. In this embodiment, the event words are killing, accused, injuring, and gunning down, which are replaced with [MASK] during the pre-training process. The priority masking strategy of event words enables the model to focus on the logical association between events and enhance the semantic feature extraction of event words for the context, while the random masking of non-event words enhances the model's generalization ability of text semantics while maintaining the mask ratio.

[0078] Preferably, step S2 includes:

[0079] By designing a contrastive loss function through contrastive learning strategy, the semantic distance between the cause event and its coreference event, and between the result event and its coreference event are shortened, while the semantic distance between non-coreference event pairs is shortened.

[0080] Construct positive sample co-reference pairs (e i ,e j )∈C, randomly combined from non-coreferential events A projection layer is introduced to map events into a new semantic space, and cosine similarity is used to calculate the contextual relevance of co-referencing event pairs, satisfying the relationship:

[0081]

[0082] Among them, W and b represent the weight matrix and bias of the projection layer respectively, represents the contextual semantic features of the i-th word, z i Denotes the embedding vector h i The projection vector of is the loss function of contrast training, s(z i ,s j ) represents z i With z j The similarity between them is τ, which represents the temperature coefficient.

[0083] Positive coreference pairs are coreference event pairs constructed from external knowledge base and original text events, for example (e i ,e j )∈C, where e i and e j are coreference events, and negative pairs are event pairs randomly combined from non-coreference events, e.g. where e i 、e k Instead of co-referencing events, the projection layer is a neural network layer that maps events to a new semantic space, performing linear transformation and nonlinear activation through the weight matrix W and bias b.

[0084] like Figure 3 As shown in the lower part of (building coreference sample pairs based on external knowledge base), first, the events in the text are formed into target event pairs (e i ,e k ), and divided them into cause events and result events, selecting killing-injuring as the target event pair, and then introducing the external knowledge base MRVEN-ERE, combining events with a coreference relationship as coreference positive sample pairs, and randomly combining events without a coreference relationship to obtain non-coreference negative sample pairs, where the coreference events of killing are murdering and demonstrate, and the coreference events of injuring are harming and damage. When the projection layer is introduced, the events are mapped to a new semantic space. Through the weight matrix W and bias b of the projection layer, the contextual semantic features of the events are obtained. is transformed into a new projection vector z i , To embed the vector, the sigmoid activation function is used to perform nonlinear processing on the features after linear transformation. This mapping method can enhance the model's ability to represent semantic features and improve the model's discrimination performance. When calculating similarity, cosine similarity is used to calculate the similarity between co-referencing event pairs. For example, when calculating z i 、z j The cosine similarity of the two vectors is used to evaluate their semantic relevance. The cosine similarity metric can more accurately reflect the directional similarity between vectors, improving the accuracy of similarity calculation. A contrastive loss function is then used to minimize the semantic distance of co-referencing event pairs while maximizing the semantic distance of non-co-referencing event pairs. For example, the temperature coefficient τ is used to control the distribution of similarity. The introduction of the contrastive loss function can effectively optimize the training process of the causal inference model, enabling it to better learn the semantic relationships between events and enhance its generalization ability when processing complex text.

[0085] Preferably, in step S3, the process includes extracting potential indirect events associated with the cause / result events from an external knowledge base, and identifying and extracting the causal relationships between the events using entity linking and relationship extraction techniques;

[0086] The causal relationship is serialized and then semantically encoded using the parameter-frozen EMBert model to satisfy the relationship:

[0087] Input=[CLS]e0[indirect]e1[indirect]…[indirect]e n [effect];

[0088]

[0089] Among them, Input represents the input sequence, [indirect] represents the indirect causal relationship between events, and [effect] represents the final result event of the causal chain. Indicates event e i The embedding vector in sequence k.

[0090] like Figure 3As shown in the upper part of the figure (constructing an external knowledge base causal chain), a causal graph containing coreference relationships is constructed using the external knowledge base CauseNet. In order to extract potential indirect events related to cause events and result events from the external knowledge base, the present invention specifically starts from the target cause event and gradually explores subsequent events related to the cause event by traversing the event network in the knowledge base. In each search step, the current event is regarded as a new cause event, and the downstream exploration continues until the target result event is successfully located. At the same time, the indirect causal relationships between coreference events are included, thereby constructing a complete causal relationship chain. Among them, the causal relationship chain is part of the causal graph, and each link is a path in the causal graph. For example, for the target event pair killing-injuring, the causal chain constructed by the external knowledge base CauseNet is as follows: [CLS]killing[indirect]protest[indirect]clash[indirect]injuring[effect], [CLS]demonstrate[indirect]protest[indirect]escalation[indirect]injury[effect]. Serialization is the process of converting causal links into linear sequences for model input and processing. Serialization typically introduces specific tags, such as [indirect] and [effect], to identify indirect causal relationships between events and the final outcome event of the causal chain. The parameter-frozen EMBert model, with fixed parameters after pre-training, is used to semantically encode the serialized causal links. Parameter freezing preserves model stability and the general semantic information learned during pre-training, while reducing computational overhead during training.

[0091] Preferably, in step S3, the process includes calculating the embedding vectors of the event nodes between the event nodes of the causal graph. and Combine the attention weights to fuse the weights of the same nodes and edges in different sequences to satisfy the relationship:

[0092]

[0093]

[0094] in, Represents event node e in sequence k i To event node e j The weight, W g is a learnable matrix, Represents event node e in sequence k i With e jThe attention weight between T is a learnable vector, is the weight between global event nodes.

[0095] Specifically, by calculating the embedding vector of the caret between event nodes in the causal graph and Combine the attention weights to fuse the weights of the same nodes and edges in different sequences. Specifically, first use the learnable matrix W g Perform a linear transformation on the embedding vector of the caret and the embedding vector of the final result event of the causal chain, and calculate the event node e in the sequence k through the LeakyReLU activation function i To event node e j Weight Then, by calculating the attention weight Perform weighted fusion on the weights between pairs of identical event nodes in different sequences to obtain the weights between global event nodes. This process aims to dynamically adjust the contribution of weights between event nodes in different sequences through the attention mechanism, so that the model can more accurately capture the causal relationship between events and improve the accuracy and robustness of causal inference. g The matrix used to linearly transform the embedding vector of the caret and the embedding vector of the final result event of the causal chain. It adjusts the representation of the vector through learning to better capture the weight relationship between event nodes; the LeakyReLU activation function is an activation function used to introduce nonlinear characteristics, enabling the model to learn more complex feature representations; attention weight Indicates that in sequence k, event node e i With E j The attention weights between them reflect the relative importance between the same event node pairs in different sequences.

[0096] Therefore, the causal inference model can more accurately calculate the weights between event nodes, dynamically adjust the contribution of the weights between event nodes in different sequences, and obtain the weight representation between global event nodes, so that the causal inference model performs well in processing complex text and multi-hop reasoning tasks, can be more reliably applied to actual causal analysis scenarios, and can provide strong support for causal inference tasks in the field of natural language processing.

[0097] Preferably, in step S3, mean pooling is used to fuse the features of the same event nodes in multiple different sequences into a unified global feature, satisfying the relationship:

[0098]

[0099] in Indicates event e iEmbedding vectors in the sequence, Indicates nodes with the same event e i the number of sequences, To fuse the global feature representation of each sequence event node.

[0100] Specifically, mean pooling is used to fuse the features of the same event nodes in multiple different sequences into a unified global feature. The global feature representation is obtained by calculating the average value of the embedding vectors of the same event node in multiple sequences. This mean pooling operation aims to integrate feature information from different sequences and reduce the dimension of the features. At the same time, it enhances the robustness and generalization ability of the causal inference model for event node features, providing higher quality input features for subsequent graph neural network processing.

[0101] Preferably, in step S3, the attention coefficient between each node and its adjacent nodes is calculated according to the graph attention mechanism, and the features of the adjacent nodes are fused. At the same time, the edge weight information is fused into the node features using a graph convolutional neural network to satisfy the relationship:

[0102]

[0103] in, Represents the attention coefficient between adjacent nodes, A is the attention coefficient matrix, H (0) Represents the node features of the initial layer, W GAT represents the learnable feature transformation matrix, H (1) Represents the features after the fusion of the first-layer graph attention mechanism, H (2) Represents the features after fusion of the second layer of graph convolutional neural network, represents the degree matrix, Represents the normalized adjacency matrix of the graph, W GCN Represents the weight between global event nodes, each element is

[0104] The attention coefficient between each node and its adjacent nodes is calculated through the graph attention mechanism (GAT), and the features of the adjacent nodes are fused. At the same time, the edge weight information is fused into the node features using the graph convolutional neural network (GCN). Specifically, the learnable feature transformation matrix W is first used. GAT For the node features H of the initial layer (0) Transform, and then calculate the attention coefficient between each node pair through the attention mechanism These attention coefficients can reflect the relative importance between nodes and are used to perform weighted fusion of the features of adjacent nodes to obtain the feature H after fusion of the first-layer graph attention mechanism. (1)Next, we use the graph convolutional neural network (GCN) to integrate the edge weight information into the node features, and then use the normalized adjacency matrix Sum degree matrix Propagate the graph structure information and obtain the feature H after the fusion of the second layer graph convolutional neural network (2) This process aims to fully utilize the structured semantic information of the causal graph through the feature fusion capability of the graph neural network, improve the causal inference model's ability to reason about complex causal logic between events, and enhance the accuracy and robustness of causal inference.

[0105] Preferably, step S4 satisfies the relationship:

[0106]

[0107] f=[y1;y2;…;y n ];

[0108] in Represents the node features of the structured semantic graph, z i Represents the coreference feature of the event node, y i It represents the fusion result of the node features and coreference features of the structured semantic graph, and f represents the feature concatenation vector of each coreference event of the cause event or result event.

[0109] In step S4, in order to fuse the node features extracted in step S3 with the co-reference features extracted in step S2, a multi-layer perceptron (MLP) is used to perform nonlinear transformation on the features of each event node. Specifically, the node features of the structured semantic graph are transformed into and the coreference feature z of the event node i Splicing is performed to form a joint feature vector, and then nonlinear transformation is performed through MLP to obtain the fused feature representation y i Finally, the fusion features y of all event nodes are i The nodes and co-references of the structured semantic graph are combined to form a feature vector f, which represents the feature concatenation vector of each co-reference event of the cause event or result event. This process aims to integrate the node features and co-reference features of the structured semantic graph to form a more comprehensive and richer feature representation, providing higher quality input features for subsequent causal inference. The fusion of the node features and co-reference features of the structured semantic graph makes the feature representation more comprehensive and richer, which can better reflect the semantic information and co-reference relationship of the event nodes; and through the nonlinear transformation of the multi-layer perceptron (MLP), it can more accurately capture the complex relationship between features and improve the effect of feature fusion; the fused feature representation y i The feature concatenation vector f can more effectively support subsequent causal inference tasks and improve the accuracy and robustness of the causal inference model.

[0110] Preferably, step S5 satisfies the relationship:

[0111]

[0112] M=[head1;…;head h ]W O ;

[0113] Among them, C cause with C effect Respectively represent the concatenated cause event and result event feature vectors, Wi i Q , Represents the weight matrices of Q, K, and V, respectively. O is the weight matrix of the multi-head attention output mapping, M is the comprehensive feature representation of the cause event and the result event, The contextual semantic features of the cause event in the text, Represents the contextual semantic features of the result event in the text, f cause The external knowledge base features representing the cause event, f effect Represents the external knowledge base features of the result event.

[0114] In step S5, in order to fuse the feature vectors of the cause event and the result event, a multi-head attention mechanism is used. Specifically, the contextual semantic features of the cause event are first With the external knowledge base feature f cause Splice into feature vector C cause , the contextual semantic features of the result event With the external knowledge base feature f effect Splice into feature vector C effect Then, the outputs of multiple attention heads are calculated through the multi-head attention mechanism, and these outputs are concatenated and passed through an output mapping matrix W O The final comprehensive feature representation M is obtained. This process aims to capture the various interactive relationships between cause events and result events through the multi-head attention mechanism, and improve the causal inference model's ability to comprehensively analyze the causal relationship between events. i Q , The query, key, and value weight matrices represent the i-th attention head, respectively, and are used to linearly transform feature vectors into different representation spaces. By concatenating contextual semantic features and external knowledge base features, information from the text and external knowledge bases can be integrated, making feature representation more comprehensive and richer. The multi-head attention mechanism can capture the various interactions between cause and effect events, improving the causal inference model's ability to understand causal relationships. The nonlinear transformations of the multi-head attention mechanism can more accurately express complex relationships between features, improving the causal inference model's feature expression capabilities and the accuracy of causal inference.

[0115] Furthermore, step S5 includes inputting the comprehensive feature representation into the prediction layer, determining whether there is a causal relationship between the event pairs, and calculating the loss, to obtain the final prediction result of the event pair relationship, satisfying the relationship:

[0116] p i =sigmoid(WM+b);

[0117]

[0118] Among them, p i Expressed as predicted probability, y i is the true label, N is the number of samples of event pairs, Expressed as a binary cross entropy loss function, is the comprehensive loss, M is the comprehensive feature representation of the cause event and the result event, b is the bias of the projection layer, and λ is the weight coefficient of the loss function. represents the loss function for contrastive training.

[0119] In step S5, the comprehensive feature representation M is input to the prediction layer, which determines whether there is a causal relationship between the event pairs and calculates the corresponding loss function to optimize the model. Specifically, the prediction layer uses the sigmoid() function to map the comprehensive feature representation to the probability value p i , indicating the possibility of a causal relationship between event pairs, through the binary cross entropy loss function Measure the difference between the predicted probability and the true label, and combine it with the loss function of contrastive learning Constructing comprehensive losses The purpose of this process is to comprehensively optimize the predictive ability and comparative learning ability of the causal inference model, and to improve the model's recognition accuracy and generalization ability of causal relationships by minimizing the comprehensive loss.

[0120] To this end, the binary cross entropy loss function And the loss function of contrast training The causal inference model can more accurately predict the causal relationship between event pairs and improve the prediction accuracy. Through the optimization of comprehensive loss, the causal inference model not only learns the ability to predict causal relationships during training, but also enhances the sensitivity to the semantic similarity between event pairs and improves the generalization ability. The introduction of enables the causal inference model to distinguish between coreference event pairs and non-coreference event pairs, which can enhance the model's discriminative ability and help improve robustness.

[0121] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative uses of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0122] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A causal inference method that integrates coreference relationships and external knowledge enhancement, characterized by: The causal inference method comprises the following steps: S1: Event words are extracted from document-level text as masked target words and non-event words are randomly masked. The improved BERT model is used to predict masked event words through context reconstruction training in the pre-training phase, so that event words can fully learn contextual semantic features. S2: Randomly extract event pairs from the text and label them as cause events and result events, extract coreference event pairs between the cause events and the result events, obtain labeled coreference event positive samples from an external knowledge base, and simultaneously generate non-coreference event pairs as negative samples. Utilize a contrastive learning mechanism to design a contrastive loss function to shorten the semantic distance between the coreference event pairs, and obtain the coreference features of the coreference events. S3: Extract potential indirect events associated with cause / result events from the external knowledge base to construct an external knowledge base causal graph. A structured semantic graph is constructed with events in the external knowledge base causal graph as nodes and causal relationships as edges. A graph attention mechanism and a graph convolutional neural network are used to fuse node features. The potential indirect events include indirect events between cause events and result events. S4: Fuse the coreference features in step S2 and the node features fused in step S3, and perform nonlinear transformation through a fully connected layer to obtain the external knowledge base features of the cause event and the result event; S5: A multi-head attention mechanism is used to fuse the contextual semantic features and the external knowledge base features to obtain a comprehensive feature representation, which is input into the prediction layer to obtain the final prediction result of the event pair relationship.

2. The causal inference method according to claim 1, characterized in that: In step S1, the relationship is satisfied: Where X=[x1,x2,…,x i ,…x N ] represents the word sequence of the input text, x i represents the i-th word, E represents the event word set, N text Indicates the length of the word sequence of the text; P mask (x i ) represents the probability of masking the i-th word, α and β are the masking probabilities, Represents the contextual semantic features of the i-th word, that is, the embedding vector of the event word.

3. The causal inference method according to claim 1, characterized in that: In step S2, it includes: By designing a contrastive loss function through contrastive learning strategy, the semantic distance between the cause event and its coreference event, and between the result event and its coreference event are shortened, while the semantic distance between non-coreference event pairs is shortened. Construct positive sample co-reference pairs (e i ,e j )∈C, randomly combined from non-coreferential events A projection layer is introduced to map events into a new semantic space, and cosine similarity is used to calculate the contextual relevance of co-referencing event pairs, satisfying the relationship: Among them, W and b represent the weight matrix and bias of the projection layer respectively, represents the contextual semantic features of the i-th word, z i Denotes the embedding vector h i The projection vector of is the loss function of contrast training, s(z i ,z j ) represents z i With z j The similarity between them is τ, which represents the temperature coefficient.

4. The causal inference method according to claim 1, characterized in that: In step S3, potential indirect events associated with cause / effect events are extracted from the external knowledge base, and the causal relationships between the events are identified and extracted using entity linking and relationship extraction techniques; The causal relationship is serialized and then semantically encoded using the parameter-frozen EMBert model to satisfy the relationship: Input=[CLS]e0[indirect]e1[indirect]…[indirect]e n [effect]; Among them, Input represents the input sequence, [indirect] represents the indirect causal relationship between events, and [effect] represents the final result event of the causal chain. Indicates event e i The embedding vector in sequence k.

5. The causal inference method according to claim 4, characterized in that: In step S3, the embedding vectors of the event nodes between the event nodes are calculated by and Combine the attention weights to fuse the weights of the same nodes and edges in different sequences to satisfy the relationship: in, Represents event node e in sequence k i To event node e j The weight, W g is a learnable matrix, Represents event node e in sequence k i With e j The attention weight between T is a learnable vector, is the weight between global event nodes.

6. The causal inference method according to claim 5, characterized in that: In step S3, mean pooling is used to fuse the features of the same event nodes in multiple different sequences into a unified global feature, satisfying the relationship: in Indicates event e i Embedding vectors in the sequence, Indicates nodes with the same event e i the number of sequences, To fuse the global feature representation of each sequence event node.

7. The causal inference method according to claim 6, characterized in that: In step S3, the attention coefficient between each node and its adjacent nodes is calculated according to the graph attention mechanism, and the features of the adjacent nodes are fused. At the same time, the edge weight information is fused into the node features using the graph convolutional neural network to satisfy the relationship: in, Represents the attention coefficient between adjacent nodes, A is the attention coefficient matrix, H (0) Represents the node features of the initial layer, W GAT represents the learnable feature transformation matrix, H (1) Represents the features after the fusion of the first-layer graph attention mechanism, H (2) Represents the features after fusion of the second layer of graph convolutional neural network, represents the degree matrix, Represents the normalized adjacency matrix of the graph, W GCN Represents the weight between global event nodes, each element is 8. The causal inference method according to claim 1, characterized in that: Step S4 satisfies the relationship: f=[y1;y2;…;y n ]; in Represents the node features of the structured semantic graph, z i Represents the coreference feature of the event node, y i It represents the fusion result of the node features and coreference features of the structured semantic graph, and f represents the feature concatenation vector of each coreference event of the cause event or result event.

9. The causal inference method according to claim 1, characterized in that: Step S5 satisfies the relationship: M=[head1;…;head h ]W O ; Among them, C cause with C effect Respectively represent the concatenated cause event and result event feature vectors, Represents the weight matrices of Q, K, and V, respectively. O is the weight matrix of the multi-head attention output mapping, M is the comprehensive feature representation of the cause event and the result event, The contextual semantic features of the cause event in the text, Represents the contextual semantic features of the result event in the text, f cause The external knowledge base features representing the cause event, f effect Represents the external knowledge base features of the result event.

10. The causal inference method according to claim 9, characterized in that: Step S5 includes inputting the comprehensive feature representation into the prediction layer, determining whether there is a causal relationship between the event pairs, and calculating their losses to obtain the final prediction result of the event pair relationship, which satisfies the relationship formula: p i =sigmoid(WM+b); Among them, p i Expressed as predicted probability, y i is the true label, N is the number of samples of event pairs, Expressed as a binary cross entropy loss function, is the comprehensive loss, M is the comprehensive feature representation of the cause event and the result event, b is the bias of the projection layer, and λ is the weight coefficient of the loss function. represents the loss function for contrastive training.

Citation Information

Patent Citations

  • Financial event-oriented hybrid causal relationship discovery method

    CN111026852A

  • Sentiment classification model and text sentiment analysis method applied by same

    CN117708328A

  • Event causal relationship identification method fusing argument and structural information

    CN117993508A

  • Dangerous chemical accident dynamic affair map construction method based on multi-source data fusion

    CN119129712A

  • Event knowledge pre-training enhanced cross-language event causal relationship identification method and device

    CN119179780A

Cited By

  • Text-oriented violent word abbreviation detection method and device, equipment and storage medium

    CN121212140A