Event causal relationship identification method based on iterative graph prompt learning

By constructing and optimizing causal graphs through iterative graph prompting learning, the problem of insufficient directionality and structured reasoning in causal relationship identification in existing technologies is solved, and highly accurate and robust event causal relationship identification is achieved.

CN121303302APending Publication Date: 2026-01-09HUAZHONG NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511332614.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing methods for identifying causal relationships are inadequate in terms of directional discrimination, structured reasoning, and activation of implicit knowledge, making it difficult to effectively identify causal directions between events and cross-sentence causal chains.

Method used

An iterative graph prompting learning approach is adopted to construct an initial causal graph of certainty. Causal relationships are generated by inputting multi-path information into a generative language model. The graph structure constraint mechanism is used for multiple rounds of iterative optimization, dynamically selecting edges to update the causal graph until the iteration termination condition is met.

Benefits of technology

It improves the accuracy and structural integrity of event causal relationship identification, is applicable to complex text and cross-sentence causal chain scenarios, has causal direction modeling capabilities and multi-path semantic fusion capabilities, and significantly improves the performance of identification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121303302A_ABST
    Figure CN121303302A_ABST
Patent Text Reader

Abstract

The invention discloses an event causal relationship identification method based on iterative graph prompt learning, and belongs to the technical field of natural language processing. Comprising the following steps: constructing an initial definite causal graph; based on the initial definite causal graph, performing context modeling on a to-be-recognized text, extracting multi-path information of each event pair in the to-be-recognized text, and inputting the multi-path information into a generative language model to perform causal relationship generation to obtain an initial causal relationship result; based on the initial causal relationship result, adopting a graph structure constraint mechanism to carry out multiple rounds of iterative optimization on the initial definite causal graph, dynamically selecting an edge according to confidence, updating the initial definite causal graph, judging whether an iteration termination condition is met or not, and stopping iteration when the iteration termination condition is met to obtain an optimized definite causal graph; and obtaining an event causal relationship result of the to-be-identified text based on the optimized definite causal graph. According to the method, the event causal relationship identification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of natural language processing technology, and in particular relates to an event causal relationship recognition method based on iterative graph prompting learning. Background Technology

[0002] Most current mainstream causal relationship identification methods rely on causal graphs to construct models for text classification, typically simplifying the task into a binary classification problem of determining whether two events are causally related. These methods neglect the directionality of causal relationships themselves, failing to effectively identify the essential difference between "event A causes event B" and "event B causes event A," thus limiting the model's ability to express causal semantics. Furthermore, these methods generally focus on intra-sentence event pairs, lacking systematic modeling of document-level cross-sentence causal chains, resulting in limited performance when dealing with complex, multi-hop, and non-explicit causal logic in real-world texts. How to construct a unified framework with clear causal direction discrimination capabilities, integrating event context and graph structure information, and supporting multi-turn causal reasoning and knowledge activation has become a core problem urgently needing breakthroughs in current research on document-level event causal relationship identification.

[0003] To alleviate these problems, some studies have introduced graph neural networks, treating events as nodes in a graph and semantic relationships between events as edges, modeling the dependencies between events in text through graph structures. However, these methods typically construct static graph structures, lacking directional modeling of edges and dynamic evolution mechanisms, making it difficult to adapt to the hierarchical expansion and semantic drift of causal chains in semantic scenarios. Furthermore, the construction of static graphs often relies on rule templates or external resources, lacking task adaptability and failing to fully utilize the latent causal knowledge inherent in pre-trained models.

[0004] Therefore, existing causal relationship identification methods still have significant shortcomings in terms of directionality discrimination, structured reasoning, and implicit knowledge activation. Existing technologies suffer from difficulties in expressing structural relationships between events, weak causal directionality modeling, and the inability of generated results to effectively support structural optimization. Summary of the Invention

[0005] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes an event causality recognition method based on iterative graph cue learning, which improves the accuracy of event causality recognition.

[0006] In a first aspect, this application provides a method for identifying event causal relationships based on iterative graph cue learning, the method comprising: Construct an initial confidence causal graph; Based on the initial certainty causal graph, context modeling is performed on the text to be identified, multi-path information of each event pair in the text to be identified is extracted, and the multi-path information is input into the generative language model to generate causal relationships and obtain the initial causal relationship results. Based on the initial causal relationship results, a graph structure constraint mechanism is used to perform multiple rounds of iterative optimization on the initial confident causal graph. Edges are dynamically selected according to the confidence level, the initial confident causal graph is updated, and it is determined whether the iteration termination condition is met. When the iteration termination condition is met, the iteration stops, and the optimized confident causal graph is obtained. Based on the optimized conviction causal graph, the event causal relationship results of the text to be identified are obtained.

[0007] According to one embodiment of this application, constructing the initial confidence causal graph includes: The text to be identified is input into the trained causal graph construction model. Through the causal directionality prompt template, the event causality of the event pairs in the text to be identified is determined, and the causal relationship prediction results between the event pairs in the text to be identified are obtained. Based on the causal relationship prediction results between event pairs in the text to be identified, a set of causal edges is formed by filtering credible edges through a set confidence threshold, and an initial confident causal graph about the text to be identified is constructed.

[0008] According to one embodiment of this application, the multi-path information includes event context information, graph structure information, and document global information. The event context information includes the semantics of the preceding and following sentences between event pairs. The graph structure information includes the adjacency structure between event pairs in the causal graph. The document global information includes the overall semantic encoding of the text to be identified.

[0009] According to one embodiment of this application, the step of inputting multi-path information into a generative language model to generate causal relationships and obtain initial causal relationship results includes: The event context information is converted into context text encoding and input into the context encoder to obtain local semantic features; The graph structure information is converted into causal graph text encoding and input into a causal graph encoder to obtain structure-aware features. The global information of the document is converted into document text encoding and input into the document encoder to obtain global semantic features; The hidden state vector is obtained by fusing local semantic features, structure-aware features, and global semantic features. The hidden state vector is input into the decoder, and the causal features are dynamically extracted through the sequence generation capability of the decoder to obtain the initial causal relationship results.

[0010] According to one embodiment of this application, the step of performing multi-round iterative optimization of the initial confident causal graph based on the initial causal relationship result using a graph structure constraint mechanism, dynamically selecting edges according to the confidence level, updating the initial confident causal graph, and determining whether the iteration termination condition is met, stopping the iteration when the iteration termination condition is met, and obtaining the optimized confident causal graph includes: Based on the initial causal relationship results, event pairs with causal confidence greater than or equal to the current dynamic threshold are extracted and transformed into corresponding directed causal edges, which are then added as new credible edges and added to the causal edge set. The confidence level of the causal edge set in the initial confidence causal graph is filtered, and all credible edges with a causal confidence level lower than the current dynamic threshold are removed. The credible edges with a confidence level greater than or equal to the current dynamic threshold are retained as credible edges. A dynamic filtering mechanism is adopted to adjust the dynamic threshold. Based on the adjusted dynamic threshold, the retained trusted edges and the newly added trusted edges are integrated to update the initial confidence causal graph. The updated confidence causal graph is obtained and it is determined whether the iteration termination condition is met. When the iteration termination condition is met, the iteration stops and the optimized confidence causal graph is obtained.

[0011] According to one embodiment of this application, the causal graph construction model is the RoBERTa model, and the generative language model is the BART generative language model.

[0012] According to one embodiment of this application, adjusting the dynamic threshold using a dynamic filtering mechanism includes: Determine whether the causal confidence of all event pairs is lower than the current dynamic threshold. If the causal confidence of all event pairs is lower than the current dynamic threshold, decrease the dynamic threshold by a preset step size and adjust the dynamic threshold until the minimum threshold is reached. If there is a causal confidence of all event pairs that is higher than the current dynamic threshold, keep the current dynamic threshold unchanged.

[0013] Secondly, this application provides an event causal relationship recognition system based on iterative graph cueing learning, the system comprising: The building block is used to construct the initial confidence causal graph; The causal relationship generation module is used to perform contextual modeling on the text to be identified based on the initial certainty causal graph, extract multi-path information of each event pair in the text to be identified, input the multi-path information into the generative language model to generate causal relationships, and obtain the initial causal relationship results. The causal graph optimization module is used to perform multiple rounds of iterative optimization on the initial confident causal graph based on the initial causal relationship results and using a graph structure constraint mechanism. It dynamically selects edges according to the confidence level, updates the initial confident causal graph, and determines whether the iteration termination condition is met. When the iteration termination condition is met, the iteration stops, and the optimized confident causal graph is obtained. The causal relationship identification module is used to obtain the event causal relationship results of the text to be identified based on the optimized confidence causal graph.

[0014] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the event causality recognition method based on iterative graph cueing learning as described in the first aspect above.

[0015] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the event causality recognition method based on iterative graph cueing learning as described in the first aspect above.

[0016] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the event causal relationship recognition method based on iterative graph cueing learning as described in the first aspect.

[0017] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the event causality recognition method based on iterative graph cueing learning as described in the first aspect above.

[0018] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application.

[0019] The present invention provides an event causal relationship identification method based on iterative graph cue learning, which has the following advantages over the prior art: By introducing causal directionality prompt templates, multi-path graph structure modeling, and iterative optimization strategies, the accuracy, structural integrity, and multi-hop reasoning capabilities of event causality recognition are effectively improved. Without relying on a fixed answer space, semantic fusion and iterative calibration guided by graph structure achieve high-precision, directional, and well-defined event causality recognition for complex text scenarios. Fully leveraging the causal knowledge embedded in the causal graph construction model, combined with the expressive power of graph structure and the contextual modeling advantages of generative language models, the accuracy and robustness of causal directionality recognition and cross-sentence causal chain modeling are enhanced. Suitable for document-level event causal analysis scenarios, by integrating the advantages of prompt learning and graph modeling, it possesses causal direction modeling capabilities, multi-path semantic fusion capabilities, and causal graph structure iterative optimization capabilities, significantly improving the performance of causal recognition tasks in complex text, long documents, and cross-sentence causal chain scenarios. Applicable to various application scenarios such as causal chain construction, document-level event understanding, and causal knowledge graph generation. Attached Figure Description

[0020] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart illustrating the event causal relationship identification method based on iterative graph cueing learning provided in the embodiments of this application; Figure 2 This is a schematic diagram of the prompting learning and causal graph construction process of the event causal relationship identification method provided in the embodiments of this application; Figure 3 This is a flowchart of graph-driven multi-path semantic modeling for the event causality identification method provided in the embodiments of this application; Figure 4 This is a graph showing the performance changes of the event causality identification method provided in this application during the iterative optimization process of graph structure; Figure 5 This is a schematic diagram of the structure of the event causal relationship recognition system based on iterative graph cue learning provided in the embodiments of this application; Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0022] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0023] The following description, in conjunction with the accompanying drawings, details the event causality recognition method based on iterative graph cue learning, the event causality recognition system based on iterative graph cue learning, the electronic device, and the readable storage medium provided in this application, through specific embodiments and application scenarios.

[0024] Among them, the event causal relationship recognition method based on iterative graph prompting learning can be applied to the terminal, specifically executed by the hardware or software in the terminal.

[0025] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).

[0026] The following embodiments describe a terminal including a display and a touch-sensitive surface. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and joystick.

[0027] The event causality recognition method based on iterative graph cue learning provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can implement the event causality recognition method based on iterative graph cue learning. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras, and wearable devices. The event causality recognition method based on iterative graph cue learning provided in this application embodiment will be described below using an electronic device as the execution subject.

[0028] Event causality identification is a crucial task in the field of natural language processing. Its goal is to identify causal relationships between events from unstructured text and further determine the direction of causality, distinguishing between "cause events" and "effect events." This technology is widely applied in practical scenarios such as medical knowledge discovery, financial information analysis, legal case reasoning, and educational feedback modeling. It serves as a vital foundation for constructing causal knowledge graphs, achieving textual logical understanding, and supporting intelligent decision-making.

[0029] Hint learning, a novel paradigm that has emerged in recent years, utilizes natural language templates to activate the implicit knowledge within causal graphs and is widely applied to few-shot learning and knowledge mining tasks. In causal recognition scenarios, some methods attempt to transform event causal judgments into cloze tests, guiding the model's causal reasoning through the design of hint templates and candidate answer words. However, existing hint learning methods are mostly static templates, lacking integration with graph structures and multi-round reasoning mechanisms. This makes it difficult to fully activate the deep causal knowledge implicit in large models and to dynamically utilize the model's intermediate outputs for structural evolution and optimization.

[0030] Figure 1 This is a flowchart illustrating the event causality recognition method based on iterative graph cueing learning provided in this application embodiment, as shown below. Figure 1As shown, the event causal relationship identification method based on iterative graph prompting learning includes steps 110, 120, 130 and 140.

[0031] Step 110: Construct an initial confidence causal graph; In some embodiments, constructing the initial confidence causal graph includes: The text to be identified is input into the trained causal graph construction model. Through the causal directionality prompt template, the event causality of the event pairs in the text to be identified is determined, and the causal relationship prediction results between the event pairs in the text to be identified are obtained. Based on the causal relationship prediction results between event pairs in the text to be identified, a set of causal edges is formed by filtering credible edges through a set confidence threshold, and an initial confident causal graph about the text to be identified is constructed. The initial confident causal graph is a set of causal edges containing credible edges.

[0032] It should be noted that the text to be identified contains multiple discrete events (such as "earthquake", "building collapse", "people evacuation"). First, these events are paired into "event pairs" (such as <earthquake, building collapse>, <building collapse, people evacuation>). All event pairs together constitute the set of processing objects for the causal graph construction model.

[0033] The process involves constructing a model from the event pairs in the text to be identified into a causal graph. By using a designed causal direction prompt template, the task of determining the causal relationship of an event is transformed into a fill-in-the-blank task. This generates causal relationship prediction results, and based on a set confidence threshold, credible edges are filtered to construct an initial confidence causal graph. In this graph, nodes represent events, edges represent predicted causal relationships, and edge weights represent causal confidence.

[0034] In some embodiments, the event pairs of the text to be identified are embedded in a directional cue template, input into a causal graph model, causal prediction results are obtained, and a high-confidence initial causal graph is constructed by filtering through a confidence threshold. This graph includes event nodes, directed causal edges, and edge weight confidence. Causal pairs below a set threshold are filtered out to ensure the accuracy and reliability of the graph structure.

[0035] For example, the prompt template adopts an event-oriented prompt format, specifically including "Event A [MASK] Event B", where the words in [MASK] are limited to "Event A [MASK] Event B". <causes>"、"<Caused by> "or" <na>"Three options; the prompt template uses a structured text format with directional guidance; the answer space corresponding to the mask position includes..." <causes>"、"<Caused by> "and" <na>", " respectively represent forward, reverse, and no causal relationship, which can explicitly model the causal direction information between events.

[0036] It is worth noting that the initial confidence causal graph is constructed based on a uniform confidence threshold. When the prediction confidence of all causal edges is lower than this threshold, the causal edge with the highest confidence is selected and added to the graph structure by the maximum probability retention rule to avoid generating an empty graph.

[0037] Each event pair in the text to be identified ( , The causal relationship prediction results include two types of key confidence scores: the confidence score that event Ei causes event Ej and the confidence score that event Ej causes event Ei. Both types of confidence scores are predicted and output by the causal graph model through causal directionality cue templates.

[0038] Based on a unified confidence threshold γ, a globally unified confidence threshold γ (γ∈(0,1)) is set for each event pair ( , The two types of confidence levels are judged separately.

[0039] Empty graph compensation involves counting the number of edges in the set of credible edges. If the number of edges is 0, the maximum causal confidence of each event pair is calculated, and the event pair with the highest maximum causal confidence is added to the set of credible edges to ensure that the initial causal graph is not empty.

[0040] The final construction of the initial confidence causal graph is achieved by using all events in the text to be identified as the set of nodes in the graph, and the set of directed edges in the graph after filtering the credible edges. The edge weights are the confidence levels of the corresponding causal relationships, thus forming the initial confidence causal graph for the text to be identified.

[0041] In some embodiments, Figure 2 This is a schematic diagram illustrating the prompting learning and causal graph construction process of the event causal relationship identification method provided in this application embodiment, such as... Figure 2 As shown, the event causal identification task is transformed into a fill-in-the-blank prediction task that can be processed by a language model. High-confidence causal relationship edges are then selected using confidence threshold constraints, ultimately generating a structured initial confidence causal graph. This provides highly reliable semantic support for subsequent inference. The construction of the initial confidence causal graph specifically includes the following steps: First, to transform the event causal identification task into a fill-in-the-blank task suitable for language models, the input text needs to be rewritten with prompts. Based on the prompt learning paradigm, by using a unified event representation format, designing standardized prompt templates, and constructing a clear causal relationship answer space, the pre-trained model is guided to focus on the semantic associations of events, thereby improving the accuracy and efficiency of causal relationship classification. By transforming the task into a fill-in-the-blank format that language models excel at, prompt learning can effectively utilize the semantic representation capabilities of the pre-trained model to achieve the classification objective.

[0042] Specifically, to enhance the model's ability to model semantic relationships between events, all event mentions in the document are first uniformly replaced with position-sensitive virtual tokens. <ei>(i is the event index), and for each virtual token <ei>Using event word tokens for average initialization allows the model to focus on the logical relationships between events, mitigating interference from text surface features.

[0043] To reduce the risk of semantic shift in causal orientation discrimination by language models, a virtual answer space is further constructed, including: <cause>: Indicates that event Ei is the cause and Ej is the effect; <Caused by> : indicates that Ej is the cause and Ei is the effect; <na>: Indicates that there is no causal relationship between the events.

[0044] This answer space serves as a constraint option for the language model's fill-in-the-blank output, thus transforming the causal discrimination task into a three-class classification problem: forward causality, backward causality, and no causal relationship.

[0045] To enable the model to effectively identify semantic relationships between events, a prompt template is designed to explicitly annotate the contextual information of event pairs to establish semantic associations between events. Specifically, for event pairs co-occurring in the same sentence, the template enhances the use of local contextual information; while for cross-sentence event pairs, multiple sentences are concatenated to retain necessary global semantic clues. The prompt template includes the event pair text, the sentence containing the event, and the location of the causal relationship to be predicted, supporting both intra-sentence and cross-sentence inference modalities, as detailed below: The prompt template is divided into two categories: The formula for intra-sentence events on the template is shown below:

[0046] The formula for the cross-sentence event template is shown below:

[0047] in, , Representing event pairs < > and < The original text containing the sentence. Intra-sentence reasoning, For cross-sentence reasoning, Token is an aggregate identifier for the global semantics of the sequence. Token, a separator for semantic units in a sequence. Placeholders for filling in blanks to predict causal relationships.

[0048] To achieve the aforementioned causal fill-in-the-blank prediction task, a causal relationship classifier is constructed based on the RoBERTa model. A masked language modeling mechanism is used to perform semantic prediction of the causal relationship between event pairs. Specifically, a prompt template containing the [MASK] placeholder is input into the RoBERTa model, and a word-level hidden state sequence is generated through a multi-layer Transformer encoder. The hidden state vector corresponding to the [MASK] position is... Incorporating contextual semantic information Projecting onto a 3D category space, the probability distribution of candidate answers is obtained through a softmax activation function. The specific process can be represented as follows:

[0049]

[0050] in, This is the weight matrix. For bias terms, For linear transformation layer, The probability distribution of the answer. This is the intermediate feature output of the model.

[0051] By predefined candidate answer space The task of determining causal relationships is transformed into a three-class classification problem based on mask positions. By utilizing the deep semantic modeling capabilities of the causal graph model, the semantic dependencies between event pairs can be effectively captured, enabling explicit discrimination of causal relationships and their directions.

[0052] Secondly, a causal relationship confidence screening mechanism is designed. To construct a high-quality initial confidence causal graph, the confidence of the causal relationship between each pair of events output by the language model is screened. In the initial confidence causal graph construction stage, a unified probability threshold strategy is adopted to achieve automatic screening of event pair causal relationships.

[0053] Specifically, for any event pair ( , First, extract the confidence scores of the two types of causal relationships from the model output. Indicates an event Caused the incident The probability, Indicates an event Caused the incident The probability of.

[0054] By setting a globally unified confidence threshold ∈ (0, 1), when the following conditions are met... or When an event pair is determined to have a corresponding causal relationship, a trustworthy edge is added to the graph structure. A trustworthy edge is defined as:

[0055]

[0056] Trusted edges Add to the initial confidence causal graph Among them For a set of event nodes, It is a set of causal edges.

[0057] Finally, the confidence graph is constructed based on the initial confidence causal graph. Considering the extreme cases that a uniform threshold might lead to, some documents generate an incorrect number of edges because the probability of event pairs is generally lower than the threshold. The "empty graph" affects subsequent graph structure analysis and model training; therefore, a targeted compensation mechanism was designed. This addresses the issue of the initial number of edges in a causal graph. The document executes the "maximum probability retention rule," iterating through all possible event pairs and comparing their... <cause>and<Caused by> The predicted probability of a relationship is used to select the causal edge with the highest probability and add it to the graph.

[0058] It should be noted that a directed edge is a connection with a directional attribute in a graph structure, used to represent a unidirectional association between two nodes (such as events). The causal edge set is the set of all directed edges with causal relationships in a document-level causal graph, denoted as L, corresponding to the edge components in the causal graph (G=(V,L). A causal relationship edge is a directed edge with a causal relationship between a single event pair, and is the basic building block of the causal edge set, which can be regarded as an instantiation of directed edges in a causal scenario. A reliable edge is a directed edge in a causal relationship whose model prediction confidence is higher than a set threshold; it is a general term for highly reliable causal relationship edges.

[0059] For example, from trusted edges to a causal graph, directed edges, the set of causal edges, causal relationship edges, and trusted edges have a clear dependency chain in the causal graph construction process, as follows: Initial trusted edge generation: Based on the prompt learning template of the RoBERTa model, the causal probability of event pairs is predicted, and causal relationship edges higher than the initial threshold are selected as initial trusted edges, forming the initial set of causal edges for the trusted graph. Causal edge set expansion: In the iteration phase, potential causal relationship edges are mined through BART generative model inference. If their predicted probability is higher than the current dynamic threshold, they are upgraded to trusted edges and added to the causal edge set, so that L gradually expands from the initial L0 to L1, L2, ..., Lt.

[0060] Cause-effect graph formation: The final set of causal edges L of the cause-effect graph consists of all reliable edges in the iteration process. The causal relationship edges contained therein are all highly reliable connections that have been screened by multiple rounds of thresholds, ensuring the cause-effect graph's ability to parse the causal logic of the document.

[0061] Model training and loss function setting: During the training phase, the RoBERTa-base model is used as the basic architecture. The Focal Loss function is used to optimize the prediction results to alleviate the causal class imbalance problem, and an L2 regularization term is introduced to suppress overfitting.

[0062] Let the total number of training samples be N, and the number of predicted categories be C=3 (corresponding to the answer space). <na> 、 <cause>,<Caused by> For the nth sample, its true label For a one-hot vector, the predicted probability distribution is as follows: The first loss function is defined as follows:

[0063] in, This is a category weight balancing factor used to adjust the weights of different categories, which can alleviate the category imbalance problem. As an adjustment factor, it controls the degree of decay of loss for easily classified samples. ≥0, The larger the value, the more significant the weight decay of easily classified samples. This represents the predicted probability of the nth sample in the true class, where N is the number of samples. This is the first loss function.

[0064] The formula for calculating the second loss function, which consists of L2 regularization terms, is shown below:

[0065] The final loss function is the sum of the two: , The sum of squares of all parameters, The regularization coefficient is . This is the second loss function.

[0066] During training, the AdamW optimizer was used to dynamically adjust the learning rate, and a gradient accumulation compensation training strategy was employed. The batch size was set to 1, and by setting the gradient accumulation step count to 16, the effective batch size was increased to 16 (aligned with model training), ensuring parameter update stability. The initial learning rate was set to... L2 regularization coefficient =0.01.

[0067] The above process completes the construction of the initial causal graph, providing a high-confidence structural input for subsequent causal relationship generation and graph structure iteration.

[0068] In some embodiments, experiments are conducted based on the EventStoryLine dataset, wherein: The EventStoryLine dataset contains English documents on 22 topics, covering 5156 event mentions and 70579 event relationship annotations. Event causal relationship annotations include three types: NONE (no causal relationship, i.e.) ), PRECONDITION (cause leading to effect, i.e.) <cause>) and FALLING_ACTION (results leading to causes, i.e. <causedby>The ratio of positive to negative samples (causal pairs / non-causal pairs) was 1:12; the ratio of causes leading to results and results leading to causes was 1:1.08; the ratio of intra-sentence event pairs to inter-sentence event pairs was 1:5.8; and the ratio of intra-sentence causal relationships (event pairs within the same sentence) to inter-sentence causal relationships (event pairs across sentences) was approximately 1:2.1. The experimental data partitioning followed the original design.

[0069] For example, the causal graph construction model uses RoBERTa-base as the backbone model for initial confidence causal graph construction, with word embedding dimension (d=768); the generative language model uses the BART-base architecture, which supports cross-modal semantic fusion and causal relationship generation.

[0070] In the initial confident causal graph construction phase, structured prior knowledge is generated using the RoBERTa-based causal graph construction model and a high-threshold screening strategy, providing a highly reliable initial causal framework for subsequent iterative optimization. To quantitatively evaluate the quality and coverage of the initial confident causal graph, two key metrics are defined: Accuracy (Acc) is the ratio of correctly predicted edges to the number of predicted initial graph edges (Acc = number of correctly predicted edges / number of predicted initial graph edges), reflecting the accuracy of edge discrimination; Coverage (C) is the ratio of the number of predicted initial graph edges to the number of true graph edges (C = number of predicted initial graph edges / number of true graph edges), measuring the scope of capturing true causal relationships.

[0071] Step 120: Based on the initial confidence causal graph, perform context modeling on the text to be identified, extract the multi-path information of each event pair in the text to be identified, input the multi-path information into the generative language model to generate causal relationships, and obtain the initial causal relationship results; Furthermore, based on the initial certainty causal graph, context modeling is performed on the text to be identified. First, multi-source heterogeneous semantic features of each event pair in the text to be identified are extracted, including event context information, graph structure information, and document global information. Then, the three types of heterogeneous features are encoded separately through a multi-encoder architecture, and the three types of encoded features are horizontally concatenated and fused to obtain a high-dimensional fused latent state vector. Finally, the fused vector is input into a generative language model, and causal features are dynamically extracted through the sequence generation capability of the decoder to complete the generation of causal relationships and output the initial causal relationship results.

[0072] In some embodiments, the multipath information includes event context information, graph structure information, and document global information. The event context information includes the semantics of the preceding and following sentences between event pairs. The graph structure information includes the adjacency structure between event pairs in the causal graph. The document global information includes the overall semantic encoding of the text to be identified.

[0073] In some embodiments, the step of inputting multi-path information into a generative language model to generate causal relationships and obtain initial causal relationship results includes: The event context information is converted into context text encoding and input into the context encoder to obtain local semantic features; The graph structure information is converted into causal graph text encoding and input into a causal graph encoder to obtain structure-aware features. The global information of the document is converted into document text encoding and input into the document encoder to obtain global semantic features; The hidden state vector is obtained by fusing local semantic features, structure-aware features, and global semantic features. The hidden state vector is input into the decoder, and the causal features are dynamically extracted through the sequence generation capability of the decoder to obtain the initial causal relationship results.

[0074] Figure 3 This is a flowchart of graph-driven multi-path semantic modeling for the event causality identification method provided in this application embodiment, such as... Figure 3 As shown, to achieve joint modeling of event context information, causal graph structure information, and global document semantics, a multi-path linearization strategy is adopted to transform the three types of heterogeneous semantic information—event context, causal graph structure, and global document—into a unified sequence input that can be processed by a generative language model. This transforms heterogeneous semantic features into a unified sequence input that can be processed by a generative language model.

[0075] First, construct three types of encoded inputs: Contextual text encoding extracts the original sentences containing event pairs, preserving their semantic details, and constructs the input sequence format as follows:

[0076] in , These represent the sentence text containing the event pair, with [SEP] serving as a separator to distinguish between different sentences. The context wrapper label defines the range of encoded input. The input sequence is .

[0077] Causal graph text encoding, the core objective of causal graph encoder is to transform graph structure information into a semantic representation that fits the sequence model. Its processing flow follows a three-level architecture of "subgraph extraction - confidence ranking - linearization encoding".

[0078] N-hop causal subgraph extraction: For each pair of target events... First, extract the causal subgraph containing its N-hop range to ensure coverage of both direct and indirect causal relationships. For the target event pair... The system will extract: The sentence in which it appears and the sentences before and after it (contextual path); In the N-hop subgraph structure (graph path) of the graph, such as ; The global summary or sentence order of the document (document path).

[0079] The above information, after being linearized, is input into a generative language model to generate a prediction of whether there is a causal relationship between the event and its direction.

[0080] Confidence-driven edge sorting: high-confidence causal edges enter the encoding process first, and low-confidence edges are placed later in sequence. This forms an encoding structure of "key information first, secondary information in sequence"; The linearized text input sequence is constructed as follows: , among which > indicates that the event was mentioned. <cause>and<Caused by> These represent positive and negative causal relationships, respectively. The semicolon is used to separate different causal pairs. The label serves as a wrapper around the causal graph to define the range of encoded inputs.

[0081] Document text encoding preserves the global semantic information of the original document, and is encoded at the sentence level, with the following format:

[0082] in, Represents the sequentially arranged sentence text in the document. As a sentence-level separator, it is used to distinguish different semantic units. and As a document wrapping tag, it defines the range of encoded input. Encodes the document text.

[0083] All virtual tokens (such as < >) Maintain consistency throughout the coding process to ensure consistency in event representation between context, graph structure, and document.

[0084] For the three types of input information mentioned above, a multi-encoder architecture is used for processing. By learning the three types of information, a high-dimensional fusion vector is output as the basis for downstream causal reasoning.

[0085] Specifically, the BART encoder uses three shared parameters to process text input separately: Context encoder The text sequence of the sentences containing the two events Mapping to local semantic features Capture word-level semantic interactions.

[0086] Cause-effect graph encoder For linearized graph sequences Encode to generate structure-aware features The confidence ranking mechanism influences the model's ability to understand semantic sequences by encoding position. Document encoder Process complete documents Output global semantic features This enables the integration of global semantic associations for events.

[0087] The encoding process employs a unified multi-layer Transformer architecture, modeling long-distance dependencies through a multi-head self-attention mechanism, extracting semantic features layer by layer, and finally outputting a hidden state containing the global semantics of the document. The feature fusion stage uses a horizontal concatenation strategy of the hidden states to achieve information fusion; the passed vector is represented as:

[0088] in, For local semantic features, This is a structurally perceptible feature. For global semantic features, This is the hidden state vector.

[0089] This vector contains multi-level semantic information, ranging from local to global and from text to graph structure, providing richer causal reasoning basis for generative language models.

[0090] In the causal relationship generation task modeling, this method constructs a joint reasoning mechanism based on a generative language model within the causal relationship generation module. The fused hidden state vectors... In the input decoder, a sequence generation task is performed. Causal features are dynamically extracted using the decoder's sequence generation capabilities, ultimately outputting the causal relationship category of event pairs. The specific steps are as follows: A structured decoder input template is constructed to guide the model to focus on causal reasoning of target event pairs. The designed decoding prompt template takes the following form: __ This template identifies the event by explicitly inserting it. and This is used to guide the decoder to focus on the task of generating causal relationships between specific event pairs.

[0091] During the decoding process, the model takes the cue template as input and is based on the hidden state vector. Initiate the generative inference process, progressively generating a token sequence, and extract the hidden state representation vector of the last generated token before decoding ends. This vector serves as a semantic representation of the causal relationship between events, comprehensively encoding context, graph structure, and document information.

[0092] Then, the vector The input is fed into the classifier layer (LM Head) for causal relationship classification. This classifier consists of a set of linear transformation layers, whose function is to map semantic vectors to a three-dimensional answer space. Its classification calculation process is as follows:

[0093]

[0094] in, The classifier weight matrix is... For bias terms, The mapped fraction vector, This represents the probability that the event belongs to the k-th class; the labels corresponding to the three classes are... <na>(No causal relationship) <cause>(Positive causality) and <causedby>(Reverse causality) Let k be the k-th element in the z-vector. Let j be the j-th element in the vector z. Is the event related to and A deep semantic representation that integrates contextual information, causal graph structure, and document information. yes The reason is The semantic feature of "cause" is reflected in the relationship between the two; if there is no obvious causal relationship between the two, the semantic feature is neutral.

[0095] Through the above mechanism, generative language models can effectively integrate multi-source semantic information, modeling and discriminating causal relationships between event pairs at a semantic depth level. Ultimately, based on vectors... The semantic representation completes the mapping and normalized probability calculation of causal relationship categories, and outputs the specific causal relationship label to which the event pair belongs.

[0096] To improve the accuracy and semantic fit of the generated results, the following strategies are used in the training phase: The causal relation generative language model based on causal graphs is based on the BART-large architecture and shares the same training strategy as the causal graph-based model, using Focal Loss and L2 regularization. The optimization objective is to minimize the joint loss function. Based on the predicted answer words in the answer space... The probability distribution and the true labels are used to construct a specific loss function, which is defined as follows: Let the total number of training samples be N, and the number of predicted categories be C=3 (corresponding to the answer space). <na> 、 <cause>,<Caused by> For the nth sample, its true label Given a one-hot vector, the predicted probability distribution is as follows: The first loss function is defined as follows:

[0097] in, This is a category weight balancing factor used to adjust the weights of different categories, which can alleviate the category imbalance problem. As an adjustment factor, it controls the degree of decay of loss for easily classified samples. ≥0, The larger the value, the more significant the weight decay of easily classified samples. This represents the predicted probability of the nth sample in the true class, where N is the number of samples. This is the first loss function.

[0098] The formula for calculating the second loss function, which is composed of L2 regularization terms, is shown below:

[0099] The final loss function is the sum of the two: , The sum of squares of all parameters, The regularization coefficient is . This is the second loss function.

[0100] During training, the AdamW optimizer was used to dynamically adjust the learning rate, and a gradient accumulation compensation training strategy was employed. The batch size was set to 1, and by setting the gradient accumulation step count to 16, the effective batch size was increased to 16 (aligned with model training), ensuring parameter update stability. The initial learning rate was set to... L2 regularization coefficient =0.01.

[0101] Through the aforementioned causal graph-driven generation mechanism, structural information and linguistic context can be better integrated, thereby enabling generative reasoning on the causal relationships between complex event pairs and providing high-quality candidate edges for iterative optimization of graph structures.

[0102] In some embodiments, the maximum length of the model input text is limited to 200 tokens for the event context information (100% data coverage), the maximum length of the linearized text of the causal subgraph is limited to 200 tokens, and the maximum length of the document text is limited to 624 tokens (80% data coverage). Any excess is truncated in sequence.

[0103] Step 130: Based on the initial causal relationship results, the initial confident causal graph is iteratively optimized in multiple rounds using a graph structure constraint mechanism. Edges are dynamically selected according to the confidence level, the initial confident causal graph is updated, and it is determined whether the iteration termination condition is met. When the iteration termination condition is met, the iteration stops, and the optimized confident causal graph is obtained. The process is straightforward: the initial causal relationship results are used to update the initial confident causal graph structure. A graph structure constraint mechanism is introduced, dynamically selecting edges with high confidence as the set of reliable edges. Multiple rounds of iterative optimization are then performed based on edge weight changes and structural adjustment strategies. After each iteration, the convergence of the graph structure is assessed, ultimately resulting in a structurally stable optimized confident causal graph from which the final event causal relationship identification results are extracted. The graph structure optimization process employs a dynamic edge set update strategy and a convergence judgment mechanism. By setting a confidence threshold and an upper limit on the number of iterations, the interpretability and stability of the final causal graph are ensured.

[0104] In some embodiments, the step of performing multi-round iterative optimization of the initial confidence causal graph based on the initial causal relationship results using a graph structure constraint mechanism, dynamically selecting edges according to the confidence level, updating the initial confidence causal graph, and determining whether the iteration termination condition is met, stopping the iteration when the iteration termination condition is met, and obtaining the optimized confidence causal graph includes: Based on the initial causal relationship results, event pairs with causal confidence greater than or equal to the current dynamic threshold are extracted and transformed into corresponding directed causal edges, which are then added as new credible edges and added to the causal edge set. The confidence level of the causal edge set in the initial confidence causal graph is filtered, and all credible edges with a causal confidence level lower than the current dynamic threshold are removed. The credible edges with a confidence level greater than or equal to the current dynamic threshold are retained as credible edges. A dynamic filtering mechanism is adopted to adjust the dynamic threshold. Based on the adjusted dynamic threshold, the retained trusted edges and the newly added trusted edges are integrated to update the initial confidence causal graph. The updated confidence causal graph is obtained and it is determined whether the iteration termination condition is met. When the iteration termination condition is met, the iteration stops and the optimized confidence causal graph is obtained.

[0105] In some embodiments, the confidence of the current edge set is first filtered. The existing edge set in the initial confident causal graph (or the causal graph of the current round in the iteration process) is filtered for confidence. All edges with causal confidence lower than the current round dynamic threshold are removed, and edges with confidence greater than or equal to the current dynamic threshold are retained as retained reliable edges to ensure the semantic reliability of the existing edges in the graph.

[0106] Furthermore, based on the initial causal relationship results, new edges are added. Based on the generated initial causal relationship results, event pairs with causal confidence greater than or equal to the current round's dynamic threshold are extracted and transformed into corresponding directed causal edges, which are then added to the causal edge set as new trusted edges.

[0107] Finally, a dynamic filtering mechanism is used to adjust the threshold. A high initial dynamic threshold is set. In each iteration, if the prediction confidence of all event pairs is lower than the current threshold, the threshold is reduced by a preset step size until the minimum threshold is reached. Based on the adjusted current dynamic threshold, the retained reliable edges and the newly added reliable edges are integrated to form the causal graph after the current round update.

[0108] In some embodiments, adjusting the dynamic threshold using a dynamic filtering mechanism includes: Determine whether the causal confidence of all event pairs is lower than the current dynamic threshold. If the causal confidence of all event pairs is lower than the current dynamic threshold, decrease the dynamic threshold by a preset step size and adjust the dynamic threshold until the minimum threshold is reached. If there is a causal confidence of all event pairs that is higher than the current dynamic threshold, keep the current dynamic threshold unchanged.

[0109] In some embodiments, based on the initial causal graph, the graph structure is updated through feedback of the causal generation results to achieve a gradual enhancement of the identification of event causal relationships, specifically including the following steps: To dynamically adjust the evolution rhythm of the graph structure and improve the accuracy and structural fit of newly added edges, an adaptive threshold mechanism and new edge extraction strategy are introduced. Based on the edge confidence output by the generative language model, high-confidence edges are dynamically selected and added to the causal graph.

[0110] The specific approach is as follows: a mechanism that dynamically adjusts with each iteration is adopted, and an initial threshold is set. In each iteration, if the predicted probabilities of all event pairs are lower than the current threshold, then a threshold reduction is triggered.

[0111] in, Indicates a decreasing step size. This represents the minimum threshold. The selection criteria are gradually relaxed as iterations progress to accommodate the decreasing confidence level of the model in later stages.

[0112] This strategy enables the edge selection criteria to be dynamically adjusted as the graph structure evolves round by round, thereby strictly controlling structural noise in the initial stage and gradually expanding the scope of edge generation in the later stage, ensuring the stability and effectiveness of causal graph updates.

[0113] The iterative construction process of the causal relationship graph employs a "dynamic filtering of unpredicted event pairs - round-by-round judgment" strategy. Each iteration only processes currently undecided event pairs, based on real-time thresholds. Perform causal relationship prediction. After predicting all data in one round, add the confirmed causal edges from that round to the causal relationship graph. Continue this process until the iteration termination condition is met; the result in the causal graph at the point of termination is the final prediction result.

[0114] To prevent infinite iteration or overfitting, iteration termination conditions are set for graph structure iterations, including the following two types of convergence determination mechanisms: When the threshold reaches its lowest point and there are no valid edges, the current threshold reaches its minimum threshold. The iteration stops when the predicted probability of all event pairs is less than the current threshold, and the predicted probability of all event pairs is less than the current threshold. When all event pairs have been processed, that is, when all event pairs have been determined to be... <na> 、 <cause>or<Caused by> If there are no unprocessed event pairs, the iteration stops.

[0115] Through this dual termination mechanism, the present invention ensures the convergence of the graph structure update process while maintaining model stability.

[0116] Figure 4 This is a graph showing the performance changes of the event causality identification method provided in this application during the iterative optimization process of graph structures. Figure 4 As shown, the trends of precision, recall, and F1 score (the harmonic mean of precision and recall) are presented at different iteration rounds for intra-sentence tasks, inter-sentence tasks, and the overall task. In the initial iterations (rounds 0-3): In the initial stage (round 0), the model used an extremely high threshold (thr0=0.99) to filter causal edges. The precision (P) for inter-sentence tasks reached 75.7%, but the recall was only 15.37%, capturing only a very small number of high-confidence causal pairs, resulting in severely insufficient coverage. As the threshold was gradually reduced, the inter-sentence recall increased to 20.3% after the first iteration (F1 score increased from 25.55% to 31.61%), and the F1 score for intra-sentence tasks increased from 41.96% to 50.63%.

[0117] Mid-Iteration (Rounds 4-10): As the threshold continues to decrease, potential causal relationships are gradually uncovered. The recall rate for the inter-sentence task increased from 30.55% (Round 4) to 40.61% (Round 10), and the F1 score increased from 40.3% to 46.12%, indicating that the iterative mechanism effectively alleviated the data sparsity problem in long-distance causal inference. By Round 10, the inter-sentence F1 score had increased by 20.57 percentage points compared to the initial stage. Meanwhile, the precision (P) for the intra-sentence task decreased from 69.73% to 64.86%, mainly because the model incorporated more potential causal edges into the prediction after the threshold was relaxed. Although the precision decreased slightly, the F1 score for the intra-sentence task increased from 56.46% to 63.28%, and the significant improvement in recall (R) (from 45.21% to 66.34%) positively impacted the overall performance.

[0118] Late iterations (rounds 11-25): As the threshold decreases to its lowest point, model performance gradually stabilizes. The inter-sentence F1 score stabilizes at 47.4% at round 25, slightly lower than the mid-term peak (47.63%). Excessive iteration may introduce low-quality causal edges, leading to an increase in the misclassification rate. The intra-sentence task F1 score fluctuates around 63%. Due to the explicitness and locality of intra-sentence causal cues, the optimization space is limited in the later iterations. The overall task F1 score gradually increases from 52.36% (round 10) to 53.13% (round 22), but the increase slows down in the later stages, and the model approaches its performance limit.

[0119] Ultimately, the event causal graph, after multiple rounds of optimization, not only has a more complete structure, but also contains edges with higher semantic confidence and directional accuracy, and can be directly used as the output result for event causal relationship identification.

[0120] In some embodiments, the performance of the proposed method is demonstrated using the EventStoryLine dataset, which is widely used in ECI (Event Causality Identification) tasks, as an example. The EventStoryLine dataset contains English documents on 22 topics, covering 5156 event mentions and 70579 event relationship annotations. Event causal relationship annotations include three types: NONE, PRECONDITION, and FALLING_ACTION. The ratio of positive to negative samples is 1:12; the ratio of cause-to-effect and result-to-cause inference is 1:1.08; the ratio of intra-sentence event pairs to inter-sentence event pairs is 1:5.8; and the ratio of intra-sentence causal relationships to inter-sentence causal relationships is approximately 1:2.1. The experimental data split follows the original design, selecting the last two topics (Topics 37 and 41) as the validation set, and using 5-fold (Fold1~Fold5) cross-validation for the remaining 20 topics. Precision (P), recall (R), and F1 score (F1) were used as performance metrics.

[0121] Model training optimizes parameters on the training set, selecting the optimal parameters based on validation set performance. The causal graph construction model uses HuggingFace's "RoBERTa-base" as the backbone model for initial confidence causal graph construction, with a word embedding dimension of d=768. The generative model uses the same platform's BART-base architecture, supporting cross-modal semantic fusion and causal relationship generation. Model training is implemented using the PyTorch framework, accelerated by an NVIDIA GeForce RTX 4090 graphics card to ensure efficiency in large-scale text encoding and graph structure reasoning.

[0122] To ensure the validity of inputs and the stability of model operation in causal relationship identification under multimodal information fusion conditions, the maximum length of various types of input information has been reasonably limited. Specifically, this includes: The maximum input length for event context information is set to 200 tokens to ensure that the semantic context of all target event pairs is covered, achieving 100% data coverage. The maximum length of the linearized text representation of the causal subgraph structure is also set to 200 tokens to adapt to the graph structure encoding capability of the model; The document text sequence allows a maximum input length of 624 tokens, balancing contextual integrity and model load capacity while maintaining 80% text coverage. For inputs exceeding the maximum length, a sequential truncation strategy is adopted, prioritizing the retention of semantic content closer to the target event pair to reduce semantic loss caused by truncation.

[0123] To improve the performance of graph-based causal relationship recognition during the model training and inference stages, a dynamic parameter monitoring and tuning mechanism is proposed to support the stable operation of iterative prediction strategies and threshold adaptive mechanisms. This mechanism mainly includes the following aspects: Key performance indicator (KPI) monitoring: During model execution, the following key KPIs are recorded and statistically analyzed in real time to dynamically evaluate the convergence and effectiveness of the model's inference process: Current average causal graph size; The number of new causal edges added in each iteration; The number of causal edges correctly identified in each round; The update trajectory and convergence rate of each event pair in the iteration rounds.

[0124] Threshold change analysis: Track and analyze the dynamic threshold adjustment process in each round of prediction, and observe its impact on the prediction results and the model convergence trend.

[0125] Optimal parameter configuration selection: Based on the above dynamic characteristics, this invention designs multiple sets of parameters. Parameter combinations are used for grid search optimization. The optimal parameter configuration is determined by evaluating the model performance of each combination on the validation set (with F1 score as the core evaluation metric), and the final performance is reported on the test set.

[0126] Through the collaborative design of the above-mentioned input management and parameter tuning mechanisms, this invention can effectively control the consumption of model computing resources and improve the controllability, stability and accuracy of the causal relationship prediction process while ensuring full utilization of input information.

[0127] Furthermore, Table 1 compares the overall performance of the IGPL (Iterative Graph Prompt Learning) of this invention with several baseline models on the ESC (Event Story Line Corpus) corpus. Table 1 shows the competitors' results for all intra-sentence, inter-sentence, and overall metrics on the ESC dataset.

[0128] Table 1. Results for all intra-sentence, inter-sentence, and overall results on the ESC dataset.

[0129] As shown in Table 1, LLaMA-2-7B generally performs poorly in intra-sentence, inter-sentence, and overall tasks, with an inter-sentence F1 score of only 10.0% and an overall F1 score of 11.5%. BERT and RoBERTa have higher intra-sentence precision (62.4% and 59.7%, respectively), but lower recall (32.6% and 38%, respectively). As a large language model, LLaMA-2-7B possesses powerful general language understanding and generation capabilities, able to handle complex semantics and generate coherent text. However, it performs poorly in identifying causal relationships in specialized domains. This model relies solely on general semantic understanding and lacks targeted modeling of event causal relationships, making it difficult to accurately capture causal directionality and cross-sentence logical connections. It also exhibits significant shortcomings in the structured modeling of document-level cross-sentence causal relationships.

[0130] BERT and RoBERTa achieved high in-sentence precision (62.4% and 59.7%, respectively), but low recall (32.6% and 38%, respectively). This reflects that traditional pre-trained language models have a strong ability to capture explicit causal cues and achieve high in-sentence precision, but insufficient coverage of causal relationships.

[0131] LONG achieves a slightly better intra-sentence F1 score of 48% than BERT / RoBERTa, while its inter-sentence performance is similar (32.7%). ERGO and SENDIR, through graph structure modeling, improve their inter-sentence F1 scores to 38.5% and 39%, respectively, but their overall performance remains below 50%. iLIF, as an iterative model, achieves the highest intra-sentence F1 score of 60% and an inter-sentence F1 score of 42.8%. IGPL performs better in both inter-sentence (48.7%) and overall (52.2%) performance. Although LONG is optimized for long texts, its reliance solely on extending the input sequence cannot effectively address cross-sentence semantic aggregation and long-distance dependency issues, limiting its application in document-level causal reasoning.

[0132] IGPL's superior performance stems from its multi-path encoding fusion and iterative graph cue learning design. IGPL achieves deep fusion of local details, graph structure dependencies, and global themes through three-path encoding: event context, causal graph structure, and global document. IGPL captures implicit associations through a causal graph encoder, achieving an intra-sentence recall rate of 59.5%, thus providing more comprehensive coverage of intra-sentence causal relationships. Simultaneously, IGPL's generative BART architecture and cue learning templates more efficiently activate pre-trained knowledge, empirically demonstrating the advantage of generative models in directly generating causal labels for capturing implicit causal logic.

[0133] IGPL achieves performance optimization in document-level event causality recognition tasks. IGPL employs an "initial confidence causal graph construction and iterative threshold adjustment" mechanism. In the first round, a high threshold is used to filter reliable causal edges and construct the inference framework. Subsequently, adaptive thresholds are used to gradually mine low-confidence associations, forming causal chain expansions from simple to complex, effectively improving the ability to resolve complex causal chains. IGPL utilizes the BART generative model architecture, transforming causal relationship recognition into an answer word generation task, directly activating the causal knowledge inherent in the language model, mining implicit causal logic in documents, and strengthening the model's reasoning on causal direction.

[0134] To verify the effectiveness of each module in the proposed iterative graph cueing learning model, five ablation experiments were designed to remove key modules and analyze their impact on performance: (1) Remove the event context module (IGPL w / o SE, Sentence Encoder): Remove the event context encoder from the input of the causal relationship prediction model, and keep only the causal graph encoder and document encoder; (2) Remove the causal graph (IGPL w / o GE, Graph Encoder): Remove the causal graph encoder from the input of the causal relationship prediction model and keep only the event context encoder and document encoder; (3) Remove the document encoder (IGPL w / o DE, Doc Encoder): In the input of the causal relationship prediction model, remove the document encoder and keep only the event context encoder and the causal graph encoder; (4) Remove the initial certainty graph (IGPL w / o ICG): Remove the initial certainty graph in the complete model framework, that is, do not load any high-confidence prior causal edge information in the initial stage of the iteration process; (5) Removal of Iterative Threshold Mechanism (IGPL w / o ITM): The dynamic iterative optimization module of the model is removed, and the adaptive threshold adjustment strategy is removed. The model directly selects the threshold based only on the single-round prediction results. <na> 、 <cause>,<Caused by> The relationship type with the highest probability among the three types of labels is used as the final determination, and no further rounds of causal graph updates are performed.

[0135] Table 2 Performance of Causal Relationship Identification

[0136] Table 2 shows the causal relationship identification performance of schemes (1)-(5). The results of the ablation experiment show: The first observation is that removing any module leads to varying degrees of performance degradation in the model. Specifically, removing the event context module (IGPL w / o SE) has the most significant impact on model performance, resulting in the most pronounced performance drop; removing the causal graph encoder (IGPL w / o GE) causes the inter-sentence accuracy (P-value) to drop sharply from 49.27% ​​to 38.39%, demonstrating its crucial role in the accuracy of inter-sentence causal relationships; removing the document encoder (IGPL w / o DE) results in a certain degree of overall performance decline.

[0137] Comparing the removal of the causal graph encoder (IGPL w / o GE) and the removal of the initial confidence causal graph (IGPL w / o ICG), IGPL w / o GE exhibits a more significant performance decline in both inter-sentence and overall performance, while IGPL w / o ICG displays a characteristic of "high precision, low recall". The absence of the causal graph encoder in IGPL w / o GE causes the model to lose its ability to model structured dependencies between events, making it unable to obtain prior knowledge such as multi-hop causal chains, hindering the effective utilization of causal graph logical connections, and significantly reducing cross-sentence reasoning ability due to the lack of graph structure support. IGPL w / o ICG, by removing the initial confidence causal graph, retains the dynamic construction capability of the causal graph encoder, allowing for the gradual discovery of causal edges through subsequent iterations, thus maintaining a relatively high level of inter-sentence accuracy.

[0138] Comparing the experimental results of removing the event context module (IGPL w / o SE) and removing the document encoder (IGPL w / o DE), the absence of the event context module has a more severe impact on intra-sentence performance. Although the document with IGPL w / o SE still contains statements with event context, the excessively long intervals between related statements make it difficult for the model to clearly identify the spoiler statements containing the event. In contrast, IGPL w / o DE only lacks document-level semantic integration capabilities. Although this affects the acquisition of long-distance information across sentences, it can still retain some local and structured information through event context and causal graph structure, resulting in a smaller performance degradation.

[0139] In summary, the event causal relationship recognition method based on iterative graph cue learning proposed in this application first constructs a model of the event pairs in the text to be identified into a causal graph. By constructing directional cue templates, the causal relationship determination task is transformed into a fill-in-the-blank task, and an initial causal graph is generated. In the graph, nodes represent events, edges represent predicted causal relationship directions, and edge weights represent confidence levels. Then, multi-path context modeling is performed based on this causal graph to extract local context information, graph structure information, and global document information of the event pairs, and this information is input into a generative language model to generate causal relationships. The generated results are then used to iteratively optimize the causal graph structure. The confidence level is used to dynamically filter edges and update the graph structure, strengthening the semantic relationships in the graph round by round, and finally extracting the optimized causal relationship recognition results.

[0140] In this embodiment, the accuracy and direction discrimination capabilities of document-level causal relationship recognition are effectively improved through a graph prompting learning framework and a multi-round graph structure optimization mechanism. The initial causal graph construction method based on the prompting learning paradigm fully activates the causal common sense inherent in the causal graph construction model, solving the problem that traditional fine-tuning models struggle to utilize implicit knowledge. The use of a generative language model to fuse graph structure and contextual information for causal relationship generation overcomes the limitations of static template methods on causal expression forms. Furthermore, the iterative graph structure update mechanism enables the construction of causal reasoning chains from simple to complex, enhancing the system's ability to parse complex cross-sentence causal structures.

[0141] In some embodiments, the causal graph construction model is the RoBERTa model, and the generative language model is the BART generative language model.

[0142] Generative models are pre-trained language generative models (such as BART) that support multi-path semantic fusion. They can aggregate and model the semantic expressions of event pairs under different structural paths, thereby improving the contextual adaptability and accuracy of causal generation.

[0143] The event causal relationship recognition method based on iterative graph hint learning provided in this application effectively improves the accuracy, structural integrity, and multi-hop reasoning ability of event causal relationship recognition by introducing causal directionality hint templates, multi-path graph structure modeling, and iterative optimization strategies. Without relying on a fixed answer space, it achieves high-precision, directional event causal relationship recognition for complex text scenarios through graph-structure-guided semantic fusion and iterative calibration.

[0144] By fully leveraging the causal knowledge embedded in the causal graph construction model and combining the expressive power of graph structures with the contextual modeling advantages of generative language models, this approach improves the accuracy and robustness of causal direction recognition and cross-sentence causal chain modeling. It is suitable for document-level event causal analysis scenarios and integrates the advantages of prompting learning and graph modeling. It possesses causal direction modeling capabilities, multi-path semantic fusion capabilities, and causal graph structure iterative optimization capabilities, significantly improving the performance of causal recognition tasks in complex text, long documents, and cross-sentence causal chain scenarios. It is applicable to various application scenarios such as causal chain construction, document-level event understanding, and causal knowledge graph generation.

[0145] The event causality recognition method based on iterative graph cue learning provided in this application can be executed by an event causality recognition system based on iterative graph cue learning. This application uses the execution of the event causality recognition method based on iterative graph cue learning by an event causality recognition system based on iterative graph cue learning as an example to illustrate the event causality recognition system based on iterative graph cue learning provided in this application.

[0146] This application also provides an event causality recognition system based on iterative graph cue learning, such as... Figure 5 As shown, the event causal relationship recognition system based on iterative graph prompting learning includes: a construction module 510, a causal relationship generation module 520, a causal graph optimization module 530, and a causal relationship recognition module 540.

[0147] Module 510 is used to construct the initial confidence causal graph; The causal relationship generation module 520 is used to perform contextual modeling on the text to be identified based on the initial certainty causal graph, extract multi-path information of each event pair in the text to be identified, input the multi-path information into the generative language model to generate causal relationships, and obtain the initial causal relationship results. The causal graph optimization module 530 is used to perform multiple rounds of iterative optimization on the initial confident causal graph based on the initial causal relationship results and adopt a graph structure constraint mechanism. It dynamically selects edges according to the confidence level, updates the initial confident causal graph, and determines whether the iteration termination condition is met. When the iteration termination condition is met, the iteration stops and the optimized confident causal graph is obtained. The causal relationship identification module 540 is used to obtain the event causal relationship results of the text to be identified based on the optimized confidence causal graph.

[0148] It is worth noting that the event causal relationship recognition system based on iterative graph prompting learning can be integrated into an event causal recognition engine that can be deployed on a graph neural network inference platform for application scenarios such as text causal reasoning, document analysis, and event tracking.

[0149] The event causality recognition system based on iterative graph cueing learning provided in this application embodiment can achieve... Figures 1 to 4 The various processes implemented in the embodiment of the event causality recognition method based on iterative graph prompting learning will not be described again here to avoid repetition.

[0150] In some embodiments, such as Figure 6 As shown, this application embodiment also provides an electronic device 600, including a processor 601, a memory 602, and a computer program stored in the memory 602 and executable on the processor 601. When the program is executed by the processor 601, it implements the various processes of the above-described embodiment of the event causal relationship recognition method based on iterative graph prompting learning, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0151] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0152] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described event causal relationship recognition method based on iterative graph prompting learning, and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0153] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0154] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described event causal relationship recognition method based on iterative graph cueing learning.

[0155] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0156] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described embodiment of the event causal relationship recognition method based on iterative graph prompting learning, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0157] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0158] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitations, an element defined by the phrase "comprising one…" does not exclude the presence of other identical elements in the process, method, article, or system that includes that element. Furthermore, it should be noted that the scope of the methods and systems in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0159] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the event causal relationship recognition method based on iterative graph prompting learning of the various embodiments of this application.

[0160] In the description of this application, "first feature" and "second feature" may include one or more of the features.

[0161] In the description of this application, "multiple" means two or more.

[0162] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0163] In the description of this specification, the references to "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0164] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.< / cause> < / na> < / cause> < / na> < / cause> < / na> < / causedby> < / cause> < / na> < / cause> < / causedby> < / cause> < / cause> < / na> < / cause> < / na> < / cause> < / ei> < / ei> < / na> < / causes> < / na> < / causes>

Claims

1. A method for identifying event causal relationships based on iterative graph cue learning, characterized in that, The method includes: Construct an initial confidence causal graph; Based on the initial certainty causal graph, context modeling is performed on the text to be identified, multi-path information of each event pair in the text to be identified is extracted, and the multi-path information is input into the generative language model to generate causal relationships and obtain the initial causal relationship results. Based on the initial causal relationship results, a graph structure constraint mechanism is used to perform multiple rounds of iterative optimization on the initial confident causal graph. Edges are dynamically selected according to the confidence level, the initial confident causal graph is updated, and it is determined whether the iteration termination condition is met. When the iteration termination condition is met, the iteration stops, and the optimized confident causal graph is obtained. Based on the optimized conviction causal graph, the event causal relationship results of the text to be identified are obtained.

2. The event causal relationship recognition method based on iterative graph cueing learning according to claim 1, characterized in that, The construction of the initial confidence causal graph includes: The text to be identified is input into the trained causal graph construction model. Through the causal directionality prompt template, the event causality of the event pairs in the text to be identified is determined, and the causal relationship prediction results between the event pairs in the text to be identified are obtained. Based on the causal relationship prediction results between event pairs in the text to be identified, a set of causal edges is formed by filtering credible edges through a set confidence threshold, and an initial confident causal graph about the text to be identified is constructed.

3. The event causal relationship recognition method based on iterative graph cueing learning according to claim 1, characterized in that, The multi-path information includes event context information, graph structure information, and document global information. The event context information includes the semantics of the preceding and following sentences between event pairs. The graph structure information includes the adjacency structure between event pairs in the causal graph. The document global information includes the overall semantic encoding of the text to be identified.

4. The event causal relationship recognition method based on iterative graph cueing learning according to claim 3, characterized in that, The step of inputting multi-path information into a generative language model to generate causal relationships and obtain initial causal relationship results includes: The event context information is converted into context text encoding and input into the context encoder to obtain local semantic features; The graph structure information is converted into causal graph text encoding and input into a causal graph encoder to obtain structure-aware features. The global information of the document is converted into document text encoding and input into the document encoder to obtain global semantic features; The hidden state vector is obtained by fusing local semantic features, structure-aware features, and global semantic features. The hidden state vector is input into the decoder, and the causal features are dynamically extracted through the sequence generation capability of the decoder to obtain the initial causal relationship results.

5. The event causal relationship recognition method based on iterative graph cueing learning according to claim 2, characterized in that, The process involves using a graph structure constraint mechanism to iteratively optimize the initial confident causal graph based on the initial causal relationship results. Edges are dynamically selected according to the confidence level, the initial confident causal graph is updated, and it is determined whether the iteration termination condition is met. The iteration stops when the termination condition is met, resulting in the optimized confident causal graph, including: Based on the initial causal relationship results, event pairs with causal confidence greater than or equal to the current dynamic threshold are extracted and transformed into corresponding directed causal edges, which are then added as new credible edges and added to the causal edge set. The confidence level of the causal edge set in the initial confidence causal graph is filtered, and all credible edges with a causal confidence level lower than the current dynamic threshold are removed. The credible edges with a confidence level greater than or equal to the current dynamic threshold are retained as credible edges. A dynamic filtering mechanism is adopted to adjust the dynamic threshold. Based on the adjusted dynamic threshold, the retained trusted edges and the newly added trusted edges are integrated to update the initial confidence causal graph. The updated confidence causal graph is obtained and it is determined whether the iteration termination condition is met. When the iteration termination condition is met, the iteration stops and the optimized confidence causal graph is obtained.

6. The event causal relationship recognition method based on iterative graph cueing learning according to claim 2, characterized in that, The causal graph construction model is the RoBERTa model, and the generative language model is the BART generative language model.

7. The event causal relationship recognition method based on iterative graph cueing learning according to claim 5, characterized in that, The method of adjusting the dynamic threshold using a dynamic filtering mechanism includes: Determine whether the causal confidence of all event pairs is lower than the current dynamic threshold. If the causal confidence of all event pairs is lower than the current dynamic threshold, decrease the dynamic threshold by a preset step size and adjust the dynamic threshold until the minimum threshold is reached. If there is a causal confidence of all event pairs that is higher than the current dynamic threshold, keep the current dynamic threshold unchanged.

8. An event causal relationship recognition system based on iterative graph cue learning, implemented using the event causal relationship recognition method based on iterative graph cue learning as described in any one of claims 1 to 7, characterized in that, The system includes: The building block is used to construct the initial confidence causal graph; The causal relationship generation module is used to perform contextual modeling on the text to be identified based on the initial certainty causal graph, extract multi-path information of each event pair in the text to be identified, input the multi-path information into the generative language model to generate causal relationships, and obtain the initial causal relationship results. The causal graph optimization module is used to perform multiple rounds of iterative optimization on the initial confident causal graph based on the initial causal relationship results and using a graph structure constraint mechanism. It dynamically selects edges according to the confidence level, updates the initial confident causal graph, and determines whether the iteration termination condition is met. When the iteration termination condition is met, the iteration stops, and the optimized confident causal graph is obtained. The causal relationship identification module is used to obtain the event causal relationship results of the text to be identified based on the optimized confidence causal graph.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the event causality recognition method based on iterative graph cueing learning as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the event causality recognition method based on iterative graph cueing learning as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Security and protection monitoring situation awareness method and device based on big data and medium

    CN121561358A

  • Large language model iterative optimization training method and system based on metamorphic test

    CN122047517A