Event causal reasoning method and electronic device for intelligence analysis

By combining knowledge graphs and large language models, a traceable causal evidence chain is generated from multi-source evidence fragments, solving the problems of low efficiency, high cost, and inconsistent conclusions in existing event causal reasoning, and realizing efficient and verifiable event causal reasoning analysis.

CN121960796BActive Publication Date: 2026-06-16DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DALIAN UNIV OF TECH
Filing Date
2026-04-02
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency, high cost, difficulty in scaling, inconsistent conclusions, and difficulty in dynamic correction in multi-source event analysis tasks. In particular, when there is insufficient evidence or conflicting materials, the model generates unverifiable causal explanations, lacks evidentiary constraints and consistency checks, and is difficult to generate traceable evidence chains.

Method used

By acquiring evidence fragments from multiple sources, constructing a knowledge graph and extracting entities, generating event nodes and causal edges, calling a large language model to generate causal assertions and performing consistency verification, constructing a traceable causal evidence chain, and combining temporal logic, source reliability, and entity association tightness to calculate confidence and conduct event causal reasoning analysis.

Benefits of technology

It enables the construction of reasonable and maintainable causal structures in multi-source texts, generating traceable chains of evidence, facilitating manual review and auditing, and improving the accuracy and timeliness of causal reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121960796B_ABST
    Figure CN121960796B_ABST
Patent Text Reader

Abstract

The application discloses an event causal reasoning method and electronic equipment for intelligence analysis. The method comprises the following steps: acquiring multiple evidence segments about query text from multiple sources, taking the multiple evidence segments as an evidence segment set; extracting entities from the evidence segments and linking the entities into a knowledge graph; extracting event elements from the evidence segments, generating event nodes and causal edges according to the event elements, and generating one or more candidate causal edges based on the events; searching for one or more binding evidences of each candidate causal edge from the evidence segment set; calling a large language model to generate a causal assertion for each candidate causal edge, performing consistency verification on the causal assertion and the binding evidence of the same candidate causal edge, combining into a traceable causal evidence chain, and performing event causal reasoning analysis. The application constructs event-level representation and generates an "event-event" candidate causal relationship, so that the reasoning object is upgraded from "entity co-occurrence" to an event causal structure that can be reasoned and maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligence analysis and intelligent reasoning, and in particular to an event causal reasoning method, electronic device, storage medium and computer program product for intelligence analysis. Background Technology

[0002] In intelligence analysis scenarios, information needs to be obtained from multiple sources. For example, in risk warning and public opinion analysis, analysts need to quickly retrieve relevant evidence from materials such as news, announcements, industry reports, social media, and internal notices, extract and align event elements, and then determine whether there is a causal relationship between events and its impact path. This type of analysis not only requires outputting conclusions, but also requires binding key assertions with verifiable evidence fragments and providing a traceable chain of evidence. At the same time, because information sources are diverse and facts continue to evolve over time, the system also needs to be able to continuously verify and correct existing causal judgments when new evidence arrives, in order to reduce the risk of misjudgment and improve the stability of results.

[0003] Current analytical processes primarily rely on manual reading, summarization, and cross-verification: analysts need to sift through large amounts of text, align different statements about the same event, identify potential causal clues, and organize the chain of evidence. While this approach offers some controllability, it suffers from low efficiency, high cost, and difficulty in scaling. Furthermore, under time pressure, it is difficult for humans to track and dynamically review the continuous influx of new evidence over the long term, leading to omissions, inconsistent conclusions, and difficulties in maintaining the chain of evidence.

[0004] To improve automation, existing technologies have proposed various machine learning-based methods. One type of method recalls and summarizes relevant materials based on keyword matching, topic clustering, or similarity retrieval; another type utilizes named entity recognition and entity linking to construct knowledge graphs, enabling cross-document aggregation and relationship queries. While these methods have some effectiveness in information retrieval and entity alignment, they primarily rely on "relevance" or "entity co-occurrence," often lacking stable modeling of event-level representations and event-to-event causal relationships. Furthermore, in situations where multi-source conflicts, clarifications, and counter-evidence are prevalent, they lack mechanisms for supporting / counter-evidence analysis and consistency verification of candidate causal relationships. This can easily solidify one-sided evidence into erroneous causal chains and makes it difficult to output traceable evidence chains that meet review requirements.

[0005] In recent years, large language models and generative artificial intelligence have developed rapidly, demonstrating outstanding performance in long text understanding, abstract induction, and natural language reasoning, providing new technical pathways for event analysis. Utilizing large language models to summarize, interpret, and generate reports from multi-source materials can reduce the burden of manual processing and improve expression efficiency. However, general-purpose large language models still suffer from insufficient adaptability when directly applied to tasks in areas such as risk warning and public opinion analysis.

[0006] (1) When there is insufficient evidence or conflicting materials, the model may generate a causal explanation that “seems reasonable but is unverifiable” and lacks evidentiary constraints on key assertions.

[0007] (2) They usually lack the ability to explicitly organize and present supporting and rebuttal evidence, and find it difficult to provide supporting evidence, rebuttal evidence and conflict points in their output at the same time.

[0008] (3) As new evidence continues to arrive, early conclusions may be overturned by authoritative clarification or counter-evidence. The general model lacks dynamic consistency verification, confidence update and conflict governance closed loop for evidence evolution, making it difficult for conclusions to self-correct as evidence changes.

[0009] More importantly, existing "knowledge graph + big model" solutions mostly adopt a fixed number of rounds or rule-triggered retrieval process, which makes it difficult to adaptively determine the next evidence collection action based on the "current evidence status". This often results in either insufficient retrieval leading to incomplete evidence or blindly conducting multiple rounds of retrieval, leading to excessive costs. In budget-constrained scenarios, it is difficult to balance coverage, accuracy and timeliness.

[0010] The aforementioned problems have hampered the accuracy, interpretability, traceability, and timeliness of multi-source event analysis tasks. Summary of the Invention

[0011] Therefore, it is necessary to address the technical problems existing in the current technology for multi-source event analysis tasks by providing an event causal reasoning method, electronic device, storage medium, and computer program product for intelligence analysis.

[0012] This invention provides a method for causal reasoning of events for intelligence analysis, comprising:

[0013] Obtain multiple pieces of evidence about the query text from multiple sources, and treat the multiple pieces of evidence as a set of evidence fragments;

[0014] Entities are extracted from the evidence fragments, and the entities are linked into a knowledge graph. Each evidence fragment is associated with and stored as an extracted entity.

[0015] Extract event elements from the evidence fragments, and generate event nodes and causal edges based on the event elements. Take the two event nodes connected by the causal edge as event pairs, and generate one or more candidate causal edges based on the event pairs.

[0016] Based on the knowledge graph, one or more binding evidences for each candidate causal edge are retrieved from the evidence fragment set;

[0017] For each candidate causal edge, a large language model is invoked to generate a causal assertion, and the consistency of the causal assertion and the binding evidence of the same candidate causal edge is verified. The causal assertion that satisfies the consistency verification and the corresponding binding evidence are combined into a traceable causal evidence chain, and event causal reasoning analysis is performed based on the traceable causal evidence chain.

[0018] Furthermore, generating one or more candidate causal edges based on the event pairs includes:

[0019] For any pair of events, perform a temporal feasibility assessment on all causal edges between the two event nodes in the pair, and remove causal edges that do not satisfy the temporal logic. The temporal logic is as follows: ,in, To relax the time window, This refers to the occurrence time of the causal event node in the event pair. This refers to the occurrence time of the result event node in the event pair;

[0020] The initial confidence level of the causal edges after elimination is calculated based on the reliability of the source, the strength of the causal trigger, the tightness of the entity association, and the time consistency. The top M causal edges with the highest initial confidence level in the result event node of each event pair are retained as candidate causal edges, where M is an integer greater than 1.

[0021] Furthermore, based on the knowledge graph, one or more binding evidences for each candidate causal edge are retrieved from the set of evidence fragments, including:

[0022] For each of the candidate causal edges, perform the following retrieval operation:

[0023] From the knowledge graph, the entity nodes of all participants involved in the cause events and result events connected by the candidate causal edges are taken as the core entity set of the candidate causal edges;

[0024] Obtain the entity nodes of each entity node in the core entity set at a preset level in the knowledge graph, as an extended entity set;

[0025] Construct a query expression to retrieve one or more first candidate evidences from the evidence fragment set. The query expression includes event keywords and the names of entity nodes in the core entity set and the extended entity set. The event keywords are the keywords of the target event included in the query text.

[0026] The first candidate evidence that meets the time requirement is filtered out as the coarse search candidate evidence;

[0027] The coarse search candidate evidence is converted into evidence vectors, the query text is converted into query vectors, and the vector similarity between each evidence vector and the query vector is calculated.

[0028] For each coarse search candidate evidence, a fine search ranking score is calculated based on vector similarity, evidence quality score, and diversity factor. The top K coarse search candidate evidences with the highest scores are selected as the binding evidence for the candidate causal edge.

[0029] Further, the step of generating a causal assertion by calling a large language model for each candidate causal edge, and performing a consistency check on the binding evidence between the causal assertion and the same candidate causal edge, and combining the causal assertion that satisfies the consistency check with the corresponding binding evidence into a traceable causal evidence chain, includes:

[0030] For each of the candidate causal edges, perform the following operation:

[0031] The large language model is invoked to generate causal assertions about the candidate causal edges, and supporting and contradictory evidence about the causal assertions is determined in the binding evidence of the same candidate causal edge.

[0032] Based on supporting and rebuttal evidence, calculate the support score, conflict score, and coverage of the causal assertion;

[0033] The causal assertion that the support score, conflict score, and / or coverage rate satisfy the consistency judgment is combined with the corresponding binding evidence, support score, conflict score, and coverage rate to form a traceable causal evidence chain.

[0034] Furthermore, the calculation of the support score, conflict score, and coverage of the causal assertion includes:

[0035] Computational support is divided into:

[0036]

[0037] in, Indicates support for the division. This represents the weight of the i-th binding evidence. This represents the strength of the large language model's judgment on the ith binding evidence supporting the causal assertion, wherein the binding evidence includes the supporting evidence and the evidence against the contrary;

[0038] Calculation conflicts are categorized as follows:

[0039]

[0040] in, Indicates conflict points, This represents the strength of the large language model's judgment on the i-th binding evidence against the causal assertion;

[0041] The calculated coverage is:

[0042]

[0043] in, Indicates coverage rate. This indicates the number of key elements of the causal assertion covered by all binding evidence. This represents all the key elements of the stated causal assertion.

[0044] Furthermore, it also includes:

[0045] In response to the updated binding evidence of the candidate causal edge, the support score and conflict score are recalculated.

[0046] As the support score increases, the confidence level of the candidate causal edge is increased;

[0047] When the conflict score increases or the support score decreases, the confidence level of the candidate causal edge is reduced.

[0048] Furthermore, it also includes:

[0049] For each causal assertion that does not meet the consistency judgment, multiple rounds of evidence collection scheduling operations are performed until the termination condition is met and the traceable causal evidence chain is output. Based on the traceable causal evidence chain, event causal reasoning analysis is performed. In each round of evidence collection scheduling operations, the following operations are performed:

[0050] Perform the selected action;

[0051] The evidence status is determined after the selected action is performed, and the evidence status is as follows: ,in, To support the score, To divide into conflict, For coverage, For the diversity index of sources, To accumulate the cost of evidence collection;

[0052] Based on the evidence status, select the action for the next round of evidence collection scheduling operation from the action set;

[0053] Determine whether the termination condition is met. If the termination condition is met, output the traceable causal evidence chain and perform event causal reasoning analysis based on the traceable causal evidence chain. Otherwise, execute the next round of evidence collection scheduling operation.

[0054] This invention provides an electronic device, comprising:

[0055] At least one processor; and,

[0056] A memory communicatively connected to at least one of the processors; wherein,

[0057] The memory stores instructions that are executed by at least one of the processors to enable the at least one of the processors to perform the event causal reasoning method for intelligence analysis as described above.

[0058] The present invention provides a storage medium that stores computer instructions, which, when executed by a computer, are used to perform all steps of the event causal reasoning method for intelligence analysis as described above.

[0059] The present invention provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the event causal reasoning method for intelligence analysis as described above.

[0060] This invention, based on a knowledge graph, retrieves one or more binding pieces of evidence for each candidate causal edge from a set of evidence fragments. For each candidate causal edge, a large language model is invoked to generate a causal assertion. Consistency checks are performed on the causal assertions and the binding evidence for the same candidate causal edge. Causal assertions that satisfy the consistency check are combined with their corresponding binding evidence to form a traceable causal evidence chain for event causal reasoning analysis. Therefore, based on entity alignment of multi-source texts, this invention constructs an event-level representation and generates "event-event" candidate causal relationships, elevating the reasoning object from "entity co-occurrence" to a reasonable and maintainable event causal structure. Simultaneously, evidence citation constraints and structured assertion outputs form a traceable evidence chain, facilitating manual review and auditing. Attached Figure Description

[0061] Figure 1 This is a flowchart illustrating the process of an event causal reasoning method for intelligence analysis according to an embodiment of the present invention.

[0062] Figure 2 This is a flowchart illustrating a method for causal reasoning of events for intelligence analysis, according to another embodiment of the present invention.

[0063] Figure 3 This is a schematic diagram of a multi-round search according to the preferred embodiment of the present invention;

[0064] Figure 4 A flowchart illustrating the workflow of an event causal reasoning method for intelligence analysis, representing a preferred embodiment of the present invention;

[0065] Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to the present invention. Detailed Implementation

[0066] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings. Identical components are indicated by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "upper," and "lower" used in the following description refer to directions in the drawings, while the terms "inner" and "outer" refer to directions toward or away from the geometric center of a specific component. These terms are used only for the convenience of describing this application and simplifying the description, and are not intended to indicate or imply that the device or element referred to has a specific orientation, or is constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0067] Unless the context otherwise requires, throughout the specification and claims, the term "comprising" is interpreted as open-ended and encompassing, meaning "including, but not limited to." In the description of the specification, terms such as "one embodiment," "some embodiments," "exemplary embodiment," "exemplary," or "some examples," etc., are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this application. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics mentioned may be included in any suitable manner in any one or more embodiments or examples.

[0068] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this application, unless otherwise stated, "a plurality of" means two or more.

[0069] In describing some embodiments, the term "connection" and its derivative expressions may be used. For example, the term "connection" may be used in describing some embodiments to indicate that two or more components have direct physical or electrical contact with each other. The embodiments claimed herein are not necessarily limited to the content of this document.

[0070] In addition, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0071] like Figure 1 The diagram shown is a flowchart of an event causal reasoning method for intelligence analysis according to an embodiment of the present invention, including:

[0072] Step S101: Obtain multiple evidence fragments about the query text from multiple sources, and treat the multiple evidence fragments as a set of evidence fragments;

[0073] Step S102: Extract entities from the evidence fragments and link the entities into a knowledge graph. Each evidence fragment is associated with and stored as an extracted entity.

[0074] Step S103: Extract event elements from the evidence fragment, generate event nodes and causal edges based on the event elements, take the two event nodes connected by the causal edge as event pairs, and generate one or more candidate causal edges based on the event pairs.

[0075] Step S104: Based on the knowledge graph, retrieve one or more binding evidences for each candidate causal edge from the evidence fragment set;

[0076] Step S105: For each candidate causal edge, call the large language model to generate a causal assertion, and perform consistency verification on the binding evidence of the causal assertion and the same candidate causal edge. Combine the causal assertion that satisfies the consistency verification with the corresponding binding evidence to form a traceable causal evidence chain, and perform event causal reasoning analysis based on the traceable causal evidence chain.

[0077] Specifically, the present invention can be applied to electronic devices with processing capabilities, such as computers.

[0078] This invention takes multi-source intelligence text as input, constructs event-level representation and event-event candidate causal relationships based on entity alignment, and ultimately achieves credible event causal reasoning through structured causal assertion generation and evidence citation using a large language model.

[0079] First, step S101 is executed, which involves obtaining multiple evidence fragments about the query text from multiple sources and using these multiple evidence fragments as a set of evidence fragments.

[0080] Specifically, multi-source data acquisition and preprocessing of the query text are performed. Data is accessed from multiple sources such as news, announcements, reports, and social media. The data is deduplicated, denoised, language standardized, and time expression normalized. It is also segmented into sentences and paragraphs to form a set of evidence fragments with source and time annotations.

[0081] The query text is the question asked by the user.

[0082] In some embodiments, the query text is text describing the causal relationship between the cause event and the result event that the user is querying.

[0083] Then, step S102 is executed to extract entities from the evidence fragments and link the entities into a knowledge graph, with each evidence fragment associated with and saved as an extracted entity.

[0084] Specifically, entity alignment and entity linking knowledge graph updates are performed: entity extraction, disambiguation, and entity linking are performed on evidence fragments, entity mentions are mapped to unified entity identifiers (IDs), and an entity linking knowledge graph including multiple entity nodes is constructed or incrementally updated based on the entity IDs and their mention relationships to form cross-document alignment capabilities.

[0085] Then, step S103 is executed to extract event elements from the evidence fragment, and generate event nodes and causal edges based on the event elements. The two event nodes connected by the causal edge are taken as event pairs, and one or more candidate causal edges are generated based on the event pairs.

[0086] Specifically, the event layer is constructed and candidate causal edges are generated: event elements are extracted from evidence fragments using a large language model to generate event candidates, and events with the same reference are merged to form event nodes; causal edges are constructed about the event nodes, and the two event nodes connected by the causal edges are taken as event pairs. One or more candidate causal edges about the event pairs are generated on the event node set, and the candidate causal edges are taken as a candidate causal edge set. Initial confidence scores are calculated for the candidate causal edges for ranking and subsequent allocation of evidence collection resources.

[0087] Then, step S104 is executed, based on the knowledge graph, to retrieve one or more binding evidences for each candidate causal edge from the evidence fragment set.

[0088] Specifically, enhanced evidence collection (coarse search + fine search) is performed: coarse and fine searches are conducted on the target event or candidate causal edge to obtain a multi-source evidence set. The coarse search aims to recall evidence, while the fine search outputs Top-K highly relevant evidence fragments as binding evidence for the candidate causal edge, and outputs their evidence span and position pointers. The target event is the event indicated by the query text entered by the user, including causal events and result events.

[0089] Then, step S105 is executed, where a large language model is called to generate a causal assertion for each candidate causal edge, and the consistency of the causal assertion and the binding evidence of the same candidate causal edge is verified. The causal assertion that satisfies the consistency verification and the corresponding binding evidence are combined into a traceable causal evidence chain, and event causal reasoning analysis is performed based on the traceable causal evidence chain.

[0090] Specifically, the evidence constraint inference and consistency verification based on the large model are as follows: Based on the binding evidence for each candidate causal edge, a large language model is invoked to generate a structured causal claim. The large model is then used to determine whether the causal edge is valid, which is equivalent to performing pairwise judgments on each bound piece of evidence and the target event. Simultaneously, cited evidence is forcibly bound; consistency verification is performed on the causal claim and cited evidence to obtain support score, conflict score, and coverage index. A traceable causal evidence chain containing supporting evidence chains and disproving / conflict location information is constructed. Finally, the large language model is invoked, and the traceable causal evidence chain is combined into prompt words and input into the large language model, instructing the large language model to perform event causal reasoning analysis based on the traceable causal evidence chain.

[0091] Among them, causal assertions are structured statements with explicit causal semantics extracted from the binding evidence of candidate causal edges by the large model, and their general form is as follows:

[0092] (cause event, causal relation, effect event), where cause event represents the causal event, effect event represents the result event, and causal relation (with direction) represents the causal relationship.

[0093] This invention, based on a knowledge graph, retrieves one or more binding pieces of evidence for each candidate causal edge from a set of evidence fragments. For each candidate causal edge, a large language model is invoked to generate a causal assertion. Consistency checks are performed on the causal assertions and the binding evidence for the same candidate causal edge. Causal assertions that satisfy the consistency check are combined with their corresponding binding evidence to form a traceable causal evidence chain for event causal reasoning analysis. Therefore, based on entity alignment of multi-source texts, this invention constructs an event-level representation and generates "event-event" candidate causal relationships, elevating the reasoning object from "entity co-occurrence" to a reasonable and maintainable event causal structure. Simultaneously, evidence citation constraints and structured assertion outputs form a traceable evidence chain, facilitating manual review and auditing.

[0094] like Figure 2 The diagram shown is a flowchart of an event causal reasoning method for intelligence analysis according to another embodiment of the present invention, including:

[0095] Step S201: Obtain multiple evidence fragments about the query text from multiple sources, and treat the multiple evidence fragments as a set of evidence fragments.

[0096] Step S202: Extract entities from the evidence fragments and link the entities into a knowledge graph. Each evidence fragment is associated with and stored with the extracted entities.

[0097] Step S203: Extract event elements from the evidence fragment, and generate event nodes and causal edges based on the event elements. For any event pair, treat the two event nodes connected by the causal edge as an event pair. Perform a temporal feasibility judgment on all causal edges between the two event nodes in the event pair, and remove causal edges that do not satisfy the temporal logic. The temporal logic is... ,in, To relax the time window, This refers to the occurrence time of the causal event node in the event pair. This refers to the occurrence time of the result event node in the event pair;

[0098] The initial confidence level of the causal edges after elimination is calculated based on the reliability of the source, the strength of the causal trigger, the tightness of the entity association, and the time consistency. The top M causal edges with the highest initial confidence level in the result event node of each event pair are retained as candidate causal edges, where M is an integer greater than 1.

[0099] Step S204: For each candidate causal edge, perform the following retrieval operation:

[0100] From the knowledge graph, the entity nodes of all participants involved in the cause events and result events connected by the candidate causal edges are taken as the core entity set of the candidate causal edges;

[0101] Obtain the entity nodes of each entity node in the core entity set at a preset level in the knowledge graph, as an extended entity set;

[0102] Construct a query expression to retrieve one or more first candidate pieces of evidence from the evidence fragment set. The query expression includes event keywords and the names of entity nodes in the core entity set and the extended entity set.

[0103] The first candidate evidence that meets the time requirement is filtered out as the coarse search candidate evidence;

[0104] The coarse search candidate evidence is converted into evidence vectors, the query text is converted into query vectors, and the vector similarity between each evidence vector and the query vector is calculated.

[0105] For each coarse search candidate evidence, a fine search ranking score is calculated based on vector similarity, evidence quality score, and diversity factor. The top K coarse search candidate evidences with the highest scores are selected as the binding evidence for the candidate causal edge.

[0106] Step S205: For each of the candidate causal edges, perform the following operation:

[0107] The large language model is invoked to generate causal assertions about the candidate causal edges, and supporting and contradictory evidence about the causal assertions is determined in the binding evidence of the same candidate causal edge.

[0108] Based on supporting and rebuttal evidence, calculate the support score, conflict score, and coverage of the causal assertion;

[0109] The causal assertion that satisfies the consistency judgment of the support score, the conflict score and / or the coverage rate is combined with the corresponding binding evidence, support score, conflict score and coverage rate to form a traceable causal evidence chain, and event causal reasoning analysis is performed based on the traceable causal evidence chain.

[0110] Step S206: In response to the updated binding evidence of the candidate causal edge, recalculate the support score and conflict score;

[0111] As the support score increases, the confidence level of the candidate causal edge is increased;

[0112] When the conflict score increases or the support score decreases, the confidence level of the candidate causal edge is reduced.

[0113] This embodiment presents a multi-round forensics method for intelligence analysis, combining a large language model with reinforcement learning. The method takes multi-source intelligence text as input, constructs event-level representations and event-event candidate causal relationships based on entity alignment, generates structured causal assertions and cites evidence through a large language model, and further, allows for multi-round forensics scheduling driven by reinforcement learning strategies. Under limited retrieval budgets, it adaptively selects actions such as in-depth retrieval, counter-evidence retrieval, and source-switching retrieval. Simultaneously, it incorporates consistency checks and confidence governance mechanisms to incrementally update and freeze / unfreeze causal edges, thereby outputting verifiable, traceable, and dynamically correctable event causal conclusions and their evidence chains.

[0114] Specifically, step S201 is first executed, which involves obtaining multiple evidence fragments about the query text from multiple sources and using these multiple evidence fragments as a set of evidence fragments.

[0115] Specifically, multi-source data acquisition and preprocessing of the query text are performed. Data is accessed from multiple sources such as news, announcements, reports, and social media. The data is deduplicated, denoised, language standardized, and time expression normalized. It is also segmented into sentences and paragraphs to form a set of evidence fragments with source and time annotations.

[0116] Then, step S202 is executed to extract entities from the evidence fragments and link the entities into a knowledge graph, with each evidence fragment associated with and saved as an extracted entity.

[0117] Specifically, entity alignment and entity linking knowledge graph updates are performed: entity extraction, disambiguation, and entity linking are performed on evidence fragments, entity mentions are mapped to unified entity identifiers (IDs), and an entity linking knowledge graph including multiple entity nodes is constructed or incrementally updated based on the entity IDs and their mention relationships to form cross-document alignment capabilities.

[0118] Entity extraction and linking can be performed based on a large language model. The large language model is invoked to perform entity extraction on the evidence fragment. Based on a preset entity type system, entities that appear explicitly or implicitly in the evidence fragment are identified to obtain entity candidates and their corresponding entity category information. The entity categories preferably include, but are not limited to, one or more of the following: people, organizations, locations, products, event names, rule / clause numbers, and behavioral objects.

[0119] Among them, the large language model can combine contextual semantic understanding capabilities to perform semantic analysis on complex text situations such as colloquial expressions, ellipsis, and polysemous words, so as to improve the completeness and accuracy of entity recognition.

[0120] The entity extraction results are further output as standardized results for the entities, which preferably include:

[0121] The entity's standardized name, historical aliases or common abbreviations, parsable key attribute information (such as time attributes, role attributes, affiliation attributes, etc.), and the entity's reference position or text span marker in the evidence fragment. By recording the corresponding positional relationship between the entity and the original evidence fragment, it provides basic support for subsequent traceability analysis and evidence backtracking.

[0122] Based on this, and combining the historical entity database, alias mapping table, contextual constraints, and entity disambiguation rules, entity disambiguation and entity linking are performed on the candidate entities. Entity disambiguation is used to distinguish entities with the same name but different referents, and entity linking is used to map multiple mentions of the same entity in different documents, different expressions, or different time segments to a unified entity identifier (entity ID), thereby achieving cross-document and cross-source entity alignment capabilities.

[0123] Then, incremental construction and deduplication updates of the entity link knowledge graph are performed. Based on the entity IDs obtained above and their mention relationships in the evidence fragments, the entity link knowledge graph is incrementally constructed and updated. The update includes the following processing:

[0124] (1) Deduplication of evidence: Deduplication is performed on newly added evidence fragments and historical evidence fragments. Deduplication is preferably performed by one or more of the following: content deduplication and reprint deduplication.

[0125] The newly added evidence fragments and historical evidence fragments are subjected to a duplicate determination. The deduplication process preferably includes one or more of the following: content deduplication and reprint deduplication.

[0126] Among them, content deduplication is based on the cosine similarity threshold of text vector representation to determine highly similar text;

[0127] Reprint deduplication is used to identify instances where the same fact is reprinted or aggregated on different platforms or from different sources. Multiple reprinted evidences are merged and marked, with the original source or authoritative source being preferred as the primary evidence, and the remaining evidence being marked as reprinted evidence, in order to avoid the same fact being counted repeatedly in subsequent analysis.

[0128] (2) Entity node update: When a new entity does not exist in the knowledge graph, a new entity node is created and the corresponding entity ID is assigned; when a new entity is determined to be the same entity as an existing entity after standard name matching, alias matching or disambiguation judgment, the new entity information is merged into the corresponding entity node, and the alias set, entity category label and related statistical attribute information of the entity node are updated synchronously.

[0129] (3) Reference relationship update: The reference relationship between evidence fragments and entities is written into the knowledge graph to form the entity-evidence fragment reference edge, and the corresponding source identifier, timestamp and evidence fragment position pointer information are recorded to support subsequent entity-based traceable retrieval, evidence playback and reasoning analysis.

[0130] (4) Conflict marker (preferred): When the entity attribute information involved in the new evidence conflicts with the existing entity attributes, it is preferred to record the conflict marker and retain the multiple version attribute values ​​or the location of the conflict for subsequent consistency verification, credibility assessment or manual review process.

[0131] Output: The updated entity link knowledge graph and its index structure. The output preferably includes a set of entity nodes, a set of entity aliases, a set of entity-evidence fragment mention relationships, and an inverted index structure from entity ID to evidence fragment.

[0132] Then, step S203 is executed: event elements are extracted from the evidence fragment, and event nodes and causal edges are generated based on the event elements. Two event nodes connected by a causal edge are considered as event pairs. For any event pair, the temporal feasibility of all causal edges between the two event nodes in the event pair is judged, and causal edges that do not satisfy the temporal logic are removed. The temporal logic is... ,in, To relax the time window, This refers to the occurrence time of the causal event node in the event pair. This refers to the occurrence time of the result event node in the event pair;

[0133] The initial confidence level of the causal edges after elimination is calculated based on the reliability of the source, the strength of the causal trigger, the tightness of the entity association, and the time consistency. The top M causal edges with the highest initial confidence level in the result event node of each event pair are retained as candidate causal edges, where M is an integer greater than 1.

[0134] Specifically, the event layer is constructed and candidate causal edges are generated: event elements are extracted from evidence fragments using a large language model to generate event candidates, and events with the same reference are merged to form event nodes; causal edges are constructed about the event nodes, and the two event nodes connected by the causal edges are taken as event pairs. One or more candidate causal edges about the event pairs are generated on the event node set, and the candidate causal edges are taken as a candidate causal edge set. Initial confidence scores are calculated for the candidate causal edges for ranking and subsequent allocation of evidence collection resources.

[0135] Specifically, a multi-constraint screening model for candidate causal edge generation is constructed:

[0136] The generation of candidate causal edges is not a simple fully connected process, but rather a selection and ranking process based on the following constraint model: When constructing the set of event-level causal candidate edges, a two-stage strategy of "feasibility filtering first, confidence estimation second, and ranking and truncation third" is preferred to avoid combinatorial explosion when the event scale is large, while ensuring that candidate edges entering the subsequent retrieval and evidence collection stages have reasonable prior credibility and interpretability. Specifically, for any event pair... Candidate causal directions ( First, a temporal feasibility assessment is performed, and candidate relationships that do not meet the basic temporal logic are directly eliminated; then, the initial confidence level of the candidate relationships that pass the temporal constraints is calculated by fusing multiple features. Based on this, the candidate edges are sorted and their size is controlled.

[0137] Timing constraint functions:

[0138] Set event The time of occurrence is ,event The time of occurrence is .

[0139] Define the timing feasibility indicator function:

[0140] If and only if .in This is a relaxation time window used to handle reporting delays or impact lags; it is typically set to... Relationships that do not meet this condition are filtered out directly.

[0141] More preferably, and The values ​​can come from: explicit time expressions in the evidence (such as "January 3rd", "last night", "this week"), time fields in structured data, and timestamps obtained through a time standardization module. For cases with ambiguous or range-based time expressions, it is preferable to use... It is represented as a comparable standardized time point or time interval, and a conservative strategy is used to determine the representative value used for comparison in order to reduce false filtering caused by time sampling errors.

[0142] By introducing The relaxed window allows the system to be more robust to real-world scenarios such as "events occurring before reporting," "delayed transmission of policy impacts," and "cross-regional diffusion": even if events... The report was later than As long as it is allowed to enter the candidate set within the relaxation window, its credibility can be further distinguished by subsequent feature scoring.

[0143] This temporal constraint function serves as a hard gate for candidate edge generation, and its output... Can be directly used as a filtering condition: when When this happens, the subsequent confidence of the event pair is no longer calculated and it is removed from the candidate edge set, thereby reducing the computational overhead of invalid edges.

[0144] Then, the initial confidence level of the fusion is calculated for the removed causal edges based on the source reliability, causal trigger strength, entity association tightness, and time consistency.

[0145] Specifically, initial confidence level The calculation integrates multiple features, and a linear weighted model is preferred:

[0146]

[0147] in, , , , The weights are non-negative and sum to 1, and can be fitted from historical data or set by domain experts.

[0148] Parameter description:

[0149] Source reliability. The weighting is assigned based on factors such as source type, historical accuracy, whether it's a first publication / reprint, and whether there is cross-verification from multiple sources. For example, authoritative sources (government / regulatory / central bank announcements, etc.) receive higher weighting, while social media or anonymous sources receive lower weighting. When the same event is supported by multiple independent sources, a weighted aggregation based on source reliability can be used to improve prior credibility.

[0150] Causal trigger strength. Preferably, it is calculated based on the frequency, position, and syntactic / semantic distance of causal trigger words (such as "cause," "initiate," "lead to," "therefore," "because," etc.) in the evidence. Furthermore, attenuation or negative weights can be applied to negative expressions (such as "did not cause," "not because of...") or hypothetical expressions (such as "may cause," "may trigger") to reduce false triggers.

[0151] The strength of causal cue words between event pairs in a text is measured; the more cue words, the more obvious the causal relationship.

[0152] First, construct a causal trigger word lexicon T, which includes, but is not limited to, “cause,” “initiate,” “lead to,” “cause,” “make,” “because,” “therefore,” and “thus.”

[0153] For example, focus on the frequency of occurrence. In the event and Count the number of times trigger words appear in the text between them:

[0154]

[0155] Entity Relationship Depth. Preferred calculations are based on the number of entities shared by events u and v, and the importance of these entities. An entity weight table or centrality indicator can be introduced to assign different weights to institutions, policy actors, key assets, geographical regions, etc. This feature encourages event pairs on the same entity / policy target / regional chain to be prioritized in the candidate set.

[0156] In one embodiment, entity affinity This measure assesses the semantic relevance of event pairs at the entity level. It can be calculated solely based on the similarity of the sets of entities contained within the event pair, for example, using Jaccard similarity.

[0157] First, from the event and Extract the corresponding entity sets from each:

[0158]

[0159]

[0160] Next, the similarity between the two entity sets is calculated. ,

[0161]

[0162] Time consistency. Ideally, under the premise of satisfying relaxed time sequence constraints, the time difference Δt between two events is mapped to a decay score in the interval (0,1]. Generally, the smaller Δt is, the higher the score. At the same time, the decay rate can be adjusted through domain parameters to avoid over-penalizing true long-lag causal relationships. If there is no time in the intelligence news, the release time can be used as the default.

[0163] Finally, sorting and size control are performed:

[0164] For all candidate edges that satisfy the temporal constraints, according to Sort in descending order.

[0165] Implement a Top-M truncation strategy (e.g., M=20) to retain only the M candidate cause edges with the highest confidence for each outcome event v, in order to control the computational cost of subsequent retrieval and evidence collection.

[0166] Calculate the initial confidence level for each candidate causal edge. This is used for candidate edge sorting and retrieval resource allocation. The initial confidence level is preferably obtained by integrating factors such as source reliability, causal triggering strength, temporal consistency, and multi-source consistency, and is then... Write the causal edge attribute as the initial confidence level.

[0167] Then, step S204 is executed, and for each candidate causal edge, the following retrieval operation is performed:

[0168] From the knowledge graph, the entity nodes of all participants involved in the cause events and result events connected by the candidate causal edges are taken as the core entity set of the candidate causal edges;

[0169] Obtain the entity nodes of each entity node in the core entity set at a preset level in the knowledge graph, as an extended entity set;

[0170] Construct a query expression to retrieve one or more first candidate evidences from the evidence fragment set. The query expression includes event keywords and the names of entity nodes in the core entity set and the extended entity set. The event keywords are the keywords of the target event included in the query text.

[0171] The first candidate evidence that meets the time requirement is filtered out as the coarse search candidate evidence;

[0172] The coarse search candidate evidence is converted into evidence vectors, the query text is converted into query vectors, and the vector similarity between each evidence vector and the query vector is calculated.

[0173] For each coarse search candidate evidence, a fine search ranking score is calculated based on vector similarity, evidence quality score, and diversity factor. The top K coarse search candidate evidences with the highest scores are selected as the binding evidence for the candidate causal edge.

[0174] Specifically, enhanced evidence collection (coarse search + fine search) is performed: coarse search and fine search are performed on the target event or candidate causal edge to obtain a multi-source evidence set. The coarse search aims to recall evidence, while the fine search outputs Top-K highly relevant evidence fragments as binding evidence for candidate causal edges and outputs their evidence span and position pointer.

[0175] Among them, in the candidate causal edge Before conducting evidence collection, it is preferable to use an entity expansion recall mechanism based on knowledge graphs to perform coarse retrieval, so as to avoid insufficient recall caused by relying solely on keyword retrieval and improve the coverage of implicitly relevant evidence.

[0176] Core entity set construction:

[0177] Given candidate causal edges Its core entity set Defined as an event With the event The union of all participating entities involved. These participating entities may include, but are not limited to: the event subject, the affected party, related institutions, policy entities, key assets, and geographical regions.

[0178] By building This can transform event-level causal hypotheses into structured retrieval clues of "entity-entity-event", providing anchor points for subsequent graph expansion and query construction.

[0179] Graph expansion algorithm:

[0180] On the entity link knowledge graph, Each entity in the array performs a K-top neighborhood expansion operation, typically taking... or To capture upstream and downstream entities that are directly or indirectly related to the core entity while controlling noise.

[0181] The types of relationships considered for expansion include, but are not limited to: synonym relationships, membership relationships, geographic location relationships, and participation relationships. By limiting the types of relationships, it is possible to avoid introducing overly generalized or weakly related entity nodes.

[0182] The expanded entity set The knowledge graph is sorted according to the strength of its relationship with the core entity. The strength of the relationship can be characterized by indicators such as the weight of the corresponding edge in the knowledge graph, the relationship confidence, or the historical co-occurrence frequency.

[0183] Furthermore, to prevent excessive expansion from introducing noise, a Top-N truncation strategy is preferred (e.g., For each core entity, only a few extended entities most closely associated with the core entity are retained, thereby achieving a balance between recall coverage and precision.

[0184] Query construction:

[0185] Will The entity specification name, alias, and event keywords in the query text are connected by an OR logical connection to construct a Boolean query expression for full-text search or index search. The event keywords are the keywords of the target events included in the query text. The target events include the cause events and result events in the query text.

[0186] At the same time, apply time window filtering to the timestamp index. ,

[0187] in and These are used to cover potential causal clues before an event occurs and to provide confirmatory reports or impacts after the event occurs, for example, to take... .

[0188] By combining entity expansion with time constraints, the coarse search stage can effectively limit the search space while ensuring the breadth of recall, providing a set of candidate evidence for subsequent fine ranking.

[0189] After a coarse search, the resulting evidence set typically still contains a certain proportion of weakly relevant or redundant text. To further improve the relevance and interpretability of the evidence, a Top-K selection based on word vector similarity is performed. Ideally, a fine-grained search and re-ranking mechanism based on vector similarity should be introduced on top of the coarse search results.

[0190] Vectorization representation: The text embedding model (Sentence-BERT) is used to vectorize the query text and evidence fragments in the coarse search result set, and semantic relevance is calculated based on vector similarity.

[0191] Similarity metric:

[0192] Calculate the cosine similarity between the query vector and the evidence vector as the basic relevance score:

[0193]

[0194] in, Represents the query vector. This refers to the evidence vector. The norm (modulus) of the query vector is preferably the Euclidean norm (also called the L2 norm). Let be the norm of the evidence vector, preferably the Euclidean norm (also called the L2 norm). By calculating vector similarity, semantic approximation relationships that are difficult to cover by traditional keyword matching can be captured, thereby improving the ability to identify implicit expressions, paraphrasing, and cross-sentence reasoning evidence.

[0195] Multi-factor reordering:

[0196] Final detailed search ranking score It does not only rely on vector similarity (sim), but also comprehensively considers the quality of evidence and the diversity of information: .

[0197] Evidence quality score is used to measure the credibility of evidence fragments. It is preferably calculated based on factors such as the authority of the source, the completeness of the text, whether it is an original report / first-hand disclosure, the time of publication, and the degree of verifiability.

[0198] Diversity factor: used to evaluate the incremental information value of candidate evidence relative to the selected evidence set. It encourages evidence from different independent sources or that provides different perspectives / mechanisms to reduce information redundancy and single-source bias caused by repeated citations.

[0199] λ1, λ2, and λ3 are harmonic weights used to balance semantic relevance, evidence quality, and information diversity. They can be set empirically or optimized based on historical cases.

[0200] Ultimately, according to The evidence fragments are sorted in descending order, and the top-K evidence fragments with the highest scores are selected to form a detailed search result set. This serves as input for subsequent causal verification, evidence chain construction, or analytical reasoning modules.

[0201] Step S205: For each of the candidate causal edges, perform the following operation:

[0202] The large language model is invoked to generate causal assertions about the candidate causal edges, and supporting and contradictory evidence about the causal assertions is determined in the binding evidence of the same candidate causal edge.

[0203] Based on supporting and rebuttal evidence, calculate the support score, conflict score, and coverage of the causal assertion;

[0204] The causal assertion that satisfies the consistency judgment of the support score, the conflict score and / or the coverage rate is combined with the corresponding binding evidence, support score, conflict score and coverage rate to form a traceable causal evidence chain, and event causal reasoning analysis is performed based on the traceable causal evidence chain.

[0205] Specifically, after completing the preliminary retrieval enhancement evidence collection and candidate evidence screening, for each candidate causal edge (causal relationship), the large language model outputs a causal assertion. Then, a consistency check is performed on the causal assertion generated by the large language model and its bound evidence set to quantify the reliability of the causal assertion under the support of existing evidence. Based on this, a structured causal evidence chain is constructed, and based on the traceable causal evidence chain, event causal reasoning analysis is performed.

[0206] Specifically, the assertions output by the large language model and their corresponding supporting and counter-evidence are modeled in a unified manner and quantified into three core indicators: support score, conflict score, and coverage rate, which are used for subsequent judgment and decision-making.

[0207] In one embodiment, calculating the support score, conflict score, and coverage of the causal assertion includes:

[0208] Computational support is divided into:

[0209]

[0210] in, Indicates support for the division. This represents the weight of the i-th binding evidence. This represents the strength of the large language model's judgment on the ith binding evidence supporting the causal assertion, wherein the binding evidence includes the supporting evidence and the evidence against the contrary;

[0211] Calculation conflicts are categorized as follows:

[0212]

[0213] in, Indicates conflict points, This represents the strength of the large language model's judgment on the i-th piece of evidence that contradicts the causal assertion;

[0214] The calculated coverage is:

[0215]

[0216] in, Indicates coverage rate. This indicates the number of key elements of the causal assertion covered by all binding evidence. This represents all the key elements of the stated causal assertion.

[0217] Specifically, support points The weighted summation of the quantity, strength, and source reliability of supporting evidence is used to characterize the degree to which an assertion is supported.

[0218] The support score is used to measure the overall strength of support of the evidence set for the causal assertion, taking into account the quantity of supporting evidence, the degree of support of the model judgment, and the reliability of the evidence source.

[0219] The support score is calculated using a weighted average method:

[0220]

[0221] in, Present evidence The weight of the evidence is related to factors such as the authority and credibility level of the evidence source; Indicating the large language model on the evidence Supports the strength of the current causal assertion.

[0222] More preferably, supporting evidence may include, but is not limited to: a clear statement of "the event". Caused the incident The direct evidence includes reports that explicitly express the causal direction through causal trigger words, and analytical texts that are highly consistent with causal assertions at both the temporal and entity levels.

[0223] By introducing evidence weight This can prevent evidence from low-credibility sources from "diluting" the supporting effect of evidence from high-credibility sources in terms of quantity, and make the support score more in line with the intuitive judgment of human experts on the credibility of evidence.

[0224] Conflict Conflict score is used to quantify the strength of evidence in the evidence set that refutes or provides alternative explanations for the causal assertion, reflecting the main risks to the internal consistency of the causal assertion.

[0225] In particular, conflict scores increase significantly when the following types of evidence are present:

[0226] Evidence that directly refutes causality, such as "the central bank denied rumors that the interest rate hike was not the main reason for the stock market decline";

[0227] Evidence providing strong alternative explanations, such as "the stock market decline was mainly due to international conflict," i.e., evidence for the outcome event. Give the main reasons for the discrepancy between the statement and the assertion.

[0228] The calculation method for conflict scores is similar to that for support scores, and evidence judged by the model as "negative", "weakening", or "alternative explanation" is weighted and aggregated.

[0229] By introducing a separate conflict sub-indicator, the situation of "more supporting evidence but key counter-evidence being ignored" can be avoided, enabling the system to explicitly identify inconsistencies or controversies within causal assertions.

[0230] Coverage The proportion of key elements in an assertion (subject, action, object, time, place, etc.) covered by evidence is used to measure the completeness of information.

[0231] Coverage is used to measure the completeness of the evidence set in a causal assertion, and is a measure of whether the evidence is "complete".

[0232] The coverage rate is defined as:

[0233]

[0234] in, Indicates coverage rate. This indicates the number of key elements of the causal assertion covered by all binding evidence. This represents all the key elements of the stated causal assertion.

[0235] Among them, key elements are preferably including but not limited to: subject, action, object, time, place and other core components of causal assertion; when the evidence can clearly correspond to or support a certain key element, it is considered that the element is covered.

[0236] More preferably, by introducing a coverage index, it is possible to distinguish between "evidence that only supports partial conclusions" and "a combination of evidence that covers the entire causal chain," thus avoiding causal judgments that are based solely on fragmented or incomplete information.

[0237] Decision thresholds and consistency assessment:

[0238] Based on the above three indicators, it is preferable to set a threshold value that supports segmentation. (e.g., 0.7), conflict score threshold (e.g., 0.4) and coverage threshold (e.g., 0.5), used for consistency determination of causal assertions.

[0239] When the following conditions are met:

[0240] , ,and ,

[0241] It is then considered that the causal relationship is strongly supported under the current evidence set and can be written into the causal evidence chain as a highly credible causal edge.

[0242] when If the causal assertion is found to be significantly conflicting, even if the support score is high, it is preferable to mark the causal relationship as "controversial" or "requires further verification" to avoid spreading potentially erroneous causal conclusions in subsequent analyses.

[0243] Output of causal evidence chain construction:

[0244] After completing the consistency check, the system preferably organizes the causal assertions that have passed the threshold judgment, along with their supporting evidence, conflicting evidence, and indicator calculation results, into a structured chain of causal evidence.

[0245] The causal evidence chain may include: causal edges ( Supporting evidence set, conflicting evidence set, and corresponding evidence set. , and Numerical values, as well as the source and citation location of evidence, enable the interpretable presentation and traceable verification of causal judgment results.

[0246] Step S206: In response to the updated binding evidence of the candidate causal edge, recalculate the support score and conflict score;

[0247] As the support score increases, the confidence level of the candidate causal edge is increased;

[0248] When the conflict score increases or the support score decreases, the confidence level of the candidate causal edge is reduced.

[0249] Specifically, causal edge confidence governance and incremental updates are performed: after candidate causal edges are updated and bound to evidence, the confidence of causal edges is incrementally updated based on the consistency verification results, and governance operations are performed on high-conflict or unstable causal edges. The governance operations include at least one or more of freezing, de-weighting, and unfreezing to suppress erroneous causal links and support continuous correction as new evidence arrives; the update and governance process records the triggering reasons and key evidence references to support review and auditing.

[0250] Specifically, after completing the consistency verification and the construction of the causal evidence chain, the system preferentially uses the consistency verification results to continuously update and manage the confidence of the causal edge, so as to reflect the dynamic reliability of the causal relationship in the constantly evolving evidence environment.

[0251] Specifically, based on changes in support and conflict scores, the confidence level of causal edges is incrementally updated: when supporting evidence strengthens, the support score increases, and the conflict score remains within an acceptable range, the confidence level of the corresponding causal edge is increased; when conflicting evidence strengthens, the conflict score increases, or the support score decreases, the confidence level of the causal edge is decreased accordingly. Simultaneously, the timestamp of the most recent consistency verification of the causal edge is preferably recorded to support subsequent timeliness analysis and governance decisions.

[0252] Furthermore, to prevent low-quality or unstable causal relationships from propagating in the system over a long period, governance operations are performed on causal edges with high conflict or large confidence fluctuations. These governance operations include at least one or more of freezing, deweighting, and unfreezing, used for fine-grained management of the usage status of causal edges.

[0253] Preferably, the freezing or de-weighting of causal edges is triggered when one or more of the following conditions occur:

[0254] Conflict Based on the calculation of the strength and authority of counter-evidence, special attention is paid to evidence that directly refutes the assertion or provides a strong alternative explanation.

[0255] Strong counter-evidence from authoritative sources emerges, substantially negating existing causal assertions;

[0256] The continuous decline in the confidence level of the causal side, falling below the preset minimum acceptable threshold, indicates that the causal relationship is not stable enough under the current evidence environment.

[0257] In the frozen or deweighted state, the causal edge selection no longer participates in subsequent automatic reasoning, decision recommendation or causal chain expansion processes, or only participates with a significantly reduced weight, so as to avoid misleading the overall analysis results.

[0258] Accordingly, when the subsequent evidence environment changes and the following conditions are met, it is preferable to trigger the unfreezing operation on frozen or downgraded causal edges:

[0259] Multiple sources of evidence consistently and stably support this causal relationship;

[0260] The conflict score decreased significantly and fell back below the threshold.

[0261] Supports the score and coverage to reach or exceed preset conditions again.

[0262] By introducing a thawing mechanism, causal edges can be prevented from being permanently rejected due to insufficient early evidence or short-term disputes, thereby supporting the dynamic correction and long-term evolution of causal relationships.

[0263] More preferably, during the process of updating and managing causal edge confidence, the system records the triggering cause of each state change and the corresponding key evidence citation information, including the source, time, and summary location of supporting or conflicting evidence.

[0264] By adopting the above-mentioned evidence status-driven multi-round evidence collection scheduling mechanism, the present invention can automatically expand and deepen evidence collection to improve coverage when evidence is insufficient, automatically trigger counter-evidence collection to locate disputes and suppress misjudgments when conflicts are too high, automatically switch sources to obtain independent corroboration when the source is biased, and achieve early cessation when the budget is limited, thereby improving the accuracy of causal reasoning, the stability of conclusions and the verifiability of the results under limited resources.

[0265] To ensure the executability and stability of the scheduling strategy, the system can pre-set an action template library (TemplateSet) and decouple "action selection" from "action execution": after the strategy network outputs the action type and parameters, the template executor generates query / prompt words and calls the coarse search and fine search modules to obtain a new set of evidence, and then enters the consistency verification and confidence governance steps to form a closed loop of traces.

[0266] In one embodiment, it further includes:

[0267] For each causal assertion that does not meet the consistency judgment, multiple rounds of evidence collection scheduling operations are performed until the termination condition is met and the traceable causal evidence chain is output. Based on the traceable causal evidence chain, event causal reasoning analysis is performed. In each round of evidence collection scheduling operations, the following operations are performed:

[0268] Perform the selected action;

[0269] The evidence status is determined after the selected action is performed, and the evidence status is as follows: ,in, To support the score, To divide into conflict, For coverage, For the diversity index of sources, To accumulate the cost of evidence collection;

[0270] Based on the evidence status, select the action for the next round of evidence collection scheduling operation from the action set;

[0271] Determine whether the termination condition is met. If the termination condition is met, output the traceable causal evidence chain and perform event causal reasoning analysis based on the traceable causal evidence chain. Otherwise, execute the next round of evidence collection scheduling operation.

[0272] Specifically, when the consistency verification result determines that the evidence is insufficient, the conflict is too high, or the coverage is incomplete, the system selects the next action from the action set {deepening search, rebuttal search, source change search, stop output} according to the scheduling strategy, and drives the execution of the corresponding evidence collection process; when selecting deepening search, rebuttal search, or source change search, the system returns to the evidence collection-verification-update closed loop of steps S204 to S206 to obtain new multi-source evidence and incrementally update and manage the causal edge confidence until the termination condition is met.

[0273] When the credible early stop condition is met or the stop output action is triggered, the system generates and outputs the event causal conclusion and its traceable explanation information. The output includes at least: causal relationship conclusion, supporting evidence chain and counter-evidence / conflict location information, supporting score / conflict score / coverage index, causal edge confidence and its update basis; and can write back the causal edge, event node and evidence reference index determined by the threshold to the graph base, and record the scheduling action, triggering reason and key evidence reference to support review and audit.

[0274] The scheduling strategy is preferably implemented by a reinforcement learning policy network. The policy network takes support score, conflict score, coverage rate, multi-source coverage and cost as state inputs, and outputs action selection results and parameter configurations under budget constraints to optimize the efficiency of multi-round evidence collection and the reliability of conclusions.

[0275] Specifically, to overcome the problems of "insufficient evidence collection / uncontrolled costs" caused by fixed rounds or static rule-triggered retrieval, this invention constructs the evidence collection process into a closed loop of "evidence collection - verification - governance - re-evidence collection": In each round, the next evidence collection action is selected based on the current evidence status. After the evidence collection is performed, the consistency verification and confidence governance steps are returned until the termination condition is met and a traceable causal evidence chain is output. The large language model is called, and the updated traceable causal evidence chain is combined into prompt words and input into the large language model. The large language model is instructed to perform event causal reasoning analysis based on the traceable causal evidence chain, and finally, a traceable conclusion is output to answer the user's question.

[0276] (1) Definition of Evidence Status

[0277] The scheduling module constructs an evidence state in each round to characterize the sufficiency, conflict, and cost constraints of candidate causal assertions within the current evidence set. The state preferably includes, but is not limited to:

[0278]

[0279] in, To support the score, For conflict points, Cov represents coverage; The Source Diversity Index measures the number of independent sources and the balance of their distribution. When obtaining evidence, there are website sources, and source diversity is the ratio of the number of evidence sources to the total number of evidence sources. This refers to the cumulative cost of evidence collection (number of searches, etc.) across multiple rounds of evidence collection.

[0280] Based on the aforementioned evidence status, the scheduling module can explicitly perceive whether the evidence is sufficient, whether the conflict is serious, whether the source is biased, and whether the budget is close to the limit, thereby making a more reasonable decision for the next round of evidence collection under limited cost.

[0281] Action sets and parameterized execution:

[0282] In each round, the next action is selected from the action set {A1 Deepen Evidence Collection, A2 Reverse Evidence Collection, A3 Source Switching Evidence Collection, A4 Stop Output}. To support fine-grained scheduling, each action can be implemented parametrically, but the specific implementation form is not limited.

[0283] A1 (Deepening Evidence Collection): Used to improve coverage and recall, preferably achieved by increasing the depth of the graph expansion, expanding the number of candidate entity truncations, increasing the number of fine searches, and widening the time window;

[0284] A2 (Evidence of Counter-evidence): Used to locate points of contention and refutation evidence. It is best achieved by selecting a counter-evidence / debunking template and switching the source set to an authoritative or fact-checking set.

[0285] A3 (Source Change Evidence): Used to reduce single-source bias and improve independent corroboration, preferably achieved by adjusting the source combination and setting constraints such as the minimum proportion of authoritative sources;

[0286] A4 (Stop Output): Triggered when there is sufficient evidence, insufficient marginal benefits, or the budget is close to the limit, it is used to terminate evidence collection and enter traceable output.

[0287] Actions A1, A2, and A3 modify specific parameters and then re-execute steps S204 to S206, while action A4 stops directly and outputs traceable results. Actions A2 and A3 will update the evidence, thereby recalculating the support score and conflict score in step S206 and updating the confidence of candidate causal edges.

[0288] (3) Triggering and Termination Conditions (Driven by Evidence Status)

[0289] Preferably, the multi-round evidence collection scheduling is triggered by the evidence status, rather than using a fixed number of rounds. Specifically, the next round of scheduling is triggered when any one or more of the following conditions occur:

[0290] The failure to reach the preset credibility threshold indicates insufficient strength or quantity of supporting evidence; or

[0291] If the value exceeds the threshold or there is strong counter-evidence from authoritative sources, it indicates that the assertion is significantly controversial; or

[0292] Below the threshold indicates insufficient coverage of key elements (subject, action, object, time, location, etc.); or

[0293] Insufficient or excessively high proportion of a single source indicates a risk of single-source bias; or

[0294] The budget has not been exhausted and the expected marginal benefits still exceed the costs.

[0295] When the trusted early stop condition is met, the budget is exhausted, or the scheduling module selects A4 to stop the output action, the multi-round evidence collection is terminated and the final conclusion is output.

[0296] Preferably, the scheduling strategy is implemented by a reinforcement learning policy network to maximize the overall benefit of "evidence gain versus cost" under budget constraints. The policy network takes the state as input and outputs the action type and its parameter configuration; it can be trained using training samples constructed from historical cases or simulation tasks, and infers online during runtime to guide the next round of evidence collection.

[0297] Reward function design

[0298] The goal of the policy network is to maximize cumulative rewards. The single-round reward design is as follows:

[0299]

[0300] in , , This represents the increment of the current verification result relative to the previous round. This is the cost incurred in this round of actions. It is a penalty item when there is a high degree of conflict in the conclusion. , , , , These are weighting coefficients used to balance accuracy, efficiency, and stability.

[0301] like Figure 3 The diagram shown is a multi-round retrieval schematic of the preferred embodiment of the present invention, including:

[0302] Step S301, Input status: Supported score, conflict score, coverage rate, and cost;

[0303] In step S302, the output action is determined through the reinforcement learning policy network. If the action is A1, A2, or A3, the re-authentication is driven. If the action is A4, the traceable result is output.

[0304] The purpose of this invention is to provide an "evidence state-driven" multi-round evidence collection method for intelligence analysis, preferably implemented through a large language model and reinforcement learning: Event nodes are constructed and candidate causal edges are generated based on an entity linking knowledge graph. Multi-source evidence is obtained through a combination of coarse and fine retrieval enhancement methods, and a verifiable and traceable causal evidence chain is constructed through consistency verification. Simultaneously, in scenarios where evidence continuously arrives and conflicts arise with multi-source narratives, evidence collection actions are adaptively scheduled according to the evidence state, and the confidence of causal edges is incrementally updated and frozen / unfrozen, enabling causal conclusions to be continuously verified and corrected as evidence evolves, thereby improving the accuracy, robustness, and interpretability of causal reasoning.

[0305] Given the continuous arrival of evidence and the conflict between multiple narratives, dynamic consistency checks are performed on candidate causal relationships using support / rebuttal evidence. Based on the check results, incremental updates of causal edge confidence and conflict management (including deweighting, freezing / unfreezing, etc.) are carried out to ensure that existing conclusions can be continuously revised as evidence evolves and maintain stability and verifiability.

[0306] This embodiment employs evidence citation constraints and structured assertion output. Each key assertion is bound to a locatable evidence fragment (span + position pointer), forming a traceable evidence chain that facilitates manual review and auditing. By supporting / refuting evidence in parallel organization and consistency verification, combined with incremental updates of causal edge confidence and freeze / unfreeze governance, it can promptly suppress erroneous links when strong refuting evidence appears and restore evaluation when new evidence supports them, reducing the risk of misjudgment. In multi-source scenarios where facts evolve over time, the system can continuously verify and correct existing conclusions as new evidence arrives, avoiding the rigidity of one-time conclusions and improving long-term stability. Through evidence state-driven multi-round evidence collection scheduling, it adaptively decides to deepen, refute, change sources, or stop early under budget constraints, avoiding cost overruns caused by insufficient retrieval or blind multi-round retrieval.

[0307] This embodiment uses the causal reasoning task of analyzing whether and how "a central bank interest rate hike in a certain country" led to "a stock market crash in a certain emerging market" as an example to illustrate the implementation process of the event causal reasoning method for intelligence analysis based on a large model of retrieval enhancement and evidence governance proposed in this invention.

[0308] Example Implementation Background and Data Preparation

[0309] The data sources used include, but are not limited to: policy announcements published on the official website of a central bank; authoritative financial news and market analysis; and market discussions on financial social media platforms (such as Twitter). To ensure complete coverage of the causes and consequences of the event, the data time window is set from 7 days before the central bank's interest rate hike to 14 days after the stock market crash.

[0310] like Figure 4 The diagram shown is a flowchart of a preferred embodiment of the present invention, a method for causal reasoning of events for intelligence analysis, comprising:

[0311] S1: Multi-source data acquisition and preprocessing.

[0312] Access text data related to "central bank interest rate hikes" and "stock market volatility" from the aforementioned data sources. Perform the following operations using the preprocessing pipeline:

[0313] Data cleaning: Remove HTML tags, special characters, and advertising content, while retaining Chinese and English characters, numbers, and punctuation marks.

[0314] Time normalization: Regular expressions are used to match various time formats (such as "2023-10-27 14:00", "10 / 27 / 2023 2:00 PM"), and the dateparser library is used to uniformly convert them into UTC timestamps.

[0315] Deduplication: Similarity deduplication can be performed using methods such as text fingerprint / hash fingerprint, SimHash, and MinHash. The threshold is configurable (preferably 0.80 to 0.95), and the first or earliest released version can be retained according to business rules to reduce the interference of the reprint chain on evidence collection and verification.

[0316] Text segmentation: Segment by paragraph or sentence boundaries, with each segment not exceeding 512 characters.

[0317] Technical effect description: The above preprocessing ensures the text quality of subsequent processing, avoids interference from duplicate data, and provides standardized timestamps for time series analysis.

[0318] S2: Entity alignment and entity linking knowledge graph update.

[0319] A large language model (such as a general-purpose model or a vertical domain model) is invoked to extract and standardize entities from the preprocessed evidence fragments, and entity linking is completed by combining entity disambiguation rules. The large language model can be any model with long text understanding and information extraction capabilities.

[0320] Entity linking algorithm: Employs a context-similarity-based disambiguation method. Specifically, it includes:

[0321] Candidate entity generation: Use ElasticSearch full-text indexing for fuzzy matching (fuzziness=2) to return Top-5 candidates;

[0322] Disambiguation decision: Context similarity is calculated based on BERT-embedding, and a weighted score is given by combining the number of times the entity is mentioned (popularity) in the knowledge graph;

[0323] Link threshold: A link is successful when the similarity is >0.7 and the overall score is >0.6; otherwise, a new entity node is created.

[0324] Knowledge graph update: Use graph databases (Neo4j) or other structured storage to implement the associated storage of entities and evidence, establish the mention relationship of "entity-evidence fragment", and record attributes such as mention_text, location pointer, confidence level, etc. to support backtracking and auditing.

[0325] Technical effect description: By entity alignment and linking, a unified representation of entities in multi-source evidence is achieved, providing a structured knowledge foundation for subsequent event layer construction and causal reasoning.

[0326] S3: Event layer construction and candidate causal edge generation.

[0327] Event Extraction: This involves using a large language model to extract events from evidence fragments. An event extraction prompt template is used.

[0328] Event merging algorithm: It can calculate semantic similarity based on the text vector representation of event summaries (such as Sentence-BERT-like embedding model or other text embedding models), and merge similar events by combining time difference threshold; the similarity threshold and time window can be configured (e.g., similarity 0.80~0.95, time difference 12~48 hours) to adapt to the granularity requirements of different businesses.

[0329] Candidate causal edge generation:

[0330] The generation of candidate edges must satisfy the following constraints:

[0331] Timing constraint: The causal event must occur earlier than the result event, and the allowed time window relaxation ΔT_relax is set to 3 days;

[0332] Entity association constraint: Two events must share at least one core entity;

[0333] The initial confidence level is calculated using the following formula:

[0334]

[0335] in, The weighting coefficients are obtained through training with historical data.

[0336] Then, a reasonable set of candidate causal edges is obtained through a time window (e.g., the initial causal edge is determined based on the occurrence or release time of the central bank's interest rate hike and the stock price increase, with the one that occurs first being initially identified as the cause), which serves as the input for subsequent confidence calculation and evidence fusion.

[0337] Technical effect description: By using multi-constraint screening and weighted confidence calculation, the computational redundancy caused by full connection is avoided, while ensuring the rationality of candidate relationships and providing a high-quality candidate set for subsequent multiple rounds of evidence collection.

[0338] S4: Enhanced forensic retrieval (coarse search + fine search).

[0339] Coarse search strategy:

[0340] Core entity set construction: given candidate causal edges Core Entity Set Let u be the union of all participating entities in events u and v.

[0341] Knowledge graph expansion: [On the knowledge graph] Each entity in the model undergoes a 1-hop neighborhood expansion, with expansion relationship types including synonym relationships, membership relationships, geographic location relationships, etc.

[0342] Query construction: The expanded entity set (Top-50 truncated) is connected with the event keyword using OR logic to form a Boolean query, and a time window filter is applied to the timestamp index.

[0343] Detailed search strategy:

[0344] Vectorization: The text-embedding-ada-002 model is used to map the query text and coarse search results into high-dimensional vectors;

[0345] Similarity calculation: Calculate cosine similarity and select the top-10 most relevant evidence segments;

[0346] Multi-factor re-ranking: The final score comprehensively considers similarity, quality of evidence (authority of sources), and diversity.

[0347] Technical effect description: The two-stage retrieval strategy ensures both the breadth of evidence recall (coarse retrieval) and the quality of evidence (fine retrieval), providing a high-quality evidence set for subsequent consistency verification.

[0348] For the causal reasoning task of "a central bank interest rate hike in a certain country" and "a stock market crash in a certain emerging market," the first step is to extract core keywords for the source and result events. The source event (central bank interest rate hike) can use keywords such as "central bank," "interest rate hike," "interest rate increase," "monetary policy tightening," and "benchmark interest rate adjustment," while the result event (emerging market stock market crash) includes keywords such as "stock market crash," "stock index decline," "emerging market stock market," "sharp drop in stock prices," and "market sell-off." Next, entities related to these events are extracted from relevant documents or knowledge bases and expanded, such as the countries or regions involved, financial institutions, and stock index names, then truncated to the Top-50. Then, these expanded entities are logically connected with the event keywords using OR to form a complete entity sequence. To ensure the temporal rationality of the causal analysis, a time window filter needs to be applied to the timestamp index, for example, obtaining relevant evidence sets around a certain number of days before and after the central bank interest rate hike or the stock market crash.

[0349] S5: Evidence constraint inference and consistency verification based on large models.

[0350] Implementation of evidence constraint mechanism:

[0351] The following methods are used to force large language models to bind their inference conclusions to specific evidence:

[0352] (1) Structured output constraints: The large language model is required to output a structured causal claim, which includes the fields of cause_event, effect_event, confidence and evi_list. The large model is used to determine whether the evidence is supporting evidence or disproving evidence by utilizing its text analysis capabilities.

[0353] (2) Evidence citation constraints: Each citation in evi_list must provide the evidence text span and position pointer (paragraph / period / character offset, etc.) to support the review of the original text.

[0354] (3) Refusal to answer and re-evidence constraints: When the output is missing a reference, the reference does not match the assertion, or the conflict is too high, the system prefers to trigger re-evidence or regeneration to avoid the solidification of "no evidence conclusion".

[0355] Large model call parameters: The generation temperature can be set to a low temperature range (0~0.3) to improve output consistency, and the upper limit of output length and call budget can be set to stably produce structured assertions and evidence citations at a controllable cost.

[0356] Consistency verification index calculation:

[0357] Support score: Calculated by weighting the amount of evidence, the strength of support, and the reliability of the source, and used to characterize the degree to which an assertion is supported.

[0358] Conflict score: Based on the quantity of evidence, strength of counter-evidence, and authority, with particular attention to evidence that directly refutes the assertion or provides a strong alternative explanation.

[0359] Coverage: The proportion of key elements (subject, action, object, time, place) in an assertion that are covered by evidence, used to measure the completeness of information.

[0360] S6: Causal edge confidence management and incremental update.

[0361] The confidence level of the causal edge is dynamically adjusted using the following incremental update formula:

[0362] Preferably, the confidence update can be expressed as

[0363] in, For the updated confidence level, The confidence level before the update. It also introduces a state governance mechanism of freezing, unfreezing, and abandonment to achieve long-term evolution management of causal relationships.

[0364] Where η is the learning rate, preferably ranging from 0.05 to 0.20; and These are the support score and conflict score obtained from the consistency check, respectively; the update magnitude can also introduce a smoothing coefficient or upper / lower confidence limits to suppress drastic fluctuations caused by short-term noise.

[0365] Governance operation state machine:

[0366] Freeze trigger condition: When a conflict occurs Two consecutive rounds above the threshold (For example, 0.6, configurable) triggers a freeze.

[0367] Unfreeze trigger condition: After freezing, when a conflict occurs... Unfreezing is triggered when the value falls below the threshold τ (e.g., 0.3) for two consecutive rounds and new strong supporting evidence emerges.

[0368] Deprecation trigger condition: When the confidence level is... Below the threshold (For example, 0.2) triggers discard, that is, if the confidence of this causal edge is lower than the threshold, then this causal edge is removed and the remaining evidence is input into the large model to answer.

[0369] Technical effect description: Through the state governance of "incremental confidence update + freeze / unfreeze / discard", the system can manage evidence conflicts and fact evolution in the long term, avoiding the long-term solidification of erroneous causal edges caused by early one-sided evidence; when strong counter-evidence appears, the relevant causal edges can be reduced in weight or frozen in time; when new independent supporting evidence appears later, they can be unfrozen and re-evaluated, thereby improving the stability and verifiability of the conclusions.

[0370] S7: Scheduling strategy decision (preferably implemented by a reinforcement learning policy network).

[0371] When the consistency verification result determines that the evidence is insufficient, the conflict is too high, or the coverage is incomplete, the system selects the next action from the action set {deepen search, rebuttal search, source change search, stop output} according to the scheduling strategy, and drives the execution of the corresponding evidence collection process. When selecting the deeper search, rebuttal search, or source change search, the system returns to execute the evidence collection-verification-update closed loop of S4–S6 to obtain new multi-source evidence and incrementally update and manage the causal edge confidence until the termination condition is met and the traceable result is output.

[0372] Execution results of the example

[0373] After multiple iterations (usually 2-4 rounds), the system outputs a final conclusion. Taking this embodiment as an example, the possible outputs are:

[0374] "There is a moderate confidence level (0.55) causal relationship between the central bank's interest rate hikes and the stock market crash, but there is significant counter-evidence pointing to international circumstances as the main alternative cause."

[0375] It also outputs a complete and traceable chain of evidence, including all supporting and disproving evidence and their sources. This causal edge, its confidence level, and the chain of evidence can be written back into the knowledge graph for subsequent analysis.

[0376] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0377] like Figure 5 The diagram shown is a hardware structure schematic of an electronic device according to the present invention, comprising:

[0378] At least one processor 501; and,

[0379] A memory 502 is communicatively connected to at least one of the processors 501; wherein,

[0380] The memory 502 stores instructions that are executed by at least one of the processors to enable the at least one of the processors to perform the event causal reasoning method for intelligence analysis as described above.

[0381] Figure 5 Take a processor 501 as an example.

[0382] The electronic device may also include an input device 503 and a display device 504.

[0383] The processor 501, memory 502, input device 503 and display device 504 can be connected by a bus or other means. The figure shows an example of connection by bus.

[0384] Memory 502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the event causal reasoning method for intelligence analysis in the embodiments of this application, for example, Figure 1 , Figure 2 The method flow is shown. The processor 501 executes various functional applications and data processing by running non-volatile software programs, instructions, and modules stored in the memory 502, thereby implementing the event causal reasoning method for intelligence analysis in the above embodiments.

[0385] Memory 502 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of event-causal reasoning methods for intelligence analysis, etc. Furthermore, memory 502 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 502 may optionally include memory remotely located relative to processor 501, and these remote memories may be connected via a network to means of performing event-causal reasoning methods for intelligence analysis. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0386] Input device 503 can receive user clicks and generate signal inputs related to user settings and function controls for event causal reasoning methods used for intelligence analysis. Display device 504 may include display devices such as a display screen.

[0387] When one or more modules are stored in the memory 502, and are run by one or more processors 501, they execute the event causal reasoning method for intelligence analysis in any of the above method embodiments.

[0388] This invention, based on a knowledge graph, retrieves one or more binding pieces of evidence for each candidate causal edge from a set of evidence fragments. For each candidate causal edge, a large language model is invoked to generate a causal assertion. Consistency checks are performed on the causal assertions and the binding evidence for the same candidate causal edge. Causal assertions that satisfy the consistency check are combined with their corresponding binding evidence to form a traceable causal evidence chain for event causal reasoning analysis. Therefore, based on entity alignment of multi-source texts, this invention constructs an event-level representation and generates "event-event" candidate causal relationships, elevating the reasoning object from "entity co-occurrence" to a reasonable and maintainable event causal structure. Simultaneously, evidence citation constraints and structured assertion outputs form a traceable evidence chain, facilitating manual review and auditing.

[0389] One embodiment of the present invention provides a storage medium that stores computer instructions, which, when executed by a computer, are used to perform all the steps of the event causal reasoning method for intelligence analysis as described above.

[0390] In the context of this disclosure, a storage medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The storage medium can be a machine-readable signal medium or a machine-readable storage medium. Optionally, the storage medium can be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), compact disc ROM (CD-ROM), magnetic tape, floppy disk, and optical data storage device.

[0391] One embodiment of the present invention provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the event causal reasoning method for intelligence analysis as described above.

[0392] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for causal reasoning of events for intelligence analysis, characterized in that, include: Multiple evidence fragments about the query text are obtained from multiple sources, and these multiple evidence fragments are used as a set of evidence fragments. Entities are extracted from the evidence fragments, and the entities are linked into a knowledge graph. Each evidence fragment is associated with and stored as an extracted entity. Extract event elements from the evidence fragments, and generate event nodes and causal edges based on the event elements. Take the two event nodes connected by the causal edge as event pairs, and generate one or more candidate causal edges based on the event pairs. Based on the knowledge graph, one or more binding evidences for each candidate causal edge are retrieved from the evidence fragment set; For each candidate causal edge, a large language model is invoked to generate a causal assertion, and the consistency of the causal assertion and the binding evidence of the same candidate causal edge is verified. The causal assertion that satisfies the consistency verification and the corresponding binding evidence are combined into a traceable causal evidence chain, and event causal reasoning analysis is performed based on the traceable causal evidence chain. The process of generating causal assertions by calling a large language model for each candidate causal edge, and performing consistency checks on the binding evidence between the causal assertions and the same candidate causal edge, and combining the causal assertions that satisfy the consistency check with the corresponding binding evidence into a traceable causal evidence chain, includes: For each of the candidate causal edges, perform the following operation: The large language model is invoked to generate causal assertions about the candidate causal edges, and supporting and contradictory evidence about the causal assertions is determined in the binding evidence of the same candidate causal edge. Based on supporting and rebuttal evidence, calculate the support score, conflict score, and coverage of the causal assertion; The causal assertion that the support score, conflict score, and / or coverage rate satisfy the consistency judgment is combined with the corresponding binding evidence, support score, conflict score, and coverage rate to form a traceable causal evidence chain.

2. The event causal reasoning method for intelligence analysis according to claim 1, characterized in that, The generation of one or more candidate causal edges based on the event pairs includes: For any pair of events, perform a temporal feasibility assessment on all causal edges between the two event nodes in the pair, and remove causal edges that do not satisfy the temporal logic. The temporal logic is as follows: ,in, To relax the time window, This refers to the occurrence time of the causal event node in the event pair. This refers to the occurrence time of the result event node in the event pair; The initial confidence level of the causal edges after elimination is calculated based on the reliability of the source, the strength of the causal trigger, the tightness of the entity association, and the time consistency. The top M causal edges with the highest initial confidence level in the result event node of each event pair are retained as candidate causal edges, where M is an integer greater than 1.

3. The event causal reasoning method for intelligence analysis according to claim 1, characterized in that, Based on the knowledge graph, one or more binding evidences for each candidate causal edge are retrieved from the evidence fragment set, including: For each of the candidate causal edges, perform the following retrieval operation: From the knowledge graph, the entity nodes of all participants involved in the cause events and result events connected by the candidate causal edges are taken as the core entity set of the candidate causal edges; Obtain the entity nodes of each entity node in the core entity set at a preset level in the knowledge graph, as an extended entity set; Construct a query expression to retrieve one or more first candidate evidences from the evidence fragment set. The query expression includes event keywords and the names of entity nodes in the core entity set and the extended entity set. The event keywords are the keywords of the target event included in the query text. The first candidate evidence that meets the time requirement is filtered out as the coarse search candidate evidence; The coarse search candidate evidence is converted into evidence vectors, the query text is converted into query vectors, and the vector similarity between each evidence vector and the query vector is calculated. For each coarse search candidate evidence, a fine search ranking score is calculated based on vector similarity, evidence quality score, and diversity factor. The top K coarse search candidate evidences with the highest scores are selected as the binding evidence for the candidate causal edge.

4. The event causal reasoning method for intelligence analysis according to claim 1, characterized in that, The calculation of the support score, conflict score, and coverage of the causal assertion includes: Computational support is divided into: in, Indicates support for the division. This represents the weight of the i-th binding evidence. This represents the strength of the large language model's judgment on the ith binding evidence supporting the causal assertion, wherein the binding evidence includes the supporting evidence and the evidence against the contrary; Calculation conflicts are categorized as follows: in, Indicates conflict points, This represents the strength of the large language model's judgment on the i-th binding evidence against the causal assertion; The calculated coverage is: in, Indicates coverage rate. This indicates the number of key elements of the causal assertion covered by all binding evidence. This represents all the key elements of the stated causal assertion.

5. The event causal reasoning method for intelligence analysis according to claim 1, characterized in that, Also includes: In response to the updated binding evidence of the candidate causal edge, the support score and conflict score are recalculated. As the support score increases, the confidence level of the candidate causal edge is increased; When the conflict score increases or the support score decreases, the confidence level of the candidate causal edge is reduced.

6. The event causal reasoning method for intelligence analysis according to claim 1, characterized in that, Also includes: For each causal assertion that does not meet the consistency judgment, multiple rounds of evidence collection scheduling operations are performed until the termination condition is met and the traceable causal evidence chain is output. Based on the traceable causal evidence chain, event causal reasoning analysis is performed. In each round of evidence collection scheduling operations, the following operations are performed: Perform the selected action; The evidence status is determined after the selected action is performed, and the evidence status is as follows: ,in, To support the score, To divide into conflict, For coverage, For the diversity index of sources, To accumulate evidence collection costs, based on the evidence status, select the action for the next round of evidence collection scheduling operation from the action set; Determine whether the termination condition is met. If the termination condition is met, output the traceable causal evidence chain and perform event causal reasoning analysis based on the traceable causal evidence chain. Otherwise, execute the next round of evidence collection scheduling operation.

7. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that are executed by at least one of the processors to enable the at least one of the processors to perform the event causal reasoning method for intelligence analysis as described in any one of claims 1 to 6.

8. A storage medium, characterized in that, The storage medium stores computer instructions that, when executed by a computer, are used to perform all steps of the event causal reasoning method for intelligence analysis as described in any one of claims 1 to 6.

9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the event causal reasoning method for intelligence analysis as described in any one of claims 1 to 6.