Potential hypothesis relation prediction method based on causal symbol network

Through the causal symbolic network method, causal symbolic networks are automatically constructed from scientific literature and multi-dimensional discrimination is performed, which solves the bias and black box problems in the hypothesis generation of large language models and achieves potential hypothesis generation with high reliability and explainability.

CN120745852AActive Publication Date: 2025-10-03NANJING UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511272308.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-10-03
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing large language models have training data bias and black-box reasoning mechanisms when generating scientific hypotheses, which may cause the generated hypotheses to tend towards mainstream views, ignore disruptive innovation paths, and make causal logical consistency difficult to ensure.

Method used

The causal symbolic network method is adopted to automatically identify and construct directed causal symbolic networks from unstructured scientific literature, and use multidimensional discrimination criteria and heuristic random walk strategies to generate potential hypothesis relationships, with clear evidence paths and credibility rankings.

Benefits of technology

The reliability and causal logic consistency of hypothesis predictions are improved. The generated hypotheses have a clear causal logic chain and interpretability, can effectively capture complex mediating and regulatory effects, and provide diversified credibility assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745852A_ABST
    Figure CN120745852A_ABST
Patent Text Reader

Abstract

The invention discloses a potential hypothesis relationship prediction method based on a causal symbol network. The method comprises the following steps: scientific hypothesis variable identification and multi-dimensional causal relationship discrimination; constructing a causal symbol network; hypothesis inspiration heuristic search is carried out; predicting a potential hypothesis relationship; ranking the credibility of the potential hypothesis relationship; according to the method, the mediation and regulation effects can be effectively captured and utilized, so that new hypotheses including the complex mechanisms can be generated, and the method is closer to the scientific problem of the real world; scientific researchers can preferentially pay attention to the assumption which is most likely to be established, and screening and verification time is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence, knowledge discovery and scientific literature mining, and in particular to a method for predicting potential hypothesis relationships based on a causal symbolic network. Background Art

[0002] Using artificial intelligence to assist scientific discovery, or AI for Science, has become a cutting-edge trend in scientific research. Existing computational methods attempt to leverage large language model technology to discover potential scientific hypotheses. However, existing approaches still suffer from numerous limitations, including: 1. Cognitive biases in training data. The power of large language models stems from their massive amounts of training data. Because the data itself may contain outdated, erroneous, or even biased knowledge, the model inevitably reproduces or even amplifies these biases during learning. This can cause the hypotheses it generates to favor mainstream or popular viewpoints while ignoring more disruptive and counterintuitive innovations. 2. Black-box nature of the reasoning mechanism. The reasoning process of large language models, based on neural networks, lacks clear logic and causal pathways. When the model generates a new hypothesis, it is difficult to determine whether it is based on rigorous logical deduction or merely superficial correlations in the data. This makes it difficult to ensure the causal consistency of the generated scientific hypotheses and may even lead to the generation of seriously misleading pseudo-hypotheses, disrupting scientific research activities. Summary of the Invention

[0003] Purpose of the invention: The purpose of the present invention is to provide a method for predicting potential hypothesis relationships based on causal symbolic networks, to achieve automated and structured extraction of scientific hypothesis information containing causal directions, effect properties and complex mediation / regulatory relationships from unstructured scientific literature, to construct a network model that can formally express the causal logic in scientific knowledge, to generate new, structurally complete hypothesis relationships, and to solve problems existing in the background technology.

[0004] Technical solution: The method for predicting potential hypothesis relationships based on a causal symbolic network described in the present invention includes the following steps: (1) Scientific hypothesis variable identification and multi-dimensional causal relationship discrimination: Automatically identify and extract structured scientific hypothesis variables from unstructured scientific literature data sources, and perform multi-dimensional discrimination on the causal relationships between variables; (2) Constructing a causal symbolic network: Based on the causal relationships identified as valid in step (1), a directed causal symbolic network with signed weights is constructed, where the nodes represent scientific variables, the directed edges represent the direction of the causal relationship, and the signs of the edges represent the nature of the causal effect; (3) Hypothesis inspiration heuristic search: For one or more target variables in the network, a heuristic random walk strategy is used to conduct exploratory search in the causal symbolic network to generate a local subgraph containing potential hypothesis inspirations; (4) Prediction of potential hypothesis relationships: Within the generated local subgraph, a set of new candidate hypothesis relationships is generated by reorganizing the knowledge of existing variables and their causal relationship paths; (5) Credibility ranking of potential hypothetical relationships: Design and calculate the comprehensive credibility score of each candidate hypothetical relationship, sort the candidate hypothetical relationships in descending order according to the score, and finally output a list of potential hypothetical relationships.

[0005] Furthermore, step (1) includes the following steps: (11) Using a large language model fine-tuned by scientific literature in a specific field, the joint task of named entity recognition and relation extraction is performed on the input full text or abstract of scientific papers; (12) The large language model automatically extracts and formalizes the structured five-tuple hypothesis relationship from the text. The structure of the five-tuple is: <independent variable, dependent variable, effect, mediating variable, moderating variable>, where the effect refers to the direction of the influence of the independent variable on the dependent variable, and its value is positive or negative; (13) Evaluate the variable relationships represented by the five-tuples one by one according to a set of pre-defined causal discriminant dimensions; the pre-defined causal discriminant dimensions include the following: domain consistency, direction rationality, mechanism rationality, causal temporality, mediation rationality, and regulatory appropriateness; (14) The evaluation is performed through expert consultation or automated rule logic. When a variable relationship meets at least two of the dimensional criteria, the variable relationship is confirmed as a valid causal relationship and used for subsequent network construction.

[0006] Furthermore, step (2) includes the following steps: (21) All identified independent and dependent variables are instantiated as nodes in the network; (22) If there is a valid causal relationship between variables A and B, where A is the independent variable and B is the dependent variable, then a directed edge is established between the node representing A and the node representing B, with the direction of the edge pointing from A to B; (23) According to the effect information in the quintuple, a sign weight is assigned to the directed edge: if the effect is positive, the sign of the edge is defined as +; if the effect is negative, the sign of the edge is defined as -; (24) The mediating variable and moderator variable information contained in the quintuple are attached as attributes and stored on the directed edge.

[0007] Furthermore, step (3) includes the following steps: (31) By calculating and comparing at least one network centrality index of each node in the network, one or more target variables with high exploration value are automatically identified and selected, denoted as v ego ; (32) From v ego Start and execute the heuristic random walk algorithm on the network, where the next node to be visited during the walk is v next From the current node v current Randomly select from the two-way neighborhood of (33) The bidirectional neighborhood of the current node N bidirectional ( v current ) is defined as the union of the node's predecessor node set and successor node set, and the mathematical expression is: ; Among them, Pre( v ) represents a node v The neighborhood set of Suc( v ) represents a node v The out-neighborhood set.

[0008] Furthermore, the termination condition of the heuristic random walk is set to satisfy any of the following conditions: the depth of the walk path, i.e., the depth from the starting point v ego The number of edges passed by the starting point reaches the preset maximum depth threshold d max The total number of unique nodes visited during the walk exceeds the pre-specified node number threshold. n max When the walk ends, all visited nodes and the original edges between them together form a local subgraph that maintains the original network structure. G sub = G [ V visited ].

[0009] Furthermore, in step (4), the potential hypothesis relationship prediction is based on v ego The local subgraph centered at G sub The following steps are performed: (41) From G sub Identify all v egoCandidate variable nodes that are path-connected but not directly adjacent v i , forming a candidate set U={ v 1 , v 2 , …, v i}; (42) For each candidate variable v i U, by calculating v ego and v i The node direction measurement on all paths between them infers the most likely causal direction between the two and quantifies it into a direction score; (43) The cumulative effect of the path between the two is calculated by multiplying the sign effects of all edges on the path, thereby determining the sign of the causal relationship in the new hypothesis, that is, positive or negative; (44) Check the connection v ego and v i Whether there are edges with mediating or moderating properties on the path to identify potential mediating variables and moderating variables in the new hypothesis.

[0010] Furthermore, the direction score is calculated according to the following formula: ; in, express G sub Middle Connection v ego and v i The set of intermediate nodes on all paths; is the size of the set, d ( x ) is a predefined node direction metric function, the positive or negative value of which indicates the relative position of the intermediate node in the path. v ego forward or backward direction.

[0011] Furthermore, the calculation of the cumulative effect follows the following formula: ; in, e j It is the first j edge; effect( e j ) is the symbolic network of the edge,k is the length of the path.

[0012] Furthermore, step (5) includes the following steps: (51) The number of common neighbors is used as the main criterion for evaluating the structural proximity of two variables, and this number is quantified as the target variable v ego Neighbor set and candidate variables v i The intersection size of the neighbor sets of ,in Representation node v The set of direct neighbors of ; (52) The frequencies of potential mediating and moderating variables identified by the candidate hypotheses are counted as additional evaluation indicators; (53) The number of common neighbors and the frequency of mediating and moderating variables are integrated into a comprehensive credibility score through a weighted sum function; (54) All candidate hypotheses h i Arrange them in descending order according to their comprehensive credibility scores to generate the final ordered set of candidate hypotheses H ={ h 1 , h 1 ,…, h i}.

[0013] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: (1) Improving the reliability and causal logic consistency of hypothesis predictions: This invention is not based on simple correlation, but extracts causal relationships with direction and nature from the literature and screens them through multi-dimensional discrimination criteria to ensure the quality of knowledge from the source. The constructed causal symbol network enables all subsequent predictions to be based on the causal logic chain, significantly improving the scientificity and reliability of the generated hypotheses; (2) Enhance the interpretability of prediction process and results: Different from the black box mechanism defects of large language models driving hypothesis prediction, each potential hypothesis generated by the present invention is accompanied by a clear evidence path, which can be traced back to which local subgraph the hypothesis is in, based on which existing causal paths, and through what kind of recombination and inference rules it is generated.

[0014] (3) Modeling and prediction of complex scientific hypotheses: Through the storage of quintuple structures and edge attributes, the present invention can effectively capture and utilize mediation and regulation effects, thereby generating new hypotheses that include these complex mechanisms and are closer to real-world scientific problems.

[0015] (4) Providing a diversified and quantifiable evaluation system: The present invention not only generates hypotheses, but also ranks the hypothetical relationships by combining the comprehensive credibility scores of network topology and relationship attributes, so that researchers can give priority to the hypotheses that are most likely to be established, saving time in screening and verification. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic diagram of the present invention. DETAILED DESCRIPTION

[0017] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0018] like Figure 1 As shown, an embodiment of the present invention provides a method for predicting potential hypothesis relationships based on a causal symbolic network, comprising the following steps: S1. Collection and preprocessing of academic literature in the target field. First, the target field was selected as enterprise digital transformation. This field lies at the intersection of disciplines such as econometrics, information technology, business administration, and organizational behavior. Academic papers in this field typically contain clear hypotheses and empirical research processes, making them suitable as a data source for this method. Second, through manual searches combined with automated crawlers, relevant academic literature in this field was captured from mainstream academic databases such as Web of Science and Scopus. The papers were retrieved and their bibliographic information (such as title, author, publication year, publication name, document ID, etc.) and abstracts were collected. To obtain more complete contextual information, APIs of open academic platforms such as OpenAlex or Semantic Scholar were used to match and retrieve the full-text content of the target documents using the acquired bibliographic information. Full-text documents in PDF format were parsed using Python libraries such as PyMuPDF and converted to plain text. For full-text web pages in HTML format, the text content was directly extracted using the Elsevier API or other academic publisher APIs. Finally, all collected academic texts (abstracts and full texts) are cleaned to remove irrelevant noise data such as headers, footers, reference lists, and figure titles, and the cleaned texts are stored in a structured manner.

[0019] S2. Extraction of Hypothetical Variable Relationships Based on a Large Language Model. Extracting hypothetical relationships from academic literature begins with identifying declarative sentences that may contain hypotheses. Academic texts typically consist of an introduction, literature review (or theory & hypothesis), methods, results, and conclusions. Declarative sentences containing hypotheses typically appear in the hypothesis, results, and conclusion sections of academic papers. The hypothesis section typically provides a specific statement based on existing theory or logical reasoning about the expected relationship between variables. The results and conclusion sections are more general, elaborating on the hypothesized results and key findings of the academic text. Alibaba's Qwen-max large language model is used to identify hypothetical relationship declarative sentences in academic texts. To improve recognition performance, examples are included in the input prompts, and LangChain is used to verify the output type to ensure that the data is returned in a structured JSON format. The structured hypothetical relationship data generated in the previous step is fed into the large language model. Using semantic parsing capabilities, the textual content of each hypothetical relationship is deeply analyzed, and preliminary classification is performed based on the effect type, such as positive, negative, moderating, and mediating. The preliminary classification results are then verified and adjusted using pre-designed customized rule templates. These templates, based on domain knowledge and common patterns of hypothesized relationships, aim to identify the characteristics and logical structure of specific types of relationships. A logical reasoning module is then used to further verify the identified relationship type, ensuring the logical consistency of the hypothesized relationship within the overall context and the rationality of the reasoning chain. After semantic parsing, rule template matching, and logical reasoning verification, the system makes a final determination of the type of each hypothesized relationship and outputs it as a tuple, annotated with the specific relationship type and related variable information.

[0020] S3. Verify multidimensional causal relationships. Each extracted five-tuple relationship is evaluated for validity. For example, for the five-tuple relationship of <Technology Proximity, Digital Transformation Maturity, Positive, Knowledge Diversity, Technological Turbulence>, the verification process is as follows: ① Domain Consistency: The enterprise leverages its existing technological foundation to drive transformation, which aligns with the resource-based view theory and is considered a pass. ② Directional Rationality: Technology Proximity is the cause, and Transformation Maturity is the effect. The direction is reasonable and the judgment is considered a pass. ③ Intermediary Rationality: Technology Proximity → Knowledge Diversity → Transformation Maturity form a logically coherent causal chain (i.e., possessing relevant technologies makes it easier to absorb new knowledge, and rich knowledge promotes transformation). The judgment is considered a pass.

[0021] S4. Causal symbolic network construction. Integrate all verified causal relationships in the field of enterprise digital transformation into a global knowledge network. GAll identified variables, such as AI technology investment, market responsiveness, supply chain management, data analysis capabilities, enterprise performance, and organizational agility, are instantiated as nodes in the network. Based on the above example, a directed edge is established between the AI ​​technology investment node and the market responsiveness node, from the former to the latter. Because the effect is positive, the sign of this edge is set to +. Supply chain management (mediating variable) and data analysis capabilities (moderating variable) are attached as metadata to this directed edge with the + sign. A program is written to repeat this process, ultimately constructing a symbolic network that describes the causal relationships between various variables in the field of enterprise digital transformation.

[0022] S5. Hypothesis Inspiration Heuristic Search. In the enterprise digital transformation causal symbolic network, identify an inspiration subgraph for generating new hypotheses. Calculate the betweenness centrality of each node in the network. You may find that nodes such as data-driven decision-making or customer experience have high centrality scores because they are key bridges connecting many different research topics. Select one of these, such as data-driven decision-making, as the target variable. v ego Starting from the data-driven decision node, a bidirectional neighborhood random walk is performed. The paths that the walk may take. The walk stops when it reaches the preset depth or node number limit. All visited nodes, such as cloud computing adoption, business process automation, employee skills training, executive support, and data-driven decision-making, and the original edges between them, together constitute a local inspiration subgraph about data-driven decision-making. G sub .

[0023] S6, prediction of potential hypothetical relationships. v ego The local subgraph centered at G sub First, start from G sub Identify all v ego Candidate variable nodes that are path-connected but not directly adjacent v i , forming a candidate set U={ v 1, v 2, …, v i}; For each candidate variable v i U, by calculating v ego and v iThe node direction measurement on all paths between them is used to infer the most likely causal direction between the two and quantify it as a direction score. Furthermore, the cumulative effect of the path between the two is calculated by multiplying the sign effects of all edges on the path to determine the sign of the causal relationship in the new hypothesis (positive or negative). The calculation of the cumulative effect follows the following formula: ; in, e j It is the first j edge; effect( e j ) is the symbolic network of the edge, k is the length of the path. Finally check the connection v ego and v i Whether there is an edge with mediating or moderating properties on the path to identify the potential mediating variables and moderating variables in the new hypothesis, the direction score is calculated according to the following formula: ; in, express G sub Middle Connection v ego and v i The set of intermediate nodes on all paths; is the size of the set, d ( x ) is a predefined node direction measurement function, and its positive or negative value indicates the relative position of the intermediate node in the path. v ego forward or backward direction.

[0024] S7. Ranking the credibility of potential hypotheses. Rank all generated new hypotheses. For example, if you calculate the intersection of the neighbor sets for data-driven decision-making and organizational learning capability, you might find that they are both directly connected to the nodes of senior management support and IT infrastructure, with a total of two common neighbors. If decision quality appears repeatedly as a mediator in multiple inference paths, its frequency score as a mediating variable will be higher. Combine the number of common neighbors (e.g., 2) and the frequency score of the mediating / moderating variable (e.g., 1.5) using a weighted formula (e.g., Score = 0.6 * CommonNeighbors + 0.4 * MetaScore) to create a single overall score. Compare this hypothesis's score with the other generated new hypotheses and sort them from high to low by score.

Claims

1. A method for predicting potential hypothesis relationships based on causal symbolic networks, characterized in that: The following steps are involved: (1) Scientific hypothesis variable identification and multi-dimensional causal relationship discrimination: Automatically identify and extract structured scientific hypothesis variables from unstructured scientific literature data sources, and perform multi-dimensional discrimination on the causal relationships between variables; (2) Constructing a causal symbolic network: Based on the causal relationships identified as valid in step (1), a directed causal symbolic network with signed weights is constructed, where the nodes represent scientific variables, the directed edges represent the direction of the causal relationship, and the signs of the edges represent the nature of the causal effect; (3) Hypothesis inspiration heuristic search: For one or more target variables in the network, a heuristic random walk strategy is used to conduct exploratory search in the causal symbolic network to generate a local subgraph containing potential hypothesis inspirations; (4) Prediction of potential hypothesis relationships: Within the generated local subgraph, a set of new candidate hypothesis relationships is generated by reorganizing the knowledge of existing variables and their causal relationship paths; (5) Credibility ranking of potential hypothetical relationships: Design and calculate the comprehensive credibility score of each candidate hypothetical relationship, sort the candidate hypothetical relationships in descending order according to the score, and finally output a list of potential hypothetical relationships.

2. The method for predicting potential hypothesis relationships based on a causal symbolic network according to claim 1, characterized in that: Step (1) includes the following steps: (11) Using a large language model fine-tuned by scientific literature in a specific field, the joint task of named entity recognition and relation extraction is performed on the input full text or abstract of scientific papers; (12) The large language model automatically extracts and formalizes the structured five-tuple hypothesis relationship from the text. The structure of the five-tuple is: <independent variable, dependent variable, effect, mediating variable, moderating variable>, where the effect refers to the direction of the influence of the independent variable on the dependent variable, and its value is positive or negative; (13) Evaluate the variable relationships represented by the five-tuples one by one according to a set of pre-defined causal discriminant dimensions; the pre-defined causal discriminant dimensions include the following: domain consistency, direction rationality, mechanism rationality, causal temporality, mediation rationality, and regulatory appropriateness; (14) The evaluation is performed through expert consultation or automated rule logic. When a variable relationship meets at least two of the dimensional criteria, the variable relationship is confirmed as a valid causal relationship and used for subsequent network construction.

3. The method for predicting potential hypothesis relationships based on a causal symbolic network according to claim 1, characterized in that: Step (2) includes the following steps: (21) All identified independent and dependent variables are instantiated as nodes in the network; (22) If there is a valid causal relationship between variables A and B, where A is the independent variable and B is the dependent variable, then a directed edge is established between the node representing A and the node representing B, with the direction of the edge pointing from A to B; (23) According to the effect information in the quintuple, a sign weight is assigned to the directed edge: if the effect is positive, the sign of the edge is defined as +; if the effect is negative, the sign of the edge is defined as -; (24) The mediating variable and moderator variable information contained in the quintuple are attached as attributes and stored on the directed edge.

4. The method for predicting potential hypothesis relationships based on a causal symbolic network according to claim 1, characterized in that: Step (3) includes the following steps: (31) By calculating and comparing at least one network centrality index of each node in the network, one or more target variables with high exploration value are automatically identified and selected, denoted as v ego ; (32) From v ego Start and execute the heuristic random walk algorithm on the network, where the next node to be visited during the walk is v next From the current node v current Randomly select from the two-way neighborhood of (33) The bidirectional neighborhood of the current node N bidirectional ( v current ) is defined as the union of the node's predecessor node set and successor node set, and the mathematical expression is: ; Among them, Pre( v ) represents a node v The neighborhood set of Suc( v ) represents a node v The out-neighborhood set.

5. The method for predicting potential hypothesis relationships based on a causal symbolic network according to claim 4, characterized in that: The termination condition of the heuristic random walk is set to satisfy any of the following conditions: the depth of the walk path, that is, the depth of the walk path from the starting point v ego The number of edges passed by the starting point reaches the preset maximum depth threshold d max The total number of unique nodes visited during the walk exceeds the pre-specified node number threshold. n max When the walk ends, all visited nodes and the original edges between them together form a local subgraph that maintains the original network structure. G sub = G [ V visited ].

6. The method for predicting potential hypothesis relationships based on a causal symbolic network according to claim 1, characterized in that: In step (4), the potential hypothesis relationship prediction is based on v ego The local subgraph centered at G sub The following steps are performed: (41) From G sub Identify all v ego Candidate variable nodes that are path-connected but not directly adjacent v i , forming a candidate set U={ v 1 , v 2 , …, v i }; (42) For each candidate variable v i U, by calculating v ego and v i The node direction measurement on all paths between them infers the causal direction between the two and quantifies it into a direction score; (43) The cumulative effect of the path between the two is calculated by multiplying the sign effects of all edges on the path, thereby determining the sign of the causal relationship in the new hypothesis, that is, positive or negative; (44) Check the connection v ego and v i Whether there are edges with mediating or moderating properties on the path to identify potential mediating variables and moderating variables in the new hypothesis.

7. The method for predicting potential hypothesis relationships based on a causal symbolic network according to claim 6, characterized in that: The direction score is calculated according to the following formula: ; in, express G sub Middle Connection v ego and v i The set of intermediate nodes on all paths; is the size of the set, d ( x ) is a predefined node direction metric function, the positive or negative value of which indicates the relative position of the intermediate node in the path. v ego forward or backward direction.

8. The method for predicting potential hypothesis relationships based on a causal symbolic network according to claim 6, characterized in that: The cumulative effect is calculated using the following formula: ; in, e j It is the first j edge; effect( e j ) is the symbolic network of edges, k is the length of the path.

9. The method for predicting potential hypothesis relationships based on a causal symbolic network according to claim 1, characterized in that: Step (5) includes the following steps: (51) The number of common neighbors is used as the main criterion for evaluating the structural proximity of two variables, and the number is quantified as the target variable v ego Neighbor set and candidate variables v i The intersection size of the neighbor sets of ,in Representation node v The set of direct neighbors of ; (52) The frequencies of potential mediating and moderating variables identified by the candidate hypotheses are counted as additional evaluation indicators; (53) The number of common neighbors and the frequency of mediating and moderating variables are integrated into a comprehensive credibility score through a weighted sum function; (54) All candidate hypotheses h i Arrange them in descending order according to their comprehensive credibility scores to generate the final ordered set of candidate hypotheses H ={ h 1 , h 1 ,…, h i }.

Citation Information

Patent Citations

  • Event causal relationship identification method and device, computer equipment and storage medium

    CN116628200A

  • Electrical load prediction method and device based on space-time correlation

    CN117175588A

  • Invariant causal discovery method for implementing evaluation for intelligent medical system

    CN117198481A

  • Temporal knowledge graph extrapolation method based on eagle history recognition

    CN118657206A

  • Scientific hypothesis atlas generation method based on hypothesis relation recognition

    CN119150979A