Directed acyclic graph construction method for medical observation data causal inference
Through systematic literature retrieval and adjustment of causal criteria, the directed acyclic graph (DAG) construction method solves the problem of inaccurate DAG construction in causal inference research, achieving higher accuracy and reliability, and supporting causal inference from medical observational data.
Patent Information
- Application Number
- CN202511683387.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-04-21
AI Technical Summary
In existing causal inference studies based on medical observational data, the construction of directed acyclic graphs (DAGs) relies on human experience, resulting in insufficient retrieval results, reduced accuracy of DAGs, and an inability to provide effective statistical strategy support.
By employing a systematic literature retrieval method, we obtained exposure and outcome variables based on the PECO framework, identified third-party causal variables, constructed an initial directed acyclic graph, and introduced causal criteria for adjustment to ensure the accuracy and scientific rigor of the graph.
It improves the accuracy and information content of directed acyclic graphs, provides more reliable quantitative evidence for causal inference studies based on medical observational data, reduces spurious associations, and enhances the scientific validity and credibility of causal relationship judgments.
Smart Images

Figure CN121905558A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of directed acyclic graph technology, specifically to a method for constructing directed acyclic graphs for causal inference of medical observational data. Background Technology
[0002] Directed acyclic graphs (DAGs), as a mathematical model for expressing the causal structure between variables, have been widely used in fields such as epidemiology, public health, and clinical research, providing a theoretical basis for causal inference and bias control in medical observational data causal inference studies.
[0003] In existing causal inference studies based on medical observational data, researchers typically construct DAGs (Directed Acyclic Graphs) based on personal experience or scattered literature. When constructing directed acyclic graphs, the concept of "evidence-based" research can be introduced. This involves systematically reviewing relevant literature evidence and generating a directed acyclic graph based on the PECO (Population, Exposure, Comparator, Outcome) framework, based on the research conclusions in the literature. However, this search strategy primarily focuses on "focal" literature evidence that simultaneously involves "exposure" and "outcome" within the PECO framework. This results in insufficient search results, reduces the accuracy of the directed acyclic graph, and fails to provide effective statistical strategy support for causal inference studies based on medical observational data. Summary of the Invention
[0004] This application provides a method for constructing directed acyclic graphs (DAGs) for causal inference from medical observational data, aiming to achieve more comprehensive literature retrieval, thereby improving the accuracy of DAGs and providing effective statistical strategy support for causal inference research from medical observational data.
[0005] Firstly, this application provides a method for constructing a directed acyclic graph (DAG) for causal inference of medical observational data. The DAG is generated based on the PECO framework. The method for constructing a DAG for causal inference of medical observational data includes: obtaining exposure variables and outcome variables under the PECO framework; retrieving a set of literature including exposure variables and outcome variables from a pre-set medical literature database; determining the respective causal and consequence variable sets for exposure variables and outcome variables in the literature set; determining causal and consequence variables that simultaneously belong to exposure variables and outcome variables in the variable set; and generating common causes, common results, and mediating variables according to the effect direction in the corresponding literature to form an initial DAG including exposure variables and outcome variables.
[0006] In the above embodiments, third-party causal variables are identified by separately retrieving literature sets related to exposure variables, outcome variables, and both, thus constructing a systematic and comprehensive framework for discovering evidence-based third-party causal variables. This method ensures broad coverage of associated variables, effectively avoids subjectivity and omissions caused by manual screening, and significantly improves the objectivity, transparency, and initial completeness of the directed acyclic graph construction.
[0007] In conjunction with some embodiments of the first aspect, in some embodiments, after retrieving a set of literature including exposure variables and outcome variables from a preset medical literature database, the method further includes: determining the literature type of the literature in the literature set; deleting literature of the target type from the literature set; wherein the target type includes at least one of the following: literature without empirical evidence, literature that does not record an effect relationship with the exposure variable or outcome variable, and literature in which the study population is not associated with the population variables under the PECO framework.
[0008] In the above embodiments, by setting a document type screening step, documents lacking empirical content, irrelevant effects, or with mismatched study populations can be automatically excluded. This effectively filters input data sources, ensuring that the evidence used to construct the directed acyclic graph is all high-quality, highly relevant empirical research, thereby improving the accuracy and reliability of the final generated graph from the source.
[0009] In conjunction with some embodiments of the first aspect, in some embodiments, an initial directed acyclic graph including exposure variables and outcome variables is formed, including: generating variable nodes corresponding to common causes, common outcomes, mediating variables, exposure variables, and outcome variables respectively; generating directed edges of the corresponding variable nodes based on the effect directions in the corresponding literature; and generating the initial directed acyclic graph based on the variable nodes and directed edges.
[0010] In the above embodiments, the directed acyclic graph (DAG) is constructed based on the "direction of effect," rather than simple variable associations. This allows the construction of edges in the graph to incorporate not only the existence relationships between variables but also information on the direction and intensity of the effect. This method greatly enhances the accuracy and information content of the DAG, providing a more reliable quantitative basis for subsequent causal relationship determination.
[0011] In conjunction with some embodiments of the first aspect, in some embodiments, after generating an initial directed acyclic graph based on variable nodes and directed edges, the method further includes: adjusting the initial directed acyclic graph based on causal criteria to obtain an adjusted directed graph, wherein the causal criteria include at least the temporal order, biological rationality, consistency and effect strength, and dose-response relationship between different variable nodes; and generating a directed acyclic graph based on the adjusted directed graph.
[0012] In the above embodiments, causal criteria are introduced to adjust the initial graph, integrating the causal inference logic recognized in the field of epidemiology into the graph construction process. This method goes beyond association mining based solely on literature data; by reviewing and correcting the relationships in the graph using scientific criteria, it effectively eliminates spurious associations and significantly improves the causal validity and scientific rigor of the generated graph.
[0013] In conjunction with some embodiments of the first aspect, in some embodiments, generating a directed acyclic graph based on the adjusted directed graph includes: identifying cyclic paths in the adjusted directed graph; adjusting the cyclic paths according to the time order of variables in the actual data to obtain a directed acyclic graph.
[0014] In the above embodiments, a step of identifying and adjusting cyclic paths is added to ensure that the final generated graph structure meets the core definition of "acyclic". This ensures the mathematical validity and structural integrity of the graph, enabling it to be directly applied to standard causal inference algorithms and avoiding analytical errors caused by logical loops.
[0015] In conjunction with some embodiments of the first aspect, after adjusting the cyclic path according to the time order of variables in the actual data to obtain the directed acyclic graph, the method further includes: identifying at least one of the mixed variable nodes, mediating variable nodes, and colliding variable nodes among the multiple variable nodes in the directed acyclic graph; and generating a minimal set of adjustment variables for causal inference studies of medical observational data based on at least one of the mixed variable nodes, mediating variable nodes, and colliding variable nodes.
[0016] In the above embodiments, the constructed directed acyclic graph is directly applied to guide statistical analysis, automatically identifying key variables such as confounding and mediating variables to generate a minimal set of adjustment variables. This method transforms complex causal structure knowledge into a clear and actionable analytical strategy, effectively helping researchers avoid confounding bias and over-adjustment, and contributing to improving the accuracy and credibility of causal inference conclusions from medical observational data studies.
[0017] In a second aspect, embodiments of this application provide a directed acyclic graph construction system for causal inference of medical observational data, which is used to perform the method described in any possible implementation of the first aspect.
[0018] Thirdly, embodiments of this application provide a directed acyclic graph (DAG) construction apparatus for causal inference of medical observational data. The DAG construction apparatus for causal inference of medical observational data includes: one or more processors and a memory; the memory is coupled to the one or more processors and is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the DAG construction apparatus for causal inference of medical observational data to perform the method described as in the first aspect and any possible implementation thereof.
[0019] Fourthly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a directed acyclic graph construction system for causal inference of medical observational data, cause the directed acyclic graph construction system for causal inference of medical observational data to perform the method described in the first aspect and any possible implementation thereof.
[0020] Fifthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a directed acyclic graph construction system for causal inference of medical observational data, cause the directed acyclic graph construction system for causal inference of medical observational data to perform the method described in the first aspect and any possible implementation thereof.
[0021] Understandably, the directed acyclic graph construction system for causal inference of medical observational data provided in the second aspect, the directed acyclic graph construction apparatus for causal inference of medical observational data provided in the third aspect, the computer program product provided in the fourth aspect, and the computer storage medium provided in the fifth aspect are all used to execute the methods provided in the embodiments of this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
[0022] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: by using three different literature retrieval strategies—retrieval based on exposure variables, retrieval based on outcome variables, and retrieval based on both exposure and outcome variables—the literature retrieval for exposure and outcome variables is more comprehensive, and the third-party variables found based on the literature are more complete, thereby improving the accuracy of directed acyclic graphs and providing effective data support for causal inference studies of medical observational data. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of an implementation process of a directed acyclic graph construction method for causal inference of medical observational data in this application embodiment; Figure 2This is another implementation flowchart of the directed acyclic graph construction method for causal inference of medical observational data in the embodiments of this application; Figure 3 This is a schematic diagram of a hybrid variable node in an embodiment of this application; Figure 4 This is a schematic diagram of an intermediate variable node in an embodiment of this application; Figure 5 This is a schematic diagram of a collision variable node in an embodiment of this application; Figure 6 This is a schematic diagram of at least a portion of the physical structure of the directed acyclic graph construction apparatus for causal inference of medical observational data in the embodiments of this application. Detailed Implementation
[0024] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification of this application, the singular expressions “a,” “an,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.
[0025] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0026] The following describes the process of constructing a directed acyclic graph (DAG) for causal inference from medical observational data provided in this implementation. The DAG is generated based on the PECO (Population, Exposure, Comparison, Outcome) framework. Please refer to [link to relevant documentation]. Figure 1 This is a flowchart illustrating a method for constructing a directed acyclic graph for causal inference of medical observational data in this application embodiment.
[0027] S101. Identify the parent and child variable sets V1 of the exposed variables.
[0028] In the embodiments of this application, the PECO framework is a structured tool for precisely defining and articulating a clinical or epidemiological research question. Exposure variables refer to the potential causes, interventions, or risk factors that the PECO framework focuses on; these are variables that researchers wish to assess the effects of, such as the use of a certain drug, a specific lifestyle, or environmental exposure. Outcome variables refer to the health conditions, disease events, or any relevant measurements that researchers wish to observe and measure that are potentially associated with the exposure variables; these are variables that researchers wish to observe and measure, such as the incidence of a specific disease, patient survival time, or changes in the level of a biomarker. Obtaining this pair of exposure and outcome variables is the starting point for constructing a directed acyclic graph. The predefined medical literature database refers to one or more recognized electronic information resources that store published medical research findings.
[0029] First, a global literature search is conducted focusing on the exposure variables in the study. By systematically retrieving literature evidence related to the exposure variables, upstream influencing factors (parent nodes) and downstream outcomes (child nodes) affected by them are identified, forming a set of exposure-related variables. This set is used to depict the antecedent and consequence relationships of the exposure variables in the causal structure, laying the foundation for the subsequent identification of outcome variables and intersection variables.
[0030] In some embodiments of this application, the process of obtaining exposure and outcome variables can automatically identify and extract potential exposure-outcome pairs from unstructured clinical research protocol text, free text of electronic medical records, or published research abstracts using a Natural Language Processing (NLP) model. For example, a Named Entity Recognition (NER) model can be trained to automatically identify "drug" or "behavior" entities as exposure variables and "disease" or "symptom" entities as outcome variables from a large amount of text.
[0031] S102, Identify the set of parent and child node variables of the outcome variable V2.
[0032] Subsequently, a systematic search was conducted on the outcome variables in the study using the same method to identify all upstream influencing factors and downstream outcome variables related to the outcome, forming a set of outcome-related variables. This set is used to reveal the direct or indirect determinants of the outcome variables in the causal structure and their extended effects, ensuring a more comprehensive description of the causal path of the outcome in the study.
[0033] In some embodiments of this application, the retrieval strategy may employ structured query statements that not only include standard medical terms for exposure and outcome variables (such as MeSH (Medical Subject Headings) entries), but also combine their synonyms, hyponyms, and related free text keywords through Boolean logic operators. This allows for a more intelligent and comprehensive capture of all evidence related to the core research question, effectively reducing information omissions caused by incomplete retrieval strategies.
[0034] S103. Identify the intersection variables V1∩V2 and classify them according to the direction of the directed edges to generate a third-party variable set.
[0035] After obtaining the sets of variables for exposure and outcome, a cross-comparison was performed to identify variables that were proven to be associated with both exposure and outcome. These variables were then structurally categorized based on the reported effects in the literature: variables affecting both exposure and outcome were classified as common causes; variables located in the causal chain from exposure to outcome were classified as intermediate variables; and variables affected by both exposure and outcome were classified as common outcomes. This classification clarifies the role of third-party variables in the causal relationship.
[0036] In some embodiments of this application, the step of identifying third-party variables (e.g., common causes, common outcomes, mediators) can be accomplished through an integrated information extraction system. This system first utilizes a relation extraction model to automatically identify sentences mentioning exposure variables, outcome variables, and other medical concepts from the abstracts or full texts of each literature collection. Next, the system applies a causal relationship identification algorithm to these sentences. This algorithm is based on a pre-trained language model (e.g., BERT (Bidirectional Encoder Representations from Transformers) or GPT (Generative Pre-trained Transformer)) and fine-tuned for medical causal expression patterns. It can determine whether a causal, association, or influence relationship exists between two entities in a sentence and determine its direction. For example, the system can automatically identify language patterns such as "A causes B" or "The risk of B increases due to A." Through comprehensive automated processing of the three literature collections, the system generates a structured list of covariates, where each covariate is labeled with its source literature, associated variable (exposure variable, outcome variable, or both), and the type and direction of the relationship. This approach automates the tedious manual document review process, ensuring the objectivity and repeatability of covariate determination, and is capable of handling a volume of documents far exceeding the capacity of human review.
[0037] S104. Based on the third-party variable set, generate a directed acyclic graph that includes exposure variables and outcome variables.
[0038] In the embodiments of this application, the process of generating a directed acyclic graph (DAG) involves using all variables (exposure variables, outcome variables, and all third-party variables) determined in the preceding steps as nodes and the established associations or causal relationships between them as directed edges, thereby constructing an initial network graph. This network graph is then checked and refined to remove any possible cyclic paths (e.g., by imposing time sequence constraints or judgments based on biological mechanisms), ultimately forming a DAG that conforms to the definition. This graph visually illustrates the complex causal network between the variables in a graphical way.
[0039] The method provided in this application is based on the inventive concept of systematically collecting evidence through a hierarchical and focused literature retrieval strategy, and distinguishingly identifying covariates that are associated only with the exposure variable, only with the outcome variable, and both. Finally, a directed acyclic graph is constructed based on these sorted and categorized variable sets, ensuring that all potential key variables are systematically considered, thereby improving the comprehensiveness and accuracy of the final generated directed acyclic graph.
[0040] Reference Figure 2 ,exist Figure 1 Based on the illustrated embodiment, a method for constructing a directed acyclic graph for causal inference of medical observational data is described below: 1. Define causal relationships.
[0041] The PECO framework (Population, Exposure, Comparator, Outcome) is used to clearly define the research objective and key variables. For example: "In a specific population, is exposure to X (as opposed to no exposure) associated with an increased / decreased risk of outcome Y?" This approach helps to break down research questions into basic elements, ensuring that subsequent searches and analyses do not deviate from the research objectives.
[0042] 2. Global evidence retrieval.
[0043] Based on the clearly defined research question, develop a search strategy targeting exposure and outcome, and conduct a global search (union of documents rather than intersection), including: a) Research surrounding the exposure of X, used to find the parent and child nodes of X; b) Research surrounding the outcome Y, used to find the parent and child nodes of Y; c) Simultaneously involving research on X and Y, directly exploring the relationship between X and Y.
[0044] The above search includes evidence of the focal relationship (i.e., X and Y), as well as evidence of the influencing factors and outcome events of X or Y respectively, ensuring a global search of "cause and effect".
[0045] 3. Evidence included in exclusion.
[0046] Inclusion criteria: The evidence to be included should include the following:
[0047] a. Based on the evidence pyramid of evidence-based medicine, research results from different levels of evidence were included; b. If relevant systematic syntheses and meta-analyses exist, they should be included first to improve the reliability of the evidence. When multiple meta-analyses exist, an umbrella review can be conducted to integrate the existing meta-analyses; c. You need to read the abstract or the full text to extract the conclusive statements related to the effects between variables. For example, the conclusion section "Exposure to X significantly increases / decreases the risk of outcome Y..." and the background section "X is a widely recognized risk factor for Y", and similar statements.
[0048] Exclusion criteria: Exclude the following document types: a. Commentary, opinion pieces, editorials, research proposals, and other literature that lacks empirical research results or only proposes theoretical hypotheses; b. There is no literature evidence clearly showing a relationship between the report and the exposure or outcome; c. Literature evidence showing a significant discrepancy between the population characteristics defined in the research question and the actual population characteristics.
[0049] 4. Information extraction.
[0050] Indicator extraction and structuring: Extract information such as effect indicators, effect size, confidence intervals, and effect direction related to the relationship between exposure and outcome. Use the PECO framework to uniformly represent different research results in PECO format, for example: "In population P, exposure E (compared to control C) significantly increases the risk of outcome O." This can be achieved using a combination of cutting-edge AI tools such as large language models and human review.
[0051] Terminology standardization: This involves unifying the naming of synonymous or different-level variables appearing in various documents. For example, "stroke" and "apoplexy" can be mapped to the same code, and nodes can be merged. Common medical terminology coding systems, such as MeSH, ICD (International List of Causes of Death), and UMLS (Unified Medical Language System), can be used to standardize synonymous or related hierarchical concepts according to the research question. This mapping allows for the merging of variables with different expressions but the same meaning, ensuring that each node in the DAG (Directed Acyclic Graph) has a unique and clear meaning.
[0052] 5. Initial construction and completion of directed acyclic graph.
[0053] Construct an initial directed acyclic graph based on literature evidence: Following the four processes of defining the research question, global evidence retrieval, evidence inclusion and exclusion, and information extraction, the causal and consequence variables of each exposure-outcome to be studied, as well as a structured table of supporting evidence, were formed. Based on this table and the following steps, an initial directed acyclic graph was constructed (see...). Figure 1 ): a) Identify the exposed set of parent and child nodes, V1; b) Identify the set of parent and child nodes V2 of the ending; c) Identify the set of variables that simultaneously belong to both V1 and V2, i.e., the intersection of the two (V1∩V2). Based on the direction of the arrows, common causes, common results, and intermediate variables can be identified (see...). Figure 1 ); d) Connect the variables supported by the literature with directed edges to construct an initial graph, and create an "edge index" for each edge, labeling the supporting literature and conclusions; e) The variables involved in the initial DAG construction should, in principle, include four main categories: 1) Demographic variables (gender, age, socioeconomic status, etc.); 2) Comorbidities or disease variables; 3) Medication variables; 4) Lifestyle and behavioral variables.
[0054] The initial DAG described above has identified the common parent node and common child node of the exposure and outcome variables, as well as the intermediate nodes from exposure to outcome. Next, the pairwise relationships between the variables involved in the initial DAG need to be completed, and the following principles are recommended.
[0055] a) Whether to add an arrowed connection between any two variables in the initial DAG can be determined based on expert experience. Since there are many pairwise combinations of variables, consulting every single literature review is practically very difficult.
[0056] b) If expert experience cannot determine the relationship between certain variables, it is advisable to use authoritative knowledge search platforms at home and abroad to obtain literature on whether there is an influence relationship between the two variables.
[0057] 6. DAG optimization.
[0058] Due to the heterogeneity of medical research, the above literature retrieval and processing procedures may lead to conflicting results or even reverse causality. Causality criteria can be used, combined with available data, for further examination and processing. a) Chronological order: Exposure occurred before the outcome; b) Biological rationale: The causal relationship is supported by biological mechanisms; c) Consistency and effect size: Multiple studies consistently support this finding, and the effect size is relatively large; d) Dose-response relationship: There is a positive correlation between the level of exposure and the degree of risk of the outcome.
[0059] Relationships that do not meet the above criteria may be temporarily deleted or marked as "uncertain" and submitted to relevant technical personnel for discussion and decision.
[0060] Cyclic paths should be handled according to the time order of variables: Cyclic paths are often caused by reverse causality (e.g., X→Y and Y→X coexist). The time order needs to be clarified according to the specific research context, and the process should be carried out with the help of actual data or clinical expert groups.
[0061] After completing the above construction, completion, and processing procedures, a directed acyclic graph suitable for the target research is formed.
[0062] 6. Identify third-party variables.
[0063] In the completed DAG, the following three types of key third-party variables need to be clearly identified (see...). Figure 2 ).
[0064] Confounding variables are common causes that simultaneously affect both exposure (e.g., starting medication after a certain condition occurs) and outcome (e.g., liver injury). Their presence can lead to spurious associations between drugs and adverse events. In a DAG, they appear as a "fork" structure pointing from the variable to both exposure and outcome, such as... Figure 3 .
[0065] Mediating variables: These are the intermediate transmission mechanisms between exposure and outcome, representing the biological or clinical pathway of the effect of the exposing factor. In a DAG, they manifest as a "chain" structure where exposure points to mediation, and mediation then points to the outcome, such as... Figure 4 .
[0066] Collision variables: Collision variables are the combined result of being affected by both exposure and outcome. In a DAG, this is represented by two arrows pointing to the same node, forming a "collision" structure, such as... Figure 5 .
[0067] 7. Generate the minimum set of adjustment variables.
[0068] After defining the roles of confounding, mediating, and collision variables, a minimal set of sufficient adjustments (i.e., the minimal set of adjusted variables) can be generated for statistical analysis. This set consists of only the covariates that must be included in the statistical model when blocking non-causal paths between exposure and outcome based on the DAG. This set aims to reduce bias caused by omissions in adjustment while avoiding estimation bias or decreased model efficiency due to over-adjustment.
[0069] In practice, researchers can use specialized tools to identify the minimum adequate adjustment set. By drawing or importing a pre-constructed directed acyclic graph (DAG), and labeling the exposures and outcomes of interest, the d-separation principle of the graph structure—that is, determining whether a given set of variables makes the exposures and outcomes conditionally independent in the graph—can automatically retrieve one or more sets of variables that can block all "backdoor paths." Researchers can then select the adjustment set that is most applicable and smallest in size from the recommended list, taking into account their specific research questions and data availability.
[0070] The adjustment strategy for the guiding variables is as follows: Confounding variables: Include the minimum adequately adjusted set in the model. Potential confounding factors that cannot be measured can be marked with "U" in the DAG to indicate residual confounding risk.
[0071] Mediating variables: If the goal is to estimate the total effect of exposure on the outcome, they are generally not adjusted; if the study is about direct effects or mechanisms of action, they can be selectively adjusted.
[0072] Collision variables: These are not included in the model to avoid opening new non-causal paths and introducing bias.
[0073] As can be seen, the embodiments of this application address the limitations of directed acyclic graph construction methods in terms of evidence retrieval scope, information extraction efficiency, structural traceability, conflict resolution capabilities, and dynamic update capabilities, proposing a series of technical innovations and achieving significant technical improvements, specifically reflected in: Expanding the scope of evidence coverage: The proposed "three-channel global evidence retrieval" strategy systematically identifies potential covariates and third-party variables, ensuring the integrity and scientific validity of the causal structure and significantly expanding the breadth of evidence coverage.
[0074] Improve information extraction efficiency: Utilizing large language models to drive automated parsing and structured extraction of evidence replaces manual reading and organization, significantly improving the efficiency and consistency of information extraction and reducing the impact of human bias on the quality of the construction.
[0075] Enhance the traceability of causal structures: By introducing the "edge-level evidence index" mechanism, each directed edge is labeled with structured information such as its evidence source, research design, and statistical results, making the construction process of the causal graph fully auditable and transparent, which facilitates subsequent verification and reuse.
[0076] Automating conflict resolution: The proposed conflict resolution mechanism based on causal criteria can automatically identify and handle evidence conflicts, reverse causality, or cyclic path problems, reducing reliance on human judgment and improving the logical consistency and stability of causal structures.
[0077] Supports dynamic evolution and version management: It supports automatic identification and incremental updates of new literature evidence, eliminating the need to rebuild the entire graph every time new evidence emerges. Simultaneously, the introduction of version management functionality allows for the recording, comparison, and tracking of historical versions of the causal graph, significantly improving the maintainability and long-term application value of the directed acyclic graph construction system used for causal inference from medical observational data.
[0078] In summary, the embodiments of this application have achieved fundamental improvements in terms of the breadth of evidence acquisition, efficiency of structure construction, traceability of causal relationships, degree of automation in conflict resolution, and dynamic evolution capability of the system, providing more systematic, scalable, and reusable technical support for fields such as epidemiological research and real-world data analysis.
[0079] The product form of this application embodiment can be used as a standalone software platform, or it can be integrated into a related statistical analysis system or knowledge management platform, and can be provided through local deployment or cloud services.
[0080] The following is a more detailed description of the process of the method provided in this implementation. Figure 1 and Figure 2 Based on the embodiment shown, this is another flowchart illustrating the method for constructing a directed acyclic graph for causal inference of medical observational data in this application.
[0081] S201. Obtain the exposure variables and outcome variables under the PECO framework.
[0082] This step can be referred to in the previous embodiments, and will not be repeated here.
[0083] S202. In a pre-defined medical literature database, retrieve a first set of literature including exposure variables, a second set of literature including outcome variables, and a third set of literature including both exposure variables and outcome variables.
[0084] The first literature set (comprising both exposure and outcome variables), the second literature set (comprising both causal and consequential variables of the exposure variable), and the third literature set (comprising both causal and consequential variables of the outcome variable) are subsets of the results from three separate search operations. The first search aims to broadly collect all literature related to the exposure variable, resulting in the first literature set, to comprehensively understand its background and potential influencing factors. The second search aims to broadly collect all literature related to the outcome variable, resulting in the second literature set, to comprehensively understand its causes and related conditions. The third search focuses on literature that directly studies the relationship between the exposure and outcome variables, resulting in the third literature set.
[0085] S203. Determine the document type of each document in the first document set, the second document set, and the third document set.
[0086] In the embodiments of this application, document type is a classification label for a medical document based on its research design, content nature, and level of evidence. Its scope covers a variety of categories, such as randomized controlled trials (RCTs) based on research design, cohort studies, case-control studies, cross-sectional studies, and non-original studies such as systematic reviews, meta-analyses, narrative reviews, guidelines, expert consensus, editorials, commentaries, and case reports. The significance of determining the document type lies in providing an objective classification basis for subsequent evidence screening.
[0087] In some embodiments of this application, the process of determining document types can be implemented using a machine learning-based automated classification system. This automated classification system constructs a deep learning text classification model, for example, using BERT (Bidirectional Encoder Representations from Transformers), pre-trained on a large corpus of medical literature, as its infrastructure. The model is then fine-tuned in a supervised manner by collecting a large-scale medical literature dataset with accurate document type labels. In practical applications, the automated classification system extracts the title, abstract, and available full-text text of each document to be classified as input. The BERT model, through deep semantic understanding, automatically identifies key patterns and structured features in the text that reflect the research design (such as "random assignment," "follow-up," "retrospective analysis," etc.) and structural features (such as the presence of "methods" and "results" sections), thereby accurately assigning one or more type labels to each document, and then determining the document type based on these type labels. Compared to relying solely on potentially incomplete or inaccurate metadata labels provided by the database, this approach achieves higher precision automated classification, ensuring the reliability of the document screening foundation.
[0088] S204. For at least one of the first document set, the second document set, and the third document set, delete documents whose document type is the target type.
[0089] In the embodiments of this application, this step is an execution process that cleanses one or more document sets based on the document type determined in the previous step. The deletion operation refers to removing documents that meet the target type criteria from the corresponding set, so that they no longer participate in the subsequent covariate extraction and graph construction processes.
[0090] The target type includes at least one of the following: literature without empirical evidence, literature that does not record an effect relationship with the exposure variable or outcome variable, and literature in which the study population is not associated with the population subjects under the PECO (Population, Exposure, Comparator, Outcome) framework.
[0091] In the embodiments of this application, the specific characteristics of the documents that need to be deleted are defined in detail: Non-empirical literature refers to literature that does not contain primary or secondary empirical research data. This primarily includes narrative reviews, editorials, opinion pieces, commentaries, conference abstracts, and papers that only describe theoretical models or hypotheses. The significance of removing such literature lies in firmly establishing the analytical foundation on objective data evidence obtained through experiments or observations, avoiding the introduction of subjective opinions or unverified speculations into causal networks.
[0092] Literature that does not document an effect relationship with the exposure or outcome variables: This refers to literature that, although it may mention exposure or outcome variables, does not focus on their association or causal effects with other variables. For example, an article describing a drug's chemical synthesis method may involve the exposure variable (drug) but does not study its clinical effects; or a literature describing the epidemiological distribution of a disease may involve the outcome variable (disease) but does not explore its risk factors. The significance of deleting such literature is to ensure that the selected set of literature is highly relevant to the goal of constructing a causal network, improving the efficiency and accuracy of subsequent information extraction.
[0093] Literature whose study population is not related to the population group defined in the PECO framework: This refers to literature whose study subjects do not match or are completely different from the population group predefined in the PECO framework in terms of key characteristics. For example, if the population group defined in this study is "adult patients with diabetes," then a study conducted only in children or non-diabetic populations, even if it studied the same exposures and outcomes, should be considered as target-type literature. Its significance lies in ensuring the consistency and applicability of evidence, ensuring that the constructed directed acyclic graph is targeted at a specific target population, thereby avoiding erroneous causal inferences due to the inclusion of studies in irrelevant populations.
[0094] In some embodiments of this application, in order to more precisely identify and delete the above-mentioned target types, a multi-label, multi-task intelligent screening system can be constructed. This intelligent screening system not only performs the aforementioned document type classification task, but also performs two additional binary classification tasks in parallel: (1) effect relationship classification, to determine whether the document describes an effect relationship; (2) population matching degree classification, to determine whether the research population of the document matches the preset PECO population. For the determination of population matching degree, the system can use named entity recognition technology to automatically extract the key features of the research population (such as age, gender, disease status, region) from the document, and perform vectorization representation and similarity calculation with the preset PECO population features. Only when a document simultaneously meets the three conditions of "is an empirical document", "contains an effect relationship" and "population matching degree is higher than the preset threshold" will it be retained. This scheme achieves in-depth filtering of document quality and relevance through multi-dimensional and automated content review, which greatly improves the purity of evidence finally included in the analysis.
[0095] S205. In the first literature set, determine the first covariate associated with the exposure variable; in the second literature set, determine the second covariate associated with the outcome variable; and in the third literature set, determine the third covariate associated with both the exposure variable and the outcome variable.
[0096] The first covariate related to the exposure variable is the causal and consequential variable of the exposure variable. The second covariate related to the outcome variable is the causal and consequential variable of the outcome variable. The third covariate related to both the exposure and outcome variables is the causal and consequential variable that belongs to both the exposure and outcome variables.
[0097] S206. In the first literature set, determine the direction of the first effect between the exposure variable and the first covariate; in the second literature set, determine the direction of the second effect between the outcome variable and the second covariate; and in the third literature set, determine the direction of the third effect between the exposure variable and the outcome variable and the third covariate.
[0098] In the embodiments of this application, the effect direction (e.g., first effect direction, second effect direction, third effect direction) refers to a numerical indicator extracted from the statistical analysis section of medical literature, used to quantify the strength and direction of the association between two or more variables. Its scope may include, but is not limited to: odds ratio (OR), relative risk (RR), hazard ratio (HR), regression coefficient (e.g., the beta coefficient of linear regression), correlation coefficient (e.g., the Pearson correlation coefficient r), etc. Typically, the effect direction is reported along with its confidence interval (CI) and p-value (a parameter used to determine the result of hypothesis testing); the former indicates the accuracy of the estimate, and the latter indicates its statistical significance. The significance of determining the effect direction lies in that it elevates the description of the relationship between variables from a qualitative "existence of association" to a quantitative "how strong the association is," providing core data for constructing a directed acyclic graph with weighted information that reflects the strength of different causal paths.
[0099] S207. Generate the variable nodes corresponding to the first covariate, second covariate, third covariate, exposure variable, and outcome variable, respectively.
[0100] In the embodiments of this application, a variable node is a graphical and data-driven abstract representation of each research variable (including exposure variables, outcome variables, and all covariates) when constructing a directed acyclic graph. All covariates include a first covariate, a second covariate, and a third covariate, and can be categorized as common causes, common outcomes, and mediating variables. Each variable node is a basic unit or vertex in the graph. In software implementation, a variable node is typically a data structure or object containing a unique identifier, a variable name (e.g., "smoking," "lung cancer," "age"), and other relevant attributes.
[0101] S208. Based on the first effect direction, the second effect direction, and the third effect direction, generate directed edges for the corresponding variable nodes.
[0102] In the embodiments of this application, a directed edge is a line segment with a clear direction connecting two variable nodes, and is the core connection element in a directed acyclic graph. It represents a potential association or causal path between two variables, supported by literature data, pointing from "cause" to "effect." The generation of directed edges is based on the effect direction extracted in previous steps; a connection is established between the corresponding variable nodes only when the effect direction shows a statistically significant relationship between the two variables. The direction of the edge is determined based on time-series information in the literature (e.g., exposure occurring before the outcome), study design, or explicit causal arguments.
[0103] In some embodiments of this application, the directed edge can be a "probability-weighted evidence composite directed edge." In this approach, when multiple documents provide the direction of the effect on the relationship between the same pair of variables, instead of simply generating an edge, the evidence is first integrated. Specifically, a mini meta-analysis is automatically performed to calculate the combined effect estimate and its confidence interval. The core properties of the directed edge are defined by this more robust merging result. The weight of the directed edge is set to the magnitude of the combined effect, thereby directly encoding the uncertainty and strength of the documentary evidence into the structure of the directed acyclic graph, allowing the generated directed edge to more accurately reflect the existing knowledge state regarding the causal relationship.
[0104] In some embodiments of this application, the specific procedure for performing a micro-meta-analysis to calculate the pooled effect estimate is as follows: Taking the ratio of odds between multiple studies on the same pair of variables (e.g., exposure variable A and covariate B) as an example, assuming three independent OR values and their 95% confidence intervals for the relationship between A and B are extracted from three studies, the data processing flow includes: Data transformation: Since the distribution of the odds ratio is skewed, a logarithmic transformation is first performed to approximate it as a normal distribution, which facilitates statistical aggregation. For each study i, its log-odds ratio y_i = ln(OR_i) is calculated.
[0105] Variance calculation: The standard error (SE) and variance (Var) of each log-ratio ratio are calculated based on the confidence interval. The standard error SE_i can be calculated from the upper and lower limits of the confidence interval (UL_i and LL_i): SE_i = (ln(UL_i) - ln(LL_i)) / (2 × 1.96). The variance Var_i = (SE_i)^2.
[0106] Weighting: Each study is assigned a weight w_i using the inverse variance method. Studies with smaller variances have more accurate results and should be assigned a larger weight. The weight calculation formula is w_i = 1 / Var_i.
[0107] Pooled effect size calculation: Calculate the pooled log-odds ratio y_pooled, which is the weighted average of the log-odds ratios of all studies. The formula is: y_pooled = (Σ(y_i × w_i)) / (Σ(w_i)).
[0108] Calculation of pooled variance and confidence interval: Calculate the standard error SE_pooled and variance Var_pooled of the pooled effect size. Var_pooled = 1 / (Σ(w_i)), therefore SE_pooled = sqrt(Var_pooled). Based on this, calculate the 95% confidence interval of the pooled effect size: ln(CI_pooled) = y_pooled ± 1.96 × SE_pooled.
[0109] Inverse transformation: The pooled log-odds ratio and its confidence interval are inversely converted back to the original odds ratio scale. The pooled odds ratio OR_pooled = exp(y_pooled), and its 95% confidence interval is [exp(ln(LL_pooled)), exp(ln(UL_pooled))].
[0110] In some embodiments of this application, the specific procedure for performing a micro-meta-analysis to calculate the pooled effect estimate is as follows: Due to the heterogeneity among different studies—meaning that the true effect sizes may differ—simple merging can obscure important information and even lead to misleading conclusions when heterogeneity is high. Therefore, a dynamic evidence synthesis and heterogeneity quantification approach based on a Bayesian hierarchical model can be adopted. This approach does not simply calculate a fixed pooled effect value, but instead constructs a Bayesian hierarchical model to dynamically synthesize evidence. This Bayesian hierarchical model can not only estimate the pooled effect but also simultaneously quantify and explain the heterogeneity among studies, assigning a probability confidence level to the existence of each edge.
[0111] Specifically, when multiple articles are identified reporting a relationship between the same pair of variables (e.g., exposure variable A and covariate B), the following data processing flow will be executed: Model Construction: Instead of using a fixed merging formula, a Bayesian hierarchical model is constructed. This Bayesian hierarchy contains two levels: The first layer (data layer) assumes that the direction of the effect extracted from each literature i (e.g., the log-odds ratio y_i) follows a normal distribution with the mean of the true effect θ_i of the study, i.e., y_i~N(θ_i, σ_i^2), where σ_i^2 is the known variance calculated based on the confidence interval (CI) reported in the literature.
[0112] The second layer (parameter layer): It is assumed that the true effect θ_i of each study is sampled from a common hyperparameter distribution, representing the overall effect distribution of all potential studies. Specifically, θ_i ~ N(μ, τ^2), where μ represents the average true effect of all studies (i.e., the desired pooled effect), and τ^2 is a key parameter representing the heterogeneity variance among studies. A larger τ^2 indicates a greater difference in the true effects among different studies.
[0113] Bayesian inference involves setting an informative prior distribution for the hyperparameters μ and τ^2, and then sampling the posterior distribution of the model using the Markov Chain Monte Carlo (MCMC) algorithm. By sampling extensively from the posterior distribution, the complete probability distributions of μ and τ^2 can be obtained, rather than just point estimates.
[0114] The merging effect and uncertainty: The median or mean of the posterior distribution is used as the best estimate of the merging effect μ, and its uncertainty is quantified by the 95% confidence interval of the posterior distribution.
[0115] Heterogeneity quantification: The τ^2 of the posterior distribution directly quantifies the magnitude of heterogeneity among studies. Based on this, the I^2 statistic (I^2=τ^2 / (τ^2+σ_pooled^2)) can be calculated, which visually shows the percentage of heterogeneity in the total variation.
[0116] Probabilistic generation of directed edges: The posterior probability P(μ>0|data) of the combined effect μ being greater than 0 (or less than 0, depending on the direction of the effect) can be calculated. The existence of this directed edge is no longer a binary decision based on whether P is less than 0.05, but is assigned a specific probability value. For example, if P(μ>0|data) = 0.98, a high-confidence directed edge can be generated; if the probability is 0.75, a low-confidence directed edge is generated, which may require further review by relevant technical personnel.
[0117] It can be seen that this scheme transforms evidence synthesis from a deterministic computational process into a probabilistic inference process. It not only provides a more robust merging effect but also incorporates inter-study heterogeneity as an endogenous, quantifiable consideration in constructing the directed acyclic graph (DAG), using posterior probabilities to determine edge generation. This allows the constructed DAG to more realistically reflect the complexity and uncertainty of current scientific evidence.
[0118] S209. Generate an initial directed acyclic graph based on the variable nodes and directed edges.
[0119] In the embodiments of this application, the initial directed acyclic graph (DAG) is a preliminary network structure formed by directly and unreviewedly combining all variable nodes and directed edges generated in the preceding steps. At this stage, all independent connections (i.e., directed edges) based on documentary evidence are presented in a unified graph. This initial DAG serves as the starting point for subsequent logical review and structural optimization; it may contain logically or biologically illogical connections, and may even contain directed cycles, therefore it is not the final DAG.
[0120] S210. Based on the causal criteria, the initial directed acyclic graph is adjusted to obtain the adjusted directed graph. The causal criteria include at least one of the following: temporal order between different variable nodes, biological rationality, consistency and effect strength, and dose-response relationship.
[0121] In the embodiments of this application, this step involves reviewing and revising the initial directed acyclic graph (DAG). The adjustment operation refers to modifying the directed edges in the initial DAG by adding, deleting, or reversing their direction according to a series of generally accepted causal inference criteria, thereby improving the scientific validity of the entire initial DAG. Causal criteria may include, for example, Hill's criteria of causation. Hill's criteria of causation are a set of guiding principles used in epidemiology to determine whether an association is likely causally related. This embodiment of the application selects the most operable parts of these criteria as the basis for adjustment: Time order: The cause must precede the effect. Each directed edge in the initial directed acyclic graph (e.g., from A to B) can be examined, and the metadata of variable nodes A and B can be queried. If the time attributes of the variable nodes (such as disease diagnosis time, drug use start time) indicate that B occurred before A, then this edge violates the time order rule and should be deleted or reversed.
[0122] Biological justification: The proposed causal relationship should conform to existing biological or medical knowledge. For example, if there is an edge in the initial directed acyclic graph pointing from "height" to "genotype", this is clearly not in line with biological principles and should be deleted.
[0123] Consistency and effect strength: If a consistent effect with a strong association (a large absolute value of the effect direction) is observed in multiple different studies, then the association is more likely to be causal. Therefore, based on this criterion, edges that are consistently supported by multiple studies and have significant combined effect sizes can be prioritized for retention, while edges that are supported only by a single, small-sample study or have weak effect sizes can have their confidence level reduced or be deleted.
[0124] Dose-response relationship: As the dose or level of exposure increases, the risk of the outcome also increases. If such information is included in the literature evidence, it can serve as strong evidence confirming the direction and existence of a directed edge.
[0125] In some embodiments of this application, the adjustment step of the initial directed acyclic graph (DAG) can be implemented using a "causal logic reasoning engine." This engine internally stores a formalized medical prior knowledge base, which encodes a large amount of generally accepted medical knowledge in the form of logical rules (e.g., "diseases cannot cause gene mutations," "treatment usually occurs after diagnosis"). When reviewing the initial DAG, the engine matches its structure against the rules in the knowledge base. For any detected conflicts (such as violations of chronological order or biological rationality), the engine marks the conflict and provides adjustment suggestions, which are then confirmed and implemented by the user; or in some embodiments, the engine automatically performs adjustments according to preset rules to obtain the adjusted directed graph.
[0126] S211. Identify cyclic paths in the adjusted directed graph.
[0127] In the embodiments of this application, a cyclic path, also known as a directed cycle or loop, refers to a sequence (e.g., A, B, C, ..., A) of three or more variable nodes in an adjusted directed graph, where there are directed edges from each node in the sequence to the next node, and the last directed edge points back from the last node of the sequence to the first node (e.g., A→B, B→C, ..., →A), forming a closed directed loop. Identifying cyclic paths is a key prerequisite for ensuring that the final graph structure satisfies the definition of a directed acyclic graph. In the context of causal inference, a cyclic path usually implies a logical contradiction (e.g., A leads to B, B leads to C, and C leads to A), which is not valid in most biomedical scenarios and therefore needs to be eliminated.
[0128] S212. Adjust the cyclic paths according to the time order of variables in the actual data to obtain a directed acyclic graph.
[0129] In the embodiments of this application, adjusting cyclic paths involves intervening in each cyclic path identified in the previous step to break the cycle, thereby transforming the entire graph structure into acyclic. Specific adjustments may include deleting one or more directed edges from the cycle, or reversing the direction of one or more directed edges. The ultimate goal is to eliminate all directed cycles in the adjusted directed graph while preserving as much of the causal structure as possible in other parts of the adjusted directed graph, so that the final directed acyclic graph best conforms to existing scientific evidence and knowledge.
[0130] In some embodiments of this application, the process of adjusting a cyclic path can be performed by an "evidence-level conflict resolution engine". Specifically, once a cyclic path is identified, the conflict resolution engine is immediately activated. It automatically retrieves the original evidence corresponding to each directed edge constituting the cyclic path and scores these pieces of evidence comprehensively. The scoring criteria are multi-dimensional and comprehensively consider: (1) the type of literature from which the evidence is sourced, i.e., the level of evidence (e.g., evidence from randomized controlled trials scores the highest, followed by prospective cohort studies, and the lowest is cross-sectional studies or case reports); (2) the statistical significance of the effect direction (e.g., the smaller the p-value, the higher the score); (3) the consistency of the effect direction (i.e., the number of studies supporting the edge and the degree of consistency of the results); and (4) whether the directed edge is consistent with pre-set, strong biological prior knowledge. After scoring all directed edges in a cyclic path, the conflict resolution engine follows the "minimum intervention principle," automatically selecting the directed edge with the lowest score (representing the weakest and most uncertain causal link in the entire cyclic path) and executing a preset adjustment operation (e.g., the default operation is to delete the directed edge). This approach transforms the complex and highly subjective task of resolving cyclic conflicts into an automated process based on quantified evidence evaluation and with clearly defined rules. This not only significantly improves the efficiency of constructing directed acyclic graphs but also makes adjustment decisions more objective, transparent, and reproducible, ensuring that the final directed acyclic graph is an optimized structure obtained while retaining the strongest evidence.
[0131] In some embodiments of this application, the specific data processing logic for the conflict resolution engine to perform a comprehensive score on each directed edge in a cyclic path is as follows: The conflict resolution engine calculates an evidence strength score S(e) for each directed edge e. This evidence strength score is a normalized value between 0 and 1; the lower the score, the weaker the evidence. The formula for calculating S(e) is: S(e)=w_1×Score_Type(e)+w_2×Score_Sig(e)+w_3×Score_Cons(e) Among them, w_1, w_2, and w_3 are preset weight coefficients that satisfy w_1+w_2+w_3=1. For example, they can be set to 0.5, 0.3, and 0.2 respectively, indicating that the importance attached to the evidence type is the highest.
[0132] The calculation logic for each item's score: Evidence Type Score_Type(e): This score is based on the research design type of the strongest source of evidence supporting the directed edge. Specifically, the directed acyclic graph construction system used for causal inference from medical observational data internally maintains a sequence list of research design types and their corresponding scores (between 0 and 1): Randomized controlled trials (RCTs): 1.0; Prospective cohort study: 0.8; Case-control study: 0.6; Cross-sectional study: 0.3; Case report / expert opinion: 0.1.
[0133] The conflict resolution engine can retrieve all documents that support edge e and take the highest score of evidence among them as Score_Type(e).
[0134] The statistical significance score, Score_Sig(e), is based on the p-value supporting the direction of the effect along the directed edge. If there are multiple references, the combined p-value is used. The score is calculated using a non-linear transformation to highlight the importance of very small p-values: Score_Sig(e) = 1 - p_value. For example, a p-value of 0.001 results in a statistical significance score of 0.999; a p-value of 0.05 results in a statistical significance score of 0.95.
[0135] Consistency score Score_Cons(e): This score measures the consistency of multiple research findings supporting the directed edge. It consists of two parts: consistency in the number of studies and consistency in the direction of the effect.
[0136] Let N be the total number of studies supporting the directed edge.
[0137] Let N_pos be the number of studies with a positive effect direction, and N_neg be the number of studies with a negative effect direction.
[0138] Consistency ratio p_cons = max(N_pos, N_neg) / N.
[0139] The consistency score Score_Cons(e) = p_cons × (1 - exp(-k × N)), where k is an adjustment coefficient (e.g., 0.5) that makes the score closer to p_cons as the number of studies increases.
[0140] Assuming a cyclic path A→B→C→A is identified, the conflict resolution engine begins to score the three directed edges: Directed edge A→B: The strongest evidence comes from a cohort study (Score_Type=0.8), with a combined p-value of 0.005 (Score_Sig=0.995), supported by three studies in the same direction (assuming Score_Cons≈0.86). S(A→B)=0.5×0.8+0.3×0.995+0.2×0.86=0.4+0.2985+0.172=0.8705.
[0141] Directed edge B→C: The strongest evidence comes from an RCT (Score_Type=1.0), with a p-value of 0.01 (Score_Sig=0.99), supported by 5 studies with highly consistent directions (assuming Score_Cons≈0.95). S(B→C)=0.5×1.0+0.3×0.99+0.2×0.95=0.5+0.297+0.19=0.987.
[0142] For the directed edge C→A: the evidence comes from only one cross-sectional study (Score_Type=0.3), with a p-value of 0.04 (Score_Sig=0.96), and only this one study (Score_Cons≈0.63). S(C→A) = 0.5×0.3 + 0.3×0.96 + 0.2×0.63 = 0.15 + 0.288 + 0.126 = 0.564. Comparing the three scores, S(C→A) is the lowest.
[0143] Therefore, the conflict resolution engine can automatically perform preset operations, such as deleting the directed edge C→A, thereby breaking the cycle and obtaining the final directed acyclic graph.
[0144] S213. Among the multiple variable nodes in a directed acyclic graph, identify at least one of the following: mixed variable node, intermediate variable node, and colliding variable node.
[0145] In the embodiments of this application, this step involves functionally classifying each node (covariate) on the already constructed directed acyclic graph representing the causal relationship structure between variables, based on the different roles each node plays in the causal path from the exposed variable to the outcome variable. Its core significance lies in transforming a static causal structure graph into a dynamic analytical tool that can guide statistical analysis and bias control.
[0146] Hybrid variable node: In a directed acyclic graph, a variable node is both a cause of the exposing variable (or shares a common cause with the exposing variable) and a cause of the outcome variable. For example... Figure 3As shown, it manifests as a directed path from the node to the exposure variable node and a directed path from the node to the outcome variable node. The confounding variable node is a key component of the "backdoor path," and if left uncontrolled, it can distort the estimation of the true causal effect between exposure and outcome.
[0147] Mediator node: A variable node located between an exposed variable and an outcome variable in a causal path. For example... Figure 4 As shown, it manifests as a directed path (exposure → mediation → outcome) from the exposure variable node to this node and from this node to the outcome variable node. The mediation variable node reveals the mechanism by which exposure affects the outcome.
[0148] Collider node: such as Figure 5 As shown, a colliding variable node is a directed edge on a path that simultaneously receives edges from two or more other nodes. When this path connects the exposure variable and the outcome variable, and the node represents the combined effect of the exposure (or its ancestor) and the outcome (or its ancestor) (exposure → collision ← outcome), it becomes a critical colliding variable node. Incorrectly adjusting colliding variable nodes in statistical analysis can artificially introduce bias (called collision bias or selection bias).
[0149] S214. Generate a minimal set of adjustment variables for causal inference studies of medical observational data, based on at least one of the confounding variable nodes, mediating variable nodes, and colliding variable nodes.
[0150] In the embodiments of this application, the minimum adjustment set refers to the minimum set of covariates that need to be adjusted in a statistical model (such as a regression model) when performing statistical analysis to estimate the total causal effect between the exposure variable and the outcome variable. The purpose of generating this set is to eliminate confounding bias while avoiding the introduction of bias or loss of precision due to unnecessary adjustments (such as adjusting mediating or colliding variables). This set provides direct and actionable guidance for researchers conducting causal inference studies using medical observational data on how to properly construct their statistical analysis models.
[0151] In some embodiments of this application, the process of generating a minimum set of adjustment variables is described. Specifically, the directed acyclic graph (DAG) construction system for causal inference of medical observational data first invokes the list of confounding variable nodes identified in the previous step. Based on the "backdoor criterion," a valid adjustment set must be able to "block" all backdoor paths connecting exposed and outcome variables. However, there may be multiple different combinations of variables that satisfy this condition. To find the "minimum" set, the DAG construction system for causal inference of medical observational data can implement a graph theory-based optimization algorithm, for example, searching for a "minimum vertex cover" that blocks all backdoor paths. The DAG construction system for causal inference of medical observational data computes all minimum sets of adjustment variables that satisfy the condition (sometimes there may be more than one). Subsequently, the DAG construction system for causal inference of medical observational data further checks whether these candidate sets contain any variables identified as mediating or colliding variable nodes. A candidate set is valid if it contains colliding variable nodes, and the colliding variable itself or its descendants are not adjusted; conversely, if colliding variable nodes that should not be adjusted are adjusted, the candidate set is flagged as potentially introducing bias. For mediating variable nodes, the directed acyclic graph (DAG) construction system for causal inference using medical observational data explicitly informs users that adjusting them will cause the analyzed effect to become a direct effect rather than a total effect. Finally, the DAG construction system for causal inference using medical observational data outputs one or more recommended minimal sets of adjustment variables, along with clear explanations and warnings. This approach not only answers "what should be adjusted," but also answers "why adjust in this way" and "what are the consequences of different adjustment strategies" by excluding inappropriate variables and providing detailed explanations. It offers researchers in causal inference studies using medical observational data in-depth statistical strategy guidance far beyond simple variable lists, significantly improving the reliability and validity of research conclusions.
[0152] In some embodiments of this application, after generating the minimum set of adjustment variables for causal inference studies using medical observational data, the following procedures may be further performed: Identification and parameterization of unmeasured confounding variables (U): All theoretically possible but missing confounding variable nodes are identified on the directed acyclic graph and denoted by node "U". Then, the user is prompted, or based on literature knowledge, to set reasonable parameter ranges for the association strength (γ) between "U" and the exposure variable (E) and the association strength (δ) between "U" and the outcome variable (O). For example, the values of γ and δ can be set to ratios between [1.2, 3.0].
[0153] Proxy variable identification and propensity score modeling: Search for "proxy variables" (i.e., other measured variables associated with U, such as a downstream biomarker P of U) for "U" on a directed acyclic graph. Construct an extended propensity score model that includes not only all measured confounding variables but also these identified proxy variables P.
[0154] E-value calibration and weighted variable selection: Simulations are performed iteratively within a defined (γ, δ) parameter grid. For each (γ, δ) combination, the weights of covariates in the propensity score model are back-calibrated using E-value theory (E-value is an indicator of how strong unmeasured confounding variables need to be to invalidate an observed association). Specifically, measured covariates more strongly associated with U (represented by proxy variables P) are given higher weights in the calibrated propensity score model. A weighted set of adjustment variables is generated for each (γ, δ) combination, where the importance of variables (whether they are included and their weights) is dynamically adjusted.
[0155] Sensitivity Curve Generation and Strategy Recommendation: After completing all simulations, a series of sensitivity analysis plots can be generated. The horizontal and vertical axes of the plot represent the effect strengths γ and δ of the unmeasured confounding variables, respectively, while the contour lines or colors indicate whether the estimated exposure-outcome effect strength (or its confidence interval) crosses null values at that intensity. Thus, the final output of the directed acyclic graph construction system for causal inference of medical observational data is not just a minimal set of adjustment variables, but a stratified adjustment strategy report, including: Core adjustment set: Variables that remain stable and important across all simulation scenarios.
[0156] Sensitivity adjustment set: Variables that need to be included only if the assumption that the unmeasured confounding variables are weak.
[0157] Robustness boundary: This clearly states that when the intensity of the unmeasured confounding variable effects (γ and δ) exceeds a certain threshold, the research conclusions will no longer be robust, even with optimal adjustments.
[0158] It can be seen that this approach goes beyond the idea of "finding the only correct adjustment set," acknowledging and addressing the problem of "unmeasured confounding variables" that is common in real-world research. By combining innovative propensity score calibration and E-value sensitivity analysis, it provides researchers with a dynamic, quantitative, and forward-looking adjustment strategy for unknown risks, thereby improving the real-world reliability of research conclusions.
[0159] The directed acyclic graph construction apparatus for causal inference of medical observational data in this application embodiment is described below from the perspective of hardware processing. Please refer to [link to relevant documentation]. Figure 6This is a schematic diagram of at least a portion of the physical structure of the directed acyclic graph construction device for causal inference of medical observational data in the embodiments of this application.
[0160] It should be noted that, Figure 6 The illustrated structure of the directed acyclic graph (DAG) construction apparatus for causal inference from medical observational data is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application. The DAG construction apparatus for causal inference from medical observational data may be integrated with a DAG construction system for causal inference from medical observational data, or integrated into a DAG construction system for causal inference from medical observational data.
[0161] like Figure 6 As shown, the directed acyclic graph construction apparatus for causal inference of medical observational data includes a CPU 601, which can perform various appropriate actions and processes according to a program stored in ROM 602 or a program loaded from storage portion 608 into RAM 603, such as performing the methods described in the above embodiments. RAM 603 also stores various programs and data required for apparatus operation. The CPU 601, ROM 602, and RAM 603 are interconnected via bus 604. I / O interface 605 is also connected to bus 604.
[0162] The following components are connected to I / O interface 605: input section 606 including audio input devices, push-button switches, etc.; output section 607 including liquid crystal display (LCD) and audio output devices, indicator lights, etc.; storage section 608 including hard disks, etc.; and communication section 609 including network interface cards such as LAN (Local Area Network) cards, modems, etc. Communication section 609 performs communication processing via a network such as the Internet. Drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 610 as needed so that computer programs read from them can be installed into storage section 608 as needed.
[0163] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by CPU 601, it performs the various functions defined in this application.
[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those shown in the drawings.
[0165] Specifically, the directed acyclic graph construction apparatus for causal inference of medical observational data in this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, it implements the directed acyclic graph construction method for causal inference of medical observational data provided in the above embodiment.
[0166] In another aspect, this application also provides a computer-readable storage medium, which may be included in the directed acyclic graph construction apparatus for causal inference of medical observational data described in the above embodiments; or it may exist independently and not assembled into the directed acyclic graph construction apparatus for causal inference of medical observational data. The storage medium carries one or more computer programs that, when executed by a processor of the directed acyclic graph construction apparatus for causal inference of medical observational data, cause the apparatus to implement the directed acyclic graph construction method for causal inference of medical observational data provided in the above embodiments.
[0167] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for constructing a directed acyclic graph (DAG) for causal inference from medical observational data, wherein the DAG is generated based on the PECO framework, characterized in that, The method for constructing a directed acyclic graph for causal inference from medical observational data includes: Obtain the exposed and end variables under the PECO framework; The set of literature including exposure variables and outcome variables was retrieved from the pre-defined medical literature database. Within this literature collection, the respective sets of causal and consequential variables for exposure and outcome variables are determined; Within this set of variables, causal and consequence variables that simultaneously belong to both exposure and outcome variables are identified. Based on the effect directions in the corresponding literature, common causes, common outcomes, and mediating variables are generated, forming an initial directed acyclic graph that includes both exposure and outcome variables.
2. The method for constructing a directed acyclic graph for causal inference of medical observational data as described in claim 1, characterized in that, After retrieving a set of literature including exposure variables and outcome variables from a pre-defined medical literature database, the process further includes: Determine the document types in the document collection; For the document collection, delete documents whose document type is the target type; The target types include at least one of the following: literature without empirical evidence, literature that does not record an effect relationship with the exposure variable or outcome variable, and literature in which the study population is not associated with the population variables under the PECO framework.
3. The method for constructing a directed acyclic graph for causal inference of medical observational data as described in claim 1, characterized in that, The formation of the initial directed acyclic graph, which includes exposure variables and outcome variables, includes: Generate variable nodes corresponding to common causes, common results, mediator variables, exposure variables, and outcome variables; Based on the effect direction in the corresponding literature, directed edges are generated for the corresponding variable nodes; Generate an initial directed acyclic graph based on variable nodes and directed edges.
4. The method for constructing a directed acyclic graph for causal inference of medical observational data as described in claim 3, characterized in that, After generating the initial directed acyclic graph based on the variable nodes and directed edges, the process also includes: Based on the causal criteria, the initial directed acyclic graph is adjusted to obtain the adjusted directed graph. The causal criteria include at least the temporal order, biological rationality, consistency and effect strength, and dose-response relationship between different variable nodes. Generate a directed acyclic graph based on the adjusted directed graph.
5. The method for constructing a directed acyclic graph for causal inference of medical observational data as described in claim 4, characterized in that, The step of generating a directed acyclic graph based on the adjusted directed graph includes: Identify cyclic paths in the adjusted directed graph; Adjusting the cyclic paths based on the time sequence of variables in the actual data yields a directed acyclic graph.
6. The method for constructing a directed acyclic graph for causal inference of medical observational data as described in claim 5, characterized in that, After adjusting the cyclic paths according to the time sequence of variables in the actual data to obtain the directed acyclic graph, the process also includes: In a directed acyclic graph, identify at least one of the following: mixed variable node, intermediate variable node, and colliding variable node; Generate a minimal set of adjustment variables for causal inference studies of medical observational data, based on at least one of the confounding variable nodes, mediating variable nodes, and colliding variable nodes.