An interpretable prediction method for clinical cardiovascular risk based on causal inference
By constructing a structural causal model and combining it with cross-center transfer learning, the problem of insufficient causal reasoning in existing cardiovascular risk prediction methods is solved. Mechanism-level path decomposition and individualized interpretable prediction are achieved, improving the robustness and cross-center consistency of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI HANXIN INTELLIGENT PROTECTION TECHNOLOGY CO LTD
- Filing Date
- 2026-05-26
- Publication Date
- 2026-07-24
AI Technical Summary
Existing clinical cardiovascular risk prediction methods lack causal reasoning mechanisms and cannot reveal the path mechanisms and mediating relationships of disease occurrence. This results in poor model transferability, insufficient generalization ability, and insufficient interpretability, making it difficult to form cross-center consistent causal estimates and population-individual level prediction loops.
By integrating causal inference and cross-center transfer learning, a structural causal model is constructed to establish an identifiable formula for intervention-type mediation effects. Combining dual robust estimation and target maximum likelihood, robust and interpretable output middleware is generated, and cross-center consistent estimation is achieved through cross-domain transferable do-calculus.
It achieves mechanism-level pathway decomposition and individualized interpretable prediction of cardiovascular risk, and has strong robustness, high transferability and excellent clinical interpretability. It can obtain reliable prediction results even in the presence of exposure-induced confounding and time-varying covariates.
Smart Images

Figure CN122455360A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cardiovascular risk prediction, and more particularly to an interpretable prediction method for clinical cardiovascular risk based on causal inference. Background Technology
[0002] Existing clinical cardiovascular risk prediction methods mainly rely on statistical regression models or machine learning algorithms. These methods extract features and identify patterns from variables such as blood lipids, blood pressure, blood glucose, and past medical history to establish risk scoring models or probabilistic prediction models. While these methods can capture correlations among multiple variables to some extent, they often remain at the correlation level, lacking causal reasoning mechanisms and failing to reveal the pathways and mediating relationships in disease development. Furthermore, differences in data distribution and collection biases between different medical centers lead to poor model transferability and insufficient generalization ability, limiting the stability and comparability of prediction results.
[0003] While research on causal inference methods has proposed modeling and estimation approaches based on causal graphs, existing solutions are mostly limited to single-center, static scenarios. They can only identify static causal relationships between variables and cannot obtain identifiable interventional mediating effects in situations with exposure-induced confounding factors and time-varying covariates. Furthermore, existing methods lack path-level effect decomposition and individual-level counterfactual explanation mechanisms, resulting in model outputs that are difficult to correlate with clinical mechanisms, insufficient interpretability, and an inability to form cross-center consistent causal estimates and population-individual level predictive loops.
[0004] Therefore, how to provide an interpretable predictive method for clinical cardiovascular risk based on causal inference is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose an interpretable prediction method for clinical cardiovascular risk based on causal inference. This invention integrates causal inference and cross-center transfer learning to achieve mechanism-level path decomposition and individualized interpretable prediction of cardiovascular risk, and has the advantages of strong robustness, high transferability and excellent clinical interpretability.
[0006] A method for predicting clinical cardiovascular risk based on causal inference according to an embodiment of the present invention includes the following steps:
[0007] Collect and standardize multimodal clinical data, construct a structural causal model, and obtain a causal diagram and data baseline;
[0008] On the causal graph, do-calculus and random substitution distribution are used to establish identifiable recipes for intervention-type mediating effects, thus obtaining a set of mechanism pathways;
[0009] By embedding the set of identifiable formulas and mechanism pathways into a causal intervention-type mediation learning algorithm, an exposure model, a mediation model, an outcome model and a time-varying covariate model are established. Robust estimation results and interpretable output middleware are obtained by combining dual robust estimation and target maximum likelihood.
[0010] By introducing cross-domain transferable do-calculus, establishing site selection nodes and performing migration correction, a cross-center consistent estimation set is obtained.
[0011] Based on the cross-center consistent estimation set, the population-level path decomposition results and individual-level counterfactual contribution map are generated, and the causal contributions and interpretable prediction results via the atherosclerotic path and non-path are output.
[0012] Optionally, the generation of the causal graph and data baseline specifically includes:
[0013] Collect multi-source data from electronic medical records, laboratory tests, medical images, follow-up records, and medication records; unify patient identification, timestamps, and examination codes; complete time alignment, spatial registration, and entity matching; and generate a multimodal clinical data set that has been identified, mapped, and standardized by location.
[0014] Data quality control and anomaly detection are performed on the multimodal clinical dataset. Outliers are identified based on the median and absolute deviation of similar observations and a quality log is output. Conditional imputation is performed on missing data, and the imputation source and confidence marker are recorded to form a standardized dataset.
[0015] The standardized dataset is scaled and normalized to map numerical features to a uniform range of values. Distribution alignment is performed within and between patients to generate a normalized feature matrix.
[0016] Semantic hierarchical analysis and feature summarization were performed based on the lipid set, atherosclerosis phenotype set, outcome event set, baseline confounding set, and time-varying covariate set to obtain a hierarchical feature view and a traceable feature dictionary;
[0017] Based on the hierarchical feature view, conditional independence test and clinical prior are performed to generate causal candidate relationship pairs and candidate edge sets, and output a list of structural constraints. The conditional independence test obtains candidate dependencies and decoupling relationships by calculating conditional mutual information under empirical distribution and making significance judgment. The clinical prior limits prohibited edges, mandatory edges and directional constraints through guideline evidence and expert consensus. The list of structural constraints includes candidate edges, directional requirements and inviolable sets.
[0018] Based on the list of structural constraints and the candidate edge set, a candidate structural causal model is constructed with a hierarchical feature view as input. The model includes a set of causal nodes and a set of directed edges. The set of causal nodes is obtained by mapping variables in the hierarchical feature view and includes exposed variable nodes, mediating variable nodes, outcome variable nodes, baseline confounding nodes and time-varying covariate nodes. A traceable feature dictionary is used as the basis for the correspondence between causal nodes and original features.
[0019] The goodness of fit of each candidate structural causal model is evaluated by the structural scoring function, and penalties are imposed according to the degree of violation of the structural constraint list. The structural score and penalty are combined to obtain the initial draft of the structural causal model, which includes the causal graph structure, conditional dependency parameters and node meaning mapping table.
[0020] Based on the initial draft of the structural causal model, distribution statistics and value solidification are performed on each causal node in the hierarchical feature view to form a snapshot of node-level marginal distribution, covariate dependency and time consistency, and generate a data baseline.
[0021] Optionally, the generation of the mechanism path set specifically includes:
[0022] Based on the causal graph and data baseline, the directed relationships between exposed variable nodes, mediating variable nodes, outcome variable nodes, baseline confounding nodes and time-varying covariate nodes are determined, candidate mechanism path pairs are locked, and a path candidate list containing path identifiers, involved nodes and directional constraints is generated.
[0023] Apply do-calculus to the exposure variable nodes on the causal graph to obtain the target expression of how the distribution of the outcome variable changes when a specified intervention is applied to the exposure. Record this as the outcome distribution after intervention and use it as the starting quantity for recipe simplification and path decomposition.
[0024] We construct a random substitution distribution to characterize the intervention-type mediation effect, and give descriptions of the conditional distribution of the mediating variable node under different exposure settings, as well as descriptions of the induced counterfactual quantity of the outcome. We define the direct intervention effect as the difference in expected outcome between different exposure settings while keeping the mediating variable at the random substitution distribution corresponding to the exposure setting. We define the indirect intervention effect as the difference in expected outcome caused by switching the random substitution distribution of the mediating variable from the distribution corresponding to one exposure setting to the distribution corresponding to another exposure setting while keeping the exposure setting unchanged. We use the target quantities of the two types of effects as the basic quantities for path decomposition.
[0025] Based on the do-calculus and sequential summation ideas, the outcome distribution after intervention is decomposed in the form of conditional expectation of the outcome. The relevant terms are expanded in the conditional form of the distribution of time-varying covariates and mediating variables. The marginal weights of the baseline confounding are incorporated into the overall expression. By traversing all possible values of the relevant variables, a systematic transformation from the intervention target quantity to the observable estimable quantity is completed, forming an identifiable formula text containing conditional expectation, conditional distribution and marginal weights.
[0026] For each candidate mechanism path in the candidate path list, an identifiability determination and recipe generation process is performed to retrieve backdoor pathways and confounding sources on the path. If there are backdoor pathways caused by undetected confounding, the condition set is adjusted and the relevant terms are re-expressed according to the blocking and simplification rules of do-calculus. After blocking and simplification, the paths that can be completely expressed by observations are retained as identifiable paths, and the corresponding computable expression text and required condition set are generated. These are then summarized to form a set of mechanism paths.
[0027] Optionally, the generation of the robust estimation results and interpretable output middleware specifically includes:
[0028] Receive a set of identifiable formulations and mechanism paths, as well as a causal graph and data baseline. Generate a list of target quantities and a set of conditions for each mechanism path to form an estimation task list.
[0029] Based on the estimation task sheet, the exposure model, mediation model, outcome model and time-varying covariate model are trained respectively to obtain the initial fitting set and cross-validation prediction set. The mapping from the model to the causal node is completed according to the traceable feature dictionary.
[0030] Substitute the initial fitted set into the identifiable formula, calculate the interpolation estimate one by one according to the mechanism path, and generate the baseline estimate and path-level intermediate results.
[0031] Based on the dual robust estimation, sample-level predicted values are extracted, the residuals between observed and predicted outcomes are calculated, and inverse probability weights are generated using the exposure propensity component and the mediation condition component. The residuals of each sample are then weighted and corrected to form a sample-level influence function. The average value of the sample-level influence function is used as the correction term for the baseline estimator, resulting in the updated estimator and residual diagnosis results after dual robust estimation.
[0032] Based on the dual robust estimation, the target maximum likelihood update is performed, a fluctuation sub-model is built for the expected outcome conditions and targeted fitting is completed, and the equilibrium conditions of the post-targeting estimator and the influence function are output.
[0033] The post-targeting estimates are backfilled into the identifiable formulation. The final path-level estimates, direct intervention effects and indirect intervention effects are calculated according to the mechanism path. Variance is estimated using the out-of-sample prediction results of the cross-validation prediction set, and interval estimates and confidence labels are generated to form a path-level robust estimation result.
[0034] Based on the path-level robust estimation result set, an individual-level counterfactual prediction function is constructed, generating a score matrix and path weight table for the individual-level counterfactual contribution map. These are then solidified together with the mechanism path set into an interpretable output middleware, which includes individual identifiers, path weights, and counterfactual scores.
[0035] Optionally, the generation of the cross-center consistent estimation set specifically includes:
[0036] Based on the causal graph and data baseline, a site selection node is established for each data source center. A selection relationship representation including site identifier, sample affiliation and collection differences is constructed to form a list of site selection nodes and a description of selection relationships.
[0037] Under the constraint of the site selection node list, the transferable components and the components that need to be corrected are identified, and the site dependencies in the exposure model, mediation model, outcome model and time-varying covariate model are labeled respectively, and cross-domain decomposition scheme and parameter partitioning table are generated.
[0038] We apply cross-domain transferable do-calculus to derive the expression of the intervention target quantity in the target center, and weight the distribution of the post-intervention outcome under each site condition with the probability of the site appearing in the target center to obtain the decomposition formula of the intervention expression in the target center, and give a list of components that need to be jointly estimated by the source center and the target center.
[0039] Based on the cross-domain decomposition scheme, the components to be corrected are reweighted and marginally replaced, and the site importance weight and covariate marginal replacement weight are constructed. The site importance weight is defined as the ratio of the probability of a site appearing in the target center under a given baseline mixed value to the probability of the corresponding site in the source center. The covariate marginal replacement weight is defined as the ratio of the baseline mixed marginal distribution of the target center to the marginal distribution of the corresponding site in the source center. The site weighted sample and the target center marginal replacement table are generated.
[0040] The site-weighted samples and the target center marginal replacement table are injected into the path-level robust estimation results and the identifiable formula. The target center estimate of the intervention expression is recalculated according to the mechanism path to obtain the path-level intermediate results after site correction.
[0041] Consistent integration is performed on the intermediate path-level results after site correction. Site weight normalization and variance stabilization are adopted. The path-level estimates of each site after correction are weighted and aggregated according to the site integration weight to generate a cross-center consistent estimate set.
[0042] Optionally, the generation of the causal contribution and the explainable prediction results specifically includes:
[0043] Receive cross-center consistent estimation set and interpretable output middleware, read mechanism path set and individual-level counterfactual prediction function, determine target population scope and path list, and generate path estimation table and individual interface list;
[0044] Calculate the total effect and the effect of each path on the path estimation table, calculate the population stratum path contribution share by the ratio of each path effect to the total effect of the population stratum, form the population stratum path decomposition baseline, synthesize uncertainty labeling and diagnostic statistics, and output the initial draft of the population stratum decomposition.
[0045] Call the individual-level counterfactual prediction function one by one according to the individual interface list, calculate the expected outcome in the observation scenario, direct intervention scenario and indirect intervention scenario respectively, obtain the individual-level effect vector and variance estimate, and generate the individual-level effect cache;
[0046] Using the cross-center consistent estimation set and the initial draft of population stratification as the consistency target, path masking consistency calibration is performed on the individual stratification effect cache, outputting the calibration matrix and generating the calibrated individual stratification effects and contribution ratios;
[0047] A score matrix and path weight table are generated based on the calibrated individual-level effects and contribution ratios, forming an individual-level counterfactual contribution map data package, and an uncertainty label and extrapolation risk label are added to each individual.
[0048] Using the initial draft of the population stratification as input, the system performs path sorting, threshold pruning, and merging display, outputting the causal contribution, contribution share, and uncertainty label of paths via atherosclerosis and those not via atherosclerosis, thus generating the population stratification path decomposition results.
[0049] The path decomposition results at the population level and the counterfactual contribution map data at the individual level are written together into the interpretable output middleware to form interpretable prediction results.
[0050] The beneficial effects of this invention are:
[0051] This invention establishes identifiable formulations of interventional mediating effects by introducing do-calculus and random substitution distributions into a causal graph, achieving path-level causal decomposition of key mechanisms such as atherosclerosis. This elevates cardiovascular risk prediction from correlation analysis to a causal inference process with interpretable mechanisms. By combining dual-robust estimation with target maximum likelihood joint optimization, this invention achieves statistically robust path-level estimation results even in the presence of exposure-induced confounding and time-varying covariates. Furthermore, the use of cross-domain transferable do-calculus and formal transfer correction methods for site selection nodes effectively addresses estimation bias caused by differences in multi-center data distribution, resulting in consistent cross-center assessments of interventional effects.
[0052] At the level of result interpretation, this invention constructs an interpretable output middleware that integrates population-level path decomposition results with individual-level counterfactual contribution maps to form mechanism-level, individualized interpretable predictive outputs. This reveals the contribution share of each mechanism pathway in overall risk formation and generates counterfactual risk differences for individual patients, providing clinicians with traceable and verifiable decision-making basis. This method ensures statistical consistency while also considering clinical understandability and transferability, significantly improving the interpretability and reliability of cardiovascular risk prediction. Attached Figure Description
[0053] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0054] Figure 1 This is a flowchart of a clinical cardiovascular risk interpretable prediction method based on causal inference proposed in this invention;
[0055] Figure 2 This is a schematic diagram of a causal intervention-type mediator learning algorithm for a clinical cardiovascular risk interpretable prediction method based on causal inference proposed in this invention. Detailed Implementation
[0056] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0057] refer to Figures 1-2 A predictive method for clinical cardiovascular risk based on causal inference, comprising the following steps:
[0058] Collect and standardize multimodal clinical data, construct a structural causal model, and obtain a causal diagram and data baseline;
[0059] On the causal graph, do-calculus and random substitution distribution are used to establish identifiable recipes for intervention-type mediating effects, thus obtaining a set of mechanism pathways;
[0060] By embedding the set of identifiable formulas and mechanism pathways into a causal intervention-type mediation learning algorithm, an exposure model, a mediation model, an outcome model and a time-varying covariate model are established. Robust estimation results and interpretable output middleware are obtained by combining dual robust estimation and target maximum likelihood.
[0061] By introducing cross-domain transferable do-calculus, establishing site selection nodes and performing migration correction, a cross-center consistent estimation set is obtained.
[0062] Based on the cross-center consistent estimation set, the population-level path decomposition results and individual-level counterfactual contribution map are generated, and the causal contributions and interpretable prediction results via the atherosclerotic path and non-path are output.
[0063] In this embodiment, the generation of the causal graph and data baseline specifically includes:
[0064] Collect multi-source data from electronic medical records, laboratory tests, medical images, follow-up records, and medication records; unify patient identification, timestamps, and examination codes; complete time alignment, spatial registration, and entity matching; and generate a multimodal clinical data set that has been identified, mapped, and standardized by location.
[0065] Data quality control and anomaly detection are performed on the multimodal clinical dataset. Outliers are identified based on the median and absolute deviation of similar observations and a quality log is output. Conditional imputation is performed on missing data, and the imputation source and confidence marker are recorded to form a standardized dataset.
[0066] The standardized dataset is scaled and normalized to map numerical features to a uniform range of values. Distribution alignment is performed within and between patients to generate a normalized feature matrix.
[0067] Semantic hierarchical analysis and feature summarization were performed based on the lipid set, atherosclerosis phenotype set, outcome event set, baseline confounding set, and time-varying covariate set to obtain a hierarchical feature view and a traceable feature dictionary;
[0068] The lipid set includes routine blood lipids and lipoprotein-related indicators; the atherosclerosis phenotype set includes quantitative indicators of thickness, plaque, and calcification derived from ultrasound and computed tomography; the outcome event set includes the timing and type of cardiovascular outcomes; the baseline confounding set includes demographic and past medical history elements; and the time-varying covariate set includes dynamic elements that arise after exposure and influence mediating and outcome events.
[0069] Based on the hierarchical feature view, conditional independence test and clinical prior are performed to generate causal candidate relationship pairs and candidate edge sets, and output a list of structural constraints. The conditional independence test obtains candidate dependencies and decoupling relationships by calculating conditional mutual information under empirical distribution and making significance judgment. The clinical prior limits prohibited edges, mandatory edges and directional constraints through guideline evidence and expert consensus. The list of structural constraints includes candidate edges, directional requirements and inviolable sets.
[0070] Based on the list of structural constraints and the candidate edge set, a candidate structural causal model is constructed with a hierarchical feature view as input. The model includes a set of causal nodes and a set of directed edges. The set of causal nodes is obtained by mapping variables in the hierarchical feature view and includes exposed variable nodes, mediating variable nodes, outcome variable nodes, baseline confounding nodes and time-varying covariate nodes. A traceable feature dictionary is used as the basis for the correspondence between causal nodes and original features.
[0071] The goodness of fit of each candidate structural causal model is evaluated by the structural scoring function, and penalties are imposed according to the degree of violation of the structural constraint list. The structural score and penalty are combined to obtain the initial draft of the structural causal model, which includes the causal graph structure, conditional dependency parameters and node meaning mapping table.
[0072] Based on the initial draft of the structural causal model, distribution statistics and value solidification are performed on each causal node in the hierarchical feature view to form a snapshot of node-level marginal distribution, covariate dependency and time consistency, and generate a data baseline.
[0073] In this embodiment, the generation of the mechanism path set specifically includes:
[0074] Based on the causal graph and data baseline, the directed relationships between exposed variable nodes, mediating variable nodes, outcome variable nodes, baseline confounding nodes and time-varying covariate nodes are determined, candidate mechanism path pairs are locked, and a path candidate list containing path identifiers, involved nodes and directional constraints is generated.
[0075] Apply do-calculus to the exposure variable nodes on the causal graph to obtain the target expression of how the distribution of the outcome variable changes when a specified intervention is applied to the exposure. Record this as the outcome distribution after intervention and use it as the starting quantity for recipe simplification and path decomposition.
[0076] We construct a random substitution distribution to characterize the intervention-type mediation effect, and give descriptions of the conditional distribution of the mediating variable node under different exposure settings, as well as descriptions of the induced counterfactual quantity of the outcome. We define the direct intervention effect as the difference in expected outcome between different exposure settings while keeping the mediating variable at the random substitution distribution corresponding to the exposure setting. We define the indirect intervention effect as the difference in expected outcome caused by switching the random substitution distribution of the mediating variable from the distribution corresponding to one exposure setting to the distribution corresponding to another exposure setting while keeping the exposure setting unchanged. We use the target quantities of the two types of effects as the basic quantities for path decomposition.
[0077] Based on the do-calculus and sequential summation ideas, the outcome distribution after intervention is decomposed in the form of conditional expectation of the outcome. The relevant terms are expanded in the conditional form of the distribution of time-varying covariates and mediating variables. The marginal weights of the baseline confounding are incorporated into the overall expression. By traversing all possible values of the relevant variables, a systematic transformation from the intervention target quantity to the observable estimable quantity is completed, forming an identifiable formula text containing conditional expectation, conditional distribution and marginal weights.
[0078] For each candidate mechanism path in the candidate path list, an identifiability determination and recipe generation process is performed to retrieve backdoor pathways and confounding sources on the path. If there are backdoor pathways caused by undetected confounding, the condition set is adjusted and the relevant terms are re-expressed according to the blocking and simplification rules of do-calculus. After blocking and simplification, the paths that can be completely expressed by observations are retained as identifiable paths, and the corresponding computable expression text and required condition set are generated. These are then summarized to form a set of mechanism paths.
[0079] In this embodiment, the generation of the robust estimation results and interpretable output middleware specifically includes:
[0080] Receive a set of identifiable formulations and mechanism paths, as well as a causal graph and data baseline. Generate a list of target quantities and a set of conditions for each mechanism path to form an estimation task list.
[0081] Based on the estimation task sheet, the exposure model, mediation model, outcome model and time-varying covariate model are trained respectively to obtain the initial fitting set and cross-validation prediction set. The mapping from the model to the causal node is completed according to the traceable feature dictionary.
[0082] Substitute the initial fitted set into the identifiable formula, calculate the interpolation estimate one by one according to the mechanism path, and generate the baseline estimate and path-level intermediate results.
[0083] Based on the dual robust estimation, sample-level predicted values are extracted, the residuals between observed and predicted outcomes are calculated, and inverse probability weights are generated using the exposure propensity component and the mediation condition component. The residuals of each sample are then weighted and corrected to form a sample-level influence function. The average value of the sample-level influence function is used as the correction term for the baseline estimator, resulting in the updated estimator and residual diagnosis results after dual robust estimation.
[0084] Based on the dual robust estimation, the target maximum likelihood update is performed, a fluctuation sub-model is built for the expected outcome conditions and targeted fitting is completed, and the equilibrium conditions of the post-targeting estimator and the influence function are output.
[0085] The volatile sub-model is constructed by equating the expected outcome after log-odds transformation to the current log-odds plus the product of the target update coefficient and the score covariate. The target update coefficient is obtained by maximizing the likelihood.
[0086] The post-targeting estimates are backfilled into the identifiable formulation. The final path-level estimates, direct intervention effects and indirect intervention effects are calculated according to the mechanism path. Variance is estimated using the out-of-sample prediction results of the cross-validation prediction set, and interval estimates and confidence labels are generated to form a path-level robust estimation result.
[0087] Based on the path-level robust estimation result set, an individual-level counterfactual prediction function is constructed, generating a score matrix and path weight table for the individual-level counterfactual contribution map, and solidified together with the mechanism path set into an interpretable output middleware, including individual identifiers, path weights and counterfactual scores;
[0088] Under the constraints of the causal graph and data baseline, the individual-level counterfactual prediction function fixes the baseline confounding and time-varying covariate characteristics of each individual. It uses the targeted exposure model, mediation model, and outcome model to generate outcome expectations under the observation scenario, direct intervention scenario, and indirect intervention scenario, respectively. The individual-level direct intervention effect and indirect intervention effect are obtained by calculating the difference in outcome expectations between scenarios. Path masking operation is performed under each mechanism path, allowing only the information flow of that path to change, while keeping the observation distribution unchanged for other paths, to obtain the individual-level effect decomposed by path. With the path-level robust estimation result set as the overall constraint, the individual-level effect is normalized and calibrated within the path and within the individual, generating the individual-level score matrix and path weight table.
[0089] In this embodiment, the generation of the cross-center consistent estimation set specifically includes:
[0090] Based on the causal graph and data baseline, a site selection node is established for each data source center. A selection relationship representation including site identifier, sample affiliation and collection differences is constructed to form a list of site selection nodes and a description of selection relationships.
[0091] Under the constraint of the site selection node list, the transferable components and the components that need to be corrected are identified, and the site dependencies in the exposure model, mediation model, outcome model and time-varying covariate model are labeled respectively, and cross-domain decomposition scheme and parameter partitioning table are generated.
[0092] We apply cross-domain transferable do-calculus to derive the expression of the intervention target quantity in the target center, and weight the distribution of the post-intervention outcome under each site condition with the probability of the site appearing in the target center to obtain the decomposition formula of the intervention expression in the target center, and give a list of components that need to be jointly estimated by the source center and the target center.
[0093] Based on the cross-domain decomposition scheme, the components to be corrected are reweighted and marginally replaced, and the site importance weight and covariate marginal replacement weight are constructed. The site importance weight is defined as the ratio of the probability of a site appearing in the target center under a given baseline mixed value to the probability of the corresponding site in the source center. The covariate marginal replacement weight is defined as the ratio of the baseline mixed marginal distribution of the target center to the marginal distribution of the corresponding site in the source center. The site weighted sample and the target center marginal replacement table are generated.
[0094] The site-weighted samples and the target center marginal replacement table are injected into the path-level robust estimation results and the identifiable formula. The target center estimate of the intervention expression is recalculated according to the mechanism path to obtain the path-level intermediate results after site correction.
[0095] Consistent integration is performed on the intermediate path-level results after site correction. Site weight normalization and variance stabilization are adopted. The path-level estimates of each site after correction are weighted and aggregated according to the site integration weight to generate a cross-center consistent estimate set.
[0096] In this embodiment, the generation of the causal contribution and the explainable prediction results specifically includes:
[0097] Receive cross-center consistent estimation set and interpretable output middleware, read mechanism path set and individual-level counterfactual prediction function, determine target population scope and path list, and generate path estimation table and individual interface list;
[0098] Calculate the total effect and the effect of each path on the path estimation table, calculate the population stratum path contribution share by the ratio of each path effect to the total effect of the population stratum, form the population stratum path decomposition baseline, synthesize uncertainty labeling and diagnostic statistics, and output the initial draft of the population stratum decomposition.
[0099] Call the individual-level counterfactual prediction function one by one according to the individual interface list, calculate the expected outcome in the observation scenario, direct intervention scenario and indirect intervention scenario respectively, obtain the individual-level effect vector and variance estimate, and generate the individual-level effect cache;
[0100] Using the cross-center consistent estimation set and the initial draft of population stratification as the consistency target, path masking consistency calibration is performed on the individual stratification effect cache, outputting the calibration matrix and generating the calibrated individual stratification effects and contribution ratios;
[0101] Among them, the calibrated individual-level effect refers to the effect of each individual on each mechanism path after proportional adjustment, so that the sum of the individual effects of the same path on all individuals is equal to the effect value of that path in the cross-center consistent estimation set and the initial draft of population stratification. The calibrated contribution ratio refers to the ratio of the calibrated individual effect of each individual on each mechanism path to the sum of the calibrated individual effects of that individual on all mechanism paths.
[0102] A score matrix and path weight table are generated based on the calibrated individual-level effects and contribution ratios, forming an individual-level counterfactual contribution map data package, and an uncertainty label and extrapolation risk label are added to each individual.
[0103] Using the initial draft of the population stratification as input, the system performs path sorting, threshold pruning, and merging display, outputting the causal contribution, contribution share, and uncertainty label of paths via atherosclerosis and those not via atherosclerosis, thus generating the population stratification path decomposition results.
[0104] The path decomposition results at the population level and the counterfactual contribution map data at the individual level are written together into the interpretable output middleware to form interpretable prediction results.
[0105] Example 1:
[0106] To verify the feasibility of this invention in practice, it was applied to a real clinical follow-up scenario at a provincial cardiovascular disease prevention and treatment center. This center has been collecting electronic medical records, laboratory reports, ultrasound images, medication records, and outcome follow-up data from patients with coronary heart disease, hyperlipidemia, and atherosclerosis for a long period. However, due to issues such as varying sampling frequencies, inconsistent examination items, missing timelines, and inconsistent collection standards among different hospitals, traditional statistical methods cannot directly establish cross-center causal path models, making it difficult to accurately quantify the association between lipid levels, atherosclerosis progression, and cardiovascular events. This invention aims to address the difficulties in causal identification and the lack of mechanistic explanation caused by this data heterogeneity.
[0107] In practical deployment, the process begins by performing time alignment and spatial registration on multimodal clinical data from previous years. Combined with patient identifier mapping, a unified, standardized feature matrix is obtained. Based on this matrix, a structural causal model is constructed, incorporating lipid indicators, atherosclerotic imaging features, outcome events, and medication factors. Subsequently, do-calculus is used to identify interventional mediating pathways between lipid levels and events via atherosclerosis, forming identifiable formulas that are embedded into a causal interventional mediation learning algorithm. During the learning phase, the algorithm jointly trains the exposure model, mediation model, outcome model, and time-varying covariate model. Double robust estimation corrects for sample bias, and target maximum likelihood is used for targeted updates, resulting in robust and consistent path-level estimation results.
[0108] To address the differences in samples from different hospitals, the algorithm further introduces a cross-domain transferable do-calculus approach. This involves constructing site selection nodes and reweighting the marginal distribution, enabling the estimation results from various centers to be transferred and fused within a unified causal framework. Finally, a counterfactual prediction function at the individual level is used to calculate the expected outcome difference for each patient under direct and indirect intervention scenarios, generating a counterfactual contribution map. Path decomposition is then performed at the population level, outputting the causal contribution shares via the atherosclerosis pathway and those not via that pathway.
[0109] During continuous validation at the center, this invention achieved unified modeling of cross-year and cross-hospital data, enabling physicians to clearly see the mediating effect of lipid levels through atherosclerosis and to distinguish the relative impact of different mechanistic pathways on event risk. Results showed that the model not only outperformed conventional machine learning models in predictive accuracy but also provided quantifiable pathway contributions and individualized counterfactual explanations at the causal explanation level, significantly improving the operability of clinical risk communication and intervention decisions. This embodiment demonstrates that this invention can achieve mechanism-level causal inference and interpretable prediction in complex clinical environments, possessing significant clinical practical value and potential for widespread application.
[0110] Table 1. Performance comparison between causal inference-based predictive methods for explainable clinical cardiovascular risk and traditional methods.
[0111] Data type fusion Single phenotype and blood lipid data Image + test data Phenotypic + Clinical Variables Multimodal fusion Cross-center adaptation none weak middle powerful AUC (Prediction Accuracy) 0.78 0.84 0.82 0.91 Path contribution explanation consistency rate not applicable 0.41 0.55 0.83 Individual-level counterfactual consistency not applicable 0.36 0.47 0.79 Clinical interpretability score 2.1 / 5 3.2 / 5 3.8 / 5 4.6 / 5 Model robustness (external validation) 0.65 0.72 0.69 0.88
[0112] As can be seen from Table 1, the causal intervention-type mediation learning algorithm proposed in this invention significantly outperforms existing methods in many core metrics.
[0113] In terms of prediction accuracy, traditional logistic regression models, which can only handle static linear relationships, have an AUC of only 0.78. While conventional deep learning models can handle nonlinear mappings, they do not introduce causal structure constraints, resulting in significant overfitting during cross-center validation. This invention introduces intervention-type mediation effect modeling under do-calculus constraints to decompose the causal path between blood lipids, atherosclerosis phenotypes, and outcome events, thereby achieving a balance between structural recognizability and prediction accuracy, and improving the AUC to 0.91.
[0114] At the interpretability level, while structural equation modeling can output estimates of indirect effects, it cannot robustly handle exposure-induced mediator-outcome confounding, with a path explanation consistency rate of only 0.55. This invention eliminates estimation bias caused by model specification errors by combining dual robust estimation with target maximum likelihood updates, and generates individual-level effect explanations using counterfactual prediction functions, achieving a path contribution explanation consistency rate of 0.83, significantly improving the consistency between causal inference and actual predictions.
[0115] Regarding cross-center adaptability, traditional models exhibit significant performance differences across different hospital datasets. This invention introduces a cross-domain transferable do-calculus approach, using site weighting and marginal substitution to achieve unified correction of the estimation results, thereby improving the model's robustness index in external validation to 0.88. This means the model can maintain high generalization ability even in heterogeneous medical data environments.
[0116] Furthermore, in the subjective scoring of the model output results, clinical experts believed that the interpretable path diagram and individual-level counterfactual contribution card output by this invention can intuitively present the causal chain of "blood lipids-atherosclerosis-event", transforming risk prediction from a "black box" to "inference through traceable mechanisms". The clinical interpretability score reached 4.6 out of 5, which greatly improved doctors' decision-making trust and risk communication efficiency.
[0117] In summary, this invention achieves multi-dimensional improvements in mechanism identification, cross-domain robustness, and interpretability, fully demonstrating the effectiveness and technological leadership of the causal intervention-based mediator learning algorithm in clinical cardiovascular risk prediction.
[0118] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for predicting clinical cardiovascular risk based on causal inference, characterized in that, Includes the following steps: Collect and standardize multimodal clinical data, construct a structural causal model, and obtain a causal diagram and data baseline; On the causal graph, do-calculus and random substitution distribution are used to establish identifiable recipes for intervention-type mediating effects, thus obtaining a set of mechanism pathways; By embedding the set of identifiable formulas and mechanism pathways into a causal intervention-type mediation learning algorithm, an exposure model, a mediation model, an outcome model and a time-varying covariate model are established. Robust estimation results and interpretable output middleware are obtained by combining dual robust estimation and target maximum likelihood. By introducing cross-domain transferable do-calculus, establishing site selection nodes and performing migration correction, a cross-center consistent estimation set is obtained. Based on the cross-center consistent estimation set, the population-level path decomposition results and individual-level counterfactual contribution map are generated, and the causal contributions and interpretable prediction results via the atherosclerotic path and non-path are output.
2. The clinical cardiovascular risk interpretable prediction method based on causal inference according to claim 1, characterized in that, The generation of the cause-effect graph and data baseline specifically includes: Collect multi-source data from electronic medical records, laboratory tests, medical images, follow-up records, and medication records; unify patient identification, timestamps, and examination codes; complete time alignment, spatial registration, and entity matching; and generate a multimodal clinical data set that has been identified, mapped, and standardized by location. Data quality control and anomaly detection are performed on the multimodal clinical dataset. Outliers are identified based on the median and absolute deviation of similar observations and a quality log is output. Conditional imputation is performed on missing data, and the imputation source and confidence marker are recorded to form a standardized dataset. The standardized dataset is scaled and normalized to map numerical features to a uniform range of values. Distribution alignment is performed within and between patients to generate a normalized feature matrix. Semantic hierarchical analysis and feature summarization were performed based on the lipid set, atherosclerosis phenotype set, outcome event set, baseline confounding set, and time-varying covariate set to obtain a hierarchical feature view and a traceable feature dictionary; Based on the hierarchical feature view, conditional independence test and clinical prior are performed to generate causal candidate relationship pairs and candidate edge sets, and output a list of structural constraints. The conditional independence test obtains candidate dependencies and decoupling relationships by calculating conditional mutual information under empirical distribution and making significance judgment. The clinical prior limits prohibited edges, mandatory edges and directional constraints through guideline evidence and expert consensus. The list of structural constraints includes candidate edges, directional requirements and inviolable sets. Based on the list of structural constraints and the candidate edge set, a candidate structural causal model is constructed with a hierarchical feature view as input. The model includes a set of causal nodes and a set of directed edges. The set of causal nodes is obtained by mapping variables in the hierarchical feature view and includes exposed variable nodes, mediating variable nodes, outcome variable nodes, baseline confounding nodes and time-varying covariate nodes. A traceable feature dictionary is used as the basis for the correspondence between causal nodes and original features. The goodness of fit of each candidate structural causal model is evaluated by the structural scoring function, and penalties are imposed according to the degree of violation of the structural constraint list. The structural score and penalty are combined to obtain the initial draft of the structural causal model, which includes the causal graph structure, conditional dependency parameters and node meaning mapping table. Based on the initial draft of the structural causal model, distribution statistics and value solidification are performed on each causal node in the hierarchical feature view to form a snapshot of node-level marginal distribution, covariate dependency and time consistency, and generate a data baseline.
3. The clinical cardiovascular risk interpretable prediction method based on causal inference according to claim 1, characterized in that, The generation of the mechanism path set specifically includes: Based on the causal graph and data baseline, the directed relationships between exposed variable nodes, mediating variable nodes, outcome variable nodes, baseline confounding nodes and time-varying covariate nodes are determined, candidate mechanism path pairs are locked, and a path candidate list containing path identifiers, involved nodes and directional constraints is generated. Apply do-calculus to the exposure variable nodes on the causal graph to obtain the target expression of how the distribution of the outcome variable changes when a specified intervention is applied to the exposure. Record this as the outcome distribution after intervention and use it as the starting quantity for recipe simplification and path decomposition. We construct a random substitution distribution to characterize the intervention-type mediation effect, and give descriptions of the conditional distribution of the mediating variable node under different exposure settings, as well as descriptions of the induced counterfactual quantity of the outcome. We define the direct intervention effect as the difference in expected outcome between different exposure settings while keeping the mediating variable at the random substitution distribution corresponding to the exposure setting. We define the indirect intervention effect as the difference in expected outcome caused by switching the random substitution distribution of the mediating variable from the distribution corresponding to one exposure setting to the distribution corresponding to another exposure setting while keeping the exposure setting unchanged. We use the target quantities of the two types of effects as the basic quantities for path decomposition. Based on the do-calculus and sequential summation ideas, the outcome distribution after intervention is decomposed in the form of conditional expectation of the outcome. The relevant terms are expanded in the conditional form of the distribution of time-varying covariates and mediating variables. The marginal weights of the baseline confounding are incorporated into the overall expression. By traversing all possible values of the relevant variables, a systematic transformation from the intervention target quantity to the observable estimable quantity is completed, forming an identifiable formula text containing conditional expectation, conditional distribution and marginal weights. For each candidate mechanism path in the candidate path list, an identifiability determination and recipe generation process is performed to retrieve backdoor pathways and confounding sources on the path. If there are backdoor pathways caused by undetected confounding, the condition set is adjusted and the relevant terms are re-expressed according to the blocking and simplification rules of do-calculus. After blocking and simplification, the paths that can be completely expressed by observations are retained as identifiable paths, and the corresponding computable expression text and required condition set are generated. These are then summarized to form a set of mechanism paths.
4. The clinical cardiovascular risk interpretable prediction method based on causal inference according to claim 1, characterized in that, The generation of the robust estimation results and interpretable output middleware specifically includes: Receive a set of identifiable formulations and mechanism paths, as well as a causal graph and data baseline. Generate a list of target quantities and a set of conditions for each mechanism path to form an estimation task list. Based on the estimation task sheet, the exposure model, mediation model, outcome model and time-varying covariate model are trained respectively to obtain the initial fitting set and cross-validation prediction set. The mapping from the model to the causal node is completed according to the traceable feature dictionary. Substitute the initial fitted set into the identifiable formula, calculate the interpolation estimate one by one according to the mechanism path, and generate the baseline estimate and path-level intermediate results. Based on the dual robust estimation, sample-level predicted values are extracted, the residuals between observed and predicted outcomes are calculated, and inverse probability weights are generated using the exposure propensity component and the mediation condition component. The residuals of each sample are then weighted and corrected to form a sample-level influence function. The average value of the sample-level influence function is used as the correction term for the baseline estimator, resulting in the updated estimator and residual diagnosis results after dual robust estimation. Based on the dual robust estimation, the target maximum likelihood update is performed, a fluctuation sub-model is built for the expected outcome conditions and targeted fitting is completed, and the equilibrium conditions of the post-targeting estimator and the influence function are output. The post-targeting estimates are backfilled into the identifiable formulation. The final path-level estimates, direct intervention effects and indirect intervention effects are calculated according to the mechanism path. Variance is estimated using the out-of-sample prediction results of the cross-validation prediction set, and interval estimates and confidence labels are generated to form a path-level robust estimation result. Based on the path-level robust estimation result set, an individual-level counterfactual prediction function is constructed, generating a score matrix and path weight table for the individual-level counterfactual contribution map. These are then solidified together with the mechanism path set into an interpretable output middleware, which includes individual identifiers, path weights, and counterfactual scores.
5. The clinical cardiovascular risk interpretable prediction method based on causal inference according to claim 1, characterized in that, The generation of the cross-center consistent estimation set specifically includes: Based on the causal graph and data baseline, a site selection node is established for each data source center. A selection relationship representation including site identifier, sample affiliation and collection differences is constructed to form a list of site selection nodes and a description of selection relationships. Under the constraint of the site selection node list, the transferable components and the components that need to be corrected are identified, and the site dependencies in the exposure model, mediation model, outcome model and time-varying covariate model are labeled respectively, and cross-domain decomposition scheme and parameter partitioning table are generated. We apply cross-domain transferable do-calculus to derive the expression of the intervention target quantity in the target center, and weight the distribution of the post-intervention outcome under each site condition with the probability of the site appearing in the target center to obtain the decomposition formula of the intervention expression in the target center, and give a list of components that need to be jointly estimated by the source center and the target center. Based on the cross-domain decomposition scheme, the components to be corrected are reweighted and marginally replaced, and the site importance weight and covariate marginal replacement weight are constructed. The site importance weight is defined as the ratio of the probability of a site appearing in the target center under a given baseline mixed value to the probability of the corresponding site in the source center. The covariate marginal replacement weight is defined as the ratio of the baseline mixed marginal distribution of the target center to the marginal distribution of the corresponding site in the source center. The site weighted sample and the target center marginal replacement table are generated. The site-weighted samples and the target center marginal replacement table are injected into the path-level robust estimation results and the identifiable formula. The target center estimate of the intervention expression is recalculated according to the mechanism path to obtain the path-level intermediate results after site correction. Consistent integration is performed on the intermediate path-level results after site correction. Site weight normalization and variance stabilization are adopted. The path-level estimates of each site after correction are weighted and aggregated according to the site integration weight to generate a cross-center consistent estimate set.
6. The clinical cardiovascular risk interpretable prediction method based on causal inference according to claim 1, characterized in that, The generation of the causal contribution and the explainable prediction results specifically includes: Receive cross-center consistent estimation set and interpretable output middleware, read mechanism path set and individual-level counterfactual prediction function, determine target population scope and path list, and generate path estimation table and individual interface list; Calculate the total effect and the effect of each path on the path estimation table, calculate the population stratum path contribution share by the ratio of each path effect to the total effect of the population stratum, form the population stratum path decomposition baseline, synthesize uncertainty labeling and diagnostic statistics, and output the initial draft of the population stratum decomposition. Call the individual-level counterfactual prediction function one by one according to the individual interface list, calculate the expected outcome in the observation scenario, direct intervention scenario and indirect intervention scenario respectively, obtain the individual-level effect vector and variance estimate, and generate the individual-level effect cache; Using the cross-center consistent estimation set and the initial draft of population stratification as the consistency target, path masking consistency calibration is performed on the individual stratification effect cache, outputting the calibration matrix and generating the calibrated individual stratification effects and contribution ratios; A score matrix and path weight table are generated based on the calibrated individual-level effects and contribution ratios, forming an individual-level counterfactual contribution map data package, and an uncertainty label and extrapolation risk label are added to each individual. Using the initial draft of the population stratification as input, the system performs path sorting, threshold pruning, and merging display, outputting the causal contribution, contribution share, and uncertainty label of paths via atherosclerosis and those not via atherosclerosis, thus generating the population stratification path decomposition results. The path decomposition results at the population level and the counterfactual contribution map data at the individual level are written together into the interpretable output middleware to form interpretable prediction results.