T2DM risk factor differentiable causal discovery method fusing weighting mechanism and multi-granularity search

By constructing a bowless ADMG search space and introducing an improved scoring function with a weighting mechanism and a multi-granularity search strategy, the problems of insufficient causal identification accuracy and network structure integrity in the processing of T2DM hazard factor data by existing methods are solved, and more efficient causal relationship discovery is achieved.

CN121301701APending Publication Date: 2026-01-09LINGNAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511620906.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing methods, when processing T2DM hazard factor data, rely on sparse penalty or simple fitting terms for the scoring function, which fails to adequately suppress confounding interference and small sample noise, resulting in decreased causal identification accuracy and insufficient network structure integrity.

Method used

We construct an ADMG search space without bows, introduce an improved scoring function based on a weighted mechanism and a multi-granularity search strategy. By introducing the Pearson correlation coefficient to improve the scoring function, and combining RRICF and coarse-fine multi-granularity search strategies, we achieve synchronous convergence of the causal structure.

Benefits of technology

It effectively discovers potential confounding causal relationships among T2DM risk factors, improves the accuracy of causal identification and the integrity of network structure, and provides a more accurate T2DM prevention and control solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301701A_ABST
    Figure CN121301701A_ABST
Patent Text Reader

Abstract

The invention provides a T2DM risk factor differentiable causal discovery method fusing a weighting mechanism and multi-granularity search. Preprocessing and standardizing the original data; constructing a linear bow-free ADMG causal structure search space by using a structural equation model (SEM) and bow-free ADMG differentiable algebraic constraint; converting a discrete search problem into a continuous optimization task, and introducing a Pearson's correlation coefficient to improve a scoring function so as to correct mining deviation of potential causal edges; a regularization residual error iteration condition fitting method and a thickness multi-granularity search strategy are fused to avoid falling into local optimum, synchronous convergence of structure learning and strength estimation is achieved through an alternating iteration optimization structure and causal strength parameters, and finally a high-precision and stable ADMG causal structure is obtained. The method provides more reliable and interpretable theoretical support for T2DM causal relationship research, and can provide a new thought for diabetes prevention and treatment and research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical informatics and relates to a method for discovering differentiable causality of risk factors for T2DM by integrating weighted mechanisms and multi-granularity search. Background Technology

[0002] With the continuous rise in the incidence of type 2 diabetes mellitus (T2DM), its complex etiology has attracted increasing attention. T2DM has multiple and strongly coupled risk factors, and the difficulty in long-term patient follow-up leads to fragmented clinical data. These factors severely impact research on the potential causal effects of confounding risk factors in T2DM. Therefore, developing causal discovery methods for T2DM risk factors that can identify confounding causal relationships is of great significance for elucidating the pathogenesis of T2DM and improving prevention and treatment.

[0003] Currently, causal discovery research on T2DM risk factors mainly includes two aspects: binary causal relationship analysis and causal network modeling. The former relies on inference methods such as Mendelian randomization (MR) to provide a theoretical basis for subsequent causal structure modeling; the latter is mostly based on methods such as directed acyclic graphs (DAGs) to analyze the causal relationships between variables. However, when unobserved confounding factors exist, DAGs are difficult to reflect the true causal structure. Against this background, directed acyclic mixture graphs (ADMGs) and their three different types (ancestor type, drought type, and bowless type) have been proposed to characterize causal relationships when unobserved confounding factors exist. In recent years, differentiable causal discovery methods have been widely applied to bowless ADMG structure learning, but when processing T2DM risk factor data, existing methods often rely on sparse penalty or simple fitting terms in the scoring function, failing to adequately suppress confounding interference and small sample noise, easily getting trapped in local optima, leading to decreased causal identification accuracy and insufficient network structure integrity. Therefore, this invention proposes a differentiable causal discovery method for T2DM risk factors by constructing an ADMG search space without bows, introducing an improved scoring function based on a weighted mechanism, and embedding an improved optimization algorithm with a multi-granularity search strategy, thus providing a new approach for the prevention and treatment of diabetes. Summary of the Invention

[0004] To address the problems existing in the background technology, this invention proposes a differentiable causal discovery method for T2DM risk factors that integrates a weighted mechanism and multi-granularity search.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A method for discovering differentiable causality of risk factors in T2DM that integrates weighted mechanisms and multi-granularity search includes:

[0007] S1. Preprocess and standardize the raw T2DM data;

[0008] S2. Using structural equation modeling and differentiable algebraic constraints of bowless ADMG, construct the linear bowless ADMG causal structure search space;

[0009] S3. Transform the discrete search problem into a continuous optimization task, introduce the Pearson correlation coefficient to improve the scoring function, and correct the mining bias of potential causal edges;

[0010] S4, combining RRICF and coarse-grained multi-granularity search strategies;

[0011] S5. By using alternating iterative optimization of structure and causal strength parameters, the structure learning and strength estimation are simultaneously converged, and the final causal structure of T2DM risk factors is output.

[0012] Furthermore, the linear bowless ADMG causal structure search space is the intersection of the SEM data generation model and the bowless ADMG space;

[0013] The SEM data generation model uses a linear SEM data generation model to constrain the causal relationships between variables.

[0014] set up for The set of observed risk factor variables is defined as follows: Each pair of causal relationships is defined as follows:

[0015] (1)

[0016] in, Let be a noise vector, and its covariance matrix be... ; and for The real number matrices describe the unidirectional causal strength and the bidirectional causal strength, respectively;

[0017] ADMG without bow is represented as ,in For cause-and-effect diagrams, For unidirectional edge causality strength, The two-way edge causality strength, both of which take values ​​in the range [0,1].

[0018] In causality diagram In the middle, two adjacency matrices are used. and These represent unidirectional and bidirectional edge structures, respectively. and hour, There exists a one-way edge; when and hour, There is a bidirectional edge.

[0019] Furthermore, the bowless ADMG is an algebraic constraint, constructed based on the bowless ADMG model, that is, it prohibits the simultaneous existence of unidirectional and bidirectional edges between any pair of variables. Its differentiable algebraic constraint is specifically as follows:

[0020] ①Directed acyclic graph constraints

[0021] (2)

[0022] in, It is a matrix The exponential function, The number of nodes;

[0023] ②Unbowed ADMG constraint

[0024] (3)

[0025] in, It is a matrix and The Hadamard product.

[0026] Furthermore, the transformation of the discrete search problem into a continuous optimization task includes:

[0027] By introducing differentiable algebra, causal discovery is transformed from a discrete search into a continuous optimization problem. The specific formula is as follows:

[0028] (4)

[0029] in, For the reason Induced graph structure, For the structure parameter vector, For the unbowed ADMG set, A scoring function to measure the goodness of fit of a model. It is a differentiable algebraic function;

[0030] Furthermore, the scoring function is improved by introducing the Pearson correlation coefficient to correct the bias in the discovery of potential causal edges, including:

[0031] By using the augmented Lagrange formula, the continuous optimization problem is transformed into an unconstrained optimization problem with a quadratic penalty term, thus yielding the scoring function:

[0032] (5)

[0033] The formula includes a data fitting term. Structural penalty items The two parts, specifically meaning:

[0034] Data fitting term Introducing the approximate Bayesian information criterion, the specific formula is as follows:

[0035] (6)

[0036] in, For model parameters, Let be the likelihood function. The regularization strength parameter is... To control the approximate accuracy parameter; regularization is used to balance the model's fitting ability and complexity;

[0037] Introducing the Pearson correlation coefficient r as a weighting term improves the data fitting term. The specific formula is as follows:

[0038] (7)

[0039] Among them, the correlation coefficient The calculation formula is:

[0040] (8)

[0041] In the formula, , These are the observed value and the predicted value, respectively. It reflects the trend similarity between the predicted value and the actual value, and the value range is [−1,1].

[0042] Structural penalty item The expression is:

[0043] (9)

[0044] in, For weight hyperparameters, For Lagrange multipliers. When When the parameter is not zero, it indicates that the current graph structure violates some constraints. and As it changes dynamically, the penalty term drives the optimization process. It gradually approaches zero.

[0045] Furthermore, the fusion of RRICF and coarse-grained multi-granularity search strategies includes:

[0046] An optimization method combining RRICF and multi-granularity search strategies is employed. By iteratively updating structural parameters and minimizing the scoring function, the method gradually converges to the optimal solution.

[0047] (10)

[0048] Based on the high-dimensionality of risk factors and the characteristics of small sample data, the causal structure optimization process is divided into two stages: global coarse-grained search and local fine-grained search.

[0049] Furthermore, the two stages of global coarse-grained search and local fine-grained search include:

[0050] In the coarse-grained stage, the algorithm operates globally with a learning rate. We will conduct preliminary exploration, explore different solution space regions globally, and provide a better search starting point for subsequent optimization;

[0051] Parameters are updated using the RRICF method within each coarse-grained range, with the specific formula as follows:

[0052] (11)

[0053] in, For the initial structure, This is the intermediate structure after coarse-grained optimization. This is the learning rate for the coarse-grained stage. Indicates the current graph structure The gradient of the total loss function is used to indicate the optimal direction of the structure update;

[0054] The fine-grained stage builds upon the coarse-grained optimization by using a step size... Further exploration of potential areas is conducted, and parameters are updated using the RRICF method at each fine-grained level. The optimization formula is as follows:

[0055] (12)

[0056] in This is the final optimized cause-effect graph structure. This represents the learning rate for the fine-grained stage.

[0057] Furthermore, by using alternating iterative optimization of the structure and causal strength parameters, the simultaneous convergence of structure learning and strength estimation is achieved, outputting the final causal structure of T2DM risk factors, specifically:

[0058] Based on the graph structure induced by the current parameters, a multi-granularity search strategy and RRICF are used to jointly minimize the scoring function in each iteration, and the structural parameters and causal strength matrix are updated synchronously.

[0059] After each iteration, it is determined whether the scoring function has been minimized or the number of iterations has been reached. If the conditions are not met, the constraint penalty is increased and the search strategy is adjusted. The iteration continues until convergence to the optimal causal structure, and the optimal causal structure is output as the final causal structure of T2DM risk factors.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] This invention proposes a differentiable causal discovery method for T2DM risk factors by constructing a bowless ADMG search space, introducing an improved scoring function based on a weighted mechanism, and embedding an improved optimization algorithm with a multi-granularity search strategy. This method can effectively discover potential confounding causal relationships among T2DM risk factors and provides a solution for T2DM prevention and research. Attached Figure Description

[0062] Figure 1 The overall flowchart for implementing the method of this invention is shown below;

[0063] Figure 2 A schematic diagram of the principle of a differentiable causal discovery model for T2DM risk factors that integrates weighting mechanisms and multi-granularity search;

[0064] Figure 3 A simulated causal structure diagram for observing the data;

[0065] Figure 4 The causal structure of the 768 PIMA dataset for Algorithm 1;

[0066] Figure 5 The causal structure of the 768 PIMA dataset for Algorithm 2;

[0067] Figure 6 The causal structure of the 768 PIMA dataset for Algorithm 3;

[0068] Figure 7 The causal structure of the 768 PIMA dataset for Algorithm 4;

[0069] Figure 8 The causal structure of the LMCH dataset for Algorithm 1;

[0070] Figure 9 The causal structure of the LMCH dataset for Algorithm 2;

[0071] Figure 10 The causal structure of the LMCH dataset for Algorithm 3;

[0072] Figure 11 The causal structure of the LMCH dataset for Algorithm 4;

[0073] Figure 12 The proportion of various causal edges in the causal structure of the 768 PIMA dataset;

[0074] Figure 13 This represents the proportion of various causal edges in the causal structure of the LMCH dataset. Detailed Implementation

[0075] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0076] Example 1:

[0077] like Figures 1-2 As shown, a differentiable causal discovery method for T2DM risk factors that integrates weighted mechanisms and multi-granularity search includes the following steps:

[0078] Perform raw data preprocessing and standardization;

[0079] Using structural equation modeling (SEM) and differentiable algebraic constraints of bowless ADMG, a linear bowless ADMG causal structure search space was constructed.

[0080] The discrete search problem is transformed into a continuous optimization task, and the Pearson correlation coefficient is introduced to improve the scoring function, thereby correcting the mining bias of potential causal edges.

[0081] By combining RRICF and coarse-grained multi-granularity search strategies, we can avoid getting trapped in local optima.

[0082] By using alternating iterative optimization of structure and causal strength parameters, the synchronous convergence of structure learning and strength estimation is achieved, and the final causal structure of T2DM risk factors is output.

[0083] The fusion weighting mechanism and multi-granularity search method for the differential causal discovery of T2DM risk factors are used for raw data preprocessing and standardization.

[0084] A causal discovery method is introduced to generate a causal structure of risk factors for type 2 diabetes. This method comprises four stages: constraining the search space, establishing the scoring function, improving the optimization method, and optimizing the causal structure. The details are as follows:

[0085] (1) Constraint Structure Search Space. Using the differentiable algebraic constraints of SEM and bowless ADMG, a linear bowless ADMG causal structure search space was constructed. The search range of the causal structure is the intersection of the SEM data generation model and the bowless ADMG space. The details are as follows:

[0086] ①SEM Data Generation Model. To simultaneously uncover the causal relationships between manifest and latent variables, a linear SEM data generation model is used to constrain the causal relationships between variables. Let... for The set of observed risk factor variables is defined as follows: Each pair of causal relationships is defined as follows:

[0087] (1)

[0088] in, Let be a noise vector, and its covariance matrix be... ; and for The real number matrices describe the unidirectional causal strength and the bidirectional causal strength, respectively.

[0089] like Figure 3 As shown, the bowless ADMG constructed in this invention can be represented as follows: ,in For cause-and-effect diagrams, For unidirectional edge causality strength, For bidirectional edge causality strength, both values ​​range from [0-1]. In the graph model... In the middle, two adjacency matrices are used. and These represent unidirectional and bidirectional edge structures, respectively. and hour, There exists a one-way edge; when That hour, There is a bidirectional edge, that is and There exists a common parent variable that has not been observed.

[0090] ② Differentiable algebraic constraints for bowless ADMG. To obtain more causal information and reduce the difficulty of equivalence class disambiguation, thereby improving model stability and inference accuracy, this invention focuses on the construction of bowless ADMG models, specifically prohibiting the simultaneous existence of both unidirectional and bidirectional edges between any pair of variables. The differentiable algebraic constraints are as follows:

[0091] 1) Directed acyclic graph constraint

[0092] (2)

[0093] in, It is a matrix The exponential function, This represents the number of nodes.

[0094] 2) ADMG constraint without bow

[0095] (3)

[0096] in, It is a matrix and The Hadamard product.

[0097] The above two constraints ensure that the ADMG mined in this invention is acyclic and bowless.

[0098] (2) Establish a scoring function. The discrete search problem is transformed into a continuous optimization task. The Pearson correlation coefficient is introduced to improve the scoring function, thereby correcting the bias in uncovering potential causal edges. The specific details are as follows:

[0099] By introducing differentiable algebraic constraints, causal discovery can be transformed from a discrete search into a continuous optimization problem, which can be formally equivalent to:

[0100] (4)

[0101] in, For the reason Induced graph structure, For the structure parameter vector, For the unbowed ADMG set, A scoring function to measure the goodness of fit of a model. It is a differentiable algebraic function.

[0102] To efficiently solve the above optimization problem, this invention employs the augmented Lagrange formula, transforming the original problem into an unconstrained optimization problem with a quadratic penalty term. This yields the following scoring function:

[0103] (5)

[0104] It includes two parts: a data fitting term and a structure penalty term, the specific meanings of which are as follows:

[0105] ① Data Fitting Term. This term measures the difference between the model's predicted causal relationships and the actual observed data. The Bayesian Information Criterion (BIC) is a commonly used choice. However, BIC contains a non-differentiable indicator function, making it difficult to use directly for continuous optimization. Therefore, the Approximate Bayesian Information Criterion (ABIC) is introduced:

[0106] (6)

[0107] in, For model parameters, Let be the likelihood function. The regularization strength parameter is... To control the approximate accuracy parameter; regularization is used to balance the model's fitting ability and complexity.

[0108] To more accurately capture the linear dependencies and potential confounding factors among variables, the Pearson correlation coefficient r is introduced as a weighting term, and the ABIC model is improved as follows:

[0109] (7)

[0110] Among them, the correlation coefficient The calculation formula is:

[0111] (8)

[0112] In the formula, , These are the observed value and the predicted value, respectively. It reflects the trend similarity between the predicted value and the actual value, and the value range is [−1,1].

[0113] ② Structural penalty term. This is used to ensure that the cause-effect graph structure satisfies the aforementioned structural constraints. By penalizing solutions that do not meet the constraints, the cause-effect graph is guaranteed to meet the bowless ADMG requirement. The mathematical expression for the structural penalty term is:

[0114] (9)

[0115] in, For weight hyperparameters, For Lagrange multipliers. When When the parameter is not zero, it indicates that the current graph structure violates some constraints. and This changes dynamically. During the optimization process, the penalty term drives... It gradually approaches zero, thereby ensuring that structural constraints are met.

[0116] The improved scoring function can correct the bias in the discovery of potential causal edges during the causal discovery process. Its formula is as follows:

[0117] (10)

[0118] (3) Improved optimization method. By integrating RRICF and coarse-grained multi-granularity search strategies to avoid getting trapped in local optima, and by iteratively optimizing the structure and causal strength parameters alternately, the structure learning and strength estimation are simultaneously converged to obtain the final causal structure of T2DM risk factors. The details are as follows:

[0119] An optimization method combining RRICF and a multi-granularity search strategy is employed. By iteratively updating the structural parameters and minimizing the scoring function, the method gradually converges to the optimal solution.

[0120] (11)

[0121] The optimization algorithm proposed in this invention comprehensively considers the high-dimensional characteristics of risk factors and the features of small sample data, and divides the causal structure optimization process into two stages: global coarse-grained search and local fine-grained search, as described below:

[0122] ① Coarse-grained stage. The optimization algorithm operates globally with a large learning rate. Preliminary exploration helps to avoid local extrema traps, globally explore different solution space regions, and provide a better search starting point for subsequent optimization; parameters are updated using the RRICF method within each coarse-grained range, the expression of which is:

[0123] (12)

[0124] in, For the initial structure, This is the intermediate structure after coarse-grained optimization. This is the learning rate for the coarse-grained stage. Indicates the current graph structure The gradient of the total loss function is used to indicate the optimal direction for structural updates.

[0125] ② Fine-grained stage. Based on the coarse-grained optimization, a smaller step size is adopted. By delving deeper into the potential region and refining the local solution, parameters are updated using the RRICF method at each fine-grained level. The optimization formula is as follows:

[0126] (13)

[0127] in This is the final optimized cause-effect graph structure. This represents the learning rate for the fine-grained stage.

[0128] (4) Optimize the causal structure. Alternating iterative optimization of the structure and causal strength parameters is used to achieve synchronous convergence of structure learning and strength estimation, yielding the final causal structure of T2DM risk factors. The specific method is described below: Based on the graph structure induced by the current parameters, a multi-granularity search strategy and RRICF are used to jointly minimize the scoring function during each iteration, synchronously updating the structure parameters and causal strength matrix. After each iteration, the algorithm determines whether the scoring function has been minimized or the iteration count has been reached. If the conditions are not met, the constraint penalty is increased and the search strategy is adjusted, continuing iteratively until convergence to the optimal causal structure. Finally, the optimal causal structure of T2DM is output.

[0129] Example 2:

[0130] This example presents a differentiable causal discovery method for risk factors in type 2 diabetes mellitus (T2DM) that integrates a weighted mechanism and multi-granularity search. Experimental data includes three typical T2DM datasets: the Pima Indians (PIMA) dataset with 768 samples, and the Iraqi dataset with 947 samples released by the Medical City Hospital Laboratory (LMCH). To verify the effectiveness of the proposed method, this section uses the following four methods to compare and analyze the algorithm's performance on the PIMA and LMCH datasets:

[0131] (1) A linear causal discovery method based on reinforcement learning (Algorithm 1);

[0132] (2) Original differentiable causal discovery method (Algorithm 2);

[0133] (3) Differentiable causal discovery method incorporating Pearson weighted mechanism (Algorithm 3);

[0134] (4) A differentiable causal discovery method that integrates Pearson weighting mechanism and multi-granularity search strategy (Algorithm 4, the algorithm of this invention).

[0135] Algorithm 1 can only identify unidirectional causal relationships, while Algorithms 2-4 can identify both unidirectional and bidirectional causal relationships. Experiments with Algorithms 1 and 2-4 are used to compare the accuracy and completeness of causal structures; experiments with Algorithms 2 and 3-4 are used to verify the advantages of the improved algorithms.

[0136] Comparative Experiments on Causal Discovery on the PIMA Dataset

[0137] The PIMA dataset includes eight risk factors: number of pregnancies. 2-hour glucose concentration diastolic blood pressure triceps skin fold thickness 2-hour insulin Body Mass Index (BMI) Diabetes pedigree function ,age .

[0138] In the causal discovery experiment on the 768 PIMA dataset, by Figure 4 It can be seen that Algorithm 1 only identified 3 pairs of unidirectional causal relationships: insulin affects plasma glucose concentration ( The thickness of the triceps skin folds affects BMI. Number of pregnancies affects age () ).Depend on Figure 5 It can be seen that Algorithm 2 identifies 3 more pairs of unidirectional causal relationships and 3 more pairs of bidirectional causal relationships than Algorithm 1: BMI affects the thickness of triceps skin folds ( Age affects the number of pregnancies. Diastolic blood pressure affects age ( Insulin and BMI share a common latent parent variable. Plasma glucose concentration and age share a common latent parent variable. Diastolic blood pressure and BMI share a common latent parent variable. ).Depend on Figure 6 It can be seen that Algorithm 3... Figure 5 In Identified as And the remaining causal edges and Figure 5 Consistent. By Figure 7 It can be seen that Algorithm 4 discovered a new pair of bidirectional causal relationships compared to Algorithm 3: plasma glucose concentration and BMI share a common latent parent variable. ).

[0139] Comparative Experiments on Causal Discovery in the LMCH Dataset

[0140] The LMCH dataset includes 11 risk factors: sex ,age urea Creatinine ratio HbA1c ,cholesterol Triglycerides High-density lipoprotein (HDL) Low-density lipoprotein (LDL) Very low density lipoprotein (VLDL) and BMI .

[0141] Depend on Figure 8 It can be seen that Algorithm 1 identifies two pairs of unidirectional causal relationships: urea affects the creatinine ratio ( ) and LDL affect cholesterol ( ).Depend on Figure 9 It can be seen that Algorithm 2 not only identified 3 pairs of unidirectional causal relationships: BMI affects HbA1c ( Urea affects the creatinine ratio. HbA1c affects age () It also identified 7 pairs of bidirectional causal relationships: gender and VLDL ( VLDL and BMI VLDL and age ( HbA1c and triglycerides HbA1c and cholesterol ), cholesterol and triglycerides ), cholesterol and LDL ( ), cholesterol and HDL ( ), and HDL and LDL ( ) and other variables all share a common potential parent variable. Figure 10 It can be seen that Algorithm 3 will Identified as The remaining causal edges and Figure 9 Consistent. By Figure 11It can be seen that Algorithm 4 has discovered a new pair of bidirectional causal relationships: gender and creatinine ratio share a common latent parent variable. ).

[0142] Discussion of the causal structure of the PIMA dataset

[0143] (1) A correct causal relationship. This is explained as follows:

[0144] 1) Insulin affects plasma glucose concentration. BMI affects the thickness of triceps skin folds. Age affects the number of pregnancies. Age affects diastolic blood pressure. Existing research has confirmed the validity of these relationships.

[0145] 2) Insulin and BMI are influenced by common underlying factors. Studies have shown that both are affected by FTO gene variations, and the interaction between insulin and BMI is one of the typical mechanisms of T2DM development. Carrying the FTO risk allele can affect energy intake and fat accumulation, making individuals more prone to weight gain (increased BMI); at the same time, adipose tissue dysfunction exacerbates insulin resistance.

[0146] 3) Diastolic blood pressure and BMI are influenced by common underlying factors. Studies have shown that both are affected by insulin resistance. Insulin resistance reduces glucose uptake by adipocytes, leading to more energy being stored as fat, thus causing an increase in BMI. Simultaneously, high insulin levels increase peripheral vascular resistance by promoting tubular reabsorption and activating the sympathetic nervous system, leading to an increase in diastolic blood pressure.

[0147] 4) Both blood glucose and BMI are influenced by common underlying factors. Studies have shown that both are affected by chronic low-grade inflammation. Obesity causes adipose tissue to secrete various inflammatory factors (such as IL-6 and CRP). These inflammatory factors not only impair the insulin signaling pathway but also promote insulin resistance and elevated blood glucose, further driving fat accumulation and causing a sustained increase in BMI.

[0148] (2) Causal relationship to be verified. and Blood glucose and age, as well as diastolic blood pressure and age, are both influenced by common underlying factors. Studies have shown that certain unobservable temporal variables (such as generation effects, early environmental exposures, growth environment, and early health conditions) can simultaneously affect indicators such as age, diastolic blood pressure, and blood glucose. These temporal factors have long-term effects on blood pressure developmental trajectories, hypertension risk, and overall metabolic health. Accurately identifying these temporal variables is crucial for revealing age-related risk factors for type 2 diabetes mellitus (T2DM) and helps avoid misattribution or confounding bias caused by temporal latent variables. Therefore, the plausibility of these two pairs of causal relationships requires further medical validation.

[0149] (3) False causal relationship. Figure 4 of and The thickness of the triceps skin folds affects BMI, and the number of pregnancies affects age. Figure 5 of Studies have shown that the relationship between diastolic blood pressure and age is illogical.

[0150] In summary, the analysis shows that for the 768 PIMA experiment, in Figure 4 In addition to All edges other than the correct ones are incorrect; Figure 5 middle, This is an incorrect edge. The edge shown is to be verified; all others are correct edges. Figure 6 Eliminated Figure 5 Error causal edge in And identify more edges to be verified. ; Figure 7 exist Figure 6 Based on this, more correct causal edges were explored. .

[0151] Discussion of the causal structure of the LMCH dataset

[0152] (1) A correct causal relationship. This is explained as follows:

[0153] 1) BMI affects HbA1c. Studies have shown that high BMI is often accompanied by insulin resistance, which prevents the body from effectively utilizing glucose, leading to elevated blood sugar levels and affecting HbA1c levels.

[0154] 2) To illustrate that the sex-to-creatinine ratio is influenced by common underlying factors, research indicates that both are affected by differences in sex chromosomes. Sex chromosomes not only determine sex but also influence kidney metabolism and physiological function by regulating sex hormones (such as testosterone and estrogen), thereby affecting serum creatinine levels. Furthermore, the expression of related genes on sex chromosomes also affects creatinine metabolism.

[0155] 3) High triglycerides and HbA1c are influenced by common underlying factors. Studies have shown that both are affected by insulin resistance, which reduces cellular sensitivity to insulin, leading to elevated insulin levels and impaired glycemic regulation, thereby increasing HbA1c. Insulin resistance also disrupts lipid metabolism, prompting the liver to synthesize and secrete more VLDL. During the breakdown of VLDL, triglycerides are released, resulting in elevated triglyceride levels in the blood.

[0156] 4) Triglycerides and cholesterol are influenced by common underlying factors. Studies have shown that genes closely related to lipoprotein synthesis, breakdown, and metabolism (such as genes related to apolipoproteins and lipoprotein lipases) have a significant impact on triglyceride and cholesterol levels.

[0157] 5) Both HbA1c and cholesterol are influenced by common underlying factors. Studies have shown that both are affected by insulin resistance. Insulin resistance reduces glucose uptake in peripheral tissues, leading to persistently elevated blood glucose levels, which in turn increases HbA1c levels. The imbalance in insulin's regulation of lipid metabolism both promotes hepatic VLDL synthesis and reduces lipoprotein lipase activity, resulting in elevated total cholesterol levels.

[0158] 6) Sex and VLDL are influenced by common underlying factors. Studies have shown that both are affected by gene expression regulated by sex chromosomes and genetic polymorphisms in lipid metabolism-related genes. Sex chromosome regulation of lipid metabolism genes and differences in the sex-specific distribution of gene polymorphisms can both lead to differences in VLDL levels between different sexes.

[0159] 7) Cholesterol and HDL are influenced by common underlying factors. Studies have shown that both are affected by insulin resistance and metabolic abnormalities. Insulin resistance leads to increased cholesterol synthesis in the liver while inhibiting HDL synthesis; metabolic abnormalities promote cholesterol accumulation, reduce HDL-mediated cholesterol reversal transport capacity, and thus increase total cholesterol while decreasing HDL.

[0160] 8) This indicates that HDL and LDL are influenced by common underlying factors. Studies have shown that not only can polymorphisms in apolipoprotein genes such as APOE, APOA1, and APOB simultaneously affect the concentrations of HDL and LDL, but many genes (such as LDLR and PCSK9) are also closely related to LDL levels and influence HDL regulation.

[0161] 9) This indicates that VLDL and BMI are influenced by common underlying factors. Studies have shown that unhealthy lifestyle habits are a common cause of high BMI and abnormal lipid levels (such as elevated VLDL). Consuming high-calorie, high-fat foods easily leads to weight gain and increases VLDL synthesis in the liver; lack of exercise not only promotes fat accumulation and reduces energy expenditure but may also cause insulin resistance, thus affecting VLDL levels.

[0162] 10) Both cholesterol and LDL are influenced by common underlying factors. Studies have shown that both are affected by insulin resistance. Insulin resistance leads to increased VLDL secretion, which is converted into LDL in the bloodstream, thereby raising serum LDL and total cholesterol levels, a key step in the occurrence and progression of T2DM. Furthermore, insulin resistance also inhibits lipoprotein lipase activity, reducing the breakdown and clearance of LDL.

[0163] (2) The causal relationship to be verified is as follows:

[0164] 1) Urea levels affect the creatinine ratio. Urea is an end product of protein breakdown, while creatinine mainly comes from muscle metabolism and meat intake; both are commonly used as indicators of kidney function. Related studies have shown that urinary protein, blood urea, and serum creatinine have significant clinical implications for type 2 diabetes mellitus (T2DM), and the urea-creatinine ratio becomes abnormal as the disease progresses. However, the direct causal effect of urea levels on the creatinine ratio lacks sufficient empirical evidence and requires further clinical validation.

[0165] 2) and VLDL and age, as well as HbA1c and age, are both influenced by common underlying factors. Studies have shown that certain unobservable temporal variables (such as growth environment, early nutrition, and lifestyle) can affect population age structure, VLDL concentration, and HbA1c distribution. These temporal factors can influence lipid profiles by regulating the expression of key lipid metabolism genes (such as APOE and APOB), thus having a long-term impact on population age structure and metabolic health. Therefore, the plausibility of these two causal relationships requires further medical verification.

[0166] (3) The incorrect causal relationship is as follows:

[0167] 1) It is clearly unreasonable to suggest that HbA1c affects age.

[0168] 2) LDL affects cholesterol. Studies have shown that total cholesterol levels determine the cholesterol content carried by various lipoproteins (such as LDL). LDL is only the main transport carrier of cholesterol, and its level varies depending on total cholesterol and metabolic status.

[0169] In summary, the analysis shows that for the LMCH experiment, in Figure 8 middle, To verify the edges, For incorrect edges; in Figure 9 middle, This is an incorrect edge. and The edge shown is to be verified; all others are correct edges. Figure 10 Eliminated Figure 9 Error edges, and identify more edges to be verified. ; Figure 11 exist Figure 10 Based on this, more correct causal edges were explored. .

[0170] Comprehensive comparative analysis

[0171] To systematically evaluate the causal structure recognition capability of the proposed method on different datasets, this paper presents a quantitative analysis of the aforementioned experimental results. Figure 12 and Figure 13 It shows the proportion of various causal edges in the causal structure.

[0172] Our proposed method outperforms other methods in both the number of correct edges and the number of edges to be verified, with the increase in the number of edges to be verified demonstrating its strong potential for causal exploration. Furthermore, on the 768 PIMA dataset, Algorithm 4 achieves accuracy improvements of 41.7%, 3.6%, and 3.6% compared to Algorithms 1, 2, and 3, respectively. On the LMCH dataset, Algorithm 4 achieves accuracy improvements of 76.9%, 1.9%, and 1.9% compared to Algorithms 1, 2, and 3, respectively.

[0173] Depend on Figure 12 It can be seen that on the 768 PIMA dataset, the percentages of erroneous edges for Algorithms 1 and 2 are 66.7% and 14.3%, respectively, while the percentages of unverified edges for Algorithms 2 and 4 are 14.3%, 28.6%, and 25%, respectively. However, Algorithm 4 has the highest percentage of correct edges (75%). Figure 13 It can be seen that on the LMCH dataset, the percentages of erroneous edges for Algorithms 1 and 2 are 50% and 8.3% respectively, the percentages of unverified edges for Algorithms 2 and 4 are 16.7%, 25% and 23.1% respectively, and Algorithm 4 has the highest percentage of correct edges (76.9%) and no erroneous edges.

[0174] Based on the data above, we can conclude that:

[0175] (1) In the 768 PIMA dataset, when Algorithm 1 performs the causal discovery task on a small sample dataset, ignoring potential confounding factors will result in low causal recognition accuracy and incomplete structure; although Algorithm 2 introduces a potential factor processing mechanism, it fails to fully integrate risk factor information, which limits the model's exploration ability and thus limits the recognition accuracy. Algorithm 4 fully solves the above problems and improves the recognition accuracy and topological completeness of the causal structure.

[0176] (2) Comparison of experiments on the PIMA and LMCH datasets shows that, in cross-domain data environments, the weighted mechanism consistently demonstrates robust causal mining advantages and effectively suppresses the generation of erroneous causal edges. Based on this, the multi-granularity structure search strategy enhances the decoupling ability of implicit causal topological structures by balancing exploration depth and breadth, revealing more non-explicit causal relationships. These two collaborative innovations achieve a systematic optimization of the causal network identification quality.

[0177] In conclusion, the following conclusions can be drawn:

[0178] (1) On the PIMA dataset, increasing the sample size has a significant impact on the performance of algorithms 1 and 4, but no significant impact on algorithms 2 and 3. This indicates that the proposed method has sample size adaptability and can maintain the robustness of network construction.

[0179] (2) By comparing the causal structure across heterogeneous datasets such as PIMA and LMCH, this method consistently demonstrates superior generalization performance and causal topological robustness, empirically proving its domain adaptability.

[0180] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for discovering differentiable causal relationships of risk factors in T2DM that integrates weighted mechanisms and multi-granularity search, characterized in that, include: S1. Preprocess and standardize the raw T2DM data; S2. Using structural equation modeling and differentiable algebraic constraints of bowless ADMG, construct the linear bowless ADMG causal structure search space; S3. Transform the discrete search problem into a continuous optimization task, introduce the Pearson correlation coefficient to improve the scoring function, and correct the mining bias of potential causal edges; S4, combining RRICF and coarse-grained multi-granularity search strategies; S5. By using alternating iterative optimization of structure and causal strength parameters, the structure learning and strength estimation are simultaneously converged, and the final causal structure of T2DM risk factors is output.

2. The method for discovering differentiable causal relationships of T2DM risk factors by fusing weighted mechanisms and multi-granularity search according to claim 1, characterized in that, The linear bowless ADMG causal structure search space is the intersection of the SEM data generation model and the bowless ADMG space; The SEM data generation model uses a linear SEM data generation model to constrain the causal relationship between variables. set up for The set of observed risk factor variables is defined as follows: Each pair of causal relationships is defined as follows: (1) in, Let be a noise vector, and its covariance matrix be... ; and for The real number matrices describe the unidirectional causal strength and the bidirectional causal strength, respectively; ADMG without bow is represented as ,in For cause-and-effect diagrams, For unidirectional edge causality strength, The two-way edge causality strength, both of which take values ​​in the range [0,1]. In causality diagram In the middle, two adjacency matrices are used. and These represent unidirectional and bidirectional edge structures, respectively. and hour, There exists a one-way edge; when and hour, There is a bidirectional edge.

3. The method for discovering differentiable causal relationships of T2DM risk factors by fusing weighted mechanisms and multi-granularity search according to claim 2, characterized in that, The bowless ADMG is an algebraic constraint, constructed based on the bowless ADMG model, that is, prohibiting the simultaneous existence of unidirectional and bidirectional edges between any pair of variables. Its differentiable algebraic constraint is as follows: ①Directed acyclic graph constraints (2) in, It is a matrix The exponential function, The number of nodes; ②Unbowed ADMG constraint (3) in, It is a matrix and The Hadamard product.

4. The method for discovering differentiable causal relationships of T2DM risk factors by integrating weighted mechanisms and multi-granularity search as described in claim 1, characterized in that, The process of transforming a discrete search problem into a continuous optimization task includes: By introducing differentiable algebra, causal discovery is transformed from a discrete search into a continuous optimization problem. The specific formula is as follows: (4) in, For the reason Induced graph structure, For the structure parameter vector, For the unbowed ADMG set, A scoring function to measure the goodness of fit of a model. It is a differentiable algebraic function.

5. The method for discovering differentiable causality of T2DM risk factors by integrating weighted mechanisms and multi-granularity search as described in claim 4, characterized in that, Introducing the Pearson correlation coefficient to improve the scoring function and correct for biases in the discovery of potential causal edges includes: By using the augmented Lagrange formula, the continuous optimization problem is transformed into an unconstrained optimization problem with a quadratic penalty term, thus yielding the scoring function: (5) The formula includes a data fitting term. Structural penalty items The two parts, specifically meaning: Data fitting term Introducing the approximate Bayesian information criterion, the specific formula is as follows: (6) in, For model parameters, Let be the likelihood function. Here is the regularization strength parameter. To control the approximate accuracy parameter; regularization is used to balance the model's fitting ability and complexity; Introducing the Pearson correlation coefficient r as a weighting term improves the data fitting term. The specific formula is as follows: (7) Among them, the correlation coefficient The calculation formula is: (8) In the formula, , These are the observed value and the predicted value, respectively. It reflects the trend similarity between the predicted value and the actual value, and the value range is [−1,1]. Structural penalty item The expression is: (9) in, For weight hyperparameters, For Lagrange multipliers. When When the parameter is not zero, it indicates that the current graph structure violates some constraints. and As it changes dynamically, the penalty term drives the optimization process. It gradually approaches zero.

6. The method for discovering differentiable causal relationships of T2DM risk factors by integrating weighted mechanisms and multi-granularity search as described in claim 1, characterized in that, The fusion of RRICF and coarse-grained multi-granularity search strategies includes: An optimization method combining RRICF and multi-granularity search strategies is employed. By iteratively updating structural parameters and minimizing the scoring function, the method gradually converges to the optimal solution. (10) Based on the high-dimensionality of risk factors and the characteristics of small sample data, the causal structure optimization process is divided into two stages: global coarse-grained search and local fine-grained search.

7. The method for discovering differentiable causality of T2DM risk factors by integrating weighted mechanisms and multi-granularity search as described in claim 6, characterized in that, The two stages of global coarse-grained search and local fine-grained search include: In the coarse-grained stage, the algorithm operates globally with a learning rate. We will conduct preliminary exploration, explore different solution space regions globally, and provide a better search starting point for subsequent optimization; Parameters are updated using the RRICF method within each coarse-grained range, with the specific formula as follows: (11) in, For the initial structure, This is the intermediate structure after coarse-grained optimization. This is the learning rate for the coarse-grained stage. Indicates the current graph structure The gradient of the total loss function is used to indicate the optimal direction of the structure update; The fine-grained stage builds upon the coarse-grained optimization by using a step size... Further exploration of potential areas is conducted, and parameters are updated using the RRICF method at each fine-grained level. The optimization formula is as follows: (12) in This is the final optimized cause-effect graph structure. This represents the learning rate for the fine-grained stage.

8. The method for discovering differentiable causality of T2DM risk factors by integrating weighting mechanisms and multi-granularity search as described in claim 7, characterized in that, By employing alternating iterative optimization of structure and causal strength parameters, simultaneous convergence of structure learning and strength estimation is achieved, outputting the final causal structure of T2DM risk factors, specifically: Based on the graph structure induced by the current parameters, a multi-granularity search strategy and RRICF are used to jointly minimize the scoring function in each iteration, and the structural parameters and causal strength matrix are updated synchronously. After each iteration, it is determined whether the scoring function has been minimized or the number of iterations has been reached. If the conditions are not met, the constraint penalty is increased and the search strategy is adjusted. The iteration continues until convergence to the optimal causal structure, and the optimal causal structure is output as the final causal structure of T2DM risk factors.