A bivariate robust causal dose-response curve estimation method based on high-dimensional covariates

By constructing a dual weighting method using a penalty weighting function and a distance correlation coefficient, important covariates are screened out. Combined with a DR estimator, the variable selection problem for causal inference in high-dimensional data is solved, achieving high-precision and efficient dose-response curve estimation with dual robustness.

CN116206777BActive Publication Date: 2025-11-25SHANXI MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310274302.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-21
Publication Date
2025-11-25
Estimated Expiration
2043-03-21

AI Technical Summary

Technical Problem

Under high-dimensional data conditions, existing techniques struggle to effectively screen out important confounding variables and prognostic covariates, leading to bias and variance inflation in dose-response function estimation during causal inference. In particular, under high-dimensional covariate conditions, existing methods have failed to effectively address the challenges of variable selection and conditional probability density function specification.

Method used

A doubly robust causal dose-response curve estimation method based on high-dimensional independent variables (GOALDeR) is adopted. By constructing a penalized weight function and distance correlation coefficient, a modified adaptive LASSO method is used to screen covariates, and the optimal λn is selected by double-weighted distance correlation coefficient. Combined with the DR estimator, the dose-response curve is estimated to achieve unbiased estimation.

Benefits of technology

It improves the accuracy and precision of causal inference, reduces sensitivity to the correlation structure of covariates and the ratio of sample size to covariate dimensionality, and has dual robustness, enabling unbiased and efficient dose-response curve estimation under high-dimensional data conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116206777B_ABST
    Figure CN116206777B_ABST
Patent Text Reader

Abstract

The application is a kind of double robust causal dose-response curve estimation method based on high-dimensional independent variables, comprising the following steps: 1) constructing a new target function based on the modified adaptive LASSO method to realize dimension reduction; 2) constructing a double weighted distance correlation coefficient DWDC to select the optimal lambda n ; 3) using DR estimator to estimate the dose-response curve. The method mainly aims at the characteristics of the potential confounding variable set contained in the health medical big data, such as high dimension, nonlinear relationship between health outcome and (or) exposure factor, etc. In the framework of GOAL method, a double robust estimation method of causal dose-response curve is provided, which is called GOALDeR method. A large number of statistical simulations show that the estimation accuracy and precision of GOALDeR method are better than those of existing methods, and the method has double robustness and is less affected by the correlation structure between covariates and the n / p (sample size / covariate dimension) ratio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of biotechnology and data analysis technology, specifically a method for estimating bi-robust causal dose-response curves based on high-dimensional independent variables. Background Technology

[0002] In non-randomized studies, the propensity score (PS) method for binary exposure factors is widely used to control for measured confounding factors in order to achieve an unbiased and efficient estimate of the treatment effect. Imai et al. (2004) extended the PS method, proposing the Generalized Propensity Score (GPS), which makes it possible to obtain an unbiased and efficient dose-response function (DRF) / curve between continuous exposure factors and health outcomes from non-randomized studies. GPS is defined as the unbiased and efficient estimate of the treatment effect given a covariate X when an individual is exposed to a specific level (…). T=t The conditional probability density of (GPS) can be used to estimate the treatment effect. Utilizing the equilibrium characteristics of GPS, unbiased estimation of the treatment effect can be achieved through matching, stratification, regression adjustment, and inverse probability weighting. The unbiased and efficient estimation of DRF depends on the consistent estimation of GPS, which in turn depends primarily on the following two points:

[0003] (1) Correct model specification, i.e., the setting of the functional form of covariates entering the GPS model and the conditional probability density function. Currently, GPS methods developed to address the potential for model misspecification can be divided into two categories: model-based GPS methods and equilibrium-based GPS methods. The former still carries the risk of misspecification because it requires manual specification of the conditional probability density function; the latter, although having a clear mechanism to ensure that covariates achieve equilibrium across different exposure levels and bypassing the difficulties of setting the GPS model and conditional probability density function, does not consider the issue of variable selection, and therefore is not suitable for causal inferences containing high-dimensional covariates.

[0004] (2) Selecting appropriate covariates. The GPS method is sensitive to the covariates included in the model. For example, including instrumental variables (covariates related to the exposure factor but not to the health outcome) will lead to variance inflation of the effect estimate without reducing bias. Previous simulation studies have found that in addition to including confounding variables to control confounding bias, the optimal GPS model should also include prognostic covariates (covariates unrelated to the exposure factor but related to the health outcome) to improve the variance of the estimate. With the increase of available data, such as omics data and the accumulation of electronic medical records, we can collect a large number of even high-dimensional covariates in actual medical research. How to screen out important confounding factors and prognostic covariates to achieve unbiased and efficient estimation of DRF is an urgent problem to be solved in causal inference based on the GPS method. The high-dimensional GPS method has thus been developed. Zhu et al. (2015) and Guo Wei (2018) used machine learning methods to fit the GPS model and then estimated the GPS based on the Gaussian density function. Su et al. (2019) and Antonelli et al. (2022), within the framework of doubly robust estimation (DR), used flexible modeling strategies such as Gaussian processes and deep neural networks to fit the outcome model and the GPS model. The advantage of the DR estimator is that as long as one of the outcome model and the GPS model is correctly specified, a consistent estimate of the treatment effect can be obtained. However, the above methods do not consider the influence of instrumental variables on the variance of the effect estimate, and still require the conditional probability density function to be specified. To address this, Gao et al. (2021) proposed the GOAL (Generalized outcome-adaptive LASSO) method, but this method relies on the linear assumption of the outcome model, i.e., it does not have doubly robustness, and is also susceptible to the influence of the correlation structure among covariates. Summary of the Invention

[0005] This invention considers that the GOAL method can both achieve the selection of causal inference variables and bypass the difficulties of setting GPS models and conditional probability density functions. Therefore, in order to address the shortcomings of the GOAL method, a double robust causal dose-response curve estimation method based on high-dimensional independent variables is proposed, which is also known as the GOALDeR (generalized outcome-adaptive LASSO and doubly robust estimation) method. This method uses the distance correlation coefficient, which can measure any correlation between covariates, to construct a penalty weight function and evaluate the balance of covariates under different exposure levels.

[0006] This invention is achieved through the following technical solution:

[0007] A biroleus causal dose-response curve estimation method based on high-dimensional independent variables includes the following steps:

[0008] 1) A new objective function is constructed based on the modified adaptive LASSO method to achieve dimensionality reduction;

[0009] Construct a GPS model and use a modified adaptive LASSO method to screen the covariates that need to be balanced or included in the GPS model;

[0010] Specifically, the GPS model is represented as:

[0011] (1)

[0012] The objective function is:

[0013] (2)

[0014] in, , γ > 1 indicates the penalty weight. dcor ( X j ,Y | T () represents the covariates given the exposure factors. X j The conditional distance correlation coefficient between the conditional distance and health outcomes ẑ It is a constant used to standardize the coefficient; λ n >0 is an adjustment parameter, when λ n satisfy and At that time, the objective function (2) can filter out confounding variables and prognostic covariates according to probability 1; therefore, a set of variables that meet the requirements is set. and alternative conditions λ n and according to λ n Select a set of candidate covariates;

[0015] 2) Construct a dual-weighted distance correlation coefficient (DWDC) and select the optimal one. λ n ;

[0016] The GPS method for obtaining unbiased estimates of dose-response curves relies on the degree of equilibrium of covariate distribution across different exposure levels, meaning there is no arbitrary correlation, including linear and nonlinear, between the covariates and the exposure factors. Based on this, a distance correlation coefficient (DWDC) is constructed using this coefficient as an evaluation index of equilibrium, and the optimal DWDC is selected by minimizing the DWDC. λ n :

[0017] (3)

[0018] in, ŵ λn The equilibrium weights are estimated by the DCOWs method;

[0019] 3) Use a DR estimator to estimate the dose-response curve;

[0020] First, calculate the false ending using the following formula:

[0021] (4)

[0022] in, w i The balanced weights are based on optimality. λ n The selected covariates were estimated using the DCOWs method; The predicted values ​​of the outcome variables are represented by the SL method for estimation; then, the calculated pseudo-outcomes and treatment factors are used to construct univariate linear or nonlinear models to estimate the dose-response curve.

[0023] Furthermore, in (1), the constructed GPS model is a model with exposure factor T as the dependent variable.

[0024] This invention addresses the characteristics of high-dimensionality and non-linear relationships between potentially confounding variables in health and medical big data and health outcomes and / or exposure factors. Within the framework of the GOAL method, it provides a doubly robust estimation method for causal dose-response curves, termed the GOALDeR method. Extensive statistical simulations demonstrate that the GOALDeR method outperforms existing methods in terms of estimation accuracy and precision, exhibits doubly robust properties, and is less affected by the correlation structure among covariates and the n / p (sample size / covariate dimensionality) ratio. Attached Figure Description

[0025] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the technical description of the present invention will be briefly introduced below. The accompanying drawings are used to provide further explanation of the present invention and constitute a part of this application.

[0026] Figure 1 This is a schematic diagram of the method of the present invention.

[0027] Figures 2 to 7 This diagram illustrates the performance of the method of the present invention in screening correction covariates under different simulation scenarios.

[0028] Figures 8 to 13 This is a distribution diagram of the causal parameter estimates under different simulation scenarios for the method of this invention, as well as the existing GOAL and SL-DR methods.

[0029] Figure 14 A schematic diagram illustrating the covariate balance before and after weighting each dataset. Implementation

[0030] The technical solution of the present invention will now be described more clearly and completely with reference to the accompanying drawings.

[0031] A biroleus causal dose-response curve estimation method based on high-dimensional independent variables includes the following steps:

[0032] (1) A new objective function is constructed based on the modified adaptive LASSO method to achieve dimensionality reduction;

[0033] A GPS model is constructed with exposure factor T as the dependent variable. A modified adaptive LASSO method is used to screen covariates that need to be balanced or included in the GPS model. The penalized weight function here takes into account the conditional correlation between covariates and outcomes.

[0034] Specifically, the GPS model is represented as:

[0035] (1)

[0036] The objective function is:

[0037] (2)

[0038] in, , γ > 1 indicates the penalty weight. dcor ( X j ,Y | T () represents the covariates given the exposure factors. X j The conditional distance correlation coefficient between the conditional distance and health outcomes ẑ It is a constant used to standardize the coefficient; this penalty weighting function does not depend on the outcome model and achieves the selection of confounding and prognostic covariates by imposing a larger penalty on covariates that are unrelated to the outcome. λ n >0 is an adjustment parameter, when λ n satisfy and At that time, the objective function (2) can filter out confounding variables and prognostic covariates according to probability 1; therefore, a set of variables that meet the requirements is set. and alternative conditions λ n and according to λ nSelect a set of candidate covariates;

[0039] (2) Constructing the dual-weighted distance correlation coefficient (DWDC) and selecting the optimal one. λ n ;

[0040] The GPS method for obtaining unbiased estimates of dose-response curves relies on the degree of equilibrium of covariate distribution across different exposure levels, meaning there is no arbitrary correlation, including linear and nonlinear, between the covariates and the exposure factors. Based on this, a dual-weight distance correlation (DWDC) is constructed using the distance correlation coefficient, which can measure arbitrary correlations between variables, as an evaluation index of equilibrium. The optimal DWDC is selected by minimizing the DWDC. λ n :

[0041] (3)

[0042] in, ŵ λn The equilibrium weights are estimated by the DCOWs method (Huling, 2021). The DCOWs method uses the distance correlation coefficient between the covariate and the exposure factor as the loss function. Under the constraint that the marginal distribution of the exposure factor and the covariate remains unchanged before and after weighting, the equilibrium weights are directly estimated. It has less dependence on the form of the GPS model and the conditional probability density function.

[0043] (3) Use the DR estimator to estimate the dose-response curve;

[0044] First, calculate the false ending using the following formula:

[0045] (4)

[0046] in, w i The balanced weights are based on optimality. λ n The selected covariates were estimated using the DCOWs method; The predicted values ​​of the outcome variables are estimated using the Super Learner (SL) method, which integrates the predictions from linear and nonlinear models. Subsequently, the calculated pseudo-outcomes and treatment factors are used to construct univariate linear or nonlinear models to estimate the dose-response curve. For ease of simulation comparison, this embodiment uses a linear regression model.

[0047] Figure 1This is a schematic diagram of the GOALDeR method of this invention. Within the framework of the GOAL method, it is implemented in three steps. First, a GPS model is constructed, and dimensionality reduction is achieved using outcome-adaptive LASSO, with penalty weights independent of the outcome model. Then, based on the selected variables, the DCOWs method is used to calculate the equilibrium weights, and the optimal values ​​of the adjustment parameters are determined using the minimum DWDC criterion, thus achieving variable selection for causal inference. Finally, the SL method is used to fit the outcome model, and the dose-response curve is estimated using a DR estimator combined with the equilibrium weights calculated in the previous step.

[0048] Figures 2 to 7 This invention demonstrates the performance of screening correction covariates under different simulation scenarios. Figure 2 When the outcome model and GPS model are linear and confounding variables are strongly correlated with the outcome and exposure factors (SoSt), the GOALDeR method selects confounding variables at different sample sizes (N=200, 500, 1000; p=20). X 1 and X 2 ), prognostic covariates ( X 3 and X 4 ), instrumental variables X 5 and X 6 Frequency distribution plots of spurious covariates; Figure 3 When the outcome model and GPS model are linear and the correlation between confounding variables and the outcome variable is greater than their correlation with the exposure factor (SoWt), the GOALDeR method selects confounding variables at different sample sizes (N=200, 500, 1000; p=20). X 1 and X 2 ), prognostic covariates ( X 3 and X 4 ), instrumental variables X 5 and X 6 Frequency distribution plots of spurious covariates; Figure 4 When the outcome model and GPS model are linear and the association between confounding variables and health outcomes is weaker than their association with the exposure factors (WoSt), the GOALDeR method selects confounding variables at different sample sizes (N=200, 500, 1000; p=20). X 1 and X 2 ), prognostic covariates ( X3 and X 4 ), instrumental variables X 5 and X 6 (and frequency distribution plots of spurious covariates). Figure 5 This indicates that when the outcome model is correctly specified (linear) but the GPS model is misspecificated (CoMt), the GOALDeR method selects confounding variables under different sample sizes (N=200, 500, 1000; p=20). X 1 , X 2 、X 3 and X 4 ), prognostic covariates ( X 5 and X 6 ), instrumental variables X 7 and X 8 Frequency distribution plots of spurious covariates; Figure 6 This indicates that when the outcome model is misspecificated but the GPS model is correctly specified (MoCt), the GOALDeR method selects confounding variables under different sample sizes (N=200, 500, 1000; p=20). X 1 , X 2 、X 3 and X 4 ), prognostic covariates ( X 5 and X 6 ), instrumental variables X 7 and X 8 Frequency distribution plots of spurious covariates; Figure 7 When both the outcome model and the GPS model are misspecificated (MoMt), the GOALDeR method selects confounding variables under different sample sizes (N=200, 500, 1000; p=20). X 1 , X 2 、X 3 and X 4 ), prognostic covariates ( X 5 and X6 ), instrumental variables X 7 and X 8 The frequency distribution plots of confounding variables and prognostic covariates were presented. The results show that when both the outcome model and the GPS model are linear, the GOALDeR method can accurately identify confounding variables and prognostic covariates, and this performance becomes more pronounced with increasing sample size; however, when model misspecification exists, the GOALDeR method performs poorly in variable selection, but this does not affect the accuracy and precision of the estimates (see [link to relevant documentation]). Figures 11 to 13 ).

[0049] Figures 8 to 13 This is a distribution graph of the causal parameter estimates under different simulation scenarios for the GOALDeR method of this invention, as well as the existing GOAL method and SL-DR method; the true value of the causal parameter is 2. Figure 8 The distribution plots of the causal parameters estimated by the GOALDeR, GOAL, and SL-DR methods under different sample sizes (N=200, 500) and covariate dimensions (P=100, 200), and different covariate correlation structures (ρ=0, 0.2, 0.5), indicating that the outcome model and GPS model are linear and the confounding variables are strongly correlated with the outcome and exposure factors (SoSt); Figure 9 The distribution plots of the causal parameters estimated by the GOALDeR, GOAL, and SL-DR methods under different sample sizes (N=200, 500) and covariate dimensions (P=100, 200), and different covariate correlation structures (ρ=0, 0.2, 0.5), indicating that the outcome model and GPS model are linear and the correlation between confounding variables and the outcome variable is greater than that with the exposure factors (SoWt). Figure 10 The distribution plots of causal parameters estimated by the GOALDeR, GOAL, and SL-DR methods under different sample sizes (N=200, 500) and covariate dimensions (P=100, 200), and different covariate correlation structures (ρ=0, 0.2, 0.5), indicating that the outcome model and GPS model are linear and the correlation of confounding variables with the outcome is weaker than with the exposure factors (WoSt); Figure 11 The distribution plots of the causal parameters estimated by the GOALDeR, GOAL, and SL-DR methods under different combinations of sample sizes (N=200, 500) and covariate dimensions (P=100, 200), and different covariate correlation structures (ρ=0, 0.2, 0.5), indicating that the outcome model is correctly specified (linear) but the GPS model is misspecificated (CoMt), the distribution plots of the causal parameters estimated by the GOALDeR, GOAL, and SL-DR methods are shown. Figure 12The distribution plots of the causal parameters estimated by the GOALDeR, GOAL, and SL-DR methods under different combinations of sample sizes (N=200, 500) and covariate dimensions (P=100, 200), and different covariate correlation structures (ρ=0, 0.2, 0.5), when the outcome model is misspecificated but the GPS model is correctly specified (MoCt), are shown. Figure 13 The figures represent the distributions of causal parameters estimated by the GOALDeR, GOAL, and SL-DR methods under different combinations of sample sizes (N=200, 500) and covariate dimensions (P=100, 200), and different covariate correlation structures (ρ=0, 0.2, 0.5) when both the outcome model and the GPS model are misspecificated (MoMt). The results show that the GOALDeR method has high estimation accuracy and precision, exhibiting a double robustness property: as long as either the outcome model or the exposure model is correctly specified, relatively ideal causal parameter estimates can be obtained. Furthermore, the GOALDeR method is less affected by the covariate correlation structure and the n / p ratio, compensating for the shortcomings of the GOAL method. Compared with the SL-DR method, the GOALDeR method has higher estimation precision. Moreover, when the outcome model is correctly specified but the GPS model is misspecificated, the SL-DR method's estimation accuracy and precision are significantly inferior to the GOALDeR method.

[0050] The following is a specific embodiment to further illustrate the above technical solution:

[0051] This study illustrates the application of the GOALDeR method by exploring the dose-response relationship between accelerated epigenetic aging in the temporal cortex (TC) region of the human brain and Alzheimer's disease (AD). Data was obtained from the GEO (Gene Expression Omnibus) database, including five datasets and a total of 861 participants, of whom 489 were AD patients. Methylated age, an indicator of biological age, is calculated based on methylation data. Accelerated epigenetic aging is defined as the residual between the methylated age and actual age regression models; a residual greater than 0 indicates accelerated aging, while a residual less than 0 indicates delayed aging. The corrected covariate set considered known age, sex, and neuron proportions, while the latent covariate set considered genome-wide methylation sites. Preliminary dimensionality reduction was performed using EWAS-meta-analysis. Results showed that the weighted sample balance was improved, as... Figure 14 As shown in Table 1, there is no statistically significant linear dose-response relationship between apparent aging acceleration and AD status.

[0052] Table 1

[0053]

Claims

1. A method for b-robust causal dose-response curve estimation based on high-dimensional covariates, characterized by, Comprising the following steps: 1) Construct a new objective function to achieve dimension reduction based on the modified adaptive LASSO method; Construct a GPS model, and use the modified adaptive LASSO method to screen the covariates that need to be balanced or included in the GPS model; Specifically, the GPS model is represented as: (1) The objective function is: (2) wherein, , , γ > 1, denotes a penalty weight, dcor ( X j ,Y ︱ T ) denotes the conditional distance correlation coefficient between the covariate X j and the health outcome given the exposure factor, ẑ is a constant used to normalize the coefficient; λ n > 0 is an adjustment parameter, when λ n satisfies and , the objective function (2) can screen out the confounding variables and prognostic covariates with probability 1; therefore, a set of candidate and satisfying λ n is set, and a set of candidate covariates is screened out according to λ n ; 2) Constructing double-weighted distance correlation coefficient (DWDC) to select the optimal λ n ; The unbiased estimation of dose-response curve by GPS method depends on the balance degree of covariate distribution among different exposure levels, that is, there is no arbitrary correlation between covariate and exposure factor, including linear and nonlinear. Based on this, the distance correlation coefficient of arbitrary correlation between measurable variables is used as the evaluation index of balance to construct DWDC, and the optimal λ n : (3) wherein, ŵ λn is the equilibrium weight estimated by the DCOWs method; 3) Estimate the dose-response curve using the DR estimator; First, calculate the pseudo outcome, and the calculation formula is as follows: (4) wherein, w i denotes the balancing weights, which are based on the optimal λ n The selected covariates are estimated using the DCOWs method; denotes the predicted value of the outcome variable, which is estimated using the SL method; subsequently, a univariate linear or nonlinear model is constructed using the calculated pseudo-outcome and the treatment factor to estimate the dose-response curve.

2. The method of claim 1, wherein the method is based on high-dimensional arguments. In step (1), the GPS model constructed is a model with exposure factor T as the dependent variable.

Citation Information

Patent Citations

  • System for enhancing data quality of dispense data sets

    CN112955968A

  • Construction method of kidney transplantation anti-infection drug dosage prediction model

    CN113035369A