Intensity scoring algorithm integrator for high-dimensional control of hybrid variables and construction method of tendency scoring algorithm integrator
By combining Evalue, adaptive lasso regression and confounding variable balance generalized propensity score, a high-dimensional propensity score algorithm integrator is designed to control confounding variables, which solves the problem of mixed variable screening and control in high-dimensional data, and achieves the balance of robustness and efficiency, improving the accuracy and efficiency of causal inference.
Patent Information
- Application Number
- CN202411979544.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
The existing causal inference method is difficult to achieve both robustness and efficiency in high-dimensional data scenarios, especially in the selection and control of mixed variables.
By combining Evalue, adaptive lasso regression method and generalized propensity score of confounding variables, a propensity score algorithm integrator is designed for high-dimensional control of confounding variables. The integrator includes an input device, an Evalue estimator, an adaptive lasso regression variable filter, an equalization confounder variable filter, a weight calculator, and an inverse probability weighted dose-response function estimator, which is used to efficiently screen important confounders and build a robust causal inference model.
It realizes efficient screening of important confounding variables in a high-dimensional variable environment, ensures the robustness of the model and the balance of confounding variables, adapts to complex data scenarios, and improves the accuracy and efficiency of causal inference.
Smart Images

Figure CN119940535A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of causal inference, and more particularly to a propensity score algorithm integrator for controlling confounding variables in high dimensions and a construction method thereof. Background Art
[0002] In recent years, the Evaoue method proposed by Ding and VanderWeel provides a quantitative measurement tool for sensitivity analysis of unmeasured confounding variables. Evalue evaluates the impact of unmeasured confounding variables on causal effects by measuring the strength of association between unmeasured confounding variables and exposure and outcomes. The larger the Evalue value, the stronger the association of unmeasured confounding variables is needed to overturn the existing effect estimate. Although Evalue has been widely used in sensitivity analysis, it is not directly used for the screening of confounding variables and the modeling of causal effect estimation.
[0003] When the propensity score method draws causal conclusions, it relies on the assumption that there are no unmeasured confounding variables, which is an untestable assumption. It is generally believed that the more confounding variables included, the more reasonable it is. However, the causal effect estimate is sensitive to the confounding variables included in the propensity score model, and the unnecessary inclusion of confounding variables will lead to reduced accuracy and precision of the estimate. Although the inclusion of all confounding factors is crucial for unbiased estimation of causal effects, studies have shown that the inclusion of irrelevant variables may reduce the efficiency of the estimate. In other words, ignoring the characteristics of various confounding factors in the PS model will lead to estimation bias, while only including variables related to exposure will increase the variance of the estimate. Some studies have pointed out that the optimal propensity score model should include variables related to the outcome. Therefore, in the case of high-dimensional variables, in order to reduce bias, it is necessary to consider the inclusion of more confounding factors; at the same time, adjusting variables that are only related to the outcome can help improve the validity of the estimate. It is particularly important to adopt variable selection methods and include all confounding variables and variables related to the outcome in the propensity score model.
[0004] In high-dimensional data scenarios, the adaptive lasso regression method introduces regularization parameters to screen confounding variables. Its advantage is that it can effectively handle a large number of high-dimensional variables and select variables that are closely related to the results. In addition, this method evaluates the balance of confounding variables between the treatment group and the control group by calculating the weighted absolute mean difference, thereby controlling the confounding bias to a certain extent. However, the adaptive lasso regression method also has some limitations, mainly reflected in the failure to fully consider the impact of the strength of confounding variables on the causal effect estimation, and the lack of comprehensive evaluation of model sensitivity analysis.
[0005] At the same time, the propensity score model based on non-parametric settings focuses on maintaining the distribution balance of confounding variables through weight adjustment. This method shows a certain robustness in dealing with model misspecification problems, but when faced with high-dimensional variables, it lacks an effective variable screening mechanism and is difficult to cope with the screening needs of confounding variables in complex data structures.
[0006] In summary, existing causal inference methods still have certain limitations when facing problems such as model misspecification, confounding variable screening and control, especially in high-dimensional data scenarios, where it is difficult to achieve a balance between robustness and efficiency. Therefore, there is an urgent need for an integrated algorithm that can efficiently screen important confounding variables in high-dimensional variables while ensuring the robustness of the model and the balance of confounding variables. Summary of the invention
[0007] The purpose of the present invention is to provide a propensity score algorithm integrator for controlling high-dimensional confounding variables and a construction method thereof. By combining Evalue, adaptive lasso regression method and generalized propensity score for confounding variable balance, a new causal inference integrated algorithm in the context of high-dimensional confounding variables is developed.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] A propensity score algorithm integrator for controlling high-dimensional confounding variables, the propensity score algorithm integrator comprising: an input device, an Eva1ue estimator, an adaptive lasso regression variable filter, a balanced confounding variable filter, a weight calculator and an inverse probability weighted dose-response function estimator;
[0010] The input device is used to input the data required for the full outcome model before and after the removal of specific confounding variables;
[0011] The Evalue estimator is used to calculate the Evalue of the outcome full model before and after removing specific confounding variables, and calculate the high-dimensional confounding variable intensity measurement index;
[0012] The adaptive lasso regression variable filter is used to filter out the confounding variable set according to the penalty weight function constructed by the confounding intensity test index, and use it for causal inference;
[0013] The balanced confounding variable filter is used to determine the optimal adjustment parameters of the adaptive lasso regression based on the confounding intensity measurement index and the balanced weight of the confounding variables, and to filter out important confounding variables;
[0014] The weight calculator is used to calculate the robust weights of the screened important confounding variables;
[0015] The inverse probability weighted dose-response function estimator is used to construct a regression model based on robust weights to obtain the dose-response effect of the exposure variable on the outcome variable in the regression model.
[0016] The present invention also provides a method for constructing a propensity score algorithm integrator for controlling confounding variables in a high-dimensional manner, comprising the following steps:
[0017] S 1. In the Evalue estimator, compare the Evalue values of the full outcome model before and after removing specific confounding variables to obtain a measure of confounding intensity;
[0018] S2. Use the confounding intensity measurement index to construct a penalty weight function, and use the adaptive lasso regression variable filter to filter the confounding variables in causal inference;
[0019] S3. Under the condition of weak equilibrium, calculate the double weighted coefficient of the confounding intensity measurement index and the equilibrium weight of the confounding variable, and select the corresponding λ when the double weighted coefficient is taken n As the optimal adjustment parameter, to construct a balanced confounding variable filter;
[0020] S4. Screen out important confounding variables based on the optimal adjustment parameters, calculate robust weights, and construct a regression model to obtain the dose-response effect of the exposure variable on the outcome variable in the regression model;
[0021] S5. Integrate all the steps in S1-S4 through R shiny to build a propensity score algorithm integrator for high-dimensional control of confounding variables.
[0022] Furthermore, the full outcome model is specifically:
[0023] The regression model of the outcome variable Y on the exposure variable T and the confounding variable X;
[0024] S1, Eva1ue value before and after the full outcome model removes specific confounding variables, where the function expression of the full outcome model before removing specific confounding variables is:
[0025] Y=ηT+f full (X 1 , …X p )
[0026] In the above formula, η is the effect value of exposure, T is the exposure variable entered in the Evalue estimator, and f full (~) is a regression model that includes all confounding variables.
[0027] Furthermore, the full outcome model after removing specific confounding variables includes:
[0028] The functional expression of the regression model of the outcome variable Y on the exposure variable T and after removing specific confounding variables is:
[0029]
[0030] In the above formula, is the effect value of the exposure, T is the exposure variable entered in the Evalue estimator, X 1 , …, X j-1 , X j+1 , …X p is the confounding variable that needs to be entered in the Evalue estimator, f j (~) is to remove specific confounding variables X j The subsequent regression model.
[0031] Furthermore, S1 compares the Evalue values before and after removing specific confounding variables from the full outcome model to obtain the confounding intensity measurement index ΔEvalue j , expressed as:
[0032]
[0033] In the above formula, Evalue full and Evalue j They represent the Evalue before and after the specific confounding variables are removed from the full outcome model; η is the effect value of exposure, that is is the effect value of the exposure in the full model of the outcome including various confounding variables, To remove a specific confounding variable X j The effect size of the subsequent exposure on the outcome.
[0034] Furthermore, S2 uses the confounding intensity measurement index to construct a penalty weight function, and based on the adaptive lasso regression variable filter, implements the screening of confounding variables in causal inference, specifically:
[0035] The penalty weight function is constructed using the confounding intensity measure index, and the specific confounding variables that need to be included in the propensity score virtualizer are screened out based on the adaptive lasso regression variable filter;
[0036] Among them, given a specific confounding variable X j Under the condition of , the conditional expectation of the exposure variable T is:
[0037]
[0038] In the above formula, a 0 is the exposure variable and the specific confounding variable X j The intercept term in the full model relationship of the outcome, a jFor a specific confounding variable X j The effect size of
[0039] The expression of the objective function for screening confounding variables is:
[0040]
[0041] In the above formula, Among them, γ>1, j=1,…p, is the penalty weight function, λ n >0 is the adjustment parameter.
[0042] Furthermore, in S3, under the weak equilibrium condition, the double weighted coefficients of the mixed intensity measurement index and the balanced weights of the mixed variables are calculated, and the corresponding λ is selected when the double weighted coefficients are taken. n As the optimal adjustment parameter, it is expressed as:
[0043]
[0044] In the above formula, refers to the equilibrium weight of the confounding variable, |ΔEvalue| is the confounding intensity measurement index, T i is the exposure variable, X i is a confounding variable.
[0045] Furthermore, the method for obtaining the balancing weight is specifically as follows:
[0046] The propensity score method is used to take the empirical likelihood as the loss function, and the equilibrium weights are directly estimated by minimizing the loss function under the constraint that the marginal distributions of confounding variables and exposure variables remain unchanged before and after weighting.
[0047] Furthermore, in S4, important confounding variables are screened out based on the optimal adjustment parameters, and robust weights are calculated, wherein the relationship between the exposure variable and the important confounding variable is expressed as follows:
[0048] T=β 0 +β 1 X 1 +β 2 X 2 +…+β n X n
[0049] In the above formula, T is the exposure variable, β 0 is the intercept term of the regression model of exposure variables and confounding variables, β 1 , β 2 ,...β n is the effect value of the important confounding variables screened out, X n =(X 1 , X 2, ...) are the important confounding variables screened out;
[0050] The expression of S4, robust weight is:
[0051]
[0052] In the above formula, W is the robust weight, X n is an important confounding variable, T is the exposure variable;
[0053] The regression model is constructed to obtain the dose-response effect of the exposure variable on the outcome variable in the regression model, specifically:
[0054] The inverse probability weighted dose-response function estimator was used to construct a regression model of the exposure variable T on the outcome variable Y, and the dose-response function of the exposure variable T on the outcome variable Y was obtained by analysis.
[0055] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0056] (1) The present invention quantifies the sensitivity of the full outcome model to unmeasured confounding by introducing a high-dimensional confounding intensity measure, and screens confounding variables through a weight function so that important confounding variables can be effectively controlled.
[0057] (2) By combining the Evalue and balance double-weighted correlation coefficient indicators and considering the balance and confounding intensity of confounding variables at the same time, it is possible to more accurately identify important variables when screening confounding variables and avoid bias caused by missetting the full outcome model.
[0058] (3) Combining the advantages of the adaptive lasso regression method and the generalized propensity score for balancing confounding variables, the integrator not only demonstrates robustness to full model misspecification of outcomes in a high-dimensional confounding variable environment, but also maintains the ability to balance confounding variables, enabling it to adapt to more complex data scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0060] The following is a further description of the propensity score algorithm integrator and its construction method using high-dimensional control of confounding variables of the present invention in conjunction with the accompanying drawings;
[0061] Figure 1It is a structural schematic diagram of a propensity scoring algorithm integrator based on high-dimensional controlled confounding variables provided by the present invention. DETAILED DESCRIPTION
[0062] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0063] In order to better understand the purpose, structure and function of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings.
[0064] like Figure 1 As shown, the present invention provides a propensity score algorithm integrator for controlling confounding variables in high dimensions, comprising:
[0065] (1) The Evalue estimator, which is a measure of the intensity of confounding variables constructed based on Evalue, prepares for the construction of the adaptive lasso regression variable filter and the balanced confounding variable filter in the next step.
[0066] (2) Confounding variable selection is achieved based on adaptive lasso regression variable filter.
[0067] (3) At the same time, the balanced confounding variable filter selects the optimal adaptive lasso regression adjustment parameters based on the double-weighted correlation coefficient of Eva1ue and balanced weight to filter out confounding variables.
[0068] (4) Use the weight calculator and the inverse probability weighted dose-response function estimator to obtain the causal effect. Based on the optimal adjustment parameters, the confounding variables are screened and the weights are calculated. On this basis, the inverse probability weighted dose-response function estimator is used to construct a weighted linear or nonlinear (such as spline regression, local linear regression, etc.) regression model of exposure to outcome to obtain the association between exposure variables and outcome variables.
[0069] (5) The above algorithms are integrated through R language and shiny is used to construct an algorithm integrator for use on various computer platforms.
[0070] The present invention also provides the following specific embodiments:
[0071] S1. In the Evalue estimator, compare the Evalue values before and after removing the confounding variables in the full outcome model to obtain the confounding intensity measurement index;
[0072] It should be noted that the confounding variable strength measurement index is calculated based on the Evalue estimator through common computer input devices such as keyboard, mouse, and touch screen, in preparation for the construction of the penalty weight function later. The Evalue estimator outputs |ΔEvalue|. Specifically, Evalue is defined as the minimum strength of association between an unmeasured confound and the treatment and outcome on the risk ratio scale when it can fully explain a treatment-outcome association of a certain size after adjusting for measured confounding variables. The larger the E-value, the stronger the association strength (association with exposure and outcome) that the unmeasured confound needs to have to fully explain the observed association. This indicates that the effect of unmeasured confounding on the observed results may be smaller, and therefore the reliability of causal inference is higher. A small E-value means that the observed exposure-outcome association is more sensitive to the interference of unmeasured confounding.
[0073] S2. Use the confounding intensity measurement index to construct a penalty weight function, and use the adaptive lasso regression variable filter to filter the confounding variables in causal inference;
[0074] S3. Under the condition of weak equilibrium, calculate the double weighted coefficient of the confounding intensity measurement index and the equilibrium weight of the confounding variable, and select the corresponding λ when the double weighted coefficient is taken n As the optimal adjustment parameter, to construct a balanced confounding variable filter;
[0075] S4. Screen out important confounding variables based on the optimal adjustment parameters, calculate robust weights, and construct a regression model to obtain the dose-response effect of the exposure variable on the outcome variable in the regression model;
[0076] S5. Integrate all the steps in S1-S4 through R shiny to build a propensity score algorithm integrator for high-dimensional control of confounding variables.
[0077] The full model of the outcome is as follows:
[0078] The regression model of the outcome variable Y on the exposure variable T and the confounding variable X;
[0079] S1, Eva1ue value before and after the full outcome model removes specific confounding variables, where the function expression of the full outcome model before removing specific confounding variables is:
[0080] Y=ηT+f full (X 1 , …X p ) (1)
[0081] In the above formula, η is the effect value of exposure, T is the exposure variable entered in the Evalue estimator, and f full(~) is a regression model that includes all confounding variables.
[0082] The full outcome model after removing specific confounding variables includes:
[0083] The functional expression of the regression model of the outcome variable Y on the exposure variable T and after removing specific confounding variables is:
[0084]
[0085] In the above formula, is the effect value of the exposure, T is the exposure variable entered in the Evalue estimator, X 1 , …, X j-1 , X j+1 , …X p is the confounding variable that needs to be entered in the Evalue estimator, f j (~) is to remove specific confounding variables X j The subsequent regression model.
[0086] S1 compares the Evalue of the outcome model before and after removing specific confounding variables to obtain the confounding intensity measurement index ΔEvalue j , expressed as:
[0087]
[0088] In the above formula, Evalue full and Evalue j They represent the Evalue before and after the specific confounding variables are removed from the full outcome model; η is the effect value of exposure, that is is the effect value of the exposure in the full model of the outcome including various confounding variables, To remove a specific confounding variable X j The effect size of the subsequent exposure on the outcome.
[0089] It should be noted that: remove the X i The larger the change in η, the greater the difference in Evalue between the two models, and X j The importance of |ΔEvalue obtained by formula (3) j |Proportional to the value.
[0090] S2 uses the confounding intensity measurement index to construct a penalty weight function, and based on the adaptive lasso regression variable filter, implements the screening of confounding variables in causal inference, specifically:
[0091] The penalty weight function is constructed using the confounding intensity measure index, and the specific confounding variables that need to be included in the propensity score virtualizer are screened out based on the adaptive lasso regression variable filter;
[0092] Among them, given a specific confounding variable X j Under the condition of , the conditional expectation of the exposure variable T is:
[0093]
[0094] In the above formula, a 0 is the exposure variable and the specific confounding variable X j The intercept term in the full model relationship of the outcome, a j For a specific confounding variable X j The effect size of
[0095] The expression of the objective function for screening confounding variables is:
[0096]
[0097] In the above formula, Among them, γ>1, j=1,…p, is the penalty weight function, λ n >0 is the adjustment parameter.
[0098] It should be noted that the penalty weight function is constructed by using the ΔEvalue output in S1 through the computer input device keyboard, mouse, and touch screen, and the causal inference confounding variable screening is implemented based on the adaptive lasso regression variable filter. The purpose is to build the propensity score model R shiny virtualizer, and use the adaptive lasso regression variable filter to screen the confounding variables that need to be included in the propensity score virtualizer. Specifically, it is run in the adaptive lasso regression variable filter.
[0099] in, is the penalty weight function, whose size is inversely proportional to the intensity of the mixture, that is, X j (j=1, ...p) The greater the confounding intensity of the effect estimate, the smaller the penalty imposed by the model, and the more conducive it is for the confounding variable to enter the propensity score algorithm integrator. n >0 is the adjustment parameter. As with the lasso regression method, a set of parameters is set to satisfy and Alternative lambda for conditions n , each λ n Each corresponds to a set of candidate confounding variables.
[0100] S3, under the condition of weak equilibrium, calculates the double weighted coefficients of the mixed intensity measurement index and the balanced weight of the mixed variable, and selects the corresponding λ when the double weighted coefficient is taken. n As the optimal adjustment parameter, it is expressed as:
[0101]
[0102] In the above formula, refers to the equilibrium weight of the confounding variable, |ΔEvalue| is the confounding intensity measurement index, T i is the exposure variable, X i is a confounding variable.
[0103] It should be noted that the balanced confounding variable filter is based on the Evalue and balance-based dual weighted coefficient (EBDWC) to select the adjustment parameter λ n The optimal value of EBDWC. EBDWC is calculated by multiplying and summing two parts: the first part is the weighted correlation coefficient, which reflects the balance of each confounding variable at different exposure levels; the second part is the function of ΔEvalne, which reflects the confounding intensity of each confounding variable. The purpose of multiplying the two parts is to enhance the impact of the imbalance of confounding variables on the EBDWC value, and weaken the impact of the imbalance of other confounding variables on the EBDWC value. Select the λ corresponding to the minimum EBDWC value. n As the optimal adjustment parameter, it can balance the confounding variables as much as possible. The balancing weight in the first part of EBDWC is based on each λ n The selected set of candidate confounding variables is estimated using the propensity score method. The propensity score method uses the empirical likelihood as the loss function and directly estimates the equilibrium weights by minimizing the loss function under the constraint that the marginal distributions of the confounding variables and the exposure variables remain unchanged before and after weighting. It is robust to the misspecification of the propensity score model.
[0104] The method for obtaining the balancing weight is specifically as follows:
[0105] The propensity score method is used to take the empirical likelihood as the loss function, and the equilibrium weights are directly estimated by minimizing the loss function under the constraint that the marginal distributions of confounding variables and exposure variables remain unchanged before and after weighting.
[0106] The S4 screens out important confounding variables based on the optimal adjustment parameters and calculates robust weights, wherein the relationship between the exposure variable and the important confounding variable is expressed as:
[0107] T=β 0 +β 1 X 1 +β 2 X 2 +…+β n X n (7)
[0108] In the above formula, T is the exposure variable, β 0 is the intercept term of the regression model of exposure variables and confounding variables, β 1 , β 2 ,...β n is the effect value of the important confounding variables screened out, X 1 , X 2 , ..., X n Important confounding variables were screened out;
[0109] In formula (7), T is exposure, β 0 is the intercept term of the regression model of exposure variables and confounding variables, β 1 , β 2 ,...β n In order to screen out the effect value of important confounding variables, X 1 , X 2 , ..., X n To screen out important confounding variables. According to formula (7), T and X can be used to represent the probability observation value of exposure under given confounding variables. This observation value is the weight. At the same time, in order to improve statistical efficiency and obtain better confidence interval coverage, robust weights are calculated as follows:
[0110]
[0111] In the above formula, W is the robust weight, X n is the confounding variable, T is the exposure variable;
[0112] The regression model is constructed to obtain the dose-response effect of the exposure variable on the outcome variable in the regression model, specifically:
[0113] The inverse probability weighted dose-response function estimator was used to construct a regression model of the exposure variable T on the outcome variable Y, and the dose-response function of the exposure variable T on the outcome variable Y was obtained by analysis.
[0114] In formula (8), W is the robust weight, X is the confounding variable, and T is the exposure variable. Then, the sample size is added by weight, so that the sample size of the population with the same propensity score in the two groups is similar, achieving the effect of "quasi-randomization". Finally, the inverse probability weighted dose-response function estimator can obtain the dose-response function of the exposure variable to the outcome variable.
[0115] In summary, the present invention has developed a new causal inference integrated algorithm for processing confounding variables in the context of high-dimensional variables by innovatively combining three methods: Evalue, adaptive lasso regression method and generalized propensity score for confounding variable balance.
[0116] The algorithm integrator is proposed to solve the following key problems:
[0117] 1. Confounding control under high-dimensional confounding variables: The confounding intensity measure is introduced through ΔEvalue, which enhances the sensitivity of the model to unmeasured confounding, and the confounding variables are screened through the weight function, so that important confounding variables can be effectively controlled.
[0118] 2. Confounding variable screening and balance optimization: The new method considers the balance and confounding intensity of confounding variables at the same time based on the dual weighted correlation coefficient index of Evalue and balance, so as to more accurately identify important variables when screening confounding variables in the context of high-dimensional variables and avoid bias caused by model missetting.
[0119] The present invention also has the following technical effects:
[0120] 1. Evalue method and high-dimensional hybrid intensity measurement:
[0121] In the present invention, Evalue is no longer used only for sensitivity analysis, but is used to evaluate the contribution of each variable to unmeasured confounding. The ΔEvalue before and after removing a confounding variable from the outcome model is used to measure the confounding intensity of the confounding variable, and a penalty weight function is constructed. The weight is inversely proportional to the confounding intensity. The greater the confounding intensity, the smaller the corresponding penalty, which is conducive to the retention of important confounding variables in the subsequent propensity score model.
[0122] 2. Enhancement of adaptive lasso:
[0123] The adaptive lasso method uses regularization parameters to screen confounding variables, but its criteria rely on the balance index of confounding variables. The new method introduces EBDWC, which not only considers the balance of confounding variables at different exposure levels, but also considers the influence of confounding intensity. By selecting the optimal parameters through EBDWC, the confounding variables can be balanced as much as possible, which enhances the robustness of the model in the screening of confounding variables.
[0124] 3. Advantages of combined models:
[0125] The confounding variable balance propensity score model is good at handling continuous exposure variables and maintaining the balance of confounding variables. In the new method, the nonparametric confounding variable balance idea of the confounding variable balance propensity score model is combined with Evalue, which further enhances the robustness of the model in handling continuous exposure and complex confounding variables.
[0126] Therefore, the propensity score algorithm integrator for controlling confounding variables in the high-dimensional variable background based on Evalue is a new causal inference integrated algorithm in the high-dimensional variable background, which cleverly combines the characteristics of Evalue in evaluating the confounding intensity of each variable and the adaptive lasso regression method by selecting the optimal regularization parameter λ.n The ability to screen variables creatively combines the two to achieve effective screening of confounding variables. Furthermore, the integrator utilizes the confounding variable balancing ability of the generalized propensity score when processing continuous variables, and constructs a propensity score algorithm integrator based on Evalue to control confounding variables in the context of high-dimensional variables.
[0127] This integrated algorithm, through a progressive design, fully combines the respective advantages of E-value, adaptive lasso regression method and propensity score method for balanced confounding variables, and shows significant superiority in screening confounding variables and estimating causal effects. Specifically, on the one hand, this algorithm inherits the ability of adaptive lasso regression method to efficiently screen variables in high-dimensional data, and achieves accurate identification of confounding variables through flexible adjustment of penalty terms; on the other hand, it incorporates the propensity score method, using its unique advantages in balanced variable distribution, and effectively reduces the bias that may be caused by the missetting of the propensity score model.
[0128] At the same time, the introduction of E-value provides a quantitative basis for the robustness assessment of causal effects, further improving the reliability of result interpretation. Compared with a single method, the integrated algorithm overcomes the limitations of the adaptive lasso regression method and the propensity score method in design. For example, adaptive lasso regression may not perform well when dealing with nonlinear relationships or complex confounding structures, while the propensity score method may lead to increased bias when the model is misdefined. By organically combining the two, the algorithm retains the flexibility of variable screening and ensures the robustness of variable balance. In addition, the introduction of E-value enables the algorithm to quantitatively evaluate the potential impact of unmeasured confounding, providing a more comprehensive analytical framework for causal inference. Each part is optimized and forms an operational interface. At the same time, the entire process is integrated through R shiny to form a causal inference integrated algorithm, which can be used by various computers.
[0129] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A propensity score algorithm integrator for high-dimensional control of confounding variables, characterized in that: include: Importer, Evalue estimator, adaptive lasso regressor variable filter, balanced confounding variable filter, weight calculator, and inverse probability weighted dose-response function estimator; The input device is used to input the data required for the full outcome model before and after the removal of specific confounding variables; The Evalue estimator is used to calculate the Evalue of the outcome full model before and after removing specific confounding variables, and calculate the high-dimensional confounding variable intensity measurement index; The adaptive lasso regression variable filter is used to filter out the confounding variable set according to the penalty weight function constructed by the confounding intensity test index, and use it for causal inference; The balanced confounding variable filter is used to determine the optimal adjustment parameters of the adaptive lasso regression based on the confounding intensity measurement index and the balanced weight of the confounding variables, and to filter out important confounding variables; The weight calculator is used to calculate the robust weights of the screened important confounding variables; The inverse probability weighted dose-response function estimator is used to construct a regression model based on robust weights to obtain the dose-response effect of the exposure variable on the outcome variable in the regression model.
2. A method for constructing a propensity score algorithm integrator for controlling high-dimensional confounding variables, applied to the propensity score algorithm integrator for controlling high-dimensional confounding variables according to claim 1, characterized in that: The following steps are involved: S1. In the Evalue estimator, the Evalue values of the full outcome model before and after removing specific confounding variables are compared to obtain a measure of confounding intensity; S2. Use the confounding intensity measurement index to construct a penalty weight function, and use the adaptive lasso regression variable filter to filter the confounding variables in causal inference; S3. Under the condition of weak equilibrium, calculate the double weighted coefficient of the confounding intensity measurement index and the equilibrium weight of the confounding variable, and select the corresponding λ when the double weighted coefficient is taken n As the optimal adjustment parameter, to construct a balanced confounding variable filter; S4. Screen out important confounding variables based on the optimal adjustment parameters, calculate robust weights, and construct a regression model to obtain the dose-response effect of the exposure variable on the outcome variable in the regression model; S5. Integrate all the steps in S1-S4 through R shiny to build a propensity score algorithm integrator for high-dimensional control of confounding variables.
3. The method for constructing a propensity score algorithm integrator for controlling high-dimensional confounding variables according to claim 2, characterized in that: The full model of the outcome is as follows: The regression model of the outcome variable Y on the exposure variable T and the confounding variable X; S1, the Evalue of the outcome full model before and after removing specific confounding variables, where the function expression of the outcome full model before removing specific confounding variables is: Y=ηT+f full (X1,…X p ) In the above formula, η is the effect value of exposure, T is the exposure variable entered in the Evalue estimator, and f full (~) is a regression model that includes all confounding variables.
4. The method for constructing a propensity score algorithm integrator for controlling high-dimensional confounding variables according to claim 3, characterized in that: The full outcome model after removing specific confounding variables includes: The functional expression of the regression model of the outcome variable Y on the exposure variable T and after removing specific confounding variables is: In the above formula, is the effect value of the exposure, T is the exposure variable entered in the Evalue estimator, X1,…,X j-1 , X j+1 , …X p is the confounding variable that needs to be entered in the Evalue estimator, f j (~) is to remove the specific confounding variable X j The subsequent regression model.
5. The method for constructing a propensity score algorithm integrator for controlling high-dimensional confounding variables according to claim 4, characterized in that: S1 compares the Evalue of the outcome model before and after removing specific confounding variables to obtain the confounding intensity measurement index ΔEvalue j , expressed as: In the above formula, Evalue full and Evalue j They represent the Evalue before and after the specific confounding variables are removed from the full outcome model; η is the effect value of exposure, that is is the effect value of the exposure in the full model of the outcome including various confounding variables, To remove a specific confounding variable X j The effect size of the subsequent exposure on the outcome.
6. The method for constructing a propensity score algorithm integrator for controlling confounding variables in high dimensions according to claim 5, characterized in that: S2 uses the confounding intensity measurement index to construct a penalty weight function, and based on the adaptive lasso regression variable filter, implements the screening of confounding variables in causal inference, specifically: The penalty weight function is constructed using the confounding intensity measure index, and the specific confounding variables that need to be included in the propensity score virtualizer are screened out based on the adaptive lasso regression variable filter; Among them, given a specific confounding variable X j Under the condition of , the conditional expectation of the exposure variable T is: In the above formula, a0 is the exposure variable and the specific confounding variable X j The intercept term in the full model relationship of the outcome, a j For a specific confounding variable X j The effect size of The expression of the objective function for screening confounding variables is: In the above formula, Among them, γ>1, j=1,…p, is the penalty weight function, λ n >0 is the adjustment parameter.
7. The method for constructing a propensity score algorithm integrator for controlling high-dimensional confounding variables according to claim 6, characterized in that: S3, under the condition of weak equilibrium, calculates the double weighted coefficients of the mixed intensity measurement index and the balanced weight of the mixed variable, and selects the corresponding λ when the double weighted coefficient is taken. n As the optimal adjustment parameter, it is expressed as: In the above formula, refers to the equilibrium weight of the confounding variable, |ΔEvalue| is the confounding intensity measurement index, T i is the exposure variable, X i is a confounding variable.
8. The method for constructing a propensity score algorithm integrator for controlling high-dimensional confounding variables according to claim 6, characterized in that: The method for obtaining the balancing weight is specifically as follows: The propensity score method is used to take the empirical likelihood as the loss function, and the equilibrium weights are directly estimated by minimizing the loss function under the constraint that the marginal distributions of confounding variables and exposure variables remain unchanged before and after weighting.
9. The method for constructing a propensity score algorithm integrator for controlling confounding variables in high dimensions according to claim 7, characterized in that: In S4, important confounding variables are screened out based on the optimal adjustment parameters, and robust weights are calculated, wherein the relationship between the exposure variable and the important confounding variable is expressed as follows: T=β0+β1X1+β2X2+…+β n X n In the above formula, T is the exposure variable, β0 is the intercept term of the regression model of the exposure variable and the confounding variable, β1, β2, ...β n is the effect value of the important confounding variables screened out, X n =(X1, X2, ...) are the important confounding variables screened out; The expression of S4, robust weight is: In the above formula, W is the robust weight, X n is an important confounding variable, T is the exposure variable; The regression model is constructed to obtain the dose-response effect of the exposure variable on the outcome variable in the regression model, specifically: The inverse probability weighted dose-response function estimator was used to construct a regression model of the exposure variable T on the outcome variable Y, and the dose-response function of the exposure variable T on the outcome variable Y was obtained by analysis.