An intervention screening method and system based on bidirectional causal effect estimation
Patent Information
- Application Number
- CN202311329688.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-13
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-10-13
AI Technical Summary
[0004]鉴于上述的分析,本发明实施例旨在提供一种基于双向因果效应估计的干预筛查方法和系统,用以解决现有筛查和干预技术中缺乏考虑健康功能和行为活动之间的异质性循环因果的问题
[0047]与现有技术相比,本发明通过收集不同个体的基线协变量数据以及每个方向上的干预前结局数据、干预措施属性和结局数据构建数据集,基于潜在结果框架分方向计算第三方变量条件下的异质性因果效应,从而可以区分个体间的异质性循环因果效应,基于潜在结果框架可以描述和评估复杂非线性因果效果,为提供细致准确的筛查和干预提供基础,并且不需要人为指定参数,避免关键参数指定错误带来的误导性放入干预和筛查,有助于提升干预和筛查的科学性和合理性。
Smart Images

Figure CN117275731B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of health intervention assessment technology, and in particular to an intervention screening method and system based on bidirectional causal effect estimation. Background Technology
[0002] There is a strong correlation between an individual's health functions (including cognition, motor skills, nutrition, vision, hearing, and mental health) and behavioral activities (including physical and social activities). However, interventions targeting behavioral activities often fail to achieve the desired results. Furthermore, screening programs for declining health functions do not yet consider behavioral activities as a key indicator. One reason for this is that the design and evaluation of screening and intervention programs neglect the circular causal relationship between health functions and behavioral activities. Specifically, behavioral activities also influence an individual's health functions, and conversely, an individual's health functions influence their behavioral activities, thus forming a two-way circular causal relationship. For example, participating in physical activity can help improve physical health functions such as cardiovascular function, muscle strength, and flexibility, while good physical health functions facilitate participation in physical activities; similarly, participating in social activities helps maintain and promote an individual's cognitive health functions, while good cognitive health functions facilitate participation in social activities.
[0003] Existing techniques for studying circular causality are primarily based on panel data models, making them difficult to apply to the evaluation and design of personalized screening and intervention programs. First, existing techniques cannot identify and estimate heterogeneous causal effects. Circular causal relationships between health function and behavioral activities often exhibit individual variability; however, existing methods struggle to capture and reflect this individual variability. Second, existing techniques lack universality. Health function and behavioral activities typically exhibit non-linear causal relationships, while existing methods are only applicable to linear function models, unsuitable for analyzing such non-linear relationships. Furthermore, existing techniques suffer from poor operability and reliability in practical applications. In practical use, existing techniques require specifying the model's time lag parameter; however, this parameter cannot be obtained through scientifically reliable means, easily leading to errors in program design and evaluation. Summary of the Invention
[0004] In view of the above analysis, the embodiments of the present invention aim to provide an intervention screening method and system based on bidirectional causal effect estimation to solve the problem that existing screening and intervention technologies lack consideration of heterogeneous circular causality between health functions and behavioral activities.
[0005] On one hand, embodiments of the present invention provide an intervention screening method based on bidirectional causal effect estimation, comprising the following steps:
[0006] Baseline covariate data on health function and behavioral activities of different individuals, as well as pre-intervention outcome data, intervention attributes and outcome data in each direction, were collected as sample sets for each direction;
[0007] The variables of interest in the baseline covariates are extracted as third-party variables. Based on the potential outcome framework, the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction are estimated according to the sample set in the target direction.
[0008] Based on the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction, intervention screening measures based on third-party variables are determined.
[0009] Based on further improvements to the above methods, the heterogeneous causal effects between health function and behavioral activities under third-party variable conditions in the target direction are calculated using a potential outcome framework based on a sample set in the target direction, including:
[0010] Fill in the data on unobserved potential outcomes in different directions;
[0011] Based on the imputed sample set, nonparametric regression is used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of a third-party variable in the target direction.
[0012] Based on further improvements to the above method, data on unobserved potential outcomes are filled in from different directions, including:
[0013] For each individual with missing potential outcome data in the target direction, M similar individuals are found based on their intervention attributes, and the imputation value of the missing potential outcome data for that individual is calculated based on the outcome data of the M individuals.
[0014] Based on further improvements to the above method, M individuals similar to it are identified according to the attributes of its intervention measures, including:
[0015] If the current individual's intervention attribute is intervention, then find the M most similar individuals in the non-intervention group; if the current individual's intervention attribute is no intervention, then find the M most similar individuals in the intervention group.
[0016] Using formula Calculate the similarity Sim(i,j) between the i-th individual and the j-th individual; where X i Let X represent the baseline covariate for the i-th individual. j Let the baseline covariate of the j-th individual be denoted as . This represents the pre-intervention outcome data for the i-th individual in the target direction. This represents the pre-intervention outcome data for the j-th individual in the target direction.
[0017] Based on further improvements to the above methods, nonparametric regression is used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of a third-party variable in the target direction, based on the imputed sample set. This includes:
[0018] If the third-party variable is a continuous variable, the heterogeneous causal effect estimator is calculated according to the following formula:
[0019]
[0020] If the third-party variable is a discrete variable, then according to the formula...
[0021]
[0022] Calculate the causal effect estimator for heterogeneity;
[0023] in, Z represents the heterogeneous causal effect estimator, z represents the third-party variable, and Z represents the third-party variable. i The value of the third-party variable representing the i-th individual. This represents the potential outcome data for the intervention attribute of the i-th individual. Let K(·) represent the potential outcome data of the i-th individual when the intervention attribute is no intervention, K(·) represent the kernel function, h represent the window width of the kernel function, N represent the number of individuals in the sample set, and I{·} represent the indicator function.
[0024] The standard deviation and confidence interval of the heterogeneous causal effect estimator are calculated using a subsampling method.
[0025] Based on further improvements to the above method, a subsampling method is used to calculate the standard deviation and confidence interval of the heterogeneous causal effect estimator, including:
[0026] Samples are randomly drawn from the samples corresponding to each intervention attribute in the sample set to form a subsample set, and B subsample sets are constructed.
[0027] For each subsample set, calculate the heterogeneous causal effect estimator between health function and behavioral activity under the third-party variable condition in the target direction;
[0028] Calculate the standard deviation of the heterogeneous causal effect estimators corresponding to the B subsets to obtain the standard deviation and confidence interval of the heterogeneous causal effect estimators.
[0029] Based on further improvements to the above methods, intervention screening measures based on third-party variables are determined according to the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction. These measures include:
[0030] Determine whether the causal effect corresponding to the value of the third variable is significant based on the confidence interval of the heterogeneous causal effect estimator;
[0031] Intervention or screening should be conducted on individuals corresponding to third-party variable values that have significant causal effects.
[0032] On the other hand, embodiments of the present invention provide an intervention screening system based on bidirectional causal effect estimation, comprising the following modules:
[0033] The sample construction module is used to collect baseline covariate data corresponding to the health function and behavioral activities of different individuals, as well as pre-intervention outcome data, intervention attributes and outcome data in each direction, as a sample set in each direction;
[0034] The heterogeneous causal effect calculation module extracts the variables of interest from the baseline covariates as third-party variables, and estimates the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction based on the sample set in the target direction according to the potential outcome framework.
[0035] The intervention screening module is used to determine intervention screening measures based on third-party variables, considering the heterogeneous causal effects between health functions and behavioral activities under the condition of third-party variables in the target direction.
[0036] Based on further improvements to the above system, the heterogeneity causal effect calculation module, using a potential outcome framework, calculates the heterogeneity causal effect between health function and behavioral activity under the condition of a third-party variable in the target direction, including:
[0037] Fill in the data on unobserved potential outcomes in different directions;
[0038] Based on the imputed sample set, nonparametric regression is used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of a third-party variable in the target direction.
[0039] Based on further improvements to the above system, nonparametric regression is used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of a third-party variable in the target direction, based on the imputed sample set. This includes:
[0040] If the third-party variable is a continuous variable, the heterogeneous causal effect estimator is calculated according to the following formula:
[0041]
[0042] If the third-party variable is a discrete variable, then according to the formula...
[0043]
[0044] Calculate the causal effect estimator for heterogeneity;
[0045] in, Z represents the heterogeneous causal effect estimator, z represents the third-party variable, and Z represents the third-party variable. i The value of the third-party variable representing the i-th individual. This represents the potential outcome data for the intervention attribute of the i-th individual. Let K(·) represent the potential outcome data of the i-th individual when the intervention attribute is no intervention, K(·) represent the kernel function, h represent the window width of the kernel function, N represent the number of individuals in the sample set, and I{·} represent the indicator function.
[0046] The standard deviation and confidence interval of the heterogeneous causal effect estimator are calculated using a subsampling method.
[0047] Compared with existing technologies, this invention constructs a dataset by collecting baseline covariate data of different individuals, as well as pre-intervention outcome data, intervention attributes, and outcome data in each direction. Based on the potential outcome framework, it calculates heterogeneous causal effects under third-party variable conditions in each direction, thereby distinguishing heterogeneous circular causal effects among individuals. The potential outcome framework can describe and evaluate complex nonlinear causal effects, providing a foundation for providing detailed and accurate screening and intervention. Furthermore, it does not require manual parameter specification, avoiding misleading inclusion of interventions and screenings due to incorrect specification of key parameters, and helps to improve the scientificity and rationality of interventions and screenings.
[0048] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0049] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.
[0050] Figure 1 This is a flowchart of an intervention screening method based on bidirectional causal effect estimation, as described in an embodiment of the present invention.
[0051] Figure 2 This is a schematic diagram of the simulation results of an embodiment of the present invention;
[0052] Figure 3 This is a block diagram of an intervention screening system based on bidirectional causal effect estimation, according to an embodiment of the present invention. Detailed Implementation
[0053] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0054] A specific embodiment of the present invention discloses an intervention screening method based on bidirectional causal effect estimation, such as... Figure 1 As shown, it includes the following steps:
[0055] S1. Collect baseline covariate data on health function and behavioral activities of different individuals, as well as pre-intervention outcome data, intervention attributes and outcome data in each direction as sample sets for each direction;
[0056] S2. Extract the variables of interest from the baseline covariates as third-party variables, and estimate the heterogeneous causal effects between health function and behavioral activities under the third-party variable condition in the target direction based on the potential outcome framework and the sample set in the target direction.
[0057] S3. Determine intervention screening measures based on third-party variables based on the heterogeneous causal effects between health functions and behavioral activities under the condition of third-party variables in the target direction.
[0058] By collecting baseline covariate data from different individuals, as well as pre-intervention outcome data, intervention attributes, and outcome data in each direction, a dataset is constructed. Based on the potential outcome framework, heterogeneous causal effects under third-party variable conditions are calculated in each direction. This allows for the differentiation of heterogeneous circular causal effects among individuals. The potential outcome framework can describe and evaluate complex nonlinear causal effects, providing a foundation for providing detailed and accurate screening and intervention. Furthermore, it eliminates the need for manually specifying parameters, avoiding misleading inclusion of interventions and screenings due to incorrect specification of key parameters, and contributing to the scientific rigor and rationality of interventions and screenings.
[0059] During implementation, the baseline covariate data for individuals include personal characteristics and health-related data such as age, gender, body mass index (BMI), and blood pressure at baseline.
[0060] In the causal relationship between health function and behavioral activity, pre-intervention outcome data represents the individual's behavioral activity at baseline, while outcome data represents the individual's behavioral activity during the follow-up period. Similarly, in the causal relationship between behavioral activity and health function, pre-intervention outcome data represents the individual's health function at baseline, while outcome data represents the individual's health function during the follow-up period. Intervention attributes include intervention and no intervention, using D... i D indicates i =1 indicates intervention, D i =0 indicates no intervention.
[0061] For example, in studying the circular causal effect between cognitive impairment and social activity, in the direction of the causal effect of cognitive impairment on social activity, the pre-intervention outcome data represents the individual's participation in social activities at baseline. In this scenario, intervention means having cognitive impairment (exposure to cognitive impairment), while no intervention means not having cognitive impairment (no exposure to cognitive impairment). The outcome data represents the individual's participation in social activities during the follow-up period. In the direction of the causal effect of social activity on cognitive impairment, the pre-intervention outcome data represents the individual's cognitive impairment at baseline. Intervention means participation in social activities, while no intervention means non-participation in social activities. The outcome data represents the individual's cognitive impairment during the follow-up period.
[0062] In practice, an individual's behavioral activity and health function can be assessed using a scale that evaluates the individual's behavioral activity level and health function level.
[0063] After establishing the sample set, the heterogeneous causal effects between health function and behavioral activities under the third-party variable condition in the target direction are calculated based on the potential outcome framework and the sample set in the target direction, including:
[0064] S21. Fill in the unobserved potential outcome data by direction;
[0065] S22. Based on the imputed sample set, nonparametric regression is used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction.
[0066] Each individual has two potential outcome data points in each direction, Y i (1) indicates D i =1 represents the outcome data at the time of intervention, Y i (0) indicates D i =0 represents the outcome data without intervention, while in actual observational data from the real world, only one scenario Y can be observed. i Y i =D i ·Y i (1)+(1-D i )·Y i (0), that is, only Y can be observed. i (0) or Y i One of (1). Therefore, it is necessary to fill in the unobserved potential outcome data. In the potential outcome framework, observed outcome data are directly filled into the corresponding potential outcome data, while the filling of unobserved potential outcome data requires data computation and then filling. For example, for the i-th individual, if several preventive measures attributes are "no intervention", then Y can be observed. i (0), the potential outcome data of the i-th individual without intervention. Potential outcome data at the time of intervention need to be further filled in through calculations. If the intervention attribute is "intervention," then Y is observed. i (1) Potential outcome data of the i-th individual at the time of intervention Potential outcome data without intervention Further calculations are needed to fill in the gaps.
[0067] Specifically, in step S21, the unobserved potential outcome data is filled in by direction, including:
[0068] For each individual with missing potential outcome data in the target direction, M similar individuals are found based on their intervention attributes, and the imputation value of the missing potential outcome data for that individual is calculated based on the outcome data of the M individuals.
[0069] Finding similar individuals based on their intervention attributes, that is, finding similar individuals based on whether the intervention attribute is intervention or no intervention, specifically including:
[0070] If the current individual's intervention attribute is intervention, then find the M most similar individuals in the non-intervention group; if the current individual's intervention attribute is no intervention, then find the M most similar individuals in the intervention group.
[0071] Using formula Calculate the similarity Sim(i,j) between the i-th individual and the j-th individual; where X i Let X represent the baseline covariate for the i-th individual. j Let the baseline covariate of the j-th individual be denoted as . This represents the pre-intervention outcome data for the i-th individual in the target direction. This represents the pre-intervention outcome data for the j-th individual in the target direction.
[0072] The measurement of inter-individual similarity is based on baseline covariate data and pre-intervention outcome data, thereby controlling for the effects of confounding factors and circular causal effects, controlling for the role of reverse causal effects, and avoiding bias in the estimation of causal effects.
[0073] After identifying the M individuals most similar to the current individual, the average of the outcome data from these M individuals is used as the imputation value for the missing potential outcome data of that individual. Therefore, the final outcome data for each individual is represented as follows:
[0074]
[0075] After the potential outcome data is filled in, it can be This is considered an estimate of the causal effect for each individual. When it is necessary to assess the causal effect in the target direction under the condition of the variable of interest, the variable of interest can be treated as a third-party variable. The heterogeneous causal effect between health function and behavioral activity under the condition of the third-party variable is then expressed as... Third-party variables can be one or more combinations of baseline covariates. Let Z represent the expectation, z represent a given third variable, and Z i Let represent the third-party variable for the i-th individual.
[0076] For example, when studying the impact of social activities on cognitive impairment, D i =1 indicates that the i-th individual participates in social activities, D i =0 indicates non-participation in social activities; Y i τ(z) represents the cognitive impairment status of individual i during the follow-up period; if z is age, then τ(z) represents the causal effect of participation in social activities on cognitive impairment at different age levels; if z is BMI, then τ(z) represents the causal effect of participation in social activities on cognitive impairment at different BMI levels; if z is gender, then τ(z) represents the causal effect of participation in social activities on cognitive impairment in different gender groups.
[0077] To ensure the identifiability of τ(z), we adopted the unconfounded assumption, common in causal inference. Under this assumption, potential outcomes and interventions are independent given baseline covariates and pre-intervention outcome variables, i.e.
[0078]
[0079] Where the symbol ⊥ represents independence, X i Let represent the baseline covariate for the i-th individual. Under this assumption, τ(z) is identifiable.
[0080] Specifically, based on the imputed sample set, nonparametric regression is used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of a third-party variable in the target direction, including:
[0081] If the third-party variable is a continuous variable, the heterogeneous causal effect estimator is calculated according to the following formula:
[0082]
[0083] If the third-party variable is a discrete variable, then according to the formula...
[0084]
[0085] Calculate the causal effect estimator for heterogeneity;
[0086] in, Z represents the heterogeneous causal effect estimator, where z represents a given third variable. i The value of the third-party variable representing the i-th individual. This represents the potential outcome data for the intervention attribute of the i-th individual. Let K(·) represent the potential outcome data for the i-th individual when the intervention attribute is no intervention. K(·) represents the symmetric kernel function, h represents the window width of the kernel function, I{·} represents the indicator function, and N represents the number of individuals in the subsample. It should be noted that the kernel function is any positive function satisfying ∫K(t)dt=1 and ∫K(t)tdt=0. Commonly used kernel functions include the Gaussian kernel. and Epanechnikov kernel K(t)=3(1-t) 2 )I{|t|≤1}.
[0087] The standard deviation and confidence interval of the heterogeneous causal effect estimator are calculated using a subsampling method.
[0088] In practice, the window width h can be automatically selected in a data-driven manner using the dpill function in the R package KernSmooth.
[0089] If z is a discrete variable, such as gender, for each subset, it can be calculated using the formula... Calculate the causal effect estimate in the target direction for males and females.
[0090] If z is a continuous variable, such as age, then for each subsample set, the formula calculates a function of the causal effect as a function of z.
[0091] For heterogeneous causal effect estimators, uncertainty inference is required, i.e., calculating the standard deviation and confidence interval. In practice, a subsampling method is used to calculate the standard deviation and confidence interval of the heterogeneous causal effect estimator, specifically including:
[0092] Samples are randomly drawn from the samples corresponding to each intervention attribute in the sample set to form a subsample set, and B subsample sets are constructed.
[0093] For each subsample set, calculate the heterogeneous causal effect estimator between health function and behavioral activity under the third-party variable condition in the target direction;
[0094] Calculate the standard deviation of the heterogeneous causal effect estimators corresponding to the B subsets to obtain the standard deviation and confidence interval of the heterogeneous causal effect estimators.
[0095] During implementation, the original sample set in the target direction includes samples with the attribute of intervention (i.e., intervention group) and samples with the attribute of no intervention (non-intervention group). The sample size of several pre-groups is N1, and the sample size of the non-intervention group is N0. Then, samples are randomly drawn from the intervention group. A sample was randomly selected from the non-intervention group. Each set of samples forms a subset. Construct B subsets in the same way. In practice, r can be 2 / 3, and B should be greater than 200.
[0096] For each subset of samples, calculate the heterogeneous causal effect estimate between health function and behavioral activity under the third-party variable condition in the target direction, denoted as . We obtain B heterogeneous causal effect estimators corresponding to B subsets. The standard deviations of these B heterogeneous causal effect estimators are denoted as . As an estimate of the standard deviation of the heterogeneous causal effect estimator, a 95% confidence level can be used in practice. The confidence interval of the heterogeneous causal effect estimator is expressed as follows:
[0097] After obtaining the confidence intervals of the heterogeneous causal effect estimator, intervention screening measures based on the third-party variable are determined based on the heterogeneous causal effect between health function and behavioral activity under the third-party variable condition in the target direction. These measures specifically include:
[0098] Determine whether the causal effect corresponding to the value of the third variable is significant based on the confidence interval of the heterogeneous causal effect estimator;
[0099] Intervention or screening should be conducted on individuals corresponding to third-party variable values that have significant causal effects.
[0100] By using confidence intervals to determine whether the causal effect corresponding to the value of a third-party variable is significant, interventions or screenings can be conducted on individuals with significant causal effects. This allows for the development of precise screening and intervention measures based on changes in the third-party variable to address heterogeneous causal effects.
[0101] For example, when studying the impact of social activities on cognitive impairment, if age is the third variable, then... This indicates the causal effect of participation in social activities on cognitive impairment at different age levels. Therefore, the significance of the causal effect at different age values can be determined by using the confidence interval of the heterogeneous causal effect estimator, i.e., by substituting the age values into the equation. In the calculation formula, we get Value, and then based on the confidence interval The presence or absence of a zero indicates the significance of the causal effect; if zero is present, the effect is not significant; otherwise, it is significant. Intervention is administered to age groups where the causal effect is significant; no intervention is administered to age groups where no significant effect is observed.
[0102] For example, when studying the impact of cognitive impairment on social activities, if age is the third variable, then τ(z) represents whether there is a causal effect of cognitive impairment on participation in social activities at different age levels. Therefore, screening measures can be established based on the obtained dataset to target age variations, achieving precise screening for cognitive impairment. For age groups where the effect is significant, social activity levels can be used for cognitive impairment screening; for age groups where the effect is not significant, social activity levels are not used for cognitive impairment screening.
[0103] It should be noted that the third-party variable can be a multivariate variable. For example, when the third-party variable is a multivariate variable consisting of age, gender, and BMI, with gender as a discrete variable and age and BMI as continuous variables, the causal effects of different combinations of age and BMI can be calculated separately for men and women to determine whether they are significant. Interventions or screenings can then be conducted for significant combinations. This allows for the development of precise intervention measures for social activities and precise screening measures for cognitive levels based on the changes in age, gender, and BMI across the three dimensions obtained from the dataset.
[0104] This invention incorporates pre-intervention outcome variables at baseline under the no-confusion assumption. This decouples circular causality into two non-circular causal relationships. This variable plays a crucial role in controlling for reverse causal effects; that is, when assessing the causal effect of an intervention on the outcome, it ensures that the pre-intervention outcome distribution falls within D. i =1 group and D i There was no significant difference between the 0 group and the 0 group. This method eliminates estimation bias caused by the influence of the outcome variable on the intervention. Furthermore, incorporating the pre-intervention outcome variable... It can also avoid estimation bias caused by some unobserved confounding factors.
[0105] Of particular note is that, throughout the estimation of τ(z), the intervention screening method based on bidirectional causal effect estimation proposed in this invention does not require any parametric model assumptions regarding the functional relationship between the outcome variable Y and the baseline covariate X or the third-party variable Z. All estimation steps are nonparametric, therefore the proposed estimation process is applicable to the construction of nonlinear causal relationship models. Furthermore, only a manually specified parameter M is required throughout the entire process. Research has demonstrated that the estimation... The convergence rate depends only on the choice of h, the dimension of z, and the dimension of X, and not on the choice of M. Extensive simulation experiments show that the estimation... The choice of parameter M is robust; furthermore, through sensitivity analysis, i.e., analyzing actual data under different values of M (e.g., setting M to 1, 2, ..., 10 respectively), the results show that the final estimate is robust. The changes are very small, so the final conclusion remains almost unchanged. In summary, theory, simulation, and modeling all verify that the bidirectional causal effect estimation of this invention is robust to M, does not depend on operator-specified key model parameters, thus eliminating the defects of operator-specified parameters and assisting in the design and evaluation of schemes in a more scientific and reliable way. It enables comprehensive and dynamic capture of causal relationships changing with environmental factors and is applicable to the description and evaluation of complex nonlinear causal relationships, enabling the formulation of scientific and personalized health screening and intervention strategies.
[0106] This embodiment provides experimental data from artificial simulation to illustrate that the present invention can accurately estimate heterogeneous causal effects in the target direction.
[0107] Assume a sample size of 2000. For each individual, the baseline covariate... It is a three-dimensional vector, where It follows a uniform distribution on the interval [-0.5, 0.5]. Each value is taken from the set {0, 1, 2} with equal probability. It follows a standard normal distribution. Let Y be the pre-intervention outcome data in one direction. pre , from the model ξ follows a standard normal distribution. The intervention attribute variable D follows the following logistic regression model.
[0108]
[0109] We set up three different simulation scenarios, where the heterogeneous causal effect τ(z) is a quadratic function, a polynomial function, and a composite function of z, respectively. Let... That is, the third-party variable is The data generation mechanisms for the potential outcomes in these three scenarios are as follows:
[0110] Scenario I: in ∈0 and ∈1 are independent of each other and both follow a standard normal distribution.
[0111] Scenario II: Where g(X,Y) pre The data generation mechanism for ∈0 and ∈1 is the same as in case I.
[0112] Scenario III: in ∈0 and ∈1 are independent of each other and both follow a standard normal distribution.
[0113] In the three simulation scenarios described above, for a certain value of the third variable z, the heterogeneous causal effect τ(z) is respectively 2z2 , z(1+2z) 2 (z-1) 2 And cos(2z)log(z+2)exp(z). It is worth noting that the resulting model has a relatively complex form, making it difficult to accurately specify these models, and parametric models often perform poorly. Our method, however, is nonparametric, avoiding the problem of model specification. 1000 repeated experiments were performed for each simulation study. Figure 2 The simulation results for scenarios I through III are presented. The solid line represents the true heterogeneous causal effect on z, and the dashed line represents the average of the heterogeneous causal effects on z estimated from 1000 simulations. The gray shaded area represents the 95% confidence interval estimated from the same 1000 simulations. In scenario I, most of the 95% confidence intervals cover the 0 point, indicating that the causal effect is not significant at different z levels. In scenario II, the causal effect is not significant when z is small; when z is large (greater than 0.25), the 95% confidence intervals are all above 0, indicating a significant causal effect, thus suggesting active intervention or screening for this population. In scenario III, only when z is in the middle range (between -0.3 and 0.4) are the 95% confidence intervals all above 0, indicating a significant causal effect, thus suggesting active intervention or screening for this population as well.
[0114] This invention provides an intervention screening system based on bidirectional causal effect estimation, such as... Figure 3 As shown, it includes the following modules:
[0115] The sample construction module is used to collect baseline covariate data corresponding to the health function and behavioral activities of different individuals, as well as pre-intervention outcome data, intervention attributes and outcome data in each direction, as a sample set in each direction;
[0116] The heterogeneous causal effect calculation module extracts the variables of interest from the baseline covariates as third-party variables, and estimates the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction based on the sample set in the target direction according to the potential outcome framework.
[0117] The intervention screening module is used to determine intervention screening measures based on third-party variables, considering the heterogeneous causal effects between health functions and behavioral activities under the condition of third-party variables in the target direction.
[0118] Preferably, the heterogeneity causal effect calculation module calculates the heterogeneity causal effect between health function and behavioral activity under the condition of a third-party variable in the target direction based on the potential outcome framework and the sample set in the target direction, including:
[0119] Fill in the data on unobserved potential outcomes in different directions;
[0120] Based on the imputed sample set, nonparametric regression is used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of a third-party variable in the target direction.
[0121] Preferably, based on the imputed sample set, nonparametric regression is used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of a third-party variable in the target direction, including:
[0122] If the third-party variable is a continuous variable, the heterogeneous causal effect estimator is calculated according to the following formula:
[0123]
[0124] If the third-party variable is a discrete variable, then according to the formula...
[0125]
[0126] Calculate the causal effect estimator for heterogeneity;
[0127] in, Z represents the heterogeneous causal effect estimator, z represents the third-party variable, and Z represents the third-party variable. i The value of the third-party variable representing the i-th individual. This represents the potential outcome data for the intervention attribute of the i-th individual. Let K(·) represent the potential outcome data of the i-th individual when the intervention attribute is no intervention, K(·) represent the kernel function, h represent the window width of the kernel function, N represent the number of individuals in the sample set, and I{·} represent the indicator function.
[0128] The standard deviation and confidence interval of the heterogeneous causal effect estimator are calculated using a subsampling method.
[0129] The above-described method and system embodiments are based on the same principles, and their related aspects can be referenced from each other to achieve the same technical effects. For specific implementation processes, please refer to the foregoing embodiments, which will not be repeated here.
[0130] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0131] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. An intervention screening method based on bidirectional causal effect estimation, characterized in that, Includes the following steps: Baseline covariate data on health function and behavioral activities of different individuals, as well as pre-intervention outcome data, intervention attributes and outcome data in each direction, were collected as sample sets for each direction; The variables of interest in the baseline covariates are extracted as third-party variables. Based on the potential outcome framework, the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction are estimated according to the sample set in the target direction. Based on the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction, intervention screening measures based on third-party variables are determined; Based on the potential outcomes framework, heterogeneous causal effects between health function and behavioral activities under third-party variable conditions in the target direction are calculated using a sample set in the target direction, including: Fill in the data on unobserved potential outcomes in different directions; Based on the imputed sample set, nonparametric regression was used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction; Based on the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction, intervention screening measures based on third-party variables are determined, including: Determine whether the causal effect corresponding to the value of the third variable is significant based on the confidence interval of the heterogeneous causal effect estimator; Intervention or screening should be conducted on individuals corresponding to third-party variable values that have significant causal effects.
2. The intervention screening method based on bidirectional causal effect estimation according to claim 1, characterized in that, Fill in the data for unobserved potential outcomes in different directions, including: For each individual with missing potential outcome data in the target direction, M similar individuals are found based on their intervention attributes, and the imputation value of the missing potential outcome data for that individual is calculated based on the outcome data of the M individuals.
3. The intervention screening method based on bidirectional causal effect estimation according to claim 2, characterized in that, Based on the attributes of its intervention, find M individuals similar to it, including: If the current individual's intervention attribute is intervention, then find the M most similar individuals in the non-intervention group; if the current individual's intervention attribute is no intervention, then find the M most similar individuals in the intervention group. Using formula Calculate the similarity between the i-th individual and the j-th individual. ;in, Let i represent the baseline covariates of the i-th individual. Let the baseline covariate of the j-th individual be denoted as . This represents the pre-intervention outcome data for the i-th individual in the target direction. This represents the pre-intervention outcome data for the j-th individual in the target direction.
4. The intervention screening method based on bidirectional causal effect estimation according to claim 1, characterized in that, Based on the imputed sample set, nonparametric regression is used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of a third-party variable in the target direction, including: If the third-party variable is a continuous variable, the heterogeneous causal effect estimator is calculated according to the following formula: If the third-party variable is a discrete variable, then according to the formula... Calculate the causal effect estimator for heterogeneity; in, Let z represent the estimator of heterogeneous causal effects, and z represent the third-party variable. The value of the third-party variable representing the i-th individual. This represents the potential outcome data for the intervention attribute of the i-th individual. Let K(•) represent the potential outcome data of the i-th individual when the intervention attribute is no intervention, where K(•) represents the kernel function, h represents the window width of the kernel function, N represents the number of individuals in the sample set, and I{•} represents the indicator function. The standard deviation and confidence interval of the heterogeneous causal effect estimator are calculated using a subsampling method.
5. The intervention screening method based on bidirectional causal effect estimation according to claim 4, characterized in that, The standard deviation and confidence interval of the heterogeneous causal effect estimator are calculated using a subsampling method, including: Samples are randomly drawn from the samples corresponding to each intervention attribute in the sample set to form a subsample set, and B subsample sets are constructed. For each subsample set, calculate the heterogeneous causal effect estimator between health function and behavioral activity under the third-party variable condition in the target direction; Calculate the standard deviation of the heterogeneous causal effect estimators corresponding to the B subsets to obtain the standard deviation and confidence interval of the heterogeneous causal effect estimators.
6. An intervention screening system based on bidirectional causal effect estimation, characterized in that, Includes the following modules: The sample construction module is used to collect baseline covariate data corresponding to the health function and behavioral activities of different individuals, as well as pre-intervention outcome data, intervention attributes and outcome data in each direction, as a sample set in each direction; The heterogeneous causal effect calculation module extracts the variables of interest from the baseline covariates as third-party variables, and estimates the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction based on the sample set in the target direction according to the potential outcome framework. The intervention screening module is used to determine intervention screening measures based on third-party variables, considering the heterogeneous causal effects between health functions and behavioral activities under the condition of third-party variables in the target direction. The heterogeneous causal effect calculation module, based on the potential outcome framework, calculates the heterogeneous causal effects between health function and behavioral activities under the condition of a third-party variable in the target direction, according to the sample set in the target direction, including: Fill in the data on unobserved potential outcomes in different directions; Based on the imputed sample set, nonparametric regression was used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction; Based on the heterogeneous causal effects between health function and behavioral activities under the condition of third-party variables in the target direction, intervention screening measures based on third-party variables are determined, including: Determine whether the causal effect corresponding to the value of the third variable is significant based on the confidence interval of the heterogeneous causal effect estimator; Intervention or screening should be conducted on individuals corresponding to third-party variable values that have significant causal effects.
7. The intervention screening system based on bidirectional causal effect estimation according to claim 6, characterized in that, Based on the imputed sample set, nonparametric regression is used to estimate the heterogeneous causal effects between health function and behavioral activities under the condition of a third-party variable in the target direction, including: If the third-party variable is a continuous variable, the heterogeneous causal effect estimator is calculated according to the following formula: If the third-party variable is a discrete variable, then according to the formula... Calculate the causal effect estimator for heterogeneity; in, Let z represent the estimator of heterogeneous causal effects, and z represent the third-party variable. The value of the third-party variable representing the i-th individual. This represents the potential outcome data for the intervention attribute of the i-th individual. Let K(•) represent the potential outcome data of the i-th individual when the intervention attribute is no intervention, where K(•) represents the kernel function, h represents the window width of the kernel function, N represents the number of individuals in the sample set, and I{•} represents the indicator function. The standard deviation and confidence interval of the heterogeneous causal effect estimator are calculated using a subsampling method.
Citation Information
Patent Citations
Disease prediction and early warning system based on causal network uncertainty reasoning
CN115862869A
Methods, systems, and articles of manufacture for the management and identification of causal knowledge
US20160292248A1