Causal inference for personalized decision making
The wTCP-DR method addresses the challenge of hidden confounders and imbalanced data in causal inference by learning density ratios, providing efficient and accurate confidence intervals for improved decision-making in high-stake applications.
Patent Information
- Application Number
- PCT/SG2024/050241
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-09
- Publication Date
- 2025-10-16
AI Technical Summary
Existing machine learning models for causal inference, particularly in high-stake applications, lack reliable uncertainty quantification and confidence intervals, especially when dealing with hidden confounders and imbalanced data sets, leading to potential inaccuracies in decision-making.
The proposed method, weighted Transductive Conformal Prediction with Density Ratio estimation (wTCP-DR), combines observational and interventional data to provide confidence intervals with guaranteed marginal coverage, even under the presence of hidden confounding, by learning the density ratio between distributions using probabilistic classification.
This approach offers more efficient and accurate confidence intervals, reducing the width of prediction intervals and ensuring reliable decision-making by accounting for distribution shifts and hidden confounders, thus enhancing the reliability of machine learning models in causal inference.
Smart Images

Figure SG2024050241_16102025_PF_FP_ABST
Abstract
Description
CAUSAL INFERENCE FOR PERSONALIZED DECISION MAKINGField
[0001] The disclosed example embodiments relate generally to machine learning and, more particularly, to a method, apparatus, device and computer readable storage medium for causal inference for personalized decision making.
[0002] In the growing area of machine learning for causal inference, various practical problems can be solved. For example, estimating the heterogeneous causal effects of an intervention (e.g., a medicine) on an important outcome (e.g., health status) of different individuals is a fundamental problem in a variety of influential research areas, including economics, healthcare and education This problem can be casted as estimating individual treatment effect (ITE). By developing machine learning models, the point estimate of ITE may be improved.
[0003] In a first aspect of the present disclosure, there is provided a method for model training. The method comprises: generating a first regression model at least based on a first observational dataset associated with a first treatment, the first observational dataset comprising a plurality of observed samples, an observed sample comprising observed feature information of an observed object and an observed outcome of the observed object, the first treatment being assigned to the object in the first observational dataset based on the observed feature information; determining respective conformality scores of the plurality of observed samples based on differences between predicted outcomes determined by the first regression model based on the observed feature information of the observed objects and observed outcomes in the first observational dataset; determining respective weights for the plurality of observed samples at least based on a first interventional dataset associated with the first treatment, the first interventional dataset comprising a plurality of intervened samples, an intervened sample comprising intervened feature information of an intervened object and an intervened outcome of the intervened object, the first treatment being randomly assigned to the intervened objects in the first intervened dataset; generating an iempirical distribution of the respective conformality scores at least based on the respective conformality scores and the respective weights for the plurality of observed samples; and for a test object, determining a first confidence interval for a predicted outcome of the test object with the first treatment assigned based on the empirical distribution and test feature information of the test object.
[0004] In a second aspect of the present disclosure, there is provided an apparatus for model training. The apparatus comprises: a model generating module configured to generate a first regression model at least based on a first observational dataset associated with a first treatment, the first observational dataset comprising a plurality of observed samples, an observed sample comprising observed feature information of an observed object and an observed outcome of the observed object, the first treatment being assigned to the object in the first observational dataset based on the observed feature information; a conformality determining module configured to determine respective conformality scores of the plurality of observed samples based on differences between predicted outcomes determined by the first regression model based on the observed feature information of the observed objects and observed outcomes in the first observational dataset; a weight determining module configured to determine respective weights for the plurality of observed samples at least based on a first interventional dataset associated with the first treatment, the first interventional dataset comprising a plurality of intervened samples, an intervened sample comprising intervened feature information of an intervened object and an intervened outcome of the intervened object, the first treatment being randomly assigned to the intervened objects in the first intervened dataset; a distribution generating module configured to generate an empirical distribution of the respective conformality scores at least based on the respective conformality scores and the respective weights for the plurality of observed samples; and a confidence determining module configured to, for a test object, determine a first confidence interval for a predicted outcome of the test object with the first treatment assigned based on the empirical distribution and test feature information of the test object.
[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions executable by the at least one processing unit. The instructions, upon execution by the at least one processing unit, cause the deviceto perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program which, when executed by a processor, causes the method of the first aspect to be implemented.
[0007] It would be appreciated that the content described in the Summary section of the present invention is neither intended to identify key or essential features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily envisaged through the following description.Brief Description of the Drawings
[0008] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent in combination with the accompanying drawings and with reference to the following detailed description. In the drawings, the same or similar reference symbols refer to the same or similar elements, where:
[0009] FIG. 1 illustrates a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;
[0010] FIG 2 illustrates an example causal graph with hidden confounding;
[0011] FIG. 3 illustrates a schematic diagram of a causal inference system in accordance with some embodiments of the present disclosure;
[0012] FIG. 4 illustrates a flow chart of a process for causal inference for personalized decision making in accordance with some embodiments of the present disclosure;
[0013] FIG. 5 illustrates a block diagram of an apparatus for causal inference for personalized decision making according to some embodiments of the present disclosure; and
[0014] FIG 6 illustrates a block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented.Detailed Description
[0015] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the presentdisclosure are shown in the drawings, it would be appreciated that the present disclosure may be implemented in various forms and should not be interpreted as limited to the embodiments described herein. On the contrary, these embodiments are provided for a more thorough and complete understanding of the present disclosure. It would be appreciated that the drawings and embodiments of the present disclosure are only for the purpose of illustration and are not intended to limit the scope of protection of the present disclosure.
[0016] Tn the description of the embodiments of the present disclosure, the term "including" and similar terms would be appreciated as open inclusion, that is, "including but not limited to". The term "based on" would be appreciated as "at least partially based on". The term "one embodiment" or "the embodiment" would be appreciated as "at least one embodiment". The term "some embodiments" would be appreciated as "at least some embodiments". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the matching degree between various data. For example, the above matching degree can be obtained based on various technical solutions currently available and / or to be developed in the future.
[0017] It will be appreciated that the data involved in this technical proposal (including but not limited to the data itself, data acquisition or use) shall comply with the requirements of corresponding laws, regulations and relevant provisions.
[0018] It will be appreciated that before using the technical solution disclosed in each embodiment of the present disclosure, users should be informed of the type, the scope of use, the use scenario, etc. of the personal information involved in the present disclosure in an appropriate manner in accordance with relevant laws and regulations, and the user’s authorization should be obtained.
[0019] For example, in response to receiving an active request from a user, a prompt message is sent to the user to explicitly prompt the user that the operation requested operation by the user will need to obtain and use the user's personal information. Thus, users may select whether to provide personal information to the software or the hardware such as an electronic device, an application, a server or a storage medium that perform the operation of the technical solution of the present disclosure according to the prompt information.
[0020] As an optional but non-restrictive implementation, in response to receiving the user's active request, the method of sending prompt information to the user may be, forexample, a pop-up window in which prompt information may be presented in text. In addition, pop-up windows may also contain selection controls for users to choose “agree” or “disagree” to provide personal information to electronic devices.|0021| It will be appreciated that the above notification and acquisition of user authorization process are only schematic and do not limit the implementations of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0022] As used herein, the term "model" can learn a correlation between respective inputs and outputs from training data, so that a corresponding output can be generated for a given input after training is completed. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural networks model is an example of a deep learning-based model. As used herein, "model" may also be referred to as "machine learning model", "learning model", "machine learning network", or "learning network", and these terms are used interchangeably herein.
[0023] ‘ ‘Neural networks” are a type of machine learning network based on deep learning.Neural networks are capable of processing inputs and providing corresponding outputs, typically comprising input and output layers and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically comprise many hidden layers, thereby increasing the depth of the network. The layers of neural networks are sequentially connected so that the output of the previous layer is provided as input to the latter layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network comprises one or more nodes (also known as processing nodes or neurons), each of which processes input from the previous layer.
[0024] Usually, machine learning can roughly comprise three stages, namely training stage, test stage, and application stage (also known as inference stage). During the training stage, a given model can be trained using a large scale of training data, iteratively updating parameter values until the model can obtain consistent inference from the training data that meets the expected objective. Through the training, the model can be considered to learn the correlation between input and output (also known as input-to-output mapping) from thetraining data. The parameter values of the trained model are determined. In the test stage, test inputs are applied to the trained model to test whether the model can provide correct outputs, thereby determining the performance of the model. In the application stage, the model can be used to process actual inputs and determine corresponding outputs based on the parameter values obtained from training."
[0025] FIG. 1 illustrates a block diagram of an example environment 100 in which various embodiments of the present disclosure may be implemented. In the environment 100 of FIG. 1, a computer system 110 applies a causal inference model 105 to perform causal inference. The causal inference model 105 is configured to process feature information 112 of an object in question to generate a predicted outcome 114 of the object with a treatment assigned to the object.
[0026] The feature information input to the causal inference model 105, the concerned treatments, and the predicted outcome output from the causal inference model 105 may be designed according to the tasks to be performed.
[0027] In an example of causal inference for individual treatment effect (ITE) of healthcare, the feature information may include information related to a subject, the concerned treatments may include applying an intervention (e.g., a medicine) on the subject or not applying the intervention on the subject, and the predicted outcome is related to a health status of the subject with or without the intervention applied (e.g., represented as a score range).
[0028] In an example of causal inference for real-world recommendation, the feature information may include information related to a user-item pair, the concerned treatments may include whether the item is recommended to the user or not, and the predicted outcome is defined as the user’s rating (e.g., within a rating range).
[0029] In addition to the ITE and recommendation scenarios, there may be various of other scenarios where effects of different treatments are evaluated on individuals. Estimation of individual treatment effect has been the key for individual decision making in economics, healthcare, education, etc.
[0030] In FIG. 1 , the computer system 1 10 may include any computing system with computing capability, such as various computing devices / systems, terminal devices, servers,etc. Terminal devices may include any type of mobile terminals, fixed terminals, or portable terminals, including mobile phones, desktop computers, laptops, netbooks, tablets, media computers, multimedia tablets, or any combination of the aforementioned, including accessories and peripherals of these devices or any combination thereof. Servers include but are not limited to mainframe, edge computing nodes, computing devices in cloud environment, etc.
[0031] It should be understood that the structure and function of each element in the environment 100 is described for illustrative purposes only and does not imply any limitations on the scope of the present disclosure.
[0032] As mentioned above, in the growing area of machine learning for causal inference, estimating the heterogeneous causal effects of an intervention has been casted as estimating ITE and most existing work focuses on developing machine learning models to improve the point estimate of ITE. However, point estimates is not enough to ensure safe and reliable decision-making in high-stake applications where failures are costly or may endanger human lives, and hence uncertainty quantification and confidence intervals allow machine learning models to express confidence in the correctness of their predictions.
[0033] Pioneering work provides confidence intervals for ITEs through Bayesian machine learning models such as Bayesian Additive Regression Trees and Gaussian Process. These approaches are known to have asymptotic coverage guarantees (i.e. they require infinite number of samples) and depend on the specific choice of regression models. Thus, these approaches cannot be easily generalized to popular machine learning models for causal inference on various input data types, including but not limited to text and graphs.
[0034] Recently, built upon conformal prediction, a conformal prediction method is proposed for counterfactual outcomes and ITEs, which may provide confidence intervals with guaranteed marginal coverage in a model-agnostic fashion. This means that, given any machine learning model that estimates the potential outcomes under treatment, conformal prediction acts as a post-hoc wrapper that provides confidence intervals guaranteed to contain the ground truth of potential outcomes and ITEs above a specified probability under marginal distribution. Unfortunately, however, it requires the assumption of strong ignorability that excludes the possibility of hidden confounders, which cannot be verified given data and may be violated in many real -world applications. For example, the socio-economic status of a patient, which is likely to be unavailable due to privacy concerns, is a common unobserved confounding factor that affects both patient's access to treatment and one's health condition. Similarly, under the strong ignorability assumption, it is proposed to use meta-learners in conformal prediction of ITEs. Recently, hidden confounding is taken into consideration for conformal prediction of ITEs from a sensitivity analysis aspect. However, their method needs access to the upper and lower bounds of the density ratio between the observational distribution and the interventional distribution to characterize the covariate shift from observational to interventional distribution.
[0035] For ease of understanding, the following will describe the problem setting and provide necessary background on conformal prediction.
[0036] The standard potential outcome (PO) framework with a binary treatment (recommendation) is considered. Let T £ [0, 1} be the treatment indicator, be the observed feature information (e.g., observed covariates), andbe the observed outcome of interest. For each object / , let (Y / (0), Y;(l)) be the pair of potential outcomes under control of treatment 7=0 and treatment T=\, respectively. It is assumed that the data generating process satisfies the following widely used assumptions:- Consistency: Y, = Y,(7j), which means the observed outcome Y, is the same as the potential outcome Y,(7i) with the observed treatment Ti.- Positivity: () <which means that any object has a positive chance to get treated and controlled.
[0037] It is noted that strong ignorability is not assumed in the standard PO framework. That is, there might exist potential hidden confounding U that affects the treatment T and the outcome Y at the same time. FIG. 2 illustrates an example causal graph with hidden confounding, where X represents observed covariates, U represents hidden confounders, T represents the observed treatment, and Y represents the outcome. The directed edges denote causal relations and the bidirectional edge signifies possible correlation. It can be seen that if there is some hidden confounding other than the observed covariates that could affect the treatment and the outcome, causal inference modelling on the observed covariates may not lead to accurate treatment and outcome.
[0038] Under this framework, the joint distribution under intervention do(T=f) is &at f°robservational data is- Note that the difference between conditional distributionhidden confounding, and the difference between’s^ue t0intervention. Throughout this work, the notation of probability density (mass) functions is used instead of probability measures. A superscript / is used for interventional distribution and O is used for observational distribution. For a given treatment t E {0, 1}, it is assumed that there are n observational and m interventional samples:
[0039] Given a predetermined target coverage rate of 1 - a, the goal is to construct confidence interval (j? for potential outcome under treatment t at a new test sampleensures marginal coverage:where the probability is over
[0040] Conformal prediction (CP) is a distribution-free framework that provides finite- sample marginal coverage guarantees. Transductive and split CP are two approaches to conformal prediction and both are briefly introduced since they will be used below.
[0041] Given a dataset 25 ™ ConformalPrediction (SCP) starts by splittinginto two disjoint subsets: a training set ,23^ , and a calibration setThen, a regression estimator is trained on 25conformityscoresare computed forwhere typicallyThe empirical distribution of the conformity scores are defined as p ^ieconfidence interval for the targetsamplewhere Qp has been proved thatunder exchangeability of XXJ' SCP(Xn-i-i ) is guaranteed to satisfy marginal coverage. Furthermore, if ties between conformity scores occur with probability zero, then
[0042] Note that the upper bound ensures that the confidence interval is nonvacuuous, i.e., the interval width does not go to infinity.
[0043] Given a same dataset XJ*asabove, Transductive Conformal Prediction (TCP) takes a different approach by looping over all possible valuor I / £ J / , TCP first constructs an augmented datasetU (A i < i / j .Then, a regression estimator is trained onand the conformity scoresWith empirical distribution defined as he interval for the target sample -4^4-1 ishe same lower and upper bound guarantee as
[0044] TCP is computationally more expensive as it requires fitting 0 for every fixed€ J? . The discretization of J / comes as a tradeoff between computational costs and accuracy of the conformal interval. For these reasons, SCP is more widely used due to its simplicity, however, SCP is less sample efficient by splitting the dataset into a training set and a calibration set. Cross-conformal prediction may be used to improve efficiency for SCP.
[0045] When calibration and test data are independent yet not drawn from the same distribution, a weighted version of conformal prediction is proposed. Herein, a more specific setting of where the dataset is merged from two different distributions is discussed,test sampleA'n+mT-l is sampled from. Define the density ratio dP as are weighted exchangeable withweight functionsif ( X , I / )For y g J / , define the normalized weights py as:where the summations are taken over permutations £7 of L * * > ,4" W-F l.Here in Eq.(6), an abuse of notation thatis used for symmetry reason. With the conformity scores <x computed in the same way as TCP and the weighted empirical distribution of the conformity scores defined as ’’PfH-m-rl Ax? , the conformal interval for the target sample is:where The lower bound guarantee and upper bound areproven under extra assumptions. When m=0, pi becomeswhich is more commonly used in the literature. When m>l, the computational cost of pf is
[0046] To address the above limitations and provide confidence intervals that have finite- sample guarantees even without the strong ignorability assumption, embodiments of the present disclosure propose an improved solution, which is referred to as weighted Transductive Conformal Prediction with Density Ratio estimation (wTCP-DR) that is based on weighted transductive conformal prediction. With less restrictive assumptions, this solution needs access to both observational and a fraction of interventional data (e.g., data collected from randomized control trials). In contrast to the weighted conformal prediction method which uses propensity score as the reweighting function, this solution computes the reweighting function by learning the density ratio of the interventional and observational distribution using the data provided.
[0047] There are benefits of the proposed solution. First, it does not require strong ignorability assumption and provides a confidence interval with coverage guarantee even under the presence of confounding. Second, it works well under an imbalanced number of interventional and observational data, i.e., when interventional data is of smaller size than observational data due to the higher cost of collecting interventional data.
[0048] Example embodiments of the present disclosure will be descried in the following, which provide a confidence interval on counterfactual outcomes at an individual level with marginal coverage guarantee.
[0049] FIG. 3 illustrates a schematic diagram of a causal inference system 300 in accordance with some embodiments of the present disclosure. The causal inference system 300 may be implemented at the computer system 110 of FIG. 1, to implement the causal inference.
[0050] According to embodiments of the present disclosure, observational data and interventional data are both considered for conformal prediction with density ratio, with arestrictive assumption that hidden confounding exists.
[0051] Different treatments (e.g., T=0 and 7= 1 ) are considered separately for the casual inference. In FIG. 3, an observational dataset 310-1 and an interventional dataset 320-1 associated with a first treatment are used for casual inference associated with the first treatment. As will be described in the following, a regression model 305-1 is generated based on the observational dataset 310-1 (and probably based on the interventional dataset 320-1). Further, an observational dataset 310-2 and an interventional dataset 320-2 associated with a second treatment are used for casual inference associated with the second treatment. As will be described in the following, a regression model 305-2 is generated based on the observational dataset 310-2 (and probably based on the interventional dataset 320-2).
[0052] The observational dataset 310-1 comprises a plurality of observed samples, an observed sample comprising observed feature information of an observed object and an observed outcome of the observed object, the first treatment being assigned to the object in the first observational dataset based on the observed feature information. The interventional dataset 320-1 comprises a plurality of intervened samples, an intervened sample comprising intervened feature information of an intervened object and an intervened outcome of the intervened object, the first treatment being randomly assigned to the intervened objects in the first intervened dataset. Similarly, the observational dataset 310-2 comprises a plurality of observed samples, an observed sample comprising observed feature information of an observed object and an observed outcome of the observed object, the second treatment being assigned to the object in the second observational dataset based on the observed feature information. The interventional dataset 310-2 comprises a plurality of intervened samples, an intervened sample comprising intervened feature information of an intervened object and an intervened outcome of the intervened object, the second treatment being randomly assigned to the intervened obj ects in the second intervened dataset.
[0053] In an example, the first treatment (e.g., 7=1) may indicate applying a medical intervention (e.g., a medicine) on a subject and the second treatment (e.g., 7=0) may indicate not applying a medical intervention (e g., a medicine) on a subject (or vice versa). In this scenario, the observed samples in the observational dataset 310-1 or 310-2 may be collected from the subjects that are diagnosed as requiring a medicine or not based on the symptoms and other information related to the subjects. The intervened samples in the interventionaldataset 310-1 or 310-2 may be collected from the subjects in clinical trial of medicine.
[0054] In an example, the first treatment (e g., 7=1) may indicate recommending a specific item to a user and the second treatment (e.g., 7=0) may indicate not applying a specific item to a user. In this scenario, the observed samples in the observational dataset 310-1 or 310-2 may be collected from the users with or without the items recommended through existing recommendation strategies (which may be constructed by considering various information related to the user and item). The intervened samples in the interventional dataset 310-1 or 310-2 may be collected from the users that are randomly recommended or not recommended with the items in an A / B test for the item recommendation.
[0055] During causal inference, a test sample, e g., test feature information 302 of a test object is input to either or both of the regression models 305-1 and 305-2, to determine a confidence interval 330-1 for a predicted outcome of the test object with the first treatment assigned and / or a confidence interval 330-2 for a predicted outcome of the test object with the second treatment assigned.
[0056] In the following, for the purpose of discussion, the observational dataset 310-1 and observational dataset 310-2 are collectively or individually referred to as observational datasets 310; the interventional dataset 320-1 and interventional dataset 320-2 are collectively or individually referred to as interventional datasets 320; the regression model 305-1 and regression model305-2 are collectively or individually referred to as regression models 305; and the confidence interval 330-1 and confidence interval 330-2 are collectively or individually referred to as confidence intervals 330.
[0057] Since different treatments (e g., T=0 and 7"=1) are considered separately, T=t is fixed and the dependence is dropped on Tin Eq. (1) for simplicity of notations. Recall there are n observed samples and m intervened samples and a test sample is x„-m+iThe distributions of the observed samples, intervened samples, and the test sample are represented as follows:
[0058] A straightforward method (naive method) is first introduced. The method comprises: constructing a confidence interval for the potential outcome only from the interventional datausing standard split conformal prediction of Eq. (3) as the intervended samplescomes from the same distribution as the test sample X1n-f-m-r 1 '
[0059] The algorithm is detailed in Algorithm 1 (as shown in Table 1). From Eq. (4) it is known that
[0060] This approach may be inefficient because it completely ignores n observed samples in the observational data and typically n is larger than m.Table 1
[0061] In embodiments of the present disclosure, it is expected to combine both the interventional data and observational data. In this case, it is necessary to take distributionshift into consideration. Therefore, weighted conformal prediction of Eq. (7) is naturally suitable for such tasks, and the key challenge is to identify the normalized weights in Eq.(6), i.e., to identify the density ratio as:
[0062] However, it is still assumed thatequalsA', t) soI / ) is as simple as estimating the propensity score p(A: ) / p( .Y j f ) . When hidden confounding exists, propensity score is not enough to account for the distribution shift. It is proposed to learn the density ratiol / J from data, as detailed next.
[0063] The key to the weighted conformal prediction is the density ratio / '( X, I / ), and fortunately there exists a rich literature of density ratio estimation, including moment matching, probabilistic classification, and ratio matching. Since probabilistic classification using neural networks is more flexible and better exploits nonlinear relations in the data, so only probabilistic classification is introduced herein.
[0064] By assigning labels z=l to an observed sampleand assigning labels z=0 to an intervened samplea new dataset is constructed for learning the density ratio as follows.
[0065] For any nonlinear binary classification algorithm like logistic regression with nonlinear features, random forests or neural networks that output estimated probabilities of class membership p(z = 1 | x, t / ) and / ?( Z “ 0 j X, I / ), the density ratio may be approximated by: / H :?™ 1 1
[0066] Since is a constant and will cancel out when computing thenormalized weights in Eq. (6), f (.t . U ! ~ is denoted as the estimateddensity ratio, so the corresponding estimated normalized weights of Eq. (6) are:
[0067] Since Eq. (13) requirestimes of evaluating the density ratiowhich is computationally impractical for m>l, in some embodiments, only the observational data are used when computing the normalized weights (i.e. m=l) and use interventional data for computing the density ratio f As a result, the estimated normalized weights become:
[0068] With the above analysis of calculating the weights by combining both the the interventional data and observational data, the conformal casual inference according to the embodiments of the present disclosure can be described as follows.
[0069] This process involves generating a regression model 305 / / based on an observational dataset 3 10associated with a specific treatment (T=t, with t= 0 or 1), where an observed sample ( v ( \ ) includes observed feature information i y) of an observed object i and an observed outcome I / O. / of the observed object with the treatment T=t applied. The regression model p* may represent the distribution of the outcomes and the observed feature information of the objects.
[0070] Then the observed feature information of the observed objects in the observational dataset 310( xb , t / b' Ey, are input to the regression model 305, to obtain predicted outcomes determined by the regression model 305. Respectiveconformality scores of the plurality of observed samples based on differences between the predicted outcomes determined by the regression model 305 and the observed outcomes in the observational dataset 310, represented as .
[0071] Further, respective weightsfor the plurality of observed samples are computed at least based on the interventional dataset 320=> associated with the specific treatment where an intervened sample includesintervened feature information of an intervened object i and an intervened outcomeof the intervened object with the treatment T=t applied. The weights may becalculated as normalized weights according to Eq. (14) above. There are various ways to determine the normalized weights which will be described below.
[0072] Then an empirical distribution of the respective conformality scores at leastbased on the respective conformality scores and the respective weights for the pluralityof observed samples. A quantile valueof the empirical distribution p corresponding to a predetermined probability is determined through a quantile function The probability may be determined through a level a, and calculated as (1 -a). For example, a may be set as 10%, and then the predetermined probability for the quantile value is 90%. Of course, the predetermined probability may be set as any suitable value from 0 to 100% as required.
[0073] For the test sample (n+m+1), test feature information of the test object is input to the regression modelto obtain a predicted outcome of the test objectwith the specific treatment assigned. Then a confidence interval for the predicted outcome of the test object is determined based on the quantile valueand the predicted outcomeof the test object ■ The confidence interval indicates a margin coverageguarantee for the predicted outcome of the test object n( - x r . ) . n-hnH-I ■[0074| In some embodiments, to allow identify the confidence interval, a complete weighted transductive conformal prediction with density ratio estimation is proposed.|0075| Respective density ratiosfor the n observed samples and a test density ratiot / / ) with 7=«+ / M+1 for the test sample based on the interventional dataset 310and the observational dataset 320associated with a specific treatment. The density ratios may be estimated as discussed above.
[0076] A confidence range J / is divided into a plurality of confidence units, 0 € J / For example, if the confidence range J / is from 0 to 1, then the confidence units may be constructed as {0, 0.1 }, {0.1, 0.2}, {0.2, 0.3}, ..., {0.9. 1.0}. The specific values of the confidence units may depend on the confidence range, and the granularities of the confidence unit division may be configured as required.
[0077] For each ofthe plurality of confidence units, 1 / € J / , the following operations are performed iteratively:- Constructing an augmented datasetincluding the observational datasetand an additional test sampleconsisting of the test feature information and the confidence unit;- Generating a regression model 305based on the augmented dataset- Computing conformity scoresn an<^- Determining respective weightsfor the plurality of observed samples and a weightfor the test sample as in Eq. (14) (whereis replaced with I / );- Generating a weighted empirical distribution of conformity scores based on weighting the respective conformality scores with the respective weights- Determining the quantile value of the empirical distribution corresponding to a predetermined probability as F) •
[0078] Through the above iterative process, a plurality of quantile values^p may be obtained for the plurality of confidence units 1 / CThen for each of the plurality of first regression models, a test conformity scoreof the test sample is determined based on a difference between a predicted output determined by a regression modeland that is used to generate the regression model, — t) If the test conformity score is lower than or equal tothe quantile value for the regression model, i.e., -Sn-HyH-1> then the current confidence unit '( / is combined as a part of the confidence interval for the test object. The confidence interval is represented as™ 0 is initialized as an empty set.
[0079] Through the above, by traversing the plurality of confidence units, 1 / € M the suitable confidence units are combined to form the confidence interval for the test sample.
[0080] A complete description of the above iterative process is summarized as Algorithm 2 as shown in Table 2 below.Table 2
[0081] By using estimated normalized weights pj rather than the oracle normalized weights pj to reweight the empirical distribution of conformity scores p, the approach of the present disclosure introduces an extra source of error, as quantified below.
[0082] Proposition 1 . Under the assumptions thatareabsolutely continuous with each other and thatthenthe confidence interval (. * w7< ' / - / )A constructed from Algorithm 2 satisfieswhere c is a constant is theapproximation error of the density ratio.|0083| By comparing Eq. (15) and Eq. (9), it can be seen that when the oracle density ratio f ( I / }, i.e.,™ 0, has been accessed, then the approach of the present disclosure obtains a tighter upper bound than the naive method, as typically the number of observational data n is much larger than the number of interventional data m in causal inference, due to the higher cost of randomized controlled trails. Unfortunately, oracle density ratio r(x, j) is usually unavailable, and the estimation error of density ratio is of orderfor moment matching or ratio matching and of orderfor probabilistic classification It seems that the solution of the present disclosure (represented as wTCP-DR) has spent a huge amount of effort while achieving a worse result in the end.
[0084] However, the efficiency of conformal prediction methods is quantified by the width of the confidence interval, not by the difference between the probability upper and lower bound. An upper bound strictly lower than 1 guarantees that the confidence interval is not arbitrarily large, however there is no guarantee that a smaller upper bound results in a smaller confidence interval. Intuitively, the method has a smaller interval compared to the naive method, because the regression model |7 of the solution of the present disclosure is trained on n observational data while the regression modelof naive method is trained on m interventional data. Intuitively, there is a higher chance that the conformity scores of wTCP-DR are smaller than the conformity scores of the naive method, which means that the confidence intervalcalculated through Algorithm 2 is a smallerinterval than £, naive calculated through Algorithm 1. The above intuition in the following section is formalized for additive Gaussian noise model.
[0085] An additive Gaussian noise model is considered, which is a simple yet popular setting in causal inference. Recall that T=t is fixed and the dependence is dropped on T for simplicity of notations. Specifically, the following assumptions are made:Al Additive Gaussian noise,and t / ~), where <p represents the (learned) features of interventional and observational data;A2 Gaussian features.A3 Upper bounds on the difference between oracle density ratioI / ) and estimated density ratio f ( x, I / )'A4 Boundeddivergence between:se assumptions, the effect of hidden confounding is reflected from they { A , ( ) p^ (t / [ X, t) through the difference of (Aand is dependent of hidden confounding u whereasis independent of u due to intervention. Before showing our main theoretical result, the implications of these assumptions may be first discussed:Al It is assumed that interventional and observational data share the same feature <p, a commonly used setting in causal inference especially when <p is learned with neural networks. It is assumed the same noise scale for observational and interventional dataJ only for simplicity, which may be relaxed to the more general case thatand jp have different noise scalesKI This assumption is satisfied when either the features are designed to have Gaussian distribution, or the features are learned from wide enough neural networks,A3 This assumption requires that the error of density ratio estimation is upper bounded, and given that a is typically 0.1 or 0.05, this assumption is usually satisfied in practice; andA4 This assumption ensures thatand p^ share the same support over <V X J / , and is required such that the central limit theorem may be used in the proof.
[0087] The following will give the main theoretical result of the present disclosure.
[0088] THEOREM 1. Assume the above assumptions hold, with probability at least jobtained from Algorithm 2 will be smaller than the interval obtained fromAlgorithm 1 up towith d j , Og, dj, O4 being the following:where fA ~ represents the dissimilarity distance betweenand erfis the inverse error function,and Gx are constants that only depend onis the effective sample size defined as below:
[0089] The implications of Theorem 1 may be summarized as below.(1) dj quantifies the number of observational data needed to contain sufficient information about the interventional distribution. Ifare very close, which means that the distributionsA l-' / ' ) are very similar, the exponent is bigger so fewer observational data (smaller «) would containsufficient information of the interventional distributions.(2) $2 quantifies the stability of the estimator used. Since the least squared estimator is used which is known to be stable when n>d and m>d, having more n would entail smaller $2 .(3) $3 andquantifies the ratio of the effective sample size nCff and the interventional sample size m. nes was first defined in covariate shift literature and an intuition is given that the performance of weighted conformal prediction should depend on neff, the theorem is the first to quantitatively show that «eff rather than n is the key to measure the performance of weighed conformal prediction when compared against standard conformal prediction.
[0090] From Theorem 1, it can be seen that the method in Algorithm 2 as proposed in the present disclosure is more efficient than the naive method in Algorithm 1 in terms of width of confidence interval provided, when the interventional distribution is close to the observational distribution, when the dimension d is relatively high compared to the number of interventional samples m, and when the effective sample size weir is larger than m.
[0091] Embodiments in the following will propose a more implementable variant of wTCP-DR in Algorithm 2.
[0092] In practice, although the transductive conformal prediction in Algorithm 2 is theoretically well-grounded, it is notoriously expensive to compute, compared to splitconformal prediction. The reason that split conformal prediction cannot be used in Algorithm 2 is the density ratio f evaluated at test sample, which requires the knowledge of both test covariate xn-m- and test target value j’n-m-i but unfortunately jn+m+i is inaccessible. In some embodiments of the present disclosure, it is proposed two-stage split conformal prediction which is computationally more efficient than transductive conformal prediction Algorithm 2 and the same marginal coverage guarantee may be achieved.
[0093] In the first stage, since the interventional labels(intervened outcomes of the intervened samples) are accessible, the density ratios normalized conformal weights in, . erefore, split weighted conformal prediction may be used to construct confidence intervals intervened samplesge guarantee. In the second r stage, by noticing that the test sample 3 , „„ . ., shares the same distribution as the intervened samples A^ , - • ■ , A^^,, a standard split conformal prediction may be used to construct confidence interval for the test samplewith marginal coverage guarantee.
[0094] Details of this process are described below.
[0095] In a first stage, for each of the plurality of intervened samples / p 1 / / te £ <r / J , the following operations are performed iteratively:- Generating a regression model or in some embodiments, usinga same regression generated on a regression model p on- Determining conformity scores of the observed samples- Determining a weight for the intervened sample and the respective weights for the observed samples- Constructing a weighted empirical distribution of the conformity scores based on the determined weights- Determining the quantile value of the empirical distribution corresponding to a predJetermined probability a nd- Determining a plurality of candidate confidence intervals based on the plurality of quantile values and the predicted outcome of the test object, a candidate confidence interval is defined as a lower bound(that is calculated as the predicted output minus the quantile value) and an upper bound(that is calculated as the predicted output plus the quantile value).
[0096] Through the above iterative process, a plurality of lower bounds and upper bounds are determined. Then in a second stage, the following operations are performed iteratively:- Determining a lower bound regression modelbased on the observed feature information of the observed objects in the observational dataset and a plurality of lower bounds of the plurality of candidate confidence intervals-i- iMb- Determining an upper bound regression modelbased on the observed feature information of the observed objects in the observational dataset and a plurality of upper bounds of the plurality of candidate confidence intervals- Determining, using the lower bound regression model, a predicted lower boundp based on the test feature information of the test object and determining, using the upper bound regression model, a predicted upper boundbasedon the test feature information of the test object, and- Determining the confidence interval for the predicted outcome of the test object by the predicted lower bound and the predicted upper bound as
[0097] The two-stage process described above is summarized in Algorithm 3 (as shown in Table 3).Table 3
[0098] Tn some embodiments, another two-stage process is provided (Algorithm 4). The computational cost of Algorithm 4 may be reduced by directly fitting a regressor overthe interval lower boundsan<^ fitingaregressor over the interval upper bounds in the secondstage. This two-stage process is described above and is summarized in Table 4.
[0099] The first stage is the same as Algorithm 3
[0100] In a second stage, the interventional dataset 320 £)‘fis split into a training and a calibration dataset
[0101] Then a lower bound regression modelis fitted based on the observed feature information of the observed objects in the training dataset and a plurality of lower bounds of the plurality of candidate confidence intervals,fitted based on the observed feature information of the observed objects in the training dataset and a plurality of upper bounds of the plurality of candidate confidence intervals,
[0102] Then a conformal prediction process is applied using the lower bound regression model, the upper bound regression model, and the calibration dataset, to obtain a lower bound and an upper bound of the confidence interval for the predicted outcome of the test object. The conformal prediction process comprises respective intervened conformalityscores of a plurality of intervened samples in the calibration dataset based on a maximum one of the following for each of the plurality of intervened samples in the calibration datasetwhere m (x ) --(,7 represents a difference between a predicted lower bound determined by the lower bound regression model based on intervened feature information of the intervened sample and an expected lower bound covering the intervened outcome of the intervened sample, and < v — ??? (A< ) represents a difference between an expected upper bound covering the intervened outcome of the intervened sample and a predicted upper bound determined by the upper bound regression model based on intervened feature information of the intervened sample.
[0103] Then an intervened empirical distribution of the respective intervened conformality scores at least based on the respective intervened conformality scores, p — . . 1 . . • An intervened quantile value ot the intervened empiricaldistribution corresponding to a predetermined probability is calculated as■
[0104] Then the lower bound regression modelis used to produce a predicted lower bound) based on the test feature information of the test obi ect and the upper bound regression modelis used to produce a predicted upper bound Zir based on the test feature information of the test object. The confidenceinterval for the test object is defined with a lower bound equal to the predicted lower bound minus the intervened quantile value and an upper bound equal to the predicted upper bound plus the intervened quantile value, represented as•
[0105] This two-stage process in Algorithm 4 is summarized in Table 4.Table 4
[0106] The above focuses on casual inference for counterfactual outcomes 1 ( 1 ) and Y (0). The confidence interval for a predicted outcome of the test object with a specific treatment is described above. As shown in FIG. 3, for two different treatments, two different groups of observational dataset 310 and interventional dataset 320 may be applied to generate a confidence interval 330-1 for a predicted outcome of the test object with the first treatment assigned and / or a confidence interval 330-2 for a predicted outcome of the test object with the second treatment assigned.
[0107] Further, offering confidence intervals for individual treatment effects may holdgreater practical significance. The algorithms proposed herein may predict confidence intervals that has marginalcoverage guarantee for the potential outcomeunder treatment £ ~~ 1 (or under control t “ ()).
[0108] In some embodiments, a confidence interval for selection of a treatment for the test object may be determined based on the confidence interval 330-1 and the confidence interval 330-2, for the ITE purposes. The confidence interval for selection of a treatment for the test object comprises an individual treatment effect (ITE) for the test object. In some examples, the lower bound of this confidence interval is calculated as a lower boundof one of the confidence interval 330-1 or 330-2 which is associated with a treatment of applying a medical intervention minus an upper bounder - ’I? of the other one of the confidence interval 330-1 or 330-2 which is associated with a treatment of not applying the medical intervention The upper bound of this confidence interval is calculated as an upper bound C.p of one of the confidence interval 330-1 or 330-2 which is associated with a treatment of applying a medical intervention minus a lower bound CQ of the other one of the confidence interval 330-1 or 330-2 which is associated with a treatment of not applying the medical intervention.
[0109] The naive way of constructing intervals for ITE is to use Bonferroni correction, i.e., . The empirical result isrdemonstrated using the naive way for fair comparison among methods that infer counterfactual outcomes, and the results are introduced where intervals for ITE are constructed using the nested methods.
[0110] According to embodiments of the present disclosure, there is proposed an improved solution that provides confidence intervals for predicting counterfactual outcomes and individual treatment effects with guaranteed marginal coverage, even under hidden confounding. This solution is strictly advantageous to the naive method that only uses interventional data or the method that only uses observational data. In some embodiments, there is also proposed a two-stage variant with the same guarantee at a lower computationalcost.
[0111] FIG. 4 illustrates a flowchart of a process 400 for casual inference in accordance with some embodiments of the present disclosure. The process 400 may be implemented at the computer system 110 of FIG. 1.
[0112] At block 410, the computer system 110 generates a first regression model at least based on a first observational dataset associated with a first treatment, the first observational dataset comprising a plurality of observed samples, an observed sample comprising observed feature information of an observed object and an observed outcome of the observed object, the first treatment being assigned to the object in the first observational dataset based on the observed feature information.
[0113] At block 420, the computer system 110 determines respective conformality scores of the plurality of observed samples based on differences between predicted outcomes determined by the first regression model based on the observed feature information of the observed objects and observed outcomes in the first observational dataset.
[0114] At block 430, the computer system 110 determines respective weights for the plurality of observed samples at least based on a first interventional dataset associated with the first treatment, the first interventional dataset comprising a plurality of intervened samples, an intervened sample comprising intervened feature information of an intervened object and an intervened outcome of the intervened object, the first treatment being randomly assigned to the intervened objects in the first intervened dataset.
[0115] At block 440, the computer system 110 generates an empirical distribution of the respective conformality scores at least based on the respective conformality scores and the respective weights for the plurality of observed samples.
[0116] At block 450, for a test object, the computer system 110 a first confidence interval for a predicted outcome of the test object with the first treatment assigned based on the empirical distribution and test feature information of the test object.
[0117] In some embodiments, determining a first confidence interval for a predicted outcome of the test object comprises: determining a quantile value of the empirical distribution corresponding to a predetermined probability; determining, using the first regression model, a predicted outcome of the test object with the first treatment assignedbased on test feature information of the test object; and determining a first confidence interval for the predicted outcome of the test object based on the quantile value and the predicted outcome of the test object.|00118| In some embodiments, determining respective weights for the plurality of observed samples based on a first interventional dataset associated with the first treatment comprises: determining respective density ratios for the plurality of observed samples and a test density ratio for the test sample based on the first interventional dataset and the first observational dataset; and determining a weight for each of the plurality of observed samples by calculating a ratio of a density ratio of the observed sample to a sum of the respective density ratios for the plurality of observed samples and the test density ratio.
[0119] In some embodiments, generating the empirical distribution of the respective conformality scores comprises: generating the empirical distribution of the respective conformality scores based on weighting the respective conformality scores with the respective weights.
[0120] In some embodiments, generating an empirical distribution of the respective conformality scores comprises: for each of the plurality of intervened samples, determining a weight for the intervened sample at least based on the first interventional dataset; and generating an empirical distribution for the intervened sample by weighting the respective conformality scores with the respective weights and weighting a predetermined score corresponding to an infinite number of samples with the weight for the intervened sample, to obtain a plurality of empirical distributions for the plurality of intervened samples.
[0121] In some embodiments, determining a first confidence interval for the predicted outcome of the test object based on the quantile value and the predicted outcome of the test object comprises: determining a plurality of quantile values of the plurality of empirical distributions corresponding to the predetermined probability, determining a plurality of candidate confidence intervals based on the plurality of quantile values and the predicted outcome of the test object, a candidate confidence interval being defined by a lower bound and an upper bound; determining the first confidence interval for the predicted outcome of the test object based on the plurality of candidate confidence intervals.
[0122] In some embodiments, determining the first confidence interval for the predicted outcome of the test object based on the plurality of candidate confidence intervals comprises:determining a lower bound regression model based on the observed feature information of the observed objects in the first observational dataset and a plurality of lower bounds of the plurality of candidate confidence intervals; determining an upper bound regression model based on the observed feature information of the observed objects in the first observational dataset and a plurality of upper bounds of the plurality of candidate confidence intervals; determining, using the lower bound regression model, a predicted lower bound based on the test feature information of the test object and determining, using the upper bound regression model, a predicted upper bound based on the test feature information of the test object, wherein the first confidence interval for the predicted outcome of the test object being defined by the predicted lower bound and the predicted upper bound.
[0123] In some embodiments, determining the first confidence interval for the predicted outcome of the test object based on the plurality of candidate confidence intervals comprises: splitting the first interventional dataset into a training dataset and a calibration dataset; determining a lower bound regression model based on the observed feature information of the observed objects in the training dataset and a plurality of lower bounds of the plurality of candidate confidence intervals; determining an upper bound regression model based on the observed feature information of the observed objects in the training dataset and a plurality of upper bounds of the plurality of candidate confidence intervals; and determining the first confidence interval for the predicted outcome of the test object by applying a conformal prediction process using the lower bound regression model, the upper bound regression model, and the calibration dataset, to obtain a lower bound and an upper bound of the first confidence interval.
[0124] Tn some embodiments, determining the first confidence interval for the predicted outcome of the test object by applying the conformal prediction process comprises: determining respective intervened conformality scores of a plurality of intervened samples in the calibration dataset based on a maximum one of the following for each of the plurality of intervened samples in the calibration dataset: a difference between a predicted lower bound determined by the lower bound regression model based on intervened feature information of the intervened sample and an expected lower bound covering the intervened outcome of the intervened sample, and a difference between an expected upper bound covering the intervened outcome of the intervened sample and a predicted upper bound determined by the upper bound regression model based on intervened feature informationof the intervened sample; generating an intervened empirical distribution of the respective intervened conformality scores at least based on the respective intervened conformality scores; determining an intervened quantile value of the intervened empirical distribution corresponding to a predetermined probability; determining, using the lower bound regression model, a predicted lower bound based on the test feature information of the test object and determining, using the upper bound regression model, a predicted upper bound based on the test feature information of the test object; and determining the first confidence interval with a lower bound equal to the predicted lower bound minus the intervened quantile value and an upper bound equal to the predicted upper bound plus the intervened quantile value.
[0125] In some embodiments, a confidence range is divided into a plurality of confidence units. In some embodiments, generating the first regression model comprises: for each of the plurality of confidence units, generating a first regression model based on the first observational dataset and an additional test sample consisting of the test feature information and the confidence unit, to obtain a plurality of first regression models corresponding to the plurality of confidence units; wherein for each of the plurality of first regression models, the determining of the respective conformality scores, the determining of the respective weights for the plurality of observed samples, and the generating of the empirical distribution are performed iteratively, to obtain a plurality of quantile values for the plurality of first regression model.
[0126] In some embodiments, determining the first confidence interval for the predicted outcome of the test object based on the quantile value and the predicted outcome of the test object comprises: for each of the plurality of first regression models, determining a test conformity score of the test sample based on a difference between a predicted output determined by the first regression model and a confidence unit used to generate the first regression model; in accordance with a determination that the test conformity score is lower than or equal to the quantile value for the first regression model, combing the confidence unit as a part of the first confidence interval
[0127] In some embodiments, the process 400 further comprises: generating a second regression model at least based on a second observational dataset associated with a second treatment, the second observational dataset comprising a plurality of observed samples, an observed sample comprising observed feature information of an observed object and anobserved outcome of the observed object, the second treatment being assigned to the object in the second observational dataset based on the observed feature information; determining respective conformality scores of the plurality of observed samples based on differences between predicted outcomes determined by the second regression model based on the observed feature information of the observed objects and observed outcomes in the second observational dataset, determining respective weights for the plurality of observed samples at least based on a second interventional dataset associated with the second treatment, the second interventional dataset comprising a plurality of intervened samples, an intervened sample comprising intervened feature information of an intervened object and an intervened outcome of the intervened object, the second treatment being randomly assigned to the intervened objects in the second intervened dataset; generating a further empirical distribution of the respective conformality scores at least based on the respective conformality scores and the respective weights for the plurality of observed samples; and determining a second confidence interval for a predicted outcome of the test object with the second treatment assigned based on the further empirical distribution and the test feature information of the test object.
[0128] In some embodiments, the process 400 further comprises: determining a confidence interval for selection of a treatment for the test object based on the first confidence interval and the second confidence interval.|00129| In some embodiments, the confidence interval for selection of a treatment for the test object comprises an individual treatment effect (ITE) for the test object.
[0130] In some embodiments, one of the first and second treatments comprises one of applying a specific medical treatment on a subject and not applying the specific medical treatment on the subject, and the other one of the first and second treatments comprises the other one of applying a specific medical treatment on a subject and not applying the specific medical treatment on the subject.
[0131] FIG 5 shows a block diagram of an apparatus 500 for model training in accordance with some embodiments of the present disclosure. The apparatus 500 may be implemented, for example, or included at the computer system 110 of FIG. 1. Various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0132] As shown, the apparatus 500 includes a model generating module 510 configured to generate a first regression model at least based on a first observational dataset associated with a first treatment, the first observational dataset comprising a plurality of observed samples, an observed sample comprising observed feature information of an observed object and an observed outcome of the observed object, the first treatment being assigned to the object in the first observational dataset based on the observed feature information.
[0133] The apparatus 500 includes a conformality determining module 520 configured to determine respective conformality scores of the plurality of observed samples based on differences between predicted outcomes determined by the first regression model based on the observed feature information of the observed objects and observed outcomes in the first observational dataset.
[0134] The apparatus 500 further includes a weight determining module 530 configured to determine respective weights for the plurality of observed samples at least based on a first interventional dataset associated with the first treatment, the first interventional dataset comprising a plurality of intervened samples, an intervened sample comprising intervened feature information of an intervened object and an intervened outcome of the intervened object, the first treatment being randomly assigned to the intervened objects in the first intervened dataset.
[0135] The apparatus 500 further includes a distribution generating module 540 configured to generate an empirical distribution of the respective conformality scores at least based on the respective conformality scores and the respective weights for the plurality of observed samples.
[0136] The apparatus 500 further includes a confidence determining module 570 configured to for a test object, determine a first confidence interval for a predicted outcome of the test object with the first treatment assigned based on the empirical distribution and test feature information of the test object.
[0137] The apparatus 500 may further comprises corresponding modules that are configured to perform the operations of the process 400 and other embodiments as described herein.
[0138] FIG. 6 illustrates a block diagram of an electronic device 600 in which one ormore embodiments of the present disclosure can be implemented. It would be appreciated that the electronic device 600 shown in FIG. 6 is only an example and should not constitute any restriction on the function and scope of the embodiments described herein. The electronic device 600 may be used, for example, to implement the computer system 110 of FIG 1. The electronic device 600 may also be used to implement the apparatus 500 of FIG. 5.
[0139] As shown in FIG. 6, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be an actual or virtual processor and can execute various processes according to the programs stored in the memory 620. In a multiprocessor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 600.
[0140] The electronic device 600 typically includes a variety of computer storage medium Such medium may be any available medium that is accessible to the electronic device 600, including but not limited to volatile and non-volatile medium, removable and non-removable medium. The memory 620 may be volatile memory (for example, a register, cache, a random access memory (RAM)), a non-volatile memory (for example, a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory) or any combination thereof. The storage device 630 may be any removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data (such as training data for training) and can be accessed within the electronic device 600.
[0141] The electronic device 600 may further include additional removable / non- removable, volatile / non-volatile, transitory / non-transitory storage medium. Although not shown in FIG. 6, a disk driver for reading from or writing to a removable, non-volatile disk (such as a "floppy disk"), and an optical disk driver for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each driver may be connected to the bus (not shown) by one or more data medium interfaces. The memory 620 may include a computer program product 625, which has one or more program modules configured to perform various methods or acts of various embodiments of the presentdisclosure.
[0142] The communication unit 640 communicates with a further computing device through the communication medium. In addition, functions of components in the electronic device 600 may be implemented by a single computing cluster or multiple computing machines, which can communicate through a communication connection. Therefore, the electronic device 600 may be operated in a networking environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0143] The input device 650 may be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as required. The external device, such as a storage device, a display device, etc., communicate with one or more devices that enable users to interact with the electronic device 600, or communicate with any device (for example, a network card, a modem, etc.) that makes the electronic device 600 communicate with one or more other computing devices. Such communication may be executed via an input / output (I / O) interface (not shown).
[0144] According to example implementation of the present disclosure, a computer- readable storage medium is provided, on which a computer-executable instruction or computer program is stored, where the computer-executable instructions or the computer program is executed by the processor to implement the method described above. According to example implementation of the present disclosure, a computer program product is also provided. The computer program product is physically stored on a non-transient computer- readable medium and includes computer-executable instructions, which are executed by the processor to implement the method described above.
[0145] Various aspects of the present disclosure are described herein with reference to the flow chart and / or the block diagram of the method, the device, the equipment and the computer program product implemented in accordance with the present disclosure. It would be appreciated that each block of the flowchart and / or the block diagram and the combination of each block in the flowchart and / or the block diagram may be implemented by computer-readable program instructions.
[0146] These computer-readable program instructions may be provided to the processing units of general-purpose computers, special computers or other programmable data processing devices to produce a machine that generates a device to implement the functions / acts specified in one or more blocks in the flow chart and / or the block diagram when these instructions are executed through the processing units of the computer or other programmable data processing devices. These computer-readable program instructions may also be stored in a computer-readable storage medium. These instructions enable a computer, a programmable data processing device and / or other devices to work in a specific way. Therefore, the computer-readable medium containing the instructions includes a product, which includes instructions to implement various aspects of the functions / acts specified in one or more blocks in the flowchart and / or the block diagram.
[0147] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operational steps can be performed on a computer, other programmable data processing apparatus, or other devices, to generate a computer-implemented process, such that the instructions which execute on a computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks in the flowchart and / or the block diagram.
[0148] The flowchart and the block diagram in the drawings show the possible architecture, functions and operations of the system, the method and the computer program product implemented in accordance with the present disclosure. In this regard, each block in the flowchart or the block diagram may represent a part of a module, a program segment or instructions, which contains one or more executable instructions for implementing the specified logic function. In some alternative implementations, the functions marked in the block may also occur in a different order from those marked in the drawings. For example, two consecutive blocks may actually be executed in parallel, and sometimes can also be executed in a reverse order, depending on the function involved. It should also be noted that each block in the block diagram and / or the flowchart, and combinations of blocks in the block diagram and / or the flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by the combination of dedicated hardware and computer instructions.
[0149] Each implementation of the present disclosure has been described above Theabove description is example, not exhaustive, and is not limited to the disclosed implementations. Without departing from the scope and spirit of the described implementations, many modifications and changes are obvious to ordinary skill in the art. The selection of terms used in this article aims to best explain the principles, practical application or improvement of technology in the market of each implementation, or to enable other ordinary skill in the art to understand the various embodiments disclosed herein.
Claims
WHAT IS CLAIMED IS:
1. A method for casual inference comprising: generating a first regression model at least based on a first observational dataset associated with a first treatment, the first observational dataset comprising a plurality of observed samples, an observed sample comprising observed feature information of an observed object and an observed outcome of the observed object, the first treatment being assigned to the object in the first observational dataset based on the observed feature information; determining respective conformality scores of the plurality of observed samples based on differences between predicted outcomes determined by the first regression model based on the observed feature information of the observed objects and observed outcomes in the first observational dataset; determining respective weights for the plurality of observed samples at least based on a first interventional dataset associated with the first treatment, the first interventional dataset comprising a plurality of intervened samples, an intervened sample comprising intervened feature information of an intervened object and an intervened outcome of the intervened object, the first treatment being randomly assigned to the intervened objects in the first intervened dataset; generating an empirical distribution of the respective conformality scores at least based on the respective conformality scores and the respective weights for the plurality of observed samples, and for a test object, determining a first confidence interval for a predicted outcome of the test object with the first treatment assigned based on the empirical distribution and test feature information of the test object.
2. The method of claim 1, wherein determining a first confidence interval for a predicted outcome of the test object comprises: determining a quantile value of the empirical distribution corresponding to a predetermined probability; determining, using the first regression model, a predicted outcome of the test object with the first treatment assigned based on test feature information of the test object; anddetermining a first confidence interval for the predicted outcome of the test object based on the quantile value and the predicted outcome of the test object.
3. The method of claim 1, wherein determining respective weights for the plurality of observed samples based on a first interventional dataset associated with the first treatment comprises: determining respective density ratios for the plurality of observed samples and a test density ratio for the test sample based on the first interventional dataset and the first observational dataset; and determining a weight for each of the plurality of observed samples by calculating a ratio of a density ratio of the observed sample to a sum of the respective density ratios for the plurality of observed samples and the test density ratio4 The method of claim 3, wherein generating the empirical distribution of the respective conformality scores comprises: generating the empirical distribution of the respective conformality scores based on weighting the respective conformality scores with the respective weights.
5. The method of claim 1, wherein generating an empirical distribution of the respective conformality scores comprises: for each of the plurality of intervened samples, determining a weight for the intervened sample at least based on the first interventional dataset; and generating an empirical distribution for the intervened sample by weighting the respective conformality scores with the respective weights and weighting a predetermined score corresponding to an infinite number of samples with the weight for the intervened sample, to obtain a plurality of empirical distributions for the plurality of intervened samples.
6. The method of claim 5, wherein determining a first confidence interval for the predicted outcome of the test object of the test object comprises: determining a plurality of quantile values of the plurality of empirical distributions corresponding to the predetermined probability;determining a plurality of candidate confidence intervals based on the plurality of quantile values and the predicted outcome of the test object, a candidate confidence interval being defined by a lower bound and an upper bound; and determining the first confidence interval for the predicted outcome of the test object based on the plurality of candidate confidence intervals.
7. The method of claim 6, wherein determining the first confidence interval for the predicted outcome of the test object based on the plurality of candidate confidence intervals comprises: determining a lower bound regression model based on the observed feature information of the observed objects in the first observational dataset and a plurality of lower bounds of the plurality of candidate confidence intervals; determining an upper bound regression model based on the observed feature information of the observed objects in the first observational dataset and a plurality of upper bounds of the plurality of candidate confidence intervals; and determining, using the lower bound regression model, a predicted lower bound based on the test feature information of the test object and determining, using the upper bound regression model, a predicted upper bound based on the test feature information of the test object, wherein the first confidence interval for the predicted outcome of the test object being defined by the predicted lower bound and the predicted upper bound.
8. The method of claim 6, wherein determining the first confidence interval for the predicted outcome of the test object based on the plurality of candidate confidence intervals comprises: splitting the first interventional dataset into a training dataset and a calibration dataset, determining a lower bound regression model based on the observed feature information of the observed objects in the training dataset and a plurality of lower bounds of the plurality of candidate confidence intervals; determining an upper bound regression model based on the observed feature information of the observed objects in the training dataset and a plurality of upper bounds of the plurality of candidate confidence intervals; anddetermining the first confidence interval for the predicted outcome of the test object by applying a conformal prediction process using the lower bound regression model, the upper bound regression model, and the calibration dataset, to obtain a lower bound and an upper bound of the first confidence interval.
9. The method of claim 8, wherein determining the first confidence interval for the predicted outcome of the test object by applying the conformal prediction process comprises: determining respective intervened conformality scores of a plurality of intervened samples in the calibration dataset based on a maximum one of the following for each of the plurality of intervened samples in the calibration dataset: a difference between a predicted lower bound determined by the lower bound regression model based on intervened feature information of the intervened sample and an expected lower bound covering the intervened outcome of the intervened sample, and a difference between an expected upper bound covering the intervened outcome of the intervened sample and a predicted upper bound determined by the upper bound regression model based on intervened feature information of the intervened sample; generating an intervened empirical distribution of the respective intervened conformality scores at least based on the respective intervened conformality scores; determining an intervened quantile value of the intervened empirical distribution corresponding to a predetermined probability; determining, using the lower bound regression model, a predicted lower bound based on the test feature information of the test object and determining, using the upper bound regression model, a predicted upper bound based on the test feature information of the test object; and determining the first confidence interval with a lower bound equal to the predicted lower bound minus the intervened quantile value and an upper bound equal to the predicted upper bound plus the intervened quantile value.
10. The method of claim 1, wherein a confidence range is divided into a plurality of confidence units, and wherein generating the first regression model comprises: for each of the plurality of confidence units, generating a first regression model based on the first observational dataset and an additional test sample consisting of the test feature information and the confidenceunit, to obtain a plurality of first regression models corresponding to the plurality of confidence units, and wherein for each of the plurality of first regression models, the determining of the respective conformality scores, the determining of the respective weights for the plurality of observed samples, and the generating of the empirical distribution are performed iteratively, to obtain a plurality of quantile values for the plurality of first regression model.
11. The method of claim 10, wherein determining the first confidence interval for the predicted outcome of the test object comprises: for each of the plurality of first regression models, determining a test conformity score of the test sample based on a difference between a predicted output determined by the first regression model and a confidence unit used to generate the first regression model; and in accordance with a determination that the test conformity score is lower than or equal to the quantile value for the first regression model, combing the confidence unit as a part of the first confidence interval.
12. The method of claim 1, further comprising: generating a second regression model at least based on a second observational dataset associated with a second treatment, the second observational dataset comprising a plurality of observed samples, an observed sample comprising observed feature information of an observed object and an observed outcome of the observed object, the second treatment being assigned to the object in the second observational dataset based on the observed feature information; determining respective conformality scores of the plurality of observed samples based on differences between predicted outcomes determined by the second regression model based on the observed feature information of the observed objects and observed outcomes in the second observational dataset; determining respective weights for the plurality of observed samples at least based on a second interventional dataset associated with the second treatment, the second interventional dataset comprising a plurality of intervened samples, an intervened sample comprising intervened feature information of an intervened object and an intervenedoutcome of the intervened object, the second treatment being randomly assigned to the intervened objects in the second intervened dataset; generating a further empirical distribution of the respective conformality scores at least based on the respective conformality scores and the respective weights for the plurality of observed samples; and determining a second confidence interval for a predicted outcome of the test object with the second treatment assigned based on the further empirical distribution and the test feature information of the test object.
13. The method of claim 12, further comprising: determining a confidence interval for selection of a treatment for the test object based on the first confidence interval and the second confidence interval14 The method of claim 13, wherein the confidence interval for selection of a treatment for the test object comprises an individual treatment effect (ITE) for the test object.
15. The method of claim 1, wherein one of the first and second treatments comprises one of applying a specific medical treatment on a subject and not applying the specific medical treatment on the subject, and the other one of the first and second treatments comprises the other one of applying a specific medical treatment on a subject and not applying the specific medical treatment on the subject.
16. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions executable by the at least one processing unit, the instructions, upon execution by the at least one processing unit, causing the device to perform: generating a first regression model at least based on a first observational dataset associated with a first treatment, the first observational dataset comprising a plurality of observed samples, an observed sample comprising observed feature information of an observed object and an observed outcome of the observed object, the first treatment being assigned to the object in the first observational dataset based on the observed feature information;determining respective conformality scores of the plurality of observed samples based on differences between predicted outcomes determined by the first regression model based on the observed feature information of the observed objects and observed outcomes in the first observational dataset; determining respective weights for the plurality of observed samples at least based on a first interventional dataset associated with the first treatment, the first interventional dataset comprising a plurality of intervened samples, an intervened sample comprising intervened feature information of an intervened object and an intervened outcome of the intervened object, the first treatment being randomly assigned to the intervened objects in the first intervened dataset; generating an empirical distribution of the respective conformality scores at least based on the respective conformality scores and the respective weights for the plurality of observed samples, and for a test object, determining a first confidence interval for a predicted outcome of the test object with the first treatment assigned based on the empirical distribution and test feature information of the test object.
17. The device of claim 16, wherein determining respective weights for the plurality of observed samples based on a first interventional dataset associated with the first treatment comprises: determining respective density ratios for the plurality of observed samples and a test density ratio for the test sample based on the first interventional dataset and the first observational dataset; and determining a weight for each of the plurality of observed samples by calculating a ratio of a density ratio of the observed sample to a sum of the respective density ratios for the plurality of observed samples and the test density ratio18 The device of claim 16, wherein generating an empirical distribution of the respective conformality scores comprises: for each of the plurality of intervened samples, determining a weight for the intervened sample at least based on the first interventional dataset; andgenerating an empirical distribution for the intervened sample by weighting the respective conformality scores with the respective weights and weighting a predetermined score corresponding to an infinite number of samples with the weight for the intervened sample, to obtain a plurality of empirical distributions for the plurality of intervened samples, and wherein determining a quantile value of the empirical distribution corresponding to a predetermined probability comprises: determining a plurality of quantile values of the plurality of empirical distributions corresponding to the predetermined probability.
19. The device of claim 18, wherein determining a first confidence interval for the predicted outcome of the test object based on the quantile value and the predicted outcome of the test object comprises: determining a plurality of candidate confidence intervals based on the plurality of quantile values and the predicted outcome of the test object, a candidate confidence interval being defined by a lower bound and an upper bound; and determining the first confidence interval for the predicted outcome of the test object based on the plurality of candidate confidence intervals.
20. A computer-readable storage medium, having a computer program stored thereon which, upon execution by an electronic device, causes the device to perform the method according to any of claims 1 to 15.
Citation Information
Patent Citations
Dual-robust clinical trial data processing method and system using historical contrast data
CN112863622A
Causal effect evaluation method and system
CN116720543A