Counterfactual inference under rank preservation
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING TECH & BUSINESS UNIV
- Filing Date
- 2024-06-21
- Publication Date
- 2026-07-29
AI Technical Summary
Existing counterfactual inference methods face challenges in identifying counterfactual outcomes due to stringent assumptions like homogeneity and strict monotonicity, which are often unmet in real-world scenarios, leading to unidentifiable results even in randomized controlled trials.
A novel approach that relaxes these assumptions by using rank preservation and a novel loss function to estimate counterfactual outcomes, allowing for unbiased estimation without requiring identical quantile values for each individual, and employing kernel smoothing for improved accuracy.
The proposed method effectively identifies counterfactual outcomes with relaxed identifiability assumptions, providing accurate predictions under diverse data types and reducing the complexity of bi-level optimization.
Smart Images

Figure CN2024100791_26122025_PF_FP_ABST
Abstract
Description
COUNTERFACTUAL INFERENCE UNDER RANK PRESERVATIONField
[0001] The disclosed example embodiments relate generally to machine learning and, more particularly, to a method, apparatus, device and computer readable storage medium for counterfactual inference under rank preservation.Background
[0002] Generally, it is challenging to answer higher-level queries based on lower-level information. Among these, counterfactual inference represents the most difficult level. It involves exploring the potential impact of a treatment on an outcome, given knowledge about a different observed treatment and the actual outcome. For instance, consider a patient who has not taken medication previously and is now experiencing a headache. The counterfactual query would be to determine whether the headache would have occurred if the patient had taken the medication from the start. Addressing such counterfactual queries can offer valuable insights in scenarios such as credit assignment, root-cause analysis, and ensuring fair algorithmic decision-making.Summary
[0003] In a first aspect of the present disclosure, there is provided a method for counterfactual inference. The method comprises: obtaining an observational dataset associated with a plurality of treatments and a plurality of observed objects, wherein the observational dataset comprises a plurality of observed samples, an observed sample corresponds to an observed object of the plurality of observed objects assigned with a treatment of the plurality of treatments and comprises an observed feature of the observed object and an observed outcome of the observed object; for a target observed object of the plurality of observed objects, determining, from the observational dataset, a first set of observed outcomes of a first set of observed objects assigned with a first treatment of the plurality of treatments, wherein the target observed object is assigned with the first treatment; determining, from the observational dataset, a second set of observed outcomes of a second set of observed objects assigned with a second treatment different from the first treatment; and generating a predicted outcome of the target observed object assigned with the second treatment based on the first set of observed outcomes, the second set of observed outcomes and a target observed outcome of the target observed object.
[0004] In a second aspect of the present disclosure, there is provided an apparatus for counterfactual inference. The apparatus comprises: a dataset obtaining module configured to obtain an observational dataset associated with a plurality of treatments and a plurality of observed objects, wherein the observational dataset comprises a plurality of observed samples, an observed sample corresponds to an observed object of the plurality of observed objects assigned with a treatment of the plurality of treatments and comprises an observed feature of the observed object and an observed outcome of the observed object; a first outcome determination module configured to, for a target observed object of the plurality of observed objects, determine, from the observational dataset, a first set of observed outcomes of a first set of observed objects assigned with a first treatment of the plurality of treatments, wherein the target observed object is assigned with the first treatment; a second outcome determination module configured to determine, from the observational dataset, a second set of observed outcomes of a second set of observed objects assigned with a second treatment different from the first treatment; and a predicted outcome generation module configured to generate a predicted outcome of the target observed object assigned with the second treatment based on the first set of observed outcomes, the second set of observed outcomes and a target observed outcome of the target observed object.
[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions executable by the at least one processing unit. The instructions, upon execution by the at least one processing unit, cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer readable storage medium is provided. The computer readable storage medium stores a computer program which, when executed by a processor, causes the method of the first aspect to be implemented.
[0007] It would be appreciated that the content described in the Summary section of the present invention is neither intended to identify key or essential features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily envisaged through the following description.Brief Description of the Drawings
[0008] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent in combination with the accompanying drawings and with reference to the following detailed description. In the drawings, the same or similar reference symbols refer to the same or similar elements, where:
[0009] FIG. 1 illustrates a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;
[0010] FIG. 2A illustrates a directed acyclic graph (DAG) in a structural causal model (SCM) ;
[0011] FIG. 2B illustrates a twin network with homogeneous and heterogeneous background variables;
[0012] FIG. 2C illustrates another twin network with homogeneous and heterogeneous background variables;
[0013] FIG. 3 illustrates a schematic diagram of a counterfactual inference architecture in accordance with some embodiments of the present disclosure;
[0014] FIG. 4 illustrates a flow chart of a process for counterfactual inference in accordance with some embodiments of the present disclosure;
[0015] FIG. 5 illustrates a block diagram of an apparatus for counterfactual inference according to some embodiments of the present disclosure; and
[0016] FIG. 6 illustrates a block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented.Detailed Description
[0017] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it would be appreciated that the present disclosure may be implemented in various forms and should not be interpreted as limited to the embodiments described herein. On the contrary, these embodiments are provided for a more thorough and complete understanding of the present disclosure. It would be appreciated that the drawings and embodiments of the present disclosure are only for the purpose of illustration and are not intended to limit the scope of protection of the present disclosure.
[0018] In the description of the embodiments of the present disclosure, the term "including" and similar terms would be appreciated as open inclusion, that is, "including but not limited to" . The term "based on" would be appreciated as "at least partially based on" . The term "one embodiment" or "the embodiment" would be appreciated as "at least one embodiment" . The term "some embodiments" would be appreciated as "at least some embodiments" . Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the matching degree between various data. For example, the above matching degree can be obtained based on various technical solutions currently available and / or to be developed in the future.
[0019] It will be appreciated that the data involved in this technical proposal (including but not limited to the data itself, data acquisition or use) shall comply with the requirements of corresponding laws, regulations and relevant provisions.
[0020] It will be appreciated that before using the technical solution disclosed in each embodiment of the present disclosure, users should be informed of the type, the scope of use, the use scenario, etc. of the personal information involved in the present disclosure in an appropriate manner in accordance with relevant laws and regulations, and the user’s authorization should be obtained.
[0021] For example, in response to receiving an active request from a user, a prompt message is sent to the user to explicitly prompt the user that the operation requested operation by the user will need to obtain and use the user's personal information. Thus, users may select whether to provide personal information to the software or the hardware such as an electronic device, an application, a server or a storage medium that perform the operation of the technical solution of the present disclosure according to the prompt information.
[0022] As an optional but non-restrictive implementation, in response to receiving the user's active request, the method of sending prompt information to the user may be, for example, a pop-up window in which prompt information may be presented in text. In addition, pop-up windows may also contain selection controls for users to choose “agree” or “disagree” to provide personal information to electronic devices.
[0023] It will be appreciated that the above notification and acquisition of user authorization process are only schematic and do not limit the implementations of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0024] As used herein, the term “model” can learn a correlation between respective inputs and outputs from training data, so that a corresponding output can be generated for a given input after training is completed. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural networks model is an example of a deep learning-based model. As used herein, “model” may also be referred to as “machine learning model” , “learning model” , “machine learning network” , or “learning network” , and these terms are used interchangeably herein.
[0025] “Neural networks” are a type of machine learning network based on deep learning. Neural networks are capable of processing inputs and providing corresponding outputs, typically comprising input and output layers and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically comprise many hidden layers, thereby increasing the depth of the network. The layers of neural networks are sequentially connected so that the output of the previous layer is provided as input to the latter layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network comprises one or more nodes (also known as processing nodes or neurons) , each of which processes input from the previous layer.
[0026] Usually, machine learning can roughly comprise three stages, namely training stage, test stage, and application stage (also known as inference stage) . During the training stage, a given model can be trained using a large scale of training data, iteratively updating parameter values until the model can obtain consistent inference from the training data that meets the expected objective. Through the training, the model can be considered to learn the correlation between input and output (also known as input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the test stage, test inputs are applied to the trained model to test whether the model can provide correct outputs, thereby determining the performance of the model. In the application stage, the model can be used to process actual inputs and determine corresponding outputs based on the parameter values obtained from training.
[0027] The term “predicted outcome” , denoted by Yx, refers to an outcome that would occur for each individual if they were assigned to a specified treatment x. For example, with two treatments x and x’, each individual has two predicted outcomes Yx and Yx’.
[0028] The term “counterfactual” refers to the denial and re-representation of past facts to construct a possible hypothetical outcome. In other word, the term “counterfactual” refers to a hypothetical scenario that contradicts known past facts. For example, for each individual in a group with the treatment x, the observed outcome is Yx, while the unobserved counterfactual outcome is Yx’. Conversely, for another group with the treatment x’, the observed outcome is Yx’, and the unobserved counterfactual outcome is Yx.
[0029] FIG. 1 illustrates a block diagram of an example environment 100 in which various embodiments of the present disclosure may be implemented. In the environment 100 of FIG. 1, a computer system 110 obtains an observational dataset 120 associated with a plurality of treatments and a plurality of observed objects. The observational dataset 120 comprises a plurality of observed samples. An observed sample 125 corresponds to an observed object of the plurality of observed objects 105. The computer system 110 processes the plurality of observed samples 125 to generate counterfactual information 130. The counterfactual information 130 may include a predicted outcome of the observed object 105.
[0030] The observed sample 125 corresponds to the observed object 105 assigned with a treatment 121 of the plurality of treatments and comprises an observed feature 123 of the observed object 105 and an observed outcome 122 of the observed object 105.
[0031] An example of observed object 105 may be a subject such as a patient. An example of treatment 121 may be a medical treatment. In some embodiments, the plurality of treatments 121 may include a first treatment and a second treatment, one of the first and second treatments may include one of applying a specific medical treatment on a subject and not applying the specific medical treatment on the subject. The other one of the first and second treatments may include the other one of applying a specific medical treatment on a subject and not applying the specific medical treatment on the subject.
[0032] As an example, in medical research, an operation or treatment X indicates whether a patient receives aspirin, an outcome Y indicates whether the patient’s headache disappears, and a feature Z may include information such as the patient’s age, gender, blood pressure, etc. The first set of observed samples may include the features Z of multiple patients who received aspirin, as well as whether their headache disappeared after taking the medication; whereas the second set of observed samples may include the features Z of multiple patients who did not receive aspirin, as well as whether their headache disappeared. Based on such observed samples, the computer system 110 may generate a model and use the model to predict whether the headache of the target object will disappear with and without the intake of aspirin and to predict the treatment effect of aspirin on the target object to determine or assist in determining whether the target object should receive treatment with aspirin.
[0033] As another example, a treatment may be recommending a first media content to users, and another treatment may be recommending a second media content to users. An outcome may indicate a user engagement. The first set of observed samples may include the features and engagement levels of users who were recommended with the first media content. The second set of observed samples may include the features and engagement levels of users who were recommended with the second media content. Utilizing these observed samples, the computer system 110 may generate a predictive model. The model may then be employed to predict user engagement with the first media content and with the second media content. The prediction may be used to determine or assist in determining which type of media content should be recommended to which users.
[0034] In FIG. 1, the computer system 110 may include any computing system with computing capability, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices may include any type of mobile terminals, fixed terminals, or portable terminals, including mobile phones, desktop computers, laptops, netbooks, tablets, media computers, multimedia tablets, or any combination of the aforementioned, including accessories and peripherals of these devices or any combination thereof. Servers include but are not limited to mainframe, edge computing nodes, computing devices in cloud environment, etc.
[0035] It should be understood that the structure and function of each element in the environment 100 is described for illustrative purposes only and does not imply any limitations on the scope of the present disclosure.
[0036] As mentioned above, in many scenarios, it is expected to predict an outcome of an individual after receiving a certain treatment and predict the difference in outcomes of an individual receiving different treatments, that is, the Individual Causal Effect (ITE) . As such, the computer system can automatically make decisions or assist people in making decisions, that is, determine whether to perform a certain treatment on an individual or determine which of multiple treatments to perform on an individual. For example, it may be desirable to predict the likely impact of a certain drug or treatment on a patient’s condition, thereby automatically or assisting a physician in formulating a treatment plan. It may be also expected to predict the extent to which a training course can improve a student’s performance or predict the impact of advertisement on a consumer’s final purchase behavior. To make such predictions, counterfactual information needs to be known.
[0037] Unlike interventional queries, which are prospective and estimate counterfactual outcomes in a hypothetical world based solely on observations obtained before treatment (which is considered as pre-treatment variables) , counterfactual inference is retrospective and also takes into account the factual outcome in the observed world (which is considered as a post-treatment variable) . The inherent conflict between the hypothetical and the observed worlds presents a unique challenge, making the counterfactual outcome generally unidentifiable, even in randomized controlled trials (RCTs) .
[0038] A three-step procedure -abduction, action, and prediction -has been proposed for counterfactual inference to estimate counterfactual outcomes. However, this approach depends on the availability of SCMs that provide a complete description of the data-generating process. In real-world applications, the ground-truth SCM is often unknown, and estimating it necessitates additional assumptions to ensure identifiability, such as assuming linearity and additive noise. Unfortunately, these assumptions are difficult to meet in practice, which limits the applicability of the model.
[0039] To address the aforementioned issues, a variety of counterfactual learning methods have been proposed, each based on different identifiability assumptions. For instance, the identifiability of counterfactual outcomes has been established under assumptions of homogeneity and strict monotonicity. The homogeneity assumption suggests that the exogenous variable for each individual remains constant across various interventional settings, while the strict monotonicity assumption posits that the outcome is a strictly monotonic function of the exogenous variable, given the features.
[0040] In terms of counterfactual learning, the previously mentioned three-step procedure is utilized, which initially requires estimating the SCM. Furthermore, quantile regression has been proposed to estimate counterfactual outcomes, effectively bypassing the need for SCM estimation. However, this approach is contingent upon a stringent assumption: that the conditional quantile functions for different counterfactual outcomes originate from the same model. It also necessitates estimating a distinct quantile value for each individual, resulting in a complex bi-level optimization problem.
[0041] To better understand the example embodiments of the present disclosure, the following will describe the problem setting and provide necessary background on conformal prediction.
[0042] Throughout, capital letters represent random variables, and lowercase letters denote their specific realizations. An SCM, denoted as consists of a causal graph and a set of structural equation models The nodes in are divided into two categories. The first category comprises exogenous variables U= (U1, ..., Up) , which represent the environment during data generation and are assumed to be mutually independent. The second category consists of endogenous variables V= {V1, ..., Vp} , which represent the relevant features needed to model a question of interest. For a variable Vj , its value is determined by a structural equation Vj=fj (PAj, Uj) , where j=1, ..., p, and PAj stands for the set of parents of Vj. SCM provides a formal language for describing how the variables interact and how the resulting distribution would change in response to certain interventions. Based on SCM, the following section will introduce the problem of counterfactual inference.
[0043] Suppose there are three sets of variables denoted by X , Y and where counterfactual inference revolves around the question: “Given evidence E=e, what would have happened if X has been set to a different value x′? ” It is proposed to answer this question using the previously mentioned three-step procedure.
[0044] At the step of abduction, the value of U is determined according to the evidence E=e. At the step of action, the model is modified by removing the structural equations for X and replacing them with X=x′ , yielding the modified model At the step of prediction, andthe value of U are used to calculate the counterfactual outcome for Y.
[0045] Herein, the focus is on estimating the counterfactual outcome for each individual. To illustrate the main ideas, the common counterfactual inference problem is formulated within the context of the backdoor criterion.
[0046] Let V= (Z, X, Y) , where X causes Y, Z affects both X and Y, and the structural equation for Y is given as: Y=fY (X, Z, UX) . (1)
[0047] FIG. 2A illustrates a typical causal graph for this setting, where the exogenous variables X and Z are omitted, as they will not be utilized throughout this work. Let Yx′ denote the potential outcome if X=x′. The counterfactual question, “Given evidence (X=x, Z=z, Y=y) for an individual, what would have happened had we set X=x′ for this individual? ” is formally expressed as estimating yx′, the realization of Yx′ for the individual. Herein, a deterministic viewpoint is adhered to, treating the value of Yx′ for each individual as a fixed constant.
[0048] According to the previously mentioned three-step procedure for counterfactual inference, given the evidence (X=x, Z=z, Y=y) for an individual, the identifiability of the counterfactual value Yx′ can be achieved by determining the structural equation fY and the value of UX for this individual. This is the key idea underlying most of the existing methods.
[0049] The following will elucidate the challenges inherent in counterfactual inference. This clarification aids in a more thorough analysis of current approaches. The main challenge in counterfactual inference is the identifiability of θ, which is generally not achievable, even in RCTs. By definition, θ involves two “different worlds” simultaneously: the observed world with (X=x, Z=z, Y=y) and the hypothetical world where X=x′. The factual outcome Yx=y is observed, but the counterfactual outcome Yx′ is never observed, representing the fundamental problem in causal inference. This inherent conflict prevents the simplification of θ into a do-calculus expression, rendering it generally unidentifiable, even in RCTs. This issue is specific to counterfactual inference and does not arise in intervention expressions, as statistics of interventional distributions such as or are defined solely in a single world where the treatment X is set to x′ by intervention.
[0050] Unlike intervention expressions, which do not present this issue, counterfactual inference requires additional assumptions beyond widely used ones such as conditional exchangeability, overlap, and consistency to ensure identifiability.
[0051] Estimating θ is essentially estimating the individual treatment effect Yx′-Yx.However, the conditional average treatment effect (CATE) , given by represents the average treatment effect (ATE) for a subpopulation with Z=z, potentially overlooking the inherent heterogeneity within this subpopulation due to noise terms such as Ux.
[0052] The existing solutions will be summarized in terms of both identifiability assumptions and estimation strategies.
[0053] For clarity, an equivalent expression of Eq. (1) is first presented, incorporating counterfactuals (Yx, Yx′) . Eq. (1) may be reformulated as the following system: Yx=fY (x, Z, Ux) , Yx′=fY (x′, Z, Ux′) , (2)
[0054] where Ux and Ux′ denote the values of UX in the observed world X=x and the hypothetical world X=x′, respectively. The difference between Ux and Ux′can be caused by variations in unmeasured background variables between the two worlds, with a detailed discussion provided in the following section.
[0055] For identification, existing solutions rely on two key assumptions: homogeneity and strict monotonicity.
[0056] Assumption 1 is about homogeneity: Ux=Ux′, which implies that the value of UX for each individual remains constant across different values of x.
[0057] Assumption 2 is about strict monotonicity: for any given x and z, f fY (x, z, Ux) is a smooth and strictly monotonic function of Ux; or Yx=fY (x, z, Ux) represents a bijective mapping from Ux to Yx. Assumption 2 implies that Yx is a strict monotonic function of Ux within the subpopulation (X=x, Z=z) . Additionally, in Assumption 2, the smoothness and strict monotonicity of fY (x, z, Ux) are analogous to a bijective mapping between Yx and Ux, serving the same purpose, and thus will not be distinguished further.
[0058] The identifiability of θ depends on the above Assumptions 1 and 2. Lemma 3 summarizes the identifiability results.
[0059] Lemma 3 is under Assumptions 1 and 2, and θ is identifiable. Lemma 3 presents identifiability assumptions that are less stringent, as they introduce additional assumptions, including variability for estimating the SCM.
[0060] For the estimation of θ, the previously mentioned three-step procedure is adopted. Quantile regression is utilized, based on the finding that θ corresponds to the -th quantile of the conditional distribution where is the quantile of y in Further details are provided in the following section.
[0061] There are limitations regarding the identification and estimation of existing solutions, which will be detailed below.
[0062] FIGS. 2B and 2C illustrate twin networks that correspond to homogeneous and heterogeneous background variables, as outlined in Assumption 1 and further discussed in the following section, respectively. As depicted in FIG. 2B, Assumption 1 imposes the restriction Ux=Ux′, which might be unrealistically stringent. The exogenous variable UX represents background and environmental information influenced by numerous unmeasured factors. Ux and Ux′ capture the heterogeneity of Yx and Yx′ in the observed and hypothetical worlds, respectively. These worlds might exhibit varying levels of noise due to unmeasured factors. Allowing Ux to vary with different values of x would better reflect unobserved, unsystematic variations among different potential outcomes, as shown in FIG. 2C, which presents the differences between Ux and Ux′.
[0063] Furthermore, in Assumption 2, the strict monotonicity of fY excludes scenarios with discrete Y that has finite support and continuous Y that may result in ties. In these cases, different individuals with distinct values of UX could produce identical Y values, thereby violating the strict monotonicity requirement. Consequently, the assumption of strict monotonicity also limits the applicability of existing methods.
[0064] To estimate θ, the three-step procedure is initially followed, which requires estimating the individual’s fY and UX. However, this approach necessitates additional assumptions, such as linearity and the presence of additive noise. Alternatively, quantile regression is employed to circumvent the challenges associated with estimating fY and UX. This method fits a single model to determine the conditional quantile functions for both counterfactual and factual outcomes. The validity of this approach hinges on the assumption that the conditional quantile functions for different treatment groups originate from the same model. Furthermore, it requires estimating a unique quantile value for each individual prior to deriving the counterfactual outcomes, leading to a complex bi-level optimization challenge.
[0065] To address at least some of the above limitations, embodiments of the present disclosure propose an improved solution for counterfactual inference under rank preservation. In the solution, an observational dataset associated with a plurality of treatments and a plurality of observed objects is obtained. The observational dataset comprises a plurality of observed samples. An observed sample corresponds to an observed object of the plurality of observed objects assigned with a treatment of the plurality of treatments and comprises an observed feature of the observed object and an observed outcome of the observed object. For a target observed object of the plurality of observed objects, from the observational dataset, a first set of observed outcomes of a first set of observed objects assigned with a first treatment of the plurality of treatments is determined. The target observed object is assigned with the first treatment. From the observational dataset, a second set of observed outcomes of a second set of observed objects assigned with a second treatment different from the first treatment is determined. A predicted outcome of the target observed object assigned with the second treatment is generated based on the first set of observed outcomes, the second set of observed outcomes and a target observed outcome of the target observed object. Advantages of the proposed solution will be apparent from the example embodiments described below.
[0066] Example embodiments of the present disclosure will be descried with the reference to the drawings.
[0067] Reference is now made to FIG. 3, which illustrates a schematic diagram of a counterfactual inference architecture 300 in accordance with some embodiments of the present disclosure. A target observed object 304 is an example of the observed object 105 in FIG. 1. The counterfactual inference architecture 300 may be implemented in the environment 100 of FIG. 1. The following will describe with reference to FIG. 1.
[0068] The target observed object 304 is assigned with a first treatment (denoted as x) . A target observed feature 302 (denoted as z) of the target observed object 304 may be obtained from the observational dataset 120. The target observed feature 302 may include any suitable information of the target observed object 304, which can be used to characterize the target observed object 304. As a result of applying the first treatment, a target observed outcome 310 (denoted as y) of the target observed object 304 is obtained.
[0069] The target observed object 304 is assigned with the first treatment. From the observational dataset 120, a first set of observed outcomes 320 (denoted as Yx) of a first set of observed objects also assigned with the first treatment is determined.
[0070] Some other observed objects are assigned with a second treatment (denoted as x’) . The second treatment is different from the first treatment. From the observational dataset 120, a second set of observed outcomes 330 (denoted as Yx’) assigned with the second treatment is determined.
[0071] Assume that the target observed object 304 was assigned with the second treatment. A predicted outcome 340 (denoted as t) of the target observed object 304, if the target observed object 304 was assigned with the second treatment x’, is generated based on the first set of observed outcomes 320, the second set of observed outcomes 330 and the target observed outcome 310.
[0072] To better understand prediction of the outcome 340, some fundamental aspects are now described. The identifiability assumptions, namely Assumptions 1 and 2, are relaxed below. From a high-level perspective, identifying θ essentially involves establishing the relationship between Yx and Yx′. The three-step procedure achieves this by estimating fY and UX based on the observed data. The relaxed identifiability assumption is grounded in a rank correlation coefficient, which is defined as follows.
[0073] Definition 1 is described below. Let (x1, y1) , ....., (xn, yn) be a set of observations of two random variables (X, Y) , such thatall the values of xi and yiare unique (for simplicity, ties are neglected) . Any pair (xi, yi) and (xj, yj) , if (xj-xi) (yj-yi) >0, is said to be concordant; otherwise, they are discordant. The Kendall rank correlation coefficient is defined as:
[0074] where n (n-1) / 2 is the total number of pair combinations and sign (t) =-1, 0, 1 for t<0 , ,t=0 , ,t>0 , respectively.
[0075] The coefficient ρ (X, Y) can also be expressed as 2 (Nc-Nd) / n (n-1) , where Nc is the number of concordant pairs and Nd is the number of discordant pairs. It is clear that -1≤ρ (X, Y) ≤1, and if there is perfect agreement between the two rankings (i.e., perfect concordance) , ρ (X, Y) =1.
[0076] Assumption 3 is about rank preservation: ρ (Yx, Yx′|Z) =1.
[0077] For an individual with the observation (X=x, Z=z, Y=y) , the pair (yx=y, yx′) is denoted as the true values of (Yx, Yx′) . Assumption 3 implies that for this individual, the rankings of yx and yx′ are consistent across the distributions and respectively. Consequently, there is:
[0078] Since yx=y is observed, and the distributions and can be identified with and respectively, according to the backdoor criterion (i.e., Based on this, there is the following Proposition 1.
[0079] Proposition 1 is under Assumption 3. θ is identified as the -th quantile of where is the quantile of y in the distribution of i.e.,
[0080] Proposition 1 shows that Assumption 3 can serve as a substitute for Assumptions 1 and 2 in identifying θ. Furthermore, it is demonstrated that Assumption 3 is strictly weaker than Assumptions 1 and 2, as shown in Proposition 2.
[0081] Proposition 2 states that Assumptions 1 and 2 imply Assumption 3, but not vice versa.
[0082] Proposition 2 is intuitive, as correlation (Assumption 3) does not necessarily imply identity (Assumption 1) . In fact, Assumptions 1 and 2 are special cases of Assumption 3. Below, an example is provided that violates Assumption 1 while Assumption 2 still holds.
[0083] Assumption 3 mitigates Assumption 1. Next, the focus is on further relaxing Assumption 2. When the outcome Y is discrete or continuous variables with tied observations, ρ (Yx, Yx′) will always be less than 1 by Definition 1. To accommodate such cases, a modified version of the Kendall rank correlation coefficient given below is introduced.
[0084] Definition 2 is described below. Let (x1, y1) , ..., (xn, yn) be the observations of two random variables (X, Y) , the modified Kendall rank correlation coefficient is defined as:
[0085] where Tx is the number of tied pairs in {x1, ..., xn} and Ty is the number of tied pairs in {y1, ..., yn} .
[0086] By comparison of Definition 2 and Definition 1, it can be seen that ρm (X, Y) adjusts ρ (X, Y) by eliminating the ties in the denominator, and ρm (X, Y) reduces to ρ (X, Y) if there are no ties.
[0087] Assumption 4 is about rank preservation: ρm (Yx, Yx′|Z) =1.
[0088] Assumption 4 is less restrictive than Assumption 3 as it accommodates broader data types of Y. To illustrate, consider a dataset with four individuals where the true values of (Yx, Yx′) are (1, 1) , (2, 1.5) , (2, 1.5) , (3, 2.5) . In this scenario, ∑1≤i<j≤nsign ( (yi, x-yj, x) (yi, x′-yj, x′) =5, resulting in ρ (Yx, Yx′) =5 / 6 and ρm (Yx, Yx′) =5 / (√6-1·√6-1) =1.
[0089] In addition, Assumption 4 also guarantees the identifiability of θ, as shown in Proposition 3.
[0090] Proposition 3 is under Assumption 4: the conclusion in Proposition 1 also holds.
[0091] The following will describe a novel estimation method. Suppose that { (xi, zi, yi) : i=1, ..., N} is asample consisting of N realizations of random variables (X, Z, Y) . For an individual, given its evidence (X=x, Z=z, Y=y) , its counterfactual outcome yx′ is estimated, which is the realization of Yx′ for this individual.
[0092] The counterfactual learning problem is formulated as the following bi-level optimization problem:
[0093] where is the check function, the upper-level optimization is to estimate the quantile of y in the distribution and the lower level optimization is to estimate the conditional quantile function
[0094] for a given Then yx′ is estimated using
[0095] To reveal the limitations of the estimation method above, let the following two equations be the conditional quantile regression functions for Yx and Yx′, respectively:
[0096] By Eq. (4) , yx′ can be expressed as with being the quantile of y in the distribution of i.e., The following Proposition 4 shows the rationale behind employing the check function as the loss for estimating conditional quantiles.
[0097] Proposition 4 states that:
[0098] for any given x; and
[0099]
[0100] There are several potential concerns with the aforementioned estimation method. First, it only fits a single quantile regression model for to obtain estimates of and If the two conditional quantile functions and originate from different models, this method may yield inaccurate estimates. Second, it requires estimating a distinct for each individual before estimating the counterfactual outcome. Additionally, the upper-level optimization does not guarantee a convex problem, which may lead to instability in the bi-level optimization process.
[0101] A simple improvementis to estimate and separately, employing the inverse propensity weighting method. For instance, for the associated loss function is given by:
[0102] where is the propensity score. Similarly, can be defined by replacing x with x′. The estimation procedure for yx′involves four steps. At step (1) , the propensity score is estimated. At step (2) , is estimated by minimizing for a range of candidate values of At step (3) , in the candidate set of that corresponds to the quantile of y in the distribution is identified. At step (4) , yx′ is estimated using where is obtained by minimizing
[0103] Although this four-step estimation method allows and to come from different models, it still requires estimating a different for each individual and can be cumbersome.
[0104] Referring again to FIG. 3, in some embodiments, a first metric may be determined by comparing each observed outcome in the first set of observed outcomes 320 with the target observed outcome 310 of the target observed object 304. The first metric may indicate rank preservation of different outcomes. A second metric may be determined. The second metric may indicate an accumulated difference between the second set of observed outcomes 320 and the target observed outcome 310 of the target observed object 304. Then, the predicted outcome 340 (denoted as Rx’) may be derived based on the first metric and the second metric.
[0105] The first and second metrics may be used in any suitable manner to derive to the predicted outcome 340. In some embodiments, a loss with respect to the predicted outcome 340 may be established based on the first metric and the second metric. The predicted outcome 340 may be optimized by minimizing the loss.
[0106] For example, to address this issue and further refine the estimation, a novel loss function is proposed that directly yields an unbiased estimator of yx′ for an individual with evidence (X=x, Z=z, Y=y) . The proposed ideal loss is constructed as
[0107] where the equation (12) is a function of t and the expectation operator is taken on the random variable of (X, Yx, Yx′) given Z=z . The item is an example of the first metric and the item is an example of the second metric. The proposed estimation method is based on Theorem 1.
[0108] Theorem 1 states that the loss Rx′ (t|x, z, y) is convex with respect to t and is uniquely minimized at t*, where t* is the solution satisfying:
[0109] Theorem 1 implies that, given the evidence (X=x, Z=z, Y=y) for an individual, the counterfactual outcome yx′ (a realization of yx′ for this individual) satisfies yx′=argmintRx′ (t|x, z, y) under Assumption 4. Importantly, the loss Rx′ (t|x, z, y) neither estimates the SCM a priori nor restricts and to stem from the same model, and it does not require explicitly estimating a different quantile value for each individual.
[0110] In some embodiments, in response to an observed outcome in the first set of observed outcomes 320 exceeding the target observed outcome 310 of the target observed object 301, a positive effect may be allocated to the first metric. Alternatively, or in addition, in response to a second observed outcome in the first set of observed outcomes 320 being below the target observed outcome 310 of the target observed object 304, a negative effect may be allocated to the first metric. For example, according to the sign () function in the item of equation (12) , if Yx exceeds y, the value of the sign () function is a predetermined positive value (such as 1) , and if Yx is below y, the value of the sign () function is a predetermined negative value (such as -1) .
[0111] In some embodiments, observed features of the first and second sets of objects may be the same as an observed feature of the target observed object 304. A first expectation over the first set of observed outcomes 320 may be derived based on a comparison result for each observed outcome in the first set of observed outcomes 320. The first expectation may be determined as the first metric. A second expectation over the second set of observed outcomes 330 may be derived based on respective differences between the second set of observed outcomes 330 and the target observed outcome 310. The second expectation may be determined as the second metric. For example, the item is an example of expectation over the first set of observed outcomes 320 and the item is an example of expectation over the second set of observed outcomes 330.
[0112] In some cases, due to the limited number of observed samples, it is difficult to obtain sufficient observed samples with the same observed features as the observed feature of the target observed object. To this end, a smoothing mechanism may be applied such that observed samples with observed features different from the observed feature of the target observed object can be used.
[0113] As an example, to optimize the ideal loss Rx′ (t; x, z, y) , it needs to be estimated first. Specifically, a kernel-smoothing-based estimator is proposed for the ideal loss, which is given as
[0114] where h is a bandwidth (smoothing) parameter, Kh (u) =K (u / h) / h, and K (·) is a symmetric kernel function that satisfies ∫K (u) du=1 and ∫uK (u) dt=1, such as the Epanechnikov kernel and the Gaussian kernel Then yx′ can be estimated by directly minimizing
[0115] In some embodiments, for each first observed outcome in the first set of observed outcomes 320, a first weight to a comparison result for the first observed outcome and the target observed outcome 310 may be applied based on a first observed feature of a first observed object corresponding to the first observed outcome. The first metric is determined based on respective weighted comparison results for the first set of observed outcomes 320. For example, the second item in the equation (14) is an example of the first metric determined based on the respective weighted comparison results.
[0116] In some embodiments, the first weight may be determined based on a first score for the first observed feature and a difference between the first observed feature and a target observed feature of the target observed object 304. The first score may indicate a probability that an object with the first observed feature is assigned with the first treatment. For example, the zk-z in the second item of equation (14) is an example of the difference and the item px (zk) is an example of the first score.
[0117] Alternatively, or in addition, in some embodiments, for each second observed outcome in the second set of observed outcomes 330, a second weight to a difference between the second observed outcome and the predicted outcome 340 may be applied based on a second observed feature of a second observed object corresponding to the second observed outcome. The second metric is determined based on respective weighted differences for the second set of observed outcomes 330. For example, the first item in the equation (14) is an example of the second metric determined based on the respective weighted differences for the second set of observed outcomes 330.
[0118] In some embodiments, the second weight may be determined based on a second score for the second observed feature and a difference between the second observed feature and the observed feature of the target object. The second score may indicate a probability that an object with the second observed feature is assigned with the second treatment. For example, the item zk-z in the first item of equation (14) is an example of the difference between the second observed feature and the observed feature of the target object, and the item px′ (zk) is an example of the second score.
[0119] Some analysis is now described to demonstrate the validity of the proposed smoothing mechanism. Proposition 5 states that if h→0 as N→∞, and the density function of Z is twice differentiable, then
[0120] where means convergence in probability.
[0121] Proposition 5 indicates that is an asymptotically unbiased estimator of Rx′ (t|x, z, y) , demonstrating the validity of the estimator of the ideal loss.
[0122] In some embodiments, the target observed object may comprise each object assigned with the first treatment. A target model based on the observational dataset and the predicted outcome for each observed object assigned with the first treatment may be generated. The target model may be used for determining, from the first treatment and the second treatment, a target treatment for a test object with a test feature.
[0123] According to embodiments of the present disclosure, there is proposed a novel counterfactual learning approach with relaxed identifiability assumptions and theoretically guaranteed estimation methods. On one hand, for identifiability assumptions, the rank preservation assumption is introduced, positing that an individual's factual and counterfactual outcomes have the same rank in the corresponding distribution containing factual and counterfactual outcomes for all individuals with the same features. The identifiability of counterfactual outcomes under the proposed rank preservation assumption is proved, which is shown to be strictly weaker than the monotonicity assumption. Notably, such identifiability results hold even when the homogeneity assumption of the exogenous variables is violated.
[0124] On the other hand, a theoretically guaranteed method is further proposed for unbiased estimation of counterfactual outcomes. The proposed estimation method enjoys several desirable merits. First, it does not necessitate a prior estimation of SCMs and thus relies on fewer assumptions. Second, in contrast to the quantile regression method, the approach neither restricts conditional quantile functions for different counterfactual outcomes to originate from the same model, nor does it require estimating a different quantile value for each individual. Third, the previous learning approaches are enhanced to adopt a convex loss for estimating counterfactual outcomes, which leads to a unique solution.
[0125] FIG. 4 illustrates a flowchart of a process 400 for casual inference in accordance with some embodiments of the present disclosure. The process 400 may be implemented at the computer system 110 of FIG. 1.
[0126] At block 410, the computer system 110 obtains an observational dataset associated with a plurality of treatments and a plurality of observed objects, wherein the observational dataset comprises a plurality of observed samples, an observed sample corresponds to an observed object of the plurality of observed objects assigned with a treatment of the plurality of treatments and comprises an observed feature of the observed object and an observed outcome of the observed object.
[0127] At block 420, the computer system 110, for a target observed object of the plurality of observed objects, determines, from the observational dataset, a first set of observed outcomes of a first set of observed objects assigned with a first treatment of the plurality of treatments, wherein the target observed object is assigned with the first treatment.
[0128] At block 430, the computer system 110 determines, from the observational dataset, a second set of observed outcomes of a second set of observed objects assigned with a second treatment different from the first treatment.
[0129] At block 440, the computer system 110 generates a predicted outcome of the target observed object assigned with the second treatment based on the first set of observed outcomes, the second set of observed outcomes and a target observed outcome of the target observed object.
[0130] In some embodiments, the computer system 110 determines a first metric indicating rank preservation of different outcomes by comparing each observed outcome in the first set of observed outcomes with the target observed outcome of the target observed object; determines a second metric indicating an accumulated difference between the second set of observed outcomes and the target observed outcome of the target observed object; and derives the predicted outcome based on the first metric and the second metric.
[0131] In some embodiments, the computer system 110 establishes a loss with respect to the predicted outcome based on the first metric and the second metric; and optimizes the predicted outcome by minimizing the loss.
[0132] In some embodiments, the computer system 110, in response to an observed outcome in the first set of observed outcomes exceeding the target observed outcome of the target observed object, allocates a positive effect to the first metric. Alternatively, or in addition, the computer system 110, in response to a second observed outcome in the first set of observed outcomes being below the target observed outcome of the target observed object, allocates a negative effect to the first metric.
[0133] In some embodiments, observed features of the first and second sets of objects are the same as an observed feature of the target observed object, and the computer system 110 derives, as the first metric, a first expectation over the first set of observed outcomes based on a comparison result for each observed outcome in the first set of observed outcomes, and the computer system 110 derives, as the second metric, a second expectation over the second set of observed outcomes based on respective differences between the second set of observed outcomes and the target observed outcome.
[0134] In some embodiments, the computer system 110, for each first observed outcome in the first set of observed outcomes, applies a first weight to a comparison result for the first observed outcome and the target observed outcome based on a first observed feature of a first observed object corresponding to the first observed outcome; and determines the first metric based on respective weighted comparison results for the first set of observed outcomes.
[0135] In some embodiments, the first weight is determined based on: a first score for the first observed feature, the first score indicating a probability that an object with the first observed feature is assigned with the first treatment; and a difference between the first observed feature and a target observed feature of the target observed object.
[0136] In some embodiments, the computer system 110, for each second observed outcome in the second set of observed outcomes, applies a second weight to a difference between the second observed outcome and the predicted outcome based on a second observed feature of a second observed object corresponding to the second observed outcome; and determines the second metric based on respective weighted differences for the second set of observed outcomes.
[0137] In some embodiments, the second weight is determined based on: a second score for the second observed feature, the second score indicating a probability that an object with the second observed feature is assigned with the second treatment; and a difference between the second observed feature and the observed feature of the target object.
[0138] In some embodiments, the target observed object comprises each object assigned with the first treatment, and the computer system 110 further generates a target model based on the observational dataset and the predicted outcome for each observed object assigned with the first treatment, wherein the target model is used for determining, from the first treatment and the second treatment, a target treatment for a test object with a test feature.
[0139] In some embodiments, one of the first and second treatments comprises one of applying a specific medical treatment on a subject and not applying the specific medical treatment on the subject, and the other one of the first and second treatments comprises the other one of applying a specific medical treatment on a subject and not applying the specific medical treatment on the subject.
[0140] FIG. 5 shows a block diagram of an apparatus 500 for counterfactual inference in accordance with some embodiments of the present disclosure. The apparatus 500 may be implemented, for example, or included at the computer system 110 of FIG. 1. Various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0141] As shown, the apparatus 500 includes a dataset obtaining module 510 configured to obtain an observational dataset associated with a plurality of treatments and a plurality of observed objects, wherein the observational dataset comprises a plurality of observed samples, an observed sample corresponds to an observed object of the plurality of observed objects assigned with a treatment of the plurality of treatments and comprises an observed feature of the observed object and an observed outcome of the observed object.
[0142] The apparatus 500 further includes a first outcome determination module 520 configured to, for a target observed object of the plurality of observed objects, determine, from the observational dataset, a first set of observed outcomes of a first set of observed objects assigned with a first treatment of the plurality of treatments, wherein the target observed object is assigned with the first treatment.
[0143] The apparatus 500 further includes a second outcome determination module 530 configured to determine, from the observational dataset, a second set of observed outcomes of a second set of observed objects assigned with a second treatment different from the first treatment.
[0144] The apparatus 500 further includes a predicted outcome generation module 540 configured to generate a predicted outcome of the target observed object assigned with the second treatment based on the first set of observed outcomes, the second set of observed outcomes and a target observed outcome of the target observed object.
[0145] The apparatus 500 may further comprises corresponding modules that are configured to perform the operations of the process 400 and other embodiments as described herein.
[0146] FIG. 6 illustrates a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure can be implemented. It would be appreciated that the electronic device 600 shown in FIG. 6 is only an example and should not constitute any restriction on the function and scope of the embodiments described herein. The electronic device 600 may be used, for example, to implement the computer system 110 of FIG. 1. The electronic device 600 may also be used to implement the apparatus 500 of FIG. 5.
[0147] As shown in FIG. 6, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be an actual or virtual processor and can execute various processes according to the programs stored in the memory 620. In a multiprocessor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 600.
[0148] The electronic device 600 typically includes a variety of computer storage medium. Such medium may be any available medium that is accessible to the electronic device 600, including but not limited to volatile and non-volatile medium, removable and non-removable medium. The memory 620 may be volatile memory (for example, a register, cache, a random access memory (RAM) ) , a non-volatile memory (for example, a read-only memory (ROM) , an electrically erasable programmable read-only memory (EEPROM) , a flash memory) or any combination thereof. The storage device 630 may be any removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data (such as training data for training) and can be accessed within the electronic device 600.
[0149] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile, transitory / non-transitory storage medium. Although not shown in FIG. 6, a disk driver for reading from or writing to a removable, non-volatile disk (such as a "floppy disk" ) , and an optical disk driver for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each driver may be connected to the bus (not shown) by one or more data medium interfaces. The memory 620 may include a computer program product 625, which has one or more program modules configured to perform various methods or acts of various embodiments of the present disclosure.
[0150] The communication unit 640 communicates with a further computing device through the communication medium. In addition, functions of components in the electronic device 600 may be implemented by a single computing cluster or multiple computing machines, which can communicate through a communication connection. Therefore, the electronic device 600 may be operated in a networking environment using a logical connection with one or more other servers, a network personal computer (PC) , or another network node.
[0151] The input device 650 may be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as required. The external device, such as a storage device, a display device, etc., communicate with one or more devices that enable users to interact with the electronic device 600, or communicate with any device (for example, a network card, a modem, etc. ) that makes the electronic device 600 communicate with one or more other computing devices. Such communication may be executed via an input / output (I / O) interface (not shown) .
[0152] According to example implementation of the present disclosure, a computer-readable storage medium is provided, on which a computer-executable instruction or computer program is stored, where the computer-executable instructions or the computer program is executed by the processor to implement the method described above. According to example implementation of the present disclosure, a computer program product is also provided. The computer program product is physically stored on a non-transient computer-readable medium and includes computer-executable instructions, which are executed by the processor to implement the method described above.
[0153] Various aspects of the present disclosure are described herein with reference to the flow chart and / or the block diagram of the method, the device, the equipment and the computer program product implemented in accordance with the present disclosure. It would be appreciated that each block of the flowchart and / or the block diagram and the combination of each block in the flowchart and / or the block diagram may be implemented by computer-readable program instructions.
[0154] These computer-readable program instructions may be provided to the processing units of general-purpose computers, special computers or other programmable data processing devices to produce a machine that generates a device to implement the functions / acts specified in one or more blocks in the flow chart and / or the block diagram when these instructions are executed through the processing units of the computer or other programmable data processing devices. These computer-readable program instructions may also be stored in a computer-readable storage medium. These instructions enable a computer, a programmable data processing device and / or other devices to work in a specific way. Therefore, the computer-readable medium containing the instructions includes a product, which includes instructions to implement various aspects of the functions / acts specified in one or more blocks in the flowchart and / or the block diagram.
[0155] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operational steps can be performed on a computer, other programmable data processing apparatus, or other devices, to generate a computer-implemented process, such that the instructions which execute on a computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks in the flowchart and / or the block diagram.
[0156] The flowchart and the block diagram in the drawings show the possible architecture, functions and operations of the system, the method and the computer program product implemented in accordance with the present disclosure. In this regard, each block in the flowchart or the block diagram may represent a part of a module, a program segment or instructions, which contains one or more executable instructions for implementing the specified logic function. In some alternative implementations, the functions marked in the block may also occur in a different order from those marked in the drawings. For example, two consecutive blocks may actually be executed in parallel, and sometimes can also be executed in a reverse order, depending on the function involved. It should also be noted that each block in the block diagram and / or the flowchart, and combinations of blocks in the block diagram and / or the flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by the combination of dedicated hardware and computer instructions.
[0157] Each implementation of the present disclosure has been described above. The above description is example, not exhaustive, and is not limited to the disclosed implementations. Without departing from the scope and spirit of the described implementations, many modifications and changes are obvious to ordinary skill in the art. The selection of terms used in this article aims to best explain the principles, practical application or improvement of technology in the market of each implementation, or to enable other ordinary skill in the art to understand the various embodiments disclosed herein.
Claims
1.A method of counterfactual inference comprising:obtaining an observational dataset associated with a plurality of treatments and a plurality of observed objects, wherein the observational dataset comprises a plurality of observed samples, an observed sample corresponds to an observed object of the plurality of observed objects assigned with a treatment of the plurality of treatments and comprises an observed feature of the observed object and an observed outcome of the observed object;for a target observed object of the plurality of observed objects,determining, from the observational dataset, a first set of observed outcomes of a first set of observed objects assigned with a first treatment of the plurality of treatments, wherein the target observed object is assigned with the first treatment;determining, from the observational dataset, a second set of observed outcomes of a second set of observed objects assigned with a second treatment different from the first treatment; andgenerating a predicted outcome of the target observed object assigned with the second treatment based on the first set of observed outcomes, the second set of observed outcomes and a target observed outcome of the target observed object.2.The method of claim 1, wherein generating a predicted outcome of the target observed object assigned with the second treatment comprises:determining a first metric indicating rank preservation of different outcomes by comparing each observed outcome in the first set of observed outcomes with the target observed outcome of the target observed object;determining a second metric indicating an accumulated difference between the second set of observed outcomes and the target observed outcome of the target observed object; andderiving the predicted outcome based on the first metric and the second metric.3.The method of claim 2, wherein deriving the predicted outcome based on the first metric and the second metric comprises:establishing a loss with respect to the predicted outcome based on the first metric and the second metric; andoptimizing the predicted outcome by minimizing the loss.4.The method of claim 2, wherein determining the first metric comprises at least one of:in response to an observed outcome in the first set of observed outcomes exceeding the target observed outcome of the target observed object, allocating a positive effect to the first metric, orin response to a second observed outcome in the first set of observed outcomes being below the target observed outcome of the target observed object, allocating a negative effect to the first metric.5.The method of claim 2, wherein observed features of the first and second sets of objects are the same as an observed feature of the target observed object, anddetermining the first metric comprises:deriving, as the first metric, a first expectation over the first set of observed outcomes based on a comparison result for each observed outcome in the first set of observed outcomes, andwherein determining the second metric comprises:deriving, as the second metric, a second expectation over the second set of observed outcomes based on respective differences between the second set of observed outcomes and the target observed outcome.6.The method of claim 2, wherein determining the first metric comprises:for each first observed outcome in the first set of observed outcomes, applying a first weight to a comparison result for the first observed outcome and the target observed outcome based on a first observed feature of a first observed object corresponding to the first observed outcome; anddetermining the first metric based on respective weighted comparison results for the first set of observed outcomes.7.The method of claim 6, wherein the first weight is determined based on:a first score for the first observed feature, the first score indicating a probability that an object with the first observed feature is assigned with the first treatment; anda difference between the first observed feature and a target observed feature of the target observed object.8.The method of claim 2, wherein determining the second metric comprises:for each second observed outcome in the second set of observed outcomes, applying a second weight to a difference between the second observed outcome and the predicted outcome based on a second observed feature of a second observed object corresponding to the second observed outcome; anddetermining the second metric based on respective weighted differences for the second set of observed outcomes.9.The method of claim 8, wherein the second weight is determined based on:a second score for the second observed feature, the second score indicating a probability that an object with the second observed feature is assigned with the second treatment; anda difference between the second observed feature and the observed feature of the target object.10.The method of claim 1, wherein the target observed object comprises each object assigned with the first treatment, and the method further comprises:generating a target model based on the observational dataset and the predicted outcome for each observed object assigned with the first treatment, wherein the target model is used for determining, from the first treatment and the second treatment, a target treatment for a test object with a test feature.11.The method of claim 1, wherein one of the first and second treatments comprises one of applying a specific medical treatment on a subject and not applying the specific medical treatment on the subject, andthe other one of the first and second treatments comprises the other one of applying a specific medical treatment on a subject and not applying the specific medical treatment on the subject.12.An electronic device, comprising a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements a method for counterfactual inference, the method comprising:obtaining an observational dataset associated with a plurality of treatments and a plurality of observed objects, wherein the observational dataset comprises a plurality of observed samples, an observed sample corresponds to an observed object of the plurality of observed objects assigned with a treatment of the plurality of treatments and comprises an observed feature of the observed object and an observed outcome of the observed object;for a target observed object of the plurality of observed objects,determining, from the observational dataset, a first set of observed outcomes of a first set of observed objects assigned with a first treatment of the plurality of treatments, wherein the target observed object is assigned with the first treatment;determining, from the observational dataset, a second set of observed outcomes of a second set of observed objects assigned with a second treatment different from the first treatment; andgenerating a predicted outcome of the target observed object assigned with the second treatment based on the first set of observed outcomes, the second set of observed outcomes and a target observed outcome of the target observed object.13.The device of claim 12, wherein generating a predicted outcome of the target observed object assigned with the second treatment comprises:determining a first metric indicating rank preservation of different outcomes by comparing each observed outcome in the first set of observed outcomes with the target observed outcome of the target observed object;determining a second metric indicating an accumulated difference between the second set of observed outcomes and the target observed outcome of the target observed object; andderiving the predicted outcome based on the first metric and the second metric.14.The device of claim 13, wherein deriving the predicted outcome based on the first metric and the second metric comprises:establishing a loss with respect to the predicted outcome based on the first metric and the second metric; andoptimizing the predicted outcome by minimizing the loss.15.The device of claim 13, wherein determining the first metric comprises at least one of:in response to an observed outcome in the first set of observed outcomes exceeding the target observed outcome of the target observed object, allocating a positive effect to the first metric, orin response to a second observed outcome in the first set of observed outcomes being below the target observed outcome of the target observed object, allocating a negative effect to the first metric.16.Th device of claim 13, wherein observed features of the first and second sets of objects are the same as an observed feature of the target observed object, anddetermining the first metric comprises:deriving, as the first metric, a first expectation over the first set of observed outcomes based on a comparison result for each observed outcome in the first set of observed outcomes, andwherein determining the second metric comprises:deriving, as the second metric, a second expectation over the second set of observed outcomes based on respective differences between the second set of observed outcomes and the target observed outcome.17.The device of claim 13, wherein determining the first metric comprises:for each first observed outcome in the first set of observed outcomes, applying a first weight to a comparison result for the first observed outcome and the target observed outcome based on a first observed feature of a first observed object corresponding to the first observed outcome; anddetermining the first metric based on respective weighted comparison results for the first set of observed outcomes.18.The device of claim 13, wherein determining the second metric comprises:for each second observed outcome in the second set of observed outcomes, applying a second weight to a difference between the second observed outcome and the predicted outcome based on a second observed feature of a second observed object corresponding to the second observed outcome; anddetermining the second metric based on respective weighted differences for the second set of observed outcomes.19.The device of claim 12, wherein the target observed object comprises each object assigned with the first treatment, and the method further comprises:generating a target model based on the observational dataset and the predicted outcome for each observed object assigned with the first treatment, wherein the target model is used for determining, from the first treatment and the second treatment, a target treatment for a test object with a test feature.20.A computer readable storage medium, having a computer program stored thereon which, upon execution by an electronic device, causes the device to perform a method for counterfactual inference, the method comprising:obtaining an observational dataset associated with a plurality of treatments and a plurality of observed objects, wherein the observational dataset comprises a plurality of observed samples, an observed sample corresponds to an observed object of the plurality of observed objects assigned with a treatment of the plurality of treatments and comprises an observed feature of the observed object and an observed outcome of the observed object;for a target observed object of the plurality of observed objects,determining, from the observational dataset, a first set of observed outcomes of a first set of observed objects assigned with a first treatment of the plurality of treatments, wherein the target observed object is assigned with the first treatment;determining, from the observational dataset, a second set of observed outcomes of a second set of observed objects assigned with a second treatment different from the first treatment; andgenerating a predicted outcome of the target observed object assigned with the second treatment based on the first set of observed outcomes, the second set of observed outcomes and a target observed outcome of the target observed object.