Training of a model for causal inference

The DPPL method for causal inference addresses the limitations of existing reward balancing methods by decomposing the optimization problem into subproblems, using strong ignorability and missing mechanism assumptions, to achieve accurate and optimal trade-offs between short-term and long-term rewards.

WO2026102592A1PCT designated stage Publication Date: 2026-05-21BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING YOUZHUJU NETWORK TECH CO LTD
Filing Date
2024-11-12
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing methods for balancing multiple short-term and long-term rewards in causal inference are limited, as they can only find optimal solutions in convex regions and fail when rewards are interrelated, leading to sub-optimal outcomes.

Method used

A machine learning model is trained using a decomposition-based Pareto policy learning (DPPL) method that generates candidate model parameter sets to minimize or decrease both short-term and long-term reward values under specific constraints, employing assumptions of strong ignorability and missing mechanism to identify and estimate causal effects, and decomposes the multi-objective optimization problem into subproblems for Pareto optimal solutions.

Benefits of technology

This approach effectively balances short-term and long-term rewards, providing diverse Pareto optimal policies that improve the accuracy of predicted treatments by addressing confounding bias and missing data issues, ensuring optimal trade-offs in non-convex regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024131657_21052026_PF_FP_ABST
    Figure CN2024131657_21052026_PF_FP_ABST
Patent Text Reader

Abstract

A solution for training a model for causal inference is provided. A method comprises: determining, using a machine learning model, a predicted treatment for an object based on feature information of the object (410); obtaining a first set of short-term outcomes and a second set of long-term outcomes for the object under the predicted treatment (420); determining respective first reward values for the first set of short-term outcomes, and respective second reward values for the second set of long-term outcomes (430); generating a plurality of candidate model parameter sets for the machine learning model based on a plurality of reward optimization objectives conditioned on a model parameter set of the machine learning model (440); and updating the model parameter set of the machine learning model based on the plurality of candidate model parameter sets (450).
Need to check novelty before this filing date? Find Prior Art

Description

TRAINING OF A MODEL FOR CAUSAL INFERENCEField

[0001] The disclosed example embodiments relate generally to machine learning and, more particularly, to a method, apparatus, device and computer readable storage medium for training of a model for causal inference.Background

[0002] In the growing area of machine learning for causal inference, various practical problems can be solved. For example, estimating the heterogeneous causal effects of an intervention (e.g., a medical treatment) on an important outcome (e.g., health status) of different individuals is a fundamental problem in a variety of influential research areas.Summary

[0003] In a first aspect of the present disclosure, there is provided a method for training a model for causal inference. The method comprises: determining, using a machine learning model, a predicted treatment for an object based on feature information of the object; obtaining a first set of short-term outcomes and a second set of long-term outcomes for the object under the predicted treatment; determining respective first reward values for the first set of short-term outcomes based on the predicted treatment and the first set of short-term outcomes, and respective second reward values for the second set of long-term outcomes based on the predicted treatment and the second set of long-term outcomes; generating a plurality of candidate model parameter sets for the machine learning model based on a plurality of reward optimization objectives conditioned on a model parameter set of the machine learning model, the plurality of reward optimization objectives being configured to minimize or decrease the respective first reward values and the respective second reward values under a plurality of constraints, wherein a constraint is constructed based on respective importances of the respective first reward values and the respective second reward values; and updating the model parameter set of the machine learning model based on the plurality of candidate model parameter sets.

[0004] In a second aspect of the present disclosure, there is provided an apparatus for training a model for causal inference. The apparatus comprises: a predicted treatment determining module configured to determine, using a machine learning model, a predicted treatment for an object based on feature information of the object; an outcome obtaining module configured to obtain a first set of short-term outcomes and a second set of long-term outcomes for the object under the predicted treatment; a reward value determining module configured to determine respective first reward values for the first set of short-term outcomes based on the predicted treatment and the first set of short-term outcomes, and respective second reward values for the second set of long-term outcomes based on the predicted treatment and the second set of long-term outcomes; a candidate model parameter set generating module configured to generate a plurality of candidate model parameter sets for the machine learning model based on a plurality of reward optimization objectives conditioned on a model parameter set of the machine learning model, the plurality of reward optimization objectives being configured to minimize or decrease the respective first reward values and the respective second reward values under a plurality of  constraints, wherein a constraint is constructed based on respective importances of the respective first reward values and the respective second reward values; and a model parameter set updating module configured to update the model parameter set of the machine learning model based on the plurality of candidate model parameter sets.

[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions executable by the at least one processing unit. The instructions, upon execution by the at least one processing unit, cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The medium stores a computer program which, when executed by a processor, causes the method of the first aspect to be implemented.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product comprises a computer program which, when executed by a processor, causes the method of the first aspect to be implemented.

[0008] It would be appreciated that the content described in the Summary section of the present invention is neither intended to identify key or essential features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily envisaged through the following description.Brief Description of the Drawings

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent in combination with the accompanying drawings and with reference to the following detailed description. In the drawings, the same or similar reference symbols refer to the same or similar elements, where:

[0010] FIG. 1 illustrates a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;

[0011] FIG. 2 illustrates a schematic diagram of training the machine learning model in accordance with some embodiments of the present disclosure;

[0012] FIG. 3 illustrates a decomposition-based policy learning algorithm in accordance with some embodiments of the present disclosure;

[0013] FIG. 4 illustrates a flowchart of a process for training a model for causal inference in accordance with some embodiments of the present disclosure;

[0014] FIG. 5 shows a block diagram of an apparatus for training a model for causal inference in accordance with some embodiments of the present disclosure; and

[0015] FIG. 6 illustrates a block diagram of an electronic device in which one or more embodiments of the  present disclosure can be implemented.Detailed Description

[0016] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it would be appreciated that the present disclosure may be implemented in various forms and should not be interpreted as limited to the embodiments described herein. On the contrary, these embodiments are provided for a more thorough and complete understanding of the present disclosure. It would be appreciated that the drawings and embodiments of the present disclosure are only for the purpose of illustration and are not intended to limit the scope of protection of the present disclosure.

[0017] In the description of the embodiments of the present disclosure, the term "including" and similar terms would be appreciated as open inclusion, that is, "including but not limited to" . The term "based on" would be appreciated as "at least partially based on" . The term "one embodiment" or "the embodiment" would be appreciated as "at least one embodiment" . The term "some embodiments" would be appreciated as "at least some embodiments" . Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the matching degree between various data. For example, the above matching degree can be obtained based on various technical solutions currently available and / or to be developed in the future.

[0018] It will be appreciated that the data involved in this technical proposal (including but not limited to the data itself, data acquisition or use) shall comply with the requirements of corresponding laws, regulations and relevant provisions.

[0019] It will be appreciated that before using the technical solution disclosed in each embodiment of the present disclosure, users should be informed of the type, the scope of use, the use scenario, etc. of the personal information involved in the present disclosure in an appropriate manner in accordance with relevant laws and regulations, and the user’s authorization should be obtained.

[0020] For example, in response to receiving an active request from a user, a prompt message is sent to the user to explicitly prompt the user that the operation requested operation by the user will need to obtain and use the user's personal information. Thus, users may select whether to provide personal information to the software or the hardware such as an electronic device, an application, a server or a storage medium that perform the operation of the technical solution of the present disclosure according to the prompt information.

[0021] As an optional but non-restrictive implementation, in response to receiving the user's active request, the method of sending prompt information to the user may be, for example, a pop-up window in which prompt information may be presented in text. In addition, pop-up windows may also contain selection controls for users to choose “agree” or “disagree” to provide personal information to electronic devices.

[0022] It will be appreciated that the above notification and acquisition of user authorization process are only schematic and do not limit the implementations of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0023] As used herein, the term "model" can learn a correlation between respective inputs and outputs from training data, so that a corresponding output can be generated for a given input after training is completed. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural networks model is an example of a deep learning-based model. As used herein, "model" may also be referred to as "machine learning model" , "learning model" , "machine learning network" , or "learning network" , and these terms are used interchangeably herein.

[0024] “Neural networks” are a type of machine learning network based on deep learning. Neural networks are capable of processing inputs and providing corresponding outputs, typically comprising input and output layers and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically comprise many hidden layers, thereby increasing the depth of the network. The layers of neural networks are sequentially connected so that the output of the previous layer is provided as input to the latter layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network comprises one or more nodes (also known as processing nodes or neurons) , each of which processes input from the previous layer.

[0025] Usually, machine learning can roughly comprise three stages, namely training stage, test stage, and application stage (also known as inference stage) . During the training stage, a given model can be trained using a large scale of training data, iteratively updating parameter values until the model can obtain consistent inference from the training data that meets the expected objective. Through the training, the model can be considered to learn the correlation between input and output (also known as input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the test stage, test inputs are applied to the trained model to test whether the model can provide correct outputs, thereby determining the performance of the model. In the application stage, the model can be used to process actual inputs and determine corresponding outputs based on the parameter values obtained from training.

[0026] FIG. 1 illustrates a block diagram of an example environment 100 in which various embodiments of the present disclosure may be implemented. In the environment 100 of FIG. 1, a computer system 110 applies a machine learning model 105 to perform causal inference. The machine learning model 105 may sometimes be referred to as a causal inference model. The machine learning model 105 is configured to process feature information 112 of an object to generate a predicted output 114 for the object.

[0027] The feature information 112 input to the machine learning model 105 and the predicted output 114 generated from the machine learning model 105 may be designed according to the tasks to be performed. The feature information 112 may include features in different aspects. For example, if the object is a patient, the feature information may include physical status, physiological reaction, disease type and the like. The machine learning model 105 may generate a treatment plan (as an example of the predicted output 114) for the patient.

[0028] In some embodiments, the feature information 112 of the object may include physiological features and psychological features of a patient. The predicted output 114 may include applying a specific medical treatment on the patient or not applying the specific medical treatment on the object.

[0029] In some embodiments, the feature information 112 of the object may include age, gender, historical behavior and the like of a user. The predicted output 114 may include recommending an item to the user or not recommending the item to the user.

[0030] In some embodiments, the feature information 112 of the object may include behavioral features and social features of a group of customer. The predicted output 114 may include implementing incentive strategies to the group of customers or not implementing incentive strategies to the group of customers.

[0031] In addition to the medical treatment, recommending and consumption scenarios, there may be various of other scenarios where effects of different treatments are evaluated on individuals. Estimation of individual treatment effect has been the key for individual decision making in economics, healthcare, education, etc.

[0032] In FIG. 1, the computer system 110 may include any computing system with computing capability, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices may include any type of mobile terminals, fixed terminals, or portable terminals, including mobile phones, desktop computers, laptops, netbooks, tablets, media computers, multimedia tablets, or any combination of the aforementioned, including accessories and peripherals of these devices or any combination thereof. Servers include but are not limited to mainframe, edge computing nodes, computing devices in cloud environment, etc.

[0033] It should be understood that the structure and function of each element in the environment 100 is described for illustrative purposes only and does not imply any limitations on the scope of the present disclosure.

[0034] As mentioned above, machine learning for causal inference is applied in various areas. Different rewards (e.g., short-term and long-term) may be assigned to different causal outcomes. Learning an optimal policy for balancing multiple short-term and long-term rewards holds extensive applications across various domains. For instance, content providers may optimize recommendations to avoid short-term clickbait strategies, ensuring sustained user engagement and revenue growth. Information technology companies may design web pages catering to immediate user preferences while enhancing long-term engagement and satisfaction. People may want to explore the effects of early childhood interventions on lifetime earnings, seeking optimal policies (e.g., class size) maximizing short-term test scores and long-term earnings of a child simultaneously. Policymakers may improve job training program design, considering both immediate income impacts and subsequent employment status improvements. Medical practitioners may refine drug prescriptions, considering short-term alleviation and long-term outcomes in chronic diseases. People may want to optimize incentive strategies to positively influence customer behavior in both short-term consumption and long-term consumption.

[0035] The optimal policy may be learned for balancing multiple short-term and long-term rewards. Despite the importance of balancing multiple short-term and long-term rewards, policy learning methods in this area remain largely unexplored. Recent works employ a linear weighting method to achieve this goal. It combines multiple rewards into a single surrogate reward by weighted summation, which is optimized to learn the optimal policy. However, this strategy has several limitations. First, it can only find optimal solutions in convex regions of objective space and cannot obtain the optimal solutions in non-convex regions. Second, it achieves the optimal  solution only when the rewards are independent of each other. When some of the rewards are interrelated, it can only achieve sub-optimal solutions. Consequently, although the linear weighting method is easy to implement, the optimality of its solution cannot be guaranteed when balancing multiple objectives.

[0036] To address at least some of the above issues, embodiments of the present disclosure propose an improved solution for training a model for causal inference. In this solution, a machine learning model is used to determine a predicted treatment for an object based on feature information of the object. A first set of short-term outcomes and a second set of long-term outcomes for the object under the predicted treatment are obtained. Respective first reward values for the first set of short-term outcomes are determined based on the predicted treatment and the first set of short-term outcomes, and respective second reward values for the second set of long-term outcomes are determined based on the predicted treatment and the second set of long-term outcomes. A plurality of candidate model parameter sets for the machine learning model are generated based on a plurality of reward optimization objectives conditioned on a model parameter set of the machine learning model. The plurality of reward optimization objectives are configured to minimize or decrease the respective first reward values and the respective second reward values under a plurality of constraints, and a constraint is constructed based on respective importances of the respective first reward values and the respective second reward values. The model parameter set of the machine learning model is updated based on the plurality of candidate model parameter sets.

[0037] With these embodiments of the present disclosure, a plurality of reward optimization objectives with corresponding constraints are constructed based on respective importances of the respective first reward values and the respective second reward values, and for each reward optimization objective, a candidate model parameter set can be generated to update the machine learning model. In this way, the balancing of reward values for short-term outcomes and reward values for long-term outcomes may be improved.

[0038] In order to better understand embodiments of the present disclosure, notations to delineate short-term and long-term causal effects are introduced first. Let A denote the binary treatment indicator, where A= 1 represents the treated group (i.e., the group assigned with a treatment) and A=0 represents the control group (i.e., the group not assigned with a treatment) . X represents the features observed (e.g., features of an object) ,  and represent the vector of short-term and long-term outcomes, respectively. A duration between a time point when the short-term outcome is obtained and a time point when the treatment A is assigned is less than a duration between a time point when a long-term outcome is obtained and a time point when the treatment A is assigned. Both short-term and long-term outcomes are observed after the treatment A, and associations among them may exist.

[0039] Utilizing a potential outcome framework, S (a) = (S1 (a) , …, SI (a) ) and Y (a) = (Y1 (a) , …, YJ (a) ) for a=0, 1 are denoted as the potential short-term and long-term outcomes under treatment A=a, respectively. It is assumed that larger short-term and long-term outcomes are preferable. The observed short-term and long-term outcomes S and Y correspond to the potential outcomes of the actual treatment, that is, S=S (A) and Y=Y (A) .

[0040] In real-world applications, long-term outcomes often suffer from missing due to prolonged follow-up periods and budget constraints. In contrast, collecting short-term outcomes is more manageable. Therefore, it is presumed that all short-term outcomes S are observable, while long-term outcomes Y may be subject to missing. Let R= (R1, …, RJ) ∈ {0, 1} J denote the indicator for observing the long-term outcome Y, where Rj=1 indicates that Yj is observed and Rj=0 indicates that Yj is missing. The missingness of Y may lead to identifiability and estimation problems.

[0041] Example embodiments of the present disclosure will be described in the following. FIG. 2 illustrates a schematic diagram of training the machine learning model 105 in accordance with some embodiments of the present disclosure. As shown in FIG. 2, the machine learning model 105 for causal inference is used to determine a predicted treatment 205 for an object (e.g., a patient) based on feature information 112 of the object. In some embodiments, a policy (sometimes also referred to as the machine learning model 105) may be denoted as π∶ χ→ {0, 1} and the policy maps from the individual context (also referred to as the feature information 112) X=x to a treatment space {0, 1} including the predicted treatment 205. In an example, the feature information 112 may include physiological features and psychological features of object.

[0042] Depending on the feature information of the object to be predicted, the predicted treatment 205 may include a first treatment of applying a specific medical treatment on the object. Taking the treatment space being {0, 1} as an example, the first treatment may be assigned with the value 1. In other cases, the predicted treatment 205 may include a second treatment of not applying the specific medical treatment on the object. Taking the treatment space being {0, 1} as an example, the second treatment may be assigned with the value 0.

[0043] After the predicted treatment is determined, a first set of short-term outcomes 210 and a second set of long-term outcomes 215 for the object under the predicted treatment 205 are obtained. In some examples, if the predicted treatment 205 is the first treatment (i.e., applying a specific medical treatment on the object) , a first set of short-term outcomes 210 and a second set of long-term outcomes 215 for the object who is assigned with the specific medical treatment may be obtained.

[0044] In some embodiments, a duration between a time point when a respective one of the first set of short-term outcomes 210 is obtained and a time point when the predicted treatment is assigned to the object is less than a duration between a time point when a respective one of the second set of long-term outcomes 215 is obtained and the time point when the predicted treatment is assigned to the object. In an example, a first duration between a time point when a short-term outcome is obtained and a time point when the predicted treatment is assigned to the object may be 7 days. A second duration between a time point when a long-term outcome is obtained and the time point when the predicted treatment is assigned to the object may be 2 months. The first duration is less than the second duration.

[0045] After the first set of short-term outcomes 210 and second set of long-term outcomes 215 are obtained, respective first reward values 220 for the first set of short-term outcomes 210 are determined based on the predicted treatment 205 and the first set of short-term outcomes 210. Respective second reward values 225 for the second set of long-term outcomes 215 are determined based on the predicted treatment 205 and the second  set of long-term outcomes 215. In some embodiments, for a given policy π (θ) =π (X, θ) parameterized by θ, the policy values (also referred to as reward values) for the i-th short-term outcome Si and the j-th long-term outcome Yj are defined as follows:

[0046] which are the i-th short-term reward value and the j-th long-term reward value induced by the policy π (θ) . π (θ) represents the predicted treatment 205 output by the machine learning model 105 and θ represents a model parameter set 230 of the machine learning model 105. If the predicted treatment 205 is the first treatment, the value of π (θ) may be 1. Otherwise, the value of π (θ) may be 0. Si (1) represents the i-th short-term outcome under the first treatment, Si (0) represents the i-th short-term outcome under the second treatment, Yj (1) represents the j-th long-term outcome under the first treatment and Yj (0) represents the j-th long-term outcome under the second treatment.

[0047] In some embodiments, the respective first reward values 220 may be negatively correlated with the first set of short-term outcomes 210 and the respective second reward values 225 may be negatively correlated with the second set of long-term outcomes 215. In some examples, maximization problems may be converted to minimization problems. Let and represents the i-th first reward value for i-th short-term outcome, where the i-th first reward value is negatively correlated with i-th short-term outcome.  represents the j-th second reward value for the j-th long-term outcome, where the j-th second reward value is negatively correlated with the j-th long-term outcome.

[0048] After the respective first reward values 220 and the respective second reward values 225 are determined, a plurality of candidate model parameter sets 235 for the machine learning model 105 are generated based on a plurality of reward optimization objectives conditioned on a model parameter set 230 of the machine learning model 105. The plurality of reward optimization objectives are configured to minimize or decrease the respective first reward values 220 and the respective second reward values 225 under a plurality of constraints. A constraint is constructed based on respective importances of the respective first reward values and the respective second reward values. In some examples, the trade-off among multiple correlated short-term rewards (also referred to as first reward values) and long-term rewards (also referred to as second reward values) may be formulated as a multi-objective optimization (MOP) problem given by

[0049] where M= I+J and the symbol means “denoted as” . There is no single solution that can simultaneously optimize all objectives (i.e., respective first reward values 220 and respective second reward values 225) in the MOP problem represented by Eq. (2) and thus the MOP problem may be decomposed into a plurality of subproblems (also referred to as the plurality of reward optimization objectives) . In addition, the concept of Pareto optimality may be employed to define the optimal solutions for the MOP problem. In the following, some definitions about Pareto optimality will be introduced.

[0050] Pareto dominance refers to that for two points θ1 and θ2, θ1 dominates θ2 if and only if  and Pareto optimality refers to θ* is a Pareto optimal point if there is no other solution that dominates θ*. Pareto optimality refers to a condition where improving one objective comes at the expense of worsening other objectives. The collection of Pareto optimal solutions is called the Pareto set. The goal of the present disclosure is to derive the set of Pareto optimal solutions or Pareto optimal policies (e.g., the plurality of candidate model parameter sets 235) , each of them providing a distinct optimal trade-off among all objectives.

[0051] The long-term and short-term rewards are causal parameters that cannot be identified without imposing causal assumptions. Therefore, before seeking the Pareto optimal solutions for balancing multiple long-term and short-term rewards, it is necessary to consider the identification and estimation of long-term and short-term rewards. Embodiments of the present disclosure are based on assumption 1 and assumption 2 in the following. Assumption 1 (also referred to as the assumption of strong ignorability) includes (a)  and (b)  Assumption 1 (a) suggests that, given the feature X, treatment assignment A is independent of the potential outcomes S (a) and Y (a) . This implies that confounding bias between the treatment A and the short or long-term outcomes (S (a) , Y (a) ) may be eliminated by conditioning on X. Assumption 1 (b) ensures that for the subpopulation of X=x, units (e.g., the object) with both A = 1 and A = 0 exist. These assumptions are widely used in causal inference.

[0052] In addition to confounding bias, the selection bias induced by the missingness of long-term outcomes may also need to be addressed. Therefore, assumption (2) is invoked in the following.

[0053] For a=0, 1 and j=1, …, J, assumption 2 (also referred to as the assumption of missing mechanism of long-term outcome ) includes (a)  and (b)  Assumption 2 (a) may be reformulated as which means that Rj relies only on the observed variables (X, S, A) . This assumption also ensures that This implies that available data can be utilized to draw conclusions about the missing long-term outcome. Assumption 2 (b) assumes that the long-term outcome for each unit has a non-zero probability of being observed. Assumptions 1 and 2 ensures the identifiability of  and Therefore, for i=1, …, I and j=1, …, J, under assumption 1, the reward value of the i-th short-term outcome is identifiable and under assumptions 1 and 2, the reward value of the j-th long-term outcome is identifiable.

[0054] Because there is no single solution that can simultaneously optimize all objectives in the MOP problem, in some embodiments, a decomposition-based Pareto policy learning (DPPL) method may be introduced to decompose the MOP problem into a plurality of subproblems (also referred to as the plurality of reward optimization objectives) with corresponding constraints, and then a set of Pareto optimal policies (e.g., the plurality of candidate model parameter sets 235) may be obtained by solving these subproblems in parallel. The constraint is constructed based on respective importances of the respective first reward values and the respective second reward values.

[0055] In some embodiments, the respective importances may be represented by a preference vector, the number of dimensions of the preference vector is equal to a total number of the respective first reward values and the respective second reward values, and each dimension of the preference vector indicates an importance of a corresponding first reward value or second reward value. For obtaining the Pareto optimal policy for balancing M short-term and long-term objectives, a set of K preference vectors {u1, u2, …, uk} in are determined, where K is equal to M. Each element (i.e., each dimension) of a preference vector specifies the importance of the corresponding short-term or long-term reward. The MOP problem may be decomposed into subproblems based on the set of preference vectors. For each preference vector uk, the corresponding subproblem is given as:

[0056] where the constraint means that objective space of the subproblem is restricted in the subregion Ωk, which is defined by Geometrically speaking, Ωk represents the set of v that forms the smallest acute angle with uk, which means that the optimal solution of the subproblem may be obtained by only searching the subregion. The preference vectors divide the objective space into different subregions.

[0057] In some embodiments, the plurality of constraints may be constructed based on a set of preference vectors. For each of the plurality of reward optimization objectives, an initial model parameter set for the reward optimization objective may be determined. In order to solve the subproblem, a reasonable initial solution (also referred to as the initial model parameter set) θ0 may be found.

[0058] In some embodiments, in order to determine the initial model parameter set, a random model parameter set for the machine learning model 105 may be generated. The random model parameter set may be updated based on a second descent direction and the step size until a second predetermined condition is satisfied, to obtain the initial model parameter set. The second predetermined condition indicates that a result of the reward optimization objective conditioned on the initial model parameter set to be located within a set of vectors for the corresponding constraint. A similarity between the set of vectors and a target preference vector in the set of preference vectors corresponding to the constraint is greater than a similarity between the set of vectors and preference vectors other than the target preference vector in the set of preference vectors. In an example, a random solution (i.e., the random model parameter set) θr may be first generated randomly in a full decision space and then the random solution may be iteratively updated with the rule where ηr represents a step size for updating the model parameter set 230. For a given adescent direction (also called as the second descent direction)  may be updated by solving the following:

[0059] The second descent direction may be derived from the constraint of Eq. (3) , where is index set of all activated constraints, which means is not in the  subregion Ωk. Eq. (4) aims to find the descent direction for each iteration t and then obtain the initial solution (also referred to as the initial model parameter set) θ0 such that is in the subregion Ωk.  represents a result of the reward optimization objective conditioned on the initial model parameter set θ0 and Ωl represents a set of vectors for a corresponding constraint. The initial model parameter set θ0 causes the result of the reward optimization objective to be located in the set of vectors. According to the definition of the subregion Ωk, an angle formed by the set of vectors (denoted as v) and a target preference vector (denoted as uk) in the set of preference vectors is smaller than an angle formed by the set of vectors and preference vectors (denoted as uk′) other than the target preference vector in the set of preference vectors. Because the angle formed by two vectors represent the similarity between the two vectors, the similarity between the set of vectors and the target preference vector in the set of preference vectors is greater than the similarity between the set of vectors and preference vectors other than the target preference vector in the set of preference vectors.

[0060] The process of determining the initial model parameter set (denoted as θ0) is restricted in a subregion of the subproblem represented by Eq. (3) , and once a feasible solution is found or a predetermined number of iterations is reached, the process stops. After the initial model parameter set is determined, the initial model parameter set may be updated based on a first descent direction and the step size until a first predetermined condition is satisfied, to obtain a candidate model parameter set. The first predetermined condition indicates that a result of the reward optimization objective conditioned on the candidate model parameter set is less than results of reward optimization objectives conditioned on model parameter sets other than the candidate model parameter set. The first descent direction may be derived based on the constraint corresponding to the reward optimization objective. In some examples, the initial model parameter set may be iteratively updated with the rule to obtain the candidate model parameter set. A descent direction (also called as the first descent direction)  may be obtained by:

[0061] The first descent direction  may be derived from the constraint of Eq. (3) , where and the threshold ε is a slack variable used to deal with the solutions near the constraint boundary. Eq. (5) may further be transformed into a dual problem which will greatly reduce the dimension of decision space. Based on the Karush-Kuhn-Tucker (KKT) conditions,  may be obtained. Therefore, the dual problem may be given as:

[0062] where λm≥0 and βk′≥0 represent the Lagrange multipliers for the linear inequality constraints.

[0063] In some examples, it is assumed that (dt, αt) is the solution to the t-th iteration of the subproblem. If θt is Pareto optimal restricted on the subregion Ωλ, then and αt=0. At the t-th iteration, no direction (dt=0) may simultaneously improve the performance for all objectives (e.g., reduce or minimize the respective first reward values and the respective second reward values) , confirming that the solution θt satisfies Pareto optimality. If θt is not Pareto optimal restricted on the subregion Ωk, then

[0064] As indicated by Eq. (7) , if θt does not meet Pareto optimality, then the descent direction dt≠0 serves as the descent direction for all objectives, such that the solution of the next iteration is closer to the Pareto optimal solution. Therefore, Pareto optimal solutions for each subproblem may be attained using the update rule

[0065] In some embodiments, by iteratively updating the initial solution with the first descent direction, an optimal solution (also referred to as candidate model parameter set) θ* for the subproblem (also referred to as the reward optimization objective) may be obtained. The candidate model parameter set is Pareto optimal restricted on the subregion Ωk and the result of the reward optimization objective conditioned on the candidate model parameter set is less than results of reward optimization objectives conditioned on model parameter sets other than the candidate model parameter set. By solving all subproblems (also referred to as the plurality of reward optimization objectives) , a diverse set of Pareto optimal solutions or policies (also referred to as the plurality of candidate model parameter sets) confined to different subregions may be acquired, even when the multiple objectives are correlated. In this way, the candidate model parameter set which is the optimal model parameter set under a constraint may be determined and thus the training effects of the machine learning model may be improved based in the candidate model parameter set.

[0066] The above process of determining the initial model parameter set and the candidate model parameter set may be summarized in FIG. 3, which illustrates a DPPL algorithm in accordance with some embodiments of the present disclosure. As shown in FIG. 3, for a subproblem, at step 305, the random model parameter set is generated. At step 310, the initial model parameter set is obtained based on the random model parameter set using a gradient-based method. Ath step 315, the candidate model parameter set is obtained based on the initial model parameter set using the gradient-based method.

[0067] After the plurality of candidate model parameter sets are generated 235, the model parameter set 230 of the machine learning model 105 may be updated based on the plurality of candidate model parameter sets. With the updated model parameter set, the machine learning model 105 may determine a more accurate predicted  treatment for the object, and reward values for short-term outcomes and reward values for long-term outcomes may be better balanced.

[0068] In some embodiments, the plurality of reward optimization objectives may be configured to minimize or decrease a weighted sum determined by weighting the respective first reward values and the respective second reward values under the plurality of constraints with a plurality of predetermined weights. In an example, in order to find the optimal policy for balancing short-term rewards (also referred to as first reward values) and long-term rewards (also referred to as second reward values) , a linear weighting method may be adopted, which may be formulated as follows:

[0069] where ωm represents a predetermined weight for the m-th objective (e.g., a first reward value or a second reward value) . The linear weighting method simply combines multiple objectives into a single surrogate objective through weighted summation.

[0070] As described above, decomposing the MOP problem into a plurality of subproblems requires a set of pre-specified preference vectors, posing challenges in practical applications where selecting suitable preference vectors is non-trivial. To mitigate this problem, embodiments of the present disclosure provide a method for decision-makers to select appropriate preference vectors by theoretically establishing the connection between the DPPL method and a ε-constraint problem. In some embodiments, a plurality of upper limits for the respective first reward values and the respective second reward values are obtained respectively. In an example, the plurality of upper limits may be set by a user in advance. The ε-constraint problem may be defined as follows:

[0071] where εm represents a pre-specified threshold (also referred to as an upper limit) . Compared to the MOP problem and the linear weighting objective, an advantage of the ε-constraint problem is its interpretation on the threshold εm, which represents the maximum acceptable value (i.e., the acceptable worst-case scenario) for the m-th objective. If the connection between ε= (ε1, …, εM) and the preference vector uk may be established, powerful guidance for choosing appropriate preference vectors may be provided.

[0072] In some embodiments, a first association between the plurality of upper limits and the plurality of predetermined weights for the respective first reward values and the respective second reward values, and a second association between the plurality of predetermined weights and the respective importances (represented by a preference vector) may be obtained. For the preference vector uk= (uk1, …, ukm) representing respective importances of the respective first reward values and the respective second reward values, the weights ω= (ω1, …, ωM) , and the thresholds ε, the connection (also referred to as the first association) between ε and ω may be given as follows:

[0073] Eq. (10) shows how to estimate the threshold ε for given weights ω, where τm (X) represents the conditional average causal effects for the m-th short or long-term outcome, which may be represented as follows:

[0074] In Eq. (10) ,   is the indicator function, and hm (X) may be represented as follows:

[0075] The connection between ω and uk may be given as follows:

[0076] Eq. (13) shows how to assign weights ω via preference vectors uk, where λm and βk′are defined in Eq. (6) , and is defined in Eq. (4) .

[0077] After the first association and the second association are obtained, a third association between the plurality of upper limits and the respective importances may be determined. Then, the respective importances (represented by a preference vector) may be determined based on the third association and the plurality of upper limits. In an example, a link (also referred to as the third association) between the preference vector uk and the threshold ε may be established through weights ω involving multiple long-term and short-term objectives. In this way, for the subproblem determined by the preference vector uk, an intuitive interpretation of the preference vector uk may be offered based on the threshold (also referred to as the plurality of upper limits) ε.

[0078] With these embodiments of the present disclosure, decision-makers may be assisted in better understanding and selecting preference vectors in practical applications. In practice, a set of preference vectors {u1, u2, …, uk} in may be initially pre-specified, then the weights ω corresponding to each preference vector uk may be derived through Eq. (13) , and finally the obtained weight ω may be substituted into Eq. (10) to calculate the threshold ε. Leveraging the intuitive interpretability of the threshold ε, decision-makers may select the appropriate preference vectors according to their specific requirements. Furthermore, guidance for specifying the threshold ε in the ε-constraint problem in Eq. (9) may be provided. Inappropriate selection of ε for this problem may result in an empty feasible region, yielding empty solutions. By utilizing a set of preference vectors, some reasonable choices of ε may be efficiently screened out and the cumbersome trial-and-error process of testing different ε may be reduced.

[0079] FIG. 4 illustrates a flowchart of a process 400 for training a model for causal inference in accordance with some embodiments of the present disclosure. The process 400 may be implemented at the computer system 110 of FIG. 1.

[0080] At block 410, the computer system 110 determines, using a machine learning model, a predicted treatment for an object based on feature information of the object.

[0081] At block 420, the computer system 110 obtains a first set of short-term outcomes and a second set of long-term outcomes for the object under the predicted treatment.

[0082] At block 430, the computer system 110 determines respective first reward values for the first set of short-term outcomes based on the predicted treatment and the first set of short-term outcomes, and respective second reward values for the second set of long-term outcomes based on the predicted treatment and the second set of long-term outcomes.

[0083] At block 440, the computer system 110 generates a plurality of candidate model parameter sets for the machine learning model based on a plurality of reward optimization objectives conditioned on a model parameter set of the machine learning model. The plurality of reward optimization objectives are configured to minimize or decrease the respective first reward values and the respective second reward values under a plurality of constraints, where a constraint is constructed based on respective importances of the respective first reward values and the respective second reward values.

[0084] At block 450, the computer system 110 updates the model parameter set of the machine learning model based on the plurality of candidate model parameter sets.

[0085] In some embodiments, the respective importances are represented by a preference vector, a number of dimensions of the preference vector is equal to a total number of the respective first reward values and the respective second reward values, and each dimension of the preference vector indicates an importance of a corresponding first reward value or second reward value.

[0086] In some embodiments, the plurality of constraints are constructed based on a set of preference vectors and wherein generating the plurality of candidate model parameter sets comprises: for each of the plurality of reward optimization objectives, determining an initial model parameter set for the reward optimization objective; and updating the initial model parameter set based on a first descent direction and a step size for updating the model parameter set until a first predetermined condition is satisfied, to obtain a candidate model parameter set, wherein the first predetermined condition indicates that a result of the reward optimization objective conditioned on the candidate model parameter set is less than results of reward optimization objectives conditioned on model parameter sets other than the candidate model parameter set and the first descent direction is derived based on the constraint corresponding to the reward optimization objective.

[0087] In some embodiments, determining the initial model parameter set comprises: generating a random model parameter set for the machine learning model; and updating the random model parameter set based on a second descent direction and the step size until a second predetermined condition is satisfied, to obtain the initial model parameter set, wherein the second predetermined condition indicates that a result of the reward optimization objective conditioned on the initial model parameter set is located within a set of vectors for a corresponding constraint, a similarity between the set of vectors and a target preference vector in the set of preference vectors being greater than a similarity between the set of vectors and preference vectors other than  the target preference vector in the set of preference vectors.

[0088] In some embodiments, the plurality of reward optimization objectives are configured to minimize or decrease a weighted sum determined by weighting the respective first reward values and the respective second reward values under the plurality of constraints with a plurality of predetermined weights.

[0089] In some embodiments, the process 400 further comprises obtaining a plurality of upper limits for the respective first reward values and the respective second reward values, respectively; obtaining a first association between the plurality of upper limits and the plurality of predetermined weights for the respective first reward values and the respective second reward values, and a second association between the plurality of predetermined weights and the respective importances; determining a third association between the plurality of upper limits and the respective importances based on the first association and the second association; and determining the respective importances based on the third association and the plurality of upper limits.

[0090] In some embodiments, a duration between a time point when a respective one of the first set of short-term outcomes is obtained and a time point when the predicted treatment is assigned to the object is less than a duration between a time point when a respective one of the second set of long-term outcomes is obtained and the time point when the predicted treatment is assigned to the object.

[0091] In some embodiments, the predicted treatment comprises one of: a first treatment of applying a specific medical treatment on the object, or a second treatment of not applying the specific medical treatment on the object.

[0092] In some embodiments, the respective first reward values are negatively correlated with the first set of short-term outcomes, and wherein the respective second reward values are negatively correlated with the second set of long-term outcomes.

[0093] FIG. 5 shows a block diagram of an apparatus 500 for training a model for causal inference in accordance with some embodiments of the present disclosure. The apparatus 500 may be implemented, for example, or included at the computer system 110 of FIG. 1. Various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0094] As shown, the apparatus 500 includes a predicted treatment determining module 510 configured to determine, using a machine learning model, a predicted treatment for an object based on feature information of the object; an outcome obtaining module 520 configured to obtain a first set of short-term outcomes and a second set of long-term outcomes for the object under the predicted treatment; a reward value determining module 530 configured to determine respective first reward values for the first set of short-term outcomes based on the predicted treatment and the first set of short-term outcomes, and respective second reward values for the second set of long-term outcomes based on the predicted treatment and the second set of long-term outcomes; a candidate model parameter set generating module 540 configured to generate a plurality of candidate model parameter sets for the machine learning model based on a plurality of reward optimization objectives conditioned on a model parameter set of the machine learning model, the plurality of reward optimization objectives being configured to minimize or decrease the respective first reward values and the respective second reward values  under a plurality of constraints, wherein a constraint is constructed based on respective importances of the respective first reward values and the respective second reward values; and a model parameter set updating module 550 configured to update the model parameter set of the machine learning model based on the plurality of candidate model parameter sets.

[0095] In some embodiments, the respective importances are represented by a preference vector, a number of dimensions of the preference vector is equal to a total number of the respective first reward values and the respective second reward values, and each dimension of the preference vector indicates an importance of a corresponding first reward value or second reward value.

[0096] In some embodiments, the plurality of constraints are constructed based on a set of preference vectors and the candidate model parameter set generating module 540 is further configured to, for each of the plurality of reward optimization objectives, determine an initial model parameter set for the reward optimization objective; and update the initial model parameter set based on a first descent direction and a step size for updating the model parameter set until a first predetermined condition is satisfied, to obtain a candidate model parameter set, wherein the first predetermined condition indicates that a result of the reward optimization objective conditioned on the candidate model parameter set is less than results of reward optimization objectives conditioned on model parameter sets other than the candidate model parameter set and the first descent direction is derived based on the constraint corresponding to the reward optimization objective.

[0097] In some embodiments, the candidate model parameter set generating module 540 is further configured to generate a random model parameter set for the machine learning model; and updating the random model parameter set based on a second descent direction and the step size until a second predetermined condition is satisfied, to obtain the initial model parameter set, wherein the second predetermined condition indicates that a result of the reward optimization objective conditioned on the initial model parameter set is located within a set of vectors for a corresponding constraint, a similarity between the set of vectors and a target preference vector in the set of preference vectors being greater than a similarity between the set of vectors and preference vectors other than the target preference vector in the set of preference vectors.

[0098] In some embodiments, the plurality of reward optimization objectives are configured to minimize or decrease a weighted sum determined by weighting the respective first reward values and the respective second reward values under the plurality of constraints with a plurality of predetermined weights.

[0099] In some embodiments, the apparatus 500 further includes an importances obtaining module configured to obtain a plurality of upper limits for the respective first reward values and the respective second reward values, respectively; obtain a first association between the plurality of upper limits and the plurality of predetermined weights for the respective first reward values and the respective second reward values, and a second association between the plurality of predetermined weights and the respective importances; determine a third association between the plurality of upper limits and the respective importances based on the first association and the second association; and determine the respective importances based on the third association and the plurality of upper limits.

[0100] In some embodiments, a duration between a time point when a respective one of the first set of short-term outcomes is obtained and a time point when the predicted treatment is assigned to the object is less than a duration between a time point when a respective one of the second set of long-term outcomes is obtained and the time point when the predicted treatment is assigned to the object.

[0101] In some embodiments, the predicted treatment comprises one of: a first treatment of applying a specific medical treatment on the object, or a second treatment of not applying the specific medical treatment on the object.

[0102] In some embodiments, the respective first reward values are negatively correlated with the first set of short-term outcomes, and wherein the respective second reward values are negatively correlated with the second set of long-term outcomes.

[0103] FIG. 6 illustrates a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure can be implemented. It would be appreciated that the electronic device 600 shown in FIG. 6 is only an example and should not constitute any restriction on the function and scope of the embodiments described herein. The electronic device 600 may be used, for example, to implement the computer system 110 of FIG. 1. The electronic device 600 may also be used to implement the apparatus 500 of FIG. 5.

[0104] As shown in FIG. 6, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be an actual or virtual processor and can execute various processes according to the programs stored in the memory 620. In a multiprocessor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 600.

[0105] The electronic device 600 typically includes a variety of computer storage medium. Such medium may be any available medium that is accessible to the electronic device 600, including but not limited to volatile and non-volatile medium, removable and non-removable medium. The memory 620 may be volatile memory (for example, a register, cache, a random access memory (RAM) ) , a non-volatile memory (for example, a read-only memory (ROM) , an electrically erasable programmable read-only memory (EEPROM) , a flash memory) or any combination thereof. The storage device 630 may be any removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data (such as training data for training) and can be accessed within the electronic device 600.

[0106] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage medium. Although not shown in FIG. 6, a disk driver for reading from or writing to a removable, non-volatile disk (such as a "floppy disk" ) , and an optical disk driver for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each driver may be connected to the bus (not shown) by one or more data medium interfaces. The memory 620 may include a computer program product 625, which  has one or more program modules configured to perform various methods or acts of various embodiments of the present disclosure.

[0107] The communication unit 640 communicates with a further computing device through the communication medium. In addition, functions of components in the electronic device 600 may be implemented by a single computing cluster or multiple computing machines, which can communicate through a communication connection. Therefore, the electronic device 800 may be operated in a networking environment using a logical connection with one or more other servers, a network personal computer (PC) , or another network node.

[0108] The input device 650 may be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as required. The external device, such as a storage device, a display device, etc., communicate with one or more devices that enable users to interact with the electronic device 600, or communicate with any device (for example, a network card, a modem, etc. ) that makes the electronic device 600 communicate with one or more other computing devices. Such communication may be executed via an input / output (I / O) interface (not shown) .

[0109] According to example implementation of the present disclosure, a computer-readable storage medium is provided, on which a computer-executable instruction or computer program is stored, where the computer-executable instructions or the computer program is executed by the processor to implement the method described above. According to example implementation of the present disclosure, a computer program product is also provided. The computer program product is physically stored on a non-transient computer-readable medium and includes computer-executable instructions, which are executed by the processor to implement the method described above.

[0110] Various aspects of the present disclosure are described herein with reference to the flow chart and / or the block diagram of the method, the device, the equipment and the computer program product implemented in accordance with the present disclosure. It would be appreciated that each block of the flowchart and / or the block diagram and the combination of each block in the flowchart and / or the block diagram may be implemented by computer-readable program instructions.

[0111] These computer-readable program instructions may be provided to the processing units of general-purpose computers, special computers or other programmable data processing devices to produce a machine that generates a device to implement the functions / acts specified in one or more blocks in the flow chart and / or the block diagram when these instructions are executed through the processing units of the computer or other programmable data processing devices. These computer-readable program instructions may also be stored in a computer-readable storage medium. These instructions enable a computer, a programmable data processing device and / or other devices to work in a specific way. Therefore, the computer-readable medium containing the instructions includes a product, which includes instructions to implement various aspects of the functions / acts specified in one or more blocks in the flowchart and / or the block diagram.

[0112] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operational steps can be performed on a computer, other programmable data processing apparatus, or other devices, to generate a computer-implemented process, such that the instructions which execute on a computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks in the flowchart and / or the block diagram.

[0113] The flowchart and the block diagram in the drawings show the possible architecture, functions and operations of the system, the method and the computer program product implemented in accordance with the present disclosure. In this regard, each block in the flowchart or the block diagram may represent a part of a module, a program segment or instructions, which contains one or more executable instructions for implementing the specified logic function. In some alternative implementations, the functions marked in the block may also occur in a different order from those marked in the drawings. For example, two consecutive blocks may actually be executed in parallel, and sometimes can also be executed in a reverse order, depending on the function involved. It should also be noted that each block in the block diagram and / or the flowchart, and combinations of blocks in the block diagram and / or the flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by the combination of dedicated hardware and computer instructions.

[0114] Each implementation of the present disclosure has been described above. The above description is an example, not exhaustive, and is not limited to the disclosed implementations. Without departing from the scope and spirit of the described implementations, many modifications and changes are obvious to ordinary skill in the art. The selection of terms used in this article aims to best explain the principles, practical application or improvement of technology in the market of each implementation, or to enable other ordinary skill in the art to understand the various embodiments disclosed herein.

Claims

1.A method for training a model for causal inference, comprising:determining, using a machine learning model, a predicted treatment for an object based on feature information of the object;obtaining a first set of short-term outcomes and a second set of long-term outcomes for the object under the predicted treatment;determining respective first reward values for the first set of short-term outcomes based on the predicted treatment and the first set of short-term outcomes, and respective second reward values for the second set of long-term outcomes based on the predicted treatment and the second set of long-term outcomes;generating a plurality of candidate model parameter sets for the machine learning model based on a plurality of reward optimization objectives conditioned on a model parameter set of the machine learning model, the plurality of reward optimization objectives being configured to minimize or decrease the respective first reward values and the respective second reward values under a plurality of constraints, wherein a constraint is constructed based on respective importances of the respective first reward values and the respective second reward values; andupdating the model parameter set of the machine learning model based on the plurality of candidate model parameter sets.2.The method of claim 1, wherein the respective importances are represented by a preference vector, a number of dimensions of the preference vector is equal to a total number of the respective first reward values and the respective second reward values, and each dimension of the preference vector indicates an importance of a corresponding first reward value or second reward value.3.The method of claim 2, wherein the plurality of constraints are constructed based on a set of preference vectors and wherein generating the plurality of candidate model parameter sets comprises:for each of the plurality of reward optimization objectives,determining an initial model parameter set for the reward optimization objective; andupdating the initial model parameter set based on a first descent direction and a step size for updating the model parameter set until a first predetermined condition is satisfied, to obtain a candidate model parameter set, wherein the first predetermined condition indicates that a result of the reward optimization objective conditioned on the candidate model parameter set is less than results of reward optimization objectives conditioned on model parameter sets other than the candidate model parameter set and the first descent direction is derived based on the constraint corresponding to the reward optimization objective.4.The method of claim 3, wherein determining the initial model parameter set comprises:generating a random model parameter set for the machine learning model; andupdating the random model parameter set based on a second descent direction and the step size until a second predetermined condition is satisfied, to obtain the initial model parameter set, wherein the second predetermined condition indicates that a result of the reward optimization objective conditioned on the initial model parameter set is located within a set of vectors for a corresponding constraint, a similarity between the set of vectors and a target preference vector in the set of preference vectors being greater than a similarity between the set of vectors and preference vectors other than the target preference vector in the set of preference vectors.5.The method of claim 1, wherein the plurality of reward optimization objectives are configured to minimize or decrease a weighted sum determined by weighting the respective first reward values and the respective second reward values under the plurality of constraints with a plurality of predetermined weights.6.The method of claim 5, further comprising:obtaining a plurality of upper limits for the respective first reward values and the respective second reward values, respectively;obtaining a first association between the plurality of upper limits and the plurality of predetermined weights for the respective first reward values and the respective second reward values, and a second association between the plurality of predetermined weights and the respective importances;determining a third association between the plurality of upper limits and the respective importances based on the first association and the second association; anddetermining the respective importances based on the third association and the plurality of upper limits.7.The method of claim 1, wherein a duration between a time point when a respective one of the first set of short-term outcomes is obtained and a time point when the predicted treatment is assigned to the object is less than a duration between a time point when a respective one of the second set of long-term outcomes is obtained and the time point when the predicted treatment is assigned to the object.8.The method of claim 1, wherein the predicted treatment comprises one of:a first treatment of applying a specific medical treatment on the object, ora second treatment of not applying the specific medical treatment on the object.9.The method of claim 1, wherein the respective first reward values are negatively correlated with the first set of short-term outcomes, andwherein the respective second reward values are negatively correlated with the second set of long-term outcomes.10.An electronic device, comprising:at least one processing unit; andat least one memory coupled to the at least one processing unit and storing instructions executable by the at least one processing unit, the instructions, upon execution by the at least one processing unit, causing the device to perform acts comprising:determining, using a machine learning model, a predicted treatment for an object based on feature information of the object;obtaining a first set of short-term outcomes and a second set of long-term outcomes for the object under the predicted treatment;determining respective first reward values for the first set of short-term outcomes based on the predicted treatment and the first set of short-term outcomes, and respective second reward values for the second set of long-term outcomes based on the predicted treatment and the second set of long-term outcomes;generating a plurality of candidate model parameter sets for the machine learning model based on a plurality of reward optimization objectives conditioned on a model parameter set of the machine learning model, the plurality of reward optimization objectives being configured to minimize or decrease the respective first reward values and the respective second reward values under a plurality of constraints, wherein a constraint is constructed based on respective importances of the respective first reward values and the respective second reward values; andupdating the model parameter set of the machine learning model based on the plurality of candidate model parameter sets.11.The electronic device of claim 10, wherein the respective importances are represented by a preference vector, a number of dimensions of the preference vector is equal to a total number of the respective first reward values and the respective second reward values, and each dimension of the preference vector indicates an importance of a corresponding first reward value or second reward value.12.The electronic device of claim 11, wherein the plurality of constraints are constructed based on a set of preference vectors and wherein generating the plurality of candidate model parameter sets comprises:for each of the plurality of reward optimization objectives,determining an initial model parameter set for the reward optimization objective; andupdating the initial model parameter set based on a first descent direction and a step size for updating the model parameter set until a first predetermined condition is satisfied, to obtain a candidate model parameter set, wherein the first predetermined condition indicates that a result of the reward optimization objective conditioned on the candidate model parameter set is less than reward results of optimization objectives conditioned on model parameter sets other than the candidate model parameter set and the first descent direction is derived based on the constraint corresponding to the reward optimization objective.13.The electronic device of claim 12, wherein determining the initial model parameter set comprises:generating a random model parameter set for the machine learning model; andupdating the random model parameter set based on a second descent direction and the step size until a second predetermined condition is satisfied, to obtain the initial model parameter set, wherein the second predetermined condition indicates that a result of the reward optimization objective conditioned on the initial model parameter set is located within a set of vectors for a corresponding constraint, a similarity between the set of vectors and a target preference vector in the set of preference vectors being greater than a similarity between the set of vectors and preference vectors other than the target preference vector in the set of preference vectors.14.The electronic device of claim 10, wherein the plurality of reward optimization objectives are configured to minimize or decrease a weighted sum determined by weighting the respective first reward values and the respective second reward values under the plurality of constraints with a plurality of predetermined weights.15.The electronic device of claim 14, the acts further comprising:obtaining a plurality of upper limits for the respective first reward values and the respective second reward values, respectively;obtaining a first association between the plurality of upper limits and the plurality of predetermined weights for the respective first reward values and the respective second reward values, and a second association between the plurality of predetermined weights and the respective importances;determining a third association between the plurality of upper limits and the respective importances based on the first association and the second association; anddetermining the respective importances based on the third association and the plurality of upper limits.16.The electronic device of claim 10, wherein a duration between a time point when a respective one of the first set of short-term outcomes is obtained and a time point when the predicted treatment is assigned to the object is less than a duration between a time point when a respective one of the second set of long-term outcomes is obtained and the time point when the predicted treatment is assigned to the object.17.The electronic device of claim 10, wherein the predicted treatment comprises one of:a first treatment of applying a specific medical treatment on the object, ora second treatment of not applying the specific medical treatment on the object.18.The electronic device of claim 10, wherein the respective first reward values are negatively correlated with the first set of short-term outcomes, andwherein the respective second reward values are negatively correlated with the second set of long-term outcomes.19.A computer-readable storage medium, having a computer program stored thereon which, upon execution by an electronic device, causes the device to perform the method according to any of claims 1 to 9.20.A computer program product being embodied on a computer-readable medium and comprising computer-executable instructions which are executed by a processor to perform the method according to any of claims 1 to 9.