Strategy optimization method, apparatus and device, and readable storage medium
By introducing attribute correlation information and transformation factors in the covariate offset scenario, combining propensity scores and sampling scores to optimize the reward estimation of the strategy model, the problem of insufficient accuracy and generalization ability of target domain strategy optimization in strategy transfer learning is solved, and more accurate strategy estimation and better model performance are achieved.
Patent Information
- Application Number
- CN202510514679.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-08
AI Technical Summary
In the covariate offset scenario, strategy transfer learning faces identification challenges, and the existing technology is difficult to accurately optimize the strategy model in the target domain, especially in the case of scarcity of resources and large distribution differences, the existing methods lack generalization ability.
By introducing attribute correlation information and transformation factors, the strategy model is optimized based on the causal relationship, and the correlation information of the source data set and the target data set are used, combined with tendency scores and sampling scores, a semi-parameter effective reward estimation method is constructed, and the rewards of the strategy model are optimized to update the strategy model.
It improves the accuracy and generalization ability of the strategy model in the target domain, can more accurately estimate rewards and generate strategies that are close to the theoretical optimality, and improves the model's performance in the target domain.
Smart Images

Figure CN120450080A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and readable storage media for policy optimization. Background Art
[0002] In the context of covariate shift, transfer learning has been widely studied for its application in prediction tasks and models. Simultaneously, research on policy learning under covariate shift is also underway. Policy learning aims to identify which individuals should receive treatment, therapy, or intervention based on their characteristics by maximizing rewards. This approach has broad application prospects in fields such as recommender systems, precision medicine, and reinforcement learning. Summary of the Invention
[0003] In a first aspect of the present disclosure, a policy optimization method is provided. The method includes: determining attribute association information based on a source data set corresponding to a first group of objects, the source data set including attributes of each object in the first group of objects, an indication of whether to apply a target action to the object, and observation results related to the object, the attribute association information indicating associations between object observation data obtained according to the source data set and the object attributes; determining a transformation factor for transforming the attribute association information from the source data set to a target data set based on the number of objects in the first group and the number of objects in a second group corresponding to the target data set, the target data set including attributes of each object in the second group of objects; determining a reward for a policy model for sample objects in the first group of objects and the second group of objects based on sample decisions of whether to apply the target action to the sample objects, the attribute association information, and the transformation factor, the sample decisions being generated using the policy model; and updating the policy model using the reward.
[0004] In a second aspect of the present disclosure, a policy optimization apparatus is provided. The apparatus includes: an attribute association information determination module configured to determine attribute association information based on a source dataset corresponding to a first group of objects, the source dataset including attributes of each object in the first group, an indication of whether a target action is applied to the object, and observations related to the object, the attribute association information indicating associations between object observation data obtained from the source dataset and the object attributes; a transformation factor determination module configured to determine a transformation factor for transforming the attribute association information from the source dataset to a target dataset based on the number of objects in the first group and the number of objects in a second group corresponding to a target dataset, the target dataset including attributes of each object in the second group; a reward determination module configured to determine, for a sample object in the first and second groups of objects, a reward for a policy model based on a sample decision of whether to apply the target action to the sample object, the attribute association information, and the transformation factor, the sample decision being generated using the policy model; and a policy model update module configured to update the policy model using the reward.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.
[0007] In a fifth aspect of the present disclosure, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, the method of the first aspect is implemented.
[0008] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram illustrating an example environment in which embodiments of the present disclosure can be implemented;
[0011] Figure 2A schematic diagram illustrating an architecture for policy optimization according to some embodiments of the present disclosure is shown;
[0012] Figure 3 A flowchart of a method for policy optimization according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A schematic structural block diagram of an apparatus for policy optimization according to some embodiments of the present disclosure is shown; and
[0014] Figure 5 A block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented is shown. DETAILED DESCRIPTION
[0015] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0016] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.
[0017] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.
[0018] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0019] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0020] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium and other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.
[0021] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0022] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0023] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.
[0024] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs. It typically includes an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the previous layer.
[0025] Generally speaking, machine learning can be roughly divided into three stages: the training stage, the testing stage, and the application stage (also known as the inference stage). During the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also known as input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to test whether the model can provide correct outputs, thereby determining the model's performance. In the application stage, the model can be used to process actual inputs based on the parameter values obtained through training to determine the corresponding outputs.
[0026] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. Environment 100 includes an electronic device 110, which can invoke a model 120 to perform policy optimization tasks. Although shown as being deployed within electronic device 110, model 120 can be deployed entirely or partially on electronic device 110 or on a remote device. Model 120 can include a machine learning-based policy model or a policy network.
[0027] Electronic device 110 receives dataset 101 and dataset 102. Dataset 101 corresponds to one group of subjects, while dataset 102 corresponds to another group of subjects. Such subjects include, but are not limited to, patients and robots. For sample subjects in both groups, electronic device 110 may determine at least one parameter for training model 120, and then use this at least one parameter to train model 120.
[0028] The electronic device 110 can be any type of device with computing capabilities, including a terminal device or a server device. The terminal device can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. The server device can include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like.
[0029] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.
[0030] In many real-world scenarios, labeled data is often scarce due to budget constraints and time-consuming collection processes, which significantly limits the generalization ability of the model. For example, in medical research, collecting labeled data involves extensive clinical trials and follow-up periods, which is both expensive and time-consuming. In the field of autonomous driving, obtaining labeled data requires manually annotating large amounts of sensor data, which is both laborious and expensive. Similar problems exist in the field of robotics. To address such problems and enhance the performance of models in unlabeled target domains, one of the active research areas is transfer learning, which aims to improve the performance of the target learner on the target domain by transferring knowledge contained in different but related source domains.
[0031] As briefly mentioned above, research on policy learning has begun in the context of covariate shift. Unlike transfer learning for predictive models, policy transfer faces identification challenges due to its counterfactual nature. Rather than directly predicting outcomes based on observed data, policy transfer requires considering what happens when different actions are applied, which complicates the process.
[0032] In an embodiment of the present disclosure, an improved scheme for policy optimization is provided. The goal of the scheme is to use a dataset from a source domain (also referred to as a source dataset) and a dataset from a target domain (also referred to as a target dataset) to optimize the policy of a target dataset. The source dataset includes, for example, the attributes of each individual (also referred to as an object), the actions applied to the individual, and the related observations, while the target dataset contains only the attributes of the individuals. It can be assumed that the source dataset satisfies the non-confounding and overlap assumptions, while imposing fewer restrictions on the target dataset. Significant differences in the attribute distributions between the source and target datasets (also referred to as covariate shift) are allowed, while assuming that the conditional distributions of potential outcomes for a given covariate (e.g., object attributes) are the same.
[0033] In this solution, attribute association information is determined based on a source data set corresponding to a first group of objects. The source data set includes attributes of each object in the first group of objects, an indication of whether a target action is applied to the object, and observation results related to the object. The attribute association information indicates the association between object observation data obtained based on the source data set and the object attributes. The electronic device determines a transformation factor for transforming the attribute association information from the source data set to the target data set based on the number of objects in the first group and the number of objects in the second group corresponding to the target data set. The target data set includes attributes of each object in the second group of objects. For sample objects in the first group of objects and the second group of objects, the electronic device determines a reward for a policy model based on a sample decision of whether to apply the target action to the sample object, the attribute association information, and the transformation factor. The sample decision is generated using the policy model. The electronic device updates the policy model using the reward.
[0034] In this way, by introducing reasonable identifiability assumptions and effective estimation methods, the policy under object attribute offset is optimized from the perspective of causality. The present disclosure can more accurately estimate rewards and produce a policy close to the theoretically optimal policy.
[0035] The following will describe in detail some exemplary embodiments of the present disclosure through theoretical analysis and with reference to the examples in the accompanying drawings. It should be understood that some embodiments are described using the medical field as an example, while other embodiments are described using robotic application scenarios. However, the present disclosure is not limited to these. The solutions proposed in this disclosure can be applied to any suitable application scenario or field, such as recommendation systems, precision medicine, medical research, artificial intelligence, reinforcement learning, etc.
[0036] Figure 2 FIG2 shows a schematic diagram of a policy optimization architecture 200 according to some embodiments of the present disclosure. The architecture 200 may be implemented in Figure 1 The architecture 200 involves a source dataset 201, a target dataset 202, and a policy model 220. The source dataset 201 and the target dataset 202 can be respectively considered as Figure 1 The example of dataset 101 and dataset 102 in FIG, and the policy model 220 can be regarded as Figure 1 The following will refer to the example of model 120 in FIG. Figure 1 Provide a description.
[0037] Source dataset 201 (e.g., represented as ) corresponds to the first group of objects. The target data set 202 (e.g., represented as ) corresponds to the second group of objects. In some embodiments, the first group of objects may include a first group of robots operating in a first environment, and the second group of objects may include a second group of robots operating in a second environment different from the first environment. The policy model 220 can be configured to provide a decision on whether to apply a planned action to the robot. That is, in this scenario, a robot policy model optimized in one environment can be migrated to another environment. For example, the real data observed by the robot operating in Building A is migrated to the robot operating in Building B. In this way, the optimization efficiency of the robot policy can be greatly reduced.
[0038] The source dataset 201 includes attributes of each object in the first group of objects (e.g., represented as X), an indication of whether a target action (e.g., represented as A) is applied to the object (e.g., applied or not applied), and an observation result related to the object (e.g., represented as Y). For example, the source dataset 201 comes from the medical field and includes attributes of a group of patients (i.e., objects), an indication of whether treatment is applied to the patients (i.e., application of the target action) (applied or not applied), and treatment results related to the patients (i.e., observation results). For another example, the source dataset 201 comes from the field of robotic applications and includes attributes of a group of robots (e.g., surrounding static obstacles, surrounding dynamic obstacles, production date, model, version, etc.), an indication of whether a planned action is applied to the robot, and observation results related to the robot after the planned action is applied.
[0039] The electronic device 110 determines attribute association information 204 based on the source dataset 201. The attribute association information 204 indicates the association between the object observation data obtained from the source dataset 201 and the object attributes. For example, the attribute association information 204 may indicate the association between the observation data of the robot's behavior and its version after the robot performs a task.
[0040] In some embodiments, the attribute association information 204 may include first association information between the result and the attribute, for example, represented as μ1(X). The first association information μ1(X) may indicate how the first average observation result of the objects in the first group of objects changes with the object attribute X when the target action A is applied, for example, how the average task completion degree of the robot changes with the version when a certain action is applied.
[0041] Additionally or alternatively, the attribute association information 204 may include second association information between the result and the attribute (e.g., represented as μ0(X)). The second association information μ0(X) may indicate how the second average observation result of the objects in the first group of objects changes with the object attribute X when the target action A is not applied, for example, how the average task completion of the robot changes with the version when a certain action is not applied.
[0042] Additionally or alternatively, the attribute association information 204 may include third association information between the decision and the attribute (also referred to as a propensity score, for example, represented as e1(X)). The third association information e1(X) may indicate how the probability of the target action A being applied to an object in the first group of objects changes as the object attribute X changes.
[0043] Additionally or alternatively, the attribute association information 204 may include fourth association information between the dataset and the attribute (also referred to as a sampling score, for example, represented as s(X)). The fourth association information s(X) may indicate how the probability of the object corresponding to the source dataset varies with the object attribute.
[0044] The target dataset 202 includes the attributes (i.e., X) of each object in the second set of objects. Based on the number of objects in the first set and the number of objects in the second set, the electronic device 110 determines a transformation factor 206 (e.g., denoted as q) for transforming the attribute association information 204 from the source dataset 201 to the target dataset 202.
[0045] A sample decision 210 is generated using the policy model 220. For a sample object 208 (e.g., denoted as i) in the first and second groups of objects, the electronic device 110 determines a reward 212 (e.g., denoted as R(π)) for the policy model 220 based on the sample decision 210 (e.g., denoted as π(X)) regarding whether to apply the target action to the sample object 208, the attribute association information 204, and the transformation factor 206. The reward 212 is then used to update the policy model 220 (also referred to as a policy or policy network, denoted as π, for example).
[0046] Combination of the above Figure 2 An example architecture 200 is described. The potential outcome framework in causal inference is used below to construct a strategy for transforming a source dataset 201 to a target dataset 202.
[0047] Assuming target action The binary indicator indicates whether the target action is applied to the object (e.g., treatment is applied to the patient), for example, A=1 indicates that the object receives treatment, and A=0 indicates that the object does not receive treatment. This random vector represents the p-dimensional attributes of the subject measured before receiving treatment. Denote the observation of interest. Assume that a large number of observations are expected. In the framework of potential observations, let Y(a) denote the observations that would be obtained if the target action A is set to a (for ) is the potential outcome that will be observed when . According to the consistency assumption, the observed outcome Y satisfies Y=Y(A)=AY(1)+(1-A)Y(0).
[0048] For example, there are two datasets involved: the source dataset and target dataset Source dataset Including the first group of objects, the target dataset Includes a second set of objects. That is, the two data sets each include representative samples. Let G∈{0,1} represent the indicator of the data source, for example, G=1 represents the source domain and G=0 represents the target domain. The observed data is represented as follows:
[0049]
[0050] Among them, the source dataset Include attribute X of each of n1 objectsi , whether to apply target action A to the object i The indication and the observation Y associated with the object i Target dataset Include attribute X for each of n0 objects i .
[0051] This is due to the scarcity of results or labeled data, which is common in real life. For example, in medical research, patient characteristics are observed, but long-term follow-up is required to obtain results. Let n = n0 + n1. Table 1 shows the structure of the observed data, where "√" indicates observed data and "×" indicates unobserved data.
[0052] Table 1
[0053]
[0054] Assuming complete data (including source dataset and target dataset Data in){(X i ,A i ,Y i (0),Y i (1),G i ), i=1,...,n} are independent and have The same distribution, then and represent the source domain and target domain respectively. Expressed as The expected operator, and let the transformation factor
[0055] We can formulate the goal of learning the optimal policy in the target domain. Specifically, let Represents mapping the object attribute X=x to the treatment space The sample decision π(X) is a set of sample decisions. The sample decision π(X) includes, for example, a treatment rule that determines whether the subject receives treatment (A=1) or does not receive treatment (A=0). For a given sample decision π(X) applied to the target domain, the average reward is defined as follows:
[0056]
[0057] The goal of this disclosure is to obtain the optimal strategy π * , which is defined as where Π is a pre-specified policy classification. For example, π(X) can be modeled using learnable parameters θ via methods such as logistic regression or multilayer perceptrons, where each value of θ corresponds to a different policy.
[0058] The optimal strategy to maximize Equation (1) has an obvious form. Let is the conditional average treatment effect (CATE) in the target domain, then
[0059]
[0060] The last equation follows from the law of iterated expectation.
[0061] The optimal strategy contained in Lemma 1 is as follows:
[0062]
[0063] where max π It is to take the maximum value among all possible strategies without being restricted to π.
[0064] For an object represented by X=x in the target domain, Lemma 1 states that the strategy of receiving treatment (A=1) should be based on the symbol τ(x). Recommend treatments to subjects who are expected to receive positive benefits, thereby optimizing the overall reward in the target domain. The target policy π is equal to the optimal policy in Lemma 1 Otherwise, they may not be equal, and the difference between them is a systematic error caused by the finite assumption space of π.
[0065] To learn the optimal policy π * , we first need to solve the identifiability problem of reward R(π), as this forms the basis for policy evaluation. Containing only attribute X, due to the lack of information about target action A and observation result Y, the reward R(π) cannot be identified from the target data alone. In order to identify the reward R(π), it is necessary to impose several assumptions from the source dataset. Borrowing information.
[0066] Assumption 1 includes, for all attributes X in the source domain, (i) no confounding: G = 1; (ii) overlap: where e1(X) is the propensity score.
[0067] Assumption 1(i) states that in the source domain, the target action A is independent given attribute X, meaning that all confounding factors that influence the target action A (e.g., treatment) and the observed outcome Y are accounted for by the observations. Assumption 1(ii) asserts that any object in the source domain represented by attribute X has a positive probability of taking the target action A. Assumption 1 is a standard assumption for identifying causal effects in the source domain. However, Assumption 1 is not sufficient to identify causal effects in the target domain. Therefore, Assumption 2 is further invoked.
[0068] Assumption 2 concerns transferability, including: (i) for all attributes (ii) For all attributes in the source domain where s(X) is the sampling score.
[0069] Assumption 2 is widely used in causal inference to estimate causal effects by combining data. Assumption 2(ii) states that all objects in the source domain have a positive probability of being in the target domain. Note
[0070]
[0071] Assumption 2(ii) also implies that the supported attribute X in the target domain needs to be more than the attribute X in the source domain. Otherwise, there may be and , which leads to a violation of Assumption 2(ii).
[0072] Assumption 2(i) implies that, for This ensures the transferability of CATE from the source domain to the target domain. Under Assumptions 1 and 2, we can conclude that
[0073]
[0074] The third equation in formula (2) is Following Assumption 1, Equation (2) leads to the identifiability of τ(X), i.e., CATE in the target domain. Therefore, under Assumptions 1 and 2, the reward R(π) can be identified as
[0075]
[0076] Assumption 2 allows for the existence of covariate shift (i.e., attribute shift), i.e., the distribution of attribute X in the source domain may be significantly different from the distribution of X in the target domain.
[0077] The method proposed in this disclosure for learning an optimal policy in a target domain may include two steps: (a) policy evaluation, which estimates the reward R(π) for a given policy π; and (b) policy learning, which learns the optimal policy based on the estimated reward R(π). The reward R(π) for the policy model π may include a reward component corresponding to the sample object i, and the reward component may be determined by the method discussed below.
[0078] In some embodiments, the reward component may be determined in the following manner (also referred to as a direct method). Specifically, the electronic device 110 may determine the reward component in response to the sample object i corresponding to the source data set. (ie, the sample object i is from the first group of objects), the reward component is determined to be a predetermined value (eg, 0). The electronic device 110 may respond to the sample object i corresponding to the target data set. (That is, the sample object i comes from the second group of objects), based on the attribute X of the sample object i i , transformation factor q, sample decision π(X), first associated information μ1(X) and second associated information μ0(X) to determine the reward component.
[0079] In some embodiments, based on the attribute X of the sample object i i , transformation factor q, sample decision π(X), first associated information μ1(X) and second associated information μ0(X) to determine the reward component includes the following operations. Specifically, the electronic device 110 can determine the reward component based on the attribute X of the sample object i. i , the first associated information μ1(X) and the second associated information μ0(X), to determine the first prediction result (for example, represented as ) and the second prediction result (e.g., represented as ). The electronic device 110 may make a decision based on the sample π(X), the first prediction result and the second prediction result Determine the predicted effect of the target action on the sample object i (e.g., expressed as ). The electronic device 110 may derive a reward component based on the transformation factor q and the predicted effect.
[0080] For example, based on formula (3), the reward component of reward R(π) is estimated by direct method, and the direct estimated value is obtained as follows:
[0081]
[0082] in Represents the estimated regression function μ defined in formula (2) a (X). This can be achieved by regressing X on Y using a source data set with A=a. G i = 0. At this time, the reward value is 0.
[0083] Direct estimate The unbiasedness of depends on the accuracy of the estimated regression function. is μ a (X) is a biased estimate, then will also be a biased estimate of R(π). In addition, since It is trained using a source dataset with A=a, but when applied to the entire target dataset, the generalization performance of the direct method is usually poor. and target dataset When the attribute distributions between the two models are significantly different, the direct method suffers from model extrapolation problems, resulting in poor actual performance.
[0084] In addition to the direct method, propensity scores can also be used and sampling score To construct the inverse probability weighted (IPW) estimate of the reward R(π). At this point,
[0085]
[0086] In some embodiments, the reward component may be determined in the following manner (also referred to as IPW method). Specifically, the electronic device 110 may determine the reward component in response to the sample object i corresponding to the target data set. (ie, the sample object i is from the second group of objects), the reward component is determined to be a predetermined value (eg, 0). The electronic device 110 may respond to the sample object i corresponding to the source data set. (That is, the sample object i comes from the first group of objects), based on the attribute X of the sample object i i , transformation factor q, sample decision π(X), the third associated information (ie, propensity score e1(X)) and the fourth associated information (ie, sampling score s(X)) are used to determine the reward component.
[0087] In some embodiments, the electronic device 110 may be configured to identify the sample object i based on its attribute X. i , propensity score e1(X) and sampling score s(X), determine the first predicted probability of the sample object i being applied the target action (for example, expressed as ) and sample object i and source dataset The corresponding second predicted probability (e.g., expressed as ). The electronic device 110 may make a decision based on the sample π(X), the first prediction probability The second predicted probability And the observation result Y of sample object i i , determine the predicted effect of the target action on the sample object i. For example, when the target action is applied, the predicted effect can be expressed as When the target action is not applied, the prediction effect can be expressed as The electronic device 110 may derive a reward component based on the transformation factor q and the predicted effect.
[0088] For example, based on formula (4), the IPW method is used to estimate the reward component of reward R(π), and the IPW estimation value is obtained. as follows:
[0089]
[0090] in is the estimate of the propensity score e1(X), is an estimate of the sampling score s(X). For the target dataset G i = 0. At this time, the reward value is 0.
[0091] exist and When the propensity score e1(X) and the sampling score s(X) are the exact estimates, the IPW estimate is an unbiased estimate of the reward R(π), that is, and However, a limitation of the IPW estimator is that it is not very efficient, which means that it tends to have a large variance. This will be shown in the experiments.
[0092] The limitations of direct and IPW methods are primarily due to their inadequate use of observational data. Direct methods fail to exploit information about the data indicator G and the target action A, while IPW methods fail to extract the association between attribute X and the observed outcome Y. To fully utilize the observed data, semiparametric efficiency theory can be employed to derive efficient bounds on the effective influence function and reward R(π), thereby obtaining a semiparametrically efficient estimate of the reward R(π). This semiparametrically efficient estimate is considered optimal because it achieves the semiparametric efficiency bound, meaning that it has the minimum asymptotic variance under several regularity conditions.
[0093] In some embodiments, the reward component may be determined in the following manner (also referred to as SE method). Specifically, the electronic device 110 may determine the reward component in response to the sample object i corresponding to the source data set. Based on the attribute X of sample object i i , transformation factor q, sample decision π(X), first association information μ1(X), second association information μ0(X), third association information (i.e., propensity score e1(X)) and fourth association information (i.e., sampling score s(X)), to determine the reward component. The electronic device 110 may determine the reward component in response to the sample object i corresponding to the target data set. Based on the attribute X of sample object i i , transformation factor q, sample decision π(X), first associated information μ1(X) and second associated information μ0(X) to determine the reward component.
[0094] As discussed above, the electronic device 110 may determine the attribute X of the sample object i based on the attribute X of the sample object i. i, the first associated information μ1(X) and the second associated information μ0(X), to determine the first prediction result when the sample object i is applied with the target action And the second prediction result when the sample object i is not applied with the target action The electronic device 110 may make a decision based on the sample π(X), the first prediction result and the second prediction result Determine the predicted effect of the target action on sample object i The electronic device 110 may derive a reward component based on the transformation factor q and the predicted effect.
[0095] In some embodiments, the electronic device 110 may be configured to identify the sample object i based on its attribute X. i , propensity score e1(X) and sampling score s(X), determine the first predicted probability of sample object i being applied the target action As well as the sample object i and the source dataset The corresponding second predicted probability The electronic device 110 may be based on the attribute X of the sample object i i , the first association information μ1(X) and the second association information μ0(X), determine the first prediction result when the sample object i is applied with the target action And the second prediction result when the sample object i is not applied with the target action The electronic device 110 may determine a first prediction result and the observation result Y of sample object i i The first result difference between ), and the second prediction result and the observation result Y of sample object i i The second result difference between ). The electronic device 110 may make a decision based on the sample π(X), the first prediction probability The second predicted probability And the first result difference or second result difference The predicted effect of the target action on the sample object i is determined by a result difference in . Then, the electronic device 110 can derive the reward component based on the transformation factor q and the predicted effect.
[0096] For example, Theorem 1 involves the efficiency bound of the reward R(π). Under Assumptions 1 and 2, the effective influence function of the reward R(π) is as follows
[0097]
[0098] where the semiparametric efficiency bound of the reward R(π) is
[0099] Theorem 1 proposes the effective influence function and semiparametric efficiency bound of the reward R(π) under Assumptions 1 and 2. Based on Theorem 1, we can construct an estimate of the semiparametric efficiency (SE) of the reward R(π), which is expressed as
[0100]
[0101] Next, we analyze the SE estimates Theoretical nature.
[0102] Proposition 1 involves SE estimates Double robustness. Under Assumptions 1 and 2, if one of the following conditions is met, the SE estimate is an unbiased estimate of the reward R(π): (i) That is, for a=0,1, Accurately estimate μ a (x); (ii) and that is, and Accurately estimate e(x) and s(x).
[0103] Proposition 1 shows that Double robustness, that is, for a = 0, 1, if the resulting regression function μ a (X) can be accurately estimated, or the propensity score e1(X) and the sampling score s(X) can be accurately estimated, then is an unbiased estimate of R(π). Compared with the direct method, the latter requires Compared with the IPW method, the latter requires and Double robustness is achieved by mitigating interference parameters e1(X), s(X) and μ a (X) to provide more reliable results by eliminating the inductive bias caused by the inaccurate model.
[0104] Theorem 2 involves Under assumptions 1 and 2, for all and a∈{0,1}, if and So is a consistent estimate of the reward R(π) and satisfies
[0105]
[0106] in is the semiparametric efficiency bound of R(π), Indicates that the distribution converges.
[0107] Theorem 2 establishes the proposed estimate Furthermore, it shows that is semiparametrically efficient, i.e., it reaches the semiparametric efficiency bound. These desirable properties are achieved when the rate of estimation of the disturbance parameter is faster than n -1 / 4 These conditions are easily satisfied by many flexible machine learning methods and are commonly used in causal inference.
[0108] After estimating the reward R(π), the optimal policy can be learned. For a given hypothesis space π, the target policy is defined as π * (x) = arg ma xπ∈Π R(π). By optimizing different estimates of reward R(π), different π can be obtained. * The estimated value of It is defined as
[0109]
[0110] in Can be or Table 2 shows the learning π summarized in Algorithm 1. * process.
[0111] Table 2
[0112]
[0113] pass It is worth noting that the definition of
[0114]
[0115] This means that the direct method essentially uses a plug-in method to approximate the optimal policy.
[0116] When optimizing a policy by optimizing the estimated reward, better generalization performance can be achieved if the estimated reward is more effective. It is most effective under Assumptions 1 and 2, as shown in Theorem 1. The following text focuses on exploring its finite sample properties and the learning strategy obtained by optimizing it. Similar results can be obtained for the direct method and IPW method. In finite samples, it is allowed to and are inaccurate, i.e. they may differ from μ a (X), e1(X) and s(X) are different.
[0117] Proposition 2 involves For a=0,1, and Given the learned Then for any given policy π, The deviation is as follows
[0118]
[0119] where e0(X i )=1-e1(Xi), and
[0120] From Theorem 2, we can see that for a=0,1, The deviation is the estimation error and It is obvious that when Close to μ a (X), or and close to s(X) and e1(X), will be close to R(π).
[0121] Next, we will show the generalization error bound (or regret) of the learned policy. For clarity, we can define
[0122]
[0123] This is achieved through optimization The learned strategy.
[0124] Theorem 3 concerns the generalization error bound, which includes that for any finite hypothesis space Π, (i) with probability 1-η at least, in equal in (ii) with probability 1-η at least,
[0125] Theorem 3(i) provides the learned policy The generalization error bound of π is given by Theorem 3(ii), which shows that the learned policy and the optimal policy π * Note that when When bounded, as the number of samples n tends to infinity, will converge to 0. Therefore, for a sufficiently large number of samples n, if the perturbation parameters are estimated sufficiently accurately, the generalization bound of the learned policy will be well approximated by the estimated reward. Moreover, the generalization bound of the learned policy will be close to the optimal policy boundaries.
[0126] In summary, this paper proposes a principled approach to learning an optimal policy in a target domain based on two datasets: one containing complete information from the source domain, and the other containing only covariates from the target domain. The problem is formalized from the perspective of causal inference, and an assumption of the identifiability of rewards under a given policy is proposed. Furthermore, an effective influence function and efficiency bounds for the rewards are derived. Thus, a doubly robust and semi-parametrically efficient reward estimator is constructed, and the optimal policy is then learned by optimizing the estimated rewards. Experiments show that the method proposed in this paper not only estimates rewards more accurately, but also produces a policy that is close to the theoretically optimal policy.
[0127] In addition, experiments show that the direct method, IPW method, and SE method proposed in this disclosure can produce larger rewards, smaller policy errors, and larger welfare changes. The proposed SE method achieves the highest reward, lowest policy error, smallest regret, and largest welfare change. The standard deviations of the direct method and SE method are significantly smaller than those of the IPW method, indicating the instability (larger variance) of the IPW method. In summary, the SE method outperforms the direct method and IPW method due to its desirable properties such as dual robustness and semi-parametric efficiency.
[0128] Figure 3 FIG. 3 is a flow chart of a method 300 for policy optimization according to some embodiments of the present disclosure. The method 300 may be implemented at the electronic device 110. Figure 1 Method 300 is described.
[0129] In box 310, the electronic device 110 determines attribute association information based on a source data set corresponding to the first group of objects, the source data set including attributes of each object in the first group of objects, an indication of whether a target action is applied to the object, and observation results related to the object, the attribute association information indicating the association between the object observation data obtained according to the source data set and the object attributes.
[0130] In block 320 , the electronic device 110 determines a transformation factor for transforming the attribute association information from the source dataset to the target dataset based on the number of the first set of objects and the number of the second set of objects corresponding to the target dataset, the target dataset including the attribute of each object in the second set of objects.
[0131] In box 330, the electronic device 110 determines a reward for the policy model based on a sample decision of whether to apply the target action to the sample object, the attribute association information and the transformation factor for the sample object in the first group of objects and the second group of objects, where the sample decision is generated using the policy model.
[0132] At block 340 , the electronic device 110 updates the policy model using the reward.
[0133] In some embodiments, the attribute association information includes at least one of the following: first association information between results and attributes, the first association information indicates how the first average observation result of objects in the first group of objects when the target action is applied changes with the object attributes, second association information between results and attributes, the second association information indicates how the second average observation result of objects in the first group of objects when the target action is not applied changes with the object attributes, third association information between decisions and attributes, the third association information indicates how the probability of objects in the first group of objects being applied the target action changes with the object attributes, or fourth association information between data sets and attributes, the fourth association information indicates how the probability of objects corresponding to the source data set changes with the object attributes.
[0134] In some embodiments, the reward for the policy model includes a reward component corresponding to the sample object, and the reward component is determined as follows: in response to the sample object corresponding to the source data set, the reward component is determined to be a predetermined value; and in response to the sample object corresponding to the target data set, the reward component is determined based on the attributes of the sample object, the transformation factor, the sample decision, the first association information, and the second association information.
[0135] In some embodiments, the reward for the policy model includes a reward component corresponding to the sample object, and the reward component is determined as follows: in response to the sample object corresponding to the target data set, the reward component is determined to be a predetermined value; and in response to the sample object corresponding to the source data set, the reward component is determined based on the attributes of the sample object, the transformation factor, the sample decision, the third association information, and the fourth association information.
[0136] In some embodiments, the reward for the policy model includes a reward component corresponding to the sample object, and the reward component is determined as follows: in response to the sample object corresponding to the source data set, the reward component is determined based on the attributes, transformation factor, sample decision, first association information, second association information, third association information and fourth association information of the sample object; and in response to the sample object corresponding to the target data set, the reward component is determined based on the attributes, transformation factor, sample decision, first association information and second association information of the sample object.
[0137] In some embodiments, determining the reward component based on the attributes of the sample object, the transformation factor, the sample decision, the first association information, and the second association information includes: determining a first prediction result when the target action is applied to the sample object and a second prediction result when the target action is not applied to the sample object based on the attributes of the sample object, the first association information, and the second association information; determining a predicted effect of the target action on the sample object based on the sample decision, the first prediction result, and the second prediction result; and deriving the reward component based on the transformation factor and the prediction effect.
[0138] In some embodiments, determining the reward component based on the attributes, transformation factors, sample decisions, third association information, and fourth association information of the sample object includes: determining the first predicted probability that the target action is applied to the sample object and the second predicted probability that the sample object corresponds to the source data set based on the attributes, third association information, and fourth association information of the sample object; determining the predicted effect of the target action on the sample object based on the sample decision, the first predicted probability, the second predicted probability, and the observation results of the sample object; and deriving the reward component based on the transformation factor and the predicted effect.
[0139] In some embodiments, determining the reward component based on the attributes, transformation factors, sample decisions, first association information, second association information, third association information, and fourth association information of the sample object includes: determining a first predicted probability that the target action is applied to the sample object and a second predicted probability that the sample object corresponds to the source data set based on the attributes, third association information, and fourth association information of the sample object; determining a first predicted result when the target action is applied to the sample object and a second predicted result when the target action is not applied to the sample object based on the attributes, first association information, and second association information of the sample object; determining a first result difference between the first predicted result and the observed result of the sample object, and a second result difference between the second predicted result and the observed result of the sample object; determining a predicted effect of the target action on the sample object based on the sample decision, the first predicted probability, the second predicted probability, and one of the first result difference or the second result difference; and deriving the reward component based on the transformation factor and the predicted effect.
[0140] In some embodiments, the first group of objects includes a first group of robots operating in a first environment, the second group of objects includes a second group of robots operating in a second environment different from the first environment, and the policy model is configured to provide a decision on whether to apply the planned action to the robots.
[0141] Figure 4 4 shows a schematic structural block diagram of an apparatus 400 for policy optimization according to some embodiments of the present disclosure. The apparatus 400 may be implemented in or included in the electronic device 110. Each module / component in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0142] As shown, apparatus 400 includes an attribute association information determination module 410 configured to determine attribute association information based on a source dataset corresponding to a first group of objects, the source dataset including attributes of each object in the first group of objects, an indication of whether to apply a target action to the object, and observations related to the object. The attribute association information indicates associations between object observations obtained from the source dataset and the object attributes. Apparatus 400 also includes a transformation factor determination module 420 configured to determine a transformation factor for transforming the attribute association information from the source dataset to a target dataset based on the number of objects in the first group and the number of objects in a second group corresponding to the target dataset. The target dataset includes attributes of each object in the second group of objects. Apparatus 400 also includes a reward determination module 430 configured to determine, for sample objects in the first and second groups of objects, a reward for a policy model based on sample decisions regarding whether to apply the target action to the sample objects, the attribute association information, and the transformation factor. The sample decisions are generated using the policy model. Apparatus 400 also includes a policy model update module 440 that updates the policy model using the reward.
[0143] In some embodiments, the attribute association information includes at least one of the following: first association information between results and attributes, the first association information indicates how the first average observation result of objects in the first group of objects when the target action is applied changes with the object attributes, second association information between results and attributes, the second association information indicates how the second average observation result of objects in the first group of objects when the target action is not applied changes with the object attributes, third association information between decisions and attributes, the third association information indicates how the probability of objects in the first group of objects being applied the target action changes with the object attributes, or fourth association information between data sets and attributes, the fourth association information indicates how the probability of objects corresponding to the source data set changes with the object attributes.
[0144] In some embodiments, the reward for the policy model includes a reward component corresponding to the sample object, and the reward component is determined as follows: in response to the sample object corresponding to the source data set, the reward component is determined to be a predetermined value; and in response to the sample object corresponding to the target data set, the reward component is determined based on the attributes of the sample object, the transformation factor, the sample decision, the first association information, and the second association information.
[0145] In some embodiments, the reward for the policy model includes a reward component corresponding to the sample object, and the reward component is determined as follows: in response to the sample object corresponding to the target data set, the reward component is determined to be a predetermined value; and in response to the sample object corresponding to the source data set, the reward component is determined based on the attributes of the sample object, the transformation factor, the sample decision, the third association information, and the fourth association information.
[0146] In some embodiments, the reward for the policy model includes a reward component corresponding to the sample object, and the reward component is determined as follows: in response to the sample object corresponding to the source data set, the reward component is determined based on the attributes, transformation factor, sample decision, first association information, second association information, third association information and fourth association information of the sample object; and in response to the sample object corresponding to the target data set, the reward component is determined based on the attributes, transformation factor, sample decision, first association information and second association information of the sample object.
[0147] In some embodiments, determining the reward component based on the attributes of the sample object, the transformation factor, the sample decision, the first association information, and the second association information includes: determining a first prediction result when the target action is applied to the sample object and a second prediction result when the target action is not applied to the sample object based on the attributes of the sample object, the first association information, and the second association information; determining a predicted effect of the target action on the sample object based on the sample decision, the first prediction result, and the second prediction result; and deriving the reward component based on the transformation factor and the prediction effect.
[0148] In some embodiments, determining the reward component based on the attributes, transformation factors, sample decisions, third association information, and fourth association information of the sample object includes: determining the first predicted probability that the target action is applied to the sample object and the second predicted probability that the sample object corresponds to the source data set based on the attributes, third association information, and fourth association information of the sample object; determining the predicted effect of the target action on the sample object based on the sample decision, the first predicted probability, the second predicted probability, and the observation results of the sample object; and deriving the reward component based on the transformation factor and the predicted effect.
[0149] In some embodiments, determining the reward component based on the attributes, transformation factors, sample decisions, first association information, second association information, third association information, and fourth association information of the sample object includes: determining a first predicted probability that the target action is applied to the sample object and a second predicted probability that the sample object corresponds to the source data set based on the attributes, third association information, and fourth association information of the sample object; determining a first predicted result when the target action is applied to the sample object and a second predicted result when the target action is not applied to the sample object based on the attributes, first association information, and second association information of the sample object; determining a first result difference between the first predicted result and the observed result of the sample object, and a second result difference between the second predicted result and the observed result of the sample object; determining a predicted effect of the target action on the sample object based on the sample decision, the first predicted probability, the second predicted probability, and one of the first result difference or the second result difference; and deriving the reward component based on the transformation factor and the predicted effect.
[0150] In some embodiments, the first group of objects includes a first group of robots operating in a first environment, the second group of objects includes a second group of robots operating in a second environment different from the first environment, and the policy model is configured to provide a decision on whether to apply the planned action to the robots.
[0151] Figure 5 1 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 5 The illustrated electronic device 500 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown may include or be implemented as Figure 1 electronic device 110 or Figure 4 device 400.
[0152] like Figure 5 As shown, electronic device 500 is in the form of a general electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of electronic device 500.
[0153] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.
[0154] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 5As shown in FIG, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. Memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0155] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0156] Input device 550 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 560 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 500 may also communicate with one or more external devices (not shown) via communication unit 540 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with electronic device 500, or with any device that allows electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0157] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0158] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0159] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0160] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0161] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0162] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A strategy optimization method, comprising: determining attribute association information based on a source data set corresponding to a first group of objects, the source data set including an attribute of each object in the first group of objects, an indication of whether a target action is applied to the object, and an observation result related to the object, the attribute association information indicating an association between object observation data obtained from the source data set and the object attribute; determining, based on the number of the first set of objects and the number of the second set of objects corresponding to a target dataset, a transformation factor for transforming the attribute association information from the source dataset to the target dataset, the target dataset including the attribute of each object in the second set of objects; determining, for a sample object in the first group of objects and the second group of objects, a reward for a policy model based on a sample decision of whether to apply the target action to the sample object, the attribute association information, and the transformation factor, wherein the sample decision is generated using the policy model; as well as The policy model is updated using the reward.
2. The method according to claim 1, wherein the attribute association information includes at least one of the following: first association information between results and attributes, wherein the first association information indicates changes in first average observation results of objects in the first group of objects when the target action is applied as a function of object attributes; second association information between the result and the attribute, wherein the second association information indicates changes in the second average observation result of the objects in the first group of objects when the target action is not applied as the object attribute changes; third association information between the decision and the attribute, the third association information indicating that the probability of the target action being applied to an object in the first group of objects changes with the attribute of the object, or Fourth association information between a data set and an attribute, the fourth association information indicating a change in a probability that an object corresponds to the source data set as a function of the object attribute.
3. The method according to claim 2, wherein the reward for the policy model includes a reward component corresponding to the sample object, and the reward component is determined by: In response to the sample object corresponding to the source data set, determining the reward component to a predetermined value; and In response to the sample object corresponding to the target data set, the reward component is determined based on the attribute of the sample object, the transformation factor, the sample decision, the first association information, and the second association information.
4. The method according to claim 2, wherein the reward for the policy model includes a reward component corresponding to the sample object, and the reward component is determined by: In response to the sample object corresponding to the target data set, determining the reward component to a predetermined value; and In response to the sample object corresponding to the source data set, the reward component is determined based on the attribute of the sample object, the transformation factor, the sample decision, the third association information, and the fourth association information.
5. The method according to claim 2, wherein the reward for the policy model includes a reward component corresponding to the sample object, and the reward component is determined by: In response to the sample object corresponding to the source data set, determining the reward component based on the attribute of the sample object, the transformation factor, the sample decision, the first association information, the second association information, the third association information, and the fourth association information; and In response to the sample object corresponding to the target data set, the reward component is determined based on the attribute of the sample object, the transformation factor, the sample decision, the first association information, and the second association information.
6. The method according to claim 3 or 5, wherein determining the reward component based on the attribute of the sample object, the conversion factor, the sample decision, the first association information, and the second association information comprises: Determining, based on the attributes of the sample object, the first association information, and the second association information, a first prediction result when the target action is applied to the sample object and a second prediction result when the target action is not applied to the sample object; Determining a predicted effect of the target action on the sample object based on the sample decision, the first prediction result, and the second prediction result; as well as The reward component is derived based on the transformation factor and the predicted effect.
7. The method according to claim 4, wherein determining the reward component based on the attribute of the sample object, the conversion factor, the sample decision, the third association information, and the fourth association information comprises: Determining, based on the attributes of the sample object, the third association information, and the fourth association information, a first predicted probability that the sample object is applied to the target object and a second predicted probability that the sample object corresponds to the source dataset; Determining a predicted effect of the target action on the sample object based on the sample decision, the first predicted probability, the second predicted probability, and an observation result of the sample object; as well as The reward component is derived based on the transformation factor and the predicted effect.
8. The method according to claim 5, wherein determining the reward component based on the attribute of the sample object, the conversion factor, the sample decision, the first association information, the second association information, the third association information, and the fourth association information comprises: Determining, based on the attributes of the sample object, the third association information, and the fourth association information, a first predicted probability that the sample object is applied to the target object and a second predicted probability that the sample object corresponds to the source dataset; Determining, based on the attributes of the sample object, the first association information, and the second association information, a first prediction result when the target action is applied to the sample object and a second prediction result when the target action is not applied to the sample object; determining a first result difference between the first prediction result and the observation result of the sample subject, and a second result difference between the second prediction result and the observation result of the sample subject; determining a predicted effect of the target action on the sample subject based on the sample decision, the first predicted probability, the second predicted probability, and one of the first outcome difference or the second outcome difference; as well as The reward component is derived based on the transformation factor and the predicted effect.
9. The method of claim 1 , wherein the first group of objects comprises a first group of robots operating in a first environment, the second group of objects comprises a second group of robots operating in a second environment different from the first environment, and the policy model is configured to provide a decision on whether to apply a planned action to the robots.
10. A device for policy optimization, comprising: an attribute association information determination module configured to determine attribute association information based on a source data set corresponding to a first group of objects, the source data set including an attribute of each object in the first group of objects, an indication of whether a target action is applied to the object, and an observation result related to the object, the attribute association information indicating an association between object observation data obtained from the source data set and the object attribute; a transformation factor determination module configured to determine a transformation factor for transforming the attribute association information from the source dataset to the target dataset based on the number of the first set of objects and the number of the second set of objects corresponding to the target dataset, the target dataset including the attribute of each object in the second set of objects; a reward determination module configured to determine, for a sample object in the first group of objects and the second group of objects, a reward for a policy model based on a sample decision of whether to apply the target action to the sample object, the attribute association information, and the transformation factor, wherein the sample decision is generated using the policy model; as well as A policy model updating module uses the reward to update the policy model.
11. An electronic device comprising: at least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processor. 12 . A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to claim 1 .
13. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Decision model training method, live broadcast decision determination method and device
CN121585837A