Methods and Systems for Supporting Policy Learning
By supporting the strategy learning method, using existing solutions as a black box, combining generalized value functions and total control strategies, the problems of training time-consuming and sparse rewards of existing robot control systems in complex environments are solved, and end-to-end efficient learning and generalized task solutions are achieved.
Patent Information
- Application Number
- CN202080100507.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-15
- Filing Date
- 2020-12-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-12-31
AI Technical Summary
Existing robot control systems have difficulty promoting existing solutions to be able to solve more similar tasks in the same environment, especially in complex problem environments such as time-consuming training, sparse rewards, and catastrophic forgetting, making it difficult for RL agents to learn effectively in state and action space.
Adopting the Support Policy Learning (SPL) method, by receiving the main policy, learning generalized value function (GVF) and total control policy, selecting to execute the main policy or support policy to achieve end-to-end complex task solutions in the state space, using existing solutions as a black box to reduce the impact of catastrophic forgetting.
Improves the learning efficiency and success rate of RL agents in complex environments, reduces training time, enhances the versatility and adaptability of existing solutions, and avoids catastrophic forgetting.
Smart Images

Figure CN115552430B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for supporting policy learning in a robot equipped with a reinforcement learning (RL) agent, and more particularly to supporting policy learning for improving the versatility and scope of use of existing policies. Background Art
[0002] Currently, robot control systems that have existing algorithms or solutions to solve specific tasks may not be able to generalize the existing solutions to be able to solve more similar tasks in the same environment.
[0003] Existing solutions can include machine learning solutions implemented using reinforcement learning (RL). In the context of artificial intelligence (AI), reinforcement learning has historically been implemented through dynamic programming, which uses a sequence of rewards to learn a function. Typically, an agent executing a reinforcement learning (RL) algorithm in a robot (hereafter referred to as an RL agent) is good at solving a task from scratch (tabular rasa) by exploring the environment, collecting states, performing actions in the environment according to a policy, receiving changes in states and corresponding rewards in the environment, and improving the policy to maximize its reward return. However, as the complexity of the problem increases, as in the case of generalized solutions, RL agents may begin to fail and become increasingly difficult to train.
[0004] Some challenges may include the following. Large or infinite state and action spaces are characteristic of complex problem environments and may be difficult for RL agents to explore. Sampling inefficiency may also be a challenge, where training an RL agent may be time consuming due to inefficiencies in sampling possible states. Sparse rewards may be a challenge, where not enough different rewards are sampled to improve the behavior of the RL agent across a range of different states. Credit assignment may be a challenge, where for long-time horizon tasks that require solving long action sequences, it is often difficult to associate rewards with the source task from which the improvement occurred. Transfer learning may be another challenge, where it is difficult to apply learned policies to related problems or the same problem in a different environment, including simulation to real world (sim-to-real) transfer.
[0005] Common approaches to handling large or infinite state and action spaces are to apply function approximation, such as deep learning, to learn features that compactly represent states. However, since deep neural networks typically require many samples for effective training, the problem of sampling inefficiency is often exacerbated.
[0006] Another common approach that attempts to address some of the above challenges is to apply curriculum learning methods to RL agents to derive learned solutions, which are particularly applicable to complex tasks with large state and action spaces, long-horizon tasks, and sparse reward tasks. A well-designed curriculum provided by an expert has several advantages. For example, curriculum learning typically breaks down a task into a series of smaller tasks to be solved in increasing order of complexity, which allows the RL agent to focus on solving simple tasks before moving on to complex tasks. Correspondingly, since the curriculum guides the agent to solve simple tasks before complex tasks, the RL agent learns faster. The key point of the curriculum is that the solution to a complex problem can reuse the knowledge of previous simple problems rather than starting from scratch. Using curriculum learning is an instance of transfer learning, where a series of increasingly complex problems are discovered, and thus the agent must transfer the knowledge of the solutions to early tasks to later tasks.
[0007] However, one challenge in using transfer learning is catastrophic forgetting. Catastrophic forgetting, or catastrophic interference, occurs when the parameters of the solution to a task in one domain are updated to optimize the solution to a new task in another domain, but the updated solution fails to solve or "forgets" how to solve the source task. One way to mitigate this catastrophic forgetting problem is to use progressive networks, which implement transfer by training the agent, the fixed network, and the shared features learned to speed up the training of the parallel network on the real task in a simulated environment.
[0008] Many and similar solutions described above involve leveraging existing solutions to simple tasks to speed up learning of more complex tasks. However, existing solutions typically perform well only when certain conditions and assumptions are met. Existing solutions generally cannot solve all problems, especially when the conditions and assumptions are not satisfied.
[0009] It is desirable to achieve end-to-end learning, where the RL agent learns a generalized solution to solve a given problem without (or with minimal) conditions and assumptions. Given the many challenges in applying RL to the complex problems described above, many problems do not yet have generalized end-to-end solutions. SUMMARY OF THE INVENTION
[0010] The present invention describes methods and systems for end-to-end RL solutions that can achieve complex tasks in the same action space and state space of an environment by efficiently reusing existing solutions, regardless of whether the existing solutions are RL-learned or hand-designed.
[0011] In at least one aspect, the present invention relates to a method for supporting policy learning (SPL). Specifically, existing solutions for one or more simple tasks, whether RL-learned or hand-designed, are treated as one or more black boxes and reused to quickly and efficiently solve more and more complex tasks, despite any limitations or assumptions of the one or more existing solutions. In some examples, SPL may be less susceptible to catastrophic forgetting because the existing solutions are retained and fixed for reuse.
[0012] In some exemplary aspects, the present invention describes a method performed by an agent in a robot that controls the robot's interaction with the environment. The method includes: receiving a primary policy, wherein the primary policy generates an action to be performed by the robot based on the state of the robot, and the performance of the agent in executing the primary policy is measured by an accumulated success value; using a policy evaluation algorithm to learn a generalized value function of the primary policy, wherein the generalized value function predicts an accumulated success value representing the future performance of the agent in executing the primary policy in a given state of the environment, the given state being throughout the state space; obtaining a master control policy, wherein the master control policy selects an action based on the predicted accumulated success value obtained from the generalized value function; when the predicted accumulated success value is an acceptable value, the action selected by the master control policy causes the primary policy to be executed so that the robot performs a primary action generated by the primary policy according to the given state in the state space; when the predicted accumulated success value is not an acceptable value, the action selected by the master control policy causes a support policy to be learned using a reinforcement learning algorithm, wherein the support policy generates a support action to be performed by the robot based on the given state, and the support action causes the robot to transition from the given state to a new state where the predicted accumulated success value has an acceptable value.
[0013] In some exemplary aspects, the present invention describes a processing unit in a robot. The processing unit executes machine-executable instructions to implement an agent to control the robot to interact with the environment, and the instructions cause the agent to perform the following operations: receive a primary policy, wherein the primary policy generates an action to be executed by the robot according to the state of the robot, and the performance of the agent executing the primary policy is measured by an accumulated success value; use a policy evaluation algorithm to learn a generalized value function of the primary policy, wherein the generalized value function predicts an accumulated success value representing the future performance of the agent executing the primary policy in a given state of the environment, and the given state is in the entire state space; obtain a master control policy, wherein the master control policy selects an action according to the predicted accumulated success value obtained from the generalized value function; when the predicted accumulated success value is an acceptable value, the action selected by the master control policy causes the primary policy to be executed, so that the robot executes the primary action generated by the primary policy according to the given state in the state space; when the predicted accumulated success value is not an acceptable value, the action selected by the master control policy causes a support policy to be learned using a reinforcement learning algorithm, wherein the support policy generates a support action to be executed by the robot according to the given state, and the support action causes the robot to transfer from the given state to a new state where the predicted accumulated success value has an acceptable value.
[0014] In some exemplary aspects, the present invention describes a computer-readable medium storing instructions. When the instructions are executed by an agent that controls the robot to interact with the environment in the robot, the instructions cause the agent to perform the following operations: receive a primary policy, wherein the primary policy generates an action to be executed by the robot according to the state of the robot, and the performance of the agent executing the primary policy is measured by an accumulated success value; use a policy evaluation algorithm to learn a generalized value function of the primary policy, wherein the generalized value function predicts an accumulated success value representing the future performance of the agent executing the primary policy in a given state of the environment, and the given state is in the entire state space; obtain a master control policy, wherein the master control policy selects an action according to the predicted accumulated success value obtained from the generalized value function; when the predicted accumulated success value is an acceptable value, the action selected by the master control policy causes the primary policy to be executed, so that the robot executes the primary action generated by the primary policy according to the given state in the state space; when the predicted accumulated success value is not an acceptable value, the action selected by the master control policy causes a support policy to be learned using a reinforcement learning algorithm, wherein the support policy generates a support action to be executed by the robot according to the given state, and the support action causes the robot to transfer from the given state to a new state where the predicted accumulated success value has an acceptable value.
[0015] In any of the above aspects, the learned generalized value function may include: performing a plurality of iterations, where each iteration includes: sampling an action generated by the primary policy according to the current state in the state space, where the action is executed by the agent so that the robot executes the action; after executing the action, sampling a next state in the state space; after executing the action, calculating a cumulative quantity by transferring from the current state to the next state, where the cumulative quantity represents the success value of the agent in the current state; storing at least the cumulative quantity associated with the current state, the action, and the next state; and updating the generalized value function using temporal difference learning.
[0016] In any of the above aspects, the generalized value function is updated using temporal difference learning or Monte Carlo estimation.
[0017] In any of the above aspects, the support policy is learned based on a reward, where the reward is based on the predicted cumulative success value obtained from the generalized value function, and the generalized value function considers a plurality of states sampled from the state space.
[0018] In any of the above aspects, obtaining the overall control policy may include: determining a threshold; the overall control policy is defined to select to execute the primary policy when the success value output by the generalized value function is greater than the threshold, and is also defined to select to learn the support policy when the success value output by the generalized value function is not greater than the threshold.
[0019] In any of the above aspects, obtaining the overall control policy may include: learning the support policy while learning the overall control policy, where the overall control policy is learned according to an overall control policy reward, the support policy is learned according to a support policy reward, and both the overall control policy reward and the support policy reward are based on the predicted cumulative success value obtained from the generalized value function.
[0020] In any of the above aspects, the generalized value function, the overall control policy, and the support policy may be learned simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] To more fully understand the exemplary embodiments and their advantages, reference is now made to the following detailed description taken in conjunction with the accompanying drawings.
[0022] Figure 1 is a schematic diagram of a robot configured to support policy learning provided by an exemplary embodiment.
[0023] Figure 2 is that which can be by Figure 1Flowchart of an exemplary support policy learning method implemented by the shown robot.
[0024] Figure 3 is a flowchart of an exemplary temporal chaining learning method for learning a generalized value function that can be implemented in step 220 of Figure 2
[0025] Figure 4 is a flowchart of an exemplary Monte Carlo supervised learning method for learning a generalized value function that can be implemented in step 220 of Figure 2
[0026] Figure 5 is a flowchart of an exemplary reinforcement learning method for learning a support policy that can be implemented in step 240 of Figure 2
[0027] Figure 6 is a flowchart of an exemplary deployment method that can be implemented by the RL agent shown in Figure 1
[0028] Figure 7 is a flowchart of another exemplary support policy learning method implemented by the robot shown in Figure 1 which combines overall control policy learning and support policy learning into a single parallel step.
[0029] Figure 8 is a flowchart of an exemplary reinforcement learning method for simultaneously learning an overall control policy and a support policy that can be implemented in step 730 of Figure 7
[0030] Figure 9 is a flowchart of yet another exemplary support policy learning method implemented by the robot shown in Figure 1 which combines generalized value function learning, overall control policy learning, and support policy learning into a single parallel step.
[0031] Figure 10 is a flowchart of an exemplary reinforcement learning method for simultaneously learning a generalized value function, an overall control policy, and a support policy that can be implemented in step 920 of Figure 9
[0032] Figure 11 is a schematic diagram of a robot configured for support policy learning provided by another exemplary embodiment.
[0033] Figure 12 is a flowchart of an exemplary support policy learning method that can be implemented by the robot shown in Figure 11
[0034] Like reference numerals are used in the different figures to denote like components. DETAILED DESCRIPTION
[0035] The present invention may use the following definitions:
[0036] Action: A control decision implemented by an actuator for interacting with the environment.
[0037] Action space: The set of all possible actions.
[0038] Action value: The expected reward obtained by an agent based on a given state, the next action, and the policy followed thereafter.
[0039] ADAS: Advanced Driver-Assistance System.
[0040] Discount: An exponentially decaying factor that measures the importance of future rewards.
[0041] GVF: General Value Function.
[0042] MC: Monte Carlo estimation.
[0043] MDP: Markov Decision Process defined by a state space, an action space, a transition model, and a reward.
[0044] Observation: A description of the environment captured by sensors or generated by other sources.
[0045] POMDP: Partially Observable Markov Decision Process defined by a state space, an action space, a transition model, a reward, an observation space, and an observation distribution, where the state distribution is a conditional probability, i.e., a mapping from a state to an observation.
[0046] Sim-to-real: Transferring a policy learned in simulation to the real world.
[0047] State: A description of the environment, plus an action, is sufficient to predict the future without any other information, i.e., without the historical state.
[0048] State space: The set of all possible states.
[0049] TD: Temporal Difference estimation.
[0050] Trajectory: A sequence of transitions in the environment, starting from an initial state, the action executed in the initial state, the reward received, the next state received, the next action executed, until the last state is received.
[0051] Transfer: Obtaining knowledge in the solution of one task and then reusing that knowledge in another task.
[0052] Transition: A set of state, action, reward, and next state.
[0053] Policy: A decision rule that specifies an action given a state.
[0054] Return: The sum of future rewards when executing a policy in the environment.
[0055] Reward: A signal received by the agent in the environment when interacting with the environment, which provides feedback on the quality of the policy.
[0056] RL: Reinforcement Learning.
[0057] Value: The expected return of the agent based on a given state and the policy followed.
[0058] Exemplary embodiments generally relate to a robot that includes an RL agent that controls the robot's interaction in an environment. To interact with the environment, the RL agent receives the current state of the environment; uses a generalized value function learned for the primary policy to calculate a predicted cumulative success value that represents the future performance of the primary policy in the current state; and, based on the decision of the overall control policy, selects (1) to execute the primary policy and perform the primary action generated by the primary policy based on the current state, or (2) to execute a support policy and perform the support action generated by the support policy based on the current state.
[0059] In some embodiments, the environment is a simulated environment, and the RL agent is implemented as one or more computer programs that interact with the simulated environment. For example, if the simulated environment is a video game, the RL agent can be a simulated user playing the video game. As another example, if the simulated environment is a motion simulation environment, such as a driving simulation or a flight simulation, the RL agent is a simulated driver who navigates through the motion simulation. In these implementations, the actions can be points in the possible control input space that control the simulated user or the simulated driver.
[0060] In some other examples, the environment is a real environment, and the RL agent is a mechanical agent that interacts with the real environment. For example, the RL agent can be a robot that interacts with the environment to perform a specific task. As another example, the RL agent can be an autonomous or semi-autonomous vehicle that navigates in the environment. In these implementations, the actions can be points in the possible control input space that control the robot or the autonomous vehicle.
[0061] Figure 1 FIG. is a schematic diagram of an exemplary robot 100 including an RL agent 102 provided by the present invention. The RL agent 102 can be implemented by one or more physical processing units in the robot 100, for example, by one or more processing units that execute computer-readable instructions (which can be stored in the memory of the robot 100) to perform the methods described herein. It should be noted that Figure 1 includes the elements shown by the dashed lines. These elements can only be implemented during the training phase of the RL agent 102 (for example, when the RL agent 102 is trained to learn the GVF, learn the support policy, and optionally learn the overall control policy, as described in detail below), and cannot be implemented when the RL agent 102 is deployed (for example, when the RL agent 102 is running in the inference phase). The RL agent 102 is configured with a main policy 104, and the main policy 104 generates main actions to be executed by the robot 100 based on the state. During the training phase, the RL agent 102 is used to learn a generalized value function 106, and the generalized value function 106 determines the predicted cumulative success value of the main policy 104 in a given state in the state space S. The RL agent 102 is also used to learn a support policy 108, and the support policy 108 generates support actions based on the state to transfer the robot 100 from a state where the main policy 104 may not be successful to a state where the main policy 104 may be successful. The RL agent 102 is also used to implement or learn an overall control policy 110, and the purpose of the overall control policy 110 is to maximize the success rate of the main policy 104 by selecting the main policy 104 or the support policy 108 according to the predicted cumulative success value in the current state.
[0062] Generally, the main policy 104 can be a routine or a program. When executed by the RL agent 102, the routine or program receives the current state associated with the cumulative success value, generates a main action based on the current state, executes the main action, and repeats these steps when the next state becomes the current state, where the main action causes the robot 100 to transition from the current state to a new state associated with a new cumulative success value. As detailed below, the main policy 104 is associated with acceptable cumulative success values in a set of states that define a subset of the entire state space (where the robot 100 can be in any state in the state space). To implement the present invention, the RL agent 102 executing the main action means sending the main action to the controller 116 in the RL agent 102. The controller 116 generates control signals sent to one or more actuators 118 in the robot 100, and these control signals cause the robot 100 to perform an action in the environment to cause the robot 100 to transition from the current state to a new state.
[0063] The robot 100 can be any mechanical device for performing specific actions in the environment. For example, the robot 100 can be a robotic arm responsible for picking components from a box, or an autonomous or semi-autonomous vehicle responsible for performing driving actions such as parking, or a robotic entity responsible for navigating in a specific environment.
[0064] As Figure 1 shown, the robot 100 includes sensors 112 and a state processor 114. The robot 100 also includes one or more controllers 116 that send control signals to one or more actuators 118.
[0065] The sensors 112 are respectively used to sense the environment and provide observation data representing the environmental observation results at a specific time point to the state processor 114. In some embodiments, the sensors 112 include cameras, one or more 3D laser scanning sensors (such as Light Detection and Ranging (LiDAR)), one or more radars, one or more accelerometers, one or more gyroscopes, one or more thermometers, etc. The state processor 114 respectively receives the observation data from the sensors 112, processes all the received observation data to generate the state s of the environment, and outputs the state s. In some embodiments, the sensors 112 themselves may generate the state s, and the state processor 114 may simply forward the state s received from the sensors 112. In some embodiments, some of the sensors 112 may output the raw observation data (which requires further processing) to the state processor 114, while other sensors 112 may output data that does not require further processing. The state processor 114 may process the raw observation data and simply add or concatenate the processing results with the other data that does not need to be processed. As described above, the set of all possible states s in the environment is called the state space S.
[0066] In some embodiments, the observation data received from one or more of the sensors 112 may include low-dimensional features characterizing the environmental observation results. In these embodiments, the state processor 114 may perform feature extraction on the observation data received from each of the one or more sensors 112 and output a low-dimensional feature vector (which can be more easily processed by the RL agent 102). In these embodiments, different dimensions of the low-dimensional feature vector may have different value ranges.
[0067] In some embodiments, the observation data of the environment may include digital images characterizing the environmental observation results, such as images of the simulated environment or images captured by one or more sensors 112 (such as cameras) when the robot 100 interacts with the real environment. In these embodiments, the state processor 114 may perform feature extraction on the digital images included in the observation data and output a high-dimensional feature vector (which can be more easily processed by the RL agent 102).
[0068] A given state s may be represented by a combination of different data with different dimensions, different formats, and / or different degrees of processing (such as data or processed feature vectors).
[0069] To enable the robot 100 to perform an action, the RL agent 102 implements a master control policy 110. The master control policy 110 determines whether to execute the main policy 104 or the support policy 108 (or learn the support policy 108 during the training phase detailed below) to execute the main action 122 generated by the main policy 104 or the support action 124 generated by the support policy 108. Executing the main action 122 or the support action 124 causes the RL agent 102 to send the corresponding action to one or more controllers 116, thereby causing the robot 100 to perform the action. Whether the master control policy selects the main policy 104 or the support policy 108 depends on the predicted cumulative success value of the main policy 104 determined by the GVF 106. The GVF 106 representing the predicted cumulative success value of the main policy 104 in a given state s can be denoted as G M (s). One or more controllers 116 are used to process each corresponding action received from the RL agent 102 and send corresponding control signals to one or more actuators 118 to cause the robot 100 to perform the corresponding action (e.g., motor control, electrical activation, mechanical movement). For example, one or more controllers 116 may include a processing unit (e.g., a microprocessor) that converts the action received from the RL agent 102 into a control signal for controlling the actuator 118. For example, if the action received from the RL agent 102 is to increase acceleration, one or more controllers 116 may convert the action into a control signal (e.g., a voltage signal) that increases the rotation of the actuator 118 such as a motor. Generally, in some exemplary embodiments, the RL agent 102 is used to extend or generalize existing or known knowledge (learned knowledge) required to solve one task to solve another task in a generalized context. Reusing existing knowledge or learned knowledge can be referred to as transfer learning. For example, one task is a problem that needs to be solved in an environment to achieve a certain goal, and this goal can be measured by maximizing the predicted success rate.
[0070] In some embodiments, the main policy 104 in the RL agent 102 is an existing solution to a task. The main policy 104 is denoted as π M(s) and provided to the RL agent 102. The main policy 104 can be manually designed (e.g., manually developed by humans through empirical experience and / or trial and error, etc.) or a learned solution (e.g., learned through reinforcement learning using a small number of assumptions in a large state space S), and this solution is used to solve a task or achieve a goal within a subset L of states in the entire state space S, where L ∈ S. In other words, the main policy 104 can represent a solution to a simple task in the same environment. For example, the main policy 104 can represent a solution with a high success probability for performing a simple task (i.e., within the subset L of states) but a low success probability for performing a more general task (i.e., in the remaining states in the state space S).
[0071] According to the present invention, the main policy 104 can be regarded as a "black box", which can advantageously reuse existing solutions without any tabular rasa learning. In addition, catastrophic forgetting can be suppressed by: maintaining the main policy 104 without attempting to improve the main policy by learning to extend the utility of the overall control policy to the entire state space S. In addition, it is not necessary to know in advance the subset L of states where the main policy 104 may succeed. It should be noted that the subset L of states can be different from the simple set of states for formulating the main policy 104. For example, the main policy 104 can be manually designed to succeed in a very limited expected scenario, yet the main policy 104 can actually succeed within a subset L of states larger than the expected scenario. Even if the geometry of the state space L where the main policy 104 succeeds may be highly irregular or even decomposed into several disjoint regions, no assumptions need to be made about its structure. It is not necessary to assume whether the main policy 104 is learned through reinforcement learning or the like or is manually designed.
[0072] The overall control policy 110 is used to maximize the success rate of the main policy 104. Specifically, the overall control policy 110 selects the support policy 108 or the main policy 104 according to the predicted cumulative success value through the GVF 106 in the current state s ∈ S, by explicitly constructing the overall control policy 110 (e.g., using manually defined rules) or by learning the overall control policy 110.
[0073] Since the overall control strategy 110 decides whether to execute the main strategy 104 or the support strategy 108 based on the predicted cumulative success value, embodiments of the present invention can use a success threshold to define a subset of states L in which the main strategy 104 may succeed. To this end, a threshold success value representing an acceptable success value can be defined. The GVF 106 can evaluate the main strategy 104 in the entire state space S and learn the subset of states L based on when the predicted cumulative success value is greater than the acceptable threshold. Therefore, the overall control strategy 110 can execute the main strategy 104 within the subset of states L (depending on when the predicted cumulative success value is greater than the threshold) and can execute the support strategy 108 in all other states outside L.
[0074] In some embodiments, more complex decisions can be made by the overall control strategy 110. For example, in a multi-objective optimization problem, multiple GVF functions (where N is an integer, N>1) (which may include different cumulative amounts and / or different conversion factors) may be required to evaluate the success rate of the main strategy 104 in achieving multiple objectives. Additionally, in addition to the main strategy 104 and the support strategy 108 that the overall control strategy 110 can choose to execute, there may be one or more other strategies. In such cases, an RL algorithm can be used to learn the overall control strategy 110, and auxiliary information can be included when learning the overall control strategy 110.
[0075] Therefore, in at least one aspect, the goal of the RL agent 102 is to learn the support strategy 106, denoted as π H (s), where the support strategy 106 generates support actions to transfer the RL agent 102 from a first state where the main action 122 generated by the main strategy 104 may result in a failure outcome (as predicted by the GVF 106) to a second state (s ∈ L) where the main action 122 generated by the main strategy 104 may result in a success outcome (as predicted by the GVF 106). The overall control strategy 110's simultaneous selection of the support strategy 106 and the main strategy 104 can provide a more general solution to tasks applied to a larger state space S.
[0076] In an exemplary embodiment, it can be assumed that the value function Q M (s,a) that is typically used to evaluate the reward value associated with a specific state s and the main action 122 of the main strategy 104 does not exist or is unknown. Even in embodiments where such a value function is available, since existing value functions can only be trained or designed within a smaller subset of problem states L, existing value functions may be inaccurate in the entire state space s ∈ S, at least for this reason, so existing value functions can be ignored.
[0077] In some embodiments, the RL agent 102 is used to learn the GVF G using a policy evaluation algorithmM (s) 106, rather than learning the value function Q M (s,a). GVF G M (s) 106 predicts the future long-term performance (e.g., the performance can be based on executing the policy starting from the current state) of the primary action 122 generated by the RL agent 102 executing the primary policy 104 in a given state s in the state space S as the cumulative success value. Functionally, since GVF G M (s) 106 notifies the RL agent 102 when to use the primary policy 104 or the support policy 108, so GVF G M (s) 106 can be regarded as representing the initiation set of the primary policy 104. The initiation set is the set of all states from which an option can be called, and the termination function of this option outputs the probability of termination in a given state. In some examples, the GVF G of the future success rate of the RL agent 102 executing the primary policy 104 M (s) 106 prediction can also be used to terminate the primary policy 104 when the predicted success rate is unacceptable, and the unacceptable predicted success rate can be determined in various ways. For example, when the predicted success rate does not meet a preset threshold, the primary policy 104 can be terminated. In other words, GVF 106 can provide an output indicating when to initiate the primary policy 104 and when to terminate the primary policy 104. An option is defined by a policy, a termination function, and an initiation set, and it is a policy that can be executed for at least one time step before terminating according to the termination probability output by the termination function in the current state and switching to another option whose initiation set includes the current state. Learning GVF G M (s) 106 generates a function describing a subset of states L, where, for example, a larger value of GVF G M (s) 106 in a given state s can indicate that this state s is part of the solution subset L. Essentially, learning GVF G M (s) 106 can be a form of policy evaluation of the primary policy 104, but in a state space larger than the state space in which the primary policy 104 is designed or trained. The details of learning GVF G M (s) 106 are described in detail below.
[0078] Although the RL agent 102 can succeed in the first state within the subset of states L by executing the primary policy 104, when encountering a second state outside the subset of states L when executing the primary policy 104, it will produce undesirable results, such as generating non-optimal actions, generating constant or random actions, triggering exceptions, and / or task failures.
[0079] Therefore, as Figure 1As shown, the RL agent 102 includes a support policy processor 126 that is used to execute an RL algorithm to update or learn the parameters of the support policy 124 (denoted as π H (s)) 108 to transfer from one state outside the state subset L to another state within the state subset L. The support policy 124 maps states outside the state subset L to support actions and can be modeled as a neural network or the like. Executing the support action generated by the support policy 124 causes the robot 100 to transfer from one state outside the state subset L to another state within the state subset L. Specifically, the support policy processor 126 can be used to maximize the reward received from the reward processor 128 under the state s received from the state processor 114 as shown in Figure 1 . It should be noted that although the present invention takes the state processor 114, the support policy processor 126, and the reward processor 128 as separate processors, these components are not necessarily different physical processors. For example, the support policy processor 126 and the reward processor 128 can be implemented as software executed by a single physical processing unit within the RL agent 102.
[0080] In Figure 1 some embodiments as shown, the reward processor 128 receives the predicted cumulative success value output from the GVF 106. By learning the parameters of the support policy 108 at least partially based on the predicted success value of the main policy 104, the parameters of the support policy 108 are learned considering the success of the main policy 104. Therefore, by maximizing the reward of the support policy 108, the support policy 108 is learned to maximize the success rate of the main policy 104.
[0081] More specifically, the RL agent 102 receives data representing the observed state s t of the environment and the reward r t at time t. In response to each observed state s t , the RL agent 102 selects an action a t from the action space and executes it. After one time step (t + 1), partly due to the RL agent 102 executing the action a t , the RL agent 102 receives data representing the reward r t+1 at the next time step and the new state s t+1 of the environment. The RL agent 102 learns the support policy 108 based on the state transition tuples, where each transition tuple includes the state s t , the action a t , the reward r t and the next state s t+1, and uses the support policy 108 to output the support action 124 in the current state to maximize the cumulative reward of the predicted cumulative success value based on the main policy 104. The details of learning the support policy 108 are described below.
[0082] Figure 2 is a flowchart of an exemplary method 200 for support policy learning (e.g., learning of the support policy 108) performed by the RL agent 102 in the robot 100 provided by an exemplary embodiment.
[0083] In step 210, the RL agent 102 receives the existing main policy 104. The main policy 104 denoted as π M (s) maps the state s to the main action a. By generating the main action to be executed by the robot 100 in the environment according to the current state of the robot 100, the main policy 104 is successful within the state subset L, where the subset is a subset of the entire state space S.
[0084] As described above, the main policy 104 can be regarded as a "black box". For example, the main policy 104 can be a constructed or learned solution. For example, the main policy 104 can be constructed by rules (e.g., based on empirical experience) for generating the main action 122 defined manually. The performance of the RL agent 102 in executing the main policy 104 can be evaluated by the cumulative success value. It should be understood that at least because the value function Q M (s,a) can only be trained in a limited state space and cannot be successfully applied to the entire state space S, the cumulative success value is not determined by or related to this value function.
[0085] In step 220, a policy evaluation algorithm is used to learn the GVF 106 denoted as G M (s) to identify the state subset L. The value of the GVF 106 is the predicted cumulative success value, where a larger value of the GVF 106 indicates that the state s can be part of the solution state subset L. For example, the GVF 106 can be learned by function approximation (e.g., deep learning) based on the sampling performance of the main policy 104 at multiple states s sampled from the entire state space S (including the state subset L and other states outside L). The GVF 106 is used to: if the overall control policy 110 always executes the main action generated by the main policy 104 at a given state in the state space S, then predict the cumulative success value representing the future performance of the RL agent 102 through the cumulative quantity, where the cumulative quantity can be a measure of the success of the main policy 104. The cumulative quantity can be considered an indication of success at a given time and can be used by the GVF 106 as the basis for predicting the cumulative future success rate. For example, in some embodiments, when the main policy is executed at a given state, the GVF GM (s t ) 106 can predict the discounted sum of cumulants:
[0086]
[0087] where γ M ∈ [0, 1] is a discount factor similar to the support policy 108 (or other RL algorithms), and controls how far into the future the GVF 106 predicts the cumulative success value. Conceptually, when the main policy 104 is executed with a discount factor that "fades out" the success rate in the distant future, the GVF 106 predicts the cumulative success value by considering the sum of all future success rates (indicated by the cumulant c over all future time steps).
[0088] In some embodiments, the GVF 106 and the support policy 108 are learned separately. It should be understood that the GVF 106 can be learned by any number of suitable machine learning techniques, including Temporal Difference (TD) estimation and Monte Carlo (MC) estimation.
[0089] Figure 3 is a flowchart of an exemplary method 320 for learning the GVF 106 using TD estimation in step 220.
[0090] In step 322, the main policy 104, the initial state distribution the discount factor γ ∈ [0, 1], and the cumulative function f(·) are received. A computer-readable memory buffer B can be initialized to be empty, and the GVF 106 represented as G M (s; θ) and parameterized by θ is initialized to map the state s to the cumulative success value. The parameter θ can be manually selected (e.g., designed by a knowledgeable machine learning engineer) according to the specific task to be learned by the RL agent 102, etc.
[0091] In step 323, a trajectory and an initial state are initialized. The time step is initialized to t = 0, and the state of the environment is initialized to the initial state s0 ∈ S, where,
[0092] By evaluating the current state s t , the action a selected at the current time step t , the next state s at the next time step t+1 and the cumulant c associated with the transition from the current state s t to the next state s t+1 , the method 320 iteratively learns the GVF represented as G t+1 M The GVF 106 of (s; θ).
[0093] Specifically, starting from step 324, at the current time step t, sample the current state s of the environment from the state space. t . According to the current state s t , sample the action a from the action space t , such that a t ~ π M (·|s t ), where the action space includes the main actions generated by the main policy 104 in the states of the state space.
[0094] In step 326, at the next time step (t + 1), after executing the sampled action a t , sample a new state s from the state space t+1 .
[0095] In step 328, calculate the cumulative quantity c representing the success rate of the RL agent 102 achieving the goal when executing the sampled action a t at the given state as follows t+1 :
[0096] c t+1 = f(s t , a t , s t+1 ).
[0097] Here, f(·) can be any function for the state to transfer from s t to s t+1 . The cumulative quantity s t+1 can represent the success rate of the sampled action a t .
[0098] In step 330, store the state transition tuple including (s t , a t , c t+1 , s t+1 ) in the buffer B.
[0099] In step 332, update the GVF G M (s; θ) 106 according to a suitable TD learning algorithm. In Figure 3 shown in an exemplary embodiment, the TD learning algorithm using mini - batch gradient descent is used. Specifically, use the gradient averaged over all state transitions in a mini - batch (i.e., a part) of the data sampled from the buffer B M to perform gradient ascent update on the GVF 106 represented as G M
[0100] Gradient descent is an optimization algorithm commonly used to find the weights or coefficients of machine learning algorithms such as artificial neural networks and logistic regression. Generally speaking, gradient descent works by having the model make predictions on the training data and using the prediction error to update (and thus learn) the model in a way that reduces the error. Mini-batch gradient descent is a variant of the gradient descent algorithm that splits the training data set into small batches of data used to calculate the model error and update the model. In some examples, the gradients can be summed over the small batches of data, which can further reduce the variance of the gradients. It should be noted that although this invention describes examples where a buffer is used to perform mini-batch gradient descent, this is illustrative only and not intended to be limiting. In some examples, the entire buffer (instead of sampling small batches of data) is used to perform updates using gradient descent. In some examples, a buffer may not be used at all. Instead, the gradients of the most recently sampled state transitions can be calculated to update the GVF 106 (or policy) using a suitable RL algorithm. Although this invention describes embodiments that use a buffer in certain ways, it should be understood that other methods for collecting and storing samples and using those samples to perform updates can be used.
[0101] In some embodiments, for convenience, the parameter θ is removed and the GVF 106 is denoted as G M (s t ) and satisfies the Bellman Equation such that the expression c t+1 +γG M (s t+1 ) estimates the target value of the GVF G M (s t )106. Thus, the TD error can be calculated as:
[0102]
[0103] Here, γ is the future discount value applied to the predicted cumulative success value in the new state (t + 1), and γ ∈ [0, 1]. Then, the TD error can be backpropagated to update the GVF G M (s t )106.
[0104] In step 334, it is determined whether the convergence condition is satisfied. One convergence condition can be whether the predefined number of updates to the GVF 106 (in step 332) is satisfied (or exceeded). Another possible convergence condition can be whether a predefined desired performance level is reached.
[0105] When the convergence condition is satisfied, the learned GVF 106 is output, and in step 336, the RL agent 102 stores the learned GVF 106.
[0106] If the convergence condition is not met, method 320 proceeds to step 335 to determine whether the completion condition is met. By way of illustrative example, the completion condition may be that the robot 100 successfully completes a certain task in the environment, such as successfully parking or the robotic arm successfully picking the required parts from a box. If the completion condition is not met, method 320 returns to step 324. If the completion condition is met, method 320 returns to step 323 to reset the initial state and trajectory (e.g., reset the scenario environment).
[0107] It should be noted that if the environment is a non-scenario environment (which can be considered a special case of a scenario environment), then it is always determined in step 335 that the completion condition is not met, and thus the trajectory is never reset (i.e., it does not return to step 323).
[0108] Optionally, the GVF 106 can be learned through supervised learning (e.g., using MC estimation), where each state visited in the trajectory list is used to manually accumulate the cumulant. The supervised learning of the GVF 106 may be advantageous for tasks with a final reward (e.g., success / failure) and does not require a discount factor. Thus, in some embodiments, the learning step 220 can be implemented as supervised learning.
[0109] Figure 4 An exemplary method 420 for learning the GVF 106 using the MC estimation technique is shown. Generally speaking, the MC technique relies on repeated sampling of random simulations to estimate system properties. Method 420 can be applicable to cases where the cumulant is only final (e.g., for a scenario environment, received at the end of the scenario), such as a success or failure signal. However, since the goal is the sampled return rather than the expected return estimated using bootstrapped value estimates compared to method 320, it is known that method 420 has no bias but a relatively large variance.
[0110] In step 422, the primary policy 104, the initial state distribution the discount factor γ ∈ [0, 1], and the cumulant function f(·) are received. The computer-readable memory buffer B can be initialized to be empty, and the GVF 106 represented as G M (s; θ) and parameterized by θ is initialized to map the state s to the cumulative success value. Similar to step 322, the parameter θ can be manually designed.
[0111] In step 423, the trajectory and the initial state are used. The time step is initialized to t = 0, and the state of the environment is initialized to the initial state s0 ∈ S, where, different from method 320, at the initial time t = 0, the trajectory is stored in τ initialized to τ = [s0].[[]END]]
[0112] In step 424, at the current time step t, sample the current state s of the environment. According to the current state s, sample an action a from the action space generated by the main policy 104 t , such that a t ~π M (·|s t ).
[0113] In step 426, at the next time step (t + 1), after the robot 100 executes the sampled action in the environment, sample a new state s t+1 .
[0114] In step 428, calculate the cumulant c representing the performance of the sampled action as follows t+1 :[[]]
[0115] c t+1 = f(s t , a t , s t+1 ).
[0116] Here, f(·) can be any function of the state transition based on the state transitioning from s t to s t+1 . The cumulant c t+1 can represent the success rate of the sampled action a t .
[0117] In step 430, for example, by adding the state transition tuple to the trajectory list τ, i.e., τ + [a t , c t+1 , s t+1 , accumulate the cumulant in the trajectory list τ.
[0118] Steps 424 to 430 can be iteratively repeated until a completion condition is met. The completion condition can be, for example, the robot 100 successfully completing a certain task in the environment.
[0119] In step 432, calculate the cumulative reward R t (also known as the return) for each time step in the trajectory list as follows:
[0120]
[0121] where T is the number of time steps in an episode.
[0122] In step 434, store the tuple including the state and cumulative reward (s t , R t ) for each time step t = 0... T - 1 in the buffer B.
[0123] For the k-th iteration, in each iteration, a mini-batch of n tuples is sampled from buffer B in step 436. The parameters k and n can be manually selected. k determines the number of updates used to learn the GVF after collecting the sample trajectories.
[0124] In each iteration, in step 438, the GVF 106 represented as G M (s t ; θ) is updated using gradient descent, where the gradient is The gradient descent step updates the parameters θ in the differentiable function G M (s t ; θ), which minimizes the error between the predicted cumulative success value G M (s t ; θ) determined by the GVF 106 and the target cumulative success value R t collected by interacting with the environment. It should be noted that, different from method 320, in method 420, the entire sample trajectory is collected before updating the GVF.
[0125] In step 439, it is determined whether the convergence condition is satisfied. One convergence condition can be whether the predefined number of updates for the GVF 106 (in step 438) is satisfied (or exceeded). If the convergence condition is not satisfied, method 420 returns to step 423 to reset the trajectory and state (e.g., reset the episodic environment).
[0126] When the convergence condition is satisfied, the RL agent 102 stores the learned GVF 106 in step 440 (i.e., the RL agent 102 stores the GVF 106 with the learned parameters θ).
[0127] It should be noted that method 420 advances to step 432 according to the satisfaction of the completion condition. When the environment is a non-episodic environment (which can be considered a special case of the episodic environment), the completion condition will not be satisfied. Therefore, the above method 320 may be more suitable for non-episodic environments.
[0128] The above examples describe some methods for learning the GVF 106 based on the sampling performance of the main policy 104 at multiple states sampled from the state space. Specifically, this includes sampling states in the state space, where the main policy 104 is not designed or trained to achieve acceptable performance. In addition, it should be noted that the cumulant used to train the GVF 106 can be the same as or different from the reward calculated for the main policy 104 (in the case where the main policy 104 also uses RL learning).
[0129] In some embodiments of method 320 or method 420, the GVF 106 can perform off-policy learning, independent of the actions performed by the RL agent 102, especially when the known behavior and πM (·|s t ) In the case of probability, various methods can be used for off-policy learning of the GVF 106. For example, one off-policy method is to use the importance sampling ratio ρ given below:
[0130]
[0131] where μ(a|s) is the policy used by the RL agent 102 to sample actions in step 424. μ(a|s) can also be referred to as the behavior policy and can be predefined. The gradient is multiplied by the importance sampling ratio ρ to learn G M (s t ; θ).
[0132] Another off-policy method is to learn the GVF 106, which is a function of state and action, i.e., G M (s t , a t ; θ). The GVF can be recovered using the following equation:
[0133] G M (s t ) = G M (s t , a; θ)
[0134] where a ∼ π M (a|s) is the action sampled from the main policy. The TD error in gradient descent can be slightly modified by the following equation:
[0135]
[0136] where a ∼ π M (a|s) is the action sampled from the main policy.
[0137] Returning to Figure 2 , in step 230, the overall control policy 110 is obtained. In an example of the method 200, the overall control policy 110 can be constructed (e.g., using manually defined rules) instead of being learned. In other examples detailed below, the overall control policy 110 can be learned. The overall control policy 110 is used to select whether to execute the main policy 104 (such that the robot 100 executes the main action 122) or the learning support policy 108 based on the predicted cumulative success value determined by the GVF 106 in a given state.
[0138] Generally, when the predicted cumulative success value is an acceptable value, the overall control policy 110 selects to execute the main policy so that the robot executes the main action 122 according to the given state in the state space.
[0139] When the predicted cumulative success value is not an acceptable value, the overall control strategy causes the RL algorithm to be used to learn the support policy 108, and the support policy 108 generates a support action 124 to be executed by the robot 100 according to a given state. Executing the support action causes the robot 100 to execute the support action to transfer from the given state to a new state where the predicted cumulative success value has an acceptable value. More details are provided below.
[0140] Since the subset L of states where the main policy 104 may succeed is unknown, the subset L can be constructed mathematically as follows:
[0141]
[0142] Here, β is a defined threshold representing the acceptable value of the predicted cumulative success value determined by the GVF 106. Thus, the state subset L includes the states where the main action 122 generated by the main policy 104 achieves a predicted cumulative success value greater than the defined acceptable threshold. Accordingly, the overall control policy 110 can be defined as:
[0143]
[0144] Here, M represents the main policy 104 and H represents the support policy 108.
[0145] In step 240, the RL algorithm executed by the support policy processor 126 is used to learn the parameters of the support policy 108. The support policy 108 maps states to support actions. The support policy 108 that maps states to support actions 124 can be modeled as a neural network. Executing the support action 124 causes the robot 100 to execute the support action to transfer from the first state in the failure subspace to the second state within the success subset L. Specifically, the support policy 108 is learned using the reward generated from the reward processor 128, which is a function of the predicted cumulative success value generated according to the learned GVF 106. In other words, by using the reward based on the success of the main policy 104 to learn the parameters of the support policy 108, the RL agent 102 using the support policy 108 can make decisions that positively impact the long-term success of the main policy 104.
[0146] Figure 5 is a flowchart of an exemplary method 530 for performing the support policy learning step 240.
[0147] In step 532, the main policy π M (s)104, the GVF G M (s; θ)106, the overall control policy π(s)110, the initial state distribution (which may be different from those in steps 322 and 422), the discount factor γ H∈[0,1], reward function f H (·) and termination function h H The computer-readable memory buffer B can be initialized to be empty. Also as part of step 532, the support policy and the action-value function (also known as the Q-function) Support policy 108 is used to map the state s to the current support policy π H The generated action, where the action can be parameterized using a parameter set represented as . The function is the action-value function that predicts the future cumulative reward of the support policy 108. The support policy 108 selects the action that maximizes .
[0148] In step 533, the trajectory and the state are initialized. The time step is initialized to t = 0, and the state of the environment is initialized to the initial state s0 ∈ S, where
[0149] Then, the support policy processor 126 executes the RL algorithm to iteratively update or learn the parameters of the support policy 108, and the support policy 108 maps the state to the support action that maximizes the cumulative reward.
[0150] In each iteration, in step 534, at the current time step t, the current state s of the environment is sampled t ∈ S. According to the sampled current state, the master control policy 110 selects the primary policy 104 or the support policy 108. This can be mathematically expressed as α t ~ π(·|s t ), where α t ∈ {H, M}
[0151] Before the termination condition is satisfied, in step 536, at the next time step (t + 1) after the action is executed, a new state s is sampled t+1 . The executed action a t is the primary action 122 generated by the primary policy 104 or the support action 124 generated by the support policy 108, depending on the policy selected by the master control policy 110 in step 534. This can be mathematically expressed as follows
[0152] If α t = H, then a t ~ π H (·|s t )
[0153] If α t = M, then at to π M (·|s t )。
[0154] In step 538, the overall control strategy determines the next policy α t+1 based on the newly sampled state s t+1 (mathematically represented as α t+1 to π(·|s t+1 ))。
[0155] In step 540, the reward represented as is calculated according to the reward function. The reward function represented as f H (·) uses the state transition tuple (represented as the tuple (s t , a t , α t , s t+1 , α t+1 )) and the predicted cumulative success values of the initial state (represented as s t ) and the next state (represented as s t+1 ) (calculated by the GVF 106 for example) to calculate the reward. The reward function can be mathematically represented as follows:
[0156]
[0157] It should be understood that the reward function f H (·) can also be a function of other features in the state transition tuple (s t , a t , s t+1 ), and these features can include reward shaping and other terms to improve the support for policy learning. Here, there is an explicit dependency relationship between the predicted cumulative success values of the main policy 104 in two consecutive states G M (s t ) and G M (s t+1 ). As a non-limiting example, the reward function f H (·) can be selected from the following equations:
[0158]
[0159]
[0160] When G M (s t , a t ) is conditional on the action, or
[0161]
[0162] There may be other examples of reward functions that may also be based on the predicted cumulative success value of the primary policy 104. In some other embodiments, the reward function may be modified by reward shaping specific to the support policy. Reward shaping adds small rewards or penalties to the reward of the RL agent 102 in order to guide the agent 102 into the desired final state. For example, assume that the RL agent 102 is learning to park. Just from random behavior, it is unlikely to achieve the expected result of successful parking. Reward shaping prompts the agent 102 to get closer to achieving its goal, such as rewarding the agent 102 for approaching the parking space or penalizing the agent 102 for moving away from the parking space.
[0163] In step 542, a termination variable is determined using a termination function. The termination function h H (·) is as follows:
[0164]
[0165] Here, if the selected policy is the primary policy 104 (which may indicate that the RL agent 102 has transitioned to a state where the master control policy 110 believes the primary policy 104 may be successful), then the support policy 108 is terminated. However, it should be noted that the agent 102 still uses the primary policy 104 to interact with the environment until the scenario in the environment ends (e.g., because the goal has been achieved), or until the selected policy changes to the support policy 108, in which case the primary policy 104 is terminated (which may indicate that the master control policy 110 believes the primary policy 104 is unlikely to be successful).
[0166] Otherwise (e.g., the master control policy 110 believes that the support policy 108 needs to be further executed in the current state), as the iteration continues, the termination function is set to a discount factor that can be used to discount the rewards as described above.
[0167] In step 544, the state transition tuple including is stored in the buffer B.
[0168] In step 546, the support policy processor 126 updates the support policy 108. Specifically, any off-policy RL algorithm (e.g., Q-learning) is used to calculate the support policy 108, the action-value function represented as and the reward r H on a mini-batch of data sampled from the buffer B.
[0169] In step 548, it is determined whether a convergence condition is satisfied. One convergence condition can be whether a predefined number of updates (in step 546) is satisfied (or exceeded). Another possible convergence condition can be whether a predefined desired performance level is reached. When the convergence condition is satisfied, the RL agent 102 stores the learned support policy 108 in step 550.
[0170] If the convergence condition is not satisfied, the method 530 proceeds to step 549 to determine whether a completion condition is satisfied. By way of illustrative example, the completion condition can be that the robot 100 successfully completes a certain task in the environment, such as successfully parking or the robotic arm successfully picking the required part from a box.
[0171] If the completion condition is not satisfied, the time step t is updated such that the current state is now the newly sampled previous state (i.e., set t = t + 1), and then the method 530 returns to step 534.
[0172] If the completion condition is satisfied, the method 530 returns to step 533 to reset the trajectory and state (e.g., reset the scenario environment).
[0173] It should be noted that if the environment is a non-scenario environment (which can be considered a special case of a scenario environment), then it is determined in step 549 that the completion condition will never be satisfied, and thus it will not return to step 533 to reset the trajectory.
[0174] Back Figure 2 In step 250, after learning the GVF and the support policy, the RL agent 102 can be deployed. Figure 6 is a flowchart of an exemplary deployment method 600.
[0175] In step 602, the main policy 104, the learned support policy 108, the implemented overall control policy 110, and the initial state s0 at the initial time t = 0 are received.
[0176] In step 603, data representing the current state of the environment at time t is obtained. For example, the state s t can be received from the state processor 114.
[0177] In step 604, at the current state at time t, the overall control policy 110 makes a decision α t such that α t ~π(·|s t ), where α t∈ {H, M}. As described above, it is determined whether to execute the main policy 104 or the support policy 108 at least in part based on the predicted cumulative success value of the main policy 104 in a specific state st. Therefore, step 604 may include using the learned GVF to determine the predicted cumulative success value of the main policy 104.
[0178] Executing the selected policy generates an action based on the current state and causes the action to be executed by the robot 100 in the environment. If the main policy 104 (α t = M) is selected, the action a t may be the main action 122 (a t ~π H (·|s t )) or, if the support policy 108 (α t = H) is selected, it may be the support action 124 (a t ~π M (·|s t ))
[0179] In step 606, the RL agent 102 executes the action a t in the given state s t . As described above, the RL agent 102 executes the action by outputting the action to one or more controllers 116, and one or more controllers 116 generate one or more control signals to the actuator 118 to cause the robot 100 to execute the action.
[0180] In an example where the overall control policy 110 is not learned, the overall control policy 110 can make a choice between the main policy 104 and the support policy 108 by comparing the predicted cumulative success value with a predefined threshold. When the comparison result indicates that the predicted cumulative success value has an acceptable value, the overall control policy 110 causes the main action generated by the main policy to be output. When the comparison result indicates that the predicted cumulative success value has an unacceptable value, the overall control policy 110 causes the support action generated by the support policy to be output.
[0181] In step 608, after executing the action a t (e.g., after the robot 100 executes the action), a new state is sampled from the state space. The time step is also updated to t = t + 1.
[0182] The above examples can extend the utility and generality of the fixed main strategy. Different from known transfer learning methods that rely on the details of at least partially known main strategies (i.e., white box or gray box), according to the present invention, the SPL can be implemented by a black box main strategy. This is advantageous because a hybrid system can be constructed that can fully utilize the constructed and learned solutions. For the learned main strategy, this may also be an advantage because it is easier to learn the strategy in a smaller problem space before expanding the agent to a more complex and larger problem space.
[0183] As described above, in some examples, the overall control strategy 110 can be learned. Figure 7 FIG. 700 is a flowchart of an exemplary method for SPL for the robot 100 provided by another exemplary embodiment of the present invention, wherein the overall control strategy 110 is learned.
[0184] In addition to learning the overall control strategy 110 and the support strategy 108, the method 700 can be similar to the method 200. More specifically, instead of learning the overall control strategy 110 (e.g., a rule-based overall control strategy 110) and the support strategy 108 sequentially, the two strategies 108, 110 are learned simultaneously. In this exemplary method 700, the step 710 for receiving the main strategy, the step 720 for learning the GVF 106, and the step 740 for deploying the RL agent can be respectively similar to the steps 210, 220, and 250 of the method 200, and for the sake of brevity, they will not be described herein again.
[0185] In step 730, the overall control strategy 110 and the support strategy 108 are learned simultaneously. Figure 8 FIG. 800 is a flowchart of an exemplary method that can be used to execute step 730.
[0186] Refer to Figure 8 , in step 802, receive the main strategy 104 represented as π M (s) and the learned GVF 106 represented as G M (s; θ) (from step 720) and the initial state distribution (which can be different from that used in step 720). Also receive the support strategy discount factor γ H ∈ [0,1] and the overall control strategy discount factor γ ∈ [0,1]. Set the computer-readable memory buffer B to be empty. The received parameters include the support strategy reward function f H (·), the overall control strategy reward function f π (·), and the support strategy termination function h H (·). After receiving, initialize the parameters π(s; θ π ), Q π (s,α; θQ (e.g., initialized to a random number).
[0187] In step 803, the trajectory and state are initialized. Initialize the time step to t = 0, and initialize the state of the environment to the initial state s0 ∈ S, where,
[0188] Generally speaking, the overall control policy π(s) can be a function of multiple GVFs and states, including partial or all of the observations of the environment at time t, historical observations, or cumulative success prediction results of a set (possibly including different cumulative amounts and discount factors for each GVF) rather than just one.
[0189] In step 804, at the current time step t, sample the current state s t . Based on the current state s t , the overall control policy determines the decision α t to be α t ~π(·|s t ), where the decision α t is whether to execute the main policy 104 or the learning support policy 108 (i.e., α t ∈{H,M}). If the overall control policy 110 selects to execute the main policy 104 (α t = M), then execute the main policy 104 to generate the main action 124 to be executed by the robot 100 (a t ~π M (·|s t )). Optionally, if the overall control policy 110 selects to learn the support policy 108 (α t = H), then execute the support policy 108 to generate the support action 126 to be executed by the robot 100 (a t ~π H (·|s t )).
[0190] In step 806, after executing the decision action α t , sample a new state at the next time step t+1.
[0191] In step 808, according to the sampled new state s t+1 , the overall control policy 110 determines the new decision α t+1 ~π(·|s t+1 ).
[0192] In step 810, the reward processor 128 uses the support policy reward function f H (·) to calculate the support policy reward associated with the state transition as follows including the initial state st and the next state s t+1 for the predicted cumulative success value of:
[0193]
[0194] In step 812, the overall control policy reward function f π (·) calculates the overall control policy reward associated with the state transition as follows including the initial state s t and the next state s t+1 for the predicted cumulative success value of:
[0195]
[0196] It should be understood that the reward functions f H (·) and f π (·) are both functions of the predicted cumulative success values determined by the GVF 106 for the primary policies. The overall control policy reward function f π (·) can be different from and independent of the support policy reward function f H (·). Thus, there can be separate reward processors for the individual reward functions. For example, instead of Figure 1 the separate reward processor 128 shown, there can be two separate reward processors (or two instances of a reward processor) to generate rewards for the support policy 108 and the primary policy 104 respectively.
[0197] It should also be understood that the reward functions f H (·) and f π (·) can be functions of other features (s t , a t , s t+1 ) in the state transition, and these features can include reward shaping and other terms to improve the corresponding policy learning. For example, f π (·) can include a reward term for reducing the frequency of switching between the support policy 108 and the primary policy 104. Exemplary functions for f π (·) include (among other possibilities):
[0198]
[0199] or
[0200]
[0201] or
[0202]
[0203] In step 814, the support policy termination function Calculated as Terminate function h H An example of (·) can be similar to that described in step 542 of method 530.
[0204] In step 816, the state transition tuple is stored in buffer B.
[0205] In step 818, on the mini-batch data sampled from buffer B, using the reward r H and the termination function for the support policy 108 represented as and its action value function perform off-policy updates. This can be achieved by any off-policy RL algorithm (e.g., Q-learning, or off-policy gradient methods such as Deep Deterministic Policy Gradient (DDPG) and Soft-Actor Critic (SAC) methods).
[0206] In step 820, on the mini-batch data sampled from buffer B, using the reward r π for the overall control policy 100 represented as π(·|s t ; θ π ) and its action value function Q π (s,α; θ Q ) are updated. For example, this can be achieved by any suitable RL algorithm.
[0207] In step 822, determine whether the convergence condition is satisfied. One convergence condition can be whether the predefined number of updates (in steps 818 and 820) is satisfied (or exceeded). Another possible convergence condition can be whether a predefined desired performance level is reached (for the support policy or the overall control policy, or both). When the convergence condition is satisfied, the RL agent 102 stores the learned support policy 108 and the learned overall control policy 110 in step 826.
[0208] If the convergence condition is not satisfied, method 800 proceeds to step 824 to determine whether the completion condition is satisfied. By way of illustrative example, the completion condition can be that the robot 100 successfully completes a certain task in the environment, such as successfully parking or the robotic arm successfully picking the required parts from a box.
[0209] If the completion condition is not satisfied, update the time step t such that the current state is now the newly sampled state from before (i.e., set t = t + 1), and then method 800 returns to step 804.
[0210] If the completion condition is satisfied, method 800 returns to step 803 to reset the trajectory and state (e.g., reset the scenario environment).
[0211] It should be noted that if the environment is a non-scenario environment (which can be considered a special case of the scenario environment), it is determined in step 824 that the completion condition will never be satisfied, so it will not return to step 803 to reset the trajectory.
[0212] In some examples, off-policy methods for learning policies have been described (e.g., in combination with methods 520 and 820). It should be understood that on-policy methods can be used instead of off-policy methods. For example, to use an on-policy method to learn a support policy, samples can be stored only in the iterations of executing the support policy, and the policy can be updated after collecting sufficient samples in one or more trajectories (i.e., the list of trajectories can replace the buffer). Similar methods can be used for on-policy methods to learn the overall control policy.
[0213] Returning Figure 7 , as described above, the deployment step 740 can be similar to step 250 of method 200.
[0214] In addition to the possible advantages described above, examples of learning the overall control policy 110 can not set a success threshold β, which can make the overall control policy decision better to maximize the success rate of the main policy. In addition, such a method can also achieve more complex decisions, such as avoiding wavering (e.g., between selecting the main policy or the support policy) due to some local noise or inaccuracy in the learned GVF 106 (denoted as G M (s)). By learning the overall control policy 110 from a set of partial or all observations, historical observations, or prediction results (which may include different cumulants and discount factors for multi-objective tasks), more complex decisions can be made to help ensure the selection of the optimal support policy, thereby maximizing the success rate of the main policy.
[0215] Figure 9 is a flowchart of another method 900 for the SPL of the robot 100 provided by another exemplary embodiment of the present invention, wherein the overall control policy 110 is learned.
[0216] In addition to the ways of learning the GVF 106, the overall control policy 110, and the support policy 108, method 900 can be similar to method 200. The step 910 for receiving the main policy and the step 930 for deploying the RL agent can be similar to steps 210 and 250 of method 200 respectively. For the sake of brevity, they will not be elaborated here.
[0217] In step 920, the GVF 106, the overall control policy 110, and the support policy 108 are learned simultaneously, rather than separately learning the GVF 106, the overall control policy 110, and the support policy 108 as in methods 200 and 700. Figure 10 is a flowchart of an exemplary method 1000 that can be used to perform step 920.
[0218] Refer to Figure 10 , in step 1002, receive the main policy 104 represented as π M (s) and the initial state distribution Also receive the discount factor γ ∈ [0, 1]. Set the computer-readable memory buffer B to be empty. The received parameter θ includes the cumulative function f(·), the support policy reward function f H (·) and the overall control policy reward function f π (·). Initialize the GVF G M (s,a;θ), the support policy support policy action value function the overall control policy π(·|s;θ π ) and the overall control policy action value function Q π (s,α;θ Q )(e.g., initialized to random numbers).
[0219] In step 1003, initialize the trajectory and the state. Initialize the time step to t = 0, and initialize the state of the environment to the initial state s0 ∈ S, where,
[0220] Similar to method 700, the overall control policy 110, more simply represented as π(·|s), can be a function of multiple GVFs and states, including some or all of the observations of the environment at time t, historical observations, or the set of cumulative success prediction results (possibly including different cumulative amounts and discount factors for each GVF) rather than just one.
[0221] In step 1004, at the current time step t, sample the current state s from the entire state space S t . According to the current state s t , the overall control policy 110 will determine the decision α t as α t ~π(·|s t ), where the decision is to choose the main policy or the support policy (i.e., α t ∈{H,M}) to execute. The RL agent 102 executing the selected policy includes: generating an action through the selected policy and executing the action. If the main policy 104 is selected as the selected policy (α t= M), then the action can be the main action 124(a generated by the main policy 104 according to the current state t ~π M (·|s t ), or, if the support policy 108(α t = H) is selected, then the action can be the support action 126(a generated by the support policy 108 according to the current state t ~π H (·|s t ).
[0222] In step 1006, when the decision-making action α t is executed in the environment, a new state is sampled at the next time step t+1.
[0223] In step 1008, according to the sampled new state s t+1 , the overall control policy 110 determines a new decision α t+1 ~π(·|s t+1 ).
[0224] In step 1010, the cumulative quantity c representing the performance of the sampled action is calculated as follows t+1 :[[]]END]]
[0225] c t+1 = f(s t ,a t ,s t+1 ).
[0226] Here, f(·) can be any function for the state to transfer from s t to s t+1 . The cumulative quantity c t+1 can represent the success rate of the sampled action a t . The sampled action a t is sampled from the action space generated according to the decision of the overall control policy 110 from the main policy 104(i.e., a t ~π M (·|s t )) or the support policy 108(a t ~π H (·|s t ). The reward function f H (·) and f π (·) can also be functions of other features(s t ,a t ,s t+1 ) in the state transition. These features can include reward shaping and other terms to improve the corresponding policy learning. For example, f π(·) may include a reward item that reduces the frequency of the switching support policy and the overall control policy (e.g., avoiding wavering).
[0227] In step 1012, the state transition tuple (s t , a t , α t , c t+1 , s t+1 , α t+1 ) is stored in buffer B.
[0228] In step 1014, a mini-batch of data is sampled from buffer B.
[0229] For each state transition tuple within the mini-batch of data, in step 1016, the reward of the support policy is calculated as a function of the prediction cumulative success values of the main policy 104 and the decision-making policy determined by the overall control policy, as follows:
[0230]
[0231] Here, represents the main action 122 generated by the main policy 104 in state s t such that the main action 122 generated at the next time step t + 1 is such that
[0232] In addition, for each state transition tuple within the mini-batch of data, in step 1018, the termination function of the support policy 108 is calculated as follows:
[0233]
[0234] Furthermore, for each state transition tuple within the mini-batch of data, in step 1020, the reward associated with the overall control policy 110 is calculated as follows:
[0235]
[0236] After completing steps 1016 to 1020 for all transition tuples in the mini-batch of data, in step 1022, each of the GVF 106, the support policy 108, and the overall control policy 110 is updated.
[0237] Specifically, using the gradient calculated by averaging over all transition tuples in the mini-batch of data denoted as G M (s t , a t;Gradient ascent update is performed on the GVF 106 of (θ). Backpropagation can be used to adjust the TD error of the GVF, which can be calculated as:
[0238]
[0239] Here, is the next action sampled from the main policy.
[0240] The reward r obtained in step 1016 can be used on mini - batch data through any suitable off - policy method (e.g., Q - learning, DDPG, SAC, etc.) H and the termination function obtained in step 1018 to perform off - policy updates on the support policy 108 represented as and its corresponding action - value function for the support policy 108.
[0241] The reward r calculated in step 1022 can be used on mini - batch data through any suitable RL algorithm π to update the overall control policy 110 represented as π(·|s; θ π ) and its action - value function Q π (s,α; θ Q ).
[0242] In step 1024, it is determined whether the convergence condition is met. One convergence condition can be whether the predefined number of updates (in step 1022) for the GVF 106, support policy 108, and overall control policy 110 is satisfied (or exceeded). Another possible convergence condition can be whether a predefined desired performance level is reached (for the support policy or the overall control policy, or both). When the convergence condition is met, the RL agent 102 stores the learned GVF 106, learned support policy 108, and learned overall control policy 110 in step 1028.
[0243] If the convergence condition is not met, method 1000 proceeds to step 1026 to determine whether the completion condition is met. By way of illustrative example, the completion condition can be that the robot 100 successfully completes a certain task in the environment, such as successfully parking or the robotic arm successfully picking the required part from a box.
[0244] If the completion condition is not met, the time step t is updated such that the current state is now the newly sampled state from before (i.e., set t = t + 1), and then method 1000 returns to step 1004.
[0245] If the completion condition is met, method 1000 returns to step 1003 to reset the trajectory and state (e.g., reset the scenario environment).
[0246] Note that if the environment is a non-episodic environment (which can be considered a special case of an episodic environment), the completion condition will never be satisfied in step 1026, and thus the process will not return to step 1003 to reset the trajectory.
[0247] Returning Figure 9 , as described above, the deployment step 930 can be similar to step 250 of method 200.
[0248] In addition to the possible advantages discussed above, method 900 combines all learning (of GVF, support policy, and overall control policy) into a single algorithm, which advantageously makes the learning algorithm more data-efficient. Note that since the rewards of the support policy continuously change as the support policy is learned, it may be more difficult to adjust the learning of the support policy and the overall control policy. Since the GVF 106 represented as G M (s) is only a policy evaluation using a fixed policy, it can be learned faster.
[0249] As described above, in some examples, the present invention can be applied to multi-objective optimization problems, where multiple GVF functions and (possibly including different cumulants or different discount factors) are needed to evaluate the success rate of the main policy in achieving multiple objectives.
[0250] In some embodiments, the overall control policy 110 can be based on a learned threshold. In these embodiments, the overall control policy 110 can be considered a learned overall control policy 110 rather than a purely constructed overall control policy 110, where the overall control policy 110 is a hybrid policy based on a threshold-based construction policy (e.g., combined Figure 5 as described) and a fully learned policy (e.g., combined Figure 10 as described).
[0251] The hybrid overall control policy 110 can be defined as follows:
[0252]
[0253] where the learning represents a second GVF represented as G π (s t ). The second GVF can also be referred to as the overall control policy GVF to distinguish it from the GVF 106 learned for the main policy as described above (which can now also be referred to as the overall control policy GVF 106, represented as G M (s t ). The overall control policy GVF 106 learns to predict the future cumulative success value of executing the overall control policy 110 in a given state, rather than learning the entire overall control policy 110. When the overall control policy 110 is executed in a given state sampled from the state space, when the overall control policy GVF GM (s t ) has a predicted cumulative success value greater than or equal to that of the overall control policy GVF G π (s t ), the overall control policy 110 selects and executes the main policy 104 in the given state, and selects and learns the support policy 108 in all other states. The overall control policy GVF G π (s t ) predicts the cumulative success value of the overall control policy 110. For example, when in the given state, the predicted cumulative success value of the overall control policy GVFG M (s t ) is greater than or equal to that of the overall control policy GVF G π (s t ), the overall control policy GVF can predict the cumulative success value of executing the support policy 108 (denoted as H), and then switch to executing the main policy 104 (denoted as M). Therefore, the overall control policy GVF G π (s t ) provides a different prediction result from that of the overall control policy GVF 106 (G M (s t )) (which predicts the future cumulative success value of the main policy 104). In the above equation, the parameter ε ≥ 0 is a very small fixed value, which is used to explain the fact that computers usually cannot accurately represent floating-point numbers.
[0254] The overall control policy GVF G π (s t ) is learned using a cumulative quantity similar to the overall control policy reward discussed above in Figure 10 . However, the cumulative quantity denoted as used to learn the overall control policy GVF G π (s t ) is different (therefore, the cumulative quantity used to learn the overall control policy GVF G should not be considered equivalent to the reward discussed in π (s t ). The cumulative quantity used to learn the overall control policy GVF G Figure 10 can be called the overall control policy cumulative quantity. π (s t )
[0255] For example, the overall control policy cumulative quantity can be mathematically expressed as follows:
[0256] where,
[0257] where, is the support policy termination function calculated in step 1018, and also represents the termination function predicted by the overall control policy GVF. By introducing When the overall control policy 110 selects to execute the main policy 104 in a given state, the overall control policy GVF G π (s t ) predicts the same cumulative success value as the overall control policy GVF G M (s t ). In this way, the overall control policy GVF G π (s t ) is compared with the overall control policy GVF 106(G M (s t ). Another possible overall control policy cumulative quantity for learning the overall control policy GVF G π (s t ) can be mathematically expressed as follows:
[0258] where a ~ π M (·|s t+1 ). The overall control policy GVF 106G M (s t ) can be learned in a similar way (as described above) by sampling actions from the action space of the main policy 104. The main policy 104 can be learned using the TD error: π (s t ) can be learned in a similar way (as described above) by sampling actions from the action space of the main policy 104. The main policy 104 can be learned using the TD error:
[0259]
[0260] where a ~ π H (·|s t+1 ) is an action sampled from the action space defined as the possible main actions generated by the main policy 104.
[0261] Using a similar method, the following TD error can be used to learn the overall control policy GVF 106G π (s t ):
[0262]
[0263] where a ~ π H (·|s t+1 ) is an action sampled from the support policy 108, calculated in step 1018.
[0264] Typically, we may hope that the overall control policy GVF G π (s t) When collecting under the overall control strategy 110, approximate it with the cumulative quantity calculated in step 1010 for easy comparison with the GVF 106 of the main strategy 104 (where the overall control strategy GVF 106G of the main strategy 104 π (s t ) When collecting under the main strategy 104, approximate it with the same cumulative quantity calculated in step 1010 as described above).
[0265] It should be noted that different from the above-mentioned main strategy GVF 106G M (s t )(used to predict the cumulative success value of the main strategy 104), the overall control strategy GVF G π (s t ) is used to predict the cumulative success value of the overall control strategy 110, including switching to execute the support strategy 108 or the main strategy 104.
[0266] An exemplary method for learning the above-mentioned hybrid overall control strategy can be considered as Figure 10 a variant of method 1000 in. For ease of understanding, only the differences from method 1000 are discussed here.
[0267] In step 1002, in addition to the content described above in combination with Figure 10 , receive the parameter ε (used to define the hybrid overall control strategy), and initialize the overall control strategy GVF.
[0268] Steps 1003 to 1018 can be executed in a similar manner as described above in combination with Figure 10 .
[0269] In step 1020, calculate the overall control strategy cumulative quantity for learning the overall control strategy GVF ( as described above), rather than calculating the reward of the overall control strategy.
[0270] In step 1022, update the overall control strategy GVF, rather than updating the overall control strategy 110 (and updating the GVF 106 of the support strategy 108 and the main strategy 104, as described in combination with Figure 10 ).
[0271] Steps 1024 and 1026 can be executed in a similar manner as described above in combination with Figure 10 .
[0272] In step 1028, store the learned overall control strategy GVF G π (s t), rather than storing the learned overall control policy 110 (and storing the learned GVF 106 for the learned support policy 108 and main policy 104, as described in conjunction with Figure 10 ). Then, the learned overall control policy GVF G π (s t ) can be used as the "learned threshold" in the hybrid overall control policy.
[0273] In another possible embodiment, the overall control policy cumulative quantity can be defined to be the same as the cumulative quantity calculated in step 1010, i.e.:
[0274] c t+1 = f(s t , a t , s t+1 ).
[0275] Then, the overall control policy GVF G π (s t ) can be updated using the following equation:
[0276]
[0277] where a is the next main action (generated by the main policy 104) or the next support action (generated by the support policy 108), depending on the selection of the overall control policy 110 in the next state s t+1 .
[0278] In this embodiment, since the cumulative quantity calculated in step 1010 is used, step 1020 is no longer needed to calculate the overall control policy cumulative quantity. Using this method can conceptually be understood as representing that the overall control policy GVF G π (s t ) predicts the performance of the overall control policy 110.
[0279] Referring to Figure 11 , a schematic diagram of another exemplary robot 100(a) is shown. Similar to the robot 100 shown in Figure 1 , the robot 100(a) includes sensors 112, a state processor 114, and one or more controllers 116 that send control signals to one or more actuators 118. The robot 100(a) also includes an RL agent 102(a), and the RL agent 102(a) includes an overall control policy 110(a) that performs complex decision-making according to the methods described herein. The RL agent 102(a) includes a plurality of main policies 104(a), a plurality of GVFs 106(a), and the overall control policy 110(a). Similar to Figure 1Similar to the RL agent 102 shown, the RL agent 102(a) further includes a support policy processor 126, a support policy 108, and a reward processor 128.
[0280] The master policy 110(a) in the RL agent 102(a) can select among the best policies of multiple main policies 104(a). In Figure 11 the exemplary robot 100(a) shown, the RL agent 102(a) includes K main policies (denoted as The RL agent 102(a) learns K successful prediction values (denoted as Each successful prediction value corresponds to one of the K main policies 104 - 1 (e.g., the best main policy 104(a) (e.g., ) is selected according to the main policy 104(a) ( ) by the maximum successful prediction value, where the master policy 110(a) can execute the selected best main policy 104(a), where is greater than or equal to the successful prediction values of other main policies 104(a) and the threshold β. The master policy 110(a) can execute the support policy 108 in all states, where is less than the threshold β, k = 1...K.
[0281] Now referring to Figure 12 , a flowchart of an exemplary method 1200 for support policy learning (e.g., learning of the support policy 108) performed by the RL agent 102(a) of the robot 100(a) in Figure 11 is shown. The method 1200 starts at step 1210. In step 1210, the RL agent 102 receives K main policies 104(a). The K main policies 104(a) (denoted as ) each map the state s to a main action a. The K main policies 104(a) are all successful within a subset of states L k (where k = 1...K) by generating main actions to be performed by the robot 100(a) in the environment according to the current state of the robot 100(a), where each subset L k is a subset of the entire state space S.
[0282] As described above, the K main policies 104(a) can each be constructed or learned solutions. For example, the K main policies 104(a) can each be constructed by manually defining rules (e.g., based on empirical experience) that control the generation of the main actions 122. The performance of the RL agent 102 executing the K main policies 104(a) can be evaluated by accumulating success values. It should be understood that at least due to the value function Q M (s,a) can only be trained in a finite state space and cannot be successfully applied to the entire state space S, so the accumulated success value is not determined by or related to this value function.
[0283] Then, method 1200 proceeds to step 1220. In step 1220, K GVFs 106(a) represented as (where k = 1...K) are learned simultaneously. The value of each GVF 106(a) is a predicted accumulated success value, where a larger value of the GVF 106 indicates that the state s can be a solution state subset L (where k = 1...K) of the main policy k . For example, the GVF 106 can be learned by function approximation (e.g., deep learning) of the sampling performance under multiple states s sampled from the entire state space S (including the state subset L k and other states outside of L k , where k = 1...K) according to the main policy 104. The K GVFs 106(a) are used to: if the master control policy 110 always executes the main actions generated by the main policy 104(a) at a given state in the state space S, predict the accumulated success value representing the future performance of the RL agent 102(a) by accumulating quantities, where the accumulated quantity can be a measure of the success of the corresponding main policy 104(a). The accumulated quantity can be considered an indication of success at a given time and can be used by the K GVFs 106(a) as a basis for predicting the accumulated future success rate. For example, in some embodiments, when executing the main policy at a given state, the K GVFs 106(a) can predict the discounted sum of the accumulated quantities:
[0284]
[0285] where γ M∈[0,1] is a discount factor similar to the support policy 108 (or other RL algorithms), and controls how far into the future the K GVFs 106(a) predict the cumulative success value. Conceptually, when executing the optimal corresponding main policy among the K main policies 104(a) that "fade out" the success rate in the very distant future, the K GVFs 106(a) predict the cumulative success value by considering the sum of all future success rates (indicated by the cumulative amount c within all future time steps).
[0286] In some embodiments, the K GVFs 106(a) and the support policy 108 are learned separately. It should be understood that the K GVFs 106(a) can be learned by any number of suitable machine learning techniques, including Temporal Difference (TD) estimation and Monte Carlo (MC) estimation.
[0287] After step 1220, the method 1200 proceeds to step 1230. In step 1230, the overall control policy 110(a) is obtained. The overall control policy 110(a) can be constructed (e.g., using manually defined rules) instead of being learned. In other examples detailed below, the overall control policy 110(a) can be learned. The overall control policy 110(a) is used to select whether to execute the optimal main policy among the K main policies 104(a) or learn the support policy 108 based on the optimal predicted cumulative success value determined by the K GVFs 106(a) in a given state.
[0288] Typically, when the optimal predicted cumulative success value is an acceptable value, the overall control policy 110(a) causes the execution of the optimal main policy among the K main policies 104(a). Executing the optimal overall control policy includes: the optimal main policy among the K main policies 104(a) generates a main action 122 based on a given state in the state space, and the robot 100(a) executes the main action 122.
[0289] When the optimal predicted cumulative success value is not an acceptable value, the overall control policy 110(a) selects an action to learn the support policy 108 using an RL algorithm. The support policy 108 generates a support action 124 to be executed by the robot 100(a) based on a given state. The robot 100(a) executing the support action causes the robot 100(a) to transfer from the given state to a new state where at least one of the predicted cumulative success values has an acceptable value.
[0290] After step 1230, method 1200 proceeds to step 1240. In step 1240, the parameters of the support policy 108 are learned using the RL algorithm executed by the support policy processor 126. The support policy 108 maps states to support actions. The support policy 108 that maps states to support actions 124 can be modeled as a neural network. The RL agent 102(a) observes the first state of the robot 100(a) and executes the support policy 108 until the robot 100(a) enters a second state within the subset L of successful states of the best primary policy among the K primary policies 104(a). k Specifically, the support policy 108 is learned using the reward generated from the reward processor 128, which is a function of the predicted cumulative success value generated from the K learned GVFs 106. In other words, by using the reward based on the success of the best primary policy among the K primary policies 104(a) to learn the parameters of the support policy 108, the RL agent 102(a) using the support policy 108 can make decisions that positively impact the long-term success of the best primary policy among the K primary policies 104.
[0291] After step 1240, method 1200 proceeds to step 1250. In step 1250, after learning the K GVFs and the support policy 108, the RL agent 102(a) can be deployed into the robot 100(a) and used to control the robot 100(a).
[0292] In various examples, the present invention describes methods and systems for support policy learning. The RL agent (which can be implemented in a robot) is configured with an existing solution in the form of a primary policy 104 that generates primary actions to be executed by the robot to produce a desired result when solving a task or achieving a goal in an environment. The performance of the RL agent in executing the primary policy is measured by a cumulative success value. The RL agent is used to learn a generalized value function that is used to predict the cumulative success value representing the performance of the agent in executing the primary policy at a given state in the state space. The RL agent is also used to learn a support policy that is used to transfer from a state where the cumulative success value is less than or equal to an acceptable value to a state where the cumulative success value is greater than the acceptable value. The support policy can be learned from state transitions (represented by a tuple including a state, an action, a reward, and a next state) to maximize the cumulative reward, where the reward is a function of the cumulative success value of the primary policy.
[0293] By leveraging existing primary policies and learned supporting policies (based on the selection of the overall control policy), the RL agent can advantageously extend known solutions to end-to-end solutions in a larger state space within the environment.
[0294] In some examples, since existing solutions in the form of primary policies for simple tasks are migrated and reused for complex tasks in the same environment, learning from scratch (tabular rasa) is avoided.
[0295] Since existing technical solutions are retained and fixed, the RL agent implemented in the examples disclosed herein can be immune to catastrophic forgetting.
[0296] The RL agent implemented in the examples disclosed herein can flexibly adopt existing solutions as black boxes, without assuming whether the solutions are constructed (e.g., manually designed) or learned. Additionally, there is no need to assume the structure of the subset of the state space where the known solution is successful.
[0297] By being able to work with the black-box primary policy, the examples in the present invention can enable the RL agent to utilize constructed and learned solutions to build a hybrid system.
[0298] In some exemplary embodiments, the RL agent makes the reward associated with supporting policy learning a function of the predicted cumulative success value of the primary policy determined by the generalized value function. Thus, the way of implementing the policy in the present invention is different from traditional hierarchical reinforcement learning (HRL). Generally, traditional HRL divides complex tasks into many simple and independent subtasks with the aim of learning the optimal order for executing the subtasks. In contrast, the overall control policy provided by the present invention is at least in one aspect to maximize the performance of the primary policy. Therefore, the SPL provided by the examples in the present invention can construct a policy where there is an imposed dependency relationship between the supporting policy and the overall control policy that does not exist in HRL. These dependencies can advantageously adopt the black-box primary policy, thereby enabling seamless transfer between the supporting policy and the primary policy. Since there is no explicit policy dependency in HRL, there is no explicit inter-policy support.
[0299] In some embodiments, the SPL provided in the present invention is to learn how to effectively utilize the primary policy from states that were not initially part of the design protocol or the learning environment of the primary agent. In at least one aspect, the SPL is to improve the performance, generality, and efficiency of the primary policy by learning the supporting policy as a function of the success of the overall control policy, thereby establishing a direct dependency on the supporting policy so that the supporting policy supports the execution of the primary policy.
[0300] In some embodiments, the RL agent alone has a GVF and a support policy.
[0301] In some embodiments, the support policy and the overall control policy are learned simultaneously. By learning the overall control policy (instead of a rule-based overall control policy), the cumulative success value threshold may not be set. Additionally, for example, the learned overall control policy can achieve more complex decision-making by avoiding wavering due to some local noise or inaccuracy in the learned GVF.
[0302] In some embodiments, the GVF, the support policy, and the overall control policy are learned simultaneously, which can achieve higher data efficiency.
[0303] In some exemplary aspects, an exemplary method is described. The method may be performed by an agent in a robot that controls the robot's interaction with the environment. The method includes: receiving a given state of the environment; using a generalized value function learned for a given main policy, determining a predicted cumulative success value representing the future performance of the main policy in the given state; using an overall control policy to determine whether to select the main policy or the learned support policy by comparing the predicted cumulative success value with a predefined threshold: when the comparison result indicates that the predicted cumulative success value has an acceptable value, executing the main policy to cause the robot to perform a main action generated by the main policy according to the given state; or, when the comparison result indicates that the predicted cumulative success value has an unacceptable value, executing the support policy to cause the robot to perform a support action generated by the support policy according to the given state, wherein executing the support policy causes the robot to transition from the given state to a new state where the predicted cumulative success value has an acceptable value.
[0304] Although the present invention describes methods and processes by steps performed in a certain order, one or more steps in the methods and processes may be appropriately omitted or changed. Where appropriate, one or more steps may be performed in an order other than the described order.
[0305] Although the present invention has been described at least in part in terms of methods, those of ordinary skill in the art will understand that the present invention also pertains to various components for performing at least some aspects and features of the described methods, whether through hardware components, software, or any combination thereof. Accordingly, the technical solution of the present invention can be embodied in the form of a software product. A suitable software product can be stored in a pre-recorded storage device or other similar non-volatile or non-transitory computer-readable medium, including DVDs, CD-ROMs, USB flash drives, removable hard drives, or other storage media, etc. The software product includes instructions tangibly stored thereon that enable a processing device (e.g., a personal computer, server, or network device) to execute examples of the methods disclosed herein.
[0306] Without departing from the subject matter of the claims, the present invention may be embodied in other specific forms. The described exemplary embodiments are illustrative and not restrictive in all respects. Features selected from one or more of the above-described embodiments may be combined to create alternative embodiments not explicitly described, and features suitable for such combinations can be understood to be within the scope of the present invention.
[0307] All values and sub-ranges within the disclosed ranges are also disclosed. Additionally, although the systems, devices, and processes disclosed herein may include a specific number of elements / components, the systems, devices, and assemblies may be modified to include more or fewer such elements / components. For example, although any of the disclosed elements / components may be referred to in a singular quantity, the embodiments disclosed herein may be modified to include multiple such elements / components. The subject matter described herein is intended to cover and encompass all suitable technical variations.
Claims
1. A method performed by an agent in a robot that controls the interaction of the robot with the environment, characterized in that, The method includes: Receiving a main policy, where the main policy generates an action to be executed by the robot according to the state of the robot, and the performance of the agent executing the main policy is measured by an accumulated success value; Using a policy evaluation algorithm to learn the generalized value function of the main policy, where the generalized value function predicts the accumulated success value representing the future performance of the agent executing the main policy in a given state of the environment, and the given state is in the entire state space; Obtaining a master control policy, where the master control policy selects an action according to the predicted accumulated success value obtained from the generalized value function; When the predicted accumulated success value is an acceptable value, the action selected by the master control policy causes the main policy to be executed, so that the robot executes the main action generated by the main policy according to the given state in the state space; When the predicted accumulated success value is not an acceptable value, the action selected by the master control policy causes a support policy to be learned using a reinforcement learning algorithm, where the support policy generates a support action to be executed by the robot according to the given state, and the support action causes the robot to transfer from the given state to a new state where the predicted accumulated success value has an acceptable value.
2. The method according to claim 1, wherein The learning of the generalized value function using the policy evaluation algorithm includes: Performing multiple iterations, where each iteration includes: Sampling the action generated by the main policy according to the current state in the state space, where the action is executed by the agent so that the robot executes the action; After executing the action, sampling the next state in the state space; After executing the action, calculating an accumulated quantity by transferring from the current state to the next state, where the accumulated quantity represents the success value of the agent in the current state; Storing at least the accumulated quantity associated with the current state, the action, and the next state; Updating the generalized value function to predict the accumulated success value.
3. The method according to claim 2, wherein The generalized value function is updated using temporal difference learning or Monte Carlo estimation.
4. The method according to any one of claims 1 to 3, characterized in that, The support policy is learned according to a reward, where the reward is based on the predicted accumulated success value obtained from the generalized value function, and the generalized value function considers multiple states sampled from the state space.
5. The method according to any one of claims 1 to 3, characterized in that, The obtaining of the master control policy includes: determining a threshold; the master control policy is defined to select to execute the main policy when the success value output by the generalized value function is greater than the threshold, and is also defined to select to learn the support policy when the success value output by the generalized value function is not greater than the threshold.
6. The method according to claim 4, wherein The obtaining of the master control policy includes: determining a threshold; the master control policy is defined to select to execute the main policy when the success value output by the generalized value function is greater than the threshold, and is also defined to select to learn the support policy when the success value output by the generalized value function is not greater than the threshold.
7. The method according to any one of claims 1 to 3 and 6, characterized in that, The obtaining of the overall control policy includes: learning the support policy while learning the overall control policy, where the overall control policy is learned according to the overall control policy reward, the support policy is learned according to the support policy reward, and both the overall control policy reward and the support policy reward are based on the predicted cumulative success value obtained from the generalized value function.
8. The method according to claim 4, wherein The obtaining of the overall control policy includes: learning the support policy while learning the overall control policy, where the overall control policy is learned according to the overall control policy reward, the support policy is learned according to the support policy reward, and both the overall control policy reward and the support policy reward are based on the predicted cumulative success value obtained from the generalized value function.
9. The method according to claim 5, characterized in that The obtaining of the overall control policy includes: learning the support policy while learning the overall control policy, where the overall control policy is learned according to the overall control policy reward, the support policy is learned according to the support policy reward, and both the overall control policy reward and the support policy reward are based on the predicted cumulative success value obtained from the generalized value function.
10. The method according to claim 7, wherein The generalized value function, the overall control policy, and the support policy are learned simultaneously.
11. The method according to claim 8 or 9, characterized in that, The generalized value function, the overall control policy, and the support policy are learned simultaneously.
12. A processing unit in a robot, characterized in that, The processing unit executes machine-executable instructions to implement an agent to control the interaction of the robot with the environment, and the instructions cause the agent to perform the following operations: Receive a primary policy, where the primary policy generates an action to be executed by the robot according to the state of the robot, and the performance of the agent executing the primary policy is measured by the cumulative success value; Use a policy evaluation algorithm to learn the generalized value function of the primary policy, where the generalized value function predicts the cumulative success value representing the future performance of the agent executing the primary policy in a given state of the environment, and the given state is in the entire state space; Obtain an overall control policy, where the overall control policy selects an action according to the predicted cumulative success value obtained from the generalized value function; When the predicted cumulative success value is an acceptable value, the action selected by the overall control policy causes the primary policy to be executed, so that the robot executes the primary action generated by the primary policy according to the given state in the state space; When the predicted cumulative success value is not an acceptable value, the action selected by the overall control policy causes a support policy to be learned using a reinforcement learning algorithm, where the support policy generates a support action to be executed by the robot according to the given state, and the support action causes the robot to transfer from the given state to a new state where the predicted cumulative success value has an acceptable value.
13. The processing unit according to claim 12, characterized in that, The instructions cause the agent to learn the generalized value function in the following manner: Execute multiple iterations, where each iteration includes: Sample the action generated by the primary policy according to the current state in the state space, where the action is executed by the agent so that the robot executes the action; After executing the action, sample the next state in the state space; After performing the action, a cumulative quantity is calculated by transitioning from the current state to the next state, where the cumulative quantity represents the success value of the agent in the current state; At least store the cumulative quantity associated with the current state, the action, and the next state; Update the generalized value function using temporal difference learning.
14. The processing unit according to claim 13, wherein The generalized value function is updated using temporal difference learning or Monte Carlo estimation.
15. The processing unit according to any one of claims 12 to 14, characterized in that, The support policy is learned based on a reward, where the reward is based on the predicted cumulative success value obtained from the generalized value function, and the generalized value function considers multiple states sampled from the state space.
16. The processing unit according to any one of claims 12 to 14, characterized in that The instruction causes the agent to obtain the overall control policy by determining a threshold; the overall control policy is defined as selecting to execute the main policy when the success value output by the generalized value function is greater than the threshold, and is also defined as selecting to learn the support policy when the success value output by the generalized value function is not greater than the threshold.
17. The processing unit according to claim 15, characterized in that, The instruction causes the agent to obtain the overall control policy by determining a threshold; the overall control policy is defined as selecting to execute the main policy when the success value output by the generalized value function is greater than the threshold, and is also defined as selecting to learn the support policy when the success value output by the generalized value function is not greater than the threshold.
18. The processing unit according to any one of claims 12 to 14 and 17, characterized in that The instruction causes the agent to obtain the overall control policy by learning the support policy while learning the overall control policy, where the overall control policy is learned based on an overall control policy reward, the support policy is learned based on a support policy reward, and both the overall control policy reward and the support policy reward are based on the predicted cumulative success value obtained from the generalized value function.
19. The processing unit according to claim 15, wherein The instruction causes the agent to obtain the overall control policy by learning the support policy while learning the overall control policy, where the overall control policy is learned based on an overall control policy reward, the support policy is learned based on a support policy reward, and both the overall control policy reward and the support policy reward are based on the predicted cumulative success value obtained from the generalized value function.
20. The processing unit according to claim 16, wherein The instruction causes the agent to obtain the overall control policy by learning the support policy while learning the overall control policy, where the overall control policy is learned based on an overall control policy reward, the support policy is learned based on a support policy reward, and both the overall control policy reward and the support policy reward are based on the predicted cumulative success value obtained from the generalized value function.
21. The processing unit according to claim 18, wherein The instruction causes the agent to learn the generalized value function, the overall control policy, and the support policy simultaneously.
22. The processing unit according to claim 19 or 20, characterized in that, The instruction causes the agent to learn the generalized value function, the overall control policy, and the support policy simultaneously.
23. A computer-readable medium storing instructions for implementing an agent to control the interaction of a robot with the environment, characterized in that, When executed by a processing unit in the robot, the instruction causes the agent to perform the method according to any one of claims 1 to 11.
24. A method performed by an agent in a robot that controls the interaction between the robot and the environment, characterized in that: The method includes: Receive multiple main policies, where each corresponding main policy among the multiple main policies generates an action to be executed by the robot according to the state of the robot, and the performance of the agent executing the corresponding main policy is measured by an accumulated success value; Use a policy evaluation algorithm to learn the generalized value function of each corresponding main policy among the multiple main policies, where the generalized value function predicts the accumulated success value representing the future performance of the agent executing the corresponding main policy in a given state of the environment, and the given state is in the entire state space; Obtain a master control policy, where the master control policy selects an action according to the best predicted accumulated success value obtained from multiple generalized value functions; When the best predicted accumulated success value is an acceptable value, the action selected by the master control policy causes the best main policy to be executed, so that the robot executes the main action generated by the best main policy according to the given state in the state space; When the predicted accumulated success value is not an acceptable value, the action selected by the master control policy causes a support policy to be learned using a reinforcement learning algorithm, where the support policy generates a support action to be executed by the robot according to the given state, and the support action causes the robot to transfer from the given state to a new state in which at least one of the predicted accumulated success values has an acceptable value.
25. A processing unit in a robot, characterized in that, The processing unit executes machine-executable instructions to implement an agent to control the interaction between the robot and the environment, and the instructions cause the agent to perform the following operations: Receive multiple main policies, where each corresponding main policy among the multiple main policies generates an action to be executed by the robot according to the state of the robot, and the performance of the agent executing the corresponding main policy is measured by an accumulated success value; Use a policy evaluation algorithm to learn the generalized value function of each corresponding main policy among the multiple main policies, where the generalized value function predicts the accumulated success value representing the future performance of the agent executing the corresponding main policy in a given state of the environment, and the given state is in the entire state space; Obtain a master control policy, where the master control policy selects an action according to the best predicted accumulated success value obtained from multiple generalized value functions; When the best predicted accumulated success value is an acceptable value, the action selected by the master control policy causes the best main policy to be executed, so that the robot executes the main action generated by the best main policy according to the given state in the state space; When the predicted accumulated success value is not an acceptable value, the action selected by the master control policy causes a support policy to be learned using a reinforcement learning algorithm, where the support policy generates a support action to be executed by the robot according to the given state, and the support action causes the robot to transfer from the given state to a new state in which at least one of the predicted accumulated success values has an acceptable value.
Citation Information
Patent Citations
Robot real time control method based on environmental interaction
CN107292344A
Reinforcement learning method and device
US20190302708A1