Large-scale Hierarchical Reinforcement Learning
The hierarchical controller system addresses the challenges of hierarchical agents in complex environments by using target conditioning and enabling goal-conditioned behavior training, resulting in effective performance comparable to flat techniques.
Patent Information
- Application Number
- JP2024554200
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-07
- Filing Date
- 2023-06-07
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-06-07
AI Technical Summary
Existing hierarchical agents struggle to perform effectively in visually complex and partially observable 3D environments, and they rely on expert-generated data for goal representations in goal-conditioned reinforcement learning.
A hierarchical controller system that includes a low-level controller neural network and a high-level controller neural network, utilizing a target conditioning technique to effectively control agents in complex environments, and enabling the training of goal-conditioned behavior from any experience generated through both controllers.
The system achieves performance comparable to or exceeding that of flat techniques in difficult real-world tasks, allows for training across multiple tasks, and can be scaled for large industrial tasks while generalizing to new tasks.
Smart Images

Figure 2025518439000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Application No. 63 / 349,968, filed Jun. 7, 2023. The disclosure of the prior application is considered part of the disclosure of this application and is incorporated herein by reference.
[0002] This specification generally describes a system implemented as a computer program on one or more computers in one or more locations that controls an agent that interacts with an environment and performs tasks in that environment.
Background Art
[0003] A machine learning model receives an input and generates an output, such as a predicted output, based on the received input. Some machine learning models are parametric models that generate an output based on the received input and the values of the model's parameters.
[0004] Some machine learning models are deep models that utilize multiple layers of models to generate an output for the received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non - linear transformation to the received input to generate an output.
Prior Art Documents
Non - Patent Documents
[0005]
Non - Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] This specification generally describes a system implemented as a computer program on one or more computers in one or more locations that controls an agent that interacts with an environment and performs tasks in that environment.
[0007] At each time step, the agent receives the input observation and executes one of the actions in the set of actions. For example, the set of actions may include a fixed number of actions or may be a continuous action space.
[0008] Generally, the system controls the agent using a hierarchical controller. The hierarchical controller includes a low-level controller neural network and a high-level controller neural network.
[0009] The hierarchical controller generates options that can be used to control the agent. Generally, an option is a generalization of an action. For example, an option may include an action (as specified above), a high-level action that affects the selection of the action by the hierarchical controller, a target output (a "target option") that characterizes the target state of the environment that the agent should reach, and any other generalizable behavior.
[0010] This specification also describes a target conditioning technique for training the low-level controller neural network and the high-level controller neural network so that the hierarchical controller can be used to effectively control the agent using a target representation generated from the target output produced by the high-level controller neural network.
[0011] Particular embodiments of the subject matter described in this specification may be implemented to realize one or more of the following advantages.
[0012] Unlike existing hierarchical agents, the hierarchical system described in this specification can exhibit performance comparable to or exceeding that of other flat (non-hierarchical) techniques for difficult real-world tasks, such as visually complex partially observable 3D environments.
[0013] In addition, the described system can train a hierarchical controller to learn goal-conditioned behavior from any experience generated through a low-level controller, a high-level controller, or both training processes. This represents an advance in the field of goal-conditioned reinforcement learning, which has hitherto relied on expert-generated data within the environment to train an agent with respect to goal representations.
[0014] This specification also describes techniques for training a low-level controller for multiple tasks at once, which enhances the hierarchical controller's ability to generalize across multiple tasks. This is a specific example of the potential for abstraction, transfer, and skill reuse, which are characteristics of hierarchical reinforcement learning in challenging environments.
[0015] Furthermore, this specification describes techniques that enable a hierarchical controller to be used "at scale," i.e., trained and effectively used for large-scale industrial tasks, and in some cases, generalized to new industrial tasks that were not encountered during the training of the low-level controller, the high-level controller, or both.
[0016] To effectively control an agent using a hierarchical controller, this specification describes various techniques that can be used together or separately to improve the performance of an agent in such complex environments.
Means for Solving the Problems
[0017] In one exemplary method described herein, a method for controlling an agent that interacts with an environment to perform a task includes, in each of a plurality of first time steps from a plurality of time steps, receiving an observation characterizing the state of the environment at the first time step, determining a target representation for the first time step characterizing a target state of the environment that the agent is to reach, and using a low-level controller neural network to process the observation and the target representation to generate a low-level policy output that defines an action to be performed by the agent in response to the observation. The low-level controller neural network includes a representation neural network configured to process the observation to generate an internal state representation of the observation, and a low-level policy head configured to process the observation state representation and the target representation to generate the low-level policy output. The method includes controlling the agent using the low-level policy output.
[0018] The step of determining a target representation for a first time step characterizing a target state of the environment to be reached by the agent may comprise determining whether criteria for generating a new target representation are met at the first time step. When the criteria are not met, the method may use the target representation from the preceding first time step as the target representation for the first time step. In one exemplary implementation, when the criteria are met, the method comprises generating a high-level observation result for the first time step, processing a high-level input comprising the high-level observation result using a high-level controller neural network to generate a high-level policy output comprising a target output characterizing the target state, and processing the target output using a target encoder neural network to generate a target representation. The high-level policy output may further comprise an indication of whether to control the agent using (i) a low-level controller or (ii) a basic action specified by the high-level policy output. In one exemplary implementation, the step of determining a target representation for the first time step, the step of processing the observation result and the target representation using a low-level controller neural network, and the step of controlling the agent using the low-level policy output may be performed only in response to a determination to control the agent using a low-level controller based on the indication. The method comprises, for each of one or more second time steps from a plurality of time steps, generating a high-level observation result for the second time step, processing the high-level observation result for the second time step using a high-level controller neural network to generate a high-level policy output for the second time step, determining to control the agent using a basic action specified by the high-level policy output based on an indication of whether to control the agent using (i) a low-level controller or (ii) a basic action specified by the high-level policy output, and in response thereto, controlling the agent using the basic action specified by the high-level policy output.
[0019] The high-level observation result for the first time step may comprise the observation result received at the first time step. The high-level observation result for the first time step may comprise data identifying the number of time steps since the most recent time step at which the criterion was met. The high-level observation result for the first time step may comprise data characterizing the reward received from the environment since the most recent time step at which the criterion was met. The high-level observation result for the first time step may comprise data characterizing the observation results received at time steps since the most recent time step at which the criterion was met. A criterion may be met when any one of a set of criteria is met. The high-level observation result for a time step may comprise data identifying which criteria were met at the first time step.
[0020] The target output characterizing the target state may be a target vector characterizing the target state. The target output characterizing the target state may be a text string describing the target state. The high-level policy output may comprise hyperparameters defining some aspect of the training. The step of controlling the agent using the low-level policy output may comprise the step of generating an adjusted policy output by applying the hyperparameters to the low-level policy output and the step of selecting an action using the adjusted policy output. The high-level policy may use a temperature parameter as a hyperparameter to adjust the low-level policy output.
[0021] The step of determining whether a criterion for generating a new target representation is met at the first time step may comprise, when the first time step is the first time step in a task episode, the step of determining that the criterion for generating a new target representation is met. The step of determining whether a criterion for generating a new target representation is met at the first time step may comprise the step of determining that the maximum number of time steps have elapsed since the most recent time step at which the criterion was met.
[0022] The low-level controller neural network may further comprise a value head configured to process an observation representation and a target representation to generate a value estimate of the value of an environment in a state characterized by the observation that a target state characterized by the target representation has been reached. The step of determining whether a criterion for generating a new target representation is satisfied at a first time step determines that the criterion for generating a new target representation is satisfied when the value estimate generated by the value head by processing the observation representation of the observations at the preceding time step and the target representation for the preceding time step is less than an unreachability threshold.
[0023] The step of determining whether a criterion for generating a new target representation is satisfied at a first time step determines that the criterion for generating a new target representation is satisfied when the value estimate generated by the value head by processing the observation representation of the observations at the preceding time step and the target representation for the preceding time step are greater than an achieved threshold indicating that the target characterized by the target representation for the preceding time step has been achieved. The achieved threshold may be included in the high-level policy output.
[0024] The step of determining whether a criterion for generating a new target representation is satisfied at a first time step comprises using a classifier neural network to process (i) the observations at the preceding time step and (ii) classifier inputs characterizing the target state for the preceding time step to generate a classifier output indicating how many time steps remain until the target state is achieved, and determining that the criterion for generating a new target representation is satisfied when the classifier output indicates that zero time steps remain until the target state is achieved.
[0025] In another exemplary method described herein, a method of training a low-level controller neural network includes obtaining one or more trajectories of observation-action pairs and respective goals for each trajectory; for each trajectory, generating respective rewards for each pair in the trajectory based on whether the goal is achieved in a state characterized by the observations in the pair; and training the low-level controller neural network on the one or more trajectories and the one or more rewards to optimize a reinforcement learning objective, the reinforcement learning objective comprising: (i) being trained on a behavior imitation loss; and (ii) for a given observation and a given goal representation, processing the given observation representation and the given goal representation for the given observation generated by a representation neural network to generate a behavior imitation policy output for the observation, and imposing a proximity regularization term that penalizes the low-level policy output generated by the low-level controller neural network for deviating from the behavior imitation policy output generated by a behavior imitation head configured to generate the behavior imitation policy output.
[0026] The behavior imitation policy output and the low-level policy output can each define a probability distribution over a plurality of actions. The proximity regularization term, for each observation in each trajectory, is based on the divergence between (i) the low-level policy output generated by the low-level policy head by processing the observation representation of the observation generated by the representation neural network and the target representation of the target for the trajectory, and (ii) the behavior imitation policy output generated by the behavior imitation policy head by processing the observation representation of the observation generated by the representation neural network and the target representation of the target for the trajectory. For example, this divergence can be the KL divergence. The reinforcement learning objective can further comprise, for each observation-action pair in each trajectory, the product of (i) a term based on the reward for the pair and (ii) the ratio between (a) the probability assigned to the action in the pair by the low-level policy output generated by the low-level policy head by processing the observation representation of the observation generated by the representation neural network and the target representation of the target for the trajectory and (b) the probability assigned to the action in the pair by the behavior imitation policy output generated by the behavior imitation policy head by processing the observation representation of the observation generated by the representation neural network and the target representation of the target for the trajectory. For example, the term based on the reward for the pair can be the V-Trace policy gradient term.
[0027] The method can further comprise training the behavior imitation policy head for one or more trajectories to minimize the behavior imitation loss.
[0028] For one or more of the trajectories, the goal can be an image of the environment in the goal state. The training step may further comprise processing the image using an image goal encoder neural network to generate a goal representation to be processed by a low-level policy head and an imitation behavior policy head. For one or more of the trajectories, the goal can be text describing the goal state of the environment. The training step may further comprise processing the text using a text goal encoder neural network to generate a goal representation to be processed by a low-level policy head and an imitation behavior policy head.
[0029] One or more trajectories can be sampled from offline data generated while the agent was being controlled by one or more different behavior policies. The low-level controller can be trained only on the offline data.
[0030] Alternatively, the low-level controller can be trained on both the offline data and online data generated by controlling the agent using both the high-level policy output generated by the high-level controller and the low-level policy output generated by the low-level controller.
[0031] The method may further comprise training the high-level controller neural network through reinforcement learning to maximize the expected reward for a task generated as a result of controlling the agent based on the high-level policy output generated by the high-level controller neural network and the low-level policy output generated by the low-level controller neural network. In some exemplary implementations, no gradients are passed to the low-level controller neural network as a result of training the high-level controller neural network.
[0032] One or more trajectories may include one or more online trajectories generated while controlling an agent during training of a high-level controller or a low-level controller. In an alternative implementation, one or more trajectories do not include any online data generated while controlling an agent during training of a high-level controller.
[0033] The reinforcement learning objective may comprise one or more auxiliary loss terms. Each auxiliary loss term may correspond to a respective auxiliary task. Each of the auxiliary tasks may require generating a prediction characterizing a target distribution conditioned on an output generated by a representational neural network.
[0034] The method may further comprise, for each of one or more trajectories of observation-action pairs, generating a respective target. The generating step may comprise, for one or more of the trajectories, selecting a target that describes a state of the environment characterized by the last observation in the trajectory.
[0035] The method may further comprise, for each of one or more trajectories of observation-action pairs, generating a respective target. The generating step may comprise, for one or more of the trajectories, selecting a target that describes a state of the environment not characterized by any of the observations in the trajectory.
[0036] The agent may be a mechanical agent and the environment may be a real-world environment. For example, the agent may be a robot. The environment may be a real-world environment of a service facility comprising a plurality of electronic devices, and the agent may be an electronic agent configured to control the operation of the service facility. The environment may be a real-world manufacturing environment for manufacturing a product, and the agent may comprise an electronic agent configured to control a manufacturing unit or machine operating to manufacture the product.
[0037] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
Brief Description of the Drawings
[0038]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0039] Like reference numerals and designations in the various drawings indicate like elements.
[0040] FIG. 1 shows an exemplary action selection system 100. The action selection system 100 is an example of a system implemented as a computer program on one or more computers at one or more locations where the systems, components, and techniques described below are implemented.
[0041] The action selection system 100 controls an agent 104 that interacts with an environment 106 to perform a task by selecting an action 108 to be executed by the agent 104 at each of a plurality of time steps during the execution of an episode of the task.
[0042] As a general example, the task may include one or more of, for example, moving to a specified location in the environment 106, identifying a particular object in the environment 106, manipulating a particular object in a specified manner, controlling a plurality of devices to meet a criterion, allocating resources to a plurality of devices, and the like. More generally, the task is specified by a received reward 130, that is, the episodic return is maximized when the completion of the task is successful. Rewards and returns are described in more detail below. Examples of agents, tasks, and environments are also given below.
[0043] An “episode” of a task is a series of interactions in which an agent 104 attempts to execute a single task starting from some initial state of the environment 106. In other words, each task episode begins with the environment 106 in an initial state, e.g., a fixed initial state or a randomly selected initial state, and ends when the agent 104 succeeds in completing the task or when some termination criterion is met, e.g., the environment 106 enters a state designated as a termination state, or when the agent 104 executes an action 108 a threshold number of times without succeeding in completing the task.
[0044] At each time step during any given task episode, system 100 receives an observation 110 that characterizes the current state of environment 106 at the time step, and in response, selects an action 108 to be executed by agent 104 at the time step. After agent 104 executes action 108, environment 106 transitions to a new state, and system 100 receives a reward 130 from environment 106.
[0045] In general, reward 130 is a scalar value that characterizes the progress of agent 104 with respect to completing the task.
[0046] As a specific example, reward 130 can be a sparse binary reward that is 0 unless the action results in a successful completion of the task, i.e., it is non-zero, e.g., equal to 1, only if the action results in a successful completion of the task.
[0047] As another specific example, reward 130 can be a dense reward that measures the progress of the agent with respect to completing the task at each individual observation time point received during an episode of attempting to execute the task, i.e., non-zero rewards may be received and frequently received before a successful completion of the task.
[0048] During the execution of any given task episode, system 100 selects action 108 in an attempt to maximize the return received over the course of the task episode.
[0049] That is, at each time step during the episode, system 100 selects an action 108 that attempts to maximize the return received for the remainder of the task episode starting from the time step.
[0050] In general, at any given time step, the return received is a composition of the rewards received at time steps that are after the given time step in the episode.
[0051] For example, at a certain time step t, the return is Σ i γ i-t-1 r i which can satisfy, where i ranges over all time steps after t in the episode or any fixed number of time steps after t in the episode, γ is a discount factor greater than 0 and less than or equal to 1, and r i is the reward at time step i.
[0052] To control agent 108, at each time step in the episode, the action selection subsystem 102 of system 100 processes the observation result 110 using the hierarchical controller 130 and selects the action 108 to be executed by agent 104 at the time step.
[0053] The hierarchical controller 130 includes a high-level controller (HLC) 126 and a low-level controller (LLC) 122.
[0054] LLC 122 is a neural network configured to receive an input including the observation result 110 and a target output characterizing the target state of the environment 106 that agent 104 should reach when interacting with the environment 106, and this input is processed into a target representation ("target") 124. The LLC processes the input and generates a low-level output that defines the action 108 to be executed by agent 104 in response to the observation result 110 conditioned by the target representation 124.
[0055] HLC 126 is a neural network configured to process a high-level input and generate a high-level policy output including a target option output.
[0056] In some implementations, the high-level policy output includes only the target output.
[0057] In some other implementations, the high-level policy output includes additional data.
[0058] As a specific example, the high-level policy output can also specify the basic behavior.
[0059] The basic behavior can be used to advance the system among multiple goals 124 or to generate training data for the low-level controller. In this case, the high-level policy output includes an indication of whether to control the agent using (i) the low-level controller or (ii) the basic behavior specified by the high-level policy output.
[0060] As another specific example, the high-level policy output can also specify a high-level behavior that changes how the policy output is processed to select an action. Specifically, this can be done to assist in agent exploration during training.
[0061] The components of the HLC neural network 126 are shown in FIG. 3, and the components of the LLC neural network 122 are shown in FIG. 4.
[0062] To select an action 108 at a given time step during an episode using the controller 130, the system determines whether the criteria for generating a new option are met at the time step. Determining whether the criteria are met is described in more detail below with respect to FIG. 2A.
[0063] In some cases where the criteria are not met, the system can use the goal representation 124 from a previous time step in the episode as the goal 124 for the time step.
[0064] The system then uses the LLC neural network 122 to select an action 108 conditioned on the goal representation 124.
[0065] If the criteria are met, the system generates a high-level policy output using the HLC neural network 126.
[0066] In implementations where the high-level policy output includes only the target output, the system generates a new target representation 124 using the target output and uses the new target representation as the target representation for the time step. The system then selects an action 108 conditioned on the target representation 124 using the LLC neural network 122.
[0067] In implementations where the high-level policy output also indicates whether to control the agent using (i) the LLC 122 or (ii) the basic actions specified by the high-level policy output, the hierarchical controller 130 can select an action 108 to be performed by the agent 104 at a time step based on the indication using the policy output of either the HLC 126 or the LLC 122.
[0068] If the controller 130 decides to control the agent using the LLC 122, the controller 130 generates a new target representation 124 using the target output and uses the new target representation as the target representation for the time step. The system then selects an action 108 conditioned on the target representation 124 using the LLC neural network 122.
[0069] If the controller 130 decides to control the agent using the HLC 126, the controller 130 selects the basic actions from the high-level output as the actions 108 for the time step.
[0070] In some cases, the HLC 126 can output a series of consecutive basic actions to advance the system before selecting a new target output.
[0071] In certain examples, at any given step within an episode, the action selection subsystem 102 can use the HLC 126 to process the observation results 110 and generate a target output that is sent to the LLC 122 to be processed into the target representation 124. The time step at which the HLC 126 generates the target output used to generate the target representation 124, and the time step at which the LLC 122 acts using the target representation 124, are referred to herein as the "first time step".
[0072] In this example, the LLC 122 can then process the observation results 110 to generate a policy output that is conditioned on the target representation 124. The system 100 can then select an action 108 using the policy output.
[0073] Furthermore, at another time step within the episode, the low-level controller 110 can capture the observation results 110 received after the first action 108 and send it to the high-level controller 126 in the form of a high-level input.
[0074] In some examples, this high-level input can include a high-level observation result that summarizes the actions 108 of the agent 104 in the environment 106 since the last time the HLC 126 provided the target output used to generate the target representation 124. The criteria for generating a new target representation are discussed in FIG. 2A.
[0075] The high-level controller 126 can process this high-level input to generate a high-level policy output that characterizes whether to control the agent 104 using the LLC 122 with a target output that is processed into the target representation 124 in successive time steps.
[0076] At some time steps, the action selection subsystem 102 can use the HLC 126 to process high-level inputs and generate basic action options that use the high-level policy output to control the agent 104 in the environment 106. That is, a "basic action" is one of the set or space of actions that can be executed by the agent, and is called "basic" because it is generated as part of, or defined by, the high-level policy output. The time steps at which the HLC generates basic action options are referred to herein as "second time steps".
[0077] As another example, at some time steps, the action selection subsystem 102 can use the high-level controller 126 to generate high-level action options that change how actions are selected from the policy output. Examples of high-level action options are treated in more detail in FIG. 3.
[0078] Thus, the agent executes under the options provided by the hierarchical controller, i.e., the options can specify directly selecting an action, changing how an action is selected, or providing a target output that the low-level controller 122 processes into the target representation 124 to control the agent 104.
[0079] Possible criteria for generating options using the hierarchical controller are treated in more detail in FIG. 2A.
[0080] The controllers 122, 126 can use either the same or different methods to process the policy output into the action 108. As discussed below in the context of possible processing methods, "policy output" can refer to either the low-level controller policy output or the portion of the high-level output that can be used to identify basic actions.
[0081] In one example, the policy output may include respective numerical probability values for each action in a fixed set. System 102 can select action 108, for example, by sampling actions according to the probability values for the action index or by selecting the action with the highest probability value.
[0082] In another example, the policy output may include respective Q-values for each action in a fixed set. System 102 can process the Q-values (e.g., using a soft-max function) to generate respective probability values for each action that can be used to select action 108 (as described previously), or can select the action with the highest Q-value.
[0083] The Q-value for an action is an estimated value of the return that results from the agent 104 performing action 108 in response to the current observation 110 and then selecting future actions 108 to be performed by the agent 104 according to the current values of the parameters of the controllers 122, 126.
[0084] As another example, when the action space is continuous, the policy output may include parameters of a probability distribution over the continuous action space, and System 102 can select action 108 by sampling from the probability distribution or by selecting the mean action. A continuous action space is an action space that includes an infinite number of actions, i.e., each action is represented as a vector having one or more dimensions, and for each dimension, the action vector can take on any value within the range for the dimension, with the only constraint being the precision of the numerical format used by the system 100.
[0085] As yet another example, when the action space is continuous, the policy output may include a regressed action, i.e., a regression vector representing an action from the continuous space, and System 102 can select the regressed action as action 108.
[0086] Before using controllers 122 and 126 to control an agent, a training system 190 within system 100 or another training system can train controllers 122 and 126.
[0087] Specifically, training system 190 can train HLC 126 to select target options that direct the interaction of LLC 122 with environment 106 to effectively execute tasks by the agent, and can train LLC 122 to effectively select actions under a certain target representation.
[0088] Training hierarchical controller 130 using training system 190 will be discussed in more detail below with reference to FIG. 2A.
[0089] In some implementations, the environment is a real-world environment and the agent is a mechanical agent that interacts with the real-world environment, such as a robot that operates or moves through the environment, or an autonomous or semi-autonomous land, air, or sea vehicle, and the actions are actions taken by the mechanical agent in the real-world environment to perform a task. For example, the agent can be a robot that interacts with the environment to perform a specific task, such as finding an object of interest in the environment, or moving an object of interest to a specified location in the environment, or traveling to a specified destination in the environment.
[0090] In these implementations, the observation results can include, for example, one or more of an image, object position data, and sensor data for capturing the observation results as the agent interacts with the environment, such as from an image sensor, a distance sensor, or a position sensor, or sensor data from an actuator. For example, in the case of a robot, the observation results can include data characterizing the current state of the robot, such as one or more of joint position, joint velocity, joint force, torque, or acceleration, such as gravity-compensated torque feedback, and the global or relative pose of an article held by the robot. In the case of a robot or other mechanical agent or vehicle, the observation results can similarly include one or more of position, linear velocity or angular velocity, force, torque or acceleration, and the global or relative pose of one or more parts of the agent. The observation results may be defined in one, two, or three dimensions and may be absolute and / or relative observations. The observation results can also include, for example, detected electronic signals such as motor current or temperature signals, and / or, for example, image or video data from a camera or LIDAR sensor, such as data from the agent's sensors or data from sensors located separately from the agent in the environment.
[0091] In these implementations, the actions can be control signals for controlling a robot or other mechanical agent, such as torques for the joints of a robot or high-level control commands, or control signals for controlling an autonomous or semi-autonomous land, air, or sea vehicle, such as control surfaces or other control elements, such as torques to steering control elements of a vehicle, or high-level control commands. The control signals can include, for example, position, velocity, or force / torque / acceleration data for one or more joints of a robot or parts of another mechanical agent. The control signals can further or alternatively include electronic control data, such as motor control data, or more generally, data for controlling one or more electronic devices in an environment where the control has an impact on the observed state of the environment. For example, in the case of an autonomous or semi-autonomous land, air, or sea vehicle, the control signals can define actions for controlling navigation, such as steering, and actions for controlling movement, such as braking and / or accelerating a vehicle.
[0092] In some implementations, the environment is a simulation of the real-world environment described above, and the agent is implemented as one or more computers that interact with the simulated environment. For example, the simulated environment can be a simulation of a robot or vehicle, and the reinforcement learning system can be trained on the simulation and, once trained, be used in the real world.
[0093] In some implementations, the environment is a real-world manufacturing environment for manufacturing products such as chemical, biological, or mechanical products, or food. As used herein, "manufacturing" a product includes purifying raw materials to make the product or processing raw materials to produce, for example, a cleaned or recycled product by removing contaminants. A manufacturing plant may include a plurality of manufacturing units such as containers for chemical or biological substances or machines for processing solids or other materials, such as robots. The manufacturing units are configured such that intermediate versions or components of the product can move between manufacturing units during the manufacture of the product, for example via piping or mechanical conveyance. As used herein, "manufacturing" a product also includes manufacturing food by a kitchen robot.
[0094] An agent may include an electronic agent configured to control manufacturing units or a machine such as a robot that operates to manufacture a product. That is, an agent may include a control system configured to control the manufacture of chemical, biological, or mechanical products. For example, the control system may be configured to control one or more of the manufacturing units or machines or to control the movement of intermediate versions or components of the product between manufacturing units or machines.
[0095] As an example, tasks performed by an agent may include tasks for manufacturing a product or an intermediate version or component thereof. As another example, tasks performed by an agent may include tasks for controlling the use of resources, such as minimizing the consumption of power, or water, or any material or consumable used in the manufacturing process.
[0096] Actions may include control actions for controlling the use of machines or manufacturing units for processing solid or liquid materials in order to manufacture a product or an intermediate version or component thereof, or for controlling the movement of intermediate versions or components of a product within a manufacturing environment, for example, between manufacturing units or machines. Generally, an action can be any action that affects an observed state of the environment, for example, an action configured to adjust any of the sensed parameters described below. These can include actions for adjusting the physical or mechanical conditions of a manufacturing unit, or actions for controlling the movement of the mechanical parts of a machine or the joints of a robot. Actions can include actions that impose operating conditions on a manufacturing unit or machine, or actions that result in a change in settings to adjust, control, or turn on or off the operation of a manufacturing unit or machine.
[0097] Reward or return may relate to a measure of task execution. For example, in the case of a task for manufacturing a product, the measure can include the quantity of the product manufactured, the quality of the product, a measure of the speed of manufacture of the product, or the physical cost of executing the manufacturing task, for example, the quantity of energy, materials, or other resources used to execute the task. In the case of a task for controlling the use of resources, the measure can include any measure of the amount of resources used.
[0098] In another example, the agent may be a human or animal agent, and controlling the agent may comprise outputting instructions or signals configured to cause the agent to perform an action. For example, the instructions or signals may be output to the agent using an output device such as a display device, a speaker, or a tactile device. In this case, the task may be a task in a real-world environment, which may be any real-world environment including the examples described above. In some implementations, the user may be a user of a digital assistant such as a smart speaker, a smart display, or other device. And the digital assistant may be used to instruct the user to perform an action. For example, controlling the agent may comprise instructing the digital assistant to issue instructions to a human user via the digital assistant for the actions the user should perform at each of a plurality of time steps. The instructions may be generated in the form of natural language, transmitted as, for example, sound and / or text on a screen, based on actions selected by a reinforcement learning system. The actions may be selected to contribute to the execution of the task. An observation result capture subsystem, such as a monitoring system including a video camera or an audio capture system, may be provided to capture visual and / or auditory observation results of the user performing the task. This may be used to monitor if any, the actions the user actually performs at each time step, and to provide observation results characterizing the state of the environment.
[0099] In some of the implementations described above, the environment may include a human or an animal. For example, the agent may be an autonomous vehicle in an environment where there is a human, such as a pedestrian or a driver / passenger of another vehicle and / or an animal, or the autonomous vehicle itself may carry a human. As another example, the environment may include at least one room, such as in a dwelling, in which one or more humans are present. The human or animal may be an element of the environment involved in the task.
[0100] In a further example, the environment may comprise a user who interacts with an agent in the form of a user device, such as a computer, a mobile device, or a digital assistant. The user device provides a user interface between the user and a computer system, which may be the same computer system or a different computer system that implements the controller. The user interface may enable the user to input data into and / or receive data from the computer system, and the agent may perform information transfer tasks related to the user, such as providing information to the user about a topic and / or enabling the user to specify components of a task to be performed by the computer system, under the control of the controller. For example, the information transfer task may be to teach the user skills such as how to speak a language or how to proceed around a geographical location. As another example, the task may be to enable the user to define a shape for the computer system such that the computer system can control an additional manufacturing (3D printing) system to produce an object having that shape. The action may comprise, for example, presenting information to the user in a certain format, at a certain rate, and / or configuring the interface to receive input from the user. For example, the action may comprise setting questions for the user to perform with respect to the skill, such as asking the user to choose from multiple options regarding the correct usage of a language, or asking the user to read aloud a passage of a language, and / or receiving input from the user, such as registering a selection of one of the options or recording a spoken passage of the language using a microphone.
[0101] Generally, the observation results of the state of the environment may comprise any electronic signals representing the function of electronic and / or mechanical devices. For example, the representation of the state of the environment may be derived from the observation results made by sensors that detect the state of the manufacturing environment, such as sensors that detect the state or configuration of a manufacturing unit or machine, or sensors that detect the movement of raw materials between manufacturing units or machines. As some examples, such sensors may be configured to detect mechanical movement or force, pressure, temperature, electric current, voltage, frequency, impedance and other electrical conditions, the quality, level, flow rate / movement speed or flow path / movement path of one or more raw materials, physical or chemical conditions, such as physical state, shape, or configuration, or chemical state such as pH, the configuration of a unit or machine such as the mechanical configuration of a unit or machine, or the configuration of a valve, may be an image or video sensor for capturing an image or video observation result of a manufacturing unit or machine or movement, or may be any other suitable type of sensor. In the case of a machine such as a robot, the observation results from the sensors may include the observation results of the position, linear velocity or angular velocity, force, torque or acceleration, or posture of one or more parts of the machine, for example, data characterizing the current state of the machine or robot, or an article held or processed by the machine or robot. The observation results may also include, for example, detected signals such as motor current or temperature signals, or image or video data from, for example, a camera or LIDAR sensor. Such sensors may be part of an agent in the environment or may be placed separately therefrom.
[0102] In some implementations, the environment is the real-world environment of a service facility that includes multiple electronic devices, such as a server farm or data center, e.g., a remote communication data center, or a computer data center for storing or processing data, or any service facility. The service facility may include auxiliary control devices for controlling the operating environment of its devices, such as environmental control devices for temperature control, e.g., cooling devices, or air flow control or air conditioning devices. The task may include a task for controlling the use of resources, e.g., minimizing, such as a task for controlling power consumption or water consumption. The agent may include an electronic agent configured to control the operation of its devices or to control the operation of auxiliary control devices, such as environmental control devices.
[0103] Generally, an action can be any action that affects the observed state of the environment, e.g., an action configured to adjust any of the detected parameters described below. These can include actions for controlling or imposing operating conditions on the devices or auxiliary control devices, e.g., actions that result in a change in settings to adjust, control, or turn on or off the operation of a device or an auxiliary control device.
[0104] Generally, the observation result of the state of the environment can include any electronic signal that represents the function of the facility or the devices within the facility. For example, the representation of the state of the environment can be derived from observations made by any sensor that detects the state of the physical environment of the facility or from observations made by any sensor that detects the state of one or more devices or one or more auxiliary control devices. These include sensors configured to detect electrical conditions such as current, voltage, power, or energy, the temperature of the facility, the flow rate, the temperature or pressure within the facility or within the cooling system of the facility, or the physical facility configuration such as whether a vent is open or not.
[0105] A reward or return may relate to a measure of task performance. For example, in the case of a task for controlling the use of a resource, such as a task for controlling the use of power or water, for example to minimize it, the measure may comprise any measure of resource use.
[0106] In some implementations, the environment is the real-world environment of a power generation facility, such as a renewable power generation facility like a solar power plant or a wind power plant. The task may comprise a control task for controlling the power generated by the facility, for example for controlling the delivery of power to a power distribution system, for example to meet demand, or to reduce the risk of mismatches between elements of the system, or to maximize the power generated by the facility. The agent may comprise an electronic agent configured to control the generation of power by the facility or the coupling of the generated power to the system. The actions may comprise actions for controlling the electrical or mechanical configuration of a generator, such as the electrical or mechanical configuration of one or more renewable power generation elements, for example a wind turbine or a configuration of one or more solar panels or mirrors, or the electrical or mechanical configuration of a rotating generator. The mechanical control actions may comprise, for example, actions for controlling the conversion of energy input to electrical energy output, for example the efficiency or degree of coupling of the conversion of energy input to electrical energy output. The electrical control actions may comprise, for example, actions for controlling one or more of the voltage, current, frequency, or phase of the generated power.
[0107] A reward or return may relate to a measure of task execution. For example, in the case of a task for controlling the delivery of power to a power distribution system, the measure may relate to a measure of the power delivered, or a measure of electrical mismatches between the power generation facility and the system, such as mismatches in voltage, current, frequency, or phase, or a measure of power or energy losses in the power generation facility. In the case of a task for maximizing the delivery of power to a power distribution system, the measure may relate to a measure of the power or energy delivered to the system, or a measure of power or energy losses in the power generation facility.
[0108] Generally, the observation results of the state of the environment may comprise any electronic signal that represents the electrical or mechanical function of a generator in a power generation facility. For example, the representation of the state of the environment may be the observation results made by any sensor that detects the physical or electrical state of equipment in a power generation facility that is generating power, or may be derived from the physical environment of such equipment, or the state of auxiliary equipment that supports the generator. Such sensors may include sensors configured to detect the electrical state of equipment such as current, voltage, power, or energy, the temperature or cooling of the physical environment, the flow rate, or the physical configuration of the equipment, and the observation results of the electrical state of the system from, for example, local or remote sensors. The observation results of the state of the environment may also comprise one or more predictions regarding the future operating conditions of a generator, such as predictions of future wind levels or solar irradiance, or predictions of the future electrical state of the system.
[0109] As another example, the environment may be a chemical synthesis or protein folding environment in which each state is the respective state of a protein chain or one or more intermediates or precursor chemicals, and the agent is a computer system for determining how to fold the protein chain or synthesize the chemical substance. In this example, the action is a possible folding action for folding the protein chain or an action for assembling the precursor chemical substance / intermediate, and the result to be achieved may include, for example, folding the protein so that the protein is stable and achieves a specific biological function, or providing an effective synthetic route for the chemical substance. As another example, the agent may be a mechanical agent that executes or controls protein folding actions or chemical synthesis steps automatically selected by the system without interaction with a human. The observation results may comprise direct or indirect observation results of the state of the protein or chemical substance / intermediate / precursor, and / or may be derived from a simulation.
[0110] Similarly, the environment may be a drug discovery environment in which each state is a respective state of a compound that may have drug activity, and the agent is a computer system for determining elements of a compound having drug activity and / or a synthetic route of a compound having drug activity. The drug / synthesis may be designed, for example in a simulation, based on a reward derived from a target of the drug. As another example, the agent may be a mechanical agent that executes or controls the synthesis of a drug.
[0111] In some further application examples, the environment is a real-world environment and the agent manages the allocation of tasks across computing resources, for example on a mobile device and / or in a data center. In these implementations, the action may include assigning a task to a particular computing resource.
[0112] As a further example, the action may include presenting an advertisement, the observation result may include an advertisement impression or click-through count or rate, and the reward may characterize a previous selection of an item or content made by one or more users.
[0113] In some cases, the observation result may include a textual or spoken command given to the agent by a third party (e.g., the operator of the agent). For example, the agent may be an autonomous vehicle, and the user of the autonomous vehicle may provide a textual or spoken command to the agent (e.g., to proceed to a particular location).
[0114] As another example, the environment can be an electrical, mechanical, or electromechanical design environment, e.g., an environment in which the design of an electrical, mechanical, or electromechanical entity is simulated. The simulated environment can be a simulation of a real-world environment in which the entity is intended to operate. The task can be to design the entity. The observations can include observations that characterize the entity, i.e., observations of the mechanical shape, or the electrical, mechanical, or electromechanical configuration of the entity, or observations of the parameters or properties of the entity. The actions can include actions that modify the entity, e.g., modify one or more of the observations. The reward or return can include one or more measures of the performance of the design of the entity. For example, the reward or return can relate to one or more physical properties of the entity, such as weight or strength, or one or more electrical properties of the entity, such as a measure of the efficiency of performing a particular function for which the entity is designed. The design process can include outputting a design for manufacturing, e.g., in the form of computer-executable instructions for manufacturing the entity. The process can include creating the entity according to the design. Thus, the design of the entity can be optimized, e.g., by reinforcement learning, and then output as an optimized design for manufacturing the entity, e.g., as computer-executable instructions. The entity can then be manufactured using the optimized design.
[0115] As described previously, the environment can be a simulated environment. Generally, in the case of a simulated environment, the observation results may include one or more simulated versions of the observation results described previously, or the types of observations and actions may include one or more simulated versions of the actions or types of actions described previously. For example, the simulated environment may be a motion simulation environment, such as a driving simulation or a flight simulation, and the agent may be a simulated vehicle that progresses through the motion simulation. In these implementations, the action can be a control input for controlling the simulated user or the simulated vehicle. Generally, the agent can be implemented as one or more computers that interact with the simulated environment.
[0116] The simulated environment can be a simulation of a particular real-world environment and agents. For example, the system may be used to select actions in a simulated environment during training or evaluation of the system, and after training, evaluation, or both are complete, it may be deployed to control real-world agents in the particular real-world environment that was the subject of the simulation. This can avoid unnecessary wear, tear, and damage to the real-world environment or real-world agents, and can enable a control neural network to be trained and evaluated for rare situations, or situations that are difficult or unsafe to reproduce in the real world. For example, the system may be partially trained using a simulation of a mechanical agent in a simulation of a particular real-world environment and then deployed to control a real mechanical agent in the particular real-world environment. Thus, in such a case, the observation results of the simulated environment are related to the real-world environment, and the selected actions in the simulated environment are related to the actions to be performed by the mechanical agent in the real-world environment.
[0117] Optionally, in any of the above implementations, the observation results at any given time step may include data from previous time steps that may be useful in characterizing the environment, such as actions executed at previous time steps, rewards received at previous time steps, or both.
[0118] FIG. 2A is a flowchart of an exemplary process 200 of a target representation training subsystem for a hierarchical controller showing how new options are generated by a high-level controller. For convenience, process 200 is described as being executed by a system of one or more computers located at one or more locations. For example, an action selection system appropriately programmed according to this specification, such as action selection system 100 of FIG. 1, can execute process 200.
[0119] The system can execute process 200 at some of the time steps during a series of time steps within an episode, for example, at each time step (step 202) when some criteria for generating a new option are met.
[0120] The criteria for generating a new option may also be referred to as the criteria for generating a new target representation.
[0121] The set of criteria may include one or more of various criteria. Some specific criteria that may be included in the set of one or more criteria are described herein.
[0122] As an example, these criteria may be that a previous option has been executed.
[0123] For example, when the previous option was a basic action, the system can determine that the criteria are met at the immediately following time step (even if none of the other criteria described below are met).
[0124] As another example, the system can determine that the criterion is met when the goal represented by the current goal representation has been fulfilled (or achieved). Exemplary techniques for determining when the current goal has been fulfilled are described in more detail below.
[0125] As another example, these criteria for generating new options can be accompanied by the same termination criteria indicating the end of a task episode. Specifically, they can correspond to the end of an episode in which the task has been completed, i.e., the goal specified by the previous goal representation has been met, or a threshold number of steps have been taken while attempting to complete the task.
[0126] As another example, these criteria can define new options every N environmental time steps, where N is an integer greater than or equal to 1.
[0127] As another example, these criteria can be based on a measure corresponding to goal achievement, such as determining whether it is possible to achieve the specified goal representation in a discrete number of future time steps from the state of the environment at the current time step. That is, the criterion can be met when the system determines that it is not possible to achieve the goal represented by the current goal representation within that number of future time steps.
[0128] In some examples, the achievability of a goal can be determined by comparing a value estimate provided by a low-level controller to a threshold. If the value estimate is higher than the achieved threshold, this indicates that the goal characterized by the goal representation has been fulfilled. If the value estimate is less than the unachievable threshold, this indicates that the goal representation is not achievable within the selected number of time steps. These value thresholds are discussed in more detail below with respect to FIGS. 4 and 6.
[0129] In some other examples, auxiliary calculations such as a logic-based system or a learnable method, such as a network or a linear mapping, that is trained to classify whether a goal is achievable in a set number of future steps, can be used to determine the unachievability of the goal at the current time step. These auxiliary calculations are dealt with in more detail in FIG. 4. Once it is determined that the goal is unachievable, the execution of the agent under the selected goal representation can be terminated early.
[0130] After determining that the criteria for generating a new option are met, the system generates a new option using the high-level controller (step 204).
[0131] The high-level controller output may include one or more types of options, including one-step or multi-step options, and criteria for selecting from them.
[0132] In some examples, the output may include only goal options. A goal option is a multi-step option that characterizes the goal state of the environment that the agent should reach. The goal option can be represented as text, an image, or any other type of data that can express the goal that the agent should fulfill.
[0133] Specifically, the goal option may include a short text describing a goal state such as "a ball on the table".
[0134] In another specific example, the goal option may include an image characterizing the goal state, such as an image showing a ball on the table.
[0135] In some other examples, the output may include other option types.
[0136] In some examples, the output may include a basic action, a one-step option that directly controls the agent in the environment using the high-level policy output to select an action for the agent.
[0137] In another example, the output may include a high-level action, a one-step option that can affect how the hierarchical controller selects an action, such as by setting a criterion for option termination, e.g., a maximum option time length threshold, or by applying noise or a function to the policy output used to select the agent's action. In some examples, the low-level policy output is affected by these high-level actions.
[0138] In a particular example that allows more than one type of option, the high-level controller output may also include an indication for selecting the type of option it outputs at that time step (step 206).
[0139] Specifically, the high-level policy output may include an indication of whether to (i) control the agent using a low-level controller or (ii) control the agent using the basic action specified by the high-level policy output.
[0140] In this case, the high-level controller can use a learned model or a logic-based rule system to determine which option to select.
[0141] In some examples, the indication of which output to select can be determined using a learned multi-class classification.
[0142] Specifically, the high-level controller may include a high-level option decision network that is trained to output an option probability distribution used to select from a plurality of option classes. In this case, either the option with the highest probability can be selected or the option can be sampled from the distribution.
[0143] As an example, a high-level controller can generate target options, basic options, and high-level actions, and (i) output target options that a low-level controller can process into a target representation to generate a policy output that controls the agent's actions regarding this target, (ii) output one-step basic actions that direct the agent to directly select actions using the high-level policy, and (iii) decide to output high-level actions that can change how actions are determined from the low-level policy output.
[0144] Once options are generated, the system controls the agent's interaction with the environment using the new options (step 208).
[0145] For example, the system can select a target option, use a low-level policy to process the target option into a target representation, and then select actions that are conditioned on the target representation.
[0146] In this case, the high-level controller can decide to reuse the target options of multiple steps from previous steps until the above criteria are met in consecutive time steps. This decision substantially delegates control of the agent to the low-level controller for a number of time steps until it is determined that the target has been fulfilled or is impossible to achieve, in which case a new target representation or another option can be generated as detailed above. As another example, the system can select a basic action and then directly select action 108 using the high-level policy.
[0147] As yet another example, the system can select a certain high-level action and use that high-level action to change the threshold for option termination.
[0148] As yet another example, the system can select a certain high-level action and use that high-level action to change how action 108 is selected by the low-level policy output.
[0149] In some cases, this change may only affect the selection of one action.
[0150] In other cases, this high-level action can give a permanent or non-cancellable change to how the low-level policy output selects actions.
[0151] Specifically, the high-level action can define parameters that affect the investigation over consecutive time steps until another high-level action sets new parameters.
[0152] For example, the high-level action can define a change to the temperature parameter τ used in a temperature-dependent softmax formula. This type of high-level action is treated in more detail in Figure 3.
[0153] The system continues to execute process 200 until a certain number of iterations for training the hierarchical controller have been performed, or until an end criterion for training the hierarchical controller is met, for example, until the task is successfully executed, until the environment reaches a specified end state, or until the maximum number of time steps has elapsed during an episode.
[0154] Figure 2B shows in more detail how HLC126 and LLC122 interact.
[0155] HLC126 includes an observation encoder 301, an RNN core 302, and a policy network 303. These components are described in more detail with respect to Figure 3.
[0156] LLC122 includes a target encoder 408, an observation encoder 405A, one or more target models 406, an RNN core 405B, and policy and value networks 409. These components are described in more detail with respect to FIG. 4.
[0157] In this example, HLC126 outputs either a target option for LLC122 or a basic action that directly interacts with the environment 106, i.e., is directly used as an action to be performed by the agent 104 in the environment 106.
[0158] When HLC126 outputs a target option to LLC122, the target option is processed into a target representation 124 that is used to condition the policy output of LLC122. LLC122 controls the agent 104 in the environment 106 under this target representation 124 until the criteria handled in FIG. 2A are met.
[0159] In this particular case, when the criteria are met, LLC122 can generate a summarized experience output that functions as an input to HLC126. This summarized output is used to induce HLC126 to select future target options for LLC122 and is handled in more detail in FIG. 4. That is, LLC122 generates an input to HLC126 that summarizes the agent's interaction with the environment since the time before HLC126 generated the target output.
[0160] In some examples, this summarized output may include the reason for the end of the target representation 124. This is handled in more detail in FIG. 6.
[0161] When HLC126 outputs a basic action, HLC126 controls the agent 104 in the environment 106 to perform this action 108.
[0162] In this particular case, the criteria of Figure 2A can be met immediately after action 108 is taken. As addressed in Figure 3, HLC126 can process the input of observation result 110 to generate a new option output.
[0163] The replay buffer 510 shown here is used by the training system 190 to execute an online-offline training protocol for HLC126 and LLC122, and is addressed in more detail in Figure 5. As part of this training, a target model 406 ("target predictor") can be used to generate auxiliary predictions 407 that can improve the training.
[0164] Figure 3 shows an exemplary HLC126 and its components in more detail.
[0165] At any given time step in which HLC126 is used, HLC126 receives a high-level policy input 310.
[0166] As an example, the input 310 can be the observation result 110 coming directly from the environment. For example, the input 310 can be the observation result received at a given time step.
[0167] As another example, the input 310 can be the high-level observation result provided by LLC122.
[0168] In a particular example, LLC122 can provide HLC126 with a series of observation results 110 collected while attempting to achieve a target representation 124 over a number of time steps.
[0169] In another example, the LLC can reduce or summarize a series of observation results into high-level observation results.
[0170] In a particular example, the high-level observation result can be created using the average for each element over this series of observation results.
[0171] In other examples, this high-level observation result can be created by encoding a series of observations by performing a series of observations through a recurrent neural network (RNN) or another suitable type of machine learning model.
[0172] Accordingly, the high-level input summarizes the agent's interaction with the environment since a previous time when the criteria described with respect to Figure 2B were met.
[0173] As another example, the high-level input may include additional context information.
[0174] For example, the high-level input may include data specifying the number of time steps since the most recent target option was generated, i.e., the current duration of the executed option, data characterizing the reward received from the environment during this time, and data specifying which criteria for generating a new option were met at this time step.
[0175] In some cases, this criterion for generating a new option may be to compare the value estimate from the LLC 122 with respect to the target representation 124 to a threshold.
[0176] Specifically, if the unreachability threshold is exceeded, the target may be terminated early, in which case a new target is needed.
[0177] Similarly, if the achievability threshold is exceeded, the target is considered achieved and a new target 124 is needed.
[0178] Other possible reasons for early termination are described below with respect to Figure 6.
[0179] In the example of Figure 3, the high-level policy input 310 is first processed using a neural network that processes the input to the observation representation.
[0180] In some cases, this neural network is an encoder. An encoder is a model that generates an encoded representation of the input, which means that the encoder transforms the input information into a latent space of outputs of a defined fixed shape but different dimensions.
[0181] Specifically, the observation encoder neural network 301 can be used to encode the input 310. In this case, the observation encoder 301 takes in the input 310 and outputs an observation latent space representation.
[0182] This encoded output is then passed to a recurrent neural network (RNN) core 302.
[0183] The RNN core 302 can be a long short-term memory (LSTM) network, a stacked LSTM, a gated recurrent unit (GRU), or any other variant of a recurrent neural network.
[0184] In other examples, the RNN 302 can be replaced by a Transformer model.
[0185] The output of the RNN 302 is then passed to a high-level policy network 303.
[0186] The high-level policy network 303 processes the output of the RNN 302 to generate a high-level policy output.
[0187] In some examples where there are more than one option, the high-level policy network 303 can include one or more option encoders 303A trained to embed potential options and a high-level option decision network 303B trained to provide an indication of which option to use to generate the high-level policy output 304.
[0188] Specifically, there can be one option encoder for each type of option available for selection by the HLC126.
[0189] In some examples, the available options for the HLC output 304 may be limited by regularizing or constraining the potential output of the option encoder 303A.
[0190] The option encoder 303A functions similarly to the observation encoder 301. The option encoder 303A processes the observation representation transferred by the RNN 302 to generate an option latent space representation regarding the type of option it encodes.
[0191] For example, if there is one option encoder for each type of option, there may be a target option encoder and a basic action encoder.
[0192] The various components of the HLC 126 can be learned together using reinforcement learning to maximize the expected reward 130 that measures the performance of the neural network when controlling the agent 104 for a given task. In this case, gradients are passed between the sub-network components during the learning process.
[0193] As a particular example, the HLC 122 may be trained according to the online Muesli protocol, which incorporates a measure from the learning of the high-level policy network 303A as an additional loss in the reinforcement learning process.
[0194] More specifically, the Muesli protocol provides flexibility in the output of both continuous options (multiple steps) and discrete options (one step) that can be used to control the low-level controller 122.
[0195] This high-level policy output 304 may include any number of different types of options as well as indicia for determining among them.
[0196] In some examples, the policy output 304 may include only the target output option sub-component.
[0197] In other examples, the policy output 304 may include the target output option sub-component, the basic action sub-component option, and a binary indication of which option should be selected at that time step provided by the high-level option decision network 303B.
[0198] In this particular case, the binary indication controls the LLC neural network 122 to determine whether the HLC 126 directly executes a basic action in the environment 106 or selects an action 108 that is conditioned on the target output returned in the output 304 that is processed into the target representation 124.
[0199] In further examples, the policy output 304 may additionally include high-level action options. In this case, the indication of which option should be selected is a multi-class indication of which option should be selected at that time step provided by the high-level option decision network 303B.
[0200] In particular, the high-level action may include an action taken by the HLC 126 to further control the selection by the LLC 122 of an action.
[0201] As an example, the high-level action may include a parameter that adjusts the LLC 122 policy output z.
[0202] In particular, the high-level action is
[0203]
Number
[0204] It is possible to specify a change to the temperature parameter τ used in a temperature-dependent softmax formula as defined, where z(i) is the score for action i in the low-level policy output, and the sum is over actions j in the set of actions that can be executed by the agent. Thus, adding the temperature parameter τ to the exponent argument in the standard softmax formula produces different probability distributions over the set of potential actions.
[0205] Thus, in this example, when there are multiple possible options, the high-level controller 130, or more generally the action selection system 100, can analyze the high-level policy output and determine which option to select at any given time step based on the indication.
[0206] Although not shown in this example, the HLC 126 can alternatively be implemented as a large language model (LLM) neural network, such as a text-based LLM neural network, or a visual language model (VLM) neural network that receives both images and text as inputs. The LLM may be either pre-trained or fine-tuned, and can take in the high-level policy input 310 and output the high-level policy output 304 as natural language instructions. That is, in this example, the high-level policy output can specify the goal as natural language text, or can specify the basic actions as one or more text tokens when the basic actions are to be executed.
[0207] In particular, the HLC 126-LLM can process an exemplary text high-level policy input 310 to generate the output 304. As another example, the HLC 126-VLM can process an exemplary image high-level policy input 310 or an exemplary high-level policy input that includes both an image and text to generate the output 304. As a further example, a multimodal Transformer-based HLC 126 can be implemented to process and generate other data modalities of the input 310 or output 304.
[0208] Figure 4 shows the exemplary low-level controller 122 and its components in more detail. This figure separates the policy execution 122A from the auxiliary task 122B into a separate branch to provide a complete overview of the functionality of the LLC 122 both during (i) policy execution 122a for choosing action 108 and (ii) training involving both policy execution 122A and auxiliary task 122B.
[0209] During training, both the policy execution branch 122A and the auxiliary task branch 122B are executed together, passing information as shown by the dotted lines in Figure 4 to update the parameters of the policy and value network 409, which includes a policy head and a value head that use the auxiliary output 407.
[0210] The policy head includes one or more policy neural networks trained to learn the probabilities of the actions to be taken for the observations 110 at that time step.
[0211] The value head includes one or more value neural networks trained to learn the value of the observations 110 at that time step. The value of the observations 110 at that time step represents the value of the environment in the state characterized by the observation reaching the target state characterized by the target representation. That is, the value is a score representing the value of the environment in the state characterized by the observation reaching the target state characterized by the target representation.
[0212] After training, as shown by the solid line in Figure 4, only the policy execution branch 122A functions to produce the low-level policy output without the auxiliary task branch 122B as parameter updates are not required.
[0213] As described above, in some implementations, the criteria for generating a new target representation may include one or more criteria that depend on target achievability, target unachievability, or both. In this case, target achievability or unachievability may be determined using a threshold corresponding to the output of the value.
[0214] LLC122 receives and encodes a low-level policy input 410 that may include options such as target option 424 or high-level actions, and an observation-action pair that includes data from the current time step O t and previous action A t-1 from.
[0215] This input 410 may either directly originate from the environment 106 during online training or after training, or may be derived from sampling from observations 110 recorded in the replay buffer during offline training.
[0216] In the case of sampling from the replay buffer during offline training, LLC122 can select and encode its own target representation 124 using hindsight goal selection. This process is dealt with in more detail in FIG. 6.
[0217] In the case of target option 424 as input 410 during training or an option proposed by the HLC after training, the target encoder 408 within the auxiliary task branch 122B encodes the received target 124 into an embedded target representation G t 124 into the latent target space.
[0218] By embedding the target into the latent space, it becomes possible for the target encoder 207 to be agnostic to the target type. This setup allows for the use of any target modality considering that the encoder can be trained from the target to the shared embedding space of the target.
[0219] Target representation G t124 can be represented by text, images, or any other type of data that can be embedded into a shared target space.
[0220] In some examples, the target encoder can be an autoencoder-like network that outputs the parameters of a multivariate normal distribution with diagonal covariance as the shared embedding space of the target. The system can then sample the target representation 124 from this multivariate normal part.
[0221] Target representation G t 124 can be used as an input to the policy and value network 409.
[0222] In this case, for an observation that includes data from the current time step and previous actions as input 410, the observation representation B used in the training of the policy and value network 409 t is passed as a pair to the representation network 405 that generates it.
[0223] In some examples, the representation network 405 can include an observation encoder 405A and an RNN 405B that generate the observation representation B t This observation encoder 405A can have an architecture similar to or different from the observation encoder 301 described above the illustration of HLC126 in FIG. 3, but has the same function.
[0224] Similarly, the RNN 405B can have an architecture similar to or different from the RNN 302 described above the illustration of HLC126 in FIG. 3, but has the same function.
[0225] The policy execution branch 122A includes a policy and value network 409 that is trained to define high-value actions 108 that the agent 104 should take in the environment 106.
[0226] The policy execution branch 122A includes a policy and value network 409 that is trained to define high-value actions 108 that the agent 104 should take in the environment 106.
[0227] Policy and value network 409 includes a low-level target policy μ LLC , a low-level behavior policy π LLC , and a value estimator V LLC .
[0228] The low-level target policy μ LLC is a policy trained to learn a target conditional probability distribution over multiple actions using the reward 130 received for the action 108. This policy serves as an expert policy that can improve the effectiveness of the training of the behavior policy π LLC and is trained using behavior imitation from offline data so as to play the role of an expert policy. In other words, the system trains the low-level target policy through behavior imitation for trajectories from an offline dataset, such as the same dataset used to train the behavior policy or a different offline dataset.
[0229] The behavior policy π LLC is a policy configured to process a given observation 110 and target representation 124 to generate a policy output that can select an action for the agent.
[0230] In some examples, the behavior policy output can be affected by high-level options that change how the action 108 is selected from the behavior policy output.
[0231] Specifically, during the training of π LLC , the system uses the behavior-imitated trained target policy π LLC to regularize the training of π LLC .
[0232] Specifically, a proximity regularization term can be incorporated into the loss function of the behavior policy to train π LLC .
[0233] In some examples, this proximity regularization term is between π LLC and μ LLCcan be based on the deviation from.
[0234] After training, this can produce a behavior policy π that can select action 108 according to a target-conditioned probability distribution over multiple actions. LLC to produce.
[0235] Value estimator V LLC is based on learning the expected cumulative reward for the observation results 110 over consecutive time steps, and generates the value V for the observation results 110. The reward at a given time step represents whether the current goal was achieved at that time step. For example, the system can train the value estimator using the regression loss against the ground truth value for that time step. LLC
[0236] Behavior policy π LLC can be trained using any suitable actor-critic method.
[0237] In particular, the behavior policy π LLC can be trained offline using the IMPALA (Importance Weighted Actor-Learner Architecture) protocol (Espeholt et al., arXiv:1802.01561), which can be scaled using distributed computing to consist of one or more decoupled learners.
[0238] In addition, IMPALA incorporates the use of V-trace terms to correct for policy lag mismatches in an offline configuration with multiple learners when training π. LLC
[0239] Specifically, the behavior policy π LLC can be optimized by following the V-Trace policy gradient.
[0240] The V-Trace term is a benefit term (advantage estimate) based on the reward 130 for the observation-action pair compared to the average of the actions that could have been taken under that observation 110, and (ii) (a) the observation representation b of the observation 110 t is processed by the low-level policy head π LLC to obtain the probability a assigned to the action 108 in the pair t and (b) the observation representation b of the target for the trajectory t and the embedded target representation g124 are processed by the behavior imitation policy head μ LLC to obtain the probability assigned to the action 108 in the pair by the behavior imitation policy output generated by μ
[0241]
Number
[0242] is the product of the ratio to. This term is
[0243]
Number
[0244] can be expressed as, and Adv t is the advantage estimate at time t, calculated using the V-Trace return and value estimate for time step t. This is explained in more detail in Espeholt et al., arXiv:1802.01561.
[0245] In particular, the total loss may include the above V-Trace term and a regularization term weighted by a hyperparameter α that penalizes the mismatch between π LLC and the target policy μ LLC and.
[0246]
Number
[0247] In the modified V-Trace algorithm, gradients can be taken only with respect to the parameters of π LLC That is, the system can train the components of the low-level controller offline using the above losses.
[0248] That is, the system can train the components of the low-level controller offline using the above losses.
[0249] The auxiliary task branch 122B may also include a target predictor 406 to enhance the performance of the LLC 122 with respect to the target representation 124.
[0250] In the auxiliary task branch 122B, the output C of the observation encoder 405A t may be transferred to the target evaluator 406, which uses the intermediate observation representation C t and the embedded target representation G t 124 to evaluate the target 124 of interest.
[0251] These target predictors 406 can be learned mappings such as neural networks or other functions and metrics that can be used, without limitation, to evaluate the target achievability and target availability with respect to previously trained targets.
[0252] In some examples, the target predictor 406 may include a target achievement evaluator 406A that predicts whether the target 124 can be achieved within one or more predetermined integer steps.
[0253] In other examples, the predictor 406 may include a target similarity score calculator 406B that uses a high-dimensional vector distance criterion to reject targets 124 that are too close to previously trained targets.
[0254] In some examples, the high-dimensional vector distance criterion can be a cosine similarity metric.
[0255] The auxiliary output 407 of the target evaluator 406 can serve as an auxiliary task for the training of the policy execution branch 122A. This means that when training the policy and value network 109, the output 407 can also be used as an additional loss.
[0256] Once training is complete, these auxiliary tasks are no longer necessary. In this case, goal achievement is determined by the value network, i.e., value estimation is used to determine when to generate new goal options.
[0257] Specifically, the criterion for generating a new goal representation is satisfied when the value estimation generated by the value head 409 by processing the observation representation and the goal representation 124 for the preceding time step exceeds an achieved threshold indicating that the goal 124 for the preceding time step has been achieved.
[0258] Similarly, the criterion for generating a new goal representation 124 is satisfied when the value estimation generated by the value head 409 by processing the observation representation and the goal representation 124 for the preceding time step falls below an unreachability threshold indicating that the goal 124 for the preceding time step has not been achieved in a discrete number of time steps.
[0259] FIG. 5 shows an exemplary training system 190, specifically an offline-online training system 500, that can be used to train the hierarchical controller 130 to perform a new task on unorganized data, i.e., data collected throughout the training process using policies with varying capabilities.
[0260] In such a system, the HLC 126 is trained online via an online protocol 501 (represented here by the dashed arrow), and the LLC 122 is trained offline via an offline protocol (stored here in the solid-lined rounded box). Online training means that the HLC (126) is trained using an online dataset. For example, its network gradient is updated after receiving immediate feedback in the form of the next observation 110 and reward 130 from time t+1 immediately after it acts on the environment 106, or in the form of the reduced observation 124 received from the LLC 122 at time t. Offline training means that the LLC (122) is trained using an offline dataset. For example, its network gradient is updated via sampling of past observations recorded in the replay buffer 510.
[0261] Here, the replay buffer 510 can be any data structure capable of storing experience objects. For example, the replay buffer 510 can store a set of past experience trajectory objects 515.
[0262] Each trajectory object 515 can be a series of observation-action pairs collected as a result of the interaction of the agent 104 with the environment 106 while being controlled using different policies, including policies with varying capabilities.
[0263] As an example, the different policies can be a fixed control policy, a learned policy, or can be controlled by an expert, such as a human user.
[0264] This offline-online training system 500 can be trained on unorganized data using target hindsight selection, which is dealt with in more detail in FIG. 6.
[0265] In this particular example, HLC126 can output either the target option 424 or the basic action 511.
[0266] If the output of HLC126 is the basic action 511, the basic action 511 bypasses LLC122 and directly interacts with the environment 106 to advance it, resulting in the basic observation 512.
[0267] Specifically, the basic observation 512 may include the next observation 110 and reward 130 from time t + 1 immediately after it acts on the environment 106 at time t.
[0268] As an exemplary interaction, HLC126 can output the basic action 511 only at the beginning and end of an episode.
[0269] As another exemplary interaction, HLC126 may output the target option 424 that bypasses the environment 106 and is sent directly to LLC122. In this case, the target option 424 is processed into the target representation 124 used for training.
[0270] In yet another example, the target hindsight selection enables the training of the hierarchical controller 130 for any data related to multiple targets in the replay buffer 510, particularly data not generated by an expert, which is dealt with in more detail in FIG. 6.
[0271] At the end of training, LLC122 can provide the reduced observation 514 to HLC126. In this case, LLC122 is the cause of what HLC126 observes from the environment.
[0272] The reduced observation 514 constitutes a component of the information that HLC126 needs to determine which option to execute next, i.e., which option to indicate in the policy output 304.
[0273] This reduced observation result 514 may include any one of the environmental observation result 110, the reason for early termination, the compressed history of the observation result, or any other information sufficient to summarize the interaction of the LLC 122 with the environment 106 during offline training 502.
[0274] As an exemplary interaction, during a single episode, the HLC 126 may receive an initial observation result 110 of the LLC 122 at the beginning of the episode and output a basic action 511. Upon receiving the basic online observation result 512, the HLC 126 may then be able to send a target option 424 to the LLC 122, thereby initiating sampling by the LLC 122 of the past experience trajectory 515.
[0275] In some cases, this sampling can use the process for hindsight choice of goals addressed in FIG. 6.
[0276] In the second-to-last step of the episode, the LLC 122 may provide the HLC 126 with a reduced observation result 514 that includes the history of the observation results, at which point the HLC 126 starts the online training protocol 501.
[0277] As another example, as the episode progresses, the HLC 126 may receive a reason for early termination for the target option 424 selected at the beginning of the episode as part of the reduced observation result 514. Possible reasons for early termination are addressed in FIG. 6.
[0278] In some implementations, the LLC network 122 may also be frozen during the training of the HLC, such as when the high-level controller is trained separately from the LLC. Keeping a neural network "frozen" during training means not passing gradients for updates, i.e., keeping the parameter values fixed while changing the parameter values of another neural network.
[0279] This exemplary offline-online training system 500 may be deployed "at scale", i.e., the system may be distributed across multiple agents and learners as defined by the IMPALA protocol.
[0280] In particular, system 500 may be deployed such that replay buffer 510, from which LLC122 samples in offline protocol 502, is filled by one or more agents interacting in one or more environments.
[0281] Furthermore, system 500 may be used to train hierarchical controller 130 so that agent 104 can be controlled to generalize to new industrial tasks not encountered during the training of LLC122, HLC126, or both. This facilitates the training of autonomous agent 105, which can learn how to initially generate and use its goal representation 124 using its unorganized experiences.
[0282] FIG. 6 is a block diagram of an exemplary process 600 for offline hindsight goal selection using LLC122 under the online-offline training protocol 500 shown in FIG. 5. For convenience, process 600 is described as being executed by a system of one or more computers located at one or more locations. For example, a suitably programmed training system, such as training system 190 of FIG. 1, can execute process 600.
[0283] In this case, LLC122 trains offline on a plurality of tasks sampled from replay buffer 510. This multi-task training has the advantage of providing HLC126 with reduced observations corresponding to a plurality of target expressions 124, thereby inducing online learning of HLC126 for LLC122 as to which option to select. In particular, these reduced observations may include reasons for early termination of the target.
[0284] Specifically, LLC122 can sample a past experience trajectory 515 from the observations recorded in replay buffer 510, and generate a plurality of tasks 610 from this sampled past experience trajectory 515. These tasks 610 form components of segment partial trajectories, and each partial trajectory is a subset of a given observation 515 from start time t start to end time t end Typically, the target is fixed for the task, i.e., the target is achieved at t end
[0285] By sampling a plurality of partial trajectories, LLC122 can be trained on a plurality of tasks 610 at once.
[0286] In particular, this pruning of the sampled past experience trajectory 515 can be optimized in some way to assist in the training of LLC122.
[0287] For example, tasks 610 can be improved to increase the frequency of segments that are non-zero reward components, or to produce specific instances of tasks desired by the user.
[0288] In some examples, non-goals, i.e., tasks without a target state of the target to be achieved, may also be included to assist in agent exploration.
[0289] In yet another example, tasks can be sampled randomly from the past experience trajectory 515.
[0290] LLC122 then performs training over a plurality of tasks 610, created by completing a loop over all tasks 610 generated from the sampled past experience 515.
[0291] In some cases, this loop over multiple tasks can be a component of a single episode of training of LLC122.
[0292] In other cases, this loop over multiple tasks can be a component of a certain integer number of steps in the environment, which is one or more.
[0293] Within the loop, each task 610 serves as an input to the target encoder 408 and the value head 409.
[0294] The target encoder 408 outputs a target representation G t 124 that is fixed for its sampled partial trajectory task 610. This output is then evaluated using the target evaluator 406, specifically the goal achievement evaluator 406A and the similarity score calculator 406B.
[0295] If the target 124 is determined by the goal achievement evaluator 406A to be too temporally distant, an early termination reason may optionally be provided to the HLC126 as part of the reduced observation result 514.
[0296] In some implementations, the goal achievement evaluator can be a classifier that directly learns an end function.
[0297] If the target 124 of the task 610 is determined by the distance similarity score calculator 406B to be too close to other targets 124 within the episode, an early termination reason may optionally be provided to the HLC126 as part of the reduced observation result 514.
[0298] In some implementations, the similarity score calculator 406B can be a measure of distance, such as cosine similarity, that is evaluated to exclude training all targets up to a certain size of similarity defined by the user.
[0299] As an example, a target 124 with a similarity greater than 60% according to this measure can be rejected from training.
[0300] For each task 610, the value estimate V from the value network 409 LLC is compared to a pre-set unreachability threshold V using a value-based logic rule that rejects targets 124 whose value is below the threshold. unreachability threshold and can be compared.
[0301] If the value estimate is too low, the HLC 126 can receive an early termination reason indicating that the target 124 was impossible to achieve.
[0302] This threshold can be provided by the HLC 126 or set to a fixed value.
[0303] Other implementations can include early termination reasons not explicitly stated here that simplify the training of the LLC 122.
[0304] If the task 610 is not rejected by any of the aforementioned early termination reasons or any reason not previously mentioned, the training of the LLC 122 is performed.
[0305] The training can be performed through any suitable reinforcement learning method.
[0306] In particular, the training can include the regularized offline V-Trace algorithm 620 described above.
[0307] Figure 7 shows the performance of the offline-online hierarchical system (H2O2) of FIGS. 5 and 6 using the average episode return for different tasks taken from the DeepMind Hard Eight task suite when compared to the criteria of the latest flat agents.
[0308] Figure 7 shows six plots corresponding to a subset of six tasks from the DeepMind Hard Eight task suite, namely the baseball task, drawbridge task, navigation cubes task, push blocks task, wall sensor task, and wall sensor stack task. These tasks are components of tasks in visually complex and partially observable 3D environments, which have been difficult for hierarchical agents until now.
[0309] The results for the throw across task and remember sensor task are excluded because neither agent showed evolution for these tasks during testing.
[0310] For most of the tasks, the performance of the agents trained using the described hierarchical training system is higher than or comparable to other agents. In particular, H2O2 achieves better results in the baseball task, wall sensor task, and also the navigation cubes task.
[0311] Although it does not achieve better results in the drawbridge task or push blocks task, H2O2 is still on par with the criteria of the latest flat agents, representing an attempt to address previous challenges in hierarchical reinforcement learning, such as demonstrating performance comparable to other flat (non-hierarchical) techniques in visually complex and partially observable 3D environments.
[0312] In practice, the reason that hierarchical agents may perform worse than flat criteria for these two tasks could be that the semi-Markov decision process (SMDP) experienced by HLC is substantially changed by the LLC design. If the SMDP is easier to solve than the original Markov decision process (MDP) of the task, the hierarchical agent may be no more than equivalent to the flat agent.
[0313] In addition, Figure 7 is important because the performance of H2O2 does not depend on the dataset organized by the expert agent, but rather stems from learning the target representation behavior offline from any experience generated by the agent.
[0314] In Figure 8, the power of including the KL regularization term in the modified V-Trace algorithm is demonstrated for the offline-online implementation form of H2O2 for an exemplary task. In this exemplary task, removing the penalty could cause the learned policy to deviate significantly from the behavior policy, thus resulting in rather poor performance.
[0315] This specification uses the term "configured" in relation to components of a system and computer program. That one or more computer systems are configured to perform a particular operation or action means that the system has installed in it software, firmware, hardware, or a combination thereof that causes the system to perform that operation or action during operation. That one or more computer programs are configured to perform a particular operation or action means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operation or action.
[0316] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more combinations of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. A computer-readable storage medium may be, or may include, a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations of them. Alternatively or in addition, the program instructions may be encoded in an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus.
[0317] The term “data processing apparatus” refers to data processing hardware and includes, by way of example, all kinds of devices, apparatus, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus may be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus may optionally include code that creates an execution environment for computer programs, e.g., code that constitutes one or more components of a processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations of them.
[0318] A computer program, which may also be called or described as a program, software, software application, app, module, software module, script, or code, may be written in any form of programming language, including a compiled language or an interpreted language, or a declarative language or a procedural language, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not correspond to a file in a file system. The program may be stored in a portion of a file that holds other programs or data, such as one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple cooperating files, such as files that hold one or more modules, subprograms, or portions of code. The computer program may be deployed to be executed on one computer or on multiple computers located in one place or distributed across multiple places and interconnected by a data communication network.
[0319] As used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. In general, an engine is implemented as one or more software modules or components installed on one or more computers located at one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines may be installed and executed on the same one or more computers.
[0320] The processes and logical flows described in this specification can be implemented by one or more programmable computers that execute one or more computer programs to manipulate input data and generate output. The processes and logical flows can also be implemented by, for example, a special-purpose logic circuit such as an FPGA or ASIC, or by a combination of special-purpose logic circuits and one or more programmed computers.
[0321] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. In general, a central processing unit receives instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a central processing unit for performing or executing instructions, and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuits. In general, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to receive data from, or transfer data to, or both of these. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive.
[0322] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0323] To enable interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other kinds of devices may also be used to enable interaction with a user. For example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input received from the user may be in any form, including acoustic input, speech input, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user, such as by sending a web page to a web browser of the user's device in response to a request received from the web browser of the user's device. Also, the computer can interact with the user by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message in reply from the user.
[0324] A data processing apparatus for implementing a machine learning model may also include, for example, a dedicated hardware accelerator unit for processing machine learning training or general and computationally intensive parts of a product, i.e., inference, workloads.
[0325] A machine learning model can be implemented and deployed using a machine learning framework, such as the TensorFlow framework.
[0326] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components, such as a data server, or middleware components, such as an application server, or front-end components, such as a graphical user interface, a web browser, or a client computer having an application with which a user can interact with the implementation of the subject matter described in this specification, or any combination of one or more such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0327] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, a server, for example, sends data, such as an HTML page, to a user device for the purpose of displaying data to a user who interacts with a device that operates as a client and receiving user input from the user. For example, data generated at a user device as a result of user interaction can be received at the server from the device.
[0328] This specification includes details of many specific implementations, which should not be construed as limitations on the scope of the invention or what can be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Some features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, features may be described above as operating in certain combinations and even initially claimed as such, but one or more features from the claimed combination may in some cases be omitted from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.
[0329] Similarly, operations are shown in the drawings and described in the claims in a particular order, but this should not be understood as requiring that such operations be performed in the particular order or sequential order shown in order to achieve the desired result, or that all of the shown operations be performed. In some situations, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated together in a single software product or packaged into multiple software products.
[0330] Certain embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the acts recited in the claims may be performed in a different order and still achieve desirable results. As one example, the processes shown in the accompanying figures do not necessarily require the particular order or sequential order shown to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Explanation of Signs
[0331] 100 Action Selection System 102 Action Selection Subsystem 104 Agent 106 Environment 108 Action 110 Observation Result 122 Low-Level Controller 124 Target Representation 126 High-Level Controller 130 Reward 190 Training System 301 Observation Result Encoder 302 RNN Core 303 Policy Network 405A Observation Result Encoder 405B RNN Core 406 Target Model 407 Auxiliary Prediction 408 Target Encoder 409 Policy and Value Network 410 Low-Level Policy Input 411 LLC Output 424 Target Option 500 Offline-Online Training System 501 Online Protocol 502 Offline Protocol 510 Replay Buffer 511 Basic Action 512 Basic OBS 514 Reduced OBS 515 Past experience trajectory 610 Task 611 Value estimation 620 Regularized offline V-Trace learning
Claims
1. A method for controlling an agent that interacts with an environment to execute a task, wherein in each of a plurality of first time steps from a plurality of time steps, receiving an observation characterizing the state of the environment at the first time step; determining a target representation for the first time step characterizing the target state of the environment to be reached by the agent; using a low-level controller neural network to process the observation and the target representation to generate a low-level policy output that defines an action to be performed by the agent in response to the observation, wherein the low-level controller neural network comprises a representation neural network configured to process the observation to generate an internal state representation of the observation; and a low-level policy head configured to process the observation result state representation and the target representation to generate the low-level policy output and a step; controlling the agent using the low-level policy output; A method comprising:
2. The step of determining the target representation for the first time step characterizing the target state of the environment to be reached by the agent comprises determining whether a criterion for generating a new target representation is satisfied at the first time step; when the criterion is not satisfied, using the target representation from the preceding first time step as the target representation for the first time step; The method according to claim 1, comprising:
3. When the criterion is satisfied, generating a high-level observation result for the first time step; Processing the high-level input with the high-level controller neural network that includes the high-level observation result, and generating a high-level policy output that includes a target output characterizing the target state; Processing the target output with the target encoder neural network to generate the target representation; The method according to claim 2, further comprising. **Claim 4** The high-level policy output further comprises an indication of whether to control the agent (i) using the low-level controller or (ii) using the basic actions specified by the high-level policy output; The step of determining the target representation for the first time step, the step of processing the observation result and the target representation using the low-level controller neural network, and the step of controlling the agent using the low-level policy output are executed only in response to the determination to control the agent using the low-level controller based on the indication. The method according to claim 3. **Claim 5** In each of one or more second time steps from the plurality of time steps, Generating a high-level observation result for the second time step; Processing the high-level observation result for the second time step using the high-level controller neural network to generate a high-level policy output for the second time step; Determining to control the agent using the basic actions specified by the high-level policy output based on the indication of whether to control the agent (i) using the low-level controller or (ii) using the basic actions specified by the high-level policy output; In response thereto, controlling the agent using the basic actions specified by the high-level policy output; The method according to claim 4, comprising
6. The method according to any one of claims 3 to 5, wherein the high-level observation result for the first time step comprises the observation result received at the first time step.
7. The high-level observation result for the first time step The method according to any one of claims 3 to 6, comprising data for identifying the number of time steps after the most recent time step when the criterion was met.
8. The high-level observation result for the first time step The method according to any one of claims 3 to 7, comprising data characterizing the reward received from the environment after the most recent time step when the criterion was met.
9. The high-level observation result for the first time step The method according to any one of claims 3 to 8, comprising data characterizing the observation results received at time steps after the most recent time step when the criterion was met.
10. When any criterion in a set of criteria is met, the criterion is met and the high-level observation result for the time step The method according to any one of claims 3 to 9, comprising data for identifying which criterion was met at the first time step.
11. The method according to any one of claims 3 to 10, wherein the target output characterizing the target state is a target vector characterizing the target state.
12. The method according to any one of claims 3 to 10, wherein the target output characterizing the target state is a text string describing the target state.
13. The step of using the high-level policy output to control the agent, where the high-level policy output comprises hyperparameters defining some aspect of the training, comprises generating an adjusted policy output by applying the hyperparameters to the low-level policy output; selecting an action using the adjusted policy output; and is the method according to any one of claims 3 to 12. **Claim 14** The method according to claim 13, wherein the high-level policy uses a temperature parameter as the hyperparameter to adjust the low-level policy output. **Claim 15** determining whether a criterion for generating a new goal representation is satisfied at the first time step; and when the first time step is the first time step in a task episode, determining that the criterion for generating a new goal representation is satisfied, and is the method according to any one of claims 2 to 14. **Claim 16** determining whether a criterion for generating a new goal representation is satisfied at the first time step; comprising determining that a maximum number of time steps have elapsed since the most recent time step at which the criterion was satisfied, and is the method according to any one of claims 2 to 15. **Claim 17** The method according to any one of claims 2 to 16, further comprising a value head, wherein the low-level controller neural network is configured to process an observation representation and the goal representation to generate a value estimate of the value of the environment in a state characterized by the observation that the goal state characterized by the goal representation has been reached. **Claim 18** determining whether a criterion for generating a new goal representation is satisfied at the first time step; The method according to claim 17, comprising the step of determining that the criterion for generating a new target representation is satisfied when the value estimate generated by the value head is less than an unreachable threshold by processing the observation result representation of the observation result in the preceding time step and the target representation for the preceding time step.
19. The step of determining whether the criterion for generating a new target representation is satisfied in the first time step is The method according to any one of claims 17 or 18, comprising the step of determining that the criterion for generating a new target representation is satisfied when the value estimate generated by the value head and the target representation for the preceding time step, by processing the observation result representation of the observation result in the preceding time step, are greater than an achieved threshold indicating that the target characterized by the target representation for the preceding time step has been achieved.
20. The method according to claim 19 when dependent on claim 3, wherein the achieved threshold is included in the high-level policy output.
21. The step of determining whether the criterion for generating a new target representation is satisfied in the first time step is Using a classifier neural network to process (i) the observations in the preceding time step and (ii) the classifier input characterizing the target state for the preceding time step to generate a classifier output indicating how many time steps remain until the target state is achieved; The step of determining that the criterion for generating a new target representation is satisfied when the classifier output indicates that 0 time steps remain until the target state is achieved; and the method according to any one of claims 2 to 19.
22. A method for training a low-level controller neural network according to any one of claims 1 to 21, comprising: obtaining one or more trajectories of observation result-action pairs and respective goals for each trajectory; for each trajectory, generating respective rewards for each pair in the trajectory based on whether the goal is achieved in the state characterized by the observation result in the pair; training the low-level controller neural network with respect to the one or more trajectories and the one or more rewards to optimize a reinforcement learning objective, the reinforcement learning objective comprising: (i) being trained with respect to a behavior imitation loss; and (ii) for a given observation result and a given goal representation, processing the given observation result representation for the given observation result and the given goal representation generated by the representation neural network to generate a behavior imitation policy output for the observation result, and imposing a penalty on the low-level policy output generated by the low-level controller neural network for deviation from the behavior imitation policy output generated by a behavior imitation head configured to generate the behavior imitation policy output, the behavior imitation head being configured to generate the behavior imitation policy output, the behavior imitation policy output and the low-level policy output each defining a probability distribution over a plurality of actions, and the proximity regularization term being based on, for each observation result in each trajectory, the deviation between (i) the low-level policy output generated by the low-level policy head by processing the observation result representation of the observation result and the goal representation of the goal for the trajectory generated by the representation neural network, and (ii) the behavior imitation policy output generated by the behavior imitation policy head by processing the observation result representation of the observation result and the goal representation of the goal for the trajectory generated by the representation neural network; A method comprising the above steps. Claim 23 The method according to claim 22, wherein the behavior imitation policy output and the low-level policy output each define a probability distribution over a plurality of actions, and the proximity regularization term is based on, for each observation result in each trajectory, the deviation between (i) the low-level policy output generated by the low-level policy head by processing the observation result representation of the observation result and the goal representation of the goal for the trajectory generated by the representation neural network, and (ii) the behavior imitation policy output generated by the behavior imitation policy head by processing the observation result representation of the observation result and the goal representation of the goal for the trajectory generated by the representation neural network. Claim 24 The method according to claim 23, wherein the deviation is a KL divergence.
25. For each observation - action pair in each trajectory, the reinforcement learning objective further comprises a term that is the product of (i) a term based on the reward for the pair and (ii) the ratio of (a) the probability assigned to the action in the pair by the low - level policy output generated by the low - level policy head by processing the observation representation of the observation and the target representation of the target for the trajectory, and (b) the probability assigned to the action in the pair by the imitation policy output generated by the imitation policy head by processing the observation representation of the observation and the target representation of the target for the trajectory, which are generated by the representation neural network. The method according to any one of claims 23 to 24.
26. The method according to claim 25, wherein the term based on the reward for the pair is a V - Trace policy gradient term.
27. The method according to any one of claims 22 to 26, further comprising the step of training the imitation policy head and the representation neural network for one or more of the trajectories to minimize the imitation loss.
28. For one or more of the trajectories, the target is an image of the environment in a target state, and the training step further comprises the step of using an image target encoder neural network to process the image to generate a target representation to be processed by the low - level policy head and the imitation policy head. The method according to any one of claims 22 to 27.
29. For one or more of the trajectories, the target is text describing the target state of the environment, and the training step The method according to any one of claims 22 to 27, further comprising the step of processing the text using a text target encoder neural network to generate a target representation to be processed by the low-level policy head and the behavior imitation policy head.
30. The method according to any one of claims 22 to 29, wherein the one or more trajectories are sampled from offline data generated while the agent was controlled by one or more different behavior policies.
31. The method according to claim 30, wherein the low-level controller is trained only on the offline data.
32. The method according to claim 30, wherein the low-level controller is trained on both the offline data and online data generated by controlling the agent using the low-level policy output generated by the low-level controller.
33. The method according to any one of claims 22 to 32 when dependent on claim 3, further comprising the step of training the high-level controller neural network through reinforcement learning to maximize the expected reward for the task generated as a result of controlling the agent based on the high-level policy output generated by the high-level controller neural network and the low-level policy output generated by the low-level controller neural network.
34. The method according to claim 33, wherein as a result of the training of the high-level controller neural network, no gradient is passed to the low-level controller neural network.
35. The method according to claim 33 or 34, wherein the one or more trajectories include one or more trajectories sampled from online data generated while controlling the agent during the training of the high-level controller.
36. The method according to claim 33 or 34, wherein none of the one or more trajectories are sampled from online data generated while controlling the agent during the training of the high-level controller.
37. The method according to any one of claims 23 to 36, wherein the reinforcement learning objective comprises one or more auxiliary loss terms, each auxiliary loss term corresponding to a respective auxiliary task, and each of the auxiliary tasks requires generating a prediction characterizing a target distribution conditioned on an output generated by the representational neural network.
38. Further comprising, for each of the one or more trajectories of the observation-action pairs, generating a respective target, the generating step comprising, for one or more of the trajectories, selecting a target that describes the state of the environment characterized by the last observation in the trajectory. The method according to any one of claims 23 to 37.
39. Further comprising, for each of the one or more trajectories of the observation-action pairs, generating a respective target, the generating step comprising, for one or more of the trajectories, selecting a target that describes the state of the environment not characterized by any of the observations in the trajectory. The method according to any one of claims 23 to 37.
40. The method according to any one of claims 1 to 39, wherein the agent is a mechanical agent and the environment is a real-world environment.
41. The method according to claim 40, wherein the agent is a robot.
42. The method according to any one of claims 1 to 41, wherein the environment is a real-world environment of a service facility including a plurality of electronic devices, and the agent is an electronic agent configured to control the operation of the service facility.
43. The method according to any one of claims 1 to 42, wherein the environment is a real-world manufacturing environment for manufacturing a product, and the agent comprises an electronic agent configured to control a manufacturing unit or machine that operates to manufacture the product.
44. One or more computers, One or more storage devices communicatively coupled to the one or more computers, the one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of each of the methods according to any one of claims 1 to 43.
45. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of each of the methods according to any one of claims 1 to 43.
Citation Information
Patent Citations
Method and system for controlling an electrified vehicle
DE102020120367A1
Device and method for controlling a robot
US20210341904A1