Reinforcement Learning Using Proxy Courses
By introducing the ‘curriculum’ concept of agents in reinforcement learning systems, using less complex agent strategies to train more complex agent strategies, the problem of training complex agent strategies in the existing technology requires a large amount of computing resources and long-term training, and the rapid and high-performance learning of complex agent strategies is achieved.
Patent Information
- Application Number
- CN201980032894.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-05-18
- Filing Date
- 2019-05-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-05-20
AI Technical Summary
When training complex proxy strategies, existing reinforcement learning systems require a large amount of computing resources and long-term training, resulting in inefficiency.
By introducing the ‘course’ concept of agents, the knowledge of less complex agent strategies is used to train more complex agent strategies, and the joint training and weight adjustment methods are adopted to gradually support more complex agent strategy neural networks.
The rapid and high-performance learning of complex proxy strategies in reinforcement learning tasks is realized, reducing the total training time and improving training efficiency.
Smart Images

Figure CN112154458B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to reinforcement learning. Background Art
[0002] In a reinforcement learning system, an agent interacts with an environment by performing actions selected by the reinforcement learning system in response to receiving observation data that characterizes the current state of the environment.
[0003] Some reinforcement learning systems select actions to be performed by an agent based on the output of a neural network in response to receiving given observation data.
[0004] A neural network is a machine learning model that uses one or more layers of non-linear units to predict an output for received inputs. Some neural networks are deep neural networks, which include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as the input to the next layer in the network (i.e., the next hidden layer or the output layer). Each layer of the network generates an output based on the current values of a corresponding set of parameters and the received inputs. Summary of the Invention
[0005] This specification describes a system implemented as a computer program on one or more computers in one or more locations that trains a final action policy neural network for selecting actions to be performed by a reinforcement learning agent that interacts with an environment.
[0006] The system trains the final action policy neural network (i.e., the neural network that will be used to control the reinforcement learning agent after training) as part of a set of candidate agent policy neural networks. The final action policy neural network generally defines the most complex policy among any of the networks in the set, i.e., the action selection policy defined by at least one other action policy neural network in the set is less complex than the policy defined by the final action policy neural network.
[0007] At the start of training, the system initializes mixed data that assigns corresponding weights to each of the candidate agent policy neural networks in the set.
[0008] Then, the system jointly trains these candidate agent policy neural networks to perform a reinforcement learning task. In particular, during training, the system uses a combined action selection policy that is a combination (according to the weights in the mixed data) of the individual action selection policies generated by the candidate networks in the set.
[0009] During training, the system repeatedly adjusts the weights in the mixed data to favor higher-performance candidate agent policy neural networks, e.g., by assigning them greater weights.
[0010] Because the different networks in the ensemble define action selection strategies with different levels of complexity, and because the weights between different networks are adjusted throughout training, the ensemble of surrogate policy neural networks is also referred to as the agent's "curriculum".
[0011] The combined action selection strategy can be used to select the actions to be performed by the agent. However, reinforcement learning can be performed on-policy or off-policy. That is, the training of the candidate surrogate policy neural network can be performed online, or based on previously generated training data (generated using an older version of the candidate surrogate policy neural network parameters) stored in a replay memory.
[0012] As described in more detail later, "more complex" in this document generally refers to the complexity of training. Thus, a more complex action selection strategy can refer to one that takes longer to train (i.e., requires more training steps) to achieve the same performance, such as the average reward over multiple episodes, or is less robust for different hyperparameter settings (e.g., learning rate, objective function weights, mini-batch size, etc.), than another action selection strategy (e.g., the action selection strategy of another candidate surrogate policy neural network). In some implementations, a more complex action selection strategy can correspond to a more complex architecture, such as a deeper or larger (with more units and / or parameters) surrogate policy neural network, or one with more different types of layers, such as including recurrent layers. In some implementations, a more complex action selection strategy can correspond to a strategy that operates in a larger action space (i.e., has more actions to choose from while learning to perform the same task).
[0013] In some implementations, the candidate surrogate policy neural network is trained to generate an action selection strategy that aligns with other action selection strategies generated by other candidate surrogate policy neural networks by processing the same training network input. For example, the reinforcement learning loss can include the cost of a consistent policy, such as a cost that depends on the difference between the policies (e.g., a measure of the difference between the policy output distributions depending on the type of reinforcement learning).
[0014] As the weights of the final surrogate policy neural network increase, the system can reduce the influence of training candidate surrogate policy neural networks to generate a consistent action selection policy. That is, the system can gradually switch from using multiple candidate surrogate policy neural networks to using the final surrogate policy neural network, and in the limit, can rely solely on the final surrogate policy neural network to select actions. This can be achieved by adjusting the weights assigned to the hybrid update as training progresses.
[0015] In an implementation, generating a combined action selection policy can include using each of the candidate surrogate policy neural networks to process the training network input to generate a corresponding action selection policy (output) for each candidate surrogate policy neural network, and combining the action selection policies according to the weights at that training iteration to generate a combined action selection policy.
[0016] In principle, the weights can be adjusted manually or by using an appropriate annealing strategy. However, in some implementations, a population of combinations of candidate surrogate policy neural networks is trained. Then, during training, the weights can be adjusted by adjusting the weights used by lower-performance combinations based on the weights used by higher-performance combinations. For example, as described below, population-based training techniques can be used such that poorly performing combinations (as measured by a performance metric of the combined action selection policy) replicate the neural network parameters of stronger combinations and perform local modifications to their hyperparameters such that the poorly performing combinations are used to explore the hyperparameter space. Any convenient performance metric that depends on the quality of the combined policy outputs generated during training can be used, such as the reward over k episodes.
[0017] Specific embodiments of the subject matter described in this specification can be implemented so as to achieve one or more of the following advantages.
[0018] By using a curriculum for an agent during training as described in this specification (i.e., by adjusting weights as described in this specification), a complex agent can learn (i.e., can train a complex agent policy selection neural network) to perform a reinforcement learning task using fewer computational resources and less training time than traditional methods. In particular, by leveraging the knowledge of less complex agents in the curriculum, a more complex agent can quickly achieve high performance in a reinforcement learning task, i.e., much faster than if the complex agent were trained in an independent manner on a particular task. In fact, in some cases, by using an agent curriculum, a complex agent can quickly achieve high performance on a task even if the agent cannot learn the task from scratch when trained in an independent manner. In other words, a more complex agent can bootstrap from the solutions found by simpler agents to learn a task that it otherwise would not be able to learn, or to learn the task with fewer training iterations than would otherwise be required. Additionally, by training as described in the distribution and adjusting the weights as described in this specification, the total training time can be reduced relative to training only a single final agent, even when multiple agents are jointly trained.
[0019] Details of one or more embodiments of the subject matter described in the specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 An example reinforcement learning system is shown.
[0021] Figures 2A - 2C is a diagram showing examples of various candidate agent policy neural networks.
[0022] Figure 3 is a flowchart of an example process for training a set of candidate agent policy neural networks.
[0023] Figure 4 is a flowchart of an example process for performing a training iteration.
[0024] Like reference numerals and names in the different figures represent like elements. DETAILED DESCRIPTION
[0025] Figure 1 An example reinforcement learning system 100 is shown. The reinforcement learning system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.
[0026] The reinforcement learning system 100 trains an agent policy neural network through reinforcement learning to control the agent 102 to perform a reinforcement learning task when interacting with the environment 104.
[0027] Specifically, at each time step during training, the reinforcement learning system 100 receives data representing the current state of the environment 104. The data representing the state of the environment will be referred to as observation data 106 in this specification. In response to the observation data, the system 100 selects an action to be performed by the agent 102 and causes the agent 102 to perform the selected action. Once the agent 102 has performed the selected action, the environment 104 transitions to a new state, and the system 100 receives a reward 110.
[0028] Generally, the reward 110 is a numerical value. The reward 110 can indicate whether the agent 102 has completed the task or the agent 102's progress towards completing the task. For example, if the task specifies that the agent 102 should navigate through the environment to a target location, the reward at each time step can have a positive value once the agent reaches that target location, and zero otherwise. As another example, if the task specifies that the agent should explore the environment, the reward at that time step can have a positive value if the agent navigates to a previously unexplored location at that time step, and zero otherwise.
[0029] In some implementations, the environment is a real-world environment, and the agent is a mechanical agent that interacts with the real-world environment, such as a robot navigating in the environment or an automated or semi-automated land, air, or sea vehicle.
[0030] In these implementations, the observation data can include, for example, one or more of images, object position data, and sensor data, to capture the observation data as the agent interacts with the environment, such as sensor data from an image, distance, or position sensor or from an actuator.
[0031] For example, in the case of a robot, the observation data can include data representing the current state of the robot, such as one or more of the following: joint position, joint velocity, joint force, torque, or acceleration, such as gravity-compensated torque feedback, and the overall or relative pose of the robot holding an item.
[0032] In the case of a robot or other mechanical agent or vehicle, the observation data can similarly include one or more of position, linear velocity or angular velocity, force, torque, or acceleration, and the overall or relative pose of one or more parts of the agent. The observation data can be defined as one-dimensional, two-dimensional, or three-dimensional, and can be absolute and / or relative observation data.
[0033] For example, the observed data may also include sensed electrical signals, such as motor current or temperature signals; and / or for example, image or video data from a camera or LIDAR sensor, such as data from a sensor of an agent or data from a sensor placed separately from an agent in the environment.
[0034] In these implementations, the action can be a control input for controlling a robot, such as the torque of a robot joint or a higher-level control command; or a control input for controlling an autonomous or semi-autonomous land, air, or sea vehicle, such as the torque or a higher-level control command for a control surface or other control element of the vehicle.
[0035] In other words, for example, the action may include position, velocity, or force / torque / acceleration data of one or more joints of a robot or parts of another mechanical agent. Additionally or alternatively, the action data may include electronic control data, such as motor control data, or more generally, data for controlling one or more electronic devices in the environment, and the control of the electronic devices has an impact on the observed environmental state. For example, in the case of an autonomous or semi-autonomous land, air, or sea vehicle, the action may include actions for controlling navigation (e.g., steering) and movement (e.g., braking and / or accelerating of the vehicle).
[0036] In some implementations, the environment is a simulated environment, and the agent is implemented as one or more computers that interact with the simulated environment.
[0037] The simulated environment can be a motion simulation environment, such as a driving simulation or a flight simulation, and the agent can be a simulated vehicle that navigates in the motion simulation. In these implementations, the action can be a control input for controlling the simulated user or the simulated vehicle.
[0038] In another example, the simulated environment can be a video game, and the agent can be a simulated user who plays the video game. Generally, in the case of a simulated environment, the observed data may include a simulated version of one or more of the previously described observed data or observed data types, and the action may include a simulated version of one or more of the previously described actions or action types.
[0039] In the case of an electronic agent, the observed data may include data from one or more sensors monitoring parts of a factory or a service facility, such as current, voltage, power, temperature, and other sensors and / or electronic signals representing the functions of electronic and / or mechanical equipment components.
[0040] In some other applications, the agent can control actions in a real environment including equipment components, e.g., in a data center, in an electric / water rationing system, or in a manufacturing plant or service facility. Observation data can be related to the operation of the factory or facility. For example, the observation data can include observation data on the electrical or water usage of equipment, or observation data on power generation or distribution control, or observation data on resource usage or waste generation. The actions can include actions to control or impose operating conditions on the equipment components of the factory / facility, and / or actions that result in a change in settings in the operation of the factory / facility, such as adjusting or turning on / off components of the factory / facility.
[0041] System 100 trains a final action policy neural network (i.e., the neural network that will be used to control the reinforcement learning agent after training) as part of a set of candidate agent policy neural networks. In Figure 1 the example, the neural networks in the set are denoted as π1 to π K , where π K represents the final action policy neural network.
[0042] Generally, each action policy neural network in the set receives a network input including observation data and generates a network output defining an action selection policy for selecting an action to be performed by the agent in response to the observation data.
[0043] In some implementations, the network output defines a likelihood distribution of actions in a set of possible actions. For example, the network output can include corresponding numerical likelihood values for each action in the set of possible actions. As another example, the network output can include corresponding numerical values of parameters defining a parametric probability distribution (e.g., the mean and standard deviation of a normal distribution). In this example, the set of possible actions can be a continuous set (e.g., a continuous range of real numbers). In some of these implementations, System 100 selects an action to be performed by the agent by sampling an action from the set of possible actions based on the likelihood distribution.
[0044] In some implementations, the network output identifies an action from the set of possible actions. For example, if the agent is a robotic agent, the network output can identify the torques to be applied to the joints of the agent. In some of these implementations, System 100 selects the action identified by the network output as the action to be performed by the agent, or adds noise to the identified action and selects the noisy action as the action to be performed.
[0045] In some implementations, the network input includes both observed data and a given action from a set of possible actions, and the network output is an estimate of the return that the system will receive if the agent executes the given action in response to the observed data. The return refers to the cumulative measure of the rewards received by the system when the agent interacts with the environment over multiple time steps. For example, the return can refer to the long-term time-discounted rewards received by the system. In some of these implementations, system 100 can select the action with the highest return as the action to be executed, or can apply an epsilon-greedy action selection strategy.
[0046] Although the policy neural networks all receive the same type of network input and generate the same type of network output, the final policy neural network is generally the most complex neural network in the set. In other words, the final agent policy neural network defines an action selection strategy for the agent that is more complex than the action selection strategy defined by at least one other candidate agent policy neural network.
[0047] As used in this specification, the complexity of an action selection strategy refers to the training complexity, i.e., how difficult it is to train a neural network from scratch so that the agent uses the action selection strategy generated by the neural network to perform a reinforcement learning task. For various reasons, for a given reinforcement learning task, one neural network may be more complex than another.
[0048] For example, one neural network can generate an output that defines a larger action space for the agent than other networks. In particular, other candidate networks in the set can be constrained to generate a policy that assigns non-zero likelihood of being selected only to a finite number of possible actions that can be executed by the agent, while the output of the final agent policy neural network is not so constrained.
[0049] As another example, one neural network can have a more complex neural network architecture than another neural network. For example, the final agent policy neural network may have many more parameters than other networks. As another example, the final agent policy neural network can include certain types of layers that are not included in other networks. As a specific example, the final agent policy neural network can include layers that are generally difficult to train to convergence (such as recurrent neural network layers) and are not present in other candidate neural networks.
[0050] As another example, the reinforcement learning task can be a combination of multiple different individual tasks, and one neural network can be a multi-task neural network that generates corresponding outputs for each of the different individual tasks, while other neural networks generate outputs for only one individual task.
[0051] Figures 2A - 2CA diagram showing examples of various candidate agent policy neural networks with different levels of complexity.
[0052] In Figure 2A the example, the system is using actor-critic reinforcement learning techniques to train the candidate neural network. Thus, the combined output includes the combined policy output π mm and the combined value output v mm both.
[0053] The combined value output assigns a value to the current state of the environment characterized by the received observation data "obs". In particular, this value is an estimate of the expected return the system will receive if actions are selected starting from the environment of the current state according to the current policy.
[0054] The combined policy output defines the actions the agent takes in response to the observation data. For example, the combined policy output can be a probability distribution over a set of possible actions to be taken by the agent, and the system can select an action by sampling from the probability distribution.
[0055] Specifically, Figure 2A shows two architectures 210 and 220, where architecture 220 is more complex than architecture 210, i.e., architecture 220 is more difficult to train from scratch on the reinforcement learning task. Architectures 210 and 220 can be the architectures of two agent policy neural networks included in a set of candidate agent policy neural networks. Although in Figure 2A the example, these are the only two neural networks in the set, in reality, the set can also include Figure 2A other candidate agent policy neural networks not shown in
[0056] In Figure 2A the example, both architectures 210 and 220 receive the observation data (obs) and process this observation data to generate the corresponding policy outputs π1 and π2. Both architectures include a convolutional encoder neural network followed by one or more long short-term memory (LSTM) layers. In fact, in some implementations, these parts of architectures 210 and 220 are shared, i.e., the parameter values are constrained to be the same between the two architectures.
[0057] However, architecture 210 includes a linear layer followed by a masking operation that sets the probabilities assigned to a subset of the possible actions in the set to zero. Thus, the policy output generated by architecture 210 can assign non-zero likelihoods only to a finite number of the possible actions that can be executed by the agent. On the other hand, architecture 220 includes a linear layer without a subsequent masking operation, so the policy output generated by architecture 220 can assign non-zero likelihoods to any of the possible actions that can be executed by the agent. Thus, the policy output generated by architecture 220 defines a larger action space for the agent. Although Figure 2A only the linear layer of architecture 220 is shown generating a value output, in reality, the linear layer of architecture 210 can also generate a value output, which is mixed (combined) with the value output of architecture 220 to generate a combined value output.
[0058] Figure 2B Two architectures 230 and 240 are shown. Architecture 230 includes a convolutional neural network encoder followed by one or more linear layers and then a final linear layer that generates a policy output and a value output. However, architecture 240 includes the same convolutional encoder but is followed by one or more LSTM layers and then a final linear layer that generates a policy output and a value output. Thus, architecture 240 is recurrent while architecture 230 is not. This increases the training complexity of architecture 240 relative to architecture 230, i.e., because recurrent layers are more difficult to train than feedforward linear layers.
[0059] Figure 2C Three architectures 250, 260, and 270 are shown. In Figure 2C the example, the reinforcement learning task includes two individual tasks i and j. Architecture 250 generates an output only for task i and architecture 270 generates an output only for task j. On the other hand, architecture 260 generates an output for both task i and task j. Thus, although architectures 250, 260, and 270 are similar in terms of the number of parameters and include the same types of neural network layers, architecture 270 is more complex to train because it has to be trained on both task i and task j while the other architectures are trained on only a single task.
[0060] Returning to the pair of Figure 1In the description, at the beginning of training, system 100 initializes the mixed data, which assigns corresponding weights to each of the candidate proxy policy neural networks in the set. Generally, the weights initially assigned to the least complex neural network in the set are much higher than those assigned to the most complex neural network in the set. As a specific example, the system can initially assign a weight of 1 (or a value close to 1) to the least complex neural network in the set, while assigning a weight of 0 (or a value close to 0) to each other neural network in the set.
[0061] System 100 then jointly trains the candidate proxy policy neural networks to perform a reinforcement learning task. In particular, during training, the system uses the combined action selection policy π mm to select the actions to be performed by the agent 102, and the combined action selection policy π mm is a combination of the individual action selection policies (according to the weights in the mixed data) generated by the candidate networks in the set.
[0062] Specifically, this specification will describe that the system combines the action selection policies by calculating the weighted sum of the individual action selection policies generated by the policy neural networks (i.e., weighted according to the weights in the mixed data). In an alternative implementation, the system can sample the policy neural networks according to the weights in the mixed data, and then use the output generated by the sampled policy networks as the combined action selection policy.
[0063] During training, system 100 repeatedly adjusts the parameter values of the proxy policy neural networks using reinforcement learning.
[0064] Specifically, system 100 adjusts the parameter values of the proxy policy neural networks through reinforcement learning so that the combined action selection policy generated as a result of combining ("mixing") the individual action selection policies generated by the policy networks exhibits improved performance on the reinforcement learning task.
[0065] In addition, during training, system 100 also trains the candidate proxy policy neural networks to generate action selection policies that are consistent with the other action selection policies generated by other candidate proxy policy neural networks by processing the same training network input. This is called "matching".
[0066] In addition, system 100 repeatedly adjusts the weights in the mixed data to gradually support more complex proxy policy neural networks, including the final proxy policy neural network.
[0067] Because the weights initially favor the least complex networks, and the least complex networks can quickly improve their performance on reinforcement learning tasks, more complex agent policy neural networks can initially (through matching updates during training) bootstrap from solutions found by simpler networks to help the more complex networks learn the task. However, although less complex networks can readily and quickly identify some solutions to the task, the solutions are generally limited due to the limited capacity of the less complex networks, e.g., due to the limited action space, limited architecture capacity, etc., of the less complex networks.
[0068] As training progresses, by increasing the weights assigned to the more complex networks, the more complex networks find better solutions because the combined policy output becomes less dependent on the simple solutions found by the simple networks.
[0069] After training, other candidate networks in the ensemble can be discarded, and the final policy neural network can be used to control the agent. Alternatively, the system can provide the final trained values of the parameters of the final policy neural network to another system for use in controlling the agent.
[0070] Figure 3 is a flowchart of an example process 300 for training candidate policy neural networks. For convenience, process 300 will be described as being performed by a system of one or more computers located at one or more locations. For example, a suitably programmed reinforcement learning system (such as Figure 1 reinforcement learning system 100) can perform process 300.
[0071] The system initializes the mixed data (step 302). In particular, as described above, the system initializes the mixed data to assign higher weights to the less complex policy networks than to the more complex policy networks.
[0072] The system trains the action policy neural networks in the ensemble according to the mixed data (step 304). In particular, the system performs one or more training iterations to update the parameter values of the policy networks in the ensemble. During training, the system updates the parameter values of the policy networks to (1) generate a combined action selection policy that results in improved performance on the reinforcement learning task, and (2) generate an action selection policy that is consistent with other action selection policies generated by other candidate agent policy neural networks by processing the same training network input. The execution of the iterations for training the policy neural networks will be described in more detail below with reference to Figure 4 more detail.
[0073] The system adjusts the weights in the mixed data (step 306).
[0074] In some implementations, the system uses a predetermined annealing schedule to adjust the weights to increase the weights assigned to the more complex policy networks. For example, the annealing schedule can specify that as training progresses, the weights assigned to the more complex policy networks increase linearly, while the weights assigned to the less complex policy networks decrease linearly.
[0075] In other implementations, the system employs population-based training techniques to update the weights in the hybrid data. In this technique, the system trains a population of candidate surrogate policy neural network ensembles in parallel, i.e., trains multiple different, identical candidate surrogate policy neural network ensembles. During this training, the system periodically adjusts the weights in the hybrid data used by the lower-performance combination (population) using the weights used by the higher-performance combination (population) based on the population-based training technique.
[0076] In other words, the system trains the population of ensembles in parallel, and these ensembles periodically query each other to check how they are performing relative to other ensembles. The poorly performing ensembles copy the weights (neural network parameters) of the stronger ensembles, and the poorly performing ensembles adopt hyperparameters that are local modifications of the hyperparameters of the stronger ensembles. In this way, the poorly performing ensembles are used to explore the hyperparameter space.
[0077] The use of population-based training for training and techniques for using population-based training to copy parameters and explore hyperparameters (including hybrid weights) are described in more detail in the article “Population based training of neural networks” by Jaderberg, Max, Dalibard, Valentin, Osindero, Simon, Czarnecki, Wojciech M., Donahue, Jeff, Razavi, Ali, Vinyals, Oriol, Green, Tim, Dunning, Iain, Simonyan, Karen, Fernando, Chrisantha, and Kavukcuoglu, Koray, CoRR, 2017.
[0078] To evaluate the performance of a given set of policy networks, the system can (i) evaluate based on the quality of the combined policy outputs generated by the set during training, or (ii) evaluate performance based only on the quality of the policy outputs generated by the final agent policy neural network in the set and not based on the policy outputs generated by the other agent policy neural networks in the set. As an example, the evaluation function can measure (i) the reward for the last k episodes of the task when using the combined policy to control, or (ii) the reward for the last k episodes of the task if only the final policy is used to control the agent. When the model is considered to have a significant benefit (in terms of performance) in switching from a simple model to a more complex model, using (i) to evaluate performance can yield good results. When it is not known whether this will be the case, using (ii) to evaluate performance may give better results than using (i) to evaluate.
[0079] For the exploration function of weights in the mixed data, the system can randomly add or subtract a fixed value (truncated between 0 and 1).
[0080] Thus, using population-based training, once there is a significant benefit in switching to a more complex training, the switch will occur automatically as part of the exploitation / exploration process.
[0081] The system can repeat steps 304 and 306 to update the parameters of the neural network and adjust the weights in the mixed data until certain criteria are met, such as a certain number of training iterations have been performed, or the performance of the final network meets certain criteria.
[0082] Figure 4 is a flowchart of an example process 400 for performing a training iteration. For convenience, process 400 will be described as being performed by a system of one or more computers located at one or more locations. For example, a suitably programmed reinforcement learning system (such as Figure 1 reinforcement learning system 100) can perform process 400.
[0083] When the system uses population-based training techniques, the system can perform process 400 in parallel for each candidate set in the population.
[0084] The system determines a reinforcement learning update to the current values of the parameters of the policy neural network (step 402).
[0085] The system can use any reinforcement technique suitable for the kind of network output that the policy network is configured to generate to determine the reinforcement learning update.
[0086] In particular, the reinforcement learning technique can be an on-policy technique or an off-policy technique.
[0087] When the technique is an on-policy technique, the system generates training data by controlling an agent according to the current value of the parameters of a policy network, i.e., by using the combined policy output generated according to the current value to control the agent, and then trains a neural network on the training data.
[0088] More specifically, to generate training data, the system may repeatedly cause the agent to act in an environment until a threshold amount of training data is generated. To cause the agent to act in the environment, the system receives observation data and uses each of the candidate agent policy neural networks, using each policy to process the network input including the observation data, to generate a corresponding action selection policy for each candidate agent policy neural network. Then, the system combines the action selection policies according to the weights in the mixed data at that training iteration to generate a combined action selection policy, i.e., by computing a weighted sum of the action selection policies and then selecting an action to be performed by the agent according to the combined action selection policy.
[0089] To train the neural network, the system computes the gradient of a reinforcement learning loss function that is adapted to the type of network output that the policy network is configured to generate and that promotes the combined policy to exhibit improved performance on a reinforcement learning task. Examples of reinforcement learning loss functions for on-policy reinforcement learning include the SARSA loss function and the on-policy actor-critic loss function. In particular, as part of computing the gradient, the system backpropagates through the combined policy output to each neural network in the set in order to compute an update to the network parameters.
[0090] When the technique is an off-policy technique, the system decouples the actions in the environment to generate training data from training on the training data.
[0091] Specifically, the system generates training data by causing the agent to act in the environment as described above and then stores the training data in a replay memory.
[0092] Then, the system samples training data from the replay memory and uses the sampled training data to train the neural network. Thus, the training data used in any given training iteration may have been generated using parameter values different from the current value of the given training iteration. Nevertheless, the training data is generated by controlling the agent using the combined control policy.
[0093] To train the neural network, the system computes the gradient of a policy-agnostic reinforcement learning loss function that is adapted to the kind of network output the policy network is configured to generate and that promotes the combined policy to exhibit improved performance on the reinforcement learning task. When computing the gradient, the system uses the combined policy and computes the policy that is an input to the reinforcement loss function according to the current weights in the mixed data. Examples of the reinforcement learning loss function for policy-agnostic reinforcement learning include the Q-learning loss function and the policy-agnostic actor-critic loss function. In particular, as part of computing the gradient, the system backpropagates through the combined policy output to each neural network in the set in order to compute updates to the network parameters.
[0094] The system determines a matching update to the current values of the parameters of the policy neural network (step 404). Generally, the matching update makes the action selection policies generated by the policy networks in the set consistent with each other. In some implementations, as the weights of the final agent policy neural network increase, i.e., as training proceeds, the system reduces the influence of training the candidate agent policy neural networks to generate consistent action selection policies.
[0095] In particular, the system obtains an observation data set that is received during interaction with the environment, i.e., that is received as a result of controlling an agent using the combined action selection policy. The received observation data may be the same as the observation data used in computing the reinforcement learning update or may be a different observation data set. For example, when the reinforcement learning technique is a policy-based technique, the observation data may be the same as the observation data in the generated training data. As another example, when the reinforcement learning technique is a policy-agnostic technique, the system may obtain the observation data set from a memory buffer that stores only the most recently encountered observation data, i.e., rather than from a replay memory that stores the observation data encountered over a longer period.
[0096] Then, the system computes the matching update by determining the gradient of a matching cost function that measures the difference in the policy outputs generated by the policy networks in the set. In particular, the matching cost function satisfies:
[0097]
[0098] where K is the total number of networks in the set and D is a function that measures the difference in the policy outputs generated by the policy networks π i and π j given (i) the current values of the parameters of two policy networks θ i and θ j and (ii) the current weights α in the mixed data for an observation data set.
[0099] As a special example, the function D between the policy networks π1 and π2 in the set can satisfy:
[0100]
[0101] where S is the observation data set, s is the trajectory of the observation data in the set, |s| is the number of observation data in the trajectory, |S| is the number of observation data in the set, and D KL is the K-L divergence, and the symbol (1 - α) represents 1 minus the weight assigned to the final policy network in the mixed data. In this example, due to the inclusion of the (1 - α) term, as the weight of the final agent policy neural network increases, the system reduces the impact of training the candidate agent policy neural network to generate a consistent action selection policy.
[0102] The system updates the current value of the parameters of the policy neural network (step 406). That is, the system determines the final update based on the reinforcement learning update and the matching update, and then adds the final update to the current value of the parameters. For example, the final update can be the sum or weighted sum of the reinforcement learning update and the matching update. Equivalently, the matching cost function can be added to the reinforcement learning loss function to form the overall loss function for training.
[0103] The system can continue to repeat process 400 until the criteria for updating the weights in the mixed data are met, for example, a certain amount of time has passed, a certain amount of training iterations have been performed, or until the final policy network reaches an acceptable accuracy level on the reinforcement learning task.
[0104] This specification uses the term "configured to" in connection with systems and computer program components. For a system consisting of one or more computers, being configured to perform a particular operation or action means that the system has software, firmware, hardware, or a combination thereof installed thereon, which in operation causes the system to perform the operation or action. For one or more computer programs configured to perform a particular operation or action, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.
[0105] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents thereof, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus.
[0106] The term “data processing apparatus” refers to data processing hardware and includes all kinds of devices, equipment, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The apparatus may also be, or further include, special purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). In addition to hardware, the apparatus may optionally include code that creates an execution environment for the computer program, such as code that forms processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0107] A computer program (which may also be referred to as or described as a program, software, a software application, an app, a module, a software module, a script, or code) can be written in any form of programming language (including a compiled or interpreted language, a declarative or procedural language), and a computer program can be deployed in any form, including as a stand-alone program or as a module, a component, a subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., stored in a markup language document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code)) in one or more scripts. A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0108] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or at all, and the data can be stored on storage devices in one or more locations. Thus, for example, an index database can include multiple collections of data, each of which can be organized and accessed differently.
[0109] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. In general, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and run on the same one or more computers.
[0110] The processes and logical flows described in this specification can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, or by a combination of, special purpose logic circuitry, such as an FPGA or ASIC, or special purpose logic circuitry and one or more programmed computers.
[0111] Computers suitable for executing computer programs can be based on general or special purpose microprocessors or both, or any other type of central processing unit. In general, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. In general, a computer will also include, or be operatively coupled to, one or more mass storage devices (such as magnetic disks, magneto-optical disks, or optical disks) for storing data, from which it can receive data or to which it can transfer data, or both. However, a computer need not have such devices. In addition, a computer can be embedded in another device (such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name just a few examples).
[0112] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices); magnetic disks (such as internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM.
[0113] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (such as a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user, and a keyboard and a pointing device (such as a mouse or a trackball) by which the user can provide input to the computer. Other types of devices may also be used to provide for interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on the user device in response to a request received from the web browser. Further, the computer may interact with the user by sending a text message or other form of message to a personal device (such as a smartphone running a messaging application) and receiving a response message from the user.
[0114] The data processing apparatus for implementing a machine learning model may further include, for example, a dedicated hardware accelerator unit for processing common and computationally intensive portions (i.e., inference, workload) of machine learning training or production.
[0115] The machine learning model may be implemented and deployed using a machine learning framework (such as the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework).
[0116] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component (such as a data server), or includes a middleware component (such as an application server), or includes a frontend component (such as a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0117] The computing system can include a client and a server. The client and the server are generally far apart from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on the respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (such as an HTML page) to a user device, for example, in order to display data to a user who acts as a client and interacts with the device and to receive user input from the user. Data generated at the user device, such as the result of a user interaction, can be received at the server from the device.
[0118] Although this specification contains many specific implementation details, these should not be construed as limitations on any invention or the scope of what is claimed, but rather as descriptions of specific features of particular embodiments of a particular invention. Certain features described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination can be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination.
[0119] Similarly, although operations are described in a specific order in the figures and the claims, this should not be understood as requiring that the operations be performed in the specific order shown or sequentially, or that all of the shown operations be performed, to obtain a desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above embodiments should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0120] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims can be performed in a different order and still achieve the desired result. As one example, the processes described in the figures need not be in the particular order shown or sequential to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for training a final agent policy neural network for selecting actions to be performed by an agent, where the agent interacts with an environment by performing actions selected using the final agent policy neural network to perform a reinforcement learning task, the method comprising: Maintaining data specifying a plurality of candidate agent policy neural networks, where the plurality of candidate agent policy neural networks includes the final agent policy neural network, and where the final agent policy neural network defines an action selection policy that is more complex than the action selection policies defined by at least one other candidate agent policy neural network, where the more complex action selection policy is an action selection policy that requires more training steps to train; Initializing mixed data that assigns corresponding weights to each of the candidate agent policy neural networks; Jointly training the plurality of candidate agent policy neural networks to perform the reinforcement learning task, including, in each of a plurality of training iterations: Obtaining a training network input including observation data of the environment, where the observation data is one or more of image, object position data, sensor data, and electronic signals, or a simulated version of one or more of image, object position data, sensor data, and electronic signals; Using the candidate agent policy neural networks and according to the weights in the mixed data at that training iteration, generating a combined action selection policy using the training network input; Using the combined action selection policy to select an action to be performed by the agent in the environment, such that the agent interacts with the environment by performing the action selected using the combined action selection policy; and Using reinforcement learning techniques to train the candidate agent policy neural networks to generate a combined action selection policy that results in improved performance on the reinforcement learning task; And During training, repeatedly adjusting the weights in the mixed data according to a defined performance metric to support higher-performance candidate agent policy neural networks, the defined performance metric being based on a measurement of the rewards received by the agent in the last k rounds of the task when controlling the agent using the combined action selection policy or only using the final agent policy neural network.
2. The method according to claim 1, wherein, Jointly training the plurality of candidate agent policy neural networks to perform the reinforcement learning task further includes: Training each of the candidate agent policy neural networks to generate an action selection policy that is consistent with other action selection policies generated by each of the other candidate agent policy neural networks by processing the same training network input, and in the reinforcement learning technique, using a reinforcement learning cost that depends on the difference between each of the action selection policies and each of the other action selection policies.
3. The method according to claim 2, wherein, Training the candidate agent policy neural networks to generate consistent action selection policies includes: As the weights of the final agent policy neural network increase, reducing the influence of training the candidate agent policy neural networks to generate consistent action selection policies by reducing the reinforcement learning cost that depends on the difference between each of the action selection policies and each of the other action selection policies.
4. The method according to any one of claims 1-3, wherein The final surrogate policy neural network has more parameters than at least one other candidate surrogate policy neural network.
5. The method according to any one of claims 1-4, wherein Compared to at least one other candidate surrogate policy neural network, the final surrogate policy neural network generates an output that defines a larger action space for the surrogate.
6. The method according to any one of claims 1-5, wherein, Using the candidate surrogate policy neural networks and based on the weights in the mixed data at the training iteration, generating a combined action selection policy using the training network input, includes: Using each of the candidate surrogate policy neural networks to process the training network input to generate a corresponding action selection policy for each candidate surrogate policy neural network; and Combining the action selection policies based on the weights at the training iteration to generate a combined action selection policy.
7. The method according to any one of claims 1-6, wherein, Jointly training the plurality of candidate surrogate policy neural networks to perform a reinforcement learning task includes: Training a combined population of candidate surrogate policy neural networks, and wherein the weights in the mixed data are repeatedly adjusted according to the performance metric to favor higher performance candidate surrogate policy neural networks, including: During training, using a population-based training technique, based on the weights used by higher performance combinations according to the performance metric, to adjust the weights in the mixed data used by lower performance combinations according to the performance metric.
8. The method according to claim 7, wherein, The performance of the combination is based on the quality of the combined policy output generated during training.
9. The method according to claim 7, wherein The performance of the combination is based only on the quality of the policy output generated by the final surrogate policy neural network in the combination and not on the policy output generated by other surrogate policy neural networks in the combination.
10. One or more computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the corresponding operations of the method of any one of claims 1-9.
11. A system comprising one or more computers and a storage device storing instructions that, when executed by one or more computers, cause the one or more computers to perform the corresponding operations of the method of any one of claims 1-9.