Deep reinforcement learning using fast update recurrent neural networks and slow update recurrent neural networks
By using the combination method of time-hierarchical neural network and recurrent neural network in the reinforcement learning system, latent variables are generated and additional information is used to improve the reward signal, the problems of environmental uncertainty and sparse rewards are solved, and the task completion performance of the agent is improved.
Patent Information
- Application Number
- CN202510131927.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-05-29
- Filing Date
- 2019-05-29
- Publication Date
- 2025-06-13
AI Technical Summary
Existing reinforcement learning systems are difficult to effectively train neural networks when dealing with environmental uncertainty and sparse rewards, resulting in poor performance of agents when tasks are completed.
The time-hierarchical neural network system is adopted, combining the rapid update of the recurrent neural network and the slow update of the recurrent neural network to generate latent variables and adjust action selection to solve the uncertainty problem; at the same time, additional information is used to generate improved reward signals to improve the efficiency of neural network training.
Effective action selection and neural network training in the case of environmental uncertainty and sparse rewards are realized, and the performance of the agent on reinforcement learning tasks is improved.
Smart Images

Figure CN120146141A_ABST
Abstract
Description
[0001] This application is a divisional application of the PCT patent application with the filing date of May 29, 2019, application number 201980032994.6, and invention title "Deep Reinforcement Learning Using Fast-Update Recurrent Neural Networks and Slow-Update Recurrent Neural Networks". Technical Field
[0002] This disclosure relates to reinforcement learning, and more particularly, to deep reinforcement learning using fast-update recurrent neural networks and slow-update recurrent neural networks. Background Art
[0003] This specification relates to reinforcement learning.
[0004] In a reinforcement learning system, an agent interacts with an environment by performing actions selected by the reinforcement learning system in response to observations representing the current state of the environment.
[0005] Some reinforcement learning systems select actions to be performed by an agent in response to a given observation based on the output of a neural network.
[0006] A neural network is a machine learning model that uses one or more layers of non-linear units to predict an output for a received input. Some neural networks are deep neural networks that include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as the input to the next layer in the network (i.e., the next hidden layer or the output layer). Each layer of the network generates an output based on the received input according to the current values of a corresponding set of parameters. Summary of the Invention
[0007] This specification generally describes a reinforcement learning system that selects actions to be performed by a reinforcement learning agent that interacts with an environment.
[0008] Specific embodiments of the subject matter described in this specification may be implemented to achieve one or more of the following advantages.
[0009] Aspects related to the temporal extended behavioral patterns of agents involved in automatic recognition and environmental interaction can contribute to the evaluation of the performance of an agent by analyzing the behavior of the agent including a neural network system. More specifically, the described methods and systems are capable of identifying high-level behavioral patterns that might otherwise be obscured by the complexity of the system. This in turn enables humans to evaluate the agent based on these patterns, such as determining the efficiency with which the agent is achieving its goals and the techniques it is using. Thus, these techniques can represent the very complex internal state of the agent in an understandable manner, which helps in the evaluation and thus in the implementation of the control of real-world or simulated systems by autonomous or semi-autonomous agents as described above. For example, the described methods and systems can help humans understand how or why an agent behaves when controlling a system, and thus can help in customizing the agent for a particular task.
[0010] Certain described aspects relate to using a time hierarchical neural network system to select actions. These aspects allow the actions selected at each time step to be consistent with the long-term plan of the agent. The agent can then achieve improved performance on tasks where the selection of actions depends on data received in observations at a large number of time steps prior to the current time step. Because the neural networks are jointly trained, the improvement in performance can be achieved without overly increasing the amount of computational resources consumed during the training of recurrent neural networks. By using a time hierarchical neural network system to generate latent variables and then conditioning action selection on the latent variables, the system can effectively perform tasks even when the observations received do not fully characterize the true state of the environment, i.e., when the true state of the environment is not clear even considering all the observations the agent has received so far. For example, many tasks involve multiple agents operating independently in an environment, such as multiple autonomous vehicles or multiple robots. In these tasks, the observations received by a single agent cannot fully characterize the true state of the environment because the agent does not know the control strategies of the other agents or whether these control strategies will change over time. Thus, there is inherent uncertainty because the agent cannot know how the other agents will behave or react to changes in the environmental state. The system can address this uncertainty by sampling latent variables from a posterior distribution that depends on long-term and short-term data received from the environment using a time hierarchical neural network.
[0011] Certain described aspects allow a system to learn rewards for neural network training. These aspects allow the system to effectively train a neural network, i.e., train the neural network such that it can be used to enable an agent to have acceptable performance on a reinforcement learning task, even in cases where the rewards associated with task performance are very sparse and very delayed. In particular, even when the data specifying how the extracted data relates to task performance is not specified before the start of training, the system can utilize additional information that can be extracted from the environment to generate an improved reward signal.
[0012] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 An example reinforcement learning system is shown.
[0014] Figure 2 Processing performed by the reinforcement learning system at a given time step is shown.
[0015] Figure 3A and Figure 3B An example architecture of a neural network used by the reinforcement learning system is shown.
[0016] Figure 4 is a flowchart of an example process for selecting an action to be performed by an agent.
[0017] Figure 5 is a flowchart of an example process for updating parameter values of a candidate neural network and a reward mapping.
[0018] Figure 6 An example of a user interface rendering of agent behavior is shown.
[0019] Figure 7 Another example of a user interface rendering of agent behavior is shown.
[0020] Like reference numerals and names in the figures represent like elements. DETAILED DESCRIPTION
[0021] This specification describes a reinforcement learning system that controls an agent interacting with an environment by processing data (i.e., "observations") representing the current state of the environment at each of a plurality of time steps to select an action to be performed by the agent.
[0022] At each time step, the environmental state at that time step depends on the environmental state at the previous time step and the action performed by the agent at the previous time step.
[0023] In some embodiments, the environment is a real-world environment and the agent is a mechanical agent that interacts with the real-world environment, e.g., a robot navigating in the environment or an autonomous or semi-autonomous land, air, or sea vehicle. In particular, the environment can be a real-world environment in which multiple agents (i.e., multiple autonomous vehicles or multiple robots) are operating. In these cases, control strategies for controlling other agents are typically not available to the system.
[0024] In these embodiments, the observations can include, for example, one or more of images, object position data, and sensor data to capture observations as the agent interacts with the environment, e.g., sensor data from an image, distance, or position sensor or from an actuator.
[0025] For example, in the case of a robot, the observations can include data characterizing the current state of the robot, e.g., one or more of joint positions, joint velocities, joint forces, torques, or accelerations (e.g., gravity-compensated torque feedback), and the global or relative pose of an item held by the robot.
[0026] In the case of a robot or other mechanical agent or vehicle, the observations can similarly include one or more of position, linear or angular velocity, force, torque, or acceleration, and the global or relative pose of one or more components of the agent. The observations can be defined as one-dimensional, two-dimensional, or three-dimensional and can be absolute and / or relative observations.
[0027] The observations can also include, for example, sensed electrical signals such as motor current or temperature signals; and / or image or video data from, e.g., a camera or LIDAR sensor, e.g., data from a sensor of the agent or from a sensor placed separately from the agent in the environment.
[0028] In these embodiments, the actions can be control inputs for controlling the robot (e.g., torques on the robot joints or higher-level control commands), or control inputs for controlling an autonomous or semi-autonomous land, air, sea vehicle (e.g., torques on the control surfaces of the vehicle or other control elements or higher-level control commands).
[0029] In other words, the actions can include, for example, position, velocity, or force / torque / acceleration data of one or more joints of a robot or components of another mechanical agent. The action data can additionally or alternatively include electronic control data (such as motor control data), or more generally, data for controlling one or more electronic devices within the environment, the control of which has an impact on the observed state of the environment. For example, in the case of an autonomous or semi-autonomous land, air, or sea vehicle, the actions can include actions for controlling the driving (e.g., steering) and movement (e.g., braking and / or accelerating) of the vehicle.
[0030] In some other applications, the agent can control actions in the real-world environment including various equipment, such as actions in a data center, a power / water distribution system, or a manufacturing plant or service facility. The observations can then be related to the operation of the plant or facility. For example, the observations can include observations of power usage or water usage of the equipment, or observations of power generation control or power distribution control, or observations of resource usage or waste generation. These actions can include actions for controlling or regulating the operating conditions of the various devices in the plant / facility, and / or actions that result in changes in the settings during the operation of the plant / facility, such as actions for adjusting or turning on / off components of the plant / facility.
[0031] In the case of an electronic agent, the observations can include data from one or more sensors (such as current, voltage, power, temperature, and other sensors) that monitor a part of the manufacturing plant or service facility and / or electronic signals representing the functions of the various electronic and / or mechanical equipment.
[0032] Figure 1 An example reinforcement learning system 100 is shown. The reinforcement learning system 100 is an example of a system implemented as a computer program on one or more computers located at one or more locations, where the systems, components, and techniques described below are implemented.
[0033] The system 100 controls an agent 102 that interacts with an environment 104 by selecting an action 106 to be performed by the agent 102 and then causing the agent 102 to perform the selected action 106. The actions repeatedly performed by the agent typically cause the environmental state to repeatedly transition to new states.
[0034] The system 100 includes a policy neural network 110 and a training engine 116, and maintains a set of model parameters 118 of the policy neural network 110.
[0035] At each of a plurality of time steps, the policy neural network 110 is configured to process an input including a current observation 120 representing the current state of the environment 104 according to the model parameters 118 to generate an action selection output 122 ("action selection policy").
[0036] System 100 uses the action selection output 122 to select the action 106 to be executed by the agent at the current time step. Several examples of using the action selection output 122 to select the action 106 to be executed by the agent 102 are described next.
[0037] In one example, the action selection output 122 can define a probability distribution of actions in the set of possible actions that can be executed by the agent. For example, the action selection output 122 can include corresponding numerical probability values for each action in the set of possible actions that the agent can execute. As another example, the action selection output 122 can include parameters of the distribution of the set of possible actions, e.g., parameters of a multivariate normal distribution of the set of possible actions when the set of possible actions is represented as a continuous space. For example, system 100 can select the action to be executed by the agent by sampling actions according to the probability values of the actions or by selecting the action with the highest probability value.
[0038] In another example, the action selection output 122 can directly define the action to be executed by the agent, e.g., by defining the torque values that should be applied to the joints of the robotic agent.
[0039] In another example, the action selection output 122 can include corresponding Q-values for each action in the set of possible actions that can be executed by the agent. System 100 can process the Q-values (e.g., using the softmax function) to generate corresponding probability values for each possible action, which can be used to select the action to be executed by the agent (as described above). System 100 can also select the action with the highest Q-value as the action to be executed by the agent.
[0040] The Q-value of an action is an estimate of the "return", which is generated by the agent executing the action in response to the current observation 120 and then selecting future actions to be executed by the agent 102 according to the current values of the policy neural network parameters.
[0041] The return refers to the cumulative measure of the "reward" 124 received by the agent, e.g., the time-discounted sum of the rewards. The agent can receive a corresponding reward 124 at each time step, where the reward 124 is specified by a scalar numerical value and characterizes, for example, the progress of the agent in completing the assigned task.
[0042] In some cases, system 100 may select an action to be performed by an agent according to an exploration strategy. For example, the exploration strategy may be an ∈-greedy exploration strategy, where system 100 selects an action to be performed by the agent according to the action selection output 122 with a probability of 1 - ∈, and randomly selects an action with a probability of ∈. In this example, ∈ is a scalar value between 0 and 1. As another example, the system may add randomly sampled noise to the action selection output 122 to generate a noisy output, and then use the noisy output instead of output 122 to select an action.
[0043] The policy neural network 110 includes a time hierarchical recurrent neural network 112. The time hierarchical recurrent neural network 112 in turn includes two recurrent neural networks: a fast-update recurrent neural network (RNN) that updates its hidden state at each time step and a slow-update recurrent neural network that updates its hidden state at time steps less than all time steps.
[0044] In this specification, the hidden state of the fast-update recurrent neural network will be referred to as the "fast-update hidden state", and the hidden state of the slow-update recurrent neural network will be referred to as the "slow-update hidden state".
[0045] As will be described in more detail below, the fast-update RNN is called "fast-update" because the fast-update hidden state is updated at each time step. In contrast, the slow-update RNN is called "slow-update" because the slow-update hidden state is not updated at each time step, but only when a specific criterion is met.
[0046] At each time step, the fast-update recurrent neural network receives an input including the observation 120 and the slow-update hidden state of the slow-update recurrent neural network, and uses this input to update the fast-update hidden state.
[0047] The policy neural network 110 then uses the updated fast-update hidden state to generate the action selection output 122. By utilizing the time hierarchical RNN 112, system 100 can select an action consistent with the agent's long-term plan at each time step, and system 100 can effectively control the agent, even on tasks where the action that may be required for the user to select at any given time step depends on the data received in observations at a large number of time steps before the given time step.
[0048] Refer to the following Figures 2 - 4 The operations and architectures of the policy neural network 110 and the time hierarchical recurrent neural network 112 are described in more detail.
[0049] The training engine 116 is configured to train the policy neural network 110 by repeatedly updating the model parameters 118 of the policy neural network 110 based on the interaction of the agent with the environment (i.e., using the observations 120 and rewards 124 received due to the interaction of the agent with the environment).
[0050] The training engine 116 can use the data collected from the interaction of the agent with the environment and use reinforcement learning techniques to train the policy neural network 110 to increase the return received by the agent (i.e., the cumulative measure of rewards). Since the return measures the progress of the agent in completing the task, training the policy neural network 110 to increase the return enables the agent to successfully complete the specified task while being controlled by the policy neural network 110.
[0051] However, for some tasks, the external reward 124 may not be sufficient to generate a high-quality learning signal that allows the training engine 116 to train the policy neural network 110 to have high performance on the task. For example, the reward 124 may be very sparse and very delayed. That is, even if a given action performed at a given time step contributes to the successful completion of the task, the corresponding reward may not be received for many time steps after the given time step.
[0052] As a special example, some tasks may require a large number of operations to be successfully completed. However, the reward 124 may be non-zero only when the task reaches an end state. For example, it may be positive only when the environment reaches the final state of task completion. This may make it difficult to train the policy neural network 110 using only the sparse reward 124 as the learning signal.
[0053] In these cases, even when the data specifying how the extracted additional data 152 is related to task performance is not specified before the start of training, the system 100 can utilize the additional data 152 that can be extracted from the environment to generate an improved reward signal. In particular, the reward mapping engine 160 can map the additional data 152 (and optionally the reward 124 in the non-zero case) to a learning reward value 162 at each time step, and this learning reward 162 can be used by the training engine 116 to update the model parameters 118 (i.e., instead of relying solely on the sparse reward 124). The learning reward is described in more detail below with reference to Figure 2 and Figure 5 The learning reward is described in more detail.
[0054] Additionally, system 100 can generate data identifying the time-extended behavioral patterns of the agents interacting with the environment and provide it to the user for presentation. That is, system 100 can generate user interface data and provide the user interface data to the user for presentation on the user device. This can facilitate the evaluation of the agents' performance by allowing the user to analyze the agents' behavior while being controlled by the policy neural network. The following references Figure 6 and Figure 7 describe the generation of such user interface data.
[0055] Figure 2 illustrates the operation of the reinforcement learning system at time step t during the interaction of the agent with the environment.
[0056] At Figure 2 's example, the observation 120 received at time step t is, for example, an image of the environment captured by the agent's camera sensor at time step t. However, as described above, in addition to or instead of an image, the observation can also include other data characterizing the state of the environment.
[0057] Figure 2 's example illustrates how the fast-update hidden state 210 and the slow-update hidden state 220 are updated over a sequence of time steps including time step t. As can be seen from Figure 2 's example, the fast-update RNN updates the fast-update hidden state at each time step using an input that includes the observation at that time step. On the other hand, the slow-update RNN does not update the slow-update hidden state at each time step. In Figure 2 , the time steps at which the hidden state is updated are the time steps at which the arrow representing the hidden state is connected to a black circle, and the time steps at which the hidden state is not updated are the time steps without a black circle. For example, at the time steps before time step t in the sequence, there is no update to the state of the slow-update RNN.
[0058] At any given time step, if the criteria for updating the slow-update hidden state are met, the system uses the slow-update RNN to process a slow-update input that includes the fast-update hidden state (i.e., the final fast-update hidden state after the time step before the given time step) to update the slow-update hidden state. Thus, at time step t, the input to the slow-update RNN can be the fast-update hidden state after time step t-1, and the slow-update RNN can use this input to update the slow-update hidden state at time step t.
[0059] The slow-update RNN can have any suitable recurrent architecture. For example, the slow-update RNN can be a stack of one or more long short-term memory (LSTM) layers. When there is more than one layer in the stack, the slow-update hidden state can be the hidden state of the last LSTM layer in the stack.
[0060] When the criteria are not met, the system avoids updating the slow-update hidden state at that time step, i.e., the slow-update hidden state is not updated until it is used for further processing at that time step.
[0061] At each time step, the system generates the parameters of the prior distribution 222 of the possible values of the latent variable based on the slow-update hidden state, e.g., by applying a linear transformation to the slow-update hidden state. Typically, the latent variable is a vector with a fixed dimension, so the prior distribution is a multivariate distribution of the possible values of the latent variable. For example, the prior distribution can be a normal distribution, so the parameters can be the mean and covariance of the multivariate normal distribution. When the slow-update hidden state is not updated at a given time step, the system can reuse the most recently computed prior parameters instead of having to recompute the prior parameters based on the same hidden state as in the previous time step.
[0062] Then, the system uses the fast-update RNN to process the fast-update input to update the fast-update hidden state, where the fast-update input includes the observation 120, the slow-update hidden state, and the parameters 222 of the prior distribution.
[0063] Like the slow-update RNN, the fast-update RNN can have any suitable recurrent architecture. For example, the fast-update RNN can be a stack of one or more long short-term memory (LSTM) layers. When there is more than one layer in the stack, the fast-update hidden state can be the hidden state of the last LSTM layer in the stack.
[0064] When the observation is an image (as in the example of Figure 2 ), or other high-dimensional data, the fast-update RNN can include a convolutional neural network (CNN) or other encoder neural network that encodes the observation before its encoded representation is processed by the recurrent layers of the fast-update RNN.
[0065] In addition, in some embodiments, the fast-update RNN and the slow-update RNN are enhanced using a shared external memory, i.e., as part of updating their respective hidden states, both RNNs read from and write to the same shared external memory. An example architecture for enhancing an RNN using an external memory that can be used by a system is the Differentiable Neural Computer (DNC) memory architecture, which is described in more detail in the following reference: Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwinska, Sergio Gomez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al. Hybrid computing using a neural network with dynamic external memory. Nature, 538(7626):471, 2016.
[0066] Once the fast-update hidden state has been updated at time step t, the system uses the fast-update hidden state, e.g., by applying a linear transformation to the slow-update hidden state, to generate the parameters of the posterior distribution 212 of the possible values of the latent variable.
[0067] The system then samples the latent variable 232 from the posterior distribution 212, i.e., samples a possible value of the latent variable according to the probabilities in the posterior distribution 212.
[0068] The system then uses the sampled latent variable 232 to select the action 106 to be performed in response to the observation 120.
[0069] In particular, the system uses one or more policy neural network layers to process the sampled latent variable to generate a policy output (i.e., the action selection output 122), and then uses the policy output to select the action 106.
[0070] The fast-update input processed by the fast-update RNN at any given time step may also include additional inputs (in addition to the observation, the slow-update hidden state, and the prior parameters). For example, the fast-update hidden state at time step t may satisfy:
[0071]
[0072] where u t is the observation x at time step t tThe encoded representation, a t-1 is the action executed at the previous time step, r t-1 is the reward (external or learned) from the previous time step, is the slow-update hidden state at time t, is the previous fast-update hidden state, and are the mean and covariance of the prior distribution generated using the slow-update hidden state, z t-1 is the latent variable sampled at the previous time step.
[0073] By sampling latent variables from the distribution instead of directly using the fast-update hidden state to select actions, the system can effectively address the uncertainty in cases where the received observations do not fully characterize the true state of the environment, i.e., the uncertainty in cases where the true state is not clear even considering all the observations received by the agent so far.
[0074] When time step t occurs during learning (or when time step t occurs after learning but the reward is part of the input to the fast-update RNN) and the system is using the learned reward to augment the external reward, the system also receives additional data 152 extracted from the environment and applies the current reward mapping to the additional data 152 to generate the learned (or "internal") reward 162. This learned reward can then be used to update the policy neural network 110. The application mapping and update mapping data will be described in more detail below.
[0075] In particular, the system can train the neural network 110 to maximize the received return using reinforcement learning.
[0076] As a specific example, in some cases, the neural network 110 also generates a value output (also referred to hereinafter as the "baseline output") based on the sampled latent variables, which is an estimate of the return generated from the environment in the current state. In these cases, the system can use actor-critic reinforcement learning techniques to train the neural network 110.
[0077] During training, the system uses a prior distribution generated by prior parameters to regularize the latent variables. This helps ensure that both the slow-update RNN and the fast-update RNN capture long-term temporal correlations and promotes the memorization of the information received in the observations. In particular, during training, the system enhances the reinforcement learning technique by minimizing the divergence (e.g., KL divergence) between the prior distribution and the posterior distribution by training a fast-update recurrent neural network and a slow-update recurrent neural network. In some cases, to prevent the two distributions from matching, the system also includes a term in the loss function during training that penalizes the KL divergence between the prior distribution and a multivariate Gaussian distribution with a mean of 0 and a standard deviation of 0.1 or some other small fixed value.
[0078] As will be described in more detail below, during training, the system can also use the hidden state of the fast-update recurrent neural network to generate a corresponding auxiliary output for each of one or more auxiliary tasks and train the neural network on the auxiliary tasks based on the auxiliary outputs of the auxiliary tasks.
[0079] Figure 3A and Figure 3B is a diagram showing an example architecture of various components of the policy neural network 110. Figure 3A The legend shown, identifying the various neural network components represented by the symbols in the figure, also applies to Figure 3B the figures in
[0080] Figure 3A shows an example of the high-level architecture 310 of the policy neural network 110. In particular, the architecture shows a visual embedding (convolutional) neural network 312 that processes input observations to generate an encoded representation, and a recurrent processing block 314 that includes a slow-update RNN and a fast-update RNN that generate latent variables.
[0081] The architecture 310 also shows: a policy neural network layer 322 that generates an action selection output conditioned on the latent variables; a baseline neural network layer 320 that uses the latent variables to generate a baseline score, which is used for training when using reinforcement learning techniques that require a baseline state value score, such as actor-critic based techniques, such as the Importance Weighted Actor-Learner (IMPALA) technique described in the following literature: IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures, Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, Koray Kavukcuoglu. As described above, the system can enhance this training by adding additional loss terms that use a prior distribution to regularize the latent variables.
[0082] The architecture 310 also includes a plurality of auxiliary components that can be used to improve the training of the policy neural network. For example, a reward prediction neural network 316 that attempts to predict the next reward to be received, and a pixel control neural network 318 that predicts how much the pixels of the current observation will change in the next observation. These neural networks can be jointly trained with the rest of the policy neural network, and gradients can be backpropagated from these neural networks into the policy neural network to improve the representation learned by the policy neural network during training.
[0083] Figure 3A The detailed architecture 330 of a visual embedding neural network is also shown, which receives an observation and generates an encoded representation that is provided as part of the input to the recurrent processing block 314.
[0084] Figure 3B The detailed architecture 340 of the policy neural network, the detailed architecture of the baseline neural network 350, and the detailed architecture of the recurrent processing block 360 are shown, and the recurrent processing block 360 includes a VU that is described in more detail in the detailed variational unit (VU) architecture 370.
[0085] Figure 4FIG. 400 is a flow chart of an example process 400 for selecting an action to be performed by an agent. For convenience, process 400 will be described as being performed by a system of one or more computers located at one or more locations. For example, a suitably programmed reinforcement learning system (e.g., Figure 1 reinforcement learning system 100) can perform process 400.
[0086] The system determines whether the criteria for updating the slow-update hidden state are met (step 402). As described above, the criteria are typically met at time steps less than all time steps.
[0087] For example, the criteria can be met only at predetermined intervals during the interaction of the agent with the environment. That is, the criteria can be met at every N time steps, where N is an integer greater than 1. For example, N can be an integer in the range of 5 to 20 (inclusive of 5 and 20).
[0088] As another example, the system can determine a difference metric between the observation at the current time step and the observation at the previous time step at which the hidden state of the slow-update recurrent neural network was updated. In this example, the criteria are met at a given time step only if the difference metric meets a threshold, i.e., when the observation is sufficiently different from the observation at the previous time step at which the hidden state was updated. For example, the distance metric can be the cosine similarity between the encoded representations generated by the encoder neural network, and the criteria are likely to be met only if the cosine similarity is below a threshold similarity.
[0089] When the criteria are met, the system updates the slow-update hidden state as described above (step 404). When the criteria are not met, the system avoids updating the slow-update hidden state.
[0090] The system receives an observation at the current time step (step 406).
[0091] The system processes the fast-update input as described above to update the fast-update hidden state (step 408).
[0092] The system uses the fast-update hidden state to select an action (step 410). For example, the system can use the fast-update hidden state to sample a latent variable as described above, and then use one or more policy neural network layers to process the sampled latent variable to generate a policy output.
[0093] As described above, in some embodiments, instead of or in addition to the actual external reward value received from the environment, the system also enhances the training of the neural network by using a learned reward value.
[0094] In particular, the system uses a reward mapping that maps data extracted from the environment to reward values. Generally, the extracted data includes the respective numerical values of each of one or more reward signals, which can be directly observed from the environment (e.g., directly measured by the agent's sensors), determined based on the agent's sensor measurements, or determined based on other information accessible to the system about the environment and the agent. As a specific example, one or more reward signals can identify the position of the agent in the environment. As another specific example, one or more reward signals can identify the distance of the agent relative to known objects or positions in the environment. As yet another example, one or more reward signals can be external rewards received from the environment, e.g., rewards indicating that a task has been successfully completed or unsuccessfully terminated.
[0095] Assuming that the reward signals can be mapped to reward values, the reward mapping data identifies, for each reward signal, the partitioning of the possible numerical values of the reward signal into multiple partitions and then maps each partition to a corresponding reward value of the reward signal. Thus, for a given reward signal, the data maps different partitions to different reward values.
[0096] When there is more than one reward signal extracted from the environment at a given time step, the system can combine the reward values of the reward signals (e.g., average or sum) to generate a final reward value.
[0097] In the manner in which the system learns the reward mapping, the system can maintain multiple candidate neural networks during training and can use the set of multiple candidates to jointly update the parameter values and learn the reward mapping.
[0098] In particular, during training, the system can maintain, for each of these candidate neural networks, data identifying the parameter values of the candidate neural network and data identifying the reward mapping used during the training of the candidate neural network.
[0099] During training, for each of the candidate neural networks, the system repeatedly updates the parameter values and the reward mapping used by the candidate neural network by performing in parallel the training operations described below with reference to Figure 5 Once training is complete, i.e., after repeatedly performing the training operations, the system selects the trained network parameter values from the parameter values in the maintenance data of the candidate neural networks. For example, after repeatedly performing the training operations, the system can select the maintained parameter values of the candidate neural network that has the best performance on the reinforcement learning task. Thus, the system can continuously improve the reward mapping used by the candidate neural networks in order to enhance the training process and generate a higher quality trained neural network than can be achieved using only external rewards.
[0100]
[0101] Figure 5 is a flowchart of an example process 500 for updating the parameter values and reward mappings of candidate neural networks. For convenience, process 500 will be described as being performed by a system of one or more computers located at one or more locations. For example, a suitably programmed reinforcement learning system (e.g., Figure 1 the reinforcement learning system 100) can perform process 500.
[0102] The system can repeatedly and parallelly perform process 500 for all candidate neural networks in the set. In some cases, the system asynchronously performs process 500 for all candidate neural networks.
[0103] The system uses the maintained reward mapping of the candidate neural network to train the candidate neural network (step 502). That is, the system uses reinforcement learning techniques (e.g., the techniques described above) on the interaction between the agent and the environment to train the candidate neural network to update the parameter values of the candidate neural network. During training, the system uses the data extracted from the environment and the reward mapping of the candidate neural network to generate the rewards used in training.
[0104] For example, the system can train the candidate neural network within a predetermined number of time steps or within a predetermined number of task episodes or a predetermined amount of time to complete step 502.
[0105] The system determines a quality metric for the trained candidate neural network, which represents to what extent the candidate neural network can control the agent to perform tasks (relative to other candidate neural networks) (step 504). An example of a quality metric can be the task segments successfully completed when the agent is controlled by the trained candidate neural network. In some cases, the system does not use external rewards as part of the reward mapping, but can use the average external reward received when using the candidate neural network to control the agent as the quality metric.
[0106] The system uses the quality metric to determine new network parameter values and a new reward mapping for the candidate neural network (step 506).
[0107] Generally, the system can update the new network parameter values, the new reward mapping, and optionally other hyperparameters of the training (e.g., the interval N as described above), such that the weaker neural networks replicate the parameters of the stronger neural networks while exploring the space of possible reward mappings and optionally other hyperparameters.
[0108] For example, the system can sample another candidate from the set and compare the quality metrics of the two candidates.
[0109] If a candidate quality metric is found to be worse than another candidate by more than a threshold amount, the system can copy the parameters, reward mapping, and hyperparameters of the better candidate for the worst-performing candidate. The system can then explore new reward mappings and hyperparameters for the worst-performing candidate. For example, the system can determine for each reward signal partition whether to modify the reward value mapped by the partition with a predetermined probability and change the value mapped by the partition in response to determining to modify the reward value mapped by the partition, e.g., randomly perturb the value in either direction by a fixed fraction of the current value
[0110] As another example, the system can sample one or more other candidate neural networks and determine whether the quality metric indicates that the candidate neural network performs better than all of the one or more other candidate neural networks.
[0111] In response to determining that the candidate neural network does not perform better than all of the one or more other candidate neural networks, the system (1) sets the new network parameter values to the maintained network parameter values of the best-performing sampled other candidate neural network, and (2) sets the new reward mapping to a modified version of the maintained reward mapping of the best-performing other candidate neural network. For example, the system can generate the modified version by determining for each reward signal partition whether to modify the reward value mapped by the partition with a predetermined probability and changing the value mapped by the partition in response to determining to modify the reward value mapped by the partition (e.g., randomly perturb the value in either direction by a fixed fraction of the current value).
[0112] In response to determining that the candidate neural network does perform better than all of the one or more other candidate neural networks, the system sets the new network parameter values to the updated network parameter values after training in step 502 and does not modify the reward mapping.
[0113] By repeatedly executing process 500 for all candidate neural networks, the system trains the neural network while exploring the space of possible reward mappings to discover a reward mapping that leads to better agent performance on the task.
[0114] As described above, in some embodiments, the system can provide certain data derived from the interaction of the agent with the environment to the user for presentation. Thus, the system can automatically identify temporal extended behavioral patterns of the agent interacting with the environment and provide the data identifying these patterns as output to the user device.
[0115] In some embodiments, to discover these patterns, the system can allow the agent to interact with the environment and then, while the agent is interacting with the environment, determine multiple time series of features characterizing the state of the environment.
[0116] Then, the system can use the captured time series to train a variational autoencoder, specifically, a variational sequence autoencoder. The variational sequence autoencoder can include a recurrent neural network encoder coupled to a recurrent neural network decoder. The recurrent neural network encoder is used to encode an input data sequence into a set of latent variables, and the recurrent neural network decoder is used to decode the set of latent variables to generate an output data sequence, that is, to reproduce the output data sequence according to the input data sequence. During training, the latent variables are constrained to approximate a defined distribution, such as a Gaussian distribution.
[0117] In some embodiments, when the agent interacts with the environment, the system also captures a further sequence of observations of the environment. Then, the system can use the trained variational sequence autoencoder to process the further sequence of observations to determine a sequence of the set of latent variables. Optionally, the system can process the sequence of the set of latent variables to identify a sequence of clusters in the space of the latent variables, where each cluster represents a time-extended behavior pattern of the agent.
[0118] The system can automatically identify and present to the user data that identifies the complex high-level behaviors of the agent. In an embodiment, the method represents these behaviors in a human-readable manner, as further described below. For example, the clusters identified by the system can correspond to high-level behavior patterns that are more easily recognizable by humans compared to the individual actions that make up the high-level behavior patterns.
[0119] The features that characterize the environmental state for identifying behaviors can be agent-centric features, that is, features that are related to the high-level behaviors performed by the agent and represented by the clusters. Thus, these features can include, for example, features that characterize the position or state of the agent relative to the environment. For example, they can be features defined relative to the agent, such as the position or orientation of the agent or a part of the agent, or the attributes of the agent, or the state of the agent or the environment related to the short-term or long-term rewards of the agent.
[0120] In some embodiments, the features can be manually defined and provided as an input to the method. In the sense that the features can be defined as macroscopic properties of the agent-environment system, these features can be high-level features, especially at a high enough level to be interpretable by humans in the context of the task. For example, the features may be related to the definition or achievement of goals at a high level, and / or the features may be expected to be meaningful components of some behavior patterns. The features can be captured directly from the environment through observations that are separate from the observations of the agent, such as where the properties of the environment can be directly obtained because they are in a simulated environment, and / or they can be derived from the observations of the agent, and / or they can be learned.
[0121] The variational sequential autoencoder processes a sequence of features on a time scale suitable for the desired behavior pattern. For example, if it is desired to characterize behavior on a time scale of one to several seconds, then the captured time series used to train the variational sequential autoencoder can have a similar time scale. The variational sequential autoencoder learns to encode or compress the sequence of features into a set of latent variables that represent the sequence. The latent variables can be constrained to have a Gaussian distribution; they can include a mean and a standard deviation vector. The recurrent neural network encoder and decoder can each include an LSTM (long short-term memory) neural network.
[0122] Thus, the set of latent variables of the variational sequential autoencoder can provide a latent representation of the observation sequence based on high-level features. A mixture model can be used to process the sequence of sets of latent variables derived from further observations to identify clusters. For example, a mixture model can map each set of latent variables to one of K components or clusters, which can be a Gaussian mixture model. Each component or cluster can correspond to a different temporally extended behavior derived from features representing high-level characteristics of the agent-environment system. These can be considered prototypical behaviors of the agent, each behavior being extended over the time range of the sequences used to train the autoencoder. Thus, processing a further sequence of observations ultimately results in determining, at each of a series of times, which one of the set of prototypical behaviors the agent is engaged in.
[0123] Optionally, these behaviors (more specifically, the clusters) can be represented in combination with the corresponding representations of the agent and / or the environment for some or all of the clusters, for example, graphically. The representation of a cluster can be, for example, a chart that has time on one axis and the identified behavior patterns on the same axis or another axis. The representation of the agent and / or the environment can be in any convenient form, such as an image and / or a set of features. There may be an association between the clusters and the corresponding states, for example, indicating that these clusters and corresponding states are combined with each other, or hovering over the representation of a cluster to provide a representation of the state.
[0124] Figure 6 An example of a user interface rendering 600 of agent behavior is shown. In particular, the user interface rendering 600 shows the behaviors that have been discovered using the above techniques. In particular, the rendering 600 shows the behavior clusters along axis 602 and the number of steps during the periods in which the behavior is engaged along axis 604. Thus, the height of the bar in the rendering 600 representing a given behavior represents the frequency with which the agent performs that behavior during that period. The system can allow the user to access more information about the behavior by hovering, selecting, or otherwise interacting with the bars in the rendering 600, for example, a representation of the state of the environment when the agent is engaged in that behavior.
[0125] Figure 7 Another example of a user interface rendering 700 of agent behavior is shown. In particular, the user interface rendering 700 is a time graph that shows the frequency with which specific behavior clusters are executed as the period of a task progresses, i.e., shows when certain behavior clusters are executed within a period (“period time”) in terms of time. The rendering 700 includes a respective row for each of 32 different behavior clusters, where a white bar represents a time period during which the agent is performing one or more behaviors within the cluster. For example, white bar 702 shows that the agent is performing a behavior within cluster 12 at a particular time period during the period, white bar 704 shows that the agent is performing a behavior within cluster 14 at a different, later time period during the period, and white bar 706 shows that the agent is performing a behavior within cluster 32 at yet another, even later time period during the period.
[0126] As with representation 600, the system can allow the user to access more information about the behavior by hovering, selecting, or otherwise interacting with the bars in the presentation 700, e.g., a representation of the state of the environment when the agent was involved in the behavior.
[0127] Although the subject technology has been described primarily in the context of a physical, real-world environment, it will be understood that the techniques described herein can also be utilized in the case of a non-real-world environment. For example, in some embodiments, the environment can be a simulated environment, and the agent can be implemented as one or more computers that interact with the simulated environment.
[0128] The simulated environment can be a motion simulation environment (e.g., a driving simulation or a flight simulation), and the agent can be a simulated vehicle that travels through the motion simulation. In these embodiments, the actions can be control inputs for controlling the simulated user or the simulated vehicle.
[0129] In another example, the simulated environment can be a video game, and the agent can be a simulated user who plays the video game. Generally, in the case of a simulated environment, the observations can include a simulated version of one or more of the previously described observations or types of observations, and the actions can include a simulated version of one or more of the previously described actions or types of actions.
[0130] This specification uses the term "configured" in connection with systems and computer program components. For a system of one or more computers, being configured to perform particular operations or actions means that the system has software, firmware, hardware, or a combination thereof installed on the system that, in operation, cause the system to perform those operations or actions. For one or more computer programs, being configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform those operations or actions.
[0131] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more of their combinations. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus.
[0132] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of devices, equipment, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The apparatus may also be or further include special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to hardware, the apparatus may optionally include code that creates an execution environment for the computer program, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0133] A computer program, which may also be referred to as or described as a program, software, software application, applet, module, software module, script, or code, can be written in any form of programming language, including compiled or interpreted languages or declarative or procedural languages; and the computer program can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), can be stored in a single file dedicated to the program, or can be stored in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). The computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0134] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or structured at all, and it can be stored on storage devices located at one or more locations. Thus, for example, an indexed database can include multiple collections of data, each of which can be organized and accessed differently.
[0135] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers located at one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and run on the same one or more computers.
[0136] The processes and logical flows described in this specification can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuitry (e.g., FPGA or ASIC), or by a combination of dedicated logic circuitry and one or more programmed computers.
[0137] A computer suitable for executing a computer program can be based on a general-purpose microprocessor, a special-purpose microprocessor, or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory, a random access memory, or both. The basic elements of a computer are a central processing unit for executing or running instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to, one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, from which it can receive data or to which it can transfer data, or both. However, a computer need not have such devices. Additionally, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a gaming console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few examples.
[0138] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; CD-ROM and DVD-ROM disks.
[0139] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including voice, speech, or tactile input. Additionally, a computer can interact with a user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on the user device in response to a request received from the web browser. Moreover, a computer can interact with a user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving, in return, response messages from the user.
[0140] The data processing apparatus for implementing a machine learning model can also include, for example, a dedicated hardware accelerator unit for processing common and computationally intensive portions of machine learning training or production, i.e., inference, workload.
[0141] Machine learning models can be implemented and deployed using machine learning frameworks (e.g., TensorFlow framework, Microsoft Cognitive Toolkit framework, Apache Singa framework, or Apache MXNet framework).
[0142] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computing device having a graphical user interface, a web browser, or an application through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0143] The computing system can include clients and servers. Clients and servers are typically remote from each other and typically interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML web page) to a user device, e.g., to display data to and receive user input from a user interacting with the device acting as a client. Data generated at the user device (e.g., the result of a user interaction) can be received at the server from the device.
[0144] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what is claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although the features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination can be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination.
[0145] Similarly, although operations are depicted in the drawings in a particular order and recited in the claims, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0146] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the acts recited in the claims can be performed in a different order and still obtain the desired result. As one example, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for training a neural network, the neural network having a plurality of network parameters and being used to select actions to be performed by an agent interacting with an environment to perform a reinforcement learning task, the method comprises: maintaining a plurality of candidate neural networks, and for each of the candidate neural networks, data specifying: (i) corresponding values of the network parameters of the candidate neural network and (ii) a corresponding reward mapping of the candidate neural network, wherein the reward mapping maps data extracted from the environment to a reward value; for each of the plurality of candidate neural networks, repeatedly performing the following training operations: using reinforcement learning to train the candidate neural network until a termination criterion is met to determine updated values of the network parameters of the candidate neural network according to the maintained network parameter values of the candidate neural network, wherein using reinforcement learning to train the candidate neural network includes generating, according to the maintained reward mapping of the candidate neural network, a reward value for reinforcement learning parameter update based on data extracted from the environment, determining a quality metric of the candidate neural network according to the updated values of the network parameters of the candidate neural network, wherein the quality metric measures the performance of the candidate neural network relative to at least one other candidate neural network in the reinforcement learning task, based on the quality metric of the candidate neural network, determining new network parameter values of the candidate neural network and a new reward mapping of the candidate neural network, and updating the maintained data of the candidate neural network to specify the new network parameter values and the new reward mapping; and after repeatedly performing the training operations, selecting the trained network parameter values from the parameter values in the maintained data.
2. The method according to claim 1, further comprises: providing the trained network parameter values for processing new inputs of the neural network.
3. The method according to claim 1, wherein, based on the maintained quality metric of the candidate neural network, selecting the trained network parameter values from the parameter values in the maintained data includes: after repeatedly performing the training operations, selecting the maintained parameter values of the candidate neural network having the best performance in the reinforcement learning task.
4. The method according to claim 1, wherein, repeatedly performing the training operations includes repeatedly performing the training operations in parallel for each candidate neural network.
5. The method according to claim 4, wherein, repeatedly performing the training operations includes repeatedly performing the training operations for each candidate neural network asynchronously with respect to performing the training operations for each other candidate neural network.
6. The method according to claim 5, wherein, determining the new reward mapping includes: in response to determining that the candidate neural network does not perform better than all one or more other candidate neural networks, modifying the maintained reward mapping of the best performing other candidate neural network; and setting the new reward mapping to the modified reward mapping.
7. The method according to claim 1, wherein, based on the quality metric of the candidate neural network, determining the new network parameter values of the candidate neural network includes: Determine whether the quality metric indicates that the candidate neural network performs better than all other candidate neural networks; and In response to determining that the candidate neural network does not perform better than all other candidate neural networks, set the new network parameter value to the maintained network parameter value of the other candidate neural network that performs best.
8. The method according to claim 7, further comprising: In response to determining that the quality metric does not perform better than other candidate neural networks, set the new network parameter value to the updated value of the network parameter.
9. The method according to claim 1, determining the new network parameter value of the candidate neural network based on the quality metric of the candidate neural network comprising: Determine whether the quality metric is worse than the quality metrics of other candidate neural networks by more than a threshold; and In response to determining that the candidate neural network is worse than the quality metrics of other candidate neural networks by more than a threshold, set the new network parameter value to the maintained network parameter value of the other candidate neural network with the best performance.
10. The method according to claim 9, further comprising: In response to determining that the quality metric is not worse than the quality metrics of other candidate neural networks by more than a threshold, set the new network parameter value to the updated value of the network parameter.
11. The method according to claim 9, wherein, Determining the new reward mapping includes: In response to determining that the quality metric is worse than the quality metrics of other candidate neural networks by more than a threshold, modify the maintained reward mapping of the other candidate neural network with the best performance; and Set the new reward mapping to the modified reward mapping.
12. The method according to claim 1, wherein, The extracted data includes the corresponding numerical values of each of one or more reward signals, and wherein each reward mapping: For each reward signal: Identify the partition of the numerical value of the reward signal to which the possible numerical values belong, and Map the partition to the reward value of the reward signal, and When there is more than one reward signal, combine the reward values of the reward signals to generate a reward.
13. The method according to claim 12, wherein, Modifying the maintained reward mapping includes, for each of one or more reward signals and for each partition of the possible values of the reward signal: Determine whether to modify the reward value mapped by the partition; and In response to determining to modify the reward value mapped by the partition, change the numerical value mapped by the partition.
14. The method according to claim 13, wherein, Determining whether to modify the reward value mapped by the partition includes determining to modify the reward value with a predetermined probability.
15. The method according to claim 12, wherein, One or more reward signals identify the position of the agent in the environment.
16. The method according to claim 12, wherein, One or more reward signals identify the distance of the agent relative to an object or position in the environment.
17. A system includes one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for training a neural network that has a plurality of network parameters and is configured to select actions to be performed by an agent interacting with an environment to perform a reinforcement learning task, the operations including: maintaining a plurality of candidate neural networks, and for each of the candidate neural networks, data specifying: (i) corresponding values of the network parameters of the candidate neural network and (ii) a corresponding reward mapping of the candidate neural network, where the reward mapping maps data extracted from the environment to a reward value; for each of the plurality of candidate neural networks, repeatedly performing the following training operations: using reinforcement learning to train the candidate neural network until a termination criterion is met to determine updated values of the network parameters of the candidate neural network based on the maintained network parameter values of the candidate neural network, where using reinforcement learning to train the candidate neural network includes generating, based on the maintained reward mapping of the candidate neural network, a reward value for reinforcement learning parameter updates according to data extracted from the environment, determining a quality metric of the candidate neural network based on the updated values of the network parameters of the candidate neural network, where the quality metric measures the performance of the candidate neural network relative to at least one other candidate neural network on the reinforcement learning task, based on the quality metric of the candidate neural network, determining new network parameter values of the candidate neural network and a new reward mapping of the candidate neural network, and updating the maintained data of the candidate neural network to specify the new network parameter values and the new reward mapping; and after repeatedly performing the training operations, selecting the trained network parameter values from the parameter values in the maintained data.
18. The method according to claim 17, wherein determining new network parameter values of the candidate neural network based on the quality metric of the candidate neural network includes: determining whether the quality metric indicates that the candidate neural network performs better than all other candidate neural networks; and in response to determining that the candidate neural network does not perform better than all other candidate neural networks, setting the new network parameter values to the maintained network parameter values of the other candidate neural network that performs best.
19. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations for training a neural network that has a plurality of network parameters and is configured to select actions to be performed by an agent interacting with an environment to perform a reinforcement learning task, the operations including: maintaining a plurality of candidate neural networks, and for each of the candidate neural networks, data specifying: (i) corresponding values of the network parameters of the candidate neural network and (ii) a corresponding reward mapping of the candidate neural network, where the reward mapping maps data extracted from the environment to a reward value; For each of the plurality of candidate neural networks, the following training operations are repeatedly performed: Use reinforcement learning to train the candidate neural network until a termination criterion is met to determine an updated value of the network parameters of the candidate neural network based on the network parameter values maintained by the candidate neural network, wherein training the candidate neural network using reinforcement learning includes generating a reward value for reinforcement learning parameter update according to data extracted from the environment based on the reward mapping maintained by the candidate neural network, Determine a quality metric of the candidate neural network based on the updated value of the network parameters of the candidate neural network, wherein the quality metric measures the performance of the candidate neural network relative to at least one other candidate neural network in a reinforcement learning task, Based on the quality metric of the candidate neural network, determine a new network parameter value of the candidate neural network and a new reward mapping of the candidate neural network, and Update the data maintained by the candidate neural network to specify the new network parameter value and the new reward mapping; and After repeatedly performing the training operations, select the trained network parameter value from the parameter values in the maintained data.
20. The one or more computer storage media according to claim 19, wherein, Based on the quality metric of the candidate neural network, determining a new network parameter value of the candidate neural network includes: Determine whether the quality metric indicates that the candidate neural network performs better than all other candidate neural networks; and In response to determining that the candidate neural network does not perform better than all other candidate neural networks, set the new network parameter value to the network parameter value maintained by the other candidate neural network that performs best.