Controlling agents using Q-TRANSFORMER neural networks

By using Transformer neural networks for Q learning and combining it with Monte Carlo returns, the problem of poor agent control performance in existing technologies is solved, and efficient training and improved generalization capabilities are achieved on data sets of varying quality.

CN120641914APending Publication Date: 2025-09-12GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480010528.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-03
Filing Date
2024-02-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing neural networks have difficulty effectively processing datasets of varying quality when controlling intelligent agents, and tend to overestimate the Q-values ​​of actions that are not well represented during training, resulting in poor performance in controlling the intelligent agent.

Method used

The Transformer neural network is used for Q learning, combined with offline dataset training and autoregressive technology, combined with Monte Carlo returns and conservative regularization, to improve the generalization ability and training efficiency of the policy system.

Benefits of technology

Through autoregressive Q-learning and conservative regularization, the control performance of the agent is improved, enabling it to better generalize to new tasks and accelerate learning progress even with uneven data quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120641914A_ABST
    Figure CN120641914A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for controlling agents that interact with an environment using Transform neural networks.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This specification relates to the use of neural networks to control intelligent agents.

[0002] A neural network is a machine learning model that uses one or more layers of nonlinear units to predict an output for a received input. In addition to the output layer, some neural networks also include one or more hidden layers. The output of each hidden layer is used as the input to the next layer in the network (i.e., the next hidden layer or output layer). Each layer of the network generates an output from the received input based on the current value input of the corresponding set of parameters. Summary of the Invention

[0003] This specification describes a system implemented as a computer program on one or more computers in one or more locations that controls an agent, such as a robot, that interacts in an environment by selecting actions for the agent to perform and then causing the agent to perform the actions.

[0004] The subject matter described in this specification can be implemented in specific embodiments to realize one or more of the following advantages.

[0005] This specification describes techniques for providing a scalable representation of Q-functions (i.e., functions that generate a Q-value for an action given a current observation and one or more previous observations). Specifically, by discretizing each action dimension of a given action and representing the Q-value for each action dimension as a separate token, a policy system can apply efficient, high-capacity sequence modeling techniques for Q-learning. Specifically, a Transformer neural network (also known as a "Q-Transformer neural network") can be used to autoregressively generate Q-values ​​for sub-actions along different action dimensions. By utilizing Transformer neural networks and generating Q-values ​​autoregressively, a policy system can control an intelligent agent, such as a robot, more efficiently than other approaches. In other words, the described techniques allow for improved control of robots, thereby improving the field of robotics. Furthermore, utilizing Transformer neural networks allows the system to efficiently incorporate natural language instructions into the neural network's inputs, thereby allowing the system to effectively control an intelligent agent to perform multiple different tasks (i.e., when the current task is specified via natural language instructions) using the same Transformer neural network while retraining the Transformer neural network.

[0006] In addition, this specification also describes techniques for improving the training of Transformer neural networks to further improve the performance of policy systems. For example, the system can train neural networks on large offline datasets collected from multiple different sources (e.g., both expert demonstrations and autonomously collected data) through offline Q-learning, even if the data from different sources are of mixed quality. "Mixed quality" data refers to a dataset that includes a large number of high-quality trajectories (i.e., trajectories that successfully perform the corresponding task and receive high rewards) and a large number of low-quality trajectories (i.e., trajectories where the agents interact randomly (so the rewards for any given task are low) or trajectories where the agents fail to perform the corresponding task). This allows the policy system to better generalize to new tasks after training.

[0007] As another example, the system can train a Transformer neural network via autoregressive Q-learning and can incorporate conservative regularization to prevent the policy system from overestimating the Q-values ​​of actions that are not well represented in the training dataset, thereby improving the performance of the system after training.

[0008] As another example, the system can incorporate Monte Carlo (MC) rewards into training to improve training efficiency. For example, by incorporating MC rewards, the system can accelerate learning progress during training, especially when the quality of the dataset is uneven.

[0009] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 An example policy system and an example control system are shown.

[0011] Figure 2 is a diagram of the architecture of an example policy system.

[0012] Figure 3 is a flow chart of an example process for controlling an intelligent agent that interacts with an environment.

[0013] Figure 4 is a flowchart of an example process for training a Transformer neural network included in a policy system.

[0014] Figure 5 is a diagram of training a Transformer neural network on experience tuples.

[0015] Figure 6 Quantitative examples of the performance gains that can be achieved by using the policy system described in this specification are shown.

[0016] Like reference numbers and designations throughout the various drawings refer to like elements. DETAILED DESCRIPTION

[0017] Figure 1 Shown are example policy system 100 and example control system 101. Policy system 100 and control system 101 are examples of a system implemented as a computer program on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.

[0018] Policy system 100 and control system 101 can control an agent 102 (e.g., a robot) to complete any of a wide variety of tasks in environment 104. To control an agent 102 that is interacting in environment 104 to complete a task, policy system 100 selects an action 144 for agent 102 to perform, and control system 101 then causes agent 102 to perform the selected action 144.

[0019] As general examples, tasks may include, for example, one or more of: navigating to a specified location in the environment 104 , identifying a specific object in the environment 104 , manipulating a specific object in a specified manner, controlling an item of equipment to meet a standard, etc. To complete such tasks, the agent 102 moves (e.g., navigates) within the environment 104 and / or changes its configuration.

[0020] Typically, the control system 101 is local to the agent 102. For example, the control system 101 may be on-board the agent 102, such as on one or more computers, local workstations, or local servers on the agent 102 with relatively small processing and memory resources.

[0021] In some implementations, policy system 100 is local to agent 102. For example, like control system 101, policy system 100 may be on-board agent 102. Furthermore, in some of these implementations, policy system 100 may be part of control system 101 that causes agent 102 to perform action 144.

[0022] In other implementations, policy system 100 is remote from agent 102. For example, unlike control system 101, policy system 100 may be hosted in a data center, which may be a distributed computing system with hundreds or thousands of computers in one or more locations. That is, control system 101 may receive data identifying action 144 from an external source, for example, rather than generating such data on-board agent 102.

[0023] In these implementations, policy system 100 and control system 101 may be connected via a data communication network, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof.

[0024] In these implementations, the control system 101 of the agent 102 interacts with a remote policy system 100 hosted in a data center that has significantly more computing and other resources than those available on the agent 102 machine to reduce latency in selecting an action 144, reduce consumption of the limited power supply of the agent 102 when selecting an action 144, or both.

[0025] In some implementations, policy system 100, control system 101, or both may expose one or more application programming interfaces (APIs) or other data interfaces that facilitate control of agent 102. For example, a user of agent 102 may use an API made available through action selection system 100 to provide a natural language text sequence 108 that characterizes a task for the agent to perform. As another example, policy system 100 and control system 101 may interact via an API between the two systems; for example, control system 101 may use an API to provide observation image 106 to policy system 100, and policy system 100 may use an API to provide data specifying a selected action 144 to control system 101.

[0026] Specifically, policy system 100 and control system 101 control the agent based on policy outputs generated by a Transformer neural network 140, which has been trained to control the agent in response to observations representing the environment and, optionally, natural language instructions 108 representing tasks to be performed by the agent. Natural language instructions 108 are sequences of natural language text that represent tasks to be performed by the agent 102 in the environment 104.

[0027] For example, the observations may each include one or more observation images 106 captured by a camera sensor of the agent 102 or by a camera sensor located in the environment 104. The camera sensor may be, for example, a still camera or a video camera.

[0028] Natural language text sequence 108 may be received from another agent in environment 104 or from control system 101 of agent 102. For example, another agent in environment 104 may speak an instruction, and control system 101 or another system may transcribe the instruction into natural language text sequence 108 and then provide the transcription to policy system 100. As another example, control system 101 may receive an instruction entered by a user (e.g., text-based input, selection-based input, or audio-based input) specifying natural language text sequence 108 and then provide the instruction to policy system 100.

[0029] More specifically, while controlling the agent, policy system 100 maintains historical data 120 representing observations that characterize the state of the environment at previous time steps. For example, at any given time step, historical data 120 may include data representing observations at the k most recent time steps, where k is an integer greater than or equal to one. The data representing an observation may include, for example, the observation itself or a set of one or more tokens that have been generated from the observation.

[0030] At each time step, policy system 100 obtains a current observation that characterizes the state of environment 104 at the time step.

[0031] A lemma system 130 within the system 100 generates an input sequence 132 of input lemmas from at least the current observation and the observations represented in the historical data 120. Specifically, the system “lemma”s the current observation and the observations in the historical data 120 so that each observation is represented as one or more lemmas, and includes these lemmas in the input sequence 132.

[0032] The system processes the input sequence 132 using a Transformer neural network 140 to select an action 144 for the agent 102 to perform in response to the current observation.

[0033] In general, an action 144 includes a corresponding sub-action for each of a plurality of action dimensions in an action dimension sequence. The action dimension sequence will also be referred to as a "dimension sequence" to distinguish it from the input sequence of the Transformer.

[0034] More specifically, each action that the agent 102 can perform includes corresponding sub-actions for each of the multiple action dimensions. For example, different action dimensions can correspond to different controls of the agent, different coordinates of a given control of the agent, or some combination. For example, the agent can be controlled by specifying the 3D positioning of the agent or a controllable element of the agent. In this example, the 3D positioning can be represented as three action dimensions, with each of these 3D coordinates corresponding to one action dimension. As another example, the agent can be controlled by specifying the 3D orientation of the agent or a controllable element of the agent. In this example, the 3D orientation can be represented as three action dimensions, with each of these 3D coordinates corresponding to one action dimension. As another example, the agent can be controlled by specifying the closure of the agent's gripper. In this example, gripper closure can be specified as a single dimension. In some cases, one or more of these action dimensions can correspond to additional actions, such as a "no action" action in which no input is provided to the agent, a "terminate" action that terminates the current task episode, and so on.

[0035] When some or all of the action dimensions have continuous sub-actions, the system can discretize the continuous space of sub-actions for each continuous action dimension to generate a candidate set of sub-actions for the action dimension. Discretizing the continuous space allows the Transformer neural network 140 to effectively model the action dimensions.

[0036] In some implementations, the system 100 represents each action dimension using a candidate set with the same number of sub-actions (i.e., the same number of "bins"). For example, for a continuous action dimension, the system 100 may discretize all continuous action dimensions into the same number of bins. For discrete action dimensions, the system 100 may represent these sub-actions using a fixed number of sub-actions through padding. Using the same number of sub-actions allows the Transformer neural network 140 to efficiently model each action dimension.

[0037] Thus, for each action dimension in the plurality of action dimensions, processing includes processing a combined sequence using a Transformer neural network 140, the combined sequence comprising the input sequence followed by a sub-action for each previous action dimension that precedes the action dimension in the sequence of dimensions, the sub-action selected for the action dimension, to generate a corresponding Q-value 142 for each sub-action in the set of candidate sub-actions for the action dimension.

[0038] Thus, for the first action dimension, the combined sequence includes only the input sequence 132 , while for each subsequent action dimension, the combined sequence includes the input sequence followed by one or more preceding sub-actions.

[0039] The system 100 then uses the corresponding Q-values ​​of the sub-actions in the set of candidate sub-actions for the action dimension to select a sub-action for the action dimension.

[0040] That is, the system 100 autoregressively selects sub-actions for action dimensions according to the dimension sequence, such that the sub-actions for each action dimension depend on the sub-actions for the previous action dimension in the dimension sequence.

[0041] To select a sub-action for the first action dimension in a sequence, the system 100 processes the input sequence using the Transformer neural network 140 to generate a corresponding Q-value for each sub-action in a set of candidate sub-actions for the first action dimension, and then uses the corresponding Q-values ​​of the candidate sub-actions to select a sub-action for the action dimension.

[0042] At a given time step, the Q-value of a given candidate sub-action for a given action dimension represents an estimate of the reward of executing the corresponding action at the given time step (and continuing to select actions using the Transformer neural network 140 at subsequent time steps).

[0043] In general, at any given time step, the reward to be received is a combination of the rewards to be received at the time step after the given time step in the task round or at a predetermined number of time steps after the given time step. For example, at time step t, the reward may satisfy:

[0044] ,

[0045] where i ranges from all time steps after t in the episode or a fixed number of time steps after t in the episode, is the discount factor, and is the reward at time step i. As can be seen from the above formula, the higher the value of the discount factor, the longer the time range of the reward calculation, that is, the rewards of time steps that are farther away from time step t in time are given greater weight in the reward calculation.

[0046] Generally speaking, rewards are scalar values ​​and represent the agent's progress towards completing a task.

[0047] As a specific example, the reward may be a sparse binary reward that is zero unless the task is successfully completed and one if the task is successfully completed due to the action being performed.

[0048] As another specific example, the reward can be a dense reward that measures the agent's progress towards completing the task as each observation is received during an episode of attempting to perform the task, i.e., such that a non-zero reward can be and often is received before the task is successfully completed.

[0049] For any given candidate sub-action for the last action dimension, the corresponding action is an action that includes: the given candidate sub-action for the last action dimension; and, for each previous action dimension in the sequence of dimensions, the sub-action that has been selected for that action dimension.

[0050] For any given candidate sub-action for any given action dimension other than the last action dimension, the corresponding action is an action that includes: (i) the sub-action that has been selected for any previous action dimension in the dimension sequence; (ii) for a given action dimension, the given candidate sub-action; and (iii) for each subsequent action dimension in the dimension sequence, the sub-action that would be selected for the given action dimension by autoregressively selecting sub-actions using the Transformer neural network 140, assuming that the given candidate sub-action was selected for the given action dimension, that is, the sub-action that would be selected by continuing to select actions using the above-mentioned autoregressive process, assuming that the candidate sub-action was selected for the given action dimension.

[0051] After having selected action 144 to be performed by agent 102 at the time step, policy system 100 provides data identifying selected action 144 to control system 101. In implementations where policy system 100 is remote from agent 102, providing data identifying selected action 144 may, for example, include transmitting the data identifying selected action 144 over a data communications network connecting policy system 100 and control system 101.

[0052] The control system 101 then causes the agent 102 to perform the selected action 144. For example, the control system 101 may do this by generating instructions to the agent 102 that, when executed, will cause the agent 102 to perform the selected action 144, by submitting control inputs directly to appropriate controls of the agent, or by using another appropriate control technique.

[0053] In some implementations, the environment 104 is a real-world environment, and the agent 102 is a mechanical agent that interacts with the real-world environment. For example, the agent can be a robot that interacts with the environment to accomplish a goal, such as locating an object of interest in the environment, moving an object of interest to a specified location in the environment, physically manipulating an object of interest in a specified manner in the environment, or navigating to a specified destination in the environment; or the agent can be an autonomous or semi-autonomous land, air, or sea vehicle that navigates through the environment to reach a specified destination in the environment.

[0054] Action 144 can be a control input to control a robot (e.g., torque to a joint of the robot, or a higher-level control command), or a control input to control an autonomous or semi-autonomous land, air, or sea vehicle (e.g., torque to a control surface or other control element of the vehicle, or a higher-level control command).

[0055] In other words, actions 144 may include, for example, positioning, velocity, or force / torque / acceleration data of one or more joints of a robot or a component of another mechanical agent. Actions may additionally or alternatively include electronic control data, such as motor control data, or more generally data for controlling one or more electronic devices within an environment, the control of which has an effect on the observed state of the environment. For example, in the case of an autonomous or semi-autonomous land, air, or sea vehicle, actions may include actions to control the navigation (e.g., steering) and movement (e.g., braking and / or acceleration) of the vehicle.

[0056] In some implementations, the environment 104 is a simulated environment, and the agent 102 is implemented as one or more computer programs that interact with the simulated environment. For example, the environment may be a computer simulation of a real-world environment, and the agent may be a simulated mechanical agent that navigates through the computer simulation.

[0057] For example, the simulated environment can be a motion simulation environment (e.g., a driving simulation or a flight simulation), and the agent can be a simulated vehicle navigating through the motion simulation. In these implementations, actions 144 can be control inputs used to control the simulated user or the simulated vehicle. As another example, the simulated environment can be a computer simulation of a real-world environment, and the agent can be a simulated robot interacting with the computer simulation.

[0058] Typically, when the environment 104 is a simulated environment, the actions 144 may include simulated versions of one or more of the previously described actions or action types.

[0059] In some implementations, environment 104 is a suitable execution environment (e.g., a runtime environment or operating system environment) implemented on one or more computing devices (such as a smartphone, tablet computer, wearable device, automotive system, stand-alone personal assistant device, etc.), and agent 102 is a virtual agent (also known as an "automated assistant" or "mobile assistant") that a user can interact with via the computing device. The virtual agent can receive input from the user (e.g., typed or spoken natural language input) and respond with responsive content (e.g., visual and / or audible natural language output). The virtual agent can provide a wide range of functionality by interacting with various local and / or third-party applications, websites, or other agents. In these implementations, actions 144 can include any activity or operation that can be performed or initiated by a user on a computing device (e.g., within application software installed on the computing device).

[0060] In some cases, policy system 100 may be used to control an agent's interaction with a simulated environment, and policy system 100 (or another training system) may train the set of neural networks used to control agent 102 based on agent 102's (or another agent's) interaction with the simulated environment to determine trained values ​​for parameters of the set of neural networks.

[0061] After the set of neural networks is trained based on the interaction of agent 102 (or another agent) with the simulated environment, the set of trained neural networks can be used by policy system 100 to control the interaction of a real-world agent with the real-world environment, i.e., to control an agent simulated in a simulated environment.

[0062] Training a neural network based on an agent's interactions with a simulated environment (i.e., rather than a real-world environment) avoids wear and tear on the agent and reduces the likelihood that the agent may harm aspects of itself or its environment by performing poorly chosen actions.

[0063] More generally, the following references Figure 4 and Figure 5 The training of the Transformer neural network 140 and, optionally, other neural networks used to generate input sequences for the Transformer neural network 140 is described in detail.

[0064] Figure 2 is a diagram of the architecture of an example policy system 200 .

[0065] The policy system 200 receives a natural language text sequence 208. The natural language text sequence 208 represents a task to be performed by the agent in the environment. The natural language text sequence 208 may have an instruction format. For example, Figure 2It is shown that the natural language text sequence 208 is a natural language instruction describing the task "Pick sponge . . . ".

[0066] The policy system 200 processes the natural language text sequence 208 using a text encoder neural network 210 to generate an encoded representation 212 of the natural language text sequence.

[0067] exist Figure 2 In the example, the text encoder neural network 210 has a universal sentence encoder architecture and generates the encoded representation 212 as a single vector, i.e., a vector comprising a fixed number (e.g., 256, 512, or 1024) of entries, where each entry is a numeric value, e.g., a floating point value. The universal sentence encoder is described in more detail in Daniel Cer et al., Universal sentence encoder. arXiv preprint arXiv:1803.11175, 2018.

[0068] In other examples, the text encoder neural network 210 can have a different architecture and can generate embeddings with smaller or larger dimensions. Additionally, in other examples, the text encoder neural network 210 can generate embeddings that include a sequence of multiple embedding vectors.

[0069] At each of a plurality of time steps, policy system 200 obtains an observation image 206 that characterizes the state of the environment at that time step. Figure 2 In the example of , the agent performs a single action in response to each observed image 206, eg, such that a new observed image is obtained by the policy system 200 after each action performed by the agent.

[0070] As mentioned above, the policy system 100 also maintains historical data 120, which represents observations that characterize the state of the environment at previous time steps. Figure 2 In the example of , the system 100 stores the observed images 207 at the previous two time steps in the historical data 120 .

[0071] Policy system 100 uses image encoder neural network 220 to generate encoded representation 222 of observed image 206 and observed image 207 in historical data. Encoded representation 222 may include a feature map including a corresponding feature vector for each of a plurality of regions in observed image 206.

[0072] In some implementations, the image encoder neural network 220 may generally be configured as a convolutional neural network comprising one or more convolutional layers. As a specific example of this, Figure 2The image encoder neural network 220 is shown as a convolutional neural network with the EfficientNet architecture, which includes a stack of reverse residual blocks ("MBConv blocks"). Reverse residual blocks and EfficientNet are described in more detail in Mingxing Tan et al., EfficientNet: Rethinking model scaling for convolutional neural networks, in Kamalika Chaudhuri and Ruslan Salakhutdinov (eds.), Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 6105–6114, PMLR, June 9–15, 2019. URL https: / / proceedings.mlr.press / v97 / tan19a.html.

[0073] More generally, image encoder neural network 220 may have any suitable neural network architecture, such as a convolutional neural network architecture or a visual Transformer neural network architecture.

[0074] In addition, Figure 2 In the example of , the image encoder neural network 220 generates the encoded representation 222 conditioned on the encoded representation 212 of the natural language text sequence. That is, the image encoder neural network 220 receives the encoded representation 212 of the natural language text sequence and the observed image 206 as input, and processes the input to generate the encoded representation 222 of the observed image as output.

[0075] The image encoder neural network 220 uses the encoded representation 212 of the natural language text sequence as context when generating the encoded representation 222 of the observed image, i.e., such that different text sequences may result in different representations being generated for the same observed image.

[0076] To this end, the image encoder neural network 220 further includes one or more conditioning layers. The conditioning layers may be interspersed between other intermediate layers of the image encoder neural network 220 (e.g., convolutional layers (e.g., depth-wise convolutional layers), attention layers, etc.).

[0077] Each conditioning layer receives as input (i) a corresponding intermediate output of a corresponding intermediate layer of the image encoder neural network and (ii) an encoded representation 212 of a natural language text sequence, and processes the input to (i) update the corresponding intermediate output of the image encoder neural network using the encoded representation 212 of the natural language instruction and (ii) provide the updated corresponding intermediate output as input to a corresponding subsequent intermediate layer of the image encoder neural network.

[0078] As a specific example of this, Figure 2 The image encoder neural network 220 is shown to include feature-wise linear modulation (FiLM) layers interspersed between stacks of inverse residual blocks. The FiLM layer learns the function and , these functions output and As input Function:

[0079]

[0080] in and Modulate the corresponding intermediate output of the corresponding intermediate layer through a feature-by-feature affine transformation :

[0081] .

[0082] function and It can, but need not, be implemented as a neural network, such as a multilayer perceptron (MLP) or a convolutional neural network.

[0083] The FiLM layer is described in more detail in Ethan Perez et al., Film: Visual reasoning with a general conditioning layer, Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), April 2018, doi: 10.1609 / aaai.v32i1.11671.

[0084] For example, the image encoder neural network 220 may include a FiLM layer arranged between a first reverse residual block and a second reverse residual block in a stack. The FiLM layer receives as input (i) the corresponding intermediate output of the first reverse residual block and (ii) the encoded representation 212 of the natural language text sequence, and processes the input to (i) update the corresponding intermediate output of the first reverse residual block using the encoded representation 212 of the natural language instruction and (ii) provide the updated corresponding intermediate output as input to the second reverse residual block. Thus, instead of receiving as input the corresponding intermediate output of the first reverse residual block, the second reverse residual block receives the updated corresponding intermediate output that has been updated by the FiLM layer using the encoded representation 212 of the natural language text sequence.

[0085] Policy system 200 generates input word-gram sequence 232 based on encoded representation 222 of observed image.

[0086] As described above, encoded representation 222 may include a feature map including a respective feature vector for each of a plurality of regions in observed image 206 , or a respective feature map for observed image 206 and for each observed image 207 from historical data 120 .

[0087] In this example, the policy system 200 generates an input sequence by flattening the feature map into a sequence of feature vectors.

[0088] In some other examples, the policy system 200 can generate an initial input sequence by flattening the feature map into a sequence of feature vectors, and then process the initial input sequence of feature vectors using a learned module that maps the input sequence into a reduced sequence comprising a smaller number of tokens. TokenLearner is described in more detail in: Michael Ryoo et al., Tokenlearner: Adaptive space-time tokenization for videos, Advances in Neural Information Processing Systems, 34:12786–12797, 2021.

[0089] When used, the learned module can be, but need not be, a neural network. When configured as a neural network, the learned module can include any suitable type of neural network layers (e.g., convolutional layers, fully connected layers, attention layers, pooling layers, etc.) connected in any suitable number (e.g., 1 layer, or 5 layers, or 10 layers) and in any suitable configuration (e.g., as a directed graph of layers).

[0090] As a specific example, the learned module can be a TokenLearner neural network module. TokenLearner is described in more detail in Michael Ryoo et al., Tokenlearner: Adaptive space-time tokenization for videos, Advances in Neural Information Processing Systems, 34:12786–12797, 2021. As another example, the word unit neural network 130 can have a visual transformer (ViT) architecture including one or more attention layers. As another example, the word unit neural network 130 can have a convolutional neural network architecture including one or more convolutional layers.

[0091] In some implementations, to speed up inference by avoiding repeated computation, the policy system 200 may store feature vectors included in the input sequence generated at each time step and reuse them at subsequent time steps. That is, the system may store feature vectors of historical observation images at each time step and then use all or part of the stored feature vectors during generation.

[0092] also, Figure 2 The strategy system 200 is shown to adopt a positional encoding scheme. Specifically, the strategy system 200 adds a corresponding positional encoding 233 to each word in the input sequence 232 of word-grams.

[0093] The positional code 233 may be determined, for example, according to a sinusoidal positional coding scheme or another coding scheme, so as to uniquely identify, for each image word-mem, the corresponding time step among a plurality of time steps at which the image word-mem was generated.

[0094] The policy system 200 then provides the input word sequence 232 as input to the Transformer neural network 240. As a specific example, Figure 2 The Transformer neural network 240 is shown having a decoder-only Transformer neural network architecture including a plurality (e.g., 4, 8, 16, or other suitable number) of self-attention layer blocks. In this example, each self-attention layer block applies a causal self-attention mechanism. For example, each self-attention layer block can use an attention mask that zeros out the attention contribution from future tokens and optionally also zeros out the action from the previous time step. Zeroing out the action from the previous time step can allow the neural network to be trained in parallel on a trajectory that includes transitions from multiple time steps.

[0095] In other examples, the Transformer neural network 240 may have a different Transformer-based architecture, such as an encoder-decoder Transformer neural network architecture that includes more or fewer layers with each layer having the same or different attention mechanisms.

[0096] Specifically, the policy system 200 uses the Transformer neural network 240 to autoregressively select an action for the agent to perform in response to the current observation 206.

[0097] Specifically, for each action dimension, the system uses a Transformer neural network 240 to process a combined sequence that includes an input sequence followed by a corresponding word element for each previous action dimension that precedes the action dimension in the dimension sequence, the corresponding word element identifying the sub-action selected for the action dimension, so as to generate a corresponding Q-value 242 for each sub-action in the candidate sub-action set for the action dimension, and then use the corresponding Q-values ​​of the sub-actions in the candidate sub-action set for the action dimension to select the sub-action for the action dimension.

[0098] exist Figure 2 In the example of , system 200 generates a combined sequence for a given action dimension by concatenating additional tokens ("action embeddings") 250 to the token input sequence for the first action dimension, or to the combined sequence for the previous action dimension for each subsequent action dimension. The system can generate additional tokens by applying a learned embedding function to data identifying the selected sub-action, where the learned embedding function is learned as part of training the Transformer neural network.

[0099] Specifically, to generate Q-values ​​for the sub-actions in the sub-action set for the action dimension, the system 200 processes the combined sequence using a self-attention layer, and then processes the output of the last self-attention layer using a sigmoid function 242 to generate a corresponding Q-value 244 for each sub-action (“action interval”), and then selects the argmax sub-action 246 based on the corresponding Q-value. That is, the system selects the candidate sub-action with the highest corresponding Q-value.

[0100] The system then generates a one-hot encoding 248 of the selected sub-action and uses the one-hot encoding to generate an action embedding that will be used as part of the combined sequence for the next action dimension.

[0101] Although Figure 2An example of generating a word-gram input sequence using an image encoder neural network conditioned on a natural language text sequence using a conditioning layer is shown, but more generally, the system can generate a word-gram input sequence in any of a variety of ways, from at least the current observation and observations represented in historical data (and optionally natural language instructions).

[0102] For example, when no natural language sequence is available, the system can generate an input sequence by processing the current observation and observations from historical data using an appropriate encoder neural network.

[0103] As another example, when there is a natural language sequence, the system can use an appropriate encoder neural network to process the current observation and observations in historical data to generate a first subsequence, and use another encoder neural network to process the natural language text sequence to generate a second subsequence, and then concatenate the two subsequences to generate the input sequence.

[0104] As described above, for the last action dimension in the dimension sequence, the corresponding Q-value for each sub-action in the set of candidate sub-actions for that action dimension represents an estimate of the reward that the agent will receive in response to performing an action that includes: the candidate sub-actions for the last action dimension; and, for each previous action dimension in the dimension sequence, the sub-actions that have been selected for that action dimension.

[0105] Similarly, for each given candidate sub-action for each given action dimension except the last action dimension, the corresponding Q-value for the given candidate sub-action represents an estimate of the reward that would be received in response to the agent performing an action comprising: (i) the sub-actions that would have been selected for the given action dimension for any previous action dimension preceding the given action dimension in the sequence of dimensions; (ii) the candidate sub-actions for the given action dimension; and (iii) the sub-actions that would have been selected for the given action dimension by autoregressively selecting sub-actions using a Transformer neural network, given that the given candidate sub-action was selected for the given action dimension.

[0106] Figure 3 is a flow chart of an example process 300 for controlling an agent interacting with an environment. For convenience, process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a policy system (e.g., Figure 1 The policy system 100) can execute process 300.

[0107] The system controls the agent to complete a task in the environment by repeatedly executing iterations of process 300 at each of a plurality of time steps (hereinafter referred to as the "current" time step).

[0108] Alternatively, the task to be performed by the agent can be represented by a natural language text sequence. For example, prior to the first iteration of process 300, the system receives a natural language text sequence representing the task to be performed by the agent in the environment and generates an encoded representation of the natural language text sequence. For example, the encoded representation includes an embedding of the natural language text sequence, and the system can generate the embedding by processing the natural language text sequence using a text encoder neural network.

[0109] The system maintains historical data representing observations characterizing the state of the environment at previous time steps (step 302).

[0110] The system obtains an observation that characterizes the state of the environment at the current time step (step 304).

[0111] The system generates an input sequence of input tokens from at least the current observation and the observations represented in the historical data (step 306).When the system also receives a natural language text sequence, the system can also generate the input sequence from an embedding of the natural language text sequence.

[0112] The system processes the input sequence of input tokens using a Transformer neural network to select an action for the agent to perform in response to the current observation (step 308).

[0113] As described above, the action includes a corresponding sub-action for each of a plurality of action dimensions in a dimensional sequence of action dimensions. To select an action, the system autoregressively selects a corresponding sub-action for the action dimension in the dimensional sequence. That is, for each action dimension in the plurality of action dimensions, the system uses a Transformer neural network to process a combined sequence comprising an input sequence followed by a corresponding word element for each previous action dimension preceding the action dimension in the dimensional sequence, the corresponding word element identifying the sub-action selected for the action dimension so as to generate a corresponding Q-value for each sub-action in a set of candidate sub-actions for the action dimension, and then selects a sub-action for the action dimension using the corresponding Q-values ​​of the sub-actions in the set of candidate sub-actions for the action dimension.

[0114] The system then causes the agent to perform the selected action (step 310), for example, by submitting control input directly to the agent or by transmitting instructions or other data, for example, via a data communications network, to the agent's control system that will cause the agent to perform the selected action.

[0115] When controlling an agent to perform a task in which the action that should be performed (e.g., the action that will result in progress toward completing the task) is unknown, some or all of the steps of process 300 may be performed. The steps of process 300 may also be performed as part of selecting an action to be performed by the agent based on processing observation images derived from a set of training data sets (i.e., observation images in response to which the action that should be performed by the agent is known in order to train the set of neural networks to determine trained values ​​for the parameters of the set of neural networks).

[0116] Figure 4 is a flow chart of an example process 400 for training a set of neural networks included in a policy system. For convenience, process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a policy system (e.g., Figure 1 Process 400 may be performed by a strategy system 100) or another training system.

[0117] Typically, the system trains a Transformer neural network on a training dataset consisting of trajectories of experience tuples.

[0118] In some implementations, the system also trains one or more of the other neural networks used by the policy system, such as one or more of an image encoder neural network, a text encoder neural network, or a learned module that reduces the number of tokens in a sequence, jointly with the training of the Transformer neural network. For any neural network not trained jointly with the training of the Transformer neural network, the system can leave the neural network unchanged during training.

[0119] Regardless of whether the other neural networks are jointly trained with the Transformer, some or all of the other neural networks can be pre-trained before training the Transformer neural network.

[0120] In some cases, the text encoder neural network can be pre-trained on a text processing task, such as a text representation learning task, prior to joint training of the set of neural networks, e.g., as part of a larger text processing neural network. In some of these cases, the text encoder neural network is then fine-tuned during the joint training, while in other of these cases, the text encoder neural network is kept frozen during the joint training, i.e., the joint training of the neural networks on the training dataset does not adjust the pre-trained parameter values ​​of the pre-trained text encoder neural network.

[0121] In some cases, an image encoder neural network may be pre-trained on an image processing task such as an image classification or segmentation task, for example as part of a larger image processing neural network, and then fine-tuned during joint training of this set of neural networks on a training dataset.

[0122] In some of these cases, the image encoder neural network does not include one or more conditioning layers used for pre-training (the one or more conditioning layers are added after pre-training). For example, the image encoder neural network may be trained as part of a neural network that is trained to classify images into a set of categories without being conditioned on any natural language text sequence as context.

[0123] Furthermore, in some of these cases, because inserting a conditioning layer as a new layer into such a pre-trained image encoder neural network may result in perturbations to the intermediate outputs of the neural network, prior to joint training, the system initializes each conditioning layer to act as an identity transformation for the corresponding intermediate outputs. This can be achieved, for example, by setting at least some of the parameter values ​​associated with each conditioning layer to zero.

[0124] In some cases, the training dataset includes expert interaction data representing the interactions of one or more expert agents with a corresponding environment. An expert agent can be any agent that, in response to observing an image, selects an action according to an action selection policy that enables the expert agent to make effective progress toward completing a task. For example, the expert agent can be an agent controlled by another trained policy system, a human skilled in the task to be performed by the agent, and so on.

[0125] In some of these cases, the expert interaction data includes simulated data, wherein a simulated expert agent performs one or more tasks in a simulated environment. In other of these cases, the expert interaction data includes real-world data, wherein a real-world expert agent performs one or more tasks in a real-world environment. In still other cases, the expert interaction data includes both simulated data and real-world data.

[0126] In some cases, the training dataset includes training examples generated from (and therefore having the same physical properties) one or more robots of the same model when they perform the same or different tasks, such as one of the tasks mentioned in Table 1 above, or other tasks.

[0127] In other cases, the training dataset may be a hybrid training dataset that includes training examples generated from multiple robots that are not of the same model, located at the same site, or even manufactured by the same manufacturer. For example, the hybrid training data may be generated from dozens or hundreds of different robots that have different physical characteristics and are different models. Furthermore, the hybrid training data need not be generated from physical robots. For example, the hybrid training data may include data generated from simulations of physical robots.

[0128] The system obtains an experience tuple (step 402).

[0129] An experience tuple includes (i) data representing a historical set of observations, (ii) a training observation, (iii) a training action performed in response to the training observation, (iv) a reward received in response to the training action being performed, and (v) the next observation received in response to the training action being performed.

[0130] For example, when training a Transformer neural network using offline Q-learning, training actions can be performed by a different agent (e.g., an expert agent or another agent being trained) in response to training observations.

[0131] Typically, experience tuples are extracted from a trajectory of experience tuples, which consists of a sequence of multiple experience tuples generated when the agent interacts with the environment.

[0132] The system then trains a Transformer neural network on the experience tuples.

[0133] As part of training the Transformer neural network, the system generates a corresponding Q-value for each candidate sub-action for each action dimension using the Transformer neural network and based on the current values ​​of the parameters of the Transformer neural network, given the training observations and the historical observations (step 404).

[0134] Specifically, due to the causal masking of the self-attention layer block within the Transformer, the system can generate corresponding Q-values ​​for all candidate sub-actions for all action dimensions in parallel (i.e., in one forward propagation of the Transformer neural network).

[0135] The system then generates, for each action dimension, corresponding target Q-values ​​for the sub-actions in the training action for the action dimension using the rewards in the experience tuple (step 406).

[0136] Typically, the system can determine the corresponding target Q-value for each action dimension in the action dimension by applying autoregressive Q-objective maximization using the corresponding input sequence. Applying autoregressive Q-objective maximization means that for a given input sequence, a Transformer neural network is used to process the input sequence so as to generate a corresponding Q-value for each sub-action.

[0137] Specifically, for the last action dimension, the system determines a first target Q-value by applying an autoregressive Q-objective maximization to an input sequence generated from at least (i) one or more of the historical observations, (ii) the training observations, and (iii) the next observation.

[0138] For example, applying an autoregressive Q-objective maximization to an input sequence may include selecting an action by processing the input sequence using a Transformer neural network as described above to set parameter values ​​to target values ​​(e.g., values ​​constrained to change more slowly than current values ​​during training). For example, the target value may be maintained as an exponential moving average (EMA) of actual training values ​​throughout training.

[0139] Thus, for the last action dimension, the first target Q-value may be equal to the Q-value assigned to the chosen action (optionally multiplied by a discount factor) plus the reward in the tuple.

[0140] For any given action dimension except the last action dimension, the system determines the maximum Q-value assigned to any sub-action in the next action dimension by processing the input sequence using a Transformer neural network, the input sequence including corresponding sub-actions from training actions for any previous action dimension in the sequence of dimensions and for the given action dimension, and uses this as the first target Q-value. For example, the system can determine this maximum Q-value based on the target values ​​of the network parameters.

[0141] As a specific example, given the current observation and historical observations within the time steps between time step t and time step t-w, the first target Q-value of action dimension i at time step t can be calculated as:

[0142]

[0143] in refers to the sub-action for action dimension i+1 at time step t, Refers to the training action for action dimensions 1 to i at time step t The sub-actions in refers to the sub-action for action dimension 1 at time step t+1, is the total number of action dimensions, is the discount factor, is the reward at time step t (also called R t , i.e., the reward in the experience tuple, is the set of observations from time step tw to time step t, is the set of observations from time step t-w+1 to time step t+1, Transformer neural network is a process that processes The generated input sequence is The Q value generated by processing the generated input sequence, and Transformer neural network processes The generated input sequence is The generated Q value.

[0144] In some implementations, the system uses the first target Q-value as the target Q-value for the action dimension.

[0145] In some other implementations, the system further determines the second target Q-value as a Monte Carlo return starting from the rewards in the tuple. The Monte Carlo return is the time-discounted sum of rewards received during the trajectory to which the experience tuple belongs, starting from the rewards in the tuple.

[0146] The system can then generate a target Q-value for the dimension as the maximum of the first target Q-value and the second target Q-value.

[0147] The system trains a Transformer neural network based on a target that, for each action dimension, measures the time difference error between the corresponding Q-values ​​of the sub-actions in the training actions for that dimension and the target Q-values ​​of the sub-actions in the training actions for that dimension (step 408).

[0148] In some implementations, for each action dimension, this objective encourages sub-actions that are not in the training actions for that dimension to have corresponding Q-values ​​equal to zero. For example, when training a Transformer neural network using offline learning, including this "conservative" term can improve training by preventing the Transformer neural network from overestimating the Q-values ​​of actions that are not present (or appear infrequently) in the training dataset.

[0149] For example, for each action dimension, the objective may include a term that measures the square of the Q-value assigned to at least one of the sub-actions that is not in the training actions for that dimension.

[0150] As a specific example of this, the system can sample actions based on the probabilities generated using the behavioral policy used to generate the experience tuples, where actions that the behavioral policy assigns a higher likelihood are assigned a lower probability. Then, for each action dimension of the selected actions, a "conservative" term can measure the square of the Q-values ​​of the sub-actions in the sampled actions assigned to that dimension (effectively encouraging the Q-value to be zero).

[0151] As another example of this, the system can approximate the behavioral policy by taking the average, for each action dimension, of the squares of the Q-values ​​assigned to all sub-actions in the training actions that are not in that dimension.

[0152] As described above, in some implementations, the other neural networks used to generate the input sequence for the Transformer neural network are pre-trained and remain unchanged during the training of the Transformer neural network. In some other implementations, the system also trains one or more neural networks in the neural network, for example, by backpropagating the gradient of the target through the Transformer neural network.

[0153] Although Figure 4 The system is described as training a Transformer neural network on a single tuple, but in practice, the system may train a Transformer neural network on a set of multiple experience tuples (e.g., multiple experience tuples from the same trajectory, or different tuples sampled from different trajectories) in each iteration of process 400. In this case, the overall objective may be the average or weighted average of the objectives for each tuple.

[0154] Figure 5 An example 500 of training a Transformer neural network on an experience tuple at time step t, which includes data representing a set of historical observations and a current observation (collectively denoted in the figure as ).

[0155] For each action dimension, the sub-actions included in the training action in the experience tuple are marked with solid black boxes in the figure.

[0156] As described above, for each sub-action included in the training action, the target measures the time difference error between the corresponding Q value of the sub-action in the training action for that dimension and the corresponding target Q value ("Q target") of the sub-action in the training action for that dimension.

[0157] like Figure 5 As shown, the system uses conservative Q-updates for each action dimension, where for each action dimension, the objective encourages sub-actions that are not in the training actions for that dimension (i.e., including “ ” marked box) has a corresponding Q value equal to zero.

[0158] As shown in the figure, to calculate the target Q value for a given action dimension, the system determines the first target Q value by applying the autoregressive Q objective maximization using the corresponding input sequence. Figure 4 The description is carried out, where for the last action dimension (using R t ) is calculated differently from the previous action dimension.

[0159] The system also determines the second target Q value as the Monte Carlo return (MC) starting from the reward in the tuple t:T ) 540. As described above, the Monte Carlo reward is the time-discounted sum of the rewards received during the trajectory to which the experience tuple belongs, starting from the reward in the tuple (discounted by the discount factor γ as described above). For example, the MC reward can be calculated using the corresponding rewards from time step t to time step T (which can be a fixed number of time steps in the future relative to time step t or the last time step in the trajectory).

[0160] For each action dimension, the system can then generate a target Q-value as the maximum of the first target Q-value and the second target Q-value for that action dimension.

[0161] from Figure 5 It can be seen that due to the autoregressive generation of actions (and the autoregressive calculation of target Q-values), there are many Q-Transformer steps 520 for each environment step (“time step”) 530.

[0162] Figure 6 Quantitative examples of the performance gains that can be achieved by using the policy system described in this specification are shown.

[0163] Specifically, Figure 6 Shows the use Figure 1 Overall performance (in terms of success rate after different numbers of training steps) of an agent controlled by policy 100 (“Q-Transformer”) 602 and an agent controlled using a baseline system in a robotics task that requires using observed images to pick up and move a specified object.

[0164] The baseline systems include the QT-Opt CQL system, the decision-making Transformer system, the AW-Opt system, the IQL system, and the RT-1 BC system.

[0165] As can be seen, the Q-Transformer significantly outperforms these baseline systems for almost all training steps. Specifically, after training progresses beyond the first 20,000 training steps, the Q-Transformer achieves a significant improvement in success rate relative to all baselines for all remaining training steps.

[0166] This specification uses the term "configuration" when referring to systems and computer program components. For a system of one or more computers to be configured to perform a particular operation or action, this means that the system has installed thereon software, firmware, hardware, or a combination thereof that, when in operation, causes the system to perform that operation or action. For one or more computer programs to be configured to perform a particular operation or action, this means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform that operation or action.

[0167] Embodiments of the subject matter and functional operations described in this specification may be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more thereof. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by the data processing device.

[0168] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of equipment, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. An apparatus may also be or further include special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, an apparatus may optionally include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.

[0169] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) may be written in any form of programming language, including compiled or interpreted languages ​​or declarative or procedural languages, and it may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communications network.

[0170] In this specification, the term "database" is used broadly to refer to any collection of data: the data need not be structured in any particular way, or at all, and may be stored on a storage device in one or more locations. Thus, for example, an index database may include multiple collections of data, each of which may be organized and accessed differently.

[0171] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific engine; in other cases, multiple engines may be installed and run on the same computer or computers.

[0172] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special-purpose logic circuitry, such as an FPGA or ASIC, or by a combination of special-purpose logic circuitry and one or more programmed computers.

[0173] A computer suitable for executing a computer program can be based on a general-purpose or special-purpose microprocessor or both, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory or random access memory or both. The basic elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory can be supplemented by or incorporated into a dedicated logic circuit system. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name a few.

[0174] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD ROM and DVD-ROM disks.

[0175] To provide for user interaction, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, as well as a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices may also be used to provide for user interaction; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. Furthermore, a computer may interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser. Furthermore, a computer may interact with a user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving responsive messages from the user in response.

[0176] A data processing device used to implement a machine learning model may also include, for example, dedicated hardware accelerator units for processing general-purpose and computationally intensive parts of machine learning training or production (i.e., inference, workloads).

[0177] Machine learning models can be implemented and deployed using a machine learning framework (e.g., the TensorFlow framework or the Jax framework).

[0178] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with implementations of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0179] A computing system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from the user. Data generated at the user device, for example, the results of the user interaction, may be received from the device at the server.

[0180] Although this specification contains many specific implementation details, these details should not be interpreted as limiting the scope of any invention or the scope of what may be claimed, but rather as descriptions of features that may be unique to a particular embodiment of a particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed as such, one or more features from the claimed combination may be deleted from the combination in some cases, and the claimed combination may involve a subcombination or a variant of a subcombination.

[0181] Similarly, although operations are depicted in the drawings and recited in the claims in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed, to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above-described embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.

[0182] This instruction manual also provides the subject matter of the following clauses:

[0183] Clause 1. A method, executed by one or more computers, for controlling an agent interacting with an environment, the method comprising, at each of a plurality of time steps:

[0184] maintaining historical data representing observations characterizing a state of the environment at previous time steps;

[0185] obtaining a current observation representing a state of the environment at the time step;

[0186] generating an input sequence of input tokens from at least the current observation and the observation represented in the historical data;

[0187] processing the input sequence of input tokens using a Transformer neural network to select an action to be performed by the agent in response to the current observation, wherein the action comprises a respective sub-action for each of a plurality of action dimensions in a dimensional sequence of action dimensions, and wherein for each of the plurality of action dimensions, the processing comprises:

[0188] processing a combined sequence using the Transformer neural network, the combined sequence comprising the input sequence followed by a corresponding word-gram for each previous action dimension preceding the action dimension in the sequence of dimensions, the corresponding word-gram identifying the sub-action selected for the action dimension, to generate a corresponding Q-value for each sub-action in a set of candidate sub-actions for the action dimension; and

[0189] selecting a sub-action for the action dimension using the corresponding Q-values ​​of the sub-actions in the set of candidate sub-actions for the action dimension; and

[0190] Causes the agent to perform the selected action.

[0191] Clause 2. The method of clause 1, wherein the environment is a real-world environment and the agent is a robot.

[0192] Clause 3. The method of clause 1 or clause 2, wherein, for one or more of the action dimensions, the set of candidate sub-actions for the action dimension represents a discretization of a continuous space of sub-actions for the action dimension.

[0193] Clause 4. The method of any preceding clause, further comprising:

[0194] Prior to a first time step in the plurality of time steps, receiving a natural language text sequence representing a task to be performed by the agent in the environment, wherein generating an input sequence of input tokens from at least the current observation and the observation represented in the historical data comprises:

[0195] The input sequence is generated from at least the current observation, the observation represented in the historical data, and the natural language text sequence.

[0196] Clause 5. The method of any preceding clause, wherein the current observation and the observations in the historical data each comprise one or more images of the environment.

[0197] Clause 6. The method of clause 5 when dependent on clause 4, further comprising:

[0198] processing the current observation using an image encoder neural network conditioned on an encoded representation of the natural language text sequence to generate an encoded representation of the current observation, wherein generating the input sequence from at least the current observation, the observation represented in the historical data, and the natural language text sequence comprises:

[0199] An input sequence of the input word-grams is generated from at least the encoded representation of the current observation and corresponding encoded representations of the observations in the history.

[0200] Clause 7. The method of Clause 6, wherein generating the input sequence of input tokens from at least the encoded representation of the current observation and corresponding encoded representations of the observations in the history comprises:

[0201] A sequence of image word-grams for the current observation is generated from the encoded representation of the current observation.

[0202] Clause 8. A method as described in Clause 7, wherein the input sequence of input word elements includes the image word element sequence for the current observation and a corresponding image word element sequence for each of the observations in the history that has been generated from the corresponding encoded representation of the observation.

[0203] Clause 9. The method of clause 8, wherein a positional encoding is applied to each image word-gram in the sequence of image words for the observed image and in a corresponding sequence of image words for one or more earlier observations.

[0204] Clause 10. The method of any one of clauses 6 to 9, wherein the encoded representation comprises a feature map comprising a respective feature vector for each of a plurality of regions in the current observation, and wherein generating a sequence of image word-grams for the observed image from the encoded representation of the observation comprises:

[0205] An initial input sequence is generated by flattening the feature map into a sequence of feature vectors.

[0206] Clause 11. The method of Clause 10, wherein generating a sequence of image word-grams for the observed image from the encoded representation of the observation comprises:

[0207] The initial input sequence of feature vectors is processed using a learned module that maps the input sequence to a reduced sequence comprising a smaller number of tokens.

[0208] Clause 12. A method as described in any of clauses 6 to 11, wherein the image encoder neural network includes one or more conditioning layers, each conditioning layer being configured to receive a corresponding intermediate output of a corresponding intermediate layer of the image encoder neural network output and the encoded representation of the natural language instruction, and (i) use the encoded representation of the natural language instruction to update the corresponding intermediate output of the image encoder neural network and (ii) provide the updated corresponding intermediate output as input to a corresponding subsequent intermediate layer of the image encoder neural network.

[0209] Clause 13. The method of Clause 12, wherein the one or more conditioning layers are feature-by-feature linear modulation FiLM layers.

[0210] Clause 14. A method as described in any of clauses 12 or 13, wherein the image encoder neural network is a convolutional neural network and the corresponding intermediate layer, the corresponding subsequent layer, or both are convolutional layers.

[0211] Clause 15. A method as described in any preceding clause, wherein the Transformer is a decoder-only Transformer comprising a plurality of self-attention layer blocks.

[0212] Clause 16. The method of any preceding clause, wherein selecting a sub-action for the action dimension comprises:

[0213] The candidate sub-action with the highest corresponding Q value is selected.

[0214] Clause 17. A method as described in any preceding clause when dependent on clause 6, wherein the image encoder neural network and the Transformer neural network have been jointly trained on a set of training data.

[0215] Clause 18. The method of clause 17, wherein prior to the joint training, the image encoder neural network has been pre-trained on an image classification task.

[0216] Clause 19. The method of clause 18 when dependent upon clause 12, wherein the image encoder neural network does not include the one or more conditioning layers used for the pre-training.

[0217] Clause 20. The method of any one of clauses 17 to 19 when dependent on clause 11, wherein each conditioning layer is initialized prior to the joint training to act as an identity transformation on the corresponding respective intermediate output.

[0218] Clause 21. The method of any one of clauses 17 to 20 when dependent on clause 11, wherein the learned module has also been trained as part of the joint training.

[0219] Clause 22. A method as described in any of clauses 17 to 21, wherein the training data includes simulated data.

[0220] Clause 23. A method as described in any of clauses 17 to 22, wherein the training data includes real-world data.

[0221] Clause 24. The method of clause 23, wherein the training data comprises both simulated data and real-world data.

[0222] Clause 25. The method of any preceding clause when dependent on clause 6, further comprising: generating the encoded representation of the natural language text sequence by processing the natural language text sequence using a text encoder neural network to generate an embedding of the encoded representation.

[0223] Clause 26. The method of Clause 25, wherein the text encoder neural network is pre-trained on a text representation learning task.

[0224] Clause 27. The method of clause 26 when dependent upon clause 17, wherein the text encoder neural network is fine-tuned during the joint training.

[0225] Clause 28. The method of clause 26 when dependent upon clause 17, wherein the text encoder neural network is kept frozen during the joint training.

[0226] Clause 29. The method of any preceding clause, wherein the Transformer has been trained on an offline dataset via offline reinforcement learning.

[0227] Clause 30. The method of Clause 29, wherein the Transformer has been trained on an offline dataset via offline Q-learning.

[0228] Clause 31. The method of Clause 30, wherein the offline Q-learning is a conservative offline Q-learning technique.

[0229] Clause 32. A method as described in any preceding clause, wherein, for the last action dimension in the sequence of dimensions, the corresponding Q-value for each of the sub-actions in the set of candidate sub-actions for the action dimension represents an estimate of the reward that will be received in response to the agent performing an action comprising: the candidate sub-actions for the last action dimension; and, for each previous action dimension in the sequence of dimensions, the sub-actions that have been selected for the action dimension.

[0230] Clause 33. The method of any preceding clause, wherein, for each given candidate sub-action for each given action dimension except the last action dimension, the corresponding Q-value for the given candidate sub-action represents an estimate of the reward that would be received in response to the agent performing an action comprising:

[0231] for any previous action dimension preceding the given action dimension in the sequence of dimensions, the sub-actions that have been selected for the action dimension;

[0232] For the given action dimension, the candidate sub-actions; and

[0233] For each subsequent action dimension after the given action dimension in the sequence of dimensions, assuming that the given candidate sub-action is selected for the given action dimension, the sub-action selected for the action dimension by autoregressively selecting sub-actions using a Transformer neural network.

[0234] Clause 34. A method as described in any preceding clause, wherein the agent is a robot and the one or more computers are on the robot.

[0235] Clause 35. A method as described in any preceding clause when dependent on clause 15, wherein each self-attention layer block applies a causal self-attention mechanism.

[0236] Clause 36. A method of controlling a robot, the method comprising, at each time step of a plurality of time steps:

[0237] obtaining, by a control system of the robot, an observation image of the environment at the time step;

[0238] providing the observed image to a strategy system via a control system of the robot;

[0239] obtaining, by the control system of the robot and from the policy system of the robot, data specifying the selected action, wherein the policy system selects the selected action by performing the operation of the corresponding method as described in any preceding clause in response to the observed image; and

[0240] The robot is caused to perform the selected action by the control system of the robot.

[0241] Clause 37. The method of clause 36, wherein the control system of the robot is onboard the robot.

[0242] Clause 38. The method of clause 37, wherein the policy system is on the robotic machine.

[0243] Clause 39. The method of Clause 37, wherein:

[0244] The policy system is remote from the robot,

[0245] Providing the observed image includes transmitting the observed image over a data communications network; and

[0246] Obtaining the data specifying the selected action comprises receiving the data specifying the selected action over the data communications network.

[0247] Clause 40. A method of training a Transformer neural network as described in any preceding clause, the method comprising:

[0248] obtaining an experience tuple comprising: (i) data representing a historical set of observations, (ii) a training observation, (iii) a training action performed in response to the training observation, (iv) a reward received in response to the action being performed, and (v) a next observation; and

[0249] Training the Transformer neural network on the experience tuples includes:

[0250] Given the training observations and the historical observations, using the Transformer neural network and based on current values ​​of parameters of the Transformer neural network, generate a corresponding Q value for each candidate sub-action for each action dimension;

[0251] For each action dimension, using the reward in the experience tuple, generating a corresponding target Q value for the sub-action in the training action for the action dimension; and

[0252] The Transformer neural network is trained based on a target that measures, for each action dimension, a temporal difference error between the corresponding Q-value of the sub-action in the training action for that dimension and the target Q-value of the sub-action in the training action for that dimension.

[0253] Clause 41. The method of clause 40, wherein, for each action dimension, the objective encourages the corresponding Q-values ​​of sub-actions that are not in the training actions for that dimension to be equal to zero.

[0254] Clause 42. The method of clause 41, wherein, for each action dimension, the target measures the square of a Q-value assigned to at least one of the sub-actions that is not in the training actions for that dimension.

[0255] Clause 43. The method of any one of clauses 40 to 42, wherein for each action dimension and using the reward in the experience tuple, a corresponding target Q-value is generated for the sub-action in the training action for the action dimension:

[0256] A first target Q-value is determined by applying an autoregressive Q-objective maximization to an input sequence generated from at least (i) one or more of the historical observations, (ii) the training observation, and (iii) the next observation.

[0257] Clause 44. A method as described in clause 43, wherein applying autoregressive Q-objective maximization to the input sequence includes selecting an action by processing the input sequence using the Transformer neural network as described above in any of clauses 1 to 35 and setting the parameter values ​​to target values.

[0258] Clause 45. The method of any one of Clauses 43 or 44, further comprising:

[0259] determining a second target Q-value as a Monte Carlo return starting from the reward in the tuple; and

[0260] The target Q-value is generated using the reward in the experience tuple and a maximum value of the first target Q-value and the second target Q-value.

[0261] Clause 46. The method of any one of clauses 40 to 45, wherein training the Transformer neural network further comprises training the image encoder neural network, the learned module, or both.

[0262] Clause 47. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the respective methods of any one of clauses 1 to 46.

[0263] Clause 48. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the respective method of any one of clauses 1 to 46.

[0264] Specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method executed by one or more computers for controlling an intelligent agent interacting with an environment, the method comprising, at each of a plurality of time steps: maintaining historical data representing observations characterizing a state of the environment at previous time steps; obtaining a current observation representing a state of the environment at the time step; generating an input sequence of input tokens from at least the current observation and the observation represented in the historical data; processing the input sequence of input tokens using a Transformer neural network to select an action to be performed by the agent in response to the current observation, wherein the action comprises a respective sub-action for each of a plurality of action dimensions in a dimensional sequence of action dimensions, and wherein for each of the plurality of action dimensions, the processing comprises: processing a combined sequence using the Transformer neural network, the combined sequence comprising the input sequence followed by a corresponding word-gram for each previous action dimension preceding the action dimension in the sequence of dimensions, the corresponding word-gram identifying the sub-action selected for the action dimension, to generate a corresponding Q-value for each sub-action in a set of candidate sub-actions for the action dimension; and selecting a sub-action for the action dimension using the corresponding Q-values ​​of the sub-actions in the set of candidate sub-actions for the action dimension; and Causes the agent to perform the selected action.

2. The method of claim 1, wherein the environment is a real-world environment and the agent is a robot.

3. The method of claim 1 or claim 2, wherein for one or more of the action dimensions, the set of candidate sub-actions for the action dimension represents a discretization of a continuous space of sub-actions for the action dimension.

4. The method of any preceding claim, further comprising: Prior to a first time step in the plurality of time steps, receiving a natural language text sequence representing a task to be performed by the agent in the environment, wherein generating an input sequence of input tokens from at least the current observation and the observation represented in the historical data comprises: The input sequence is generated from at least the current observation, the observation represented in the historical data, and the natural language text sequence.

5. A method as claimed in any preceding claim, wherein the current observation and the observations in the historical data each comprise one or more images of the environment.

6. The method of claim 5 when dependent on claim 4, further comprising: processing the current observation using an image encoder neural network conditioned on an encoded representation of the natural language text sequence to generate an encoded representation of the current observation, wherein generating the input sequence from at least the current observation, the observation represented in the historical data, and the natural language text sequence comprises: An input sequence of the input word-grams is generated from at least the encoded representation of the current observation and corresponding encoded representations of the observations in the history.

7. The method of claim 6 , wherein generating the input sequence of input tokens from at least the encoded representation of the current observation and corresponding encoded representations of the observations in the history comprises: A sequence of image word-grams for the current observation is generated from the encoded representation of the current observation.

8. A method as claimed in claim 7, wherein the input sequence of input word-grams includes the image word-gram sequence for the current observation and a corresponding image word-gram sequence for each of the observations in the history that has been generated from the corresponding encoded representation of the observation.

9. The method of claim 8, wherein a positional encoding is applied to each image word-gram in the sequence of image words for the observed image and in a corresponding sequence of image words for one or more earlier observations.

10. The method of any one of claims 6 to 9, wherein the encoded representation comprises a feature map comprising a respective feature vector for each of a plurality of regions in the current observation, and wherein generating a sequence of image word-grams for the observed image from the encoded representation of the observation comprises: An initial input sequence is generated by flattening the feature map into a sequence of feature vectors.

11. The method of claim 10, wherein generating a sequence of image word-grams for the observed image from the encoded representation of the observation comprises: The initial input sequence of feature vectors is processed using a learned module that maps the input sequence to a reduced sequence comprising a smaller number of tokens.

12. The method of any one of claims 6 to 11, wherein the image encoder neural network comprises one or more conditioning layers, each conditioning layer being configured to receive a corresponding intermediate output of a corresponding intermediate layer of the image encoder neural network and the encoded representation of the natural language instruction, and (i) update the corresponding intermediate output of the image encoder neural network using the encoded representation of the natural language instruction and (ii) provide the updated corresponding intermediate output as input to a corresponding subsequent intermediate layer of the image encoder neural network.

13. The method of claim 12, wherein the one or more conditioning layers are feature-by-feature linear modulation (FiLM) layers.

14. The method of any one of claims 12 or 13, wherein the image encoder neural network is a convolutional neural network and the corresponding intermediate layer, the corresponding subsequent layer, or both are convolutional layers.

15. A method as claimed in any preceding claim, wherein the Transformer is a decoder-only Transformer comprising a plurality of self-attention layer blocks.

16. The method of any preceding claim, wherein selecting a sub-action for the action dimension comprises: The candidate sub-action with the highest corresponding Q value is selected.

17. A method as claimed in any preceding claim when dependent on claim 6, wherein the image encoder neural network and the Transformer neural network have been jointly trained on a set of training data.

18. The method of claim 17, wherein prior to the joint training, the image encoder neural network has been pre-trained on an image classification task.

19. The method of claim 18 when dependent on claim 12, wherein the image encoder neural network does not include the one or more conditioning layers used for the pre-training.

20. A method as claimed in any one of claims 17 to 19 when dependent on claim 11, wherein each conditioning layer is initialised prior to said joint training to act as an identity transformation on said corresponding intermediate output.

21. A method as claimed in any one of claims 17 to 20 when dependent on claim 11, wherein the learned module has also been trained as part of the joint training.

22. The method of any one of claims 17 to 21, wherein the training data comprises simulated data.

23. A method as claimed in any one of claims 17 to 22, wherein the training data comprises real-world data.

24. The method of claim 23, wherein the training data comprises both simulated data and real-world data.

25. The method of any preceding claim when dependent on claim 6, further comprising: The encoded representation of the natural language text sequence is generated by processing the natural language text sequence using a text encoder neural network to generate an embedding of the encoded representation.

26. The method of claim 25, wherein the text encoder neural network is pre-trained on a text representation learning task.

27. The method of claim 26 when dependent on claim 17, wherein the text encoder neural network is fine-tuned during the joint training.

28. The method of claim 26 when appended to claim 17, wherein the text encoder neural network is kept frozen during the joint training.

29. A method as claimed in any preceding claim, wherein the Transformer has been trained on an offline dataset by offline reinforcement learning.

30. The method of claim 29, wherein the Transformer has been trained on an offline dataset via offline Q-learning.

31. The method of claim 30, wherein the offline Q-learning is a conservative offline Q-learning technique.

32. A method as claimed in any preceding claim, wherein for the last action dimension in the sequence of dimensions, the corresponding Q-value for each of the sub-actions in the set of candidate sub-actions for the action dimension represents an estimate of the reward that will be received in response to the agent performing an action comprising: the candidate sub-actions for the last action dimension; and, for each previous action dimension in the sequence of dimensions, the sub-actions that have been selected for the action dimension.

33. A method as claimed in any preceding claim, wherein for each given candidate sub-action for each given action dimension except the last action dimension, the corresponding Q-value for the given candidate sub-action represents an estimate of the reward that would be received in response to the agent performing an action comprising: for any previous action dimension preceding the given action dimension in the sequence of dimensions, the sub-actions that have been selected for the action dimension; For the given action dimension, the candidate sub-actions; and For each subsequent action dimension after the given action dimension in the sequence of dimensions, assuming that the given candidate sub-action is selected for the given action dimension, the sub-action selected for the action dimension by autoregressively selecting sub-actions using the Transformer neural network.

34. A method as claimed in any preceding claim, wherein the agent is a robot and the one or more computers are on board the robot.

35. A method as claimed in any preceding claim when dependent on claim 15, wherein each self-attention layer block applies a causal self-attention mechanism.

36. A method of controlling a robot, the method comprising, at each time step of a plurality of time steps: obtaining, by a control system of the robot, an observation image of the environment at the time step; providing the observed image to a strategy system via a control system of the robot; obtaining, by the control system of the robot and from the policy system of the robot, data specifying the selected action, wherein the policy system selects the selected action by performing the operations of the corresponding method of any preceding claim in response to the observed image; and The robot is caused to perform the selected action by the control system of the robot.

37. The method of claim 36, wherein the control system of the robot is onboard the robot.

38. The method of claim 37, wherein the policy system is on the robotic machine.

39. The method of claim 37, wherein: The policy system is remote from the robot, Providing the observed image includes transmitting the observed image over a data communications network; and Obtaining the data specifying the selected action comprises receiving the data specifying the selected action over the data communications network.

40. A method of training a Transformer neural network as claimed in any preceding claim, the method comprising: obtaining an experience tuple comprising: (i) data representing a historical observation set, (ii) a training observation, (iii) a training action performed in response to the training observation, (iv) a reward received in response to the training action being performed, and (v) a next observation; as well as Training the Transformer neural network on the experience tuples includes: Given the training observations and the historical observations, using the Transformer neural network and based on current values ​​of parameters of the Transformer neural network, generate a corresponding Q value for each candidate sub-action for each action dimension; For each action dimension, using the reward in the experience tuple, generating a corresponding target Q value for the sub-action in the training action for the action dimension; as well as The Transformer neural network is trained based on a target that measures, for each action dimension, a temporal difference error between the corresponding Q-value of the sub-action in the training action for that dimension and the target Q-value of the sub-action in the training action for that dimension.

41. The method of claim 40, wherein for each action dimension, the objective encourages the corresponding Q-values ​​of sub-actions that are not in the training actions for that dimension to be equal to zero.

42. The method of claim 41, wherein for each action dimension, the target measures the square of a Q-value assigned to at least one of the sub-actions that is not in the training actions for that dimension.

43. The method of any one of claims 40 to 42, wherein for each action dimension and using the reward in the experience tuple, generating a corresponding target Q value for the sub-action in the training action for the action dimension comprises: For the last action dimension, a first target Q-value is determined by applying an autoregressive Q-objective maximization to an input sequence generated from at least (i) one or more of the historical observations, (ii) the training observations, and (iii) the next observation.

44. The method of claim 43, wherein applying autoregressive Q-objective maximization to the input sequence comprises selecting an action by processing the input sequence using the Transformer neural network as described above in any one of claims 1 to 35 and setting the parameter values ​​to target values.

45. The method of any one of claims 43 or 44, further comprising: determining a second target Q-value as a Monte Carlo return starting from the reward in the tuple; as well as For each action dimension, the target Q value is generated as the maximum value of the first Q value and the second target Q value of the action dimension.

46. ​​The method of any one of claims 40 to 45, wherein training the Transformer neural network further comprises training the image encoder neural network, the learned module, or both.

47. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the respective methods of any one of claims 1 to 46.

48. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the respective method of any one of claims 1 to 46.