Real-world robotic control using TRANSFORMER neural networks
Through the strategy system of the multi-task neural network backbone and the Transformer neural network, the control delay and complex task execution problems of agents in an environment that have not been seen before are solved, and high-precision and smooth agent control is achieved, which is suitable for complex robot tasks and reduces the cost of data acquisition.
Patent Information
- Application Number
- CN202380085883.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-13
- Filing Date
- 2023-12-13
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively control the task execution of an agent in complex environments, especially in unseen environments or tasks, and there is a control delay problem.
A strategy system using a multi-task neural network backbone is used to realize a high-capacity and data efficient architecture through open task-independent training and large-scale data sets. The Transformer neural network is used to process observation images and natural language text sequences, and generate policy outputs to control agent movements.
It realizes high-precision and smooth agent control in unseen tasks and environments, reduces control delays, and can adapt to complex robot control tasks, reducing data acquisition costs.
Smart Images

Figure CN120303668A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims the priority of U.S. Provisional Application No. 63 / 432,373, filed on December 13, 2022. The disclosure of the prior application is considered part of the disclosure of this application and is incorporated herein by reference. Background Art
[0002] This specification relates to using neural networks to control agents.
[0003] A neural network is a machine learning model that uses one or more layers of non - linear units to predict an output for received inputs. In addition to the output layer, some neural networks also include one or more hidden layers. The output of each hidden layer is used as an input to the next layer in the network (i.e., the next hidden layer or the output layer). Each layer of the network generates an output from the received inputs based on the current values of the corresponding set of parameters. Summary of the Invention
[0004] This specification describes a system implemented as a computer program on one or more computers in one or more locations that controls an agent, e.g., a robot, that interacts in an environment by selecting an action to be performed by the agent and then causing the agent to perform the action.
[0005] The subject matter described in this specification can be implemented in particular embodiments so as to achieve one or more of the following advantages. The policy system described in this specification is a neural network system that implements a multi - task neural network backbone that can control an agent, e.g., a robot, to perform any one of a variety of real - world robot control tasks. By leveraging open - ended task - agnostic training that extends beyond robot control, a high - capacity and data - efficient architecture that can learn knowledge present in large - scale datasets, or both, the policy system is able to solve specific downstream robot control tasks to a high level of performance with zero - shot or by leveraging relatively small task - specific datasets.
[0006] In addition, the policy system can control the agent with reduced latency, e.g., can generate policy outputs that specify the actions to be performed by the agent in response to observed images at a frequency of 3 Hz or higher. This reduced latency ensures that the policy system can control the agent to move in a more natural and smoother manner, which in turn results in higher - precision agent movement, which may be key to successful task completion.
[0007] In addition, once trained, the policy system can control a robot on new tasks (involving new environments, new objects, or both) that were not seen during the training of the policy system. This generalization ability of the policy system is beneficial because it extends the applicability of the described policy system to complex robot control tasks, such as dexterous tasks, long-term tasks, etc., for which demonstration data is difficult to obtain or computationally expensive to acquire.
[0008] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 Illustrates an example policy system and an example control system.
[0010] Figure 2 Is a diagram of the architecture of an example policy system.
[0011] Figure 3 Is a diagram of operations performed by an example policy system.
[0012] Figure 4 Is a flowchart of an example process for controlling an agent interacting with an environment.
[0013] Figure 5 Is a flowchart of an example process for training a set of neural networks included in a policy system.
[0014] Figure 6 Is a diagram of training a set of neural networks on a mixed training data set.
[0015] Figure 7 Illustrates a quantitative example of the performance gain achievable by using the policy system described in this specification.
[0016] Like reference numerals and names in the various figures indicate like elements. DETAILED DESCRIPTION
[0017] Figure 1 Illustrates an example policy system 100 and an example control system 101. The policy system 100 and the control system 101 are examples of systems implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below are implemented.
[0018] The policy system 100 and the control system 101 can control an agent 102 (e.g., a robot) to perform any one of a wide variety of tasks in the environment 104. To control the agent 102 interacting in the environment 104 to perform a task, the policy system 100 selects an action 144 to be executed by the agent 102, and then the control system 101 causes the agent 102 to execute the selected action 144.
[0019] As a general example, the task can include, for example, one or more of the following: navigating to a specified location in the environment 104, identifying a specific object in the environment 104, manipulating a specific object in a specified manner, controlling an equipped item to meet a standard, etc. To complete such tasks, the agent 102 moves (e.g., navigates) and / or changes its configuration within the environment 104.
[0020] Generally, the control system 101 is located locally to the agent 102. For example, the control system 101 can be on-board the agent 102, such as implemented on one or more computers, local workstations, or local servers on the agent 102 that have relatively small processing and memory resources.
[0021] In some implementations, the policy system 100 is located locally to the agent 102. For example, like the control system 101, the policy system 100 can also be on-board the agent 102. Additionally, in some of these implementations, the policy system 100 can be part of the control system 101 that causes the agent 102 to execute the action 144.
[0022] In other implementations, the policy system 100 is remote from the agent 102. For example, different from the control system 101, the policy system 100 can be hosted within a data center, which can be a distributed computing system with hundreds or thousands of computers in one or more locations. That is, the control system 101 can receive data identifying the action 144 from an external source, rather than generating such data on the agent 102.
[0023] In these implementations, the policy system 100 and the control system 101 can be connected via a data communication network, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof.
[0024] In these implementations, the control system 101 of the agent 102 interacts with a remote policy system 100 within a data center that has much more computing resources and other resources compared to those available on the agent 102 to reduce the latency in selecting the action 144, reduce the consumption of the limited power supply of the agent 102 when selecting the action 144, or both.
[0025] In some implementations, the policy system 100, the control system 101, or both may expose one or more application programming interfaces (APIs) or other data interfaces that facilitate the control of the agent 102. For example, a user of the agent 102 may use an API that becomes available through the action selection system 100 to provide a natural language text sequence 108 that characterizes a task to be performed by the agent. As another example, the policy system 100 and the control system 101 may interact through an API between the two systems. For example, the control system 101 may use the API to provide an observation image 106 to the policy system 100, and the policy system 100 may use the API to provide data specifying the selected action 144 to the control system 101.
[0026] Specifically, the policy system 100 and the control system 101 control the agent based on policy outputs generated by a set of neural networks that have been configured through training to control the agent 102 in response to an observation image 106 that characterizes the environment 104 and a natural language text sequence 108 that characterizes a task to be performed by the agent 102.
[0027] For example, the observation image 106 may be an image captured by a camera sensor of the agent 102 or by a camera sensor located in the environment 104. The camera sensor may be, for example, a static camera or a video camera.
[0028] The natural language text sequence 108 may be received from another agent in the environment 104 or from the control system 101 of the agent 102. For example, another agent in the environment 104 may speak an instruction, and the control system 101 or another system may transcribe the instruction into the natural language text sequence 108 and then provide the transcription to the policy system 100. As another example, the control system 101 may receive an instruction (e.g., text-based input, selection-based input, or audio-based input) specifying the natural language text sequence 108 entered by a user and then provide the instruction to the policy system 100.
[0029] More specifically, the policy system 100 receives a natural language text sequence 108 that characterizes a task to be performed by the agent 102 in the environment 104 and uses a text encoder neural network 110 to process the natural language text sequence 108 to generate an encoded representation 112 of the natural language text sequence (or simply an “encoded text”).
[0030] The encoded representation 112 may be or include an embedding of the text sequence 108. As used in this specification, an “embedding” is a sequence of one or more vectors of numerical values (e.g., floating point values or other values), each vector having a predetermined dimension.
[0031] The text encoder neural network 110 can have any suitable neural network architecture that allows the neural network to map a natural language text sequence 108 to an embedding of the text sequence 108. Specific example architectures will be described further below, but more generally, the text encoder neural network 110 can include any suitable type of neural network layer (e.g., an embedding layer, a fully connected layer, etc.) connected in any suitable number (e.g., 2 layers, or 5 layers, or 10 layers) and in any suitable configuration (e.g., as a directed graph of layers).
[0032] At each of a plurality of time steps, the policy system 100 obtains an observation image 106 that characterizes the state of the environment 104 at that time step. The observation image 106 can be obtained from a camera sensor (e.g., the camera sensor of the agent 102 or a camera sensor located in the environment 104), or from the control system 101 of the agent 102. For example, the control system 101 of the agent 102 obtains the observation image 106 of the environment 104 at that time step and then provides the observation image 106 to the policy system 100.
[0033] In an implementation where the policy system 100 is remote from the agent 102, providing the observation image 106 to the policy system 100 can include transmitting data representing the observation image 106, for example, via a data communication network connecting the policy system 100 and the control system 101. As another example, providing the observation image 106 to the policy system 100 can include providing data specifying a name or network location (e.g., a uniform resource locator (URL) of a server from which the policy system 100 can obtain the observation image 106) to the policy system 100.
[0034] The policy system 100 uses the image encoder neural network 120 to process an input that includes the observation image 106 and, in some cases, an encoded representation 112 of a natural language text sequence to generate an encoded representation 122 of the observation image (or simply an "encoded image").
[0035] Although this specification generally describes observations as images, in some cases, an observation can include additional data in addition to the image, such as proprioceptive data characterizing the agent or other data captured by other sensors of the agent. In these cases, the other data can be jointly encoded with the observation image 106 by the image encoder neural network 120.
[0036] When the encoded representation 112 of the natural language text sequence is also provided by the policy system 100 as part of the input to the image encoder neural network 120, the image encoder neural network 120 generates an encoded representation 122 of the observed image conditioned on the encoded representation 112 of the natural language text sequence. That is, when generating the encoded representation 122 of the observed image, the image encoder neural network 120 uses the encoded representation 112 of the natural language text sequence as context, i.e., such that different text sequences can result in different representations being generated for the same observed image.
[0037] Specific example architectures will be described further below, but more generally, the image encoder neural network 120 can include any suitable type of neural network layer (e.g., convolutional layers, conditioning layers, attention layers, etc.) connected in any suitable number (e.g., 5 layers, or 10 layers, or 50 layers) and in any suitable configuration (e.g., as a directed graph of layers).
[0038] The policy system 100 generates an input token sequence 132 based on the encoded representation 122 of the observed image. The input token sequence 132 can include an image token sequence for the observed image 106. As used in this specification, a "token" is a vector of numerical values with a fixed dimension or other ordered set (i.e., the number of values in the ordered set is constant across different tokens).
[0039] The encoded representation 122 can include a feature map that includes corresponding feature vectors for each of a plurality of regions in the observed image 106. Thus, the policy system 100 can generate the input token sequence 132 by using the feature vectors included in the feature map or data derived from these feature vectors as the image tokens to be included in the input token sequence 132.
[0040] In some implementations, the policy system 100 applies a learned module to map the feature vectors included in the encoded representation 122 of the observed image to a smaller number of feature tokens, and then generates the input token sequence 132 by using the feature vectors generated as a result of applying the learned module as the image tokens to be included in the input token sequence 132.
[0041] The learned module can be, but need not be, a neural network. In Figure 1 the example, the learned module is implemented using a neural network ("token neural network 130"). When configured as a neural network, the learned module can include any suitable type of neural network layer (e.g., convolutional layers, fully connected layers, attention layers, pooling layers, etc.) connected in any suitable number (e.g., 1 layer, or 5 layers, or 10 layers) and in any suitable configuration (e.g., as a directed graph of layers).
[0042] A specific example architecture of the token neural network 130 will be further described below. As another example, the token neural network 130 can have a Vision Transformer (ViT) architecture including one or more attention layers. As another example, the token neural network 130 can have a convolutional neural network architecture including one or more convolutional layers.
[0043] In still other examples, the learned module can include any data values that define a learned mapping to reduce the number of feature vectors provided as input to the learned module, i.e., map a larger number of feature vectors to a smaller number of feature vectors. "Learned" means that these data values are adjusted during the joint training of the set of neural networks included in the policy system 100.
[0044] After generating the input token sequence 132, the policy system 100 then processes the input token sequence 132 using the Transformer neural network 140 to generate a policy output 142 that defines the action to be performed by the agent 102 in response to the observed image 106 received at that time step. The Transformer neural network 140 receives the input token sequence 132 and generates, for example, in an autoregressive manner, a policy output 142 consisting of a plurality of data values.
[0045] Specific example architectures will be further described below. More generally, however, the Transformer neural network 140 can have any suitable Transformer-based architecture, such as one of the architectures described in the following: “Attention is all you need,” by Ashish Vaswani et al., Advances in neural information processing systems, 30, 2017; “Generating wikipedia by summarizing long sequence,” by Peter J. Liu et al., arXiv preprint arXiv:1801.10198 (2018); “Towards a human-like open-domain chatbot,” by Daniel Adiwardana et al., CoRR, abs / 2001.09977, 2020; “Language models are few-shot learners,” by Tom B Brown et al., arXiv preprint arXiv:2005.14165, 2020; and “Palm: Scaling language modeling with pathways,” by Aakanksha Chowdhery et al., arXiv preprint arXiv:2204.02311 (2022).
[0046] The policy system 100 uses the policy output 142 to select the action 144 to be performed by the agent 102. Specific examples of the policy output 142 and how the policy system 100 uses such policy output 142 to make this action selection will be further described below at Figure 2 the place.
[0047] After the action 144 to be performed by the agent 102 at this time step has been selected, the policy system 100 provides the data identifying the selected action 144 to the control system 101. In an implementation where the policy system 100 is remote from the agent 102, providing the data identifying the selected action 144 can include, for example, transmitting the data identifying the selected action 144 via a data communication network connecting the policy system 100 and the control system 101.
[0048] Then, the control system 101 causes the agent 102 to execute the selected action 144. For example, the control system 101 can do this by generating instructions for the agent 102 that, when executed, will cause the agent 102 to execute the selected action 144, by submitting control inputs directly to the appropriate controls of the agent, or by using another suitable control technique.
[0049] In some implementations, the environment 104 is a real-world environment, and the agent 102 is a mechanical agent that interacts with the real-world environment. For example, the agent can be a robot that interacts with the environment to accomplish goals such as locating an object of interest in the environment, moving the object of interest to a specified location in the environment, physically manipulating the object of interest in a specified manner in the environment, or navigating to a specified destination in the environment; or the agent can be an autonomous or semi-autonomous land, air, or sea vehicle that navigates through the environment to a specified destination in the environment.
[0050] The action 144 can be a control input for controlling the robot (e.g., torque for the joints of the robot, or a higher-level control command), or a control input for controlling an autonomous or semi-autonomous land, air, or sea vehicle (e.g., torque to the control surfaces or other control elements of the vehicle, or a higher-level control command).
[0051] In other words, the action 144 can include, for example, position, velocity, or force / torque / acceleration data for one or more joints of the robot or components of another mechanical agent. The action can additionally or alternatively include electronic control data such as motor control data, or more generally data for controlling one or more electronic devices within the environment, the control of which has an impact on the observed state of the environment. For example, in the case of an autonomous or semi-autonomous land, air, or sea vehicle, the action can include actions for controlling the navigation (e.g., steering) and movement (e.g., braking and / or accelerating) of the vehicle.
[0052] In some implementations, the environment 104 is a simulated environment, and the agent 102 is implemented as one or more computer programs that interact with the simulated environment. For example, the environment can be a computer simulation of a real-world environment, and the agent can be a simulated mechanical agent that navigates through the computer simulation.
[0053] For example, the simulated environment can be a motion simulation environment (e.g., driving simulation or flight simulation), and the agent can be a simulated vehicle that navigates through the motion simulation. In these implementations, action 144 can be a control input used to control the simulated user or the simulated vehicle. As another example, the simulated environment can be a computer simulation of the real-world environment, and the agent can be a simulated robot that interacts with the computer simulation.
[0054] Generally, when environment 104 is a simulated environment, action 144 can include a simulated version of one or more of the actions or action types described previously.
[0055] In some implementations, environment 104 is a suitable execution environment (e.g., a runtime environment or an operating system environment) implemented on one or more computing devices (such as smartphones, tablet computers, wearable devices, automotive systems, stand-alone personal assistant devices, etc.), and agent 102 is a virtual agent (also referred to as an "automation assistant" or a "mobile assistant") with which the user can interact via the computing device. The virtual agent can receive input from the user (e.g., typed or spoken natural language input) and respond with responsive content (e.g., visual and / or audible natural language output). The virtual agent can provide a wide range of functionality by interacting with various local and / or third-party applications, websites, or other agents. In these implementations, action 144 can include any activity or operation that can be performed or initiated by the user on the computing device (e.g., within an application software installed on the computing device).
[0056] In some cases, the policy system 100 can be used to control the interaction of the agent with the simulated environment, and the policy system 100 (or another training system) can train the set of neural networks for controlling agent 102 based on the interaction of agent 102 (or another agent) with the simulated environment to determine the trained values of the parameters of the set of neural networks. This will be described in more detail below with reference to Figures 5 - 6 the training of the set of neural networks.
[0057] After training the set of neural networks based on the interaction of agent 102 (or another agent) with the simulated environment, the trained set of neural networks can be used by the policy system 100 to control the interaction of the real-world agent with the real-world environment, i.e., to control the agent simulated in the simulated environment.
[0058] Training the neural networks based on the interaction of the agent with the simulated environment (i.e., rather than the real-world environment) can avoid the wear and tear of the agent and can reduce the likelihood that the agent may damage itself or aspects of its environment by performing poorly chosen actions.
[0059] Figure 2 is an illustration of the architecture of an example policy system 200 .
[0060] The policy system 200 receives a natural language text sequence 208. The natural language text sequence 208 represents a task to be performed by an agent in an environment. The natural language text sequence 208 may have an instruction format. For example, Figure 2 It is shown that the natural language text sequence 208 is a natural language instruction describing the task "Pick apple from top drawer and place on counter".
[0061] The policy system 200 processes the natural language text sequence 208 using a text encoder neural network 210 to generate an encoded representation 212 of the natural language text sequence.
[0062] exist Figure 2 In the example of , the text encoder neural network 210 has a universal sentence encoder architecture and generates a 512-dimensional embedding, i.e., a vector comprising 512 entries, each of which is a numerical value, e.g., a floating point value. The universal sentence encoder is described in more detail in Daniel Cer et al. Universal sentence encoder. arXiv preprint arXiv:1803.11175, 2018.
[0063] In other examples, the text encoder neural network 210 may have a different architecture and may generate embeddings having smaller or larger dimensions. Additionally, in other examples, the text encoder neural network 210 may generate an embedding that includes a sequence of multiple embedding vectors.
[0064] At each of a plurality of time steps, policy system 200 obtains observation image 206 that characterizes the state of the environment at that time step. Figure 2 In the example of , the agent performs a single action in response to each observed image 206, e.g., such that a new observed image is obtained by the policy system 200 after each action performed by the agent.
[0065] Policy system 100 uses image encoder neural network 220 to generate encoded representation 222 of observed image 206. Encoded representation 222 may include a feature map including a corresponding feature vector for each of a plurality of regions in observed image 206.
[0066] For example, Figure 2The image encoder neural network 220 generates a feature map that includes 81 feature vectors corresponding to 81 regions (e.g., 81 non-overlapping subsets of pixels) arranged along the horizontal and vertical dimensions in the observed image 206, where each feature vector has 512 dimensions.
[0067] The image encoder neural network 220 can generally be configured as a convolutional neural network including one or more convolutional layers. As a specific example of this, Figure 2 shown is a stack of 26 inverted residual blocks ("MBConv blocks") in the image encoder neural network 220. The inverted residual blocks are described in more detail below: in EfficientNet: Rethinking model scaling for convolutional neural networks by Mingxing Tan et al., in Proceedings of the 36th International Conference on Machine Learning (Proceedings of the 36th International Conference on Machine Learning), volume 97 of Proceedings of Machine Learning Research (Volume 97 of Proceedings of Machine Learning Research), pages 6105–6114, PMLR, 09–15 June 2019. URL https: / / proceedings.mlr.press / v97 / tan19a.html.
[0068] In addition, in Figure 2 the example of, the image encoder neural network 220 generates an encoded representation 222 of the observed image conditioned on the encoded representation 212 of the natural language text sequence. That is, the image encoder neural network 220 receives the encoded representation 212 of the natural language text sequence and the observed image 206 as inputs, and processes this input to generate an encoded representation 222 of the observed image as an output.
[0069] When generating the encoded representation 222 of the observed image, the image encoder neural network 220 uses the encoded representation 212 of the natural language text sequence as context, i.e., such that different text sequences can result in different representations being generated for the same observed image.
[0070] To this end, the image encoder neural network 220 further includes one or more conditioning layers. The conditioning layers may be interspersed between other intermediate layers of the image encoder neural network 220 (such as convolutional layers, such as depth-wise convolutional layers).
[0071] Each conditioning layer receives (i) the corresponding intermediate output of the corresponding intermediate layer of the image encoder neural network and (ii) the encoded representation 212 of the natural language text sequence as inputs, and processes the inputs to (i) update the corresponding intermediate output of the image encoder neural network using the encoded representation 212 of the natural language instruction and (ii) provide the updated corresponding intermediate output as an input to the corresponding subsequent intermediate layer of the image encoder neural network.
[0072] As a specific example of this, Figure 2 it is shown that the image encoder neural network 220 includes 26 Feature-wise Linear Modulation (FiLM) layers, which are interspersed between a stack of 26 inverted residual blocks. The FiLM layers learn functions and which output and as functions of the input : where and modulate the corresponding intermediate output of the corresponding intermediate layer through feature-wise affine transformation : .
[0073] The functions and may or may not be implemented as neural networks.
[0074] The FiLM layers are described in more detail below: in Ethan Perez et al., Film: Visual reasoning with a general conditioning layer, Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), April 2018, doi: 10.1609 / aaai.v32i1.11671.
[0075] For example, as shown, the image encoder neural network 220 includes a FiLM layer disposed between a first inverse residual block and a second inverse residual block in a stack. The FiLM layer receives (i) the respective intermediate outputs of the first inverse residual block and (ii) the encoded representation 212 of the natural language text sequence as inputs, and processes the inputs to (i) update the respective intermediate outputs of the first inverse residual block using the encoded representation 212 of the natural language instructions and (ii) provide the updated respective intermediate outputs as inputs to the second inverse residual block. Thus, the second inverse residual block receives, as inputs, not the respective intermediate outputs of the first inverse residual block, but the updated respective intermediate outputs that have been updated by the FiLM layer using the encoded representation 212 of the natural language text sequence.
[0076] The policy system 200 generates an input token sequence 232 based on the encoded representation 222 of the observed image. The input token sequence 232 may include an image token sequence.
[0077] As described above, the encoded representation 222 of the observed image may include a feature map that includes respective feature vectors for each of a plurality of regions in the observed image 206.
[0078] In Figure 2 the example, the policy system 200 generates an initial input sequence by flattening the feature map into a sequence of feature vectors, and then processes this initial input sequence of feature vectors using the token neural network 230 to map the initial input sequence to a reduced input sequence that includes a smaller number of feature tokens. Then, the policy system 200 uses the feature tokens included in the reduced input sequence as the image tokens to be included in the input token sequence 232.
[0079] As a specific example of this, Figure 2 shows that the token neural network 230 has a TokenLearner architecture and processes an initial input sequence of 81 feature vectors (generated based on flattening a feature map that includes 81 feature vectors) by applying a spatial attention mechanism to generate a reduced input sequence that includes 8 feature tokens for the observed image 206 obtained at that time step, where each feature token has 512 dimensions. TokenLearner is described in more detail below: Michael Ryoo et al., Tokenlearner: Adaptive space-time tokenization for videos, Advances in Neural Information Processing Systems, 34:12786–12797, 2021.
[0080] In some implementations, the input token sequence 232 includes only the image tokens for the observed image 206 obtained at the current time step, that is, only the feature tokens included in the reduced input sequence that has been generated for the observed image 206 obtained at the current time step.
[0081] In other implementations, the input token sequence 232 includes not only the image tokens for the observed image 206 obtained at the current time step, but also the image tokens for one or more earlier (or, previous) observed images obtained at one or more earlier time steps, that is, the feature tokens included in the reduced input sequences that have been generated respectively for the one or more earlier observed images obtained at one or more earlier time steps.
[0082] In these implementations, to accelerate inference by avoiding redundant calculations, the policy system 200 may store the feature tokens included in the reduced input sequences that have been generated at each time step and reuse them in subsequent time steps.
[0083] In Figure 2 the example of, the input token sequence 232 includes an image token sequence that is a combination (e.g., concatenation) of: (i) the feature tokens included in the reduced input sequence that has been generated for the observed image 206 obtained at the current time step, and (ii) the feature tokens included in the reduced input sequences that have been generated respectively for five earlier observed images obtained at five earlier time steps before the current time step. Specifically, Figure 2 shows that the input token sequence 232 includes a concatenation of 48 feature tokens.
[0084] In addition, Figure 2 shows that the policy system 200 adopts a positional encoding scheme. Specifically, the policy system 200 adds the corresponding positional encoding 233 to each image token in (i) the image token sequence for the observed image obtained at the current time step, and (ii) the corresponding image token sequences for one or more earlier observed images obtained at one or more earlier time steps.
[0085] The positional encoding 233 can be determined, for example, according to the sine positional encoding scheme or another encoding scheme, so as to uniquely identify the corresponding time step among a plurality of time steps at which the image token is generated for each image token.
[0086] Then, the policy system 200 provides the input token sequence 232 as an input to the Transformer neural network 240. As a specific example, Figure 2FIG. 0 shows a Transformer neural network 240 having a decoder-only Transformer neural network architecture including eight self-attention layers. In other examples, the Transformer neural network 240 may have a different Transformer-based architecture, e.g., an encoder-decoder Transformer neural network architecture that includes more or fewer layers, each layer having the same or different attention mechanisms.
[0087] In some implementations, each possible action that can be performed by an agent is defined by a corresponding value for each of a plurality of action dimensions. In these implementations, for each of the plurality of action dimensions, the policy output 242 may define a corresponding distribution over the possible values for that action dimension.
[0088] In Figure 2 the example of FIG. 8, the agent is a robot having a base and one or more arms, where at least one of the arms has an end effector (e.g., a gripper or another tool) attached to its end. In this example, the plurality of action dimensions includes seven action dimensions for arm movement: x, y, z, roll, pitch, yaw, and the state of the end effector (e.g., the open / closed state of the gripper). The plurality of action dimensions also includes three action dimensions for base movement: x, y, yaw. The plurality of action dimensions further includes an action dimension for mode switching (e.g., for switching between controlling the robot's arm, controlling the robot's base, or terminating a round).
[0089] In other examples, the agent may be a different type of robot, or it may be a vehicle or another type of agent as described above, and each possible action that can be performed by the agent may thus be characterized by a different set of action dimensions.
[0090] In any example, the possible values for an action dimension may be discretized into a fixed number of bins, and the policy output 242 may include data values defining a distribution over the fixed number of bins for that action dimension. The distribution may be a categorical distribution (a corresponding discrete probability distribution) that assigns a corresponding probability score to each of the fixed number of bins for that action dimension.
[0091] In some of these implementations, for each action dimension, a fixed number of intervals may correspond to approximately the same number of possible values for that action dimension. For example, the possible values for a roll (or similarly, pitch or yaw) action dimension have a range from 0 to 360 degrees, which is divided into 32 intervals (each interval corresponding to a range spanning 11.25 degrees), 128 intervals (each interval corresponding to a range spanning approximately 2.81 degrees), 256 intervals (each interval corresponding to a range spanning approximately 1.41 degrees), and so on. As another example, the possible values for a mode switch action dimension are 0 (control the arm of the robot), 1 (control the base of the robot), and 2 (end the turn), which are divided into 3 intervals (each interval corresponding to the respective value), 256 intervals (each of approximately 85 intervals corresponding to the same respective value), and so on.
[0092] In some other implementations, the policy output 242 may have a different format that defines or otherwise specifies the action. For example, the policy output 242 may be a natural language description of the action to be performed by the agent, such as "move arm to position (x, y, z) (move the arm to the position (x, y, z))", "move arm to pose (x, y, z, roll, pitch, yaw) (move the arm to the pose (x, y, z, roll, pitch, yaw))", or "open gripper (open the gripper)". As another example, the policy output 242 may be a sequence of data elements representing the action to be performed by the agent, such as an identifier for the action.
[0093] To select the action to be performed by the agent at that time step, the policy system 200 then uses the respective distribution for each of one or more of the action dimensions to select a corresponding value within the possible values for that action dimension. For example, the policy system 200 may greedily select the interval with the highest score, or may sample an interval from the respective distribution defined by the policy output 242 for the action dimension using, for example, kernel sampling or another sampling technique, and then select the value corresponding to the selected interval - for example, falling within the selected interval - as the selected value for that action dimension.
[0094] Figure 3 is an illustration of the operations performed by an example policy system 300. The policy system 300 may be the same as or similar to Figure 2 the policy system 200 therein. The policy system 300 may control the agent to complete a task in the environment by repeatedly executing one iteration of these operations at each of a plurality of time steps to select the action to be performed by the agent in response to obtaining an observation image of the environment or the environment captured at that time step.
[0095] At a high level, at each of a plurality of time steps, the operation involves receiving data derived from an observed image sequence and a natural language text sequence as input and processing the data to generate a policy output that defines an action to be performed by the agent at that time step.
[0096] Specifically, when configured to have the Figure 2 architecture described in, the policy system 300 can repeat these operations at a frequency of 3 Hz or higher. That is, the policy system is capable of selecting three or more actions to be performed by the agent per second (in response to obtaining three or more or more observed images), which is comparable to the control frequency achieved by, for example, a person skilled in the task to be performed by the agent.
[0097] As shown, the policy system 300 receives a natural language text sequence 308 and uses a text encoder neural network to process the natural language text sequence 308 to generate an encoded representation of the natural language text sequence. The natural language text sequence 308 can have an instruction format that defines a task to be performed by the agent in the environment.
[0098] At each of a plurality of time steps, the policy system 300 obtains an observed image 306 that characterizes the state of the environment at that time step and then provides the observed image 306 to the set of neural networks included in the policy system 300 for further processing to select an action to be performed by the agent in response to the observed image 306.
[0099] The set of neural networks includes an image encoder neural network 320. The image encoder neural network 320 processes both the observed image 306 and the encoded representation of the natural language text sequence to generate an encoded representation of the observed image. The encoded representation includes a feature map that includes a corresponding feature vector for each of a plurality of regions in the observed image 306.
[0100] The set of neural networks also includes a token neural network 330. The policy system 300 flattens the feature map into a sequence of feature vectors and uses the token neural network 330 to process the initial input sequence of feature vectors to generate a reduced input sequence that includes a smaller number of feature tokens.
[0101] The policy system 300 generates an input token sequence from the feature tokens generated for the current observed image 306 obtained at the current time step and optionally from the feature tokens generated for one or more earlier observed images obtained at one or more earlier time steps.
[0102] This set of neural networks further includes a Transformer neural network 340. The Transformer neural network 340 processes the input sequence of tokens to generate a policy output 342 that defines the actions to be performed by the agent in response to the observed image 306.
[0103] The policy system 300 uses the policy output to select the actions to be performed by the agent and then causes the agent to perform the selected actions.
[0104] The policy system 300 can repeat these operations at a frequency of 3 Hz or above until a certain termination condition is met, such as until the policy system generates a specific policy output 342 that defines a terminal action, until a predetermined number of iterations of these operations have been performed, or until a predetermined length of time has elapsed.
[0105] Figure 4 is a flowchart of an example process 400 for controlling an agent that interacts with an environment. For convenience, process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a policy system (e.g., Figure 1 the policy system 100) appropriately programmed according to this specification can execute process 400.
[0106] The system controls the agent to complete a task in the environment by repeatedly executing iterations of process 400 at each of a plurality of time steps (hereinafter referred to as the "current" time step).
[0107] The task to be performed by the agent is characterized by a sequence of natural language text. For example, before the first iteration of process 400, the system receives a sequence of natural language text that characterizes the task to be performed by the agent in the environment and generates an encoded representation of the sequence of natural language text. For example, the encoded representation includes an embedding of the sequence of natural language text, and the system can generate the embedding by processing the sequence of natural language text using a text encoder neural network.
[0108] The system obtains an observed image that characterizes the state of the environment at the current time step (step 402). For example, the observed image can be captured by a camera sensor of the agent or by a camera sensor located in the environment.
[0109] The system uses an image encoder neural network conditioned on the encoded representation of the sequence of natural language text to process the observed image to generate an encoded representation of the observed image (step 404). That is, the image encoder neural network receives the encoded representation of the sequence of natural language text and the observed image as inputs and processes the inputs to generate an encoded representation of the observed image as an output.
[0110] The system generates an input token sequence from at least an encoded representation of an observed image (step 406). The input token sequence may include an image token sequence for the current observed image obtained at the current time step. Optionally, the input token sequence may also include an image token sequence for each of one or more earlier observed images obtained at one or more earlier time steps.
[0111] For example, the encoded representation may include a feature map that includes a respective feature vector for each of a plurality of regions in the observed image. In this example, the system may generate an initial input sequence by flattening the feature map into a sequence of feature vectors, and then use a learned module to process this initial input sequence of feature vectors, the learned module mapping this initial input sequence to a reduced input sequence that includes a smaller number of feature tokens. Then, the feature tokens included in the reduced input sequence may be used as the image tokens to be included in the input token sequence.
[0112] The system uses a Transformer neural network to process the input token sequence to generate a policy output that defines the actions to be performed by the agent in response to the observed image obtained at the current time step (step 408). Each possible action that can be performed by the agent is defined by a respective value for each of a plurality of action dimensions. To define the action to be performed by the agent, the policy output generated by the Transformer neural network includes a respective categorical distribution over the possible values for that action dimension for each of the plurality of action dimensions.
[0113] The system uses the policy output to select the action to be performed by the agent (step 410). This selection can be made by selecting the respective values for one or more of the plurality of action dimensions using the respective categorical distributions defined by the policy output of the Transformer neural network.
[0114] The system causes the agent to perform the selected action (step 412), for example, by directly submitting control input to the agent or by transmitting instructions or other data to a control system of the agent that will cause the agent to perform the selected action, such as via a data communication network.
[0115] Process 400 may be performed when controlling an agent to perform a task where the actions that should be performed (e.g., actions that will result in progress towards completing the task) are unknown. Process 400 may also be performed as part of selecting the action to be performed by the agent based on processing observed images derived from a set of training data (i.e., a set of observed images for which the actions that should be performed by the agent in response to the observed images are known in order to train the set of neural networks to determine the trained values of the parameters of the set of neural networks).
[0116] Figure 5 is a flow chart of an example process 500 for training a set of neural networks included in a policy system. For convenience, process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a policy system (e.g., Figure 1 Process 500 may be performed by a strategy system 100) or another training system.
[0117] To train the set of neural networks, the system obtains a set of training data (step 502). The set of training data may include training data generated based on the interaction of the agent (or another agent) with the environment.
[0118] In one example, the training dataset includes N training examples Each training example corresponds to a corresponding episode spanning multiple time steps, where each time step starts from Start and in e.g. Figure 1 When the agent 102 or another agent successfully completes the task End at; in each training example, i represents a natural language text sequence, represents the observed image obtained at time step t, and represents the action performed by the agent at time step t. Different episodes can have different lengths, i.e., can include different numbers of time steps.
[0119] For example, the training dataset may include Figure 1 Training examples collected by agent 102 or another agent when performing one or more of the following tasks listed in Table 1 below. Table 1.
[0120] In some cases, the training data set includes expert interaction data characterizing the interaction of one or more expert agents with a corresponding environment. An expert agent may be any agent that selects an action according to an action selection strategy in response to observing an image, and the action selection strategy enables the expert agent to make effective progress toward completing a task. For example, an expert agent may be an agent controlled by another already trained policy system, a person skilled in the task to be performed by the agent, etc.
[0121] In some of these cases, the expert interaction data includes simulation data, where simulated expert agents perform one or more tasks in a simulated environment. In other of these cases, the expert interaction data includes real-world data, where real-world expert agents perform one or more tasks in a real-world environment. In still other cases, the expert interaction data includes both simulation data and real-world data.
[0122] In some cases, the training dataset includes training examples generated (and thus having the same physical characteristics) from one or more robots of the same model when the one or more robots of the same model perform the same or different tasks, such as one of the tasks mentioned in Table 1 above, or other tasks.
[0123] In other cases, the training dataset can be a mixed training dataset that includes training examples generated from multiple robots that are not of the same model, not located at the same site, or even not made by the same manufacturer. For example, the mixed training data can be generated from dozens or hundreds of different robots having different physical characteristics and being of different models. Additionally, the mixed training data does not need to be generated from physical robots. For example, the mixed training data can include data generated from simulations of physical robots.
[0124] Figure 6 is an illustration of a set for training a neural network on a mixed training dataset. Figure 6 Shows that the system obtains a first set of training examples generated from a first robot performing a first task, and a second set of training examples generated from a second robot performing a second task. The first robot and the second robot have different physical characteristics and are made by different manufacturers. Then this set of the neural network can be trained to select actions to be performed by the first robot to perform both the first task and the second task.
[0125] Generally, the more diverse the training dataset is, the better the system can generalize to unseen robot control tasks once it is trained on the training dataset. For example, tasks characterized by unseen instructions, tasks involving selecting actions in response to observed images collected within or about an unseen environment, tasks involving unseen objects, etc. The increased diversity of the training dataset can also enhance the system's robustness to possible distractions such as new obstacles, new background scenes, etc. that may occur in tasks previously seen during training.
[0126] The system trains a set of neural networks on this set of training data (step 504). The set of neural networks can include Figure 1 the text encoder neural network 110, the image encoder neural network 120, the token neural network 130, and the Transformer neural network 140.
[0127] To train this set of neural networks, the system selects training examples from the training dataset and, for each training example selected from the training dataset, generates a training policy output that defines the actions to be performed by the agent based on processing the natural language text sequence and the observed image included in the training example. For example, step 502 may involve performing multiple iterations in process 400.
[0128] The system updates the values of the parameters of the neural network based on using machine learning training techniques (e.g., gradient descent training techniques with backpropagation to optimize an objective function using a suitable optimizer (e.g., stochastic gradient descent, RMSprop, or Adam optimizer)).
[0129] For example, the objective function may be a cross-entropy objective function or another objective function that, for each training example selected from the training dataset, evaluates the difference between (i) the training policy output generated by this set of neural networks from processing the natural language text sequence and the observed image included in the training example and (ii) the true value policy output that defines the actions included in the training example.
[0130] During training, the system may incorporate any number of techniques to improve the speed, effectiveness, or both of the training process.
[0131] For example, instead of training each neural network in this set of neural networks from scratch - e.g., from initial parameter values - the system may start training this set of neural networks from pre-trained parameter values in some of the neural networks.
[0132] In some cases, the text encoder neural network may be pre-trained, for example, as part of a larger text processing neural network on a text processing task such as a text representation learning task before the joint training of this set of neural networks. In some of these cases, the text encoder neural network is then fine-tuned during the joint training, while in other of these cases, the text encoder neural network is kept frozen during the joint training, i.e., the joint training of the neural network on the training dataset does not adjust the pre-trained parameter values of the pre-trained text encoder neural network.
[0133] In some cases, the image encoder neural network may be pre-trained, for example, as part of a larger image processing neural network on an image processing task such as image classification or segmentation, and then fine-tuned during the joint training of this set of neural networks on the training dataset.
[0134] In some of these cases, the image encoder neural network does not include one or more conditioning layers for pre-training. For example, the image encoder neural network can be trained as part of a neural network that is trained to classify images into a set of categories without conditioning on any natural language text sequence as context.
[0135] Furthermore, in some of these cases, since inserting a conditioning layer as a new layer into such a pre-trained image encoder neural network may cause disruption to the intermediate outputs of the neural network, the system initializes each conditioning layer to act as an identity transformation on the corresponding intermediate output before joint training. For example, this can be achieved by setting at least some of the parameter values associated with each conditioning layer to zero.
[0136] As another example, the joint training of the set of neural networks can include imitation learning. For example, when the training dataset includes expert interaction data, the system can train the set of neural networks by behavior cloning the expert interaction data to generate a training policy output from which actions can be selected that highly imitate the actions performed by an expert agent.
[0137] Figure 7 A quantitative example showing the performance gain achievable by using the policy system described in this specification. Specifically, Figure 7 shows the Figure 1 overall performance of agents controlled by Policy 100 (“RT-1”) and agents controlled by a baseline system across the following aspects: seen tasks, generalization ability to unseen tasks, and robustness to distractors and backgrounds, in terms of success rate.
[0138] The baseline systems include the Gato system (described in Reed, Scott et al., “A generalist agent.” arXiv preprint arXiv:2205.06175 (2022)) and the BC-Z system (described in Jang, Eric et al., “Bc-z: Zero-shot task generalization with robotic imitation learning. Conference on Robot Learning. PMLR, 2022”), as well as the BC-Z XL system (a BC-Z system with a larger number of parameters).
[0139] In Figure 7Among them, "seen tasks" are tasks seen during training, i.e., tasks on which the policy system has been trained; "unseen tasks" are tasks involving instructions and objects that have been seen individually in the training dataset but are combined in a new way; "distractor tasks" are tasks involving distractor objects not seen previously during training; and "background tasks" are tasks involving environments with backgrounds not seen previously (e.g., backgrounds with different scenes or different lighting).
[0140] It can be understood that RT-1 outperforms these baseline systems by a large margin on all these tasks. Specifically, RT-1 has high overall performance on seen tasks (a 97% success rate compared to 72% for BC-Z (the highest success rate among all baseline systems)), and additionally has an impressive degree of generalization ability (a 76% success rate compared to 52% for Gato) and robustness (an 83% success rate on distractor tasks compared to 47% for BC-Z and a 59% success rate on background tasks compared to 41% for BC-Z).
[0141] This specification uses the term "configured" in connection with systems and computer program components. For a system of one or more computers configured to perform a particular operation or action, it means that software, firmware, hardware, or a combination thereof has been installed on the system that, in operation, causes the system to perform that operation or action. For one or more computer programs configured to perform a particular operation or action, it means that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform that operation or action.
[0142] Embodiments of the subject matter and functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing device. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by the data processing device.
[0143] The term "data processing device" refers to data processing hardware and encompasses all kinds of devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The device can also be or further include dedicated logic circuitry, such as, for example, an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to the hardware, the device can optionally also include code that creates an execution environment for a computer program, such as, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0144] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program can, but does not have to, correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (such as one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (such as files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0145] In this specification, the term "database" is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and can be stored on a storage device in one or more locations. Thus, for example, an indexed database can include multiple collections of data, each of which can be organized and accessed differently.
[0146] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and run on the same one or more computers.
[0147] The processes and logical flows described in this specification can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, for example, special-purpose logic circuitry such as an FPGA or ASIC, or by a combination of special-purpose logic circuitry and one or more programmed computers.
[0148] Computers suitable for executing computer programs can be based on general or special-purpose microprocessors or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data from one or more mass storage devices or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name just a few.
[0149] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices); magnetic disks (such as internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0150] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including voice, speech, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from the devices used by the user; for example, by sending a web page to a web browser in response to a request received from the web browser on the user's device. In addition, the computer may interact with the user by sending a text message or other form of message to a personal device (e.g., a smart phone running a messaging application) and receiving a responsive message from the user in response.
[0151] The data processing device for implementing a machine learning model may also include, for example, a dedicated hardware accelerator unit for processing the general and computationally intensive parts of machine learning training or production (i.e., inference, workload).
[0152] A machine learning model may be implemented and deployed using a machine learning framework (e.g., the TensorFlow framework or the Jax framework).
[0153] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface, a web browser, or an app through which the user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN) and a wide area network (WAN), such as the Internet.
[0154] A computing system may include a client and a server. The client and the server are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs that run on respective computers and have a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a user device, e.g., for the purpose of displaying data to and receiving user input from a user interacting with the device acting as the client. Data generated at the user device, e.g., the result of a user interaction, may be received at the server from the device.
[0155] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features that are described in this specification in the context of separate embodiments may also be implemented in combination within a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination may be excluded from the combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.
[0156] Similarly, although operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood to require that such operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the above-described embodiments should not be understood to require such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0157] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims may be performed in a different order and still achieve desirable results. As one example, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method, executed by one or more computers and for controlling an agent interacting with an environment, the method comprising: Receiving a sequence of natural language text characterizing a task to be performed by the agent in the environment; Generating an encoded representation of the sequence of natural language text; And At each of a plurality of time steps: Obtaining an observation image characterizing the state of the environment at that time step; Using an image encoder neural network conditioned on the encoded representation of the sequence of natural language text to process the observation image to generate an encoded representation of the observation image; Generating an input token sequence from at least the encoded representation of the observation image; Using a Transformer neural network to process the input token sequence to generate a policy output, the policy output defining an action to be performed by the agent in response to the observation image; Using the policy output to select an action to be performed by the agent; and Causing the agent to perform the selected action.
2. The method of claim 1, wherein the environment is a real-world environment and the agent is a robot.
3. The method of any one of the preceding claims, wherein generating an input token sequence from at least the encoded representation of the observation image comprises: Generating an image token sequence for the observation image from the encoded representation of the observation image.
4. The method of claim 3, wherein generating an input token sequence from at least the encoded representation of the observation image comprises: Generating the input token sequence by combining the image token sequence for the observation image with corresponding image token sequences for one or more earlier observation images received at one or more earlier time steps.
5. The method according to claim 4, wherein combining the image token sequence for the observation image with corresponding image token sequences for each of one or more earlier observation images received at one or more earlier time steps comprises: Applying positional encoding to each image token in the image token sequence for the observation image and the corresponding image token sequences for the one or more earlier observation images.
6. The method of any one of claims 3-5, wherein the encoded representation comprises a feature map, the feature map comprising a corresponding feature vector for each of a plurality of regions in the observation image, and wherein generating an image token sequence for the observation image from the encoded representation of the observation image comprises: Generating an initial input sequence by flattening the feature map into a sequence of feature vectors.
7. The method of claim 6, wherein generating an image token sequence for the observation image from the encoded representation of the observation image comprises: Using a learned module to process the initial input sequence of feature vectors, the learned module mapping the initial input sequence to a reduced input sequence comprising a smaller number of feature tokens.
8. The method according to any one of the preceding claims, wherein the image encoder neural network comprises one or more conditioning layers, each conditioning layer being configured to receive a respective intermediate output of a respective intermediate layer of the image encoder neural network and the encoded representation of the natural language instruction, and (i) use the encoded representation of the natural language instruction to update the respective intermediate output of the image encoder neural network and (ii) provide the updated respective intermediate output as an input to a respective subsequent intermediate layer of the image encoder neural network.
9. The method according to claim 8, wherein the one or more conditioning layers are feature-wise linear modulation (FiLM) layers.
10. The method according to any one of claims 8 or 9, wherein the image encoder neural network is a convolutional neural network and the respective intermediate layer, the respective subsequent layer, or both are convolutional layers.
11. The method according to any one of the preceding claims, wherein the Transformer is a decoder-only Transformer.
12. The method according to any one of the preceding claims, wherein the policy output comprises, for each of a plurality of action dimensions, a respective categorical distribution over the possible values for that action dimension.
13. The method according to claim 12, wherein using the policy output to select an action to be performed by the agent comprises using the respective categorical distributions to select respective values for one or more of the action dimensions.
14. The method according to any one of the preceding claims, wherein the image encoder neural network and the Transformer neural network have been jointly trained on a set of training data.
15. The method according to claim 14, wherein prior to the joint training, the image encoder neural network has been pre-trained on an image classification task.
16. The method according to claim 15 when dependent on claim 8, wherein the image encoder neural network does not comprise the one or more conditioning layers for the pre-training.
17. The method according to any one of claims 14-16 when dependent on claim 8, wherein each conditioning layer is initialized prior to the joint training to act as an identity transformation on the corresponding respective intermediate output.
18. The method according to any one of claims 14-17 when dependent on claim 7, wherein the learned module has also been trained as part of the joint training.
19. The method according to any one of the preceding claims when dependent on claim 14, wherein the joint training comprises training by imitation learning, and the training data comprises expert interaction data representing the interaction of one or more expert agents with a corresponding environment.
20. The method according to claim 19, wherein the expert interaction data comprises simulated data.
21. The method according to claim 19, wherein the expert interaction data comprises real-world data.
22. The method according to claim 20 or 21, wherein the expert interaction data includes both simulated data and real-world data.
23. The method according to any one of the preceding claims, wherein generating the encoded representation of the natural language text sequence includes using a text encoder neural network to process the natural language text sequence to generate an embedding of the natural language text sequence.
24. The method according to claim 22, wherein the text encoder neural network is pre-trained on a text representation learning task.
25. The method according to claim 23 when dependent on claim 14, wherein the text encoder neural network is fine-tuned during the joint training.
26. The method according to claim 23 when dependent on claim 14, wherein the text encoder neural network is kept frozen during the joint training.
27. The method according to any one of the preceding claims, wherein the agent is a robot and the one or more computers are on the robot.
28. A method of controlling a robot, the method comprising, at each of a plurality of time steps: obtaining, by a control system of the robot, an observation image of the environment at the time step; providing, by the control system of the robot, the observation image to a policy system; obtaining, by the control system of the robot and from the policy system of the robot, data specifying a selected action, wherein the policy system selects the selected action by performing the operations of the corresponding method according to any one of the preceding claims in response to the observation image; and causing, by the control system of the robot, the robot to perform the selected action.
29. The method according to claim 28, wherein the control system of the robot is on the robot.
30. The method according to claim 29, wherein the policy system is on the robot.
31. The method according to claim 29, wherein: the policy system is remote from the robot, providing the observation image includes transmitting the observation image via a data communication network; and obtaining the data specifying the selected action includes receiving the data specifying the selected action via the data communication network.
32. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1-31.
33. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding method according to any one of claims 1-31.
Citation Information
Cited By
Robot action generation method and system thereof, medium, equipment and program product
CN121105000A