Training and / or utilizing a machine learning model for use in robot control based on natural language

By training a goal-conditioned policy network with diverse task descriptions, the robot can effectively execute tasks based on natural language inputs, overcoming the limitations of existing systems in understanding varied user commands.

JP7683085B2Active Publication Date: 2025-05-26GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024087083
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-05-14
Filing Date
2024-05-29
Publication Date
2025-05-26
Estimated Expiration
2041-05-14

AI Technical Summary

Technical Problem

Existing robots struggle to execute specific tasks in response to free-form natural language inputs, limiting their ability to understand and respond to varied user commands.

Method used

Training a goal-conditioned policy network using multiple datasets with different task descriptions, such as goal images, natural language text, and task IDs, to generate a shared latent goal space representation, allowing the robot to perform tasks based on various input types.

Benefits of technology

Enables a robot to robustly perform tasks defined by natural language instructions without requiring extensive computing resources or human supervision, and allows the robot to follow a wide range of free-form natural language commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683085000025
    Figure 0007683085000025
  • Figure 0007683085000026
    Figure 0007683085000026
  • Figure 0007683085000027
    Figure 0007683085000027
Patent Text Reader

Abstract

To provide a training method of a goal-conditioned policy network based on datasets.SOLUTION: Multiple data sets can include: a goal image data set, where a task is captured in the goal image; a natural language instruction data set, where the task is described in the natural language instruction; a task ID data set, where the task is described by the task ID, etc. In various implementations, each of the multiple data sets has a corresponding encoder, where the encoders are trained to generate a shared latent space representation of the corresponding task description. Additional or alternative techniques are disclosed that enable control of a robot using a goal-conditioned policy network. For example, the robot can be controlled, using the goal-conditioned policy network, based on free-form natural language input describing robot task(s).SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Many robots are programmed to perform specific tasks. For example, robots on an assembly line can be programmed to recognize specific objects and perform specific operations on those specific objects.

Summary of the Invention

Problems to be Solved by the Invention

[0002] Furthermore, some robots can execute specific tasks in response to clear user interface inputs corresponding to specific tasks. For example, a vacuum cleaner robot can execute general vacuum cleaner tasks in response to the utterance "Robot, vacuum." However, usually, the user interface input for causing a robot to execute a specific task must be clearly associated with the task. Therefore, there is a possibility that a robot cannot execute a specific task in response to various free-form natural language inputs from a user who attempts to control the robot. For example, a robot may not be able to move to a target position based on free-form natural language input provided by a user. For example, a robot may not be able to move to a specific position in response to a user's request such as "Go out of the door, turn left, and pass through the door at the end of the corridor."

Means for Solving the Problems

[0003] The techniques disclosed herein are directed to training a goal conditioned policy network based on multiple datasets, where the training tasks are described in different ways in each of the datasets. For example, a robot's task can be described using a goal image, using natural language text, using a task ID, using a natural language utterance, and / or using additional or alternative task descriptions. For example, a robot can be trained to perform the task of putting a ball in a cup. An exemplary description of the task by a goal image could be a photo of a ball in a cup, a description of the task by natural language text could be the natural language instruction "put the ball in the mug", and a description of the task by the task ID could be "task id=4", where 4 is the ID associated with the task of putting a ball in a cup. In some implementations, multiple encoders (i.e., one encoder per dataset) can be trained so that each encoder can generate a shared latent goal space representation of the task by processing the task description. In other words, different ways of describing the same task (e.g., an image of a ball in a cup, the natural language instruction "put the ball in the mug", and / or "task id=4") can be associated with the same latent goal representation based on processing the task description using the corresponding encoder. The techniques described herein are directed to training a robot to perform tasks, but this is not intended to be limiting. Additional and / or alternative networks can be trained according to the techniques described herein based on multiple datasets each having a different context.

[0004] Additional or alternative implementations are directed to controlling a robot based on output generated using a goal-conditioned policy network. In some implementations, a robot can be trained using multiple datasets (e.g., using a goal image dataset and a natural language instruction dataset), and only one task description type can be used to describe a task for the robot at inference time (e.g., only natural language instructions, only goal images, only task IDs, etc. are provided to the system at inference time). For example, the system can be trained based on a goal image dataset and a natural language instruction dataset, and the system is provided with natural language instructions for describing a task for the robot at runtime. Additionally or alternatively, in some implementations, the system can be provided with multiple instruction description types at runtime (e.g., natural language instructions, goal images, and task IDs can be provided at runtime, natural language instructions and task IDs can be provided at runtime, etc.). For example, the system can be trained based on a natural language instruction dataset and a goal image dataset, and the system can be provided with natural language instructions and / or goal image instructions at runtime.

[0005] In some embodiments, the robotic agent may achieve task - independent control using a goal - conditioned policy network, in which case a single robot can reach any reachable goal state in its environment. In traditional teleoperated multi - task demonstrations, the diversity of data collected may be restricted by a prior task definition (e.g., a human operator is provided with a list of tasks to perform the demonstration). In contrast, a human operator performing a teleoperated "play" is not restricted by a set of pre - defined tasks when generating play data. In some implementations, the goal image dataset may be generated based on teleoperated "play" data. Play data may include a continuous log (e.g., a data stream) of low - level observations and actions collected while a human remotely operates the robot and engages in behavior that satisfies their own curiosity. Collecting play data, unlike collecting expert demonstrations, may not require task segmentation, labeling, or resetting to an initial state, so it is possible to quickly collect large amounts of play data. Additionally or alternatively, play data may be structured based on human knowledge about object affordances (e.g., people tend to press a button when they see it in a scene). The human operator may try multiple ways to achieve the same result and / or explore new behaviors. In some implementations, play data may be expected to naturally subsume the interaction space of the environment in ways that are not possible with expert demonstrations.

[0006] In some implementations, the target image dataset can be generated based on remotely operated play data. A segment of the play data stream (e.g., a sequence of image frames) may be selected as an imitation trajectory, and the last image in the selected segment of the data stream is the target image. In other words, the target image that describes the imitation trajectory in the target image dataset may be generated with hindsight, as opposed to generating a sequence of actions based on the target image, where the target image is determined based on a sequence of actions. In some implementations, short-horizon target image training instances can be generated quickly and / or inexpensively based on a data stream of remotely operated play data.

[0007] In some implementations, additionally or alternatively, a natural language instruction dataset can be obtained based on remotely operated play data. A segment of the play data stream (e.g., a sequence of image frames) can be selected as an imitation trajectory. Then, one or more humans may describe the imitation trajectory, thus generating natural language instructions with hindsight (as opposed to generating an imitation trajectory based on natural language instructions). In some implementations, the natural language instructions collected can include functional behaviors (e.g., "open the drawer", "press the green button", etc.), behaviors not specific to a general task (e.g., "move the hand a little to the left", "do nothing", etc.), and / or additional behaviors. In some implementations, the natural language instructions may be in free-form natural language, and there are no constraints on the natural language instructions that can be provided. In some implementations, multiple humans can describe the imitation trajectory using free-form natural language, which may result in different descriptions of the same object, behavior, etc. For example, the imitation trajectory may capture a robot lifting a wrench. Multiple human describers may provide different free-form natural language instructions for the imitation trajectory, such as "grab the tool", "lift the wrench", "hold the object", and / or additional free-form natural language instructions. In some implementations, this diversity in free-form natural language instructions may lead to a more robust goal-conditioned policy network, in which case a wider range of free-form natural language instructions can be implemented by the agent.

[0008] The target-conditioned policy network and the corresponding encoder can be trained in various ways based on an image target dataset and a free-form natural language instruction dataset. For example, the system can use the target image encoder to process the target image portion of a target image training instance to generate a latent target space representation of the target image. The latent target space representation of the target image and the initial frame of the imitation trajectory portion of the target image training instance generate a target image candidate output. Based on the target image candidate output and the target image imitation trajectory, a target image loss can be generated. Similarly, the system can process the natural language instruction portion of a natural language instruction training instance to generate a latent space representation of the natural language instruction. The natural language instruction and the initial frame of the imitation trajectory portion of the natural language instruction training instance can be processed using the target-conditioned policy network to generate a natural language instruction candidate output. Based on the natural language instruction candidate output and the imitation trajectory portion of the natural language instruction training instance, a natural language instruction loss can be generated. In some implementations, the system can generate a target-conditioned loss based on the target image loss and the natural language instruction loss. One or more portions of the target-conditioned policy network, the target image encoder, and / or the natural language instruction encoder can be updated based on the target-conditioned loss. However, this is only an example of training the target-conditioned policy network, the target image encoder, and / or the natural language instruction encoder. Additional and / or alternative training methods may be used.

[0009] In some implementations, the target-conditioned policy network can be trained using target image datasets and natural language instruction datasets of different sizes. For example, the target-conditioned policy network can be trained based on a first amount of target image training instances and a second amount of natural language instruction training instances, where the second amount is 50 percent of the first amount, less than 50 percent of the first amount, less than 10 percent of the first amount, less than 5 percent of the first amount, less than 1 percent of the first amount, and / or more or less than an additional or alternative percentage of the first amount.

[0010] Accordingly, various implementations describe techniques for learning a shared latent goal space for a number of task descriptions for use in training a single goal-conditioned policy network. In contrast, conventional techniques train multiple policy networks, training one policy network per task description type. Training a single policy network enables more diverse data to be utilized when training the network. Additionally or alternatively, the policy network can be trained using a larger number of training instances of one data type. For example, a goal-conditioned policy network can be trained using hindsight goal image training instances that can be automatically generated from an imitation learning data stream (e.g., hindsight goal image training instances are inexpensive to automatically generate compared to natural language instruction training instances that may require natural language instructions provided by humans). By training the goal-conditioned policy network using both a goal image dataset and a natural language instruction dataset, most of the training instances become automatically generated goal image training instances, and the resulting goal-conditioned policy network can robustly generate actions for a robot based on natural language instructions without requiring computing resources (e.g., processor cycles, memory, power, etc.) and / or human resources (e.g., the time required by a group of people to provide natural language instructions).

[0011] The above description is provided only as an overview of some of the implementations disclosed in this specification. These and other implementations of the technology are disclosed in more detail below.

[0012] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter that appear at the end of this disclosure are considered to be part of the subject matter disclosed herein.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Modes for Carrying Out the Invention

[0014] Natural language is a versatile and intuitive way for humans to communicate tasks to robots. Existing methods enable learning diverse robot behaviors from general-purpose sensors. However, each task has to be defined using a target image, which is not realistic in an open-world environment. Instead, the implementations disclosed herein target a simple and / or scalable way to condition policies on human language. Short robot experiences from play can be paired post hoc with the relevant human language. To do this efficiently, some implementations utilize multi-context imitation, which can enable training a single agent to follow either image or language goals, where only language conditioning is used at test time. This can reduce the cost of language pairing to a small percentage (e.g., 10%, 5%, or less than 1%) of the collected robot experiences, with the majority of control still learned via self-supervised imitation learning. At test time, a single agent trained in this way can continuously execute a number of different robot manipulation skills defined directly from natural language only, in a 3D environment (e.g., open the drawer... pick up the block... press the green button). Additionally, some implementations use techniques to transfer knowledge from large unlabeled text corpora to robot learning. The transfer can significantly improve downstream robot manipulation. It can also enable an agent to follow thousands of new commands at test time in a zero-shot manner in multiple different languages, for example.

[0015] The long-term motivation for robotic learning is the idea of a generalist robot, which is a single agent that can solve many tasks in everyday environments using only common on-board sensors. Along with the generality of the task and observation spaces, an aspect that is fundamental but less considered is the specification of common tasks, namely that an untrained user can instruct the behavior of the agent using the most intuitive and flexible mechanisms. Thus, it is difficult to imagine a true generalist robot without imagining a robot that can follow instructions expressed in natural language.

[0016] More broadly, children learn language against the backdrop of rich and relevant sensorimotor experiences. This raises questions that have persisted for years about embodied language acquisition in artificial intelligence, namely how intelligent agents can ground perception based on language understanding. The ability to associate language with the physical world has the potential to enable communication on a common basis through the shared sensory experiences of robots and humans, which can lead to a much more meaningful form of human-machine interaction.

[0017] Furthermore, language acquisition can be a highly social process, at least in humans. Infants provide actions in early interactions, and caregivers provide the relevant words. The actual learning mechanism in humans in the field is not fully understood, but the implementations disclosed herein explore what a robot can learn from similar paired data.

[0018] However, even following simple instructions can pose notoriously difficult learning problems for AI, including many long-term problems. For example, a robot given the instruction "put the block away in the drawer" must be able to associate the language with low-level perception (what does the block look like? what is a drawer?). It must perform visual reasoning (what does it mean for the block to be inside the drawer?). Additionally, it must solve complex sequential decision problems (what commands should be sent to the arm to "put away"?). These questions barely scratch the surface of a single task, but note that the generalist robot environment requires a single agent to perform a number of tasks.

[0019] In some implementations, the free-form robot manipulation environment can be combined with free-form human language conditioning. Existing techniques typically involve restricted observation spaces, such as games, 2D grid worlds, simplified actuators, such as binary pick-and-place primitives, and synthetic language data. The implementations herein target combinations of 1) human language instructions, 2) high-dimensional continuous sensor inputs and actuators, and / or 3) complex tasks such as long-horizon robotic object manipulation. At test time, in some implementations, a single agent that can perform multiple tasks in sequence can be considered, where each task can be defined by a human in natural language. For example, "open the door all the way to the right... pick up the block... press the red button... close the door". Further, the agent should be able to execute any combination of subtasks in any order. This is sometimes referred to as the "ask me anything" scenario and can test general aspects such as multi-purpose control, learning from on-board sensors, and / or general task specification.

[0020] Existing techniques can provide a starting point for learning multi-purpose skills from on-board. However, like other methods that combine relabeling with image observations, existing techniques require that the task be defined using the target image to be reached. Although trivial in a simulator, this form of task definition can be unrealistic in an open-world environment.

[0021] In some implementations, the system can extend existing techniques to a natural language environment as follows.

[0022] (1) Handle space in a remotely operated play. In some implementations, the system can collect remotely operated "play" datasets. These long temporal state-action logs can be (automatically) relabeled into a number of short demonstrations, solving the image goal.

[0023] (2) Pair play with human language. Existing techniques typically pair commands with optimal behavior. In contrast, in some implementations described herein, behavior from play can be paired post hoc with optimal commands (i.e., hindsight instruction pairing). This can generate a demonstration dataset, solving the human language goal.

[0024] (3) Multi-context imitation learning. In some implementations, a single policy can be trained to solve image goals and / or language goals. Additionally or alternatively, in some implementations, only language conditioning is used at test time. To enable this, the system can utilize multi-context imitation learning. Multi-context imitation learning can be highly data-efficient. It reduces the cost of language pairing, for example, to less than a small percentage (e.g., less than 10%, less than 5%, less than 1%) of the collected robot experiences to enable language conditioning, while most of the control is still learned through imitation with self-supervised learning.

[0025] (4) Condition with human language at test time. In some implementations, at test time, the single policy trained in this manner can continuously execute many complex robot operation skills that are fully specified directly from the image using natural language.

[0026] Additionally or alternatively, some implementations include transfer learning from unlabeled text corpora to robot operations. Transfer learning enhancement can be used, which can be applicable to any language-conditioned policy. In some implementations, this can improve downstream robot operations. Importantly, this technique can enable the agent to follow new instructions zero-shot (e.g., follow thousands of new instructions and / or instructions across multiple languages). Goal-conditioned learning can be used to train a single agent to reach any goal. This can be formulated as a goal-conditioned policy π θ (a|s,g) that outputs the next action a∈A conditioned on the current state s∈S and the task descriptor g∈G. The imitation approach is based on a dataset of expert state-action trajectories τ={(s 0 ,a 0 ),...}

[0027]

Number

[0028] This association can be learned using supervised learning over a range, and solve paired task descriptors (such as one-hot task encoding). A convenient choice for the task descriptor is some target state g = s g ∈ S. This allows any state taken during collection to be relabeled as the "reached target state", and previous states and actions are treated as optimal behavior to reach that target. When applied to some original dataset D, this produces a much larger dataset of relabeled examples

[0029]

Number

[0030] that can produce, N R >> N, and provide input to a simple maximum likelihood objective for goal directed control, namely goal conditioned behavioral cloning (GCBC).

[0031]

Number

[0032] Relabeling can automatically generate a large number of goal-directed demonstrations during training, but this may not take into account the diversity of those demonstrations that can be completely derived from the underlying data. The ability to reach any goal provided by any user is a motivation for seeking data collection methods upstream of relabeling that handle the entire state space.

[0033] The collection of "play" remotely operated by humans can directly address the issue of the handling range of the state space. In this environment, the operator no longer has to be restricted by a predefined set of tasks. Instead, the operator can participate in any available object manipulation within the scene. The motivation is to handle the entire state space using prior human knowledge of object affordances. During the collection, the on-board robot observations and the flow of actions are recorded,

[0034]

Number

[0035] resulting in an unstructured but semantically useful dataset of unsegmented segments of behavior, which may be useful in the context of relabeled imitation learning.

[0036] Learning from play can combine relabeled imitation learning with remotely operated play. First, the unsegmented play logs are relabeled using Algorithm 2. This can produce a training set

[0037]

Number

[0038] that holds a number of diverse short-term examples. In some implementations, these can be fed into a standard maximum likelihood goal conditioned imitation objective.

[0039]

Number

[0040] A limitation of learning from play, and other approaches that combine relabeling with image state space, is that during test, the behavior is not consistent with the target image s g The challenge is that the task must be conditional on the human's command. Some implementations described herein may focus on a more flexible mode of conditioning, where a human describes the task in natural language. To succeed in this, complex underlying problems may need to be solved. To address this, hindsight command pairing, a method for pairing large amounts of diverse robot sensor data with relevant human language, may be used. In some implementations, multi-context imitation learning may be used to leverage both image target datasets and language target datasets. Additionally or alternatively, language learning from play (LangLfP) may be used to tie these components together and learn a single policy that follows multiple human commands over time.

[0041] From a statistical machine learning perspective, candidates for basing human language on robot sensor data are large corpora of robot sensor data paired with associated language. One way to collect this data is to pick instructions and then collect optimal behaviors. Additionally or alternatively, some implementations can take a sample of every robot's behavior from play and collect optimal instructions, which can be called hindsight instruction pairing (Algorithm 3). Just as a hindsight goal image is an after-the-fact answer to the question "Which goal state will make this trajectory optimal?", a hindsight instruction is an after-the-fact answer to the question "Which language instruction will make this trajectory optimal?". In some implementations, these pairs can be obtained by showing a human the on-board robot sensor video and then asking the human "What instructions would you give the agent to get from the first frame to the last frame?"

[0042] The hindsight instruction pairing process is playAccess to this can be assumed and can be obtained using Algorithm 2 and a group of non-expert human supervisors. D play From this, a new dataset

[0043] [Number]

[0044] can be created, which consists of short play sequences τ paired with l ∈ L, where l is afterthought instructions provided by humans without constraints on vocabulary and / or grammar.

[0045] In some implementations, this process may be scalable because the pairing occurs afterwards, and parallelization (e.g., via crowdsourcing) is simplified. The languages collected may also be rich, as they are associated with play and are not similarly restricted by prior task definitions. This can generate instructions for functional behaviors (e.g., "open the drawer", "press the green button"), as well as behaviors not specific to general tasks (e.g., "move your hand slightly to the left" or "do nothing"). In some implementations, it may not be necessary to pair every experience from play with language in order to learn to follow instructions. This may be made possible by the multi-context imitation learning described herein.

[0046] So far, methods for creating two context imitation datasets have been described, which hold examples of afterthought target images D play , and D( play,lang ) which hold examples of afterthought instructions. In some implementations, a single policy that does not depend on any task description can be trained. This can make it possible to share statistical strength across multiple datasets during training, and / or to use only language-based specifications during testing.

[0047] Motivated by this, some implementations use multi-context imitation learning (MCIL), which is a simple and / or widely applicable generalization of context imitation across multiple heterogeneous contexts. The main idea is to represent a large number of policies by a single unified function approximator that can generalize across states, tasks, and / or task descriptions. MCIL assumes access to a plurality of imitation learning datasets D = {D 0 ,..., D K} where the ways of describing the tasks are different for each. In some implementations, each

[0048]

Number

[0049] holds a pair of state-action trajectories τ paired with some context c ∈ C. For example, D 0 may contain demonstrations paired with a one-hot task id (conventional multi-task imitation learning dataset), D 1 may contain image goal demonstrations, and D 2 may contain language goal demonstrations.

[0050] Instead of training one policy per dataset, MCIL instead trains a single latent goal-conditioned policy π θ (a t |s t , z) simultaneously across all datasets, mapping each task description type to the same latent goal space

[0051]

Number

[0052] Learn to associate with this. This latent space can be regarded as a common abstract goal representation shared across a number of imitation learning problems. To enable this, MCIL has a set of parameterized encoders, one for each dataset, each of which is responsible for associating a specific type of task descriptor with a common latent goal space, i.e.,

[0053]

Number

[0054] to a set of parameterized encoders.

[0055]

Number

[0056] This can be assumed. For example, these can be, respectively, task id embedded lookup, image encoder, language encoder, one or more additional or alternative values, and / or combinations thereof.

[0057] In some implementations, MCIL has a simple training procedure. In each training step, for each dataset D k in D, a mini-batch of trajectory-context pairs (τ k , c k ) ~ D k is sampled, the context is encoded in the latent goal space

[0058]

Number

[0059] and a simple maximum likelihood contextual imitation objective function is calculated.

[0060]

Number

[0061] The complete MCIL objective function can average the per-dataset objective function across all datasets at each training step.

[0062]

Number

[0063] And the policy and all target encoders are trained end-to-end to maximize L MCIL . See Algorithm 1 for the pseudo-code of complete mini-batch training.

[0064] In some implementations, multi-context learning has properties that can be usefully extended beyond learning from play. Here, the dataset D can be set to D = {D play , D( play.lang )}, but this approach can be used more generally to train across any set of imitation datasets with various descriptions, such as task ids, languages, human video demonstrations, utterances, etc. Context independence can enable a highly efficient training regime. That is, while learning most of the control from the cheapest data sources, it learns the conditioning of the most general form of tasks from a small number of labeled examples. In this way, multi-context learning can be interpreted as transfer learning through a shared target space. This can reduce the cost of human supervision to the extent that it is realistically applicable. Multi-context learning can enable training agents to follow human instructions, where a small percentage (e.g., less than 10%, less than 5%, less than 1%, etc.) of the collected robot experiences require paired language, and most of the control is instead learned from relabeled target image data.

[0065] In some implementations, language-conditioned learning from play (LangLfP) is a special case of multi-context imitation learning. At a high level, LangLfP trains a single multi-context policy π play ,D( play,lang )} over a dataset D = {D θ (a t |s t ,z). In some implementations, F = {g enc ,s enc} can be the association of neural network encoders from image goals and instructions respectively to the same latent visuo-lingual goal space. LangLfP can learn recognition, natural language understanding, and control end-to-end without an auxiliary loss.

[0066] Recognition module. In some implementations, τ in each example consists of

[0067]

Number

[0068] , a sequence O of on-board observation results t , and actions. Each observation result can include high-dimensional images and / or measurements of internal proprioceptive sensors. The learned recognition module P θ associates each observation result tuple with a low-dimensional embedding, e.g., s t = P θ (O t ). This recognition module may be shared with g enc , which defines the top additional network for associating the encoded goal observations s g with points in the z space.

[0069] Language module. In some implementations, the language target encoder s enc tokenizes raw text l into subwords, retrieves subword embeddings from a lookup table, and / or then summarizes the embeddings into points in the z-space. The subword embeddings are randomly initialized at the beginning of training and can be end-to-end learned by the final imitation loss.

[0070] Control module. Many architectures can be used to implement the multi-context policy π θ (a t |s t ,z). For example, Latent Motor Plans (LMP) can be used. LMP is a goal-directed imitation architecture that uses latent variables to model a large amount of multi-modality specific to a free-form imitation dataset. Specifically, it can be a sequence-to-sequence conditional variational autoencoder (seq2seq CVAE) that auto-encodes contextual demonstrations through a latent "plan" space. The decoder is a goal-conditioned policy. As a CVAE, LMP can be easily adapted to the lower bound maximum likelihood contextual imitation in a multi-context environment.

[0071] LangLfP training. LangLfP training can be contrasted with existing LfP training. In each training step, a batch of image target tasks can be sampled from D play and a batch of language target tasks can be sampled from D( play,lang ). The recognition module P θ is used to encode observations into the state space. The image target and language target can be encoded into the latent target space z using the encoders g enc and s enc . The policy π θ (a t |s t, z) can be used to compute a multi-context imitation objective function averaged across both task descriptions. In some implementations, a combined gradient step can be taken for all modules, i.e., the recognition, language, and control modules, to optimize the entire architecture end-to-end as a single neural network.

[0072] Follow human instructions during testing. At the beginning of a test episode, the agent receives as input the on-board observation O t and the natural language goal l specified by the human. The agent uses the trained sentence encoder s enc to encode l in the latent goal space z. The agent then solves the goal in a closed loop, repeatedly feeding the current observation and the goal into the learned policy π θ (a t | s t , z) to sample actions and execute them in the environment. The human operator can type a new language goal l at any time.

[0073] Large "wild" natural language corpora can reflect a significant amount of human knowledge about the world. Many recent studies have succeeded in transferring this knowledge to downstream tasks in NLP via pre-trained embeddings. Can a similar knowledge transfer be achieved for robotic manipulation in some of the implementations described herein?

[0074] This type of transfer has many advantages. First, if there is a semantic match between the source corpus and the target environment, more structured inputs can serve as a powerful prior knowledge or reference. Additionally or alternatively, language embeddings have been shown to encode many word and text similarities. This can enable the agent to follow many new instructions zero-shot as long as they are "close enough" to the instructions the agent was trained to follow. Considering the complexity of natural language, it should be noted that robots in an open-world environment may sometimes have to be able to follow synonymous instructions outside the scope of a particular training set.

[0075] Algorithm 1 Multi-Context Imitation Learning Input:

[0076]

Number

[0077] , one dataset per context type (e.g., target image, language instruction, task id), each holding a (demonstration, context) pair. Input:

[0078]

Number

[0079] , one encoder per context type, a shared latent target space for the contexts, e.g.,

[0080]

Number

[0081] associated with. Input: π θ (a t |s t , z), a single latent target-conditioned policy. Input: Parameter

[0082]

Number

[0083] Initialize randomly while True do L MCIL ←0 # Loop over the dataset. for k = 0...K do # Sample a (demonstration, context) batch from this dataset. (τ k , c k ) ~ D k # Encode the context in the shared latent goal space.

[0084]

Number

[0085] # Accumulate the imitation loss.

[0086]

Number

[0087] end for # Average the gradients over the context types.

[0088]

Number

[0089] # Train the policy and all encoders end-to-end. L MCIL Update θ by taking a gradient step with respect to end while

[0090] Algorithm 2 Create millions of target image-conditioned imitation examples from remotely operated play. Input:

[0091]

Number

[0092] , the unsegmented stream of observations and actions recorded during play. Input: D play ←{} Input: w low , w high , the limit of the hindsight window size. while True do # Get the next play episode from the stream. (s 0:t , a 0:t ) ~ S for w = w low ... w high do for i = 0..(t - w) do # Select a window of each size w. τ = (s i:i+w , a i:i+w ) # Treat the last observation in the window as the target. s g = s w (τ, s g ) to D play Add end for end for end while

[0093] Algorithm 3 Pair robot sensor data with natural language commands. Input: D play , the relabeled play dataset holding (τ, s g ) pairs. Input: D (play,lang) ←{} Input: get_hindsight_instruction(): Provide a natural language instruction after the fact for a given τ to a human supervisor. Input: K, the number of pairs to generate, K << |D play |. for 0...K do # Sample a random trajectory from the play. (τ, ) ~ D play # Ask the human for an instruction to optimize τ.

[0094]

Number

[0095] (τ, l) to D( play.lang ) Add end for

[0096] Looking at the figure here, an exemplary robot 100 is shown in FIG. 1. The robot 100 is a "robot arm" having a plurality of degrees of freedom to allow passage of a gripping end effector 102 along any of a plurality of potential paths for positioning the gripping end effector 102 at a desired position. The robot 100 further controls two opposing "claws" of the gripping end effector 102 to actuate the claws between at least an open state and a closed state (and / or optionally a plurality of "partially closed" states).

[0097] The exemplary vision component 106 is also shown in FIG. 1. In FIG. 1, the vision component 106 is mounted in a fixed orientation relative to the base of the robot 100 or other non-moving reference point. The vision component 106 includes one or more sensors that can generate an image of an object in the line of sight of the sensor and / or other vision data regarding its shape, color, depth, and / or other characteristics. The vision component 106 can be, for example, a monographic camera, a stereographic camera, and / or a 3D laser scanner. The 3D laser scanner can be, for example, a time-of-flight 3D laser scanner or a triangulation-based 3D laser scanner, and can include a position detection sensor (PDS) or other optical position sensors.

[0098] The vision component 106 has a field of view of at least a portion of the working space of the robot 100, such as a portion of the working space that includes the exemplary object 104. The surfaces on which the objects 104 are placed are not shown in FIG. 1, but those objects may be placed on a table, tray, and / or other surface. The objects 104 can include a spatula, a stapler, and a pencil. In other implementations, more objects, fewer objects, additional objects, and / or alternative objects may be provided during all or some of the grasping attempts of the robot 100 as described herein.

[0099] Although a specific robot 100 is shown in FIG. 1, additional and / or alternative robots may be utilized, including additional robot arms similar to robot 100, robots having other robot arm configurations, robots having a humanoid configuration, robots having an animal configuration, robots that move via one or more wheels (e.g., robots that balance themselves), submarine robots, unmanned aerial vehicles (“UAVs”), and the like. Also, although a specific gripping end effector is shown in FIG. 1, additional and / or alternative end effectors may be utilized, such as alternative impactive gripping end effectors (e.g., those having a gripping “plate,” those having a greater or fewer number of “fingers” / “claws”), ingressive gripping end effectors, astrictive gripping end effectors, contigutive gripping end effectors, or non-gripping end effectors. In addition, although a specific mounting of vision component 106 is shown in FIG. 1, additional and / or alternative mountings may be utilized. For example, in some implementations, the vision component may be mounted directly on the robot to a non-operational component of the robot or an operational component of the robot (e.g., an end effector or a component near the end effector). Also, for example, in some implementations, the vision component may be mounted to a non-fixed structure separate from the associated robot and / or may be mounted in a manner that is not fixed to a structure separate from the associated robot.

[0100] Data from the robot 100 (e.g., vision data captured using the vision component 106) can be utilized by the action output engine 108 to generate an action output, together with the natural language instructions 130 captured using the user interface input device 128. In some implementations, the robot 100 can be controlled to perform one or more actions based on the action output (e.g., one or more actuators of the robot 100 can be controlled). In some implementations, the user interface input device 128 can include, for example, a physical keyboard, a touch screen (e.g., implementing a virtual keyboard or other text input mechanism), a microphone, and / or a camera. In some implementations, the natural language instructions 130 can be free-form natural language instructions.

[0101] In some implementations, the potential target engine 110 can use the natural language instruction encoder 114 to process the natural language instructions 130 and generate a potential state representation of the natural language instructions. For example, the keyboard user interface input device 128 can capture the natural language instruction "Press the green button". The potential target engine 110 can use the natural language instruction encoder 114 to process the natural language instruction 130 of "Press the green button" and generate a potential target representation of "Press the green button".

[0102] In some implementations, the target image training instance engine 126 can be used to generate a target image training instance 124 based on remotely operated "play" data 122. The remotely operated "play" data 122 may be generated by a human controlling a robot in an environment, and the human controller does not have a defined task to perform. In some implementations, each target image training instance 124 may include an imitation trajectory portion and a target image portion, and the target image portion describes the task of the robot. For example, the target image may be an image of a closed drawer, which may describe the action of the robot closing the drawer. As another example, the target image may be an image of an open drawer, which may describe the action of the robot opening the door. In some implementations, the target image training instance engine 126 can select a sequence of image frames from the remotely operated play data stream. The target image training instance engine 126 can generate one or more target image training instances by storing the selected sequence of image frames as the imitation trajectory portion of the training instance and storing the last image frame of the sequence of image frames as the target image portion of the training instance. In some implementations, the target image training instance 124 can be generated according to the process 400 of FIG. 4 described herein.

[0103] In some implementations, the natural language instruction training instance engine 120 can be used to generate natural language training instances 118 using teleoperated play data 122. The natural language instruction training instance engine 120 can select a sequence of image frames from a data stream of teleoperated play data 122. In some implementations, a human describer can provide natural language instructions that describe the tasks being performed by the robot in the selected sequence of image frames. In some implementations, multiple human describers can provide natural language instructions that describe the tasks being performed by the robot in the same selected sequence of image frames. Additionally or alternatively, multiple human describers can provide natural language instructions that describe the tasks being performed in separate sequences of image frames. In some implementations, multiple human describers can provide natural language instructions in parallel. The natural language instruction training instance engine 120 can generate one or more language instruction training instances by storing the selected sequence of image frames as an imitation trajectory portion of the training instance and storing the natural language instructions provided by humans as a natural language instruction portion of the training instance. In some implementations, the natural language training instance 124 can be generated according to the process 500 of FIG. 5 described herein.

[0104] In some implementations, the training engine 116 can be used to train the goal-conditioned policy network 112, the natural language instruction encoder 114, and / or the goal image encoder 132. In some implementations, the goal-conditioned policy network 112, the natural language instruction encoder 114, and / or the goal image encoder 132 can be trained according to the process 600 of FIG. 6 described herein.

[0105] Figure 2 shows an example of generating action output 208 according to various implementation forms. Example 200 includes receiving a natural language instruction input 202 (for example, receiving a natural language instruction input via one or more user interface input devices 128 in FIG. 1). In some implementation forms, the natural language instruction input 202 can be a free-form natural language input. In some implementation forms, the natural language instruction input 202 can be a text natural language input. The natural language instruction encoder 114 can process the natural language instruction input 202 to generate a latent target space representation of the natural language instruction 204. The target-conditioned policy network 112 can process the latent target 204 together with the current instance of the vision data 206 (for example, an instance of the vision data captured via the vision component 106 in FIG. 1) and can be used to generate the action output 208. In some implementation forms, the action output 208 can describe one or more actions for the robot to perform the task commanded by the natural language instruction input 202. In some implementation forms, one or more actuators of the robot (for example, the robot 100 in FIG. 1) can be controlled based on the action output 208 so that the robot performs the task indicated by the natural language instruction input 202.

[0106] Figure 3 is a flowchart showing a process 300 of generating an output according to the implementation forms disclosed herein based on a natural language instruction using a target-conditioned policy network when controlling a robot. For convenience, the operations of the flowchart are described with reference to the system that executes the operations. This system can include various components of various computer systems, such as one or more components of the robot 100, the robot 725, and / or the computing system 810. Moreover, although the operations of the process 300 are shown in a specific order, this is not intended to be limiting. One or more operations may be rearranged, omitted, and / or added.

[0107] In block 302, the system receives natural language instructions that describe a task for the robot. For example, the system can receive natural language instructions such as "Press the red button", "Close the door", "Lift the driver", and / or additional or alternative natural language instructions that describe the task to be performed by the robot.

[0108] In block 304, the system processes the natural language instructions using a natural language encoder to generate a latent space representation of the natural language instructions.

[0109] In block 306, the system receives an instance of vision data that captures at least a portion of the robot's environment.

[0110] In block 308, the system uses a goal-conditioned policy network to generate an output based on processing at least (a) an instance of vision data and (b) the latent goal representation of the natural language instructions.

[0111] In block 310, the system controls one or more actuators of the robot based on the generated output.

[0112] The process 300 of FIG. 3 is described in relation to controlling a robot based on natural language instructions. In additional or alternative implementations, the system can control the robot based on a target image, a task ID, an utterance, etc., instead of or in addition to natural language instructions. For example, the system can control the robot based on natural language instructions and target image instructions, where the natural language instructions are processed using a corresponding natural language instruction encoder and the target image is processed using a corresponding target image encoder.

[0113] FIG. 4 is a flowchart showing a process 400 for generating a target image training instance according to an implementation disclosed herein. For convenience, the operations of the flowchart are described with reference to a system that executes the operations. This system may include various components of various computer systems, such as one or more components of robot 100, robot 725, and / or computing system 810. Moreover, although the operations of process 400 are shown in a particular order, this is not intended to be limiting. One or more operations may be rearranged, omitted, and / or added.

[0114] In block 402, the system receives a data stream that captures remotely operated play data.

[0115] In block 404, the system selects a sequence of image frames from the data stream. For example, the system can select a 1-second sequence of image frames in the data stream, a 2-second sequence of image frames in the data stream, a 10-second sequence of image frames in the data stream, and / or a segment of additional or alternative length of image frames in the data stream.

[0116] In block 404, the system determines the last image frame in the selected sequence of image frames.

[0117] In block 406, the system stores a training instance that includes (1) the sequence of image frames as an imitation trajectory portion of the training instance and (2) the last image frame as a target image portion of the training instance. In other words, the system stores the last image as a target image that describes a task captured in the sequence of image frames.

[0118] In block 410, the system determines whether to generate additional training instances. In some implementations, the system can decide to generate additional training instances until one or more conditions are met. For example, the system can continue to generate training instances until a threshold number of training instances are generated, until the entire data stream is processed, and / or until additional or alternative conditions are met. If the system decides to generate additional training instances, the system returns to block 404, selects an additional sequence of image frames from the data stream, and performs additional iterations of blocks 406 and 408 based on the additional sequence of image frames. If it does not so decide, the process ends.

[0119] FIG. 5 is a flowchart showing a process 500 for generating natural language instruction training instances, in accordance with an implementation disclosed herein. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. This system can include various components of various computer systems, such as one or more components of robot 100, robot 725, and / or computing system 810. Moreover, although the operations of process 500 are shown in a particular order, this is not intended to be limiting. One or more operations may be rearranged, omitted, and / or added.

[0120] In block 502, the system receives a data stream that captures remotely operated play data.

[0121] In block 504, the system selects a sequence of image frames from the data stream. For example, the system can select a 1-second sequence of image frames in the data stream, a 2-second sequence of image frames in the data stream, a 10-second sequence of image frames in the data stream, and / or segments of additional or alternative lengths of image frames in the data stream.

[0122] In block 506, the system receives natural language instructions that describe tasks in the selected sequence of image frames.

[0123] In block 508, the system stores a training instance that includes (1) the sequence of image frames as the mimicking trajectory portion of the training instance and (2) the received natural language instructions that describe the task as the natural language instruction portion of the training instance.

[0124] In block 510, the system determines whether to generate additional training instances. In some implementations, the system can determine to generate additional training instances until one or more conditions are met. For example, the system can continue to generate training instances until a threshold number of training instances are generated, until the entire data stream is processed, and / or until additional or alternative conditions are met. If the system determines to generate additional training instances, the system returns to block 504, selects an additional sequence of image frames from the data stream, and performs additional iterations of blocks 506 and 508 based on the additional sequence of image frames. If it does not so determine, the process ends.

[0125] FIG. 6 is a flowchart showing a process 600 for training a target-conditioned policy network, a natural language instruction encoder, and / or a target image encoder, according to an implementation disclosed herein. For convenience, the operations of the flowchart are described with reference to a system that executes the operations. The system may include various components of various computer systems, such as one or more components of robot 100, robot 725, and / or computing system 810. Moreover, although the operations of process 600 are shown in a particular order, this is not intended to be limiting. One or more operations may be rearranged, omitted, and / or added.

[0126] In block 602, the system selects a target image training instance that includes (1) an imitation trajectory and (2) a target image.

[0127] In block 604, the system processes the target image using the target image encoder to generate a latent target space representation of the target image.

[0128] In block 606, the system uses the target-conditioned policy network to process at least (1) an initial image frame of the imitation trajectory and (2) the latent space representation of the target image to generate a candidate output.

[0129] In block 608, the system determines a target image loss based on (1) the candidate output and (2) at least a portion of the imitation trajectory.

[0130] In block 610, the system selects a natural language instruction training instance that includes (1) an additional imitation trajectory and (2) a natural language instruction.

[0131] In block 612, the system processes the natural language instruction portion of the natural language instruction training instance using the natural language encoder to generate a latent space representation of the natural language instruction.

[0132] In block 614, the system processes (1) an initial image frame of an additional imitation trajectory and (2) a latent space representation of a natural language instruction using a goal-conditioned policy network to generate an additional candidate output.

[0133] In block 616, the system determines a natural language loss based on (1) the additional candidate output and (2) at least a portion of the additional imitation trajectory.

[0134] In block 618, the system generates a goal-conditioned loss based on (1) an image goal loss and (2) a natural language instruction loss.

[0135] In block 620, the system updates one or more portions of the goal-conditioned policy network, the goal image encoder, and / or the natural language instruction encoder based on the goal-conditioned loss.

[0136] In block 622, the system determines whether to perform additional training on the goal-conditioned policy network, the goal image encoder, and / or the natural language instruction encoder. In some implementations, the system can determine to perform more training if there is one or more additional unprocessed training instances and / or if one or more other criteria have not yet been met. The one or more other criteria can include, for example, whether a threshold number of epochs have occurred and / or whether a threshold length of time of training has been performed. Process 600 can be trained using non-batch learning techniques, batch learning techniques, and / or both additional or alternative techniques. If the system determines to perform additional training, the system returns to block 602, selects an additional goal image training instance, performs additional iterations of blocks 604, 606, and 608 based on the additional goal image training instance, selects an additional natural language instruction training instance in block 610, performs additional iterations of blocks 612, 614, and 616 based on the additional natural language instruction training instance, and performs additional iterations of blocks 618 and 610 based on the additional goal image training instance and the additional natural language instruction training instance. If it does not so determine, the process ends.

[0137] FIG. 7 schematically shows an exemplary architecture of a robot 725. The robot 725 includes a robot control system 760, one or more motion components 740a - 740n, and one or more sensors 742a - 742m. The sensors 742a - 742m can include, for example, vision components, light sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors, accelerometers, gyroscopes, thermometers, barometers, etc. The sensors 742a - m are shown as being integral with the robot 725, but this is not intended to be limiting. In some implementations, the sensors 742a - m can be located external to the robot 725, for example, as stand-alone units.

[0138] The motion components 740a - n can include, for example, one or more end effectors and / or one or more servo motors or other actuators to effect the movement of one or more components of the robot. For example, the robot 725 may have multiple degrees of freedom, and each of the actuators may control the operation of the robot 725 within one or more of the degrees of freedom in response to control commands. As used herein, the term actuator includes, in addition to any driver that may be associated with the actuator and that converts a received control command into one or more signals for driving the actuator, a mechanical or electrical device (e.g., a motor) that produces motion. Thus, providing a control command to an actuator can include providing the control command to a driver that converts the control command into appropriate signals for driving an electrical or mechanical device to produce the desired motion.

[0139] The robot control system 760 can be implemented in one or more processors such as the CPU, GPU, and / or other controllers of the robot 725. In some implementations, the robot 725 can include a “brain box” that can include all or a plurality of aspects of the control system 760. For example, the brain box may provide a real - time burst of data to the motion components 740a - n, and each real - time burst can include a set of one or more control commands that, inter alia, define (if any) the motion parameters for one or more of the motion components 740a - n. In some implementations, the robot control system 760 can execute one or more aspects of the processes 300, 400, 500, 600, and / or other methods described herein.

[0140] As described herein, in some implementations, all or some aspects of the control commands generated by the control system 760 when positioning the end effector to grasp an object can be based on the end effector commands generated using a target-conditioned policy network. For example, the vision components of sensors 742a - m can capture environmental state data. This environmental state data, along with the robot state data, can be part of a process that uses the policy network of a meta-learning model to generate one or more end effector control commands for controlling movement and / or for gripping the end effector of the robot. The control system 760 is shown in FIG. 7 as an integral part of the robot 725, but in some implementations, all or some aspects of the control system 760 can be implemented in a component that is separate from but communicates with the robot 725. For example, all or some aspects of the control system 760 can be implemented in one or more computing devices that communicate with the robot 725, such as a wired and / or wireless communication with the computing device 810.

[0141] FIG. 8 is a block diagram of an exemplary computing device 810 that can optionally be utilized to execute one or more aspects of the techniques described herein. The computing device 810 typically includes at least one processor 814 that communicates with several peripheral devices via a bus subsystem 812. These peripheral devices can include, for example, a storage subsystem 824 that includes a memory subsystem 825 and a file storage subsystem 826, a user interface output device 820, a user interface input device 822, and a network interface subsystem 816. The input and output devices enable user interaction with the computing device 810. The network interface subsystem 816 provides an interface to an external network and is coupled to a corresponding interface device in other computing devices.

[0142] The user interface input device 822 can include a keyboard, a mouse, a trackball, a touchpad, or a pointing device such as a graphics tablet, a scanner, a touch screen incorporated in a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 810 or a communication network.

[0143] The user interface output device 820 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem can also provide a non-visual display via an audio output device or the like. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 810 to a user or another machine or computing device.

[0144] The storage subsystem 824 stores the programming and data constructs that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 824 can include logic for performing selected aspects of the processes of FIGS. 3, 4, 5, 6, and / or other methods described herein.

[0145] These software modules are generally executed by processor 814, either alone or in combination with other processors. Memory 825 used in storage subsystem 824 may include several memories, including main random access memory (RAM) 830 for storing instructions and data during program execution and read-only memory (ROM) 832 in which fixed instructions are stored. File storage subsystem 826 can perform persistent storage of program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of some implementations may be stored in file storage subsystem 826 within storage subsystem 824 or by other machines accessible by processor 814.

[0146] Bus subsystem 812 provides a mechanism for enabling the various components and subsystems of computing device 810 to communicate with each other as intended. Although bus subsystem 812 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0147] Computing device 810 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computing device 810 shown in FIG. 8 is intended only as a specific example for the purpose of illustrating some implementations. Many other configurations of computing device 810 are possible with more or fewer components than the computing device shown in FIG. 8.

[0148] In situations where the systems described in this specification collect or may use personal information about a user (or often referred to herein as a "participant"), the user may be provided with the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographical location), and / or to control whether and / or how content related to the user is received from a content server. Alternatively, since certain data may be treated in one or more ways before being stored or used, personally identifiable information may be removed. For example, the user's identifying information may be treated so that personally identifiable information about the user cannot be determined, or so that the user's geographical location is generalized when geographical location information is obtained (such as to the city level, zip code level, or state level) so that the user's specific geographical location cannot be determined. Thus, the user can manage how information about the user is collected and / or used.

[0149] In some implementations, a method implemented by one or more processors is provided, the method including receiving free-form natural language instructions that describe a task for a robot, and free-form natural language instructions generated based on user interface inputs provided by a user via one or more user interface input devices. In some implementations, the method includes processing the free-form natural language instructions using a natural language instruction encoder to generate a latent target representation of the free-form natural language instructions. In some implementations, the method includes receiving an instance of vision data, the instance of vision data being generated by at least one vision component of the robot, the instance of vision data capturing at least a portion of the environment of the robot. In some implementations, the method includes using a target-conditioned policy network to generate an output based on at least (a) the instance of vision data and (b) the latent target representation of the free-form natural language instructions, the target-conditioned policy network being trained based on at least (i) a set of target images of training instances such that training tasks are described using the target images, and (ii) a set of natural language instructions of training instances such that training tasks are described using the free-form natural language instructions. In some implementations, the method includes controlling one or more actuators of the robot based on the generated output, controlling one or more actuators of the robot causing the robot to perform at least one action indicated by the generated output.

[0150] These and other implementations of the technology disclosed herein may include one or more of the following features.

[0151] In some implementations, the method includes receiving additional free-form natural language instructions that describe additional tasks for the robot, the additional free-form natural language instructions being generated based on additional user interface inputs provided by a user via one or more user interface input devices. In some implementations, the method includes using a natural language instruction encoder to process the additional free-form natural language instructions to generate additional potential target representations of the additional free-form natural language instructions. In some implementations, the method includes receiving additional instances of vision data generated by at least one vision component of the robot. In some implementations, the method includes using a goal-conditioned policy network to generate additional output based on processing at least (a) the additional instances of vision data and (b) the additional potential target representations of the additional free-form natural language instructions. In some implementations, the method includes controlling one or more actuators of the robot based on the generated additional output, and controlling one or more actuators of the robot causes the robot to perform at least one additional action indicated by the generated additional output.

[0152] In some implementations, the additional tasks for the robot are separate from the tasks for the robot.

[0153] In some implementations, for each training instance in a set of target images of training instances, where the training task is described using the target images, the training instance includes an imitation trajectory provided by a human and a target image that describes the training task to be performed by the robot in the imitation trajectory. In some implementations, generating each training instance in a set of target images of training instances includes receiving a data stream that captures the state of the robot and the corresponding actions of the robot while a human is controlling the robot to interact with the environment. In some implementations, the method includes, for each training instance in a set of target images of training instances, selecting a sequence of image frames from the data stream, selecting the last image frame in the sequence of image frames as a training target image that describes the training task to be performed in the sequence of image frames, and generating a training instance by storing the selected sequence of image frames as an imitation trajectory portion of the training instance and storing the training target image as a target image portion of the training instance.

[0154] In some implementations, each training instance in a set of natural language instructions for training instances, where the training is described using free-form natural language instructions, includes an imitation trajectory provided by a human and free-form natural language instructions that describe a training task to be performed by a robot in the imitation trajectory. In some implementations, generating each training instance in a set of natural language instructions for training instances includes receiving a data stream that captures the state of the robot and the corresponding actions of the robot while a human is controlling the robot to interact with the environment. In some implementations, the method includes, for each training instance in a set of natural language instructions for training instances, selecting a sequence of image frames from the data stream, providing the sequence of image frames to a human evaluator, receiving free-form training natural language instructions that describe a training task to be performed by the robot in the sequence of image frames, and generating a training instance by storing the selected sequence of image frames as the imitation trajectory portion of the training instance and the free-form training natural language instructions as the free-form natural language instruction portion of the training instance.

[0155] In some implementations, the goal-conditioned policy network selects a first training instance from a set of target images of training instances based on at least (i) a set of target images of training instances such that the training task is described using the target images, and (ii) a set of natural language instructions of training instances such that the training task is described using free-form natural language instructions. The first training instance includes a first imitation trajectory and a first target image describing the first imitation trajectory. In some implementations, the method includes generating a latent space representation of the first target image by processing the first target image portion of the first training instance using a target image encoder. In some implementations, the method includes processing, using the goal-conditioned policy network, at least (1) an initial image frame in the first imitation trajectory and (2) the latent space representation of the first target image portion of the first training instance to generate a first candidate output. In some implementations, the method includes determining a target image loss based on the first candidate output and one or more portions of the first imitation trajectory. In some implementations, the method includes selecting a second training instance from the set of natural language instructions of the training instances, where the second training instance includes a second imitation trajectory and a second free-form natural language instruction describing the second imitation trajectory. In some implementations, the method includes generating a latent space representation of the second free-form natural language instruction by processing the second free-form natural language instruction portion of the second training instance using a natural language encoder, where the latent space representation of the first target image and the latent space representation of the second free-form natural language instruction are represented in a shared latent space. In some implementations, the method includes processing, using the goal-conditioned policy network, at least (1) an initial image frame in the second imitation trajectory and (2) the latent space representation of the second free-form natural language instruction portion of the second training instance to generate a second candidate output. In some implementations, the method includes determining a natural language instruction loss based on the second candidate output and one or more portions of the second imitation trajectory.In some implementations, the method includes determining a target conditional loss based on a target image loss and a natural language instruction loss. In some implementations, the method includes updating one or more parts of a target image encoder, a natural language instruction encoder, and / or a target conditional policy network based on the determined target conditional loss.

[0156] In some implementations, the target conditional policy network is trained based on a first amount of training instances of a target image set of training instances and a second amount of training instances of a natural language instruction set of training instances, where the second amount is less than 50 percent of the first amount. In some implementations, the second amount is less than 10 percent of the first amount, less than 5 percent of the first amount, or less than 1 percent of the first amount.

[0157] In some implementations, the generated output includes a probability distribution over a robot's action space, and controlling one or more actuators based on the generated output comprises selecting at least one action based on the at least one action having the highest probability in the probability distribution.

[0158] In some implementations, generating an output based on using a target conditional policy network to process at least (a) an instance of vision data and (b) a latent target representation of a free-form natural language instruction further includes generating an output based on using the target conditional policy network to process (c) at least one action, and controlling one or more actuators based on the generated output comprises selecting the at least one action based on the at least one action meeting a threshold probability.

[0159] In some implementations, a method implemented by one or more processors is provided, the method including receiving free-form natural language instructions that describe a task for a robot, the free-form natural language instructions being generated based on user interface inputs provided by a user via one or more user interface input devices. In some implementations, the method includes processing the free-form natural language instructions using a natural language instruction encoder to generate a latent target representation of the free-form natural language instructions. In some implementations, the method includes receiving an instance of vision data, the instance of vision data being generated by at least one vision component of the robot and capturing at least a portion of the environment of the robot. In some implementations, the method includes generating an output based on using a goal-conditioned policy network to process at least (a) the instance of vision data and (b) the latent target representation of the free-form natural language instructions. In some implementations, the method includes controlling one or more actuators of the robot based on the generated output, controlling one or more actuators of the robot causing the robot to perform at least one action indicated by the generated output. In some implementations, the method includes receiving goal image instructions that describe an additional task for the robot, the goal image instructions being provided by a user via one or more user interface input devices. In some implementations, the method includes processing the goal image instructions using a goal image encoder to generate a latent target representation of the goal image instructions. In some implementations, the method includes receiving an additional instance of vision data, the additional instance of vision data being generated by at least one vision component of the robot and capturing at least a portion of the environment of the robot.In some implementations, the method includes generating additional output based on using a goal-conditioned policy network to process at least (a) additional instances of vision data and (b) latent goal representations of goal images. In some implementations, the method includes controlling one or more actuators of a robot based on the generated additional output, and controlling one or more actuators of the robot causes the robot to perform at least one additional action indicated by the generated additional output.

[0160] In some implementations, a method implemented by one or more processors is provided, the method including the step of selecting a first training instance from a set of target images of a training instance, the first training instance including a first imitation trajectory and a first target image describing the first imitation trajectory. In some implementations, the method includes the step of generating a latent space representation of the first target image by processing a first target image portion of the first training instance using a target image encoder. In some implementations, the method includes the step of processing at least (1) an initial image frame in the first imitation trajectory and (2) a latent space representation of the first target image portion of the first training instance using a target-conditioned policy network to generate a first candidate output. In some implementations, the method includes the step of determining a target image loss based on the first candidate output and one or more portions of the first imitation trajectory. In some implementations, the method includes the step of selecting a second training instance from a set of natural language instructions of a training instance, the second training instance including a second imitation trajectory and a second free-form natural language instruction describing the second imitation trajectory. In some implementations, the method includes the step of generating a latent space representation of the second free-form natural language instruction by processing a second free-form natural language instruction portion of the second training instance using a natural language encoder, the latent space representation of the first target image and the latent space representation of the second free-form natural language instruction being represented in a shared latent space. In some implementations, the method includes the step of processing at least (1) an initial image frame in the second imitation trajectory and (2) a latent space representation of the second free-form natural language instruction portion of the second training instance using a target-conditioned policy network to generate a second candidate output. In some implementations, the method includes the step of determining a natural language instruction loss based on the second candidate output and one or more portions of the second imitation trajectory. In some implementations, the method includes the step of determining a target-conditioned loss based on the target image loss and the natural language instruction loss.In some implementations, the method includes updating one or more parts of a target image encoder, a natural language instruction encoder, and / or a target-conditioned policy network based on a determined target-conditioned loss.

[0161] In addition, some implementations include one or more processors (e.g., a central processing unit (CPU)), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) of one or more computing devices, where the one or more processors are operable to execute instructions stored in associated memory, and the instructions are configured to cause execution of any of the methods described herein. Some implementations also include one or more temporary or non-temporary computer-readable storage media storing computer instructions executable by one or more processors to execute any of the methods described herein.

Description of Reference Numerals

[0162] 100 Robot 102 Gripping End Effector 104 Object 108 Action Output Engine 110 Potential Target Engine 112 Target-Conditioned Policy Network 114 NL Instruction Encoder 116 Training Engine 118 NL Instruction Training Instance 120 NL Instruction Training Instance Engine 122 Remotely Operated “Play” Data 124 Target Image Training Instance 126 Target Image Training Instance Engine 128 User Interface Input Device 130 Natural Language Instruction 202 Natural Language Instruction Input 204 Potential Target Current instance of 206 vision data 208 Action output 725 Robot 740 Motion component 742 Sensor 760 Robot control system 810 Computing system 812 Bus subsystem 814 Processor 816 Network interface 820 User interface output device 822 User interface input device 824 Storage subsystem 825 Memory subsystem 826 File storage subsystem

Claims

1. A method implemented by one or more processors, comprising: receiving free-form, natural language instructions describing a task for a robot, the free-form, natural language instructions being generated based on user interface inputs provided by a user via one or more user interface input devices; processing the free-form natural language instruction using a natural language instruction encoder to generate a latent target representation of the free-form natural language instruction; receiving an instance of vision data, the instance of vision data being generated by at least one vision component of the robot, the instance of vision data capturing at least a portion of an environment of the robot; generating an output based on processing at least (a) the instances of vision data and (b) the latent goal representations of the free-form natural language instructions using a goal-conditional policy network; controlling one or more actuators of the robot based on the generated output, where controlling the one or more actuators of the robot causes the robot to perform at least one behavior indicated by the generated output; receiving a goal image describing an additional task for the robot; processing the target image using a target image encoder to generate a latent target representation of the target image; receiving an additional instance of vision data, the additional instance of vision data being generated by the at least one vision component of the robot, the additional instance of vision data capturing at least a portion of the environment of the robot; generating additional outputs using the goal-conditional policy network based on processing at least (a) the additional instances of vision data and (b) the latent goal representations of the goal images; controlling the one or more actuators of the robot based on the generated additional output, wherein controlling the one or more actuators of the robot causes the robot to perform at least one additional behavior indicated by the generated additional output. method.

2. The method described in claim 1, wherein the additional task for the robot is separate from the task for the robot.

3. The method of claim 1, wherein the generated output comprises a probability distribution over the robot's action space, and controlling the one or more actuators based on the generated output comprises selecting the at least one action based on the at least one action having the highest probability in the probability distribution.

4. The method described in claim 3, wherein the generated additional outputs include an additional probability distribution over the action space of the robot, and controlling the one or more actuators based on the generated additional outputs includes selecting the at least one additional action based on the at least one additional action having the highest probability in the additional probability distribution.

5. The method of claim 1, wherein the target image is provided by the user via the one or more user interface input devices.

Citation Information

Patent Citations

  • Controlling a robot based on free-form natural language input

    WO2019183568A1