Natural Language Control of Robots

By using a language model to process natural language inputs and grounding the outputs with task and world grounding metrics, the solution enables robots to effectively execute tasks in response to free-form inputs, addressing the challenge of real-world experience and environmental grounding.

JP2025517823AActive Publication Date: 2025-06-11GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024557754
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-31
Filing Date
2023-03-30
Publication Date
2025-06-11
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

Existing robots struggle to execute specific tasks in response to various free-form natural language inputs due to a lack of real-world experience and grounding in the current environment.

Method used

A language model is used to process natural language inputs and generate outputs that reflect other content, while grounding the outputs using techniques that ensure the selected robot skills are executable and contextually appropriate, considering both task and world grounding metrics.

Benefits of technology

Enables robots to successfully execute tasks in response to free-form natural language inputs by ensuring the selected skills are both likely to complete the task and suitable for the current environment, thereby improving the robot's ability to understand and act upon high-level instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517823000001_ABST
    Figure 2025517823000001_ABST
Patent Text Reader

Abstract

Embodiments use a large language model to process free-form natural language (NL) instructions and generate an LLM output. Those embodiments generate a task grounding measure that reflects the probability of the skill description in the probability distribution of the LLM output, based on the LLM output and the NL skill description of the robot skills. Those embodiments further generate a world grounding measure that reflects the probability that the robot skill is successful, based on the current environmental state data, based on the robot skills and the current environmental state data. Those embodiments further determine whether to execute the robot skill, based on both the task grounding measure and the world grounding measure.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Many robots are programmed to perform a specific task. For example, a robot on an assembly line can be programmed to recognize certain objects and perform specific operations on those certain objects.

[0002] Furthermore, some robots can execute a specific task in response to an input of an explicit user interface corresponding to a specific task. For example, a cleaning robot can execute a general cleaning task in response to the utterance "Robot, clean". However, usually, the input of the user interface for causing a robot to execute a task must be explicitly mapped to that task. Therefore, there are cases where a robot cannot execute a specific task in response to various free-form natural language inputs of a user who attempts to control the robot.

Summary of the Invention

[0003] A language model (LM) has been developed that can be used to process natural language (NL) content and / or other input(s) and generate an LM output that reflects other content in response to the NL content and / or input(s). For example, large language models (LLMs) have been developed that are trained on large amounts of data to robustly process a wide range of NL inputs and generate corresponding LM outputs that accurately reflect the corresponding NL content in response to the NL inputs. An LLM can include at least several hundred million parameters and can often include at least several billion parameters, such as over 100 billion parameters. An LLM can be, for example, a sequence-to-sequence model, based on a Transformer, and / or can include an encoder and / or a decoder. One non-limiting example of an LLM is Google's Pathways Language Model (PaLM). Another non-limiting example of an LLM is Google's Language Model for Dialogue Applications (LaMDA).

[0004] Separately, efforts have been made to enable robust free-form (FF) NL control of robots. For example, it is to enable a robot to appropriately respond to various different types of, or any one of, oral instructions from a human directed at the robot. For example, in response to an FF NL instruction such as "Put the block into the toy box", it is to enable the execution of a robot task including (a) navigating to the "block", (b) picking up the "block", (c) navigating to the "toy box", and (d) placing the "block" into the "toy box".

[0005] The embodiments disclosed herein recognize that LMs such as LLMs can encode rich semantic knowledge about the world and that such knowledge can be useful to a robot when acting based on high-level, temporally extended instructions expressed in FF NL instructions. The embodiments disclosed herein further recognize that a drawback of LLMs is that they lack real-world experience such as real-world experience in robot control and / or they lack grounding to the current real-world state(s). For example, by prompting an LLM with "I spilled a drink on the table, please help", a result can be obtained where an LLM output is generated that reflects NL content describing appropriate steps (if any) for cleaning up the spill. For example, the most probable decoding of the LLM output may reflect NL content such as "How about trying using a vacuum cleaner". However, "How about trying using a vacuum cleaner" may not be applicable to a particular agent such as a robot that needs to perform this task in a particular environment. For example, the robot may not have an integrated vacuum cleaner and there may be no separate vacuum cleaner available for use in a particular environment or the robot may not be able to control a separate vacuum cleaner in a particular environment.

[0006] The embodiments disclosed in this specification describe tasks at a high level and utilize LM output to determine how to control a robot to execute a task in response to FF NL input that does not describe all (or any) of the robot skill(s) necessary to execute the task. However, in view of the recognized drawback(s) of the LM, the embodiments described herein ground the LM output using one or more techniques such that the robot skill(s) selected based on the LM output for execution in response to the FF NL input are both executable (e.g., executable by the robot) and contextually appropriate (e.g., likely to succeed if executed by the robot in a particular environment). More specifically, the embodiments ground the LM output considering not only the LM output but also robot skills executable by the robot such as pre-trained robot skills.

[0007] As some non-limiting examples of the operation of those embodiments, assume that a user gives an FF NL command of "I spilled a drink on the table, please help." An LLM prompt can be generated based on the FF NL command. For example, the LLM prompt can strictly conform to the FF NL command. As another example, as described herein, the LLM prompt may be based on the FF NL command but may not strictly conform. For example, the prompt may include some or all of the terms of the FF NL command, but additionally, a scene descriptor(s) of the current environment (e.g., an NL descriptor(s) of the object(s) detected in the environment), an explanation generated in a previous pass using the LLM (e.g., based on a previous LLM prompt such as "explain how to help when the user says 'I spilled a drink on the table, please help'"), and / or terms (e.g., including "I will do 1." at the end of the prompt) that prompt the prediction of step(s) by the LLM.

[0008] The generated LLM prompt can be processed using the LLM to generate LLM output that models the probability distribution of candidate word compositions that depend on the instructions. Continuing with the example operation, the most probable decoding of the LLM output may be, for example, "using a vacuum cleaner". However, the embodiments disclosed herein do not simply blindly utilize the probability distribution of the LLM output when determining a method for controlling a robot. Rather, the embodiments consider robot skills that can actually be executed by the robot, such as dozens, hundreds, or thousands of pre-trained robot skills, while leveraging the probability distribution of the LLM output. Some versions of those embodiments generate a corresponding task grounding scale and a corresponding world grounding scale for each robot skill considered. Further, some of those versions select a particular robot skill to perform in response to the FF NL prompt based on considering both the corresponding task grounding scale and the corresponding world grounding scale. For example, the corresponding overall scale may be generated for each of the robot skills as a function of the corresponding world grounding and task grounding scales of the robot skill, and a particular robot skill may be selected based on having the best corresponding overall scale.

[0009] When generating a task grounding metric for a robotic skill, the NL skill description of the robotic skill can be compared with the LLM output to generate a task grounding metric. For example, the task grounding metric can reflect the probability of the NL skill description in the probability distribution modeled by the LLM output. For example, continuing with the example of an operation, the task grounding metric of a first robotic skill with the NL skill description of "pick up a sponge" can reflect a higher probability than the task grounding metric of a second robotic skill with the NL skill description of "pick up a banana". In other words, the probability distribution of the LLM output can characterize that the probability of "pick up a sponge" is higher than that of "pick up a banana". Also, for example, continuing with the example of an operation, the task grounding metric of a third robotic skill with the NL skill description of "pick up a squeegee" can reflect a probability similar to that of the first robotic skill.

[0010] When generating a world grounding metric for a robotic skill, current state data such as environmental state data and / or robotic state data can optionally be considered. Environmental state data reflects the state of one or more objects (multiple possible) that exist in the environment in addition to the robot, and can include, for example, visual sensor data and / or decisions made by processing visual sensor data (e.g., object detection(s) and / or classification(s)). Robotic state data reflects the state of one or more components of the robot and can include, for example, the current position of the robot's component(s), the current velocity of the robot's component(s), and / or other state data.

[0011] In some embodiments, or for some robot skills, the description of the robot skill may also be considered. In some of those embodiments, the description of the robot skill (e.g., its word embedding) is processed, and the current state data is processed using the trained value function model to generate a value that reflects the probability that the robot skill is successful based on the current state data. The world grounding scale is generated based on (e.g., conforms to) that value. In some versions of those embodiments, candidate robot actions are also processed using the trained value function model together with the description and the current state data. In some of those versions, multiple values are generated for a given robot skill, each generated based on processing different candidate robot actions but using the same description and the same current state data. In those versions, the world grounding scale may be generated based on, for example, the generated value that reflects the highest success probability.

[0012] Continuing with the example operation, assume that the current environmental state data includes an image captured by the robot's camera, and that the image captures a nearby sponge and a nearby banana but not the squeegee. In such a situation, the world grounding scale for a first robot skill with the NL skill description "pick up the sponge" may reflect a high probability, the world grounding scale for a second robot skill with the NL skill description "pick up the banana" may also reflect a high probability, and the world grounding scale for a third robot skill with the NL description "pick up the squeegee" may reflect a low probability.

[0013] Continuing with the operation example, the first robot skill with the NL skill description of "pick up the sponge" can be selected as the implementation target based on considering both its task grounding scale and its world grounding scale, both of which reflect a high probability. It should be noted that the second robot skill is not selected because it has a low task grounding scale despite having a high world grounding scale. Similarly, the third robot skill is not selected because it has a low world grounding scale despite having a high task grounding scale.

[0014] In these and other ways, the embodiments disclosed herein can consider both the task grounding scale and the world grounding scale when selecting the robot skills to be implemented. This ensures that the selected robot skill is both (a) likely to lead to the successful completion of the task reflected by the FF NL input (as reflected by the task grounding scale), and (b) likely to succeed when implemented by the robot in the current environment (as reflected by the world grounding scale).

[0015] The above description is provided as an overview of only some of the embodiments disclosed herein. These and other embodiments are described in more detail herein, including the forms for implementing the invention and the claims.

[0016] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are intended to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter that appear at the end of this disclosure are intended to be part of the subject matter disclosed herein.

Brief Description of the Drawings

[0017]

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Figure 6

DETAILED DESCRIPTION OF THE INVENTION

[0018] Before turning to the detailed description of the drawings, a non-limiting overview of various embodiments is provided.

[0019] In various embodiments disclosed herein, the robot has a repertoire of learned robot skills for atomic behaviors that enable low-level visuomotor control. Some of those embodiments not only prompt the LLM to interpret just the FF NL high-level instructions, but also use the LLM output generated by the prompting to generate a task grounding metric that quantifies for each the likelihood that the corresponding robot skill will progress toward completing the high-level instruction. Further, a corresponding affordance function (e.g., a learned value function) for each robot skill can be used to generate a world grounding metric for the robot skill that quantifies the likelihood that the robot skill will succeed from the current state. Further, both the task grounding metric and the world grounding metric can be used in determining which next robot skill to execute in achieving the task(s) reflected by the high-level instruction. In these and other ways, embodiments leverage that the LLM output describes the probability that each skill contributes to completing the instruction, the affordance function describes the probability that each skill succeeds, and combining these two gives the probability that each skill will succeed in executing the instruction. The affordance function allows consideration of real-world grounding in addition to the task grounding of the LLM output, and by constraining completion of the skill description, allows consideration of the LLM output in a way that recognizes the robot's capabilities (e.g., in terms of the repertoire of learned robot skills). Further, from this combination, a fully explainable sequence of steps that the robot executes to complete the high-level instruction (e.g., a description of the robot skill(s) selected for execution on the subject), i.e., an interpretable plan expressed through language, is obtained.

[0020] Accordingly, embodiments utilize an LLM to provide task grounding for determining useful actions for high-level goals and utilize an affordance function (e.g., a learned function(s)) to provide world grounding for determining which action(s) are actually executable to achieve the high-level goal. Some of those embodiments utilize reinforcement learning (RL) as a way to learn a language-conditioned value function that provides an affordance of what is possible in the real world.

[0021] A language model attempts to model the probability p(W) of a text W={w 0 ,w 1 ,w 2 ,···,w n}, which is a sequence of strings w. This is typically done by factoring the probability through the chain rule such that

Number

[0022] The embodiments disclosed herein utilize the vast semantic knowledge contained in an LLM to determine useful tasks for solving high-level instructions. Some of those embodiments further attempt to accurately predict whether a skill (given by an NL descriptor) is executable in the current state of the environment of a robot or other agent. In some versions of these embodiments, or at least for some skills, time-difference (TD)-based reinforcement learning is utilized to achieve this goal. In some of those versions, a Markov decision process is defined as M=(S,A,P,R,γ), where S and A are the state space and action space,

Number

Number

Number

Number

[0023] For example, TD-based methods may be used to learn a value function such as a value function that additionally conditions on natural language descriptors of skills, and the value function can be utilized to determine whether a given skill is executable from a given state. Notably, in the case of undiscounted sparse rewards, the agent receives a reward of 1.0 at the end of the episode if successful and 0.0 otherwise, and the value function trained by RL corresponds to an affordance function that specifies whether a skill is possible in a given state.

[0024] The various embodiments disclosed herein receive a user-provided natural language instruction i that describes a task for a robot to perform. The instruction may be long, abstract, and / or ambiguous. Further, the embodiments can utilize a set of robot skills II, where each skill π ∈ II performs a short task such as picking up a particular object and has a short language description l π (e.g., "find the sponge") and an affordance function p(c π │s,l π ). The affordance function indicates the probability of successfully completing a skill with description l from state s. Intuitively, p(c π │s,l π │s,l π ) means "if I ask the robot to do l π , will the robot do so?" In RL terms, p(c π │s,l π ) is the value function of the skill where the reward is 1 for successful completion and 0 otherwise.

[0025] As described above, l π indicates the text label of skill π, and p(c π │s,l π ) indicates the probability of successfully completing when skill π with text label l π is executed from state s. Here, c πis a Bernoulli random variable. By processing a prompt based on a natural language instruction i using an LLM, an LLM output is generated that characterizes the probability p(l π │i) that the text label of the skill is the next step effective for the user's instruction. However, the embodiments disclosed herein recognize that it is important to consider the probability p(c i │i,s,l π ) that a given skill progresses successfully towards actually completing the instruction. Assuming that the successful skill progresses for i with probability p(l π │i) (i.e., the probability that the skill is appropriate) and the failing skill progresses with probability zero, this can be factored as p(c i │i,s,l π ) ∝ p(c π │s,l π )p(l π │i), which may be referred to herein as task grounding or a task grounding measure. Further, the probability that a skill is possible in the current state of the world can be factored as p(c π │s,l π ), which may be referred to herein as world grounding or a world grounding measure.

[0026] LLMs can draw out rich knowledge learned from large amounts of text, but do not necessarily break down high-level commands into low-level instructions suitable for robot execution. For example, when an LLM is asked "how can a robot bring an apple?", it may respond that "the robot can go to a nearby store and purchase an apple". This response is a reasonable completion for the prompt, but may not necessarily be effective for embodied agents such as robots, which may have a narrow, fixed set of capabilities. Thus, in order to adapt an LLM or other LM, some embodiments attempt to implicitly inform the LM that high-level commands should be broken down into a sequence of available low-level skills. One approach to accomplish this is careful prompt engineering, a technique that guides the LM to a specific response structure. Prompt engineering provides examples in the context text ("prompt") for the LM that specify the tasks and response structures that the LM should emulate.

[0027] Scoring language models opens up a means to constrained responses by outputting the probabilities assigned by the LM to fixed outputs. The LM represents a distribution p(w k │w <k ) of potential completions, where w k is the word that appears at the k-th position in the text. Typical generation applications (e.g., conversational agents) sample from this distribution or decode the maximum likelihood completion, but the embodiments disclosed herein can instead use that distribution to score candidate completions (e.g., candidate skill descriptions from a set of candidate skill descriptions) selected from a set of options. More formally, assuming a set of low-level robot skills II, their language descriptions l II , and a command i, the probability p(l π │i) of the language description of the skill l II ∈l π that progresses towards the execution of the command i can be computed. This corresponds to querying the model for potential completions. The optimal skill by the language model is,

Number

[0028] In some embodiments, the process proceeds by repeatedly selecting a skill and adding it to the instructions. In practice, this may be viewed as an interaction between the user and the robot, where the user gives a high-level instruction (e.g., "Can you bring me a can of Coke?"), and the language model responds in an explicit order (e.g., "I will, 1.", "I will, 1. find a can of Coke, 2. pick up the can of Coke, 3. bring it to you"). Along with the concept of likelihood across many possible responses, a generative response is generated, which has the additional advantage of interpretability.

[0029] Such an approach enables the effective extraction of knowledge from the language model, but a major problem remains. That is, the decoding of the instructions obtained in this way is always composed of the skills available to the robot, but these skills may not always be appropriate for performing the desired high-level task in the specific situation where the robot is currently located. For example, if the FF NL prompt is "Bring me an apple", the optimal set of skills changes if there is no apple within the robot's field of view or if the robot already has an apple in its gripper.

[0030] Therefore, embodiments attempt to ground the large language model through a value function, i.e., an affordance function that captures the log-likelihood of a particular skill being successful in the current state. Given a skill π ∈ Π, its linguistic description l π , and the completion probability p(c π │ s, l π ), of the skill described by l π in state s, assuming its corresponding value function, the affordance space {p(c π │ s, l π)}π ∈ II can be formed. This value function space captures affordances across all skills. For each skill, the affordance function can be multiplied by the probability of the LLM, and finally the most likely skill, i.e., π = argmax π∈II p(c π │s, l π )p(l π │i) is selected.

[0031] Once the most likely skill is selected, the corresponding policy is executed by the agent, and the LLM query is modified to include l π (the linguistic description of the selected skill), and the process is executed again until an end token (e.g., "end") is selected. This process is described in Algorithm 1 given below. These two mirrored processes together lead to a probabilistic interpretation, where the LLM provides the probability of skills useful for high-level commands, and the affordance provides the probability of successful execution of each skill. By combining these two probabilities, the probability that this skill further advances the execution of the high-level command commanded by the user is obtained. Algorithm 1 Assumptions: High-level command i, state s 0 , and set of skills II and their linguistic descriptions l π 1: n = 0, π = Φ 2: while l πn-1 ≠ done do 3: C = Φ 4: for π ∈ II and l π ∈ l II do 5:

Number

Number

[0032] The embodiments disclosed in this specification utilize a set of skills, each of which has a policy, a value function, and a short language description (e.g., "pick up the can") for use when an agent executes the skill. These skills, value functions, and descriptions can be obtained in a variety of different ways. As an example, individual skills can be trained using image-based behavior cloning (e.g., following the BC-Z method) or reinforcement learning (e.g., following the MT-Opt method). Regardless of how the policy of a skill is obtained, a value function such as a value function trained by TD backup can be utilized for the skill. To amortize the cost of training many skills, multi-task BC and / or multi-task RL can be utilized for one or more skills. In multi-task BC and / or multi-task RL, instead of training individual policies and value functions for each skill, a multi-task policy and model conditioned on NL skill descriptions are trained. However, note that this description only applies to low-level skills, and it is still the LLM that is used to interpret high-level instructions and split them into individual low-level skill descriptions.

[0033] ​In some embodiments, a pre-trained large-scale sentence encoder language model may be utilized to condition policies regarding language. The parameters of the sentence encoder language model may be frozen during training, and the embeddings generated by passing the text description of each skill may be the embeddings used when conditioning the policy. The text embeddings are used as inputs to the policy and value function that specify which skills should be executed. Since the language model used to generate the text embeddings is not necessarily the same as the language model used in planning, embodiments can utilize different language models that are better suited for different levels of abstraction, understanding the plan with respect to many skills rather than representing specific skills at a larger granularity.

[0034] In some embodiments, BC and / or RL policy training procedures may be used to respectively obtain one or more of the language-conditioned policies and one or more of the value functions. As described above, for skill specifications, a set of short natural language descriptions, represented as embeddings of a language model, may be utilized. Additionally, a sparse reward function, such as having a reward value of 1.0 at the end of an episode if the execution of the language command is successful and 0.0 otherwise, may be utilized. The success of the language command execution may be rated by a human, for example, and the person rating is given a video of the robot executing the skill along with a given command. If two out of three (or other threshold) people rating agree that the completion of the skill was a success, the episode is labeled with a positive reward. The action space of the policy may include, for example, for a robot similar to that shown in FIG. 1A, 6-degree-of-freedom end effector poses and gripper open / close commands, xy positions and yaw direction deltas of the robot's mobile base, as well as termination actions.

[0035] Various robot skills, such as operations using a mobile manipulator robot and navigation skills, can be utilized in the embodiments disclosed herein. Such skills can include picking, placing, and rearranging objects, opening and closing drawers, navigating to various positions, and placing objects in specific configurations.

[0036] In some embodiments, RL models utilized, such as RL policy models and / or RL value function models, can use an architecture similar to MT-Opt with minor modifications to optionally support natural language input. As one specific example, camera images (e.g., from a robot's camera) can be processed by the convolutional layers of the architecture's image tower to generate image embeddings. Skill descriptions can be embedded by an LLM (or other embedding model) and then concatenated with non-image portions of the state, such as robot actions and gripper height. In order to support asynchronous control, inference can optionally be performed while the robot is still moving from a previous action. The model is optionally given the remaining amount of the previous action to be executed. The conditioning input passes through the fully connected layers of an additional tower of the architecture and is then spatially tiled to generate additional embeddings. The additional embeddings are added / concatenated to the image embeddings before passing through additional convolutional layers. Since the output is gated through a sigmoid, the Q-values are always in [0,1].

[0037] In some embodiments, BC models utilized, such as BC policy models, can use an architecture similar to BC-Z. As one specific example, skill descriptions can be embedded by a Universal Sentence Encoder and then used to condition a Resnet-18 based architecture by FiLM. Unlike RL models, the previous action or gripper height may not be provided. Multiple FC layers can be applied to the final visual features to output each action component (arm position, arm direction, gripper, and end action).

[0038] Note that due to the flexibility of the embodiments disclosed herein, it is possible to mix and match strategies and affordances from different methods. For example, in the case of pick manipulation skills, a single multi-task language-conditioned strategy can be used, and in the case of placement manipulation skills, a scripted strategy using affordances based on the gripper state can be used. In the case of navigation strategies, a planning-based approach can be used that recognizes where certain object(s) can be found and the corresponding distance metric(s). In some embodiments, an upper limit can be set on the affordance indicating that the skill has been completed and a reward has been received in order to avoid situations where a skill is selected but is already being executed, or where there is no effect.

[0039] As described above, the embodiments disclosed herein enable the use of many different strategies and / or affordance functions through their probabilistic interface. Thus, as a skill becomes more performant or as new skills are learned, it is straightforward to incorporate such skills into the embodiments disclosed herein.

[0040] As described above, an embodiment uses an affordance function p(c π │s,l π │s,l π ) that indicates the probability of successfully completing a skill using description l from state s. Some learned strategies that can be used herein create a Q-function Q π (s,a). Given action a and state s with Q π (s,a), similar to MT-Opt, through optimization by the cross-entropy method, the value v(s)=max a Q π (s,a) is found. For simplicity, some exemplary value functions are hereinafter their skill text descriptions l π and

Number

[0041] An exemplary affordance function for the "pick" skill is as follows. The trained value function for the pick skill generally has a minimum value when the skill is impossible and a maximum value when the skill is successful. Thus, the value function is normalized to obtain the following affordance function.

Equation

[0042] An exemplary affordance function for the "goto" / "navigate" skill is as follows. The affordance function for the goto skill is based on the distance d (in meters) to that location.

Equation

[0043] An exemplary affordance function for the terminate skill is

Equation

[0044] Referring now to the figures, FIG. 1A shows an example in which a human 101 gives an exemplary robot 110 a free-form (FF) natural language (NL) command 105 of "bring a snack from the table".

[0045] The robot 110 shown in FIG. 1 is a robot with specific mobility. However, additional and / or alternative robots, such as additional robots that differ from the robot 110 shown in FIG. 1A in one or more respects, may be available for use with the technology disclosed herein. For example, in the technology described herein, a mobile forklift robot, an unmanned aerial vehicle ("UAV"), a non-mobile robot, and / or a humanoid robot may be used instead of or in addition to the robot 110.

[0046] The robot 110 includes a base 113 having wheels provided on its opposing sides for the movement of the robot 110. The base 113 may include one or more motors, for example, to drive the wheels of the robot 110 to achieve movement of the robot 110 in a desired direction, speed, and / or acceleration. The robot 110 also includes a robot arm 114 having an end effector 115 in the form of a gripper having two opposing "fingers" or "digits".

[0047] The robot 110 also includes a vision component 111 that can generate vision data (e.g., an image) related to the shape, color, depth, and / or other characteristics of an object(s) within the line of sight of the vision component 111. The vision component 111 can be, for example, a monocular camera, a stereographic camera (active or passive), and / or a 3D laser scanner.

[0048] A 3D laser scanner may include one or more lasers that emit light and one or more sensors that collect data related to the reflection of the emitted light. The 3D laser scanner may generate visual component data that is a 3D point cloud, where each point in the 3D point cloud defines the position of a point on the surface within 3D space. A monocular camera includes a single sensor (e.g., a charge-coupled device (CCD)) and may generate an image, each containing a plurality of data points that define color values and / or grayscale values based on the physical characteristics sensed by the sensor. For example, a monocular camera may generate an image that includes red, blue, and / or green channels. Each channel may define values for each of the plurality of pixels in the image, such as values from 0 to 255 for each pixel in the image. A stereographic camera may include two or more sensors that are each in a different viewing position. In some of these embodiments, the stereographic camera generates an image, each containing a plurality of data points that define depth values as well as color values and / or grayscale values based on the characteristics sensed by the two sensors. For example, a stereographic camera may generate an image that includes a depth channel and red, blue, and / or green channels.

[0049] Robot 110 may also include one or more processors. The one or more processors may, for example, use an LLM to process an LLM prompt based on the FF NL input 105 to generate an LLM output, determine, based on the LLM output, a description of a robot skill and a value function(s) for the robot skill(s), where the robot skill(s) is / are for implementation when executing a robot task, and control the robot 110 during the execution of the robot task based on the determined robot skill(s), etc. For example, the one or more processors of the robot 110 may implement all or aspects of the methods 300 and / or 400 described herein. Additional explanations of some examples of the structure and functionality of various robots are provided herein.

[0050] Referring now to FIG. 1B, a simplified bird's-eye view of an exemplary environment in which the human 101 and the robot 110 of FIG. 1A are located is shown. The human 101 and the robot 110 are represented by circles in FIG. 1B. Further, environmental features 191, 192, 193, and 194 are shown in FIG. 1B. The environmental features 191, 192, 193, and 194 indicate the outlines of various landmarks of the environment. For example, the environment may be an office kitchen or a workplace kitchen, features 191 and 192 may be counter tops, feature 193 may be a sink, and feature 194 may be a round table.

[0051] FIG. 1B also shows an example of a current visual data instance 180 that is captured within the environment and can be used, for example, when generating a world grounding scale for robot skills. For example, the robot 110 may use the vision component 111 to capture the current visual data instance 180. The visual data instance 180 captures a pear 184A and a key 184B, both of which are present on the round table represented by feature 194. Note that in the bird's-eye view, the pear and the key are shown as dots for simplicity.

[0052] Referring now to FIG. 2A, a process flow of how various exemplary components can interact when selecting an initial robot skill to perform within the environment of FIG. 1B in response to the FF NL command 150 of FIG. 1A. The exemplary components shown in FIG. 2A include the LLM engine 130, the LLM 150, the task grounding engine 132, the world grounding engine 134, value function model(s), the selection engine 130, and the execution engine 136. One or more of the illustrated components may be implemented by the robot 110 (e.g., using a processor(s) and / or its memory) and / or using a remote computing device(s) (e.g., a cloud-based server(s)) that network communicates with the robot 110.

[0053] In FIG. 2A, the LLM engine 130 generates an LLM prompt 205A based on the FF NL input 105 (“Bring me a snack from the table”). The LLM engine 130 may generate the LLM prompt 205A to strictly conform to the FF NL input 105, or may generate the LLM prompt 205A based on the FF NL input 105 but not strictly conform to the FF NL input 105. For example, as shown by the LLM prompt 205A1 which is a non-limiting example of the LLM prompt 205A, the LLM prompt may be “How would you bring me a snack from the table? I would 1.” Such an LLM prompt 205A1 includes “How would you” as a prefix and “I would 1” as a suffix. Either or both of them may facilitate the prediction of the step(s) related to the achievement of the high-level task specified by the FF NL input 105 in the LLM output.

[0054] In some embodiments, the LLM engine 130 may optionally further generate the LLM prompt 205A based on one or more of the scene descriptor(s) 202A, prompt example(s) 203A, and / or description 204A of the current environment of the robot 110.

[0055] The scene descriptor(s) 202A may include NL descriptors of the object(s) currently or recently detected in the environment having the robot 110, such as descriptors of the object(s) (plural) determined based on the processed image(s) or other visual data using the object detection and classification machine learning model(s) (plural). For example, the scene descriptor(s) 202A may include "pear", "key", "human", "table", "sink", and "countertop", and the LLM engine 130 may generate the LLM prompt 205A to incorporate one or more of such descriptors. For example, the LLM prompt 205A may be "There are pears, keys, humans, tables, sinks, and countertops nearby. How do I get a snack from the table? I'll make it 1."

[0056] Prompt example(s) 203A may include examples(s) manually designed with an optional selection of desired output style(s). For example, they may include "in a step-by-step format" or "in a numbered list format" or "in a style such as 1. First step, 2. Second step, 3. Third step, etc.". The prompt example(s) 203A may be added before the LLM prompt 205A or incorporated into the LLM prompt 205A in other ways to facilitate the prediction of content in the output style(s) in the LLM output. Explanation 204A may be an explanation generated based on processing the previous LLM prompt, and based on the FF NL input in the previous pass, and using the LLM 150. For example, the previous LLM prompt may be "Please explain how to get a snack from the table", and the explanation 204A may be generated based on the most probable decoding from the previous LLM output. For example, the explanation 204A may be "Find the table, then find the snack on the table, and then bring it". The explanation 204A may be added before the LLM prompt 205A, may replace the term(s) of the FF NL input 105 within the LLM prompt 205A, or may be incorporated into the LLM prompt 205A in other ways.

[0057] The LLM engine 130 processes the generated LLM prompt 205A using the LLM 150 to generate the LLM output 206A. As described herein, the LLM output 206A may model the probability distribution of candidate word compositions and depends on the LLM prompt 205A.

[0058] The task grounding engine 132 generates a task grounding metric 208A, and generates the task grounding metric 208A based on the LLM output 206A and the skill description 207. Each of the plurality of skill descriptions 207 describes a corresponding skill configured to be executed by the robot 110. For example, "Go to the table" may describe the skill of "Navigate to the table" that the robot can execute by using a trained navigation policy for the navigation target of "table" (or the location corresponding to "table"). As another example, "Go to the sink" may describe the skill of "Navigate to the sink" that the robot can execute by using a trained navigation policy for the navigation target of "sink" (or the location corresponding to "sink"). As yet another example, "Pick up the bottle" may describe the skill of "Grip the bottle" that the robot can execute by using grip heuristics fine-tuned for the bottle and / or by using a trained gripping network. As yet another example, "Pick up the key" may describe the skill of "Grip the key" that the robot can execute by using grip heuristics fine-tuned for the key and / or by using a trained gripping network. The skill description 207 includes descriptors for skills A to G in FIG. 2A, but the skill description 207 may include descriptors for additional skills in various embodiments (as indicated by the ellipsis). Such additional skills may correspond to alternative objects and / or be for different types of robot actions (e.g., "place", "push", "open", "close").

[0059] Each of the task grounding metrics 208A is generated based on the probability of the corresponding skill description in the LLM output 206A. For example, the task grounding metric A, "0.85", reflects the probability of the word sequence "Go to the table" in the LLM output 206A. As another example, the task grounding metric B, "0.20", reflects the probability of the word sequence "Go to the sink" in the LLM output 206A.

[0060] The world grounding engine 134 generates a world grounding scale 211A for robot skills. When generating the world grounding scale 211A for at least some of the robot skills, the world grounding engine 134 can generate the world grounding scale based on the environmental state data 209A and optionally further based on the corresponding robot state data 210A and / or skill description 207. Further, when generating at least some of the world grounding scales 211A for the robot skills, the world grounding engine 134 can utilize one or more of the value function model(s) 152.

[0061] In some embodiments, the world grounding engine 134 can generate a fixed world grounding scale for some robot skills (plural). For example, any "place" robot skill may always have a fixed scale, such as 1.0 or 0.9, or a "finish" robot skill (i.e., indicating that the task is completed) may always have a fixed scale, such as 0.1 or 0.2. In some embodiments, the world grounding engine 134 can additionally or alternatively generate a world grounding scale for some robot skills (plural) based on the corresponding one of the value function model(s) 152 that is not machine learning-based (e.g., not a neural network and / or not trained). For example, the value function model can define that any "place" robot skill should have a fixed scale, such as 1.0 or 0.9, when the environmental state data 209A and / or the robot state data 210A indicates that an object is being held by the robot 110, and should have another fixed scale, such as 0.0 or 0.1, otherwise. As another example, the value function of the "navigate to [object / location]" robot skill can define that the world grounding scale is a function of the distance between the robot and the object / location such that the world grounding scale is determined based on the environmental state data 209A and the robot state data 210. Or a "finish" robot skill (i.e., indicating that the task is completed) may always have a fixed scale, such as 0.1 or 0.2.

[0062] In some embodiments, the world grounding engine 134 may additionally or alternatively generate a world grounding scale for some robot skill(s) based on the corresponding one(s) of the value function model(s) 152 that are trained value function models. In some of those embodiments, the trained value function model may be a language-conditioned model. For example, when generating a world grounding scale for a robot skill, the world grounding engine 134 may use a language-conditioned model to process the corresponding one of the plurality of skill descriptions 207 for the robot skill, along with the environmental state data 209A and optionally the robot state data 210A, to generate a value that reflects the probability that the robot skill is successful based on the current state data. The world grounding engine may generate a world grounding scale based on the generated value (e.g., to conform to the value). In some versions of those embodiments, candidate robot actions are also processed using a language-conditioned model, along with the corresponding one of the plurality of skill descriptions 207, the environmental state data 209A, and optionally the robot state data 210A. In some of those versions, a plurality of values are generated, each by processing a different candidate robot action, but based on using the same corresponding one of the plurality of skill descriptions 207, the same environmental state data 209A, and optionally the same robot state data 210A. In those versions, the world grounding engine 134 may generate a world grounding scale based on, for example, the generated value that reflects the highest success probability. In some of those versions, different candidate robot actions may be selected, for example, using the cross-entropy method. For example, N robot actions may initially be sampled randomly, values may be generated for each, and then N additional robot actions may be sampled from about one of the first N robot actions, based on the first robot action for which the generated value is the highest.The trained value function model used by the world grounding engine 134 when generating the world grounding scale(s) of robot skills can also be utilized for the robot 110 to actually execute robot skills.

[0063] The world grounding scale 211A is generated based on the state of the robot 110 as reflected in the bird's-eye view of FIG. 1B. That is, when the robot 110 is still quite far from the pear 184A and the key 184B. Accordingly, the world grounding scales F and G for "pick up the pear" and "pick up the key" respectively are both relatively low ("0.10"). This reflects that due to the long distance between the robot 110 and the pear 184A and the key 184B, the probability of successfully grasping any of the items to be attempted is low. The "0.80" of the world grounding scale A for "go to the table" reflects a lower probability than the "0.85" of the world grounding scale B for "go to the sink". This can be based on the fact that the robot 111 is closer to the sink 193 than to the table 194.

[0064] When the selection engine 136 selects the robot skill A ("go to the table"), it considers both the world grounding scale 211A and the task grounding scale 208A, and sends the instruction 213A of the selected robot skill A to the execution engine 136. In response, the execution engine 136 controls the robot 110 based on the selected robot skill A. For example, the execution engine 136 can control the robot using a navigation strategy with a navigation target of "table" (or the location corresponding to "table").

[0065] In FIG. 2A, the selection engine 136 generates the overall scale 212A by multiplying the world grounding scale 211A and the task grounding scale 208A, and selects the robot skill A based on the fact that the robot skill A is the highest among the overall scales 212A. Note that the robot skill A has the highest overall scale even though it does not have the highest world grounding scale. FIG. 2A shows that the overall scale 212A is generated by multiplying the world grounding scale 211A and the task grounding scale 208A, but other techniques may be used when generating the overall scale 212A. For example, in the multiplication, different weightings may be applied to the world grounding scale 211A and the task grounding scale 208A. For example, the world grounding scale 211A may be weighted 90%, and the task grounding scale 208A may be weighted 100%.

[0066] Now, referring to FIG. 2B, following (or during) the implementation of the robot skill A selected in FIG. 2A, when selecting the next robot skill to be implemented in the environment of FIG. 1B in response to the FF NL instruction of FIG. 1A, a process flow of how the various exemplary components of FIG. 2A can interact.

[0067] In Figure 2B, the LLM engine 130 generates an LLM prompt 205B based on the FF NL input 105 (“Bring a snack from the table”) and further based on the skill descriptor 201B (“Go to the table”) selected for robot skill A to be provided for execution. The LLM engine 130 may generate the LLM prompt 205B to strictly conform to the FF NL input 105 and the selected skill descriptor 201B, or may generate the LLM prompt 205B based on the FF NL input 105 and the selected skill descriptor 201B but not strictly conforming. For example, as shown by the LLM prompt 205B1 which is a non-limiting example of the LLM prompt 205B, the LLM prompt may be “How do I bring a snack from the table? 1. I will go to 1. Go to the table. 2.” Such an LLM prompt 205B1 includes “How do I” as a prefix and “2.” as a suffix. Either or both of them may facilitate the prediction of steps (s) related to the achievement of the high-level task specified by the FF NL input 105 in the LLM output. Further, such an LLM prompt 205B1 includes the selected skill descriptor 201B of “Go to the table” preceded by “1.”, which may facilitate the prediction of steps (s) related to occurring after the step of “Go to the table” in the LLM output.

[0068] In some embodiments, the LLM engine 130 may optionally generate the LLM prompt 205B further based on one or more of the scene descriptor(s) 202B of the current environment of the robot 110 (which may vary with respect to Figure 2A), the example(s) 203B of the prompt (which may optionally vary from that of Figure 2A), and / or the explanation 204B (which may optionally vary from that of Figure 2A).

[0069] The LLM engine 130 processes the generated LLM prompt 205B using the LLM 150 to generate the LLM output 206B. As described herein, the LLM output 206B may model the probability distribution of candidate word compositions and is dependent on the LLM prompt 205B.

[0070] The task grounding engine 132 generates a task grounding scale 208B and generates the task grounding scale 208B based on the LLM output 206A and the skill description 207. Each of the task grounding scales 208B is generated based on the probability of the corresponding skill description in the LLM output 206B. For example, task grounding scale A, "0.00", reflects the probability of the word sequence "go to the table" in the LLM output 206B. As another example, task grounding scale B, "0.10", reflects the probability of the word sequence "go to the sink" in the LLM output 206B.

[0071] The world grounding engine 134 generates a world grounding scale 211B for the robot skills. When generating the world grounding scale 211B for at least some of the robot skills, the world grounding engine 134 may generate the world grounding scale based on the environmental state data 209B and optionally further based on the corresponding ones of the robot state data 210B and / or the plurality of skill descriptions 207. It should be noted that the environmental state data 209B and the robot state data 210B change from their counterparts in Figure 2A due to robot skill A being at least partially implemented at the time of Figure 2B. Further, when generating the world grounding scale 211B for at least some of the robot skills, the world grounding engine 134 may utilize one or more of the value function model(s) 152.

[0072] In some embodiments, the world grounding engine 134 may generate a fixed world grounding scale for some robot skills (plural). In some embodiments, the world grounding engine 134 may additionally or alternatively generate a world grounding scale for some robot skills (plural) based on the corresponding one of the value function model(s) 152 that is not machine learning-based (e.g., not a neural network and / or not trained). In some embodiments, the world grounding engine 134 may additionally or alternatively generate a world grounding scale for some robot skills (plural) based on the corresponding one of the value function model(s) 152 that is a trained value function model. In some of those embodiments, the trained value function model may be a language-conditioned model.

[0073] The world grounding scale 211B is generated based on the state of the robot 110 after the robot skill A has been at least partially executed (i.e., the robot 110 is closer to the table 194 than is reflected in the bird's-eye view of FIG. 1B). Thus, the robot 110 is closer to the pear 184A and the key 184B at the time of FIG. 2A (e.g., within the reach range of both). Thus, the world grounding scales F and G for "pick up the pear" and "pick up the key", respectively, are both relatively high ("0.90").

[0074] When selecting the robot skill F ("pick up the pear"), the selection engine 136 considers both the world grounding scale 211B and the task grounding scale 208B and transmits an instruction 213B for the selected robot skill F to the execution engine 136. In response, the execution engine 136 controls the robot 110 based on the selected robot skill F. For example, the execution engine 136 may control the robot using a grasping strategy that optionally uses parameters fine-tuned for grasping the pear.

[0075] In FIG. 2B, the selection engine 136 generates the overall scale 212B by multiplying the world grounding scale 211B and the task grounding scale 208B, and selects the robot skill F based on the fact that the robot skill F is the highest among the overall scale 212B. It should be noted that although the robot skill A does not have the highest world grounding scale and does not have the highest task grounding scale either, it has the highest overall scale. FIG. 2B shows that the overall scale 212B is generated by multiplying the world grounding scale 211B and the task grounding scale 208B, but other techniques may be used when generating the overall scale 212B. For example, in the multiplication, different weightings can be applied to the world grounding scale 211B and the task grounding scale 208B. For example, the world grounding scale 211B may be weighted 90%, and the task grounding scale 208B may be weighted 100%.

[0076] FIG. 3 is a flowchart showing an exemplary method 300 of leveraging an LLM in the implementation of robot skill(s) that execute the task(s) reflected in the FF NL instruction, according to the embodiments disclosed herein. For convenience, the operations of method 300 are described with respect to the system that performs the operations. This system may include one or more components of a robot, such as a robot processor and / or a robot control system of robot 110, robot 520, and / or other robots, and / or one or more components of a computer system, such as computer system 610. Further, although the operations of method 300 are shown in a particular order, this is not meant to be limiting. One or more operations may be changed in order, omitted, or added.

[0077] In block 352, the system identifies the FF NL instruction.

[0078] In block 354, the system uses the LLM to process an LLM prompt based on the FF NL instruction to generate an LLM output. Block 354 optionally includes sub-block 354A and / or sub-block 354B. In sub-block 354A, the system includes a scene descriptor(s) in the LLM prompt. In sub-block 354B, the system generates an explanation and includes the explanation in the LLM prompt.

[0079] In block 356, the system generates a corresponding task grounding metric based on the LLM output of block 354 and the corresponding NL skill description for each of the plurality of candidate robot skills. Block 356 optionally includes sub-block 356A, and in sub-block 356A, the system generates a task grounding metric for the candidate robot skill based on the probability of the NL skill description reflected in the probability distribution of the LLM output.

[0080] In block 358, the system generates a corresponding world grounding metric for each of the plurality of candidate robot skills based on the current state data. Although shown positionally below block 356 in FIG. 3, it should be noted that block 358 may occur in various embodiments in parallel with or even before block 356.

[0081] Block 358 optionally includes sub-block 358A and / or sub-block 358B. Optionally, sub-block 358A and / or sub-block 358B are executed only for a certain robot skill among the candidate robot skills.

[0082] In sub-block 358A, the system processes the current environmental state data and / or the current robot state data using a value function model when generating the world grounding metric for the candidate robot skill.

[0083] In sub-block 358B, the system also processes the NL description (e.g., its embedding) of the robot skill when generating the world grounding scale of the candidate robot skill using the multi-skill and language-conditioned value function model.

[0084] In block 360, the system selects the robot skill to be executed or selects an end condition (which can be regarded as a specific robot skill) based on the task grounding scale of block 356 and the world grounding scale of block 358. For example, if the task grounding scale for an end condition corresponding to "ended", "completed", or other end descriptions meets a threshold, and / or if the task grounding scale, world grounding scale, and / or overall scale for all robot skills do not meet the same or different thresholds, the end condition can be selected.

[0085] In block 362, the system determines whether a robot skill has been selected for execution in block 360. If not, the system proceeds to block 364 and ends by controlling the robot based on the FF NL instruction of block 352. If selected, the system proceeds to blocks 366 and 368.

[0086] In block 366, the system executes the selected robot skill. In block 368, the system modifies the latest LLM prompt based on the skill description of the executed skill. Then the system returns to block 354 and processes the LLM prompt as modified in block 368. The system also performs another iteration of blocks 356, 358, 360, and 362, and optionally (depending on the decision in block 362) blocks 366 and 368. This general process can continue until an end condition is selected in an iteration of block 360.

[0087] Figure 4 is a flowchart illustrating an exemplary method for generating a world grounding measure of a robot skill based on processing current state data and a skill description for the robot skill. For convenience, the operations of method 400 are described with respect to a system that performs the operations. The system may include one or more components of a robot, such as a robot processor and / or a robot control system of robot 110, robot 520, and / or other robots, and / or one or more components of a computer system, such as computer system 610. Further, the operations of method 400 are shown in a particular order, but this is not meant to be limiting. One or more operations may be changed, omitted, or added.

[0088] In block 452, the system generates an image embedding by processing the current image using an image tower of a trained language-conditioned value function model. The image can be a current image captured by a visual component of the robot.

[0089] In block 454, the system selects a candidate action. For example, the system can select a candidate action by sampling from the action space.

[0090] In block 456, the system generates an additional embedding by processing the candidate action selected in block 454, a skill description (e.g., its embedding) of the robot skill, and optionally the robot state data, using an additional tower of the trained language-conditioned value function model.

[0091] In block 458, the system generates a value based on processing a concatenation of the image embedding and the additional embedding using an additional layer of the trained language-conditioned value function model.

[0092] In block 460, the system determines whether to generate additional values for additional candidate actions. If not, the system proceeds to block 464. If so, the system proceeds to block 462 and selects an additional candidate action. In some iterations of block 462, this may involve sampling from the action space without considering the values generated so far. In some later iterations of block 462, this may involve sampling from the action space based on the values generated so far using the cross-entropy method. For example, the sampling may be around the portion of the action space corresponding to the candidate action with the highest value generated so far, or may be biased towards that portion in other ways. The system then returns to block 456 with the newly selected candidate action but using the same skill description and optionally the same robot state.

[0093] In block 464, the system generates a world grounding scale for the robot skill corresponding to the skill description (used in the iteration(s) of block 456) based on the maximum value of the values generated in the iteration(s) of block 458.

[0094] Figure 5 schematically shows an exemplary architecture of a robot 520. The robot 520 includes a robot control system 560, one or more motion components 540a - 540n, and one or more sensors 542a - 542m. The sensors 542a - 542m can include, for example, visual sensors, light sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors, accelerometers, gyroscopes, thermometers, barometers, etc. The sensors 542a - 542m are shown as being integrated with the robot 520, but this is not meant to be limiting. In some embodiments, the sensors 542a - 542m can be located external to the robot 520, for example, as a stand-alone unit.

[0095] The motion components 540a - 540n can include, for example, one or more end effectors, and / or one or more servo motors, or other actuators, to effect the movement of one or more components of a robot. For example, the robot 520 may have multiple degrees of freedom, and each of the actuators can control the operation of the robot 520 within the range of one or more degrees of freedom in response to control commands. As used herein, the term actuator includes any mechanical or electrical device that can be associated with an actuator and that converts a received control command into one or more signals for driving the actuator, in addition to any driver(s) that can convert a received control command into one or more signals for driving the actuator. Thus, providing a control command to an actuator can include providing the control command to a driver that converts the control command into appropriate signals for driving an electrical or mechanical device to effect a desired motion.

[0096] The robot control system 560 can be implemented in one or more processors such as the CPU, GPU, and / or other controller(s) of the robot 520. In some embodiments, the robot 520 can include a "brain box" that can include all or aspects of the control system 560. For example, the brain box may provide real - time bursts of data to the motion components 540a - n, and each real - time burst can include a set of one or more control commands that, in particular, indicate motion parameters (if any) for one or more of the motion components 540a - n. In some embodiments, the robot control system 560 can execute one or more aspects of the methods (plural) described herein, such as the method 300 of FIG. 3 and / or the method 400 of FIG. 4.

[0097] As described herein, in some embodiments, all or aspects of the control commands generated by the control system 560 when controlling the robot during the execution of a robotic task may be generated based on robot skill(s) determined to be related to the robotic task based on the world grounding scale and task grounding scale described herein. In some embodiments, although the control system 560 is shown in FIG. 5 as an integral part of the robot 520, all or aspects of the control system 560 may be implemented by a component that is separate from the robot 520 but communicates with the robot 520. For example, all or aspects of the control system 560 may be implemented on one or more computing devices that communicate with the robot 520, such as the computing device 610, either wired and / or wirelessly.

[0098] FIG. 6 is a block diagram of an exemplary computing device 610 that may optionally be utilized to execute one or more aspects of the techniques described herein. The computing device 610 typically includes at least one processor 614 that communicates with several peripheral devices via a bus subsystem 612. These peripheral devices may include, for example, a storage subsystem 624 that includes a memory subsystem 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices enable user interaction with the computing device 610. The network interface subsystem 616 provides an interface to an external network and is coupled to a corresponding interface device within other computing devices.

[0099] The user interface input device 622 can include a keyboard, a mouse, a trackball, a touchpad, or a pointing device such as a graphics tablet, a scanner, a touch screen incorporated in a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 610 or the communication network.

[0100] The user interface output device 620 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or any other mechanism for creating a visible image. The display subsystem can also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 610 to the user or another machine or computing device.

[0101] The storage subsystem 624 stores the programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 624 can include logic for performing selected aspects of the method 300 of FIG. 3 and / or the method 400 of FIG. 4.

[0102] These software modules are typically executed by the processor 614 alone or in combination with other processors. The memory 625 used in the storage subsystem 624 may include several memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 can provide persistent storage of program files and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. The modules implementing the functionality of an embodiment may be stored by the file storage subsystem 626 within the storage subsystem 624 or in other machines accessible by the processor(s) 614.

[0103] The bus subsystem 612 provides a mechanism for enabling the various components and subsystems of the computing device 610 to communicate with each other as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative embodiments of the bus subsystem may use multiple buses.

[0104] The computing device 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Since computers and networks are constantly changing in nature, the description of the computing device 610 shown in FIG. 6 is intended only as a specific example for explaining some embodiments. Many other configurations of the computing device 610 are possible that may have more or fewer components than the computing device shown in FIG. 6.

[0105] In some embodiments, a method is provided that is performed by one or more processors and includes identifying an instruction and processing the instruction using a language model (LM) (e.g., a large language model (LLM)) to generate an LM output. The instruction can be a free-form natural language instruction generated based on user interface inputs provided by a user via one or more user interface input devices. The generated LM output can model a probability distribution of candidate word compositions that depends on the instruction. The method further includes identifying robot skills executable by a robot and skill descriptions that are natural language descriptions of the robot skills. The method further includes generating a task grounding metric for a robot skill based on the LM output and the skill description. The task grounding metric can reflect the probability of the skill description in the probability distribution of the LM output. The method further includes generating a world grounding metric for the robot skill based on the robot skill and current environmental state data. The world grounding metric can reflect the probability that the robot skill is successful based on the current environmental state data. The current environmental state data can include sensor data captured by one or more sensor components of the robot in the current environment of the robot. The method further includes determining to execute the robot skill instead of a plurality of additional robot skills each executable by the robot based on both the task grounding metric and the world grounding metric. The method further includes causing the robot to execute the robot skill in the current environment in response to determining to execute the robot skill.

[0106] These and other embodiments of the technology disclosed herein may include one or more of the following features.

[0107] In some embodiments, the method further includes identifying one additional robot skill of a plurality of additional robot skills and an additional skill description that is a natural language description of the one additional robot skill, and generating an additional task grounding metric for the one additional robot skill based on the LM output and the additional skill description. The additional task grounding metric may reflect an additional probability of the additional skill description in the probability distribution. In those embodiments, the method further includes generating an additional world grounding metric for the one additional robot skill based on the one additional robot skill and the current environmental state data. The additional world grounding metric may reflect an additional probability that the one additional robot skill is successful based on the current environmental state data. Further, in those embodiments, determining to perform the robot skill instead of a plurality of additional robot skills each executable by the robot based on both the task grounding metric and the world grounding metric may include determining to perform the robot skill based on the task grounding metric, the world grounding metric, the additional task grounding metric, and the additional world grounding metric. For example, the method may include generating an overall metric for the robot skill based on both the task grounding metric and the world grounding metric, generating an additional overall metric for the additional robot skill based on both the additional task grounding metric and the additional world grounding metric, and determining to perform the robot skill instead of the additional robot skill based on a comparison of the overall metric and the additional overall metric.

[0108] In some embodiments, the method, in response to determining to execute a robotic skill, uses an LM to process the instructions and the skill description of the robotic skill to generate additional LM output that models a further probability distribution of candidate word compositions that depend on the instructions and the skill description; identifies an additional robotic skill of one of a plurality of additional robotic skills and an additional skill description that is a natural language description of the one additional robotic skill; generates an additional task grounding metric for the one additional robotic skill based on the additional LM output and the additional skill description; generates an additional world grounding metric for the one additional robotic skill based on the one additional robotic skill and updated current environmental state data; determines to execute the one additional robotic skill instead of the robotic skill or another additional robotic skill of the plurality of additional robotic skills each executable by the robot, based on both the additional task grounding metric and the additional world grounding metric; and in response to determining to execute the one additional robotic skill, causes the robot to execute the one additional robotic skill in the current environment. The updated current environmental state data may include updated sensor data captured by one or more sensor components after execution of the robotic skill in the current environment.

[0109] In some embodiments, the method, in response to determining to perform a robot skill, uses an LM to process the instructions and the skill description of the robot skill to generate additional LM output that models a further probability distribution of candidate word compositions that depend on the instructions and the skill description, identifies an end skill among a plurality of additional robot skills and an end description that is a natural language description that the execution of the instructions is complete, generates an end task grounding measure for the end skill based on the additional LM output and the end description, generates an end world grounding measure for the end skill based on the end skill and the updated current environmental state data, determines to perform the end skill instead of the robot skill or another additional robot skill among the plurality of additional robot skills each executable by the robot, based on both the additional task grounding measure and the additional world grounding measure, and in response to determining to perform the end skill, further stops the robot from further executing any control commands that advance the instructions. The updated current environmental state data may include updated sensor data captured by one or more sensor components after the performance of the robot skill in the current environment.

[0110] In some embodiments, generating a world grounding metric based on robot skills and current environmental state data includes using a trained value function to process the robot skills and current environmental state data to generate a value function output that includes the world grounding metric. In some versions of those embodiments, the current environmental state data includes visual data of sensor data (e.g., multi-channel images), and the visual data is captured by one or more visual components of one or more sensor components of the robot. In some additional or alternative versions of those embodiments, the trained value function is a language-conditioned value function, and using the trained value function to process the robot skills includes processing the skill descriptions of the robot skills. In some additional or alternative versions of those embodiments, the trained value function is trained to correspond to an affordance function, and the value function output specifies whether the robot skills are possible based on the current environmental state data. In some additional or alternative versions of those embodiments, the value function is a machine learning model trained using reinforcement learning.

[0111] In some embodiments, determining to perform a robotic skill instead of a plurality of additional robotic skills based on both a task grounding metric and a world grounding metric includes generating an overall metric as a function of the task grounding metric and the world grounding metric, comparing the overall metric to a corresponding plurality of additional metrics for corresponding ones of the plurality of additional robotic skills, and based on the comparison, determining to perform the robotic skill. In some versions of those embodiments, the overall metric is a weighted or unweighted combination of the task grounding metric and the world grounding metric. In some additional or alternative versions of those embodiments, the task grounding metric is a task grounding probability, the world grounding metric is a world grounding probability, and generating the overall metric includes generating a product based on multiplying the task grounding probability and the world grounding probability and using the product as the overall metric.

[0112] In some embodiments, causing a robot to perform a robotic skill in a current environment includes causing the robot to execute a language-conditioned robotic control policy conditioned by a skill description of the robotic skill. In some versions of those embodiments, the language-conditioned robotic control policy includes a machine learning model. In some versions of those embodiments, the language-conditioned robotic control policy is trained using reinforcement learning and / or imitation learning.

[0113] In some embodiments, the instructions strictly conform to natural language input provided by a user via a user interface input.

[0114] In some embodiments, the instructions do not strictly conform to natural language input provided by the user via a user interface input. In some of those embodiments, the method further includes determining that the natural language input is not in question form and, in response to determining that the natural language input is not in question form, generating instructions by modifying the natural language input to be in question form.

[0115] In some embodiments, the LM is a large language model (LLM).

[0116] Provided is a method implemented by one or more processors that includes identifying instructions that are free-form natural language instructions generated based on user interface input provided by a user via one or more user interface input devices. The method further includes using a language model (LM) to process the instructions to generate an LM output, such as an LM output that models a probability distribution of candidate word compositions that depend on the instructions. The method further includes, for each of a plurality of robot skills, each executable by a robot, generating a task grounding metric for the robot skill (e.g., one that reflects the probability of the skill description in the probability distribution) based on the LM output and the skill description of the robot skill, generating a world grounding metric for the robot skill (e.g., one that reflects the probability that the robot skill will be successful based on the current environmental state data) based on the robot skill and the current environmental state data, and generating an overall metric for the robot skill based on the task grounding metric and the world grounding metric. The current environmental state data includes sensor data captured by one or more sensor components of the robot in the robot's current environment. The method further includes selecting a given skill among the robot skills based on the overall metric for the robot skill. The method further includes, in response to selecting a given robot skill, causing the robot to perform the given robot skill in the current environment.

[0117] In some embodiments, a method is provided that is performed by one or more processors and includes identifying an instruction that is a free-form natural language instruction, such as one generated based on a user interface input provided by a user via one or more user interface input devices. The method further includes processing a language model (LM) prompt generated based on the instruction using an LM (e.g., a large language model (LLM)) to generate an LM output that models a probability distribution of candidate word compositions that depends on the LM prompt. The method further includes identifying robot-executable robot skills and skill descriptions that are natural language descriptions of the robot skills. The method further includes generating a task grounding metric for a robot skill that reflects the probability of the skill description in the probability distribution of the LM output, based on the LM output and the skill description. The method further includes generating a world grounding metric for the robot skill that reflects the probability that the robot skill will be successful based on the current environmental state data, based on the robot skill and the current environmental state data. The current environmental state data optionally includes sensor data captured by one or more sensor components of the robot in the current environment of the robot. The method further includes determining to perform the robot skill instead of a plurality of additional robot skills each executable by the robot, based on both the task grounding metric and the world grounding metric. The method further includes causing the robot to perform the robot skill in the current environment, in response to determining to perform the robot skill.

[0118] In some embodiments, a method is provided that is performed by one or more processors and includes identifying an instruction that is a free-form natural language instruction, such as one generated based on user interface inputs provided by a user via one or more user interface input devices. The method further includes generating a large language model (LLM) prompt based on the instruction. The method further includes using the LLM to process the LLM prompt to generate an LLM output that models a probability distribution of candidate word compositions that depend on the LLM prompt. The method further includes, for each of a plurality of robot skills each executable by a robot, generating a task grounding measure for the robot skill that reflects the probability of the skill description in the probability distribution based on the LLM output and the skill description of the robot skill, generating a world grounding measure for the robot skill that reflects the probability that the robot skill will be successful based on the current environmental state data based on the robot skill and the current environmental state data, and generating an overall measure for the robot skill based on the task grounding measure and the world grounding measure. The method further includes selecting a given skill among the robot skills based on the overall measure for the robot skills. The method further includes causing the robot to perform the given robot skill in the current environment in response to selecting the given robot skill.

[0119] Other embodiments may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., one or more central processing units (CPUs), one or more graphics processing units (GPUs), and / or one or more tensor processing units (TPUs)) to perform methods such as one or more of the methods described herein. Further other embodiments may include a system of one or more computers and / or one or more robots including one or more processors operable to execute stored instructions to perform methods such as one or more of the methods described herein.

Claims

1. A method performed by one or more processors, comprising: identifying an instruction, wherein the instruction is a free-form natural language instruction generated based on a user interface input provided by a user via one or more user interface input devices; using a language model (LM) to process an LM prompt generated based on the instruction and generate an LM output that models a probability distribution of candidate word compositions that depend on the LM prompt; identifying robot-executable robot skills and skill descriptions that are natural language descriptions of the robot skills; generating a task grounding metric for the robot skill that reflects the probability of the skill description in the probability distribution of the LM output, based on the LM output and the skill description; generating a world grounding metric for the robot skill that reflects the probability that the robot skill will be successful based on the current environmental state data, based on the robot skill and the current environmental state data, wherein the current environmental state data includes sensor data captured by one or more sensor components of the robot in the current environment of the robot; generating; determining, based on both the task grounding metric and the world grounding metric, to perform the robot skill instead of a plurality of additional robot skills each executable by the robot; in response to determining to perform the robot skill, causing the robot to perform the robot skill in the current environment. A method comprising the above steps.

2. identifying one of the plurality of additional robot skills and an additional skill description that is a natural language description of the one additional robot skill; generating an additional task grounding metric for the one additional robot skill that reflects an additional probability of the additional skill description in the probability distribution, based on the LM output and the additional skill description; Generating an additional world grounding measure for the additional robot skill that reflects an additional probability of success of the one additional robot skill based on the current environmental state data, based on the one additional robot skill and the current environmental state data; further comprising Based on both the task grounding measure and the world grounding measure, determining to perform the robot skill instead of a plurality of additional robot skills each executable by the robot; Generating an overall measure for the robot skill based on both the task grounding measure and the world grounding measure; Generating an additional overall measure for the one additional robot skill based on both the additional task grounding measure and the additional world grounding measure; Based on a comparison between the overall measure and the additional overall measure, determining to perform the robot skill instead of the one additional robot skill; The method according to claim 1, comprising.

3. In response to determining to perform the robot skill, Using the LM to process a further LM prompt based on the instruction and the skill description of the robot skill, and generating a further LM output that models a further probability distribution of the candidate word compositions that depends on the further LM prompt; Identifying one additional robot skill among the plurality of additional robot skills and an additional skill description that is a natural language description of the one additional robot skill; Based on the further LM output and the additional skill description, generating an additional task grounding measure for the one additional robot skill that reflects the probability of the additional skill description in the further probability distribution; Based on the one additional robot skill and the updated current environmental state data, generating an additional world grounding measure for the one additional robot skill that reflects the probability of success of the one additional robot skill based on the updated current environmental state data; The updated current environmental state data includes updated sensor data captured by the one or more sensor components after implementation of the robot skill in the current environment. Said generating; Based on both the additional task grounding metric and the additional world grounding metric, determining to implement the one additional robot skill instead of the robot skill or another one of the plurality of additional robot skills each executable by the robot. In response to determining to implement the one additional robot skill; Causing the robot to implement the one additional robot skill in the current environment. The method according to claim 1, comprising:

4. In response to determining to implement the robot skill; Using the LM to process a further LM prompt based on the instructions and the skill description of the robot skill, and generating a further LM output that models a further probability distribution of the candidate word compositions that depends on the further LM prompt. Identifying an end skill among the plurality of additional robot skills and an end description that is a natural language description that the execution of the instructions has been completed. Based on the further LM output and the end description, generating an end task grounding metric for the end skill that reflects the probability of the end skill description in the further probability distribution. Based on the additional task grounding metric, determining to implement the end skill instead of the robot skill or another one of the plurality of additional robot skills each executable by the robot. In response to determining to implement the end skill; Causing the robot to stop further implementation of any control commands that advance the instructions. The method according to claim 1, comprising:

5. Generating the world grounding metric based on the robot skill and the current environmental state data may be The method according to any one of claims 1 to 4, comprising using a trained value function to process the robot skill and the current environmental state data to generate a value function output including the world grounding metric.

6. The method according to claim 5, wherein the current environmental state data includes the visual data of the sensor data, and the visual data is captured by one or more visual components among the one or more sensor components of the robot.

7. The method according to claim 6, wherein the visual data includes a multi-channel image.

8. The trained value function is a language-conditioned value function, The method according to claim 5, wherein processing the robot skill using the trained value function includes processing the skill description of the robot skill.

9. The method according to claim 5, wherein the trained value function is trained to correspond to an affordance function, and the value function output specifies whether the robot skill is possible based on the current environmental state data.

10. The method according to claim 5, wherein the value function is a machine learning model trained using reinforcement learning.

11. Determining to perform the robot skill instead of the plurality of additional robot skills based on both the task grounding scale and the world grounding scale, Generating an overall scale as a function of the task grounding scale and the world grounding scale, Comparing the overall scale with a corresponding plurality of additional scales for the corresponding ones of the plurality of additional robot skills, Determining to perform the robot skill based on the comparison, The method according to any one of claims 1 to 10, comprising:

12. The method according to claim 11, wherein the overall scale is a weighted or unweighted combination of the task grounding scale and the world grounding scale.

13. The method according to claim 12, wherein the task grounding scale is a task grounding probability, the world grounding scale is a world grounding probability, generating the overall scale includes generating a product based on multiplying the task grounding probability and the world grounding probability, and using the product as the overall scale.

14. Causing the robot to perform the robot skill in the current environment The method according to any one of claims 1 to 14, comprising causing execution of a language-conditioned robot control policy conditioned by the skill description of the robot skill. **Claim 15** The method according to claim 14, wherein the language-conditioned robot control policy includes a machine learning model. **Claim 16** The method according to claim 15, wherein the language-conditioned robot control policy is trained using reinforcement learning and / or imitation learning. **Claim 17** The method according to any one of claims 1 to 16, wherein the LM prompt strictly conforms to a natural language input provided by the user via the user interface input. **Claim 18** The LM prompt does not strictly conform to a natural language input provided by the user via the user interface input, determining that the natural language input is not in the form of a question, generating the LM prompt by modifying the natural language input to be in the form of a question in response to determining that the natural language input is not in the form of a question. The method according to any one of claims 1 to 16, further comprising. **Claim 19** The LM prompt does not strictly conform to a natural language input provided by the user via the user interface input, The method according to any one of claims 1 to 16, further comprising generating the LM prompt to include, as a suffix, content claiming generation of the next step. **Claim 20** The LM prompt does not strictly conform to a natural language input provided by the user via the user interface input, The method according to any one of claims 1 to 16, further comprising generating the LM prompt to include one or more scene descriptors describing the current environment of the robot. **Claim 21** The method according to any one of claims 1 to 20, wherein the LM is a large language model (LLM). **Claim 22** A method implemented by one or more processors, comprising: identifying an instruction, the instruction being a free-form natural language instruction generated based on a user interface input provided by a user via one or more user interface input devices. Generating a large language model (LLM) prompt based on the command; Using the LLM to process the LLM prompt to generate an LLM output that models a probability distribution of candidate word compositions that depend on the LLM prompt; For each of a plurality of robot skills each executable by a robot, Generating a task grounding measure for the robot skill that reflects the probability of the skill description in the probability distribution based on the LLM output and the skill description of the robot skill; Generating a world grounding measure for the robot skill that reflects the probability that the robot skill is successful based on the current environmental state data, based on the robot skill and the current environmental state data, wherein the current environmental state data includes sensor data captured by one or more sensor components of the robot in the current environment of the robot; said generating; Generating an overall measure for the robot skill based on both the task grounding measure and the world grounding measure; Selecting a given skill among the robot skills based on the overall measure for the robot skill; In response to selecting the given robot skill, Causing the robot to perform the given robot skill in the current environment; A method comprising.

23. A robot, One or more actuators; An end effector; A memory storing instructions; One or more processors operable to execute the instructions for executing the method according to any one of claims 1 to 22; A robot comprising.

24. A memory storing instructions; One or more processors operable to execute the instructions for executing the method according to any one of claims 1 to 22; A system comprising.

Citation Information

Patent Citations

  • Information processor, information processing method, program, and interactive system

    JP2011054088A

  • Machine learning unit, robot system and machine learning method

    JP2020121381A

  • Speech recognition biasing

    US10438587B1