Integration of planning and video conferencing
By integrating video conferencing into robot training and management, the challenges of resource-intensive robot policy training and non-expert user management are addressed, achieving efficient data collection and intuitive task management for robots.
Patent Information
- Application Number
- PCT/GB2024/052605
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2024-10-10
- Publication Date
- 2025-06-12
AI Technical Summary
Training robot planning and control policies requires significant resources and expertise, and managing robots to perform tasks using these policies can be challenging for non-expert users.
Leveraging video conference tools to streamline machine learning model training for robot planning and control, allowing human robot supervisors, robotic planning processes, and robots to join video conference sessions, enabling efficient data collection and low-granularity management of robot task performance through natural language interactions.
This approach enables efficient, fast, and cost-effective collection of training data for machine learning models, allows non-expert users to manage robot tasks using natural language, and simplifies the deployment of augmented-reality applications for robot control.
Smart Images

Figure GB2024052605_12062025_PF_FP_ABST
Abstract
Description
INTEGRATION OF PLANNING AND VIDEO CONFERENCING Background
[0001] Training robot planning and / or control policies often requires significant resources. Expert trainers may spend significant amounts of time closely monitoring robot operation and providing feedback that is used to train the robot control policies. By necessity, this work is often performed in close proximity with robot(s), e.g., in the same room or building. Even after robot control policies are extensively trained, subsequently managing robots to perform tasks using those trained robot control policies can be difficult for non-expert users, who may lack the knowledge and / or technological capability to manage a robot outside of issuing high-level, or “long-horizon,” commands.
[0002] Separately, many video conference tools provide infrastructure for real time video streaming, multi-participants, and high-quality speech transcription for population of transcripts and / or message exchange threads provided as part of video conferences. Summary
[0003] Implementations are described herein for leveraging video conference tools to streamline machine learning model training relating to robot (and non-robot) planning and / or control, and / or for subsequent robot management. More particularly, but not exclusively, implementations are described herein for allowing human robot supervisors, robotic planning processes, robot control data generation processes, and / or robots themselves to join video conference sessions.
[0004] Techniques described herein give rise to various technical advantages and benefits. Leveraging video conferencing enables far more efficient, fast, and less costly collection of training data for training, for instance, generative language models that predict mid-level actions for carrying out high-level tasks. Leveraging video conferencing also enables relatively low granularity management of robot task performance, e.g., by allowing non-expert users to intervene in robot operation using natural language, rather than expecting the non-expert users to control the robot at a lower level. Leveraging video conferencing also simplifies deployment of augmented-reality (AR) applications, whether the relate to control robots or simply to performing tasks (e.g., a user may not know how to rebuild an engine, but may wish to receive instructions iteratively, even without using a robot).
[0005] In some implementations, a computer implemented method may be provided that includes: causing one or more video conference clients of a video conference session to render output that includes one or more sensor feeds, wherein the one or more sensor feeds capture an environment in which a robot operates; communicatively coupling a robotic planner process to the video conference session; receiving, from a video conference client operated by a first participant in the video conference session, a natural language request for the robot to perform a high-level task; causing the robotic planner process to process the natural language request using a first machine learning model to generate a plurality of natural language responses, each natural language response expressing a mid-level action to be performed by the robot to carry out a respective portion of the high-level task; causing one or more of the natural language responses to be processed to generate robot control data; and causing the robot to be operated based on the robot control data.
[0006] In various implementations, the robotic planner process may processes the natural language request over multiple iterations to generate the plurality of natural language responses, wherein during at least one of the iterations, the natural language request and a natural language response from a previous iteration are processed to generate a next natural language response. In various implementations, the communicatively coupling may include registering, as a second participant of the video conference session, the robotic planner process. In various implementations, the one or more natural language responses may be processed using a second machine learning model to generate the robot control data.
[0007] In various implementations, the one or more sensor feeds may include a video feed captured by a camera that is mounted on the robot, a camera that is deployed in the environment separately from the robot, a LIDAR feed captured by a LIDAR sensor, and / or a microphone, to name a few.
[0008] In various implementations, the robot planner process may process data from one or more of the sensor feeds in conjunction with the natural language request using the first machine learning model to generate the one or more natural language responses.
[0009] In various implementations, the method may further include: receiving, from a video conference client operated by the first participant or a second participant of the video conference session, feedback on a given natural language response of the one or more natural language responses generated using the first machine learning model; and modifying the given naturallanguage response based on the feedback to generate a modified natural language response that conveys a modified mid-level action to be performed by the robot to carry out a portion of the high-level task. In various implementations, the modified natural language response may be processed to generate at least some of the robot control data. In various implementations, the feedback may include input generated in response to actuation of a graphical element that is operable to reject the given natural language response. In various implementations, the method may further include storing the modified natural language response as training data for training the first machine learning model. In various implementations, the method may further include causing the first machine learning model and one or more edge-based machine learning models with less parameters than the first machine learning model to be jointly trained based on the training data.
[0010] In various implementations, the method may include receiving, from a video conference client operated by the first participant or a second participant of the video conference session, feedback on a given natural language response of the one or more natural language responses generated using the first machine learning model. The feedback may include acceptance or approval of the given natural language response, and the acceptance or approval is generated by the first or third participant using a graphical element.
[0011] In various implementations, the method may include receiving, from a video conference client operated by the first participant or a second participant of the video conference session, a request to alter a temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets; and based on the request, altering the temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets. In various implementations, altering the temporal frequency may include transitioning between a first state in which a sequence of natural language snippets are processed without interruption and a second state in which the sequence of natural language snippets are processed one-at-a-time based on subsequent input from the video conference client.
[0012] In various implementations, the method may include: receiving, from the planner process, a request to alter a temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets; and based on the request, altering the temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets.
[0013] In various implementations, the method may include, based on one or more metrics associated with one or more of the natural language responses generated by the robotic planner process, altering a temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets.
[0014] In various implementations, the method may include causing one or more of the video conference clients to render the one or more natural language responses as additional output. In various implementations, wherein the first machine learning model may be a generative language model such as a visual question answering (VQA) model.
[0015] In a related aspect, a method may be implemented using one or more processors to perform the following operations: registering the robotic planner process as a first participant of the video conference session; communicatively coupling the robotic planner process to one or more sensor feeds that capture an environment in which one or more robots operate; monitoring a transcript, message exchange thread, and / or audio of the video conference session; based on the monitoring, retrieving a natural language input incorporated provided by another participant of the video conference session, wherein the natural language input expresses a request for one or more of the robots to perform a high-level task; processing the natural language request and data from one or more of the sensor feeds using a first machine learning model to generate one or more natural language responses, each natural language response expressing a mid-level action to be performed by one or more of the robots to carry out a respective portion of the high-level task; and incorporating the one or more natural language responses into the transcript, message exchange thread, or audio of the video conference session, or into a message exchange thread that is provided as part of the video conference session.
[0016] In various implementations, the method may include causing one or more of the natural language responses to be processed using a second machine learning model to generate robot control data. In various implementations, the first machine learning model may include a VQA model. In various implementations, processing the natural language request and data from one or more of the sensor feeds using the first machine learning model may include processing the natural language request, the data from one or more of the sensor feeds, and a natural language response generated during a previous iteration of the first machine learning model being applied to the natural language request. In various implementations, the one or more robots may includea plurality of robots, and the mid-level action may include selecting a given robot from the plurality of robots to perform a respective portion of the high-level task.
[0017] Other implementations may include a non-transitory computer readable storage medium storing instructions executable by a processor to perform a method such as one or more of the methods described above. Yet another implementation may include a control system including memory and one or more processors operable to execute instructions, stored in the memory, to implement one or more modules or engines that, alone or collectively, perform a method such as one or more of the methods described above.
[0018] It should be appreciated that all combinations of the foregoing concepts and additional concepts described in greater detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. Brief Description of the Drawings
[0019] Fig.1 schematically depicts an example environment in which disclosed techniques may be employed, in accordance with various implementations.
[0020] Fig.2 depicts an example robot, in accordance with various implementations.
[0021] Fig.3 schematically depicts an example of how various components depicted in Fig.1 may cooperate to carry out selected aspects of the present disclosure.
[0022] Fig.4A, Fig.4B, Fig.4C, Fig.4D, Fig.4E, and Fig.4F schematically depict an example of how techniques described herein may be implemented to collect training data, in accordance with various implementations.
[0023] Fig.5A, Fig.5B, and Fig.5C schematically depict an example of how techniques described herein may be implemented to manage a robot remotely, in accordance with various implementations.
[0024] Fig.6 depicts an example method for practicing selected aspects of the present disclosure.
[0025] Fig.7 depicts an example method for practicing selected aspects of the present disclosure.
[0026] Fig.8 schematically depicts an example architecture of a computer system.Detailed Description
[0027] Implementations are described herein for leveraging video conference tools to streamline machine learning model training relating to robot planning and / or control, and / or for subsequent robot management. More particularly, but not exclusively, implementations are described herein for allowing human robot supervisors, robotic planning processes, robot control data generation processes, and / or robots themselves to join video conference sessions.
[0028] In various implementations, one or more sensor feeds that depict one or more robots operating in an environment, such as feeds provided by onboard vision sensors and / or environmental vision sensors, may be provided to the video conference session. For example, a vision sensor feed generated by a robot’s vision sensor may be registered as a participant of the video conference session. Once registered, the vision sensor feed may be presented as that participant’s video feed, e.g., so that other participants can see what the robot’s vision sensor perceives.
[0029] In various implementations, a human robot supervisor may, acting as a participant of the video conference session, provide a natural language request for the robot to perform a high- level / long-horizon task. For example, the human robot supervisor may incorporate the natural language request into transcript of the video conference session and / or a message exchange thread that is provided as part of the video conference session (e.g., that is visible to participants). For example, the human robot supervisor may speak the natural language request during the video conference session. In some further examples, a speech to text function may convert the spoken natural language request into text. For example, a transcript or captioning function of the video conference session may be utilized to convert the spoken natural language request into text. In some cases, the human robot supervisor may explicitly direct the natural language request to the robotic planner process, e.g., be prefacing the natural language request with a predetermined phrase such as “Hey robot planner,…,” actuating a graphical element such as a “raise hand” function, or the like. As used herein, in various implementations, a high-level or long-horizon task may refer to a first level task and a mid-level or “medium-horizon” action may refer to a second level action, wherein a first level task comprises two or more second level actions. Low level robot commands such as joint commands and / or Cartesian commands (e.g., to control an end effector) may be referred to in some instances as “short horizon” commands.
[0030] In some implementations, a single human robot supervisor may be permitted (e.g., at a time) as a participant in a video conference session configured with selected aspects of the present disclosure. In some implementations, a participant may declare themselves a robot supervisor (e.g., “I am taking over as robot supervisor”), in which case a previous robot supervisor may be relegated to another non-robot supervisor role, such as a regular video conference participant, a spectator, etc. In other implementations, multiple human robot supervisors may be permitted at the same time.
[0031] The robotic planner process, which may also be acting as a participant in the video conference session, may monitor the audio, captions, transcript and / or message exchange thread for such requests, e.g., by monitoring for the predetermined phrase. Once the robotic planner processor detects a natural language request for a robot to perform a high-level task, the robotic planner process may break down the high-level task into multiple mid-level actions to be performed by the robot to carry out respective portions of the high-level task. For example, the robotic planner process may process the natural language request using various types of machine learning models such as a generative language model or a visual question answering (VQA) model, to generate one or more natural language responses. Each natural language response may express a respective mid-level action.
[0032] In some implementations, robot control data may be generated iteratively. For instance, an original natural language request may be processed using the machine learning model in a first iteration, e.g., in conjunction with current sensor feed data, to generate a first natural language response expressing a first mid-level action. In some implementations, prior to the next iteration, the machine learning model may be used to process sensor feed data and the first natural language response to ensure that the mid-level action expressed in the first natural language response was carried out. Assuming the answer is yes, during a second iteration, the original natural language request may be processed in conjunction with the first natural language response (and with updated sensor feed data that reflects performance of the mid-level action expressed by the first natural language response) to generate a second natural language response. This process may continue until the high-level task is carried out.
[0033] The robotic planner process may also process data from the sensor feed that depicts an environment in which the robot operates, e.g., as contextual data for generating mid-level actions to carry out the high-level robot task. For example, features of the images (i.e., the sensorfeed(s)) may be detected and may influence how the machine learning model processes natural language requests to generate the natural language responses that express the mid-level actions for the robot to carry out portions of the high-level task. Suppose the human robot supervisor issues the high-level request, “put all this stuff in the drawer.” The sensor feed may depict various objects, such as a cup and a spoon, that are detected by the machine learning model or by another upstream model (e.g., a convolutional neural network or similar) that provides input (e.g., embeddings) to the machine learning model. Based on these detected objects, the machine learning model may generate mid-level actions that are appropriate to the context, such as “pick up the cup,” “put the cup into the drawer,” “pick up the spoon,” and “put the spoon into the drawer.”
[0034] In some implementations, the robotic planner process may incorporate the natural language responses it generates (and which express the mid-level actions) into the transcript or message exchange thread so that they are perceivable by other participants, such as the human robot supervisor, e.g., prior to those natural language responses being used to generate robot control data. In some such implementations, the human robot supervisor may have an opportunity to provide feedback on individual natural language responses (and hence, mid-level actions) prior to those being used to generate robot control data. For example, some predetermined amount of time may elapse between each natural language response being incorporated into the transcript or message exchange thread and corresponding robotic action(s) being performed, thereby giving the human robot supervisor an opportunity to accept, reject, and / or modify the natural language response (and hence, mid-level action). Alternatively, each natural language response may be incorporated into the transcript or message exchange and / or used to generate robot control data thread only after the human robot supervisor has provided feedback on the previous natural language request. In some implementations, the human robot supervisor may issue a command to transition the video conference session between these alternative approaches, and / or to alter a speed at which the natural language responses are incorporated into the thread and / or processed to generate robot control data.
[0035] In some implementations, the manner (e.g., frequency) in which natural language responses expressing mid-level actions are incorporated into the transcript or message exchange thread and / or processed to generate robot control data may be modulated based on features of the responses themselves and / or the process of generating robot control data. For instance, thenatural language responses and / or the mid-level actions they express may be assigned metrics such as confidence scores. These metrics may be influenced, for instance, by factors such as entropy and / or noise in the robot’s environment, ambiguity between perceived objects (e.g., it may not be clear whether an object is of one type or another), and so forth. In various implementations, the frequency at which natural language responses are presented and / or processed to generate robot control data may be modulated based on these metrics. If one or more natural language responses are generated with relatively low confidence scores, an amount of time that is elapsed between those natural language responses being presented, e.g., in a transcript of the video conference session and / or in a message exchange thread provided as part of the video conference session, and processed to generate robot control data may be increased, e.g., to give a participant of the video conference session an opportunity to provide feedback, alter the mid-level action, etc. Additionally or alternatively, in some implementations, if the aforementioned VQA machine learning model is unable to generate high-confidence predictions that mid-level action(s) were carried out, the frequency at which natural language responses are presented and / or processed to generate robot control data may be decreased to give a video conference participant time to weigh in, intervene, etc.
[0036] As described above, in some implementations, the natural language responses may be processed, e.g., by a robot control data generation process, using one or more additional machine learning models to generate robot control data such as joint commands and / or Cartesian commands (e.g., for robotic end effectors). In some implementations, the robot control data process may also (e.g., concurrently) apply the additional machine learning model(s) to data from the sensor feed. These robot control data may then be used to operate the robot. In some implementations, the robot control data generation process and / or the robot itself may be enrolled as participants in the video conference session, in which case data may be exchanged between the various entities using the video conference session. In other implementations, the robot control data generation process and / or the robot may not be registered participants, and instead may be communicatively coupled with each other and / or with the robotic planner process.
[0037] In other implementations, to collect training data, a human participant in the video conference session such as the human robot supervisor may manually control the robot, e.g., via teleoperation, based on the natural language responses that express mid-level actions. Forinstance, as the robotic planner process incorporates natural language responses (expressing mid- level actions) into the transcript and / or message exchange thread provided as part of the video conference session, the participant may teleoperate the robot (e.g., with the robot being registered as a participant in the video conference session) to perform the mid-level actions manually. The robot control data generated by the participants teleoperation may be collected, e.g., to train a robot control data generation machine learning model. Additionally, to the extent the user performs or does not perform (e.g., rejects) the mid-level actions expressed by the natural language responses, that feedback may be used to train the machine learning model (e.g., generative language model, VQA model) used by the robotic planner process.
[0038] Fig.1 is a schematic diagram components that can cooperate to carry out selected aspects of the present disclosure, in accordance with various implementations. The various components depicted in Fig.1, particularly those components forming a video conference system 120, a robotic planner system 140, and robot control data system 150, may be implemented using any combination of hardware and software. A robot 100 may be in communication with systems 120, 140, and / or 150, and / or all or parts of systems 120, 140, and / or 150 may be implemented onboard robot 100. The components of Fig.1 may be communicatively coupled with each other via one or more networks 199, which may include one or more personal area networks, local area networks, and / or wide area networks (e.g., the Internet).
[0039] Robot 100 may take various forms, including but not limited to a telepresence robot (e.g., which may be as simple as a wheeled vehicle equipped with a display and a camera), a robot arm, a multi-pedal robot such as a “robot dog,” an aquatic robot, a wheeled device, a submersible vehicle, an unmanned aerial vehicle (“UAV”), and so forth. One non-limiting example of a mobile robot arm is depicted in Fig.2. In various implementations, robot 100 may include logic 102. Logic 102 may take various forms, such as a real time controller, one or more processors, one or more field-programmable gate arrays (“FPGA”), one or more application-specific integrated circuits (“ASIC”), and so forth. In some implementations, logic 102 may be operably coupled with memory 103. Memory 103 may take various forms, such as random-access memory (“RAM”), dynamic RAM (“DRAM”), read-only memory (“ROM”), Magnetoresistive RAM (“MRAM”), resistive RAM (“RRAM”), NAND flash memory, and so forth. In some implementations, a robot controller may include, for instance, logic 102 and memory 103 of robot 100.
[0040] In some implementations, logic 102 may be operably coupled with one or more joints 104-1 to 104-N, one or more end effectors 106, and / or one or more sensors 108-1 to 108- M, e.g., via one or more buses 110. As used herein, “joint” 104 of a robot may broadly refer to actuators, motors (e.g., servo motors), shafts, gear trains, pumps (e.g., air or liquid), pistons, drives, propellers, flaps, rotors, or other components that may create and / or undergo propulsion, rotation, and / or motion. Some joints 104 may be independently controllable, although this is not required. In some instances, the more joints robot 100 has, the more degrees of freedom of movement it may have.
[0041] As used herein, “end effector” 106 may refer to a variety of tools that may be operated by robot 100 in order to accomplish various tasks. For example, some robots may be equipped with an end effector 106 that takes the form of a claw with two opposing “fingers” or “digits.” Such a claw is one type of “gripper” known as an “impactive” gripper. Other types of grippers may include but are not limited to “ingressive” (e.g., physically penetrating an object using pins, needles, etc.), “astrictive” (e.g., using suction or vacuum to pick up an object), or “contigutive” (e.g., using surface tension, freezing or adhesive to pick up object). More generally, other types of end effectors may include but are not limited to drills, brushes, force-torque sensors, cutting tools, deburring tools, welding torches, containers, trays, and so forth. In some implementations, end effector 106 may be removable, and various types of modular end effectors may be installed onto robot 100, depending on the circumstances. Some robots, such as some telepresence robots, may not be equipped with end effectors. Instead, some telepresence robots may include displays to render visual representations of the users controlling the telepresence robots, as well as speakers and / or microphones that facilitate the telepresence robot “acting” like the user.
[0042] Sensors 108-1 to 108-M may take various forms, including but not limited to 3D laser scanners (e.g., light detection and ranging, or “LIDAR”) or other 3D vision sensors (e.g., stereographic cameras used to perform stereo visual odometry) configured to provide depth measurements, two-dimensional cameras (e.g., RGB, infrared), light sensors (e.g., passive infrared), force sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors (also referred to as “distance sensors”), depth sensors, torque sensors, barcode readers, radio frequency identification (“RFID”) readers, radars, range finders, accelerometers, gyroscopes, compasses, position coordinate sensors (e.g., global positioning system, or “GPS”),speedometers, edge detectors, Geiger counters, and so forth. While sensors 108-1 to 108-M are depicted as being integral with robot 100, this is not meant to be limiting.
[0043] In some implementations, video conference system 120, robotic planner system 140, and / or robotic control data system 150 may include one or more computing devices cooperating to perform selected aspects of the present disclosure. An example of such a computing device is depicted schematically in Fig.8. In some implementations, one or more of systems 120, 140, and / or 150 may include one or more servers forming part of what is often referred to as a “cloud” infrastructure, or simply “the cloud.” Alternatively, one or more components of systems 120, 140, and / or 150 may be operated by logic 102 of robot 100.
[0044] Video conference system 120 may be configured to host or otherwise facilitate video conference sessions that allow participants (e.g., 160) to collect training data for training robot control policies and / or remotely manage robot operation. Video conference system 120 includes a participant engine 122, a transcript engine 124, a teleoperation engine 126, a sensor feed engine 128, a message exchange engine 130, and a rating engine 132. One or more of engines 122-132 may be combined with other engines, omitted, or implemented elsewhere.
[0045] Participant engine 122 may be configured to manage the process of registering and / or enrolling participants (e.g., 160) in video conference sessions. A participants 160 may operate a video conference client 164 that executes on a client device 162 to interact with video conference system 120. While one participant 160 is depicted in Fig.1, there may be multiple participants in a video conference session, including human participants (e.g., 160) as well as non-human participants (e.g., scripts, routines, processes). A “video conference session” may be an instance of one or more participants operating respective video conference clients 164 to exchange messages and other information and / or conduct video calls with other participants who are joined in the video conference session. While client device 162 is depicted as a tablet computer in Fig. 1, this is not meant to be limiting. Client device 162 may take numerous other forms, such as a laptop computer, desktop computer, mobile phone, head-mounted display, etc.
[0046] Transcript engine 124 may be configured to record and / or manage a transcript of a video conference session. A transcript may include, for instance, speech-to-text (STT) conversions of utterances spoken by participant(s) 160 that are captured at client device(s) 162 by video conference client(s) 164.
[0047] Teleoperation engine 126 may facilitate teleoperation of robot 100 by participant(s) of video conference sessions. In particular, teleoperation engine 126 may provide an interface (e.g., an application programming interface, or “API”) that enables video conference participants to teleoperate robot 100, e.g., using input mechanisms such as joysticks, keyboards, mice, touchscreens, etc. In some implementations, robot 100 may be registered, e.g., by participant engine 122, as a non-human participant of a video conference session.
[0048] Sensor feed engine 128 may facilitate presentation of sensor feed data capturing an environment in which robot 100 operates to participant(s) of the video conference session. For instance, one or more of sensors 108-1 to 108-M of robot 100, and / or one or more sensors external to robot 100, may generate sensor data (e.g., images, video, point clouds, audio, etc.) that sensor feed engine 128 can cause to be rendered (e.g., streamed) at video conference client(s) 164. In some implementations, sensor feed engine 128 may incorporate the one or more sensor feeds as additional, non-human participant(s) in the video conference session, and those sensor feeds would be used in a similar fashion as camera feeds that allow human participants to see each other.
[0049] Video conference systems often may provide a separate message exchange thread that allows participants to communicate with each other textually, e.g., as a “sidebar” without interrupting a spoken conversation between participants of the video conference session. Accordingly, message exchange engine 130 may be configured to maintain and / or manage a message exchange thread that is provided by video conference system 120 as part of a video conference. In some implementations, there may not be a separate transcript and message exchange thread, but instead may be a single exchange thread that any messages—spoken or otherwise—are incorporated into so that participants can communicate with each other.
[0050] A rating engine 132 may be configured to logs interactions between participants and / or determine which interactions constitute, for instance, corrections from trusted authorities (for example a human that says "I am the supervisor"). Rating engine 132 may use this data to obtain training data that is usable, e.g., by training engine 148 of robotic planner system 140 to train one or more machine learning models 144. This training data may include, for instance, pairs (e.g., erroneous mid-level action, corrected mid-level action) of labels which can be used for retraining and / or to determine metric(s) relating to how often participants have to intervene. These metrics can provide measure(s) of progress over time as machine learning model(s) 144 are trained andand / or deployed. This may be particularly beneficial during early deployment of machine learning model(s) in the real world because it avoids waiting for models to be perfect (which can take a long time). More generally, this automatic labeling performed by rating engine 132 provides considerable savings in terms of labeling time and also reduces overhead associated with generation and / or collection of training data.
[0051] Robotic planner system 140 may be configured to facilitate participation of a robotic planner process 142 in a video conference hosted by video conference system 120. Robotic planner process 142 may be configured to process natural language snippets (e.g., requests, queries, commands, etc.) using one or more machine learning models 144 and generate natural language responses. In many implementations described herein, the natural language input that is processed by robotic planner process 142 includes a natural language request for robot 100 to perform a high-level task (e.g., “put these dishes into the dishwasher”). Robotic planner process 142 may process such a natural language request using machine learning model 144 to generate a plurality of natural language responses. Each natural language response may express a mid-level action to be performed by robot 100 to carry out a respective portion of the high-level task. Various non-limiting examples will be discussed herein.
[0052] Machine learning model(s) 144 may take various forms, including generative language model(s) (sometimes referred to as “large language models,” or “LLMs”) such as PaLM, BERT, LaMDA, Meena, and / or any other generative language model, such as any other generative model that is encoder-only based, decoder-only based, sequence-to-sequence based and that optionally includes an attention mechanism or other memory. In generative language model form, machine learning model(s) 144 may have hundreds of millions, or even hundreds of billions of parameters. In some implementations, machine learning model(s) 144 may take the form of a multi-modal model such as a visual question answering (VQA) model, which can have any of the aforementioned architectures, and which can be used to process multiple modalities of data, particularly images and text, and / or images and audio for example, to generate one or more modalities of output, such as the aforementioned natural language responses.
[0053] Feedback engine 146 may be configured to receive feedback from participants and take various actions based on that feedback. To train and / or fine-tune machine learning model(s) 144, for instance, feedback engine 146 may receive feedback (e.g., participant approval or rejection of a mid-level action, failure of a mid-level action to yield a successful outcome) and provide it,e.g., along with mid-level action(s), as training data for a training engine 148. Training engine 148 may use this feedback to train and / or fine-tune machine learning model(s) 144. In some implementations, feedback engine 146 and rating engine 132 may be combined, and either may be implemented in whole or in part on robotic planner system 140 and / or video conference system 140.
[0054] Robotic control data system 150 may be configured to generate robot control data that is operable to control robot 100, e.g., by transmitting robot control data to robot 100. Robot control data may include, for instance, joint commands that control individual joints 104-1 to 104-N and / or Cartesian commands that specify Cartesian coordinates for end effector 106. In some cases, robot logic 102 may be configured to convert between joint commands and Cartesian commands, e.g., using forward and / or inverse kinematics.
[0055] Robotic control data system 150 may include a robot control data generation process 152 that processes natural language snippets, including those that express mid-level actions for carrying out portions of a higher level robot task, to generate robot control data. In some implementations, robot control data generation process may use one or more robot control data machine learning models 154 to generate robot control data. Robot control data machine learning model(s) 154 may take various forms, similar to model(s) 144, such as PaLM, BERT, LaMDA, Meena, and / or any other generative language model, such as any other generative model that is encoder-only based, decoder-only based, sequence-to-sequence based and that optionally includes an attention mechanism or other memory. An example of a robot control data machine learning model that may be used is described in “RT-1: Robotics Transformer for Real-World Control at Scale” (arXiv:2212.06817).
[0056] Feedback engine 156 may obtain feedback, e.g., from human(s) (e.g., 160) and / or from outcomes of robot 100 being operated using robot control data generated by robot control data generation process 152. Feedback engine 156 may provide feedback data to another training engine 158. Similar to training engine 148, training engine 158 may be configured to train and / or fine-tune machine learning model(s) 154 based on feedback generated by feedback engine 156.
[0057] Fig.2 depicts a non-limiting example of a robot 200 in the form of a robot arm. An end effector 206 in the form of a gripper claw is removably attached to a sixth joint 204-6 of robot 200. In this example, six joints 204-1 to 204-6 are indicated. However, this is not meant to belimiting, and robots may have any number of joints. In some implementations, robot 200 may be mobile, e.g., by virtue of a wheeled base 265 or other locomotive mechanism. Robot 200 is depicted in Fig.2 in a particular selected configuration or “pose.”
[0058] Fig.3 schematically depicts an example of how various components depicted in Fig.1 may cooperate to carry out selected aspects of the present disclosure. In Fig.3, time runs down the page. Starting at top left, video conference client 164 may be operated to, for instance, establish a video conference session, e.g., by sending a request to establish a session to video conference system 120. As a result of this video conference session being established, various entities may communicatively couple with video conference system 120. For example, in Fig.3, robotic planner system 140, robotic control data system 150, one or more sensors 108, and robot 100 itself all join the video conference session, at least some as participants. However, it is not required that these entities join the video conference session as participants. One or more of components 140, 150, 108, and / or 100 may be communicatively coupled (e.g., over network(s) 119) to other of components 140, 150, 108, and / or 100 outside of the video conference session. For instance, sensor 108 is simply communicatively coupled with video conference system 120, so that a sensor feed from sensor 108 may be rendered, in whole or in part, by video conference client 164.
[0059] Once the session is established, video conference client may transmit data indicative of a natural language request to video conference system. For instance, a participant (e.g., 160) who operates video conference client 164 may type the natural language request, or may speak an utterance that is recorded and optionally converted to text representing the natural language request using STT. The natural language request may include, for instance, a request for robot 100 to perform a high-level task, such as “load those dishes into the dishwasher,” “empty the dishwasher,” “put away the toys in this room,” “assemble these pieces into a tool,” etc.
[0060] In various implementations, video conference system 120 may provide data indicative of the natural language request (e.g., the request itself, embeddings generated therefrom) to robotic planner system 140. In some implementations, video conference system may also provide data from sensor 108 (sensor feed at T1) to robotic planner system 140. For instance, if the high-level task was “load these dishes into the dishwasher,” the sensor feed may initially depict a sink full of dishes because nothing has been loaded into the dishwasher yet.
[0061] Robotic planner system 140, e.g., by way of robotic planner process 142, may process the natural language request, and if applicable, data from sensor 108, using machine learning model(s) 144. For instance, machine learning model(s) 144 may include a VQA generative model that is usable to process embedding(s) generated from multiple modalities, including images and text and / or images and audio for example. Based on this processing, robotic planner process 142 may generate one or more natural language responses expressing a first mid-level action. Using the dishwasher example, the first mid-level action might be “load the white cup into the dishwasher,” assuming that there is a readily accessible white cup in the sink. In various implementations, robotic planner system 140 may provide data indicative of this first mid-level action (“1stMLA” in Fig.3) to video conference system 120.
[0062] In some implementations, depending on factors such as a confidence score calculated by robotic planner system 140 for the first mid-level action, or whether a participant of the video conference session has requested to approve each mid-level action before it is processed by downstream component(s), video conference client 164 may render the first mid-level action. The participant may then approve, reject, and / or modify the first mid-level action as they see fit. In Fig.3, the participant approves the mid-level action, e.g., by selecting a “thumbs up” icon or similar, and data indicative of this approval is provided to video conference system 120.
[0063] Once video conference system 120 has the approval of the participant to proceed further with the first mid-level action, in various implementations, video conference system 120 may provide data indicative of the first mid-level action to robotic control data system 150. In other implementations, robotic planner system 140 may transmit this data directly to robotic control data system 150, e.g., without sending it through video conference system 120 first.
[0064] In any case, once it has the first mid-level action, robotic control data system 150 may process the first mid-level action, e.g., using machine learning model(s) 154, to generate first robot control data. As noted previously, this first robot control data may include, for instance, joint commands and / or Cartesian coordinates for end effector 106. Additionally or alternatively, this root control data may include other commands, such as commands that would be issued to robot 100 using a joystick or other similar controls.
[0065] Robotic control data system 150 may provide the first robot control data to robot 100. While not shown in Fig.3, in implementations in which robot 100 and / or robotic control data system 150 are registered participants in the video conference session hosted by videoconference system 120, this first robot control data may be transmitted / relayed via video conference system 120. Robot 100 may then be operated based on this robot control data to perform the first mid-level action.
[0066] In various implementations, video conference system 120, e.g., by way of sensor feed engine 128, may obtain the latest sensor feed data that depicts robot 100 after / during performance of the first mid-level action. In Fig.3 this sensor feed data is referred to as “sensor feed T2.” In the working dishwasher example, this sensor feed data may portray the same sink, except without the white cup because it has been loaded into the dishwasher.
[0067] Video conference system 120 may then provide data indicative of the original natural language request and the first mid-level action, as well as the latest sensor feed data (“sensor feed T2”), to robotic planner system 140. Similar to before, robotic planner system 140 may generate, and in some cases provide to video conference system 150, data indicative of a second mid-level action (“2ndMLA”). In the ongoing dishwasher example, this second mid-level action may be, for instance, “put the blue plate into the dishwasher,” assuming the latest sensor feed indicates that a blue plate is readily accessible (e.g., not blocked by other dishes).
[0068] In Fig.3, when second mid-level action is rendered at video conference client 164, the participant (not depicted) rejects it and provides, as a modified mid-level action, a third mid-level action (“3rdMLA”). In the dishwasher example, for instance, the participant may see another dish that should be loaded into the dishwasher before the blue plate, such as a fragile wine glass, and so may provide the modified mid-level action, “load the wine glass into the dishwasher.” This third mid-level action may be provided to video conference system 120.
[0069] Video conference system 120 may provide data indicative of (e.g., embedding(s)) the third mid-level action to robotic planner system 140. Similar to before, robotic planner system 140, e.g., by way of robotic planner process 142, may process the data indicative of the third mid-level action using the machine learning model (e.g., 144) to generate second robot control data. This second robot control data may then be provided by robotic planner system 140 to robot 100, which may be operated based on the second robot control data. In the ongoing dishwasher example, robot 100 may grab the wine glass from the sink and load it into the dishwasher.
[0070] Once robot 100 is operated based on the second robot control data, e.g., by loading the wine glass into the dishwasher, video conference system may receive an updated sensor feed(“sensor feed T3”) that depicts the new state of the environment. In the dishwasher example, for instance, the wine glass will no longer be in the sink. The same process as before may then repeat. Data indicative of the original natural language request, the third mid-level action, the first mid-level action in some cases (and more generally, all mid-level actions performed thus far in furtherance of carrying out the high-level task), and the latest sensor feed data may be provided by video conference system 120 to robotic planner system 140, which in turn may generate, e.g., using machine learning model(s) 144, a fourth mid-level action (“4thMLA”). As indicated by the ellipses, this process may continue until the high-level task is completed (e.g., all dishes from the sink are loaded into the dishwasher).
[0071] The data exchanges depicted in Fig.3 are for illustration only and are not meant to be limiting. Numerous variations are contemplated. For instance, it is not required that robot 100 be involved in the process at all. Rather, a participant may act as the “robot” or “body” that will be performing the mid-level actions to cumulatively accomplish the high-level task. The sensor feed can be used to monitor the participant’s progress in carrying out the high-level task. And robotic planner process 142 may act as an “instructor” that tells the human participant what to do by outputting, as natural language, the mid-level actions described previously, instead of having those mid-level actions processed to generate robot control data.
[0072] Suppose the high-level task expressed by the participant’s original natural language request is to a rebuild a car engine, and that the sensor feed depicts a car engine in a shop. For instance, the sensor feed may be generated by a camera mounted on the participant, a camera of the participant’s mobile device, a camera mounted in the shop, etc. This high-level task may be used by robotic planner process 142 to generate a sequence of mid-level actions that can be carried out to collectively accomplish the high-level task. For example, machine learning model(s) 144 may have been trained with documents pertaining to rebuilding the engine.
[0073] As noted above, these mid-level actions may be generated and presented to the participant iteratively. For example, each mid-level action may be generated by robotic planner process 142 applying machine learning model(s) 144 to the original high-level task plus any mid- level actions that have been carried out thus far, as well as to sensor feed data that depicts the engine’s evolving physical state as it is rebuilt. At each iteration, with the knowledge of the high-level task, the steps that have already been performed, and the current state of the engine (depicted in the sensor feed), robotic planner process 142 is able to use machine learningmodel(s) 144 to predict a natural language response that expresses a next mid-level action to be performed by the participant.
[0074] In other implementations, rather than the human participant themselves performing the mid-level actions to accomplish the high-level task, the human participant may interact with teleoperation engine 126 to manually control robot 100 to perform the mid-level tasks. Otherwise, the process may work similarly as described above. The human participant can accept mid-level actions provided by robotic planner process 142 by either explicitly accepting them using a graphical element such as a “thumbs up” button, or by causing robot to interact with the environment consistent with the mid-level actions. For example, if the mid-level action is “put the spoon in the dishwasher” and the human participant operates robot 100 to do just that, that may constitute an acceptance of the mid-level action that can be used as a training example. By contrast, if the human modifies the mid-level action, that may constitute a rejection that also can be used as a training example.
[0075] Figs.4A-F depict an example graphical user interface (GUI) 466 that may be rendered by video conference client 164, in accordance with various implementations. GUI 466 includes a variety of different components, some of which are similar to those often presented as part of video conference client applications. These are not meant to be limiting, and in various implementations, other components may be presented by a GUI of video conference client 164 in other combinations, orders, etc.
[0076] While Figs.4A-F and other figures herein depict a GUI that might be rendered on a display, this is not meant to be limiting. Other output devices, such as head-mounted displays or augmented reality (AR) glasses, may be leverage to present information as described herein. For instance, a user wearing AR glasses may wish to receive a sequence of instructions for performing a task, and have those instructions presented in an AR fashion overlaying the environment in which the user wishes to perform the task.
[0077] While the examples depicted in Fig.3 and elsewhere herein include a single robot and a single robotic planner process 142, this is not meant to be limiting. In various implementations, multiple robots and / or robotic planner processes may be involved in a single video conference session. For example, multiple robotic planner processes may be involved in a single video conference session and may collaborate to determine appropriate mid-level actions. In another example, there may be an environment in which multiple robots operate, such as a warehouse,factory, etc. In such a scenario, the various participants (e.g., humans, robots, robotic planner processes) may avoid interrupting each other by, for instance, “raising their hands” in the video conference session. In addition, while the examples described herein include real robots, this is not required, and a simulated robot operating in a simulated environment is also contemplated. In such a scenario, the sensor feed may be generated by a virtual sensor based on states of the virtual environment.
[0078] In Fig.4A, various video feeds 468A-C are presented, one for each participant in the video conference session. First video feed 468A renders sensor feed data from a vision sensor that captures an environment in which a robot 400 operates. In this example, the vision sensor is mounted on robot 400 itself, thereby providing a “point of view” of robot 400, which is currently viewing a piece of furniture with a top surface 480 and two drawers 482A and 482B. There are also some items placed on top surface 480, namely, a plate, a cup, and a bowl. However, this is not required, and a vision sensor that presents robot 400 from an external position is also contemplated. And as noted above, depending on the circumstances, there may be no robot at all; instead, the sensor feed may depict all or part of a human participant who intends to leverage aspects of the present disclosure, e.g., to receive guidance in performing some high-level task.
[0079] Second video feed 468B is associated with robotic planner (RP) process 142 acting as a participant. Consequently, no video is provided, and instead, an icon indicating that robotic planner process 142 is a participant is presented in a fashion similar to a human participant with their video feed disabled. Third video feed 468C is associated with a human participant named “Mavis” and presents a video stream of the participant named Mavis. While no other participants are depicted in Figs.4A-F, this is not meant to be limiting. Any number of participants may join such a video conference session. In addition, while not true in all implementations, in some implementations, one participant may join the video conference session as a “supervisor.” In some such implementations, only one participant at a time can be supervisor, e.g., to avoid conflicts.
[0080] A transcript 470 of spoken utterances is provided, e.g., by transcript engine 124, e.g., using STT processing. A message exchange thread 472 is also provided, e.g., by message exchange engine 130, for exchanging textual messages outside of the primary spoken conversation enabled by the video conference session. At bottom are a number of icons, many that are commonly found in video conference clients, such as a microphone icon that can beactuated to toggle between muted and unmuted, a camera icon that can be used to enable and disable a camera, and so forth. Notably, an emoji icon 474 and a “hands up” icon are provided that can be used to provide affect how natural language responses expressing mid-level actions for carrying out respective portions of a high-level task are presented, whether they need approval before being processed to generate robot control data, etc.
[0081] In this example, robotic planner process 142 communicates with video conference participants by incorporating messages into the message exchange thread 472. However, this is not meant to be limiting, and in other implementations, robotic planner process 142 may communicate in other ways, such as in transcript 470.
[0082] In message exchange thread 472, robotic planner process 142 has indicated that its role is “robotic planner, as well that it can be reached using phrases such as “Hey robotic planner” or “@RP.” Then, robotic planner process 142 seeks the attention of another participant, robot 400, by generating the natural language request, “@robot, what is your role?” Robot 400 responds in message exchange thread 472, “my role is body,” which may indicate that robot 400 will be the entity that will be acting upon something (autonomously or under human control), as opposed to the acting entity being a human participant performing actions themselves, without using a robot. In response to this statement from robot 400, robotic planner incorporates additional content into message exchange thread, such as the statement, “assigned body role to robot.”
[0083] Meanwhile, the human participant Mavis utters (or types) natural language, “My name is Mavis and my roles are supervisor and operator.” The “supervisor” role may give Mavis various capabilities that other participants may be denied, such as the ability to accept, reject, and / or modify mid-level actions proposed by robotic planner process 142. The “operator” role may signal that Mavis will be controlling robot 400 to perform mid-level actions proposed by robotic planner process 142. Otherwise, the proposed mid-level actions may be processed as depicted in Fig.2 to generate robot control data that is used to control robot 400 with much less human intervention.
[0084] Referring now to Fig.4B, Mavis has provided (typed or spoken) the following natural language request: “Tell me how to put all the dishes into the top drawer.” As shown in Fig.2, this natural language request may be processed by robotic planner process 142 based on machine learning model(s) 144 (e.g., a multi-modal VQA model)—in some cases in conjunction with the sensor feed data may correspond to video feed 468A—to generate a natural language responsethat expresses a first mid-level action. In Fig.4B, robotic planner process 142 responds, “To put all the dishes into the top drawer, do the following: open the top drawer.” This is shown in message exchange thread 472.
[0085] As shown in video feed 468A, Mavis performed the mid-level by operating robot 400 to open top drawer 482A. Subsequently, in Fig.4C, Mavis types / utters the natural language statement, “I’m done,” to signify that she has completed the task. This is shown in transcript 470 in Fig.4C. This statement may be used, alone or in combination with sensor feed data that provides additional context about the environment in which robot 400 operates, as an acceptance or affirmation of the first mid-level action proposed by robotic planner process 142. In various implementations, rating engine 132 and / or feedback engine 146 may record this episode (e.g., along with an indication of a positive outcome) for purposes of training and / or fine-tuning machine learning model(s) 144.
[0086] Referring back to Fig.4C, in message exchange thread 472, robotic planner process 142 has responded to Mavis’ statement that she completed the first mid-level action by predicting a second mid-level action. As shown in Fig.2, this may include, for instance, assembling (e.g., as a new input prompt) data indicative of the original natural language request, the first mid-level action (and an indication of successful completion in some instances), and the latest sensor feed data from the vision sensor of robot 400 (depicted in video stream 468A). In Fig.4C, the next mid-level action predicted by robotic planner process 142 is “Put the plate into the top drawer.”
[0087] In Fig.4D, Mavis has operated robot 400 to put the plate into top drawer 482A and has once again provided the natural language statement, “I’m done.” In response, robotic planner process 142 predicts a third mid-level action, “put the cup into the top drawer.” In Fig.4E, Mavis has operated robot 400 to put the cup into top drawer 482A and has once again provided the natural language statement, “I’m done.” In response, robotic planner process 142 predicts a fourth mid-level action, “put the bowl into the top drawer.”
[0088] In Fig.4F, Mavis has operated robot 400 to put the bowl into top drawer 482A and has once again provided the natural language statement, “I’m done.” However, after robotic planner 142 responds (in message exchange thread 472) that it is also done, Mavis provides an additional mid-level action, “Close the top drawer.” Mavis then operates robot 400 to close the top drawer and says, “I’m done.”
[0089] In Fig.4F, an “intervention rate” of 20% is calculated, e.g., by rating engine 132. This intervention rate indicates that out of five mid-level actions that needed to be performed in order to accomplish the high-level task of putting the dishes in the top drawer, the human supervisor (Mavis) only had to intervene in one of those tasks, namely, closing top drawer 482A. This intervention rate (20%) may serve as a progress indicator for machine learning model(s) 144 used by robotic planner process 142 to predict mid-level actions from high-level tasks.
[0090] Additionally, the four successfully predicted mid-level actions and the fifth unsuccessfully predicted mid-level action may be logged, e.g., by rating engine 132, and used by training engine 148 as training examples for training machine learning model(s) 144. It may not be necessary for the human supervisor to operate a robot to perform the predicted mid-level actions in order to provide this benefit of generating training data. For example, and as mentioned earlier, a human participant could request instructions on performing a high-level task like rebuilding an engine. Robotic planner process 142 could predict (e.g., iteratively) mid-level actions for rebuilding the engine, and the human participant could follow those instructions (including modifying some as needed) without using a robot at all. The predicted mid-level actions and their outcomes may nonetheless be logged and used by training engine 148 to train machine learning model(s) 144.
[0091] Figs.5A-C depict another video conference session scenario in which a GUI 566 that shares many characteristics with GUI 466 provides information about a video conference session in which a robot 500 is tasked with performing as similar high-level task as before, “Put the dishes in the top drawer” (seen in message exchange thread 572). In this scenario, instead of a human participant operating robot 500 manually, predicted mid-level actions are processed, e.g., by robot control data generation process 152 using machine learning model(s) 154, to generate robot control data. This robot control data is then provided to robot 500, which performs the mid-level actions. Notably, in Figs.5A-C, there are only two participants: the human supervisor (whose video feed is not shown in Figs.5A-C) and robotic planner process 142 (also without a video feed shown). A sensor feed of robot 500 is shown in video feed 568A, but this does not necessarily require that robot 500 be a participant in the video conference session. For example, the robot’s vision sensor (or a vision sensor in the robot’s vicinity) may simply be communicatively coupled with (e.g., patched into) the video conference session, e.g., by sensor feed engine 128. In some implementations, the sensor data that is patched in may be a lowerresolution than the native sensor data generated at the sensor (e.g., 108). This may be acceptable because robotic planner process 142 may not necessarily require sensor data at the same resolution as robot 100. Moreover, streaming sensor data at a lower resolution may preserver computing resources such as network bandwidth.
[0092] In Fig.5A, robotic planner process 142 has predicted a first mid-level action, “open the top drawer.” As shown in message exchange thread, this statement by robotic planner process 142 may be followed by graphical elements that are operable by the human participant to provide feedback, such as a “thumbs up” to approve and a “thumbs down” to disapprove. Depending on factors such as the human participant’s preferences, confidence metrics associated with predicted mid-level actions, etc., it may or may not be required that the human participant approve the predicted mid-level action before robotic planner process 142 predicts a next mid- level action. If not required, then in some implementations, a predetermined amount of time (which may be adjusted based on similar factors) may elapse before robotic planner process moves on to having the current mid-level action processed by robot control data generation process 152 so that robot 500 can perform the action, and predicting the next mid-level action.
[0093] More generally, in various implementations, a temporal frequency at which joint commands or Cartesian commands are generated by robot control data generation process 152 based on natural language snippets may be altered. For example, video conference system 120, robotic planner process 142, and / or robot control data generation process 152 may be transitioned between a first state in which a sequence of natural language snippets are processed without interruption at a first temporal frequency, a second state in which a sequence of natural language snippets are processed without interruption at a second temporal frequency that is different than the first temporal frequency (e.g., slower or faster), and / or a third state in which the sequence of natural language snippets are processed one-at-a-time based on subsequent input from the video conference client.
[0094] In some implementations, a participant may trigger such a transition by, for instance, operating graphical element 576 (“hands up”) to indicate that they’d like to provide feedback on predicted mid-level actions before proceeding further. Additionally or alternatively, if the participant is satisfied and no longer requires approval for each mid-level action, they may operate graphical element 576 to “put their hand down,” and / or may operate graphical element 574 to indicate a general satisfaction with the process.
[0095] In Fig.5B, the human participant did not operate one of the graphical elements beneath the “open the top drawer” mid-level action in a predetermined amount of time. Consequently, that mid-level action was processed, e.g., by robot control data generation process 152, to generate robot control data that was used to operate robot 500 to open top drawer 582A.
[0096] Next, robotic planner process 142 predicts a second mid-level action, “put the plate into the top drawer.” In Fig.5B, the human supervisor rejects this predicted second mid-level action by operating the “thumbs down” icon. The human supervisor provides a modified mid-level action, “put the bowl into the top drawer” instead. Accordingly, robotic planner process 142 repeats this new mid-level action, and includes additional graphical elements to approve or reject it (this repetition is not required in all implementations).
[0097] In Fig.5C, the human supervisor has approved the modified mid-level action, which resulted in robot 500 putting the bowl in top drawer 582A. As shown in message exchange thread 572, robotic planner process 142 predicts the next mid-level action to be “put the plate into the top drawer.” The human supervisor operates the “thumbs up” graphical element to approve this mid-level action, which results in robot 500 putting the plate into top drawer 582A. Robotic planner process 142 then predicts the next mid-level action, “put the cup into the drawer.” This process may be repeated until the high-level is completed to the human supervisor’s satisfaction.
[0098] Referring now to Fig.6, an example method 600 of practicing selected aspects of the present disclosure is described. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. This system may include various components of various computer systems, including video conference system 120. Moreover, while operations of method 600 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted or added.
[0099] At block 602, the system, e.g., by way of sensor feed engine 128, may cause one or more video conference clients of a video conference session to render output that includes one or more sensor feeds (e.g., 108). As noted elsewhere herein, the one or more sensor feeds may capture an environment in which a robot (e.g., 100, 400, 500) operates. For example, a vision sensor such as a two-dimensional camera, LIDAR sensor, microphone, etc., may be mounted on the robot and / or in the robot’s environment.
[0100] At block 604, the system, e.g., by way of participant engine 122, may communicatively couple robotic planner process 142 to the video conference session, e.g., by joining robotic planner process 142 as a participant of the video conference session and / or by otherwise looping robotic planner process 142 into the video conference session.
[0101] At block 606, video conference system 120 may receive, from a video conference client 164 operated by a first participant 160 in the video conference session, a natural language request for the robot to perform a high-level task. At block 606, video conference system 120 may cause an input prompt to be assembled, e.g., by robotic planner process 142, that includes data indicative of (e.g., embeddings generated from) the natural language request. This input prompt may also include data from one or more of the sensor feeds (e.g., 108).
[0102] At block 610, video conference system 120 may cause robotic planner process 142 to process the natural language request using machine learning model(s) 144 to generate a “next” natural language response. This “next” natural language response (and others described herein) may express a “next” mid-level action to be performed by the robot to carry out a respective portion of the high-level task. At block 612, video conference system 120 may cause the next natural language response to be rendered at one or more video conference clients, e.g., so that human(s) operating those client(s) may have a chance to review the next mid-level action, monitor performance of the high-level task, provide feedback, intervene, etc.
[0103] Next, the method 600 may determine whether a score S associated with the next mid- level action satisfies a threshold T. For example the score S may be a confidence score, and the threshold T may be a minimum confidence threshold (e.g., 0.7). If the answer is no, in some implementations, that may cause method 600 to transition into a state in which individual mid- level actions need to be approved before method 600 proceeds further. In some implementations, rather than receiving a metric and performing a comparison itself, video conference system 120 may receive, from robotic planner process 142, a request to alter a temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets. Robotic planner process 142 may issue this request on a similar basis, e.g., that a confidence metric of a mid-level action fails to satisfy a threshold. Based on the request, video conference system 120 may alter the temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets.
[0104] In Fig.6, assuming the threshold was not satisfied, method 600 next determines whether a human participant (e.g., a supervisor) has approved the mid-level action. If the answer is no, then method 600 proceeds to block 616, at which point a modified next mid-level action is received from the user, as demonstrated in Fig.5B, for instance. This modified mid-level action may then be used as the “new” next mid-level action that is used to operate the robot subsequently at blocks 618-624. If the threshold T was satisfied by the score S, on the other hand, then the predicted mid-level action is used to operate the robot subsequently. In either case, method 600 proceeds to block 618.
[0105] At block 618, video conference system 120 causes the next mid-level action (predicted or participant-modified) to be processed, e.g., by robot control data generation process 152, to generate robot control data. At block 620, the system, e.g., by way of video conference system 120 or another component depicted in Fig.1, may cause the robot to be operated based on the robot control data. At block 622, the system, e.g., by way of rating engine 132, may store the next mid-level action as part of a training episode that can be used subsequently to fine tune machine learning model(s) 144.
[0106] Next, video conference system may determine whether there is more to be done to complete the high-level task, e.g., whether there are more mid-level actions to perform to complete the high-level task. If the answer is no, then method 600 ends. However, if the answer is yes, then method 600 proceeds to block 624, where a new input prompt is assembled. This new input prompt may include data indicative (e.g., embeddings, raw data) of, for instance, the original natural language request, the next mid-level action that was just used to operate the robot, and the latest sensor data showing the current state of the environment in which the robot operates. Method 600 may then proceed back to block 610, and a next iteration of the process is performed. This may proceed until the high-level task is completed and / or until a supervisor participant issues a stop command or similar.
[0107] Referring now to Fig.7, an example method 700 of practicing selected aspects of the present disclosure is described. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. This system may include various components of various computer systems, including robotic planner system 140. Moreover, while operations of method 700 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted or added.
[0108] At block 702, the system, e.g., by way of robotic planner process 142, may register robotic planner process 142 as a first participant of a video conference session. At block 704, the system, e.g., by robotic planner process 142 or sensor feed engine 128, may communicatively couple (e.g., register as a participant in the video conference or otherwise) one or more sensor feeds that capture an environment in which a robot operates to the robotic planner process.
[0109] At block 706, the system, e.g., by way of robotic planner process 142, may monitor a transcript and / or message exchange thread of the video conference session, or audio exchanged in the video conference session, e.g., for instances in which the robotic planner process is explicitly invoked, and / or for instances in which a user has provided natural language input. At block 708, the system, e.g., by way of robotic planner process 142, may, based on the monitoring, retrieve a natural language input provided by another participant of the video conference session. As before, the natural language input may express a request to perform a high-level task, e.g., by a robot or otherwise.
[0110] At block 710, the system, e.g., by way of robotic planner process 142, may assemble an input prompt that includes the natural language request, alone or in conjunction with data from one or more of the sensor feeds. At block 712, the system may process the input prompt using machine learning model(s) 144 to generate one or more natural language responses (iteratively or multiple at once). Each natural language response may express a mid-level action to be performed by the robot to carry out a respective portion of the high-level task.
[0111] At block 714, the system may incorporate the one or more natural language responses into the transcript (e.g., 470) of the video conference session, or into a message exchange thread (e.g., 472, 572) that is provided as part of the video conference session. Method 700 may next consider whether a score S associated with the mid-level action satisfies a threshold T, similar as before. If the answer is yes, method proceeds to block 718. If the answer is no, then method proceeds to block 716, similar to before, and the human participant is able to provide a modified mid-level action as the next mid-level action.
[0112] Whichever the case, at block 718, the system, e.g., by way of robotic planner process 142, may cause data indicative of the next mid-level action (e.g., embeddings, the natural language response that expresses it, etc.) to be processed, e.g., by robot control data system 150, using different machine learning model(s) 154 to generate robot control data. This robot control data may or may not be provided to a robot. Method 700 then determines whether more mid-level actions are needed to complete the high-level task. If the answer is no, then method 700 ends. If the answer is yes, however, method 700 proceeds to block 720. At block 720, a new input prompt is assembled, e.g., by robotic planner process 142, that may include any combination of the original natural language request, the next (or more recent) mid-level action, and the latest sensor data if applicable.
[0113] Although not depicted in Fig.7, in various implementations, the mid-level actions can be logged by rating engine 132 as training episodes that can be used by training engine 148 to fine- tune machine learning model(s) 144. In some implementations, training engine 148 may, in addition to training a primary generative language model (e.g., a VQA model) that resides at robotic planner system, jointly train smaller edge-based machine learning models that can be deployed at the edge, e.g., on robots. In some implementations, these smaller edge-based machine learnings may still be VQA generative models or similar, but may include fewer parameters than the primary model, which may include hundreds of millions or billions of parameters.
[0114] In some implementations, machine learning model(s) 144 the mid-level actions predicted by robotic planner process 142 may extend beyond a single robot. For instance, suppose a participant requests performance of a task without specifying a robot, in an environment in which multiple candidate robots are available. Robic planner process 142 may apply machine learning model(s) 144 to generate, e.g., as a first mid-level action, a proposal to use the nearest robot that satisfies constraint(s) of the high-level task. For example, if the high-level task requires a robot with two arms, then robotic planner process 142 may predict a natural language response that expresses the mid-level task, “It looks like your task requires two arms. Let’s look for the nearest bimanual robot.” In some cases in which robotic planner process 142 has access to an inventory of robots and their capabilities, robotic planner process 142 may even present an enumerate list of options for which robot(s) to use.
[0115] Fig.8 is a block diagram of an example computer system 810. Computer system 810 typically includes at least one processor 814 which communicates with a number of peripheral devices via bus subsystem 812. These peripheral devices may include a storage subsystem 824, including, for example, a memory subsystem 825 and a file storage subsystem 826, user interface output devices 820, user interface input devices 822, and a network interface subsystem 816. The input and output devices allow user interaction with computer system 810. Networkinterface subsystem 816 provides an interface to outside networks and is coupled to corresponding interface devices in other computer systems.
[0116] User interface input devices 822 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into the display, audio input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information into computer system 810 or onto a communication network.
[0117] User interface output devices 820 may include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem may include a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual display such as via audio output devices. In general, use of the term "output device" is intended to include all possible types of devices and ways to output information from computer system 810 to the user or to another machine or computer system.
[0118] Storage subsystem 824 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 824 may include the logic to perform selected aspects of method 600 and / or 700, and / or to implement one or more aspects of robot 100 or systems 120, 140, and / or 150. Memory 825 used in the storage subsystem 824 can include a number of memories including a main random-access memory (RAM) 830 for storage of instructions and data during program execution and a read only memory (ROM) 832 in which fixed instructions are stored. A file storage subsystem 826 can provide persistent storage for program and data files, and may include a hard disk drive, a CD-ROM drive, an optical drive, or removable media cartridges. Modules implementing the functionality of certain implementations may be stored by file storage subsystem 826 in the storage subsystem 824, or in other machines accessible by the processor(s) 814.
[0119] Bus subsystem 812 provides a mechanism for letting the various components and subsystems of computer system 810 communicate with each other as intended. Although bus subsystem 812 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0120] Computer system 810 can be of varying types including a workstation, server, computing cluster, blade server, server farm, smart phone, smart watch, smart glasses, set top box, tablet computer, laptop, or any other data processing system or computing device. Due to the ever- changing nature of computers and networks, the description of computer system 810 depicted in Fig.8 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of computer system 810 are possible having more or fewer components than the computer system depicted in Fig.8.
[0121] While several implementations have been described and illustrated herein, a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein may be utilized, and each of such variations and / or modifications is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. It is, therefore, to be understood that the foregoing implementations are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.
Claims
CLAIMS What is claimed is:
1. A method implemented using one or more processors and comprising: causing one or more video conference clients of a video conference session to render output that includes one or more sensor feeds, wherein the one or more sensor feeds capture an environment in which a robot operates; communicatively coupling a robotic planner process to the video conference session; receiving, from a video conference client operated by a first participant in the video conference session, a natural language request for the robot to perform a high-level task; causing the robotic planner process to process the natural language request using a first machine learning model to generate a plurality of natural language responses, each natural language response expressing a mid-level action to be performed by the robot to carry out a respective portion of the high-level task; causing one or more of the natural language responses to be processed to generate robot control data; and causing the robot to be operated based on the robot control data.
2. The method of claim 1, wherein the robotic planner processes the natural language request over multiple iterations to generate the plurality of natural language responses, wherein during at least one of the iterations, the natural language request and a natural language response from a previous iteration are processed to generate a next natural language response.
3. The method of claim 1, wherein the communicatively coupling comprises registering, as a second participant of the video conference session, the robotic planner process.
4. The method of claim 1, wherein the one or more natural language responses are processed using a second machine learning model to generate the robot control data.
5. The method of claim 1, wherein the one or more sensor feeds include a video feed captured by a camera that is mounted on the robot.
6. The method of claim 1, wherein the one or more sensor feeds include a video feed captured by a camera that is deployed in the environment separately from the robot.
7. The method of claim 1, wherein the one or more sensor feeds include a light detection and ranging (LIDAR) feed captured by a LIDAR sensor.
8. The method of claim 1, wherein the one or more sensor feeds include audio data captured by a microphone.
9. The method of claim 1, wherein the robot planner process processes data from one or more of the sensor feeds in conjunction with the natural language request using the first machine learning model to generate the one or more natural language responses.
10. The method of claim 1, further comprising: receiving, from a video conference client operated by the first participant or a second participant of the video conference session, feedback on a given natural language response of the one or more natural language responses generated using the first machine learning model; and modifying the given natural language response based on the feedback to generate a modified natural language response that conveys a modified mid-level action to be performed by the robot to carry out a portion of the high-level task.
11. The method of claim 10, wherein the modified natural language response is processed to generate at least some of the robot control data.
12. The method of claim 10, wherein the feedback comprises input generated in response to actuation of a graphical element that is operable to reject the given natural language response.
13. The method of claim 10, further comprising storing the modified natural language response as training data for training the first machine learning model.
14. The method of claim 10, further comprising causing the first machine learning model and one or more edge-based machine learning models with less parameters than the first machine learning model to be jointly trained based on the training data.
15. The method of claim 1, further comprising receiving, from a video conference client operated by the first participant or a second participant of the video conference session, feedback on a given natural language response of the one or more natural language responses generated using the first machine learning model, wherein the feedback comprises acceptance or approval of the given natural language response, and the acceptance or approval is generated by the first or third participant using a graphical element.
16. The method of claim 1, further comprising: receiving, from a video conference client operated by the first participant or a second participant of the video conference session, a request to alter a temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets; and based on the request, altering the temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets.
17. The method of claim 16, wherein altering the temporal frequency comprises transitioning between a first state in which a sequence of natural language snippets are processed without interruption and a second state in which the sequence of natural language snippets are processed one-at-a-time based on subsequent input from the video conference client.
18. The method of claim 1, further comprising: receiving, from the planner process, a request to alter a temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets; and based on the request, altering the temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets.
19. The method of claim 1, further comprising, based on one or more metrics associated with one or more of the natural language responses generated by the robotic planner process, altering a temporal frequency at which joint commands or Cartesian commands are generated based on natural language snippets.
20. The method of claim 1, further comprising causing one or more of the video conference clients to render the one or more natural language responses as additional output.
21. The method of claim 1, wherein the first machine learning model comprises a generative language model.
22. A method implemented using one or more processors, the method comprising executing a robotic planner process to perform the following operations: registering the robotic planner process as a first participant of the video conference session; communicatively coupling the robotic planner process to one or more sensor feeds that capture an environment in which one or more robots operate; monitoring a transcript, message exchange thread, and / or audio of the video conference session;based on the monitoring, retrieving a natural language input incorporated provided by another participant of the video conference session, wherein the natural language input expresses a request for one or more of the robots to perform a high-level task; processing the natural language request and data from one or more of the sensor feeds using a first machine learning model to generate one or more natural language responses, each natural language response expressing a mid-level action to be performed by one or more of the robots to carry out a respective portion of the high-level task; and incorporating the one or more natural language responses into the transcript, message exchange thread, or audio of the video conference session, or into a message exchange thread that is provided as part of the video conference session.
23. The method of claim 22, further comprising causing one or more of the natural language responses to be processed using a second machine learning model to generate robot control data.
24. The method of claim 22, wherein the first machine learning model comprises a visual question answering (VQA) model.
25. The method of claim 22, wherein processing the natural language request and data from one or more of the sensor feeds using the first machine learning model comprises processing the natural language request, the data from one or more of the sensor feeds, and a natural language response generated during a previous iteration of the first machine learning model being applied to the natural language request.
26. The method of claim 22, wherein the one or more robots comprise a plurality of robots, and the mid-level action comprises selecting a given robot from the plurality of robots to perform a respective portion of the high-level task.
27. A system comprising one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to execute a video conference server and a robotic planner process: wherein the video conference server is configured to: establish a video conference session between one or more video conference clients;cause one or more of the video conference clients to render output that includes one or more sensor feeds, wherein the one or more sensor feeds capture an environment in which a robot operates; receive, from video conference client operated by a first participant in the video conference session, a natural language request for the robot to perform a high-level task; cause one or more natural language responses generated by the robotic planner process to be processed to generate robot control data, wherein each natural language response expresses a mid-level action to be performed by the robot to carry out a respective portion of the high-level task; and transmit the robot control data to the robot; wherein the robotic planner process is configured to: join the video conference session as a second participant; and process the natural language request and data from one or more of the sensor feeds using a visual question answering (VQA) model to generate the one or more natural language responses.
28. The system of claim 27, wherein the robot planner process is further configured to incorporate the one or more natural language responses into a message exchange thread that is provided as part of the video conference session, or into a transcript of the video conference session, and that is rendered at one or more of the video conference clients.
29. The system of claim 27, wherein the video conference server is further configured to transmit the one or more natural language responses generated by the robot planner process to a remote computing device, wherein the one or more natural language responses are operable by the remote computing device to generate the robot control data using a machine learning model.
Citation Information
Patent Citations
Determining and utilizing corrections to robot actions
EP3628031B1
Methods and systems for using voice input to control a surgical robot
US11457983B1
Interactive cost corrections with natural language feedback
US20230271330A1
Interpreting discrete tasks from complex instructions for robotic systems and applications
US20230297074A1
Cited By
Executing a collaboration session for ai agent control of a robotic device
WO2026135786A1
Executing a collaboration session for human teleoperation of a robotic device participant
WO2026135788A1
Graphically representing an ai agent participant in a collaboration session
WO2026177787A1