Efficient robot control based on input from a remote client device

By utilizing user interface input from remote client devices to define object manipulation parameters and train machine learning models, the inefficiencies of traditional robot control methods are addressed, resulting in more efficient and flexible robotic operations.

JP7675264B2Active Publication Date: 2025-05-12GOOGLE LLC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024102342
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-26
Filing Date
2024-06-25
Publication Date
2025-05-12
Estimated Expiration
2040-07-24

AI Technical Summary

Technical Problem

Existing robot control methods are inefficient when dealing with a wide variety of actions and components, as they require significant pre-programming and reconfiguration of environments, and often rely on human guidance, leading to idle robot time.

Method used

The use of user interface input from remote client devices to control robots, where object manipulation parameters such as gripping posture and trajectory are defined and used to train machine learning models, allowing for predictive control of robotic actions.

Benefits of technology

This approach reduces the need for extensive pre-programming and human intervention, enabling robots to operate more efficiently by predicting and automatically executing object manipulation parameters, thus minimizing idle time and enhancing operational flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675264000001
    Figure 0007675264000001
  • Figure 0007675264000002
    Figure 0007675264000002
  • Figure 0007675264000003
    Figure 0007675264000003
Patent Text Reader

Abstract

To perform efficient robot control through utilization of user interface inputs from remote client devices.SOLUTION: Implementations relate to generating training instances based on object manipulation parameters, defined by instances of user interface inputs, and training machine learning models to predict the object manipulation parameters. Those implementations can subsequently utilize the trained machine learning models to reduce a quantity of instances that inputs from remote client devices are solicited in performing a given set of robotic manipulations and / or to reduce an extent of inputs from the remote client devices in performing the given set of robotic manipulations. Implementations are additionally or alternatively related to reducing idle time of robots through utilization of vision data that captures objects, to be manipulated by a robot, prior to the objects being transported to a robot workspace within which the robot can reach and manipulate the object.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to efficient control of a robot based on input from a remote client device. [Background technology]

[0002] In industrial or commercial environments, robots are often preprogrammed to perform specific tasks repeatedly. For example, a robot may be preprogrammed to repeatedly fasten a particular assembly component on an assembly line. Also, for example, a robot may be preprogrammed to repeatedly grasp and move a particular assembly component from a first fixed location to a second fixed location. In grasping an object, a robot may use a grasping end effector, such as an "impactive" end effector (e.g., using a "claw" or other finger to apply force to an area of ​​the object), an "ingressive" end effector (e.g., using pins, needles, etc. to physically penetrate the object), an "astrictive" end effector (e.g., using suction or vacuum to lift an object), and / or one or more "contigutive" end effectors (e.g., using surface tension, freezing, or adhesives to lift an object).

[0003] Such approaches may work well in environments where constrained actions are repeatedly performed on a constrained set of components. However, such approaches may not work in environments where the robot is tasked with performing a wide variety of actions and / or performing actions on a diverse set of components, optionally including new components for which the robot was not pre-programmed. Moreover, such approaches require significant engineering effort (and associated use of computational resources) to pre-program the robot. Moreover, such approaches may require significant reconfiguration of the industrial or commercial environment in order to provision the robot in the environment.

[0004] Alternatively, some human-in-the-loop approaches have been proposed in which a human repeatedly provides the same type of guidance to assist the robot in performing a task. However, such approaches may suffer from various drawbacks. For example, some approaches may result in the robot being idle while seeking and / or waiting for human guidance, which leads to inefficient operation of the robot. Also, for example, some approaches always seek human guidance and / or the same type of guidance, which limits the ability of the robot to operate more efficiently. Summary of the Invention [Means for solving the problem]

[0005] Implementations disclosed herein relate to utilizing user interface inputs from a remote client device in controlling a robot in an environment. An instance of a user interface input provided at a remote client device indicates (directly or indirectly) one or more object manipulation parameters used by the robot in manipulating at least one object. For example, the object manipulation parameters indicated by an instance of a user interface input may include one or more of a grasp pose, a placement pose, a sequence of waypoints encountered in traversing to the grasp pose, a sequence of waypoints encountered in traversing to the placement pose (after grasping the object), a complete path or trajectory (i.e., a path with velocity, acceleration, jerk, and / or other parameters) in traversing to and / or from a manipulation posture (e.g., a grasp posture or other manipulation posture), and / or other object manipulation parameters such as, but not limited to, object manipulation parameters described in further detail herein.

[0006] User interface inputs of an instance are provided with reference to a visual representation that includes an object representation of at least one object. The visual representations may also optionally include environmental representations of other environmental objects (e.g., a work surface, a container in which the at least one object is to be placed), and / or a robot representation of all or a portion of the robot. The visual representations may be rendered, for example, on a standalone display screen controlled by a remote client device or a virtual reality (VR) headset controlled by a remote client device. Input interface inputs may be provided, for example, via a mouse, a touch screen, a VR hand controller, and / or a VR glove. Further, additional descriptions of example visual representations and how they may be rendered are provided herein, including descriptions of implementations that generate the visual representations in a manner that reduces network traffic and / or reduces latency in rendering the visual representations.

[0007] Some implementations disclosed herein are directed to generating training instances based on object manipulation parameters defined by instances of user interface inputs. The implementations are further directed to training a machine learning model based on the training instances for use in predicting the object manipulation parameters. In some of the implementations, a training instance may be generated and / or labeled as a positive training instance in response to determining that a success measure of an attempted manipulation based on the corresponding object manipulation parameters meets a threshold. The success measure may be generated based on sensor data from one or more sensors and may be generated in a manner that depends on the manipulation being performed. As an example, if the manipulation is a grasp with an impact end effector, the success measure may indicate whether the grasp was successful. The success measurement may be based on one or more of, for example, sensor data from a sensor of the impact end effector (e.g., using finger position determined based on data from a position sensor and / or torque indicated by a torque sensor to determine whether the impact end effector has grasped an object), vision data from a vision sensor of the robot (e.g., to determine whether the impact end effector has grasped an object and / or whether the object has moved from its previous location), weight sensors in the environment (e.g., to determine whether the object has been lifted from a location and / or placed in another location), etc. As another example, if the operation includes grasping an object and subsequent placement of the object in a container, the success measurement of the placement operation may indicate whether the object was successfully placed in the container and / or the degree to which the placement in the container conforms to the desired placement. As yet another example, if the operation includes joining two objects, the success measurement of the placement operation may indicate whether the objects were successfully joined together and / or the degree of accuracy of the joining.

[0008] Implementations that train a machine learning model based on the generated training instances are further directed to later utilizing the trained machine learning model. Utilizing the trained machine learning model reduces the amount of instances where input is required from a remote client device when performing a given set of robotic manipulations (thereby reducing network traffic) and / or reduces the extent of input from a remote client device when performing a given set of robotic manipulations (thereby providing efficient resource utilization at the remote client device). These implementations may enable robots in an environment to operate more efficiently by reducing instances and / or durations during which the robot is idle while waiting for user interface input. These implementations may further facilitate the performance of technical tasks of controlling, supervising, and / or monitoring the operation of one or more robots by an operator of a remote client device, and may enable the operator to provide a greater amount of manipulations and / or a greater amount of input for the robot in a given time.

[0009] As one particular example, assume that one or more robots are newly deployed in a given environment to perform operations, each of which involves grasping a corresponding object from a conveyor belt and placing the object in an appropriate one of N available containers (e.g., shipping boxes). Initially, user interface inputs may be solicited for each operation to determine object manipulation parameters, including any one, any combination, or all of the following: a sequence of waypoints encountered in traversing to a grasp pose for grasping the object, a sequence of waypoints encountered in traversing to the appropriate one of the N available containers, and a placement pose for placing the object in the container. These determined manipulation parameters may be utilized to control the robot in performing the operation.

[0010] Over time, training instances may be generated for each of the one or more machine learning models based on the corresponding vision data (and / or other sensor data), one or more of the object manipulation parameters, and optionally, a success measurement. Each of the machine learning models may be trained to process the vision data and / or other sensor data in predicting one or more corresponding manipulation parameters. Additionally, the machine learning models may be trained based on the training instances. For example, assume a machine learning model that is trained for use in processing the vision data to generate a corresponding probability (e.g., of successful object manipulation, e.g., grasping) for each of N gripping poses. Positive training instances may be generated for manipulations that included successful grasps (determined based on the grasp success measurement) based on the corresponding vision data and the corresponding gripping poses defined by user interface input.

[0011] The trained machine learning model can then be at least selectively utilized in predicting one or more corresponding object manipulation parameters, which are then at least selectively utilized to control the robot. For example, the predicted object manipulation parameters can be utilized automatically (without prompting a confirmatory user interface input) and / or can be utilized after presenting an indication of the predicted object manipulation parameters (e.g., as part of a visual representation) and receiving a confirmatory user interface input in response. In these and other ways, the object manipulation parameters can be determined and utilized without requiring a user interface input (e.g., when the object manipulation parameters are utilized automatically) and / or with a reduced amount of user interface input (e.g., when a confirmatory user interface input is provided instead of a more time-consuming complete input to define the object manipulation parameters). This can reduce the duration of time required to determine the object manipulation parameters, allowing the robot and / or remote operator to work more efficiently.

[0012] In some implementations, the trained machine learning model is only utilized in predicting the object manipulation parameters to be at least selectively utilized after determining that one or more conditions are satisfied. The one or more conditions may include, for example, at least a threshold amount of training and / or validation of the trained machine learning model. Validating the trained machine learning model may include comparing predictions generated using the machine learning model--optionally for instances of vision data (and / or other sensor data) for which the machine learning model was not trained--to ground truth object manipulation parameters based on user interface inputs. In various implementations, as described herein, the trained machine learning model may continue to be trained even after it has been actively utilized in predicting the object manipulation parameters to be at least selectively utilized for the robot's operation. For example, additional training instances may be generated based on the predicted and utilized object manipulation parameters and labeled as positive or negative based on a determined measure of success. Also, for example, additional training instances may be generated based on the predicted object manipulation parameters and labeled as negative if the user interface input rejects the predicted object manipulation parameters.

[0013] As one particular example, consider again a machine learning model trained for use in processing vision data to generate corresponding probabilities for each of N grip postures. When vision data is processed using the trained machine learning model, resulting in a probability for a corresponding grip posture that exceeds a first threshold (e.g., 85% or other threshold), the grip posture may be automatically utilized without prompting a confirmatory user interface input. If the grip posture does not exceed the first threshold, but the probability for the grip posture does exceed a second threshold (e.g., 50% or other threshold), an indication of one or more of the grip postures may be presented with the object representation in the visual representation, and one grip posture may be utilized only if a confirmatory input is directed to that grip posture. If the grip posture does not exceed the first or second threshold, a user interface input may be solicited to determine the grip posture without providing any indication of the predicted grip posture. The grip postures determined based on the user interface input may be utilized to generate training instances, optionally also taking into account a measure of grip success. The training instances can then be utilized to further train the model. It is noted that such training instances are "hard negative" training instances, which can be particularly beneficial for efficiently updating parameters of a machine learning model to improve the accuracy and / or robustness of the model.

[0014] Thus, for a given deployment of a robot in an environment, initially, instances of user interface inputs may be utilized to determine object manipulation parameters utilized to control the robot in performing the manipulation. Additionally, training instances may be generated based on the object manipulation parameters determined using the instances of user interface inputs, as well as based on corresponding vision data and / or other sensor data, and optionally based on a measure of success determined based on the sensor data. The training instances may be utilized to train a machine learning model for utilization in predicting the object manipulation parameters. In response to satisfaction of one or more conditions, the trained machine learning model may then be brought “online” and utilized in generating predicted object manipulation parameters. The predicted object manipulation parameters are at least selectively automatically utilized to control the robot and / or utilized when a corresponding indication of the predicted object manipulation parameters is rendered on a remote client device and a confirmatory user interface input is received in response. Additionally, even after being brought online, the trained machine learning model may continue to be trained, increasing its accuracy and efficiency, thereby increasing the amount of instances in which predictions may be automatically utilized to control the robot and / or rendered as suggestions for confirmatory approval.

[0015] In these and other ways, the robot can be deployed and immediately utilized in new environments and / or for new tasks without requiring significant use of technology and / or computational resources prior to deployment. For example, the object manipulation parameters utilized initially upon deployment may be based firmly (or even exclusively) on user interface input from the remote device. However, over time, the user interface input from the remote device may be utilized to train machine learning models that are brought online to lower the amount and / or degree of user interface input required in manipulating the robot within the environment. This allows the robot to operate more efficiently within the environment and reduces the amount of network traffic to the remote device for a given amount of robotic manipulation. Additionally, this allows the operator of the remote client device to assist in controlling a greater amount of robotic manipulation.

[0016] Some implementations disclosed herein are additionally or alternatively directed to specific techniques for determining object manipulation parameters for manipulating a given object based on user interface input from a remote operator. Some of those implementations are directed to techniques for mitigating (e.g., reducing or eliminating) robot idle time while waiting for user interface input to be provided. Reducing robot idle time improves the overall efficiency of the robot's operations.

[0017] Some implementations attempt to mitigate robot idle time through the utilization of vision data that captures an object to be manipulated by the robot before the object is transported to a robot workspace where the robot can reach and manipulate the object. For example, a vision component (e.g., a monographic and / or stereographic camera, a LIDAR component, and / or other vision component) may have a view of a first region of an environment different from the robot workspace. The vision data from the vision component may capture characteristics of the object when it is in the first region before the object is transported to the robot workspace. For example, the first region may be a portion of a conveyor system, which transports the object from the portion to the robot workspace. The vision data capturing the object in the first region may be used to generate a visual representation including at least an object representation of the object that is generated based on the object characteristics of the object captured in the vision data.

[0018] The visual representation may be transmitted to the remote client device before the transfer of the object to the robot workspace is completed (e.g., while the object is being transported by the conveyor system, but before the object arrives at the robot workspace). Additionally, data can be received from the remote client device before the transfer of the object to the robot workspace is completed, the data being generated based on user interface inputs that are intended for the visual representation when rendered at the remote client device.

[0019] The received data directly or indirectly indicates one or more object manipulation parameters for manipulating the object within the robot workspace. Thus, the object manipulation parameters can be determined based on the data, and optionally can be determined prior to completion of the transfer of the object to the robot workspace. The determined object manipulation parameters can be utilized to control the robot to have the robot manipulate the object after the object is transferred to the robot workspace when the object is within the robot workspace. Because at least the visual representation is transmitted and the response data is received prior to completion of the transfer of the object to the robot workspace, the robot can rapidly manipulate the object based on the manipulation parameters determined based on the data once the object is within the robot workspace. For example, the robot can determine when the object is within the robot workspace based on vision data from its vision components, and act according to the object manipulation parameters in response to such determination. The robot can optionally wait for the object to be in a pose corresponding to the pose for which the object manipulation parameters are defined, or can translate the object manipulation parameters into a newly detected pose of the object within the robot workspace (e.g., when the pose changes from the pose for which the object manipulation parameters are defined). If the robot workspace itself includes a conveyor section along which objects are transported, that conveyor section may optionally be temporarily stopped while the robot manipulates the object. In other implementations, objects can be transported to the robot workspace using a conveyor or other transport means (e.g., air tubes, a separate transport robot, human hands) and the robot workspace itself may not include a conveyor section.

[0020] Optionally, when the trained machine learning model is brought online for use in predicting the object manipulation parameters, the vision data from the first region may be utilized in predicting the object manipulation parameters. This allows the object manipulation parameters to be predicted prior to completion of the transfer of the object to the robot working area. The predicted object manipulation parameters may be automatically used as part of the object manipulation parameters and / or an indication of the predicted object manipulation parameters may be provided along with the visual representation--and if the received data indicates confirmation of the predicted object manipulation parameters, one or more of the predicted object manipulation parameters may be utilized.

[0021] In some implementations, the pose of the vision honoring element in the first region and the pose of the robot vision component are known, allowing for the determination of a transformation between a frame of reference of the vision component in the first region and a robot coordinate system of the robot vision component. Using this transformation allows inputs at the remote client device to be defined directly in the robot coordinate system or first defined in the first frame of reference and then transformed to the robot coordinate system.

[0022] In some implementations, the visual representation sent to the remote client device includes an object representation of the object and, optionally, one or more object representations of other nearby dynamic objects (that are dynamic in the first region), but omits other portions of the first region that are static. In some of those implementations, only the representation of the object and, optionally, the nearby dynamic objects are rendered at the remote client device. In some other implementations, all or a portion of the robot and / or the robot workspace is also rendered at the remote client device (even though they are not captured in the vision data capturing the first region). For example, the remote client device may run a robot simulator or communicate with an additional device that runs a robot simulator. The robot simulator may simulate all or a portion of the robot and / or all or a portion of the robot workspace and may render a simulation of the object along with the simulation of the robot and / or the simulation of the robot workspace. The pose of the object relative to the simulation of the robot and / or the simulation of the robot workspace may be determined using the transformations described above. This may enable a human operator to provide user interface inputs to manipulate the simulation of the robot to define object manipulation parameters. For example, to define a gripping pose, a human operator can provide user interface inputs that adjust a simulation of the robot until it is in a desired pose, and then provide further user interface inputs to define the desired pose as a gripping pose.

[0023] Implementations simulating a robot and / or a robot workspace allow visual representations of smaller data size to be transmitted from the environment to a remote client device. This may be a result of those transmissions defining only dynamic objects and not static features of the robot workspace and / or not defining features of the robot. In addition to saving network resources, this may reduce delays in rendering visual representations at a remote device, since smaller data sizes can be transmitted to and / or rendered at a remote client device more quickly. This reduction in delays may similarly reduce idle time of the robot. Additionally, it is noted that even in implementations where object representations are generated based on robot vision data (instead of vision data from a different region), simulating a robot and / or a robot workspace may still allow visual representations of smaller data size to be transmitted--and reduce idle time of the robot.

[0024] Some implementations additionally or alternatively attempt to mitigate idle time of the robot by generating object representations that render the object with less precision than the full representation, but with a smaller data size than the full representation, of the visual representation rendered at the client device. For example, an object may be represented by one or more bounding boxes and / or other bounding shapes that approximate the surface of the object. For example, an object may be defined by multiple connected bounding boxes, each of which may be defined by a center point, a height dimension, and a width dimension--which includes significantly less data than a representation that defines the color, texture, and / or depth of each pixel or voxel that corresponds to the surface of the object. In addition to conserving network resources, the less precise object representations may mitigate delays in rendering visual representations at the remote device, since the smaller data size may be transmitted and / or rendered more quickly at the remote client device. Additionally, a less accurate object representation can obscure or remove potentially sensitive data from the object or obscure the object itself, preventing an operator of a remote device from locating the data and / or object.

[0025] Although some examples are described herein in connection with manipulations involving grasping and / or placing an object, it is understood that the techniques described herein may be utilized for a variety of robotic manipulations of objects. For example, the techniques may be utilized for manipulations involving pushing and / or pulling an object to move it to a different location and / or to mate the object with another object. Also, for example, the techniques may be utilized for manipulations including grasping a first object, grasping a second object, joining the first and second objects together, and placing the joined object in a particular location. As yet another example, the techniques may be utilized for manipulations including acting on an object with an end effector including an etching tool, a screwdriver tool, a cutting tool, and / or other tools.

[0026] The above description is provided as a summary of some implementations of the present disclosure. Further descriptions of those and other implementations are provided in more detail below.

[0027] Other implementations may include a transitory or non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform methods such as one or more of the methods described above and / or elsewhere herein. Still other implementations may include one or more computers and / or one or more robotic systems including one or more processors operable to execute the stored instructions to perform methods such as one or more of the methods described above and / or elsewhere herein.

[0028] It should be understood that all combinations of the above concepts, and additional concepts described in more detail herein, are contemplated as part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as part of the subject matter disclosed herein. [Brief description of the drawings]

[0029] [Figure 1A] FIG. 1 illustrates an exemplary environment in which the implementations described herein may be implemented. [Figure 1B] 1B illustrates an example of how the components of FIG. 1A may interact according to various implementations described herein. [Figure 2A] FIG. 13 illustrates an example of rendering a visual representation at a remote client device including an object representation of an object to be manipulated by a robot, and examples of user interface input that may be provided to define and / or confirm object manipulation parameters for the manipulation of the object by the robot. [Figure 2B] FIG. 13 illustrates an example of rendering a visual representation at a remote client device including an object representation of an object to be manipulated by a robot, and examples of user interface input that may be provided to define and / or confirm object manipulation parameters for the manipulation of the object by the robot. [Figure 2C] FIG. 13 illustrates an example of rendering a visual representation at a remote client device including an object representation of an object to be manipulated by a robot, and examples of user interface input that may be provided to define and / or confirm object manipulation parameters for the manipulation of the object by the robot. [Figure 2D] FIG. 13 illustrates an example of rendering a visual representation at a remote client device including an object representation of an object to be manipulated by a robot, and examples of user interface input that may be provided to define and / or confirm object manipulation parameters for the manipulation of the object by the robot. [Figure 2E]FIG. 13 illustrates an example of rendering a visual representation at a remote client device including an object representation of an object to be manipulated by a robot, and examples of user interface input that may be provided to define and / or confirm object manipulation parameters for the manipulation of the object by the robot. [Figure 2F] FIG. 13 illustrates an example of rendering a visual representation at a remote client device including an object representation of an object to be manipulated by a robot, and examples of user interface input that may be provided to define and / or confirm object manipulation parameters for the manipulation of the object by the robot. [Diagram 3] 1 is a flow diagram illustrating an example method for causing a robot to manipulate an object according to object manipulation parameters determined based on data generated in response to a visual representation including an object representation of the object at a remote client device. [Figure 4] 1 is a flow diagram illustrating an example method for generating training instances based on a robot's object manipulation attempts and using the training instances in training a predictive model. [Diagram 5] 1 is a flow diagram illustrating an example method for selectively utilizing a trained predictive model to determine object-manipulation parameters for use by a robot in manipulating an object. [Figure 6] 1 is a flow diagram illustrating an example method for training a predictive model, validating a predictive model, deploying a predictive model, and optionally further training the deployed predictive model. [Figure 7] FIG. 1 illustrates a schematic diagram of an exemplary architecture of a robot. [Figure 8] FIG. 1 is a schematic diagram illustrating an exemplary architecture of a computer system. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0030] FIG. 1A illustrates an example environment in which the implementations described herein may be implemented. FIG. 1A includes a first robot 170A and associated robot vision components 174A, a second robot 170B and associated robot vision components 174B, and an additional vision component 194. The additional vision component 194 may be, for example, a monographic camera (e.g., generating 2D RGB images), a stereographic camera (e.g., generating 2.5D RGB images), a laser scanner (e.g., generating a 2.5D "point cloud"), and may be operably connected to one or more systems (e.g., system 110) disclosed herein. Optionally, multiple additional vision components may be provided, and vision data from each may be utilized as described herein. The robot vision components 174A and 174B can be, for example, monographic cameras, stereographic cameras, laser scanners, and / or other vision components--and vision data therefrom can be provided to and utilized by the corresponding robots 170A and 170B as described herein. Although shown adjacent to the robots 170A and 170B in FIG. 1A, in other implementations the robot vision components 174A and 174B can alternatively be directly coupled to the robots 170A and 170B (e.g., mounted near the end effectors 172A and 172B).

[0031] The robots 170A and 170B, the robot vision components 174A and 174B, and the additional vision component 194 are all deployed within an environment, such as a manufacturing facility, a packaging facility, or other environment. The environment may include additional robots and / or additional vision components, but for simplicity, only the robots 170A and 170B and the additional vision component 194 are shown in FIG.

[0032] Robots 170A and 170B are "robot arms" having multiple degrees of freedom to enable traversal of their corresponding gripping end effectors 172A, 172B along any of multiple potential paths to position the gripping end effectors at desired locations. Robots 170A and 170B further control two opposing "claws" of their corresponding gripping end effectors 172A, 172B, respectively, to drive the claws between at least open and closed positions (and / or optionally multiple "partially closed" positions). Although specific robots 170A and 170B are shown in FIG. 1A, additional and / or alternative robots may be utilized, including additional robot arms similar to robots 170A and 170B, robots having other robot arm configurations, robots having humanoid configurations, robots having animal configurations, robots that move via one or more wheels, unmanned aerial vehicles ("UAVs"), and the like. Also, while particular gripping end effectors 172A and 172B are shown in FIG. 1A, additional and / or alternative end effectors (or even no end effectors) may be utilized, such as alternative impact-type gripping end effectors (e.g., those with gripping "plates," those with more or fewer "fingers" / "claws"), "penetrating" gripping end effectors, "convergent" gripping end effectors, or "contact" gripping end effectors, or non-gripping end effectors (e.g., welding tools, cutting tools, etc.). For example, a convergent end effector having multiple suction cups may be used to pick up and place multiple objects (e.g., four objects may be picked up and placed at once through the use of multiple suction cups).

[0033] The robot 170A, in FIG. 1A , includes sunglasses 192A on a conveyor portion 103A of a conveyor system, and can access the robot workspace 101A, which also includes a container 193A. The robot 170A can utilize the object manipulation parameters determined as described herein in grasping the sunglasses 192A and appropriately placing them in the container 193A. Just as different containers can be on the conveyor portion 103A in the robot workspace 101A at different times (e.g., containers can be placed by another system or on another conveyor system), other objects can be on the conveyor portion 103A in the robot workspace 101A at different times. For example, as the conveyor system moves, other objects can be transported into the robot workspace 101A and manipulated by the robot 170A while in the robot workspace 101A. The robot 170A can similarly utilize the corresponding object manipulation parameters to pick up, place, and / or perform other operations on such objects.

[0034] 1A includes a stapler 192B on a conveyor portion 103B of a conveyor system, and can access a robot workspace 101B that also includes a container 193B. The robot 170B can utilize the object manipulation parameters determined as described herein in grasping the stapler 192B and properly placing the stapler 192B within the container 193B. Just as different containers can be on the conveyor portion 103B in the robot workspace 101B at different times, other objects can be on the conveyor portion 103B in the robot workspace 101B at different times. The robot 170B can similarly utilize the corresponding object manipulation parameters to pick up and place such objects, and / or perform other operations on such objects.

[0035] The additional vision component 194 has a view of an area 101C, which is different from the robot workspace 101A and different from the robot workspace 101B. In FIG. 1A, the area includes a conveyor portion 103C of a conveyor system, which also includes a spatula 192C. The area 101C can be "upstream" of the robot workspace 101A and / or the robot workspace 101B in that the objects to be manipulated pass through the area 101C first before being transported to the robot workspace 101A or the robot workspace 101B. For example, the conveyor system can pass the objects through the area 101C first before they are sent by the conveyor system to either the robot workspace 101A or the robot workspace 101B. For example, in FIG. 1A, the spatula 192C is in the area 101C, but has not yet been transported to the robot workspace 101A or the robot workspace 101B.

[0036] As described in detail herein, in various implementations, the additional vision component 194 can capture vision data that captures features of the spatula 192C. Additionally, the vision data can be utilized by the system 110 (described below) to determine object manipulation parameters to enable the robot 170A or the robot 170B to manipulate (e.g., pick and place) the spatula 192C. For example, the system 110 can determine the object manipulation parameters based at least in part on user interface input from the remote client device 130 that is directed to a visual representation that is generated based at least in part on the vision data captured by the additional vision component 194 (e.g., based at least in part on object features of the vision data that captures features of the spatula 192C). By utilizing additional vision components 194 "upstream" of the robot workspaces 101A and 101B--before the spatula 192C enters the robot workspace 101A or the robot workspace 101B (i.e., before the transfer of the spatula into either of the robot workspaces 101A, 101B is completed)--a visual representation can be provided to the remote client device 130, user interface input can be given at the remote client device 130, and / or object manipulation parameters can be determined based on data corresponding to the user interface input. In these and other ways, the robots 170A and 170B can operate more efficiently because object manipulation parameters for manipulating an object can be determined quickly, optionally even before the object reaches the robot workspaces 101A and 101B.

[0037] The example environment of FIG. 1A also includes a system 110, a remote client device 130, a training data engine 143, a training data database 152, a training engine 145, and one or more machine learning models 165 (also referred to herein as “predictive models”).

[0038] The system 110 may be implemented by one or more computing devices. The one or more computing devices may be located within the environment with the robots 170A and 170B and / or may be located in a remote server farm. The system 110 includes one or more prediction engines 112, visual representation engines 114, and operation parameter engines 116. The system 110 may perform one or more (e.g., all) of the operations of the method 300 of FIG. 3 and / or the method 500 of FIG. 5, both of which are described in detail below.

[0039] The remote client devices 130 can optionally be within the environment, but in various implementations are located in different structures that can be miles away from the environment. The remote client devices 130 include a display engine 132, an input engine 134, and input devices 136. It is noted that in various implementations, multiple remote client devices 130 can access the system 110 at any given time. In those implementations, a given remote client device 130 can be selected at a given time based on various considerations such as whether a given remote client device 130 has pending requests in its queue, the amount of pending requests in its queue, and / or the expected duration to address the pending requests in its queue.

[0040] The prediction engines 112 of the system 110 can receive vision data from the vision components 194, 174A, and / or 174B, and optionally other sensor data. The prediction engines 112 can process the vision data and / or other sensor data, each utilizing a corresponding one of the machine learning models 165, to generate one or more predicted object manipulation parameters for manipulating the object captured by the vision data. For example, one of the prediction engines 112 can process vision data from the additional vision components 194 using a corresponding one of the machine learning models 165 to generate a predicted gripping pose for gripping the spatula 192C. Also, for example, one of the prediction engines 112 can additionally or alternatively process vision data from the additional vision components 194 using a corresponding one of the machine learning models 165 to generate a predicted placement pose for placing the spatula 192C. Also, for example, one of the prediction engines 112 may additionally or alternatively process vision data from additional vision components 194 using a corresponding one of the machine learning models 165 to generate predicted waypoints to encounter in traversing to a gripping pose for the spatula. As described herein, which prediction engines 112 and corresponding machine learning models 165 (if any) are online and used by the system 110 may change over time and may depend on sufficient training and / or validation of the machine learning models (e.g., by the training engine 145).

[0041] The predicted object manipulation parameters (if any) generated by the prediction engine 112 for a given object manipulation can be automatically used as manipulation parameters by the manipulation parameter engine 116, can be first presented for confirmation by the visual representation engine 114 before utilization, or can be discarded and not utilized. For example, one of the prediction engines 112 may generate predicted object manipulation parameters and a measure of confidence for the predicted object manipulation parameters. If the measure of confidence meets a first threshold, the prediction engine can specify that the predicted object manipulation parameters should be utilized by the manipulation parameter engine 116 without prompting for confirmation. If the measure of confidence fails to meet the first threshold, but meets a second threshold, the prediction engine can specify that an indication of the predicted object manipulation parameters should be included in the visual representation by the visual representation engine 114--and utilized only if a confirmatory user interface input directed to the indication is received. If the reliability measure fails to meet the first and second thresholds, the prediction engine can specify that the predicted object manipulation parameters are not utilized and prompt the visual representation engine 114 to define corresponding object manipulation parameters.

[0042] The visual representation engine 114 receives vision data from the vision components 194, 174A, and / or 174B and generates a visual representation that it transmits to the remote client device 130 for rendering by the display engine 132 of the remote client device 130. Transmission to the remote client device 130 may be via one or more networks (not shown), such as the Internet or other wide area network (WAN).

[0043] The visual representation generated by the visual representation engine 114 includes an object representation of at least one object captured by the vision data. For example, the visual representation may include an object representation of the spatula 192 captured in the vision data from the additional vision component 194. For example, the visual representation may include an object representation that is a two-dimensional (2D) image of the spatula 192. Examples of the 2D image of the spatula 192 are shown in FIG. 2D and FIG. 2E, which are described in more detail below. Also, for example, the visual representation may include an object representation that is a three-dimensional (3D) representation of the spatula 192. For example, the 3D representation of the spatula 192 may define the positions (e.g., x, y, z positions) of one or more points on the surface of the spatula, and may optionally include one or more color values ​​for each of the positions. Examples of the 3D representation of the spatula 192 are shown in FIG. 2A, FIG. 2B, and FIG. 2C, which are described in more detail below. Also, the visual representation may optionally include an indication of predicted object manipulation parameters (if any) from the prediction engine 112. 2E, which is described in more detail below. The visual representation may also optionally include an environmental representation of other environmental objects (e.g., a work surface, a container in which the at least one object will be placed), and / or a robotic representation of all or a portion of the robot.

[0044] In some implementations, the visual representation generated by the visual representation engine 114 and transmitted to the remote client device 130 includes an object representation of the object and, optionally, one or more object representations of other nearby dynamic objects, but omits other portions that are static. In some of those implementations, only the object and, optionally, nearby dynamic objects are rendered at the remote client device 130. In some implementations, all or a portion of the robot and / or robot workspace are also rendered at the remote client device 130, even though they are not captured in the vision data transmitted to the remote client device 130. For example, the display engine 132 of the remote client device may include a robot simulator. The robot simulator may simulate all or a portion of the robot and / or all or a portion of the robot workspace and may render a simulation of the object along with a simulation of the robot and / or a simulation of the robot workspace. A robot simulator may be used to simulate an environment including corresponding objects, to simulate all or a portion of a robot (e.g., at least an end effector of the robot) operating within the simulated environment, and, optionally, to simulate interactions between the simulated robot and objects of the simulated environment in response to actions of the simulated robot. Various simulators may be utilized, such as physics engines that simulate collision detection, soft body and rigid body dynamics, etc. A non-limiting example of such a simulator is the BULLET physics engine.

[0045] As one particular example, the display engine 132 of the client device can receive a visual representation that includes only a 3D object representation of the object to be manipulated. The display engine 132 can position the 3D object representation in the simulated robot workspace and / or relative to the simulated robot. For example, the robot simulator of the display engine 132 can be preloaded with a visual representation of the robot workspace and / or the robot and can position the 3D object representation relative to those objects. When the object representation is based on vision data from the additional vision component 194, a pose of the object for the simulation of the robot and / or the simulation of the robot workspace can optionally be determined using a transformation between the pose of the additional vision component 194 and a pose of a corresponding one of the robot vision components 174A, 174B. The simulated robot can be set to a default state (e.g., a start state) or, optionally, the current state of the robot (e.g., the current position of the joints) can be given a visual representation for rendering the simulated robot in the current state. Implementations that simulate a robot and / or a robot workspace allow visual representations of smaller data size to be transmitted from the system 110 to the remote client device 130.

[0046] In some implementations, the visual representation engine 114 generates an object representation that renders the object with less precision than the complete representation in the visual representation rendered at the client device, but is a smaller data size than the complete representation. For example, the visual representation engine 114 can generate an object representation that includes one or more bounding boxes and / or other bounding shapes that approximate the surface of the object. For example, the visual representation engine 114 can generate an object representation that is composed of multiple connected bounding boxes, each of which may be defined by a center point, a height dimension, and a width dimension. One non-limiting example of this is shown in FIG. 2F, which is described in more detail below. A less detailed object representation is more concise in data, thereby conserving network resources. Additionally, a less detailed object representation can reduce delays in rendering the visual representation at the remote device, and / or obscure or remove potentially sensitive data from the object, or obscure the object itself.

[0047] An operator of the remote client device 130 utilizes one or more input devices 136 of the remote client device 130 to interact with the visual representation provided by the display engine 132. The input devices 136 may include, for example, a mouse, a touch screen, a VR hand controller, and / or a VR glove. The input devices 136 may form an integral part of the remote client device (e.g., a touch screen) or may be peripheral devices coupled with the remote client device 130 using wired and / or wireless protocols.

[0048] The input engine 134 of the remote client device 130 processes user interface input provided via the input device 136 to generate data indicative (directly or indirectly) of one or more object manipulation parameters used in manipulating the object. For example, the object manipulation parameters indicated by the data generated by the input engine 134 of an instance of user interface input may include a grasp pose, a place pose, a sequence of waypoints encountered in traversing to the grasp pose, a sequence of waypoints encountered in traversing to the place pose (after grasping the object), a complete path or trajectory (i.e., a path with velocity, acceleration, jerk, and / or other parameters) in traversing to and / or from a manipulation posture (e.g., a grasp posture or other manipulation posture), and / or other object manipulation parameters. The user interface input of an instance is provided by an operator of the remote client device 130 with reference to a visual representation rendered by the display engine 132. For example, an instance of user interface input may indicate a complete trajectory utilized during assembly of a part utilizing multiple component parts.

[0049] The manipulation parameter engine 116 determines the manipulation parameters based on the data provided by the input engine 134. In some implementations, the data directly defines the object manipulation parameters, and the manipulation parameter engine 116 determines the object manipulation parameters by utilizing the object manipulation parameters defined by the data. In other implementations, the manipulation parameter engine 116 transforms and / or otherwise processes the data when determining the object manipulation parameters.

[0050] The manipulation parameter engine 116 sends the determined object manipulation parameters and / or commands generated based on the object manipulation parameters to the robot 170A or 170B. In some implementations, the manipulation parameter engine 116 sends the object manipulation parameters and / or high-level commands based on the object manipulation parameters. In those implementations, a control system of the corresponding robot converts the object manipulation parameters and / or high-level commands into corresponding low-level actions, such as control commands issued to actuators of the robot. In other implementations, the object manipulation parameters can themselves define low-level actions (e.g., when a complete trajectory is defined by user interface inputs) and / or low-level actions can be generated based on the object manipulation parameters, and the manipulation parameter engine 116 sends the low-level actions to the corresponding robot for control based on the low-level actions.

[0051] The training data engine 143 generates training instances and stores the training instances in the training data database 152. Each of the training instances is generated for a corresponding one of the machine learning models 165 and is generated based on corresponding operational parameters of the instance, vision data and / or other data related to the instance, and optionally a measure of success for the instance (also referred to herein as a “success measure”).

[0052] As an example, the training data engine 143 can receive from the operation parameter engine 116 operation parameters utilized to control one of the robots 170A, 170B in performing the operation. The operation parameters can be generated based on user interface input from the remote client device 130, predicted by one of the prediction engines 112 and confirmed based on user interface input from the remote client device, or predicted by one of the prediction engines 112 and utilized automatically. The training data engine 143 can further receive vision data for the instance, such as vision data capturing an object manipulated in the operation. The vision data can be from an additional vision component 194 or from one of the robot vision components 174A or 174B. It is noted that in some implementations, the vision data utilized by the training data engine 143 in generating the training instance can be different than the vision data utilized in generating the object manipulation parameters. For example, object manipulation parameters may be defined based on user interface inputs targeted to object representations generated based on vision data from additional vision component 194, while vision data from robot vision component 174A (which captures the object) may be used in generating training instances.

[0053] The training data engine 143 can optionally further determine a measure of success of the manipulation (as a whole and / or of the portion targeted to the object manipulation parameter) based on the vision data and / or data from other sensors 104. The other sensors 104 can include, for example, weight sensors in the environment, non-vision sensors (e.g., torque sensors, position sensors) of the robot, and / or other sensors. The training data engine 143 can then generate training instances based on the vision data, the object manipulation parameters, and, optionally, the measure of success. For example, the training instances can include the vision data and the object manipulation parameters (e.g., representations thereof) as inputs of the training instances, and include the measure of success as an output of the training instances. As another example, the training instances can include the vision data as an input of the training instances, and the object manipulation parameters as an output of the training instances, and can be labeled as a positive or negative training instance based on the measure of success. As yet another example, the training instances can include the vision data as an input of the training instances, and include values ​​corresponding to the object manipulation parameters and determined based on the measure of success as an output of the training instances.

[0054] The training engine 145 utilizes the corresponding training instances of the training data database 152 to train the machine learning model 165. The trained machine learning model can then be at least selectively utilized by one of the prediction engines 112 in predicting one or more corresponding object manipulation parameters, which are then at least selectively utilized to control the robot. In some implementations, the trained machine learning model is only utilized in predicting the at least selectively utilized object manipulation parameters after the training engine 145 determines that one or more conditions are met. The one or more conditions may include at least a threshold amount of training and / or validation of the trained machine learning model, for example, as described herein. In some implementations, the training data engine 143 and the training engine 145 may implement one or more aspects of the method 400 of FIG. 4, described in detail herein.

[0055] Turning now to FIG. 1B, an example of how the components of FIG. 1A can interact with each other is shown, according to various implementations described herein. In FIG. 1B, vision data from an additional vision component 194 is provided to the prediction engine 112 and the visual representation engine 114. For example, the vision data may capture the spatula 192 shown in FIG. 1A. The prediction engine 112 can generate predicted object manipulation parameters 113 based on processing the vision data using one or more machine learning models 165. The visual representation engine 114 generates a visual representation 115 including at least an object representation of the object, the object representation being based on the object features of the vision data. In some implementations, the visual representation 115 may also include an indication of the predicted object manipulation parameters 113 (e.g., when a corresponding confidence measure indicates that confirmation is required). Additionally or alternatively, as indicated by the dashed arrow, the predicted object manipulation parameters 113 may be immediately provided to the manipulation parameter engine 116 without including an indication thereof in the visual representation 115 or seeking confirmation (e.g., when the corresponding confidence measure indicates that confirmation is not required).

[0056] The visual representation 115 is transmitted to a display engine 132, which optionally renders the visual representation along with other simulated representations (e.g., a simulated robot and / or a simulated workspace). Input data 135 is generated by an input engine 134 in response to one or more user interface inputs provided at one or more input devices 136 and targeted to the visual representation. The input data 135 directly or indirectly indicates one or more additional object manipulation parameters and / or confirmation of any predicted object manipulation parameters shown in the visual representation 115.

[0057] The manipulation parameter engine 116 utilizes the input data and, optionally, any directly provided predicted object manipulation parameters 113 to generate object manipulation parameters 117 that are provided to the robot 170A for implementation. For example, the robot 170A can generate control commands based on the object manipulation parameters 117 and implement those control commands in response to determining that an object has entered the robot workspace of the robot 170A and / or is in a particular pose within the robot workspace. For example, the robot 170A can make such a determination based on robot vision data from the robot vision component 174A.

[0058] The training data engine 143 can generate training instances 144 based on the implemented operational parameters 117. Each of the training instances 144 can include training instance inputs based on vision data from the additional vision component 194 and / or from the robot vision component 174. Each of the training instances 144 can be further based on a corresponding one of the operational parameters 117 (e.g., the input or output of the training instance can be based on the operational parameters). Each of the training instances 144 can be further based on corresponding success measures determined by the training data engine, based on vision data from the vision components 174A and / or 194, and / or based on data from other sensors 104. The training instances 144 are stored in the training data database 152 for use by the training engine 145 (FIG. 1) in training one or more of the machine learning models 165.

[0059] 2A, 2B, 2C, 2D, 2E, and 2F, each of which illustrates an example visual representation that may be rendered at remote client device 130 (FIG. 1A) or other remote client device. Each of the visual representations includes an object representation of an object to be manipulated by a robot and illustrates an example of user interface input that may be provided to define and / or confirm object manipulation parameters for the manipulation of the object by the robot.

[0060] 2A shows a visual representation including a simulated environment with a simulation 270A of one of the robots of FIG. 1A. Additionally, an object representation 292A of the spatula 192C of FIG. 1A is shown within the simulated environment. As described herein, the pose of the object representation 292A can be determined based on vision data utilized to capture the spatula 192C and generate the object representation 292A, optionally taking into account a transformation to the robot's reference coordinate system. The visual representation of FIG. 2A can be rendered, for example, by a VR headset.

[0061] An operator provided user interface input (e.g., via a VR controller) to define a path 289A1 of the robot's end effector from a start pose (not shown) to the illustrated grasping pose. The operator may, for example, actuate a first virtual button (e.g., virtual button 282A1) or hardware button to begin the definition of path 289A1 and actuate a second virtual or hardware button to define the end of path 289A1, which also constitutes the grasping pose. It is noted that, although not shown, the simulated robot 270A may "move" during the definition of trajectory 289A1 to provide the operator with visual feedback of path 289A1 as it is being performed by the robot 270A.

[0062] Further shown in FIG. 2A is a virtual button 282A2 that may be selected by the operator to use a predefined path that has been defined for a previous instance of user interface input and then “saved” by the operator. Selecting the virtual button 282A2 may paste the predefined path into the virtual environment along with an option for the user to modify the predefined path to adapt the predefined path to a particular object. Also shown in FIG. 2A is a virtual button 282A3 that may be selected by the operator to define the path 289A1 as a “predefined path” that may be selected later. By allowing the operator to save and reuse a particular path, the amount of user interface input required to redefine a path that is a slight modification of that path or a predefined path may be reduced. Furthermore, this may allow a path for the current instance to be defined more quickly, which may reduce idle time of the robot while waiting for definition of object manipulation parameters and / or may increase operator productivity.

[0063] 2B shows a visual representation including a simulated environment with a simulation 270B of one of the robots of FIG. 1A. Additionally, an object representation 292B of the spatula 192C of FIG. 1A is shown within the simulated environment. The operator has provided user interface input (e.g., via a VR controller) to define waypoints 289B1 and 289B2 (instead of a complete path) and a gripping posture 289B3 that will be encountered in traversing to the gripping posture 289B3 and that will be utilized in gripping the spatula 192C. The operator can, for example, actuate a first hardware button (e.g., of a VR controller) in a first manner to define the waypoints 289B1 and 289B2, and can actuate the first hardware button in a second manner (or actuate a second hardware button) to define the gripping posture 289B3. Although not shown, it is noted that the simulated robot 270B can "move" during definition of waypoints 289B1, 289B2 and / or gripping pose 289B3 to provide visual feedback to the operator. Although not shown in FIG. 2B, virtual buttons may also be provided for saving waypoints 289B1 and 289B2 and / or for reusing (and possibly adapting) predefined waypoints.

[0064] FIG. 2C shows a visual representation including a simulated environment with a simulation 270C of one of the robots of FIG. 1A. Additionally, an object representation 292C of the spatula 192C of FIG. 1A is shown in the simulated environment. The operator provided a user interface input (e.g., via a VR controller) to define the gripping posture 289C1 only. The operator can, for example, activate a first hardware button (e.g., of the VR controller) to define the gripping posture 289C1. Although not shown, it is noted that the simulated robot 270C can "move" during the definition of the gripping posture 289C1 to provide visual feedback to the operator. In some implementations, a visual representation akin to FIG. 2A and / or FIG. 2B can be provided at least until a machine learning model is trained that allows for the prediction of a path or waypoint that can be automatically performed selectively (without requiring confirmation), after which a visual representation akin to FIG. 2C can be provided for the definition of the gripping posture only via a user interface input. FIG. 2C may also provide a visual indication of the predicted path and / or predicted waypoints to prompt confirmation of the predicted waypoints or path, or redefinition of the predicted waypoints or path (if not confirmed).

[0065] FIG. 2D shows a visual representation including an object representation 292D of the spatula 192C of FIG. 1A, which is a 2D image (e.g., an RGB image) of the spatula. The visual representation may be rendered, for example, on a touch screen of a remote client device. An operator of the client device is prompted by an indication 282D to swipe on the touch screen to define an antipodal grasp. In response, the operator touches the touch screen at 289D1 and swipes to 289D2, at which point the operator ceases the touch. As a result, the antipodal grasp is defined by a first point at 289D1 and a second point at 289D2. Points 289D1 and 289D2 may be converted from 2D points to 3D points using, for example, a mapping between the 2D image and the corresponding 2.5D or 3D vision data.

[0066] FIG. 2E shows a visual representation including an object representation 292E of the spatula 192C of FIG. 1A, which is a 2D image (e.g., an RGB image) of the spatula. The visual representation also includes an indication 288E of the predicted antipodal grasp. The visual representation may be rendered, for example, on a screen of a remote client device. An operator of the client device is prompted by indication 282E1 to confirm the predicted antipodal grasp of indication 288E, or alternatively (by indication 282E2) to define an alternative grasp. If the operator agrees with the predicted antipodal grasp of indication 288E, he or she can simply click / tap on indication 282E1. If the operator does not agree with the predicted antipodal grasp of indication 288E, the operator can click / tap on indication 282E2 and modify indication 288E (e.g., drag indication 288E up / down, change its width, etc.) or define a new antipodal grasp from scratch.

[0067] FIG. 2F shows a visual representation including an object representation 292F of spatula 192C of FIG. 1A, which includes three connected bounding boxes (dashed lines) that approximate the surface of spatula 192A. As described herein, object representation 292F may be more data-efficient than the representations of FIGS. 2D and 2E and / or may prevent potentially sensitive data from being viewed by an operator of the client device. The visual representation may be rendered, for example, on a touchscreen of a remote client device. An operator of the client device is prompted by indication 282F to swipe on the touchscreen to define an antipodal grasp. In response, the operator touches the touchscreen at 289F1 and swipes to 289F2, at which point the operator ceases the touch. As a result, an antipodal grasp is defined by a first point at 289F1 and a second point at 289F2.

[0068] Various examples of visual representations and interactions with the visual representations are illustrated in Figures 2A-2F, however, it is understood that additional and / or alternative visual representations and / or interactions may be utilized in various implementations disclosed herein.

[0069] Turning now to FIG. 3, an exemplary method 300 is shown for causing a robot to manipulate an object according to object manipulation parameters determined based on data generated in response to a visual representation including an object representation of the object at a remote client device. For convenience, some of the operations of method 300 are described with reference to a system that performs the operations. The system may include various components of various computer systems and / or robots, such as one or more components illustrated in FIGS. 1A and 1B. Additionally, although the operations of method 300 are illustrated in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0070] At block 352, the system receives vision data from one or more vision components capturing object features of one or more objects. In some implementations or iterations of method 300, the vision components are robot vision components viewing a robot workspace of a corresponding robot, and the vision data captures the object features when the object is in the robot workspace. In some other implementations or iterations, the vision components are in a first region of the environment that is different from the robot workspace of the environment, and the vision data captures the object features when the object is in the first region--and prior to completion of transfer of the object to the robot workspace. In some of those implementations, one or more of blocks 354, 356, 358, 360, 362, and / or 364 may be completed prior to completion of transfer of the object to the robot workspace.

[0071] In optional block 354, the system generates one or more predicted object manipulation parameters based on the vision data and the predictive model. For example, the system can process the vision data and / or other sensor data using a corresponding predictive model trained and online to generate a predicted gripping pose and, optionally, a predicted probability (e.g., of successful object manipulation) for the predicted gripping pose. As another example, the system can additionally or alternatively process the vision data and / or other sensor data using a corresponding predictive model trained and online to generate a predicted classification of the grasped object and, optionally, a predicted probability of the predicted classification (e.g., the probability that the predicted classification is correct). The predicted classification can be used to determine a predicted placement location of the object (e.g., in a particular container of a plurality of available containers that corresponds to the predicted classification). The predicted probability for the predicted classification can optionally be utilized as a probability for a corresponding predicted placement location (e.g., a predicted placement location that has a defined relationship with the predicted classification).

[0072] In optional block 356, the system determines (a) whether additional object manipulation parameters are required in addition to the predicted manipulation parameters of block 354 to manipulate the object, and / or (b) whether one or more of the predicted object manipulation parameters need to be confirmed by a remote user interface input (e.g., because the corresponding predicted probability fails to meet a threshold).

[0073] If the determination in block 356 is "no", the system proceeds directly to block 360 and causes the robot to manipulate the object according to object-manipulating parameters that in such circumstances correspond to the predicted object-manipulating parameters of block 354.

[0074] If the determination at block 356 is “yes,” the system proceeds to optional block 358 or to block 360 .

[0075] Blocks 354 and 356 are shown as optional (as indicated by dashed lines) because they may not be utilized in method 300 in various implementations and / or they may be utilized in only some iterations in other implementations. For example, in some other implementations, block 354 may be performed only when at least one predictive model has been trained and brought online, which may be contingent on satisfying one or more conditions described herein.

[0076] At optional block 358, the system selects a remote client device from the plurality of client devices. The system may select a remote client device based on a variety of considerations. For example, the system may select a remote client device in response to determining that the remote client device does not currently have any requests for object manipulation parameters in its queue. Also, for example, the system may additionally or alternatively select a remote client device in response to determining that the amount of pending requests and / or the expected duration of the pending requests for the remote client device is less than that of other candidate remote client devices (e.g., remote client devices available for use in an environment in which a robot utilized in method 300 is deployed). As yet another example, the system may select a remote client device based on a measurement of proficiency of an operator of the remote client device. The proficiency measurements may be based on past success measurements for the operation based on object manipulation parameters determined based on user interface input from the operator, and may be an overall proficiency measurement or may be specific to one or more particular operations (e.g., a first proficiency measurement for a grasp and place operation, a second proficiency measurement for a grasp and join two objects, etc.).

[0077] At block 360, the system transmits to a remote client device (e.g., the remote client device selected at block 358) a visual representation based on the vision data of block 352. The visual representation includes an object representation based on at least the object features of the vision data of block 352. In some implementations, the object representation includes less data than the object features of the vision data of block 352. For example, the object representation may define bounding shapes that each approximate a corresponding region of the object without defining the color and / or other values ​​of the individual pixels or voxels enclosed by the bounding shapes in the vision data. For example, 64 pixel or voxel values ​​of the vision data may be replaced by seven values: three values ​​that define the x, y, z coordinates of the center of the bounding box, two values ​​that collectively define the orientation of the bounding box, and two values ​​that define the width and height of the bounding box.

[0078] In some implementations, the visual representation transmitted in block 360 does not have any representation of the robot and / or does not have any representation of one or more static objects and / or other objects in the robot workspace of the robot. In some of those implementations, the client device renders the transmitted visual representation along with a simulation of the robot and / or a simulation of all or a portion of the robot workspace. For example, the remote client device can run a robot simulator that simulates the robot and the robot workspace, and can render the object representations in the robot simulator along with the simulated robot and the robot workspace. It is noted that this can conserve network resources by eliminating the need to transmit a representation of the robot and / or the robot workspace each time a visual representation is transmitted to a remote client device. It is also noted that even when the vision data of block 352 is captured in a first region different from the robot workspace, the simulated robot and / or the simulated robot workspace can be rendered along with the object representations appropriately.

[0079] Optionally, block 360 includes a sub-block 360A in which the system generates a visual representation based on the vision data and based on the predicted operation parameters (if any) of block 354. For example, if a predicted gripping pose is generated in block 354, an indication of the predicted gripping pose may optionally be included in the visual representation. For example, the indication of the predicted gripping pose may be a representation of the robot's end effector rendered at the predicted gripping pose together with the object representation. An operator of the remote client device may confirm the predicted gripping pose or suggest an alternative gripping pose (e.g., by adjusting the representation of the robot's end effector). As another example, if a set of predicted waypoints is generated in block 354, an indication of those waypoints may optionally be included in the visual representation. For example, the indication of the waypoints may be a circle or other marking of the waypoints rendered together with the object representation and / or the simulation of the robot.

[0080] In block 362, the system receives data from the remote client device generated based on user interface input directed to the visual representation transmitted in block 360. The user interface input may include user interface input that defines (directly or indirectly) the object manipulation parameters and / or user interface input that confirms predicted object manipulation parameters.

[0081] At block 364, the system determines object manipulation parameters to use to manipulate the object by the robot based on the data received at block 362. The object manipulation parameters may include object manipulation parameters based on predicted object manipulation parameters (if any) shown in the visual representation if the data indicates confirmation of those predicted object manipulation parameters. The object manipulation parameters may additionally or alternatively include object manipulation parameters defined based on user interface inputs, independent of any predicted object manipulation parameters.

[0082] In some implementations, data generated at the remote client device may directly define and be utilized as the object manipulation parameters. In some other implementations, the data may indirectly define and be further processed in determining the object manipulation parameters. As one non-limiting example, block 364 may optionally include a sub-block 364A in which the system transforms poses and / or points to a robot coordinate system of the robot. For example, poses, points (e.g., waypoints), and / or other features defined by the data received in block 362 may be defined with respect to a given coordinate system that is different from the robot coordinate system and then transformed to the robot coordinate system. For example, the given coordinate system may be a first coordinate system of a vision component of block 352 that is different from a robot vision component of the robot.

[0083] At block 360, the system causes the robot to manipulate the object according to the object manipulation parameters. The object manipulation parameters may include object manipulation parameters based on predicted object manipulation parameters and / or object manipulation parameters defined based on user interface inputs independent of any predicted object manipulation parameters. In some implementations, the system provides the robot with the object manipulation parameters and / or high-level commands based on the object manipulation parameters. In those implementations, a control system of the robot converts the object manipulation parameters and / or high-level commands into corresponding low-level actions such as control commands issued to actuators of the robot. For example, the robot may include a controller that converts the high-level commands into more detailed control commands for application to one or more actuators of the robot. The control commands may include one or more velocity control commands issued to actuators of the robot at corresponding instances to control the movement of the robot. For example, in controlling the movement of the robot, velocity control commands may be issued to each of the actuators that control the movement of an end effector of the robot. In other implementations, the object manipulation parameters may themselves define low-level actions (e.g., when a complete trajectory is defined by user interface input) and / or low-level actions may be generated based on the object manipulation parameters, and the low-level actions may be provided to the robot for control based on the low-level actions.

[0084] In implementations where the vision component is in a first region of the environment that is different from the robot workspace of the environment, block 360 may include causing the robot to further manipulate the object in response to determining that the object is in the robot workspace. In some of these implementations, the robot may determine that the object is in the robot workspace based on robot vision data from the robot's vision component. In some additional or alternative implementations, the object may be determined to be in the workspace based on data from a conveyance for the object that indicates that the object is in the workspace. For example, when the conveyance includes a conveyor system, an arrival time of the object in the robot workspace may be determined based on operation data of the conveyor system.

[0085] After block 360, the system returns to block 352. It is noted that in various implementations, multiple iterations of method 300 may be performed in parallel for a given environment, allowing visual representations to be generated, transmitted, corresponding data to be received, and / or corresponding object manipulation parameters to be determined for new objects--before completion of method 300 for a previous object (e.g., before completion of at least block 360). For example, multiple iterations of method 300 may be performed in parallel, each for a different robot in the environment. Also, for example, multiple iterations of method 300 may be performed in parallel for a given robot, allowing object manipulation parameters to be determined for each of multiple different objects before those objects reach the robot workspace of the given robot and are manipulated by the given robot.

[0086] Turning now to FIG. 4, an example method 400 of generating training instances based on a robot's object manipulation attempts and using the training instances in training a predictive model is shown. For convenience, some of the operations of method 400 are described with reference to a system that performs the operations. The system may include various computer systems and / or various components of a robot, such as one or more components shown in FIGS. 1A and 1B. Furthermore, although the operations of method 400 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0087] In block 452, the system identifies (1) object manipulation parameters utilized in the robot's object manipulation attempt and (2) vision data associated with the object manipulation attempt. For example, the object manipulation parameters may include grasp and place poses defined based on user interface inputs directed to visual representations generated based on vision data from a first region, and the vision data may be robot vision data from a robot workspace different from the first region.

[0088] In optional block 454, the system generates a success measure for the object manipulation attempt based on the sensor data from the sensors. In some implementations, the system generates a single success measure for the entire object manipulation attempt. For example, for a pick-and-place operation, the system may determine a single success measure based on whether the object was successfully placed and / or the accuracy of the placement. In some other implementations, the system generates multiple success measures for the object manipulation attempt, each corresponding to a corresponding subset of the object manipulation parameters. For example, for a pick-and-place operation, the system may determine a first success measure for the pick operation (e.g., based on whether the object was successfully grasped) and a second success measure for the place operation (e.g., based on whether the object was successfully placed and / or the accuracy of the placement). Sensors on which the success indicators may be based may include, for example, position sensors on the robot, torque sensors on the robot, robot vision data from a vision component of the robot, weight sensors in the environment, and / or other robot and / or environmental sensors.

[0089] In block 456, the system generates training instances based on the object manipulation parameters, the vision data, and optionally, the success measurement. As indicated by the arrow from block 456 to block 452, the system can continue to perform iterations of blocks 452, 454, and 456 to generate additional training instances based on additional object manipulation attempts.

[0090] As an example of block 456, consider a pick-and-place operation with operation parameters of grip pose and place pose. A first training instance may be generated based on the vision data and grip pose and based on a success measure (e.g., a success measure of grip, or an overall success measure for pick-and-place). For example, the first training instance may be for a grip prediction model that approximates a value function, which is used to process the vision data and grip pose and predict the probability of successful gripping of an object using the grip pose, given the vision data. In such a case, the input of the training instance includes the vision data and grip pose (e.g., a representation of x, y, z position and orientation), and the output of the training instance includes the success measure (e.g., "0" if the success measure indicates an unsuccessful grip, and "1" if the success measure indicates a successful grip). Also, for example, the first training instance may instead be for a prediction model that processes the vision data (without also processing the grip pose) and generates a corresponding probability for each of the N grip poses. In such a case, the input of the training instance includes the vision data, and the output of the training instance includes a "1" for an output value corresponding to the grasp pose if the success measure indicated a successful grasp, and optionally a "0" for all other values. A second training instance may be generated based on the vision data and the place pose, and based on the success measure (e.g., a grasp success measure, or an overall success measure for pick-and-place). For example, the second training instance may be for a place prediction model that approximates a value function, and is used to process the vision data and the place pose and predict the probability of successful placement of the object when using a grasp pose taking into account the vision data.In such cases, the inputs of a training instance include vision data and a placement pose (e.g., a representation of x, y, z position and orientation), and the output of a training instance includes a success measure (e.g., "0" if the success measure indicates a failed placement, "1" if the success measure indicates a successful placement, or "0.7" if the success measure indicates a successful but not completely accurate placement).

[0091] As another example of block 456, consider an operation with operational parameters including a sequence of waypoints defined based on user interface input. Training instances may be generated based on the vision data and the sequence of waypoints. For example, the training instances may be for a waypoint prediction model that approximates a value function used to process the vision data and the sequence of waypoints and predict the probability of the sequence of waypoints (e.g., the probability that the sequence of waypoints is correct) given the vision data. In such a case, the input of the training instance includes the vision data and a representation of the sequence of waypoints (e.g., an embedding of the sequence generated using a recurrent neural network model or a transformer network), and the output of the training instance includes a "1" (or other "positive" value) based on the sequence being defined based on the user interface input.

[0092] In block 458, the system uses the generated training instances to update parameters of the predictive models. If different training instances for different predictive models were generated in block 456, the appropriate training instance for the corresponding predictive model may be utilized in each iteration of block 458. For example, some iterations of block 458 may use a first type of training instance to train a first predictive model, other iterations may use a second type of training instance to train a second predictive model, and so on. Furthermore, multiple iterations of blocks 458, 460, and 462 may optionally run in parallel, each dedicated to training a corresponding predictive model.

[0093] At block 460, the system determines whether further training is required. In some implementations, this may be based on whether a threshold amount of training has occurred, whether a threshold duration of training has occurred, and / or whether one or more performance characteristics of the predictive model have been observed (e.g., high probability of prediction and / or successful operation in at least X% of operations upon use of the predictive model). In some implementations, training of the predictive model may continue indefinitely, at least periodically.

[0094] If the determination at block 460 is "yes," the system waits for another training instance to become available at block 462 and returns to block 458 based on the available training instance. If the determination at block 460 is "no," the system proceeds to block 464 and ends the training of the predictive model (although training of other predictive models may continue). The trained predictive model may be utilized in method 300 or method 500 and may optionally continue to be trained during utilization.

[0095] Turning now to FIG. 5, an exemplary method 500 is shown that selectively utilizes a trained predictive model to determine object manipulation parameters for use by a robot in manipulating an object. Method 500 illustrates several implementations of method 300. For convenience, some of the operations of method 500 are described with reference to a system that performs the operations. The system may include various computer systems and / or various components of a robot, such as one or more components illustrated in FIGS. 1A and 1B. Additionally, while the operations of method 500 are illustrated in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0096] At block 552, the system receives vision data from one or more vision components capturing object features of one or more objects. In some implementations or iterations of method 500, the vision components are robot vision components viewing a robot workspace of a corresponding robot, and the vision data captures the object features when the object is in the robot workspace. In some other implementations or iterations, the vision components are in a first region of the environment that is different from the robot workspace of the environment, and the vision data captures the object features when the object is in the first region--and prior to completion of transfer of the object to the robot workspace. In some of those implementations, one or more of the blocks prior to block 572 may be completed prior to completion of transfer of the object to the robot workspace.

[0097] In block 554, the system selects one or more object manipulation parameters of the plurality of object manipulation parameters that need to be solved for manipulation of the object by the robot.

[0098] In block 556, the system determines whether a trained model for the object manipulation parameter has been brought online as described herein for the selected object manipulation parameter. If not, the system proceeds to block 558 and prompts the object manipulation parameter to be specified by a user interface input at a remote client device. For example, the system may generate a visual representation based on the vision data of block 552, transmit the visual representation to the client device, and render a prompt at the client device to define the object manipulation parameter via a user interface input targeted to the visual representation based on block 558. Any object manipulation parameter defined by the user interface input received in response to the prompt of block 558 may then be used as the selected object manipulation parameter in block 570.

[0099] If, in block 556 , the system determines that for the selected object manipulation parameter, a trained model for the object manipulation parameter has been brought online, the system proceeds to block 560 .

[0100] At block 560, the system generates predicted object manipulation parameters and corresponding reliability measures based on the vision data and the predictive models of block 552. For example, the system may select a predictive model that corresponds to the object manipulation parameters and process the vision data and / or other data using the predictive model to generate the predicted object manipulation parameters and corresponding reliability measures.

[0101] The system then proceeds to block 562 and determines whether the reliability measure of the predicted object manipulation parameters meets one or more thresholds (e.g., 90% or other thresholds). If not, the system proceeds to block 564 and prompts confirmation of the predicted object manipulation parameters at the remote client device and / or prompts corresponding object manipulation parameters to be specified by user interface input at the remote client device. For example, the system may generate a visual representation including an indication of one or more of the predicted object manipulation parameters, transmit the visual representation to the client device, and cause a prompt to be rendered at the client device based on block 564. The prompt may ask an operator of the client device to confirm the predicted object manipulation parameters by user interface input or to define corresponding alternative object manipulation parameters by user interface input. Also, for example, the system may additionally or alternatively prompt one or more of the object manipulation parameters to be specified by user interface input at the remote client device without presenting an option to confirm the predicted object manipulation parameters. In some implementations, if the reliability measure of a given predicted object manipulation parameter does not meet the threshold of block 562 but meets an additional lower threshold (e.g., 65% or other threshold), the system may trigger a prompt for confirmation of the given predicted object manipulation parameter. In those implementations, if the reliability measure of a given predicted object manipulation parameter does not meet the additional lower threshold, the system may optionally prompt the corresponding object manipulation parameter to be defined without providing any indication of the given predicted object manipulation parameter. Any object manipulation parameters defined by user interface input received in response to block 564 may then be used as all or a portion of the selected object manipulation parameters in block 570.

[0102] If the system determines in block 562 that the reliability measure meets the threshold, the system proceeds to block 566 and uses the predicted object manipulation parameters without prompting for confirmation of the predicted object manipulation parameters.

[0103] Then, in block 568, the system determines whether there are additional object manipulation parameters that need to be resolved for the robot to manipulate the object. If so, the system returns to block 554 to select additional object manipulation parameters. If not, the system proceeds to block 572. It is noted that in instances of method 500 where the determination at block 556 or block 562 is "no" for more than one iteration of block 556 or block 562, the prompting at the client device can be a single prompt requesting that the object manipulation parameters be defined and / or confirmed for all object manipulation parameters for which a "no" determination was made at block 556 or block 562. In other words, there are not necessarily N separate prompts for each of the N iterations. Rather, there may optionally be a single prompt that encompasses a request for each of the N iterations.

[0104] At block 572, the system causes the robot to manipulate the object according to the object manipulation parameters. The object manipulation parameters may include object manipulation parameters from one or more iterations of block 566 and / or from one or more iterations of block 570. For example, the object manipulation parameters may include object manipulation parameters based on predicted object manipulation parameters (with or without confirmation) and / or object manipulation parameters defined based on user interface inputs independent of any predicted object manipulation parameters.

[0105] The system then returns to block 552. It is noted that in various implementations, multiple iterations of method 500 may be running in parallel for a given environment, allowing visual representations for new objects to be generated and transmitted, corresponding data received, and / or corresponding object manipulation parameters determined--before completion of method 500 for a previous object (e.g., before completion of at least block 572).

[0106] Turning now to FIG. 6, an example method 600 of training a predictive model, validating a predictive model, deploying a predictive model, and optionally further training a deployed predictive model is illustrated. For convenience, some of the operations of method 600 are described with reference to a system that performs the operations. The system may include various computer systems and / or various robotic components, such as one or more components illustrated in FIGS. 1A and 1B. Additionally, while the operations of method 600 are illustrated in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0107] At block 652, the system trains the predictive model based on data from operator-guided object manipulation attempts. For example, the system may train the predictive model based on the training instances generated at blocks 452, 454, and 456 of method 400.

[0108] At block 654, the system determines whether one or more conditions have been met. If not, the system returns to block 652. If met, the system proceeds to block 656. The conditions considered at block 654 may include, for example, a threshold amount of training instances utilized in the training of block 652 and / or a threshold duration of the training of block 652.

[0109] In block 656, the system attempts to validate the predictive model based on comparing predictions generated using the predictive model with operator-guided ground truth. For example, the system can compare predicted object manipulation parameters made using the model with corresponding object manipulation parameters defined based on user interface inputs (i.e., operator-guided ground truth). The system can determine an error measure for the prediction based on the comparison. The operator-guided ground truth can optionally be confirmed based on the determined success measure. In other words, the operator-guided ground truth can be considered as ground truth only if the corresponding success measure indicates an overall success of the corresponding operation and / or a success of a portion of the operation corresponding to the defined object manipulation parameters.

[0110] In block 658, the system determines whether the validation was successful. If not, the system returns to block 652 and optionally adjusts the conditions of block 654 (e.g., to require a higher level of training). Various indicators may be utilized in determining whether the validation was successful. For example, the system may determine successful validation if at least a threshold percentage of the predictions are less than a threshold error measure based on the comparison of block 656.

[0111] If the determination at block 658 is that the validation is successful, the system proceeds to block 660. At block 660, the system deploys the predictive model for use in generated, suggested, and / or automatically implemented predictions. For example, the predictive model may be deployed for use in method 300 and / or method 500.

[0112] In optional block 662, the system further trains the predictive model based on operator feedback on the suggestions during deployment and / or based on sensor-based success measurements during deployment.

[0113] 7 illustrates a schematic of an example architecture of a robot 725. The robot 725 includes a robot control system 760, one or more motion components 740a-740n, and one or more sensors 742a-742m. The sensors 742a-742m may include, for example, vision components, optical sensors, pressure sensors, pressure wave sensors (e.g., microphones), proximity sensors, accelerometers, gyroscopes, thermometers, barometers, and the like. Although the sensors 742a-742m are shown as integral to the robot 725, this is not intended to be limiting. In some implementations, the sensors 742a-742m may be located outside the robot 725, for example, as stand-alone units.

[0114] The motion components 740a-740n may include, for example, one or more end effectors and / or one or more servo motors or other actuators for implementing movement of one or more components of the robot. For example, the robot 725 may have multiple degrees of freedom, and each of the actuators may control the actuation of the robot 725 in one or more of the degrees of freedom in response to a control command. As used herein, the term actuator encompasses a mechanical or electrical device (e.g., a motor) that produces movement, in addition to any driver that may be associated with the actuator and convert a received control command into one or more signals for driving the actuator. Thus, providing a control command to an actuator may include providing the control command to a driver, which converts the control command into an appropriate signal for driving an electrical or mechanical device to produce a desired movement.

[0115] The robot control system 760 may be implemented in one or more processors, such as a CPU, GPU, and / or other controller of the robot 725. In some embodiments, the robot 725 may include a "brain box" that may include all or some aspects of the control system 760. For example, the brain box may provide real-time bursts of data to the motion components 740a-740n, each of which includes, among other things, a set of one or more control commands that dictate movement parameters (if any) for each of one or more of the motion components 740a-740n. In some implementations, the robot control system 760 may perform one or more aspects of one or more methods described herein.

[0116] As described herein, in some implementations, all or some aspects of the control commands generated by control system 760 may be generated based on object manipulation parameters generated by techniques described herein. Although control system 760 is shown in FIG. 7 as an integral part of robot 725, in some implementations, all or some aspects of control system 760 may be implemented in components separate from but in communication with robot 725. For example, all or some aspects of control system 760 may be implemented in one or more computing devices in wired and / or wireless communication with robot 725, such as computing device 810.

[0117] 8 is a block diagram of an example computing device 810 that may optionally be utilized to perform one or more aspects of the techniques described herein. For example, in some implementations, the computing device 810 may be utilized to execute the simulator 120, the sim difference engine 130, the real episode system 110, the sim training data system 140, and / or the training engine 145. Generally, the computing device 810 includes at least one processor 814 that communicates with several peripheral devices via a bus subsystem 812. These peripheral devices may include, for example, a storage subsystem 824 including a memory subsystem 825 and a file storage subsystem 826, a user interface output device 820, a user interface input device 822, and a network interface subsystem 816. The input and output devices enable user interaction with the computing device 810. The network interface subsystem 816 provides an interface to an external network and is coupled to corresponding interface devices of other computing devices.

[0118] The user interface input devices 822 may include a keyboard, a pointing device such as a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touch screen integrated into a display, a voice input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 810 or a communications network.

[0119] The user interface output devices 820 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for generating a visible image. The display subsystem may also provide a non-visual display, such as through an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and ways of outputting information from the computing device 810 to a user or to another machine or computing device.

[0120] Storage subsystem 824 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 824 may include logic for performing selected aspects of one or more methods described herein.

[0121] These software modules are generally executed by the processor 814 alone or in combination with other processors. The memory 825 used in the storage subsystem 824 may include several memories including a main random access memory (RAM) 830 for storing instructions and data during execution of the program, and a read-only memory (ROM) 832 in which certain instructions are stored. The file storage subsystem 826 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular implementation may be stored by the file storage subsystem 826 in the storage subsystem 824 or other machines that can be accessed by the processor 814.

[0122] The bus subsystem 812 provides a mechanism for allowing the various components and subsystems of the computing device 810 to communicate with each other as intended. Although the bus subsystem 812 is shown diagrammatically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0123] The computing device 810 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 810 shown in Figure 8 is intended only as a specific example intended to illustrate some implementations. Many other configurations of the computing device 810 are possible having more or fewer components than the computing device shown in Figure 8.

[0124] In some implementations, a method is provided that includes receiving vision data from one or more vision components in a first region of an environment capturing features of the first region at a first time. The captured features include object features of an object in the first region at the first time. The method further includes transmitting a visual representation generated based on the vision data to a remote client device via one or more networks prior to completion of a transfer of the object from the first region to a robot workspace of a different environment not captured by the vision data, and receiving data generated based on one or more user interface inputs from the remote client device via the one or more networks. The visual representation includes the object representation generated based on the object features. The user interface inputs are made at the remote client device and are directed to the visual representation as rendered at the remote client device. The method further includes determining, based on the data, one or more object manipulation parameters for manipulation of the object by a robot operating within the robot workspace. The method further includes, in response to detecting that the object is within the robot workspace, causing the robot to manipulate the object according to the one or more object manipulation parameters. The object is within the robot workspace at a second time subsequent to the first time after transfer of the object from the first region to the robot workspace.

[0125] These and other implementations of the technology disclosed herein may include one or more of the following features.

[0126] In some implementations, determining one or more object manipulation parameters also occurs prior to completion of transfer of the object from the first region to the robot workspace.

[0127] In some implementations, the one or more object manipulation parameters include a grasp pose for grasping the object. In those implementations, causing the robot to manipulate the object in accordance with the one or more object manipulation parameters in response to detecting the object in the robot workspace after transfer includes causing an end effector of the robot to traverse to the grasp pose and attempt to grasp the object after traversing to the grasp pose.

[0128] In some implementations, the data defines one or more poses and / or one or more points relative to a first reference coordinate system. In some of those implementations, generating the one or more object manipulation parameters includes transforming the one or more poses and / or the one or more points to a robot coordinate system that is different from the reference coordinate system and using the transformed poses and / or points in generating the object manipulation parameters.

[0129] In some implementations, the method further includes, after having the robot manipulate the object, determining a measure of success of the manipulation based on additional sensor data from one or more additional sensors, generating a positive training instance based on the measure of success meeting a threshold, and training the machine learning model based on the positive training instance. In some versions of those implementations, the one or more additional sensors include a robot vision component, a torque sensor of the robot, and / or a weight sensor in the environment. In some additional or alternative versions of those implementations, generating a training instance input of the positive training instance based on the vision data or based on robot vision data from one or more robot vision components of the robot, and / or generating an output of the training instance of the positive training instance based on the object manipulation parameters. In some additional or alternative versions of those implementations, the method further includes, after training the machine learning model based on the positive training instance, using the machine learning model to process additional vision data incorporating the additional object, generating one or more predicted object manipulation parameters for the additional object based on the processing, and having the robot manipulate the additional object according to the one or more predicted object manipulation parameters. Additionally, the method may optionally further include transmitting a visual indication of the predicted object manipulation parameters to the remote client device or the additional remote client device, and receiving an indication from the remote client device or the additional remote client device that a positive user interface input has been received in response to presenting the visual indication of the predicted object manipulation parameters. Causing the robot to manipulate the additional object in accordance with the one or more predicted object manipulation parameters may be in response to receiving an indication that a positive user interface input has been received.Optionally, the method further includes generating a reliability measure of the one or more predicted object manipulation parameters based on the processing. The step of sending a visual indication of the predicted object manipulation parameters can be in response to the reliability measure failing to meet a reliability measure threshold. Additionally or alternatively, the method may optionally further include: after training the machine learning model based on the positive training instances, using the machine learning model to process additional vision data capturing the additional object, generating one or more predicted object manipulation parameters for the additional object based on the processing, sending the visual indication of the predicted object manipulation parameters to the remote client device or the additional remote client device, receiving from the remote client device or the additional remote client device an indication of alternative object manipulation parameters defined by user interface input received in response to presentation of the visual indication of the predicted object manipulation parameters, and in response to receiving the alternative object manipulation parameters, causing the robot to manipulate the additional object according to the one or more alternative object manipulation parameters. The method may optionally further include further training the machine learning model using training instances having labeled outputs based on the alternative object manipulation parameters.

[0130] In some implementations, the method further includes receiving, from one or more vision components in the first region before the robot manipulates the object, vision data capturing features of the first region at a third time after the first time but before the second time, the vision data including new object features of a new object in the first region at the third time, transmitting to the remote client device a new visual representation generated based on the new vision data, the new visual representation including the new object representation generated based on the new object features, receiving from the remote client device new data generated based on one or more new user interface inputs at the remote client device targeted to the new visual representation as rendered at the remote client device, and determining, based on the data, one or more new object manipulation parameters for manipulation of the new object by the robot operating within the robot workspace. In some of those implementations, the method further includes, after the robot manipulates the object, in response to the robot detecting, by the one or more robot vision components, that a new object is in the robot workspace, causing the robot to manipulate the new object according to the one or more new object manipulation parameters. The new object is within the robot workspace at a fourth time subsequent to the second time after transfer of the new object.

[0131] In some implementations, the transport of the object from the first region to the robot workspace is by one or more conveyors.

[0132] In some implementations, the method further includes accessing corresponding queue data defining a quantity and / or duration of outstanding robot manipulation assistant requests for each of the plurality of remote client devices. In some of those implementations, the method further includes selecting a remote client device from the plurality of remote client devices based on corresponding query data for the remote client device. Sending the visual representation to the remote client device can be in response to selecting the remote client device.

[0133] In some implementations, the object representation is a rendering of the object, where the rendering is generated based on the object features and omits one or more features of the object that are visible in the vision data.

[0134] In some implementations, detecting that an object is within the robot workspace is by the robot based on robot vision data from one or more robot vision components of the robot.

[0135] In some implementations, a method is provided that includes receiving vision data from one or more vision components in the environment, the vision data capturing features of the environment including object features of an object in the environment. The method further includes generating predicted object manipulation parameters for the object and a confidence measure for the predicted object manipulation parameters based on processing the vision data using a machine learning model. The method further includes determining whether the confidence measure for the predicted object manipulation parameters meets a confidence measure threshold. The method further includes transmitting, to a remote client device via one or more networks, in response to determining that the confidence measure fails to meet the confidence measure threshold, (1) an object representation of the object generated based on the object features, and (2) a visual indication of the predicted object manipulation parameters, and receiving, from the remote client device via the one or more networks, data generated based on one or more user interface inputs. The user interface inputs are made at the remote client device and are responsive to rendering the object representation and the visual indication at the remote client device. The method further includes determining to utilize either the object manipulation parameters or alternative object manipulation parameters based on the data. The method further includes causing the robot to manipulate the object according to the determined object manipulation parameters or alternative object manipulation parameters. In response to determining that the confidence measure satisfies a confidence measure threshold, the method further includes causing the robot to manipulate the object according to the object manipulation parameters without sending a visual indication to any remote client device for confirmation before manipulating the object according to the object manipulation parameters.

[0136] These and other implementations of the technology disclosed herein may include one or more of the following features.

[0137] In some implementations, the vision component is within a first region of the environment, and the step of determining whether the confidence measure for the predicted object manipulation parameters satisfies a confidence measure threshold occurs prior to transporting the object to a different robot workspace of the robot. [Explanation of symbols]

[0138] 101A Robot workspace 101B Robot workspace 101C area 103A Conveyor section 103B Conveyor section 103C Conveyor section 110 System, Real Episode System 112 Prediction Engine 113 Predicted Object Manipulation Parameters 114 Visual Expression Engine 115 Visual Representation 116 Operation Parameter Engine 117 Object Manipulation Parameters 120 Simulator 130 Remote Client Device, SIM Diff Engine 132 Display Engine 134 Input Engine 135 Input Data 136 Input Devices 140 SIM Training Data System 143 Training Data Engine 144 training instances 145 Training Engine 152 Training Data Database 165 Machine Learning Models 170A First Robot 170B Second Robot 172A End Effector 172B End Effector 174A Related Robot Vision Components 174B Related Robot Vision Components 194 Additional Vision Components 192A Sunglasses 192B Stapler 192C Spatula 193A Container 193B Container Simulation of the 270A robot 270B Robot Simulation 270C Robot Simulation 282A1 Virtual Button 282A2 Virtual Button 282A3 Virtual Button 282E1 Indication 282E2 Indication 288E Indication 289A1 Route, trajectory 289B1 Midway point 289B2 Midway point 289B3 Gripping posture 289C1 Gripping posture 289D1 point 289D2 points 289F1 point 289F2 points 292A Object representation 292B Object representation 292C Object representation 292D Object representation 292E Object representation 292F Object representation 300 ways 400 ways 500 ways 600 ways 725 Robot 740a-740n Operating components 742a~742m Sensor 760 Control System 810 Computing Devices 812 Bus Subsystem 814 Processor 816 Network Interface Subsystem 820 User Interface Output Device 822 User Interface Input Devices 824 Storage Subsystem 825 Memory Subsystem 826 File Storage Subsystem 830 Main Random Access Memory (RAM) 832 Read-Only Memory (ROM)

Claims

1. 1. A method comprising: receiving vision data from one or more vision components in an environment capturing characteristics of the environment including object characteristics of objects in the environment; based on processing the vision data using a machine learning model; predicted parameters for use in controlling a robot in the environment; a measure of confidence for the predicted parameters; and generating determining whether the reliability measure for the predicted parameters satisfies a reliability measure threshold; in response to determining that the reliability measure fails to meet the reliability measure threshold; transmitting, via one or more networks to a remote client device, an object representation of the object generated based on the object features and a visual indication of the predicted parameters; transmitting a visual indication of the predicted parameters to the remote client device is further responsive to determining that the reliability measure satisfies a lower reliability measure threshold indicating a lower reliability than the reliability measure threshold; receiving at the remote client device, in response to rendering at the remote client device, data generated based on one or more user interface inputs from the remote client device via the one or more networks; controlling said robot within said environment in dependence on said received data; in response to determining that the reliability measure satisfies the reliability measure threshold, causing the robot to be controlled within the environment in accordance with the predicted parameters without transmitting the visual indication to any remote client device for confirmation before the robot is controlled in accordance with the predicted parameters; A method comprising:

2. in response to determining that the reliability measure fails to meet a lower threshold reliability measure; transmitting the object representation of the object to the remote client device without transmitting any visual indication of the predicted parameters. The method of claim 1, further comprising:

3. the received data indicates confirmation of the predicted parameters; 2. The method of claim 1, wherein causing the robot to be controlled within the environment in dependence on the received data comprises, in response to the received data indicating the confirmation of the predicted parameters, causing the robot to be controlled within the environment in accordance with the predicted parameters.

4. a confirmation user interface element is rendered at the remote client device in conjunction with the rendering of the object representation and the visual indication at the remote client device; The method of claim 3 , wherein in response to a user interface input being directed to the confirmation user interface element at the remote client device, the received data indicates the confirmation of the predicted parameters.

5. the received data indicates alternative parameters; 2. The method of claim 1, wherein causing the robot to be controlled within the environment in dependence on the received data comprises, in response to the received data indicating the alternative parameter, causing the robot to be controlled within the environment in accordance with the alternative parameter.

6. rendering alternative parameter user interface elements at the remote client device in conjunction with the rendering of the object representation and the visual indication at the remote client device; The method of claim 5 , wherein the received data indicates the alternative parameter in response to a user interface input being directed to a user interface element of the alternative parameter at the remote client device.

7. generating positive training instances based on the alternative parameters indicated by the received data; training the machine learning model based on the positive training instances; 6. The method of claim 5, further comprising:

8. The method of claim 1 , wherein both a confirmation user interface element and an alternative parameter user interface element are rendered at the remote client device along with the rendering of the object representation and the visual indication at the remote client device.

9. 1. A method comprising: receiving vision data from one or more vision components in an environment capturing characteristics of the environment including object characteristics of objects in the environment; based on processing the vision data using a machine learning model; predicted parameters for use in controlling a robot in the environment; a measure of confidence for the predicted parameters; and generating determining whether the reliability measure for the predicted parameters satisfies a reliability measure threshold; generating a visual representation for transmission over one or more networks to a remote client device; including in the visual representation an object representation of the object that is generated based on the object features; determining whether to include a visual indication of the predicted parameters in the visual representation based on whether the confidence measure meets a lower confidence measure threshold; and transmitting the visual representation to the remote client device; receiving, in response to rendering the visual representation at the remote client device, data generated at the remote client device based on one or more user interface inputs from the remote client device over the one or more networks; controlling said robot within said environment in dependence on said received data; A method comprising:

10. 10. The method of claim 9, wherein the visual representation does not include a visual indication of any of the predicted parameters based on determining that the reliability measure fails to meet a lower threshold of the reliability measure.

11. the received data indicates confirmation of the predicted parameters; 10. The method of claim 9, wherein causing the robot to be controlled within the environment in dependence on the received data comprises, in response to the received data indicating the confirmation of the predicted parameters, causing the robot to be controlled within the environment in accordance with the predicted parameters.

12. a confirmation user interface element is rendered at the remote client device in conjunction with the rendering of the object representation and the visual indication at the remote client device; The method of claim 11 , wherein in response to a user interface input being directed to the confirmation user interface element at the remote client device, the received data indicates the confirmation of the predicted parameters.

13. the received data indicates alternative parameters; 10. The method of claim 9, wherein causing the robot to be controlled within the environment in dependence on the received data comprises, in response to the received data indicating the alternative parameter, causing the robot to be controlled within the environment in accordance with the alternative parameter.

14. rendering alternative parameter user interface elements at the remote client device in conjunction with the rendering of the object representation and the visual indication at the remote client device; 14. The method of claim 13, wherein the received data is indicative of the alternative parameter in response to a user interface input being directed to a user interface element of the alternative parameter at the remote client device.

15. generating positive training instances based on the alternative parameters indicated by the received data; training the machine learning model based on the positive training instances; 14. The method of claim 13, further comprising:

16. The method of claim 9 , wherein both a confirmation user interface element and an alternative parameter user interface element are rendered at the remote client device along with the rendering of the object representation and the visual indication at the remote client device.

17. 1. A system comprising: one or more vision components in the environment; A memory for recording instructions; One or more processors wherein the one or more processors: receiving vision data from the one or more vision components capturing characteristics of the environment including object characteristics of objects within the environment; based on processing the vision data using a machine learning model; predicted parameters for use in controlling a robot in the environment; a measure of confidence for the predicted parameters; and and determining whether the reliability measure for the predicted parameters satisfies a reliability measure threshold; in response to determining that the reliability measure fails to meet the reliability measure threshold; transmitting, via one or more networks to a remote client device, an object representation of the object generated based on the object features and a visual indication of the predicted parameters; transmitting a visual indication of the predicted parameters to the remote client device, further responsive to determining that the reliability measure satisfies a lower reliability measure threshold indicating a lower reliability than the reliability measure threshold; receiving at the remote client device, in response to rendering at the remote client device, data generated based on one or more user interface inputs from the remote client device via the one or more networks; controlling the robot within the environment in dependence on the received data; in response to determining that the reliability measure satisfies the reliability measure threshold, causing the robot to be controlled within the environment in accordance with the predicted parameters without transmitting the visual indication to any remote client device for confirmation before the robot is controlled in accordance with the predicted parameters; The system is operable to execute the instructions to:

18. one or more of the processors; in response to determining that the reliability measure fails to meet a lower threshold reliability measure; Transmitting the object representation of the object to the remote client device without transmitting any visual indication of the predicted parameters.

20. The system of claim 17, further operable to execute the instructions for:

Citation Information

Patent Citations

  • deep machine learning facility for robot gripping

    DE202017106506U1

  • Handling system and controller

    JP2018118343A

  • Connector posture recognition apparatus, terminal unit holding apparatus,connector posture recognition method, and terminal unit holding method

    JP2018180756A

  • Machine learning device, robot system and machine learning method

    JP2019048365A

  • Robot system and workpiece take-out method

    JP2019058960A