Robot demonstration learning skill templates
By using a demonstration-based robot learning approach to generate customized control strategies from local data, the high computational cost, sparse rewards, and vulnerability issues in traditional robot control technologies are resolved, enabling rapid and efficient robot task adaptation and control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-10
- Publication Date
- 2026-04-07
AI Technical Summary
Existing robot control technologies suffer from high computational costs, sparse rewards, and difficulty in scaling. Traditional reinforcement learning methods are fragile and cannot quickly adapt to changes in robot models.
A demonstration-based robot learning approach is adopted, which generates customized control strategies by collecting local demonstration data and combines visual, somatosensory, and tactile data to quickly adapt to specific robot models using skill templates.
It enables high-precision control for robots to quickly adapt to specific tasks, reduces training time, improves the robustness and adaptability of robot learning, reduces computational costs, and is applicable to a variety of robot models.
Smart Images

Figure CN116133800B_ABST
Abstract
Description
Background Technology
[0001] This manual relates to robots, and more particularly to planning robot movement.
[0002] Robot control refers to controlling the physical movement of a robot to perform a task. For example, an industrial robot used in car manufacturing can be programmed to first pick up a car part and then weld it onto the car's frame. Each of these actions can be comprised of dozens or hundreds of individual movements via robot motors and actuators.
[0003] Robot planning traditionally requires extensive manual programming to precisely determine how robot components should move to accomplish a specific task. Manual programming is tedious, time-consuming, and error-prone. Furthermore, a schedule manually generated for one work cell is often incompatible with other work cells. In this specification, a work cell is the physical environment in which a robot will operate. Work cells have specific physical properties, such as physical dimensions, that impose constraints on how the robot can move within the work cell. Therefore, a manually programmed schedule for one work cell may be incompatible with work cells that have different robots, different numbers of robots, or different physical dimensions.
[0004] Some research has been conducted to use machine learning control algorithms, such as reinforcement learning, to control robots to perform specific tasks. However, robots have many drawbacks that often make traditional learning methods unsatisfactory.
[0005] First, robots inherently possess highly complex, high-dimensional, and continuous action spaces. Therefore, generating and evaluating all possible candidate actions is computationally expensive. Second, robot control operates in an environment with extremely sparse rewards, as most possible actions do not lead to the completion of a specific task. A technique called reward shaping has been used to mitigate the sparse reward problem, but it is generally not scalable for hand-designed reward functions.
[0006] Another complicating factor is that traditional techniques for robot control using robot learning are very fragile. This means that even if a workable model is successfully trained, even very small changes to the task, the robot, or the environment can render the entire model completely unusable.
[0007] All of these problems mean that traditional methods of robot control using techniques such as reinforcement learning result in computationally expensive processes that are fundamentally difficult to implement, cannot be well scaled, and cannot be generalized to other situations. Summary of the Invention
[0008] This specification describes techniques related to demonstration-based robot learning. Specifically, it describes how a robot can execute a skill template with one or more sub-tasks, where these sub-tasks employ custom control strategies learned using demonstration data.
[0009] In this specification, a task refers to a specific robot capability involving the performance of one or more sub-tasks. For example, a connector insertion task is the capability that enables a robot to insert a wire connector into a socket. This task typically includes two sub-tasks: 1) moving the robot's tool to the location of the socket, and 2) inserting the connector into a specific position within the socket.
[0010] In this specification, a subtask is an operation that will be performed by the robot using a tool. In short, when the robot has only one tool, a subtask can be described as an operation that the robot as a whole will perform. Example subtasks include welding, dispensing, part positioning, and surface polishing, to name just a few. Subtasks are typically associated with the type of tool required to perform the subtask and the location within the coordinate system of the work cell where the subtask will be performed.
[0011] In this specification, a skill template, or simply a template, is a collection of data and software that allows a robot to be adapted to perform a specific task. Skill template data represents one or more subtasks required to perform a task, as well as information describing which subtasks of the skill require local demonstration learning and which perception flows will be needed to determine success or failure. Therefore, a skill template can define demonstration subtasks that require local demonstration learning, non-demonstration subtasks that do not require local demonstration learning, or both.
[0012] These technologies are particularly advantageous for robotic tasks that are traditionally difficult to control using machine learning (such as reinforcement learning). These tasks include those involving physical contact with objects in the workspace, such as grinding, joining and inserting tasks, and wiring, to name a few.
[0013] Specific embodiments of the subject matter described in this specification can be implemented to achieve one or more of the following advantages. The learning using demonstration data described in this specification addresses the problems of sparse rewards and lack of generalization in traditional reinforcement learning methods.
[0014] The system can perform tasks using visual, proprioceptive (joint) data, tactile data, and any other features, allowing it to adapt quickly and with high precision to specific robot models. The emphasis is on "sensor-rich robot manipulation," contrary to the classic view of minimal sensing in robotics. This typically means that cheaper robots can accomplish the same task in a shorter setup time.
[0015] The techniques described below allow machine learning techniques to be rapidly adapted to any suitable robot with an appropriately installed hardware abstraction. In a typical scenario, a single non-expert can train a robot to perform a skill template in less than a day of setup time. This is a significant improvement over traditional methods that might require expert teams spending weeks designing reward functions to solve the problem and weeks of training time in very large data centers. This effectively allows machine learning-based robot control to be widely distributed across many types of robots, even those the system has never seen before.
[0016] These technologies can effectively realize robot learning as a service, thereby increasing access to the technology. This, in turn, makes the entire robotics industry safer and more efficient.
[0017] The combination of reinforcement learning, perceptual data processing and machine learning, along with advanced impedance / admittance control, will enable robotic skills to be executed with a very high success rate (even in complex tasks), as required in industrial applications.
[0018] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, drawings, and claims. Attached Figure Description
[0019] Figure 1 This is a schematic diagram demonstrating an example learning system.
[0020] Figure 2A This is a diagram of an example system used to perform subtasks using a custom control strategy based on local demo data.
[0021] Figure 2B This is a diagram of another example system used to perform subtasks using local demo data.
[0022] Figure 2C This is a diagram of another example system used to perform subtasks using residual reinforcement learning.
[0023] Figure 3A This is a flowchart of an example process for combining sensor data from multiple different sensor streams.
[0024] Figure 3B This is a diagram of a camera wristband.
[0025] Figure 3C This is another example view of the camera wristband.
[0026] Figure 3D This is another example view of the camera wristband.
[0027] Figure 4 An example skill template is shown.
[0028] Figure 5 This is a flowchart of an example process for configuring a robot to perform skills using skill templates.
[0029] Figure 6A This is a flowchart illustrating an example process for using skill templates in tasks that utilize force as a guide.
[0030] Figure 6B This is a flowchart illustrating an example process for training skill templates using a cloud-based training system.
[0031] The same reference numerals and names in different figures denote the same elements. Detailed Implementation
[0032] Figure 1 This is a schematic diagram illustrating an example demonstration learning system. System 100 is an example of a system capable of implementing the demonstration-based learning techniques described in this specification.
[0033] System 100 includes multiple functional components, including an online execution system 110, a training system 120, and a robot interface subsystem 160. Each of these components can be implemented as a computer program installed on one or more computers in one or more locations, which are coupled to each other via any suitable communication network (e.g., an intranet or the Internet or a combination of networks).
[0034] System 100 controls robots 170a-n in two basic modes: demonstration mode and execution mode.
[0035] In demonstration mode, the user can control one or more robots 170a-n to perform specific tasks or subtasks. While doing so, the online execution system 110 collects status messages 135 and online observations 145 to generate local demonstration data. The demonstration data collector 150 is a module that can generate local demonstration data 115 from the status messages 135 and online observations 145, which the online execution system 110 can then provide to the training system 120. The training system can then generate a customized control strategy 125 specific to both the task and the robot performing the task, tailored to its specific characteristics.
[0036] In this specification, a control strategy is a module or subsystem that generates one or more next actions for a robot to perform in response to a given observation input. The output of the control strategy can be a command directly output by the strategy, or a higher-level command, each consumed by multiple robot components through a mechanism of the robot control stack, to influence the movement of one or more robot components (e.g., motors or actuators). Therefore, a control strategy can include one or more machine learning models that translate environmental observations into one or more actions.
[0037] In this specification, local demonstration data is data collected when the user controls the robot to demonstrate how the robot can perform specific tasks by causing it to perform physical movements. Local demonstration data may include kinematic data, such as joint positions, orientations, and angles. Local demonstration data may also include sensor data, such as data collected from one or more sensors. Sensors may include force sensors; vision sensors, such as cameras, depth cameras, and LiDAR; electrical connection sensors; accelerometers; audio sensors; gyroscopes; contact sensors; radar sensors; and proximity sensors, such as infrared proximity sensors, capacitive proximity sensors, or inductive proximity sensors, to name just a few.
[0038] Typically, local demo data is obtained from one or more bots that are very close to the user controlling the bot in demo mode. However, close physical proximity between the user and the bot is not a requirement for obtaining local demo data. For example, a user can obtain local demo data remotely from a specific bot via a remote user interface.
[0039] Training system 120 is a computer system that can generate customized control policies 125 from local demonstration data 115 using machine learning techniques. Training system 120 typically has significantly more computing resources than online execution system 110. For example, training system 120 could be a cloud-based computing system with hundreds or thousands of computing nodes.
[0040] To generate a custom control strategy 125, the training system 120 can first obtain or pre-generate a basic control strategy for the task. The basic control strategy is one that is expected to work well enough for a particular task so that any sufficiently similar robot is relatively close to being able to perform that task. For the vast majority of tasks, a single basic control strategy is not expected to be accurate enough to reliably complete the task successfully. For example, connection and insertion tasks typically require sub-millimeter accuracy, an accuracy that is unattainable without the detail provided by local demonstration data for a specific robot.
[0041] The basic control policy for a specific task can be generated in several ways. For example, the basic control policy can be programmed manually, trained using traditional reinforcement learning techniques, or trained using the demonstration-based learning techniques described in this specification. All of these techniques are suitable for pre-generating the basic control policy before receiving local demonstration data for the task, since time is a less important factor when generating the basic control policy.
[0042] In some implementations, the training system generates basic control policies from generalized training data 165. While the local demonstration data 115 collected by the online execution system 110 is typically specific to a particular robot or a particular robot model, the generalized training data 165 can, conversely, be generated from one or more other robots that do not need to be the same model, located in the same field, or manufactured by the same manufacturer. For example, the generalized training data 165 can be generated off-site from dozens, hundreds, or thousands of different robots with different characteristics and of different models. Furthermore, the generalized training data 165 does not even need to be generated from physical robots. For example, the generalized training data can include data generated from simulations of physical robots.
[0043] Therefore, local demo data 115 is local in a sense; it is specific to the particular robot that the user can access and manipulate. Thus, local demo data 115 represents data specific to a particular robot, but it can also represent local variables, such as specific characteristics of a particular task or a particular working environment.
[0044] System demonstration data collected during the development of skill templates can also be used to define basic control strategies. For example, the engineering team associated with the entity generating the skill templates can use one or more robots to perform the demonstrations at a facility located away from and / or not associated with system 100. The robots used to generate the system demonstration data also do not need to be the same robots or the same robot models as robots 170a-n in work cell 170. In this case, the system demonstration data can be used to guide the actions of the basic control strategy. The basic control strategy can then be adapted into a customized control strategy using computationally more expensive and sophisticated learning methods.
[0045] Using local demonstration data to adapt a basic control policy has the highly desirable effect of being relatively fast compared to generating a basic control policy, for example, by collecting system demonstration data or by training using generalized training data 165. For example, the size of generalized training data 165 for a specific task is often several orders of magnitude larger than that of local demonstration data 115, so training the basic control policy is expected to take much longer than adapting it for a specific robot. For instance, training the basic control policy may require substantial computing resources; in some cases, data centers with hundreds or thousands of machines may operate for days or weeks to train the basic control policy from generalized training data. In contrast, adapting the basic control policy using local demonstration data 115 may only take a few hours.
[0046] Similarly, collecting system demonstration data to define a basic control policy may require more iterations than local demonstration data. For example, to define a basic control policy, an engineering team might demonstrate 1,000 successful tasks and 1,000 unsuccessful tasks. In contrast, a fully adapted basic control policy might only require 50 successful demonstrations and 50 unsuccessful demonstrations.
[0047] Therefore, the training system 120 can use local demo data 115 to improve the basic control policy in order to generate a customized control policy 125 for the specific robot used to generate the demo data. The customized control policy 125 adjusts the basic control policy to take into account the characteristics of the specific robot and the local variables of the task. Training the customized control policy 125 using local demo data can take much less time than training the basic control policy. For example, while training the basic control policy may take many days or weeks, a user may only spend 1-2 hours generating local demo data 115 with the robot, which can then be uploaded to the training system 120. The training system 120 can then generate the customized control policy 125 in a much shorter time than it would take to train the basic control policy, for example, perhaps only one or two hours.
[0048] In execution mode, execution engine 130 can automatically execute tasks using a custom control strategy 125 without any user intervention. Online execution system 110 can use the custom control strategy 125 to generate commands 155 to be provided to robot interface subsystem 160, which drives one or more robots, such as robots 170a-n, in work unit 170. Online execution system 110 can consume status messages 135 generated by robots 170a-n and online observations 145 performed by one or more sensors 171a-n within work unit 170. Figure 1As shown, each sensor 171 is coupled to a corresponding robot 170. However, sensors do not need to correspond one-to-one with a robot, nor do they need to be coupled to a robot. In fact, each robot can have multiple sensors, and the sensors can be mounted on fixed or movable surfaces within the work cell 170.
[0049] The execution engine 130 can use status messages 135 and online observations 145 as inputs to the customized control strategy 125 received from the training system 120. Therefore, the robots 170a-n can react in real time to complete the task based on their specific characteristics and the specific characteristics of the task.
[0050] Therefore, using local demo data to fine-tune the control strategy results in a very different user experience. From the user's perspective, training the robot to perform tasks very precisely using a customized control strategy, including generating local demo data and waiting for the customized control strategy to be generated, is a very fast process, potentially requiring less than a day of setup time. This speed comes from utilizing a pre-calculated basic control strategy.
[0051] This arrangement introduces a significant technological improvement to existing robot learning methods, which typically require weeks of testing and generating hand-designed reward functions, weeks of generating appropriate training data, and even more weeks of training, testing, and improving the model to make them suitable for industrial production.
[0052] Furthermore, unlike traditional robot reinforcement learning, using local demo data is highly robust to small perturbations in the characteristics of the robot, task, and environment. If a company purchases a new robot model, the user only needs to spend a day generating new local demo data for the new customized control strategy. This contrasts with existing reinforcement learning methods, where any change to the physical characteristics of the robot, task, or environment can require an entire process that can take weeks to complete from scratch.
[0053] To initiate a demonstration-based learning process, the online execution system can receive a skill template 105 from the training system 120. As described above, the skill template 105 can specify a sequence of one or more subtasks required to execute the skill, which subtasks require local demonstration learning, which subtasks will require which perceptual flows, and specify transition conditions for when to move from executing one subtask of the skill template to the next.
[0054] As mentioned above, skill templates can define demonstration subtasks that require local demonstration learning, non-demonstration subtasks that do not require local demonstration learning, or both.
[0055] The demonstration subtasks are implicitly or explicitly bound to a basic control policy, which, as described above, can be pre-computed from generalized training data or system demonstration data. Therefore, for each demonstration subtask in the template, the skill template may include a separate basic control policy or an identifier for the basic control policy.
[0056] For each demonstration subtask, the skill template may also include software modules required to adapt the demonstration subtask using local demonstration data. Each demonstration subtask may rely on different types of machine learning models and may be adapted using different techniques. For example, a mobile demonstration subtask may heavily rely on camera images from the local work cell environment to locate a specific task target. Therefore, the adaptation process for a mobile demonstration subtask may heavily tune the machine learning model to identify features in the camera images captured in the local demonstration data. In contrast, an insertion demonstration subtask may heavily rely on force feedback data to sense the edges of a connector socket and insert the connector into the socket using appropriate, gentle force. Therefore, the adaptation process for an insertion demonstration subtask may heavily tune the machine learning model that processes force perception and the corresponding feedback. In other words, even if the basic models for the subtasks in the skill template are the same, each subtask may have its own corresponding adaptation process for incorporating local demonstration data in different ways.
[0057] Non-demonstration subtasks may or may not be associated with a basic control strategy. For example, a non-demonstration subtask may simply specify moving to a specific coordinate position. Alternatively, a non-demonstration subtask may be associated with a basic control strategy (e.g., calculated from another robot) that uses sensor data to specify how a joint should move to a specific coordinate position.
[0058] The purpose of skill templates is to provide a generalized framework for programming robots to have specific task capabilities. Specifically, skill templates can be used to adapt robots to perform similar tasks with relatively little effort. Therefore, adapting a skill template to a specific robot and a specific environment involves performing a training process for each demonstration subtask in the skill template. For simplicity, this process may be referred to as training the skill template, even though it may involve multiple separately trained models.
[0059] For example, a user can download a connector insertion skill template that specifies the execution of a first movement subtask, followed by a connector insertion subtask. The template can also specify that the first subtask relies on a visual perception stream, such as from a camera, while the second subtask relies on a force perception stream, such as from a force sensor. Furthermore, the template can specify that only the second subtask requires local demonstration learning. This could be because moving the robot to a specific location is typically not highly dependent on the task at hand or the work environment. However, if the work environment has tight space requirements, the template can also specify that the first subtask requires local demonstration learning so that the robot can quickly learn to navigate within the tight space constraints of the work environment.
[0060] To equip the robot with connector insertion skills, the user simply guides the robot to perform subtasks that the skill template indicates require local demo data. The robot will automatically capture the local demo data, which the training system can use to improve the basic control strategy associated with the connector insertion subtask. Once the customized control strategy training is complete, the robot only needs to download the final trained customized control strategy to be equipped to perform the subtask.
[0061] It's worth noting that the same skill template can be used for many different kinds of tasks. For example, the same connector insertion skill template can be used to equip a robot to perform HDMI cable insertion, USB cable insertion, or both. All the user needs to do is demonstrate these different insertion sub-tasks to refine the basic control strategy for the demonstration sub-task being learned. As mentioned above, this process typically requires far less computing power and far less time compared to developing or learning a complete control strategy from scratch.
[0062] Furthermore, the skill template approach can be hardware-agnostic. This means that skill templates can be used to equip robots to perform tasks even if the training system has never trained a control policy for that particular robot model. Therefore, this technique addresses many of the problems associated with using reinforcement learning to control robots. Specifically, it addresses the fragility problem, where even very small hardware changes require relearning the control policy from scratch, which is expensive and repetitive work.
[0063] To support the collection of local demonstration data, system 100 may also include one or more UI devices 180 and one or more demonstration devices 190. UI devices 180 can help guide users to obtain local demonstration data that will be most beneficial for generating customized control strategies 125. UI devices 180 may include a user interface that instructs users on which actions to perform or repeat, and an augmented reality device that allows users to control the robot without physically approaching it.
[0064] The demonstration device 190 is the primary operating device of the auxiliary system 100. Generally, the demonstration device 190 is a device that allows a user to demonstrate skills to the robot without introducing external force data into the local demonstration data. In other words, the demonstration device 190 can reduce the likelihood that the user's demonstration actions will affect what the influence sensors will actually read during execution.
[0065] In operation, the robot interface subsystem 160 and the online execution system 110 can operate according to different timing constraints. In some embodiments, the robot interface subsystem 160 is a real-time software control system with strict real-time requirements. A real-time software control system is a software system that requires execution within strict timing requirements to achieve normal operation. Timing requirements typically stipulate that specific actions must be performed or outputs must be generated within a specific time window to avoid the system entering a fault state. In a fault state, the system can pause execution or take other actions to interrupt normal operation.
[0066] On the other hand, the online execution system 110 typically offers greater operational flexibility. In other words, the online execution system 110 may, but is not required to, provide commands 155 within each real-time time window under which the robot interface subsystem 160 operates. However, to provide the ability to make sensor-based responses, the online execution system 110 can still operate under strict timing requirements. In a typical system, the real-time requirements of the robot interface subsystem 160 dictate that the robot provides commands every 5 milliseconds, while the online requirements of the online execution system 110 stipulate that the online execution system 110 should provide commands 155 to the robot interface subsystem 160 every 20 milliseconds. However, even if no such command is received within the online time window, the robot interface subsystem 160 is not necessarily required to enter a fault state.
[0067] Therefore, in this specification, the term "online" refers to both the time of operation and the rigid parameters. The time window is larger than the time window of the real-time robot interface subsystem 160 and generally offers greater flexibility when timing constraints are not met. In some embodiments, the robot interface subsystem 160 provides a hardware-agnostic interface, making commands 155 issued by the field execution engine 150 compatible with multiple different versions of the robot. During execution, the robot interface subsystem 160 can report status messages 135 back to the online execution system 110, allowing the online execution system 150 to adjust robot movement online, for example, due to local failures or other unforeseen circumstances. The robot can be a real-time robot, meaning that the robot is programmed to continuously execute its commands according to a highly constrained timeline. For example, each robot can expect commands from the robot interface subsystem 160 at a specific frequency (e.g., 100 Hz or 1 kHz). If the robot does not receive the expected command, it can enter a fault mode and cease operation.
[0068] Figure 2A This is a diagram of an example system 200 for performing subtasks using a customized control strategy based on local demonstration data. Generally, data from multiple sensors 260 is fed through multiple individually trained neural networks and combined into a single low-dimensional task state representation 205. The low-dimensional representation 205 is then used as input to an adjusted control strategy 210, which is configured to generate robot commands 235 to be executed by the robot 270. Therefore, system 200 can implement a customized control strategy based on local demonstration data by modifying the basic control strategy through subsystem 280.
[0069] Sensor 260 may include a perception sensor that generates a perception data stream representing the visual characteristics of a target within the robot or robotic work cell. For example, to achieve better visual capabilities, the robotic tool may be equipped with multiple cameras, such as visible light cameras, infrared cameras, and depth cameras, to name just a few.
[0070] Different sensing data streams 202 can be processed independently by corresponding convolutional neural networks 220a-n. Each sensing data stream 202 can correspond to a different sensing sensor, such as a different camera or a camera of a different type. Data from each camera can be processed by a different corresponding convolutional neural network.
[0071] Sensor 260 also includes one or more robot state sensors, wherein the one or more robot state sensors generate robot state data streams 204 representing physical characteristics of the robot or robot components. For example, robot state data stream 204 may represent forces, torques, angles, positions, velocities, and accelerations of the robot or corresponding components, to name just a few. Each robot state data stream 204 may be processed by a corresponding deep neural network 230a-m.
[0072] Modified subsystem 280 can have any number of neural network subsystems that process sensor data in parallel. In some implementations, the system includes only one sensing stream and one robot state data stream.
[0073] The output of the neural network subsystem is a corresponding portion of the task state representation 205, which cumulatively represents the state of the subtasks performed by the robot 270. In some implementations, the task state representation 205 is a low-dimensional representation with fewer than 100 features (e.g., 10, 30, or 50 features). Having a low-dimensional task state representation means fewer model parameters need to be learned, which further improves the speed at which local demonstration data can be used to adapt to specific subtasks.
[0074] The task status representation 205 is then used as input to the adjusted control strategy 210. During execution, the adjusted control strategy 210 generates robot commands 235 from the input task status representation 205, which are then executed by the robot 270.
[0075] During training, the training engine 240 generates parameter corrections 255 using the representation of the local demonstration action 275 and the suggested commands 245 generated by the adjusted control policy 210. The training engine can then use the parameter corrections 255 to improve the adjusted control policy 210, so that the commands generated by the adjusted control policy 210 in future iterations will more closely match the local demonstration action 275.
[0076] During training, the adjusted control policy 210 can be initialized using a base control policy associated with the demonstration subtask being trained. The adjusted control policy 210 can be iteratively updated using the local demonstration action 275. The training engine 240 can use any appropriate machine learning technique to adjust the adjusted control policy 210, such as supervised learning, regression, or reinforcement learning. When the adjusted control policy 210 is implemented using a neural network, parameter correction 235 can be backpropagated through the network, making the output suggested command 245 closer to the local demonstration action 275 in future iterations.
[0077] As described above, each subtask of the skill template can have a different training priority, even if their base model architectures are the same or similar. Therefore, in some implementations, the training engine 240 may optionally take subtask hyperparameters 275, specifying how the adjusted control policy 210 should be updated, as input. For example, the subtask hyperparameters may indicate that visual sensing is critical. Therefore, the training engine 240 can more aggressively correct the adjusted control policy 210 to align with the camera data captured using the local demonstration action 275. In some implementations, the subtask hyperparameters 275 identify separate training modules to be used for each different subtask.
[0078] Figure 2B This is a diagram of another example system for performing subtasks using local demonstration data. In this example, the system includes multiple independent control policies 210a-n, rather than just a single adjusted control policy. Each control policy 210a-n can use task state representation 205 to generate corresponding robot subcommands 234a-n. The system can then combine these subcommands to generate a single robot command 235 that will be executed by robot 270.
[0079] Having multiple individually tunable control strategies is advantageous in sensor-rich environments, such as those where data from multiple sensors with different update rates can be used. For example, different control strategies 210a-n can be executed at different update rates, allowing the system to combine both simple and more complex control algorithms within the same system. For instance, one control strategy can focus on robot commands using current force data, which can be updated at a much faster rate than image data. Simultaneously, another control strategy can focus on robot commands using current image data, which may require more complex image recognition algorithms, potentially with indeterminate runtime. The result is that the system can adapt quickly to both force and image data without slowing down its adaptation to force data. During training, the subtask hyperparameters can be used to identify a separate training process for each of the individually tunable control strategies 210-an.
[0080] Figure 2C This is a diagram of another example system for performing subtasks using residual reinforcement learning. In this example, instead of a single adjusted control policy that generates robot commands, the system uses a residual reinforcement learning subsystem 212 to generate corrective actions 225, which modify the basic actions 215 generated by the basic control policy 250.
[0081] In this example, the basic control strategy 250 takes sensor data 245 from one or more sensors 260 as input and generates a basic action 215. As described above, the output of the basic control strategy 250 can be one or more commands consumed by corresponding components of the robot 270.
[0082] During execution, the reinforcement learning subsystem 212 generates a corrective action 225 from the input task state representation 205, which will be combined with the basic action 215. The corrective action 225 is corrective because it modifies the basic action 215 from the basic control policy 250. The resulting robot command 235 can then be executed by the robot 270.
[0083] Traditional reinforcement learning processes employ two phases: (1) the action phase, in which the system generates new candidate actions, and (2) the training phase, in which the weights of the model are adjusted to maximize the cumulative reward for each candidate action. As described in the background section above, traditional approaches to reinforcement learning for robotics suffer from a severe sparse reward problem, meaning that actions randomly generated in the action phase are highly unlikely to yield any type of reward through the task's reward function.
[0084] Unlike traditional reinforcement learning, using local demo data provides all the information about which action to choose during the action phase. In other words, local demo data provides a set of actions, so these actions do not need to be randomly generated. This technique greatly constrains the problem space and allows the model to converge much faster.
[0085] During training, local demonstration data is used to drive robot 270. In other words, robot commands 235 generated from correction actions 225 and basic actions 215 are needed to drive robot 270. At each time step, reinforcement learning subsystem 210 receives representations of demonstration actions for physically moving robot 270. Reinforcement learning subsystem 210 also receives basic actions 215 generated by basic control policy 250.
[0086] The reinforcement learning subsystem 210 can then generate a reconstructed correction action by comparing the demonstration action with the basic action 215. The reinforcement learning subsystem 210 can also use a reward function to generate actual reward values for the reconstructed correction action.
[0087] The reinforcement learning subsystem 210 can also generate a predicted correction action generated from the current state of the reinforcement learning model, as well as a predicted reward value already generated using the predicted correction action. The predicted correction action is a correction action that the reinforcement learning subsystem 210 has already generated for the current task state representation 205.
[0088] The reinforcement learning subsystem 210 can then use the predicted correction action, the predicted reward value, the reconstructed correction action, and the actual reward value to compute weight updates for the augmented model. During iterations of the training data, the weight updates are used to adjust the predicted correction action toward the reconstructed correction action reflected by the demonstration action. The reinforcement learning subsystem 210 can compute weight updates based on any appropriate reward maximization process.
[0089] Figures 2A-2C The architecture shown provides the ability to combine multiple different models for sensor streams with varying update rates. Some real-time robots have very demanding control loop requirements, and therefore, they can be equipped with force and torque sensors that generate high-frequency updates (e.g., 100, 1000, or 10,000 Hz). In contrast, very few cameras or depth cameras operate above 60 Hz.
[0090] Figures 2A-2C The architecture shown has multiple parallel and independent sensor streams, and optionally, multiple different control strategies that allow for the combination of these different data rates.
[0091] Figure 3A This is a flowchart illustrating an example process for combining sensor data from multiple different sensor streams. This process can be performed by a computer system having one or more computers in one or more locations, for example... Figure 1 System 100. This process will be described as being performed by a system of one or more computers.
[0092] The system selects a basic update rate (302). The basic update rate will determine the rate at which the learning subsystem (e.g., the adjusted control strategy 210) will generate commands to drive the robot. In some implementations, the system selects the basic update rate based on the robot's minimum real-time update rate. Alternatively, the system may select the basic update rate based on the sensors that generate data at the fastest rate.
[0093] The system generates the corresponding portion of the task state representation at a corresponding update rate (304). Because the neural network subsystem can operate independently and in parallel, it can repeatedly generate the corresponding portion of the task state representation at a rate determined by the rate of its respective sensors.
[0094] To enhance the system's independence and parallelism, in some implementations, different portions of the system maintenance task state representation are written to multiple separate memory devices or memory partitions. This prevents different neural network subsystems from competing for memory access while generating their outputs at a high frequency.
[0095] The system repeatedly generates the task state representation at a basic update rate (306). During each time period defined by the basic update rate, the system can generate a new version of the task state representation by reading from the most recently updated sensor data output by multiple neural network subsystems. For example, the system can read from multiple separate memory devices or memory partitions to generate a complete task state representation. It is worth noting that this means that the data generated by some neural network subsystems is generated at a rate different from the rate at which it is consumed. For example, for sensors with a slower update rate, data can be consumed at a much faster rate than it is generated.
[0096] The system repeatedly uses the task state representation at a basic update rate to generate commands for the robot (308). By using independent and parallel neural network subsystems, the system can ensure that commands are generated at a sufficiently fast update rate to power even robots with hard real-time constraints.
[0097] This arrangement also means that the system can simultaneously feed multiple independent control algorithms with different update frequencies. For example, as mentioned above... Figure 2B Unlike a system that generates a single command, a system can include multiple independent control strategies, each generating a sub-command. The system can then generate a final command by combining these sub-commands into a final hybrid robot command, which represents the output of multiple different control algorithms.
[0098] For example, vision control algorithms can enable robots to move towards identified objects more quickly. Meanwhile, force control algorithms can enable robots to track along surfaces they have already contacted. Even though vision control algorithms typically update at a much slower rate than force control algorithms, the system can still function effectively. Figures 2A-2C The architecture shown powers both simultaneously at a basic update rate.
[0099] Figures 2A-2C The architecture shown offers numerous opportunities to expand the system's capabilities without requiring a major redesign. Multiple parallel and independent data streams allow for the implementation of machine learning functions, which is advantageous for local demonstration learning.
[0100] For example, integrating sensors that take into account local environmental data can be very advantageous in order to more thoroughly adapt robots to perform in a specific environment.
[0101] One example of using local environment data is considering the functionality of electrical connections. Electrical connections can be used as a reward factor for a variety of challenging robotic tasks involving establishing current between two components. These tasks include plugging cables into jacks, inserting power plugs into power outlets, and screwing in light bulbs, to name just a few.
[0102] To integrate the electrical connection into the modification subsystem 280, an electrical sensor, such as one of the sensors 260, can be configured in the working unit to detect when a current has been established. The output of the electrical sensor can then be processed by a separate neural network subsystem, and the result can be added to the task state representation 205. Alternatively, the output of the electrical sensor can be directly provided as input to a system or reinforcement learning subsystem that implements the adjusted control strategy.
[0103] Another example of using local environment data is the functionality that considers certain types of audio data. For instance, many connector insertion tasks produce very distinctive sounds upon successful completion. Therefore, the system can use a microphone to capture the audio, and its output can be added to an audio processing neural network in the task state representation. The system can then use a functionality that considers the specific acoustic characteristics of the connector insertion sound, forcing the learning subsystem to learn what a successful connector insertion sounds like.
[0104] Figure 3B This is a schematic diagram of a camera wristband. The camera wristband is an example of a wide range of instrument types that can be used to perform high-precision demonstration learning using the architecture described above. Figure 3B This is a perspective view of the tool at the end of the robot arm closest to the observer.
[0105] In this example, the camera wrist strap is mounted on the robot arm 335, just before the tool 345 located at the far end of the robot arm 335. The camera wrist strap is mounted to the robot arm 335 via a collar 345 and has four radially mounted cameras 310a-d.
[0106] The collar 345 can have any suitable protruding shape, wherein the protruding shape allows the collar 345 to be securely mounted to the end of the robotic arm. The collar 345 can be designed to be added to a robot manufactured by a third-party manufacturer. For example, a system distributing skill templates can also distribute camera wristbands to help non-professional users quickly converge the model. Alternatively or additionally, the collar 345 can be integrated into the robotic arm by the manufacturer during the manufacturing process.
[0107] The collar 345 can be elliptical, such as circular or oval, or rectangular. The collar 345 can be formed from a single solid volume, which is secured to the end of the robotic arm before the tool 345 is fastened. Alternatively, the collar 345 can be opened and securely closed by a fastening mechanism (e.g., a hook or pin). The collar 345 can be made of any suitable material that provides a secure connection to the robotic arm, such as rigid plastic; fiberglass; fabric; or metal, such as aluminum or steel.
[0108] Each camera 310a-d has a corresponding bracket 325a-d that secures the sensor, other electronics, and corresponding lenses 315a-d to a collar 345. The collar 345 may also include one or more lights 355a-b for illuminating the volume captured by the cameras 310a-d. Generally, the cameras 310a-d are arranged to capture different corresponding views of the working volume, either from the tool 345 or just outside the tool 345.
[0109] The example camera wristband has four radially mounted cameras, but any suitable number of cameras, such as 2, 5, or 10, can be used. As described above, modifying the architecture of subsystem 280 allows any number of sensor streams to be included in the task state representation. For example, the computer system associated with the robot can implement different corresponding convolutional neural networks to process sensor data generated by each of the cameras 310a-d in parallel. The processed camera outputs can then be combined to generate a task state representation, which, as described above, can be used to power multiple control algorithms running at different frequencies. As described above, the processed camera outputs can be combined with the outputs of other networks that independently process force sensors, torque sensors, position sensors, velocity sensors, or tactile sensors, or any suitable combination of these sensors.
[0110] Generally, using a camera wristband during demonstration learning leads to faster model convergence because the system will be able to recognize reward conditions for more positions and orientations. Therefore, using a camera wristband effectively further reduces the amount of training time required to adapt the basic control policy using local demonstration data.
[0111] Figure 3C This is another example view of the camera wristband. Figure 3C Other instruments that can be used to implement a camera wristband are shown, including cables 385a-d that can be used to feed the camera's output to a corresponding convolutional neural network. Figure 3C The diagram also illustrates how an additional depth camera 375 can be mounted onto the collar 345. As described above, the system architecture allows any other sensor to be integrated into the sensing system; therefore, for example, a separately trained convolutional neural network can process the output of the depth camera 375 to generate another part of the task state representation.
[0112] Figure 3D This is another example view of the camera wristband. Figure 3D This is a perspective view of the camera wrist strap of the 317a-d camera, which has a metal collar and four radially mounted cameras.
[0113] With these basic mechanisms for using local demo data to improve control strategies, users can write tasks to build hardware-agnostic skill templates that can be downloaded and used to quickly deploy tasks on many different types of robots and in many different types of environments.
[0114] Figure 4 Example skill template 400 is shown. Generally, a skill template defines a state machine for multiple subtasks required to perform a task. It is worth noting that skill templates are hierarchically combined, meaning that each subtask can be an independent task or another skill template.
[0115] Each subtask of a skill template has a subtask ID and includes subtask metadata, including whether the subtask is a demonstration subtask or a non-demonstration subtask, or whether the subtask references another skill template that should be trained separately. The subtask metadata can also indicate which sensor streams will be used to perform the subtask. Subtasks serving as demonstration subtasks will additionally include a base policy ID, which identifies the base policy that will be combined with corrective actions learned from local demonstration data. Each demonstration subtask will also be explicitly or implicitly associated with one or more software modules that control the training process for that subtask.
[0116] Each subtask in a skill template also has one or more transition conditions, which specify the conditions under which a transition to another task within the skill template should occur. Transition conditions can also be referred to as subtask objectives of the subtask.
[0117] Figure 4 The example illustrates a skill template for performing a task notoriously difficult to achieve using traditional robot learning techniques. The task is a grasping and connecting insertion task, requiring the robot to locate a wire within a work cell and insert a connector at one end of the wire into a socket also within the work cell. This problem is difficult to generalize using traditional reinforcement learning techniques because wires come in many different textures, diameters, and colors. Furthermore, traditional reinforcement learning techniques cannot inform the robot what to do next or how to make progress if the grasping subtask of the skill fails.
[0118] Skill Template 400 includes four sub-tasks, in Figure 4 In this context, nodes are represented as the graphical representations of the state machine. In fact, Figure 4 All information can be represented in any suitable format, such as a plain text configuration file or a record in a relational database. Alternatively or additionally, the user interface device can generate a graphical skill template editor, which allows the user to define skill templates through a graphical user interface.
[0119] The first subtask in skill template 400 is the movement subtask 410. Movement subtask 410 is designed to locate a wire within a work cell, which requires moving a robot from an initial position to the desired location of the wire, such as the position placed by a previous robot in the assembly line. Moving from one location to the next typically doesn't rely heavily on the robot's local characteristics, therefore the metadata for movement subtask 410 specifies that this subtask is a non-demonstration subtask. The metadata for movement subtask 410 also specifies that a camera stream is required to locate the wire.
[0120] The moving subtask 410 also specifies the "acquire wire vision" transition condition 405, which indicates when the robot should transition to the next subtask in the skill template.
[0121] The next subtask in skill template 400 is the grasping subtask 420. Grasping subtask 420 is designed to grasp a wire in a work cell. This subtask is highly dependent on the characteristics of the wire and the robot, especially the tool used to grasp the wire. Therefore, grasping subtask 420 is designated as a demonstration subtask that requires improvement using local demonstration data. Thus, grasping subtask 420 is also associated with a base policy ID, which typically identifies the previously generated basic control policy for grasping the wire.
[0122] The capture subtask 420 also specifies that executing the subtask requires both the camera stream and the force sensor stream.
[0123] The grasping subtask 420 also includes three transition conditions. The first transition condition, namely the "loss of wire vision" transition condition 415, is triggered when the robot loses visual contact with the wire. This could happen, for example, when the wire moves unexpectedly within the work cell, such as by a person or another robot. In this case, the robot transitions back to the moving subtask 410.
[0124] The second transition condition for grasping subtask 420, namely the "grabbing failure" transition condition 425, is triggered when the robot attempts to grasp the wire but fails. In this scenario, the robot can simply return and try grasping subtask 420 again.
[0125] The third transition condition for the grabbing subtask 420, namely the "grab successful" transition condition 435, is triggered when the robot attempts to grab the wire and succeeds.
[0126] The demonstration subtask can also indicate which transformation conditions require local demonstration data. For example, a specific subtask could indicate that all three transformation conditions require local demonstration data. Therefore, users can demonstrate how a grasping operation succeeds, how a grasping operation fails, and how the robot loses its vision on a power line.
[0127] The next subtask in skill template 400 is the second movement subtask 430. Movement subtask 430 is designed to move a grasped wire to a location within the work cell near a socket. In many cases where the user expects the robot to perform connections and insertions, the socket is located in a highly constrained space, such as inside an assembled dishwasher, television, or microwave oven. Because movement in a highly constrained space is highly dependent on the subtask and the work cell, the second movement subtask 430 is designated as a demonstration subtask, even though it only involves moving from one location within the work cell to another. Therefore, a movement subtask can be either a demonstration subtask or a non-demonstration subtask, depending on the skill requirements.
[0128] Although the second movement subtask 430 is designated as a demonstration subtask, it does not specify any base policy ID. This is because some subtasks are highly dependent on local working cells, and including a base policy would only hinder model convergence. For example, if the second movement subtask 430 needs to move the robot inside an appliance in a very specific orientation, a generalized base policy for movement would be unhelpful. Therefore, the user can perform an improvement process to generate local demonstration data that demonstrates how the robot should move through the working cells to obtain the specific orientation inside the appliance.
[0129] The second movement subtask 430 includes two transition conditions 445 and 485. The first transition condition 445, "acquire socket visual contact," is triggered when the camera stream makes visual contact with the socket.
[0130] If the robot happens to drop a wire while moving towards the socket, a second "dropped wire" transition condition 485 is triggered. In this case, skill template 400 specifies that the robot will need to return to movement subtask 1 in order to restart the skill. These kinds of transition conditions in the skill template provide the robot with a level of built-in robustness and dynamic response that traditional reinforcement learning techniques cannot easily provide.
[0131] The final subtask in skill template 400 is insertion subtask 440. Insertion subtask 440 is designed to insert the connector of a grabbed wire into a socket. Insertion subtask 440 is highly dependent on the type of wire and the type of socket; therefore, skill template 400 indicates that insertion subtask 440 is a demonstration subtask associated with the basic strategy ID typically involved in insertion subtasks. Insertion subtask 440 also indicates that this subtask requires both camera stream and force sensor stream.
[0132] Insertion subtask 440 includes three transition conditions. When insertion fails for any reason, the first "Insertation Failure" transition condition 465 is triggered, specifying a retry. When the socket happens to move out of the camera's line of sight, the second "Out of Socket Visibility" transition condition 455 is triggered, specifying that the wire within the height-constrained space be moved back to the socket's location. Finally, when the wire falls during the insertion task, the "Fallen Wire" transition condition 475 is triggered. In this case, skill template 400 specifies a return to the first movement subtask 410.
[0133] Figure 4 One of the main advantages of the skill templates shown is their composability by developers. This means that new skill templates can be composed of subtasks that have already been developed. The feature also includes hierarchical composition, meaning that each subtask within a specific skill template can reference another skill template.
[0134] For example, in an alternative implementation, the insertion subtask 440 may actually reference an insertion skill template that defines a state machine of multiple finely controlled movements. For example, the insertion skill template may include a first movement subtask aimed at aligning the connector and receptacle as precisely as possible, a second movement subtask aimed at achieving contact between one side of the connector and the receptacle, and a third movement subtask aimed at achieving a full connection by using one side of the receptacle as a force guide.
[0135] Furthermore, further skill templates can be layered from skill template 400. For example, skill template 400 could be a small part of a more complex set of subtasks required to assemble electronic appliances. The overall skill template could have multiple connector insertion subtasks, where each insertion subtask references a skill template, such as skill template 400, used to implement the subtask.
[0136] Figure 5 This is a flowchart illustrating an example process for configuring a robot to perform skills using skill templates. This process can be performed by a computer system having one or more computers in one or more locations, for example... Figure 1 System 100. This process will be described as being performed by a system of one or more computers.
[0137] The system receives a skill template (510). As described above, the skill template defines a state machine with multiple subtasks and the transition conditions that define when the robot should transition from performing one task to the next. Furthermore, the skill template can define which tasks are demonstration subtasks that require improvement using local demonstration data.
[0138] The system obtains the basic control strategy (520) for the demonstration subtask of the skill template. The basic control strategy can be a generalized control strategy generated from multiple different robot models.
[0139] The system receives local demonstration data (530) for the demonstration subtask. The user can use an input device or user interface to make the robot perform the demonstration subtask over multiple iterations. During this process, the system automatically generates local demonstration data for performing the subtask.
[0140] The system trains a machine learning model (540) for the demonstration subtask. As described above, the machine learning model can be configured to generate commands to be executed by the robot in response to one or more input sensor streams, and the machine learning model can be tuned using local demonstration data. In some implementations, the machine learning model is a residual reinforcement learning model that generates corrective actions that are combined with basic actions generated by the basic control policy.
[0141] The system executes a skill template (550) on the robot. After training all the demonstration subtasks, the system can use the skill template to enable the robot to fully perform the task. During this process, the robot will use improved demonstration subtasks, which are specifically tailored for the robot's hardware and operating environment using local demonstration data.
[0142] Figure 6A This is a flowchart illustrating an example process for using skill templates in tasks that utilize force as guidance. The skill template arrangement described above provides a relatively simple way to generate very complex tasks consisting of multiple highly complex subtasks. An example of such a task in a connector insertion task uses force data as guidance. This allows the robot to achieve a higher level of precision than it could otherwise achieve. The process can be performed by a computer system having one or more computers in one or more locations, for example... Figure 1 System 100. The process will be described as being executed by a system of one or more computers.
[0143] The system receives a skill template with transition conditions, wherein the transition conditions require the establishment of physical contact forces (602) between an object held by the robot and a surface in the robot's environment. As described above, the skill template can define a state machine with multiple tasks. The transition conditions can define the transition between the first and second subtasks of the state machine.
[0144] For example, the first subtask could be a movement subtask, and the second subtask could be an insertion subtask. The transition conditions can specify that the connector held by the robot and to be inserted into the socket needs to establish physical contact with the edge of the socket.
[0145] The system receives local demonstration data for the transition (604). In other words, the system can ask the user to demonstrate the transition between the first and second subtasks. The system can also ask the user to demonstrate a failure scenario. One such failure scenario could be the loss of physical contact with the edge of the socket. If such a scenario occurs, the skill template can specify a return to the first movement subtask of the template so that the robot can re-establish the physical contact force specified by the transition conditions.
[0146] The system uses local demo data to train the machine learning model (606). As described above, through training, the system learns to avoid actions that result in loss of physical contact force and learns to select actions that are likely to maintain physical contact force throughout the second task.
[0147] The system executes a trained skill template (608) on the robot. This enables the robot to automatically perform subtasks and transitions defined by the skill template. For example, for connection and insertion tasks, local demo data can make the robot highly adaptable to inserting a specific type of connector.
[0148] Figure 6B This is a flowchart illustrating an example process for training a skill template using a cloud-based training system. Generally, the system can generate all demo data locally and then upload the demo data to the cloud-based training system to train all demo subtasks of the skill template. The process can be performed by a computer system having one or more computers in one or more locations, such as... Figure 1 System 100. The process will be described as being executed by a system of one or more computers.
[0149] The system receives a skill template (610). For example, the online execution system can download the skill template from a cloud-based training system that will demonstrate the skill template in a training subtask or from another computer system.
[0150] The system identifies one or more demonstration subtasks (620) defined by the skill template. As described above, each subtask defined in the skill template can be associated with metadata indicating whether the subtask is a demonstration subtask or a non-demonstration subtask.
[0151] The system generates a corresponding local demonstration dataset (630) for each of one or more demonstration subtasks. As described above, the system can instantiate and deploy individual task systems, each generating local demonstration data as the user manipulates the robot to perform a subtask in a local work cell. The task state representation can be generated at the basic rate of the subtask, regardless of the update rate of the sensors contributing data to the task state representation. This provides a convenient way to store and organize local demonstration data, rather than generating many different sensor datasets that must all be coordinated in some way later.
[0152] The system uploads the local demonstration dataset to a cloud-based training system (640). Most facilities employing robots for real-world tasks lack on-site data centers suitable for training complex machine learning models. Therefore, while the local demonstration data can be collected on-site by the system, which is located in the same location as the robot performing the task, the actual model parameters can be generated by a cloud-based training system, which can only be accessed via the Internet or another computer network.
[0153] As mentioned above, the size of the local demo data is expected to be several orders of magnitude smaller than the data used to train the basic control policy. Therefore, although the local demo data may be large, the upload burden is manageable within a reasonable amount of time (e.g., a few minutes to an hour of upload time).
[0154] The cloud-based training system generates corresponding training model parameters (650) for each local demonstration dataset. As described above, the training system can train the learning system to generate robot commands, which may consist, for example, corrective actions that correct basic movements generated by a basic control strategy. As part of this process, the training system can obtain the corresponding basic control strategy for each demonstration subtask locally or from another computer system, which may be a third-party computer system that publishes task or skill templates.
[0155] Cloud-based training systems typically have significantly more computing power than online execution systems. Therefore, while training each demo subtask involves a substantial computational burden, these operations can be massively parallelized on a cloud-based training system. Consequently, in typical scenarios, training a skill template from local demo data on a cloud-based training system takes no more than a few hours.
[0156] The system receives training model parameters (660) generated by a cloud-based training system. The size of the training model parameters is typically much smaller than the size of the local demo data for a specific subtask; therefore, the time spent downloading the trained parameters after the model has been trained is negligible.
[0157] The system uses training model parameters generated by a cloud-based training system to execute skill templates (670). As part of this process, the system can also download basic control policies for demonstration sub-tasks, for example, from the training system, from their original source, or from another source. The training model parameters can then be used to generate commands for the robot to execute. In a reinforcement learning system used for the demonstration sub-tasks, the parameters can be used to generate corrective actions, which modify the basic actions generated by the basic control policies. The online execution system can then repeatedly issue the obtained robot commands to drive the robot to perform specific tasks.
[0158] Figure 6B The process described can be performed by a team of one person over a day to enable a robot to perform highly precise skills in a way that is appropriate for its environment. This is a huge improvement over traditional manual programming methods and even traditional reinforcement learning methods, which require teams of many engineers to spend weeks or months designing, testing, and training models that do not generalize well to other scenarios.
[0159] In this specification, a robot is a machine having a basic position, one or more movable components, and a kinematic model, wherein the kinematic model can be used to map a desired position, orientation, or both in a coordinate system (e.g., a Cartesian coordinate system) to commands for physically moving one or more movable components to the desired position or orientation. In this specification, a tool is a device that is part of a kinematic chain of one or more movable components of the robot and attached to the end of that kinematic chain. Example tools include grippers, welding equipment, and grinding equipment.
[0160] In this specification, a task is an operation performed by a tool. In short, when the robot has only one tool, a task can be described as an operation performed by the robot as a whole. Example tasks include welding, dispensing, part positioning, and surface polishing, to name just a few. Tasks are typically associated with the type of tool required to perform the task and its location within the work cell where it will be performed.
[0161] In this specification, motion planning is a data structure that provides information for performing an action, which can be a task, a task family, or a transition. Motion planning can be fully constrained, meaning that all values of all controllable degrees of freedom of the robot are explicitly or implicitly represented; or underconstrained, meaning that some values of the controllable degrees of freedom are unspecified. In some implementations, for the action corresponding to the motion plan to be actually performed, the motion plan must be fully constrained to include all necessary values of all controllable degrees of freedom of the robot. Therefore, at some points in the planning process described in this specification, some motion plans may be underconstrained, but by the time the motion plan is actually performed on the robot, it can be fully constrained. In some implementations, motion planning represents the edges in a task graph between two configuration states of a single robot. Therefore, typically each robot has one task graph.
[0162] In this specification, the motion scan volume is the spatial region occupied by at least a portion of the robot or tool throughout the entire execution of motion planning. The motion scan volume can be generated from the collision geometry associated with the robot tool system.
[0163] In this specification, a transformation is a motion plan describing a movement to be performed between a start point and an end point. The start and end points can be represented by pose, position in a coordinate system, or the task to be performed. A transformation can be under-constrained by lacking one or more values of one or more corresponding degrees of freedom (DOF) of the robot. Some transformations represent free motion. In this specification, free motion is a transformation without any constraints on its degrees of freedom. For example, robot motion that simply moves from pose A to pose B without any constraints on how it moves between these two poses is free motion. During the planning process, the DOF variables for free motion are eventually assigned values, and the path planner can use any appropriate value for the motion that does not conflict with the physical constraints of the work cell.
[0164] The robot functions described in this specification can be implemented using a hardware-agnostic software stack, or, for the sake of brevity, only a software stack that is at least partially hardware-agnostic. In other words, the software stack can accept commands generated by the planning process described above as input, without requiring the commands to specifically relate to a particular robot model or component. For example, the software stack can be at least partially composed of… Figure 1 The field execution engine 150 and robot interface subsystem 160 are implemented.
[0165] A software stack can include multiple levels that increase hardware specificity in one direction and software abstraction in another. The lowest level of the software stack is the robot components, which include devices that perform low-level actions and sensors that report low-level states. For example, a robot may include various low-level components, including motors, encoders, cameras, actuators, grippers, application-specific sensors, linear or rotary position sensors, and other peripherals. As an example, a motor may receive a command indicating the amount of torque that should be applied. In response to receiving this command, the motor may, for example, use an encoder to report the current position of the robot joints to higher levels of the software stack.
[0166] Each next highest level in the software stack can implement interfaces that support multiple different underlying implementations. Generally, each interface between levels provides status messages from lower to higher levels and commands from higher to lower levels.
[0167] Typically, commands and status messages are generated cyclically during each control cycle; for example, one status message and one command per control cycle. Lower levels of the software stack generally have more stringent real-time requirements than higher levels. For example, at the lowest level of the software stack, the control cycle may have actual real-time requirements. In this specification, real-time means that within a specific control cycle time, commands received at a level of the software stack must be executed, and optionally, status messages are provided back to higher levels of the software stack. If this real-time requirement is not met, the robot can be configured to enter a fault state, for example, by freezing all operations.
[0168] At the next highest level, the software stack can include software abstractions of specific components, which will be referred to as motor feedback controllers. A motor feedback controller can be a software abstraction of any suitable low-level component, not just a literal motor. Therefore, the motor feedback controller receives state from lower-level hardware components via an interface and, based on higher-level commands received from higher levels in the stack, sends commands back to lower-level hardware components via the same interface. The motor feedback controller can have any suitable control rules that determine how higher-level commands should be interpreted and translated into lower-level commands. For example, a motor feedback controller can use anything from simple logic rules to more advanced machine learning techniques to translate higher-level commands into lower-level commands. Similarly, a motor feedback controller can use any suitable fault rules to determine when a fault state is reached. For example, if the motor feedback controller receives a higher-level command within a specific part of the control loop but does not receive a lower-level state, it can cause the robot to enter a fault state that stops all operations.
[0169] At the next highest level, the software stack can include actuator feedback controllers. Actuator feedback controllers can include control logic for controlling multiple robot components via their respective motor feedback controllers. For example, some robot components, such as articulated arms, can actually be controlled by multiple motors. Therefore, actuator feedback controllers can provide a software abstraction of the articulated arm by sending commands to the motor feedback controllers of multiple motors using their control logic.
[0170] At the next highest level, the software stack can include joint feedback controllers. Joint feedback controllers can represent joints mapped to logical degrees of freedom in the robot. Thus, for example, while a robot's wrist might be controlled by a complex network of actuators, a joint feedback controller can abstract away this complexity and expose that degree of freedom as a single joint. Therefore, each joint feedback controller can control an arbitrarily complex network of actuator feedback controllers. For example, a six-DOF robot could be controlled by six different joint feedback controllers, each controlling an independent network of actual feedback controllers.
[0171] Each level of the software stack can also enforce level-specific constraints. For example, if a specific torque value received by the actuator feedback controller is outside the acceptable range, the actuator feedback controller can modify it to be within the range or enter a fault state.
[0172] To drive the inputs to the joint feedback controller, the software stack can use command vectors that include command parameters for each component at lower levels, such as the positive values, torque, and speed of each motor in the system. To expose the state from the joint feedback controller, the software stack can use state vectors that include state information for each component at lower levels, such as the position, speed, and torque of each motor in the system. In some implementations, the command vectors also include constraint information about the constraints to be enforced by the controller at lower levels.
[0173] At the next highest level, the software stack can include joint collection controllers. Joint collection controllers can handle the issuance of commands and state vectors exposed as a set of component abstractions. Each component can include, for example, a kinematic model for performing inverse kinematics calculations, constraint information, and joint state vectors and joint command vectors. For example, a single joint collection controller can be used to apply different sets of policies to different subsystems at lower levels. Joint collection controllers can effectively decouple the relationship between how motors are physically represented and how control policies are associated with these components. Thus, for example, if a robotic arm has a movable base, a joint collection controller can be used to impose a set of constraint policies on how the arm moves and different sets of constraint policies on how the movable base moves.
[0174] At the next highest level, the software stack can include a joint selection controller. The joint selection controller is responsible for dynamically selecting between commands originating from different sources. In other words, the joint selection controller can receive multiple commands during a control loop and select one of them to execute. This ability to dynamically select from multiple commands during a real-time control loop allows for a significantly increased control flexibility compared to traditional robot control systems.
[0175] At the next highest level, the software stack can include joint position controllers. Joint position controllers can receive target parameters and dynamically calculate the commands needed to achieve those parameters. For example, a joint position controller can receive a position target and calculate a setpoint to achieve that target.
[0176] At the next highest level, the software stack can include a Cartesian position controller and a Cartesian selection controller. The Cartesian position controller receives the target in Cartesian space as input and uses an inverse kinematics solver to compute the output in joint position space. Then, before passing the computation in joint position space to the joint position controller at the next lowest level in the stack, the Cartesian selection controller can impose constraints on the results computed by the Cartesian position controller. For example, the Cartesian position controller can be given three independent target states in Cartesian coordinates x, y, and z. In some cases, the target state may be position, while in others it may be the desired velocity.
[0177] Therefore, these functionalities provided by the software stack offer extensive flexibility for control instructions, allowing them to be easily expressed as target states in a way that naturally aligns with the higher-level planning techniques described above. In other words, when the planning process uses process definition diagrams to generate specific actions to be taken, it is not necessary to specify actions in the low-level commands for individual robot components. Instead, they can be expressed as high-level objectives accepted by the software stack, which are translated through various levels until they finally become low-level commands. Furthermore, actions generated through the planning process can be specified in Cartesian space in a way that makes them understandable to human operators, making debugging and analysis scheduling easier, faster, and more intuitive. Moreover, actions generated through the planning process do not need to be tightly coupled to any specific robot model or low-level command format. Instead, the same actions generated during the planning process can actually be executed by different robot models, as long as they support the same degrees of freedom and have already implemented the appropriate level of control in the software stack.
[0178] Embodiments of the subject matter and functional operation described in this specification may be implemented in digital electronic circuits, tangibly embodied computer software or firmware, computer hardware, including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium, for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof. Alternatively or additionally, program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus.
[0179] The term "data processing apparatus" refers to data processing hardware and includes all kinds of devices, apparatuses, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The apparatus may also be or further include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.
[0180] A computer program (also referred to or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not need to, correspond to a file in a file system. A program may be stored as a portion of a file that holds other programs or data, for example, as one or more scripts stored in a markup language document, as a single file dedicated to the program in question, or as multiple collaborative files, for example, as a file storing one or more modules, subroutines, or code portions. A computer program can be deployed to execute on a single computer or on multiple computers located in one place or distributed across multiple locations and interconnected via a data communication network.
[0181] For a computer system configured to perform a specific operation or action, this means that the system has software, firmware, hardware, or a combination thereof installed thereon, which, in operation, cause the system to perform those operations or actions. For a computer program configured to perform a specific operation or action, this means that the program includes instructions that, when executed by a data processing device, cause that device to perform those operations or actions.
[0182] As used herein, "engine" or "software engine" refers to a software-implemented input / output system that provides outputs different from the inputs. An engine can be a coded functional block, such as a library, platform, software development kit ("SDK"), or object. Each engine can be implemented on any suitable type of computing device including one or more processors and computer-readable media, such as a server, mobile phone, tablet computer, laptop computer, music player, e-book reader, laptop or desktop computer, PDA, smartphone, or other fixed or portable device. Furthermore, two or more engines can be implemented on the same computing device or on different computing devices.
[0183] The processes and logic flows described in this specification can be executed by one or more programmable computers, wherein the one or more programmable computers execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by special-purpose logic circuitry (such as an FPGA or ASIC), or by a combination of special-purpose logic circuitry and one or more programmable computers.
[0184] A computer suitable for executing computer programs can be based on a general-purpose or special-purpose microprocessor, or both, or any other type of central processing unit. Generally, the central processing unit receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are the central processing unit for executing or running instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented or incorporated therein by special-purpose logic circuitry. Generally, a computer will also include or be operatively coupled to one or more mass storage devices (e.g., disks, magneto-optical disks, or optical disks) for storing data, to receive data from or transfer data to, or both of, these mass storage devices. However, a computer does not need to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0185] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0186] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse, trackball, or a presence-sensing display or other surface through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser on the user's device in response to a request received from a web browser. Additionally, the computer can interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smartphone), running a messaging application, and in turn receiving response messages from the user.
[0187] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes backend components, such as a data server, or middleware components, such as an application server, or frontend components, such as a client computer with a graphical user interface, a web browser, or an application through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0188] A computing system may include clients and servers. Clients and servers are typically geographically separated and usually interact via a communication network. The client-server relationship arises from computer programs running on their respective computers, and they have a client-server relationship. In some embodiments, the server sends data, such as HTML pages, to a user device, for example, to display data to a user interacting with the device acting as a client and to receive user input from that user. Data generated at the user device, such as the result of user interaction, may be received at the server from the device.
[0189] In addition to the embodiments described above, the following embodiments are also innovative:
[0190] Example 1 is a method comprising:
[0191] Receive a skill template for a task to be performed by the robot, wherein the skill template defines a state machine with multiple subtasks and one or more corresponding transition conditions between one or more of the subtasks.
[0192] The skill template indicates which subtasks are demo subtasks that need to be improved using local demo data;
[0193] Obtain the basic control strategy for the demonstration subtask of the skill template;
[0194] For the demonstration subtask of the skill template, receive local demonstration data generated from a user demonstrating how to perform the demonstration subtask using the robot;
[0195] An improved machine learning model for the demonstration subtask is configured to use one or more input sensor streams to generate commands to be executed by the robot; and
[0196] The skill template is executed on the robot, thereby enabling the robot to perform the task through transitions in the state machine defined by the skill template, including executing commands generated by the machine learning model.
[0197] Example 2 is the method according to Example 1, wherein receiving local demo data includes receiving local demo data for each of a plurality of transformation conditions for the demo subtask of the skill template.
[0198] Example 3 is the method according to any of Examples 1-2, wherein the skill template specifies that less than all conversion conditions of the demonstration subtask require local demonstration data.
[0199] Example 4 is the method described according to any one of Examples 1-3, wherein the skill template is hardware-agnostic.
[0200] Example 5 is the method according to Example 4, and further includes adapting the skill template for multiple different robot models, including generating corresponding local demonstration data from the multiple different robot models.
[0201] Example 6 is a method according to any one of Examples 1-5, wherein the skill template specifies which input sensor streams are required to complete each of the plurality of subtasks.
[0202] Example 7 is the method described in Example 6, wherein the basic control strategy is generated from demonstrations of multiple different robot models.
[0203] Example 8 is the method according to any one of Examples 1-7, wherein the first subtask of the skill template references a different second skill template having multiple second subtasks.
[0204] Example 9 is a system comprising: one or more computers and one or more storage devices storing instructions, which, when executed by the one or more computers, are operable to cause the one or more computers to perform the method according to any one of Examples 1 to 8.
[0205] Example 10 is a computer storage medium coded with a computer program, the program including instructions that, when executed by a data processing device, are operable to cause the data processing device to perform the method described in any of Examples 1 to 8.
[0206] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope that may be claimed, but rather as descriptions of features characteristic of particular embodiments of a particular invention. Certain features described in the context of independent embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations, and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.
[0207] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific or sequential order shown, or requiring all illustrated operations to be performed to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0208] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions described in the claims can be performed in different orders and the desired results can still be obtained. As an example, the processes depicted in the figures do not necessarily require the specific or sequential order shown to obtain the desired results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method executed by one or more computers, the method comprising: Receive a skill template for a task to be performed by the robot, wherein the skill template defines a state machine with multiple subtasks and one or more corresponding transition conditions between one or more of the subtasks. The skill template indicates which subtasks are demo subtasks that need to be improved using local demo data; The basic control strategy for obtaining the demonstration subtask of the skill template includes one or more machine learning models that transform environmental observations into one or more actions. For the demonstration subtask of the skill template, local demonstration data generated by a user demonstrating how to perform the demonstration subtask using the robot is received, wherein receiving local demonstration data includes receiving local demonstration data for each of a plurality of transformation conditions for the demonstration subtask of the skill template; i) using the local demonstration data to improve the basic control policy for the demonstration subtask to generate a custom control policy, wherein the custom control policy for the demonstration subtask is configured to use one or more input sensor streams to generate commands to be executed by the robot; or ii) using the local demonstration data to train a residual reinforcement learning model for the demonstration subtask, wherein the residual reinforcement learning model is configured to use one or more input sensor streams to generate corrective actions, which will be combined with the basic actions generated by the basic control policy to generate commands to be executed by the robot; and The skill template is executed on the robot, thereby enabling the robot to perform the task through transitions in the state machine defined by the skill template, including executing i) commands generated by the customized control strategy or ii) commands generated by combining corrective actions generated by the residual reinforcement learning model and basic actions generated by the basic control strategy.
2. The method according to claim 1, wherein, The skill template specifies that less than all transformation conditions for the demo subtask require local demo data.
3. The method according to claim 1, wherein, The skill templates are hardware-agnostic.
4. The method according to claim 3 further includes adapting the skill template for multiple different robot models, including generating corresponding local demonstration data using the multiple different robot models.
5. The method according to claim 1, wherein, The skill template specifies which input sensor streams are needed to complete each of the plurality of subtasks.
6. The method according to claim 5, wherein, The basic control strategy was generated from demonstrations of multiple different robot models.
7. The method according to claim 1, wherein, The first subtask of the skill template references a different second skill template that has multiple second subtasks.
8. A system comprising: One or more computers and one or more storage devices storing instructions, wherein, when executed by said one or more computers, the instructions are operable to cause said one or more computers to perform the method according to any one of claims 1 to 7.
9. One or more non-transitory computer storage media encoded with computer program instructions, which, when executed by one or more computers, cause the one or more computers to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Unsupervised detection of intermediate reinforcement learning goals
CN110168574A
Systems, apparatus, and methods for robotic learning and execution of skills
WO2020047120A1