Distributed Robot Demonstration Learning
The system uses skill templates and demonstration data to enable efficient and robust robot learning, addressing the limitations of traditional methods by allowing robots to quickly adapt to specific tasks and environments with sensor-rich manipulation techniques.
Patent Information
- Application Number
- CN202180036594.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-21
- Filing Date
- 2021-05-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-05-17
AI Technical Summary
Existing robot control technologies have problems such as high computational volume, difficulty in scaling, sparse rewards and vulnerability, which makes manual programming cumbersome, time-consuming and difficult to generalize to different workspaces and robot models.
A demonstration-based robot learning method is adopted to generate customized control strategies through skill templates and local demonstration data, and quickly adapt to visual, tactile and proprioceptive data, combined with residual reinforcement learning to optimize control strategies.
It realizes that robots can quickly adapt to specific environments and hardware, reduces setting time, improves control accuracy and robustness, can be widely used in a variety of robot models, and improves the safety and efficiency of the robot industry.
Smart Images

Figure CN115666871B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to robots, and more particularly to planning the movement of a robot. Background Art
[0002] Robot control refers to controlling the physical movement of a robot in order to perform a task. For example, an industrial robot that manufactures cars can be programmed to first pick up a car part and then weld the car part to the car's frame. Each of these actions can itself include dozens or hundreds of individual movements of the robot's motors and actuators.
[0003] Robot planning has traditionally required a great deal of manual programming in order to carefully specify how the robot components should move to complete a particular task. Manual programming is tedious, time-consuming, and error-prone. In addition, a schedule manually generated for one workcell generally cannot be used for other workcells. In this specification, a workcell is the physical environment in which a robot will operate. A workcell has specific physical characteristics (e.g., physical dimensions) that impose constraints on how the robot can move within the workcell. Thus, a schedule manually programmed for one workcell may not be compatible with a workcell that has different robots, a different number of robots, or different physical dimensions.
[0004] Some research has been conducted to use machine learning control algorithms (e.g., reinforcement learning) to control a robot to perform a particular task. However, robots have several drawbacks that make traditional learning methods generally unsatisfactory.
[0005] First, robots naturally have a very complex, high-dimensional, and continuous action space. Thus, generating and evaluating all possible candidate actions is computationally expensive. Second, robot control is an environment with extremely sparse rewards because most possible actions do not lead to the completion of a particular task. A technique called reward shaping has been used to alleviate the sparse reward problem, but it is generally not scalable for hand-designed reward functions.
[0006] Another complication is that traditional techniques for using robot learning for robot control are very fragile. This means that even if a viable model has been successfully trained, even a very small change to the task, the robot, or the environment will render the entire model completely unusable.
[0007] All of these problems mean that traditional ways of using techniques such as reinforcement learning for robot control result in computationally intensive processes that are simply difficult to work with, do not scale well, and cannot be generalized to other situations. Summary of the Invention
[0008] This specification describes techniques related to demonstration-based robot learning. In particular, this specification describes how to program a robot to perform a robotic task using a customized control strategy learned from skill templates and demonstration data.
[0009] In this specification, a task refers to the ability of a particular robot to perform one or more subtasks. For example, the connector insertion task is the ability of a robot to insert a wire connector into a socket. This task typically includes two subtasks: 1) moving the robot's tool to the location of the socket, and 2) inserting the connector into the socket at a specific location.
[0010] In this specification, a subtask is an operation performed by a robot using a tool. For simplicity, when a robot has only one tool, a subtask can be described as an operation that the robot as a whole is to perform. Example subtasks include welding, dispensing, part positioning, and surface grinding, to name just a few. Subtasks are typically associated with the type of tool required to perform the subtask and the location within the coordinate system of the workspace where the subtask is to be performed.
[0011] In this specification, a skill template (or, for simplicity, a template) is a collection of data and software that allows a robot to be adjusted to perform a specific task. The skill template data represents one or more subtasks required to perform the task, as well as information describing which subtasks of the skill require local demonstration learning and which perceptual streams are needed to determine success or failure. Thus, a skill template can define demonstration subtasks that require local demonstration learning, non-demonstration subtasks that do not require local demonstration learning, or both.
[0012] These techniques are particularly advantageous for robotic tasks that have traditionally been difficult to control using machine learning (e.g., reinforcement learning). These tasks include tasks involving physical contact with objects in the workspace, such as grinding, connecting, and inserting tasks, as well as wiring, to name just a few.
[0013] Specific embodiments of the subject matter described in this specification can be implemented to achieve one or more of the following advantages. Learning using the demonstration data described in this specification addresses the sparse reward and non-generalizability problems of traditional reinforcement learning methods.
[0014] The system can use vision, proprioceptive (joint) data, tactile data, and any other features to perform tasks, which allows the system to quickly adapt to a specific robot model with high precision. The emphasis is on "sensor-rich robot manipulation", which contrasts with the classical view of minimal sensing in robots. This generally means that the same task can be completed using a cheaper robot with less setup time.
[0015] The techniques described below allow machine learning techniques to quickly adapt to any suitable robot with a properly installed hardware abstraction. In a typical scenario, a single non-expert person can train a robot to perform a skill template in less than a day of setup time. This is a huge improvement over traditional methods, which may require a team of experts to spend weeks designing a reward function to solve the problem and weeks of training time on a very large data center. This effectively allows machine learning-based robot control to be widely distributed to multiple types of robots, even robots that the system has never seen before.
[0016] These techniques can effectively implement robot learning as a service, which leads to greater use of the technology. This in turn makes the entire robot industry safer and more efficient as a whole.
[0017] Although the tasks are complex, the combination of reinforcement learning, machine learning-based perception data processing, and advanced impedance / admittance control will enable robot skills to still be executed with a very high success rate according to the requirements of industrial applications.
[0018] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a diagram of an example demonstration learning system.
[0020] Figure 2A is a diagram of an example system for performing a subtask using a customized control strategy based on local demonstration data.
[0021] Figure 2B is a diagram of another example system for performing a subtask using local demonstration data.
[0022] Figure 2C is a diagram of another example system for performing a subtask using residual reinforcement learning.
[0023] Figure 3A is a flowchart of an example process for combining sensor data from multiple different sensor streams.
[0024] Figure 3B is a diagram of a camera wristband.
[0025] Figure 3C is another example view of the camera wristband.
[0026] Figure 3D is another example view of the camera wristband.
[0027] Figure 4Illustrates an example skill template.
[0028] Figure 5 Is a flowchart of an example process for configuring a robot to perform a skill using the skill template.
[0029] Figure 6A Is a flowchart of an example process for using the skill template for a task guided by force.
[0030] Figure 6B Is a flowchart of an example process for training the skill template using a cloud-based training system.
[0031] Like reference numerals and names in the various figures indicate like elements. Detailed Description
[0032] Figure 1 Is a diagram of an example demonstration learning system. System 100 is an example of a system that can implement the demonstration-based learning techniques described in this specification.
[0033] System 100 includes a plurality of functional components, including an online execution system 110, a training system 120, and a robot interface subsystem 160. Each of these components can be implemented as one or more computer programs installed on one or more computers coupled to each other via any suitable communication network (e.g., an intranet or the Internet, or a combination of networks).
[0034] System 100 operates in the following two basic modes to control robots 170a-n: a demonstration mode and an execution mode.
[0035] In the demonstration mode, a user can control one or more robots 170a-n to perform a specific task or subtask. While doing so, the online execution system 110 collects status messages 135 and online observations 145 to generate local demonstration data. The demonstration data collector 150 is a module that can generate local demonstration data 115 from the status messages 135 and online observations 145, and then the online execution system 110 can provide the local demonstration data 115 to the training system 120. The training system can then generate a customized control strategy 125 specific to both the task and the specific characteristics of the robot performing the task.
[0036] In this specification, a control strategy is a module or subsystem that generates one or more next actions for a robot to perform in response to a given observed input. The output of the control strategy can affect the movement of one or more robot components (e.g., motors or actuators), either as commands directly output by the strategy or as higher-level commands that are each consumed by multiple robot components through the machinery of the robot control stack. Thus, a control strategy can include one or more machine learning models that translate environmental observations into one or more actions.
[0037] In this specification, local demonstration data is data collected while a user is controlling a robot to demonstrate how the robot performs a particular task by causing the robot to perform physical movements. Local demonstration data can include kinematic data, e.g., joint positions, orientations, and angles. Local demonstration data can also include sensor data, e.g., data collected from one or more sensors. Sensors can include force sensors; vision sensors, e.g., cameras, depth cameras, and lidar; electrical connection sensors; acceleration sensors; audio sensors; gyroscopes; contact sensors; radar sensors; and proximity sensors, e.g., infrared proximity sensors, capacitive proximity sensors, or inductive proximity sensors, to name just a few examples.
[0038] Typically, local demonstration data is obtained from one or more robots that are in close proximity to the user who is controlling the robot in demonstration mode. However, a close physical distance between the user and the robot is not a requirement for obtaining local demonstration data. For example, a user can obtain local demonstration data remotely from a particular robot via a remote user interface.
[0039] A training system 120 is a computer system that can generate a customized control strategy 125 from local demonstration data 115 using machine learning techniques. The training system 120 typically has significantly more computational resources than an online execution system 110. For example, the training system 120 can be a cloud-based computing system with hundreds or thousands of compute nodes.
[0040] To generate the customized control strategy 125, the training system 120 can first obtain or pre-generate a basic control strategy for the task. A basic control strategy is a control strategy that is expected to work well enough for a particular task such that any sufficiently similar robot can perform the task with reasonable proximity. For the vast majority of tasks, a standalone basic control strategy is not expected to be precise enough to complete the task with sufficient reliability. For example, tasks such as connecting and inserting typically require sub-millimeter precision, which cannot be achieved without the details provided by local demonstration data for a particular robot.
[0041] The basic control strategy for a particular task can be generated in a variety of ways. For example, the basic control strategy can be manually programmed, trained using traditional reinforcement learning techniques, or using the demonstration-based learning techniques described in this specification. All of these techniques can be suitable for pre-generating the basic control strategy before receiving local demonstration data for the task, because less consideration is given to time when generating the basic control strategy.
[0042] In some embodiments, the training system generates the basic control strategy from the generalized training data 165. While the local demonstration data 115 collected by the online execution system 110 is typically specific to a particular robot or a particular robot model, in contrast, the generalized training data 165 can be generated by one or more other robots, which do not have to be of the same model, located in the same place, or manufactured by the same manufacturer. For example, the generalized training data 165 can be generated off-site from dozens or hundreds or thousands of different robots with different characteristics and different models. In addition, the generalized training data 165 does not even need to be generated from physical robots. For example, the generalized training data can include data generated from simulations of physical robots.
[0043] Thus, the local demonstration data 115 is local in the sense that it is specific to a particular robot that the user can access and manipulate. The local demonstration data 115 thus represents data specific to a particular robot, but can also represent local variables, e.g., specific characteristics of a particular task and specific characteristics of a particular work environment.
[0044] System demonstration data collected during the process of developing the skill template can also be used to define the basic control strategy. For example, an engineer team associated with the entity that generates the skill template can perform demonstrations using one or more robots at a facility remote from and / or not associated with the system 100. The robots used to generate the system demonstration data also need not be the same robots or the same robot models as the robots 170a-n in the work area 170. In this case, the system demonstration data can be used to guide the actions of the basic control strategy. The basic control strategy can then be adjusted to a customized control strategy using more computationally expensive and sophisticated learning methods.
[0045] Adjusting a basic control strategy using local demonstration data has highly desirable effects, namely, it is relatively fast compared to generating the basic control strategy, e.g., by collecting system demonstration data or by training using generalized training data 165. For example, the size of the generalized training data 165 for a particular task often is several orders of magnitude larger than the local demonstration data 115, and thus it is expected that training the basic control strategy takes much longer than adapting it to a particular robot. For example, training the basic control strategy may require a large amount of computing resources and, in some cases, may require a data center with hundreds or thousands of machines working for days or weeks to train the basic control strategy based on the generalized training data. In contrast, adjusting the basic control strategy using the local demonstration data 115 can take only a few hours.
[0046] Similarly, collecting system demonstration data to define the basic control strategy can require more iterations than the local demonstration data. For example, to define the basic control strategy, a team of engineers can demonstrate 1000 successful tasks and 1000 unsuccessful tasks. In contrast, fully adjusting the resulting basic control strategy can require only 50 successful demonstrations and 50 unsuccessful demonstrations.
[0047] The training system 120 can thus use the local demonstration data 115 to refine the basic control strategy to generate a customized control strategy 125 for a particular robot for which the demonstration data was generated. The customized control strategy 125 adjusts the basic control strategy to account for the characteristics of the particular robot and local variables for the task. Training the customized control strategy 125 using the local demonstration data can take much less time than training the basic control strategy. For example, while training the basic control strategy can take many days or weeks, a user may spend only 1 - 2 hours generating the local demonstration data 115 with the robot and then can upload the local demonstration data 115 to the training system 120. The training system 120 can then generate the customized control strategy 125 in much less time than it takes to train the basic control strategy, e.g., perhaps just one or two hours.
[0048] In the execution mode, the execution engine 130 can use the customized control strategy 125 to automatically execute tasks without any user intervention. The online execution system 110 can use the customized control strategy 125 to generate commands 155 to be provided to the robot interface subsystem 160, which drives one or more robots (e.g., robots 170a - n) in the workspace 170. The online execution system 110 can consume the status messages 135 generated by the robots 170a - n and the online observations 145 made by one or more sensors 171a - n observing within the workspace 170. As Figure 1As shown, each sensor 171 is coupled to a corresponding robot 170. However, the sensors do not need to be in one-to-one correspondence with the robots, nor do they need to be coupled to the robots. In fact, each robot can have multiple sensors, and the sensors can be mounted on fixed or movable surfaces in the work area 170.
[0049] The execution engine 130 can use the status messages 135 and the online observations 145 as inputs to the customized control policy 125 received from the training system 120. Thus, the robots 170a-n can react in real time according to their specific characteristics and the specific characteristics of the tasks to complete the tasks.
[0050] Therefore, using local demonstration data to adjust the control policy results in a very different user experience. From the user's perspective, training a robot to perform a task very precisely with a customized control policy (including generating local demonstration data and waiting for the customized control policy to be generated) is a very fast process, which may take less than a day of setup time. The speed comes from making full use of the pre-computed basic control policy.
[0051] This arrangement introduces a huge technical improvement over existing robot learning methods, which typically require weeks of testing and generating manually designed reward functions, weeks of generating suitable training data, and more weeks of training, testing, and refining the model so that they are suitable for industrial production.
[0052] In addition, different from traditional robot reinforcement learning, using local demonstration data is highly robust to small perturbations of the characteristics of the robot, task, and environment. If a company purchases a new robot model, then the user only needs to spend one day generating new local demonstration data for the new customized control policy. This is in contrast to existing reinforcement learning methods, in which any change in the physical characteristics of the robot, task, or environment may require the entire weeks-long process to start from scratch.
[0053] To initiate a demonstration-based learning process, the online execution system can receive the skill template 105 from the training system 120. As described above, the skill template 105 can specify a sequence of one or more subtasks required to execute the skill, which subtasks require local demonstration learning, which perceptual streams are required for which subtasks, and the transition conditions that specify when to transition from one subtask of the execution skill template to the next.
[0054] As described above, the skill template can define demonstration subtasks that require local demonstration learning, non-demonstration subtasks that do not require local demonstration learning, or both.
[0055] The demonstration subtasks are implicitly or explicitly bound to a base control policy, which, as described above, can be pre-computed from generalized training data or system demonstration data. Thus, for each demonstration subtask in the template, the skill template can include a separate base control policy or an identifier of the base control policy.
[0056] For each demonstration subtask, the skill template can also include software modules required to adjust the demonstration subtask using local demonstration data. Each demonstration subtask can rely on a different type of machine learning model and can be adjusted using different techniques. For example, a mobile demonstration subtask can heavily rely on camera images of the local workspace environment in order to find a specific task target. Thus, the adjustment process for the mobile demonstration subtask can more heavily adjust the machine learning model to identify features in the camera images captured in the local demonstration data. In contrast, an insertion demonstration subtask can heavily rely on force feedback data to sense the edges of a connecting socket and insert a connector into the socket using an appropriate gentle force. Thus, the adjustment process for the insertion demonstration subtask can more strictly adjust the machine learning model that processes force perception and corresponding feedback. In other words, even if the underlying models for the subtasks in the skill template are the same, each subtask can have its own corresponding adjustment process in order to incorporate local demonstration data in different ways.
[0057] The non-demonstration subtasks may or may not be associated with a base control policy. For example, a non-demonstration subtask can simply specify moving to a particular coordinate location. Alternatively, a non-demonstration subtask can be associated with a base control policy, e.g., as computed according to other robots, which specifies how the joints should move to a particular coordinate location using sensor data.
[0058] The purpose of the skill template is to provide a generalized framework for programming a robot to have specific task capabilities. In particular, the skill template can be used to adjust the robot to perform similar tasks with relatively little effort. Thus, adjusting the skill template for a particular robot and a particular environment involves performing a training process for each demonstration subtask in the skill template. For the sake of brevity, this process can be referred to as training the skill template, even though multiple separately trained models may be involved.
[0059] For example, a user can download a connector insertion skill template that specifies performing a first movement subtask and then a connector insertion subtask. The connector insertion skill template can also specify that the first subtask depends on a visual perception stream (e.g., from a camera), but the second subtask depends on a force perception stream (e.g., from a force sensor). The connector insertion skill template can also specify that only the second subtask requires local demonstration learning. This may be because moving the robot to a specific location typically does not highly depend on the circumstances of the task at hand or the working environment. However, if the working environment has strict spatial requirements, then the template can also specify that the first subtask requires local demonstration learning so that the robot can quickly learn to navigate within the strict spatial requirements of the working environment.
[0060] To equip a robot with the connector insertion skill, the user only needs to guide the robot to perform the subtasks that require local demonstration data as indicated by the skill template. The robot will automatically capture the local demonstration data, and the training system can use this data to refine the basic control strategy related to the connector insertion subtask. When the training of the customized control strategy is completed, to be equipped to perform the subtask, the robot only needs to download the finally trained customized control strategy.
[0061] It is worth noting that the same skill template can be used for many different kinds of tasks. For example, the same connector insertion skill template can be used to equip a robot to perform HDMI cable insertion or USB cable insertion or both. All that is required of the user is to demonstrate these different insertion subtasks in order to refine the basic control strategy for the subtasks being learned. As described above, this process typically requires less computing power and less time compared to developing or learning a complete control strategy from scratch.
[0062] In addition, the skill template approach can be hardware-independent. This means that even if the training system has never trained a control strategy for that specific robot model, the skill template can be used to equip the robot to perform the task. Thus, this technique addresses many of the problems with using reinforcement learning to control robots. In particular, it addresses the vulnerability problem in which even very small hardware changes require re-learning the control strategy from scratch, which is an expensive and repetitive task.
[0063] To support the collection of local demonstration data, system 100 can also include one or more UI devices 180 and one or more demonstration devices 190. The UI devices 180 can help guide the user to obtain local demonstration data that will be most beneficial for generating the customized control strategy 125. The UI devices 180 can include a user interface that indicates what actions the user should perform or repeat, and an augmented reality device that allows the user to control the robot without being physically close to the robot.
[0064] The demonstration device 190 is a device that aids in the primary operation of the assistance system 100. Generally speaking, the demonstration device 190 is a device that allows a user to demonstrate a skill to the robot without introducing external force data into the local demonstration data. In other words, the demonstration device 190 can reduce the likelihood that the user's demonstration actions affect what the force sensors actually read during execution.
[0065] In operation, the robot interface subsystem 160 and the online execution system 110 can operate according to different timing constraints. In some embodiments, the robot interface subsystem 160 is a real-time software control system with hard real-time requirements. A real-time software control system is a software system that is required to execute within strict timing requirements to achieve normal operation. The timing requirements often specify that certain operations or outputs must be performed or generated within a specific time window so that the system avoids entering a fault state. In a fault state, the system may stop executing or take some other action that interrupts normal operation.
[0066] On the other hand, the online execution system 110 generally has more flexibility in operation. In other words, the online execution system 110 can (but does not have to) provide the command 155 within each real-time time window in which the robot interface subsystem 160 operates. However, in order to provide the ability to make sensor-based responses, the online execution system 110 can still operate under strict timing requirements. In a typical system, the real-time requirements of the robot interface subsystem 160 require the robot to provide a command every 5 milliseconds, while the online requirements of the online execution system 110 specify that the online execution system 110 should provide the command 155 to the robot interface subsystem 160 every 20 milliseconds. However, even if such a command is not received within the online time window, the robot interface subsystem 160 does not necessarily need to enter a fault state.
[0067] Accordingly, in this specification, the term online refers to both the time and stiffness parameters for operation. The time window is larger than the time window for the real-time robot interface subsystem 160 and generally has greater flexibility when timing constraints are not met. In some embodiments, the robot interface subsystem 160 provides a hardware-independent interface such that the commands 155 issued by the field execution engine 150 are compatible with multiple different versions of the robot. During execution, the robot interface subsystem 160 can report status messages 135 back to the online execution system 110 such that the online execution system 110 can make online adjustments to the robot movement, for example, due to local failures or other unexpected conditions. The robots can be real-time robots, meaning that the robots are programmed to continuously execute their commands according to a highly constrained timeline. For example, each robot can expect commands from the robot interface subsystem 160 at a particular frequency (e.g., 100 Hz or 1 kHz). If a robot does not receive the expected commands, the robot can then enter a fault mode and stop operating.
[0068] Figure 2A FIG. is of an example system 200 for performing subtasks using a customized control strategy based on local demonstration data. Generally, data from multiple sensors 260 is fed through multiple separately trained neural networks and combined into a single low-dimensional task state representation 205. The low-dimensional representation 205 is then used as an input to a conditioned control strategy 210 that is configured to generate robot commands 235 to be executed by a robot 270. Accordingly, system 200 can implement a customized control strategy based on local demonstration data by modifying the subsystem 280 to effect modifications to the base control strategy.
[0069] The sensors 260 can include sensing sensors that generate a sensed data stream representing the visual characteristics of a target in the robot or the robot's workspace. For example, to achieve better visual capabilities, a robot tool can be equipped with multiple cameras, such as visible light cameras, infrared cameras, and depth cameras, to name a few examples.
[0070] The different sensed data streams 202 can be processed independently by corresponding convolutional neural networks 220a-n. Each sensed data stream 202 can correspond to a different sensing sensor (e.g., a different camera or different type of camera). The data from each camera can be processed by a different corresponding convolutional neural network.
[0071] The sensor 260 also includes one or more robot state sensors that generate a robot state data stream 204, which represents the physical characteristics of the robot or components of the robot. For example, the robot state data stream 204 may represent forces, torques, angles, positions, velocities, and accelerations of the robot or corresponding components of the robot, to name just a few examples. Each robot state data stream 204 may be processed by a corresponding deep neural network 230a-m.
[0072] The modification subsystem 280 may have any number of neural network subsystems that process sensor data in parallel. In some embodiments, the system includes only one perception stream and one robot state data stream.
[0073] The output of the neural network subsystem is a corresponding part of the task state representation 205, which cumulatively represents the state of the subtasks being performed by the robot 270. In some embodiments, the task state representation 205 is a low-dimensional representation having less than 100 features (e.g., 10, 30, or 50 features). Having a low-dimensional task state representation means fewer model parameters to learn, which further increases the speed at which local demonstration data can be used to adapt to a specific subtask.
[0074] The task state representation 205 is then used as an input to the adjusted control strategy 210. During execution, the adjusted control strategy 210 generates a robot command 235 from the input task state representation 205, which is then executed by the robot 270.
[0075] During training, the training engine 240 generates a parameter correction 255 by using a representation of the local demonstration actions 275 and the proposed commands 245 generated by the adjusted control strategy 210. The training engine can then use the parameter correction 255 to refine the adjusted control strategy 210 so that the commands generated by the adjusted control strategy 210 in future iterations will more closely match the local demonstration actions 275.
[0076] During the training process, the adjusted control strategy 210 can be initialized with a base control strategy associated with the demonstration subtask being trained. The adjusted control strategy 210 can be iteratively updated using the local demonstration actions 275. The training engine 240 can use any suitable machine learning technique to adjust the adjusted control strategy 210, e.g., supervised learning, regression, or reinforcement learning. When the adjusted control strategy 210 is implemented using a neural network, the parameter correction 235 can be propagated back through the network so that the proposed commands 245 output will more closely match the local demonstration actions 275 in future iterations.
[0077] As mentioned above, each subtask of a skill template can have a different training priority, even if the underlying model architectures are the same or similar. Thus, in some embodiments, the training engine 240 can optionally take as input subtask hyperparameters 275 that specify how to update the conditioned control policy 210. For example, the subtask hyperparameters can indicate that vision sensing is very important. Thus, the training engine 240 can more aggressively correct the conditioned control policy 210 to align with camera data captured by the locally demonstrated actions 275. In some embodiments, the subtask hyperparameters 275 identify separate training modules to be used for each different subtask.
[0078] Figure 2B FIG. is a diagram of another example system for performing subtasks using local demonstration data. In this example, the system includes multiple independent control policies 210a-n instead of just having a single conditioned control policy. Each control policy 210a-n can use the task state representation 205 to generate corresponding robot sub-commands 234a-n. The system can then combine the sub-commands to generate a single robot command 235 that will be executed by the robot 270.
[0079] Having multiple individually adjustable control policies can be advantageous in sensor-rich environments where, for example, data from multiple sensors with different update rates can be used. For example, the different control policies 210a-n can be executed at different update rates, which allows the system to combine simple and more sophisticated control algorithms into the same system. For example, one control policy can focus on robot commands using current force data, which updates much faster than image data. At the same time, another control policy can focus on robot commands using current image data, which can require more sophisticated image recognition algorithms that can have an indeterminate runtime. The result is a system that can quickly adapt to both force data and image data without slowing down its adaptation to force data. During training, the subtask hyperparameters can identify separate training processes for each individually adjustable control policy 210-an.
[0080] Figure 2C FIG. is a schematic diagram of another example system for performing subtasks using residual reinforcement learning. In this example, the system uses a residual reinforcement learning subsystem 212 to generate corrective actions 225 that modify the base actions 215 generated by a base control policy 250, rather than having a single regulatory control policy that generates robot commands.
[0081] In this example, the base control strategy 250 takes sensor data 245 from one or more sensors 260 as input and generates base actions 215. As described above, the output of the base control strategy 250 can be one or more commands consumed by the corresponding components of the robot 270.
[0082] During execution, the reinforcement learning subsystem 212 generates corrective actions 225 to be combined with the base actions 215 from the input task state representation 205. The corrective actions 225 are corrective in the sense that they modify the base actions 215 from the base control strategy 250. The resulting robot commands 235 can then be executed by the robot 270.
[0083] Traditional reinforcement learning processes have used two phases: (1) an action phase where the system generates new candidate actions and (2) a training phase where the weights of the model are adjusted to maximize the cumulative reward for each candidate action. As described in the background section above, traditional methods of using reinforcement learning for robots suffer from a severe sparse reward problem, meaning that actions randomly generated during the action phase are highly unlikely to receive any type of reward through the reward function for the task.
[0084] However, different from traditional reinforcement learning, using local demonstration data can provide all the information about which action to select during the action phase. In other words, local demonstration data can provide a sequence of actions, so these actions do not need to be randomly generated. This technique greatly limits the problem space and makes the convergence speed of the model much faster.
[0085] During training, local demonstration data is used to drive the robot 270. In other words, the robot commands 235 generated by the corrective actions 225 and the base actions 215 are needed to drive the robot 270. At each time step, the reinforcement learning subsystem 210 receives a representation of the actions demonstrated for physically moving the robot 270. The reinforcement learning subsystem 210 also receives the base actions 215 generated by the base control strategy 250.
[0086] The reinforcement learning subsystem 210 can then generate reconstructed corrective actions by comparing the demonstrated actions with the base actions 215. The reinforcement learning subsystem 210 can also use the reward function to generate an actual reward value for the reconstructed corrective actions.
[0087] The reinforcement learning subsystem 210 can also generate predicted corrective actions generated by the current state of the reinforcement learning model and predicted reward values that have been generated by using the predicted corrective actions. The predicted corrective actions are the corrective actions that the reinforcement learning subsystem 210 has generated for the current task state representation 205.
[0088] The reinforcement learning subsystem 210 can then use the predicted corrective actions, predicted reward values, reconstructed corrective actions, and actual reward values to calculate weight updates for the reinforcement model. During an iteration of training data, the weight updates are used to adjust the predicted corrective actions to the reconstructed corrective actions reflected by the demonstrated actions. The reinforcement learning subsystem 210 can calculate the weight updates according to any suitable reward maximization process.
[0089] Figure 2A One ability provided by the architecture shown in -C is the ability to combine multiple different models for sensor streams with different update rates. Some real-time robots have very strict control loop requirements, and thus, they can be equipped with force and torque sensors that generate high-frequency updates, e.g., at 100, 1000, or 10,000 Hz. In contrast, few cameras or depth cameras operate at frequencies above 60 Hz.
[0090] Figure 2A The architecture shown in -C with multiple parallel and independent sensor streams and optionally multiple different control strategies allows for the combination of these different data rates.
[0091] Figure 3A is a flowchart of an example process for combining sensor data from multiple different sensor streams. The process can be performed by a computer system having one or more computers at one or more locations (e.g., Figure 1 system 100). The process will be described as being performed by a system of one or more computers.
[0092] The system selects a base update rate (302). The base update rate will specify the rate at which the learning subsystem (e.g., the tuned control strategy 210) will generate commands to drive the robot. In some embodiments, the system selects the base update rate based on the minimum real-time update rate of the robot. Alternatively, the system can select the base update rate based on the sensor that generates data at the fastest rate.
[0093] The system generates corresponding portions of a task state representation at corresponding update rates (304). Because the neural network subsystems can operate independently and in parallel, the neural network subsystems can repeatedly generate corresponding portions of the task state representation at rates indicated by the rates of their respective sensors.
[0094] To enhance the independent and parallel nature of the system, in some embodiments, the system maintains multiple separate memory devices or memory partitions into which different portions of the task state representation will be written. This can prevent different neural network subsystems from competing for memory access when generating outputs at high frequencies.
[0095] The system repeatedly generates a task state representation (306) at a base update rate. During each time period defined by the base update rate, the system can generate a new version of the task state representation by reading from the most recently updated sensor data output by multiple neural network subsystems. For example, the system can read from multiple separate memory devices or memory partitions to generate a complete task state representation. Notably, this means that the generation rate of data generated by some neural network subsystems is different from the rate at which it is consumed. For example, for sensors with a slower update rate, the rate of consumption of data can be much faster than the rate at which it is generated.
[0096] The system repeatedly uses the task state representation at the base update rate to generate commands for the robot (308). By using independent and parallel neural network subsystems, the system can ensure that commands are generated at a sufficiently fast update rate to power a robot with hard real-time constraints.
[0097] This arrangement also means that the system can simultaneously feed multiple independent control algorithms with different update frequencies. For example, as described above with respect to Figure 2B Rather than the system generating a single command, the system can include multiple independent control strategies that each generate sub-commands. The system can then generate a final command by combining the sub-commands into a final hybrid robot command that represents the output of multiple different control algorithms.
[0098] For example, a vision control algorithm can cause the robot to move faster towards an identified object. At the same time, a force control algorithm can cause the robot to track along a surface it has contacted. Even though the vision control algorithm is typically updated at a much slower rate than the force control algorithm, the system can still use the architecture depicted in Figure 2A -C to power both at the base update rate.
[0099] Figure 2A The architecture shown in -C provides many opportunities to expand the capabilities of the system without significant re-engineering. Multiple parallel and independent data streams allow for the implementation of machine learning functions that are beneficial for local demonstration learning.
[0100] For example, in order to more thoroughly adapt the robot to a specific environment, it can be highly advantageous to integrate sensors that consider local environmental data.
[0101] An example of using local environmental data is a function that considers electrical connectivity. Electrical connectivity can be useful as a reward factor for various challenging robot tasks, including establishing an electric current between two components. These tasks include inserting a cable into a jack, plugging a power plug into a power socket, and screwing in a light bulb, to name just a few examples.
[0102] To integrate electrical connectivity into the modification subsystem 280, an electrical sensor, which can be one of the sensors 260 for example, can be configured in the workspace to detect when a current has been established. The output of the electrical sensor can then be processed by a separate neural network subsystem, and the result can be added to the task state representation 205. Alternatively, the output of the electrical sensor can be provided directly as an input to a system implementing a regulated control strategy or a reinforcement learning subsystem.
[0103] Another example of using local environmental data is a function that considers certain types of audio data. For example, many connector insertion tasks make a very distinctive sound when the task is successfully completed. Thus, the system can use a microphone that captures audio and an audio processing neural network whose output can be added to the task state representation. The system can then use a function that takes into account the specific acoustic characteristics of the sound of connector insertion, which forces the learning subsystem to learn what a successful connector insertion sounds like.
[0104] Figure 3B is a schematic diagram of a camera wristband. The camera wristband is an example of a rich instrument type that can be used to perform high-precision demonstration learning through the above architecture. Figure 3B is a perspective view in which the tool at the end of the robotic arm is closest to the observer.
[0105] In this example, the camera wristband is mounted on the robotic arm 335, just before the tool 345 located at the very end of the robotic arm 335. The camera wristband is mounted to the robotic arm 335 with a collar 345 and has four radially mounted cameras 310a-d.
[0106] The collar 345 can have any suitable convex shape that allows the collar 345 to be firmly mounted to the end of the robotic arm. The collar 345 can be designed to be added to a robot manufactured by a third-party manufacturer. For example, a system that distributes skill templates can also distribute camera wristbands to help non-expert users quickly converge the model. Alternatively or additionally, the collar 345 can be integrated into the robotic arm by the manufacturer during the manufacturing process.
[0107] The collar 345 can have an elliptical shape, such as circular or oval, or a rectangular shape. The collar 345 can be formed from a single solid volume that is fastened to the end of the robotic arm before the tool 345 is fastened. Alternatively, the collar 345 can be opened and firmly closed by a fastening mechanism (e.g., a buckle or a latch). The collar 345 can be made of any suitable material that provides a secure connection to the robotic arm, e.g., hard plastic; fiberglass; fabric; or metal, e.g., aluminum or steel.
[0108] Each camera 310a-d has corresponding bases 325a-d that secure a sensor, other electronics, and corresponding lenses 315a-d to a collar 345. The collar 345 may also include one or more lights 355a-b for illuminating the volume captured by the cameras 310a-d. Generally, the cameras 310a-d are arranged to capture different corresponding views of a working volume either within or outside of the tool 345.
[0109] An example camera wristband has four radially-mounted cameras, but any suitable number of cameras may be used, e.g., 2, 5, or 10 cameras. As described above, modifying the architecture of the subsystem 280 allows any number of sensor streams to be included in the task state representation. For example, a computer system associated with the robot may implement different corresponding convolutional neural networks to process the sensor data generated by each of the cameras 310a-d in parallel. The processed camera outputs may then be combined to generate a task state representation, which, as described above, may be used to power multiple control algorithms operating at different frequencies. As described above, the processed camera outputs may be combined with the outputs of other networks from independently processed force sensors, torque sensors, position sensors, velocity sensors, or tactile sensors or any suitable combination of these sensors.
[0110] Generally, using the camera wristband during the demonstration learning process results in faster convergence of the model because the system will be able to identify reward conditions in more positions and orientations. Thus, using the camera wristband can effectively further reduce the amount of training time required to adjust the base control strategy with local demonstration data.
[0111] Figure 3C is another example view of the camera wristband. Figure 3C Illustrated are further instruments that may be used to implement the camera wristband, including cables 385a-d that may be used to feed the outputs of the cameras to corresponding convolutional neural networks. Figure 3C Also illustrated is how an additional depth camera 375 may also be mounted to the collar 345. As described above, the architecture of the system allows any other sensors to be integrated into the perception system, and thus, for example, a separately trained convolutional neural network may process the output of the depth camera 375 to generate another portion of the task state representation.
[0112] Figure 3D is another example view of the camera wristband. Figure 3D is a perspective view of a camera wristband having a metal collar and four radially-mounted cameras 317a-d.
[0113] With these basic mechanisms for refining control strategies using local demonstration data, users can write tasks to build hardware-independent skill templates that can be downloaded and used to quickly deploy tasks across a variety of different types of robots and a variety of different types of environments.
[0114] Figure 4 An example skill template 400 is illustrated. Generally speaking, a skill template defines a state machine for multiple subtasks required to perform a task. Notably, skill templates are composable hierarchically, meaning that each subtask can be an independent task or another skill template.
[0115] Each subtask of a skill template has a subtask id and includes subtask metadata that includes whether the subtask is a demonstration subtask or a non-demonstration subtask, or whether the subtask references another skill template that should be trained separately. The subtask metadata can also indicate which sensor streams will be used to perform the subtask. A subtask that is a demonstration subtask will additionally include a base policy id that identifies the base policy that will be combined with corrective action learning learned from local demonstration data. Each demonstration subtask will also be explicitly or implicitly associated with one or more software modules that control the training process for the subtask.
[0116] Each subtask of a skill template also has one or more transition conditions that specify under which conditions a transition should be made to another task in the skill template. The transition conditions can also be referred to as the subtask goals of the subtask.
[0117] Figure 4 The example in illustrates a skill template for performing a task that is difficult to achieve using traditional robotics learning techniques. The task is a grasping and connecting insertion task that requires the robot to find a wire in a workspace and insert a connector at one end of the wire into a socket that is also located in the workspace. This problem is difficult to generalize with traditional reinforcement learning techniques because wires come in many different textures, diameters, and colors. Additionally, if the grasping subtask of the skill is not successful, traditional reinforcement learning techniques cannot tell the robot what to do next or how to make progress.
[0118] Skill template 400 includes four subtasks, which are represented as nodes in a graph that defines a state machine in Figure 4 In practice, Figure 4 all the information in can be represented in any suitable format, for example, as a plain text configuration file or a record in a relational database. Alternatively or additionally, a user interface device can generate a graphical skill template editor that allows the user to define skill templates through a graphical user interface.
[0119] The first subtask in skill template 400 is the movement subtask 410. The movement subtask 410 is designed to locate a wire in the workspace, which requires moving the robot from an initial position to the intended position of the wire, e.g., as placed by a previous robot in an assembly line. Moving from one position to the next is generally less dependent on the local characteristics of the robot, so the metadata for the movement subtask 410 specifies that this subtask is a non-demonstration subtask. The metadata for the movement subtask 410 also specifies that a camera stream is required to locate the wire.
[0120] The movement subtask 410 also specifies the "acquire wire vision" transition condition 405, which indicates when the robot should transition to the next subtask in the skill template.
[0121] The next subtask in skill template 400 is the grasping subtask 420. The grasping subtask 420 is designed to grasp the wire in the workspace. This subtask is highly dependent on the characteristics of the wire and the characteristics of the robot, especially the tool used to grasp the wire. Therefore, the grasping subtask 420 is designated as a demonstration subtask that requires refinement with local demonstration data. The grasping subtask 420 is thus also associated with a base policy id that identifies a previously generated base control policy for generally grasping wires.
[0122] The grasping subtask 420 also specifies that both a camera stream and a force sensor stream are required to perform the subtask.
[0123] The grasping subtask 420 also includes three transition conditions. The first transition condition, the "lost wire vision" transition condition 415, is triggered when the robot loses visual contact with the wire. For example, this can occur when the wire is accidentally moved in the workspace (e.g., by a person or another robot). In this case, the robot transitions back to the movement subtask 410.
[0124] The second transition condition of the grasping subtask 420 (the "grasping failed" transition condition 425) is triggered when the robot attempts to grasp the wire but fails. In that scenario, the robot can simply loop back and try the grasping subtask 420 again.
[0125] The third transition condition of the grasping subtask 420 (the "grasping successful" transition condition 435) is triggered when the robot attempts to grasp the wire and succeeds.
[0126] Demonstration subtasks can also indicate which transition conditions require local demonstration data. For example, a particular subtask can indicate that all three transition conditions require local demonstration data. Thus, the user can demonstrate how to grasp successfully, demonstrate a failed grasp, and demonstrate the robot losing vision of the wire.
[0127] The next subtask in the skill template 400 is the second movement subtask 430. The movement subtask 430 is designed to move the grasped wire to a position near the socket in the workspace. In many connection and insertion scenarios that the user wishes the robot to perform, the socket is located in an extremely restricted space, such as inside a dishwasher, a TV, or a microwave oven being assembled. Since moving in that extremely restricted space is highly dependent on the subtask and the workspace, the second movement task 430 is designated as a demonstration subtask, even though it only involves moving from one position to another in the workspace. Thus, a movement subtask can be either a demonstration subtask or a non-demonstration subtask, depending on the requirements of the skill.
[0128] Although the second movement subtask 430 is indicated as a demonstration subtask, the second movement subtask 430 does not specify any base policy id. This is because some subtasks are so dependent on the local workspace that including a base policy would only hinder the convergence of the model. For example, if the second movement task 430 requires moving the robot inside an appliance in a very specific orientation, a generalized base policy for movement would not be helpful. Thus, the user can perform a refinement process to generate local demonstration data that demonstrates how the robot should move through the workspace to obtain a specific orientation inside the appliance.
[0129] The second movement subtask 430 includes two transition conditions 445 and 485. The first "acquired socket vision" transition condition 445 is triggered when the camera stream makes visual contact with the socket.
[0130] If the robot happens to drop the connection while moving towards the socket, then the second "dropped connection" transition condition 485 is triggered. In that case, the skill template 400 specifies that the robot will need to return to the movement subtask 1 to restart the skill. These kinds of transition conditions in the skill template provide the robot with a level of built-in robustness and dynamic responsiveness that traditional reinforcement learning techniques simply cannot provide.
[0131] The last subtask in the skill template 400 is the insertion subtask 440. The insertion subtask 440 is designed to insert the connector of the grasped wire into the socket. The insertion subtask 440 is highly dependent on the type of wire and the type of socket. Thus, the skill template 400 indicates that the insertion subtask 440 is a demonstration subtask associated with a base policy id that generally involves insertion subtasks. The insertion subtask 440 also indicates that this subtask requires a camera stream and a force sensor stream.
[0132] The insertion subtask 440 includes three transition conditions. When the insertion fails for any reason and re - insertion is specified, the first "insertion failure" transition condition 465 is triggered. When the socket happens to move out of the camera's sight and re - moving the wire in the extremely restricted space to the socket's position is specified, the second "lost socket vision" transition condition 455 is triggered. Finally, the "dropped connection" transition condition 475 is triggered when the connection is lost during the execution of the insertion task. In this case, the skill template 400 specifies returning all the way to the first movement subtask 410.
[0133] Figure 4 One of the main advantages of the skill template shown in is developer composability. This means that new skill templates can be composed of subtasks that have already been developed. This feature also includes hierarchical composability, which means that each subtask within a particular skill template can reference another skill template.
[0134] For example, in an alternative embodiment, the insertion subtask 440 can actually reference an insertion skill template that defines a state machine for multiple fine - controlled movements. For example, the insertion skill template can include a first movement subtask aimed at aligning the connector with the socket as precisely as possible, a second movement subtask aimed at achieving sub - contact between the side of the connector and the socket, and a third movement subtask aimed at achieving full connection by using the side of the socket as a force guide.
[0135] And further skill templates can be hierarchically composed from the skill template 400. For example, the skill template 400 can be a small part of a more complex set of subtasks required to assemble an electronic appliance. The entire skill template can have multiple connector insertion subtasks, each of which references a skill template for implementing the subtask, such as the skill template 400.
[0136] Figure 5 is a flowchart of an example process for configuring a robot to execute a skill using a skill template. The process can be executed by a computer system having one or more computers at one or more locations (e.g., Figure 1 system 100). The process will be described as being executed by a system of one or more computers.
[0137] The system receives a skill template (510). As described above, the skill template defines a state machine with multiple subtasks and transition conditions that define when the robot should transition from executing one task to the next. In addition, the skill template can define which tasks are demonstration subtasks that require refinement using local demonstration data.
[0138] The system obtains a basic control strategy (520) for a demonstration subtask of a skill template. The basic control strategy can be a generalized control strategy generated from multiple different robot models.
[0139] The system receives local demonstration data (530) for the demonstration subtask. The user can use an input device or a user interface to have the robot perform the demonstration subtask over multiple iterations. During this process, the system automatically generates local demonstration data for performing the subtask.
[0140] The system trains a machine learning model for the demonstration subtask (540). As described above, the machine learning model can be configured to generate commands to be executed by the robot for one or more input sensor streams, and the machine learning model can be tuned using the local demonstration data. In some embodiments, the machine learning model is a residual reinforcement learning model that generates corrective actions to be combined with basic actions generated by the basic control strategy.
[0141] The system executes the skill template on the robot (550). After training all the demonstration subtasks, the system can use the skill template to have the robot fully perform the task. During this process, the robot will use the local demonstration data with refined demonstration subtasks tailored to the specific hardware and working environment of the robot.
[0142] Figure 6A is a flowchart of an example process for using a skill template for a task guided by force. The above skill template arrangement provides a relatively simple way to generate a very precise task consisting of multiple extremely complex subtasks. An example of such a task is a connector insertion task that uses a task guided by force data. This allows the robot to achieve a higher precision than would otherwise be possible. The process can be executed by a computer system having one or more computers at one or more locations (e.g., Figure 1 system 100). The process will be described as being executed by a system of one or more computers.
[0143] The system receives a skill template with a transition condition that requires establishing a physical contact force between an object held by the robot and a surface in the robot's environment (602). As described above, the skill template can define a state machine with multiple tasks. The transition condition can define the transition between a first subtask and a second subtask of the state machine.
[0144] For example, the first subtask can be a movement subtask and the second subtask can be an insertion subtask. The transition condition can specify that a connector held by the robot and to be inserted into a socket needs to generate a physical contact force with the edge of the socket.
[0145] The system receives local demonstration data (604) for conversion. In other words, the system can ask the user to demonstrate the conversion between the first subtask and the second subtask. The system can also ask the user to demonstrate a fault scenario. One such fault scenario can be the loss of the physical contact force with the socket edge. If this occurs, then as specified by the conversion conditions, the skill template can specify a return to the first movement subtask of the template so that the robot can re - establish the physical contact force.
[0146] The system trains a machine - learning model using the local demonstration data (606). As described above, through training, the system learns to avoid actions that lead to the loss of physical contact force and learns to select actions that can maintain the physical contact force throughout the second task.
[0147] The system executes the trained skill template on the robot (608). This enables the robot to automatically execute the subtasks and conversions defined by the skill template. For example, for a connection and insertion task, the local demonstration data can make the robot highly adaptable to inserting a specific type of connector.
[0148] Figure 6B Is a flowchart of an example process for training a skill template using a cloud - based training system. Generally, the system can generate all demonstration data locally and then upload the demonstration data to a cloud - based training system to train all the demonstration subtasks of the skill template. This process can be executed by a computer system (e.g., Figure 1 system 100) having one or more computers at one or more locations. The process will be described as being executed by a system of one or more computers.
[0149] The system receives a skill template (610). For example, an online execution system can download a skill template from a cloud - based training system that will train the demonstration subtasks of the skill template or from another computer system.
[0150] The system identifies one or more demonstration subtasks defined by the skill template (620). As described above, each subtask defined in the skill template can be associated with metadata indicating whether the subtask is a demonstration subtask or a non - demonstration subtask.
[0151] The system generates a corresponding set of local demonstration data for each of the one or more demonstration subtasks (630). As described above, the system can instantiate and deploy separate task systems, each of which generates local demonstration data when the user manipulates the robot to perform a subtask in the local workspace. The task - state representation can be generated at the base rate of the subtask, regardless of the update rate of the sensors contributing data to the task - state representation. This provides a convenient way to store and organize local demonstration data rather than generating many different sets of sensor data that must be coordinated later.
[0152] The system uploads the local demonstration data set to a cloud-based training system (640). Most facilities that use robots to perform actual tasks do not have an on-site data center suitable for training sophisticated machine learning models. Thus, while the local demonstration data can be collected on-site by a system co-located with the robot that will perform the task, the actual model parameters can be generated by a cloud-based training system that can only be accessed via the Internet or other computer networks.
[0153] As noted above, the size of the local demonstration data is expected to be several orders of magnitude smaller than the size of the data used to train the basic control strategy. Thus, while the local demonstration data can be large, the upload burden is controllable within a reasonable time, e.g., an upload time of a few minutes to an hour.
[0154] The cloud-based training system generates corresponding trained model parameters (650) for each local demonstration data set. As noted above, the training system can train the learning system to generate robot commands, which can, for example, consist of corrective actions that correct the basic actions generated by the basic control strategy. As part of this process, the training system can obtain the corresponding basic control strategy for each demonstration subtask locally or from another computer system, which can be a third-party computer system that publishes tasks or skill templates.
[0155] The cloud-based training system will typically have much more computing power than the online execution system. Thus, while training each demonstration subtask involves a large computational burden, these operations can be massively parallelized on the cloud-based training system. Thus, in a typical scenario, the time required to train a skill template from local demonstration data on the cloud-based training system does not exceed a few hours.
[0156] The system receives the trained model parameters generated by the cloud-based training system (660). The size of the trained model parameters is typically much smaller than the size of the local demonstration data for a particular subtask, and thus, after training the model, the time taken to download the trained parameters is negligible.
[0157] The system uses the trained model parameters generated by the cloud-based training system to execute the skill template (670). As part of this process, the system can also download the basic control strategy for the demonstration subtask, e.g., from the training system, from its original source, or from another source. The trained model parameters can then be used to generate commands for the robot to execute. In a reinforcement learning system for a demonstration subtask, the parameters can be used to generate corrective actions that modify the basic actions generated by the basic control strategy. The online execution system can then repeatedly issue the resulting robot commands to drive the robot to perform a specific task.
[0158] Figure 6B The process described Figure 6B can be performed by a single - person team in one day to enable a robot to perform highly precise skills in a manner suitable for its environment. This is a huge improvement over traditional manual programming methods or even traditional reinforcement learning methods, which require a team of many engineers to spend weeks or months designing, testing, and training models that do not generalize well to other scenarios.
[0159] In this specification, a robot is a machine having a base position, one or more movable components, and a kinematic model that can be used to map a desired position, orientation, or both in a coordinate system (e.g., Cartesian coordinates) to commands for physically moving one or more movable components to the desired position or orientation. In this specification, a tool is a device that is part of the kinematic chain of one or more movable components of the robot and is attached to the end thereof. Example tools include grippers, welding devices, and grinding devices.
[0160] In this specification, a task is an operation performed by a tool. For simplicity, when a robot has only one tool, the task can be described as an operation to be performed by the entire robot. Example tasks include welding, dispensing, part positioning, and surface grinding, to name just a few. A task is generally associated with the type of tool required to perform the task and the location within the workspace where the task will be performed.
[0161] In this specification, a motion plan is a data structure that provides information for performing an action, which can be a task, a cluster of tasks, or a transformation. A motion plan can be: fully - constrained, meaning that all values of all controllable degrees of freedom of the robot are explicitly or implicitly represented; or under - constrained, meaning that some values of the controllable degrees of freedom are not specified. In some embodiments, in order to actually execute the action corresponding to the motion plan, the motion plan must be fully - constrained to include all necessary values of all controllable degrees of freedom of the robot. Thus, at some points in the planning process described in this specification, some motion plans can be under - constrained, but when the motion plan is actually executed on the robot, the motion plan can be fully - constrained. In some embodiments, a motion plan represents an edge in a task graph between two configuration states for a single robot. Thus, generally each robot has one task graph.
[0162] In this specification, a motion - swept volume is the spatial region occupied by at least a portion of the robot or tool during the entire execution of a motion plan. The motion - swept volume can be generated by the collision geometry associated with the robot - tool system.
[0163] In this specification, a transformation is a motion plan that describes a movement to be performed between a start point and an end point. The start point and the end point can be represented by a pose, a position in a coordinate system, or a task to be performed. A transformation may be underconstrained by the lack of one or more values of one or more corresponding controllable degrees of freedom (DOFs) of the robot. Some transformations represent free motion. In this specification, free motion is a transformation in which none of the degrees of freedom are constrained. For example, a robot motion that simply moves from pose A to pose B with no restrictions on how to move between the two poses is free motion. During the planning process, values are ultimately assigned to the DOF variables for free motion, and the path planner can use any appropriate values for the motion that do not conflict with the physical constraints of the workspace.
[0164] The robot functions described in this specification can be implemented by a software stack that is at least partially hardware-independent, or, for brevity, simply a software stack. In other words, the software stack can accept input commands generated by the above-described planning process without requiring commands that are specifically related to a particular robot model or particular robot components. For example, the software stack can be implemented at least in part by Figure 1 a field execution engine 150 and a robot interface subsystem 160.
[0165] The software stack can include multiple levels that increase hardware specialization in one direction and software abstraction in the other direction. At the lowest level of the software stack are the robot components, which include devices that perform low-level actions and sensors that report low-level states. For example, a robot can include various low-level components, including motors, encoders, cameras, drivers, grippers, specialized sensors, linear or rotary position sensors, and other peripherals. As an example, a motor can receive a command indicating the amount of torque that should be applied. In response to receiving the command, the motor can report the current position of the robot joint to a higher level of the software stack, for example, using an encoder.
[0166] Each next-highest level in the software stack can implement an interface that supports multiple different underlying implementations. In general, each interface between levels provides status messages from the lower level to the higher level and commands from the higher level to the lower level.
[0167] Typically, commands and status messages are cyclically generated during each control cycle. For example, one status message and one command per control cycle. Lower-level software stacks generally have more stringent real-time requirements than higher-level software stacks. For example, at the lowest level of the software stack, the control cycle can have actual real-time requirements. In this specification, real-time means that commands received at one layer of the software stack must be executed, and optionally, status messages are provided back to the upper layer of the software stack within a specific control cycle time. If this real-time requirement is not met, then the robot can be configured to enter a fault state, for example, by freezing all operations.
[0168] At the next highest level, the software stack can include software abstractions of specific components, which will be referred to as motor feedback controllers. The motor feedback controller can be a software abstraction of any suitable low-level component, not just literally a motor. Thus, the motor feedback controller receives status into the lower-level hardware components through an interface and sends commands back down to the lower-level hardware components through the interface based on higher-level commands received from the higher level in the stack. The motor feedback controller can have any suitable control rules that determine how the upper-level commands should be interpreted and transformed into lower-level commands. For example, the motor feedback controller can use anything from simple logical rules to more advanced machine learning techniques to transform the upper-level commands into lower-level commands. Similarly, the motor feedback controller can use any suitable fault rules to determine when a fault state is reached. For example, if the motor feedback controller receives an upper-level command but does not receive a lower-level status within a specific part of the control cycle, then the motor feedback controller can put the robot into a fault state of stopping all operations.
[0169] At the next highest level, the software stack can include an actuator feedback controller. The actuator feedback controller can include control logic for controlling multiple robot components through their respective motor feedback controllers. For example, some robot components (e.g., joint arms) can actually be controlled by multiple motors. Thus, the actuator feedback controller can provide a software abstraction of the joint arm by sending commands to the motor feedback controllers of multiple motors using its control logic.
[0170] At the next highest level, the software stack can include a joint feedback controller. The joint feedback controller can represent joints mapped to the logical degrees of freedom in the robot. Thus, for example, although the robot's wrist may be controlled by a complex network of actuators, the joint feedback controller can abstract this complexity and expose this degree of freedom as a single joint. Thus, each joint feedback controller can control an arbitrarily complex network of actuator feedback controllers. As an example, a six-degree-of-freedom robot can be controlled by six different joint feedback controllers, each controlling a separate network of actual feedback controllers.
[0171] Each level of the software stack may also enforce level-specific constraints. For example, if a particular torque value received by the actuator feedback controller is outside the acceptable range, then the actuator feedback controller may either modify it to be within the range or enter a fault state.
[0172] To drive the inputs to the joint feedback controller, the software stack may use a command vector that includes command parameters for each component at a lower level for each motor in the system, e.g., position, torque, and velocity. To expose the state from the joint feedback controller, the software stack may use a state vector that includes state information for each component at a lower level, e.g., the position, velocity, and torque of each motor in the system. In some embodiments, the command vector also includes some limit information regarding the constraints to be enforced by the lower-level controllers.
[0173] At the next highest level, the software stack may include a joint collection controller. The joint collection controller may handle the publication of command and state vectors exposed as a collection of abstractions. Each abstraction may include a kinematic model, e.g., for performing inverse kinematics calculations, limit information, and a joint state vector and a joint command vector. For example, a single joint collection controller may be used to apply different sets of policies to different subsystems at a lower level. The joint collection controller may effectively decouple the relationship between how the motors are physically represented and how the control policies are associated with those abstractions. Thus, for example, if a robotic arm has a movable base, then the joint collection controller may be used to enforce a set of limit policies on how the arm moves and a different set of limit policies on how the movable base may move.
[0174] At the next highest level, the software stack may include a joint selection controller. The joint selection controller may be responsible for making a dynamic selection between commands issued from different sources. In other words, the joint selection controller may receive multiple commands during a control cycle and select one of the multiple commands to be executed during the control cycle. The ability to dynamically select from multiple commands during a real-time control cycle allows for a significant increase in the control flexibility of a conventional robotic control system.
[0175] At the next highest level, the software stack may include a joint position controller. The joint position controller may receive target parameters and dynamically calculate the commands required to achieve the target parameters. For example, the joint position controller may receive a position target and may calculate the setpoints for achieving the target.
[0176] At the next highest level, the software stack can include a Cartesian position controller and a Cartesian selection controller. The Cartesian position controller can receive an input target in Cartesian space and use an inverse kinematics solver to calculate an output in joint position space. The Cartesian selection controller can then enforce a limiting policy on the result calculated by the Cartesian position controller and then pass the calculated result in joint position space to the next lowest-level joint position controller in the stack. For example, the Cartesian position controller can give three separate target states in Cartesian coordinates x, y, and z. For some degrees, the target state can be a position, while for other degrees, the target state can be a desired velocity.
[0177] Thus, these functions provided by the software stack provide extensive flexibility for control instructions to be easily expressed as target states in a manner that naturally combines with the higher-level planning techniques described above. In other words, when the planning process uses a process definition diagram to generate the specific actions to be taken, there is no need to specify these actions in the low-level commands of individual robot components. More precisely, they can be expressed as high-level targets accepted by the software stack, which are translated through the various levels until they finally become low-level commands. Moreover, the actions generated by the planning process can be specified in Cartesian space in a way that human operators can understand them, which makes it easier, faster, and more intuitive to debug and analyze the plan schedule. In addition, the actions generated by the planning process do not need to be tightly coupled to any specific robot model or low-level command format. Instead, the same actions generated during the planning process can actually be executed by different robot models as long as they support the same degrees of freedom and the appropriate control levels have been implemented in the software stack.
[0178] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access storage device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus.
[0179] The term "data processing apparatus" refers to data processing hardware and encompasses all types of devices, equipment, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The apparatus may also be or further include dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to the hardware, the apparatus may optionally include code that creates an execution environment for a computer program, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0180] A computer program (which may also be referred to as or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may but need not correspond to a file in a file system. A program can be stored in a part of a file that holds other programs or data (such as one or more scripts stored in a markup language document), stored in a single file dedicated to the program in question, or stored in multiple coordinated files (such as files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one or more computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0181] A system of one or more computers configured to perform particular operations or actions means that the system has software, firmware, hardware, or a combination of them installed on it, which in operation causes the system to perform the operations or actions. One or more computer programs configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operations or actions.
[0182] As used in this specification, an "engine" or "software engine" refers to a software-implemented input / output system that provides an output different from the input. An engine can be an encoded functional block, such as a library, platform, software development kit ("SDK"), or object. Each engine can be implemented on any suitable type of computing device including one or more processors and a computer-readable medium, such as a server, mobile phone, tablet computer, notebook computer, music player, e-book reader, laptop or desktop computer, PDA, smart phone, or other fixed or portable device. Additionally, two or more of the engines can be implemented on the same computing device or on different computing devices.
[0183] The processes and logical flows described in this specification can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, or by a combination of, special purpose logic circuitry, such as an FPGA or ASIC, or special purpose logic circuitry and one or more programmed computers.
[0184] A computer suitable for executing a computer program can be based on a general or special purpose microprocessor or both, or any other type of central processing unit. In general, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for executing or running instructions and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. In general, a computer will also include one or more mass storage devices for storing data (e.g., magnetic disks, magneto-optical disks, or optical disks), or be operatively coupled to one or more mass storage devices to receive data therefrom or transfer data thereto or both. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, such as a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few examples.
[0185] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example: semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0186] To provide interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or an LCD (liquid crystal display)) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse, a trackball, or a touch-sensitive display or other surface) by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, voice, or tactile input. In addition, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user; for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser. Moreover, the computer can interact with the user by sending a text message or other form of message to a personal device (e.g., a smart phone running a messaging application) and receiving a response message from the user.
[0187] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface or a web browser or an app through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN) and a wide area network (WAN), such as the Internet.
[0188] The computing system can include a client and a server. The client and the server are typically located far apart from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs that run on their respective computers and have a client-server relationship with each other. In some embodiments, the server, for example, transmits data to a user device, such as an HTML page, for displaying data to a user who interacts with the device acting as a client and receiving user input from the user. Data generated at the user device, such as the result of a user interaction, can be received at the server from the device.
[0189] In addition to the above embodiments, the following embodiments are also innovative:
[0190] Embodiment 1 is a method, including:
[0191] Receiving, by an online execution system configured to control a robot, a skill template to be trained for the robot to perform a specific skill having a plurality of subtasks;
[0192] Identifying one or more demonstration subtasks defined by the skill template, where each demonstration subtask is an action to be refined using local demonstration data;
[0193] Generating, by the online execution system, a corresponding set of local demonstration data for each of the one or more demonstration subtasks;
[0194] Uploading, by the online execution system, the set of local demonstration data to a cloud-based training system;
[0195] Generating, by the cloud-based training system, corresponding trained model parameters for each set of local demonstration data;
[0196] Receiving, by the online execution system, the trained model parameters generated by the cloud-based training system; and
[0197] Performing the skill template using the trained model parameters generated by the cloud-based training system.
[0198] Example 2 is the method of Example 1, where generating the corresponding set of local demonstration data includes generating a task state representation for each of a plurality of time points.
[0199] Example 3 is the method of Example 2, where the task state representations each represent an output generated by observing a corresponding sensor of the robot.
[0200] Example 4 is the method of any one of Examples 1-3, where the online execution system and the robot are located in the same facility, and where the cloud-based training system can only be accessed via the Internet.
[0201] Example 5 is the method of any one of Examples 1-4, further comprising:
[0202] Receiving, by the online execution system, a basic control strategy for each of the one or more demonstration subtasks from the cloud-based training system.
[0203] Example 6 is the method of Example 5, where performing the skill template includes generating corrective actions by the online execution system using the trained model parameters generated by the cloud-based training system.
[0204] Example 7 is the method of Example 6, further comprising adding the corrective actions to basic actions generated from the basic control strategy received from the cloud-based training system.
[0205] Example 8 is a system, comprising: one or more computers and one or more storage devices storing instructions which, when executed by the one or more computers, cause the one or more computers to perform the method of any one of Examples 1 to 7.
[0206] Example 9 is a computer storage medium encoded with a computer program, the program comprising instructions which, when executed by a data processing apparatus, are operative to cause the data processing apparatus to perform the method of any one of Examples 1 to 7.
[0207] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what is claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination within a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although the above features may be described as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination can be deleted from the combination, and the claimed combination can be directed to a sub-combination or variation of a sub-combination.
[0208] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to obtain a desired result. In some cases, multitasking and parallel processing may be advantageous. Also, the separation of various system modules and components in the above embodiments should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated in a single software product or packaged into multiple software products.
[0209] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the acts recited in the claims can be performed in a different order and still obtain a desired result. As one example, the processes described in the figures do not necessarily require the particular order or sequential order shown to obtain a desired result. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for executing a skill template of a robot, the method comprising: Receiving, by an online execution system configured to control the robot, the skill template to be trained from a cloud-based training system to enable the robot to execute a specific skill having a plurality of subtasks; Identifying, by the online execution system, one or more demonstration subtasks defined by the skill template, wherein each demonstration subtask is an action to be refined using local demonstration data; Generating, by the online execution system, a corresponding set of local demonstration data for each of the one or more demonstration subtasks; Uploading, by the online execution system, the set of local demonstration data to the cloud-based training system; Generating, by the cloud-based training system, corresponding trained model parameters for each set of local demonstration data, wherein the cloud-based training system includes more computing power than the online execution system; Receiving, by the online execution system, the trained model parameters generated by the cloud-based training system from the cloud-based training system; And Executing, by the online execution system, the skill template using the trained model parameters generated by the cloud-based training system.
2. The method according to claim 1, wherein generating the corresponding set of local demonstration data includes generating a task state representation for each of a plurality of time points.
3. The method according to claim 2, wherein the task state representations each represent an output generated by observing a corresponding sensor of the robot.
4. The method according to claim 1, wherein the online execution system and the robot are located in the same facility, and wherein the cloud-based training system can only be accessed via the Internet.
5. The method according to claim 1, further comprising: Receiving, by the online execution system, a basic control strategy for each of the one or more demonstration subtasks from the cloud-based training system.
6. The method according to claim 5, wherein executing the skill template includes generating a corrective action by the online execution system using the trained model parameters generated by the cloud-based training system.
7. The method according to claim 6, further comprising adding the corrective action to a basic action generated by the basic control strategy received from the cloud-based training system.
8. A system for executing a skill template of a robot, comprising: One or more computers and one or more storage devices storing instructions, the instructions when executed by the one or more computers cause the one or more computers to perform operations including: Receiving, by an online execution system configured to control the robot, the skill template to be trained from a cloud-based training system to enable the robot to execute a specific skill having a plurality of subtasks; Identifying, by the online execution system, one or more demonstration subtasks defined by the skill template, wherein each demonstration subtask is an action to be refined using local demonstration data; The online execution system generates a corresponding local demonstration data set for each of the one or more demonstration subtasks; The online execution system uploads the local demonstration data set to the cloud-based training system; The cloud-based training system generates corresponding trained model parameters for each local demonstration data set, wherein the cloud-based training system includes more computing power than the online execution system; The online execution system receives the trained model parameters generated by the cloud-based training system from the cloud-based training system; and The online execution system executes the skill template using the trained model parameters generated by the cloud-based training system.
9. The system of claim 8, wherein generating the corresponding local demonstration data set includes generating a task state representation for each of a plurality of time points.
10. The system of claim 9, wherein the task state representations each represent an output generated by observing a respective sensor of the robot.
11. The system of claim 8, wherein the online execution system and the robot are located in the same facility, and wherein the cloud-based training system is only accessible via the Internet.
12. The system of claim 8, wherein the operation further includes: The online execution system receives a basic control strategy for each of the one or more demonstration subtasks from the cloud-based training system.
13. The system of claim 12, wherein executing the skill template includes the online execution system generating a corrective action using the trained model parameters generated by the cloud-based training system.
14. The system of claim 13, further comprising adding the corrective action to a basic action generated from the basic control strategy received from the cloud-based training system.
15. One or more non-transitory computer storage media encoded with computer program instructions that, when executed by one or more computers, cause the one or more computers to perform operations including the following: An online execution system configured to control a robot receives a skill template to be trained from a cloud-based training system to cause the robot to perform a specific skill having a plurality of subtasks; The online execution system identifies one or more demonstration subtasks defined by the skill template, wherein each demonstration subtask is an action to be refined using local demonstration data; The online execution system generates a corresponding local demonstration data set for each of the one or more demonstration subtasks; The online execution system uploads the local demonstration data set to the cloud-based training system; The cloud-based training system generates corresponding trained model parameters for each local demonstration data set, wherein the cloud-based training system includes more computing power than the online execution system; The online execution system receives the trained model parameters generated by the cloud-based training system from the cloud-based training system; and The skill template is executed by the online execution system using the trained model parameters generated by the cloud-based training system.
16. The non-transitory computer storage medium of claim 15, wherein generating the respective local demonstration data sets includes generating a task state representation for each of a plurality of time points.
17. The non-transitory computer storage medium of claim 16, wherein the task state representations each represent an output generated by observing a respective sensor of the robot.
18. The non-transitory computer storage medium of claim 15, wherein the online execution system and the robot are located in the same facility, and wherein the cloud-based training system is only accessible via the Internet.
19. The non-transitory computer storage medium of claim 15, further comprising: Receiving, by the online execution system from the cloud-based training system, a basic control strategy for each of the one or more demonstration subtasks.
20. The non-transitory computer storage medium of claim 19, wherein executing the skill template includes generating, by the online execution system, a corrective action using the trained model parameters generated by the cloud-based training system.
Citation Information
Patent Citations
Intelligent physical therapy robot system and operation method therefor
WO2018014824A1
Systems, apparatus, and methods for robotic learning and execution of skills
WO2020047120A1