System and method for learning sequences in robotic tasks to promote to new tasks
By learning techniques from the demonstration, decomposing long-distance tasks into meaningful sequences, and generating and executing these sequences using a planning module based on graph search and dynamic motion primitives, the difficulty of designing long-distance task controllers is solved, and data efficient learning and execution is achieved.
Patent Information
- Application Number
- CN202380073135.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2023-07-07
- Publication Date
- 2025-05-27
AI Technical Summary
There are difficulties in designing sequential task controllers for long-term tasks, especially because the search space is too large, making it difficult for optimization-based technologies to find solutions, reinforcement learning requires a lot of data and the reward engineering is complex.
Using Learn from Demo (LfD) technology, these sequences are generated and executed by decomposing long-term tasks into meaningful sequences and using graph search-based planning modules and dynamic motion primitives (DMPs).
It realizes efficient data learning and executes long-term sequential tasks, reduces programming burden, and improves the robot's autonomous execution ability in complex tasks.
Smart Images

Figure CN120051361A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to learning sequences in sequential tasks, and more particularly to methods and apparatus for learning sequences in robotic tasks in order to perform new robotic tasks using these learned sequences and failures observed during demonstrations. Background Art
[0002] The field of machine learning and artificial intelligence has witnessed tremendous improvements and achievements in the areas of computer vision and natural language processing. However, when used for robotic applications, these algorithms suffer in terms of data efficiency and thus become impractical for use for many robotic applications. Learning from Demonstration (LfD) is a data-efficient learning technique where a robot can learn to perform a task by first recording several demonstrations of the task and then recreating these demonstrations using an appropriate machine learning model.
[0003] In LfD, the robot is provided with one or more demonstrations of the desired task. Demonstrations can be provided for known tasks by a person or by a programmed controller. In the case where the demonstrations are provided by a person, the person can provide the demonstrations directly on the robot or by performing the task himself. In the latter case, a motion capture system or a vision system consisting of one or more cameras can be used to record human movements. On the other hand, if the person provides the demonstrations directly on the robot, the person can provide the demonstrations by using kinesthetic teaching to move the robot or by remote operation using appropriate equipment. In all these cases, the movements of the robot and the objects manipulated by the robot are observed and recorded. This data is then used to learn or represent the movements of the robot when performing the task shown to the robot.
[0004] LfD techniques are widely used to reduce the programming of robots and allow unskilled workers to demonstrate tasks on the robot. The robot can then use standard LfD techniques to recreate the task and perform it autonomously without explicit human programming. The LfD representation learned to perform a task is called a skill. However, many useful robotic tasks are sequential in nature. For example, consider the task of assembling an electronic item. Such a task would require that the robot can put all the different pieces of the electronic item together in the desired order to assemble the complete item. It is also expected that the robot will be able to use the learned skills to assemble any new electronic item using the same subset of operations in a specific order.
[0005] In order to learn, LfD techniques to perform these long-duration tasks autonomously, two key elements are required. First, the demonstrations must be sequenced into different sequences or subtasks when performing the full task. These individual sequences or subtasks can then be learned using a suitable machine learning model while being parameterized by some parameters of the task. These models of learned subtasks are called skills. Second, the robot should optimize the sequence of these skills based on the new task the robot needs to perform. By using all or a subset of the skills learned in the first part, the skills learned in a specific unknown sequence can be used to perform a new task.
[0006] Therefore, methods are needed that can automatically decompose long demonstrations into meaningful sequences and then optimally synthesize these sequences to perform new tasks. Summary of the invention
[0007] Some embodiments of the proposed disclosure are based on the recognition that it is difficult to design controllers for long-duration, sequential tasks. This is mainly because the search space of feasible solutions is too large, so optimization-based techniques cannot find solutions. Reinforcement learning (RL) may find solutions, however this will require carefully designed rewards and require a large amount of data to be able to guide the RL agent to learn a solution. This technique is very inefficient because it requires an excessive amount of data. In addition, reward engineering for complex tasks is a very difficult problem.
[0008] Some embodiments are based on the recognition that learning from demonstration (LfD) may be useful for learning efficient controllers for performing long-term, sequential tasks. The reason is that the robot can get ideas about how to perform the task from an expert, either a human or a controller. The robot can use appropriate learning methods (e.g., dynamic motion primitives, SEDS, etc.). However, when using LfD for long-term tasks for robots, there are challenges that need to be addressed. For example, if the task contains multiple steps that need to be completed for the task to be successful, it is very difficult to learn the entire task as a single motor skill. Therefore, it is important that the robot recognizes sequences in long-term tasks that have been demonstrated to the robot.
[0009] It is an object of some embodiments to provide a system and method for identifying sequences in a demonstration to perform a long-duration, sequential task.Some embodiments of the invention are based on the recognition that the segmentation of a task trajectory will depend on the feature representation of the demonstrated trajectory.
[0010] It is an object of some embodiments to provide a system and method that can detect suitable features that can be used for sequence recognition in a demonstrated trajectory. The problem is similar to feature recognition or feature selection of collected demonstrations that can be applied to a robot so that it can then be used for trajectory segmentation. The method can allow for better segmentation of demonstration trajectories.
[0011] Additionally or alternatively, some embodiments aim to provide a system and method that can detect appropriate features from the data to correctly identify different sequences and changes between different sequences. Additionally or alternatively, some embodiments aim to provide a system and method that can fit a machine learning model to each identified sequence parameterized by some parameters of the task. Additionally or alternatively, some embodiments aim to provide a system and method that can use information from demonstration attempts that resulted in failures to provide robustness to sequence detection.
[0012] Additionally or alternatively, some embodiments aim to provide a system and method that can generate an optimal sequence for executing a subset of these sequences in order to perform a new task presented to a robot. Additionally or alternatively, some embodiments aim to provide a system and method that implements a learned sequence of tasks in a feedback manner using an object-state detection framework.
[0013] According to some embodiments of the present invention, a robot controller is provided for generating a sequence of motion primitives for a sequential task of a robot with a manipulator. The robot controller may include: at least one control processor; and a memory circuit, the memory circuit storing a dictionary including motion primitives, a pre-trained learning module, and a planning module based on a graph search, the planning module based on a graph search having instructions stored thereon, the instructions when executed by at least one control processor, causing the robot controller to perform the following steps: obtaining a planning task provided by an interface device operated by a user, wherein the planning task is represented by an initial state and a target state about an object; generating a planning graph by searching a feasible path of the object for a new task using a planning module based on a graph search and selecting motion primitives in a dictionary in a pre-trained learning module, wherein the pre-trained learning module has been trained based on a demonstration task; parameterizing the feasible path represented by the motion primitive into a dynamic motion primitive (DMP) using the initial state and the target state; and implementing the parameterized feasible path as a trajectory using the manipulator of the robot according to the selected motion primitive by tracking and following the parameterization for the planning task.
[0014] In addition, some embodiments may provide a robot controller for learning a sequence of motion primitives for sequential tasks of a robot with a manipulator. In this case, the robot controller may include: at least one control processor; and a memory circuit storing a dictionary including motion primitives, and a learning module having instructions stored thereon, which instructions, when executed by at least one control processor, cause the robot controller to perform the following steps: collect demonstration data from trajectories acquired via a motion sensor, the motion sensor being configured to measure the trajectory of an object when the object is manipulated by an interface device operated by a user according to a demonstration task, wherein each trajectory corresponds to each demonstration task, wherein each demonstration task is represented by an initial state and a target state with respect to each object, wherein the collection continues until the user stops the demonstration task; for each demonstration task, segment the demonstration data into motion primitives by dividing the trajectory into primitive trajectories; and update the dictionary using the motion primitives based on the collected demonstration data.
[0015] In addition, according to some embodiments of the present invention, a robot controller is provided for generating a sequence of motion primitives for a sequential task of a robot having a manipulator. The robot controller may include: at least one control processor; and a memory circuit storing a dictionary including motion primitives, a pre-trained learning module, and a graphic search-based planning module storing instructions, which, when executed by at least one control processor, cause the robot controller to perform the following steps: via an interface controller, obtaining demonstration data of one or more demonstration tasks provided by an interface device operated by a user for a planned task, wherein the planned task is represented by an initial state and a target state of at least one object to be manipulated; by selecting features from the demonstration data based on a feature selection method and using a segmentation metric to separate each demonstration data The invention relates to a method for performing a multi-functional robot robot in a pre-trained learning module, wherein the pre-trained learning module is trained based on the collected demonstration data of the training demonstration task and is divided into a plurality of segments, wherein each segment of the plurality of segments represents a subtask; generating a planning graph by searching a feasible path of at least one object for the planning task using a planning module based on a graph search and selecting motion primitives from a dictionary in a pre-trained learning module, wherein the pre-trained learning module is trained based on the collected demonstration data of the training demonstration task; parameterizing the feasible path represented by the motion primitive into a dynamic motion primitive (DMP) using an initial state and a target state; and implementing the parameterized feasible path as a trajectory according to the selected motion primitive by tracking and following the parameterized feasible path of the planning task using a manipulator of the robot.
[0016] By way of non-limiting examples of exemplary embodiments of the present disclosure, the present disclosure is further described in the following detailed description with reference to the indicated multiple drawings, wherein like reference numerals represent similar parts throughout the several views of the drawings. The drawings shown are not necessarily drawn to scale, with emphasis generally being placed upon illustrating the principles of the presently disclosed embodiments.
[0017] Although the drawings indicated above illustrate the embodiments of the present disclosure, other embodiments are also conceivable, as noted in the discussion. The present disclosure presents illustrative embodiments by way of representation and not limitation. Those skilled in the art can devise many other modifications and embodiments that fall within the scope and spirit of the principles of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] [ Figure 1 ]
[0019] Figure 1 A schematic representation of a general learning paradigm according to an embodiment of the present invention is shown.
[0020] [ Figure 2A ]
[0021] Figure 2A A schematic diagram of a robotic system according to an embodiment of the invention is shown, wherein several different interfaces are used to control the robotic manipulator.
[0022] [ Figure 2B ]
[0023] Figure 2B A diagram illustrating components of a controller connected to an interface device according to an embodiment of the present invention is shown.
[0024] [ Figure 3 ]
[0025] Figure 3 A schematic representation showing how a specific task according to an embodiment of the present invention may be composed of different subtasks, which are different sections extracted by the proposed method during the learning process.
[0026] [ Figure 4 ]
[0027] Figure 4 Shown is a schematic representation of a state for a block stacking task and a world coordinate system for measuring the state according to an embodiment of the present invention.
[0028] [ Figure 5 ]
[0029] Figure 5 A schematic diagram showing metrics and corresponding examples for segmenting a task into different tracks according to an embodiment of the present invention, where different features are required to segment the tracks.
[0030] [ Figure 6 ]
[0031] Figure 6A schematic diagram of a Dynamic Motion Primitive (DMP) in the proposed work for learning different skill representations according to an embodiment of the present invention is shown.
[0032] [ Figure 7 ]
[0033] Figure 7 is a schematic diagram showing skills learned from an original demonstration task that is not shown during the demonstration that needs to be achieved for the desired task.
[0034] [ Figure 8 ]
[0035] Figure 8 A planning graph used in some embodiments of the invention is shown, where the initial node is the goal state of the task and additional nodes are added.
[0036] [ Fig. 9 ]
[0037] Fig. 9 A flow chart is shown indicating sequential steps implemented in some embodiments of the present invention.
[0038] [ Fig.10 ]
[0039] Fig.10 A schematic diagram showing the segmentation of a peg-in-hole demonstration task consisting of multiple steps for precise operation is shown in accordance with an embodiment of the present invention. DETAILED DESCRIPTION
[0040] In the following description, for the purpose of explanation, many specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it is apparent to those skilled in the art that the present disclosure can be practiced without these specific details. In other cases, only in order to avoid obscuring the present disclosure, devices and methods are shown in block diagram form.
[0041] As used in this specification and claims, the terms "for example," "for example," and "such as," and the verbs "include," "have," "comprise," and other verb forms thereof, when used in conjunction with a list of one or more components or other items, are each interpreted as open-ended, meaning that the list is not to be viewed as excluding other additional components or items. The term "based on" means based, at least in part, on. In addition, it should be understood that the wording and terminology employed herein are for descriptive purposes and should not be considered limiting. Any headings used in this specification are for convenience only and have no legal or limiting effect.
[0042] In robotics, designing controllers for long-duration manipulation tasks remains very challenging. There are several reasons that make the task challenging. First, it is difficult to find solutions for very long-duration control using model-based techniques or model-free RL based methods. Second, the success of the entire task depends on the success of each individual task. Therefore, these problems require careful formulation, where the complete task can be broken down into smaller sub-problems and then ensuring that each sub-problem can be completed reliably. It is also expected that in order to reduce the effort of designing these controllers, suitable learning-based methods should be used, which can be trained in a data-efficient manner and can be generalized to new tasks. The present disclosure proposes systems and methods that can be used to reduce the programming burden for performing long-duration tasks.
[0043] Reinforcement learning (RL) based methods have found great success in many robotic manipulation tasks, but they suffer from data requirements during training and are difficult to train for long-duration tasks. Therefore, the use of RL is limited to short-duration tasks, where the robot can be trained with dense rewards, otherwise the method becomes very data intensive. Learning from Demonstration (LfD) provides an alternative learning-based approach that can utilize expert or human demonstrations to learn motor skills for different tasks. The systems and methods presented in this disclosure are motivated by this requirement, where the proposed method is data efficient and reduces the effort of expert programming.
[0044] Figure 1 A schematic diagram 110 illustrating a block stacking task is shown, which is a sequential task that requires a robot to sequentially perform the placement of individual blocks based on the positions of other blocks in a scene. In this task, the robot is presented with blocks in its workspace so that the robot can observe the positions of the blocks. The basic idea of the proposed learning is shown in 100. An expert user provides multiple demonstrations for performing a task. These demonstrations can then be segmented into separate segments using appropriate feature detection. These separate segments can then be represented as dynamic motion primitives (DMPs). The robot can then use these DMPs for new instances of the problem for autonomous execution.
[0045] Some embodiments are based on the recognition that LfD methods provide a data-efficient alternative to RL-based methods for designing learning-based controllers for long-duration, multi-stage manipulation tasks. A robotic system can be equipped with a system for providing demonstrations to the robot to perform these tasks. The system can include at least one interface for moving the robot by an expert human. Some examples of such interfaces can be a 3-axis joystick, a space mouse, a virtual or augmented reality system. These interfaces can be used to move the robot remotely. Alternatively, the expert human can also demonstrate the task on the robot using a kinesthetic controller on the robot, where the robot can be moved directly by applying forces on the robot arm.
[0046] Alternatively, demonstration data may also be collected in a simulation by creating a simulated environment similar to the physical environment and collecting demonstration data by moving the robot in the simulated environment using a similar interface, such as a joystick or a virtual reality or augmented reality interface.
[0047] Figure 2A A robot system 200 is shown, which includes a controller (robot controller) 205 and a robot manipulator 210, which is configured to stack blocks 220 into a desired shape. The controller 205 of the robot system 200 is connected to an interface device 230, which is configured to be operated by a user / operator, who demonstrates the task of the robot system 200 by using the interface device 230. The robot system 200 includes a robot manipulator 210 having an actuator 2103, a sensor 2101 including a visual sensor (3D sensor) 2102 arranged on the robot manipulator 210. The motion sensor 2101 may include an acceleration sensor, a position sensor, a torque sensor, and a force sensor. The signal measured by the motion sensor 2101 is acquired by the controller 205 via an interface controller 2110B, which includes an analog / digital (A / D) signal converter and a digital / analog (D / A) signal converter. In this case, the interface device 230 includes a network interface (not shown) configured to connect to the controller 205 of the robot system 200 via the communication network 215. In some cases, the communication network 215 can be a wired or wireless network. The interface device 230 is configured to control the robot manipulator 210 via the controller 205 so that a demonstration trajectory operated by a user is provided to learn a sequence of robot tasks using signals from the sensor 2010 and the actuator 2103. For example, the interface device 230 can be any type of joystick 240, 250 or 260 configured to be used / operated by a user / operator. The operator can also use a virtual reality game engine controller to move the robot during these demonstrations. Note that these figures are not exhaustive and that humans can use other interaction modes to demonstrate tasks corresponding to the demonstration trajectory.
[0048] The kind of tasks of interest are long-duration tasks that are a combination of several subtasks. It is assumed that an expert human provides several demonstrations of such long-duration tasks. Note that during the demonstration, observations from different sensors available to the robotic system are recorded, which may include encoders on the robotic arm, a vision system for tracking objects in the robot's working environment, and force sensors for observing the forces experienced by the robot's end effector during the task demonstration. There may be other sensors of other sensing modalities that the robotic system may be equipped with, such as tactile sensors, which may provide more detailed information on the contact forces and torques at the fingers of the gripper during the demonstrated manipulation task. Therefore, the demonstration trajectory is represented by a sequence of sensor trajectories collected by the robotic system during the task demonstration. At any moment, the state of the robotic system is represented as the set of the pose of the end effector (or the gripper end of the robot) and the poses of all objects in the robot's workspace.
[0049] Figure 2A Controller 205 is shown configured to control manipulator 210 to demonstrate a sequence of steps used by manipulator arm 2155 of robotic system 200 to perform a desired manipulation task in accordance with an embodiment of the present invention. In some cases, robotic system 200 may be referred to as a robot.
[0050] The robot system 200 includes a manipulator 210 and a force sensor 2101 arranged on the manipulator 210, and a vision system 2102 (at least one camera). The force sensor 2101 (which may be referred to as at least one force sensor) is configured to detect the force implemented by the manipulator 210 on the object at the contact point between the object and the manipulator. The vision system 2102 may be at least one camera or multiple cameras, a depth camera, a distance camera, etc. The vision system 2102 is arranged at a position so that the vision system 2102 can observe the state of the object representing the positional relationship between the object, the desktop (not shown), and the additional contact surface. The vision system 2102 is configured to estimate the posture of the object on the desktop using the additional contact surface in the environment of the robot system 200.
[0051] The vision system 2102 is configured to detect and estimate the pose of an object to be manipulated on the tabletop. The controller 205 is configured to determine whether a component needs to be redirected before the component can be used for a desired task (e.g., assembly). The controller 205 is configured to calculate a series of control forces applied to the object using a two-level optimization algorithm. The robot 200 applies a series of control forces (a series of contact forces) to the object against the external contact surface according to the control signal sent from the interface device 230.
[0052] In addition, the controller 205 is configured to acquire simulation data and learning data via the communication network 215. The simulation data and learning data generated in the computer (simulation computer system) 2500 are configured to be used in the robot system 200. The collected simulation data and learning data are transmitted to the controller 205 via the communication network 215.
[0053] The controller 205 is configured to generate and send control data including instructions regarding the calculated control force sequence to a low-level robot controller (e.g., an actuator controller of a manipulator), such that the instructions cause the manipulator to apply the calculated control force sequence (contact force) on the tabletop. The robot 200 is configured to grasp and redirect the parts so that they can then be used for the desired task (assembly or packaging) on the tabletop.
[0054] Figure 2B A robot system (robot) 200 is shown for manipulating an object on a table (not shown) according to a trajectory generated by the proposed robust trajectory optimization problem according to an embodiment of the present invention. A robot control system 2100 is configured to control an actuator system 2103 of a robot 2150. The robot control system 200 may be referred to as a control system of the robot or a robot controller.
[0055] The robot control system 200 may include an interface controller 2110B, a control processor 2120 (or at least one control processor) and a memory circuit 2130B. The memory circuit may be referred to as a memory unit or a memory module, which may include one or more static random access memories (SRAMs), one or more dynamic random access memories (DRAMs), one or more read-only memories (ROMs), or a combination thereof. The memory circuit 2130B is configured to store a computer-implemented method, the method including a learning from demonstration (LfD) module and a graph-search-based planning module, the graph-search-based planning module may generate a feasible sequence of LfD skills (using the LfD module) to generate a feasible plan for a new task. The processor 2120 may be one or more processor units, and the memory circuit 2130B may be a memory device, a data storage device, etc. The interface controller (robot interface controller) 2110B may be an interface circuit that may include analog / digital (A / D) and digital / analog (D / A) converters to communicate signals / data with the sensor 2101 including the force sensor and the visual sensor 2102 and the motion controller 2150B of the robot 200. In addition, the interface controller 2110B may include a memory to store data to be used by the A / D or D / A converter. The sensor 2101 is arranged at a joint of the robot (robot arm or manipulator) or a picking object mechanism (e.g., finger, end effector) to measure the contact state with the robot. The visual sensor 2102 may be arranged at any position that provides a viewpoint for observing / measuring the object state representing the positional relationship between the object, the desktop, and the additional contact surface.
[0056] The controller 205 includes an actuator controller (device / circuit) 2150B, which includes a strategy unit 2151B to generate action parameters to control the robot 200, which controls the manipulator 210, the processing mechanism or a combination of the arm 2103 including the processing mechanisms 2103-1, 2103-2, 2103-3 and 2103-#N according to the number of joints or processing fingers. For example, the sensor 2101 may include an accelerometer, an angle sensor, a force sensor or a tactile sensor for measuring the position of the object and the force during the external period. For example, complementary constraints can be used to represent the interaction between the object and the robot arm of the robot system to capture the contact state between the object and the robot arm of the robot system. In other words, the interaction is based on the contact state, which is represented by the relationship between the sliding speed of the object on the table and the friction between the object and the table when the object is moved by the robot arm.
[0057] The interface controller 2110B is also connected to a sensor 2101 that measures / acquires the motion state of the robot mounted on the robot. The motion sensor 2101 may be configured to measure the sequence of forces applied to the robot and the position of the sensor arranged on the robot. The position is determined by Fig.10 The world coordinate system 1010 is represented in .
[0058] In some cases, when the actuator is an electric motor, the actuator controller 2150B can control the angle of the robot arm or the individual electric motors that handle the object through the handling mechanism. In some cases, the actuator controller 2150B can control the rotation of the individual motors arranged in the arm to smoothly accelerate or safely decelerate the movement of the robot in response to the strategy parameters generated from the computer-implemented method 2000 for learning the robot task sequence stored in the memory circuit 2130B, and the memory circuit includes a learning module 2101B for LfD and a graphic search-based planning module 2140B for control signals. In addition, depending on the design of the object handling mechanism, the actuator controller 2150B can control the length of the actuator in response to the strategy parameters according to the instructions generated by the computer-implemented method 2000 stored in the memory circuit 2130B.
[0059] The controller 205 is connected to an imaging device or visual sensor 2102 that provides an RGBD image. In another embodiment, the visual sensor 2102 may include a depth camera, a thermal camera, an RGB camera, a computer, a scanner, a mobile device, a webcam, or any combination thereof. In some cases, the visual sensor 2102 may be referred to as a visual system. The signals from the visual sensor 2102 are processed and used to classify, identify, or measure the state of the object 220.
[0060] Note that no labels are available for the different segments of the demonstration trajectory. The different segments represent different (sub)tasks that need to be performed sequentially, and only by completing these subtasks can the success of the entire long-term task composed of these short-term tasks be ensured. Note that each of these subtasks needs to be robustly implemented to be able to complete the entire long-term task. Figure 3 A possible sequence of subtasks that may be implemented to complete the block stacking task for a target configuration is shown.
[0061] For example, there are five subtasks in the block stacking task using the interface device 230 operated by the user, such as Figure 3As shown. In this case, the user operates the robotic manipulator using the interface device 230 as follows. In a first step 310, the robot grabs object B. In step 320, the robotic manipulator places object B next to object A in step 320. In the next step, in step 330, the robotic manipulator pushes object B close to object A so that objects A and B are in contact. Then, in step 340, the robotic manipulator grabs object C. Finally, in step 350, the stacking is completed by placing object C on objects A and B. Note that this is one possible sequence of operations that can be demonstrated during the learning phase of the robotic task. The user can demonstrate any other feasible sub-task sequence that can be used to successfully complete the task. However, the user / operator needs to provide the same demonstration multiple times. During these multiple demonstrations, the initial states of the blocks and the robot may be different.
[0062] The task can be demonstrated directly on the robot using teleoperation, or the robot can be moved using a kinesthetic controller 205 configured to move the robotic manipulator 210. For teleoperation of the robot, a human expert can move the robot 210 during a task using one of several possible joystick interfaces. Figure 2A Several different joystick interfaces are shown that can be used to move the robot during the demonstration. For example, a person can use one of the joysticks 240 or 250 to control the movement of the robot during the demonstration. A person can also use a virtual reality controller 260 with a virtual reality setting to demonstrate a trajectory for moving the robot 210 (collecting data physically or in virtual reality). These remote operation interfaces can be used to move the robot 210 to demonstrate a desired task, such as stacking a set of blocks in a certain way 220. The desired task can be referred to as a planned task. As each task is demonstrated, data about the demonstrated trajectory of each task is acquired into the memory circuit 2130B via a sensing system including a sensor 2101 and a visual sensor 2102. In another option, a person can also try to move the robot 201 by directly moving the robot arm 210 using a kinesthetic mode that may be available on the robot.
[0063] Figure 4 Schematic representation of the state and world coordinate system for a block stacking task and a measurement state according to an embodiment of the present invention is shown. Figure 4 In the case of the block stacking task shown, the state 430 of the system is a concatenation of the states of the manipulator 410 and the block 420 measured in a fixed coordinate system 440. The state of the manipulator is represented by the state x of the end effector ee Similarly, the states of blocks A, B, C, and D are represented as x A , x B ,x C , x D, which may represent the pose of the block in the fixed coordinate system 440 .
[0064] Some embodiments of the present disclosure are based on implementations without any labels for the demonstrated trajectories, and a metric must be designed that can be used to consistently segment / partition the demonstrated trajectories into different subtasks represented by the segmented trajectories. However, in order to determine the different segments of the demonstrated trajectories, a metric that can be used to segment / partition the demonstrated trajectories is designed. Note that both the number of segments and the metric used for trajectory segmentation are unknown. Therefore, in order to allow segmentation of the demonstrated trajectories, feature extraction is first performed, and then the segmentation of the trajectories into different components is performed using a metric that exploits these features.
[0065] For feature extraction in the current work, the pose data of the robot and the objects in the reference frames of different objects are simply transformed. This can be achieved by applying the correct transformation to transform the observations of all the data in different coordinate frames and use them as features.
[0066] A coordinate system is used to define a coordinate system that the robot can use to measure its own position and to know the position of objects in the robot's working environment. Features are functions of measurements or observations used to train machine learning models. Some embodiments of the present disclosure are based on the recognition that different demonstration trajectories can be transformed in a variety of different coordinate systems that can be attached to different objects in the robot's working environment. Feature selection is performed using a user-defined function or cost function that represents the purpose of feature selection. In the case of supervised learning, this can be performed using a metric such as maximum classification accuracy. However, in the present disclosure, there are no labels and feature selection is performed using an unsupervised learning cost function. This can be a function consisting of features and a maximized segmentation metric (in Figure 5 convex sum of the number of segments obtained (described in ).
[0067] Figure 5 Indicates the metric proposed in this work for segmenting demonstration trajectories. Figure 5 A metric for trajectory segmentation is shown in 520. An example of how the metric 520 is used to segment a trajectory is shown in 510. Figure 5 Also shown are different features that are useful for segmentation of the demonstration trajectory 511. As shown by metric 520, the maximum value among the different features is used to segment / partition the demonstration trajectory.
[0068] exist Figure 5 , different features of transformations in the reference systems of different blocks in the scene are shown. For example, the object A coordinate system represents data in the reference system of object (or block) A in 310. Similarly, the object B coordinate system and the object C coordinate system represent the reference systems of objects B and C. Measurements can be made directly in these coordinate systems, or the measurements can be converted after collection using transformations between the global reference system and the individual object reference systems.
[0069] Once the demonstration trajectory is segmented into different parts (the primitive trajectories correspond to dynamic motion primitives) using the metrics presented in 520, a representative motion model is fit in each segmented trajectory.
[0070] In the present disclosure, a dynamic motion primitive (Dynamic Motion Primitive) or DMP is used to represent each of the segmented trajectories. Figure 6 A schematic diagram of the Dynamic Motion Primitives (DMPs) used in the proposed work to learn different skill representations according to an embodiment of the present invention is shown. For completeness, they will be described here. The DMP is a set of two dynamic systems described by ordinary differential equations - point attractor dynamics & forcing terms.
[0071] To remove explicit time dependencies, they use a canonical system to track progress through learned behaviors:
[0072] τs · =-α_s s
[0073] Here, s=1 (and α_s>0) at the start of DMP execution and τ>0 specifies the rate of progress through the DMP.
[0074] To capture the attraction behavior of the point attractor dynamics and the forcing term, the DMP 610 uses a spring-damper system 612 (transformation system) with the addition of a nonlinear forcing term 611. Writing the DMP equations as a system of coupled first-order ordinary differential equations (ODEs) yields:
[0075]
[0076] τy · =z
[0077] where g represents the target pose. The forcing function has adjustable parameters that are learned from the motion primitive data and weight the contributions of the basis functions. The forcing term is defined as the radial basis function 620:
[0078]
[0079] ψ i (s) = exp{(-h i (sc i ) 2 )}
[0080] Among them, h i and c i denote respectively the width and center of the Gaussian basis function 630. The forcing term is learned from the demonstrations by solving a local weighted regression to fit the demonstration data given by the expert.
[0081] Using the segmentation of the trajectory into individual components, and fitting each of the individual segments, any expert demonstration of the task can be reproduced. However, if the desired task is different from the demonstrated task, the described method is not sufficient to perform that task. Figure 7 A scenario is shown in which a robot is demonstrated to perform a task 710 using an interface device 701. The goal of this task 711 is very different from the goal 721 of the desired task 720. In these cases, an algorithm is needed that can help sequence the learned subtasks so that the robot can successfully perform the desired task 720.
[0082] Some embodiments of the present disclosure are based on the recognition that a graph search based planning algorithm can be used to help plan tasks that were not demonstrated during training of a robot. Figure 8 A planning graph used in some embodiments of the present invention is shown, wherein the initial node 802 is the target state of the task and additional nodes are added. In this case, a planning method 801 based on graph search is introduced, which also infers the feasibility of finding actions that can perform a new task that was not demonstrated during training. In the planning method based on graph search, the initial node is the target node of the task. Then, edges and nodes are added to the graph from the existing set of nodes and the available actions from all such nodes. For example, in the target node 803, the robot can only take two available actions 804 and 805 that lead to states 806 and 807. Similarly, available actions from all other nodes are added and the corresponding edges and vertices are added to the set. When the initial state of the system is reached or a solution is not found, the process terminates. Note that the actions available to the robot during the graph construction process are the individual DMPs learned by the robot by segmenting the demonstration trajectory. In the planning based on graph search, the robot simply builds a feasible graph in which it can use the learned DMPs in different sequences to perform new tasks that are not visible during demonstration.
[0083] Fig. 9 The overall method for learning and task execution proposed in the present disclosure is shown. A robot system equipped with an interface for providing demonstrations and collecting demonstration data is used to provide demonstrations of different tasks on the robot 901. A sensing system including a motion sensor 2101 and a visual sensor 2102 is used to observe and record demonstration trajectories 902 as demonstration data. During the learning process, a training demonstration task is performed, and the demonstration data of the training demonstration task is collected and stored in a dictionary arranged in a memory circuit 2130B as the collected demonstration data. The robot controller uses feature selection (feature selection method) and appropriate metric selection for segmentation (segmentation metric) to segment each demonstration into different segments 903.
[0084] Figure 5The metric for trajectory segmentation used in the proposed work is shown. In particular, the metric is defined as follows:
[0085] Φ=max(|var w |-|var b |)
[0086] Among them, var w is the variance within a single demonstration, and var b is the variance between demonstrations. And the metric Φ is the maximum value of the difference between the variances. The metric is calculated for the features of different segments selected for learning demonstrations. The feature selection (feature selection method) in the present disclosure can be performed using a cost function that is a convex sum of the number of segments obtained by the features and maximizing the segmentation metric (explained above).
[0087] The robot controller creates a dictionary of executable skills (trajectories) 904 using the segmented demonstration and fitting the DMP into the segments 904. The robot controller generates a planning graph for the new task using the known goal state of the task and adds nodes to the planning graph based on the feasibility of performing the task from its current state and the dictionary of skills 905. The robot performs the new task using the planning graph, where the robot transitions between nodes of the planning graph using the learned DMP 906.
[0088] The methods presented in this disclosure can be used to perform many tasks, such as assembly, which consists of many steps that need to be completed in a specific order. Fig.10 The task of pin insertion is shown, which can be performed using the proposed method, which provides demonstrations, sorting them into different components (segmented trajectories) and then fitting a DMP into each component. Fig.10 A possible motion sequence using a robot (end effector of the robot) is shown in , where the robot may need to align a pin in the XY 1001 plane, and then align the pin in a specific axis, such as X 1002. After the robot may insert the pin 1003, the robot retracts the end effector of the robot in 1004. Note that all of these demonstrations can be recorded in a suitable coordinate system 1010. The proposed technology can be used to create a programming-free system for performing complex tasks using a robotic system.
[0089] According to an embodiment of the present invention, the above-mentioned method for learning and task execution is performed by a simulation computer system 2500. The simulation computer system 2500 is configured to create a simulation environment corresponding to the physical environment of the robot system 200, and collect demonstration data generated by moving the robot in the simulation environment to implement the above-mentioned tasks / training using an interface device including a joystick or a virtual reality or augmented reality interface. Once the simulation computer system 2500 collects the demonstration data and / or learning data, those data are transmitted to the controller 205 of the robot system 200 via the communication network 215. The robot system 200 is configured to use the data to perform the desired task / planned task, or to perform further training using the manipulator of the robot system 200 using real parts to improve the performance of the manipulation of the robot system 200.
[0090] The above-described embodiments of the present invention may be implemented in any of a variety of ways. For example, embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code may be executed on any suitable processor or set of processors, whether provided in a single computer or distributed in multiple computers. Such processors may be implemented as integrated circuits, wherein one or more processors are in an integrated circuit assembly. However, the processor may be implemented using circuits of any suitable format.
[0091] Moreover, embodiments of the present invention may be embodied as a method, examples of which have been provided. The actions performed as part of the method may be ordered in any suitable manner. Thus, embodiments may be constructed in which the actions are performed in an order different from that shown, which may include performing some actions simultaneously, even though shown as sequential actions in illustrative embodiments.
[0092] The use of ordinal terms such as "first", "second" and the like in the claims to modify claim elements does not in itself indicate the priority, precedence or order of one claim element relative to another claim element, nor does it indicate the temporal order of performance of method actions, but is merely used as a label to distinguish one claim element with a specific name from another element with the same name (but with an ordinal number) to differentiate the claim elements.
[0093] Although the present invention has been described by way of examples of preferred embodiments, it is to be understood that various other modifications and variations can be made within the spirit and scope of the invention.
[0094] Therefore, it is the object of the appended claims to cover all such changes and modifications as come within the true spirit and scope of the invention.
Claims
1. A robot controller for generating a sequence of motion primitives for a sequential task of a robot having a manipulator, the robot controller include: at least one control processor; and a memory circuit storing a dictionary including the motion primitives, a pre-trained learning module, and a graph-search-based planning module, the graph-search-based planning module having instructions stored thereon that, when executed by the at least control processor, cause the robot controller to perform the following steps: Acquiring, via the interface controller, demonstration data of one or more demonstration tasks provided by an interface device operated by a user for a planned task, wherein the planned task is represented by an initial state and a target state of at least one manipulated object; Each of the demonstration data is segmented into a plurality of segments by selecting features from the demonstration data based on a feature selection method and using a segmentation metric, wherein each segment of the plurality of segments represents a subtask; generating a planning graph by searching a feasible path of the at least one object for the planning task using the graph-search-based planning module and selecting motion primitives from a dictionary in the pre-trained learning module, wherein the pre-trained learning module has been trained based on the collected demonstration data of the training demonstration task; parameterizing a feasible path represented by the motion primitive into a dynamic motion primitive DMP using the initial state and the target state; and The parameterized feasible path of the planned task is implemented as a trajectory according to the selected motion primitives using a manipulator of the robot by tracking and following the parameterized feasible path.
2. The robot controller according to claim 1, in, The planning task is not included in the dictionary.
3. The robot controller according to claim 1, in, The dictionary is updated by adding the parameterized feasible paths according to the selected motion primitives.
4. The robot controller according to claim 1, in, The feature is detected based on a metric providing a maximum separation between the plurality of segments of the one or more presentation tasks.
5. The robot controller according to claim 1, in, The demonstration data for the one or more demonstration tasks is segmented using the features detected by the feature selection method and the segmentation metric method.
6. The robot controller according to claim 1, in, Each of the DMPs is learned for each of the segments of the demonstration task and is parameterized based on the target state and initial state of the planning task.
7. The robot controller according to claim 1, in, The planning graph is created for the planning task based on a planning target state of the planning task, wherein a state transition is created from the planning target state of the planning task based on the feasible path.
8. The robot controller according to claim 1, in, A DMP is generated by segmenting the trajectory of a demonstration task and detecting features of the segmented motion.
9. The robot controller according to claim 1, in, The tracks for the demonstration task are segmented using a metric representing the variance between different demonstrations and the variance within the same track.
10. The robot controller according to claim 1, in, Each of the DMPs is stored as a skill representation of a task, and the dictionary is updated by storing each of the DMPs of all of the segments inferred from the demonstrated task.
11. The robot controller of claim 1 , further comprising generating a control strategy for a new task using the planning graph of the planned task, and fitting a DMP between different nodes of the planning graph according to a skill dictionary.
12. The robot controller according to claim 1, in, The robot controller is connected to a simulation computer system, which is configured to generate a simulation environment corresponding to the physical environment of the robot to virtually perform a predetermined task, wherein the robot controller collects demonstration data, training data, or a combination of the demonstration data and the training data from the simulation computer system, wherein the demonstration data and the training data are generated by the simulation computer system when performing the predetermined task.
13. A computer-implemented method for learning a sequence of motion primitives for a sequential task of a robot, the robot comprising a manipulator, a robot controller having at least one control processor, a memory circuit storing a dictionary including the motion primitives, and a learning module having instructions stored thereon, the instructions, when executed by the at least one control processor, causing the at least one control processor to perform the following steps: collecting demonstration data from trajectories acquired via a motion sensor configured to measure the trajectory of an object as a user manipulates the object using an interface device operated in accordance with a demonstration task, in, Each trajectory corresponds to each demonstration task, where each demonstration task is represented by an initial state and a target state with respect to each object, and where collection continues until the user stops the demonstration task; For each demonstration task, segmenting the demonstration data into motion primitives by dividing the trajectory into primitive trajectories; and The dictionary is updated using the motion primitives based on the collected demonstration data.