Systems and methods for controlling a robot using constrained dynamic motion primitives

The conversion of DMPs to CDMPs addresses constraint incorporation in robot trajectory generation, ensuring safe and efficient operation by adapting to new environments through a non-linear optimization-based perturbation function, thus overcoming collisions and joint limit violations.

JP2025524722AActive Publication Date: 2025-07-30MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025523227
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-14
Filing Date
2023-06-16
Publication Date
2025-07-30
Estimated Expiration
2043-06-16

AI Technical Summary

Technical Problem

Conventional Dynamic Movement Primitives (DMP)-based techniques struggle with incorporating constraints during robot trajectory generation, leading to potential collisions or joint limit violations due to environment changes, and require computationally expensive corrections that are often unrealistic.

Method used

A method to convert DMPs into Constrained DMPs (CDMPs) by defining a perturbation function that satisfies motion constraints through a non-linear optimization problem, allowing adaptation to new environments without re-learning the forcing function, thus incorporating constraints within the skill.

Benefits of technology

Enables safe and efficient robot operation by ensuring compliance with collision avoidance, joint limits, and self-collision constraints, reusing legacy DMP-based control methods in new environments with simplified adaptation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524722000001_ABST
    Figure 2025524722000001_ABST
Patent Text Reader

Abstract

A controller is provided for controlling the operation of a robot to execute a task. The controller includes a memory configured to store a set of dynamic movement primitives (DMPs) associated with the task. The set of DMPs includes at least two sets of dynamic systems, a function representing point attractor dynamics, and a forcing function corresponding to a learned demonstration of the task. The controller includes a processor configured to convert the set of DMPs into a set of constrained DMPs (CDMPs) by obtaining a perturbation function associated with the forcing function. The perturbation function is associated with a set of motion constraints. The processor is further configured to solve a non-linear optimization problem for the set of CDMPs based on the set of motion constraints and, based on the solution, generate a control input for controlling a robot to execute the task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the operation and movement of robots, and more specifically to methods and systems for the operation and movement of robots with constraints.

Background Art

[0002] For performing various operations such as material handling, transportation, welding, assembly, etc., various types of robotic devices have been developed. Learning from demonstrations, also known as demonstration-based robot programming, is a technique commonly used to enable a robot to autonomously execute new tasks. These techniques are based on the premise that an appropriate robot controller can be derived by observing a human performing the desired task himself or observing a human remotely operating the robot for the desired task. Dynamic Movement Primitives (DMP)-based techniques are widely used to learn skills that can be demonstrated to a robot by an expert human or a controller. DMP is a formulation of a non-linear dynamic system that can learn complex trajectories from demonstrations by decoupling a non-linear forcing function from a nominal attracting behavior. DMP-based techniques offer generalization ability and a simple formulation. Furthermore, DMP can be re-parameterized by its start and goal positions.

[0003] However, it may be difficult to directly control the trajectory executed between the start and target positions (postures). Therefore, DMP-based techniques may cause undesirable behavior when there are constraints. Without caution, the robot may collide with itself or the environment (i.e., obstacles) due to new start or target positions, or extend beyond the joint limits of the robot. For example, if the environment in which the task is executed changes from the environment in which the task was demonstrated, causing a potential for collision, conventional DMP-based techniques tend to correct the trajectory estimated based on the skills learned from the demonstration. However, such corrections are computationally expensive and unrealistic, and may even be impossible when the generated trajectory significantly violates the constraints.

[0004] Therefore, there is a need for systems and methods for incorporating constraints during tasks such as robot trajectory generation that use learning from demonstrations in an efficient and effective manner. SUMMARY OF THE INVENTION

[0005] The object of some embodiments is to provide a system and method for constrained control of a robot. Specifically, the object of some embodiments is to provide a system and method for constrained control of a robot manipulator that uses skills learned from demonstrations. In addition to or instead of this, the object of some embodiments is to incorporate constraints during the execution of a task by a robot that uses learning from demonstrations, for example, during the generation of a robot trajectory.

[0006] In addition to or instead of this, the object of some embodiments is to provide such a system and method that can provide constrained control of a robot, specifically, constrained control of a robot manipulator, in the presence of different types of constraints such as collision avoidance with the environment, joint limits, and self-collision.

[0007] Some embodiments are based on the understanding that the skills demonstrated to the robot are executed offline in an environment that may be different from the environment during the actual control of the robot. Also, in the DMP framework, the learned skills are captured using a forcing function that represents the forces acting on the dynamic system during demonstration, enabling the dynamic system to follow the demonstrated trajectory. The forcing function is a mathematical representation of the skills demonstrated to the robot in the demonstrated trajectory. These forces can be applied to the environment, but after the forcing function is learned, the forcing function is environment-independent. Therefore, a natural way to consider constraints is external to the environment-independent forcing function. Furthermore, it is unrealistic to re-learn the forcing function for different environments of robot operation.

[0008] However, some embodiments are based on the recognition that they can be adapted to different environments without the need to re-learn the forcing function. Such adaptation enables the constraints to be within the skill rather than corrections performed outside the skill. To do so, some embodiments aim to find the minimum correction to the predefined weights of the basis functions that form the forcing function such that the new forcing function satisfies the constraints in the new environment.

[0009] Such formulation is different from learning new weights. Because the weights define the skill, while the correction to the weights defines the adaptation of the skill to the environment and / or new environment. Thus, the unknown correction is environment-dependent but skill-independent, simplifying the adaptation and enabling consideration of different types of constraints. Additionally, by finding the correction, the constraints are incorporated inside the forcing function, so that the constraints become intrinsic to the skill without re-learning the process. In this way, different legacy methods for DMP-based control can be reused in the new environment.

[0010] For example, the correction to the skill for adapting the skill to the environment can be represented as an additional parameter in the original formulation of the DMP. These additional parameters can be optimized using an optimization problem that can be solved using a commercially available solver.

[0011] Therefore, the additional parameters in the original formulation of the DMP can define the correction of the weights of the basis functions that can be estimated using an optimization method, and thus the DMP can satisfy specific constraints. Thereby, the correction of the weights is used as a perturbation in the original forcing function, and the DMP is converted into a constrained DMP, and the conversion can be represented by a non-linear optimization problem. Solve the non-linear optimization problem to identify these perturbations, and then identify the control input for the robot to execute the task under the constraint conditions.

[0012] Therefore, in one embodiment, a method for controlling a robot for executing a task is disclosed. The method is executed by a processor coupled to a memory, and the memory stores a set of dynamic movement primitives (DMPs) associated with the task. The set of DMPs includes at least two sets of dynamic systems, and the at least two sets of dynamic systems include at least a function representing point attractor dynamics associated with the task and a forcing function associated with the demonstration of the learned task. The processor includes stored instructions that, when executed by the processor, execute the steps of the method, and the method includes the step of obtaining a set of DMPs associated with the task. The obtained set of DMPs is then converted into a set of constrained DMPs (CDMPs) by defining a perturbation function associated with the learned forcing function, and the perturbation function is associated with a set of motion constraints that need to be satisfied for the execution of the task. Further, based on the motion constraints, a non-linear optimization problem for the set of CDMPs is solved. Further, based on the solution of the non-linear optimization problem for the set of CDMPs, a control input for controlling the robot for executing the task is generated.

[0013] According to another embodiment, a controller is provided for controlling the operation of a robot to perform a task. The controller includes a memory configured to store a set of dynamic movement primitives (DMPs) associated with the task, the set of DMPs including at least two sets of dynamic systems. The two sets of dynamic systems include at least a function representing point attractor dynamics associated with the task and a forcing function associated with a learned demonstration of the task. The controller further includes a processor configured to convert the set of DMPs into a set of constrained DMPs (CDMPs) by obtaining a perturbation function associated with the learned forcing function, the perturbation function being associated with a set of motion constraints that need to be satisfied for the execution of the task, and the processor is further configured to solve a non-linear optimization problem for the set of CDMPs based on the set of motion constraints and generate a control input for controlling the robot to perform the task based on the solution to the non-linear optimization problem for the set of CDMPs. The controller further includes an output interface configured to execute the generated control input and command the robot to perform the task.

[0014] According to yet another embodiment, a non-transitory computer-readable medium storing computer-executable instructions for a robot to perform a task is disclosed. The computer-executable instructions may be configured to obtain a set of DMPs associated with the task, the set of DMPs including at least two sets of dynamic systems. The at least two sets of dynamic systems include at least a function representing point attractor dynamics associated with the task and a forcing function associated with a learned demonstration of the task. The computer-executable instructions may further be configured to convert the set of DMPs into a set of constrained DMPs (CDMPs) by defining a perturbation function associated with the learned forcing function, the perturbation function being associated with a set of motion constraints that need to be satisfied for the execution of the task. The computer-executable instructions may further be configured to solve a non-linear optimization problem for the set of CDMPs based on the motion constraints. Additionally, the computer-executable instructions are configured to generate a control input for controlling a robot to perform the task based on a solution to the non-linear optimization problem for the set of CDMPs.

[0015] It should be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention claimed.

[0016] The embodiments disclosed herein are further described with reference to the accompanying drawings. The drawings shown are not necessarily to scale and instead emphasis is generally placed upon showing the principles of the embodiments disclosed herein.

Brief Description of the Drawings

[0017]

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 4A

Figure 4B

Figure 5

Figure 6A

Figure 6B

Embodiments for Carrying Out the Invention

[0018] Description of Embodiments

[0019] In the following description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, devices and methods are shown in block diagram form only to avoid obscuring the present disclosure.

[0020] As used in this specification and the claims, the terms "for example," "as an example," and "such as," as well as the verbs "comprising," "having," "including," and their other verb forms, when used in conjunction with a listing of one or more components or other items, should be construed as open-ended, meaning that the listing should not be considered as excluding other additional components or items. The term "based on" means at least partially based on. Further, it should be understood that the style and terminology used in this specification are for the purpose of description and should not be regarded as limiting. Any headings used in this specification are for convenience only and have no legal or limiting effect.

[0021] The various embodiments disclosed herein provide a conversion from DMP-based technology to CDMP-based technology for the realization of constraints in tasks performed by a robot. DMP-based technology is based on learned skills captured using a forcing function that represents the forces acting on a dynamic system representing the task during demonstration of the task by a skilled user. This enables the robot to follow the demonstrated task. The forcing function is a mathematical representation of the skills demonstrated to the robot in the demonstrated task. These forces can be applied to the environment, but once the forcing function is learned, the forcing function is environment-independent. That is, the constraints are external to the environment-independent forcing function. It is unrealistic to re-learn the forcing function for different environments of the robot.

[0022] Some embodiments are based on the recognition that they can be adapted to different environments without the need to re-learn the forcing function. Such adaptation enables the constraints to be within the skill rather than corrections performed outside the skill. Therefore, the objective of the present disclosure is to find the minimum correction to the predetermined weights of the basis functions that form the forcing function such that the new forcing function satisfies the constraints in the new environment.

[0023] Some embodiments are further based on the recognition that the above approach is different from new weight learning. This is because the weights define the skills, while the corrections to the weights define the adaptation of the skills to the environment and / or a new environment. Thus, the unknown corrections are environment-dependent but not skill-dependent. This simplifies the adaptation and allows different types of constraints to be considered. In addition, by discovering the corrections, the constraints are incorporated into the forcing function, so that the constraints are inherent in the skills without having to re-learn the process. In this way, different legacy methods for DMP-based control can be reused in a new environment. For example, the corrections to the skills for adapting the skills to a (new) environment can be represented as additional parameters defined by perturbations to the original formulation of the forcing function in the original DMP formulation. These additional parameters can be optimized using an optimization problem that can be solved using any known solver. The resulting formulation is a Constrained DMP or CDMP that provides constrained control of the robot for performing tasks within the environment.

[0024] FIG. 1A shows a block diagram of a system 100 including a controller 101 for controlling the operation of a robot 102 to perform a task, according to an embodiment of the present disclosure.

[0025] The robot 102 may include a robotic arm or robotic manipulator, etc., for which it is desired to perform a task. The task may include lifting an object, moving an object, placing an object at a desired position, moving an object from one position to another, etc. Therefore, the robot 102 may need to follow a motion trajectory to perform the desired task. For example, the robot 102 may be a food placement robot configured to place various foods at designated positions in a box or shipping carton.

[0026] In another example, the robot 102 is an assembly line robot used to lift and move objects within an industry or manufacturing unit, such as between different machines, for purposes such as transporting objects or removing defective manufacturing units.

[0027] Another example of a task is for the end effector of the robot 102 to move from an initial position in 3D space to another position in 3D space. Another example of a task is for the gripper end effector to open, go to a position in 3D space, close the gripper to grasp an object, and move to a final position in 3D space.

[0028] Generally, a task has a start condition and an end condition called a task goal. When the task goal is achieved, the task is considered complete. A task can be divided into subtasks. For example, if the task is for the robot 102 to move to a 3D position in Cartesian space and then pick up an object, the task can be decomposed into subtasks of (1) moving to the 3D position and (2) picking up the object. It is understood that more complex tasks can be decomposed into subtasks by a program. When a task is decomposed into subtasks, a task description can be provided for each subtask. In one embodiment, the task description may be provided by a human operator. In another embodiment, a program is executed to obtain the task description. The task description may also include constraints on the robot 102. Examples of such constraints are that the robot 102 cannot move faster than a speed specified by a human operator, and another example of a constraint is that the robot 102 is prohibited from entering a part of the 3D Cartesian work space specified by a human operator. The purpose of the robot 102 is to complete the task specified in the task description as quickly as possible given any constraints in the task description.

[0029] Robot 102 may include one or more sensors configured to obtain sensor data associated with one or more obstacles or objects present in the environment of robot 102, or movements executed by robot 102 itself. For example, the one or more sensors may include a vision sensor, i.e., a camera. In some embodiments, robot 102 may be communicatively coupled to the one or more sensors via a communication network.

[0030] Furthermore, robot 102 is configured to receive one or more control inputs generated by controller 101. The control inputs are configured to cause robot 102 to perform a desired task. To that end, the control inputs may be received by one or more actuators of robot 102, and the control inputs further generate signals for controlling different parts of the robot, such as the robot arm of robot 102, to move robot 102 along a desired trajectory and thus perform the task. Robot 102 receives different control inputs for different tasks from controller 101 configured to generate these different control inputs. Controller 101 controls the movement of robot 102 to complete the task by sending commands to a physical robot system such as robot 102. In another embodiment, controller 101 is incorporated into robot 102.

[0031] FIG. 1B shows a block diagram of controller 101 for controlling the operation of robot 102 to perform a task. Controller 101 includes a memory 103, a processor 105, and an output interface 106. Additionally, controller 101 may also include an input interface and other components required for controller 101 to perform the operations included in the description herein without departing from the scope of the present disclosure.

[0032] Memory 103 is configured to store computer-executable instructions that can be executed by processor 105. Processor 105 can be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. Memory 103 can include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. Processor 105 is connected via a bus to one or more input and output devices.

[0033] Memory 103 is also configured to store a set 104 of dynamic movement primitives (DMPs) associated with the task. Thus, memory 103 may be configured to store DMP 104 that represents the original task to be executed by robot 102. For example, the original task may include the movement of the end effector of robot 102 along an original trajectory. These DMPs 104 may be obtained by demonstrations performed by a human operator during a learning phase of robot 102 in which robot 102 learns various parameters associated with the task. As will be understood by those skilled in the art, DMP-based techniques are widely used to learn skills that can be demonstrated to a robot by a skilled person or a controller. Thus, using techniques for learning from demonstrations, the characteristics of the original demonstration are learned by robot 102. These characteristics include, for example, how the end effector of the robot moved during task execution.

[0034] DMP104 is a formulation of a non - linear dynamic system that can learn complex tasks (such as trajectories) from demonstrations by decoupling non - linear forcing functions from nominal pulling - in behavior. That is, DMP104 is a method of task control and planning. Therefore, DMP104 is a proposed mathematical formalization of complex subtasks, such as the movement in the case of a trajectory, which is composed of a set of primitive actions "building blocks" that are executed sequentially and / or in parallel. As further understood, since DMP104 is a non - linear dynamic system, DMP104 is different from the building blocks proposed so far.

[0035] For example, in the case of trajectory planning, DMP104 can be understood as a combination of two systems: a virtual system that plans the trajectory and a real system that actually executes the planned trajectory. DMP104 can include its own dynamics, and once the DMP is set up, a control signal that the robot 102 should follow can be obtained. This control signal forms the control input generated by the controller 101 for the robot 102 to perform the desired task. For example, the DMP control signal for the path (i.e., trajectory) that the end - effector (also called the robot manipulator) of the robot 102 should follow can include a set of forces that need to be applied to the end - effector to execute that path. The robot 102 may apply these forces by converting them into joint torques.

[0036] For the task of trajectory planning, the trajectory can be represented by the function y d (t), where t is in the set [0, T]. The trajectory can include multiple waypoints (postures and attitudes) of the end - effector in Cartesian space. When one or more trajectories y(t) are recorded for one fixed posture, the DMP learning algorithm can learn a DMP for each of the components of y(t). To remove explicit time - dependence, DMP104 uses a canonical system to track the progress of the task.

[0037] Therefore, DMP104 is in the form of two combined sets of parameterized ordinary differential equations (ODEs) that represent tasks. For example, DMP104 can generate a trajectory that moves a system such as robot 102 from an initial pose to a target pose. DMP104 can easily adapt the trajectory according to new start and target states, thus essentially constituting a closed-loop controller. Also, DMP104 can learn from a limited number of training examples, including just a single training example. Therefore, it is possible to modify the original trajectory according to changes in the initial and target poses. The formulation of DMP104 is further described in FIG. 2A.

[0038] FIG. 2A shows a schematic diagram illustrating a set 1 of DMPs 104 used by the controller 101 of FIG. 1B, according to an embodiment of the present disclosure. The set 104 of DMPs includes a set of at least two dynamic systems, a first dynamic system 107 and a second dynamic system 108. The second dynamic system 108 includes a function representing point attractor dynamics 109 associated with the task, and a forcing function 110 associated with the learned demonstration of the task.

[0039]

Number

[0040] Furthermore, the second dynamic system 108 is shown in FIG. 2B.

[0041]

Number

[0042] For example, the set of DMPs 104 is associated with a point attractor dynamics 109 parameterized by a starting spatial pose and a target spatial pose of the robot 102, and a forcing function 110 including one or more weights corresponding to basis functions associated with an original trajectory of the robot 102. The original trajectory may be configured with multiple original spatial points between the starting spatial pose and the end spatial pose. Note that the one or more predetermined weights may be configured in a first configuration in response to one or more demonstrations.

[0043] Furthermore, the forcing function 110 is represented by the term f(x,g) shown in FIG. 2C.

[0044]

number

[0045]

number

[0046] Therefore, the forcing function 110 is an adjustable parameter w i The weights 111 of the weighted combination are learned based on a performance of the task. Furthermore, one or more of the weights 111 are learned by solving a local weighted regression to learn one or more weights 111 for the basis functions 112. In the example of a trajectory generation task for a robot 102, a forcing function 110 f(x,y) is learned from a performance trajectory of the task. By learning f(x,y) from the performance, characteristics of the original performance (e.g., how the robot's end effector moved during the task) are learned. These are learned, for example, in a training phase, by solving a local weighted regression to learn weights for the basis functions 112.

[0047] In particular, the forcing function 110 may include a weighted combination of basis functions 112. For example, the basis functions 112 may be radial basis functions as shown in Figure 2D.

[0048] [Number]

[0049] Referring again to FIG. 1B, the controller 101 further includes a processor 105 configured to execute one or more computer-executable instructions. The one or more computer-executable instructions cause the processor 105 to perform one or more operations. Thus, by performing the one or more operations, the processor 105 is configured to convert the set 104 of DMPs into a set of constrained DMPs (CDMPs).

[0050] The present disclosure provides a constrained DMP (CDMP) that includes an additional function representing a perturbation to the original forcing function 110 so that the learned DMPs 104 can satisfy new motion constraints. Note that the motion constraints may depend on the environment in which the robot 102 performs the task. Therefore, different environments may have different constraints, and thus the function may only depend on the constraints that need to be satisfied during the operation of the robot 102 in the new environment.

[0051] FIG. 3A shows a block diagram illustrating a conversion 113 from a set 104 of DMPs to a set 114 of constrained DMPs (CDMPs) according to an embodiment of the present disclosure.

[0052] The conversion 113 is performed by first defining a perturbation to the initially learned forcing function 110, and the perturbation is associated with a set 115 of motion constraints related to the task performed by the robot 102. The set 115 of motion constraints may be defined by the environment in which the robot 102 is operating. More specifically, the perturbation defines an additional set of parameters that represent the motion constraints 115 in the learned forcing function 110. The additional set of parameters represents new task constraints. Therefore, the success of the task execution is determined based on the satisfaction of the constraints included in the set 115 of motion constraints.

[0053] For example, the set of motion constraints 115 may include collision avoidance constraints. The collision avoidance constraints may include conditions for ensuring that the robot 102 does not collide with obstacles that may be present in the environment of the robot 102.

[0054] In another example, the set of motion constraints may include self - collision avoidance constraints. The self - collision avoidance constraints may include conditions for ensuring that the robot 102 does not collide with itself.

[0055] In yet another example, the set of motion constraints 115 may include constraints on the joint limits of the end - effector of the robot 102. For example, consider a task where a learned DMP is re - parameterized in a new environment based on new goals and start states. However, the new trajectory for this task may not be kinematically executable for some of the joint rotations, and thus may not be possible to perform for the robot 102. Such constraints are explicitly considered in the formulation of the CDMP114, and thus there is a possibility of finding an executable solution that satisfies the joint limits.

[0056] The set of motion constraints 115 is introduced in the form of adding perturbation terms to the forcing function 110 to be learned. As a result of the transformation 113, as shown in Figure 3B, the original set 104 of DMPs is transformed into a set of CDMPs.

[0057]

Number

[0058] Furthermore, the perturbation function 116 includes new task constraints represented by an additional set of parameters as a barrier function that is at least once differentiable, which is shown in Figure 3C.

[0059]

Number

[0060]

Number

[0061] Figure 3D shows the mathematical formulation of the non - linear optimization problem 118 that is solved to obtain the value of the barrier function 117 ζ i of the non - linear optimization problem 118 that is solved to obtain the value of the barrier function 117 ζ

[0062]

Number

[0063] Next, using the solution of the non - linear optimization problem 118 for the set of CDMPs based on the set of motion constraints 115, an executable set of values for an additional parameter set for the new task executed by the robot 102 is identified, and the barrier function 117 here represents the new task constraints. The executable set of values corresponds to the values of the additional parameter sets that satisfy the constraints associated with the tasks to be executed, such as the new task constraints mentioned here, under a given set of environmental conditions. Thus, the executable set of values includes the set of all possible points of the non - linear optimization problem 118 that satisfy the problem constraints, potentially including inequalities, equalities, and integer constraints. Further, the perturbation function 116 is obtained using the solution of the non - linear optimization problem 118, and then, using the perturbation function 116 further, one or more control inputs for controlling the robot 102 to execute the task are generated based on the solution of the non - linear optimization problem 118 for the set of CDMPs 114.

[0064]

Number

[0065] In one example, the additional constraint 119 specifies a limit on the deviation amount of the CDMP 114 from the original DMP 104. The design of the CDMP 114 can be regarded as a trade-off between constraint satisfaction and the original forcing function 110. This trade-off can be controlled using a hyperparameter that constrains the maximum allowable deviation between the original DMP 104 and the CDMP 114, and the hyperparameter can be added as an additional constraint 119 to the non-linear optimization problem 118.

[0066] The solution to the non-linear optimization problem 118 may be obtained using any known commercially available solver such as IPOPT (trademark), SNOPT, etc. Thus, the conversion from the DMP 104 disclosed herein to the CDMP 114 provides a very cost-effective, computationally efficient, easily implementable, and practical solution to the problem of robot motion control in a constrained environment. Also, since the CDMP 114 is based on satisfying the motion constraints 115 including safety and collision avoidance conditions, the overall motion of the robot 102 becomes very safe and efficient.

[0067] Furthermore, the solution to the optimization problem 118 determines the corrected weights for the modified forcing function 110, and the weights are converted into control inputs, which are transmitted to the robot 102 by the output interface 106 to command the robot 102 to execute the generated control inputs to perform the task.

[0068] FIG. 4A shows a flowchart of a method executed by the controller 101 to perform a task by the robot 102 according to an embodiment of the present disclosure.

[0069] At 401, a set of DMPs associated with the task is obtained. For example, the set 104 of DMPs stored in the memory 103 is obtained by the processor 105.

[0070] Next, at 402, the set of DMPs is converted to a set of CDMPs by defining a perturbation to the first learned forcing function associated with the task. For example, DMP 104 includes the first learned forcing function 110 f(x,g), and the forcing function is converted by defining a perturbation function 116 g(x) added as an additional parameter to the learned forcing function 110 f(x,g) (113). The perturbation function 116 g(x) is associated with the motion constraints 115 that need to be satisfied for the execution of the task.

[0071] To determine the value of the perturbation function 116, at 403, a non-linear optimization problem 118 is formulated and solved by the processor 105 using the formulation of the set of CDMPs 114 and the motion constraints 115, including a set of new task constraints defined by changes in the environment of the robot executing the task. These new task constraints are represented as an additional set of parameters in the perturbation function 116. Therefore, FIG. 4B shows a flowchart of another method executed by the controller 101 to determine the solution of the non-linear optimization problem 118. At 405, a set of executable values for the barrier function 117 representing the new task constraints for the motion constraints 115 is determined. This is discussed with reference to FIG. 3D.

[0072] Next, at 406, the value of the perturbation function 116 is determined, such as by using the value of the barrier function 117 and using the modified forcing function given in Equation 7. .

[0073] Furthermore, at 407, the solution of the value of the perturbation function 116 is used to generate a control input by the controller 101 for controlling the robot 102. The methods shown in FIGS. 4A and 4B are described using an exemplary trajectory generation task as given below.

[0074] In this example, the original trajectory may be obtained using sensors based on a plurality of movements executed by an end effector associated with the robot 102 according to a demonstration. For example, the sensors may be encoders and / or vision sensors. These movements are used to generate DMPs 104 corresponding to the movements executed by the end effector associated with the robot 102. These DMPs are then stored in the memory 103 of the controller 101 and retrieved in step 401 described above.

[0075] As will be appreciated, the DMPs 104 may cause undesirable behavior if there are motion constraints 115 that are different from those during the demonstration. Since the DMPs 104 do not explicitly consider additional disturbances and simply learn the forcing function during the demonstration, the resulting generalization may only be acceptable during the demonstration. However, if there is no adaptation of the forcing function during actual operation, the resulting robot trajectory may be infeasible during actual control. For example, if the environment changes, the robot 102 may collide due to new task constraints on the allowable states of the robot that did not exist during operation, such that an object the robot 102 is operating on may exceed the scope of the robot defined by the physical constraints on the robot's structure.

[0076] The motion constraints 115 may include additional or different obstacles present in the environment of the robot 102 compared to the previous environment. For this reason, the motion constraints 115 may include the positions and configurations associated with the obstacles present in the environment of the robot 102 in a changed environment. For example, the positions and configurations may be obtained using a vision sensor that uses pose estimation techniques. Alternatively or in addition, the robot 102 may have joint limitations of the robot's end effector.

[0077] Thus, these additional constraints can form new task constraints that can be used in step 402 to define perturbations 116 to the original forcing function 110, whereby, based on one or more motion constraints 115, one or more predetermined weights 111 associated with the original trajectory of the original forcing function 110 are reconfigured in a second configuration. In some embodiments, reconfiguring one or more predetermined weights 111 may include converting one or more motion constraints 115 into one or more differentiable functions using a smoothing function or a barrier function 117. For example, the smoothing function may be a control barrier function (CBF).

[0078] In one example, the barrier function 117 is a zeroing barrier function (ZBF). One advantage of using ZBF to represent constraints is the generality provided by ZBF. That is, ZBF can prove joint limit avoidance and obstacle avoidance. Further, as a result of the above formulation, a non-linear optimization is obtained that perturbs the DMP forcing weights regressed by local weighted regression to approve the ZBF constructed by the user, and the non-linear optimization can be solved using a standard NLP solver. CDMP is subject to different constraints on the movement of the end effector such as collision avoidance, and state constraints for safety. ZBF is used to represent smooth task-specific constraints. An additional set of parameters that can be optimized to satisfy these constraints is added to the resulting formulation. The resulting optimization is cast as a non-linear program, which can be solved using a commercially available non-linear program solver such as IPOPT. By utilizing the set invariance of ZBF, constraint satisfaction for CDMP is guaranteed. Next, one or more weights can be optimized using ZBF.

[0079]

Number

[0080]

Number

[0081]

Number

[0082]

Number

[0083]

Number

[0084]

Number

[0085] One or more predetermined weights are optimized using these one or more differentiable functions. This is done by formulating and solving a non-linear optimization problem in step 403.

[0086] CDMP incorporates an existing DMP with forcing functions learned from expert trajectories and then optimizes this forcing term so that the DMP dynamic system approves a ZBF that proves that the DMP generates trajectories that remain within the safe set of the workspace. This safe set (and ZBF) is constructed by composing a signed distance field from primitive convex polyhedra. More specifically, the resulting non-linear optimization problem is described as follows.

[0087]

Number

[0088]

Number

[0089]

Number

[0090] In the formula, ζ i is the decision variable optimized for the formulated optimization problem. Note that this is just one possible way to represent the perturbation of the forcing function 110 of the original DMP104 based on the most common representation as a radial basis function in the DMP literature.

[0091]

Number

[0092] The problem formulation represented by [Equations 11 - 14] is converted into a finite - dimensional discretization problem. Next, this can be solved using a non - linear optimization solver such as IPOPT or SNOPT to generate the desired parameter set.

[0093] Furthermore, based on one or more predetermined weights configured in the second configuration, a new trajectory is generated. The new trajectory may include a plurality of new space points between the start space point and the end space point. It should be further noted that at least one of the plurality of new space points may be different from the plurality of original space points.

[0094] In some embodiments, generating a new trajectory may include formulating an optimization problem with non - linear dynamic constraints 118 using one or more predetermined weights associated with the original trajectory and one or more differentiable functions 117 corresponding to one or more motion constraints 115. Further, generating a new trajectory may include solving the optimization problem 118 with non - linear dynamic constraints by optimizing one or more predetermined weights for a radial basis function such as the basis function 112 to generate a new trajectory. The new trajectory satisfies one or more motion constraints 115 for performing the task. Thus, determining a new trajectory includes determining a correction term 116 that modifies at least a portion of the weights 111 of the forcing function 110 such that the DMP 104 with the forcing function having the corrected weights represents a new executable trajectory that satisfies the motion constraint 115. In some embodiments, the one or more predetermined weights may be optimized using a gradient - based solver, and the gradient - based solver is based on an interior point method such as IPOPT (Interior Point OPTimizer), SNOPT (Sparse Nonlinear OPTimizer), etc.

[0095] In some embodiments, to generate a new trajectory, a hyperparameter corresponding to the deviation of the new trajectory from the original trajectory may be specified. Further, based on the hyperparameter, the degree of deviation of the new trajectory from the original trajectory may be limited.

[0096]

Number

[0097]

Number

[0098]

Number

[0099] In this way, using the controller 101, the motion control of the robot 102 can be executed to perform various tasks such as trajectory generation and optimization tasks in the case of new constraints as described above. As will be understood by those skilled in the art, the examples of trajectory generation described herein are merely for illustration. Without departing from the scope of the present disclosure, any equivalent examples can be used to implement the principles of the various embodiments disclosed herein.

[0100] The controller 101 may be embodied as being within the robot 102. The controller 101 may be any general-purpose or dedicated computer system well known in the art. An example of such a computer system is described in FIG. 5.

[0101] FIG. 5 is a block diagram 500 of an exemplary computer system for implementing various embodiments. The disclosed methods and systems may be implemented on a conventional or general-purpose computer system such as a personal computer (PC) or a server computer. Referring now to FIG. 5, a block diagram 500 of an exemplary computer system 502 for implementing various embodiments is shown. The controller 101 may be implemented using the computer system 502. Alternatively, the computer system 502 may be the controller. The computer system 502 may include a central processing unit (“CPU” or “processor”) 504. The processor 504 may include at least one data processor for executing program components for executing requests generated by a user or generated by the system. The processor 504 may be equivalent to the processor 105 shown in FIG. 1B. The user may include a person, a person using a device as included in the present disclosure, or the device itself. The processor 504 may include special processing units such as an integrated system (bus) controller, a memory management control unit, a floating point unit, a graphics processing unit, a digital signal processing unit, etc. The processor 504 may include a microprocessor such as an AMD® ATHLON® microprocessor, a DURON® microprocessor or an OPTERON® microprocessor, an ARM application, embedded or secure processor, an IBM® POWERPC® processor, an Intel® CORE® processor, an ITANIUM® processor, a XEON® processor, a CELERON® processor, or other lines of processors. The processor 504 may be implemented using a mainframe, a distributed processor, a multi-core, parallel, grid, or other architecture. Some embodiments may utilize embedded technologies such as application specific integrated circuits (ASICs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), etc.

[0102] Processor 504 may be arranged to communicate with one or more input / output (I / O) devices via I / O interface 506. I / O interface 506 may employ communication protocols / methods such as, but not limited to, audio, analog, digital, monaural, RCA, stereo, IEEE-1394, serial bus, Universal Serial Bus (USB), infrared, PS / 2, BNC, coaxial, component, composite, Digital Visual Interface (DVI), High-Definition Multimedia Interface (HDMI (registered trademark)), RF antenna, S-video, VGA, IEEE802.n / b / g / n / x, Bluetooth (registered trademark), cellular (e.g., Code Division Multiple Access (CDMA), High-Speed Packet Access (HSPA+), Global System for Mobile Communications (GSM), Long Term Evolution (LTE), WiMax, etc.).

[0103] Computer system 502 may communicate with one or more I / O devices using I / O interface 506. For example, input device 508 may be an antenna, keyboard, mouse, joystick, (infrared) remote control, camera, card reader, facsimile, dongle, biometric reader, microphone, touch screen, touch pad, trackball, sensor (e.g., accelerometer, light sensor, GPS, gyroscope, proximity sensor, etc.), stylus, scanner, storage device, transceiver, video device / source, visor, etc. Output device 510 may be a printer, facsimile, video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light emitting diode (LED), plasma, etc.), audio speaker, etc. In some embodiments, transceiver 512 may be disposed in relation to processor 504. Transceiver 512 may facilitate various types of wireless transmission or reception. For example, transceiver 512 may include an antenna operably connected to a transceiver chip (e.g., TEXAS (registered trademark) INSTRUMENTS WILINK WL1286 (registered trademark) transceiver, BROADCOM (registered trademark) BCM4550IUB8 (registered trademark) transceiver, INFINEON TECHNOLOGIES (registered trademark) X-GOLD 618-PMB9800 (registered trademark) transceiver, etc.) and may provide IEEE802.6a / b / g / n, Bluetooth, FM, global positioning system (GPS), 2G / 3G HSDPA / HSUPA communication, etc.

[0104] In some embodiments, the processor 504 may be arranged to communicate with a communication network 514 via a network interface 516. The network interface 516 may communicate with the communication network 514. The network interface 516 may employ connection protocols including, but not limited to, direct connect, Ethernet® (e.g., twisted pair 50 / 500 / 5000 base T), Transmission Control Protocol / Internet Protocol (TCP / IP), token ring, IEEE802.11a / b / g / n / x, etc. The communication network 514 may include, but is not limited to, direct interconnection, local area network (LAN), wide area network (WAN), wireless network (e.g., using wireless application protocol), the Internet, etc. The computer system 502 may communicate with devices 518, 520, and 522 using the network interface 516 and the communication network 514. These devices 518, 520, and 522 may include various mobile devices such as personal computers, servers, facsimiles, printers, scanners, mobile phones, smartphones (e.g., APPLE® IPHONE® smartphones, BLACKBERRY® smartphones, ANDROID®-based phones, etc.), tablet computers, e-book readers (AMAZON® KINDLE® e-readers, NOOK® tablet computers, etc.), laptop computers, notebooks, game consoles (MICROSOFT® XBOX® game consoles, NINTENDO® DS® game consoles, SONY® PLAYSTATION® game consoles, etc.), etc., but are not limited thereto. In some embodiments, the computer system 502 itself may embody one or more of these devices 518, 520, and 522.

[0105] In some embodiments, the processor 504 may be arranged to communicate with one or more memory devices 530 (e.g., RAM 526, ROM 528, etc.) via a storage interface 524. The storage interface 524 may connect to a memory 530 including, but not limited to, a memory drive, a removable disk drive, etc., that employ connection protocols such as SATA (serial advanced technology attachment), IDE (integrated drive electronics), IEEE-1394, Universal Serial Bus (USB), Fibre Channel, SCSI (small computer systems interface), etc. The memory drive may further include a drum, a magnetic disk drive, a magneto-optical drive, an optical drive, RAID (redundant array of independent discs), a solid state memory device, a solid state drive, etc.

[0106] Memory 530 may store a collection of program or data repository components, including but not limited to, operating system 532, user interface application 534, web browser 536, mail server 538, mail client 540, user / application data 542 (e.g., any data variable or data record discussed in the present disclosure), etc. Memory 530 may be equivalent to memory 103. Operating system 532 may facilitate resource management and operation of computer system 502. Examples of operating system 532 include, but are not limited to, the APPLE (registered trademark) MACINTOSH (registered trademark) OS X (registered trademark) platform, the UNIX (registered trademark) platform, Unix-like system distributions (e.g., Berkeley Software Distribution (BSD), FreeBSD, NetBSD, OpenBS, etc.), LINUX distributions (e.g., RED HAT (registered trademark), UBUNTU (registered trademark), KUBUNTU (registered trademark), etc.), IBM (registered trademark) OS / 2 platform, MICROSOFT (registered trademark) WINDOWS (registered trademark) platform (XP, Vista / 7 / 8, etc.), APPLE (registered trademark) IOS (registered trademark) platform, GOOGLE (registered trademark) ANDROID (registered trademark) platform, BLACKBERRY (registered trademark) OS platform, etc. User interface 534 may facilitate the display, execution, interaction, operation, or action of program components through text or graphic functions. For example, user interface 534 may provide computer interaction interface elements on a display system operably connected to computer system 502, such as a cursor, icon, checkbox, menu, scroll bar, window, widget, etc.A graphical user interface (GUI) including, but not limited to, the AQUA (registered trademark) platform of the APPLE (registered trademark) Macintosh (registered trademark) operating system, the IBM (registered trademark) OS / 2 (registered trademark) platform, the MICROSOFT (registered trademark) WINDOWS (registered trademark) platform (e.g., the AERO (registered trademark) platform, the METRO (registered trademark) platform, etc.), UNIX X-WINDOWS, web interface libraries (e.g., the ACTIVEX (registered trademark) platform, the JAVA (registered trademark) programming language, the JAVASCRIPT (registered trademark) programming language, the AJAX (registered trademark) programming language, HTML, the ADOBE (registered trademark) FLASH (registered trademark) platform, etc.) may be adopted.

[0107] In some embodiments, computer system 502 may implement program components stored in web browser 536. The web browser 536 may be a hypertext browsing application such as the MICROSOFT® INTERNET EXPLORER® web browser, the GOOGLE® CHROME® web browser, the MOZILLA® FIREFOX® web browser, the APPLE® SAFARI® web browser, etc. Secure web browsing may be provided using HTTPS (Secure Hypertext Transfer Protocol), Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. The web browser may utilize features such as AJAX, DHTML, the ADOBE® FLASH® platform, the JAVASCRIPT® programming language, the JAVA® programming language, application programming interfaces (APIs), etc. In some embodiments, computer system 502 may implement program components stored in mail server 538. The mail server 538 may be an Internet mail server such as the MICROSOFT® EXCHANGE® mail server, etc. The mail server 538 may utilize features such as ASP, ActiveX, ANSI C++ / C#, the MICROSOFT®.NET® programming language, CGI scripts, the JAVA® programming language, the JAVASCRIPT® programming language, the PERL® programming language, the PHP® programming language, the PYTHON® programming language, WebObjects, etc. The mail server 538 may utilize communication protocols such as the Internet Message Access Protocol (IMAP), the Messaging Application Programming Interface (MAPI), Microsoft Exchange, the Post Office Protocol (POP), the Simple Mail Transfer Protocol (SMTP), etc.In some embodiments, computer system 502 may implement program components stored in mail client 540. Mail client 540 may be a mail viewing application such as the APPLE MAIL® mail client, the MICROSOFT ENTOURAGE® mail client, the MICROSOFT OUTLOOK® mail client, the MOZILLA THUNDERBIRD® mail client, and the like.

[0108] In some embodiments, computer system 502 may store user / application data 542, such as data, variables, records, etc., as described in this disclosure. Such a data repository may be implemented as a fault-tolerant, relational, scalable, secure data repository, such as an ORACLE® data repository or a SYBASE® data repository. Alternatively, such a data repository may be implemented using standardized data structures such as arrays, hashes, linked lists, structures, structured text files (e.g., XML), tables, or as an object-oriented data repository (e.g., using an OBJECTSTORE® object data repository, a POET® object data repository, a ZOPE® object data repository, etc.). Such a data repository may be integrated or distributed, in some cases, among the various computer systems discussed above in this disclosure. It should be understood that the structure and operation of any computer or data repository component may be combined, integrated, or distributed in any working combination.

[0109] One operational combination of such a system may include a combination of a computer system 502 operating as a controller 101 to control the robot 102 to perform a task based on a prior demonstration of a task performed by a human operator.

[0110] FIG. 6A shows a schematic diagram of a use case of a robot 602 in a process of learning from demonstration according to an embodiment of the present disclosure.

[0111] A human operator 600 performs a demonstration of a task, such as a task in which the robot 602 holds a moving object 604 with a gripper 603, moves along a trajectory 605, and places the moving object 604 in a fixed posture 607 within a stationary object 606. The demonstration may be related to an assembly task including the moving object 604.

[0112] According to an embodiment, the human operator 600 may use a teaching pendant 601 that stores the coordinates of waypoints corresponding to the original trajectory 605 in the memory of the robot 602 to instruct the robot 602 to follow the original trajectory 605. The teaching pendant 601 may be a remote control device. The remote control device may be configured to transmit a robot configuration setting (i.e., the setting of the robot) to the robot 602 to demonstrate the original trajectory 605. For example, the remote control device transmits control commands such as movement in the XYZ directions, speed control commands, joint position commands, etc. to demonstrate the original trajectory 605. In an alternative embodiment, the human operator 600 may use a joystick to instruct the robot 602 through kinesthetic feedback and the like. The human operator 600 may instruct the robot 602 to follow the original trajectory 605 multiple times for the same fixed posture 607 of the stationary object 606.

[0113] Robot 602 may be coupled to or include a controller 602a, such as controller 101 discussed in the previous embodiment. Controller 602a includes a memory that may store a DMP associated with the original trajectory 605 based on a demonstration. These DMPs may include forcing functions learned based on a demonstration of the original trajectory 605 by a human operator 600. Thus, the DMP of the original trajectory corresponds to the set 104 of DMPs shown in the previous figure, specifically FIG. 1B. The original trajectory 605 may include a plurality of original spatial points between the starting and ending postures of the robot 602. Further, the original trajectory 605 may be generated based on a demonstration corresponding to a task executed in a first set of conditions. However, during the execution of the task, a second set of conditions or a different environment may be observed by the robot 602. Note that the second set of conditions may be different from the first set of conditions.

[0114] FIG. 6B shows an example of the operation of the robot 602 in an environment different from that of FIG. 6A, having a second set of conditions. The second set of conditions may include motion constraints that form new task constraints for the environment of FIG. 6B, related to an obstacle 608 within the path of the original trajectory 605. The position and configuration of the obstacle 608 may be determined using one or more sensors, such as a vision sensor. Thus, the position and configuration associated with the obstacle 608 present in the environment of the robot 602 in the second set of conditions may be obtained.

[0115] In this scenario, first, the controller 602a may obtain a set of DMPs stored in memory, obtained using the demonstration of FIG. 6A. The controller 602a may further be configured to convert the obtained set of DMPs into a set of CDMPs by using motion constraints for obstacle avoidance for the obstacle 608 within the original trajectory 605. This conversion may be performed using the conversion 115 described in the previous embodiment. Thus, one or more predetermined weights associated with the original trajectory may be optimized using one or more differentiable functions related to the motion constraints.

[0116] Furthermore, the controller 602a can then dynamically generate a new trajectory 605a based on one or more predetermined weights configured in a second configuration. The new trajectory 605a may include a plurality of new spatial points between the start pose and the end pose. It should be further noted that at least one of the plurality of new spatial points may be different from the plurality of original spatial points.

[0117] Therefore, generating the new trajectory 605a may include formulating an optimization problem with non - linear dynamic constraints, such as the non - linear optimization problem 118 described in the previous embodiments, using one or more predetermined weights of the original trajectory and one or more differentiable functions corresponding to one or more new task constraints. Furthermore, generating the new trajectory may include solving the optimization problem with non - linear dynamic constraints by optimizing one or more predetermined weights for the radial basis functions to generate the new trajectory 605a. The new trajectory 605a satisfies one or more new task constraints for performing the task. Therefore, determining the new trajectory 605a includes determining a correction term that modifies at least a part of the weights of the forcing function so that the DMP with the forcing function having the corrected weights represents a new executable trajectory that satisfies the motion constraints. The new executable trajectory enables the robot 602 to operate smoothly in an environment different from the original environment and avoid any obstacles along the new trajectory 605a.

[0118] Furthermore, when the environment of the robot 602 changes, the object 604 being manipulated by the robot 602 may go out of the range of the robot 602, and the range may be defined by physical constraints on the structure of the robot 602.

[0119] Therefore, use the set of CDMPs to solve a non - linear optimization problem to satisfy the motion constraints for obstacle avoidance. Next, use the solution of the non - linear optimization problem to determine a new trajectory 605a for placing the object 604 within the stationary object 606. Next, the controller of the robot 602 can generate a control input that causes the end - effector of the robot 602 to follow the new trajectory 605a with the gripper 603 to successfully place the object 604 at a new target pose 607a within the stationary object 606.

[0120] Thus, without the burden of further learning calculations, the robot 602 operated using the controller 602a can successfully and safely complete the task. The robot 602 can adapt to different types of environments by converting from DMP to CDMP without incurring additional learning and demonstration costs, and by naturally resetting the weights of the basis functions and solving the non - linear optimization problem using any known off - the - shelf solver. In addition, the robot 602 can achieve higher safety compared to DMP - based task execution through enforcement and satisfaction guarantees for the constraints as defined in the transformed CDMP formulation disclosed herein.

[0121] For clarity, it will be understood that the above description has been presented with reference to different functional units and processors for embodiments of the present invention. However, it will be apparent that any suitable distribution of functionality between different functional units, processors or domains may be used without departing from the present invention. For example, functions shown to be executed by separate processors or controllers may be executed by the same processor or controller. Thus, references to specific functional units should be regarded only as references to suitable means for providing the described functionality, and not as indicating a strict logical or physical structure or organization.

[0122] Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present disclosure. A computer-readable storage medium refers to any kind of physical memory where information or data readable by a processor can be stored. Thus, a computer-readable storage medium may store instructions executable by one or more processors, including instructions for causing a processor to perform steps or stages consistent with the embodiments described herein. The term "computer-readable medium" should be understood to include tangible items and to exclude carrier waves and transitory signals, i.e., to be non-transitory. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, and any other known physical storage media.

[0123] The present disclosure and examples are considered to be merely illustrative, and the true scope and spirit of the disclosed embodiments are intended to be indicated by the appended claims.

[0124] Also as will be appreciated, the above-described techniques may take the form of processes implemented by a computer or a controller and apparatus for performing these processes. The present disclosure may also be embodied in the form of computer program code including instructions embodied in a tangible medium such as a floppy disk cassette, a solid state drive, a CD-ROM, a hard drive, or any other computer-readable storage medium, the computer program code being loaded into and executed by a computer or a controller, and when the computer program code is loaded into and executed by a computer or a controller, the computer becomes an apparatus for implementing the present invention. The present disclosure may also be embodied in the form of a computer program code or a signal, whether stored in a storage medium, loaded into and / or executed by a computer or a controller, or transmitted on some transmission medium such as on an electrical wiring or cable, through an optical fiber, or via electromagnetic radiation, etc., and when the computer program code is loaded into and executed by a computer, the computer becomes an apparatus for implementing the present invention. When implemented on a general-purpose microprocessor, the computer program code segments configure the microprocessor to create specific logic circuits.

[0125] The disclosed methods and systems may be implemented on a conventional or general-purpose computer system such as a personal computer (PC) or a server computer. For clarity, it will be understood that the above description has been presented with reference to different functional units and processors for the embodiments of the present invention. However, it will be apparent that any suitable distribution of functionality between different functional units, processors or domains may be used without departing from the present invention. For example, functions shown to be executed by separate processors or controllers may be executed by the same processor or controller. Thus, references to specific functional units are to be seen only as references to suitable means for providing the described functionality, and not as indicating any strict logical or physical structure or organization.

[0126] The robot can be understood to mean a physical robot system or a robot simulator aimed at faithfully simulating the behavior of a physical robot system, in the absence of a classification as "physical", "real", or "real-world". The robot simulator is a program consisting of a set of mathematical algorithms for simulating the kinematics and dynamics of a real-world robot. In a preferred embodiment, the robot simulator also simulates the robot controller. The robot simulator can generate data for 2D or 3D visualization of the robot, and the data can be output to a display device via a display interface.

[0127] The above description provides only embodiments as specific examples and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the following description of embodiments as specific examples will provide those skilled in the art with an explanation that enables the implementation of one or more embodiments as specific examples. Various changes may be made to the functions and configurations of the elements without departing from the spirit and scope of the disclosed subject matter recited in the appended claims.

[0128] Specific details are given in the following description in order to provide a thorough understanding of the embodiments. However, one skilled in the art will understand that the embodiments can be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may sometimes be shown as components in block diagram form in order not to obscure the embodiments with unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail in order not to obscure the embodiments. Further, like reference numbers and designations in the various drawings indicate like elements.

[0129] Also, individual embodiments may be described as a process shown as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. A flowchart may describe operations as a sequential process, but many of the operations can be performed in parallel or simultaneously. In addition, the order of the operations may be rearranged. A process may end when its operations are completed, but may have additional steps not discussed or included in the figure. Further, not all operations in any specifically described process may occur in all embodiments. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, the end of the function may correspond to returning the function to the calling function or the main function.

[0130] Furthermore, embodiments of the disclosed subject matter may be implemented, at least in part, either manually or automatically. Manual or automatic implementation may be carried out or at least assisted through the use of a machine, hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the required tasks may be stored on a machine-readable medium. A processor may perform the required tasks.

[0131] Each of the various methods or processes outlined in this specification may be encoded as software executable on one or more processors employing any one of a variety of operating systems or platforms. Additionally, such software may be described using any of a plurality of suitable programming languages and / or programming or scripting tools, and may be compiled as executable machine language code or intermediate code to be executed on a framework or virtual machine. Typically, the functionality of program modules may be combined or distributed as desired in various embodiments.

[0132] Embodiments of the present disclosure may be embodied as a method, and an example thereof is provided. The order of operations performed as part of this method may be determined in any suitable manner. Accordingly, embodiments may be configured to perform operations in an order different from that illustrated, which may include performing some operations simultaneously, even though in the exemplary embodiments they are shown as a series of operations. Although the present disclosure has been described with reference to several preferred embodiments, it should be understood that various other adaptations and modifications can be made within the spirit and scope of the present disclosure. Accordingly, it is the aspect of the appended claims to cover all such variations and modifications that fall within the true spirit and scope of the present disclosure.

Claims

1. A method for controlling a robot for performing a task, the method using a processor coupled to a memory storing a set of dynamic movement primitives (DMPs) associated with the task, the set of DMPs including a set of at least two dynamic systems, the set of at least two dynamic systems including at least a function representing point attractor dynamics associated with the task and a forcing function associated with a learned demonstration of the task, the processor being coupled to stored instructions that, when executed by the processor, perform the steps of the method, the method comprising: obtaining the set of DMPs associated with the task; converting the obtained set of DMPs into a set of constrained DMPs (CDMPs) by defining a perturbation function associated with the learned forcing function, the perturbation function being associated with a set of motion constraints that need to be satisfied for execution of the task, the method further comprising: solving a non-linear optimization problem for the set of CDMPs based on the motion constraints; generating a control input for controlling the robot for performing the task based on a solution to the non-linear optimization problem for the set of CDMPs. A method.

2. The method according to claim 1, wherein the set of at least two dynamic systems includes a combination of ordinary differential equations (ODEs) representing each dynamic system in the set of at least two dynamic systems.

3. The method according to claim 1, wherein the function of the point attractor dynamics includes parameters associated with the starting pose of the robot and the target pose of the robot.

4. The method according to claim 1, wherein the forcing function includes one or more weights corresponding to a set of basis functions associated with the task, the one or more weights being adjustable parameters associated with the learned demonstration of the task.

5. The method according to claim 4, wherein the one or more weights are learned by solving a locally weighted regression to learn the one or more weights for the basis functions.

6. The step of converting the set of DMPs into the set of CDMPs by defining a perturbation function associated with the set of motion constraints comprises: Defining the perturbation function by adding an additional parameter set representing the operation constraints in the learned forced function, wherein the new task constraint is represented using the additional parameter set as a barrier function that is at least once differentiable, the method according to claim 1.

7. The step of solving the non-linear optimization problem for the set of CDMPs based on the operation constraints, comprises determining an executable set of values for the additional parameter set such that the new task constraint represented by the barrier function is satisfied during execution of the task, the method according to claim 6.

8. Determining a solution to the non-linear optimization problem by discovering the executable set of values for the barrier function representing the new task constraint; Determining the perturbation function based on the obtained solution to the non-linear optimization problem; Generating the control input for controlling the robot based on the obtained perturbation function; and further comprising controlling the robot to execute the task based on the generated control input, the method according to claim 7.

9. The task includes operating the robot to follow a trajectory, the method according to claim 1.

10. One or more of the operation constraints are avoiding collision with one or more obstacles present in the environment of the robot, avoiding self-collision, and joint limitations of the end effector of the robot including at least one of them, the method according to claim 1.

11. A controller for controlling the operation of a robot to execute a task, wherein the controller comprises a memory configured to store a set of dynamic movement primitives (DMPs) associated with the task, the set of DMPs including at least two sets of dynamic systems, the at least two sets of dynamic systems including at least a function representing point attractor dynamics associated with the task, and a forced function associated with a learned demonstration of the task, the controller further comprises a processor, and the processor ​ configured to convert the set of DMPs into a set of Constrained DMPs (CDMPs) by obtaining a perturbation function associated with the learned forcing function, the perturbation function being associated with a set of motion constraints that need to be satisfied for the execution of the task, the processor further solving a non-linear optimization problem for the set of CDMPs based on the set of motion constraints, configured to generate a control input for controlling the robot for executing the task based on a solution to the non-linear optimization problem for the set of CDMPs, The controller further comprises an output interface configured to command the robot to execute the generated control input to execute the task. **Claim 12** The controller according to claim 11, wherein the set of at least two dynamic systems comprises a combination of ordinary differential equations (ODEs) representing each dynamic system in the set of at least two dynamic systems. **Claim 13** The controller according to claim 11, wherein the function of the point attractor dynamics comprises parameters associated with the starting pose of the robot and the target pose of the robot. **Claim 14** The controller according to claim 11, wherein the forcing function comprises one or more weights corresponding to a set of basis functions associated with the task, the one or more weights being adjustable parameters associated with the demonstration of the learned task. **Claim 15** The controller according to claim 14, wherein the one or more weights are learned by solving a locally weighted regression to learn the one or more weights for the basis functions. **Claim 16** To convert the set of DMPs into the set of CDMPs by obtaining the perturbation function associated with the set of motion constraints, the processor is configured to add an additional set of parameters representing the motion constraints in the learned forcing function to define the perturbation function, and the new task constraint is represented using the additional set of parameters as a barrier function that is at least once differentiable. **Claim 17** To solve the non-linear optimization problem for the set of CDMPs based on the motion constraints, the processor The controller according to claim 16, configured to obtain an executable set of values for the additional parameter set such that the new task constraint represented by the barrier function is satisfied during execution of the task.

18. The processor further obtains a solution to the non-linear optimization problem by discovering an executable parameter set for the barrier function representing the new task constraint, obtains the perturbation function based on the obtained solution to the non-linear optimization problem, generates the control input for controlling the robot based on the obtained perturbation function, and is configured to output the control input for controlling the robot to execute the task based on the generated control input. The controller according to claim 17.

19. The task includes operating the robot to follow a trajectory. The controller according to claim 11.

20. One or more of the motion constraints include avoidance of collision with one or more obstacles present in the environment of the robot, avoidance of self-collision, and joint limitations of the end effector of the robot The controller according to claim 11, including at least one of.

21. A non-transitory computer-readable medium storing computer-executable instructions for a robot to perform a task, the computer-executable instructions configured to obtain a set of dynamic DMPs associated with the task, the obtained set of DMPs including at least two sets of dynamic systems, the at least two sets of dynamic systems including at least a function representing point attractor dynamics associated with the task and a forcing function associated with a learned demonstration of the task, the computer-executable instructions further configured to convert the obtained set of DMPs into a set of constrained DMPs (CDMPs) by defining a perturbation function associated with the learned forcing function, the perturbation function being associated with a set of motion constraints that need to be satisfied for execution of the task, the computer-executable instructions further solving a non-linear optimization problem for the set of CDMPs based on the set of motion constraints. A non-transitory computer-readable medium configured to generate a control input for controlling the robot for performing the task based on a solution to the non-linear optimization problem with respect to the set of CDMPs.

Citation Information

Patent Citations

  • System and method for quick scripting of tasks for autonomous robotic manipulation

    US9486918B1

  • Method and system for trajectory optimization for nonlinear robotic systems with geometric constraints

    WO2021065196A1