System and method for controlling a robot using constrained dynamic motion primitives
By applying perturbation functions to DMPs to internalize constraints, the method ensures robot trajectories adhere to environmental constraints, addressing the limitations of conventional DMPs in adapting to new environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MITSUBISHI ELECTRIC CORP
- Filing Date
- 2023-06-16
- Publication Date
- 2026-04-24
AI Technical Summary
Conventional Dynamic Movement Primitive (DMP)-based techniques struggle with incorporating environmental constraints during robot trajectory generation, leading to potential collisions or joint limit violations when environments change from those in which tasks were demonstrated.
Adapt the coercive function used in DMPs by applying perturbation functions to internalize constraints without retraining, optimizing additional parameters to satisfy new environmental conditions through nonlinear optimization.
Enables constrained robot control by ensuring trajectories adhere to environmental constraints, allowing reuse of legacy DMP-based control methods in new environments efficiently and safely.
Smart Images

Figure 0007851494000024 
Figure 0007851494000025 
Figure 0007851494000026
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the operation and movement of robots, and more specifically, to methods and systems for the operation and movement of constrained robots.
Background Art
[0002] For performing various tasks such as material handling, transportation, welding, assembly, etc., various types of robotic devices have been developed. Learning from demonstrations, also known as demonstration-based robot programming, is a commonly used technique to enable robots to autonomously perform new tasks. These techniques are based on the premise that an appropriate robot controller can be derived by observing a human performing the desired task himself or observing a human remotely operating the robot for the desired task. Dynamic Movement Primitives (DMP)-based techniques are widely used to learn skills that can be demonstrated to a robot by an expert human or a controller. DMP is a formulation of a non-linear dynamic system that can learn complex trajectories from demonstrations by decoupling a non-linear forcing function from a nominal attracting behavior. DMP-based techniques offer generalization ability and a simple formulation. Furthermore, DMP can be reparameterized by its start and goal positions.
[0003] However, directly controlling the trajectory performed between the starting and target positions (orientations) can be difficult. Therefore, DMP-based techniques can lead to undesirable behavior when constraints are present. Without careful consideration, a new starting or target position could cause the robot to collide with itself or the environment (i.e., obstacles), or extend beyond the robot's joint limits. For example, if the environment in which the task is performed changes from the environment in which the task is demonstrated, potentially creating a collision, conventional DMP-based techniques tend to modify the estimated trajectory based on skills learned from the demonstration. However, such modifications are computationally expensive, impractical, and even impossible when the generated trajectory significantly violates constraints.
[0004] Therefore, there is a need for systems and methods to incorporate constraints into tasks such as robot trajectory generation using learning from demonstrations in an efficient and effective manner. [Overview of the project]
[0005] The objective of some embodiments is to provide systems and methods for constrained control of robots. Specifically, the objective of some embodiments is to provide systems and methods for constrained control of robotic manipulators using skills learned from demonstrations. In addition to or instead of this, the objective of some embodiments is to incorporate constraints during task performance by a robot using learning from demonstrations, for example, during robot trajectory generation.
[0006] In addition to or instead of the above, an objective of some embodiments is to provide such systems and methods that can provide constrained control of a robot, specifically constrained control of a robotic manipulator, when there are different kinds of constraints, such as collision avoidance with the environment, joint limitations, and self-collision.
[0007] Some embodiments are based on the understanding that the skills demonstrated by the robot are performed offline in an environment that may differ from the actual environment in which the robot is controlled. Furthermore, in the DMP framework, learned skills are captured using forcing functions that represent the forces acting on the dynamic system during demonstration, allowing the dynamic system to follow the demonstrated trajectory. The forcing function is a mathematical representation of the skills demonstrated by the robot along the demonstrated trajectory. While these forces can be applied to the environment, once the forcing function is learned, it becomes environment-independent. Therefore, a natural way to consider constraints lies outside of the environment-independent forcing function. Moreover, retraining forcing functions for different environments of robot operations is impractical.
[0008] However, some embodiments are based on the recognition that the coercive function can be adapted to different environments without the need to retrain it. Such adaptations allow the constraints to be internalized to the skill rather than being corrected outside the skill. To do so, some embodiments aim to find the smallest correction to predetermined weights of the basis functions that form the coercive function such that the new coercive function satisfies the constraints in the new environment.
[0009] Such a formulation differs from learning new weights because weights define the skill, while corrections to weights define the fit of the skill to the environment and / or the new environment. Thus, unknown corrections depend on the environment but not on the skill itself, simplifying the fit and allowing for the consideration of different types of constraints. In addition, by discovering the corrections, constraints are incorporated into the forcing function, so the constraints become inherent in the skill without having to retrain the process. In this way, different legacy methods for DMP-based control can be reused in new environments.
[0010] For example, skill adjustments to adapt skills to the environment can be expressed as additional parameters in the original DMP formulation. These additional parameters can be optimized using optimization problems that can be solved using commercially available solvers.
[0011] Therefore, the additional parameters in the original formulation of the DMP define corrections to the basis function weights that can be estimated using optimization methods, and thus the DMP can satisfy certain constraints. This allows the DMP to be transformed into a constrained DMP by using the weight corrections as perturbations in the original forcing function, and this transformation can be represented by a nonlinear optimization problem. By solving the nonlinear optimization problem, these perturbations are identified, and then the control inputs for the robot to perform the task under the constraints are determined.
[0012] Therefore, in one embodiment, a method for controlling a robot to perform a task is disclosed. The method is executed by a memory-coupled processor, which stores a set of dynamic motion primitives (DMPs) associated with a task, the set of DMPs comprising a set of at least two dynamic systems, the set of at least two dynamic systems comprising at least a function representing point attractor dynamics associated with the task and a forced function associated with the learned task performance. The processor comprises stored instructions that, when executed by the processor, perform the steps of the method, the method comprising the step of obtaining a set of DMPs associated with a task. The obtained set of DMPs is then converted into a set of constrained DMPs (CDMPs) by defining perturbation functions associated with the learned forced functions, the perturbation functions associated with a set of motion constraints that must be satisfied for the task to be performed. Furthermore, a nonlinear optimization problem is solved for the set of CDMPs based on the motion constraints. Furthermore, based on the solution to the nonlinear optimization problem for the set of CDMPs, control inputs are generated for controlling a robot to perform a task.
[0013] According to another embodiment, a controller is provided for controlling the movement of a robot to perform a task. The controller includes a memory configured to store a set of dynamic motion primitives (DMPs) associated with a task, the set of DMPs including a set of at least two dynamic systems. The set of two dynamic systems includes at least a function representing the point attractor dynamics associated with the task and a forcing function associated with the learned task performance. The controller further includes a processor configured to transform the set of DMPs into a set of constrained DMPs (CDMPs) by finding perturbation functions associated with the learned forcing functions, the perturbation functions associated with a set of motion constraints that must be satisfied for the task to be performed, and the processor is further configured to solve a nonlinear optimization problem for the set of CDMPs based on the set of motion constraints, and to generate control inputs for controlling the robot to perform the task based on the solution to the nonlinear optimization problem for the set of CDMPs. The controller further includes an output interface configured to execute the generated control inputs to instruct the robot to perform the task.
[0014] In yet another embodiment, a non-temporary computer-readable medium is disclosed for storing computer-executable instructions for a robot to perform a task. The computer-executable instructions may be configured to acquire a set of DMPs associated with a task, the set of DMPs comprising a set of at least two dynamic systems. The set of at least two dynamic systems comprises at least a function representing the point attractor dynamics associated with the task and a forcing function associated with the learned task performance. The computer-executable instructions may further be configured to transform the set of DMPs into a set of constrained DMPs (CDMPs) by defining perturbation functions associated with the learned forcing functions, the perturbation functions being associated with a set of behavioral constraints that must be satisfied for the task to be performed. The computer-executable instructions may further be configured to solve a nonlinear optimization problem for the set of CDMPs based on the behavioral constraints. In addition, the computer-executable instructions may be configured to generate control inputs for controlling a robot to perform a task based on the solution to the nonlinear optimization problem for the set of CDMPs.
[0015] It should be understood that the general description above and the detailed description below are for illustrative and explanatory purposes only and do not limit the invention described in the claims.
[0016] The embodiments disclosed herein will be further described with reference to the accompanying drawings. The drawings shown are not necessarily to exact scale and are instead exaggerated in general to illustrate the principles of the embodiments disclosed herein. [Brief explanation of the drawing]
[0017] [Figure 1A] This is a block diagram of a system including a controller for controlling the movement of a robot to perform a task, according to an embodiment of the present disclosure. [Figure 1B] This is a block diagram of the controller shown in Figure 1A according to an embodiment of the present disclosure. [Figure 2A] Schematic diagram showing a set of DMPs used by the controller of FIG. 1B, according to an embodiment of the present disclosure. [Figure 2B] Diagram showing a mathematical formulation corresponding to a dynamic system represented by a set of DMPs of FIG. 2A, according to an embodiment of the present disclosure. [Figure 2C] Diagram showing a mathematical formulation corresponding to a forcing function represented by a set of DMPs of FIG. 2A, according to an embodiment of the present disclosure. [Figure 2D] Diagram showing a mathematical formulation corresponding to a basis function represented by a set of DMPs of FIG. 2A, according to an embodiment of the present disclosure. [Figure 3A] Block diagram showing the conversion from a set of DMPs to a set of constrained DMPs (CDMPs), according to an embodiment of the present disclosure. [Figure 3B] Block diagram showing the conversion from a set of DMPs to a set of constrained DMPs (CDMPs) using a perturbation function, according to an embodiment of the present disclosure. [Figure 3C] Diagram showing a mathematical expression demonstrating the use of a barrier function in the perturbation function of FIG. 3B, according to an embodiment of the present disclosure. [Figure 3D] Diagram showing a mathematical formulation of a non - linear optimization problem solved to obtain the value of a barrier function, according to an embodiment of the present disclosure. [Figure 4A] Flow diagram of a method executed by a controller to execute a task by a robot, according to an embodiment of the present disclosure. [Figure 4B] Flow diagram of another method executed by a controller to execute a task by a robot, according to an embodiment of the present disclosure. [Figure 5] Block diagram of an exemplary computer system for implementing various embodiments. [Figure 6A] Schematic diagram of a use case of a robot system based on the controller of FIG. 1A in a first set of conditions, according to an embodiment of the present disclosure. [Figure 6B]Schematic diagram of a use case of a robot system based on the controller of FIG. 1A in a second set of conditions according to an embodiment of the present disclosure.
Embodiments for Carrying Out the Invention
[0018] Description of Embodiments
[0019] In the following description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, devices and methods are shown in block diagram form only to avoid obscuring the present disclosure.
[0020] As used in this specification and the claims, each of the terms “for example,” “as an example,” and “such as,” and the verbs “comprising,” “having,” “including,” and other verb forms thereof, when used in conjunction with a listing of one or more components or other items, should be construed as open-ended, meaning that the listing should not be considered as excluding other additional components or items. The term “based on” means at least partially based on. Further, it should be understood that the style and terminology used in this specification are for the purpose of description and should not be regarded as limiting. Any headings used in this specification are for convenience only and have no legal or limiting effect.
[0021] The various embodiments disclosed herein provide a conversion from DMP-based techniques to CDMP-based techniques for realizing constraints in tasks performed by robots. DMP-based techniques are based on learned skills captured during task demonstration by a skilled user using coercive functions that represent forces acting on the dynamic system representing the task. This enables the robot to follow the demonstrated task. The coercive function is a mathematical representation of the skills demonstrated to the robot in the demonstrated task. These forces can be applied to the environment, but once the coercive function is learned, it becomes environment-independent. In other words, the constraints are outside the environment-independent coercive function. Retraining the robot's coercive function for different environments is impractical.
[0022] Some embodiments are based on the understanding that the coercive function can be adapted to different environments without the need to retrain it. Such adaptations allow the constraints to be internalized to the skill rather than being corrected outside the skill. Therefore, the objective of this disclosure is to find the minimum correction to the predetermined weights of the basis functions forming the coercive function such that the new coercive function satisfies the constraints in the new environment.
[0023] Some embodiments further base their approach on the understanding that the above method is different from learning new weights, because while weights define the skill, corrections to weights define the adaptation of the skill to the environment and / or a new environment. Thus, unknown corrections depend on the environment but not on the skill itself. This simplifies the adaptation and allows for the consideration of different types of constraints. In addition, by discovering the corrections, constraints are incorporated into the internal coercion function, so that constraints become inherent in the skill without retraining the process. In this way, different legacy methods for DMP-based control can be reused in new environments. For example, corrections to the skill for adapting the skill to a (new) environment can be expressed as additional parameters defined by a perturbation of the coercion function in the original DMP formulation to the original formulation. These additional parameters can be optimized using an optimization problem that can be solved using any known solver. The resulting formulation is a constrained DMP or CDMP, which provides constrained control of a robot to perform a task in an environment.
[0024] Figure 1A shows a block diagram of a system 100, including a controller 101 for controlling the movement of a robot 102 to perform a task, according to an embodiment of the present disclosure.
[0025] Robot 102 may include a robotic arm or robotic manipulator, etc., which is intended to perform a task. Tasks may include lifting an object, moving an object, positioning an object in a desired location, or moving an object from one location to another. Therefore, robot 102 may need to follow a motion trajectory to perform the desired task. For example, robot 102 may be a food-distributing robot configured to place various food items in designated locations within a box or shipping carton.
[0026] In another example, robot 102 is an assembly line robot used to lift and move objects within an industrial or manufacturing unit, such as between different machines, for purposes such as transporting objects or removing defective manufacturing units.
[0027] Another example of a task is for the robot 102's end effector to move from an initial position in 3D space to another position in 3D space. Another example of a task is for the gripper end effector to open, move to a position in 3D space, close the gripper to grasp an object, and move to a final position in 3D space.
[0028] Generally, a task has a start condition and a finish condition called a task objective. When the task objective is achieved, the task is considered complete. A task can be divided into subtasks. For example, if the task is for robot 102 to move to a 3D location in Cartesian space and then pick up an object, the task can be broken down into the subtasks (1) move to the 3D location and (2) pick up the object. It is understood that more complex tasks can be broken down into subtasks by a program. When a task is broken down into subtasks, a task description may be provided for each subtask. In one embodiment, the task description may be provided by a human operator. In another embodiment, a program is executed to obtain the task description. The task description may also include constraints on robot 102. An example of such a constraint is that robot 102 cannot move faster than a speed specified by a human operator. Another example of a constraint is that robot 102 is prohibited from entering a portion of the 3D Cartesian workspace specified by a human operator. The objective of robot 102 is to complete the task specified in the task description as quickly as possible, given any constraints in the task description.
[0029] The robot 102 may include one or more sensors that can be configured to acquire sensor data associated with one or more obstacles or objects present in the robot 102's environment, or with movements performed by the robot 102 itself. For example, one or more sensors may include vision sensors, i.e., cameras. In some embodiments, the robot 102 may be communicably coupled to one or more sensors via a communication network.
[0030] Furthermore, the robot 102 is configured to receive one or more control inputs generated by the controller 101. The control inputs are configured to cause the robot 102 to perform a desired task. Therefore, the control inputs may also be received by one or more actuators of the robot 102, which further generate signals to control different parts of the robot, such as the robotic arm of the robot 102, to move the robot 102 along a desired motion trajectory and thus perform the task. The robot 102 receives different control inputs for different tasks from the controller 101, which is configured to generate these different control inputs. The controller 101 controls the motion of the robot 102 to complete the task by sending commands to a physical robotic system such as the robot 102. In another embodiment, the controller 101 is integrated into the robot 102.
[0031] Figure 1B shows a block diagram of a controller 101 for controlling the movement of the robot 102 to perform a task. The controller 101 includes memory 103, a processor 105, and an output interface 106. In addition, the controller 101 may also include input interfaces and other components required for the controller 101 to perform the operations described herein, without departing from the scope of this disclosure.
[0032] Memory 103 is configured to store computer executable instructions that can be executed by processor 105. Processor 105 can be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. Memory 103 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. Processor 105 is connected to one or more input and output devices via a bus.
[0033] Memory 103 is also configured to store a set of dynamic motion primitives (DMPs) 104 associated with a task. Thus, memory 103 may be configured to store DMPs 104 representing the original task performed by the robot 102. For example, the original task might include the motion of the robot 102's end-effector along an original trajectory. These DMPs 104 may be acquired through demonstrations performed by a human operator during the robot 102's learning phase, in which the robot 102 learns various parameters associated with a task. As those skilled in the art will understand, DMP-based techniques are widely used to teach a robot skills that can be demonstrated by a skilled person or controller. Thus, using demonstration-learning techniques, the characteristics of the original demonstration are learned by the robot 102. These characteristics include how the robot's end-effector moved during the task demonstration.
[0034] DMP104 is a formulation of a nonlinear dynamic system that allows for the learning of complex tasks (such as trajectories) from demonstrations by decoupling the nonlinear forcing function from the nominal pull-in behavior. In other words, DMP104 is a method of task control and planning. Thus, DMP104 is a proposed mathematical formalization of complex subtasks, such as the motion of a trajectory, consisting of a set of primitive action "building blocks" that are executed sequentially and / or in parallel. To be understood further, since DMP104 is a nonlinear dynamic system, DMP104 differs from building blocks proposed to date.
[0035] For example, in the case of trajectory planning, the DMP 104 can be understood as a combination of two systems: a virtual system that plans the trajectory, and a real system that actually executes the planned trajectory. The DMP 104 may include its own dynamics, and once the DMP is set up, it can obtain control signals that the robot 102 should follow. These control signals form the control inputs generated by the controller 101 for the robot 102 to perform the desired task. For example, the DMP control signals for the path (i.e., trajectory) that the robot 102's end effector (also called the robot manipulator) should follow may include a set of forces that need to be applied to the end effector in order to carry out that path. The robot 102 may apply these forces by converting them into joint torques.
[0036] For the task of orbit planning, the orbit is a function y d It can be represented by (t), where t lies in the set [0, T]. The trajectory may include multiple waypoints (orientation and attitude) of the end effector in Cartesian space. Once one or more trajectories y(t) are recorded for a fixed orientation, the DMP learning algorithm can learn a DMP for each component of y(t). To eliminate explicit time dependency, DMP104 uses a canonical system to track the progress of the task.
[0037] Therefore, DMP104 is in the form of two combined sets of parameterized ordinary differential equations (ODEs) that represent the task. For example, DMP104 can generate a trajectory that moves a system such as robot 102 from a starting position to a target position. DMP104 can easily adapt the trajectory according to the new starting and target states, and thus essentially constitutes a closed-loop controller. Furthermore, DMP104 can learn from a limited number of training examples, including even a single training example. Thus, it is possible to modify the original trajectory in response to changes in the starting and target positions. The formulation of DMP104 is further explained in Figure 2A.
[0038] Figure 2A shows a schematic diagram of a set of DMPs 104 used by the controller 101 of Figure 1B according to an embodiment of the present disclosure. The set of DMPs 104 includes a set of at least two dynamic systems, a first dynamic system 107 and a second dynamic system 108. The second dynamic system 108 includes a function representing the point attractor dynamics 109 associated with a task, and a forcing function 110 associated with the performance of the learned task.
[0039]
number
[0040] Furthermore, a second dynamic system 108 is shown in Figure 2B.
[0041]
number
[0042] For example, a set of DMPs 104 is associated with point attractor dynamics 109 parameterized by the starting and target spatial attitudes of the robot 102, and a forcing function 110 containing one or more weights corresponding to basis functions associated with the robot 102's original trajectory. The original trajectory may consist of multiple original spatial points between the starting and ending spatial attitudes. Note that the one or more predetermined weights may consist of a first configuration, depending on one or more demonstrations.
[0043] Furthermore, the forcing function 110 is represented by the term f(x,g) shown in Figure 2C.
[0044]
number
[0045]
number
[0046] Therefore, the forcing function 110 is a tunable parameter w i This includes the weights 111 of the weighted combination, which are learned based on the task demonstration. In addition, one or more weights 111 are learned by solving local weighted regression to learn one or more weights 111 for basis function 112. In the example of the robot 102 trajectory generation task, the forcible function 110 f(x,y) is learned from the task demonstration trajectory. By learning f(x,y) from the demonstration, the characteristics of the original demonstration (e.g., how the robot's end effector moved during the task) are learned. These are learned, for example, in the learning phase by solving local weighted regression to learn the weights for basis function 112.
[0047] In particular, the forcing function 110 may include weighted combinations of basis functions 112. For example, basis functions 112 can be radial basis functions as shown in Figure 2D.
[0048]
number
[0049] Referring again to Figure 1B, the controller 101 further includes a processor 105 configured to execute one or more computer executable instructions. One or more computer executable instructions cause the processor 105 to perform one or more operations. Thus, the execution of one or more operations is configured to convert a set of DMPs 104 into a set of constrained DMPs (CDMPs).
[0050] This disclosure provides a constrained DMP (CDMP) that includes an additional function representing a perturbation to the original coercion function 110, so that the learned DMP 104 can satisfy new behavioral constraints. Note that the behavioral constraints may depend on the environment in which the robot 102 performs the task. Therefore, different environments may have different constraints, and thus the function may depend only on constraints that need to be satisfied during the operation of the robot 102 in the new environment.
[0051] Figure 3A shows a block diagram illustrating a conversion 113 from a set of DMPs 104 to a set of constrained DMPs (CDMPs) 114, according to an embodiment of the present disclosure.
[0052] The transformation 113 is performed by defining a perturbation to the initially learned coercion function 110, which is associated with a set of behavioral constraints 115 related to the task performed by the robot 102. The set of behavioral constraints 115 may be defined by the environment in which the robot 102 operates. More specifically, the perturbation defines an additional set of parameters that represent the behavioral constraints 115 in the learned coercion function 110. The additional set of parameters represents the new task constraints. Thus, the success of task execution is determined based on the satisfaction of the constraints included in the set of behavioral constraints 115.
[0053] For example, the set of motion constraints 115 may include collision avoidance constraints. Collision avoidance constraints may include conditions to ensure that the robot 102 does not collide with obstacles that may be present in the robot 102's environment.
[0054] In another example, the set of motion constraints may include self-collision avoidance constraints. Self-collision avoidance constraints may include conditions to ensure that robot 102 does not collide with itself.
[0055] In yet another example, the set of motion constraints 115 may include constraints on the joint limits of the robot 102's end effector. For example, consider a task in which a learned DMP is reparameterized in a new environment based on a new goal and starting state. However, the new trajectory for this task may not be able to perform what is done to the robot 102 because some of the joint rotations may be kinematically impossible. Such constraints may be explicitly considered in the formulation of the CDMP 114, and thus a viable solution that satisfies the joint limits may be found.
[0056] The set of behavioral constraints 115 is introduced in the form of adding perturbation terms to the learned forcing function 110. As a result of transformation 113, the original set of DMPs 104 is transformed into a set of CDMPs, as shown in Figure 3B.
[0057]
number
[0058] Furthermore, the perturbation function 116 includes a novel task constraint expressed by an additional set of parameters as a barrier function that is at least once differentiable, as shown in Figure 3C.
[0059]
number
[0060]
number
[0061] Figure 3D shows the barrier function 117 ζ i The mathematical formulation of the nonlinear optimization problem 118, which is solved to find the value of , is shown.
[0062]
number
[0063] Next, using the solution to the nonlinear optimization problem 118 for the set of CDMPs based on the set of motion constraints 115, we identify a set of feasible values for an additional set of parameters for a new task to be performed by the robot 102, where the barrier function 117 represents the new task constraints. The set of feasible values corresponds to the values of the additional set of parameters that satisfy the constraints associated with the task to be performed, such as the new task constraints mentioned here, under a given set of environmental conditions. Thus, the set of feasible values includes the set of all possible points in the nonlinear optimization problem 118 that satisfy the constraints of the problem, potentially including inequalities, equalities, and integer constraints. Furthermore, we obtain a perturbation function 116 using the solution to the nonlinear optimization problem 118, and then use the perturbation function 116 to generate one or more control inputs for controlling the robot 102 to perform the task, based on the solution to the nonlinear optimization problem 118 for the set of CDMPs 114.
[0064]
number
[0065] In one example, the additional constraint 119 specifies a limit on the amount of deviation of CDMP114 from the original DMP104. The design of CDMP114 can be viewed as a trade-off between constraint satisfaction and the original forcing function 110. This trade-off can be controlled using a hyperparameter that constrains the maximum allowable deviation between the original DMP104 and CDMP114, and this hyperparameter can be added as an additional constraint 119 to the nonlinear optimization problem 118.
[0066] The solution to the nonlinear optimization problem 118 may be obtained using any known commercially available solver such as IPOPT(trademark) or SNOPT. Thus, the conversion from DMP104 to CDMP114 disclosed herein provides a highly cost-effective, computationally efficient, easily implementable, and practical solution to the problem of robot motion control in a constrained environment. Furthermore, since CDMP114 is based on satisfying motion constraints 115, including safety and collision avoidance conditions, the overall operation of the robot 102 becomes highly safe and efficient.
[0067] Furthermore, the solution to the optimization problem 118 determines the corrected weights for the modified forcing function 110, which are converted into a control input and sent to the robot 102 via the output interface 106 to instruct the robot 102 to execute the task by performing the generated control input.
[0068] Figure 4A shows a flowchart of a method performed by the controller 101 to allow the robot 102 to perform a task, according to an embodiment of the present disclosure.
[0069] In step 401, the set of DMPs associated with the task is retrieved. For example, the set of DMPs 104 stored in memory 103 is retrieved by the processor 105.
[0070] Next, in 402, the set of DMPs is transformed into a set of CDMPs by defining perturbations to the initially learned forcibly function associated with the task. For example, DMP 104 includes the initially learned forcibly function 110 f(x,g), which is transformed by defining a perturbation function 116 g(x) that is added as an additional parameter to the learned forcibly function 110 f(x,g) (113). The perturbation function 116 g(x) is associated with behavioral constraints 115 that must be satisfied for the task to be performed.
[0071] To find the values of the perturbation function 116, the nonlinear optimization problem 118 is formulated in 403 by using the formulation of the CDMP set 114 and motion constraints 115, which include a new set of task constraints defined by the changing environment of the robot performing the task, and is solved by the processor 105. These new task constraints are represented as an additional set of parameters in the perturbation function 116. Thus, Figure 4B shows a flowchart of another method performed by the controller 101 to find the solution to the nonlinear optimization problem 118. In 405, a feasible set of values is found for the barrier function 117, which represents the new task constraints for the motion constraints 115. This is discussed together with Figure 3D.
[0072] Next, in 406, the value of the perturbation function 116 is determined by using the value of the barrier function 117 and the modified forcing function given in Equation 7. ru.
[0073] Furthermore, in 407, the solution to the perturbation function 116 is used to generate the control input for the controller 101 to control the robot 102. The method shown in Figures 4A and 4B is illustrated using the exemplary trajectory generation task, as given below.
[0074] In this example, the original trajectory may be based on acquiring multiple movements performed by an end effector associated with the robot 102 using sensors, depending on the demonstration. For example, the sensors may be encoders and / or vision sensors. These movements are used to generate DMPs 104 corresponding to the movements performed by the end effector associated with the robot 102. These DMPs are then stored in the memory 103 of the controller 101 and retrieved in step 401 described above.
[0075] As can be understood, DMP104 may cause undesirable behavior if there are different operating constraints 115 than those during demonstration. Since DMP104 does not explicitly consider additional disturbances and simply learns the forcing function during demonstration, the resulting generalization may only be acceptable during demonstration. However, if there is no fit of the forcing function during actual operation, the resulting robot trajectory may be unfeasible during actual control. For example, if the environment changes, the robot 102 may collide with an object that the robot 102 is manipulating, due to a new task constraint on the robot's acceptable states that does not exist during operation, thereby causing the object the robot 102 is manipulating to go beyond the robot's range as defined by the physical constraints on the robot's structure.
[0076] The motion constraints 115 may include additional or different obstacles present in the robot 102's environment compared to the previous environment. Therefore, the motion constraints 115 may include the position and configuration associated with obstacles present in the robot 102's environment in the changed environment. For example, the position and configuration may be obtained using a vision sensor employing posture estimation techniques. Alternatively, or in addition to this, the robot 102 may have joint constraints on the robot's end effectors.
[0077] Therefore, these additional constraints may form new task constraints that can be used in step 402 to define a perturbation 116 to the original forcing function 110, thereby reconstructing in a second configuration one or more predetermined weights 111 associated with the original trajectory of the original forcing function 110 based on one or more behavioral constraints 115. In some embodiments, reconstructing one or more predetermined weights 111 may involve transforming one or more behavioral constraints 115 into one or more differentiable functions using a smoothing function or barrier function 117. For example, the smoothing function may be a control barrier function (CBF).
[0078] In one example, the barrier function 117 is a zeroing barrier function (ZBF). One advantage of using ZBF to represent constraints is the generality it provides; that is, ZBF can prove joint restriction avoidance and obstacle avoidance. Furthermore, the above formulation results in a nonlinear optimization that perturbs the DMP forced weights regressed by local weighted regression to validate the user-constructed ZBF, and this nonlinear optimization can be solved using a standard NLP solver. CDMP is subject to different constraints on the movement of end-effectors, such as collision avoidance, and state constraints for safety. ZBF is used to represent smooth task-specific constraints. An additional set of parameters that can be optimized to satisfy these constraints is added to the resulting formulation. The resulting optimization is cast as a nonlinear program, which can be solved using a commercially available nonlinear program solver such as IPOPT. By utilizing the set invariance of ZBF, constraint satisfaction for CDMP is guaranteed. Next, one or more weights can be optimized using ZBF.
[0079]
number
[0080]
number
[0081]
number
[0082]
number
[0083]
number
[0084]
number
[0085] One or more predetermined weights are optimized using one or more differentiable functions. This is done by formulating and solving a nonlinear optimization problem in step 403.
[0086] CDMP incorporates an existing DMP with a forcing function learned from expert trajectories, and then optimizes this forcing term so that the DMP dynamic system approves a ZBF that proves the DMP generates trajectories that remain within a safe set of workspaces. This safe set (and ZBF) is constructed by constructing a signed distance field from primitive convex polyhedra. More specifically, the resulting nonlinear optimization problem is described as follows:
[0087]
number
[0088]
number
[0089]
number
[0090] In the formula, ζ i This is the optimized decision variable for the formulated optimization problem. Note that this is only one possible way of representing the perturbation of the original DMP104 forcible function 110, based on the most common representation of radial basis functions in the DMP literature.
[0091]
number
[0092] The problem formulation represented by [Equations 11-14] is transformed into a finite-dimensional discretization problem. Next, this can be solved using a nonlinear optimization solver such as IPOPT or SNOPT to generate the desired parameter set.
[0093] Furthermore, a new trajectory is generated based on one or more predetermined weights configured in the second configuration. The new trajectory may contain multiple new spatial points between the start and end spatial points. It should be noted that at least one of these multiple new spatial points may be different from the multiple original spatial points.
[0094] In some embodiments, generating a new trajectory may involve formulating a nonlinear dynamically constrained optimization problem 118 using one or more predetermined weights associated with the original trajectory and one or more differentiable functions 117 corresponding to one or more operating constraints 115. Furthermore, generating a new trajectory may involve solving the nonlinear dynamically constrained optimization problem 118 by optimizing one or more predetermined weights for radial basis functions such as basis function 112 to generate a new trajectory. The new trajectory satisfies one or more operating constraints 115 for performing the task. Therefore, determining a new trajectory involves determining a correction term 116 that modifies at least some of the weights 111 of the forcing function 110 such that DMP 104, which has a forcing function with corrected weights, represents a new viable trajectory that satisfies the operating constraints 115. In some embodiments, one or more predetermined weights may be optimized using a gradient-based solver, which is based on interior-point methods such as IPOPT (Interior Point OPTimizer) or SNOPT (Sparse Nonlinear OPTimizer).
[0095] In some embodiments, hyperparameters corresponding to the deviation of the new orbit from the original orbit may be specified to generate the new orbit. Furthermore, the degree of deviation of the new orbit from the original orbit may be limited based on the hyperparameters.
[0096]
number
[0097]
number
[0098]
number
[0099] In this way, the controller 101 can be used to perform motion control of the robot 102 to perform various tasks, such as trajectory generation and optimization tasks under new constraints, as described above. As those skilled in the art will understand, the trajectory generation examples described herein are for illustrative purposes only. Any equivalent examples can be used to carry out the principles of the various embodiments disclosed herein without departing from the scope of this disclosure.
[0100] The controller 101 may be implemented as being located within the robot 102. The controller 101 may be any general-purpose or dedicated computer system known in the art. An example of such a computer system is shown in Figure 5.
[0101] Figure 5 is a block diagram 500 of an exemplary computer system for realizing various embodiments. The disclosed methods and systems may be implemented on conventional or general-purpose computer systems such as personal computers (PCs) or server computers. Referring here to Figure 5, a block diagram 500 of an exemplary computer system 502 for realizing various embodiments is shown. The controller 101 may be implemented using the computer system 502. Alternatively, the computer system 502 may be a controller. The computer system 502 may include a central processing unit ("CPU" or "processor") 504. The processor 504 may include at least one data processor for executing program components for performing user-generated or system-generated requests. The processor 504 may be equivalent to the processor 105 shown in Figure 1B. The user may include a person, a person using an apparatus such as those included in this disclosure, or such apparatus itself. The processor 504 may include specialized processing units such as an integrated system (bus) controller, a memory management control unit, a floating-point unit, a graphics processing unit, a digital signal processing unit, and so on. The processor 504 may include microprocessors such as AMD® ATHLON® microprocessors, DURON® microprocessors, or OPTERON® microprocessors, ARM® application, embedded, or secure processors, IBM® POWERPC® processors, Intel® CORE® processors, ITANIUM® processors, XEON® processors, CELERON® processors, or other lines of processors. The processor 504 may be implemented using mainframe, distributed, multicore, parallel, grid, or other architectures. Some embodiments may utilize embedded technologies such as application-specific integrated circuits (ASICs), digital signal processors (DSPs), or field-programmable gate arrays (FPGAs).
[0102] The processor 504 may be configured to communicate with one or more input / output (I / O) devices via the I / O interface 506. The I / O interface 506 may employ, but is not limited to, communication protocols / methods such as audio, analog, digital, mono, RCA, stereo, IEEE-1394, serial bus, Universal Serial Bus (USB), infrared, PS / 2, BNC, coaxial, component, composite, Digital Visual Interface (DVI), High Definition Multimedia Interface (HDMI®), RF antenna, S-video, VGA, IEEE 802.n / b / g / n / x, Bluetooth®, and cellular (e.g., Code Division Multiple Access (CDMA), High Speed Packet Access (HSPA+), Global System for Mobile Communications (GSM), Long-Term Evolution (LTE), WiMAX, etc.).
[0103] The computer system 502 may communicate with one or more I / O devices using the I / O interface 506. For example, the input device 508 may be an antenna, keyboard, mouse, joystick, (infrared) remote control, camera, card reader, facsimile, dongle, biometric reader, microphone, touchscreen, touchpad, trackball, sensor (e.g., accelerometer, light sensor, GPS, gyroscope, proximity sensor, etc.), stylus, scanner, storage device, transceiver, video device / source, visor, etc. The output device 510 may be a printer, facsimile, video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED), plasma, etc.), audio speaker, etc. In some embodiments, a transceiver 512 may be located in relation to the processor 504. The transceiver 512 can facilitate various types of wireless transmission or reception. For example, the transceiver 512 may include an antenna operably connected to a transceiver chip (e.g., a TEXAS® INSTRUMENTS WILINK WL1286® transceiver, a BROADCOM® BCM4550IUB8® transceiver, an INFINEON TECHNOLOGIES® X-GOLD 618-PMB9800® transceiver, etc.) and may provide IEEE 802.6a / b / g / n, Bluetooth, FM, Global Positioning System (GPS), 2G / 3G HSDPA / HSUPA communication, etc.
[0104] In some embodiments, the processor 504 may be configured to communicate with a communication network 514 via a network interface 516. The network interface 516 may communicate with the communication network 514. The network interface 516 may employ connection protocols including, but not limited to, Direct Connect, Ethernet® (e.g., twisted-pair 50 / 500 / 5000-based T), Transmission Control Protocol / Internet Protocol (TCP / IP), Token Ring, IEEE 802.11a / b / g / n / x, etc. The communication network 514 may include, but not limited to, Direct Interconnect, Local Area Network (LAN), Wide Area Network (WAN), Wireless Network (e.g., using Wireless Application Protocol), the Internet, etc. The computer system 502 may communicate with devices 518, 520, and 522 using the network interface 516 and the communication network 514. These devices 518, 520, and 522 may include, but are not limited to, various mobile devices such as personal computers, servers, facsimile machines, printers, scanners, mobile phones, smartphones (e.g., APPLE® IPHONE® smartphones, BLACKBERRY® smartphones, ANDROID®-based phones, etc.), tablet computers, e-readers (AMAZON® Kindle® e-readers, NOOK® tablet computers, etc.), laptop computers, notebooks, and game consoles (MICROSOFT® XBOX® game consoles, NINTENDO® DS® game consoles, SONY® PLAYSTATION® game consoles, etc.). In some embodiments, the computer system 502 itself may embody one or more of these devices 518, 520, and 522.
[0105] In some embodiments, the processor 504 may be configured to communicate with one or more memory devices 530 (e.g., RAM 526, ROM 528, etc.) via a storage interface 524. The storage interface 524 may be connected to memory 530, including but not limited to memory drives, removable disk drives, etc., employing connection protocols such as SATA (serial advanced technology attachment), IDE (integrated drive electronics), IEEE-1394, Universal Serial Bus (USB), Fibre Channel, SCSI (small computer systems interface), etc. The memory drive may further include drums, magnetic disk drives, magneto-optical drives, optical drives, RAID (redundant array of independent discs), solid-state memory devices, solid-state drives, etc.
[0106] Memory 530 may store a collection of program or data repository components, including but not limited to an operating system 532, a user interface application 534, a web browser 536, a mail server 538, a mail client 540, and user / application data 542 (for example, any of the data variables or data records discussed in this disclosure). Memory 530 may be equivalent to memory 103. The operating system 532 can facilitate resource management and operation of the computer system 502. Examples of operating systems 532 include, but are not limited to, the APPLE® MACINTOSH® OS X® platform, the UNIX® platform, Unix-like system distributions (e.g., Berkeley Software Distribution (BSD), FreeBSD, NetBSD, OpenBS, etc.), LINUX distributions (e.g., RED HAT®, UBUNTU®, KUBUNTU®, etc.), the IBM® OS / 2 platform, the MICROSOFT® WINDOWS® platform (XP, Vista / 7 / 8, etc.), the APPLE® IOS® platform, the GOOGLE® ANDROID® platform, and the BLACKBERRY® OS platform. The user interface 534 can facilitate the display, execution, interaction, manipulation, or operation of program components through text or graphic functions. For example, the user interface 534 may provide computer interaction interface elements on a display system operably connected to the computer system 502, such as a cursor, icons, checkboxes, menus, scrollers, windows, widgets, etc.A graphical user interface (GUI) may be employed, including but not limited to the following: the AQUA® platform of the APPLE® Macintosh® operating system, the IBM® OS / 2® platform, the MICROSOFT® WINDOWS® platform (e.g., the AERO® platform, the METRO® platform, etc.), UNIX X-WINDOWS, and web interface libraries (e.g., the ACTIVEX® platform, the JAVA® programming language, the JAVASCRIPT® programming language, the AJAX® programming language, HTML, the ADOBE® FLASH® platform, etc.).
[0107] In some embodiments, the computer system 502 may implement program components stored in the web browser 536. The web browser 536 may be a hypertext browsing application such as the MICROSOFT® INTERNET EXPLORER® web browser, the GOOGLE® CHROME® web browser, the MOZILLA® FIREFOX® web browser, or the APPLE® SAFARI® web browser. Secure web browsing may be provided using HTTPS (Secure Hypertext Transfer Protocol), Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. The web browser may utilize features such as AJAX, DHTML, the ADOBE® FLASH® platform, the JAVASCRIPT® programming language, the JAVA® programming language, or an Application Programming Interface (APi). In some embodiments, the computer system 502 may implement program components stored in the mail server 538. Mail server 538 may be an Internet mail server such as the MICROSOFT® EXCHANGE® mail server. Mail server 538 may utilize functions such as ASP, ActiveX, ANSI C++ / C#, MICROSOFT .NET® programming language, CGI script, JAVA® programming language, JAVASCRIPT® programming language, PERL® programming language, PHP® programming language, PYTHON® programming language, WebObjects, etc. Mail server 538 may utilize communication protocols such as Internet Message Access Protocol (IMAP), Messaging Application Programming Interface (MAPI), Microsoft Exchange, Post Office Protocol (POP), Simple Mail Transfer Protocol (SMTP), etc.In some embodiments, the computer system 502 may implement program components stored in the mail client 540. The mail client 540 may be a mail browsing application such as the APPLE MAIL® mail client, MICROSOFT ENTOURAGE® mail client, MICROSOFT OUTLOOK® mail client, or MOZILLA THUNDERBIRD® mail client.
[0108] In some embodiments, a computer system 502 may store user / application data 542, such as data, variables, records, etc., as described in this disclosure. Such a data repository may be implemented as a fault-tolerant, relational, scalable, and secure data repository, such as an ORACLE® data repository or a SYBASE® data repository. Alternatively, such a data repository may be implemented using standardized data structures such as arrays, hashes, linked lists, structures, structured text files (e.g., XML), tables, or as an object-oriented data repository (e.g., using an OBJECTSTORE® object data repository, a POET® object data repository, a ZOPE® object data repository, etc.). Such a data repository may, depending on the context, be integrated or distributed across the various computer systems discussed above in this disclosure. It should be understood that the structure and operation of any computer or data repository component may be combined, integrated, or distributed in any working combination.
[0109] One operational combination of such a system may include a combination of a computer system 502 acting as a controller 101 for controlling a robot 102 to perform a task based on a prior demonstration of the task performed by a human operator.
[0110] Figure 6A shows a schematic diagram of a use case for robot 602 in a demonstration-based learning process according to an embodiment of this disclosure.
[0111] A human operator 600 performs a task such as having a robot 602 hold a moving object 604 in a gripper 603, move along a track 605, and position the moving object 604 in a fixed position 607 within an inanimate body 606. The demonstration may also relate to an assembly task involving the moving object 604.
[0112] In one embodiment, a human operator 600 may instruct the robot 602 to track the original trajectory 605 using a teaching pendant 601 that stores the coordinates of waypoints corresponding to the original trajectory 605 in the robot 602's memory. The teaching pendant 601 may be a remote control device. The remote control device may be configured to transmit robot configuration settings (i.e., robot settings) to the robot 602 to demonstrate the original trajectory 605. For example, the remote control device transmits control commands such as movement in the XYZ directions, velocity control commands, joint position commands, etc., to demonstrate the original trajectory 605. In an alternative embodiment, the human operator 600 may instruct the robot 602 using a joystick, through kinesthetic feedback, etc. The human operator 600 may instruct the robot 602 to track the original trajectory 605 multiple times for the same fixed posture 607 of the inanimate body 606.
[0113] The robot 602 may be coupled to a controller 602a, such as the controller 101 discussed in the previous embodiment, or may include such a controller 602a. The controller 602a includes a memory that can store DMPs associated with the original trajectory 605 based on demonstrations. These DMPs may include a forcing function learned based on a demonstration of the original trajectory 605 by a human operator 600. Thus, the DMP of the original trajectory corresponds to the set of DMPs 104 shown in the previous figure, specifically Figure 1B. The original trajectory 605 may include a plurality of original spatial points between the starting and ending poses of the robot 602. Furthermore, the original trajectory 605 may be generated based on a demonstration corresponding to a task performed in a first set of conditions. However, during the execution of the task, a second set of conditions or a different environment may be observed by the robot 602. Note that the second set of conditions may differ from the first set of conditions.
[0114] Figure 6B shows an example of robot 602's operation in a different environment than that in Figure 6A, with a second set of conditions. The second set of conditions may include motion constraints that form new task constraints for the environment in Figure 6B, related to obstacles 608 in the path of the original trajectory 605. The location and configuration of obstacles 608 may be determined using one or more sensors, such as a vision sensor. Thus, the location and configuration associated with obstacles 608 present in the robot 602's environment under the second set of conditions may be obtained.
[0115] In this scenario, controller 602a may first acquire a set of DMPs stored in memory, obtained using the demonstration in Figure 6A. Controller 602a may further be configured to convert the acquired set of DMPs into a set of CDMPs by using obstacle avoidance behavioral constraints for obstacles 608 in the original trajectory 605. This conversion may be performed using the conversion 115 described in the previous embodiment. Thus, one or more predetermined weights associated with the original trajectory may be optimized using one or more differentiable functions associated with the behavioral constraints.
[0116] Furthermore, the controller 602a may then dynamically generate a new trajectory 605a based on one or more predetermined weights configured in the second configuration. The new trajectory 605a may include a plurality of new spatial points between the start and end attitudes. It should be further noted that at least one of the plurality of new spatial points may be different from the plurality of original spatial points.
[0117] Therefore, generating a new trajectory 605a may involve formulating a nonlinear dynamically constrained optimization problem, such as the nonlinear optimization problem 118 described in the previous embodiment, using one or more predetermined weights of the original trajectory and one or more differentiable functions corresponding to one or more new task constraints. Furthermore, generating a new trajectory may involve solving the nonlinear dynamically constrained optimization problem by optimizing one or more predetermined weights for the radial basis functions to generate the new trajectory 605a. The new trajectory 605a satisfies one or more new task constraints for performing the task. Therefore, determining the new trajectory 605a involves determining a correction term that modifies at least some of the weights of the forcing function such that the DMP, which has a forcing function with corrected weights, represents a new feasible trajectory that satisfies the motion constraints. The new feasible trajectory enables the robot 602 to operate smoothly in an environment different from the original environment and to move along the new trajectory 605a while avoiding any obstacles.
[0118] Furthermore, if the environment of the robot 602 changes, the object 604 being manipulated by the robot 602 may extend beyond the robot 602's range, and this range may be defined by physical constraints on the structure of the robot 602.
[0119] Therefore, a set of CDMPs is used to solve a nonlinear optimization problem to satisfy the obstacle avoidance motion constraints. Next, the solution to the nonlinear optimization problem is used to determine a new trajectory 605a for positioning object 604 within the immobile body 606. Then, the end effector of the robot 602 can generate a control input that causes the gripper 603 to follow the new trajectory 605a and successfully position object 604 in a new target posture 607a within the immobile body 606.
[0120] Therefore, without the burden of further learning computations, the robot 602, operated using the controller 602a, can successfully and safely complete tasks. The robot 602 can be adapted to different types of environments by converting from DMP to CDMP without incurring extra learning and demonstration costs, and by naturally resetting the basis function weights and solving nonlinear optimization problems using any known commercial solver. In addition, the robot 602 can achieve higher safety compared to DMP-based task execution by enforcing and satisfying guarantees of constraints as defined in the converted CDMP formulation disclosed herein.
[0121] For clarity, it will be understood that the above description illustrates embodiments of the invention with reference to different functional units and processors. However, it will be apparent that any appropriate distribution of functions between different functional units, processors, or domains can be used without departing from the invention. For example, functions indicated as being performed by separate processors or controllers may be performed by the same processor or controller. Thus, references to specific functional units should be considered not as indicating a strict logical or physical structure or organization, but only as references to appropriate means for providing the described functions.
[0122] Furthermore, one or more computer-readable storage media may be used to implement embodiments consistent with the present disclosure. Computer-readable storage media means any type of physical memory in which information or data readable by a processor can be stored. Thus, computer-readable storage media may store instructions executed by one or more processors, including instructions for causing a processor to perform steps or stages consistent with the embodiments described herein. The term “computer-readable media” should be understood to include tangible items that exclude carrier waves and transient signals, i.e., are non-transient. Examples include random-access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD-ROMs, DVDs, flash drives, disks, and any other known physical storage media.
[0123] This disclosure and examples are intended to be illustrative only, and the true scope and spirit of the disclosed embodiments are intended to be shown by the appended claims.
[0124] As can also be understood, the technologies described above may take the form of processes implemented by a computer or controller and devices for carrying out these processes. The disclosure may also take the form of computer program code, which includes instructions embodied on a tangible medium such as a floppy disk cassette, solid-state drive, CD-ROM, hard drive, or any other computer-readable storage medium, and when the computer program code is loaded into a computer or controller and executed by the computer or controller, the computer becomes a device for carrying out the invention. The disclosure may also take the form of computer program code or signals, whether stored on a storage medium, loaded into a computer or controller and / or executed by the computer or controller, or transmitted on any transmission medium such as over electrical wiring or cables, through optical fibers, or via electromagnetic radiation, and when the computer program code is loaded into a computer and executed by the computer, the computer becomes a device for carrying out the invention. When implemented on a general-purpose microprocessor, the computer program code segment configures the microprocessor to create a specific logic circuit.
[0125] The disclosed methods and systems may be implemented on conventional or general-purpose computer systems such as personal computers (PCs) or server computers. For clarity, it will be understood that the above description illustrates embodiments of the invention with reference to different functional units and processors. However, it will be apparent that any suitable distribution of functions between different functional units, processors, or domains can be used without departing from the invention. For example, functions indicated to be performed by separate processors or controllers may be performed by the same processor or controller. Thus, references to specific functional units should be considered not as indicating a strict logical or physical structure or organization, but only as references to suitable means for providing the described functions.
[0126] When there is no distinction between "physical," "real," or "real-world," the term "robot" can be understood to mean a physical robotic system, or a robot simulator aimed at faithfully simulating the behavior of a physical robotic system. A robot simulator is a program consisting of a set of mathematically based algorithms to simulate the kinematics and dynamics of a real-world robot. In a preferred embodiment, the robot simulator also simulates a robot controller. The robot simulator can generate data for 2D or 3D visualization of the robot, and this data can be output to a display device via a display interface.
[0127] The above description provides only specific embodiments and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the following description of specific embodiments will provide a description that enables the realization of one or more specific embodiments. Various modifications are intended to be made to the function and configuration of the elements without departing from the spirit and scope of the subject matter disclosed in the appended claims.
[0128] Specific details are provided in the following description to ensure a full understanding of the embodiments. However, those skilled in the art will understand that the embodiments can be carried out even without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in the form of block diagrams to avoid obscuring the embodiments with unnecessary details. In other examples, well-known processes, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments. Furthermore, similar reference numbers and names in different drawings refer to similar elements.
[0129] Furthermore, individual embodiments may be described as processes shown as flowcharts, flow diagrams, data flow diagrams, structural diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many operations can be performed in parallel or simultaneously. In addition, the order of operations may be reordered. A process may terminate when its operations are complete, but it may have additional steps that are not discussed or included in the diagrams. Moreover, not all operations in any specifically described process can occur in all embodiments. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. If a process corresponds to a function, the termination of the function may correspond to returning the function to the calling function or the main function.
[0130] Furthermore, embodiments of the disclosed subject matter may be implemented either manually or automatically, at least in part. Manual or automatic implementation may be performed, or at least assisted, through the use of a machine, hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. If implemented with software, firmware, middleware, or microcode, program code or code segments for performing the required tasks may be stored on a machine-readable medium. A processor may perform the required tasks.
[0131] The various methods or processes outlined herein may be encoded as software executable on one or more processors employing any one of a variety of operating systems or platforms. In addition, such software may be written using any of several suitable programming languages and / or programming or scripting tools, and may be compiled as executable machine language code or intermediate code to run on a framework or virtual machine. Typically, the functions of program modules may be combined or distributed as desired in various embodiments.
[0132] Embodiments of this disclosure may be embodied as methods, and an example thereof is provided. The order of operations performed as part of this method may be determined in any suitable manner. Thus, embodiments may be configured such that operations are performed in an order different from the order illustrated, which may include performing some operations simultaneously, although they are shown as a series of operations in the exemplary embodiments. While this disclosure has been described with reference to several preferred embodiments, it should be understood that various other adaptations and modifications can be carried out within the spirit and scope of this disclosure. Therefore, it is the aspect of the appended claims to cover all such variations and modifications that fall within the true spirit and scope of this disclosure.
Claims
1. A method for controlling a robot to perform a task, the method using a processor coupled to memory storing a set of dynamic motion primitives (DMPs) associated with the task, the set of DMPs comprising a set of at least two dynamic systems, the set of at least two dynamic systems comprising at least a function representing point attractor dynamics associated with the task and a coercion function associated with a learned performance of the task, the processor coupled to stored instructions that, when executed by the processor, perform the steps of the method, the method A step of obtaining the set of DMPs associated with the task, The method further comprises the steps of: transforming the acquired set of DMPs into a set of constrained DMPs (CDMPs) by defining perturbation functions associated with the learned forcing functions, wherein the perturbation functions are associated with a set of behavioral constraints that must be satisfied in order to perform the task; Based on the aforementioned operational constraints, the steps include solving a nonlinear optimization problem for the set of CDMPs, The step of generating a control input for controlling the robot to perform the task, based on the solution to the nonlinear optimization problem for the set of CDMPs, The step of converting the set of DMPs to the set of CDMPs by defining the perturbation function associated with the set of operation constraints is: A method comprising the step of defining the perturbation function by adding an additional set of parameters representing the behavioral constraints in the learned forcing function, wherein the new task constraints are represented using the additional set of parameters as barrier functions that are at least once differentiable.
2. The method according to claim 1, wherein the set of at least two dynamic systems includes a combination of ordinary differential equations (ODEs) representing each dynamic system in the set of at least two dynamic systems.
3. The method according to claim 1, wherein the function of the point attractor dynamics includes parameters associated with the starting posture of the robot and the target posture of the robot.
4. The method according to claim 1, wherein the forcing function includes one or more weights corresponding to a set of basis functions associated with the task, and the one or more weights are adjustable parameters associated with the learned performance of the task.
5. The method according to claim 4, wherein the one or more weights are learned by solving a locally weighted regression to learn the one or more weights for the basis function.
6. The step of solving the nonlinear optimization problem for the set of CDMPs based on the aforementioned operating constraints is: The method according to claim 1, further comprising the step of determining a set of feasible values for the additional set of parameters such that the new task constraint represented by the barrier function is satisfied during the execution of the task.
7. The steps include finding the solution to the nonlinear optimization problem by discovering the set of feasible values for the barrier function that represents the new task constraint, The steps include: determining the perturbation function based on the obtained solution to the nonlinear optimization problem; The steps include generating the control input for controlling the robot based on the obtained perturbation function, The method according to claim 6, further comprising the step of controlling the robot to perform the task based on the generated control input.
8. The method according to claim 1, wherein the task includes operating the robot to follow a trajectory.
9. One or more of the above operating constraints are Avoidance of collisions with one or more obstacles present in the environment of the robot, Avoiding self-collisions, and Joint limitations of the end effector of the robot The method according to claim 1, comprising at least one of the following.
10. A controller for controlling the movement of a robot to perform a task, The controller includes memory, and the memory is It is configured to store a set of dynamic motion primitives (DMPs) associated with the task, the set of DMPs comprising at least two sets of dynamic systems, and the set of at least two sets of dynamic systems comprising at least, A function representing the point attractor dynamics associated with the aforementioned task, Includes a forcing function associated with the learned performance of the task, The controller further comprises a processor, and the processor is The processor is configured to transform the set of DMPs into a set of constrained DMPs (CDMPs) by finding perturbation functions associated with the learned forcing functions, the perturbation functions being associated with a set of behavioral constraints that must be satisfied for the task to be performed, and the processor further: Based on the set of operating constraints, solve the nonlinear optimization problem for the set of CDMPs. Based on the solution to the nonlinear optimization problem for the set of CDMPs, it is configured to generate control inputs for controlling the robot to perform the task, The controller further comprises an output interface configured to execute the generated control inputs and instruct the robot to perform the task, To convert the set of DMPs to the set of CDMPs by determining the perturbation function associated with the set of operation constraints, the processor: A controller configured to define the perturbation function by adding an additional set of parameters representing the behavioral constraints in the learned forcing function, wherein the new task constraints are represented using the additional set of parameters as barrier functions that are at least once differentiable.
11. The controller according to claim 10, wherein the set of at least two dynamic systems includes a combination of ordinary differential equations (ODEs) representing each dynamic system in the set of at least two dynamic systems.
12. The controller according to claim 10, wherein the function of the point attractor dynamics includes parameters associated with the starting posture of the robot and the target posture of the robot.
13. The controller according to claim 10, wherein the coercion function includes one or more weights corresponding to a set of basis functions associated with the task, and the one or more weights are adjustable parameters associated with the learned performance of the task.
14. The controller according to claim 13, wherein the one or more weights are learned by solving a locally weighted regression to learn the one or more weights for the basis function.
15. In order to solve the nonlinear optimization problem for the set of CDMPs based on the aforementioned operating constraints, the processor: The controller according to claim 10, configured to determine a set of executable values for the additional set of parameters such that the new task constraint represented by the barrier function is satisfied during the execution of the task.
16. The aforementioned processor further, By discovering the feasible parameter set for the barrier function representing the new task constraint, the solution to the nonlinear optimization problem is obtained. Based on the obtained solution to the nonlinear optimization problem, the perturbation function is determined, Based on the obtained perturbation function, the control input for controlling the robot is generated. The controller according to claim 15, configured to output a control input for controlling the robot to perform the task based on the generated control input.
17. The controller according to claim 10, wherein the task includes operating the robot to follow a trajectory.
18. One or more of the above operating constraints are Avoidance of collisions with one or more obstacles present in the environment of the robot, Avoiding self-collisions, and Joint limitations of the end effector of the robot The controller according to claim 10, comprising at least one of the following.
19. A non-temporary computer-readable medium for storing computer-executable instructions for a robot to perform a task, wherein the computer-executable instructions are configured to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
System and method for quick scripting of tasks for autonomous robotic manipulation
US9486918B1
Method and system for trajectory optimization for nonlinear robotic systems with geometric constraints
WO2021065196A1