Apparatus and method for planning a contact interaction trajectory

By relaxing the contact model and multi-objective optimization, using virtual force and penalty cycle algorithm, the non-smooth dynamics problem of contact interaction trajectory in robot motion planning is solved, and efficient and accurate contact interaction trajectory generation is achieved.

CN115666870BActive Publication Date: 2025-10-10MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180037672.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-29
Filing Date
2021-03-23
Publication Date
2025-10-10
Estimated Expiration
2041-03-23

AI Technical Summary

Technical Problem

Existing technologies have difficulty in automatically determining feasible contact interaction trajectories in complex motion planning, especially when the robot is in contact with the environment, which leads to non-smooth dynamics problems and hinders the effective use of gradient-based optimization methods.

Method used

A relaxation contact model is adopted, and through the cyclic automatic penalty adjustment of virtual forces and relaxation parameters, combined with multi-objective optimization and penalty cycle algorithm, smooth contact interaction trajectories are generated, reducing the need for parameter adjustment.

Benefits of technology

It achieves efficient and accurate generation of contact interaction trajectories in robot motion planning, reduces sensitivity to initialization, and improves computational efficiency and physical accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115666870B_ABST
    Figure CN115666870B_ABST
Patent Text Reader

Abstract

An apparatus and method for planning a contact interaction trajectory are provided. The apparatus is a robot that accepts a contact interaction between the robot and an environment. The robot stores a dynamics model representing geometric properties, dynamic properties, and friction properties of the robot and the environment and a relaxed contact model representing a dynamic interaction between the robot and an object via virtual forces. The robot also iteratively determines, by performing an optimization, an associated control command, a trajectory, and a virtual stiffness value for controlling the robot until a termination condition is satisfied, the optimization reducing a stiffness of the virtual forces and minimizing a difference between a target pose of the object and a final pose of the object moved from an initial pose. Further, an actuator moves a robot arm of the robot in accordance with the trajectory and the associated control command.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to robotics and, more particularly, to an apparatus and method for generalized planning of multi-touch trajectories without a predefined touch schedule. Background Art

[0002] In robotic systems, motion planning is used to determine the trajectory that allows the robot to perform a task that requires reaching a target configuration (the state of the robotic system) or end-effector pose, given its current configuration. To efficiently plan robotic motion, various trajectory optimization techniques are used to find feasible trajectories that obey the robot's dynamics and the task constraints. In other words, trajectory optimization techniques aim to determine an input trajectory that minimizes a cost function that obeys a set of constraints on the robotic system's state and inputs.

[0003] Typically, trajectory optimization methods aim to avoid contact between the robot and the environment, i.e., to avoid collisions. However, contact with the environment must be exploited in various robotic manipulation and motion tasks. For this reason, contact needs to be considered in trajectory optimization. Due to the discrete nature of contact, introducing contact into the robot's motion leads to non-smooth dynamics. For example, making or breaking contact with the environment or using different friction modes for contact (e.g., sticking or sliding) changes the dynamic constraints that determine the system's motion for a given joint force. This phenomenon hinders the use of trajectory optimization to plan contact interaction trajectories.

[0004] To overcome this shortcoming, contact schedules are predefined by the user or generated by advanced heuristic planners. This approach is suitable for motion planning involving a small number of contacts. However, for complex motion planning, predefined contact schedules become computationally impractical.

[0005] In an alternative approach, contact implicit trajectory optimization (CITO) is used, which implements motion planning for contact-rich complex motions without a predefined contact schedule. CITO uses a differential model of contact dynamics to simultaneously optimize state, input, and contact force trajectories given only a high-level goal (e.g., the desired end posture of the system). A contact model in which physical contact is modeled as a smooth function is an element that enables gradient-based optimization to reason about contact. A smooth contact model that allows for a certain distance of penetration and / or contact force facilitates convergence of the optimization. However, one or more parameters of the contact model and cost function must be adjusted to accurately approximate the true contact dynamics while finding the motion to complete a given task. However, this adjustment is difficult. In addition, due to the relaxation that makes numerical optimization efficient (a certain distance of penetration and / or contact force), the contact model causes physical inaccuracies. In addition, when the task or robot changes, one or more parameters may need to be readjusted. Even for minor task changes, not readjusting one or more parameters may also lead to sudden changes in the planned motion.

[0006] Therefore, an adjustment-free contact implicit trajectory optimization technique is needed to automatically determine feasible contact interaction trajectories given a system model and task specifications. Summary of the Invention

[0007] An object of some embodiments is to plan the motion of a robot so that an object moves to a target pose. Another object of some embodiments is to plan the motion of a robot so that an object moves via physical contact between the object and the robot (e.g., between the object and the robot's gripper) without having to grasp the object. One of the challenges of this control is the lack of the ability to use various optimization techniques to determine suitable trajectories for the robot to achieve these contact interactions with the environment. For robotic manipulation, physical contact manifests as impacts, thus introducing non-smoothness in the dynamics, which in turn hinders the use of gradient-based solvers. To this end, a variety of methods have been proposed to determine trajectories by testing numerous trajectories. However, this generation of trajectories is computationally inefficient and may not produce viable results.

[0008] To this end, some embodiments aim to introduce a model that represents the contact dynamics between a robot and its environment, while allowing the use of smooth optimization techniques when using such a model to plan contact interaction trajectories. Additionally or alternatively, another object of some embodiments is to provide such a model that adds a minimal number of additional parameters, has a structure that allows efficient planning computations, yields physically accurate trajectories, and is insensitive to motion planning initialization.

[0009] In the present disclosure, such a model is referred to as a relaxed contact model, which utilizes a cyclic automatic penalty adjustment of relaxation parameters. Specifically, in addition to the physical (real) forces acting on the robot and the object, such as the robot's actuation according to the trajectory, the impact received by the object in response to the robot's touch, friction, and gravity, the relaxed contact model also utilizes virtual forces (not existing in reality) as the vehicle for modeling, and utilizes the contact between the robot's end effector and the environment (e.g., the object) for non-grasping manipulation. Therefore, the virtual forces provide a smooth relationship between the dynamics of the unactuated degrees of freedom (represented as free bodies, such as the object to be manipulated or the torso of a humanoid robot) and the configuration of the system including the robot and the free bodies through contact.

[0010] In this disclosure, the contact geometry between the robot and the environment is defined by the user, taking into account the task. For example, the robot's gripper or linkage, the surface of the object to be manipulated, or, in the case of locomotion, the floor. Furthermore, the geometry is paired by the user so that a given task can be accomplished using one or more contact pairs. For example, the robot's gripper and all surfaces of an object can be paired for non-grasping manipulation applications, or the robot's feet and the floor can be paired for locomotion applications. Furthermore, each contact pair is assigned a free body and a nominal virtual force direction, which will rotate based on the system configuration. For example, a contact pair between the robot's gripper and the surface of an object can generate a virtual force acting perpendicular to the object's contact surface on the object's center of mass. For example, using a pair including the object's front face (i.e., facing the robot) can generate a virtual force on the object that pushes forward, while using the right face can generate a virtual force that pushes leftward at the object's center of mass. In the case of locomotion, using a pair including the humanoid robot's feet and the floor can generate a virtual force on the torso calculated by projecting the virtual force perpendicular to the floor at the contact point onto the torso's center of mass.

[0011] In addition, in some embodiments, the magnitude of the virtual force is expressed as a function of the distance between the contact geometries in the contact pair. The virtual force is penalized and gradually reduced during the optimization process so that at the end of the optimization, the virtual force no longer exists. As the optimization converges, this produces a robot motion that solves the task using only physical contact. Discovering the contact (i.e., virtual force) with this separate relaxation allows minimizing only the virtual force without considering frictional rigid body contact as a decision variable. Therefore, this representation allows the relaxed contact model to annotate the physical forces acting on the free body with only a new independent parameter (the penalty for relaxation).

[0012] Therefore, some embodiments utilize a relaxed contact model in underactuated dynamics of frictional rigid body contact to describe robotic manipulation and locomotion, and replace the general determination of trajectories for robotic control with a multi-objective optimization over at least two objective terms (i.e., the pose of an object moved by the robot and a virtual force, and the magnitude of the virtual force). Specifically, the multi-objective optimization minimizes a cost function to generate a trajectory that penalizes the difference between the target pose of an object and the final pose of the object, the object being placed by the robot moving along the trajectory estimated by the underactuated dynamics and the relaxed contact model for frictional rigid body contact mechanics while penalizing the virtual force.

[0013] Some embodiments perform the optimization by appropriately adjusting the posture deviation from the target posture and the penalty of the virtual force so that both the virtual force and the posture deviation converge to zero at the end of the optimization without changing the penalty value. These embodiments require readjustment for each task. Some embodiments are based on the recognition that the optimization can be performed iteratively when the trajectory obtained from the previous iteration initializes the current iteration while adjusting the penalty. In addition, the magnitude of the virtual force is reduced for each iteration. In this way, previous trajectories with larger virtual forces are processed without optimization to reduce the virtual force and improve the trajectory in each iteration. In some implementations, iterations are performed until a termination condition is met, such as the virtual force reaches zero or the optimization reaches a predetermined number of iterations.

[0014] Thus, one embodiment discloses a robot configured to perform a task involving moving an object from an initial pose of the object to a target pose of the object in an environment, the robot including an input interface configured to receive contact interactions between the robot and the environment. The robot also includes a memory configured to store a dynamic model representing one or more of geometric properties, dynamic properties, and frictional properties of the robot and the environment, and a relaxed contact model representing the dynamic interactions between the robot and the object via virtual forces generated by one or more contact pairs associated with geometric shapes on the robot and geometric shapes on the object, wherein a virtual force acting on the object at a distance in each contact pair is proportional to a stiffness of the virtual force. The robot also includes a processor configured to iteratively determine associated control commands for controlling the robot, a trajectory, and a virtual stiffness value for moving the object according to the trajectory by performing an optimization until a termination condition is satisfied, wherein the optimization reduces the stiffness of the virtual force and reduces the difference between the target pose of the object and a final pose of the object after being moved from the initial pose by the robot controlled according to the control commands via the virtual forces generated according to the relaxed contact model.

[0015] To perform at least one iteration, the processor is configured to: determine a current trajectory, a current control command, and a current virtual stiffness value for a current penalty value of the virtual force stiffness by solving an optimization problem initialized using a previous trajectory and a previous control command determined during a previous iteration with a previous penalty value of the virtual force stiffness; update the current trajectory and the current control command to reduce the distance between each contact pair to generate an updated trajectory and an updated control command to initialize the optimization problem in a next iteration; and update the current value of the virtual force stiffness for optimization in the next iteration. The robot also includes an actuator configured to move a robotic arm of the robot according to the trajectory and the associated control command.

[0016] Another embodiment discloses a method for performing a task involving moving an object from an initial pose of the object to a target pose of the object by a robot, wherein the method uses a processor coupled to instructions for implementing the method, wherein the instructions are stored in a memory. The memory stores a dynamic model representing one or more of geometric, dynamic, and frictional properties of the robot and its environment, and a relaxed contact model representing the dynamic interaction between the robot and the object via virtual forces generated by one or more contact pairs associated with geometric shapes on the robot and geometric shapes on the object, wherein the virtual force acting on the object at a distance in each contact pair is proportional to the stiffness of the virtual force. When executed by the processor, the instructions perform the steps of the method, comprising: obtaining a current state of the interaction between the robot and the object; and iteratively determining associated control commands for controlling the robot, a trajectory, and a virtual stiffness value for moving the object according to the trajectory by performing an optimization until a termination condition is satisfied, wherein the optimization reduces the stiffness of the virtual force and reduces the difference between the target pose of the object and a final pose of the object after being moved from the initial pose by the robot controlled according to the control commands via the virtual forces generated according to the relaxed contact model.

[0017] In addition, in order to perform at least one iteration, the method also includes: determining a current trajectory, a current control command and a current virtual stiffness value for a current penalty value of the virtual force stiffness by solving an optimization problem, wherein the optimization problem is initialized using a previous trajectory and a previous control command determined with a previous penalty value of the virtual force stiffness during a previous iteration; updating the current trajectory and the current control command to reduce the distance between each virtual active contact pair to generate an updated trajectory and an updated control command to initialize the optimization problem in a next iteration; updating the current value of the virtual force stiffness for optimization in the next iteration; and causing the robot arm of the robot to move according to the trajectory and the associated control command.

[0018] As non-limiting examples of exemplary embodiments of the present disclosure, the present disclosure is further described in the following detailed description with reference to the accompanying drawings, in which like reference numerals represent like parts throughout the several views of the drawings. The drawings shown are not necessarily to scale, emphasis generally being placed on illustrating the principles of the presently disclosed embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] [ Figure 1A ] Figure 1A An environment is shown in which a robot according to an embodiment of the present disclosure performs a task involving moving an object from an initial pose of the object to a target pose of the object.

[0020] [ Figure 1B ] Figure 1B A block diagram illustrating a robot according to an embodiment of the present disclosure is shown.

[0021] [ Figure 1C ] Figure 1C Shows the execution result of the traction controller according to some embodiments, the distance changes.

[0022] [ Figure 2A ] Figure 2A 1 and 2. The steps performed by the penalty loop algorithm when the current trajectory does not satisfy the pose constraints according to an embodiment of the present disclosure are shown.

[0023] [ Figure 2B ] Figure 2B 1 and 2. The steps performed by the penalty algorithm when the current trajectory satisfies the pose constraints according to an embodiment of the present disclosure are shown.

[0024] [ Figure 2C ] Figure 2C 1 and 2. The steps performed by the penalty algorithm when the updated trajectory does not satisfy the pose constraint according to an embodiment of the present disclosure are shown.

[0025] [ Figure 2D ] Figure 2D Steps performed during post-processing according to an embodiment of the present disclosure are shown.

[0026] [ Figure 2E ] Figure 2E Steps of a method performed by a robot for performing a task involving moving an object from an initial pose of the object to a target pose of the object are shown according to an embodiment of the present disclosure.

[0027] [ Figure 3A ] Figure 3A Controlling a 1-degree-of-freedom (DOF) pushrod-slider system based on optimized trajectories and associated control commands according to an example embodiment of the present disclosure is shown.

[0028] [ Figure 3B ] Figure 3B Controlling a 7-DOF robot based on optimized trajectories and associated control commands according to an example embodiment of the present disclosure is shown.

[0029] [ Figure 3C ] Figure 3C Controlling a mobile robot with a cylindrical holonomic base based on optimized trajectories and associated control commands according to an example embodiment of the present disclosure is shown.

[0030] [ Figure 3D ] Figure 3D Controlling a 2-DOF humanoid robot with a prismatic torso and cylindrical arms and legs based on optimized trajectories and associated control commands according to an example embodiment of the present disclosure is shown.

[0031] While the above figures set forth embodiments of the presently disclosed implementations, other implementations can also be contemplated as discussed in the Discussion. The present disclosure presents illustrative implementations as representative of the principles of the presently disclosed implementations. Numerous other modifications and implementations can be devised by those skilled in the art that will fall within the scope and spirit of the presently disclosed implementations. DETAILED DESCRIPTION

[0032] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, to one skilled in the art that the present disclosure can be practiced without these specific details. In other instances, devices and methods are shown in block diagram form in order to avoid obscuring the present disclosure.

[0033] As used in this specification and claims, the terms "for example," "e.g.," and "e.g.," and the like shall not be construed as limiting the scope of what is described and claimed, but rather as merely providing examples, and only some of the examples that can be possible. The term "based on" means at least partially based on. Further, it will be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. Any headings used herein are for convenience only and should not be construed as limiting the disclosure.

[0034] Figure 1A An environment 100 is shown in which a robot 101 according to an implementation of the present disclosure performs a task involving moving an object 107 from an initial pose of the object 107 to a target pose 113 of the object 107. Further, Figure 1B A block diagram of a robot 101 according to an implementation of the present disclosure is shown. Reference is made to Figure 1A In connection with Figure 1B The execution of the task performed by the robot 101 is described in detail. As Figure 1AAs shown, the robot 101 includes a robot arm 103, which is used to perform a non-grasping task, such as pushing an object 107 from an initial pose of the object to a target pose 113. The target pose 113 can be an intended pose to which a user wants to move the object 107. In some embodiments, the robot 101 can perform a grasping task, such as grasping the object 107 to perform a task such as moving the object 107. To this end, the robot 101 includes an end effector 105 on the robot mechanism 103, which is actuated to move along a trajectory 111 to contact the surface of the object 107 to exert a virtual force on the object in order to move the object 107 from the initial pose to the target pose 113. The object 107 has a center of mass (CoM) 109. In addition, in this case, there are four contact pairs between the robot end effector 105 and at least one of four contact candidates 107a, 107b, 107c, or 107d on the surface of the object 107 in the environment 100. Each contact pair has a distance and its associated stiffness (k).

[0035] Furthermore, the mobility of the robot 101 or the number of degrees of freedom (DOF) of the robot 101 is defined as the number of independent joint variables required to specify the position of all links of the robot 101 (e.g., robot arm 103, robot end effector 105) in space. It is equal to the minimum number of joints that can be actuated to control the robot 101. Figure 1A As observed in , robot 101 has two degrees of freedom (DOF), and both DOF are actuated. Figure 1B As seen in FIG, the robot 101 includes an input interface 115 configured to receive contact interactions between the robot 101 and the object 107. The input interface 115 may include a proximity sensor, etc. The input interface 115 may be connected to other components of the robot 101 (e.g., a processor 117, a memory 119, etc.) via a bus 121. The processor 117 is configured to execute instructions stored in the memory 119. The processor 117 may be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory 119 may be a random access memory (RAM), a read-only memory (ROM), a flash memory, or any other suitable memory system. The processor 117 may be connected to other components of the robot 101 via a bus 121.

[0036] Some embodiments are based on the recognition that a challenge in controlling a robot 101 to perform grasping or non-grasping tasks is the lack of ability to use various optimization techniques to determine a suitable trajectory (e.g., trajectory 111) for the robot end effector 105 to achieve the desired control. For robotic manipulation, physical contact manifests as a jerky motion, thus introducing non-smooth dynamics that hinder the use of gradient-based solvers. Some embodiments are based on the recognition that a trajectory can be determined by testing numerous trajectories. However, such trajectory generation is computationally inefficient and may not yield optimal results.

[0037] To avoid these consequences, in some embodiments, the storage 119 of the robot 101 is configured to store a dynamic model 143 of the robot 101 and the environment 100. The dynamic model 143 represents the geometric, dynamic, and frictional properties of the robot 101 and the environment 100. The storage is also configured to store a relaxed contact model 123 of the dynamics of the interaction between the robot 101 and the object 107 via virtual forces generated at one or more contact pairs associated with the robot end effector 105 and a surface (107a, 107b, 107c, or 107d) in the environment 100. The virtual forces are generated by the respective contact pairs at a distance The virtual force acting on the object 107 at a certain distance (i.e., without physical contact) is represented by the virtual force at the object 107, where the virtual force is proportional to the stiffness. The relaxed contact model 123 relates the configuration of the robot 101 and the object moved by the robot 101 through the virtual force applied from a distance (i.e., without physical contact), while allowing the use of optimization techniques when controlling the robot 101 using this model. In addition, the relaxed contact model 123 adds a minimal number of additional parameters. The relaxed contact model 123 has a structure that allows for efficient control calculations, resulting in accurate control trajectories, and is insensitive to the initialization of the robot control.

[0038] Furthermore, the processor 117 is configured to determine a trajectory 111 and associated control commands for controlling the robot 101 to move the object 107 to a target pose 113 according to the trajectory 111 by performing an optimization that reduces the stiffness of the virtual force. Furthermore, the optimization reduces the difference between the target pose 113 of the object 107 and a final pose of the object 107 moved from an initial pose by the robot 101 controlled according to the control commands via the virtual forces generated according to the relaxed contact model 123. The final pose of the object 107 may be a pose to which the object 107 was moved earlier (e.g., in a previous iteration), where the final pose was still far from the target pose 113. The robot 101 is configured to use the optimization to reduce this difference between the target pose 113 and the final pose of the object 107.

[0039] To reduce the stiffness of the virtual force and the difference between the target pose 113 of the object 107 and the final pose of the object 107 moved by the robot 101 from the initial pose, the processor 117 can be further configured to perform a multi-objective optimization of the cost function. Multi-objective optimization aims to achieve multiple competing objectives by performing optimization over at least two parameters. According to embodiments, the multi-objective optimization reduces the stiffness and the difference between the target pose 113 and the final pose of the object 107. Further, the cost function is a combination of a first cost that determines the positioning error of the final pose of the object 107 moved by the robot 101 with respect to the target pose 113 of the object 107 and a second cost that determines the cumulative stiffness of the virtual force.

[0040] Further, the processor 117 can be configured to iteratively determine the trajectory 111 until a termination condition 139 is satisfied. To this end, the trajectory in each iteration is analyzed to check whether the trajectory satisfies the pose constraints 131 and the termination condition 139. The trajectory that satisfies the pose constraints 131 and the termination condition 139 can be used as the trajectory 111 by which the robot 101 moves the robot arm 103 to move the object 107. The termination condition 139 can be satisfied in case the number of iterations is greater than a first threshold or the virtual force is reduced to zero. The first threshold can be determined by the robot 101 based on the possible number of iterations, distance, time, etc. that can be required to move the object 107 from the initial pose to the target pose 113. In an example embodiment, the first threshold can be manually defined by a user.

[0041] To perform at least one iteration, the processor 117 is further configured to determine a current trajectory and a current control command for a current value of the virtual force stiffness by solving an optimization problem initialized with a previous trajectory and a previous control command. The previous trajectory and the previous control command are determined during a previous iteration at a previous value of the virtual force stiffness. The optimization problem focuses on optimizing the current trajectory to determine an optimal trajectory (e.g., the trajectory 111) such that the robot 101 moves the object 107 from the initial pose or a pose between the initial pose and the target pose 113 (in case the object 107 has already been moved from the initial pose but has not yet reached the target pose 113) to the target pose 113. The concept of trajectory optimization in the presence of contacts can be formulated as finding contact locations, timings, and forces and robot control inputs given a high-level task.

[0042] The processor 117 is further configured to update the current trajectory and the current control command to reduce the distance in each contact pair (i.e., contact between the robot end effector 105 and at least one of the surfaces 107a, 107b, 107c, or 107d of the object 107) to generate an updated trajectory and an updated control command to initialize the optimization problem in the next iteration and update the current value of the virtual force stiffness for optimization in the next iteration. ​

[0043] The robot 101 includes an actuator 125 configured to move the robotic arm 103 of the robot 101 according to the trajectory 111 and associated control commands. The actuator 125 communicates with the processor 117 and the robotic arm 103 via the bus 121 .

[0044] In an embodiment, the virtual force generated according to the relaxed contact model 123 corresponds to at least one of the four contact pairs between the robot end effector 105 and the at least one surface 107a, 107b, 107c, or 107d. The virtual force may be based on a virtual stiffness, a curvature associated with the virtual force, and a signed distance between the robot end effector 105 and the surfaces 107a, 107b, 107c, and 107d of the object 107 associated with the contact pair. At various moments during the interaction, the virtual force is directed to the projection of the contact surface normal on the object 107 onto the CoM of the object 107 .

[0045] Some embodiments are based on the recognition that introducing contact into trajectory optimization problems leads to non-smooth dynamics, thus hindering the use of gradient-based optimization methods in various robotic manipulation and locomotion tasks. To address this issue, in some embodiments, the robot 101 of the present disclosure can use a relaxed contact model 123 implemented by a processor 117 to generate a contact interaction trajectory 111 using the relaxed contact model 123 and a trajectory optimization module 127.

[0046] Some embodiments are based on the recognition that to solve trajectory optimization problems with reliable convergence properties, a continuous convexification (SCVX) algorithm, a special type of sequential quadratic programming, can be used. In this method, a convex approximation to the original optimization problem is obtained by linearizing the dynamic constraints on the previous trajectory, and a convex subproblem is solved within a trust region. The radius of the trust region is adjusted based on the similarity between the convex approximation and the actual dynamics. In some embodiments, the robot 101 can use the SCVX algorithm in the trajectory optimization module 127 to efficiently calculate the trajectory 111.

[0047] Some embodiments are based on the recognition that a smooth contact model can be used to determine the solution to an optimization problem. In a smooth model, the contact force is a function of distance, thereby allowing the dynamic motion of the robot 101 to be planned. The smooth model facilitates convergence of the iterations required to determine the optimal trajectory 111. However, the smooth model results in physical inaccuracies and is difficult to tune.

[0048] To address this issue, in some embodiments, the robot 101 of the present disclosure may be configured to use a variable smooth contact model (VSCM), in which virtual forces acting at a certain distance are used to discover contacts, while a physics engine is used to simulate the existing contact mechanics. The virtual forces are minimized throughout the optimization. Thus, physically accurate motion is obtained while maintaining fast convergence. In these embodiments, the relaxed contact model 123 corresponds to the VSCM. Using the VSCM in conjunction with SCVX significantly reduces the sensitivity and adjustment burden of the initial guess of the trajectory by reducing the number of adjustment parameters to one (i.e., the penalty of the virtual stiffness). However, the robot 101 using the VSCM and SCVX may still need to readjust the relaxation penalty when the task or robot changes; without additional adjustment, even small task modifications can cause abrupt changes in the planned motion. In addition, due to the structure of the contact model, the resulting contact is often impulsive.

[0049] To address this issue, in some embodiments, the robot 101 may include a penalty loop module 129 that implements a specific penalty loop algorithm for at least one iteration associated with determining the trajectory 111. In the specific penalty loop algorithm, a penalty for a relaxation parameter (e.g., a virtual stiffness associated with a virtual force) included in the relaxed contact model 123 is iteratively changed based on the pose constraint 131.

[0050] To this end, the processor 117 is configured to execute the penalty loop module 129 to assign a first penalty value as an updated penalty value to the virtual stiffness associated with the virtual force, wherein if the pose constraint 131 is satisfied, the assigned penalty value is greater than the penalty value assigned in the previous iteration. On the other hand, if the pose constraint 131 is not satisfied, a second penalty value is assigned as an updated penalty value to the virtual stiffness associated with the virtual force, wherein the assigned penalty value is less than the penalty value assigned in the previous iteration. The pose constraint 131 includes information about the position error and the orientation error associated with the trajectory 111. More specifically, if the value of the position error is below a threshold and the value of the orientation error is below a threshold (e.g., the normalized position error is below 30% and the orientation error is below 1 rad), the pose constraint 131 is satisfied.

[0051] Additionally, the processor 117 determines a current trajectory that satisfies the pose constraints 131 , associated control commands and virtual stiffness values, and a residual virtual stiffness that indicates the location, timing, and magnitude of the physical forces used to perform the task.

[0052] Some embodiments are based on the recognition that the average stiffness associated with the current trajectory, calculated by the penalty loop module 129, can generate impact contact forces on the object 107 by the robot arm 103 contacting the object 107. Such impact contact forces on the object 107 can undesirably cause the object 107 to displace. To address this issue, in some embodiments, the robot 101 utilizes a post-processing module 133. Post-processing is performed on the current trajectory to attract geometric shapes on the robot to corresponding geometric shapes in the environment using a traction controller 135 to facilitate physical contact. To this end, the processor 117 is configured to utilize information associated with a residual virtual stiffness variable that indicates the location, timing, and magnitude of the force required to complete the task. The traction controller 135 is executed when the average value of the virtual stiffness for the current trajectory is greater than a virtual stiffness threshold.

[0053] Figure 1C As a result of the execution of the traction controller 135 according to some embodiments, the distance The processor 117 uses the information associated with the residual virtual stiffness variable to execute the traction controller 135 on the current trajectory to determine the traction force The traction force attracts the virtual active contact geometry on the robot 101 associated with a non-zero virtual stiffness value toward the corresponding contact geometry on the object 107. For example, the distance between the end effector 105 and the contact candidate 107a is Change to distance Less than distance For this purpose, the traction force is given by

[0054]

[0055] in is the distance vector from the geometric center of the contact candidate on the robot (eg, end effector 105) to the geometric center of the contact candidate in the environment (eg, contact candidate 107a), and k is the virtual stiffness value in the current trajectory.

[0056] To this end, when applying traction Afterwards, the distance However, the virtual stiffness value k may lead to an excessively large virtual force, because the magnitude of the normal virtual force is inversely proportional to the distance, that is, To avoid these excessive virtual forces, the stiffness values are reduced by a hill climbing search such that the task constraints are satisfied. To this end, the processor 117 is configured to use a hill climbing search implemented by a hill climbing search (HCS) module 137. The hill climbing search reduces the non-zero stiffness values one by one as long as the positioning error is reduced. In an embodiment, the hill climbing search uses a fixed step size. In an alternative embodiment, the hill climbing search uses an adaptive step size. The traction controller 135 and the HCS module 137 include a post-processing module 133. Thus, the post-processing includes executing the traction controller and the hill climbing search. Depending on the embodiment, the post-processing outputs the associated control commands and the trajectory that satisfy the pose constraints for performing the task. Thus, the post-processing reduces the number of iterations to determine the trajectory. In other words, the post-processing improves the convergence. Moreover, the trajectory resulting from the post-processing is better than the current trajectory as it facilitates the physical contact and explicitly reduces the virtual stiffness values without degrading the task performance.

[0057] In some embodiments, the robot 101 further includes a damping controller 141. The processor 117 is configured to execute the damping controller 141 to prevent the traction controller 135 from generating sudden motions for large stiffness values.

[0058] Some embodiments are based on the recognition that by adjusting the weight (or penalty value), the virtual forces disappear as the optimization converges, thereby only using the physical contact to produce the motions that solve the task. In these embodiments, the penalty of the virtual stiffness can be applied in the process of reducing the virtual forces by adjusting the penalty value, where a small penalty value can lead to physically inconsistent motions due to the remaining virtual forces. Moreover, if the penalty value is too large, then it can not be possible to find the motions that complete the task. Although this adjustment of the penalty is quite straightforward, it hinders the generalization of the method for a wide range of tasks and robots. To address this issue, a penalty loop algorithm that automatically adjusts the penalty is utilized in some embodiments of the present disclosure. Moreover, the determined trajectory, the associated control commands, and the virtual stiffness values are improved by the post-processing stage described above after each iteration.

[0059] Figure 2A Steps performed by the penalty loop algorithm when the current trajectory does not satisfy the pose constraints 131 are shown in accordance with embodiments of the present disclosure. In some embodiments, the processor 117 can be configured to perform the steps by the penalty loop algorithm.

[0060] At step 201, an initial state vector representing the state of the robot 101, an initial penalty value adjusting the relaxation parameters (e.g., virtual stiffness) included by the relaxation contact model 123, and an initial control trajectory to be optimized to obtain an optimal trajectory 111 that moves the object 107 with zero virtual forces can be obtained.

[0061] At step 203, the trajectory optimization module 127 can be executed based on the initial state vector, the penalty and the initialized values of the control trajectory to determine a current trajectory, a current control command and a current virtual stiffness value. The continuously convexing algorithm can significantly reduce the sensitivity to the initial guess and the variable-smoothing contact model can reduce the number of tuning parameters to one, i.e., the virtual stiffness penalty.

[0062] At step 205, performance parameters can be evaluated, where the performance parameters can include a position error, an orientation error, a maximum stiffness value k max and an average stiffness value k avg The pose constraints 131 include the position error and the orientation error.

[0063] At step 207, it can be checked whether the current trajectory satisfies the pose constraints 131. When the determined trajectory satisfies the pose constraints 131, control is passed to step 211. On the other hand, when the current trajectory does not satisfy the pose constraints 131, control is passed to step 209.

[0064] At step 209, the penalty values assigned to the relaxation parameters can be reduced by half of the previous change and these values can be fed back to step 203, where the trajectory optimization module 127 is executed with these values to determine a new optimized trajectory that satisfies the pose constraints 131 in the next iteration at step 207.

[0065] Figure 2B Steps performed by the penalty algorithm when the current trajectory satisfies the pose constraints 131 are shown according to an embodiment of the present disclosure.

[0066] At step 211, when the determined trajectory satisfies the pose constraints 131, it can be checked whether the average stiffness value k avg is less than a threshold stiffness value k threshold . When the average stiffness value k avg is determined to be less than k threshold , control is passed to step 213. Otherwise, control is passed to step 223.

[0067] At step 213, based on the average stiffness value k avg being less than the threshold stiffness value k threshold , post-processing can be performed on the determined trajectory. The post-processing step improves the current trajectory and the associated current control command by adjusting the relaxation parameters (e.g., virtual stiffness associated with virtual forces) in the previous iteration with contact information implied based on the penalty values. In addition, the post-processing step uses the traction controller 135 to attract the robot link (or the robot end effector 105 of the robot arm 103) associated with the non-zero stiffness value towards the corresponding contact candidate in the environment 100. Thus, performing the post-processing 213 results in an updated trajectory, an updated control command and an updated virtual stiffness value.

[0068] At step 215 , it may be determined whether the trajectory, associated control commands, and virtual stiffness values ​​updated during the post-processing step 213 satisfy the pose constraints 131 . If the pose constraints 131 are satisfied, control passes to step 217 . Otherwise, control passes to step 223 .

[0069] In step 217, it may be determined whether the termination condition 139 is satisfied. If the termination condition 139 is not satisfied, control passes to step 219. Otherwise, control passes to step 221. The termination condition 139 may be satisfied if the number of iterations is greater than a first threshold or the virtual stiffness value decreases to zero.

[0070] In step 219 , when the termination condition 139 is not satisfied, the penalty value may be increased by a fixed step size.

[0071] At step 221, the robot 101 may be controlled according to the current trajectory and the associated current control commands and the current virtual stiffness value. In some embodiments, the actuator 125 is controlled according to the trajectory and the associated control commands to move the robot arm 103. After executing step 219, execution of the steps of the penalty algorithm ends.

[0072] Figure 2C 1 and 2. The steps performed by the penalty algorithm when the updated trajectory does not satisfy the pose constraint 131 according to an embodiment of the present disclosure are shown.

[0073] In the event that the updated trajectory obtained after post-processing (at step 215 ) does not satisfy the pose constraint 131 , control passes to step 223 .

[0074] At step 223 , the previous iteration trajectory and associated control commands may be used as the optimal solution.

[0075] In step 217, it may be determined whether the termination condition 139 is satisfied. If the termination condition 139 is not satisfied, control passes to step 219, otherwise, control passes to step 221.

[0076] In step 219 , when the termination condition 139 is not satisfied, the penalty value may be increased by a fixed step size.

[0077] At step 221, the robot 101 may be controlled according to the current trajectory and associated current control commands. In some embodiments, the actuator 125 is controlled according to the trajectory and associated control commands to move the robot arm 103. After executing step 219, execution of the steps of the penalty algorithm ends.

[0078] Figure 2DThe steps performed during post-processing according to an embodiment of the present disclosure are shown. Post-processing is performed when the average stiffness value for the current trajectory determined by the penalty algorithm is less than the threshold stiffness value required to move the robotic arm 103 according to the current trajectory. After obtaining the current trajectory, the associated current control command, and the current virtual stiffness value, the method begins at step 225.

[0079] At step 225, the traction controller 135 may be executed to attract the virtual active robot end effector 105 toward the corresponding contact candidate among the candidates 107a, 107b, 107c, and 107d on the object 107 in the environment 100 to facilitate physical contact. The traction controller 135 is executed based on the current trajectory, the current control command, and the current virtual stiffness value. Furthermore, as the distance between the robot end effector 105 and the contact candidates 107a, 107b, 107c, and 107d decreases, the stiffness value considered for executing the current trajectory may result in excessive virtual force. To overcome this situation, control passes to step 227.

[0080] At step 227, a hill climbing search (HCS) operation may be performed. As long as the nonlinear pose error decreases, the HCS operation results in a decrease in the non-zero stiffness value by normalizing the change in the final cost by dividing the previous change. Reducing the non-zero stiffness value results in a clear suppression of the virtual force. Thus, the current trajectory, associated control commands, and virtual stiffness values ​​are refined through post-processing to generate an updated trajectory, updated control commands, and updated virtual stiffness values. The updated trajectory is further analyzed in a penalty loop algorithm to check whether the updated trajectory satisfies the pose constraints 131.

[0081] In some embodiments, post-processing is used to improve the output trajectory and associated control commands determined by the penalty loop algorithm explicitly by exploiting the contact information implied by the relaxation. For example, for a contact pair p and control cycle i, from the distance vector d(p,i)∈R 3 and the associated virtual stiffness value k(p,i) to calculate the traction force f(p,i)∈R 3 , where the traction force is given by:

[0082] f(p,i)=k(p,i)d(p,i), (1)

[0083] Where d is the vector from the center of mass of the contact candidate on the robot 101 (i.e., the end effector 105) to a point in the environment 100 that is offset from the center of the contact candidate. In an example embodiment, this offset can be arbitrarily initialized to 5 cm for the first penalty iteration and obtained for subsequent iterations by dividing the initial offset by the number of successful penalty iterations. This offset helps reach occluding surfaces in the environment 100. In another embodiment, a potential field method can be used with repulsive forces on surfaces 107a, 107b, 107c, and 107d with a zero stiffness value.

[0084] Corresponding generalized joint force vector It can be calculated by the following formula:

[0085]

[0086] in is the translation Jacobian matrix of the center of mass of the contact candidate on the robot 101.

[0087] To prevent the traction forces from generating sudden motions for large stiffness values, a damping controller 141 may be applied to keep the joint velocity close to the planned motion, where the damping force may be given as:

[0088]

[0089] in is a positive finite gain matrix, is the deviation of the joint velocity from the planned velocity, is the damped generalized joint force. In some embodiments, this calculation can be performed efficiently using a sparse form of the inertia matrix.

[0090] Some embodiments are based on the recognition that the relaxed contact model 123 can be used to implement a contact implicit trajectory optimization framework that can plan contact interaction trajectories (e.g., trajectory 111) for different robot architectures and tasks using trivial initial guesses without any parameter tuning. a Actuation DOF and n u The mathematical representation of the dynamics of an underactuated system with unactuated DOF can be given as follows:

[0091]

[0092] in is the configuration vector; is the mass matrix; represents the Coriolis, centrifugal, and gravitational terms; is the selection matrix for the actuation DOF, is the selection matrix for the unactuated DOF; is the vector of generalized joint forces; is n c The vector of the generalized contact force at the contact point, is the Jacobian matrix that maps joint velocities to Cartesian velocities at the contact points, is the vector of generalized contact forces for the unactuated DOFs generated by the contact model.

[0093] In another embodiment, for n in a special Euclidean (SE) group (e.g., SE(3)), f free bodies (e.g., free objects or humanoid torsos), n u =6n f The state of the system is determined by Indicates that n=2(n a +n u There are two types of contact mechanics in this system: (i) contact forces due to actual contact in the simulation world (i.e., contact detected by the physics engine), which are effective on all DOFs; and (ii) virtual forces due to the contact model and applied only to the non-actuated DOFs.

[0094] In another embodiment, the generalized joint force is decomposed into in and yes and λ c estimates; and is a vector of control variables associated with the joint forces. This helps focus the optimization problem in terms of the joint forces, which means that even in the presence of external contact, the control terms are linearly related to the acceleration.

[0095] Some embodiments are based on the recognition that contact-implicit manipulation is used to define a manipulation task (grasping or non-grasping) as an optimization problem, where contact schedules and corresponding forces are found as a result of trajectory optimization.The choice of contact model is crucial.

[0096] In the present disclosure, a relaxed contact model 123 is used to facilitate the convergence of the gradient-based solver. The relaxed contact model 123 takes into account the contact relationships between the robot 101 (e.g., the end effector link 105) and the environment 100 (e.g., the surfaces 107a, 107b, 107c, and 107d of the object 107). p For each contact pair, the magnitude of the virtual force γ∈R normal to the surface is given by γ(q)=ke using the signed distance φ∈R between the contact candidates, the virtual stiffness k and the curvature α. -αφ(q)Calculate the corresponding generalized virtual force λ acting on the free body associated with the contact pair v ∈R 6 By λ v (q) = γ(q)[I3-I] T n(q) is calculated, where I3 is the 3×3 identity matrix, l is the vector from the center of mass of the free body to the closest point on the contact candidate on the robot 101, is the skew-symmetric matrix form of I that performs the cross product, n∈R 3 is the contact surface normal. The net virtual force acting on a free body is the sum of the virtual forces corresponding to the contact candidates associated with that body. As a result, the virtual force provides a smooth relationship between the dynamics of the free body and the configuration of the system.

[0097] In addition, in VSCM, the virtual stiffness value is the decision variable to be optimized Therefore, the vector of control variables is Where m = n a +n p .

[0098] Some embodiments are based on the recognition that trajectory optimization can be used to determine the optimal trajectory 111 for a given high-level task. To this end, the robot 101 determines a trajectory 111 based on a trajectory optimization method that minimizes the virtual force while satisfying the posture constraints 131 associated with the determined trajectory 111. The finite-dimensional trajectory optimization problem for N time steps can be based on the state trajectory and control trajectory Final cost item C F and comprehensive cost item C I ; and the lower control boundary u L , upper control boundary u U , lower state boundary x L and the upper state boundary x U To write:

[0099]

[0100] obey:

[0101] x i+1 =f(x i ,u i ), for i=1,…,N (5b)

[0102] u L ≤u 1,…,N ≤u U ,x L ≤x 1,…,N+1 ≤x U (5c)

[0103] where xi+1 =f(x i ,u i ) describes the evolution of the nonlinear dynamics over the control period i.

[0104] Some embodiments define actions and non-grasp manipulation tasks based on a desired torso / object configuration. To this end, a weighted quadratic final cost is used, which is based on the deviations of the free body's position and orientation from the desired pose, pe and θe:

[0105]

[0106] Where w1 and w2 are weights. In order to suppress all virtual forces, the L of the virtual stiffness variable in the comprehensive cost is 1 Norm is penalized:

[0107] C I =ω‖k i ‖1 (7)

[0108] Furthermore, the penalty ω is adjusted by a penalty loop module 129 comprising instructions corresponding to the steps of the penalty loop algorithm.

[0109] Some embodiments are based on the recognition that non-convexity (or nonlinearity in the dynamics associated with the trajectory) can originate from the objective function, the state or control constraints, or the nonlinear dynamics. The first case is generally easy to manage, as the non-convexity can be transferred from the objective to the constraints by changing the variables. For the second case, it is necessary to convert the non-convex constraints (state or control constraints, nonlinear dynamics) into convex constraints while ensuring an optimal solution. Continuous Convexification (SCVX) is an algorithm for solving optimal control problems with non-convex constraints or dynamics by iteratively creating and solving a sequence of convex problems. This algorithm is described below.

[0110] The SCVX algorithm is based on sequentially repeating three main steps: (i) linearizing non-convex constraints (e.g., nonlinear dynamics) on the trajectory from the previous sequence; (ii) solving the resulting convex subproblem subject to trust region constraints that avoid artificial unboundedness due to the linearization; and (iii) adjusting the trust region radius based on the fidelity of the linear approximation.

[0111] The convex subproblem is given by:

[0112]

[0113] obey:

[0114] For i=1,…,N, (8b)

[0115] For i=1,…,N+1 (8c)

[0116] For i=1,…,N, (8d)

[0117] ‖δX‖1+‖δU‖1≤r s (8e) Where (X s ,U s ) is the trajectory from sequence s; r is the trust region radius. Additionally, dummy controls can be added to the problem to prevent artificial infeasibility due to linearization.

[0118] The convex subproblem is a simultaneous problem and therefore has a larger size but a sparse structure that can be exploited by a suitable solver. After solving the convex subproblem, only the change in control is applied, rather than applying changes in both state and control. The state trajectory is then recalculated by unfolding the dynamics. This modification prevents the drawbacks that can occur when using a penalty method in the original method (i.e., f(x i ,u i )-x i+1 ) accumulation, and improves the convergence rate by allowing a larger trust region in our experiments. As a result, the modified method combines the numerical efficiency of the direct method with the accuracy of the shooting method.

[0119] Figure 2E The steps of a method performed by the robot 101 for performing a task involving moving the object 107 from an initial pose to a target pose 113 according to an embodiment of the present disclosure are shown. The method is performed by the processor 117 of the robot 101. The method starts at step 229.

[0120] At step 229, the initial state vector, penalty value, and control trajectory may be initialized to determine the optimal state and control trajectory. The initial value may correspond to the starting value of the control trajectory. In some embodiments, the initial value may be predefined for the robot 101. In other embodiments, the initial value may be manually provided by a user. Furthermore, the current state of the interaction between the robot 101 and the object 107 is obtained via the input interface 115 to determine the complete trajectory.

[0121] At step 231 , an SCVX algorithm may be executed to solve the trajectory optimization problem, which is non-convex due to nonlinear dynamics, in a numerically efficient manner.

[0122] At step 233 , performance measurement parameters associated with the trajectory (e.g., position error, orientation error, average stiffness value, and maximum stiffness value) may be evaluated. Pose constraints 131 include position error and orientation error. The performance measurement parameters may be used to optimize the trajectory to obtain an optimized trajectory 111 . To this end, control passes to step 235 .

[0123] At step 235, a penalty loop algorithm may be executed. The penalty loop is executed to iteratively determine the trajectory 111 of the object 107 and the associated control commands for controlling the robot 101 to move the object 107 according to the trajectory 111 by performing an optimization that reduces (e.g., minimizes) the stiffness of the virtual force and reduces (e.g., minimizes) the difference between the target pose 113 of the object 107 and the final pose of the object 107 moved by the robot 101 from the initial pose until a termination condition 139 is satisfied. The termination condition 139 may be satisfied when the number of iterations is greater than a first threshold or the virtual force is reduced to zero. The robot 101 is controlled based on the control commands via the virtual force generated according to the relaxed contact model 123.

[0124] To this end, different penalties are assigned to relaxation parameters, such as virtual stiffness, based on a determination of whether the calculated trajectory satisfies the pose constraint 131. By dynamically varying the penalties on the relaxation parameters based on the pose constraint 131, the penalty algorithm gradually reduces the virtual force to zero, thereby performing the task using only physical forces. Furthermore, a determination is made as to whether the average stiffness value is less than a threshold stiffness value required to move the robot arm 103 according to the trajectory determined by the penalty loop algorithm. Control then passes to step 237.

[0125] At step 237, post-processing may be performed on the current trajectory to attract the robot link (or the robot end effector 105 of the robot arm 103) associated with the non-zero stiffness value toward the corresponding contact candidate in the environment 100 using the traction controller 135. To this end, the processor 117 is configured to utilize information associated with the residual virtual stiffness variable, which indicates the position, timing, and magnitude of the force required to accomplish the task of moving the object from the object's initial pose to the object's target pose 113. Furthermore, control passes to step 239.

[0126] At step 239, an optimal trajectory and associated control commands may be determined based on a determination of whether the trajectory updated using post-processing in step 237 satisfies the termination condition 139, wherein the termination condition 139 may be satisfied if the number of iterations is greater than a first threshold or the virtual force decreases to zero. In some embodiments, the optimal trajectory satisfies both the pose constraint 131 and the termination condition 139.

[0127] In some embodiments, an optimal trajectory can be determined based on a determination that the trajectory determined in the penalty loop algorithm includes an average stiffness value less than a threshold stiffness value and that the trajectory satisfies the termination condition 139. Furthermore, the robot 101 can use control commands to move the robot end effector 105 along the optimal trajectory to move the object 107 to the target pose 113.

[0128] Thus, the object 107 moves from the initial pose to the target pose 113 according to the optimized trajectory.

[0129] In an example embodiment, Figure 3A 、 Figure 3B 、 Figure 3C and Figure 3D Trajectory optimization using the penalty loop method and post-processing is implemented in four different robotics applications shown.

[0130] Figure 3A FIG3 illustrates controlling a 1-degree-of-freedom (DOF) push-slider system 301 based on an optimized trajectory and associated control commands according to an example embodiment of the present disclosure. The system 301 is configured to perform a push task with a single control time step of 1 second (sec). The system 301 performing the push task includes a contact pair comprising a tip 309 of a push rod 311 and a front face 313 of a slider 303. The system 301 may include a relaxed contact model 123 to determine the optimized trajectory and associated control commands.

[0131] Furthermore, the system 301 pushes the slider 303 (20 cm) in a single direction (eg, forward direction 305 ) to reach a target pose 307 of the slider (eg, box 303 ) based on the optimized trajectory and associated control commands determined by the relaxed contact model 123 .

[0132] Figure 3B Controlling a 7-DOF robot 315 based on optimized trajectories and associated control commands according to an example embodiment of the present disclosure is shown. In an example embodiment, the 7-DOF robot 315 may be a Sawyer robot. In addition to pushing the box 319 forward, the 7-DOF robot 315 may perform side and diagonal pushes. The 7-DOF robot 315 has four contact pairs between the sides 321a and 321b of the box 319 and the cylindrical end effector flange 311. In an example embodiment, the 7-DOF robot 315 performs three forward push tasks to move the box 319. An optimized trajectory and associated control commands are determined using the relaxed contact model 123 to perform a slight or jerk motion to move the box 319 out of the workspace of the 7-DOF robot 315.

[0133] Figure 3C Controlling a mobile robot having a cylindrical holonomic base 323 based on optimized trajectories and associated control commands according to an example embodiment of the present disclosure is shown. In an example embodiment, the mobile robot 323 may be a human support robot (HSR) having a cylindrical holonomic base for contacting the environment.

[0134] In order to use the velocity control of the HSR 323 to control the holo base 327 to perform the task of pushing the box 325, the optimized trajectory and control commands are determined for the HSR 323 by relaxing the contact model 123. Figure 3C As shown, there are four contact pairs between the sides of the box 325 and the cylindrical base 327 of the HSR 323. Since the translational and rotational speed limits are ±2 m / s and ±2 rad / s, a longer simulation time of 5 seconds and a larger control sampling period of 0.5 seconds are used to perform the different tasks. The forward push task of moving the box 325 50 cm and two diagonal push tasks are performed by the HSR 323. It is observed that when the default friction coefficient of the physics engine (μ = 1) is used, the HSR 323 relies heavily on friction to perform the diagonal push, which seems unrealistic. To avoid this problem, the task is repeated using μ = 0.1.

[0135] Figure 3D A 2-DOF humanoid robot 329 having a prismatic torso and cylindrical arms 331a, 331b and legs 331c, 331d is shown being controlled based on optimized trajectories and associated control commands according to an example embodiment of the present disclosure.

[0136] A planar humanoid robot 329 is controlled according to an optimal trajectory and associated control commands for a motion application, wherein the humanoid robot 329 can make and break multiple contacts simultaneously. The environment containing the humanoid robot 329 has zero gravity, which avoids stability constraints. The task is specified based on the desired torso pose that the humanoid robot 329 can achieve using four stationary bricks in the environment, such as Figure 3D However, since the motion is frictionless, the humanoid robot 329 can use contact to slow down or stop. Figure 3D As shown, there are 8 contact candidates for the front and back faces of the brick and 4 contact candidates for the end links of the arms and legs on the humanoid robot 329. These candidates are paired based on the side faces, resulting in a total of 16 contact candidates. Leg contacts are not required to complete the task, however, they are included as contact candidates to demonstrate that additional or unnecessary contact pairs do not hinder the performance of the proposed method.

[0137] The various methods or processes outlined herein may be encoded as software that can be executed on one or more processors using any of a variety of operating systems or platforms. Additionally, such software may be written using any of a variety of suitable programming languages ​​and / or programming or scripting tools, and may also be compiled into executable machine language code or intermediate code that is executed on a framework or virtual machine. Generally, in various embodiments, the functionality of the program modules may be combined or distributed as desired.

[0138] In addition, the embodiments of the present disclosure can be embodied as a method, an example of which has been provided. The actions performed as part of the method can be sorted in any suitable manner. Therefore, an embodiment in which the actions are performed in an order different from that shown can be constructed, which may include performing some actions simultaneously, although shown as sequential actions in the illustrative embodiments. In addition, the use of ordinal numbers such as "first" and "second" in the claims to modify claim elements does not itself imply any priority or order of one claim element over another claim element or the time order in which the method actions are performed, but is merely used as a label to distinguish a claim element with a specific name from another element with the same name (but using ordinal numbers) to distinguish the claim elements.

[0139] Although the present disclosure has been described with reference to certain preferred embodiments, it will be understood that various other adaptations and modifications may be made within the spirit and scope of the present disclosure. Therefore, it is intended that the appended claims encompass all such changes and modifications as fall within the true spirit and scope of the present disclosure.

Claims

1. A robot configured to perform a task in an environment involving moving an object from an initial pose of the object to a target pose of the object, the robot comprising: an input interface configured to receive contact interactions between the robot and the environment; a memory configured to store a dynamic model and a relaxed contact model, the dynamic model representing one or more of geometric properties, dynamic properties, and friction properties of the robot and the environment, the relaxed contact model representing a dynamic interaction between the robot and the object via a virtual force generated by one or more contact pairs associated with a geometric shape on the robot and a geometric shape on the object, wherein the virtual force acting on the object at a distance in each contact pair is proportional to a stiffness of the virtual force; a processor configured to iteratively determine associated control commands for controlling the robot, a trajectory, and a virtual stiffness value for causing the object to move according to the trajectory by performing an optimization until a termination condition is satisfied, the optimization reducing the stiffness of the virtual force and reducing a difference between the target pose of the object and a final pose of the object moved from the initial pose by the robot controlled according to the control commands via the virtual force generated according to the relaxed contact model, wherein, to perform at least one iteration, the processor is configured to: determining a current trajectory, a current control command, and a current virtual stiffness value for the current penalty value of the stiffness of the virtual force by solving an optimization problem initialized with a previous trajectory and a previous control command determined during a previous iteration with a previous penalty value of the stiffness of the virtual force; updating the current trajectory and the current control command to reduce the distance between each contact pair to generate an updated trajectory and an updated control command to initialize the optimization problem in a next iteration; and updating the current value of the stiffness of the virtual force for use in the optimization in the next iteration; and an actuator configured to move a robotic arm of the robot according to the trajectory and the associated control commands, The memory is further configured to store a pulling controller that uses a virtual force remaining after calculating the current trajectory to attract a geometric shape on the robot toward a corresponding geometric shape on the object in the environment to facilitate physical contact.

2. The robot according to claim 1, in, The virtual force corresponding to the contact pair is based on one or more of a stiffness of the virtual force, a curvature associated with the virtual force, and a signed distance between a geometry on the robot associated with the contact pair and a geometry on the object in the environment.

3. The robot according to claim 1, in, At each moment during the dynamic interaction, the virtual force is directed to the projection of the contact surface normal onto the center of mass of the object.

4. The robot according to claim 1, in, The optimization corresponds to a multi-objective optimization of a cost function, wherein the processor is further configured to perform the multi-objective optimization of the cost function, and The cost function is a combination of the following: determining a first cost of a positioning error of the final pose of the object moved by the robot relative to the target pose of the object, and A second cost of the cumulative stiffness of the virtual force is determined.

5. The robot according to claim 1, in, In order to perform the at least one iteration, the processor is further configured to: Use continuous convexification to perform trajectory optimization problems; assigning a first penalty value as an updated penalty value to a stiffness associated with the virtual force, wherein the assigned penalty value is greater than a penalty value assigned in a previous iteration if a pose constraint is satisfied, and wherein the pose constraint includes information about a position error and an orientation error associated with the trajectory; determining the current trajectory, the current control command, and the current virtual stiffness value that satisfy the pose constraints and a residual stiffness indicating the location, timing, and magnitude of physical forces used to perform the task; and The traction controller is executed on the current trajectory to determine a traction force for pulling contact pairs on the robot associated with non-zero stiffness values ​​toward corresponding contact pairs in the environment.

6. The robot according to claim 5, in, The processor is further configured to: assigning a second penalty value as an updated penalty value to the stiffness associated with the virtual force, wherein the assigned penalty value is less than the penalty value assigned in a previous iteration if the pose constraint is not satisfied; and The trajectory optimization problem is performed using the continuous convexification.

7. The robot according to claim 5, in, The processor is further configured to execute the traction controller based on the average value of the stiffness when the average value of the stiffness is greater than a stiffness threshold.

8. The robot according to claim 5, in, To determine the traction force, the processor is further configured to execute the traction controller based on a previous stiffness, and Wherein the previous stiffness indicates the location, timing and magnitude of the physical force associated with the previous iteration.

9. The robot according to claim 5, in, The memory also stores a hill climbing search, and The processor is further configured to perform the hill climbing search to reduce the non-zero stiffness value to eliminate excessive virtual force.

10. The robot according to claim 1, in, The task includes at least one of a non-gripping operation or a gripping operation.

11. The robot according to claim 1, in, When the number of iterations is greater than a first threshold or the virtual stiffness value decreases to zero, the termination condition is met.

12. A method of performing, by a robot, a task involving moving an object from an initial pose of the object to a target pose of the object, wherein: The method uses a processor coupled with instructions for implementing the method, wherein the instructions are stored in a memory, wherein the memory stores a dynamic model and a relaxed contact model, the dynamic model representing one or more of geometric properties, dynamic properties, and friction properties of the robot and the environment, and the relaxed contact model representing a dynamic interaction between the robot and the object via a virtual force generated by one or more contact pairs associated with a geometric shape on the robot and a geometric shape on the object, wherein the virtual force acting on the object at a certain distance in each contact pair is proportional to a stiffness of the virtual force, and The instructions, when executed by the processor, perform the steps of the method, which includes the following steps: obtaining a current state of interaction between the robot and the object; and Iteratively determining associated control commands for controlling the robot, a trajectory, and a virtual stiffness value for moving the object according to the trajectory by performing an optimization until a termination condition is satisfied, the optimization minimizing the stiffness of the virtual force and minimizing the difference between the target pose of the object and a final pose of the object moved from the initial pose by the robot controlled according to the control commands via the virtual force generated according to the relaxed contact model, wherein, to perform at least one iteration, the method further comprises the following steps: determining a current trajectory, a current control command, and a current virtual stiffness value for the current penalty value of the stiffness of the virtual force by solving an optimization problem initialized with a previous trajectory and a previous control command determined during a previous iteration with a previous penalty value of the stiffness of the virtual force; updating the current trajectory and the current control command to reduce the distance between each contact pair to generate an updated trajectory and an updated control command to initialize the optimization problem in a next iteration; and updating the current value of the stiffness of the virtual force for use in the optimization in the next iteration; and moving the robot arm of the robot according to the trajectory and the associated control command, The memory also stores a pulling controller that uses a virtual force remaining after calculating the current trajectory to attract a geometric shape on the robot toward a corresponding geometric shape on the object in the environment to facilitate physical contact.

13. The method according to claim 12, further comprising the following steps in order to perform the at least one iteration: Use continuous convexification to perform trajectory optimization problems; assigning a first penalty value as an updated penalty value to a stiffness associated with the virtual force, wherein the assigned penalty value is greater than a penalty value assigned in a previous iteration if a pose constraint is satisfied, and wherein the pose constraint includes information about a position error and an orientation error associated with the trajectory; determining the current trajectory and the control commands that satisfy the pose constraints and residual stiffness indicating the location, timing, and magnitude of physical forces used to perform the task; and A traction controller is executed on the current trajectory to determine a traction force for pulling contact pairs on the robot associated with non-zero stiffness toward corresponding contact pairs in the environment.

14. The method according to claim 12, further comprising the steps of: assigning a second penalty value as an updated value to the stiffness associated with the virtual force, and wherein the assigned penalty value is less than a penalty value assigned in a previous iteration if a pose constraint is not satisfied, wherein the pose constraint includes information about a position error and an orientation error associated with the trajectory; and Use continuous convexification to perform trajectory optimization problems.

15. A non-transitory computer-readable storage medium having a program embodied thereon, the program being executable by a processor for performing a method of moving an object from an initial pose of the object to a target pose of the object, wherein: The non-transitory computer-readable storage medium stores a dynamic model and a relaxed contact model, the dynamic model representing one or more of geometric properties, dynamic properties, and friction properties of a robot and an environment, the relaxed contact model representing a dynamic interaction between the robot and the object via virtual forces generated by one or more contact pairs associated with a geometric shape on the robot and a geometric shape on the object, wherein the virtual force acting on the object at a certain distance in each contact pair is proportional to a stiffness of the virtual force, the method comprising the following steps: obtaining a current state of interaction between the robot and the object; and Iteratively determining associated control commands for controlling the robot, a trajectory, and a virtual stiffness value for moving the object according to the trajectory by performing an optimization until a termination condition is satisfied, the optimization minimizing the stiffness of the virtual force and minimizing the difference between the target pose of the object and a final pose of the object moved from the initial pose by the robot controlled according to the control commands via the virtual force generated according to the relaxed contact model, wherein, to perform at least one iteration, the method further comprises the following steps: determining a current trajectory, a current control command, and a current virtual stiffness value for a current penalty value of the stiffness of the virtual force by solving an optimization problem initialized with a previous trajectory and a previous control command determined during a previous iteration with a previous penalty value of the stiffness of the virtual force; updating the current trajectory and the current control command to reduce the distance between each contact pair to generate an updated trajectory and an updated control command to initialize the optimization problem in a next iteration; and updating the current value of the stiffness of the virtual force for use in the optimization in the next iteration; and moving the robot arm of the robot according to the trajectory and the associated control command, The non-transitory computer-readable storage medium also stores a pulling controller that uses a virtual force remaining after calculating the current trajectory to attract geometric shapes on the robot toward corresponding geometric shapes on the object in the environment to facilitate physical contact.