Systems and methods for adaptive compliance-based robot assembly

By installing force sensors and nonlinear compliant controllers on the robot's wrist and combining them with machine learning, the accuracy problem caused by changes in part position during robot assembly was solved, realizing a low-cost, high-efficiency adaptive assembly strategy that adapts to changes in part posture.

CN117062693BActive Publication Date: 2025-10-28MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280023077.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-25
Filing Date
2022-01-14
Publication Date
2025-10-28
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

In existing robotic assembly operations, it is difficult to accurately perform assembly tasks when faced with changes in the position of parts, especially when the part position determined by the industrial vision camera is not accurate enough, requiring complex and costly dedicated computer programs for adjustment.

Method used

An adaptive compliant control strategy is adopted. By installing force sensors on the robot's wrist and combining nonlinear compliant controllers and machine learning, the robot learns the mapping between force and posture changes and automatically adjusts its trajectory to adapt to changes in part position, reducing task-specific programming.

Benefits of technology

This enables robots to adaptively complete assembly operations under low-precision measurement conditions, reducing reliance on high-precision measuring devices, lowering deployment costs, and improving assembly accuracy and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117062693B_ABST
    Figure CN117062693B_ABST
Patent Text Reader

Abstract

A robot for performing assembly operations is provided. The robot includes: a processor configured to determine a control law for controlling a plurality of motor-driven manipulator arms based on an original trajectory; execute a self-exploration program to generate training data indicative of the original trajectory in space; and learn a nonlinear compliant control law using the training data, the nonlinear compliant control law comprising a nonlinear mapping that maps measurements from the robot's force sensors to a corrected direction of the original trajectory defining the control law. The processor transforms the original trajectory according to a new target posture to generate a transformed trajectory, updates the control law based on the transformed trajectory to generate an updated control law, and commands the plurality of motors to control the manipulator arms according to the updated control law corrected by the compliant control law.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to robot assembly, and more specifically, to adaptive compliant robot assembly based on providing corrective measures. Background Technology

[0002] Various types of robotic devices have been developed to perform a variety of operations such as material handling, transportation, welding, and assembly. Assembly operations can correspond to connecting, coupling, or positioning a pair of parts in a specific configuration. Robotic devices include various components designed to assist the robotic device in interacting with its environment and performing operations. These components can include robotic arms, actuators, end effectors, and various other robotic manipulators.

[0003] In robotic assembly operations, parts are typically placed together or assembled using one or more robotic arms. For example, Cartesian and SCARA robots with three or four degrees of freedom (DoF) are used for this purpose; however, these robots have limitations on the permissible positions of the parts. In industrial applications, the strict repeatability of the part positions and / or orientations is assumed, so open-loop programs can be executed easily and correctly. Under this assumption, a fixed sequence of operations to be performed by the robotic device can be taught to the robot by a human operator using a teaching pendant. The teaching pendant stores the coordinates of via points in the robot's memory and can traverse these via points during runtime without any modifications. In many cases, no further programming of the robotic device is required. Summary of the Invention

[0004] [Technical Issues]

[0005] However, complexity arises when the positions of the parts involved in the assembly operation change between repetitions. This can occur when parts are placed on a surface by a feeder and end up in a different position each time; it can also occur when parts arrive at a moving conveyor belt. In such cases, industrial vision cameras can be used to determine the position of the parts. However, the part positions determined by industrial vision cameras are generally inaccurate. In other words, the part positions determined by industrial vision cameras are not accurate enough for assembly operations during various industrial robot applications.

[0006] Furthermore, even if the robotic device knows the exact location of the part, its path (defined by the pass points) still needs to be modified to adapt to the change in the part's position. In practice, this modification is performed using a dedicated computer program that takes the changed position of the part as input and outputs a new path for the robot. However, developing such a dedicated computer program is typically very difficult and laborious, and is currently one of the major components of the high cost of deploying and upgrading robotic devices for new assembly operations.

[0007] Therefore, a system is needed to accurately perform different assembly operations by utilizing the variable positions of the parts to be assembled without requiring specific programming of the operation.

[0008] [Solution to the problem]

[0009] One objective of some implementations is to provide a method for learning adaptive assembly strategies (AAS) for a wide variety of tasks from both human demonstrations and the robot's own self-experimentation, reducing or avoiding task-specific robot programming. Alternatively, another objective of some implementations is to provide such learned AAS that adapts to changes in one or a combination of the robot's initial and end-effector postures required for successful assembly operations.

[0010] Additionally or alternatively, another objective of some embodiments is to provide an AAS that can adapt to attitude modifications even when the initial and / or target attitude is not precisely known. As used herein, the attitude is not precisely known when the accuracy of the measured or estimated attitude is lower than the accuracy required for the assembly operation. Therefore, another objective of some embodiments is to provide an AAS suitable for end-effector attitude modification, which includes a change in at least one or a combination of a new initial attitude and a new target attitude of the wrist of the robotic arm, measured by a measuring device with an accuracy lower than the tolerance of the assembly operation.

[0011] To overcome this limitation of the measuring device, the robot is equipped with additional force / torque sensors mounted on the robot's wrist or below the platform holding immovable parts used in the assembly. For example, one implementation aims to control the robot to follow a trajectory of the robot's end-effector modified by both the target posture and possible forces generated by contact encountered along the trajectory. In this way, the force can be used to correct inaccuracies in posture estimation.

[0012] Without loss of generality, some implementations follow the following example. The target orientation of an end effector, such as a gripper, is implicitly determined by the placement of a fixed object A attached to a work surface. A robot is holding a second object B in its gripper, and the purpose of the assembly operation is to place the two objects (typically in close contact) together, for example, by inserting object B into object A. At the successful completion of the assembly operation, the orientation of the end effector is considered to be in the target orientation. By this definition, achieving the target orientation of the end effector for a given position of fixed object A is equivalent to the successful execution of the assembly operation. Furthermore, for different executions of this assembly operation, one or a combination of objects A and B may be in a new, different orientation. This can occur, for example, when objects are placed onto a surface by a feeder and each time ends in a different orientation; additionally, it can occur when object A arrives on a moving conveyor belt. In this case, various measuring devices, such as vision cameras, can be used to determine the orientation of objects A and B. However, the accuracy of the measurements provided by such devices is lower than the accuracy (tolerance) specified by the assembly operation.

[0013] Some implementations are based on the understanding that a primitive trajectory of motion for a gripper performing the assembly operation of objects A and B can be designed. Examples of such a trajectory include one or a combination of gripper posture and gripper velocity that vary over time. An example of a control law that tracks the trajectory is... in, It is to achieve the desired trajectory y d (t) the required velocity (the change in relative position per unit time step), and The actual speed is achieved by a low-level robot controller. However, due to imperfections in the robot's control and measurement devices, the gripper cannot always be controlled precisely along the original trajectory. For example, almost all industrial robot controllers introduce small errors while following the desired trajectory.

[0014] Therefore, some implementations are based on the understanding that control laws can be combined with compliant control laws to compensate for imperfections in the robot's control and measuring devices. In this case, the actuator can use a measured force value to move the gripper linearly in the direction opposite to the force. Here, an example of a control law is... Here, τ is the force measured by a force sensor, and K is a linear diagonal matrix with a predetermined value depending on the required compliance of the gripper against encountered obstacles. However, some implementations are based on the understanding that such a linear compliance control law is insufficient when the inaccuracy of the measuring device exceeds the accuracy of the assembly operation. For example, in the scenario of inserting a bolt into a hole, if the bolt experiences a vertical force due to collision with the edge of the hole, the stiffness control law with the diagonal matrix K cannot generate horizontal movement toward the center of the hole. In this case, it is necessary to actively interpret the measured force and generate corrective motion based on the measured force.

[0015] To this end, some implementations use nonlinear compliant controllers that map the forces experienced by the robot to changes in attitude and velocity in a nonlinear manner to modify linear compliant control. In this example, the control law takes the following form. Here, H is a nonlinear mapping that generates a correction for the robot's velocity. Some implementations are based on the understanding that such a control law, combining a trajectory with a nonlinear compliant controller, can be determined for a specific assembly operation along a particular trajectory and repeated an arbitrary number of times by the same type of robot for the same assembly operation. However, the control law should be modified accordingly when the starting or target posture of the assembly operation changes, which is challenging without additional learning. In other words, the aim of some implementations is to adapt the original control law learned for the original trajectory of the robot assembly operation in response to changes in the starting and / or target posture of the robot assembly operation. Transform into To be used based on the transformed trajectory Force mapping H new (τ) is used for control.

[0016] Some implementations are based on the understanding that, for various practical applications, an affine mapping of the original trajectory can be used to transform it into a transformed trajectory connecting a new initial and target pose. For example, the original trajectory can be represented using dynamic motion primitives (DMPs). A DMP is a set of parameterized ordinary differential equations (ODEs) that can generate trajectories that transform a system such as a robot from an initial pose to a target pose. A DMP can easily adjust the trajectory according to the new initial and target states, thus essentially constituting a closed-loop controller. Furthermore, a DMP can be learned from a finite number of training examples (even a single example). Therefore, the original trajectory can be modified in response to changes in the initial and target poses.

[0017] However, adapting a nonlinear mapping learned for the original trajectory to the modified trajectory is challenging and may even be impossible in online setups for assembly operations. Some implementations are based on the understanding that if the original trajectory is modified according to changes in the initial and / or target attitude, the nonlinear mapping learned for the original trajectory is valid for the transformed trajectory without any additional adjustments. This understanding can be explained by the nature of the forces generated due to contact between objects. The sign and magnitude of the force depend entirely on the relative positions of the two objects, not their absolute positions in space. For this purpose, if one object moves to a new position (undergoing an affine rigid body transformation of its coordinates) and another object approaches it along a similarly transformed trajectory, the same force can be generated.

[0018] Therefore, this understanding allows some implementations to determine the control law offline, making the offline control law suitable for online adjustment. Specifically, some implementations determine the original trajectory and a nonlinear mapping for the original trajectory offline, and modify the original trajectory online (i.e., during assembly operations) to adapt to a new starting or target posture, and control the robot based on the transformed trajectory and the nonlinear mapping learned for the original trajectory. In this case, the control law is... In this way, various implementations can be adapted to changes in the initial and / or target attitude of the assembly measurement using a measurement method with lower accuracy than the assembly operation.

[0019] One implementation aims to provide such a control law with minimal task-specific robot programming. Some implementations are based on the understanding that a DMP (Distributed Manipulation Program) for an initial trajectory can be learned through demonstration. For example, assuming the object's position is fixed for the initial trajectory, a human operator can teach the robot a fixed sequence of operations to perform using a teaching pendant or joystick with an appropriate number of degrees of freedom. This teaching pendant or joystick stores the coordinates of points in the robot's memory and can traverse these points at runtime without any modification. A DMP can be generated (learned) from these points without any further programming of the robot, resulting in relatively fast and low-cost deployment. However, it is also necessary to determine the nonlinear force mapping H(τ) with minimal human involvement.

[0020] For example, some implementations learn the nonlinear mapping of the controller via machine learning (e.g., deep learning). In this way, in cases where new assemblies need to be deployed to adapt to new starting or ending postures measured with insufficient accuracy, trajectories can be determined by human demonstration, while the nonlinear mapping can be learned through training using a self-exploration process, thereby minimizing human involvement.

[0021] While the fixed object A remains in its original posture (for which the human operator has specified at least one trajectory y for successfully completing the assembly operation),... d While learning the nonlinear mapping H(τ) through self-experimentation, the robot intentionally introduces random variations into the trajectory y. d In (t), the displacement d(t) relative to the original trajectory is obtained, while following the associated velocity profile. Repeatedly follow the trajectory y d (t). If time t is recorded. k Contact force τ k =τ(t) k When the trajectory changes by d k =d(t) k When ), then under certain conditions, force τ can be used. k To infer the displacement d relative to the correct trajectory k .

[0022] Some implementations are based on the understanding that they rely on the ability of a human demonstrator to guide the robot safely during demonstrations, using the original trajectory y demonstrated by the demonstrator. d (t) can be assumed to be safe and collision-free, for the original safe trajectory y d (t) The deliberately perturbed modified trajectory y d (t)+d(t), but this is not the case. Some implementations are based on the further understanding that when following the modified trajectory y d (t)+d(t) instead of the original trajectory y d At time (t), the object to be assembled may collide or become blocked. Therefore, some implementations are based on traversing the modified trajectory y in a safe manner without damaging the robot or the object being assembled. d The objective of (t)+d(t).

[0023] In some implementations, robots equipped with force sensors include safety measures that shut down the robot when the sensed force exceeds a threshold, thereby protecting the robot from damage due to collisions. However, when the object being assembled is a precision part (e.g., an electronic component), the threshold can be higher to protect the object being assembled. Some implementations are based on the understanding that modified trajectories can be safely executed by using a compliant controller that responds to and minimizes the forces experienced. For example, a linear compliance law with a diagonal matrix K can be used.

[0024] In implementation, the terms of the diagonal matrix can be determined based on the maximum force that can be safely applied in a specific direction. For example, if the maximum force along the z-direction / axis is f zmaxAnd the maximum velocity along the z-direction is (where the desired velocity along the z-direction is) Then the elements k of the diagonal matrix K z The value can be These elements of the diagonal matrix K ensure that, in the case of an obstacle along the z-direction, the desired velocity along the z-direction is achieved. Size always obeys At that time, due to the force f experienced along the z-direction z (t) leads to correction k z f z (t)(if the desired speed) If it is positive, then it should normally be negative, and vice versa. This allows the robot to stop at 150 as desired. Where |f z (t)|≤f zmax Here, f z It is the component of the vector τ(t) corresponding to the force sensed along the z-direction, and the remaining terms of the diagonal matrix K corresponding to the other two linear contact forces (along x and y) and the three moments around these axes can be determined similarly.

[0025] According to the implementation method, this can be achieved by providing a compliant (impedance or admittance) controller with a stiffness matrix K at discrete time t. k =kΔt (where Δt is the control step size) is the target position of a series of commands. To achieve linear compliance on the robotic arm The execution of the command. To calculate the target position y. c,k Calculate the position of the initial command. The position of the initial command is given as y. c,0 =y d (0), that is, it is related to the original trajectory y. d The initial position of (t) coincides. Additionally, in each subsequent control step k, the actual realized position y of the robot (or the robot's end effector) is measured. r,k =y(t) k Additionally, the target position y of the next control step command. c,k+1 Calculated as Here, Δy k The time step k is introduced into the original trajectory y. d The change / displacement of (t). This is determined by the change in y. c,k+1 The calculation uses y r,k Instead of y c,k The robot follows the velocity profile Instead of implicit location profile (trajectory) y d (t kThis ensures that when the robot's movement stops due to a collision, the error between the actual position and the commanded target position does not accumulate over time. Alternatively, if y r,k If the position y of each new command remains constant due to obstacles, collisions, or blockages, then... c,k+1 Only relative to y r,k small relative displacement Furthermore, based on the diagonal matrix K specifying the degree of motion compliance, the compliance controller can tolerate the absence of displacement. Without reaching excessive contact force.

[0026] In applying linear compliance law During this period, the measured position y r,k The time series indicates the robot's actual position in each control step. This is achieved by measuring the position y. r,k With the robot (or end effector) based on the original trajectory y d (t) should be compared with the position, and the displacement in each control step is calculated as d. k =d(t) k )=y(t k )-y d (t k )=y r,k -y d (t k ).

[0027] The self-experiment described above can be followed multiple times, each time starting from the original trajectory y. d Starting from the same initial position in (t), different displacements are applied at different time points. The displacements can be systematic, for example, at a single moment, before contact occurs between the objects, while the robot is still in free space, and in a plane perpendicular to the robot's motion at that moment, only one displacement is introduced. This displacement yields a displacement relative to the original trajectory y. d (t) The corrected trajectory with constant offset. In some implementations, the displacement can also be random, i.e., achieved by adding small random variations sampled from a probability distribution (e.g., a Gaussian distribution) at each time step.

[0028] As the original trajectory y is traversed multiple times with different displacements d The results of (t) collect data relating the direction and magnitude of displacement to the forces experienced as a result. When the robot is moving in free space without contact, the force relative to the original trajectory y cannot be inferred from the contact force. d The displacement of (t) is such that the force acting on it is zero. Therefore, τ k The moment when τ = 0 is discarded. For each of the remaining cases, i.e., when τ = 0, the time is discarded. k When ≠, for (τ)i , d i ) form of training examples are added to the database of training examples, where τ i = τ k and d i = d k , where i is the index of the pair in the database.

[0029] When a sufficient number N of training examples are collected, a supervised machine learning algorithm can be used to learn the mapping between the forces and displacements that led to them. However, since the mapping from displacement to force is typically many-to-one (multiple displacements can sometimes result in the same force), the inverse mapping can be one-to-many, i.e., not a function that can be learned with machine learning. This ambiguity in the mapping challenges the possibility of learning a non-linear compliant controller.

[0030] However, some implementations are based on the recognition that it is not necessary to recover the exact magnitude of the displacement to successfully perform a corrective action, and furthermore, as long as the magnitude of the displacement does not exceed the radius R of the object B being inserted, multiple displacements can only produce the same force if the signs of the multiple displacements are the same. Based on this recognition, a supervised machine learning algorithm is used to learn the mapping sign(d<s i ) ≤ R for all examples i = 1, N such that ‖d i ‖ = H0(τ i ), where ‖d i ‖ is the L2 norm of the displacement d i . When the radius of the object B being inserted is known, it can be provided to the supervised machine learning algorithm. When it is unknown, the radius of the object B can be found by searching for the maximum value of R that gives a good fit to the training examples with the magnitude constraint ‖d i ‖ ≤ R. Thus, a non-linear mapping that maps the measurements of the force sensor to the direction of correction of the original trajectory is learned. After learning the mapping H0(τ), the desired mapping H(τ) can be obtained by scaling it with a suitable velocity constant v0: H(τ) = v0H0(τ), where the value of v0 is pre-determined by the application designer. In some implementations, the value of v0 does not exceed a value determined based on the radius R of the object B being inserted. For example, the velocity and the radius have different measurement units m and m / s, making a direct comparison impractical or at least inconvenient. Therefore, some implementations compare the radius R with the distance v0*dt traveled per control step, where dt is the duration of the control step, and ensure that v0*dt < R, so that the movement does not exceed the hole.

[0031] Therefore, one embodiment discloses a robot comprising: a robotic arm including a wrist having multiple degrees of freedom of motion, wherein, during robot operation, force sensors are arranged to generate measurements indicating the forces experienced by the end-effector of the robotic arm during operation; a plurality of motors configured to change the motion of the robotic arm according to commands generated according to a control law; at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the processor to receive data of the original trajectory of the robotic arm motion instructing it to change the posture of the end-effector from an initial posture to a target posture to perform an assembly operation. Afterwards: A control law for controlling multiple motors to move the robotic arm according to the original trajectory is determined; a self-exploration program is executed to explore the space of the original trajectory by: controlling multiple motors according to the control law while introducing different displacements relative to the original trajectory into the state of the robotic arm; and recording the measurements of paired force sensors and the corresponding displacement values ​​relative to the original trajectory when a force is detected on the end-effector of the robotic arm to generate training data indicating the space of the original trajectory; and using the training data to learn a nonlinear compliant control law including a nonlinear mapping that maps the force sensor measurements to the corrected direction of the original trajectory defining the control law. Instructions executed by at least one processor cause the processor, upon receiving an end-effector posture modification of the original trajectory including at least one or a combination of a new starting posture and a new target posture of the end-effector of the robotic arm measured with a lower accuracy than the assembly operation: transforming the original trajectory according to the end-effector posture modification to generate a transformed trajectory; updating the control law according to the transformed trajectory to generate an updated control law; and commanding multiple motors to control the robotic arm according to the updated control law corrected by the compliant control law learned for the original trajectory.

[0032] The currently disclosed embodiments will be further described with reference to the accompanying drawings. The drawings shown are not necessarily to scale, but generally focus on illustrating the principles of the currently disclosed embodiments. Attached Figure Description

[0033] [ Figure 1A ]

[0034] Figure 1A The configuration of the robot according to some implementation methods is shown.

[0035] [ Figure 1B ]

[0036] Figure 1B An exemplary assembly operation according to an implementation method is shown.

[0037] [ Figure 1C ]

[0038] Figure 1C The results of the assembly operation are shown according to some embodiments, due to the inaccurate determination of the object's orientation by the imaging device.

[0039] [ Figure 2 ]

[0040] Figure 2 A schematic diagram is shown, according to some implementations, for determining control laws based on adaptive compliant control learning to eliminate ambiguity in nonlinear mappings.

[0041] [ Figure 3 ]

[0042] Figure 3 A block diagram of a robot for performing assembly operations according to some embodiments is shown.

[0043] [ Figure 4 ]

[0044] Figure 4 A schematic diagram of a Dynamic Motion Primitive (DMP) that learns an original trajectory by demonstration, according to some implementations, is shown.

[0045] [ Figure 5A ]

[0046] Figure 5A The execution of a robot’s self-exploration procedure according to some implementations is illustrated.

[0047] [ Figure 5B ]

[0048] Figure 5B A schematic diagram of the target location of a calculation command according to some implementations is shown.

[0049] [ Figure 5C ]

[0050] Figure 5C A schematic diagram illustrating an overview of the learning of a nonlinear compliant control law according to some implementations is shown.

[0051] [ Figure 6 ]

[0052] Figure 6 A schematic diagram is shown illustrating the use of a nonlinear compliant control law to correct the transformed trajectory according to some implementations.

[0053] [ Figure 7A ]

[0054] Figure 7A The calculation of displacement is illustrated under the alignment condition where the bottom of a movable object contacts the edge of an immovable object, according to some embodiments.

[0055] [ Figure 7B ]

[0056] Figure 7B The displacement under alignment conditions where the edge of a movable object contacts the surface of a fixed object according to some embodiments is illustrated.

[0057] [ Figure 7C ]

[0058] Figure 7C The displacement of a movable object within a fixed object under alignment conditions according to some embodiments is illustrated.

[0059] [ Figure 8 ]

[0060] Figure 8 Examples are illustrated of robots configured to perform assembly operations in industrial facilities according to some embodiments. Detailed Implementation

[0061] In the following description, numerous specific details are set forth for illustrative purposes in order to provide a thorough understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure may be practiced without these specific details. In other instances, apparatus and methods are shown in block diagram form in order to avoid obscuring this disclosure.

[0062] As used in this specification and claims, the terms “for example,” “e.g.,” and “such as,” as well as the verbs “comprising,” “having,” “including,” and other verb forms thereof, when used in conjunction with a list of one or more parts or other articles, shall be interpreted as open-ended, meaning that the list shall not be considered to exclude other additional parts or articles. The term “based on” means at least partially based on. Furthermore, it should be understood that the words and terms used herein are for descriptive purposes and should not be considered limiting. Any headings used in this specification are for convenience only and have no legal or limiting effect.

[0063] Figure 1A The diagram illustrates a mechanical configuration 100 of a robot 150 according to some embodiments. The robot 150 includes a robotic arm 101 for performing assembly operations. The robotic arm 101 includes a wrist 102 for ensuring multiple degrees of freedom for moving an object. In some implementations, the wrist 102 is provided with an end-tool 104 for holding an object 103 and / or for performing any other robotic operation such as assembly. For example, the end-tool 104 may be a gripper. Hereinafter, "end-tool" and "gripper" may be used interchangeably. According to embodiments, the purpose of the assembly operation is to place two parts (typically in close contact) together. For example, inserting an object along a trajectory into another object to assemble a product.

[0064] Figure 1B An exemplary assembly operation according to an embodiment is shown. Combined with, for example... Figure 1A The robot 150 shown is used to explain Figure 1B Robot 150 is configured to perform assembly operations, such as inserting object 103 into another object along a trajectory. As used herein, a trajectory corresponds to a path defining the movement of object 103 held by gripper 104 for performing the assembly operation. In simple scenarios, the trajectory may only specify the vertical movement of wrist 102. However, since wrist 102 includes multiple degrees of freedom, the trajectory may include a motion profile across multidimensional space.

[0065] An object's pose refers to the combination of its position and orientation. Gripper 104 initially holds a movable object 103 (e.g., a nail) in a starting pose 111. The pose of gripper 104 corresponding to the starting pose 111 is called the starting pose of gripper 104. According to an embodiment, the purpose of the insertion operation is to insert the movable object 103 into a fixed object 112 in a pose 115, wherein object 112 includes a hole for receiving object 103. The pose 115 of object 112 can refer to the position and / or orientation of object 112. Robot 150 is configured to move gripper 104 along trajectory 113 to insert and place object 103 in the hole of object 112 in a pose 114. The pose 114 of object 103 in the hole of object 112 is called the goal pose. The pose of gripper 104 corresponding to the goal pose is called the goal pose of gripper 104.

[0066] The target orientation of the gripper 104 is determined based on the position of the object 112. Upon successful completion of the insertion operation, the orientation of the gripper 104 of the robot arm 101 is considered to have achieved the target orientation. Therefore, achieving the target orientation of the gripper 104 is equivalent to successful execution of the insertion operation. According to the embodiment, the trajectory 113 is defined based on the initial and target orientations of the gripper 104 and the orientation 115 of the object 112. Furthermore, this assembly operation can be repeatedly performed by the robot 150.

[0067] Some implementations are based on the understanding that the orientation of object 103 and object 112 involved in the assembly operation may change between repetitions of the assembly operation, thereby positioning one or a combination of objects 103 and 112 in different orientations. For example, when object 112 arrives on the moving conveyor belt, object 112 may not always arrive at the moving conveyor belt in a specific orientation (e.g., orientation 115). Therefore, object 112 may end in different orientations. Thus, the change in the orientation (orientation and position) of object 112 involved in the assembly operation results in at least one or a combination of a new starting orientation and a new target orientation, which is referred to as end-position modification. Since the trajectory is defined based on the starting and target orientations of gripper 104 and the orientation 115 of object 112, trajectory 113 cannot be used for different assembly operations involving orientations other than those mentioned above. In this case, various measurement assemblies are used to determine the orientations of objects 103 and 112. According to some implementations, measurement assemblies determine the new starting orientation and the new target orientation of gripper 104. The measuring device includes an imaging device 106, such as an industrial vision camera. In some implementations, a single imaging device may be used.

[0068] However, the accuracy of the poses of object 103 and object 112 determined by such a camera is lower than the accuracy of the assembly operation. For example, unless expensive imaging devices are used, industrial vision cameras have an error of about 1-2 mm in pose determination. Such an error is at least an order of magnitude larger than the tolerance required in precision insertion operations (which can be about 0.1 mm). Therefore, due to the significant inaccuracy in the determined poses of objects 103 and 112, the object to be inserted (e.g., 103) may collide with a portion of another object (e.g., 112) involved in the assembly operation.

[0069] Figure 1C The following diagram illustrates the results of assembly operations due to inaccurate determination of the orientation of object 103 by the imaging device, according to some embodiments. (Combined with...) Figure 1A and Figure 1B The robot 150 shown in the image is used to explain... Figure 1C For example, the orientation 115 of object 112 ( Figure 1B (As shown in the image) may change, and imaging device 106 may determine that attitude 115 has changed to attitude 116. In particular, imaging device 106 may determine that object 112 is located at position 116. When the position 115 of object 112 changes to position 116, the target attitude 114 ( Figure 1B(As shown in the diagram) can be changed to the target pose 118. Based on pose 116 and target pose 118, trajectory 113 is transformed into trajectory 117. However, if the true position of object 112 is not accurately determined and is a certain distance 119 from the determined position 116, the trajectory 117 does not lead to correct insertion, and a collision may occur between object 103 and a portion of object 112 (e.g., edge 120). As a result, object 103 is displaced, and object 103 may remain in an incorrect pose 121. Additionally, due to this collision, the gripper 104 of the robotic arm 101 may be subjected to forces specific to pose 121.

[0070] Therefore, some implementations recognize that the posture determined solely by the imaging device 106 is insufficient to successfully perform the assembly operation. To overcome this limitation of the imaging device 106, an adaptive assembly strategy (AAS) 107 is used. AAS 107 is based on the understanding that the forces experienced during the assembly operation can be used to correct inaccuracies in the posture determination of the imaging device 106. For this purpose, the robot 150 is equipped with a force sensor. For example, a force sensor 105 is operatively connected to the wrist 102 or end-effector of the robotic arm 101. The force sensor 105 is configured to generate a measurement 108 (also referred to as force sensor measurement 108) of the force and / or torque experienced by the end-effector (gripper 104) of the robot 150 during the assembly operation. In some implementations, the robot 150 is equipped with a torque sensor for measuring the torque experienced by the end-effector 104. Some implementations are based on the understanding that the force sensor measurement 108 can be used to correct the trajectory 117 to achieve the target posture 118.

[0071] To this end, a nonlinear mapping 109 is determined for trajectory 113. The nonlinear mapping maps force sensor measurements 108 to corrections for trajectory 117 in a nonlinear manner. In other words, the nonlinear mapping provides corrections for trajectory 117 of robot 150 during assembly operations along trajectory 117. This correction may include displacements of object 103 that allow for the achievement of a new target posture. For this purpose, the nonlinear mapping provides a mapping between force and displacement. In alternative embodiments, this correction may correspond to posture and / or velocity corrections. Trajectory 113 is referred to as the “original trajectory.” As explained below, the original trajectory is the trajectory used to determine the nonlinear mapping.

[0072] Some implementations are based on the understanding that a nonlinear mapping can be determined for a specific assembly operation along a particular trajectory (e.g., trajectory 113) and repeated any number of times for the same assembly operation performed by the same robot as robot 150. However, when the initial and / or target poses involved in the assembly operation change, the original trajectory 113 is transformed accordingly to produce a transformed trajectory. Subsequently, the nonlinear mapping determined for the original trajectory 113 may need to be modified based on the transformed trajectory (e.g., trajectory 117).

[0073] However, some implementations are based on the understanding that if the original trajectory 113 is transformed according to changes in the initial and / or target attitude, the nonlinear mapping determined for the original trajectory 113 without any additional adjustments 110 is valid for the transformed trajectory. For example, this understanding is true because the sign and magnitude of the force depend entirely on the relative positions of the two objects (e.g., object 103 and object 112), not on their absolute positions in space. Therefore, if one of object 103 and object 112 moves to a different position and the other object approaches it along a similarly transformed trajectory, the same force can be generated.

[0074] Therefore, this understanding allows some implementations to determine the original trajectory (e.g., trajectory 113) and the nonlinear mapping for the original trajectory offline, i.e., in advance, and to transform the original trajectory online, i.e., during the assembly operation, to accommodate changes in the initial / or target posture and to control the robot 150 based on the transformed trajectory and the nonlinear mapping determined for the original trajectory. In this way, various implementations can be adapted to changes in the initial and / or target posture measured using imaging devices 106, such as cameras, which have lower accuracy than the assembly operation. As a result, it allows the use of economical cameras during the assembly operation. In addition, it minimizes task-specific robot programming because the nonlinear mapping determined for the original trajectory can be retained for the transformed trajectory.

[0075] Nonlinear mappings can be determined through training. For example, supervised machine learning algorithms can be used to learn the mapping between force and displacement resulting from the force. This mapping is learned offline. The mapping from displacement to force is often many-to-one, i.e., multiple displacements can sometimes result in the same force. During online, i.e., real-time assembly operations, the inverse mapping can be used for corrections in the assembly operation. However, the inverse mapping can be one-to-many, i.e., the measured force can be mapped to multiple displacements, which is not a function that can be learned by machine learning. This ambiguity in the mapping challenges the possibility of learning nonlinear mappings. Some implementations are based on the understanding that adaptive compliant control learning can be used in AAS to eliminate ambiguity in the mapping of nonlinear compliant controllers.

[0076] Figure 2 A schematic diagram is shown, according to some embodiments, for determining a control law based on adaptive compliant control learning to eliminate ambiguity in nonlinear mappings. Some embodiments are based on the understanding that a trajectory 200 (e.g., trajectory 113) can be designed for the motion of the gripper 104 performing an assembly operation. Trajectory 113 includes one or a combination of the time-varying posture of the gripper 104 and the time-varying velocity of the gripper 104. A control law is determined to track trajectory 113. An example of such a control law is:

[0077]

[0078] in, It is to achieve the desired trajectory y d (t) the required velocity (the change in relative position per time step), and The actual speed is achieved by a low-level robot controller.

[0079] However, errors in the robot 150's control mechanism (such as actuators) and measuring devices make it difficult to precisely control the gripper 104 along trajectory 113. For example, in practice, industrial robot controllers often experience at least a small error while following the desired trajectory. To address this, some implementations are based on the understanding that a control law can be combined with compliant control to adjust for errors in the robot 150's control mechanism and measuring devices. In this case, a stiffness actuator can use force measurements from force sensor 105 to linearly move the gripper 104 in the direction opposite to the force. For this purpose, the control law can be given, for example, by the following equation:

[0080]

[0081] Wherein, τ is the force and / or torque measured by the force sensor 105, and K is a linear diagonal matrix with a predetermined value that depends on the degree of compliance required by the gripper 104 for the encountered obstacle.

[0082] However, this compliant control law is insufficient when the inaccuracy of the measuring device exceeds the accuracy of the assembly operation. For example, in an insertion operation where object 103 is inserted into a hole in object 112, if object 103 experiences a vertical force due to collision with the edge of the hole in object 112, the rigid control law with diagonal matrix K cannot generate horizontal movement toward the center of object 112. In this case, active interpretation of the measured force is required, and corrective motion must be generated based on the measured force.

[0083] Therefore, some implementations modify the control law (1) into a nonlinear compliant control law 201. The nonlinear compliant control law is obtained by using a nonlinear mapping of the control law (1). Therefore, the nonlinear compliant control law can be given by the following equation.

[0084]

[0085] Here, H is a nonlinear mapping (function) that generates a correction for the velocity of robot 150.

[0086] For a specific assembly operation along a specific trajectory, a control law (2) combining the trajectory with a nonlinear compliant control law can be determined. Therefore, in the event of a change in the initial and / or target attitude of the assembly operation, the original trajectory 113 is transformed according to the change in the initial and / or target attitude to generate a transformed trajectory. Furthermore, the control law based on the original trajectory 113... Transformed into

[0087]

[0088] It is used for control based on the transformed trajectory.

[0089] For example, in the context of Figures 1A to 1B As described in the description, if the original trajectory 113 is transformed according to the change in the initial and / or target attitude, the nonlinear mapping learned for the original trajectory 113 holds true for the transformed trajectory without any additional adjustments. Therefore, control law 202 (3) is modified according to the transformed trajectory and the nonlinear mapping learned for the original trajectory 113. In other words, control law 202 (3) is modified for the new attitude (the changed initial and / or target attitude) without changing the nonlinear mapping learned for the original trajectory 113. For this purpose, for example, the control law is updated to

[0090]

[0091] According to some implementations, the original trajectory 113 can be transformed into a transformed trajectory using an affine mapping of the original trajectory 113. In other implementations, the original trajectory 113 can be represented by a dynamic motion primitive (DMP). A DMP is a collection of parameterized ordinary differential equations (ODEs) that can generate a trajectory (e.g., trajectory 113) for implementing assembly operations. The DMP can easily adjust the original trajectory based on new starting and target poses, thus forming a closed-loop controller. In other words, the DMP of the original trajectory can accept new starting poses and new target poses to generate the transformed trajectory. In addition, the DMP can be learned from several training examples (or even a single example). Therefore, the control law (3) can be written as

[0092]

[0093] However, as in the case of Figures 1A to 1B As described in the description, the ambiguity in nonlinear mappings challenges the learning of nonlinear compliant control laws.

[0094] According to some implementations, adaptive compliant control learning is used to overcome ambiguity. Adaptive compliant control learning is based on the understanding that it is not necessary to recover the precise magnitude of the displacement for successful correction, and furthermore, multiple displacements can only produce the same force R if they have the same sign, provided the magnitude of the displacement does not exceed the radius R of the object being inserted (i.e., object 103). Based on this understanding, in adaptive compliant control learning, for all examples i = 1, N, ||d|| i ||≤R, use a supervised machine learning algorithm to learn the mapping sign(d) i )=H0(τ i )‖d i ||≤R. If the radius of object 103 is known, it can be provided to a supervised machine learning algorithm. If it is unknown, a search can be performed to find the radius with the constraint ||d. i The radius of object 103 is found by finding the maximum value of R for a well-fitted training example where || ≤ R. Therefore, a nonlinear mapping is learned to map the force sensor measurements to a correction direction of the original trajectory. After learning the mapping H0(τ) according to the implementation method, the mapping Hτ can be obtained by scaling it with an appropriate correction amplitude based on the velocity constant v0. For this purpose,

[0095] H(τ) = v0H0(τ),

[0096] Wherein, the value of v0 is a predetermined value. In some implementations, the value of v0 does not exceed the radius R of the object being inserted. Therefore, the nonlinear compliant control law for correcting the mapping of the force sensor 105 measurements to the original trajectory 113 is configured to use a predetermined amplitude (v0) of the correction v0, and to determine the direction of correction by a nonlinear function of the force measurements trained on the original trajectory 113. Thus, this implementation of learning and modifying the mapping eliminates the ambiguity that exists. Therefore, the control law (5) is updated to

[0097]

[0098] The control law (6) can also be written as

[0099]

[0100] To this end, AAS, including adaptive compliant control learning, eliminates ambiguity or problems in learning nonlinear mappings. Furthermore, this AAS can be applied to perform contact-rich assembly operations with variable start and target poses, provided that the accuracy of the determined position of the object being assembled is lower than the accuracy of the assembly operation. Robot 150 controls one or more of the robotic arm 101, wrist 102, or gripper 104 according to the updated control law (i.e., control law (5)) to perform the assembly operation.

[0101] Figure 3 A block diagram of a robot 150 for performing assembly operations according to some embodiments is shown. The robot 150 includes an input interface 300 configured to receive data indicating a raw trajectory (e.g., trajectory 113) for manipulator movements that transition the attitude of an end-effector 104 from a starting attitude to a target attitude to perform the assembly operation. The input interface 300 may also be configured to accept end-effector attitude modifications. End-effector attitude modifications include at least one or a combination of a new starting attitude of the end-effector 104 and a new target attitude of the end-effector 104, measured with a lower accuracy than that of the assembly operation. In some embodiments, the input interface 300 is configured to receive measurements indicating the forces experienced by the end-effector 104 during the assembly operation. These measurements are generated by a force sensor 105. The measurements may be raw measurements received from the force sensor or any derivative of the measured values ​​representing the forces experienced.

[0102] Robot 150 may have multiple interfaces for connecting robot 150 to other systems and devices. For example, robot 150 is connected to imaging device 106 via bus 301 to receive new starting and target postures via input interface 300. Alternatively, in some implementations, robot 150 includes a human-machine interface 302 connecting processor 305 to keyboard 303 and indicating device 304, wherein indicating device 304 may include a mouse, trackball, touchpad, joystick, pointer, stylus, or touchscreen, and others. In some embodiments, robot 150 may include motor 310 or more motors configured to change the movement of the robotic arm according to commands generated according to a control law. Additionally, robot 150 includes a controller 309. Controller 309 is configured to operate motor 310 to change the movement of robotic arm 101 according to a control law.

[0103] Robot 150 includes a processor 305 configured to execute stored instructions and a memory 306 storing instructions executable by the processor 305. The processor 305 may be a single-core processor, a multi-core processor, a computing cluster, or any other configuration. The memory 306 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. The processor 305 is connected to one or more input interfaces and other devices via a bus 301.

[0104] Robot 150 may also include a storage device 307 suitable for storing different modules of instructions executable by storage processor 305. Storage device 307 stores the original trajectory 113 of the motion of robotic arm 101 that transforms the posture of end tool 104 from a starting posture to a target posture to perform assembly operations. The original trajectory 113 is stored in 307 in the form of dynamic motion primitives (DMPs) including ordinary differential equations (ODEs).

[0105] Storage device 307 also stores a self-exploration program 308 for generating training data indicative of the space of the original trajectory 113. Storage device 307 may be implemented using a hard disk drive, optical drive, thumb drive, drive array, or any combination thereof. Processor 305 is configured to determine a control law for controlling multiple motors to move the robotic arm according to the original trajectory and to execute self-exploration program 308, which explores the space of the original trajectory by controlling multiple motors according to the control law while introducing different displacements relative to the original trajectory into the state of the robotic arm, and by recording the measurements of pairs of force sensors and the corresponding displacement values ​​relative to the original trajectory when a force is detected on the end-effector of the robotic arm to generate training data indicative of the space of the original trajectory. Processor 305 is also configured to use the training data to learn a nonlinear compliant control law including a nonlinear mapping that maps the force sensor measurements to a correction direction of the original trajectory defining the control law.

[0106] Additionally, in some embodiments, processor 305 is also configured to transform the original trajectory based on end-effector posture modification to generate a transformed trajectory, and to update the control law based on the transformed trajectory to generate an updated control law. Processor 305 is also configured to command multiple motors to control the robotic arm according to the updated control law, corrected using a compliant control law learned for the original trajectory.

[0107] Figure 4A schematic diagram illustrating a DMP (Discretionary Manipulator) for learning an original trajectory 113 through demonstration, according to some embodiments, is shown. Some embodiments are based on the understanding that a DMP for learning an original trajectory 113 can be learned through demonstration. In the demonstration, a fixed object (i.e., object 112) is fixed in pose 115. Furthermore, a human operator 400 performs the demonstration by guiding a robot 150 holding object 103 in its gripper 104 along the original trajectory 113, which has been successfully assembled. According to embodiments, the human demonstrator can guide the robot 150 to follow the original trajectory using a teaching pendant 401, which stores the coordinates of the points corresponding to the points traversed on the original trajectory 113 in the robot 150's memory 306. The teaching pendant 401 can be a remote control device. This remote control device can be configured to send robot configuration settings (i.e., robot settings) to the robot 150 for demonstrating the original trajectory 113. For example, the remote control device sends control commands such as movement in the XYZ directions, speed control commands, joint position commands, etc., for demonstrating the original trajectory 113. In an alternative implementation, a human operator 400 may guide the robot 150 using a joystick, kinetic feedback, or similar means. The human operator 400 may guide the robot 150 to repeatedly track the original trajectory 113 for the same fixed object 112 in a fixed posture 115.

[0108] Trajectory 113 can be represented as y d (t), where t is in [0,T], is the point (attitude and pose) through which the end-effector of robot 150 passes in Cartesian space.

[0109] After recording one or more trajectories y(t) for the same fixed pose 115, processor 305 is configured to apply a DMP learning algorithm to learn a separate DMP for each component of y(t). This DMP is in the form of two coupled ODEs, for example,

[0110] and

[0111] Where f(x,g) is the forced function and can be given by the following equation:

[0112]

[0113] With the help of parameter w i This is used to parameterize the forced function. According to some implementations, the parameters w are obtained through least-squares regression from the trajectory y(t). i In this way, a set of DMPs is determined by applying the DMP learning algorithm. Given a new target pose g... dBy integrating the ODE of the DMP forward in time from the starting position without any additional demonstration or programming, this set of DMPs can generate new desired trajectories y. new (t).

[0114] Some implementations aim to determine the nonlinear mapping with minimal human intervention. To this end, some implementations aim to learn the nonlinear mapping of the controller via training (e.g., deep learning). In this way, in cases where a new insert assembly needs to be deployed to adapt to new end-effector poses measured with insufficient accuracy, the trajectory can be determined via human demonstration, while the nonlinear mapping can be learned through training using a self-exploration procedure, thereby minimizing human intervention. Specifically, robot 150 receives the original trajectory 113 as input. In response to receiving the original trajectory 113, robot 150 executes a self-exploration procedure.

[0115] Figure 5A The execution of a self-exploration procedure by a robot 150 according to some embodiments is illustrated. The end effector 104 of the robotic arm 101 is configured to track an original trajectory y by controlling multiple motors according to a control law. d (t)113, to insert object 103 into fixed object 112. The execution of the self-exploration procedure includes... (The sentence is incomplete and requires more context to translate accurately.) d The displacement of (t)113 is introduced into the state of the robotic arm while exploring the original trajectory y. d (t)113 space y d (t). For example, a displacement d(t)500 relative to the original trajectory 113 is introduced at the end tool 104 of the robotic arm 101. Therefore, the end tool 104 may experience a force τ. The force experienced by the end tool 104 is measured by a force sensor (e.g., force sensor 105) arranged at the end tool 104. In addition, the robot 150 records the measurements of the paired force sensors and the corresponding values ​​relative to the original trajectory y. d The value of the displacement (t)113.

[0116] Some implementations are based on the understanding that they rely on the ability of a human demonstrator to ensure safety while guiding the robot 150 during the demonstration, based on the original trajectory y demonstrated by the demonstrator. d (t)113 can be assumed to be safe and collision-free, but for the original safe trajectory y d (t)113 The deliberately disturbed modified trajectory y d (t)+d(t), but this is not the case. Some implementations are based on the further understanding that when following the modified trajectory y d (t)+d(t) instead of the original trajectory y d(t) At this time, the objects to be assembled (e.g., objects 103 and 112) may collide or become blocked. Therefore, some implementations are based on traversing the modified trajectory y in a safe manner without damaging the robot 150 or the objects being assembled. d The objective of (t)+d(t).

[0117] In some implementations, the robot 150 equipped with force sensor 105 includes a safety measure that shuts down the robot 150 when the sensed force exceeds a threshold, thereby protecting the robot 150 from damage due to collision. However, when the object being assembled is a precision part (e.g., an electronic component), the threshold may be high to protect the object being assembled. Some implementations are based on the understanding that modified trajectories can be safely practiced by using a compliant controller that responds to and minimizes the forces experienced. For example, a linear compliance law with a diagonal matrix K (also known as a stiffness matrix) can be used.

[0118] In implementation, the terms of the diagonal matrix can be determined based on the maximum force that can be safely applied in a specific direction. For example, if the maximum force along the z-direction / axis is f zmax And the maximum velocity along the z-direction is The desired velocity along the z-direction is Then the elements k of the diagonal matrix K z The value can be These elements of the diagonal matrix K ensure that, in the case of an obstacle along the z-direction, the desired velocity along the z-direction is achieved. Size always obeys At that time, due to the force f experienced along the z-direction z (t) leads to correction k z f z (t)(if the desired speed) If positive, then normal is negative, and vice versa. This can stop robot 150 as desired, that is, Where |f z (t)|≤f zmax Here, f z It is the component of the vector τ(t) corresponding to the force sensed along the z-direction, and similarly, the remaining terms of the diagonal matrix K corresponding to the other two linear contact forces (along x and y) and the three moments around these axes can be determined.

[0119] According to the implementation method, this can be achieved by providing a compliant (impedance or admittance) controller with a stiffness matrix K at discrete time t. k =kΔt (where Δt is the control step size) is the target position of a series of commands. To achieve linear compliance on robotic arm 101 The execution.

[0120] Figure 5B The target position y of the calculation command according to some embodiments is shown. c,k The diagram illustrates this. In box 501, the position of the initial command is calculated. The position of the initial command is given by y. c,0 =y d (0), that is, it is related to the original trajectory y. d The initial position of (t) coincides. In box 502, in each subsequent control step k, the robot (or end effector 104) y is measured. r,k =y(t) k The actual implementation location y r,k =y(t) k Additionally, in box 503, the target position y for the next control step command is specified. c,k+1 Calculated by processor 305 as Here, Δy k This involves introducing the time step k into the change / displacement of the original trajectory yd(t). c,k+1 The calculation uses y r,k Instead of y c,k Robot 150 follows the velocity profile Instead of implicit location profile (trajectory) yd(t) k This ensures that when robot 150 stops due to a collision, the error between its actual position and the commanded target position does not accumulate over time. Alternatively, if y r,k If the position y of each new command remains constant due to obstacles, collisions, or blockages, then... c,k+1 Only relative to y r,k small relative displacement Based on a diagonal matrix K specifying the degree of motion compliance, the compliance controller can tolerate the absence of displacement. Without reaching excessive contact force.

[0121] In applying linear compliance law During this period, the measured position y r,k The time series indicates the actual position of robot 150 in each control step. This is achieved by measuring the position y. r,k The robot (or end effector 104) follows the original trajectory y. d By comparing the position that (t) should be in, the processor 304 can calculate the displacement in each control step as d. k =d(t) k )=y(t k )-yd (t k )=y r,k -y d (t k ).

[0122] The above is for Figure 5A and Figure 5B The described process can be followed multiple times, each time starting from the original trajectory y. d Starting from the same initial position in (t), different displacements are applied at different time points. The displacements can be systematic; for example, at a single moment, before contact occurs between the objects, when robot 150 is still in free space, and in a plane perpendicular to the motion of robot 150 at that moment, only one displacement is introduced. This displacement results in a displacement relative to the original trajectory y. d (t) is the corrected trajectory with a constant offset. In some implementations, the displacement can also be random, i.e., by adding small random variations sampled from a probability distribution (e.g., a Gaussian distribution) at each time step.

[0123] As the original trajectory y is traversed multiple times with different displacements d The results of (t) collect data relating the direction and magnitude of the displacement to the forces experienced as a result. When robot 150 is moving in free space without contact, the force relative to the original trajectory y cannot be inferred from the contact force. d The displacement of (t)113 is such that the force experienced is zero. Therefore, τ k The moment when τ = 0 is discarded. For each of the remaining cases, i.e., when τ = 0... k When ≠0, for (τ) i ,d i Training examples of the form ) are added to the database of training examples, where τ i =τ k And d i =d k , where i is the index of the pair in the database.

[0124] To this end, robot 150 records the measurements from multiple pairs of force sensors and their corresponding values ​​relative to the original trajectory y. d The value of the displacement of (t)113. The recorded pair indicates the original trajectory y. d The training data for (t)113 in space. This training data can be used to learn a nonlinear mapping that maps the force sensor measurements to the correction direction of the original trajectory 113.

[0125] Figure 5CA schematic diagram is shown outlining an overview of learning a nonlinear compliant control law, including a nonlinear mapping, based on training data according to some embodiments. In step 504, the fixed object (i.e., object 112) is fixed in its original pose, i.e., pose 115.

[0126] In step 505, processor 305 is configured to generate training data based on the execution of a self-exploration program (as referenced above). Figure 5A and Figure 5B (Detailed description). The training data includes measurements from multiple pairs of force sensors and the corresponding displacement values ​​relative to the original trajectory 113.

[0127] In step 506, the processor is configured to apply a supervised machine learning method to the training data to learn a nonlinear compliant control law that includes a nonlinear mapping of force sensor measurements to a correction direction of the original trajectory 113. Depending on the implementation, the supervised machine learning method may include, for example, Gaussian process regression (GPR) or a deep neural network (DNN). Alternatively, a nonlinear compliant control law may be used to correct the trajectory to complete the assembly operation.

[0128] Figure 6 A schematic diagram is shown illustrating the use of a nonlinear compliant control law to correct the transformed trajectory 600 according to some embodiments. When the position of the stationary object 112 changes, the initial and / or target orientation of the end effector 104 of the robot 150 changes. For example, for a new position 602 of the stationary object 112, there exists a new target orientation of the end effector 104 of the robot 150, such that an object 103 with orientation 601 can be inserted into the stationary object 112. The change in position of the stationary object 112 is determined by the imaging device 106. For example, the imaging device 106 can determine the position 602 of the stationary object 112.

[0129] The new target pose and / or new starting pose of the end-effector 104 are referred to as end-effector pose modification. The end-effector pose modification can be received by the robot 150. Upon receiving the end-effector pose modification, the processor 305 is configured to use the DMP to transform the original trajectory 113 according to the end-effector pose modification to generate a transformed trajectory 600. Additionally, the processor 305 is configured to update the control law (e.g., Equation (1)) according to the transformed trajectory 600 to generate an updated control law.

[0130] Some implementations are based on the understanding that the position 602 of the fixed object 112 determined by the imaging device 106 may be inaccurate. For example, the imaging device 106 may determine the position 602 of the fixed object 112, but the actual position of the fixed object 112 may be located at a distance 603 from the determined position 602. Due to this inaccuracy, the execution of the transformed trajectory 600 may result in a collision between the object 103 and the edge 604 of the fixed object 112. Consequently, the end effector of the robot 150 is subjected to force. In response to the subjected force, the processor 305 uses a nonlinear compliant control law learned for the original trajectory 113 to provide correction to the transformed trajectory 600. For example, the processor 305 is configured to add displacement to the transformed trajectory 600 based on the nonlinear compliant control law. As a result, a new modified trajectory 605 is generated. The new modified trajectory 605 is not generated at the moment of collision; instead, the displacement relative to the transformed trajectory 600 is gradually calculated and added to the transformed trajectory 600. Therefore, a nonlinear compliant control law is used to correct the updated control law. Additionally, the processor 305 is configured to command multiple motors of the robot 150 to control the robotic arm 101 according to the updated control law corrected by the nonlinear compliant control law, in order to complete the assembly operation.

[0131] Figure 7A The calculation of displacement under alignment conditions where the bottom of a movable object contacts the edge of a fixed object, according to some embodiments, is illustrated. The movable object (i.e., object 103) is tilted to the left, where the tilt in the opposite direction is symmetrical. The gripper 104 of the robotic arm 101 is holding object 103 such that the centerline 700 of object 103 is at an angle to the centerline 701 of the fixed object (i.e., the hole of object 112). A force 703 applied at the wrist 102 of the robot 150 is sensed by means of a force sensor 105 mounted on the wrist 102 of the robot 150. Due to the force 703, a torsional torque 704 is sensed at the wrist 102. The torsional torque 704 is the product of the force 703 and the arm 705 of that force. The arm 705 is the distance from the point of contact to the direction of the force 703. Therefore, the sensed torsional torque 704 depends precisely on the position where the bottom of object 103 contacts the edge of the hole of object 112.

[0132] Therefore, the torsional torque 704 depends on the contact configuration. Additionally, another force acting on object 103 is generating an additional torsional torque, but this is not detected by force sensor 105.

[0133] According to the embodiment, another force is caused by the weight of the gripper 104 and the object 103. The additional torsional torque also depends on the contact configuration. Therefore, for example... Figure 7AThe object 103 shown is aligned, and the magnitude of the sensed torsional torque 704 depends on the contact configuration, which in turn depends on the amount of misalignment.

[0134] Figure 7B The displacement of a movable object 103 under alignment conditions, according to some embodiments, with its edge contacting the surface 706 of a fixed object 112, is illustrated. The centerline 708 of the object 103 forms an angle with the centerline 707 of the fixed object (i.e., the hole in the object 112). Here, for Figure 7B The alignment conditions shown indicate that the torsional torque sensed at the wrist does not depend precisely on the position of the surface 706 on the side where the edge of object 103 contacts the hole in object 112. However, regardless of the contact point, the sign of the torsional torque is the same, and it is related to... Figure 7A The signs of the alignments are reversed. Therefore, for such torsional moments, a corrective step of constant magnitude in the positive x-direction can be learned to improve... Figure 7B The alignment of object 103 shown in the figure.

[0135] Figure 7C The displacement of a movable object 103 within a fixed object 112 under alignment conditions, according to some embodiments, is illustrated. The centerline 709 of the object 103 forms an angle with the centerline 710 of the fixed object (i.e., the hole in the object 112). Here, the sensed torsional torque depends on how far the object is (i.e., its z-coordinate) and the misalignment angle. In this case, an inverse mapping between the sensed torsional torque and the rotation to be applied to improve alignment is learned.

[0136] Figure 8 An example is illustrated of a robot 150 configured to perform assembly operations in an industrial facility, according to some embodiments. The industrial facility includes a conveyor belt 800 configured to move one or more objects, such as empty boxes 801, 802, and 803, in direction 804. The robot 150 is configured to grasp objects from a stack 806 via a robotic arm 101 and insert them sequentially into the objects moving on the conveyor belt 800. For example, the robotic arm 101 may grasp object 805 from the stack 806 and insert it into empty box 801. The robot 150 performs this assembly operation according to a trajectory defined by a starting posture and / or a target posture. For this purpose, the robotic arm 101 can be controlled based on a control law.

[0137] Additionally, the robotic arm 101 can grasp an object 807 from the stack 806 to insert the object 807 into the empty box 802. Since the orientation of the empty box 802 differs from that of the empty box 801, the initial and / or target postures change. The processor 305 of the robot 150 is configured to transform the trajectory based on the changed initial and / or target postures without any additional assembly-specific programming, thus generating a transformed trajectory. Furthermore, the control law can be updated based on the transformed trajectory, and the robotic arm 101 can be controlled according to the updated control law. Therefore, the robot 150 performs the assembly operation via the robotic arm controlled based on the updated control law, according to the transformed trajectory, to insert the object 807 into the empty box 802. Since the transformed trajectory is generated without any additional assembly-specific programming, the high costs of deploying and updating the robot device for new assembly operations are eliminated. Therefore, the robot 150 can perform different assembly operations with the variable position of the object to be assembled without operation-specific programming.

[0138] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary embodiments will provide those skilled in the art with a description enabling the implementation of one or more exemplary embodiments. Various changes in the function and arrangement of elements are contemplated without departing from the spirit and scope of the subject matter set forth in the appended claims.

[0139] Specific details are set forth in the following description to provide a thorough understanding of the embodiments. However, those skilled in the art will understand that embodiments can be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary detail. In other cases, well-known processes, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments. Furthermore, the same reference numerals and designations in the various figures refer to the same elements.

[0140] Additionally, various implementations may be described as processes depicted as flowcharts, diagrams, data flow graphs, structural diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many of these operations may be executed in parallel or simultaneously. Furthermore, the order of operations may be rearranged. A process may terminate upon completion of its operations, but the process may have additional steps not discussed or included in the diagram. Moreover, not all operations in any particularly described process appear in all implementations. A process may correspond to a method, function, procedure, routine, subroutine, etc. When a process corresponds to a function, the termination of the function may correspond to the function returning to the calling function or the main function.

[0141] Furthermore, implementations of the disclosed subject matter can be carried out, at least partially, manually or automatically. They can be performed, or at least assisted by, using machines, hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, for manual or automatic implementation. When implemented with software, firmware, middleware, or microcode, program code or code segments for performing the necessary tasks can be stored in a machine-readable medium. The processor can then perform the necessary tasks.

[0142] The various methods or processes outlined herein can be encoded as software that can execute on one or more processors employing any of a variety of operating systems or platforms. Furthermore, such software can be written using any of a variety of suitable programming languages ​​and / or programming or scripting tools, and it can also be compiled into executable machine language code or intermediate code that executes on a framework or virtual machine. Typically, in various implementations, the functionality of program modules can be combined or distributed as desired.

[0143] The embodiments of this disclosure can be implemented as the methods exemplified therein. Actions performed as part of the method can be ordered in any suitable manner. Therefore, embodiments can be constructed in which actions are performed in a different order than those illustrated, and may include performing some actions simultaneously, even those shown as sequential in the illustrated embodiments. Although this disclosure has been described with reference to certain preferred embodiments, it should be understood that various other changes and modifications can be made within the spirit and scope of this disclosure. Therefore, aspects of the appended claims cover all such changes and modifications falling within the true spirit and scope of this disclosure.

Claims

1. A robot comprising: A robotic arm, the robotic arm including an end-effector having multiple degrees of freedom of motion, wherein, during operation of the robot, force sensors are arranged to generate measurements indicating the forces experienced by the end-effector of the robotic arm during the operation; Multiple motors are configured to change the movement of the robotic arm according to commands generated according to a control law; At least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the processor to, upon receiving data of the original trajectory of the movement of the robotic arm instructing it to change the orientation of the end-effector from a starting orientation to a target orientation to perform an assembly operation: Determine the control law used to control the plurality of motors to move the robotic arm according to the original trajectory; A self-exploration program is executed to explore the space of the original trajectory by: controlling the multiple motors according to the control law while introducing different displacements relative to the original trajectory into the state of the robotic arm; and recording the measured values ​​of the pairs of force sensors and the corresponding displacement values ​​relative to the original trajectory when a force is detected on the end tool of the robotic arm, so as to generate training data indicating the space of the original trajectory. The training data is used to learn a nonlinear compliant control law that includes a nonlinear mapping of the force sensor's measurements to a corrected direction of the original trajectory that defines the control law. The instructions executed by the at least one processor further cause the processor to, upon receiving an end-effector posture modification of the original trajectory including at least one or a combination of a new starting posture of the end-effector of the robotic arm and a new target posture of the end-effector, measured with an accuracy lower than that of the assembly operation: The original trajectory is transformed according to the end-effector attitude modification to generate a transformed trajectory; The control law is updated based on the transformed trajectory to generate an updated control law; and The commands instruct the multiple motors to control the robotic arm according to the updated control law, which is corrected by the nonlinear compliant control law learned for the original trajectory.

2. The robot according to claim 1, wherein, During the self-exploration process, the processor is also configured to minimize the forces experienced by the end-effector of the robotic arm based on a linear compliance law with a diagonal matrix.

3. The robot according to claim 2, wherein, The elements of the diagonal matrix are based on the maximum safe force that can be applied in different directions.

4. The robot according to claim 1, wherein, The different displacements introduced into the state of the robotic arm relative to the original trajectory are either symmetrical or random.

5. The robot according to claim 2, wherein, The processor is also configured to calculate the target position of a series of commands at discrete moments for the execution of the linear compliance law, wherein the processor is further configured to calculate the target position of the command at a given moment based on the displacement relative to the original trajectory in the state of the robotic arm at a given moment, the velocity profile corresponding to the original trajectory, and the position of the end effector at that moment.

6. The robot according to claim 1, wherein, The processor is also configured to apply a supervised machine learning method to the training data to learn the nonlinear compliant control law.

7. The robot according to claim 1, wherein, The original trajectory is in the form of a Dynamic Motion Element (DMP), which includes an Ordinary Differential Equation (ODE) that accepts the values ​​of the initial attitude and the target attitude as inputs. The processor is further configured to submit the end-position modification to the DMP of the original trajectory to generate the transformed trajectory.

8. The robot according to claim 1, wherein, The nonlinear mapping is trained to produce the direction of the correction scaled according to a predetermined magnitude of the correction.

9. The robot according to claim 1, wherein, The nonlinear mapping is trained to generate the corrected direction based on the velocity scaling of the end tool.

10. The robot according to claim 1, wherein, The end-effector attitude modification is received from one or more imaging devices.

11. The robot according to claim 10, wherein, The one or more imaging devices include industrial vision cameras with an accuracy on the order of 1 mm, while the accuracy tolerance of the robot's operation is on the order of 0.1 mm.

Citation Information

Patent Citations

  • Active compliance control strategy of Stewart platform

    CN108445764A

  • Robot abrasive belt grinding constant force control method and device based on one-dimensional force sensor

    CN109664295A