Multi-robot imitation learning assembly method based on dual-loop force-position coupling collaborative control

Through the multi-robot assembly method of dual-loop force-position coupling collaborative control and imitation learning, the contact force problem in the assembly of large workpieces is solved, high-precision assembly is achieved and the robustness of assembly is improved.

CN116100550BActive Publication Date: 2025-09-16HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310111409.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2025-09-16
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

Existing industrial robots have large contact forces and are prone to damaging workpieces or tools when performing precision assembly of large workpieces. Traditional methods also take a long time to deploy, are complex to program, and have poor adaptability to the environment, making it difficult to achieve high-precision assembly.

Method used

A multi-robot imitation learning assembly method based on dual-loop force-position coupling collaborative control is adopted. The workpiece is clamped by multiple robotic arms and an imitation learning model is established to adjust the position of the workpiece to achieve high-precision assembly. Multi-robot assembly is performed by combining dual-loop force-position coupling collaborative control and imitation learning.

Benefits of technology

It achieves high-precision assembly of large workpieces in a tightly coupled state, prevents damage caused by synchronization errors of the robotic arms, and is robust to uncertainties in the collaborative assembly of multiple robotic arms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116100550B_ABST
    Figure CN116100550B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-robot imitation learning assembly method based on dual-loop force-position coupling collaborative control, comprising: obtaining the position information of workpiece 2; performing dual-loop force-position coupling collaborative control on multiple robotic arms, wherein the multiple robotic arms maintain a predetermined contact internal force when clamping workpiece 1, and carry workpiece 1 to the top of workpiece 2, so that workpiece 1 and workpiece 2 are in contact and reach a predetermined contact external force; continuously adjusting the position of workpiece 1 through an imitation learning model until a sudden decrease in the contact external force is detected, and workpiece 1 and workpiece 2 are aligned in position, and workpiece 1 and workpiece 2 are assembled to a specified installation position by clamping the multiple robotic arms; the multiple robotic arms collaboratively clamp workpiece 1 back to the initial position to complete an imitation learning round; and continuously performing a set number of imitation learning rounds to complete training. The present invention combines dual-loop force-position coupling collaborative control and imitation learning to perform multi-robot assembly, and is suitable for high-precision assembly of large workpieces under a tightly coupled state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a multi-robot imitation learning assembly method based on dual-loop force-position coupling collaborative control. Background Art

[0002] Manufacturing technology is at the core of economic competition, and manufacturing processes are becoming increasingly automated. Industrial robots are widely used in the manufacturing sector to improve production efficiency and product quality. Currently, industrial robots are primarily used for unconstrained tasks such as handling and painting, where the movement of the robot's end-tool is unrestricted. For constrained tasks, such as assembly, which involve contact with the workpiece, general precision assembly tasks require certain tolerances and precision, and small assembly clearances can easily cause assembly jams. Industrial robots based on position or velocity control have high contact stiffness and are prone to generating large contact forces during contact, which can easily damage the workpiece or tool, making them virtually incapable of completing precision assembly tasks.

[0003] Especially when assembling large workpieces, due to their considerable weight, multiple robotic arms often need to coordinate. Traditional industrial robots typically rely on teach-point programming or offline programming to complete constrained tasks. These robots suffer from numerous drawbacks, including long deployment times, complex algorithms and programming, high operator requirements, limited use in structured environments, and poor adaptability. Therefore, there is an urgent need for a high-precision assembly method suitable for large workpieces to achieve automated assembly. Summary of the Invention

[0004] The main purpose of the present invention is to provide a multi-robot imitation learning assembly method based on dual-loop force-position coupling collaborative control to achieve high-precision assembly of large workpieces.

[0005] To achieve the above main objectives, the present invention provides a multi-robot imitation learning assembly method based on dual-loop force-position coupling collaborative control, which is used for high-precision assembly of large workpieces under tight coupling, comprising the following steps:

[0006] Step 1: Get the location information of workpiece 2;

[0007] Step 2: Dual-loop force-position coupling collaborative control is performed on the multiple manipulators in the multi-robot system. The multiple manipulators maintain a predetermined internal contact force when gripping workpiece 1 and move workpiece 1 to above workpiece 2. Workpieces 1 and 2 then come into contact and achieve a predetermined external contact force.

[0008] Step 3: Establish an imitation learning model and continuously adjust the position of workpiece 1 through the imitation learning model until a sudden decrease in the contact force is detected, at which point workpiece 1 and workpiece 2 are aligned. Use a multi-manipulator arm to clamp workpiece 1 and workpiece 2 and assemble them to the specified installation position.

[0009] Step 4: Multiple robotic arms collaborate to grip the workpiece and return it to its initial position, completing one round of imitation learning. Repeat the set number of imitation learning rounds to complete the training.

[0010] According to another specific embodiment of the present invention, the imitation learning model in step 3 is as follows:

[0011] The imitation learning model includes an agent, an environment, a state space, an action space, and a reward strategy. The multi-robot system acts as the agent. Based on the contact state between workpieces 1 and 2, the multi-robot system selects actions within the action space and obtains a certain reward.

[0012] The state space S is expressed as:

[0013] S=[W Ecx W Ecy W Ecz T Ex T Ey C]

[0014] The action space A is expressed as:

[0015]

[0016] Among them, W Ecx 、W Ecy 、W Ecz represents the external force along the x, y, and z axes; T Ex 、T Ey represents the position of workpiece 1 in the xoy plane, and ΔT is the step length of workpiece 1’s movement;

[0017] The reward obtained each time is expressed as:

[0018] R t =r t +γr t+1 +…+γ n-t r n =r t +R t+1

[0019] Among them, r, γ (0<γ<1) and t represent the current reward and the number of learning steps respectively;

[0020] The strategy for obtaining rewards is expressed as:

[0021]

[0022] Among them, t max is the maximum number of learning steps; the reward values ​​corresponding to different actions in a certain state will be recorded in the Q table, and the Q table is updated as follows:

[0023] Q(s,a)←Q(s,a)+δ(r+γmax a′ Q(s′,a′)-Q(s,a))

[0024] Among them, s′ and a′ are the next state and action, δ represents the learning rate; the strategy for selecting the action is:

[0025]

[0026] Where ε∈(0,1) is the greedy coefficient and β∈(0,1) is a random number.

[0027] According to another specific embodiment of the present invention, in step 4, when the number of learning steps in the imitation learning round tends to be stable, the training is completed.

[0028] According to another specific embodiment of the present invention, the method for performing dual-loop force-position coupling coordinated control on multiple robotic arms in a multi-robot system in step 2 is as follows:

[0029] 1) Establish kinematic constraints

[0030]

[0031] in, Represent the homogeneous transformation matrices of the manipulator base relative to the world coordinate system, the object coordinate system relative to the world coordinate system, the manipulator end coordinate system relative to the object coordinate system, and the manipulator end relative to the base coordinate system, respectively. i represents the i-th manipulator;

[0032] The motion of workpiece 1 can be further mapped to the joint space of the robot arm:

[0033]

[0034] Among them, q ri 、f IK (.) represent the joint vector of the i-th robot arm and the inverse kinematics of the robot arm respectively;

[0035] 2) Establish force constraints

[0036] The dynamic equation of workpiece 1 is established through the Newton-Euler equation:

[0037]

[0038] Where m, I, and g represent the mass, inertia matrix, and gravitational acceleration of workpiece 1, respectively; represents the force and torque applied by the robotic arm; f e , τ e Indicates the force and torque applied by the environment; v and ω represent the linear velocity and angular velocity of the workpiece;

[0039] W r It can be obtained by capturing the matrix G:

[0040] W r =GW C

[0041] in, is the force applied by the robot arm on the workpiece;

[0042]

[0043] but:

[0044] in, is the generalized inverse matrix, w s Indicates internal force,

[0045] The force exerted by the robot arm on the workpiece is decomposed into the internal force W I and external force W E :

[0046]

[0047] In summary, the force-position coupling equation of the multi-manipulator is as follows:

[0048]

[0049]

[0050]

[0051]

[0052]

[0053]

[0054] Among them, W Ec 、W Ed 、T Ec 、T Ed 、M E 、B E , K E , α E and β E They represent the contact force between the workpiece and the environment, the expected external force, the actual trajectory, the expected trajectory, the mass matrix, the damping parameter matrix, the stiffness parameter, the sampling period and the update rate respectively. The letter “E” represents the outer loop; W Ic 、W Id 、T Ic 、T Id 、MI 、B I , K I , α I and β I They represent the contact force between workpiece 1 and the robot arm, the expected external force, the actual trajectory, the expected trajectory, the mass matrix, the damping parameter matrix, the stiffness parameter, the sampling period and the update rate respectively. The letter “I” represents the inner loop.

[0055] According to another specific embodiment of the present invention, a force sensor is provided at the end of each robotic arm.

[0056] According to another specific embodiment of the present invention, multiple robotic arms in the multi-robot system are distributed in a circular array.

[0057] According to another specific embodiment of the present invention, in step 1, the position information of the workpiece 2 is obtained by using an RGB-D camera.

[0058] The present invention has the following beneficial effects:

[0059] The present invention can perform high-precision assembly of large workpieces, combine dual-loop force-position coupling collaborative control and imitation learning to perform multi-robot assembly, and is suitable for high-precision assembly of large workpieces under a tightly coupled state; when multiple robotic arms in a multi-robot system collaboratively clamp workpieces, it can effectively prevent the huge internal force caused by the synchronization error of the robotic arms from damaging the workpiece, and has strong robustness against the interference of uncertain factors during the collaborative assembly of multiple robotic arms.

[0060] In order to more clearly illustrate the purpose, technical solutions and advantages of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a schematic diagram of the multi-robot imitation learning assembly method of the present invention;

[0062] Figure 2 This is a simulation diagram of multiple rounds of imitation learning training in the present invention;

[0063] Figure 3 This is an experimental diagram of multiple simulation learning rounds of training in the present invention;

[0064] Figure 4 This is a framework diagram of the dual-loop force-position coupling collaborative control of multiple robotic arms in the present invention. DETAILED DESCRIPTION

[0065] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.

[0066] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0067] An embodiment of the present invention provides a multi-robot imitation learning assembly method based on dual-loop force-position coupling collaborative control, which is used to perform high-precision assembly of large workpieces in a tightly coupled state, including the following steps:

[0068] Step 1: Obtain the position information of workpiece 2 through the RGB-D camera;

[0069] The position information here is rough and approximate to a certain extent. It is not the precise position information that can be directly assembled. It only provides range position information for assembly.

[0070] Step 2: Perform dual-loop force-position coupling collaborative control on the multiple manipulators in the multi-robot system to prevent the huge internal forces caused by the manipulator synchronization errors from damaging the workpiece and solve the dual coupling problem of force and position when the multiple manipulators work together to move objects. The multiple manipulators maintain a predetermined contact internal force when clamping workpiece 1 and move workpiece 1 to above workpiece 2. Then, workpiece 1 and workpiece 2 contact and reach a predetermined contact external force.

[0071] Specifically, such as Figure 1 As shown, the multi-robot system in this embodiment of the present invention includes three robotic arms, each arranged in a circular array. In other embodiments, more robotic arms may be provided, which will not be discussed further here. Each robotic arm is equipped with a force sensor at its end to detect applied force.

[0072] Step 3: Establish an imitation learning model and continuously adjust the position of workpiece 1 through the imitation learning model until a sudden decrease in the contact force is detected, at which point workpiece 1 and workpiece 2 are aligned. Use a multi-manipulator arm to clamp workpiece 1 and workpiece 2 and assemble them to the specified installation position.

[0073] The imitation learning model is as follows:

[0074] The imitation learning model includes an agent, an environment, a state space, an action space, and a reward strategy. The multi-robot system acts as the agent. Based on the contact state between workpieces 1 and 2, the multi-robot system selects actions within the action space and obtains a certain reward.

[0075] The state space S is expressed as:

[0076] S=[W Ecx W Ecy W Ecz T Ex T Ey C]

[0077] The action space A is expressed as:

[0078]

[0079] Among them, W Ecx 、W Ecy 、W Ecz represents the external force along the x, y, and z axes; T Ex 、T Ey represents the position of workpiece 1 in the xoy plane, and ΔT is the step length of workpiece 1’s movement;

[0080] The reward obtained each time is expressed as:

[0081] R t =r t +γr t+1 +…+γ n-t r n =r t +R t+1

[0082] Among them, r, γ (0<γ<1) and t represent the current reward, learning rate and number of learning steps respectively;

[0083] The strategy for obtaining rewards is expressed as:

[0084]

[0085] Among them, t max is the maximum number of learning steps; the reward values ​​corresponding to different actions in a certain state will be recorded in the Q table, and the Q table is updated as follows:

[0086] Q(s,a)←Q(s,a)+δ(r+γmax a′ Q(s′,a′)-Q(s,a))

[0087] Among them, s′ and a′ are the next state and action, δ represents the learning rate; the strategy for selecting the action is:

[0088]

[0089] Where ε∈(0,1) is the greedy coefficient and β∈(0,1) is a random number.

[0090] The imitation learning process in the embodiment of the present invention is as follows:

[0091]

[0092]

[0093] Step 4: Once the multiple robotic arms coordinate to grip the workpiece and return it to its initial position, one imitation learning round is completed. This is repeated for the set number of imitation learning rounds to complete the training. Training is complete when the number of learning steps in each imitation learning round stabilizes.

[0094] like Figure 2 As shown in the figure, after simulation, after 10 imitation learning rounds, the number of learning steps in subsequent imitation learning rounds tends to stabilize at less than 20 steps, and after 19 imitation learning rounds, the number of learning steps in subsequent imitation learning rounds tends to stabilize at 12 steps; the optimal assembly process is determined according to the level of assembly accuracy requirements.

[0095] like Figure 3 As shown in the figure, after experiments, after 20 imitation learning rounds, the number of learning steps in subsequent imitation learning rounds tends to stabilize at less than 100 steps, and after 19 imitation learning rounds, the number of learning steps in subsequent imitation learning rounds tends to stabilize in the range of 20-25 steps; the optimal assembly process is determined according to the level of assembly accuracy requirements.

[0096] like Figure 4 As shown, the method for performing dual-loop force-position coupling collaborative control on multiple robotic arms in a multi-robot system in step 2 of the embodiment of the present invention is as follows:

[0097] 1) Establish kinematic constraints

[0098]

[0099] in, Represent the homogeneous transformation matrices of the manipulator base relative to the world coordinate system, the object coordinate system relative to the world coordinate system, the manipulator end coordinate system relative to the object coordinate system, and the manipulator end relative to the base coordinate system, respectively. i represents the i-th manipulator;

[0100] The motion of workpiece 1 can be further mapped to the joint space of the robot arm:

[0101]

[0102] Among them, q ri 、f IK (.) represent the joint vector of the i-th robot arm and the inverse kinematics of the robot arm respectively;

[0103] 2) Establish force constraints

[0104] In the process of multi-manipulator collaborative handling, the load distribution problem also needs to be solved; the dynamic equation of workpiece 1 is established through the Newton-Euler equation:

[0105]

[0106] Where m, I, and g represent the mass, inertia matrix, and gravitational acceleration of workpiece 1, respectively; represents the force and torque applied by the robotic arm; f e , τ e Indicates the force and torque applied by the environment; v and ω represent the linear velocity and angular velocity of the workpiece;

[0107] W r It can be obtained by capturing the matrix G:

[0108] W r =GW C

[0109] in, is the force applied by the robot arm on the workpiece;

[0110]

[0111] but:

[0112] in, is the generalized inverse matrix, w s Indicates internal force,

[0113] The force exerted by the robot arm on the workpiece is decomposed into the internal force W I and external force W E :

[0114]

[0115] In summary, the force-position coupling equation of the multi-manipulator is as follows:

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122] Among them, W Ec 、W Ed 、T Ec 、T Ed 、M E 、B E , K E , α E and β E They represent the contact force between the workpiece and the environment, the expected external force, the actual trajectory, the expected trajectory, the mass matrix, the damping parameter matrix, the stiffness parameter, the sampling period and the update rate respectively. The letter “E” represents the outer loop; W Ic 、W Id 、T Ic 、T Id 、M I 、B I , K I , α I and β I They represent the contact force between workpiece 1 and the robot arm, the expected external force, the actual trajectory, the expected trajectory, the mass matrix, the damping parameter matrix, the stiffness parameter, the sampling period and the update rate respectively. The letter “I” represents the inner loop.

[0123] Although the present invention has been disclosed above with reference to specific embodiments, these embodiments are not intended to limit the scope of the present invention. Any person skilled in the art may make changes or modifications without departing from the scope of the present invention. Therefore, any equivalent changes or modifications made in accordance with the present invention are intended to be covered by the scope of protection of the present invention.

Claims

1. A multi-robot imitation learning assembly method based on dual-loop force-position coupling collaborative control is used for high-precision assembly of large workpieces under tight coupling conditions, characterized by: The following steps are involved: Step 1: Get the location information of workpiece 2; Step 2: Dual-loop force-position coupling collaborative control is performed on the multiple manipulators in the multi-robot system. The multiple manipulators maintain a predetermined internal contact force when gripping workpiece 1 and move workpiece 1 to above workpiece 2. Workpieces 1 and 2 then come into contact and achieve a predetermined external contact force. Step 3: Establish an imitation learning model and continuously adjust the position of workpiece 1 through the imitation learning model until a sudden decrease in the contact force is detected, at which point workpiece 1 and workpiece 2 are aligned. Use a multi-manipulator arm to clamp workpiece 1 and workpiece 2 and assemble them to the specified installation position. Among them, the imitation learning model in step 3 is as follows: The imitation learning model includes an agent, an environment, a state space, an action space, and a reward strategy. The multi-robot system acts as the agent. Based on the contact state between workpieces 1 and 2, the multi-robot system selects actions within the action space and obtains a certain reward. The state space S is expressed as: S=[W Ecx W Ecy W Ecz T Ex T Ey C] The action space A is expressed as: Among them, W Ecx 、W Ecy 、W Ecz represents the external force along the x, y, and z axes; T Ex 、T Ey represents the position of workpiece 1 in the xoy plane, and ΔT is the step length of workpiece 1’s movement; The reward obtained each time is expressed as: R t =r t +γr t+1 +L+γ n-t r n =r t +R t+1 Among them, r, γ (0 < γ < 1) and t represent the current reward, discount coefficient and number of learning steps respectively; The strategy for obtaining rewards is expressed as: Among them, t max is the maximum number of learning steps; the reward values ​​corresponding to different actions in a certain state will be recorded in the Q table, and the Q table is updated as follows: Q(s,a)←Q(s,a)+δ(r+γmax a′ Q(s′,a′)-Q(s,a)) where s′ and a′ are the next state and action, and δ represents the learning rate. The strategy for selecting the action is: Where ε∈(0,1) is the greedy coefficient and β∈(0,1) is a random number; Step 4: Multiple robotic arms collaborate to grip the workpiece and return it to its initial position, completing one round of imitation learning. Repeat the set number of imitation learning rounds to complete the training.

2. The multi-robot imitation learning assembly method based on dual-loop force-position coupling coordinated control according to claim 1, characterized in that: In step 4, when the number of learning steps in the imitation learning round tends to be stable, the training is completed.

3. The multi-robot imitation learning assembly method based on dual-loop force-position coupling coordinated control according to claim 1, characterized in that: A force sensor is provided at the end of each robotic arm.

4. The multi-robot imitation learning assembly method based on dual-loop force-position coupling coordinated control according to claim 1, characterized in that: The multiple robotic arms in the multi-robot system are distributed in a circular array.

5. The multi-robot imitation learning assembly method based on dual-loop force-position coupling coordinated control according to claim 1, characterized in that: In step 1, the position information of workpiece 2 is obtained through the RGB-D camera.

Citation Information

Patent Citations

  • Multi-mechanical-arm force-position coupling cooperative control method and welding method

    CN114473323A

  • Robot shaft hole assembling method based on deep reinforcement learning and admittance control

    CN115674204A