Robotic compliant assembly system and method based on integrated transfer learning

By combining a weighted integration multi-source domain strategy with a force controller, the problems of poor migration effect and part damage in robot assembly are solved, achieving efficient and compliant assembly and improving the robot's adaptability and control accuracy in the target domain.

CN117103274BActive Publication Date: 2026-01-09SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311210967.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-19
Publication Date
2026-01-09
Estimated Expiration
2043-09-19

AI Technical Summary

Technical Problem

In the field of compliant assembly in robots, the weight acquisition method in the current technology for integrating multi-source domain learning strategies does not take into account the differences between the source domain and the target domain, resulting in poor transfer effect. Furthermore, the end effector control of the robotic arm does not take compliance into account, which may lead to damage to parts.

Method used

Multiple source domain strategies are integrated using a weighted average method, combined with the force controller output, and a cooperative controller is used to achieve compliant assembly at the end of the robotic arm. The test success rate of the source domain strategies in the target domain is used to set weights, irrelevant strategies are eliminated, and contact force is reduced.

Benefits of technology

It improves the migration efficiency and compliance of robot assembly tasks, reduces part damage, and enhances adaptability and control accuracy in the target domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117103274B_ABST
    Figure CN117103274B_ABST
Patent Text Reader

Abstract

The application discloses a robot compliant assembly system and method based on integrated transfer learning, comprising: at least one source domain controller configured to acquire position information and contact force information of the end of a mechanical arm, utilize a source domain control model, and output an assembly action control strategy of the robot; a strategy integration module configured to perform weighted average integration on the output of each source domain controller to obtain an integrated assembly action control strategy of the end of the robot mechanical arm; a force controller configured to acquire a difference between a desired contact force and an actual contact force, and output the assembly action control strategy of the end of the mechanical arm; and a cooperative controller configured to combine the integrated assembly action control strategy of the robot with the assembly action control strategy output by the force controller to obtain an assembly action control strategy actually executed by the mechanical arm at the next moment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot assembly, and in particular to a robot compliant assembly system and method based on integrated transfer learning. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute the prior art.

[0003] In the field of robot compliant assembly, the assembly objects have characteristics such as various shapes and easy damage and deformation. When various assembly skills learned in known assembly tasks are transferred to new assembly tasks, the problem of how to efficiently transfer often arises.

[0004] In shaft hole assembly, electrical connector assembly, flexible circuit board assembly and other assembly tasks, although the assembly objects have complex and various shapes, they have similar assembly mechanisms, including common processes such as search and insertion feeding. However, how to utilize similar skills in multiple source domain tasks to reduce the damage of useless knowledge in the source domain to the target domain assembly task is a problem to be solved.

[0005] In the prior art, multiple source domain learning strategies are integrated to control the end of the robot arm to complete the assembly task of the target domain. However, when multiple source domain learning strategies are integrated, the weights are usually obtained using the average method or the Bayesian parameter optimization method, and the influence of the difference between the source domain and the target domain on the weight parameters is not considered, so the transfer effect is poor.

[0006] In addition, in the prior art, the control of the end of the robot arm only relies on the output of the integrated learning strategy, and the compliant control of the end of the robot arm is not considered, which may cause the contact force between the end of the robot arm and the assembly object to be too large, causing damage to the parts. SUMMARY

[0007] To solve the above problems, the present application proposes a robot compliant assembly system and method based on integrated transfer learning, which integrates the strategies of multiple source domains in a weighted average manner to form a composite strategy. The integrated strategy can be directly applied to the target task without further learning.

[0008] In some embodiments, the following technical solutions are adopted:

[0009] A robot compliant assembly system based on integrated transfer learning, comprising:

[0010] At least one source domain controller configured to obtain position information and contact force information of the end of the robot arm, utilize a source domain control model, and output an assembly action control strategy of the robot;

[0011] A policy integration module is configured to integrate outputs of the source domain controllers by weighted average to obtain an integrated assembly action control policy of the robot manipulator end;

[0012] A force controller is configured to obtain a difference between the desired contact force and the actual contact force, and output the assembly action control policy of the manipulator end;

[0013] A cooperative controller is configured to combine the integrated assembly action control policy of the robot and the assembly action control policy output by the force controller to obtain an assembly action control policy actually executed by the manipulator at the next moment, so as to control the manipulator to perform assembly work according to the assembly action control policy actually executed.

[0014] As a further scheme, the source domain control model in each source domain controller is learned by a serial mode; the serial mode is to learn the next source domain control model on the basis of the previous source domain control model.

[0015] As a further scheme, the source domain control model in each source domain controller is learned by a parallel mode; the parallel mode is that each source domain control model learns respectively.

[0016] As a further scheme, when the outputs of the source domain controllers are integrated by weighted average, the weight of the output of the i th source domain controller is specifically:

[0017]

[0018] Wherein, n is the total number of source domain controllers, δ i is the test success rate of the current i th source domain strategy in the target field, δ j is the test success rate of the j th source domain strategy in the target field.

[0019] As a further scheme, when the integrated assembly action of the robot is combined with the assembly action output by the force controller, the integrated assembly action of the robot is interpolated to make the integrated assembly action of the robot consistent with the assembly action output by the force controller in frequency, and then the interpolated assembly action and the assembly action output by the force controller are added to obtain the assembly action actually executed by the manipulator at the next moment.

[0020] In some other embodiments, the following technical scheme is adopted:

[0021] A robot compliant assembly method based on integrated transfer learning, comprising:

[0022] The position information and the contact force information of the end of the robot arm are acquired as inputs of a source domain control model in each source domain controller, and the source domain control model in each source domain controller outputs a robot assembly action control strategy respectively;

[0023] The output of each source domain controller is weighted and averaged to obtain an integrated robot assembly action control strategy of the end of the robot arm;

[0024] The difference between the expected contact force and the actual contact force is acquired as an input of a force controller, and an assembly action control strategy of the end of the robot arm output by the force controller is obtained;

[0025] The integrated robot assembly action control strategy and the assembly action control strategy output by the force controller are combined to obtain an assembly action control strategy actually executed by the end of the robot arm at the next moment;

[0026] The robot arm is controlled by using the actually executed assembly action control strategy to realize the robot compliant assembly operation.

[0027] As a further scheme, the source domain control model in each source domain controller is learned by a serial mode; the serial mode is that a next source domain control model is learned on the basis of a previous source domain control model;

[0028] Alternatively,

[0029] The source domain control model in each source domain controller is learned by a parallel mode; the parallel mode is that each source domain control model is learned respectively.

[0030] As a further scheme, when the outputs of each source domain controller are weighted and averaged, the weight of the output of the i th source domain controller is specifically:

[0031]

[0032] Wherein, n is the total number of source domain controllers, δ i is the test success rate of the current i th source domain strategy in the target field, δ j is the test success rate of the j th source domain strategy in the target field.

[0033] As a further scheme, when the integrated robot assembly action and the assembly action output by the force controller are combined, the integrated robot assembly action is interpolated to make the integrated robot assembly action consistent with the frequency of the assembly action output by the force controller, and then the interpolated assembly action and the assembly action output by the force controller are added to obtain an assembly action actually executed by the robot arm at the next moment.

[0034] In other embodiments, the following technical solutions are adopted:

[0035] A robot comprising the robot compliant assembly system based on integrated transfer learning described above; or, a robot compliant assembly method based on integrated transfer learning described above is adopted to perform an assembly operation.

[0036] Compared with the prior art, the robot compliant assembly system based on integrated transfer learning has the following advantages:

[0037] (1) In the present application, each source domain controller outputs a control strategy, the control strategies output by multiple source domain controllers are integrated by weighting to obtain an integrated transfer learning strategy, and the weights of the source strategies in the integrated transfer learning strategy are set by considering the size of the field evaluation (i.e. the success rate of the source domain strategy in the target field) when the multiple control strategies are integrated. Compared with the existing average integration method, the similarity between the source domain and the target domain can be better utilized to extract useful information about the target domain in the source domain, and these useful information can be directly tested without further learning. The useful knowledge in each source domain strategy is fully utilized to accelerate the transfer to the target field and improve the final learning effect of the robot.

[0038] (2) In the present application, the weighted weights of the source domain strategies are obtained by testing in the target field, which can eliminate the strategies of irrelevant fields and enhance the adaptability of the integrated strategy in the target field.

[0039] (3) In the present application, the integrated transfer learning strategy is combined with the control strategy output by the force controller, and through the collaborative control of the two, the contact force in the assembly process can be further reduced to realize compliant assembly.

[0040] Other features and advantages of the additional aspects of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0041] Figure 1 A flowchart of the robot compliant assembly method based on integrated transfer learning in the embodiments of the present application;

[0042] Figure 2 A schematic diagram of the learning process of each source domain control model in the embodiments of the present application;

[0043] Figure 3 A schematic diagram of the integration of multiple source domain control strategies in the embodiments of the present application;

[0044] Figure 4 A schematic diagram of the control process of the force controller in the embodiments of the present application. DETAILED DESCRIPTION

[0045] It should be noted that the following detailed description is exemplary in nature and is intended to provide further description of the application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0046] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments consistent with the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, components, and / or groups thereof, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0047] Embodiment One

[0048] In one or more embodiments, a robot compliant assembly system based on integrated transfer learning is disclosed, specifically comprising:

[0049] At least one source domain controller configured to acquire position information and contact force information of the end of the robot arm, and output an assembly action control strategy of the robot using a source domain control model;

[0050] A strategy integration module configured to perform weighted average integration on the outputs of the respective source domain controllers to obtain an integrated assembly action control strategy of the end of the robot arm;

[0051] A force controller configured to acquire a difference between the desired contact force and the actual contact force, and output an assembly action control strategy of the end of the robot arm;

[0052] A cooperative controller configured to combine the integrated assembly action of the robot with the assembly action output by the force controller to obtain an assembly action control strategy actually executed by the robot arm at the next time; and control the robot arm to perform assembly work according to the assembly action actually executed.

[0053] In this embodiment, the hardware of the system mainly includes a robot arm, a vision sensor, an end six-axis force sensor, and an assembly object, etc.

[0054] In combination Figure 1In the integration migration strategy stage, the system collects position information and contact force information of the end of the robot arm as inputs of the source domain control model in each source domain controller, each source domain control model outputs an assembly action control strategy of the end of the robot arm, the assembly action control strategy output by each source domain control model is integrated by weighted average to obtain an integrated assembly action control strategy of the end of the robot arm, wherein the weight of each source domain control model output is obtained by domain evaluation when the weighted average is calculated, wherein the domain evaluation is to obtain the difference between each source domain and the target domain, in the embodiment, the success rate δ of the source domain strategy in the target domain is used to represent, the higher the success rate δ, the smaller the difference, the lower the success rate δ, the greater the difference.

[0055] The input of the force controller is the difference between the expected contact force and the actual contact force, and the output is the assembly action control strategy of the end of the robot arm, the control strategy output by the force controller is combined with the integrated control strategy as the control strategy of the actual execution action of the robot arm at the next moment, and the output of the force control is combined to make the contact force between the end of the robot arm and the assembly object smaller, so that the parts are not easily damaged, and compliant control is realized.

[0056] As a specific implementation, the source domain control model in each source domain controller is a model after learning and training, Figure 2 The learning and training process of each source domain control model is given, which specifically includes the following processes:

[0057] 1) initialize each network, including Q network V network and policy network π θ ,π θ' ;

[0058] Q represents the QCritic network, and the output is q(s, a), which represents the value estimation of the action-state pair; V represents the VCritic network, and the output is v(s), which represents the state value estimation; , respectively, represent the network parameters of the Q network and the V network, and π θ ,π θ' represent the new and old policies, respectively.

[0059] 2) the robot arm executes action a t in the current state s t , obtains the reward value r t at the current moment and the state s t+1 at the next moment, stores (s t , a t , r t , s t+1 ) in the experience pool buffer, at this time, it is the 0th round of training, and n = 0.

[0060] 3) Collect batch data from buffer (experience pool for storing data in interactive process) t t t t+1 where batch is a constant, and the embodiment takes 64, i.e. batch = 64, let n = n + 1;

[0061] 4) Calculate TD error (time difference error):

[0062]

[0063] where y q is the target Q value, r is the reward, γ is the discount factor, and α is the entropy coefficient.

[0064] Calculate the loss of Q network:

[0065] Update the parameters of Q network using gradient descent:

[0066]

[0067]

[0068] 5) Update V network:

[0069]

[0070] 6) Calculate the loss of policy network:

[0071]

[0072] where E represents expectation, i.e. average, and s ~ buffer represents that s is obtained from buffer. The above formula means that for each state s obtained from buffer, calculate once Finally, take the average.

[0073] Update the parameters of policy network:

[0074]

[0075] where θ represents the parameters of policy network, ξ represents the parameters of V ξ network, β is the coefficient of gradient update, and a ξ is the action obtained in the state s

[0076] 7) Update ψ1,

[0077] ψ1←ρψ1+(1-ρ)ψ​​​

[0078] wherein p represents the weight of the old parameter in the updating process.

[0079] 8) If n=N, end the training, otherwise go back to step 3, N represents the total number of training rounds.

[0080] In the embodiment, each source domain control model can be obtained through a serial learning mode or a parallel learning mode. The serial learning mode is to sequentially learn a next source domain control model based on a previous source domain control model in the same task field to obtain multiple learning strategies by using the dependency between the models. Specifically, a base model is obtained in a certain specific field through learning The base model is stored and continues to learn to obtain a base model and so on until the mth base model is obtained The parallel learning mode fuses learning strategies of multiple fields, and each strategy is weighted and averaged to reduce model error. Each strategy in the parallel integration learns independently to learn n strategies

[0081] Each strategy in the parallel learning mode learns in different fields and does not affect each other. Each strategy in the serial learning mode learns in the same field, and each strategy is learned based on the previous strategy.

[0082] In combination Figure 3 , each source domain control model obtains its own control strategy through the serial or parallel learning mode, and then integrates each control strategy through the weighted average mode.

[0083] Suppose there are multiple source domains {D1, D2,..., D n}, and the target domain is {D t}. The test success rates of the models obtained from the source domains on the target domain are {δ1, δ2,..., δ n}.

[0084] The integrated migration strategy model is obtained by weighting multiple models:

[0085]

[0086] The weight of each model in the source field in the final integrated strategy is:

[0087]

[0088] wherein n is the total number of source domain controllers, δ i is the test success rate of the current ith source domain strategy in the target field, and δ jThe jth source domain strategy in the target field test success rate.

[0089] The embodiment calculates the weight through the test success rate parameter of the source domain strategy in the target field, extracts useful information of the source domain for the target field, directly tests by using the useful information, does not need to learn again, and improves the speed of integrated migration.

[0090] In the embodiment, the force controller adopts a PI controller design method, which adopts a position control and force control respectively control mode. Figure 4 The position control is PD control, the input is the expected position, and the output is the six-dimensional motion of the robot end.

[0091]

[0092] The input is the difference between the expected contact force F and the actual contact force F ext , and the output is the six-dimensional motion of the robot end.

[0093]

[0094] The output frequency of the force controller is 500Hz, the frequency of the integrated migration strategy is 20Hz, the output motion of the integrated migration strategy is interpolated to convert the frequency to 500Hz, and the motion of the two strategies is added to obtain the actual motion to be executed by the robot at the next time:

[0095]

[0096] Wherein represents the actual motion to be executed by the robot at the next step, represents the output motion of the integrated migration strategy, represents the output motion of the force controller.

[0097] The integrated strategy is used to execute various assembly tasks in the new target field without further training.

[0098] Embodiment two

[0099] In one or more embodiments, a robot compliant assembly method based on integrated migration learning is disclosed, specifically including the following processes:

[0100] (1) Obtain the position information and contact force information of the robot end as the input of the source domain control model in each source domain controller, and the source domain control model in each source domain controller outputs the assembly motion control strategy of the robot;

[0101] (2) the outputs of each source domain controller are integrated by weighted average to obtain an integrated assembly action control strategy of the robot manipulator end;

[0102] (3) the difference between the expected contact force and the actual contact force is obtained as the input of the force controller, and an assembly action control strategy of the robot manipulator end output by the force controller is obtained;

[0103] (4) the integrated assembly action control strategy of the robot and the assembly action control strategy output by the force controller are combined to obtain an assembly action control strategy actually executed by the robot manipulator end at the next moment;

[0104] (5) the robot is controlled by using the actually executed assembly action control strategy to realize the compliant assembly operation of the robot.

[0105] The source domain control model in each source domain controller is learned by a serial mode; the serial mode is that the next source domain control model is learned on the basis of the previous source domain control model.

[0106] Alternatively,

[0107] The source domain control model in each source domain controller is learned by a parallel mode; the parallel mode is that each source domain control model is learned respectively.

[0108] In the embodiment, the weight of the output of the i th source domain controller when the outputs of each source domain controller are integrated by weighted average is specifically:

[0109]

[0110] wherein n is the total number of source domain controllers, δ i is the test success rate of the current i th source domain strategy in the target field, δ j is the test success rate of the j th source domain strategy in the target field.

[0111] In the embodiment, when the integrated assembly action of the robot and the assembly action output by the force controller are combined, the integrated assembly action of the robot is interpolated to make the integrated assembly action of the robot and the assembly action output by the force controller consistent in frequency, and then the interpolated assembly action and the assembly action output by the force controller are added to obtain the assembly action actually executed by the robot manipulator at the next moment.

[0112] The specific implementation mode of the above process is the same as that in embodiment one, and will not be described in detail here.

[0113] Embodiment three

[0114] In one or more embodiments, a robot is disclosed, comprising the robot compliant assembly system based on integrated transfer learning described in embodiment one;

[0115] Alternatively, the assembly work is performed using the robot compliant assembly method based on integrated transfer learning described in embodiment two.

[0116] The above describes the specific embodiments of the present application in conjunction with the accompanying drawings, but is not a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications or variations made by those skilled in the art on the basis of the technical solutions of the present application without creative labor are still within the scope of protection of the present application.

Claims

1. A robot compliant assembly system based on integrated transfer learning, characterized by, The method comprises the following steps: at least one source domain controller is configured to obtain position information and contact force information of the end of the robot arm, use a source domain control model to output an assembly action control strategy of the robot, and integrate the outputs of the source domain controllers by weighted average to obtain an integrated assembly action control strategy of the end of the robot arm; the weight of the output of the i th source domain controller is specifically: a force controller is configured to obtain a difference between the expected contact force and the actual contact force, and output an assembly action control strategy of the end of the robot arm; Wherein, n is the total number of source domain controllers, is the test success rate of the current i-th source domain strategy in the target field, is the test success rate of the j-th source domain strategy in the target field. a cooperative controller is configured to combine the integrated assembly action control strategy of the robot and the assembly action control strategy output by the force controller to obtain an assembly action control strategy actually executed by the robot arm at the next moment, so as to control the robot arm to perform assembly work according to the assembly action control strategy actually executed. The source domain control model in each source domain controller is learned by a serial mode; the serial mode is to learn the next source domain control model based on the previous source domain control model.

2. The robot compliant assembly system based on integrated transfer learning of claim 1, wherein, The source domain control model in each source domain controller is learned by a parallel mode; the parallel mode is that each source domain control model is learned respectively.

3. The robot compliant assembly system based on integrated transfer learning of claim 1, wherein, When the integrated assembly action of the robot and the assembly action output by the force controller are combined, the integrated assembly action of the robot is interpolated to make the integrated assembly action of the robot and the assembly action output by the force controller consistent in frequency, then the interpolated assembly action and the assembly action output by the force controller are added to obtain the assembly action actually executed by the robot arm at the next moment.

4. The robot compliant assembly system based on integrated transfer learning of claim 1, wherein, The method comprises the following steps:

5. A robot compliant assembly method based on integrated transfer learning, characterized in that, obtain position information and contact force information of the end of the robot arm as input of the source domain control model in each source domain controller, and the source domain control model in each source domain controller outputs an assembly action control strategy of the robot; integrate the outputs of the source domain controllers by weighted average to obtain an integrated assembly action control strategy of the end of the robot arm; the weight of the output of the i th source domain controller is specifically: obtain a difference between the expected contact force and the actual contact force as input of the force controller to obtain an assembly action control strategy of the end of the robot arm output by the force controller; Wherein, n is the total number of source domain controllers, is the test success rate of the current i-th source domain strategy in the target field, is the test success rate of the j-th source domain strategy in the target field. combine the integrated assembly action control strategy of the robot and the assembly action control strategy output by the force controller to obtain an assembly action control strategy actually executed by the end of the robot arm at the next moment; control the robot arm by using the assembly action control strategy actually executed to realize compliant assembly work of the robot. The source domain control model in each source domain controller is learned by a serial mode; the serial mode is to learn the next source domain control model based on the previous source domain control model.

6. The robot compliant assembly method based on integrated transfer learning of claim 5, wherein, Alternatively, The source domain control model in each source domain controller is learned by a parallel mode; the parallel mode is that each source domain control model is learned respectively. ​ 7. The robot compliant assembly method based on integrated transfer learning of claim 5, wherein, When the integrated robot assembly action is combined with the assembly action output by the force controller, the integrated robot assembly action is interpolated so that the integrated robot assembly action is consistent with the frequency of the assembly action output by the force controller, and then the interpolated assembly action and the assembly action output by the force controller are added to obtain the assembly action actually executed by the robot arm at the next time.

8. A robot, characterized in that The robot compliant assembly system based on integrated transfer learning according to any one of claims 1-4; or, the robot compliant assembly method based on integrated transfer learning according to any one of claims 5-7 is used to perform an assembly operation.

Citation Information

Patent Citations

  • Robot assembly method and system based on feature adaptive migration reinforcement learning

    CN115481688A

  • Virtual-real transfer learning method and device for robot operation skills and storage medium

    CN115533905A