Industrial robot shaft hole assembly system and active compliance control method

By dynamically optimizing the parameters of the impedance controller and using the network model of joint reinforcement learning, the flexible control problem of industrial robots in the shaft hole assembly and insertion stage is solved, achieving high-precision assembly effect and operation stability.

CN115922746BActive Publication Date: 2025-05-02QINGDAO UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211594626.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2025-05-02
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

Existing industrial robots are difficult to achieve high-precision and compliant control during the shaft hole assembly and insertion stage, resulting in failure of assembly tasks or damage to the robot, and fixed impedance parameters cannot ensure stable interaction in the assembly process.

Method used

The parameters of the impedance controller are dynamically optimized by using a segmented strategy, and the damping parameters are adjusted in real time through the industrial robot shaft hole assembly network model based on joint reinforcement learning to achieve flexible control.

Benefits of technology

The flexible control of the shaft hole assembly and insertion stage is achieved, which improves assembly quality and speed, reduces the requirements for precise modeling of the assembly process, and enhances the operation stability of the robot in non-structural environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115922746B_ABST
    Figure CN115922746B_ABST
Patent Text Reader

Abstract

The present invention relates to an industrial robot shaft hole assembly system and an active compliance control method in the field of intelligent manufacturing technology. The system comprises a computer, an industrial robot, a six-dimensional force / torque sensor, shaft parts, and a sensor data acquisition box. An industrial robot shaft hole assembly network model based on joint reinforcement learning is run on the computer, and impedance parameters are output to control the dynamic change relationship between the robot assembly force and the robot terminal motion. A segmented strategy is used to dynamically optimize the parameters of the impedance controller to achieve compliance control in the shaft hole assembly insertion stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an industrial robot shaft hole assembly system and an active compliance control method, belonging to the technical field of intelligent manufacturing. Background Art

[0002] The advancement of robotics technology has promoted the development of industrial automation. Today, industrial robots have been widely used in non-contact tasks such as handling, palletizing, spraying, welding, etc., greatly improving work efficiency and intelligence. With the diversification of application scenarios and the improvement of automation levels, industrial robots need to perform operating tasks in more and more contact-rich non-structured environments. Automated assembly is a typical contact-rich robot operation task. As the final process in modern manufacturing, assembly is a key stage in production and manufacturing, and assembly performance will directly affect product quality.

[0003] Assembly tasks with rich contacts can usually be abstracted as shaft-hole assembly tasks. Shaft-hole assembly is usually divided into two stages: a search and approach stage and a compliant insertion stage. The present invention is mainly aimed at the compliant control of the shaft-hole assembly insertion stage. General industrial robots are usually rigid position control robots. Even a very small posture error during the assembly process will generate a large contact force. Excessive contact force will cause the assembly task to fail or even damage the robot. It is usually difficult to complete high-precision shaft-hole assembly tasks using only position-based control methods. Impedance control is an effective force control method. The relationship between force and displacement is processed to make the dynamic behavior of the robot appear as a system with mass, damping and stiffness, thereby achieving stable interaction with the environment. However, the shaft-hole assembly process is a nonlinear dynamic process, and fixed impedance parameters cannot guarantee stable interaction during the assembly process.

[0004] In response to the above problems, this case provides an industrial robot shaft-hole assembly system and an active compliance control method, which adopts a segmented strategy to dynamically optimize the parameters of the impedance controller to achieve compliance control during the shaft-hole assembly insertion stage. Summary of the invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides an industrial robot shaft hole assembly system and an active compliance control method.

[0006] The technical solution of the present invention is as follows:

[0007] An industrial robot shaft hole assembly system, including a computer, an industrial robot, a six-dimensional force / torque sensor, shaft parts, and a sensor data acquisition box;

[0008] The six-dimensional force / torque sensor is installed at the end of the industrial robot and communicates with the sensor data acquisition box through a data line;

[0009] The sensor data acquisition box is connected to the computer to transmit the acquired six-dimensional force / torque to the computer;

[0010] The six-dimensional force / torque sensor, industrial robot, and assembly shaft parts have their own coordinate systems;

[0011] The computer and the industrial robot perform two-way communication, the computer controls the industrial robot to assemble shaft parts, and the six-dimensional force / torque sensor records and collects force / torque information, and the sensor data acquisition box processes the force / torque information and transmits it to the computer;

[0012] The computer processes the force / torque information and the robot posture information through a python interface; the industrial robot shaft-hole assembly network model based on joint reinforcement learning is run on the computer, and the impedance parameters are output to control the dynamic change relationship between the robot assembly force and the robot movement to achieve compliant control.

[0013] The industrial robot shaft-hole assembly network model based on joint reinforcement learning includes a state judgment module M01, a parameter optimization module M02 and an impedance control module M03;

[0014] The input of the state judgment module M01 is the judgment of the assembly contact state, and the contact state is output to the parameter optimization module M02; the parameter optimization module M02 optimizes the damping parameters according to the optimization strategy corresponding to the contact state, and assigns the output parameters to the diagonal elements of the diagonal damping matrix in the impedance control module M03.

[0015] The state judgment module M01 includes a state judgment Q network, a state judgment target Q network, and a state judgment experience replay pool; the state input S of the state judgment DQN is an eight-dimensional vector composed of six-dimensional force / torque information, the current Z-direction depth z, and the Z-direction offset dz in the previous training step:

[0016] S=[F x ,F y ,F z ,T x ,T y ,T z ,z,dz]

[0017] The output Stage is a one-dimensional number 0, 1, 2, representing the three states of shaft-hole single-point contact, three-point contact, and two-point contact respectively;

[0018] The state judgment DQN reward function is divided into two parts:

[0019]

[0020] in They are the positive and negative rewards for ending a training set;

[0021] Set to:

[0022]

[0023] k represents the current step number, k max It is to manually set a maximum number of steps in a training set;

[0024] Introducing pose error and insertion depth into negative reward values In the calculation of:

[0025]

[0026] Where z is the current depth in the Z direction (mm), and θ represents the magnitude of the posture error.

[0027] Among them, the parameter optimization module M02 includes a parameter optimization Q network, a parameter optimization target Q network and a parameter optimization experience replay pool; the parameter optimization module M02 optimizes the diagonal parameters of the diagonal positive definite damping matrix in the impedance control, and these parameters correspond to the six motion units of movement and rotation in the XYZ direction in the motion space of the robot end effector.

[0028] The state input of the parameter optimization module is the nine-dimensional vector S composed of the state input S of the state judgment module and the action output Stage q :

[0029] S q =[F x ,F y ,F z ,T x ,T y ,T z ,z,dz,Stage]

[0030] The output is a three-dimensional action sequence A, and the three elements in A are used as the damping coefficients of the impedance controller in the Z, Rx, and Ry directions respectively:

[0031] A=[B z ,B rx ,B ry ]

[0032] Similarly, the reward value r of parameter optimization DQN q Positive Rewards and negative rewards The parameter optimization module M02 is positively rewarded Negative Rewards A fuzzy reward system is used for calculation, and the input values ​​of the fuzzy set are divided into five triangular memberships, namely very bad (VB), bad (B), normal (N), good (G), and very good (VG).

[0033] The impedance control module M03 includes a variable impedance controller and an industrial robot physical environment; the selected impedance control law is:

[0034]

[0035] Among them, M d , B d are the diagonal positive definite matrices of system inertia and damping, V d , V are the expected and actual Cartesian velocity vectors of the robot end, F d is the desired external force vector, and F is the external force vector actually received by the end of the robot.

[0036] The inertia diagonal matrix M of the mass system d Set it to a constant and set the damping matrix B d The diagonal elements of are used as the optimization objects of joint reinforcement learning. By changing B d , the conversion relationship between the force error signal and the speed signal is adjusted to achieve smooth control in the insertion stage.

[0037] The present invention has the following beneficial effects:

[0038] 1. The method of the present invention adopts reinforcement learning for autonomous exploration, which saves a lot of prior knowledge and eliminates the requirement for accurate modeling of the assembly process.

[0039] 2. The present invention adopts a segmented assembly strategy of state judgment and parameter optimization to improve parameter search and utilization efficiency, and improve assembly quality and speed.

[0040] 3. The present invention designs a new reward function, taking into account the impact of posture error in the insertion phase on compliant control. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A schematic diagram of an industrial robot shaft hole assembly system of the present invention;

[0042] Figure 2 A schematic diagram of a state judgment and parameter optimization network model of the active compliance control method of the present invention;

[0043] Figure 3 It is a flow chart of the active compliance control method of the present invention;

[0044] Figure 1 The markup is represented by:

[0045] 10. Computer; 20. Industrial robot; 30. Six-dimensional force / torque sensor; 40. Shaft parts; 50. Sensor data acquisition box. DETAILED DESCRIPTION

[0046] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0047] like Figure 1 An industrial robot shaft hole assembly system based on joint reinforcement learning includes a computer 10, an industrial robot 20, a six-dimensional force / torque sensor 30, a shaft part 40, and a sensor data acquisition box 50; the six-dimensional force / torque sensor 30 is installed at the end of the industrial robot 20 and communicates with the sensor data acquisition box 50 through a data line; the sensor data acquisition box 50 is connected to the computer 10 and transmits the collected six-dimensional force / torque to the computer 10; the six-dimensional force / torque sensor 30, the industrial robot 20, and the assembled shaft part 40 have their own coordinate systems; the computer 10 and the industrial robot 20 perform two-way communication; the computer 10 controls the industrial robot 20 to assemble the shaft part 40, and the six-dimensional force / torque sensor 30 records and collects force / torque information, and the sensor data acquisition box 50 processes the force / torque information and transmits it to the computer 10. The computer 10 processes the force / torque information and robot posture information through a python interface. The computer 10 runs an industrial robot shaft-hole assembly network model based on joint reinforcement learning, outputs impedance parameters to control the dynamic relationship between the robot assembly force and the robot movement, and realizes compliant control.

[0048] Combination Figure 3 An industrial robot shaft-hole assembly model based on joint reinforcement learning includes a state judgment module M01, a parameter optimization module M02 and an impedance control module M03. The input of the state judgment module M01 is the judgment of the assembly contact state, and the contact state is output to the parameter optimization module M02. The parameter optimization module M02 optimizes the damping parameters according to the optimization strategy corresponding to the contact state, and assigns the output parameters to the diagonal elements of the diagonal damping matrix in the impedance control module M03.

[0049] (1) Status judgment module M01:

[0050] The state judgment module M01 includes the state judgment Q network, the state judgment target Q network, and the state judgment experience replay pool. The state input S of the state judgment DQN is an eight-dimensional vector composed of six-dimensional force / torque information, the current Z-direction depth z, and the Z-direction offset dz in the previous training step:

[0051] S=[F x ,F y ,Fz ,T x ,T y ,T z ,z,dz]

[0052] The output Stage is a one-dimensional number 0, 1, 2, representing the three states of shaft-hole single-point contact, three-point contact, and two-point contact, respectively.

[0053] The state judgment DQN reward function is divided into two parts:

[0054]

[0055] in They are the positive reward and negative reward for ending a training set. Set to:

[0056]

[0057] k represents the current step number, k max It is a manually set maximum step number constant for a training set. The present invention introduces posture error and insertion depth into the negative reward value In the calculation of:

[0058]

[0059] Where z is the current depth in the Z direction (mm), and θ represents the magnitude of the posture error.

[0060] Module process: ① Initialize state judgment DQN: first initialize the state judgment experience replay pool capacity Memorys = N; initialize parameter optimization Q network, randomly generate weight ωs; initialize parameter optimization target Q network, weight ω = -ωs; set the number of training episodes episodes = M, the number of steps in each training episode steps = T; ② Obtain DQN state input: the force information of the six-dimensional force / torque sensor and the robot position information constitute the system state input S; ③ Generate strategy: select a random state judgment action Stage with probability ∈, or select the action with the largest Q value with probability 1-∈, that is, Stage = maxaQ(S, a; ω), and the probability ∈ decreases with the increase of training sets.

[0061] (2) Impedance parameter optimization module M02:

[0062] The impedance parameter optimization module M02 includes parameter optimization Q network, parameter optimization target Q network, and parameter optimization experience playback pool. This module optimizes the diagonal parameters of the diagonal positive definite damping matrix in impedance control, which correspond to the six motion units of movement and rotation in the XYZ direction in the motion space of the robot end effector.

[0063] The state input of the parameter optimization module is the nine-dimensional vector S composed of the state input S of the state judgment module and the action output Stage q :

[0064] S q =[F x ,F y ,F z ,T x ,T y ,T z ,z,dz,Stage]

[0065] The output is a three-dimensional action sequence A, and the three elements in A are used as the damping coefficients of the impedance controller in the Z, Rx, and Ry directions respectively:

[0066] A=[B z ,B rx ,B ry ]

[0067] Similarly, the reward value r of parameter optimization DQN q Positive Rewards and negative rewards This module is rewarded Negative Rewards A fuzzy reward system is used for calculation, and the input values ​​of the fuzzy set are divided into five triangular memberships, namely very bad (VB), bad (B), normal (N), good (G), and very good (VG).

[0068] Final Output As shown in the following table:

[0069]

[0070] Module flow: ① Initialize parameter optimization DQN: First, initialize the parameter optimization experience pool capacity Memoryq = N; initialize the parameter optimization Q network and randomly generate weights ωq; initialize the parameter optimization target Q network with weights ω = -ωq; set episodes, steps and state judgment DQN to use the same settings; ② Get DQN state input: Use the input and output combination of the state judgment module as the input of the parameter optimization DQN; ③ Generate strategy: Select a random action sequence output with probability ∈, or select the action with the largest Q value with probability 1-∈, that is, A = maxQ(S q ,a;ω), the probability ∈ decreases as the number of training sets increases; ④ Loop through episodes=1,2,…,M; Loop through steps=1,2,…,T.

[0071] (3) Impedance control module M03:

[0072] The impedance control module includes a variable impedance controller and an industrial robot physical environment. The impedance control law selected by the present invention is:

[0073]

[0074] Among them, M d , B d are the diagonal positive definite matrices of system inertia and damping, V d , V are the expected and actual Cartesian velocity vectors of the robot end, F d is the desired external force vector, and F is the external force vector actually received by the end of the robot.

[0075] The present invention uses the inertia diagonal matrix M of the mass system as d Set it to a constant and set the damping matrix B d The diagonal elements of are used as the optimization objects of joint reinforcement learning. By changing B d , the conversion relationship between the force error signal and the speed signal is adjusted to achieve smooth control in the insertion stage.

[0076] Combination Figure 1 For the hardware environment, please refer to Figure 3 , introduces the specific steps of the reinforcement learning training process of an active compliant control method for industrial robot shaft hole assembly based on joint reinforcement learning:

[0077] Step S01: Obtain the initial state input of the reinforcement learning system, that is, the state input of the state judgment DQN. At the beginning of a step, the robot is given an initial assembly speed parallel to the hole axis until the axis contacts the hole to produce the initial state S, which is an eight-dimensional vector composed of six-dimensional force / torque information and the robot's Z-direction position information.

[0078] Step S02: Get the action output of the state judgment DQN. The state judgment Q network obtains the input state S, outputs the action Stage to judge the three states of single-point contact, three-point contact, and two-point contact, and selects different parameter optimization strategies

[0079] Step S03: Obtain the input of the parameter optimization module, which is the combination of the state input of the state judgment DQN and its output Stage S q

[0080] Step S04: Obtain the action output of the parameter-optimized DQN. The parameter-optimized Q network selects a three-dimensional action output from the parameter optimization strategy, representing the diagonal elements in the damping matrix that control the motion in the Z, Rx, and Ry directions.

[0081] Step S05: The three-dimensional action sequence output by the parameter optimization DQN is assigned to the diagonal elements in the diagonal damping matrix that control the movement of the robot end in the Z, Rx, and Ry directions, generating a compliant assembly action and obtaining a new state.

[0082] Step S06: The new state is the next state S_ of the state judgment DQN, and the reward value rs of the state judgment DQN and the reward value rq of the parameter optimization DQN are calculated respectively through the reward function

[0083] Step S07: Store rs and S_ into the state judgment experience replay pool. Input S_ into the state judgment Q network and combine it with the output Stage_ to form the next state S of the parameter-optimized DQN. q _.

[0084] Step S08: rq and S q _ Input parameters to optimize the experience replay pool and set S q _ is used as the state of the next step. The training is repeated until a single assembly task is completed or the maximum number of assembly steps is reached. This is considered the end of an episode, and the robot returns to the initial state and enters the next episode.

[0085] Step S09: Repeat the above training set cycle until episodes = N, the experience pool is full, and start iteratively updating the network parameters of the two DQN Q networks and the target Q network, and finally converge to the optimal strategy.

[0086] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. An industrial robot shaft hole assembly system, characterized by: It includes a computer (10), an industrial robot (20), a six-dimensional force / torque sensor (30), shaft parts (40), and a sensor data acquisition box (50); The six-dimensional force / torque sensor (30) is installed at the end of the industrial robot (20) and communicates with the sensor data acquisition box (50) via a data line; The sensor data acquisition box (50) is connected to the computer (10) to transmit the acquired six-dimensional force / torque to the computer (10); The six-dimensional force / torque sensor (30), the industrial robot (20), and the assembly shaft parts (40) have their own coordinate systems; The computer (10) and the industrial robot (20) perform two-way communication, the computer (10) controls the industrial robot (20) to assemble the shaft parts (40), and the six-dimensional force / torque sensor (30) records and collects force / torque information, and the sensor data acquisition box (50) processes the force / torque information and transmits it to the computer (10); The computer (10) processes the force / torque information and the robot posture information through a python interface; the computer (10) runs an industrial robot shaft-hole assembly network model based on joint reinforcement learning, outputs impedance parameters to control the dynamic change relationship between the robot assembly force and the robot movement, and realizes compliant control; The industrial robot shaft-hole assembly network model based on joint reinforcement learning includes a state judgment module M01, a parameter optimization module M02 and an impedance control module M03; The input of the state judgment module M01 is the judgment of the assembly contact state, and the contact state is output to the parameter optimization module M02; the parameter optimization module M02 optimizes the damping parameters according to the optimization strategy corresponding to the contact state, and assigns the output parameters to the diagonal elements of the diagonal damping matrix in the impedance control module M03; The state judgment module M01 includes a state judgment Q network, a state judgment target Q network and a state judgment experience playback pool; The state input S of the state judgment DQN is an eight-dimensional vector composed of six-dimensional force / torque information, the current Z-direction depth z, and the Z-direction offset dz in the previous training step: S=[F x ,F y ,F z ,T x ,T y ,T z ,z,dz] The output Stage is a one-dimensional number 0, 1, 2, representing the three states of shaft-hole single-point contact, three-point contact, and two-point contact respectively; The state judgment DQN reward function is divided into two parts: in They are the positive and negative rewards for ending a training set; Set to: k represents the current step number, k max It is to manually set a maximum number of steps in a training set; Introducing pose error and insertion depth into negative reward values In the calculation of: Where z is the current depth in the Z direction (mm), and θ represents the magnitude of the posture error.

2. An industrial robot shaft hole assembly system as claimed in claim 1, characterized in that: The parameter optimization module M02 includes a parameter optimization Q network, a parameter optimization target Q network, and a parameter optimization experience playback pool; the parameter optimization module M02 optimizes the diagonal parameters of the diagonal positive definite damping matrix in the impedance control, and these parameters correspond to the six motion units of movement and rotation in the XYZ direction in the motion space of the robot end effector; The state input of the parameter optimization module is the nine-dimensional vector S composed of the state input S of the state judgment module and the action output Stage q : S q =[F x ,F y ,F z ,T x ,T y ,T z ,z,dz,Stage] The output is a three-dimensional action sequence A, and the three elements in A are used as the damping coefficients of the impedance controller in the Z, Rx, and Ry directions respectively: A=[B z ,B rx ,B ry ] Similarly, the reward value r of parameter optimization DQN q Positive Rewards and negative rewards The parameter optimization module M02 is positively rewarded Negative Rewards A fuzzy reward system is used for calculation, and the input values ​​of the fuzzy set are divided into five triangular memberships, namely very bad (VB), bad (B), normal (N), good (G), and very good (VG).

3. The industrial robot shaft hole assembly system according to claim 1, characterized in that: The impedance control module M03 includes a variable impedance controller and an industrial robot physical environment; the selected impedance control law is: Among them, M d , B d are the diagonal positive definite matrices of system inertia and damping, V d , V are the expected and actual Cartesian velocity vectors of the robot end, F d is the desired external force vector, and F is the external force vector actually received by the robot end; Let M be the inertia diagonal matrix of the mass system. d Set it to a constant and set the damping matrix B d The diagonal elements of are used as the optimization objects of joint reinforcement learning. By changing B d , the conversion relationship between the force error signal and the speed signal is adjusted to achieve smooth control during the insertion phase.

4. The active compliance control method of the shaft hole assembly system of an industrial robot as claimed in claim 1, characterized in that: The following steps are involved: Step S01: Obtain the initial state input of the reinforcement learning system, that is, the state input of the state judgment DQN; Step S02: Obtain the action output of the state judgment DQN; Step S03: Obtain the input of the parameter optimization module, which is the combination of the state input of the state judgment DQN and its output Stage S q ; Step S04: Obtaining the action output of the parameter-optimized DQN; Step S05: The three-dimensional action sequence output by the parameter optimization DQN is assigned to the diagonal elements in the diagonal damping matrix that control the movement of the robot end in the Z, Rx, and Ry directions, generating a compliant assembly action and obtaining a new state; Step S06: The new state is the next state S_ of the state judgment DQN, and the reward value rs of the state judgment DQN and the reward value rq of the parameter optimization DQN are calculated respectively through the reward function; Step S07: store rs and S_ into the state judgment experience replay pool; Step S08: rq and S q _ Input parameters to optimize the experience replay pool and set S q _As the state of the next step, the training is cyclically repeated until a single assembly task is completed or the maximum assembly step is reached. This is considered the end of an episode, and the robot returns to the initial state and enters the next episode. Step S09: Repeat the above training set cycle until episodes = N, the experience pool is full, and start iteratively updating the network parameters of the two DQN Q networks and the target Q network, and finally converge to the optimal strategy.

5. The active compliance control method of the shaft hole assembly system of an industrial robot according to claim 4, characterized in that: In step S01, at the beginning of a step, the robot is given an initial assembly speed parallel to the hole axis until the axis contacts the hole to produce an initial state S, where S is an eight-dimensional vector composed of six-dimensional force / torque information and the robot's position information in the Z direction.

6. The active compliance control method of the shaft hole assembly system of an industrial robot according to claim 4, characterized in that: In step S02, the state judgment Q network obtains the input state S, and the output action Stage judges the three states of single-point contact, three-point contact, and two-point contact, and selects different parameter optimization strategies.

7. The active compliance control method of the shaft hole assembly system of an industrial robot as claimed in claim 4, characterized in that: In step S04, the parameter optimization Q network selects a three-dimensional action output from the parameter optimization strategy, representing the diagonal elements in the damping matrix that control the motion in the Z, Rx, and Ry directions.

8. The active compliance control method of the shaft hole assembly system of an industrial robot as claimed in claim 4, characterized in that: In step S07, S_ is input into the state judgment Q network and combined with the output Stage_ to form the parameter optimization DQN's next state S q _.

Citation Information

Patent Citations

  • Control system and method for learning variable impedance

    CN108153153A

  • Robot staged force guide assembling method and system based on deep reinforcement learning

    CN112847235A