A robot rigidity optimization weld seam polishing method and system based on proximal policy optimization multi-task reinforcement learning

By dividing the robotic weld grinding process into motion and posture adjustment stages, and using proximal strategies to optimize multi-task reinforcement learning, specific motion and state spaces for each stage are set, the vibration problem caused by weak stiffness in robotic weld grinding is solved, thus improving grinding stability and surface quality.

CN121132632BActive Publication Date: 2026-07-24HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2025-08-29
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies for robotic weld grinding suffer from vibrations due to weak stiffness, which affects the surface quality and force control, and also shortens the lifespan of the weld.

Method used

A proximal strategy optimization multi-task reinforcement learning method is adopted, which divides the robot end effector motion into motion stage and posture adjustment stage, sets control objectives for each stage, optimizes the robot end effector stiffness through reinforcement learning neural network, sets stage-specific motion space and state space, and designs relevant reward functions to improve stiffness.

Benefits of technology

It reduces vibration during weld grinding, improves grinding stability and surface quality, and extends robot lifespan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121132632B_ABST
    Figure CN121132632B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of robot motion and force control, and discloses a robot rigidity optimization weld seam polishing method and system based on proximal policy optimization multi-task reinforcement learning. The method comprises the following steps: dividing the movement of the robot from the starting position to the position to be polished into a movement stage and a posture adjustment stage; the control target of the movement stage is to move the grinding disc to the polishing starting point and maximize the average end rigidity of the robot in the initial state of the posture adjustment stage while avoiding collision; the control target of the posture adjustment stage is to maximize the average end rigidity of the robot in the polishing process while avoiding collision and making the robot end reach the polishing point; under the condition of simultaneously satisfying the two control targets, the movement trajectory and the posture adjustment trajectory of the robot end from the starting position to the position to be polished are calculated. Through the present application, the problem of vibration of the robot in the welding seam polishing process due to its weak rigidity state is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot motion and force control, and more specifically, relates to a robot stiffness optimization weld grinding method and system based on proximal policy optimization multi-task reinforcement learning. Background Technology

[0002] The current main method for optimizing the stiffness of robot grinding involves setting constraints and joint optimization parameters to determine the objective function for stiffness optimization. This transforms the robot stiffness optimization into a typical multivariable, single-objective, constrained optimization problem. Solving this nonlinear optimization problem yields the optimal trajectory under the highest stiffness. Solution methods include the projection gradient method, interior point method, sequential quadratic programming method, and genetic algorithm. However, their application in weld grinding has limited effectiveness.

[0003] To achieve stable robot-environment interaction and improve the quality of robot tasks, especially in weld grinding, stable force control at the robot's end effector is becoming increasingly essential. However, during processing, robots are prone to vibration due to their low stiffness, which can affect the surface quality of the ground surface, the effectiveness of force control, and the robot's lifespan.

[0004] Therefore, a method is needed to solve the above problems. Summary of the Invention

[0005] This invention provides a robot stiffness optimization weld grinding method based on proximal policy optimization multi-task reinforcement learning, which solves the problem of vibration caused by the robot being in a weak stiffness state during weld grinding.

[0006] To achieve the above objectives, according to one aspect of the present invention, a robot stiffness optimization weld grinding method based on proximal policy optimization multi-task reinforcement learning is provided, the method comprising the following steps: The movement of the robot end effector from the starting position to the position to be polished is divided into a motion phase and an attitude adjustment phase. The control objective for the motion phase is set as follows: while avoiding collisions, the grinding disc is moved to the grinding starting point and the average end effector stiffness of the robot in the initial state of the posture adjustment phase is maximized. The control objective for the attitude adjustment phase is set as follows: while avoiding collisions, the robot end effector reaches the point to be polished and the average end effector stiffness of the robot is maximized during the polishing process. Under the control objectives of both the motion phase and the posture adjustment phase, the motion trajectory and posture adjustment trajectory of the robot end effector from the starting position to the position to be polished are calculated.

[0007] More preferably, the state space of the motion phase is: The action space is: Where q is the joint angle of the robot. Position the robot along its axis. For the redundant angle around the grinding disc axis, This is the position vector of the weld grinding start point relative to the grinding disc. This represents the current length of the weld to be ground. This represents the change in the robot's end-effector position. This represents the change in redundant angle around the grinding disc axis. This represents the change in the robot's axial position.

[0008] More preferably, the reward function for the movement phase is:

[0009] in, 、 for 、 The reward coefficient, Penalties for collisions and joints exceeding their movement limits. The reward is given when the grinding disc approaches the starting point of grinding and eventually reaches a state with good robot average end stiffness.

[0010] More preferably, the values ​​of r1 and r2 are as follows:

[0011]

[0012] in, Let j be the distance between joint j and the workpiece. For safe distance threshold, The distance between the grinding disc and the starting point of the grinding process before the action is performed. The distance between the grinding wheel and the starting point of the grinding process after the action is performed. A value less than 1cm is considered reached. To maintain a constant robot axis position Redundancy angle around the grinding disc axis Average end effector stiffness of the robot performing the grinding process.

[0013] More preferably, the state space of the attitude adjustment phase is: The motion space is: Where q is the joint angle of the robot. Position the robot along its axis. For the redundant angle around the grinding disc axis, This is the position vector of the weld grinding start point relative to the grinding disc. This represents the current length of the weld to be ground. This represents the change in redundant angle around the grinding disc axis. This represents the change in the robot's axial position.

[0014] More preferably, the reward function for the attitude adjustment phase is:

[0015] in, 、 for 、 The reward coefficient, Penalties for collisions and joints exceeding their movement limits. The average end effector stiffness of the robot is rewarded.

[0016] More preferably, the The calculation formula is as follows:

[0017] in, This represents the total number of weld path points. The index of the path point where the inverse solution does not exist. This represents the average end effector stiffness index of the robot.

[0018] More preferably, the formula for calculating the average end effector stiffness of the robot is as follows:

[0019]

[0020] in, is the average end effector stiffness of the robot, m is the total number of weld points, and i is the weld point number. This refers to the overall stiffness performance index of the robot when grinding the i-th weld point. It is the end force. It is the displacement compliance matrix.

[0021] More preferably, the motion trajectory and attitude adjustment trajectory of the computational robot end effector from the starting position to the grinding position are performed using a reinforcement learning neural network. When calculating the motion trajectory, the input of the reinforcement learning neural network is the state space of the motion phase, and the output is the action space of the motion phase. The goal is to maximize the reward function of the motion phase. When calculating the attitude adjustment trajectory, the input of the reinforcement learning neural network is the state space of the attitude adjustment phase, and the output is the action space of the attitude adjustment phase. The goal is to maximize the reward function of the attitude adjustment phase.

[0022] According to another aspect of the present invention, a robot stiffness optimization weld grinding system based on proximal policy optimization multi-task reinforcement learning is provided. The system includes an actuator for performing the above-described robot stiffness optimization weld grinding method based on proximal policy optimization multi-task reinforcement learning.

[0023] In summary, the technical solutions conceived by this invention have the following beneficial effects compared with the prior art: 1. This invention divides the movement of the robot end effector from the starting position to the position to be polished into two stages, and sets the objectives for each stage. This achieves two tasks: maximizing the average end effector stiffness of the robot during the polishing process and rapidly moving to the position to be polished. This reduces the vibration of the polishing force during the robot's weld polishing process and improves the stability of the polishing process.

[0024] 2. In this invention, action space and state space are set separately for the motion phase and the posture adjustment phase to meet the control requirements of different phases. In the motion phase, it is necessary to control the displacement of the mobile platform and the movement of the robot joints. In the posture adjustment phase, the position of the mobile platform is fixed, and it is only necessary to control the movement of the robot joints. Setting action space and state space separately can improve the network's targeting of the tasks in the two phases and improve the training effect.

[0025] 3. The reward functions set in this invention for both the motion phase and the attitude adjustment phase are related to the average end-effector stiffness of the robot. This is used to continuously optimize the end-effector stiffness of the robot in both phases. Optimizing stiffness in the motion phase can ensure that the initial stiffness of the attitude adjustment phase is relatively high, thereby improving the final stiffness optimization effect. Attached Figure Description

[0026] Figure 1 This is a flowchart of a robot stiffness optimization and polishing process based on proximal policy optimization multi-task reinforcement learning constructed according to a preferred embodiment of the present invention.

[0027] Figure 2 This is a joint stiffness identification platform according to a preferred embodiment of the present invention.

[0028] Figure 3 It is a strategy network framework constructed according to a preferred embodiment of the present invention.

[0029] Figure 4 This is an evaluation network framework constructed according to a preferred embodiment of the present invention.

[0030] Figure 5 This is a training process for a robot stiffness optimization weld grinding strategy based on a multi-task reinforcement learning algorithm, according to a preferred embodiment of the present invention.

[0031] Figure 6These are reward curves for the training process of a multi-task reinforcement learning algorithm constructed according to a preferred embodiment of the present invention. Wherein, (a) is the reward curve for task 1, and (b) is the reward curve for task 2.

[0032] Figure 7 This is a robot stiffness optimization weld grinding experimental platform according to a preferred embodiment of the present invention.

[0033] Figure 8 The welded workpiece is an experimental process according to a preferred embodiment of the present invention.

[0034] Figure 9 The diagram shows the control effect of robot weld grinding according to a preferred embodiment of the multi-task reinforcement learning method of the present invention. Among them, (a) is a diagram showing the change in the distance between the grinding disc and the grinding start point when the control disc reaches the grinding start point, and (b) is a diagram showing the change in the average end effector stiffness of the robot during the posture optimization and adjustment stage after reaching the grinding start point.

[0035] Figure 10 This is a tracking diagram of the actual robot weld grinding force according to a preferred embodiment of the present invention. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0037] like Figure 1 As shown, a robot stiffness optimization grinding method based on proximal policy optimization multi-task reinforcement learning is presented. The method includes the following steps: S1 optimizes the stiffness of the robot during grinding. First, the robot's stiffness is modeled and identified, and stiffness performance indicators are proposed for subsequent stiffness optimization.

[0038] When modeling the stiffness of a robot, the main task is to derive the relationship between the deformation of the robot's end effector and the force it experiences during the grinding task. The following three assumptions are made: (1) The deformation of the robot under stress comes from the joints; (2) The deformation of the robot's joints satisfies Hooke's Law; (3) The robot's deformation is small, and its Jacobian matrix will not change before and after being subjected to force.

[0039] Joint stiffness This indicates that, based on assumption 2, the joint torque can be obtained. With joint deformity The relationship is:

[0040] in, Joint stiffness matrix Joint torque Joint deformity .

[0041] Robot end-effector deformation D and joint deformation The relationship is:

[0042] in, for Jacobian matrix, , For end translational deformation, This is due to torsional deformation at the end.

[0043] Ignoring the effect of joint friction, joint torque The relationship between the force vector F acting on the end and the force vector F is as follows:

[0044] in, for The Jacobian matrix, , As the end force, This is the end-effector torque.

[0045] From this, we can obtain the relationship between the end deformation and the applied six-dimensional force vector F:

[0046] in, for The joint flexibility matrix, It is a six-dimensional column vector. Since it is a diagonal matrix, therefore:

[0047] in, This is the column vector of joint flexibility.

[0048] End-effector torsional deformation caused by end force vector action It is very small and can be ignored, so it can be simplified to:

[0049]

[0050]

[0051] in, for matrix Then, a force loading experiment was conducted to identify the joint flexibility column vector C. The identification platform is as follows: Figure 2 As shown. By measuring the applied forces and deformations under n different robot postures, a set of statically indeterminate equations is obtained:

[0052] The joint flexibility column vector C is obtained by fitting the joint flexibility using the least squares method, and the joint stiffness is obtained by taking the reciprocal of the joint flexibility.

[0053] A robot stiffness performance index is established to quantitatively evaluate the end effector stiffness performance, serving as an indicator for subsequent robot stiffness optimization. The comprehensive stiffness performance index is defined as follows:

[0054] in, The magnitude of the end force. denoted by k, represents the magnitude of the translational deformation at the end effector. k measures the robot's ability to resist deformation caused by end effector forces.

[0055] Will terminal compliance matrix Write it in the following form:

[0056] in, Here is the displacement compliance matrix. Here is the coupling flexibility matrix. Here is the rotational compliance matrix.

[0057] During robotic grinding, the torque acting on the robot's end effector spindle is very small, and the end effector torque and torsional deformation can be ignored. The relationship between the robot's end effector deformation and the force vector can be simplified as follows:

[0058] Therefore, we can obtain:

[0059] Assume end force Assuming a unit force, the robot's end effector stiffness performance index can be obtained as follows:

[0060] The above formula represents the stiffness performance of a robot with a single configuration, specifically for robot weld grinding tasks. To comprehensively evaluate the overall stiffness performance of the polishing process, an average end effector stiffness index for the robot is proposed:

[0061] in For the first i The robot's overall stiffness performance index when grinding individual weld points.

[0062] S2 analyzes the requirements for stiffness optimization and collision-free tasks, and designs the state space, action space and reward function for reinforcement learning.

[0063] The robot is mounted on an AGV that moves along the weld seam axis. When the robot moves to grind the weld seam, it first needs to move the grinding disc to the grinding starting point, and then adjust its posture. Stiffness optimization is continuously performed during these two stages to improve the robot's stiffness during grinding. The control phase before grinding begins is divided into a motion phase and a posture adjustment phase. The control objective of the motion phase is to move the grinding disc quickly to the grinding starting point while avoiding collisions, and to ensure that the initial state of the posture adjustment phase has a high average end effector stiffness. The control objective of the posture adjustment phase is to position the robot axially while avoiding collisions. and redundant angle around the grinding disc axis Adjustments were made to use redundant values. , During grinding, the grinding disc can reach the grinding point, and the robot has a high average end effector stiffness.

[0064] (1) Movement phase The main task of the motion phase is to control the movement of the AGV and the robot's end effector, enabling the grinding disc to move quickly to the grinding starting point while avoiding collisions, and ensuring that the initial state of the attitude adjustment phase has good average end effector stiffness. The state space is set as a 12-dimensional vector:

[0065] Where q represents the angles of the robot's six joints. Position the robot along its axis. For the redundant angle around the grinding disc axis, This is the position vector of the weld grinding start point relative to the grinding disc. This represents the current length of the weld seam to be ground.

[0066] The action space is set as a 5-dimensional vector:

[0067] in, This represents the change in the robot's end-effector position. This represents the change in redundant angle around the grinding disc axis. This represents the change in the robot's axial position.

[0068] The reward function is designed as follows:

[0069] in, 、 for 、 The reward coefficient, Penalties for collisions and joints exceeding their movement limits.

[0070] In one embodiment of the present invention, The possible values ​​are as follows:

[0071] in, Let j be the distance between joint j and the workpiece. This is the safe distance threshold.

[0072] In one embodiment of the present invention, The possible values ​​are as follows: The reward for the grinding disc approaching the grinding start point and eventually reaching a state of good robot average end-effector stiffness is as follows: the closer the grinding disc is to the grinding start point, the greater the reward value. When the grinding disc reaches the grinding start point, a large reward related to the robot's average end-effector stiffness is given.

[0073]

[0074] in The distance between the grinding disc and the starting point of grinding before the action is performed; The distance between the grinding disc and the starting point of grinding after the action is performed; A value less than 1cm is considered reached. To maintain a constant robot axis position Redundancy angle around the grinding disc axis Average end effector stiffness of the robot performing the grinding process.

[0075] (2) Posture adjustment stage The main task of the attitude adjustment phase is to adjust the robot's axial positioning. The redundant angle α around the grinding disc axis is used to optimize the robot's average end effector stiffness while avoiding collisions, ensuring the grinding disc can reach the processing point and selecting a better value for the grinding process. α. The state space is the same as the motion phase, and the action space is set as a 2-dimensional vector:

[0076] in This represents the change in redundant angle around the grinding disc axis. This represents the change in the robot's axial position.

[0077] The reward function is designed as follows:

[0078] in 、 for 、 The reward coefficient, The penalties for collisions and joints exceeding their movement limits are the same as those for the movement phase; The robot's average end effector stiffness and accessibility are rewarded: the entire posture optimization and adjustment process involves 60 control steps to complete the robot's axial positioning. and redundant angle around the grinding disc axis α The adjustment, the adjusted 、α Used in subsequent grinding; accessibility can be determined by solving for whether there is an inverse kinematic solution at each weld path point.

[0079]

[0080] in This represents the total number of weld path points. The index of the path point where the inverse solution does not exist. This represents the average end effector stiffness index of the robot.

[0081] S3 constructs and trains a proximal policy optimization multi-task reinforcement learning neural network.

[0082] Proximal policy optimization multi-task reinforcement learning consists of two parts: a policy network and an evaluation network. The policy network outputs the action space, and the evaluation network outputs the state-value function. The hidden layers of both the policy and evaluation networks use the tanh function as the activation function, mapping the network outputs to the range (-1, 1). The policy network outputs a Gaussian-mean distribution. The learnable parameter is the variance of a Gaussian distribution. The final output action follows a Gaussian distribution. The sampled values. Two network models are used for the motion phase and the posture adjustment phase, respectively. There are common parts and special parts between the two models. The common parts are used to learn a general policy, and the special parts are used to learn a specific policy.

[0083]

[0084] Policy networks such as Figure 3 As shown, the input is a 12-dimensional state vector. The shared part consists of a hidden layer h1 of size 256 and a hidden layer h2 of size 128. The hidden layer size is 5 during the motion phase and 2 during the attitude adjustment phase.

[0085] Evaluating networks such as Figure 4As shown, the input is a 12-dimensional state vector, with a shared part consisting of a hidden layer h1 of size 256 and a hidden layer h2 of size 128. The hidden layer size is 1 in both the motion phase and the attitude adjustment phase.

[0086] After the network is built, a multi-task reinforcement learning algorithm is used to train the robot stiffness optimization weld grinding strategy. The process is as follows: Figure 5 As shown, the reward curve during the training process is as follows: Figure 6 As shown.

[0087] To verify the effectiveness of the robot stiffness optimization weld grinding method proposed in this invention, control performance was verified using an embodiment. In this embodiment, a robot weld grinding task platform was built for experimental verification, such as... Figure 7 As shown, it includes a workpiece, weld, UR16e robot, electric spindle, ATI force sensor, distance sensor, and workstation. Figure 8 The weld seam is for grinding and is straight, extending along the direction of AGV movement. Figure 9 The diagram shows the control effect when grinding one of the welds using the method proposed in this invention. The stiffness after optimization is significantly improved compared to before optimization. Figure 10 The image shows a comparison of the force tracking performance when grinding a weld using the method proposed in this invention and a conventional reinforcement learning method. The contact force vibration amplitude is smaller when grinding the weld using the method proposed in this invention compared to the conventional reinforcement learning method, demonstrating the effectiveness of the method proposed in this invention in optimizing robot stiffness.

[0088] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A robot stiffness optimization weld grinding method based on proximal policy optimization multi-task reinforcement learning, characterized in that, The method includes the following steps: The movement of the robot end effector from the starting position to the position to be polished is divided into a motion phase and an attitude adjustment phase. The control objective for the motion phase is set as follows: while avoiding collisions, the grinding disc is moved to the grinding starting point and the average end effector stiffness of the robot in the initial state of the posture adjustment phase is maximized. The control objective for the attitude adjustment phase is set as follows: while avoiding collisions, the robot end effector reaches the point to be polished and the average end effector stiffness of the robot is maximized during the polishing process. Under the control objectives of both the motion phase and the posture adjustment phase, calculate the motion trajectory and posture adjustment trajectory of the robot end effector from the starting position to the position to be polished. The state space of the motion phase is: The motion space is: Where q is the joint angle of the robot. Position the robot along its axis. For the redundant angle around the grinding disc axis, This is the position vector of the weld grinding start point relative to the grinding disc. This represents the current length of the weld to be ground. This represents the change in the robot's end-effector position. This represents the change in redundant angle around the grinding disc axis. This represents the change in the robot's axial position. The reward function for the aforementioned movement phase is: in, 、 for 、 The reward coefficient, Penalties for collisions and joints exceeding their movement limits. The reward is given when the grinding disc approaches the starting point of grinding and eventually reaches a state with good average robot end stiffness. The state space for the attitude adjustment phase is: The motion space is: Where q is the joint angle of the robot. Position the robot along its axis. For the redundant angle around the grinding disc axis, This is the position vector of the weld grinding start point relative to the grinding disc. This represents the current length of the weld to be ground. This represents the change in redundant angle around the grinding disc axis. This represents the change in the robot's axial position. The reward function for the attitude adjustment phase is: in, 、 They are respectively 、 The reward coefficient, Penalties for collisions and joints exceeding their movement limits. The average end effector stiffness of the robot is rewarded.

2. The robot stiffness optimization weld grinding method based on proximal policy optimization multi-task reinforcement learning as described in claim 1, characterized in that, The values ​​of r1 and r2 are as follows: in, Let j be the distance between joint j and the workpiece. For safe distance threshold, The distance between the grinding disc and the starting point of the grinding process before the action is performed. The distance between the grinding wheel and the starting point of the grinding process after the action is performed. A value less than 1cm is considered reached. To maintain a constant robot axis position Redundancy angle around the grinding disc axis Average end effector stiffness of the robot performing the grinding process.

3. The robot stiffness optimization weld grinding method based on proximal policy optimization multi-task reinforcement learning as described in claim 1, characterized in that, The The calculation formula is as follows: in, This represents the total number of weld path points. The index of the path point where the inverse solution does not exist. This represents the average end effector stiffness index of the robot.

4. A robot stiffness optimization weld grinding method based on proximal policy optimization multi-task reinforcement learning as described in claim 1 or 3, characterized in that, The formula for calculating the average end effector stiffness of the robot is as follows: in, is the average end effector stiffness of the robot, m is the total number of weld points, and i is the weld point number. This refers to the overall stiffness performance index of the robot when grinding the i-th weld point. It is the end force. It is the displacement compliance matrix.

5. A robot stiffness optimization weld grinding method based on proximal policy optimization multi-task reinforcement learning as described in claim 1 or 3, characterized in that, The motion trajectory and attitude adjustment trajectory of the computational robot end effector from the starting position to the grinding position are calculated using a reinforcement learning neural network. When calculating the motion trajectory, the input of the reinforcement learning neural network is the state space of the motion phase, and the output is the action space of the motion phase. The goal is to maximize the reward function of the motion phase. When calculating the attitude adjustment trajectory, the input of the reinforcement learning neural network is the state space of the attitude adjustment phase, and the output is the action space of the attitude adjustment phase. The goal is to maximize the reward function of the attitude adjustment phase.

6. A robot stiffness optimization weld grinding system based on proximal policy optimization multi-task reinforcement learning, characterized in that, The system includes an actuator for performing a robot stiffness optimization weld grinding method based on proximal policy optimization multi-task reinforcement learning as described in any one of claims 1-5.