A robot constant force control polishing method based on PPO reinforcement learning

By combining PPO reinforcement learning and impedance controller, real-time estimation of environmental parameters, adjustment of grinding and polishing trajectory and addition of force closed-loop feedback, the problem of contact force control in unknown environments during the robot grinding and polishing process is solved, and high-precision and robust constant force control effect is achieved.

CN117226613BActive Publication Date: 2025-10-03HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311444136.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-10-03
Estimated Expiration
2043-10-31

AI Technical Summary

Technical Problem

It is difficult to achieve stable force control during existing robotic grinding and polishing processes, especially in unknown environments. Traditional impedance controllers lack accurate environmental parameter information, resulting in poor contact force control.

Method used

A PPO reinforcement learning-based method is used in combination with an impedance controller to estimate the environmental stiffness and position in real time. The grinding and polishing trajectory is adjusted through reinforcement learning, and force closed-loop feedback is added to optimize the force control of the robot end.

Benefits of technology

The contact force control accuracy and robustness during the grinding and polishing process are improved, the steady-state error of force tracking is reduced, and the autonomous adjustment and high-precision tracking of the robot's constant force control are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117226613B_ABST
    Figure CN117226613B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field related to robot grinding and polishing, and discloses a robot constant force control grinding method based on PPO reinforcement learning. The method includes: S1 for a three-dimensional model or point cloud model of a workpiece to be processed, obtaining the original grinding and polishing trajectory for processing the workpiece to be processed; S2 selecting an impedance control method for constant force grinding control, and thereby constructing an impedance controller containing unknown parameters and corresponding constraints; S3 calculating the environmental stiffness and position of the robot end in real time, using the calculated environmental stiffness and position to calculate the robot normal control instruction, and adjusting the normal displacement of the original trajectory in real time according to the normal control instruction, so that the actual grinding and polishing force is equal to the preset expected grinding and polishing force; S4 solving the unknown parameters in the impedance controller and then determining the impedance controller, and performing constant force grinding on the robot according to the impedance controller. Through the present invention, the problem of how to achieve constant force control of the grinding and polishing force during the grinding and polishing process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field related to robot grinding and polishing, and more specifically, relates to a robot constant force control grinding method based on PPO reinforcement learning. Background Art

[0002] To achieve stable interaction between robots and their environments, there is an increasing need for stable force control at the robot's end. Position-based impedance controllers are used to receive contact force signals to track a constant desired force. The dynamic parameters of most robots are often difficult to identify. Furthermore, due to safety concerns, these robots are not highly open, typically not providing underlying control interfaces. They only provide position control modes and lack direct access to joint currents. When performing force control on this type of robot, it is necessary to control the robot's mechanical impedance characteristics by generating a reference trajectory for the existing position controller, a process known as position-based impedance control.

[0003] Traditional impedance control has a simple structure and is widely used in robotic force control, often to achieve compliant control of robots. However, when working in an unknown environment, the lack of precise environmental parameter information leads to poor contact force control. To reduce the steady-state error in contact force, the surface position and stiffness of the environment must be known in advance. Therefore, a method for achieving steady-state contact force control during grinding and polishing is urgently needed. Summary of the Invention

[0004] In response to the above defects or improvement needs of the prior art, the present invention provides a robot constant force control grinding method based on PPO reinforcement learning to solve the problem of how to achieve constant force control of the grinding and polishing force during the grinding and polishing process.

[0005] To achieve the above objectives, according to one aspect of the present invention, a robot constant force control polishing method based on PPO reinforcement learning is provided, the method comprising the following steps:

[0006] S1: for a three-dimensional model or a point cloud model of a workpiece to be processed, obtaining an original grinding and polishing trajectory for processing the workpiece to be processed;

[0007] S2 selects the impedance control method for constant force grinding control, and uses it to construct an impedance controller with unknown parameters and corresponding constraints;

[0008] S3 calculates the environmental stiffness and position of the robot end in real time, calculates the robot normal control instruction using the calculated environmental stiffness and position, and adjusts the normal displacement of the original trajectory in real time according to the normal control instruction so that the actual grinding and polishing force is equal to the preset expected grinding and polishing force;

[0009] S4: solving unknown parameters in the impedance controller to determine the impedance controller, and performing constant force grinding on the robot according to the impedance controller.

[0010] Further preferably, in step S2, the impedance controller is performed according to the following formula:

[0011]

[0012] in, is the coefficient of inertia, is the damping coefficient, is the stiffness coefficient, is the actual position of the robot end With the expected position The error, and They are The first and second derivatives of is the expected contact force Actual contact force The error, is the proportionality coefficient, is the integration coefficient.

[0013] Further preferably, in step S2, the constraint condition is obtained according to the following steps:

[0014] S21 calculates the initial stiffness and damping of the grinding environment, where the environment is the grinding tool and workpiece as a whole;

[0015] S22 constructs the constraints using the obtained initial stiffness and damping.

[0016] Further preferably, in step S21, the initial stiffness and damping are performed according to the following formula:

[0017]

[0018] in, , is the parameter obtained by iteratively using the recursive augmented least squares method, , are the identified initial damping and initial stiffness, It is a time period.

[0019] Further preferably, in step S22, the constraint conditions are as follows:

[0020]

[0021] in, , , is the coefficient in the impedance equation, , is an intermediate variable, is the environmental stiffness.

[0022] Further preferably, in step S3, the normal control instruction of the robot is performed according to the following formula:

[0023]

[0024] in, is the desired position, It is the power of expectation, is the environmental location, is the environmental stiffness.

[0025] Further preferably, the environmental stiffness and position are calculated according to the following formula:

[0026]

[0027] in, is the initial ambient stiffness, is the initial environment position, , is the measured contact force, is the estimated ambient stiffness, yes The estimated environmental position at each moment, yes The estimated environmental position at each moment, yes The robot's end position at the moment, is the robot's movement time, , , is a constant and satisfies , is the differential of the original trajectory generated by the robot.

[0028] Further preferably, in step S4, the unknown parameters are solved using a reinforcement learning method.

[0029] Further preferably, the reinforcement learning method is performed according to the following steps:

[0030] S41 Construct the reward function for reinforcement learning and set the action space and state space;

[0031] S42 uses the parameters of the state space as input and the parameters of the action space as output to construct a reinforcement learning strategy neural network, assigns initial values ​​to the unknown parameters, and stops training when the training reward value converges stably. The current corresponding unknown parameter value is the required unknown parameter value.

[0032] Further preferably, in step S41, the reward function is performed according to the following formula:

[0033]

[0034] The working space is carried out according to the following formula:

[0035] a=

[0036] The state space is carried out according to the following formula:

[0037]

[0038] in, is the deviation between the actual force and the expected force, is the robot terminal velocity, It is a positive number.

[0039] In general, the above technical solutions conceived by the present invention have the following beneficial effects compared with the prior art:

[0040] 1. The impedance controller constructed by the present invention includes ,Compared to the traditional impedance controller, it adds force closed-loop feedback in the impedance ,control, and adopts PI control to ensure higher contact force control ,precision;

[0041] 2. In step S3, the present invention compensates for the normal displacement in the original polishing trajectory by estimating the environmental position and stiffness, so that the actual polishing force is equal to the preset expected polishing force. This overcomes the inherent defects of the impedance controller, achieves compensation for the actual polishing force, and improves polishing accuracy.

[0042] 3. The solution of unknown parameters in step S4 of the present invention is obtained and , realizing closed-loop control of grinding and polishing force, by feeding back the real-time monitored contact force to the control system, the control system makes real-time adjustments based on the monitoring results, thus reducing the fluctuation of grinding and polishing force during the grinding and polishing process;

[0043] 4. This invention applies PPO reinforcement learning to the control of robotic constant-force polishing. Using the Lyanov stability determination method, the environment's position and stiffness are estimated online, adjusting the robot's reference trajectory and reducing the steady-state error in force tracking. Using reinforcement learning and incorporating closed-loop force control, this approach eliminates the need for prior models of control parameters and polishing force errors, improving the robustness of constant-force tracking.

[0044] 5. To improve force tracking performance, the present invention uses reinforcement learning to make adjustments. This approach does not require expert knowledge or a priori understanding of the underlying complex world. It can autonomously discover optimal behavior in the process of repeated interaction with the environment. The proposed method aims to combine force control with RL to learn contact constant-force polishing tasks when using a position-controlled robot. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a flowchart of a robot constant force control polishing method based on PPO reinforcement learning constructed according to a preferred embodiment of the present invention;

[0046] Figure 2 1 is a structural diagram of a workpiece to be processed according to a preferred embodiment of the present invention, wherein (a) is an inclined workpiece and (b) is a curved workpiece;

[0047] Figure 3 are the initial stiffness and damping of the environment obtained by the recursive least square method based on a variable forgetting factor according to a preferred embodiment of the present invention, wherein (a) is the estimated stiffness, and (b) is the estimated damping;

[0048] Figure 4 1 is a design diagram of a reinforcement learning value function network structure and a policy network structure according to a preferred embodiment of the present invention, wherein (a) is a schematic diagram of the reinforcement learning value function network structure, and (b) is a design diagram of the policy network structure;

[0049] Figure 5 This is a block diagram of a robot grinding and polishing constant force control based on reinforcement learning constructed according to a preferred embodiment of the present invention;

[0050] Figure 6 is a reward value image trained 100 times by reinforcement learning according to a preferred embodiment of the present invention;

[0051] Figure 7 1 is an actual grinding force tracking diagram according to a preferred embodiment of the present invention, wherein (a) is a bevel grinding force tracking diagram after robot reinforcement learning training, and (b) is a curved surface grinding force tracking diagram after robot reinforcement learning training. DETAILED DESCRIPTION

[0052] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0053] like Figure 1As shown in FIG, a robot constant force control polishing method based on PPO reinforcement learning specifically includes the following steps:

[0054] S1 obtains the original trajectory data of the robot grinding and polishing through the three-dimensional model or point cloud model of the workpiece.

[0055] The robot's initial grinding and polishing trajectory is generated through the three-dimensional model or point cloud of the workpiece using the trajectory generation method.

[0056] S2 estimates the initial stiffness and damping of the environment through recursive augmented least squares method, designs the robot impedance controller, and selects appropriate impedance parameters. In this embodiment, the environment refers to the grinding tool and workpiece as a whole.

[0057] S21 uses recursive augmented least squares to make an initial estimate of the equivalent stiffness of the robot end and the workpiece, using a variable forgetting factor , ,in is the attenuation coefficient, is the difference between the current output and the previous output, is the minimum forgetting factor.

[0058] Assume that the contact dynamics between the grinding tool and the workpiece is , after z-transform discretization, we can get , the meaning of each parameter is:

[0059]

[0060] in, represents the contact force with the environment, represents the end position of the robot, Indicates the location of the environment. , denote the damping and stiffness of the environment, represents the sensor noise, , , represents the intermediate variable, represents a time series, Indicates the difference between the robot's current position and the previous position.

[0061] The parameters that need to be estimated using the RELS recursive formula are as follows:

[0062]

[0063] where the gain vector Calculated as follows:

[0064]

[0065]

[0066] Finally, the initial stiffness and damping of the environment are obtained:

[0067]

[0068] The final experimental results are as follows Figure 3 shown.

[0069] in represents the estimated ambient damping, represents the estimated environmental stiffness.

[0070] S22 uses impedance control to perform constant force grinding control. The existing impedance control equation is as follows: ,in, Expectation and actual contact force The difference, The actual contact force is measured by the sensor in real time. is the inertia coefficient, is the damping coefficient, is the stiffness coefficient, is the actual position of the robot end, is the desired position. Based on the estimated initial parameters of the environment, the following constraints are constructed:

[0071]

[0072] Within this range, select the impedance parameter and select , Parameters, then Can be achieved through Calculated.

[0073] S23 constructs an impedance controller, , adding force closed-loop feedback to impedance control, and using PI control to ensure higher contact force control accuracy, Indicates the actual position of the robot end With expected position The error, and Respectively The first and second derivatives of Denotes the expected contact force Actual contact force The error, represents the proportionality coefficient, Represents the integral coefficient.

[0074] S3 uses the Lyapunov stability method to estimate the environmental position and stiffness parameters, adjust the original trajectory of the robot grinding and polishing, and reduce the steady-state error.

[0075] By the Lyapunov stability method, the estimated equations for the ambient stiffness and position are:

[0076]

[0077] in, is the initial value of the environmental stiffness, is the initial value of the environment position, , represents the measured contact force, represents the estimated environmental stiffness, express The estimated environmental position at each moment, express The estimated environmental position at each moment, express The robot's end position at the moment, represents the robot's movement time, , , is a small constant, and there is . The differential of the original trajectory generated by the robot. The estimated environmental stiffness and position of the robot are obtained in real time. The normal displacement of the original trajectory of the robot is adjusted, that is, the normal control instruction of the robot is set to , follow this instruction to adjust the grinding position of the grinding tool and reduce , which can reduce the steady-state error of robot force tracking.

[0078] S4 Calculation and .

[0079] S41 analyzes the influencing factors of constant force control, constructs reinforcement learning reward function, and sets action space and state space.

[0080] The goal of robot training is to make the actual contact force at the end of the robot smaller than the expected contact force and to make the normal velocity of the end of the robot smaller.

[0081] The reward function is set to , is the deviation between the actual force and the expected force, is the robot terminal velocity, Is a positive number.

[0082] Set the action space of reinforcement learning to be a= , is the real-time changing proportional coefficient, is the real-time changing integral coefficient, and is the output value of the reinforcement learning model. The state space is set to , for real-time changes 、 , and , Considering that the actual force measured by the force sensor often has a lot of noise, it is not advisable to directly differentiate the force error. and environmental parameters can be considered constants in a short period of time, then , so the state space is set to , .

[0083] S42 builds a reinforcement learning strategy neural network, trains it based on the PPO reinforcement learning method, and uses the trained model to perform constant force grinding control of the robot.

[0084] like Figure 4 Figure 2 shows the design of a deep neural network for reinforcement learning training, including the design of the policy network and the value function network. Because the training parameters are not complex, a three-layer neural network is used for training, with 128 nodes per layer. The activation function between each hidden layer is the Tanh activation function, and the output of the policy network is a Gaussian-distributed sample value.

[0085] The training data is obtained by setting the initial value 、 Then let the robot perform cyclic grinding on the workpiece to be ground, and record the contact force and the end position of the robot during the grinding process, which are input into the model as training data, and then the training data is continuously updated. 、 parameter.

[0086] Normalize the training data, normalize the input state quantity and output state quantity of the neural network, and divide them by their corresponding upper limits so that the input and output value ranges of the neural network are both [-1, 1]. The control instruction obtained by the impedance equation is , the robot's joint angles are obtained by inverse kinematics to control the robot. In order to avoid excessive contact force during training on a real robot, the contact force is always monitored. and the inverse joint angles ,When the contact force and joint angle change too greatly, the robot is ,directly moved to a safe position and the training is terminated.

[0087] In this embodiment, the measured six-dimensional force sensor data also needs to be gravity compensated. A grinding tool is set at the end of the robot, and the sensor is set between the end of the robot and the grinding tool. The results displayed by the sensor include the gravity of the grinding tool and the contact force with the workpiece during the grinding process. The gravity compensation is to subtract the gravity of the grinding tool to obtain the contact force between the grinding tool and the workpiece. Specifically: the center of mass of the robot end tool is calculated by the least squares method ,gravity , sensor zero drift Through coordinate transformation, the measured force data is fed forward and compensated to eliminate the influence of the end gravity and obtain the actual contact force of the robot end. Figure 2 As shown in the figure, this figure shows the artifacts used during the experiment.

[0088] The PPO reinforcement learning method is used for constant force control of robot grinding and polishing. The machining process contains a total of 750 control cycles with a control frequency of 50 Hz and a force sensor frequency of 125 Hz. The first 200 control cycles are the robot approach stage.

[0089] like Figure 5 As shown, the entire robot control steps are as follows:

[0090] Set the initial trajectory and expected force to initialize the PPO reinforcement learning algorithm.

[0091] Estimate the stiffness and position of the environment, adjust the robot's reference motion trajectory, and calculate the force error by moving the robot along the reference trajectory. , Expressing expectation, represents the measured ambient contact force.

[0092] Select Action ( 、 Adjustment amount of control parameters) , Indicates that the action selection follows a high standard deviation distribution, calculate the robot through the impedance equation 、 The adjustment amount of the control parameter is then substituted into the equation Calculate joint displacement instructions , the robot follows the joint displacement instruction sports;

[0093] Calculate current 、 Correspondingly, the reward value is obtained according to the reward function, the state of the robot end is obtained, and the , update the actor and critic neural networks in the PPO reinforcement learning model every n steps, and finally evaluate the average reward and performance of the training model. represents the state space of reinforcement learning at the current moment, Represents the action space of reinforcement learning at the current moment, Represents the reward value of reinforcement learning at the current moment, Represents the state space of the reinforcement learning model at the next moment. The goal of the training is to meet the goal of the policy network structure converging to a stable state. Figure 6 As shown in the figure, in this embodiment, it is the reinforcement learning reward value after 100 reinforcement learning trainings. After the PPO reinforcement learning training model converges, the trained model can be directly used as a controller for constant force control of grinding and polishing. In the actual experiment, the parameters used by the robot are as follows: the grinding wheel speed is 2000rpm, the expected force is 20N, the feed speed is 0.035m / s, and the maximum threshold of the grinding and polishing force is 30N. Figure 7 As shown in Figure 1, the contact force variation diagram of the inclined workpiece and the curved workpiece after reinforcement learning training is used. CAC (Constant Admittance Control) represents constant impedance control, and R-AC (Reinforcement Learning Applied in Admittance Control) represents the method proposed in this article.

[0094] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A robot constant force control polishing method based on PPO reinforcement learning, characterized in that: The method comprises the following steps: S1: for a three-dimensional model or a point cloud model of a workpiece to be processed, obtaining an original grinding and polishing trajectory for processing the workpiece to be processed; S2 uses impedance control to perform constant force grinding control, and uses this to construct an impedance controller with unknown parameters and corresponding constraints; The impedance controller is performed according to the following formula: in, is the coefficient of inertia, is the damping coefficient, is the stiffness coefficient, is the actual position of the robot end With the expected position The error, and They are The first and second derivatives of is the expected contact force Actual contact force The error, is the proportionality coefficient, is the integration coefficient; S3 calculates the environmental stiffness and position of the robot end in real time, calculates the robot normal control instruction using the calculated environmental stiffness and position, and adjusts the normal displacement of the original trajectory in real time according to the normal control instruction so that the actual grinding and polishing force is equal to the preset expected grinding and polishing force; S4: solving unknown parameters in the impedance controller to determine the impedance controller, and performing constant force grinding on the robot according to the impedance controller.

2. A robot constant force control polishing method based on PPO reinforcement learning as claimed in claim 1, characterized in that: In step S2, the constraint conditions are obtained according to the following steps: S21 calculates the initial stiffness and damping of the grinding environment, where the environment is the grinding tool and workpiece as a whole; S22 constructs the constraints using the obtained initial stiffness and damping.

3. A robot constant force control polishing method based on PPO reinforcement learning as claimed in claim 2, characterized in that: In step S21, the initial damping and initial stiffness are calculated according to the following formula: in, , is the parameter obtained by iteratively using the recursive augmented least squares method, , are the identified initial damping and initial stiffness, It is a time period.

4. A robot constant force control polishing method based on PPO reinforcement learning as claimed in claim 2, characterized in that: In step S22, the constraints are as follows: in, , , is the coefficient in the impedance equation, , is an intermediate variable, is the environmental stiffness.

5. The robot constant force control polishing method based on PPO reinforcement learning as claimed in claim 1, characterized in that: In step S3, the normal control instructions of the robot are carried out as follows: in, is the desired position, It is the power of expectation, is the estimated environmental position, is the estimated ambient stiffness.

6. A robot constant force control polishing method based on PPO reinforcement learning as described in claim 1 or 5, characterized in that: The environmental stiffness and position are as follows: in, is the initial ambient stiffness, is the initial environment position, is the measured contact force, is the estimated ambient stiffness, yes The estimated environmental position at each moment, yes The estimated environmental position at each moment, yes The robot's end position at the moment, is the robot's movement time, , , is a constant and satisfies , is the differential of the original trajectory generated by the robot.

7. The robot constant force control polishing method based on PPO reinforcement learning as claimed in claim 1, characterized in that: In step S4, the unknown parameters are solved using a reinforcement learning method.

8. A robot constant force control polishing method based on PPO reinforcement learning as claimed in claim 7, characterized in that: The reinforcement learning method is performed according to the following steps: S41 Construct the reward function for reinforcement learning and set the action space and state space; S42 uses the parameters of the state space as input and the parameters of the action space as output to construct a reinforcement learning strategy neural network, assigns initial values ​​to the unknown parameters, and stops training when the training reward value converges stably. The current corresponding unknown parameter value is the required unknown parameter value, where the action space is , the state space is , is the proportionality coefficient, is the integration coefficient; is the expected contact force Actual contact force The error, is the robot terminal velocity.

9. A robot constant force control polishing method based on PPO reinforcement learning as claimed in claim 8, characterized in that: In step S41, the reward function is performed as follows: The action space is carried out according to the following formula: The state space is carried out according to the following formula: in, is the expected contact force Actual contact force The error, is the robot terminal velocity, It is a positive number.

Citation Information

Patent Citations

  • Grinding constant force control method based on deep reinforcement learning PPO algorithm

    CN114660940A

  • Robot grinding control method based on model

    CN116909141A