A Robot Constant Force Tracking Method Based on Teaching and Learning Modes
Through the robot constant force tracking method of teaching and learning modes, combined with NURBS trajectory planning and reinforcement learning algorithm, the robot motion trajectory is optimized, and the constant force tracking problem affected by surface uncertainty in unknown environments is solved, and high-precision constant force contact operation is achieved.
Patent Information
- Application Number
- CN202310103955.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-13
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-02-13
AI Technical Summary
In an unknown environment, the uncertainty of the surface position and stiffness affects the accuracy of the robot's constant force tracking, and cannot meet the constant force tracking scenarios with high accuracy requirements.
Using a teaching and learning mode-based method, NURBS trajectory planning and position-force hybrid control, combined with impedance control and reinforcement learning algorithms, the robot motion trajectory is optimized, the end contact force error is recorded and corrected in real time, and the ε-greedy algorithm is used to select behavior and evaluate the return function to optimize the trajectory.
It improves the robot's constant force tracking accuracy in unknown environments, reduces environmental uncertainty and fitting trajectory errors, and is suitable for contact work tasks with high accuracy requirements.
Smart Images

Figure CN116107204B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot control, and particularly relates to a robot constant force tracking method based on teaching and learning modes. Background Art
[0002] At present, in most cases, the stiffness and position of the curved surface in the environment are uncertain for the proposed constant force tracking strategies, which easily affect the accuracy of the robot during constant force tracking and cannot meet the requirements of high-precision constant force tracking scenarios. In the scenario of robot contact operations with high precision requirements, it is necessary to strictly control the fluctuation of the contact force at the end of the robot.
[0003] Since the position and stiffness of the curved surface in the unknown environment are both uncertain, it is easy to affect the constant force tracking effect of the robot. Therefore, it is necessary to improve the traditional constant force tracking strategy based on compliant control. In view of the above problems, a robot constant force tracking method based on teaching and learning modes is proposed. The tracking trajectory of the end of the robot is corrected by learning to indirectly ensure the accuracy and stability of constant force tracking. Summary of the Invention
[0004] The purpose of the present invention is to provide a robot constant force tracking method based on teaching and learning modes for the existing problems, so as to solve the stability of the contact force at the end of the robot during contact operations.
[0005] The present invention is realized by the following technical solutions: A robot constant force tracking method based on teaching and learning modes, comprising the following steps:
[0006] S1. Fit the motion trajectory of the robot based on the velocity controllable NURBS trajectory planning algorithm, fuse the impedance control algorithm and the trajectory planning algorithm through the position-force hybrid control algorithm framework, and use different control strategies for different directions at the end of the robot to achieve the constant force tracking of the robot for complex curved surfaces in the unknown environment;
[0007] S2. Based on the position-force hybrid control strategy, initially perform a constant force traversal of the complex curved surface in the unknown environment, and record the position, attitude, and end contact force information of the end of the robot in real time during the traversal process. The information recorded during the traversal process of the robot on the curved surface is used as the input quantity of the reinforcement learning algorithm;
[0008] S3. Compensate the motion trajectory tracking error of the end of the robot through the difference Δf between the actual contact force and the expected contact force, select the next action through the ε-greedy algorithm, and evaluate the reward of the taken action through the reward function to optimize the motion trajectory during the constant force tracking of the robot, so that the error of the constant force tracking can be minimized.
[0009] Preferably, the S1 includes the following steps:
[0010] S101. Collect the profile points of the complex surface in the environment, calculate the NURBS trajectory passing through the profile points by the NURBS trajectory planning algorithm, and use the velocity interpolation algorithm to perform velocity planning on the fitted trajectory to fit a robot motion trajectory X with controllable velocity. nurbs ;
[0011] S102. Determine the compliant force control direction of the robot through the selection matrix, and perform position control on other directions of the robot, so that the robot can perform constant force tracking on the unknown environment. The robot motion trajectory equation based on the position-force hybrid control framework is:
[0012] X robot = H·X nurbs +(I - H)·X c
[0013] Among them, is the selection matrix, h i ∈[0, 1], I is the identity matrix, X robot is the actual motion trajectory sent to the robot, X nurbs is the trajectory fitted by the NURBS trajectory planning algorithm with controllable velocity, X c is the correction amount of the compliant control algorithm for the robot motion trajectory.
[0014] Preferably, the S2 includes the following steps:
[0015] S201. When the robot initially traverses the complex surface in the unknown environment based on the position-force hybrid control, record the actual motion trajectory X m of the robot end, the end attitude matrix R m and the end contact force F e in real time;
[0016] S202. The Q-learning algorithm is:
[0017] newQ S,A =(1 - α)Q S,A +α(R S,A +γ·maxQ′(s′, a′))
[0018] Among them, newQ S,A is the new Q value based on the state and action; Q S,A is the current Q value; R S,A is the reward based on the state and action; maxQ′(s′, a′) is the maximum future reward under the given new state and action; (1 - α)Q S,A is the proportion of the old Q value in newQ S,A ; (R S,A+γ·maxQ′(s′, a′)) is the reward brought by this action itself and the potential future rewards;
[0019] S203. Take the recorded actual motion trajectory and actual end contact force of the robot as the input quantities of the Q-learning algorithm. That is, the difference Δf between the actual contact force and the desired contact force at the end of the robot at each moment is used as the state quantity, and the position correction amount obtained by compliant control is used as the action quantity.
[0020] Preferably, the S3 includes the following steps:
[0021] S301. The ε-greedy search strategy is as follows:
[0022]
[0023] S302. After determining the action, it is necessary to evaluate the reward function R of the action taken:
[0024]
[0025] Among them, δ1 and δ2 respectively represent the weight values of the force error and the position error; f d , p d respectively represent the desired force and the desired position, f and p represent the actual contact force obtained and the actual position of the robot. The reward of the action taken is evaluated through the reward function, so that the error can be minimized.
[0026] The beneficial effects of the present invention are:
[0027] Based on the force / position hybrid control framework, constant force tracking of position-force hybrid control is realized. After traversing the complex surface in the unknown environment, the motion trajectory of the robot is optimized through the learning algorithm, reducing the problem of poor constant force tracking accuracy caused by environmental uncertainty and fitting trajectory error, so that it can be applied to the constant force contact operation task of the robot in the position environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is the structural schematic diagram of the present invention;
[0029] Figure 2 is the flow chart of the force control algorithm for Q-learning of the present invention;
[0030] Figure 3 is the schematic diagram of the actual contact force and the desired contact force between the end of the robot and the complex surface in the environment based on the teaching and learning mode of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0031] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0032] Embodiment:
[0033] Please refer to Figures 1-3 As shown, the present invention provides a technical solution: a robot constant force tracking method based on the teaching and learning mode, including the following steps:
[0034] S1. Fit the motion trajectory of the robot based on the velocity-controllable NURBS trajectory planning algorithm, fuse the impedance control algorithm and the trajectory planning algorithm through the position-force hybrid control algorithm framework, and use different control strategies for different directions of the robot end to achieve the constant force tracking of the complex surface of the unknown environment by the robot;
[0035] S2. Based on the position-force hybrid control strategy, initially perform a constant force traversal of the complex surface in the unknown environment, and record the position, attitude, and end contact force information of the robot end in real time during the traversal process. The information recorded during the robot's traversal of the surface is used as the input quantity of the reinforcement learning algorithm;
[0036] S3. Compensate the motion trajectory tracking error of the robot end through the difference Δf between the actual contact force and the desired contact force, select the next behavior through the ε-greedy algorithm, and evaluate the reward of the behavior taken through the reward function to optimize the motion trajectory during the robot's constant force tracking, so that the error of the constant force tracking can be minimized.
[0037] According to the above-mentioned robot constant force tracking method based on the teaching and learning mode, the motion trajectory of the robot for constant force tracking can be determined more accurately, and the constant force tracking effect of the complex surface in the environment can be achieved.
[0038] Obtain the control points of the surface in the environment. Generally, select the inflection points of the surface or the points with large curvature changes. Calculate the surface contour of the surface through the NURBS trajectory planning algorithm, and use it as the motion constraint trajectory of the robot end in the Cartesian space. Further, perform velocity planning on the fitted trajectory through the T-type velocity planning algorithm to ensure that the velocity of the robot end is stable and controllable. Finally, determine the fitted trajectory X of the robot end with controllable velocity nurbs ;
[0039] While controlling the Z-axis direction of the robot end through position control, the actual contact force f between the robot and the environment is obtained in real time through the six-axis force sensor at the end e , and set the contact force f between the robot end and the environmentd , the trajectory error X caused by the change of the end contact force is corrected by using the admittance control strategy c , and the trajectory of the robot end is corrected in real time, where the admittance control equation is:
[0040]
[0041] where M is the mass coefficient, generally set to 1, B is the damping coefficient, K is the stiffness coefficient, d represents the expected value, c represents the control quantity, and F e represents the input external force. The stability condition of the equation is
[0042] The fitted robot trajectory X nurbs and the trajectory correction amount X c for realizing constant force tracking based on admittance control are fused as the motion trajectory when the robot actually tracks the surface:
[0043] X robot = H·X nurbs +(I - H)·X c
[0044] where, is the selection matrix, h i ∈[0, 1], h i = 0 means that the trajectory of this dimension is controlled by force, h i = 1 means that the trajectory of this dimension is controlled by the fitted trajectory, h i ∈(0, 1) means that it is controlled by force while controlling the trajectory, I is the identity matrix, X robot is the actual motion trajectory sent to the robot, X nurbs is the trajectory fitted by the velocity-controllable NURBS trajectory planning algorithm, and X c is the correction amount of the robot motion trajectory by the compliance control algorithm.
[0045] When the robot traverses the surface in the environment through the received X robot trajectory, the motion trajectory points P of the robot end and the contact force F between the robot end and the surface at the corresponding moment are synchronously recorded in real time e .
[0046] The iterative equation of the Q-learning algorithm is:
[0047] newQ S,A = (1 - α)Q S,A + α(R S,A + γ·maxQ′(s′, a′))
[0048] where newQ S,Ais the new Q-value based on the state and action; Q S,A is the current Q-value; R S,A is the reward based on the state and action; maxQ′(s′, a′) is the maximum future reward given the new state and action; (1-α)Q S,A is the proportion of the old Q-value in newQ S,A ; (R S,A +γ·maxQ′(s′, a′)) is the reward brought by this action itself and the potential future rewards; α represents the learning rate, which defines the proportion of the old Q-value learning the new Q-value from the new Q-value. The learning rate α determines the speed at which the reinforcement learning converges to the optimal value. γ is called the discount factor, with a value range of 0 to 1, which determines the impact degree of the time distance on the return. A value of 0 means only considering short-term rewards, and a value of 1 means paying more attention to long-term rewards;
[0049] Interpolate the recorded actual contact force and the desired contact force at the end of the robot ΔF = F e -F d as the state quantity of the Q-learning algorithm, and record the actual trajectory point P at the end of the robot as the action quantity;
[0050] Balance the relationship between exploration and exploitation through the ε-greedy search strategy, explore with a probability of ε, and exploit with a probability of 1-ε. Its exploration distribution is shown as follows, and its equation is:
[0051]
[0052] After determining the behavior of the robot, evaluate the return of the action taken through the return function:
[0053]
[0054] where δ1 and δ2 represent the weights of the force error and the position error respectively. If the force tracking is dominant, δ1 can be increased; f d , p d represent the desired force and the desired position respectively. The return function R is negated in order to finally select the behavior with the maximum return, optimize the motion trajectory during the constant force tracking of the robot, and minimize the error during the constant force tracking;
[0055] Set the discount factor γ to meet the conditions and the conditions for updating the target, perform iterative learning of the Q-leaming learning algorithm, optimize the trajectory of the constant force tracking at the end of the robot, and use the optimized trajectory as the motion trajectory X of the robot during the constant force tracking new .
[0056] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0057] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A robot constant force tracking method based on the teaching and learning mode, characterized in that It includes the following steps: S1. Based on the speed-controllable NURBS trajectory planning algorithm, fit the motion trajectory of the robot. Through the position-force hybrid control algorithm framework, fuse the impedance control algorithm and the trajectory planning algorithm, and use different control strategies for different directions of the robot end to achieve the constant force tracking of the complex surface of the unknown environment by the robot; S2. Based on the position-force hybrid control strategy, initially perform a constant force traversal of the complex surface in the unknown environment, and record the position, attitude, and end contact force information of the robot end in real time during the traversal process. The information recorded during the robot's traversal of the surface is used as the input quantity of the reinforcement learning algorithm; S3. Compensate the motion trajectory tracking error of the robot end through the difference Δf between the actual contact force and the desired contact force. Select the next behavior through the ε-greedy algorithm, and evaluate the reward of the behavior taken through the reward function to optimize the motion trajectory during the constant force tracking of the robot, so that the error of the constant force tracking can be minimized.
2. The robot constant force tracking method based on the teaching and learning modality according to claim 1, wherein The S1 includes the following steps: S101. Collect the profile points of the complex surface in the environment, calculate the NURBS trajectory passing through the profile points by the NURBS trajectory planning algorithm, and use the velocity interpolation algorithm to perform velocity planning on the fitted trajectory to fit the robot motion trajectory X with controllable velocity nurbs ; S102. Determine the compliant force control direction of the robot through the selection matrix, and perform position control on other directions of the robot, so that the robot can perform constant force tracking of the unknown environment. The robot motion trajectory equation based on the position-force hybrid control framework is: X robot = H·X nurbs +(I - H)·X c Among them, is a selection matrix, h i ∈ [0, 1], I is the identity matrix, X robot is the actual motion trajectory sent to the robot, X nurbs is the trajectory fitted by the velocity-controllable NURBS trajectory planning algorithm, X c is the correction amount of the robot motion trajectory by the compliance control algorithm.
3. A robot constant force tracking method based on teaching and learning modalities according to claim 1, characterized in that, The S2 includes the following steps: S201. When the robot first traverses the complex surface in the unknown environment based on position - force hybrid control, the actual motion trajectory X of the robot end - effector is recorded in real - time m , the end - effector attitude matrix R m and the end - effector contact force F e ; S202. The Q-learning algorithm is: newQ S,A =(1 - α)Q S,A +α(R S,A +γ·maxQ′(s′, a′)) Among them, newQ S,A is the new Q-value based on the state and action; Q S,A is the current Q-value; R S,A is the reward based on the state and action; maxQ′(s′, a′) is the maximum future reward given the new state and action; (1 - α)Q S,A is the proportion of the old Q-value in newQ S,A ; (R S,A + γ·maxQ′(s′, a′)) is the reward brought by this action itself and the potential future reward; S203. Take the recorded actual motion trajectory and actual end contact force of the robot as the input quantity of the Q-learning algorithm, that is, the difference Δf between the actual contact force and the desired contact force at the robot end at each moment is used as the state quantity, and the position correction amount obtained by the compliant control is used as the behavior quantity.
4. A method for a robot's constant force tracking based on teaching and learning modalities according to claim 1, characterized in that The S3 includes the following steps: S301. The ε-greedy search strategy is: S302. After determining the behavior, it is necessary to evaluate the reward function R of the behavior taken: where δ1 and δ2 represent the weights of the force error and the position error respectively; f d , p d represent the desired force and the desired position respectively, f and p represent the actual contact force obtained and the actual position of the robot, and the reward of the behavior taken is evaluated through the reward function so that the error can be minimized.
Citation Information
Patent Citations
Method for tracking constant force surface of robot based on fuzzy iterative algorithm
CN108972545A
Robot machining work normal constant force tracking method and device
CN110948504A