Robot tail end contact force control method and system based on reinforcement learning

Through reinforcement learning training of the robot strategy model and adjusting the learning rate in combination with reward trends, the problem that robots are difficult to adapt to the dynamic environment in traditional methods is solved, efficient and adaptive contact force control is achieved, and the control accuracy and stability of the robots in complex tasks is improved.

CN120244976APending Publication Date: 2025-07-04HEFEI UNIV OF TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510575798.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional robot contact force control methods are difficult to adapt to dynamic environmental changes, resulting in a decrease in control performance and safety, and the large model has poor applicability among different robots.

Method used

A multi-task optimization framework that uses the robot end contact force control method based on reinforcement learning is adopted, and a robot strategy model is trained, and learning rate is adjusted in combination with reward trends to realize adaptive force control and optimize position, attitude and force control.

Benefits of technology

It improves the control accuracy and stability of the robot in complex tasks, can quickly adapt to environmental changes, automatically adjust control strategies, and improves the robot's ability to adapt to the environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244976A_ABST
    Figure CN120244976A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent robots and force control, in particular to a robot tail end contact force control method and system based on reinforcement learning. The robot strategy model is trained through reinforcement learning, and actions are decided according to the state of the robot; the states comprise the current state and the target state, and the current state comprises the joint angle, the joint speed, the tail end position, the posture and the tail end contact force; the target state comprises a target position, a target attitude and a target contact force; in the model training process, the learning rate is adjusted in real time according to reward feedback, and if the average reward declines, the learning rate is reduced; and if the average reward is in a rising trend for multiple consecutive rounds, the learning rate is increased. The robot strategy model provided by the invention can quickly adapt to the dynamic change of the environment and realize efficient self-adaptive force control, so that the control strategy is quickly, accurately and automatically adjusted, the control precision of the contact force at the tail end of the robot is improved, and the control precision and stability of the robot in complex tasks are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent robots and force control, and in particular to a method and system for controlling the contact force at the end of a robot based on reinforcement learning. Background Art

[0002] With the development of robot technology, multi-degree-of-freedom robots are widely used in industrial automation, medical treatment, assembly and other fields. In actual tasks, robots often need to contact the environment, and at this time, the interaction force between the robot and the external environment needs to be considered. Traditional position control is no longer sufficient to meet the requirements, and force control needs to be adopted to ensure accuracy and safety.

[0003] In traditional contact force control methods, impedance control has received extensive attention because of its simple implementation, intuitive design based on the dynamic model, and ability to adapt to the requirements of various contact tasks. Impedance control adjusts the impedance parameters so that the robot can exhibit the desired dynamic characteristics in different contact environments. However, a significant disadvantage of impedance control is that these impedance parameters need to be manually adjusted, which usually depends on the operator's experience and understanding of the environment. In a dynamically changing environment, the efficiency and accuracy of manual parameter adjustment will decrease significantly, resulting in the inability to guarantee the control performance and safety of the robot when performing tasks.

[0004] In addition, some scholars have proposed to train a large model through machine learning to make decisions about the movement of the robot. However, the physical characteristics of different robots make it difficult for the large model trained on a fixed data set to be applicable to different robots; and it is also difficult for the pre-trained large model to achieve the expected control accuracy.

[0005] Therefore, how to achieve adaptive force control according to the dynamic changes of the environment, automatically adjust the impedance parameters or control strategies, and improve the adaptability of the robot to environmental uncertainties is a technical problem. Summary of the Invention

[0006] In order to overcome the defect that the robot in the above-mentioned prior art is difficult to adapt to environmental uncertainties, the present invention proposes a method for controlling the contact force at the end of a robot based on reinforcement learning, solves the problem of controlling the contact force at the end in a dynamic environment, and improves the stability and environmental adaptability of the robot in complex tasks.

[0007] To achieve the above object, the present invention adopts the following technical solutions, including:

[0008] A method for controlling the end contact force of a robot based on reinforcement learning proposed by the present invention first trains a robot policy model based on a reinforcement learning algorithm, which decides actions according to the state of the robot; the state includes the current state and the target state, and the current state includes: joint angle, joint velocity, end position, attitude and end contact force; the target state includes: target position, target attitude and target contact force.

[0009] During the training process of the robot policy model, the learning rate of the robot policy model is adjusted in combination with the reward trend; the learning rate adjustment method is: after each round of decision-making, calculate the average reward of the most recent N rounds.

[0010] If the average reward decreases, the learning rate is decreased.

[0011] If the average reward shows an upward trend for multiple consecutive rounds, the learning rate is increased.

[0012] Preferably, the optimization objective of the robot policy model is: minimizing the error between the end contact force and the target contact force.

[0013] The decision constraint of the robot policy model is: minimizing the error between the robot end position and the target position, and keeping the vertical state when the robot end contacts the contact surface.

[0014] Preferably, the reward is the weighted sum of the position reward, the attitude reward and the control force reward, and the weights of the three decrease in turn.

[0015] Preferably, the attitude reward reward o The calculation formula is:

[0016]

[0017] Q error =Q T Q d

[0018] where arccos is a trigonometric function, max is a function to take the maximum value, and Q error is the error matrix between the current rotation matrix Q and the target rotation matrix Q d .

[0019] Preferably, the position reward is the opposite of the unknown error, and the control force reward is the opposite of the end contact force error.

[0020] Preferably, when the average reward decreases, the learning rate α is updated to A×α, A<1; when the average reward remains unchanged or increases for 3 consecutive moments or more, the learning rate α is updated to B×α, 1<B<2.

[0021] Preferably, A = 0.9 and B = 1.1.

[0022] Preferably, the robot policy model is first trained on the initialized experience pool. After the training is completed, it is used to make decisions on actions based on the robot state. The robot executes actions in the current state to achieve the target state;

[0023] During the decision-making process of the robot policy model, the motion samples composed of the robot state, actions, observation states, and rewards are added to the experience pool; during the decision-making process, the learning rate is continuously adjusted, and the robot policy model is continuously updated based on the latest experience pool and learning rate.

[0024] Preferably, the initialized experience pool is obtained by combining the simulation of the robot dynamics model. The robot dynamics model is:

[0025]

[0026] where, θ ∈ R n×1 is the robot joint position vector, is the derivative of, is the derivative of θ; M(θ) ∈ R n×n is the robot inertia matrix; is the robot centripetal force and Coriolis force vector term, G(θ) ∈ R n×1 is the robot gravity term vector, is the robot joint friction force vector; τ ∈ R n×1 is the robot joint torque; f c ∈ R n×n is the robot Coulomb friction coefficient, f v ∈ R n×n is the robot viscous friction coefficient, f b ∈ R n×1 is the robot friction offset value; sgn is a function.

[0027] A robot end-effector contact force control system based on reinforcement learning proposed by the present invention includes a memory and a processor. A computer program is stored in the memory, and the processor is connected to the memory. The processor is used to execute the computer program to implement the above-mentioned robot end-effector contact force control method based on reinforcement learning.

[0028] The advantages of the present invention are as follows:

[0029] In the model training process of the robot end-effector contact force control method proposed by the present invention, the learning rate is adjusted in real time according to the reward feedback, so that the robot policy model can quickly adapt to the dynamic changes of the environment, realize efficient adaptive force control, and thus quickly and accurately automatically adjust the control strategy, improve the robot end-effector contact force control accuracy, and thereby improve the control accuracy and stability of the robot in complex tasks.

[0030] Based on the robot dynamics model, the present invention transforms the problem of end-effector contact force control into a multi-task optimization framework that combines position control, attitude control, and force control, and adjusts the joint torques of the robot through a reinforcement learning algorithm. By optimizing the control strategy, accurate control of the end-effector contact force can be achieved on the basis of ensuring the accuracy of trajectory tracking and attitude control.

[0031] The reward design of the present invention can achieve coordinated control of position and attitude while controlling the end-effector contact force. However, since the rewards for position, attitude, and force have different numerical ranges and dimensions, directly adding them may cause the terms with larger numerical values to contribute too much to the reward function, making it impossible to optimize other objectives and difficult to achieve balanced optimization. Therefore, each reward term is weighted and summed through weights to normalize its numerical scale, ensuring that the contributions of the three objectives of position, attitude, and force to the reward function are balanced, thereby stably and efficiently optimizing all control objectives and meeting the requirements of the task.

[0032] When the reinforcement learning algorithm of the present invention controls the end-effector contact force, it avoids the problem of manually adjusting impedance parameters in the traditional impedance control model. The end-effector contact force control method of the present invention can adaptively learn and adjust the control strategy according to the actual environment and task requirements, improving the adaptability and robustness of the robot in a dynamic environment.

[0033] Compared with the traditional impedance control method, the reinforcement learning algorithm adopted by the present invention improves the adaptability of the robot to the environment, enabling it to autonomously adjust the control strategy in a changing dynamic environment. Brief Description of the Drawings

[0034] Figure 1 is a flow chart of the robot end-effector contact force control method based on reinforcement learning proposed by the present invention;

[0035] Figure 2 is a schematic diagram of the result of the end-effector contact force in the embodiment;

[0036] Figure 3 is a schematic diagram of the error result of the end-effector contact force in the embodiment;

[0037] Figure 4 is a comparison chart of the reward increase curves of the variable learning rate PPO algorithm and the traditional PPO algorithm in the present invention;

[0038] Figure 5 is the logical structure diagram of the present invention; where P d , R d , F d are the desired position, desired attitude, and desired force, and τ f is the disturbance;

[0039] Figure 6 Flow chart for training and application of the robot decision-making model in the embodiments. Detailed implementation manners

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0041] Referring to Figure 1 , a method for controlling the end contact force of a robot based on reinforcement learning: First, train a robot policy model based on the reinforcement learning algorithm, which makes decisions on actions according to the state of the robot; the state includes the current state and the target state, and the current state includes: joint angle, joint velocity, end position, attitude, and end contact force; the target state includes: target position, target attitude, and target contact force; the robot executes actions in the current state to achieve the target state;

[0042] The robot policy model drives the movement of the robot, and the robot policy model is continuously updated during the decision-making process.

[0043] During the training and updating of the robot policy model, the learning rate of the robot policy model is adjusted in combination with the reward trend;

[0044] The learning rate adjustment method is: after each round of decision-making, calculate the average reward of the most recent N rounds

[0045] If the average reward decreases, then decrease the learning rate;

[0046] If the average reward shows an upward trend for multiple consecutive rounds, then increase the learning rate.

[0047] In the present invention, a robot policy model is constructed based on the robot dynamics model, and the robot dynamics model includes variables such as robot joint angles, joint velocities, and joint torques.

[0048] The dynamic model of an n-degree-of-freedom robot is established as follows:

[0049]

[0050] where, θ ∈ R n×1 is the robot joint position vector, is the derivative of, is the derivative of θ; M(θ) ∈ R n×n is the robot inertia matrix; is the centripetal force and Coriolis force vector terms of the robot, and G(θ) ∈ R n×1 is the vector of the robot's gravity term is the vector of the robot's joint friction force; τ ∈ R n×1 is the robot's joint torque; f c ∈ R n×n is the Coulomb friction coefficient of the robot, f v ∈ R n×n is the viscous friction coefficient of the robot, f b ∈ R n×1 is the friction offset value of the robot

[0051] sgn is a function where x is an algebraic expression

[0052] That is

[0053] is the derivative of θ

[0054] During the decision-making process, the robot's policy model is constrained by the auxiliary task to execute the main task

[0055] The main task is the force control task, and its ultimate goal is to achieve precise regulation of the end contact force, minimize the error of the end contact force, and ensure that the contact force between the robot's end effector and the target object meets the predetermined goal

[0056] The auxiliary tasks include the position control task and the attitude control task

[0057] The goal of the position control task is to ensure that the robot's end moves along the desired trajectory, that is, to minimize the error between the robot's end position and the target position, so as to achieve force control on the target trajectory

[0058] The goal of the attitude control task is to ensure that the robot's end is perpendicular when contacting the contact surface, so as to improve the force control accuracy

[0059] That is, the actions output by the robot's policy model during the decision-making process need to meet the constraints of the auxiliary task to achieve precise regulation of the end contact force

[0060] In the present invention, the PPO algorithm is used to construct the robot's policy model, which makes decisions based on the state, constructs rewards based on observations, and adjusts the learning rate according to the rewards

[0061] The state space is denoted as S = {s1, s2, s3.... s t ...}, where s tDenote the set of states of each robot at the $t$-th moment. The robot state includes the current joint angles, joint velocities, end effector positions, end effector postures, end effector contact forces, as well as target positions, target postures, and target contact forces. Among them, the current joint angles $\theta$ of the robot belong to $\mathbb{R}$. n×1 and joint velocities The end effector position $P$ belongs to $\mathbb{R}$. 3×1 and end effector posture, end effector contact force $F$ belonging to $\mathbb{R}$ constitute the current state for describing the current motion of the robot. The target position $P$ d belongs to $\mathbb{R}$. 3×1 The target posture $Q$ d belongs to $\mathbb{R}$. 3×3 and the target contact force $F$ d belonging to $\mathbb{R}$ constitute the target state for describing the robot's task objective.

[0062] Among them, the end effector posture is represented by the rotation matrix $Q$ obtained through forward kinematics, which belongs to $\mathbb{R}$. 3×3 and satisfies the orthogonality constraint $Q$ T $Q^TQ = I$, where $I$ represents the identity matrix. The end effector contact force $F$ belonging to $\mathbb{R}$ is the force along the normal direction of the contact surface. The target position, target posture, and target contact force are set according to the task requirements.

[0063] The action space is denoted as $A=\{a_1,a_2,a_3,\cdots,a$ t $\cdots\}$, where $a$ t belongs to $\mathbb{R}$. n×1 Here, $a$ t represents the control command at the $t$-th moment, which is specifically manifested as the control torques of each joint of the robot, and its magnitude is limited by the maximum output of the robot's motors. The essence of the action is to control the motion of the robot by adjusting the torques of each joint.

[0064] The observation space is denoted as $O=\{o_1,o_2,o_3,\cdots,o$ t $\cdots\}$, where $o$ t represents the set of states and errors of the robot at the $t$-th moment, including the current joint angles, joint velocities, end effector positions, end effector postures, end effector contact forces, as well as position errors, posture errors, and contact force errors, providing a decision-making basis for the PPO algorithm to optimize the policy. The position error is the error between the target position at the current moment and the current end effector position, the posture error is the error between the target posture at the current moment and the current posture, and the contact force error is the error between the target contact force at the current moment and the current end effector contact force.

[0065] The reward space is denoted as $R=\{r_1,r_2,r_3,\cdots,r$ t $\cdots\}$, where $r$ tIt represents the learning feedback of the robot at the t-th moment, that is, the reward; the reward is calculated based on the position error, attitude error, and contact force error at the end of the robot, and is used to drive the PPO algorithm to optimize the control strategy.

[0066] The reward is weighted for the reward functions under three tasks to ensure that each task can be optimized. The reward at a certain moment is calculated according to the following formula:

[0067] reward = 7 * reward d + 1 * reward o + 0.01 * reward f

[0068] Among them, reward d represents the position reward, reward o represents the attitude reward, reward f represents the control force reward.

[0069]

[0070] Among them, x e , y e respectively represent the actual positions of the end of the robot in the xy plane, x d , y d respectively represent the target positions of the end of the robot in the xy plane.

[0071]

[0072] Q error = Q T Q d

[0073] Among them, arccos is a trigonometric function, max is a function to take the maximum value, and Q error is the error matrix between the current rotation matrix Q and the target rotation matrix Q d .

[0074]

[0075] Among them, F represents the actual contact force between the end of the robot and the contact surface, and F d represents the target contact force between the end of the robot and the contact surface.

[0076] In this embodiment, the PPO algorithm is used to achieve force control under the multi-task optimization framework. A suitable state space, action space, observation space, and reward function are designed. The PPO algorithm selects actions according to task requirements, and adjusts the torques of each joint of the robot through the reward function to achieve the desired end contact force control. The reward function is optimized according to the task objective to drive the robot to learn the optimal strategy.

[0077] The maximum value of the reward is 0. The PPO algorithm maximizes the reward, reduces the position error, attitude error, and force error, continuously optimizes the strategy, and tracks the target contact force.

[0078] In this embodiment, when using the PPO algorithm to optimize the strategy, a learning rate adjustment mechanism based on the reward trend is introduced to adapt to the learning needs of different training stages and ensure the convergence speed and stability.

[0079] Referring to Figure 6 , the robot end contact force control method based on reinforcement learning proposed in this embodiment specifically includes the following steps.

[0080] S1. Construct a robot dynamics model and a robot policy model. The robot policy model makes decisions on actions according to the state of the robot;

[0081] S2. Initialize the robot policy model and the learning rate; the initial value range of the learning rate α is 1e-4 < α < 1e-3;

[0082] S3. After training the robot policy model on the experience pool, during the training process, update the learning rate according to the reward; the experience pool samples are denoted as {s, a, o, r}, where s is the robot state at the current moment, a is the decision action at the current moment, o is the observation state of the robot after executing action a in state s, and r is the reward at the current moment; the initial experience pool can be obtained by simulating the robot dynamics model.

[0083] The update rule of the learning rate is: calculate the average reward of the robot in the most recent N moments at each moment t

[0084] If the reward drops and the strategy may be unstable, then reduce the learning rate α and update α to 0.9×α;

[0085] If then judge whether the average reward in the most recent K moments continues to rise or remain unchanged. If so, judge that the reward in the most recent K times shows a stable upward trend and the learning rate needs to be increased. The learning rate α can be updated to 1.1×α;

[0086] In other cases, the learning rate remains unchanged.

[0087] Represents the average reward at the previous moment of time t.

[0088] S4. Make a decision on the action through the converged robot policy model, load the decision action on the robot dynamics model, generate the data sample {s, a, o, r} at this moment according to the execution result of the robot action, and add it to the experience pool; where s is the robot state at the current moment, a is the decision action at the current moment, o is the observed state after the robot executes action a in state s, and r is the reward at the current moment.

[0089] S5. Continuously update the learning rate. Every T moments, combined with the latest learning rate, let the robot policy model learn and update on the experience pool.

[0090] The above robot end contact force control method based on the reinforcement learning algorithm is simulated and verified in combination with specific embodiments below.

[0091] In this embodiment, a 6-degree-of-freedom robot is used, that is, n = 6; the initial learning rate α = 0.5×10 -3 .

[0092] In this embodiment, K = 4 is set, that is, during the training process of the robot policy model:

[0093] If Let α be updated to 0.9×α;

[0094] If Let the learning rate α be updated to 1.1×α;

[0095] In other cases, the learning rate α remains unchanged.

[0096] In this embodiment, a robot policy model is constructed based on the reinforcement learning algorithm PPO. According to the dynamic changes of the environment, the control strategy is automatically adjusted, and the principle of realizing the adaptive control of the robot end contact force is as Figure 5 shown.

[0097] In this embodiment, the robot end contact force control method based on the reinforcement learning algorithm provided by the present invention (referred to as the method of the present invention) and the existing impedance model-based method are respectively used for end contact force control, and the target contact force at the robot end is used as the reference value.

[0098] Finally, the control curves and error curves of the end contact force by the method of the present invention and the impedance model method are as Figure 2 and Figure 3As shown, it can be seen that the control effects of the method of the present invention and the impedance model method are not ideal within the first 0.5 s. However, as time progresses, both the method of the present invention and the impedance model method can track the reference value, and the end contact force control error of the method of the present invention is smaller, with higher accuracy. The reward increase curves of the variable learning rate PPO algorithm and the traditional PPO algorithm of the present invention are as Figure 4 shown. It can be seen that the variable learning rate PPO algorithm can converge faster and the training process is more stable. It can be seen that the method of the present invention can better control the end contact force of the robot.

[0099] Of course, for those skilled in the art, the present invention is not limited to the details of the above exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0100] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0101] The technologies, shapes, and structures not detailedly described in the present invention are all well-known technologies.

Claims

1. A method for controlling the end - effector contact force of a robot based on reinforcement learning, characterized in that, First, train a robot policy model based on a reinforcement learning algorithm, which decides actions according to the state of the robot; the state includes the current state and the target state, and the current state includes: joint angles, joint velocities, end - effector positions, postures, and end - effector contact forces; the target state includes: target positions, target postures, and target contact forces. During the training process of the robot policy model, adjust the learning rate of the robot policy model in combination with the reward trend. The learning rate adjustment method is as follows: After each round of decision-making, calculate the average reward in the most recent N rounds If the average reward decreases, then decrease the learning rate. If the average reward shows an upward trend for multiple consecutive rounds, then increase the learning rate.

2. The method for controlling the end-effector contact force of a robot based on reinforcement learning according to claim 1, wherein The optimization objective of the robot policy model is: to minimize the error between the end - effector contact force and the target contact force. The decision constraint of the robot policy model is: to minimize the error between the robot end - effector position and the target position, and to keep the robot end - effector perpendicular when it contacts the contact surface.

3. The method for controlling the end contact force of a robot based on reinforcement learning according to claim 2, wherein, The reward is the weighted sum of the position reward, the posture reward, and the control - force reward, and the weights of the three decrease in turn.

4. The method for controlling the end contact force of a robot based on reinforcement learning according to claim 3, wherein Attitude reward o The calculation formula is as follows: Q error = Q T Q d Among them, arccos is a trigonometric function, max is a function for taking the maximum value, and Q error is the error matrix between the current rotation matrix Q and the target rotation matrix Q d ​ 5. The method for controlling the end - effector contact force of a robot based on reinforcement learning according to claim 3, wherein, The position reward is the opposite of the unknown error, and the control - force reward is the opposite of the end - effector contact - force error.

6. The method for controlling the end - effector contact force of a robot based on reinforcement learning according to claim 1, wherein, When the average reward decreases, update the learning rate α to A×α, where A < 1; when the average reward remains unchanged or increases for 3 or more consecutive moments, update the learning rate α to B×α, where 1 < B < 2.

7. The method for controlling the end contact force of a robot based on reinforcement learning according to claim 6, wherein, A = 0.9, B = 1.

1.

8. The method for controlling the end - effector contact force of a robot based on reinforcement learning according to any one of claims 1 - 7, characterized in that, The robot policy model is first trained on the initialized experience pool. After training, it is used to decide actions based on the robot state. The robot executes actions in the current state to achieve the target state. During the decision - making process of the robot policy model, add the motion samples composed of the robot state, action, observation state, and reward to the experience pool; continuously adjust the learning rate during the decision - making process, and continuously update the robot policy model based on the latest experience pool and learning rate.

9. The method for controlling the end - effector contact force of a robot based on reinforcement learning according to claim 7, wherein, The initialized experience pool is obtained by simulating the robot dynamics model, and the robot dynamics model is: where θ ∈ R n×1 is the robot joint position vector, is the derivative of, and is the derivative of θ; M(θ) ∈ R n×n is the robot inertia matrix; is the robot centripetal force and Coriolis force vector term, G(θ) ∈ R n×1 is the robot gravity term vector, is the robot joint friction force vector; τ ∈ R n×1 is the robot joint torque; f c ∈ R n×n is the robot Coulomb friction coefficient, f v ∈ R n×n is the robot viscous friction coefficient, f b ∈ R n×1 is the robot friction offset value; sgn is a function.

10. A robot end-effector contact force control system based on reinforcement learning, characterized in that, It includes a memory and a processor. The memory stores a computer program, and the processor is connected to the memory. The processor is used to execute the computer program to implement the robot end - effector contact - force control method based on reinforcement learning as described in any one of claims 1 - 9.

Citation Information

Patent Citations

  • Target detection network FPGA acceleration method based on sparse coding improvement

    CN116720558A

  • Bone grinding reinforcement learning system and method based on near-end strategy optimization

    CN119871420A

  • Off-line learning for robot control using a reward prediction model

    US20230256593A1