A control system and method for mechanical arm contact force zero steady state error tracking

By combining reinforcement learning algorithms and admittance controllers, the learning rate and loss function are optimized to achieve accurate tracking of contact forces in robotic arms. This solves the steady-state error problem in the contact force tracking process of robotic arms and is suitable for remote operations in industrial and service sectors.

CN118809599BActive Publication Date: 2026-02-24HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410977389.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-02-24
Estimated Expiration
2044-07-22

AI Technical Summary

Technical Problem

Existing robotic arms suffer from steady-state errors during contact force tracking, and their control systems and strategies are imperfect, making it impossible to effectively achieve precise matching between the expected contact force and the actual contact force.

Method used

By employing a reinforcement learning algorithm combined with an admittance controller, a mapping relationship is established between the error between the actual contact force and the expected contact force and the expected position adjustment amount. The learning rate is optimized using a learning rate scheduler, and the loss function is optimized by combining a deep deterministic policy gradient algorithm and a compensation term, thereby achieving precise adjustment of the end effector position of the robotic arm.

Benefits of technology

It achieves a steady-state error of 0 in the contact force between the robotic arm and the environment, is simple to operate, has high safety, and is suitable for remote operation in hazardous environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118809599B_ABST
    Figure CN118809599B_ABST
Patent Text Reader

Abstract

The application provides a control system for mechanical arm contact force zero steady-state error tracking, comprising: a collection device, a control device and an execution device; wherein the control device adopts a mobility controller fused with a reinforcement learning algorithm; the reinforcement learning algorithm obtains an expected position adjustment amount based on the actual contact force between the mechanical arm and the object, the actual position of the mechanical arm end and the force error between the actual contact force and the preset expected contact force collected by the collection device; and adjusts the preset expected position using the expected position adjustment amount; and the mobility controller outputs the updated position of the mechanical arm end according to the input force error after the expected position is adjusted. The application fuses the reinforcement learning into the mechanical arm contact force tracking system, not only ensures the rationality and accuracy of the expected position, but also realizes that the steady-state error of the system is 0, and ensures that the contact force between the mechanical arm and the object is equal to the set expected contact force.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a control system and method for tracking contact force with zero steady-state error in a robotic arm, belonging to the field of artificial intelligence technology. Background Technology

[0002] With the widespread application of robotic arms in industrial and service sectors, the demand for contact force control of robotic arms is constantly increasing. In certain tasks, precise control of the contact force between the robotic arm and external objects is crucial for safety, accuracy, and efficiency. For example, critical subsea oil pipelines require regular inspection and maintenance to maintain their normal operation and prevent leaks. In this environment, robotic arms need to perform anomaly detection and maintenance tasks based on contact force control. Furthermore, during assembly processes, robotic arms need to grasp and install parts with appropriate contact force to avoid damage or assembly errors. Therefore, developing technologies capable of contact force tracking has become a focus of research and industry attention.

[0003] Since the robotic arm itself can be regarded as a PID control system, there is a steady-state error in the process of tracking the desired contact force. At the same time, realizing contact force tracking requires comprehensive consideration of the dynamics, mechanical characteristics and control strategy of the robotic arm. Current control algorithms may not be able to fully solve the problem of complex contact force tracking. It can be seen that designing an effective control strategy suitable for different tasks and environments is a challenging task. The existing technology for tracking the desired contact force between the robotic arm and the environment has the following defects: (1) the steady-state error of tracking the desired contact force is not 0 or is too large; (2) the control system and strategy are imperfect. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a control system and method for tracking the contact force of a robotic arm with zero steady-state error, which can achieve a steady-state error of 0 in the control system and ensure that the contact force between the robotic arm and the environment is equal to the preset expected contact force.

[0005] To achieve the above objectives, the present invention is implemented using the following technical solution:

[0006] On one hand, the present invention provides a control system for tracking zero steady-state error of contact force of a robotic arm, comprising: a data acquisition device for acquiring the actual contact force between the robotic arm end effector and the environment and the actual position of the robotic arm end effector; a control device for acquiring an updated position of the robotic arm end effector based on the acquired actual contact force and actual position of the robotic arm end effector, as well as a preset desired contact force and desired position; and an execution device for receiving the updated position of the robotic arm end effector and manipulating the robotic arm to a corresponding state based on the updated position of the robotic arm end effector. The control device includes an admittance controller incorporating a reinforcement learning algorithm; the reinforcement learning algorithm takes the force error between the actual contact force and the preset desired contact force and the actual position of the robotic arm end effector as input, and obtains a desired position adjustment amount as output; the preset desired position is adjusted using the desired position adjustment amount, and after the desired position is adjusted, the admittance controller outputs the updated position of the robotic arm end effector based on the input force error.

[0007] Furthermore, the optimizer of the reinforcement learning algorithm includes a learning rate scheduler, which is used to establish the correlation between the rate of change of contact force and the learning rate.

[0008] Furthermore, the correlation between the rate of change of the contact force and the learning rate includes: the learning rate The expression is:

[0009] ;

[0010] in, It is the initial learning rate. It refers to adjusting parameters; The rate of change of contact force;

[0011] The rate of change of the contact force The expression is:

[0012] ;

[0013] in, and These are the actual contact forces at the current time step and the previous time step, respectively. This indicates the time taken to complete the current time step.

[0014] Furthermore, the loss function of the reinforcement learning algorithm includes a compensation term, the expression of which is:

[0015] ;

[0016] in, It is the entropy of the action output by the Actor network. It is the policy function of the Actor network; It is the square of the norm of the Actor network spatial parameters. It is a constant term, and s is the state. It is a function of the action with respect to the state.

[0017] Furthermore, the parameters of the loss function The update method is as follows:

[0018] ;

[0019] in, and These are the parameters in the loss functions for the next time step and the current time step, respectively. It is a small positive constant used to ensure that the divisor is not zero. and These are the first-order moment estimates and second-order moment estimates after gradient correction, respectively.

[0020] Furthermore, the reinforcement learning algorithm employs a deep deterministic policy gradient algorithm.

[0021] Furthermore, the control system is built in Simulink.

[0022] Furthermore, the robotic arm includes a seven-degree-of-freedom robotic arm.

[0023] Furthermore, the execution device includes a computer running the Ubuntu system; the computer is equipped with an inverse kinematics algorithm.

[0024] On the other hand, the present invention provides a control method for tracking zero steady-state error of contact force in a robotic arm, comprising:

[0025] Collect the actual contact force between the robotic arm's end effector and the environment, and the actual position of the robotic arm's end effector;

[0026] The updated position of the robotic arm end is obtained based on the collected actual contact force and actual position of the robotic arm end, as well as the preset expected contact force and expected position.

[0027] The updated position of the robotic arm's end effector is sent to the execution device, so that the execution device controls the robotic arm to the corresponding state according to the updated position of the robotic arm's end effector.

[0028] The step of obtaining the updated position of the robotic arm end effector based on the collected actual contact force and actual position of the robotic arm end effector, as well as the preset desired contact force and desired position, includes:

[0029] The force error between the actual contact force and the preset expected contact force, and the actual position of the robotic arm end effector are used as inputs to the reinforcement learning algorithm. The expected position adjustment amount is obtained as the output of the reinforcement learning algorithm. The preset expected position is adjusted using the expected position adjustment amount. After the expected position is adjusted, the updated position of the robotic arm end effector is output according to the input force error.

[0030] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0031] (1) The control system provided by the present invention adopts an admittance control system that integrates reinforcement learning algorithm. The reinforcement learning algorithm establishes a mapping relationship between the error between the actual contact force and the expected contact force and the expected position adjustment amount. The expected position adjustment amount is used to adjust the preset expected position to make it more reasonable. The position of the end of the robotic arm is gradually adjusted to the optimal position so that the actual contact force between the robotic arm and the object is equal to the expected contact force. The steady-state error of the system is 0.

[0032] (2) The reinforcement learning algorithm used in this invention establishes the correlation between the rate of change of contact force and the learning rate by designing a learning rate scheduler in the optimizer. The learning rate realizes the influence of the rate of change of contact force on the parameter update of the loss function, so that the obtained expected position adjustment value is more accurate.

[0033] (3) The control method provided by the present invention only requires the operator to sit in front of the control device and set the desired contact force. The robotic arm can operate remotely in dangerous environments, which is simple, safe and convenient. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the structure of Embodiment 1 of the control system for zero steady-state error tracking of contact force of the robotic arm according to the present invention;

[0035] Figure 2 This is the control flowchart in Embodiment 4 of the control method for tracking zero steady-state error of contact force of the robotic arm according to the present invention. Detailed Implementation

[0036] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention. Example 1

[0037] like Figure 1As shown in Embodiment 1, a control system for tracking zero steady-state error of contact force in a robotic arm is provided. This system includes: a host computer, a slave computer, and a force sensor installed at the end of the Baxter robotic arm, all connected to the same local area network. The host computer is an improved admittance control system incorporating reinforcement learning algorithms, built in Simulink. The slave computer is a laptop running Ubuntu and equipped with inverse kinematics algorithms. The slave computer controls the robotic arm by performing kinematic modeling and inverse kinematics algorithm design on the slave computer. Data transmission between the slave and host computers is via UDP. The force sensor is installed at the end of the robotic arm and communicates with the slave computer via USB.

[0038] In existing admittance control systems, both the model admittance parameters and environmental stiffness are positive semi-definite diagonal matrices, which do not introduce coupling between models. Therefore, the admittance control model of the robotic arm can be considered decoupled in each direction. This also facilitates formula derivation. Thus, when considering only the robotic arm subjected to a force in a single spatial direction, the admittance control formula can be simply regarded as:

[0039] ;

[0040] in, , , These represent the acceleration, velocity, and position of the robotic arm's end effector in one direction, respectively, output by the admittance control system. , , These represent the preset expected acceleration, velocity, and position of the robotic arm end effector in one direction, respectively.

[0041] In this embodiment, the reinforcement learning algorithm takes the force error between the actual contact force and the preset expected contact force, and the actual position of the robotic arm end effector as input, and the expected position adjustment amount as output, to perform reinforcement learning training. Each training session involves the robotic arm making contact with objects in the environment, and observing and learning the error between the actual contact force and the expected contact force in a stable state. This continues until the training results converge, obtaining the learned model parameters and the expected position adjustment value. Then, the expected position adjustment amount is used to adjust the preset expected acceleration, velocity, and position of the robotic arm end effector in one direction in the existing admittance controller. The adjusted expected position is used to obtain the output acceleration, velocity, and position of the robotic arm end effector in one direction, which is then used as the updated position of the robotic arm end effector. Example 2

[0042] Based on Example 1, this embodiment sets the reinforcement learning algorithm as the Deep Deterministic Policy Gradient Algorithm, namely the DDPG algorithm.

[0043] The limitations of the adaptive learning rate and local gradient information inherent in the traditional Adam algorithm can easily lead to the agent getting trapped in local optima. Furthermore, the actual contact force in this invention varies significantly and rapidly during the initial training phase, which may prevent the learning rate from adjusting in time to adapt to the rate and magnitude of these changes. Therefore, this embodiment aims to avoid the reinforcement learning process from getting trapped in local optima by optimizing the learning rate.

[0044] In this embodiment, the optimizer of the DDPG algorithm includes a learning rate scheduler. The learning rate scheduler is used to establish the correlation between the rate of change of contact force and the learning rate, specifically including:

[0045] Learning rate The expression is:

[0046] ;

[0047] in, It is the initial learning rate. It refers to adjusting parameters; The rate of change of contact force;

[0048] The rate of change of the contact force The expression is:

[0049] ;

[0050] in, and These are the actual contact forces at the current time step and the previous time step, respectively. This indicates the time taken to complete the current time step. Example 3

[0051] This embodiment, based on embodiment 2, adds a compensation term to the loss function of the DDPG algorithm. The expression for the compensation term is:

[0052] ;

[0053] in, It is the entropy of the action output by the Actor network. It is the policy function of the Actor network; It is the square of the norm of the Actor network spatial parameters. It is a constant term, and s is the state. It is a function of the action with respect to the state.

[0054] In this embodiment, the parameters of the loss function The update method is as follows:

[0055] ;

[0056] in, and These are the parameters in the loss functions for the next time step and the current time step, respectively. It is a small positive constant used to ensure that the divisor is not zero. and These are the first-order moment estimates and second-order moment estimates after gradient correction, respectively.

[0057] Example 4

[0058] like Figure 2 As shown, this embodiment provides a control method for tracking zero steady-state error of contact force of a robotic arm, including: a host computer collecting the actual position of the end effector of the Baxter robotic arm and the actual contact force between it and the environment; obtaining an updated position of the end effector of the Baxter robotic arm based on the collected actual contact force and actual position, as well as preset desired contact force and desired position; and sending the updated position of the end effector of the Baxter robotic arm to a slave computer, so that the slave computer controls the Baxter robotic arm to the corresponding state based on the updated position of the end effector of the Baxter robotic arm.

[0059] The process of obtaining the updated position of the Baxter robotic arm end effector based on the collected actual contact force and actual position, as well as the preset expected contact force and expected position, includes: using the force error between the actual contact force and the preset expected contact force and the actual position of the Baxter robotic arm end effector as input to the reinforcement learning algorithm, obtaining the output quantity of the reinforcement learning algorithm—the expected position adjustment quantity, and using the expected position adjustment quantity to adjust the preset expected position; after the expected position is adjusted, the updated position of the Baxter robotic arm end effector is output based on the input force error.

[0060] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A control system for tracking contact force with zero steady-state error in a robotic arm, characterized in that, include: The data acquisition device is used to collect the actual contact force between the end effector of the robotic arm and the environment, as well as the actual position of the end effector of the robotic arm. A control device is used to obtain the updated position of the robotic arm end based on the collected actual contact force and actual position of the robotic arm end, as well as the preset desired contact force and desired position. An actuator is used to receive the updated position of the robotic arm's end effector and to control the robotic arm to a corresponding state based on the updated position of the robotic arm's end effector. The control device includes an admittance controller that incorporates a reinforcement learning algorithm; The reinforcement learning algorithm takes the force error between the actual contact force and the preset expected contact force and the actual position of the robotic arm end as input, and obtains the expected position adjustment amount as output; the preset expected position is adjusted using the expected position adjustment amount, and after the expected position is adjusted, the admittance controller outputs the updated position of the robotic arm end based on the input force error. The optimizer of the reinforcement learning algorithm includes a learning rate scheduler, which is used to establish the correlation between the rate of change of contact force and the learning rate. The relationship between the rate of change of contact force and the learning rate includes: the learning rate The expression is: ; in, It is the initial learning rate. It refers to adjusting parameters; The rate of change of contact force; The rate of change of the contact force The expression is: ; in, and These are the actual contact forces at the current time step and the previous time step, respectively. This indicates the time taken to complete the current time step; The loss function of the reinforcement learning algorithm includes a compensation term, the expression of which is: ; in, It is the entropy of the action output by the Actor network. It is the policy function of the Actor network; It is the square of the norm of the Actor network spatial parameters. It is a constant term, and s is the state. It is a function of the action with respect to the state; The parameters of the loss function The update method is as follows: ; in, and These are the parameters in the loss functions for the next time step and the current time step, respectively. It is a small positive constant used to ensure that the divisor is not zero. and These are the first-order moment estimates and second-order moment estimates after gradient correction, respectively. The reinforcement learning algorithm employs a deep deterministic policy gradient algorithm.

2. The control system as described in claim 1, characterized in that, The control system was built in Simulink.

3. The control system as described in claim 1, characterized in that, The robotic arm includes a seven-degree-of-freedom robotic arm.

4. The control system as described in claim 1, characterized in that, The execution device includes a computer running the Ubuntu system; the computer is equipped with an inverse kinematics algorithm.

5. A control method for tracking zero steady-state error of contact force in a robotic arm, applied to the control system described in any one of claims 1-4, characterized in that, include: Collect the actual contact force between the robotic arm's end effector and the environment, and the actual position of the robotic arm's end effector; The updated position of the robotic arm end is obtained based on the collected actual contact force and actual position of the robotic arm end, as well as the preset expected contact force and expected position. The updated position of the robotic arm's end effector is sent to the execution device, so that the execution device controls the robotic arm to the corresponding state according to the updated position of the robotic arm's end effector. The step of obtaining the updated position of the robotic arm end effector based on the collected actual contact force and actual position of the robotic arm end effector, as well as the preset desired contact force and desired position, includes: The force error between the actual contact force and the preset expected contact force, and the actual position of the robotic arm end effector are used as inputs to the reinforcement learning algorithm. The expected position adjustment amount is obtained as the output of the reinforcement learning algorithm. The preset expected position is adjusted using the expected position adjustment amount. After the expected position is adjusted, the updated position of the robotic arm end effector is output according to the input force error.

Citation Information

Patent Citations

  • Tracking differentiator-based cooperative control method for tail end position of mechanical arm and force-variable impedance

    CN116079739A

  • Implementation method and device of machine learning model

    CN116225687A