An Adaptive Impedance Control Method and System for a Robotic Arm Based on Inner-Loop Performance Feedback and Energy Tank

By combining the inner and outer loop collaborative control framework with the energy tank, the problem of unstable control force of the robotic arm in complex environments is solved, realizing compliant and stable control of the robotic arm in unknown or variable stiffness environments, and improving the safety and stability of human-machine collaboration.

CN121374658BActive Publication Date: 2026-03-06SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511973932.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-06
Estimated Expiration
2045-12-25

AI Technical Summary

Technical Problem

Existing adaptive impedance control methods for robotic arms struggle to maintain stable control force in complex environments, lacking inner-loop performance feedback and energy constraints, leading to oscillations or joint saturation and posing safety risks.

Method used

By adopting an inner and outer loop collaborative control framework, combined with reinforcement learning and an energy tank, data is acquired through force/torque sensors and pose sensors to construct a state vector, define a reward function, optimize impedance parameters, and generate actual control torque through a sliding mode controller, thereby achieving compliant interactive control of the robotic arm.

Benefits of technology

It improves training efficiency and control stability, prevents energy overshoot, ensures that the robotic arm maintains a compliant and low-vibration response in unknown or variable stiffness environments, and enhances the safety and stability of human-robot collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121374658B_ABST
    Figure CN121374658B_ABST
Patent Text Reader

Abstract

This invention discloses an adaptive impedance control method and system for a robotic arm based on inner-loop performance feedback and an energy tank. It constructs a collaborative impedance control framework with inner and outer loops: the outer loop employs a reinforcement learning-based impedance controller, while the inner loop uses a sliding mode controller. The dynamic performance indicators of the inner-loop sliding mode controller are fed back to the outer-loop reinforcement learning agent in real time, guiding the adaptive adjustment of impedance parameters. Simultaneously, force / torque sensors detect the robotic arm's interactive power, and an energy tank mechanism is used to constrain the energy output of the reinforcement learning, preventing energy overshoot and system instability during training and execution. This method achieves force / position hybrid control performance while maintaining the robotic arm's compliance, safety, and energy stability in complex or unknown environments, improving the robustness and control accuracy of the robotic arm's interaction with the environment, and forming a comprehensive control closed loop with dynamic performance constraints, energy constraints, and adaptive learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot intelligent control and energy constraint control, and particularly relates to an adaptive impedance control method and system for a robotic arm based on inner loop performance feedback and an energy tank. Background Technology

[0002] Existing active compliance control methods for collaborative robotic arms mainly fall into three categories: traditional impedance control, adaptive impedance control, and reinforcement learning-based impedance control. However, these methods suffer from the following drawbacks when facing real-world human-robot collaboration:

[0003] 1. In compliant human-machine contact tasks, the environment is complex, time-varying and uncertain. Therefore, in complex environments, fixed impedance parameters and manually set control laws are difficult to maintain stable control force and cannot achieve compliant control under human-machine collaboration.

[0004] 2. Existing impedance control methods based on reinforcement learning often set the inner loop controller as an ideal unit without considering the performance of the true inner loop (such as the amplitude of the sliding surface chattering, the error convergence speed, etc.). Furthermore, in the early training stage of reinforcement learning, the policy has not converged and is not combined with feedback from real physical sensors, lacking physical energy constraints. This results in an incomplete state space for the reinforcement learning agent, which affects training efficiency, reduces the policy convergence speed, and may generate control outputs that exceed the capabilities of the inner loop, causing the robotic arm to oscillate or the joints to saturate, posing potential safety risks.

[0005] 3. Adaptive impedance control and reinforcement learning-based methods can achieve dynamic parameter adjustment, but reinforcement learning lacks physical constraints and inner-loop feedback in the early stages of training, often resulting in energy overshoot or policy non-convergence. The traditional energy tank method can constrain energy input and output to ensure system stability, but it is not combined with reinforcement learning or adaptive control and also lacks inner-loop performance feedback. Summary of the Invention

[0006] Purpose of the Invention: This invention provides an adaptive impedance control method and system for a robotic arm based on inner-loop performance feedback and an energy tank. It aims to address the problems of existing technologies, such as neglecting inner-loop performance, lack of integration with real physical sensor feedback, poor environmental adaptability, energy overshoot, or strategy non-convergence. By detecting the performance indicators of the inner-loop sliding mode controller and adding force / torque sensors, the inner-loop performance and real physical state of the robotic arm are introduced into the state space of reinforcement learning. While adjusting impedance parameters through reinforcement learning, the system output is constrained in parallel, thereby achieving high safety and compliant interactive control of the robotic arm in complex environments.

[0007] Technical solution: This invention provides an adaptive impedance control method for a robotic arm based on inner-loop performance feedback and an energy tank, comprising:

[0008] A collaborative control framework for the inner and outer loops of a robotic arm and an energy tank are constructed. The collaborative control framework includes a reinforcement learning agent, an outer loop, and an inner loop. The outer loop uses an impedance controller, and the inner loop uses a sliding mode controller.

[0009] The system acquires force / torque sensor data and pose sensor data from the end effector during the joint movement of the robotic arm, and simultaneously acquires the tracking error and sliding surface variables of the inner loop of the sliding mode controller. Based on the force / torque sensor data and pose sensor data, the system calculates the interaction power between the robotic arm and the external environment, and updates the energy state of the energy tank based on the interaction power. The system then normalizes the sensor data, sliding surface variables, and energy state of the energy tank to construct a state vector characterizing the current operating state of the robotic arm.

[0010] A reward function is defined for the reinforcement learning agent. The current state vector of the robotic arm is used as the current state space of the reinforcement learning agent. The agent's policy network generates agent actions based on the current state space. The agent actions include impedance parameters and desired contact force. The reward value of the agent actions is calculated according to the reward function. The reward value is used to optimize the policy network parameters of the reinforcement learning agent. Finally, the modulated agent actions are output.

[0011] The modulated agent motion input impedance controller generates the desired joint torque;

[0012] The desired joint torque is converted into the actual control torque used to control the movement of the robotic arm joints by a sliding mode controller.

[0013] Furthermore, the acquisition of force / torque sensor data and pose sensor data of the end effector during the joint movement of the robotic arm includes:

[0014] Based on a six-dimensional force / torque sensor and a position / velocity sensor of the robotic arm's end effector, the interaction force between the robotic arm and the environment is measured in real time. and torque End linear velocity With angular velocity ;

[0015] The interactive power of the robotic arm is calculated based on force, torque, linear velocity, and angular velocity. The formula is:

[0016] ;

[0017] The acquisition of the tracking error and sliding surface variables of the inner loop of the sliding mode controller includes:

[0018] The differential rate of the sliding surface of the sliding mode controller at adjacent sampling points is taken as the chattering speed of the sliding surface. The formula is:

[0019] ;

[0020] in, Let k be the difference rate at time k. The difference rate at time k-1 The interval time, Let be the chattering speed of the sliding surface at time k.

[0021] Furthermore, updating the energy state of the energy tank includes:

[0022] Set contact force threshold When the magnitude of the applied force exceeds the contact force threshold At that time, based on the interaction power Update the energy tank's energy status The formula is:

[0023] ;

[0024] in, To control the cycle, Let k be the energy state of the energy tank at time k+1. Let be the interaction power at time k;

[0025] When the magnitude of the applied force does not exceed the contact force threshold At that time, the energy state of the energy tank remains unchanged.

[0026] Furthermore, the normalization process for the sensor data, sliding surface variables, and energy state of the energy tank, to construct a state vector characterizing the current operating state of the robotic arm, includes:

[0027] The chattering velocity of the sliding surface is exponentially smoothed to obtain the smoothed performance feedback value. The formula is:

[0028] ;

[0029] in, For smoothing coefficients;

[0030] Performance feedback value The inner loop performance index data were obtained through normalization. The formula is:

[0031] ;

[0032] in, The performance feedback value is taken as the 5th percentile of the training / experimental data during offline calibration. The performance feedback value is taken as the 95th percentile of the training / experimental data during offline calibration;

[0033] Calculate the actual position x of the robotic arm's end effector and the preset desired position. Positional error between Calculate the actual speed of the robotic arm's end effector. With the preset expected speed speed error between The actual contact force at the end is obtained from the six-dimensional force / torque sensor. Obtain the normalized energy state of the energy tank. and normalized inner loop performance index data ;

[0034] Based on position error Speed ​​error Actual contact force at the end Energy state of the energy tank and inner ring performance index data Obtain the current state vector of the robotic arm. .

[0035] Furthermore, the modulated agent action output specifically includes:

[0036] Define the reward function for the reinforcement learning agent as follows:

[0037] ;

[0038] in, As a reward value;

[0039] The formula for location tracking reward is: Where x is the actual position of the robotic arm's end effector. The desired position is preset for the end effector of the robotic arm;

[0040] To track rewards, the formula is: ,in To enhance the expected contact force calculated by the learning agent, This represents the actual contact force at the end.

[0041] The formula for the inner loop smoothing performance bonus is: ;in, For inner ring performance index data;

[0042] For energy-constrained rewards, the formula is:

[0043] ;

[0044] in, The preset safety threshold, For interactive power, This represents the lower limit threshold for the energy tank's storage capacity. The energy state of the energy tank;

[0045] for The weighting coefficients, for The weighting coefficients, for Weighting coefficients;

[0046] A reinforcement learning agent is constructed using the Deep Deterministic Policy Gradient (DDPG) algorithm, which includes an Actor policy network and a Critic value network.

[0047] Input the current state vector of the robotic arm To the state space of the reinforcement learning agent;

[0048] Actor Policy Network Output continuous actions based on the current state. The formula is:

[0049] ;

[0050] in, For end stiffness parameters, For damping parameters, Desired contact force; , For impedance parameters;

[0051] action After execution, calculate the reward value. Obtain the next state vector of the robotic arm. And construct the ancestor of experience ;

[0052] experience tuples The Critic value network stores the data in the experience replay pool by minimizing the temporal difference error. Update network parameters The formula is:

[0053] ;

[0054] in, For the Critic value network at the current parameters Next state and actions Value assessment, For the target Critic value network under the current parameters Next state and actions The formula for assessing the value of [is] is:

[0055] ;

[0056] in, As a discount factor, and These are the target Critic value network and the target Actor policy network, used to train the Critic value network and the Actor policy network, respectively, where E is the mathematical expectation;

[0057] The output of the Critic value network is the expected cumulative reward of the action output by the Actor policy network. The formula is:

[0058] ;

[0059] in, The cumulative reward represents the total revenue earned from the current time t to the end time T. As a discount factor, As a reward value;

[0060] The Actor policy network updates its policy gradients based on the Critic value network, updating the network parameters. The formula is:

[0061] ;

[0062] The action output by the Actor policy network Through energy modulation function Modulation is performed to obtain the modulated impedance parameters and the desired contact force. That is, the modulated agent action, the formula is:

[0063] ;

[0064] The modulation function formula is:

[0065] ;

[0066] Where T represents the energy state of the energy tank. , This is the smoothing constant.

[0067] Furthermore, the step of generating the desired joint torque from the modulated agent action input impedance controller includes: converting the modulated parameters... The input impedance controller generates the desired interaction force at the end, as shown in the formula:

[0068] ;

[0069] in, This represents the actual position of the robotic arm's end effector. The desired position of the robotic arm's end effector. This represents the actual speed at the end effector of the robotic arm. The desired speed at the end effector of the robotic arm;

[0070] Using the Jacobian matrix of the robotic arm Will Mapped to desired joint torque The formula is:

[0071] ,

[0072] in, This is the actual joint position vector.

[0073] Furthermore, the conversion of the desired joint torque into the actual control torque for controlling the movement of the robotic arm joints via the sliding mode controller includes:

[0074] Define joint position error as Constructing the sliding surface The formula is:

[0075] ;

[0076] in, This is the actual joint position vector. To pass the joint position error The joint velocity error is obtained by differentiating with respect to time t, where t represents time t. The approximation coefficient;

[0077] Based on sliding surface To achieve the desired joint torque The input is the actual control torque used to drive the joint. The formula is:

[0078] ;

[0079] in: For linear feedback gain matrix, For sliding mode switching gain matrix, It is a saturation function. These are the boundary layer parameters.

[0080] This invention also provides an adaptive impedance control system for a robotic arm based on inner-loop performance feedback and an energy tank, comprising:

[0081] The framework construction module is used to build the inner and outer loop collaborative control framework and energy tank of the robotic arm. The inner and outer loop collaborative control framework includes a reinforcement learning agent, an outer loop, and an inner loop. The outer loop adopts an impedance controller, and the inner loop adopts a sliding mode controller.

[0082] The data acquisition module is used to acquire force / torque sensor data and pose sensor data of the end effector during the joint movement of the robotic arm, and simultaneously acquire the tracking error and sliding surface variables of the inner loop of the sliding mode controller; calculate the interaction power between the robotic arm and the external environment based on the force / torque sensor data and pose sensor data, and update the energy state of the energy tank based on the interaction power; normalize the sensor data, sliding surface variables and energy state of the energy tank to construct a state vector characterizing the current operating state of the robotic arm;

[0083] The reinforcement learning module is used to define the reward function of the reinforcement learning agent. The current state vector of the robotic arm is used as the current state space of the reinforcement learning agent. The agent's policy network generates agent actions based on the current state space. The agent actions include impedance parameters and expected contact force. The reward value of the agent actions is calculated according to the reward function, and the reward value is used to optimize the policy network parameters of the reinforcement learning. Finally, the modulated agent actions are output.

[0084] The desired torque generation module is used to generate the desired joint torque from the modulated intelligent agent motion input impedance controller;

[0085] The actual torque conversion module is used to convert the desired joint torque into the actual control torque for controlling the movement of the robotic arm joints via a sliding mode controller.

[0086] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0087] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method.

[0088] Beneficial effects: Compared with the prior art, the present invention has the following outstanding advantages:

[0089] 1. This invention quantifies the sliding mode chattering rate and introduces it into the reinforcement learning state space, so that the policy output matches the actual execution capability, thereby improving training efficiency and control stability.

[0090] 2. The system uses an energy tank to monitor the interactive power in real time, and embeds the energy tank into the reinforcement learning state and reward function. Combined with the energy modulation function, the control action is subjected to physical-level hard constraints to prevent energy overshoot and end force over-force.

[0091] 3. Construct a collaborative architecture of outer-loop reinforcement learning, inner-loop sliding mode control, and energy closed-loop feedback to enable the robotic arm to maintain a compliant, stable, and low-jitter response in unknown or variable stiffness environments. Attached Figure Description

[0092] Figure 1 This is a flowchart illustrating the implementation of the control method of the present invention.

[0093] Figure 2 This is a flowchart of the reinforcement learning training process based on the DDPG algorithm.

[0094] Figure 3 This is a flowchart for calculating the performance indicators of the inner ring.

[0095] Figure 4 This is a flowchart of the energy tank's workflow. Detailed Implementation

[0096] like Figure 1 As shown, the adaptive impedance control method for a robotic arm based on inner-loop performance feedback and an energy tank, as described in this invention, includes:

[0097] S1. Construct a robotic arm inner and outer loop collaborative control framework and an energy tank. The inner and outer loop collaborative control framework includes a reinforcement learning agent, an outer loop, and an inner loop. The outer loop uses an impedance controller, and the inner loop uses a sliding mode controller.

[0098] S2. Acquire force / torque sensor data and pose sensor data of the end effector during the joint movement of the robotic arm, and simultaneously acquire the tracking error and sliding surface variables of the inner loop of the sliding mode controller;

[0099] S3. Determine whether the robotic arm is in contact with the external surface: If not, keep the energy state of the energy tank unchanged; if it is in contact, calculate the interaction power between the robotic arm and the external environment based on the force / torque sensor data and the pose sensor data, and update the energy state of the energy tank based on the interaction power.

[0100] S4. Normalize the sensor data, sliding surface variables, and energy state of the energy tank to construct a state vector that characterizes the current operating state of the robotic arm.

[0101] S5. Establish the reward function for the reinforcement learning agent and input the current state vector of the robotic arm into the reinforcement learning agent;

[0102] S6. After the policy network of the reinforcement learning agent infers, the impedance parameters are updated, and the policy network is updated according to the reward value calculated by the reward function, and the impedance parameters are modulated.

[0103] S7. Input the updated impedance parameters into the impedance controller to generate the desired joint torque;

[0104] S8. Convert the desired joint torque into the actual control torque using a sliding mode controller;

[0105] S9. Apply the actual control torque to the robotic arm execution layer to control the movement of the robotic arm joints;

[0106] S10. Determine whether the current joint movement of the robotic arm has met the expectations. If it has, end the control; otherwise, return to S2 and update the sensor data of the robotic arm and the inner loop performance index data of the sliding mode controller.

[0107] In this embodiment,

[0108] S1 specifically includes:

[0109] S1-1 Constructing the dynamics and kinematics model of the robotic arm;

[0110] Before initiating real-time control, the geometric and kinematic model of the robotic arm is established by acquiring its DH parameters, and the dynamic equations are derived from the Lagrange equations:

[0111] ;

[0112] in, The inertia matrix, For the Coriolis and centrifugal terms, For gravity compensation, To control the input torque, For external disturbance torque, Joint angle, The joint angular velocity vector. Let be the joint angular acceleration vector. Based on the dynamic equations, the joint space equations are transformed into Cartesian space to obtain a system description of the force-position coupling at the end effector, providing a foundation for subsequent impedance control and energy modeling.

[0113] S1-2 establishes a collaborative control framework for the robotic arm's inner and outer loops, including a reinforcement learning agent, an outer loop, and an inner loop. The outer loop uses an impedance controller to generate the desired joint torque based on the impedance parameters. The inner loop uses a sliding mode controller to generate the actual control torque based on the desired joint torque. The reinforcement learning agent is used to adaptively update the impedance parameters. The energy tank uses the energy state of the energy tank to constrain the policy network of the reinforcement learning agent.

[0114] This invention simplifies the relationship between the end effector of the robotic arm and the assembly plane into a spring-mass-damping model, the mathematical model of which is as follows:

[0115] ;

[0116] in, Indicates external contact force. Let be the desired inertia matrix of the impedance controller. Let be the desired stiffness matrix of the impedance controller. Let be the desired damping matrix of the impedance controller. This represents the actual position of the robotic arm's end effector. This represents the actual speed at the end effector of the robotic arm. This represents the actual acceleration at the end effector of the robotic arm. The desired position of the robotic arm's end effector. The desired speed at the end effector of the robotic arm. The desired acceleration at the end effector of the robotic arm; selected in the impedance controller. As a control variable Set to a fixed value of 1.

[0117] like Figure 3 As shown, to ensure that the outer loop impedance output can track quickly and stably at the execution layer, the inner loop uses a sliding mode control law to control the joint space. Let the inner loop sliding surface be... In classic form:

[0118] ;

[0119] in This refers to the joint position error (or an appropriate trajectory error). express time, This is the actual joint position vector. For the desired joint trajectory (desired position). The joint velocity error is obtained by differentiating the joint position error. As the approach coefficient, this control structure can effectively suppress joint chattering and ensure trajectory tracking accuracy.

[0120] S1-3 Initialize the energy state of the energy tank , The initial energy reference for the system, ranging from 0.5 to 0.8, represents the proportion of available energy when the control system starts up, and sets the upper limit of the energy storage capacity of the energy tank after normalization. lower limit Normalization can eliminate differences between equipment and environment, making the energy state of the energy tank comparable and controllable under different experimental conditions.

[0121] S2 specifically includes:

[0122] S2-1 acquires sensor data from the end effector of the robotic arm during joint movement;

[0123] Based on a six-dimensional force / torque sensor and a position / velocity sensor at the end effector of the robotic arm, the interaction force between the robotic arm and the environment is measured in real time. and torque End linear velocity With angular velocity ;

[0124] The interactive power of the robotic arm is calculated based on force, torque, linear velocity, and angular velocity. The formula is:

[0125] ;

[0126] The interactive power of the robotic arm is input to the energy tank T in real time to describe the dynamic state of the system's energy release and consumption. When the robotic arm comes into contact with or collides with a high-rigidity environment, The energy state of the energy tank increases rapidly. Approaching the upper threshold causes the impedance parameter The controller automatically reduces the output stiffness K or increases the damping D to absorb some energy and avoid oscillations, thus achieving closed-loop regulation of energy constraint and compliant control.

[0127] S2-2 Obtain the tracking error and sliding surface variables of the inner loop of the sliding mode controller;

[0128] The differential rate of the sliding surface of the sliding mode controller at adjacent sampling points is taken as the chattering speed of the sliding surface. The approximate value is given by the formula:

[0129] ;

[0130] in, Let k be the difference rate at time k. The difference rate at time k-1 The interval time, Let be the chattering speed of the sliding surface at time k.

[0131] like Figure 4 As shown, S3 specifically includes:

[0132] Set contact force threshold When the magnitude of the force Exceeding the contact force threshold At that time, based on the interaction power The formula for updating the energy tank's energy state is:

[0133] ;

[0134] in, To control the cycle (e.g., 1ms or 5ms). To update the energy state of the energy tank at time k+1, Let be the interaction power at time k;

[0135] When the magnitude of the applied force does not exceed the contact force threshold hour( The energy state of the energy tank remains unchanged, and the impedance control parameters are directly transmitted without affecting the control output.

[0136] The energy tank only begins to update and adjust the control output after physical contact is made at the end of the robotic arm; it retains its initial value during the free space phase. The impedance parameters output by the reinforcement learning are directly transmitted to the controller. When the robotic arm's end effector begins to make contact with the environment, the energy state variable T of the energy tank is input to the reinforcement learning module in real time to characterize the system's current energy level.

[0137] like Figure 3 As shown, S4 specifically includes:

[0138] S4-1. Perform exponential smoothing on the chatter speed to obtain the smoothed performance feedback value. The formula is:

[0139] ;

[0140] in, For smoothing coefficients, For example, 0.9 uses exponential smoothing to suppress high-frequency noise and improve the stability of the index;

[0141] S4-2. Obtain the inner loop performance index data by normalizing the performance feedback values. The formula is:

[0142] ;

[0143] in, The performance feedback value is taken as the 5th percentile of the training / experimental data during offline calibration. The performance feedback value is taken as the 95th percentile of the training / experimental data during offline calibration;

[0144] Normalized upper and lower limits , It can be set according to system calibration or experience to limit the performance feedback value Φ[k] at time k to the range of [0,1], which is convenient for reinforcement learning agents to process;

[0145] In addition, the advantages of normalization settings are that they have simple computational requirements, respond quickly when the robotic arm interacts with the outside world, are sensitive to chattering, and can quickly reflect the performance of the inner loop; exponential smoothing can remove noise and has low latency; performance indicators are monitored and quantified in real time by the system and used as feedback signals to be input to the upper-level reinforcement learning module to describe the control stability and dynamic performance of the current system.

[0146] S4-3. Calculate the actual position x of the robotic arm's end effector and the preset desired position. Positional error between Calculate the actual speed of the robotic arm's end effector. With the preset expected speed speed error between The actual contact force at the end is obtained from the six-dimensional force / torque sensor. Obtain the normalized energy state of the energy tank. and normalized inner loop performance index ;

[0147] Position error Speed ​​error Actual contact force at the end Energy state of the energy tank and inner ring performance indicators Combined into the current state vector of the robotic arm .

[0148] S5 specifically includes:

[0149] S5-1. Define the state space and action space;

[0150] The state space of a reinforcement learning agent is defined as follows: Action space is defined as ;in For end stiffness parameters, For damping parameters, These correspond one-to-one with the desired stiffness matrix and the desired damping matrix in the impedance model, respectively. For the desired contact force, A deterministic policy network parameterized by an Actor network;

[0151] S5-2, Design the reward function;

[0152] reward function Provided at each time step for immediate action evaluation. The advantages and disadvantages of [the system] are comprehensively considered, taking into account position accuracy, force control error, and energy safety constraints. The formula is:

[0153] ;

[0154] in, Let be the reward value at time t;

[0155] The formula for location tracking reward is: Where x is the actual position of the robotic arm's end effector. The desired position is preset for the end effector of the robotic arm; the position tracking reward encourages the end effector to accurately track the desired trajectory, and imposes a greater penalty for large position errors;

[0156] To track rewards, the formula is: ,in To enhance the expected contact force calculated by the learning agent, The actual contact force at the end point; force tracking reward ensures precise force control and is a core performance indicator in interactive tasks;

[0157] The formula for the inner loop smoothing performance bonus is: ;in, The data represents the inner loop performance indicators; when the sliding mode chattering speed is high, a penalty is imposed to guide the agent to generate a smoother, lower-jitter action strategy.

[0158] For energy-constrained rewards, the formula is:

[0159] ;

[0160] in, The preset safety threshold; For interactive power, This represents the lower limit threshold for the energy tank's storage capacity. The energy state of the energy tank; when the system interactive power Exceeding the safety threshold Penalties are incurred when the energy tank's reserves fall below a threshold. The system generates penalties to guide the reinforcement learning agent to proactively reduce stiffness or increase damping to ensure safety.

[0161] like Figure 2 As shown, this invention employs the Deep Deterministic Policy Gradient (DDPG) algorithm to achieve smooth, stable, and adaptive adjustment of impedance parameters in a continuous action space. It introduces the real-time state of the energy tank as the input to the policy network and embeds an energy constraint term into the reward function, enabling reinforcement learning to optimize control performance within the energy safety boundary, thereby achieving collaborative control of intelligence and physical safety.

[0162] By using the energy state of the energy tank and the input of the inner loop performance feedback, the agent can sense the changes in system energy and the dynamic performance of the execution layer, thereby adaptively adjusting the impedance parameters to achieve safe and stable policy output.

[0163] S6 specifically includes:

[0164] S6-1. Initialize the reinforcement learning agent environment; Critic value network within the DDPG framework. Actor policy network used to evaluate state-action values. Used to output actions.

[0165] S6-2, Input the current state vector of the robotic arm To the state space of the reinforcement learning agent, It includes information on the physical state and energy safety of the robotic arm;

[0166] ;

[0167] S6-3, Actor Policy Network Output continuous actions based on the current state. The output is the optimal impedance parameter predicted by the reinforcement learning policy in this state, but it has not yet been subjected to security modulation.

[0168] ;

[0169] in, For end stiffness parameters, For damping parameters, Desired contact force; , For impedance parameters;

[0170] S6-4, Calculate the reward;

[0171] Execute action Then, the system calculates the reward. Obtain the next state vector of the robotic arm. And construct the ancestor of experience ; Empirical tuples Stored in the experience replay pool for calculating returns on the critic value network;

[0172] S6-5, Strategy Optimization;

[0173] Critic value network minimizes temporal difference error Update network parameters The formula is:

[0174] ;

[0175] in, For the Critic value network at the current parameters Next state and actions Value assessment, For the target Critic value network under the current parameters Next state and actions The formula for assessing the value of [is] is:

[0176] ;

[0177] in, As a discount factor, and These are the target Critic value network and the target Actor policy network, used to train the Critic value network and the Actor policy network, respectively, where E is the mathematical expectation;

[0178] The output of the Critic value network is the expected cumulative reward of the action output by the Actor policy network. The cumulative reward formula is:

[0179] ;

[0180] in, The cumulative reward represents the total revenue earned from the current time t to the end time T. As a discount factor, As a reward value;

[0181] The Actor policy network updates its policy gradient based on feedback from the Critic value network to maximize cumulative rewards, gradually learning to generate better impedance parameters under different states.

[0182] The Actor policy network updates its parameters through policy gradients. The formula is:

[0183] ;

[0184] Through multiple iterations, the Actor policy network gradually learns how to output the optimal impedance parameters under energy constraints and chattering constraints during the learning process.

[0185] S6-6, Energy Modulation: When the energy stored in the energy tank (the physical state quantity of the energy tank) is too low, the control output is automatically limited to prevent instability. To achieve dynamic energy regulation, this invention designs an energy modulation function. :

[0186] ;

[0187] Energy modulation function Based on the current energy state of the energy tank Adaptive adjustment of control output intensity:

[0188] when When it is large, When the value is close to 1, the system maintains normal output.

[0189] when When approaching 0, When the output approaches zero, the controller output automatically weakens, thereby limiting the energy release rate and preventing non-passive behavior, as well as preventing instantaneous over-force and system oscillation.

[0190] The action output by the Actor policy network Through energy modulation function Modulation is performed to obtain the modulated impedance parameters and the desired contact force. That is, the modulated agent action, the formula is:

[0191] ;

[0192] The reinforcement learning agent, based on inner-loop performance feedback, introduces an end-effector force / torque sensor. This sensor, combined with the outer-loop end-effector force error and the energy state of the energy tank, adjusts the impedance parameters. Adaptive adjustments are made to output impedance parameters that meet expectations, achieving better human-machine interaction. The modulated parameters are input to the robotic arm controller, which executes joint torque commands to generate actual end-effector motion. The new end-effector force and velocity are collected by sensors and the energy state of the energy tank is updated, realizing a periodic closed loop. This ensures that the reinforcement learning strategy output is executed under physical energy safety constraints, achieving safe and compliant control.

[0193] S7 specifically includes:

[0194] modulated parameters The input impedance controller generates the desired interaction force at the end. The formula is:

[0195] ;

[0196] in, This represents the actual position of the robotic arm's end effector. The desired position of the robotic arm's end effector. This represents the actual speed at the end effector of the robotic arm. The desired speed at the end effector of the robotic arm;

[0197] Using the Jacobian matrix of the robotic arm Will Mapped to desired joint torque The formula is:

[0198] .

[0199] in, This is the actual joint position vector;

[0200] Through the aforementioned inner and outer loop collaborative control structure, rapid and stable tracking of the outer loop impedance control output is achieved at the execution layer. Using the robotic arm's inner and outer loop collaborative control framework and energy tank, the policy parameters generated by reinforcement learning no longer directly affect physical execution. Instead, they are fused with the control law through energy state modulation from the energy tank to achieve physical consistency constraints. This mechanism establishes an energy closed loop between the algorithm layer and the power layer, ensuring that the control output meets energy conservation constraints, the reinforcement learning training process is physically interpretable, the system automatically reduces output during energy stress or sudden contact changes, avoids stiffness abrupt changes, oscillations, or mechanical shocks, and achieves a balance between intelligent adjustment and physical safety.

[0201] S8 specifically includes:

[0202] The desired joint torque is used as the feedforward input to the inner loop controller. The inner loop employs a joint space sliding mode controller to perform joint tracking control, and the joint position error is defined as... Constructing the sliding surface The formula is:

[0203] ;

[0204] in, This is the actual joint position vector. To pass the joint position error The joint velocity error is obtained by differentiating with respect to time t, where t represents time t. The approximation coefficient;

[0205] Based on sliding surface To achieve the desired joint torque As input, the inner loop controller outputs the actual control torque used to drive the joint actuator. The formula is:

[0206] ;

[0207] in: For linear feedback gain matrix, For sliding mode switching gain matrix, It is a saturation function. These are boundary layer parameters used to suppress chattering.

[0208] S9 specifically includes:

[0209] The actual control torque is applied to the robotic arm's execution layer (servo driver) to control the movement of the robotic arm's joints, thereby realizing the power output of the physical layer.

[0210] S10 specifically includes:

[0211] Determine if the current joint movement of the robotic arm has achieved the desired result. If it has, end the control; otherwise, return to S2, and the force / torque sensor will collect new contact forces in real time. With power It is used to update the energy state and the next strategy.

[0212] The present invention discloses an adaptive impedance control system for a robotic arm based on inner-loop performance feedback and an energy tank, comprising:

[0213] The framework construction module is used to build the inner and outer loop collaborative control framework and energy tank of the robotic arm. The inner and outer loop collaborative control framework includes a reinforcement learning agent, an outer loop, and an inner loop. The outer loop adopts an impedance controller, and the inner loop adopts a sliding mode controller.

[0214] The data acquisition module is used to acquire force / torque sensor data and pose sensor data of the end effector during the joint movement of the robotic arm, and simultaneously acquire the tracking error and sliding surface variables of the inner loop of the sliding mode controller; calculate the interaction power between the robotic arm and the external environment based on the force / torque sensor data and pose sensor data, and update the energy state of the energy tank based on the interaction power; normalize the sensor data, sliding surface variables and energy state of the energy tank to construct a state vector characterizing the current operating state of the robotic arm;

[0215] The reinforcement learning module is used to define the reward function of the reinforcement learning agent. The current state vector of the robotic arm is used as the current state space of the reinforcement learning agent. The agent's policy network generates agent actions based on the current state space. The agent actions include impedance parameters and expected contact force. The reward value of the agent actions is calculated according to the reward function, and the reward value is used to optimize the policy network parameters of the reinforcement learning. Finally, the modulated agent actions are output.

[0216] The desired torque generation module is used to generate the desired joint torque from the modulated intelligent agent motion input impedance controller;

[0217] The actual torque conversion module is used to convert the desired joint torque into the actual control torque for controlling the movement of the robotic arm joints via a sliding mode controller.

[0218] The computer device of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above method.

[0219] The computer-readable storage medium of the present invention stores a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.

Claims

1. A method for adaptive impedance control of a robot arm based on inner loop performance feedback and energy tank, characterized in that, The method comprises the following steps: An inner-outer ring collaborative control framework and an energy tank are built, the inner-outer ring collaborative control framework comprising a reinforcement learning agent, an outer ring and an inner ring, the outer ring adopting an impedance controller and the inner ring adopting a sliding mode controller; Force / torque sensor data and pose sensor data of an end effector during joint movement of the robot arm are acquired, and tracking error and sliding surface variables of the inner ring of the sliding mode controller are acquired; interaction power between the robot arm and the external environment is calculated according to the force / torque sensor data and the pose sensor data, and the energy state of the energy tank is updated based on the interaction power; the sensor data, the sliding surface variables and the energy state of the energy tank are normalized to construct a state vector for representing the current operating state of the robot arm; A reward function of the reinforcement learning agent is defined, the current state vector of the robot arm is taken as a current state space of the reinforcement learning agent, a policy network of the agent generates an agent action based on the current state space, the agent action comprises impedance parameters and a desired contact force, a reward value of the agent action is calculated according to the reward function, the reward value is used to optimize the policy network parameters of the reinforcement learning, and finally a modulated agent action is outputted; The modulated agent action is inputted into the impedance controller to generate a desired joint torque; The desired joint torque is converted into an actual control torque for controlling the joint movement of the robot arm through the sliding mode controller; The acquisition of the force / torque sensor data and the pose sensor data of the end effector during the joint movement of the robot arm comprises: Six-dimensional force / torque sensor and position and velocity sensor based on the end effector of a robot arm, which measure the force between the robot arm and the environment in real time and torque , end linear velocity and angular velocity ; Interaction power of a robot arm is calculated based on force, torque, linear velocity, and angular velocity , the formula is: ; The acquisition of the tracking error and the sliding surface variables of the inner ring of the sliding mode controller comprises: The differential rate of the sliding mode surface of the sliding mode controller at adjacent sampling points is taken as the chattering speed of the sliding mode surface , the formula is: ; wherein, is the differential rate at time k, is the differential rate at time k-1, is the interval time, is the chattering speed of the sliding surface at time k. The updating of the energy state of the energy tank comprises: Set contact force threshold When the magnitude of the applied force exceeds the contact force threshold At that time, based on the interaction power Update the energy tank's energy status The formula is: ; wherein, is the control period, is the energy state of the energy tank at time k+1, is the interaction power at time k. When the modulus of the force does not exceed the contact force threshold the energy state of the energy tank remains unchanged; The normalization of the sensor data, the sliding surface variables and the energy state of the energy tank to construct the state vector for representing the current operating state of the robot arm comprises: The chattering speed of the sliding surface is subjected to exponential smoothing to obtain a smoothed performance feedback value , the formula is: ; wherein is a smoothing coefficient; The performance feedback value is obtained The inner loop performance index data is obtained by normalization processing The formula is: ; wherein, is the 5th percentile of the performance feedback values taken from the training / test data at the offline calibration, is the 95th percentile of the performance feedback values taken from the training / test data at the offline calibration; calculating a position error between an actual position x of the end of the robot arm and a preset desired position calculating a velocity error between an actual velocity of the end of the robot arm and a preset desired velocity acquiring an actual contact force of the end from the six-dimensional force / torque sensor acquiring an energy state of the normalized energy tank and normalized inner loop performance index data ;​​​ based on a position error , a velocity error , an actual end contact force , an energy state of an energy tank , and an inner loop performance index data , a current state vector of the robot arm is obtained .

2. The adaptive impedance control method for a robot manipulator based on inner loop performance feedback and energy tank according to claim 1, wherein, The output of the modulated agent action specifically comprises: The reward function of the reinforcement learning agent is defined, and the formula is: ; wherein, is a reward value; For position tracking reward, the formula is: ; where x is the actual position of the end of the robot arm, is the preset desired position of the end of the robot arm; For force tracking reward, the formula is: where is the expected contact force calculated by reinforcement learning agent, is the actual contact force at the end. The inner loop is rewarded for smooth performance, with the formula: ; where is the inner loop performance index data; For the energy constraint reward, the formula is: ; wherein, is a predetermined safety threshold, is an interaction power, is a lower threshold of the energy tank reserves, is an energy state of the energy tank; is a weight coefficient, is a weight coefficient, is a weight coefficient, is a weight coefficient, is a weight coefficient, is a weight coefficient; A deep deterministic policy gradient algorithm DDPG is adopted to construct the reinforcement learning agent, the reinforcement learning agent comprising an Actor policy network and a Critic value network; Input current state vector of the robot arm to the state space of the reinforcement learning agent; Actor policy network Output continuous action from current state , which is ; wherein, is an end stiffness parameter, is a damping parameter, is a desired contact force; , is an impedance parameter; Actions After execution, the reward value is calculated , the next state vector of the robot arm is obtained , and the experience tuple is constructed ; empirical tuples into an experience replay pool, the Critic value network updates its network parameters by minimizing the temporal difference error updating network parameters , which is ; wherein, is the value estimate of the Critic value network under the current parameters for the state and action , is the value estimate of the Target Critic value network under the current parameters for the state and action , and is given by: ; wherein, is a discount factor, and are a target Critic value network and a target Actor policy network, respectively, for training the Critic value network and the Actor policy network, and E is a mathematical expectation; The output of the Critic value network is the expected cumulative reward of the action output by the Actor policy network, the cumulative reward The formula is: ; wherein, is a cumulative reward, representing the total reward obtained from the current time t to the termination time T, is a discount factor, is a reward value; The Actor policy network is updated by the Critic value network according to a policy gradient, and the network parameters are updated , and the formula is: ; the action output by the actor policy network by the energy modulation function modulating, obtaining the modulated impedance parameter and the expected contact force , i.e., the modulated agent action, is ; The formula of the modulation function is: ; where T is the energy state of the energy tank, , is a smoothing constant.

3. The adaptive impedance control method for a robot manipulator based on inner loop performance feedback and energy tank according to claim 2, wherein, The inputting the modulated agent action into the impedance controller to generate a desired joint torque comprises: inputting the modulated parameter into the impedance controller to generate the desired joint torque The input impedance controller generates an end effector desired interaction force, and a formula is as follows: ; wherein, is the actual position of the end of the robot arm, is the desired position of the end of the robot arm, is the actual velocity of the end of the robot arm, is the desired velocity of the end of the robot arm; Through the robot's Jacobian matrix Will be mapped to the desired joint torques , which are given by , wherein, is the actual joint position vector.

4. The adaptive impedance control method for a robot arm based on inner loop performance feedback and energy tank according to claim 3, wherein, The conversion of the desired joint torque into the actual control torque for controlling the joint movement of the robot arm through the sliding mode controller comprises: The joint position error is defined as , a sliding surface is constructed , and the formula is: ; wherein, is the actual joint position vector, is the joint position error is the joint velocity error obtained by differentiating t with respect to t, t indicating the time t, is the approach coefficient; Based on a sliding surface In terms of desired joint torque The actual control torque for driving the joint execution is output The formula is: ; wherein: is a linear feedback gain matrix, is a sliding mode switching gain matrix, is a saturation function, is a boundary layer parameter.

5. A system for adaptive impedance control of a manipulator based on inner loop performance feedback and energy tank according to any one of claims 1-4, characterized in that, The method comprises the following steps: A framework building module is configured to build an inner-outer ring collaborative control framework and an energy tank, the inner-outer ring collaborative control framework comprising a reinforcement learning agent, an outer ring and an inner ring, the outer ring adopting an impedance controller and the inner ring adopting a sliding mode controller; A data acquisition module is configured to acquire force / torque sensor data and pose sensor data of an end effector during joint movement of the robot arm, and acquire tracking error and sliding surface variables of the inner ring of the sliding mode controller; interaction power between the robot arm and the external environment is calculated according to the force / torque sensor data and the pose sensor data, and the energy state of the energy tank is updated based on the interaction power; the sensor data, the sliding surface variables and the energy state of the energy tank are normalized to construct a state vector for representing the current operating state of the robot arm; The reinforcement learning module is configured to define a reward function of a reinforcement learning agent, take a current state vector of the robot arm as a current state space of the reinforcement learning agent, generate an agent action by a policy network of the agent based on the current state space, the agent action including impedance parameters and a desired contact force, calculate a reward value of the agent action according to the reward function, and optimize a policy network parameter of the reinforcement learning by using the reward value, and finally output a modulated agent action. The desired torque generation module is configured to input the modulated agent action into an impedance controller to generate a desired joint torque. The actual torque conversion module is configured to convert the desired joint torque into an actual control torque for controlling the joint movement of the robot arm by a sliding mode controller.

6. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the method of any one of claims 1-4 when executing the computer program.

7. A computer-readable storage medium having stored thereon a computer program, characterized in that The computer program, when executed by the processor, implements the steps of the method of any one of claims 1-4.

Citation Information

Patent Citations

  • Sliding form control method of flexible joint mechanical arm based on disturbance observer

    CN102591207A

  • Fractional order sliding mode optimization control method for flexible joint mechanical arm

    CN109683478A