Quadruped robot active compliance control method based on hierarchical reinforcement learning

By employing a hierarchical reinforcement learning-based active compliant control method, quadruped robots can effectively adapt to continuous external disturbances, solving the problems of high motion stiffness and insufficient compliance in existing technologies, and achieving improvements in stability and energy efficiency.

CN120972536APending Publication Date: 2025-11-18SUN YAT SEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511109149.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing quadruped robot control methods lack compliance with sudden disturbances, resulting in high motion stiffness, which can easily lead to hardware damage. Furthermore, they are difficult to adapt to continuous external interference and lack an effective compliant response mechanism.

Method used

An active compliant control method based on hierarchical reinforcement learning is adopted. The lower-level motion policy network estimates external forces and trunk velocity, and the upper-level compliant control policy network generates modified trunk velocity commands to achieve adaptive control against continuous disturbances.

Benefits of technology

This technology improves the stability and compliance of quadruped robots under continuous external interference, reduces the risk of hardware damage, enhances motion stability and energy efficiency, and improves human-computer interaction friendliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120972536A_ABST
    Figure CN120972536A_ABST
Patent Text Reader

Abstract

The invention relates to a quadruped robot active compliance control method based on hierarchical reinforcement learning. The method comprises a bottom-layer motion strategy network and an upper-layer compliance control strategy network, the bottom layer motion strategy network comprises a basic encoder, a force estimation head, a speed estimation head and a basic strategy network module; the basic encoder is used for outputting a low-dimensional feature vector; the force estimation head is used for estimating the current external force borne by the quadruped robot; the speed estimation head is used for estimating the current trunk speed of the quadruped robot; the basic strategy network module is used for generating a current expected joint position of the quadruped robot; and the upper-layer compliance control strategy network is used for generating a current residual speed instruction and adding the current residual speed instruction and the current trunk speed instruction to obtain a corrected current trunk speed instruction. By adopting the method, high coupling of disturbance estimation and a compliance control module can be realized, and active compliance adaptive to continuous interference can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control technology, and in particular to an active compliant control method for quadruped robots based on hierarchical reinforcement learning. Background Technology

[0002] With the rapid development of robotics technology, quadruped robots, due to their excellent terrain adaptability, have shown great potential in fields such as industrial inspection, disaster relief, and military reconnaissance. However, existing methods mostly focus on robust motion control in unstructured terrain, lacking compliance to cope with sudden disturbances. This overemphasis on robustness can easily lead to excessively high motion stiffness, causing the robot to produce high-frequency trembling movements when encountering sudden disturbances. Such violent movements can cause motor torque over-limits or even hardware damage. A safe and reliable robotic system should be able to actively adjust its posture and speed in response to external forces, rather than violently resisting them; this is what is known as compliant behavior.

[0003] Active compliant movement offers significant advantages in operational safety, energy efficiency, and human-robot interaction. This preference for compliance is prevalent in the animal kingdom. These adaptive behaviors demonstrate a delicate balance between robustness to minor perturbations and compliance to sustained perturbations. Incorporating similar mechanisms into robots can enhance their ability to cope with external disturbances, reduce the risk of injury, and improve overall performance.

[0004] To improve the compliant movement capabilities of quadruped robots, researchers have primarily employed two methods: model-based control and reinforcement learning (RL)-based approaches. However, existing methods generally suffer from the following problems: 1) Passive compliance: Most methods achieve passive compliance only through reward function design (such as torque penalty, energy efficiency optimization) or small joint PD control parameters, failing to dynamically adjust the response strategy according to the disturbance intensity. 2) Module coupling: Disturbance estimation is highly coupled with the compliant control module, making system expansion or parameter adjustment difficult. 3) Insufficient adaptability to continuous disturbances: Existing methods are mostly designed for instantaneous impacts, lacking an effective compliant response mechanism for continuous external forces (such as lateral tension, ramp thrust, and human traction). Summary of the Invention

[0005] Therefore, it is necessary to provide an active compliant control method for quadruped robots based on hierarchical reinforcement learning, which can adapt to continuous disturbances and has a high degree of coupling between disturbance estimation and compliant control module, in order to address the above-mentioned technical problems.

[0006] The technical solution of the present invention is as follows: An active compliance control model, the model comprising: a bottom-level motion policy network and an upper-level compliance control policy network; The underlying motion strategy network includes: a basic encoder, a force estimation head, a velocity estimation head, and a basic strategy network module. The basic encoder is connected to the force estimation head and the velocity estimation head, respectively, and the force estimation head and the velocity estimation head are connected to the basic strategy network module, respectively. The basic encoder is used to encode the current and historical first observation data into a low-dimensional feature vector; the first observation data includes the body sensor feedback data of the quadruped robot, the trunk velocity command, and the expected joint position at the previous moment; The force estimation head is used to estimate the external force currently acting on the quadruped robot from the low-dimensional feature vector; the velocity estimation head is used to estimate the current trunk velocity of the quadruped robot from the low-dimensional feature vector. The basic strategy network module is used to generate the current expected joint position of the quadruped robot based on the latent features obtained by fusing the low-dimensional feature vector at the current moment, the estimated external force at the current moment, and the current trunk velocity, as well as the first observation data at the current moment. The upper-level compliant control strategy network is used to generate the current residual velocity command based on the second observation data, and to add the current residual velocity command and the current trunk velocity command to obtain the corrected current trunk velocity command. The second observation data includes: current body sensor feedback data, current trunk velocity command, expected joint position at the previous moment, residual velocity command at the previous moment, low-dimensional feature vector at the previous moment, and potential features obtained by fusing the estimated current external force and current trunk velocity.

[0007] This invention also proposes an active compliant control method for quadruped robots based on hierarchical reinforcement learning, which applies the above-mentioned active compliant control model and includes the following steps: S1: Obtain the first observation data and input the first observation data into the underlying motion strategy network of the active compliant control model. The underlying motion strategy network outputs the current expected joint position and simultaneously outputs the potential features obtained by fusing the low-dimensional feature vector, the estimated current external force, and the current trunk velocity, thereby updating the second observation data. S2: Obtain the second observation data, and use the second observation data to construct the upper-level compliant control strategy network of the active compliant control model. The upper-level compliant control strategy network outputs the correction amount of the current torso speed command. S3: Update the current torso speed command to the corrected current torso speed command, and update the command information in the first observation data accordingly, so that the underlying motion strategy network of the quadruped robot tracks the corrected torso speed command and realizes active compliant control at the current moment. S4: Repeat steps S1 to S3 until active compliance control is no longer needed, then stop repeating.

[0008] Compared with the prior art, the beneficial effects of the technical solution of the present invention are: This invention adopts a hierarchical design, separating external force estimation from compliant control. Through the decoupling of the "perception-planning-execution" paradigm, it optimizes disturbance estimation and compliant behavior generation respectively. The bottom-level motion strategy network focuses on robust motion control, while the upper-level compliant control strategy network focuses on dynamically adjusting motion speed commands based on external force estimation, realizing hierarchical compliant response. This achieves a high degree of coupling between disturbance estimation and compliant control modules, enabling active compliance to adapt to continuous disturbances. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of the active compliance control model in Example 1; Figure 2 This is a schematic diagram of the pulley system for generating a constant external force in Example 3; Figure 3 This is a schematic diagram illustrating the performance of different controllers in Example 3 under impact and continuous lateral force. Figure 4 This is a schematic diagram of the pulley system compliance test in Example 3; Figure 5 This is a schematic diagram of the multi-terrain robustness test in Example 3; Figure 6 This is a schematic diagram of the human-computer interaction test in Example 3. Detailed Implementation

[0010] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this embodiment. To better illustrate this embodiment, some parts in the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions; It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings.

[0011] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments. Example 1 This embodiment proposes an active compliance control model. Figure 1 This is a schematic diagram of the active compliance control model proposed in this embodiment.

[0012] like Figure 1 As shown, the model includes: a bottom-level motion strategy network and an upper-level compliant control strategy network; The underlying motion strategy network includes: a basic encoder, a force estimation head, a velocity estimation head, and a basic strategy network module. The basic encoder is connected to the force estimation head and the velocity estimation head, respectively, and the force estimation head and the velocity estimation head are connected to the basic strategy network module, respectively. The basic encoder is used to encode the current and historical first observation data into a low-dimensional feature vector; the first observation data includes the body sensor feedback data of the quadruped robot, the trunk velocity command, and the expected joint position at the previous moment; The force estimation head is used to estimate the external force currently acting on the quadruped robot from the low-dimensional feature vector; the velocity estimation head is used to estimate the current trunk velocity of the quadruped robot from the low-dimensional feature vector. The basic strategy network module is used to generate the current expected joint position of the quadruped robot based on the latent features obtained by fusing the low-dimensional feature vector at the current moment, the estimated external force at the current moment, and the current trunk velocity, as well as the first observation data at the current moment. The upper-level compliant control strategy network is used to generate the current residual velocity command based on the second observation data, and to add the current residual velocity command and the current trunk velocity command to obtain the corrected current trunk velocity command. The second observation data includes: current body sensor feedback data, current trunk velocity command, expected joint position at the previous moment, residual velocity command at the previous moment, low-dimensional feature vector at the previous moment, and potential features obtained by fusing the estimated current external force and current trunk velocity.

[0013] In an optional embodiment, the force estimation head, velocity estimation head, basic policy network module, and upper-level compliant control policy network are all selected from multilayer perceptrons, and the basic encoder is selected from a multilayer perceptron or a temporal convolutional network.

[0014] Example 2 This embodiment proposes an active compliant control method for quadruped robots based on hierarchical reinforcement learning. Applying the active compliant control model described in Embodiment 1, the method includes the following steps: S1: Obtain the first observation data and input the first observation data into the underlying motion strategy network of the active compliant control model. The underlying motion strategy network outputs the current expected joint position and simultaneously outputs the potential features obtained by fusing the low-dimensional feature vector, the estimated current external force, and the current trunk velocity, thereby updating the second observation data. S2: Obtain the second observation data, and use the second observation data to construct the upper-level compliant control strategy network of the active compliant control model. The upper-level compliant control strategy network outputs the correction amount of the current torso speed command. S3: Update the current torso speed command to the corrected current torso speed command, and update the command information in the first observation data accordingly, so that the underlying motion strategy network of the quadruped robot tracks the corrected torso speed command and realizes active compliant control at the current moment. S4: Repeat steps S1 to S3 until active compliance control is no longer needed, then stop repeating.

[0015] In one optional embodiment, the active compliance control model is a pre-trained model. When training the active compliance control model, the underlying motion policy network is trained first. After freezing the weights of the trained underlying motion policy network, the upper-layer compliance control policy network is trained. When training the underlying motion policy network, a training decoder is added before the basic policy network module, so that the force estimation head and velocity estimation head are connected to the decoder respectively, and the decoder is connected to the basic policy network module; wherein, the decoder is used to predict the first observation data at the next moment based on the fused output of the basic encoder, force estimation head and velocity estimation head; The training set is formed by acquiring the first observation data for training, the actual current external force corresponding to the first observation data, and the current trunk velocity. The training set is input into the active compliant control model with a decoder. During the training process, supervised learning, contrastive learning and proximal policy optimization algorithms are used for joint training. During training, the goal is to minimize the loss value of the loss function. The preset loss function is solved iteratively. When the number of iterations reaches the preset value or the loss value reaches the minimum, the training ends and the trained underlying motion policy network is obtained. Remove the decoder from the trained underlying motion policy network to obtain the trained underlying motion policy network.

[0016] In an optional embodiment, the expression for the preset loss function includes:

[0017]

[0018]

[0019]

[0020] In the formula, This represents the preset loss function. Indicates the loss of supervised learning. Indicates the contrast learning loss. This represents the loss from near-end strategy optimization; , and All represent weights; Indicates the current time; Indicates the next moment; This represents the corresponding current trunk velocity in the training set. This represents the velocity estimation head's estimate of the current torso velocity. This represents the net external force acting on the current torso in the training set. This represents the force estimate of the net external force acting on the head and torso at that moment. This represents the first observation data in the training set corresponding to the next time step. This represents the decoder's estimate of the first observation data at the next time step; Indicates the operation of natural exponents; This indicates the calculation of cosine similarity. Indicates the first A sample of the first observation history sequence containing noise. The feature vector obtained by the basic encoder, wherein This represents the sample number in a batch training set; express The feature vector obtained by encoding the corresponding positive sample, where, The corresponding positive samples are the original noise-free first observation history sequences. ; express The corresponding negative sample encoding yields the feature vector, where, This represents the index of a negative sample, excluding those in a batch of training sets. Any noisy first observation history sequence Considered a negative sample; This represents the hyperparameter of temperature coefficient; This indicates the total number of samples in the training set for this batch. To estimate the loss of the value function for the near-end policy optimization algorithm, The entropy loss used to encourage policy exploration in near-end policy optimization algorithms. For near-policy optimization algorithms, the loss is used to maximize the reward function. for The weight.

[0021] In an optional embodiment, the expression for the reward function when training the underlying motion policy network includes:

[0022] This represents the overall reward function during the training of the underlying motion policy network. This represents the reward function aimed at maximizing speed tracking performance. This represents the reward function aimed at maximizing trunk stability. This represents the reward function aimed at maximizing the smoothness of joint motion. This represents the reward function aimed at minimizing energy consumption. This represents the reward function aimed at optimizing gait. reward function The expressions include:

[0023] in This indicates the operation of the natural exponent. The linear velocity tracking error along the x and y axes in the body coordinate system. The angular velocity tracking error along the z-axis in the body coordinate system; reward function The expressions include:

[0024] in Let be the linear velocity along the z-axis in the body coordinate system. Let x be the angular velocity along the x and y axes in the body coordinate system. Let x be the projection of the unit gravity vector onto the xy plane in the body coordinate system. This is for the error in fuselage height; reward function The expressions include:

[0025] in , , These are the expected joint positions sent to the actuator at times t-2, t-1, and t, respectively. reward function The expressions include:

[0026] in For joint torque, For joint velocity, This represents the calculation of variance, and its subscript... These represent the hip joint, thigh joint, and calf joint, respectively. reward function The expressions include:

[0027] in For the first The time of swing of each foot between two consecutive touchdowns. , To account for the error in leg lift height, Let be the swing velocity of the i-th foot end in the xy plane of the machine system. This represents the calculation of variance. This represents the calculation of the mean. This refers to the duty cycle of the four legs.

[0028] In one optional embodiment, the step of freezing the weights of the lower-level motion policy network after training, and then training the upper-level compliant control policy network, includes: The weights of the trained low-level motion policy network are frozen, treated as part of the environmental dynamics, and the upper-level compliant control policy network is trained using a proximal policy optimization algorithm based on an asymmetric Actor-Critic architecture. Privileged observations are then input into the network. Input the Critic network, In the formula, Indicates the current time The corresponding second observation data, Indicates trunk speed. The environmental parameters for randomization of the representation domain include the ground friction coefficient, body mass, and motor gain. This represents the historical sequence of external forces in the torso coordinate system. This represents a height map of the terrain surrounding the quadruped robot; the output of the upper-layer compliant network obtained from the training of the proximal policy optimization algorithm is... , Indicates the current residual velocity command; The near-end policy optimization algorithm is iteratively trained with the aim of maximizing the reward value corresponding to the reward function until the number of iterations reaches a preset value or the reward value reaches the maximum, at which point the training ends and a trained upper-layer compliant control policy network is obtained. Specifically, the reward function is set based on the target task, enabling the upper-layer compliant network to obtain the functions required for the target task.

[0029] In an optional embodiment, when the objective is to respond compliantly to external forces while maintaining a stable yaw angle, the reward function includes:

[0030]

[0031]

[0032] In the formula, This represents the reward function corresponding to the target task. This represents the reward function used to track the corrected speed command. This represents the reward function aimed at achieving a compliant response to external forces. This represents the reward function aimed at maximizing trunk stability. This represents the reward function aimed at maximizing the smoothness of joint motion. This represents the reward function aimed at minimizing energy consumption. This represents the reward function aimed at optimizing gait. These represent the desired linear velocity in the xy plane and the desired angular velocity along the z-axis, respectively, in the user's original command. These represent the outputs of the upper-layer compliant network. The correction amount for the xy-plane linear velocity command and the z-axis angular velocity command. and Represents the actual linear velocity and z-axis angular velocity of the fuselage in the xy plane. This represents the vector of the net external force acting on the machine in the xy plane within the body coordinate system. The preset force threshold for transitioning from robust to compliant behavior. Virtual impedance for pre-defined compliant behavior.

[0033] In an optional embodiment, when the objective task is to make a compliant response to external forces while actively rotating the fuselage to align the yaw angle with the direction of the external forces, the expression of the reward function includes:

[0034]

[0035] In the formula, This represents the reward function corresponding to the target task; This represents the reward function aimed at aligning the yaw angle with the direction of the external force. This indicates the yaw direction of the quadruped robot's body in the xy plane of the world coordinate system. This represents the horizontal external force acting on the body of a quadruped robot in the world coordinate system.

[0036] In an optional embodiment, during training, constraints are further set, including: The modified current torso speed command is limited to a speed command range, which includes: ; The quadruped robot is subject to a maximum limit on the magnitude of the impulse and sustained force, which includes the maximum limit on the linear velocity change caused by the velocity impulse. The maximum speed is 1.5 m / s, and the angular velocity impulse causes a sudden change in the robot's angular velocity. Maximum 1.5 radians / second, continuous force Maximum 50 Newtons; Furthermore, during training, to enhance the generalization of the policy, domain randomization was applied to the environment, dynamic model, actuator parameters, and sensor noise of different agents during parallel training. This domain randomization included: Environmental randomization: Ground friction coefficient: Random, uniform sampling within the range; Ground elastic coefficient: Random, uniform sampling within the range; Randomization of the dynamic model: Airframe weight increment: Randomly added to the original weight. kg; Airframe moment of inertia scaling: Randomly multiply the original moment of inertia by... ; Fuselage center of mass shift: The center of mass shifts randomly in the x, y, and z directions. cm; Actuator randomization: Motor scaling factor scaling: Randomly multiply the original scaling factor by... ; Motor differential coefficient scaling: Randomly multiply the original differential coefficients by... ; Motor output torque scaling: Randomly multiplies the original output torque by... ; Motor delay: Randomizes torque output Output; Sensor randomization involves adding uniformly sampled random noise to the true measurement value. Angular velocity noise: rad / s; Gravity unit vector projection noise: ; Joint position noise: rad; Joint velocity noise: rad / s; Constant joint position offset: rad; Among them, the ground friction coefficient, ground elastic coefficient, fuselage mass increment, fuselage moment of inertia scaling, fuselage center of gravity offset, motor proportional coefficient scaling, motor differential coefficient scaling, motor output torque scaling, and joint position constant offset will be combined into privileged information during training. The input is fed into the critic network; In addition, during training, random linear velocity impulses are applied to the quadruped robot within a preset time period. With angular velocity impulse External forces are applied in a manner that allows for the application of force in each environment. In each environment, a velocity impulse and an angular velocity impulse are randomly applied at 10-second intervals. The magnitudes of the resulting velocity and angular velocity changes are randomly sampled within a limited range, while the direction of the impulse is randomly sampled in the world coordinate system. Simultaneously, in each environment, the magnitude and direction of the external force are randomly sampled at 4-second intervals and applied continuously for 4 seconds. For the training of the underlying strategy, a progressive interference application method based on course learning is adopted, and the maximum amplitude of impulse and persistence are both determined by task rewards. Whether the set threshold is reached determines the course progress; while when training the upper-level strategy, since the lower-level motion strategy already has strong robustness, the interference amplitude is always set to the upper limit value defined by the course.

[0037] As an example, this invention proposes a hierarchical active compliance control framework called HAC-LOCO, the overall architecture of which is as follows: Figure 1 As shown. The framework mainly consists of two key components: 1) Robust motion strategy at the bottom layer: Based on historical ontological perception information (joint position, velocity, trunk angular velocity, etc.), external forces and ontological velocities are estimated to generate joint target position commands; 2) Upper-layer compliant control strategy: Generate residual speed commands based on the underlying coding characteristics, and dynamically adjust the original speed commands to achieve active compliant control.

[0038] This framework employs a hierarchical design, separating external force estimation from compliant control. Through a "perception-planning-execution" paradigm, it decouples these components, optimizing disturbance estimation and compliant behavior generation separately. The lower-level strategy focuses on robust motion control, while the upper-level strategy focuses on dynamically adjusting motion velocity commands based on external force estimation, achieving a tiered compliant response.

[0039] As an example, the underlying strategy receives historical ontology-aware information as observation input (the first observation input), denoted as... (H=10). Specifically, the observations at each time step are recorded as follows:

[0040] Includes the following information: Body sensor feedback Including trunk angular velocity Gravity vector projection Joint position and speed The joint position and velocity at the previous moment and and clock signals (Step frequency f = 2.5Hz). Aircraft speed command: : These represent the horizontal linear velocity command and the yaw rate command, respectively. Previous time-instance strategy output. : That is, the expected joint position at the previous time step.

[0041] As an example, the network architecture of the underlying motion strategy includes the following during training: a basic encoder. : Historical observation sequence Encode as a low-dimensional feature vector Force estimation head :from Estimating external forces Speed ​​estimation head :from Estimate trunk velocity Decoder module Based on the latent features after fusion Predicting the next moment of observation Basic policy network π: Input and Generate target joint position By PD controller (proportional gain) Differential gain Track and execute.

[0042] As an example, the upper-layer compliant strategy obtains potential information from the lower-layer motion strategy encoder and generates residual velocity commands to achieve a compliant response to external force disturbances. This design follows the traditional impedance control concept—adjusting the reference velocity based on external force feedback. To avoid confusion with the lower-layer strategy, we denote the upper-layer strategy action as... The observed value is denoted as .

[0043] The observation space of the upper-level strategy includes: the robot's self-perceived state. Original torso speed command Potential features of the previous moment and the preliminary actions of upper-level and lower-level strategies. ,Right now

[0044] action This represents the residual velocity command used to achieve compliant behavior, defined as follows: The final correction speed instruction passed to the underlying strategy is generated by superimposing the original instruction and the residual instruction:

[0045] in This refers to the corrected speed instructions that the underlying strategy needs to track.

[0046] As an example, the training process of the upper-layer compliant policy network is as follows: the upper-layer policy undergoes a second stage of training after the lower-layer policy training is completed.

[0047] In this stage, the weights of the underlying policy network are frozen, treated as part of the environmental dynamics, and trained using the PPO algorithm based on an asymmetric Actor-Critic architecture. The privileged observation inputs of the Critic network... Defined as ,in (Torso speed) (Randomization domain parameters) (Historical external force sequence) and The definition of (terrain elevation map) is consistent with the underlying strategy.

[0048] In the second phase of training, the velocity tracking reward is modified, and a disturbance resistance-compliance tradeoff reward is introduced. By adjusting the overall reward function, the quadruped robot can achieve a dual capability: maintaining accurate trajectory tracking under minor external force disturbances, while exhibiting compliant behavior under larger external forces. This compliant behavior is controlled by two parameters: Force threshold : The critical value for distinguishing between precise tracking and compliant behavior, when external force Exceed Time-triggered compliant mode; virtual impedance Adjusting the compliant stiffness, relatively small Value enhances compliance, larger This makes robots more rigid.

[0049] It is important to note that the underlying motion strategy only needs to track the corrected velocity command; that is, the corrected velocity tracking reward function will be dynamically adjusted based on the residual velocity command output by the upper-level strategy.

[0050] The newly introduced disturbance resistance-compliance tradeoff reward function takes the form of:

[0051]

[0052] This design forces the upper-level strategy to operate when external forces are below a threshold. Minimize command corrections as much as possible, and when external forces exceed... Then, based on the target impedance Generate compliant behavior (i.e., promote) ≈ This method combines the interpretability of model-based control (through...). and The physical meaning of the equations and the flexibility of reinforcement learning optimize motion stability and energy efficiency while ensuring a smooth transition from precise tracking to compliant behavior.

[0053] By directly adjusting the reward function and The parameters allow for changes to compliant behavior characteristics without retraining the underlying motion policy, enabling efficient customization of motion controllers with different compliant behavior features. Furthermore, the disturbance-resistant-compliant reward function... The magnitude of the (i.e., the rotational angular velocity command correction) is penalized to maintain the robot's heading stability under external interference.

[0054] As an example, this embodiment adds a head orientation reward to the reward function. ,in The award, which defines the robot's head orientation in the world coordinate system, encourages the alignment of the robot's head with the direction of the applied external force. This allows the robot to maintain compliance while expanding to a mode where the heading is aligned with the direction of the external force, thus adapting to specific human-machine interaction applications such as human-assisted maneuvering.

[0055] The upper-level strategy uses the same simulation environment settings as the lower-level strategy, including curriculum learning and domain randomization mechanisms. For curriculum learning, a velocity curriculum and a gridded adaptive terrain curriculum are employed, enabling the strategy to gradually adapt to increasingly complex commands and environments.

[0056] Meanwhile, in order to improve the generalization of the strategy and narrow the gap between simulation and reality, domain randomization is introduced during the training process to randomize robot dynamic parameters (mass, inertia, centroid offset, ground friction coefficient and elastic coefficient, etc.), actuator parameters (PD gain, motor strength, motion delay) and sensor noise (angular velocity, gravity projection direction, joint position, joint velocity).

[0057] Furthermore, two types of external disturbances are considered simultaneously during the training process: 1) Instantaneous impact disturbance: A random linear velocity impulse is applied to the robot every 10 seconds. With angular velocity impulse ;2) Continuous interference: The continuous external force f will be continuously applied to the robot's torso, and its magnitude and direction will change randomly every 4 seconds.

[0058] The core objective of this invention is to provide a hierarchical active compliant control framework (HAC-LOCO) to address the following problems in the prior art: 1) It enables explicit estimation and hierarchical adaptation of external forces, providing different responses to transient and continuous disturbances, thereby improving the motion stability and active compliance of quadruped robots under continuous external disturbances; 2) It decouples the underlying motion control from the upper-level compliant planning, enabling module reusability, improving algorithm scalability and physical interpretability; 3) It improves energy efficiency, reduces unnecessary energy consumption, and enhances safety and human-machine interaction friendliness in dynamic disturbance environments.

[0059] This invention achieves adaptive compliant motion control of a quadruped robot under continuous external disturbances by decoupling the disturbance estimation and compliant control modules, and has the following innovative highlights: 1) Hierarchical Reinforcement Learning Architecture: A two-stage hierarchical control framework is proposed, separating external force estimation from compliant control. Hierarchical compliant response is achieved through the collaborative optimization of robust motion at the lower level and active compliance at the upper level. This framework enables the robot to effectively resist transient impact disturbances and exhibit compliant motion to continuous external force disturbances. Experimental verification shows that this framework allows the robot to generate natural compliant responses to external disturbances, thereby improving safety, stability, and energy efficiency.

[0060] 2) Learning-based state estimator: An autoencoder-based historical feature extraction network is designed to extract key features from continuous ontology perception information and explicitly estimate external forces and body velocities. This network combines supervised learning and reinforcement learning to achieve efficient feature learning and disturbance estimation. Experimental results show that the proposed estimator significantly improves the robot's compliance, stability, and energy efficiency in complex environments.

[0061] 3) Impedance-based lightweight compliant control module: Inspired by impedance control, a lightweight compliant control module was designed that can dynamically adjust velocity commands based on real-time external force estimation to achieve an impedance-adjustable compliant response. This module is quick to train and replaceable, allowing the reuse of the underlying robust motion control network while adjusting the robot's impedance characteristics without retraining the entire system, while maintaining clear physical interpretability and analyzability.

[0062] Example 3 Based on the active compliant control model of Example 1 and the active compliant control method of Example 2, this embodiment proposes the following comparative experiments: To verify that this invention achieves higher stability, compliance, safety, and lower joint torque and energy consumption compared to existing technologies under impact or continuous external force disturbance, we implemented the following state-of-the-art reinforcement learning-based quadruped robot control algorithm for method comparison, ablation experiments, and performance evaluation: Deep Compliant Control (DCC) [Reference: Deep compliant control for legged robots]: A learning-based compliant control method with an added recovery training module.

[0063] PA-LOCO [Reference: PA-LOCO: Learning perturbation-adaptive locomotion for quadruped robots]: A perturbation-adaptive motion controller designed for robust motion.

[0064] HAC: The proposed HAC-LOCO framework (which includes high-level and low-level strategies).

[0065] HAC-Low: A simplified version of the HAC-LOCO framework that only uses the underlying robust strategy.

[0066] Experimental data verification: like Figure 2 As shown, to verify the beneficial effects of the present invention, a cable-pulley system was designed to evaluate the effectiveness of the HAC-LOCO framework in achieving compliant behavior under continuous external perturbation. This system generates precise and continuous force, ensuring the repeatability and consistency of the applied force perturbation in multiple trials. The experimental setup includes two fixed pulleys, an inelastic rope, and a 4.5-liter bottle of water. The weight of the bottled water is transferred to the quadruped robot via the rope. During the experiment, a continuous lateral force is applied to the robot's torso by lifting and releasing the water tank.

[0067] The experimental scenario was set up with the quadruped robot in a stationary state, first subjected to a lateral impact, and then subjected to a continuous external force through a cable-pulley system. The compliant behavior of each controller was evaluated using three key metrics: maximum torque of all joints, trunk tilt angle, and robot power consumption. Figure 3 The results show the stability, safety, and energy efficiency of each controller under continuous disturbances.

[0068] Experiments show that while the DCC controller can withstand the initial impact, it cannot maintain stability under continuous disturbances, eventually leading to tipping over. This instability stems from its control design not considering continuous disturbances, which weakens the robot's ability to maintain balance through posture adjustments. PA-LOCO, on the other hand, demonstrates adaptability to both instantaneous and continuous external forces, achieving dynamic and aggressive motion stability through rapid leg swings, but this results in drastic fluctuations in joint torque and a significant increase in power consumption after impact.

[0069] In contrast, the underlying strategy in the HAC-LOCO framework proposed in this invention effectively mitigates external shocks, thereby reducing maximum joint torque and overall power consumption. However, when faced with sustained external forces, this strategy tends to maintain a rigid posture, holding a fixed position and lacking attitude adjustment capability. Through a higher-level strategy that generates residual velocity commands based on external force estimation, HAC-LOCO can dynamically adjust its position to accommodate sustained disturbances, rather than rigidly resisting them. This approach improves trunk attitude stability with smoother power consumption, thus exhibiting the highest level of compliance among all tested controllers and achieving an optimal balance between velocity tracking and compliance optimization.

[0070] Simulation Data Validation: We further validated the performance of each control strategy under instantaneous shocks and continuous disturbances through comprehensive simulation evaluation across multiple scenarios. In the Isaac Gym simulation environment, each strategy underwent a 20-second / 4096-channel parallel simulation test under different configurations: random pulse forces were applied every 3 seconds to evaluate the instantaneous shock response, while continuous external forces were repeatedly applied at the same intervals to test disturbance adaptability. All external force amplitudes and randomization parameters remained consistent with the training settings. The simulation results are summarized in the table below.

[0071] Table 1. Results of Multi-Scenario Integrated Simulation Evaluation

[0072] Simulation data show that: 1) The HAC-Low strategy exhibits excellent robustness against disturbances with the highest survival rate, and its response to instantaneous shocks is basically on par with the full version of the HAC-LOCO algorithm, confirming that the underlying strategy dominates the response mechanism to sudden shocks; 2) When facing continuous disturbances, the maximum joint torque of the full HAC controller is significantly reduced compared with the single underlying strategy, verifying the compliance enhancement effect of the high-level strategy under continuous external forces; 3) This conclusion is in high agreement with the physical experiment, and the effectiveness of HAC-LOCO in improving compliant motion performance is doubly verified.

[0073] In summary, through comprehensive simulations and physical experiments, the active compliant control algorithm proposed in this invention demonstrates the following significant advantages over existing technologies: 1) Improved stability and compliance: Experimental results show that, compared with existing technologies, this invention exhibits higher stability and compliance in resisting instantaneous and continuous external force disturbances, with a significantly higher survival rate and significantly lower body tilt and shaking. 2) Reduced hardware damage risk and improved energy efficiency: When subjected to impacts or continuous external force disturbances, compared with existing control algorithms, the control algorithm proposed in this invention enables the robot to significantly reduce joint torque and average power consumption while ensuring body stability. 3) Improved algorithm modularity, scalability, and interpretability: The two-stage training method and the clearly defined upper and lower layer strategy design proposed in this invention give it stronger scalability and physical interpretability compared to other existing fully end-to-end algorithms. Its lightweight upper layer strategy training only takes fifteen minutes, allowing for rapid iteration or functional addition. 4) More user-friendly for human-computer interaction: Compared with existing algorithms that require the addition of expensive force sensors to achieve contact-based human-computer interaction, this invention only requires the most basic IMU and joint encoder data to achieve natural human-computer interaction (such as traction walking).

[0074] This embodiment also provides the following specific implementation examples: Example 1: Compliance test of pulley system (training a low-impedance upper-layer policy while keeping the underlying policy unchanged) A cable-pulley system (such as) was designed. Figure 2 (As shown), this system was used to evaluate the effectiveness of the HAC-LOCO framework in achieving compliant behavior under sustained external perturbations. The system generates precise and continuous forces, ensuring the repeatability and consistency of applied force perturbations across multiple trials. The experimental setup consists of two fixed pulleys, an inelastic rope, and a 4.5-liter bottle of water. The weight of the water is transferred to the quadruped robot via the rope. During the experiment, a continuous lateral force is applied to the robot's torso by lifting and releasing the water tank.

[0075] like Figure 4 As shown, during the application of external forces, the algorithm of this invention can control the robot to remain stable and exhibit compliance with external forces.

[0076] Example 2: Multi-terrain robustness test (training a high-impedance upper-layer policy while keeping the underlying policy unchanged) Test scenarios: 1) Stairs: This invention enables the quadruped robot to smoothly climb stairs while dragging a 1.5L water bottle; 2) Grass: This invention enables the quadruped robot to drag a 25kg cart across uneven grass; 3) Garage ramp: This invention enables the quadruped robot to smoothly drag a 4.5L water bucket across a garage ramp (20° inclination angle) and maintain stability when experiencing leg rope entanglement or external impact.

[0077] Example 3: Human-Computer Interaction Testing While keeping the underlying policy unchanged, we only need to add a head-orientation reward when training the upper-level policy:

[0078] in This indicates the yaw direction of the quadruped robot's body in the xy plane of the world coordinate system. This award represents the horizontal external force acting on the body of a quadruped robot in the world coordinate system. It encourages the alignment of the robot's head with the direction of the applied force, enabling the robot to maintain compliance while expanding to a mode where the heading is aligned with the direction of the applied force, thus adapting to specific human-computer interaction applications such as human-assisted maneuvering.

[0079] like Figure 6 As shown, the modular control scheme designed in this invention can achieve contact-based human-machine interaction without sensors with simple expansion. Without any external perception or remote control, the quadruped robot can follow and turn by sensing the traction force applied by the human and can climb complex terrain under traction.

Claims

1. An active compliance control model, characterized in that, The model includes: a bottom-level motion strategy network and an upper-level compliant control strategy network; The underlying motion strategy network includes: a basic encoder, a force estimation head, a velocity estimation head, and a basic strategy network module. The basic encoder is connected to the force estimation head and the velocity estimation head, respectively, and the force estimation head and the velocity estimation head are connected to the basic strategy network module, respectively. The basic encoder is used to encode the current and historical first observation data into a low-dimensional feature vector; the first observation data includes the body sensor feedback data of the quadruped robot, the trunk velocity command, and the expected joint position at the previous moment; The force estimation head is used to estimate the external force currently acting on the quadruped robot from the low-dimensional feature vector; the velocity estimation head is used to estimate the current trunk velocity of the quadruped robot from the low-dimensional feature vector. The basic strategy network module is used to generate the current expected joint position of the quadruped robot based on the latent features obtained by fusing the low-dimensional feature vector at the current moment, the estimated external force at the current moment, and the current trunk velocity, as well as the first observation data at the current moment. The upper-level compliant control strategy network is used to generate the current residual velocity command based on the second observation data, and to add the current residual velocity command and the current trunk velocity command to obtain the corrected current trunk velocity command. The second observation data includes: current body sensor feedback data, current trunk velocity command, expected joint position at the previous moment, residual velocity command at the previous moment, low-dimensional feature vector at the previous moment, and potential features obtained by fusing the estimated current external force and current trunk velocity.

2. The active compliance control model according to claim 1, characterized in that, The force estimation head, velocity estimation head, basic strategy network module, and upper-level compliant control strategy network all use multilayer perceptrons, and the basic encoder uses a multilayer perceptron or a temporal convolutional network.

3. A method for active compliant control of a quadruped robot based on hierarchical reinforcement learning, applying the active compliant control model described in claim 1 or 2, characterized in that, Includes the following steps: S1: Obtain the first observation data and input the first observation data into the underlying motion strategy network of the active compliant control model. The underlying motion strategy network outputs the current expected joint position and simultaneously outputs the potential features obtained by fusing the low-dimensional feature vector, the estimated current external force, and the current trunk velocity, thereby updating the second observation data. S2: Obtain the second observation data, and use the second observation data to construct the upper-level compliant control strategy network of the active compliant control model. The upper-level compliant control strategy network outputs the correction amount of the current torso speed command. S3: Update the current torso speed command to the corrected current torso speed command, and update the command information in the first observation data accordingly, so that the underlying motion strategy network of the quadruped robot tracks the corrected torso speed command and realizes active compliant control at the current moment. S4: Repeat steps S1 to S3 until active compliance control is no longer needed, then stop repeating.

4. The active compliant control method for quadruped robots based on hierarchical reinforcement learning according to claim 3, characterized in that, The active compliance control model is a pre-trained model. When training the active compliance control model, the underlying motion policy network is trained first. After freezing the weights of the trained underlying motion policy network, the upper-layer compliance control policy network is trained. When training the underlying motion policy network, a training decoder is added before the basic policy network module, so that the force estimation head and velocity estimation head are connected to the decoder respectively, and the decoder is connected to the basic policy network module; wherein, the decoder is used to predict the first observation data at the next moment based on the fused output of the basic encoder, force estimation head and velocity estimation head; The training set is formed by acquiring the first observation data for training, the actual current external force corresponding to the first observation data, and the current trunk velocity. The training set is input into the active compliant control model with a decoder. During the training process, supervised learning, contrastive learning and proximal policy optimization algorithms are used for joint training. During training, the goal is to minimize the loss value of the loss function. The preset loss function is solved iteratively. When the number of iterations reaches the preset value or the loss value reaches the minimum, the training ends and the trained underlying motion policy network is obtained. Remove the decoder from the trained underlying motion policy network to obtain the trained underlying motion policy network.

5. The active compliant control method for quadruped robots based on hierarchical reinforcement learning according to claim 4, characterized in that, The expression for the preset loss function includes: In the formula, This represents the preset loss function. Indicates the loss of supervised learning. Indicates the contrast learning loss. This represents the loss from near-end strategy optimization; , and All represent weights; Indicates the current time; Indicates the next moment; This represents the corresponding current trunk velocity in the training set. This represents the velocity estimation head's estimate of the current torso velocity. This represents the net external force acting on the current torso in the training set. This represents the force estimate of the net external force acting on the head and torso at that moment. This represents the first observation data in the training set corresponding to the next time step. This represents the decoder's estimate of the first observation data at the next time step; Indicates the operation of natural exponents; This indicates the calculation of cosine similarity. Indicates the first A sample of the first observation history sequence containing noise. The feature vector obtained by the basic encoder, wherein This represents the sample number in a batch training set; express The feature vector obtained by encoding the corresponding positive sample, where, The corresponding positive samples are the original noise-free first observation history sequences. ; express The corresponding negative sample encoding yields the feature vector, where, This represents the index of a negative sample, excluding those in a batch of training sets. Any noisy first observation history sequence Considered a negative sample; This represents the hyperparameter of temperature coefficient; This indicates the total number of samples in the training set for this batch. To estimate the loss of the value function for the near-end policy optimization algorithm, The entropy loss used to encourage policy exploration in near-end policy optimization algorithms. For near-policy optimization algorithms, the loss is used to maximize the reward function. for The weight.

6. The active compliant control method for quadruped robots based on hierarchical reinforcement learning according to claim 5, characterized in that, The expression for the reward function when training the underlying motion policy network includes: This represents the overall reward function during the training of the underlying motion policy network. This represents the reward function aimed at maximizing speed tracking performance. This represents the reward function aimed at maximizing trunk stability. This represents the reward function aimed at maximizing the smoothness of joint motion. This represents the reward function aimed at minimizing energy consumption. This represents the reward function aimed at optimizing gait. reward function The expressions include: in This indicates the operation of the natural exponent. The linear velocity tracking error along the x and y axes in the body coordinate system. The angular velocity tracking error along the z-axis in the body coordinate system; reward function The expressions include: in Let be the linear velocity along the z-axis in the body coordinate system. Let x be the angular velocity along the x and y axes in the body coordinate system. Let x be the projection of the unit gravity vector onto the xy plane in the body coordinate system. This is for the error in fuselage height; reward function The expressions include: in , , These are the expected joint positions sent to the actuator at times t-2, t-1, and t, respectively. reward function The expressions include: in For joint torque, For joint velocity, This represents the calculation of variance, and its subscript... These represent the hip joint, thigh joint, and calf joint, respectively. reward function The expressions include: in For the first The time of swing of each foot between two consecutive touchdowns. , To account for the error in leg lift height, Let be the swing velocity of the i-th foot end in the xy plane of the machine system. This represents the calculation of variance. This represents the calculation of the mean. This refers to the duty cycle of the four legs.

7. The active compliant control method for a quadruped robot based on hierarchical reinforcement learning according to any one of claims 3 to 6, characterized in that, After freezing the weights of the lower-level motor policy network that has completed training, the steps for training the upper-level compliant control policy network include: The weights of the trained low-level motion policy network are frozen, treated as part of the environmental dynamics, and the upper-level compliant control policy network is trained using a proximal policy optimization algorithm based on an asymmetric Actor-Critic architecture. Privileged observations are then input into the network. Input the Critic network, In the formula, Indicates the current time The corresponding second observation data, Indicates trunk speed. The environmental parameters for randomization of the representation domain include the ground friction coefficient, body mass, and motor gain. This represents the historical sequence of external forces in the torso coordinate system. This represents a height map of the terrain surrounding the quadruped robot; the output of the upper-layer compliant network obtained from the training of the proximal policy optimization algorithm is... , Indicates the current residual velocity command; The near-end policy optimization algorithm is iteratively trained with the aim of maximizing the reward value corresponding to the reward function until the number of iterations reaches a preset value or the reward value reaches the maximum, at which point the training ends and a trained upper-layer compliant control policy network is obtained. Specifically, the reward function is set based on the target task, enabling the upper-layer compliant network to obtain the functions required for the target task.

8. The active compliant control method for a quadruped robot based on hierarchical reinforcement learning according to claim 7, characterized in that, When the objective is to respond compliantly to external forces while maintaining a stable yaw angle, the reward function includes: In the formula, This represents the reward function corresponding to the target task. This represents the reward function used to track the corrected speed command. This represents the reward function aimed at achieving a compliant response to external forces. This represents the reward function aimed at maximizing trunk stability. This represents the reward function aimed at maximizing the smoothness of joint motion. This represents the reward function aimed at minimizing energy consumption. This represents the reward function aimed at optimizing gait. These represent the desired linear velocity in the xy plane and the desired angular velocity along the z-axis, respectively, in the user's original command. These represent the outputs of the upper-layer compliant network. The correction amount for the xy-plane linear velocity command and the z-axis angular velocity command. and Represents the actual linear velocity and z-axis angular velocity of the fuselage in the xy plane. This represents the vector of the net external force acting on the machine in the xy plane within the body coordinate system. The preset force threshold for transitioning from robust to compliant behavior. Virtual impedance for pre-defined compliant behavior.

9. The active compliant control method for a quadruped robot based on hierarchical reinforcement learning according to claim 7, characterized in that, When the objective is to respond compliantly to external forces while actively rotating the fuselage to align the yaw angle with the direction of the external forces, the expression for the reward function includes: In the formula, This represents the reward function corresponding to the target task; This represents the reward function aimed at aligning the yaw angle with the direction of the external force. This indicates the yaw direction of the quadruped robot's body in the xy plane of the world coordinate system. This represents the horizontal external force acting on the body of a quadruped robot in the world coordinate system.

10. The active compliant control method for a quadruped robot based on hierarchical reinforcement learning according to claim 7, characterized in that, During training, constraints are also set, including: The modified current torso speed command is limited to a speed command range, which includes: ; The quadruped robot is subject to a maximum limit on the magnitude of the impulse and sustained force, which includes the maximum limit on the linear velocity change caused by the velocity impulse. The maximum speed is 1.5 m / s, and the angular velocity impulse causes a sudden change in the robot's angular velocity. Maximum 1.5 radians / second, continuous force Maximum 50 Newtons; Furthermore, during training, to enhance the generalization of the policy, domain randomization was applied to the environment, dynamic model, actuator parameters, and sensor noise of different agents during parallel training. This domain randomization included: Environmental randomization: Ground friction coefficient: Random, uniform sampling within the range; Ground elastic coefficient: Random, uniform sampling within the range; Randomization of the dynamic model: Airframe weight increment: Randomly added to the original weight. kg; Airframe moment of inertia scaling: Randomly multiply the original moment of inertia by... ; Fuselage center of mass shift: The center of mass shifts randomly in the x, y, and z directions. cm; Actuator randomization: Motor scaling factor scaling: Randomly multiply the original scaling factor by... ; Motor differential coefficient scaling: Randomly multiply the original differential coefficients by... ; Motor output torque scaling: Randomly multiplies the original output torque by... ; Motor delay: Randomizes torque output Output; Sensor randomization involves adding uniformly sampled random noise to the true measurement value. Angular velocity noise: rad / s; Gravity unit vector projection noise: ; Joint position noise: rad; Joint velocity noise: rad / s; Constant joint position offset: rad; Among them, the ground friction coefficient, ground elasticity coefficient, fuselage mass increment, fuselage moment of inertia scaling, fuselage center of gravity offset, motor proportional coefficient scaling, motor differential coefficient scaling, motor output torque scaling, and joint position constant offset will be combined into privileged information during training. The input is fed into the critic network; In addition, during training, random linear velocity impulses are applied to the quadruped robot within a preset time period. With angular velocity impulse External forces are applied in a manner that allows for the application of force in each environment. In each environment, a velocity impulse and an angular velocity impulse are randomly applied at 10-second intervals. The magnitudes of the resulting velocity and angular velocity changes are randomly sampled within a limited range, while the direction of the impulse is randomly sampled in the world coordinate system. Simultaneously, in each environment, the magnitude and direction of the external force are randomly sampled at 4-second intervals and applied continuously for 4 seconds. For the training of the underlying strategy, a progressive interference application method based on course learning is adopted, and the maximum amplitude of impulse and persistence are both determined by task rewards. Whether the set threshold is reached determines the course progress; while when training the upper-level strategy, since the lower-level motion strategy already has strong robustness, the interference amplitude is always set to the upper limit value defined by the course.

Citation Information

Cited By

  • Quadruped robot autonomous inspection navigation control and active three-dimensional mapping method, system and terminal

    CN121596806A