Hybrid force-position control method and device for quadruped robot

By actively and dynamically modeling the load and using hybrid force-position control, combined with reinforcement learning algorithms, anti-interference target instructions are generated, which solves the problem of load interference in quadruped robots under complex terrain and improves the balance control capability of load handling.

CN122299574APending Publication Date: 2026-06-30HUAZHONG AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG AGRI UNIV
Filing Date
2026-04-30
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

When a quadruped robot is carrying a load with active movement capabilities, it has difficulty accurately sensing the dynamic characteristics and real-time disturbances of the load, which makes it unable to effectively counteract the impact of load disturbances on the robot's balance in complex terrain.

Method used

By extracting key features using an active dynamic load model, generating anti-interference target instructions using a hybrid force-position control model and reinforcement learning algorithm, adjusting the robot's motion strategy to counteract load interference, and constructing an anti-interference compensation mechanism.

Benefits of technology

It enables quadruped robots to adaptively perceive and balance the load under active dynamic conditions in complex terrain, thereby improving their stability and anti-interference ability in load handling tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122299574A_ABST
    Figure CN122299574A_ABST
Patent Text Reader

Abstract

This invention relates to a hybrid force-position control method and apparatus for a quadruped robot. The method includes extracting features from the body perception sequence and quantifying key features of the active dynamic load carried by the quadruped robot using a preset active dynamic load model. These key features include real-time interactive interference forces, enabling adaptive perception of load changes and real-time interference forces. The method also involves actively correcting the quadruped robot's motion commands using the preset hybrid force-position control model and real-time interactive interference forces to generate anti-interference target commands, thereby guiding the robot to actively correct its motion commands. Finally, the method generates actuator motion commands to drive the quadruped robot using a preset motion control strategy network. These actuator motion commands maintain the quadruped robot's balance during the active dynamic load transport process in complex terrain, enabling adaptive adjustment of the quadruped robot and improving its balance control capability and anti-interference robustness in active dynamic load transport tasks in complex terrain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot motion control technology, and in particular to a hybrid force-position control method and device for a quadruped robot. Background Technology

[0002] Quadruped robots, with their excellent terrain adaptability, have demonstrated great potential in transportation and operational tasks in complex environments, and are widely used in logistics, search and rescue, and agriculture. Load handling is a core requirement in these fields, but the introduction of loads alters the system's dynamic characteristics, posing challenges to robot balance control. This is especially true when the load itself has active motion capabilities, such as a robotic arm operating on the robot's back or other load systems with internal degrees of freedom. This "active dynamic load" not only changes the system's mass distribution and inertial characteristics over time, causing a continuous shift in the robot's center of mass, but also generates time-varying disturbance forces acting on the robot. Specifically, when the robotic arm adjusts its position and posture during operation, it not only changes the load's mass distribution but also applies dynamically changing torques to the robot, affecting its balance and stability. Therefore, quadruped robots face two core challenges in active dynamic load scenarios: first, how to accurately perceive the active dynamic characteristics and real-time disturbances of the load to achieve adaptive recognition of load changes; and second, how to design effective resistance control strategies in complex terrain to counteract the impact of load disturbances on robot balance. This makes quadruped robot control with active dynamic loads in complex terrain a key area that urgently needs in-depth exploration.

[0003] Therefore, there is an urgent need to propose a hybrid force-position control method and device for quadruped robots to solve the technical problems of existing technologies in active dynamic load scenarios where the load cannot be accurately perceived in complex terrain and the impact of load interference on robot balance cannot be offset. Summary of the Invention

[0004] In view of this, it is necessary to provide a hybrid force-position control method and device for quadruped robots to solve the technical problems of existing active dynamic load scenarios where the load cannot be accurately perceived in complex terrain and the load interference cannot be offset from the impact on the robot's balance.

[0005] To address the aforementioned problems, in a first aspect, the present invention provides a hybrid force-position control method for a quadruped robot, comprising: Obtain the proprioception sequence of the quadruped robot; The key features of the active dynamic load carried by the quadruped robot are extracted and quantified by a preset active dynamic load model. The key features include the real-time interactive interference force generated by the active dynamic load on the quadruped robot. The motion commands of the quadruped robot are actively corrected by a preset hybrid force-position control model and the real-time interactive interference force to generate anti-interference target commands. The motion command is input into a preset motion control strategy network, and the anti-interference target command is used as the guiding benchmark for the reward function in the preset motion control strategy network to generate actuator motion commands for driving the quadruped robot. The actuator motion commands are used to maintain the balance of the quadruped robot in the process of carrying the active dynamic load in complex terrain.

[0006] In one possible implementation, the active dynamic load is a load system with active motion capability; the key characteristic quantities also include dynamic geometric features and active dynamic states, the dynamic geometric features include mass and center of mass position, and the active dynamic states include joint position, joint velocity and joint force.

[0007] In one possible implementation, the motion command includes a velocity command and a pose command; the anti-interference target command includes a target pose command and a target velocity command; the step of actively correcting the motion command of the quadruped robot through a preset hybrid force-position control model and the real-time interactive interference force to generate the anti-interference target command includes: Based on the preset hybrid force-position control model and the torque component in the real-time disturbance force, the pose command is corrected to generate the target pose command. Based on the preset hybrid force-position control model and the force components in the real-time disturbance force, the speed command is corrected to generate the target speed command.

[0008] In one possible implementation, the training process of the preset motion control policy network is as follows: Acquire the historical ontological perception observation sequence, environmental privilege information, active dynamic load characteristic information, and real disturbance force information of the quadruped robot in multiple training rounds; The historical ontology perception observation sequence is encoded using an environmental encoder network to generate environmental latent variables, base linear velocity estimates, and load disturbance force estimates at different times. Using the proprioceptive observation value at the current moment in the historical proprioceptive observation sequence, the environmental latent variable, the base linear velocity estimate, and the load disturbance force estimate as inputs to the preset motion control strategy network, a reinforcement learning algorithm is used to update the parameters of the preset motion control strategy network to obtain the trained preset motion control strategy network.

[0009] In one possible implementation, the step of using the proprioceptive observation value at the current moment in the historical proprioceptive observation sequence, the environmental latent variables, the estimated base linear velocity, and the estimated load disturbance force as inputs to the preset motion control strategy network, and then using a reinforcement learning algorithm to update the parameters of the preset motion control strategy network to obtain a trained preset motion control strategy network, includes: The environmental encoder network is optimized by using a supervised learning loss function, the base linear velocity estimate, and the load disturbance force estimate to obtain the mean square error loss. The proprioceptive observation value at the current moment in the historical proprioceptive observation sequence, the environmental latent variable, the base linear velocity estimate, and the load disturbance force estimate are input into the preset motion control strategy network, and the current action is output. The proprioceptive observation value at the current moment in the historical proprioceptive observation sequence, the environmental privilege information, the active dynamic load characteristic information, and the real interference force information are input into a preset evaluation network to obtain the current value; The current advantage estimate is obtained by calculating the current action and the current value based on a preset reward function; The parameters of the preset motion control strategy network are updated using a reinforcement learning algorithm and the current advantage estimate to obtain the reinforcement learning loss. The total policy training loss at the current moment is obtained based on the mean squared error loss and the reinforcement learning loss. The total training loss of the strategy at adjacent time points is determined by the gradient descent method to obtain the pre-set motion control strategy network after training.

[0010] In one possible implementation, the step of calculating the current advantage estimate based on the current action and the current value using a preset reward function includes: The current action is parameterized according to a preset state space to obtain the current state value; Based on a preset observation space, the robot body perception observation and collection of the current action are performed to obtain the current observation value; The current reward value is obtained by calculating the current state value and the current observation value using the preset reward function; Based on the current reward value and the current value, the current advantage estimate is obtained.

[0011] In one possible implementation, the preset reward function includes a force-position correction reward term constructed based on the preset hybrid force-position control model; the force-position correction reward term is used to compensate the motion command according to the real disturbance force information, so as to guide the preset motion control strategy network to learn anti-interference behavior.

[0012] In one possible implementation, the supervised learning loss function is:

[0013] In the formula, For mean square error loss, and These are the actual value and the estimated value of the base linear velocity, respectively. and These are the actual value and the estimated value of the load disturbance force, respectively. and They are respectively t Proprioceptive observations at time +1 and reconstructed values ​​from the environment decoder. To balance the weights of different loss terms, This is the expected operation.

[0014] In one possible implementation, the correction formulas for the target pose command and the target velocity command are as follows:

[0015] In the formula, For the target pose command, For the target speed command, X As a preset constant value, The torque component in the real-time disturbance force. For the force component in real-time disturbance force, and These are the stiffness coefficient and damping coefficient, respectively. This is the preset speed command.

[0016] Secondly, the present invention also provides a hybrid force-position control device for a quadruped robot, comprising: The perception and acquisition module is used to acquire the body perception sequence of the quadruped robot; The dynamic load quantification module is used to extract features from the ontology perception sequence and quantify the key features of the active dynamic load carried by the quadruped robot through a preset active dynamic load model; the key features include the real-time interactive interference force generated by the active dynamic load on the quadruped robot. An active correction module is used to actively correct the motion commands of the quadruped robot by using a preset hybrid force-position control model and the real-time interactive interference force, and generate anti-interference target commands. The motion execution module is used to input the motion command into a preset motion control strategy network, and use the anti-interference target command as the guiding reference for the reward function in the preset motion control strategy network to generate actuator motion commands for driving the quadruped robot, and maintain the balance of the quadruped robot in the process of carrying the active dynamic load in complex terrain through the actuator motion commands.

[0017] The beneficial effects of this invention are: acquiring the body perception sequence of a quadruped robot; extracting features from the body perception sequence using a preset active dynamic load model and quantifying the key features of the active dynamic load carried by the quadruped robot; the key features include the real-time interactive interference force generated by the active dynamic load on the quadruped robot, thereby parametrically describing the active and dynamic characteristics of the load through the preset active dynamic load model, achieving adaptive perception of load changes and real-time interference forces; actively correcting the quadruped robot's motion commands through a preset hybrid force-position control model and real-time interactive interference forces, generating anti-interference target commands, thereby achieving hybrid force-position control... Based on the perceived load interference information, the robot is guided to actively correct its motion commands and construct an effective anti-interference compensation mechanism. The motion commands are input into a preset motion control strategy network, and the anti-interference target command is used as the guiding benchmark for the reward function in the preset motion control strategy network to generate actuator motion commands to drive the quadruped robot. The actuator motion commands are used to maintain the balance of the quadruped robot in the process of actively carrying dynamic loads in complex terrain. Thus, the actuator motion commands are output through the preset motion control strategy network to realize the adaptive adjustment of the quadruped robot and improve its balance control capability and anti-interference robustness in the active dynamic load carrying task in complex terrain. Attached Figure Description

[0018] Figure 1 This is a schematic flowchart of an embodiment of the hybrid force-position control method for quadruped robots provided by the present invention; Figure 2 A schematic diagram of an embodiment of the active dynamic load modeling provided by the present invention; Figure 3 A schematic flowchart of an embodiment of the training process of the preset motion control strategy network provided by the present invention; Figure 4 For the present invention Figure 3 A schematic diagram of an embodiment of step S301; Figure 5 This is a schematic diagram of an embodiment of the hybrid force-position control device for quadruped robots provided by the present invention. Detailed Implementation

[0019] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0020] In this invention, quadruped robots demonstrate great potential for load-carrying tasks due to their terrain adaptability. A major challenge in such tasks is maintaining robot stability when the load has active and dynamic characteristics. Unlike static and predictable traditional loads, the motion of active dynamic loads alters the robot's dynamic characteristics in real time, thus interfering with its motion and posture. To address this, this invention proposes a deep reinforcement learning (DALFPDRL) framework based on active dynamic load modeling and hybrid force-position control.

[0021] like Figure 1 As shown, a specific embodiment of the present invention discloses a hybrid force-position control method for a quadruped robot, comprising: S101. Obtain the body perception sequence of the quadruped robot.

[0022] This invention can acquire the proprioception sequence of a quadruped robot, wherein the observations in the proprioception sequence are... It can be collected through body perception sensors. It is a 45-dimensional real vector, and the observations can include the body's angular velocity. Projected gravity Speed ​​command Joint angle Joint velocity and the previous step , It is a 3-dimensional real vector. It is a 12-dimensional real vector.

[0023] S102. Extract features from the ontology perception sequence and quantify the key features of the active dynamic load carried by the quadruped robot by using a preset active dynamic load model; the key features include the real-time interactive interference force generated by the active dynamic load on the quadruped robot.

[0024] In this embodiment of the invention, a preset active dynamic load model can be constructed. This preset active dynamic load model is a parameterized state representation model based on physical mechanisms. The active dynamic load is a load system with active motion capabilities, such as a robotic arm. The preset active dynamic load model is used to extract features from the body perception sequence and quantify the key features of the active dynamic load carried by the quadruped robot. These key features may include the real-time interactive interference force generated by the active dynamic load on the quadruped robot, and may also include dynamic geometric features and active dynamic states. The dynamic geometric features may include mass. Location of the center of mass Parameters such as joint position and other parameters; active dynamic state includes joint position. Joint velocity and joint forces These parameters directly affect system stability: the mass of the load changes the total mass of the system, its center of mass affects the overall center of mass shift of the system, and its active dynamic state determines the inertial force and torque transmitted to the quadruped robot. When the robotic arm is used as an active dynamic load, it is modeled as the mass of the robotic arm. Joint position ,speed Force and the quality of the end effector ,Bit Place and speed The model is used to represent the disturbance force and torque on the quadruped robot as follows: , specifically Figure 2 As shown. According to Figure 2 The process of proactive dynamic load modeling can be described as follows: Step 1: Define the static / dynamic geometric characteristics of the load; The active dynamic load carried by the quadruped robot (as shown in the "Active Dynamic Load" section of the figure) is parameterized as a mass. Location of the center of mass The origin o and its coordinates x and y relative to the quadruped robot body are defined by these parameters. These parameters describe the basic mass distribution and geometric positional relationship of the load, constituting the fundamental static characteristics of the load.

[0025] Step 2: Define the active dynamic state of the load; For load systems with internal degrees of freedom (as shown in the "Arm Load" section of the figure), further extract their active motion parameters, including the positions of each joint. Joint velocity and joint forces At the same time, the mass of the robotic arm body is clearly defined. and the quality of the end effector ,Location ,speed These parameters reflect the changes in the internal motion state of the load in real time.

[0026] Step 3: Establish the interference transmission relationship between the load and the robot; Based on the aforementioned dynamic geometric features and active dynamic state, a real-time interactive disturbance force model is established for the load on the quadruped robot body. Specifically, when the robotic arm joints press... , , During movement, changes in the pose and velocity of the end effector cause real-time shifts in the overall system's center of mass and inertia, which are then transmitted dynamically to generate time-varying disturbance forces and torques at the robot's base. This disturbance force is the input F to the subsequent hybrid force-position control model.

[0027] Step 4: Establish a parameterized representation for control; Ultimately, the active dynamic load is modeled as a set of parameters. .

[0028] The first two parameters describe the overall geometric and mass characteristics of the load, while the latter six describe its internal active motion state. This parameterized characterization provides complete input information for the subsequent steps of calculating the disturbance force compensation amount using the pre-set hybrid force-position control model.

[0029] S103. Actively correct the motion commands of the quadruped robot by using a preset hybrid force-position control model and real-time interactive interference force to generate anti-interference target commands.

[0030] In this embodiment of the invention, a preset hybrid force-position control model is constructed. The preset hybrid force-position control model uses real-time interactive interference force to construct active physical compensation for interference force, guides the strategy to learn the anti-resistance action adjustment under active dynamic load disturbance, and generates anti-interference target instructions.

[0031] S104. Input the motion command into the preset motion control strategy network, and use the anti-interference target command as the guiding benchmark of the reward function in the preset motion control strategy network to generate actuator motion commands for driving the quadruped robot, and maintain the balance of the quadruped robot in the process of carrying active dynamic loads in complex terrain through the actuator motion commands.

[0032] In this embodiment of the invention, a preset motion control strategy network is constructed. The input to the preset motion control strategy network is the motion command output by the preset hybrid force-position control model. Then, the anti-interference target command is used as the guiding benchmark for the reward function in the preset motion control strategy network to generate actuator motion commands for driving the quadruped robot, i.e., the output is the motion command. Then, the quadruped robot is controlled to execute actuator action commands to maintain the balance of the quadruped robot during the process of carrying active dynamic loads in complex terrain.

[0033] Compared with existing technologies, this embodiment provides a method for acquiring the body perception sequence of a quadruped robot; extracting features from the body perception sequence using a preset active dynamic load model and quantifying key features of the active dynamic load carried by the quadruped robot; key features include the real-time interactive interference force generated by the active dynamic load on the quadruped robot, thereby parametrically describing the active and dynamic characteristics of the load through the preset active dynamic load model, achieving adaptive perception of load changes and real-time interference forces; and actively correcting the quadruped robot's motion commands through a preset hybrid force-position control model and real-time interactive interference forces to generate anti-interference target commands, thereby achieving adaptive perception of load changes and real-time interference forces through hybrid force-position control. Based on perceived load disturbance information, the control system guides the robot to actively correct its motion commands and constructs an effective anti-interference compensation mechanism. The motion commands are input into a preset motion control strategy network, and the anti-interference target command is used as the guiding benchmark for the reward function in the preset motion control strategy network to generate actuator motion commands to drive the quadruped robot. The actuator motion commands are used to maintain the balance of the quadruped robot in the process of actively carrying dynamic loads in complex terrain. Thus, the actuator motion commands are output through the preset motion control strategy network to achieve adaptive adjustment of the quadruped robot and improve its balance control capability and anti-interference robustness in the active dynamic load carrying task in complex terrain.

[0034] In some embodiments of the present invention, the motion command includes a velocity command and a pose command; the anti-interference target command includes a target pose command and a target velocity command; step S103 includes: Based on the preset hybrid force-position control model and the torque component in the real-time disturbance force, the pose command is corrected and the target pose command is generated. Based on the preset hybrid force-position control model and the force component correction speed command in the real-time disturbance force, the target speed command is generated.

[0035] Impedance control in this invention has been proven to effectively achieve compliant and force-adaptive behavior in robots. Inspired by this success, the principle of impedance control is extended to the active dynamic load tasks of quadruped robots. First, the general problem modeling of this method is explained: given a position command relative to the robot's body coordinate system... and speed command The objective of this invention is to learn a reinforcement learning strategy to ensure the robot operates under disturbances. Under these instructions, the impedance control model is adopted as shown in equation (1): (1) In the formula, , , , These represent the robot's actual position and velocity, and its desired position and velocity, respectively. , These represent the stiffness coefficient and damping coefficient, respectively. Since the robot's motion is controlled by speed commands... ) They represent xy The axial velocity and z-axis angular velocity are used as the basis for the description of the robot's position as its pose. These represent the target altitude, pitch angle, and roll angle, respectively. (Disturbance force) Let F represent the force and torque exerted by the load on the quadrupedal base, respectively. In load resistance control, we use force to guide velocity correction and torque to guide pose correction, as shown in formula (2): (2) Among them, the actual pose command maintains a constant. Then the corrections for the target pose command and the target velocity command are as shown in formulas (3) and (4): (3) (4) In the formula, For the target pose command, For the target speed command, As a preset constant value, The torque component in the real-time disturbance force. For the force component in real-time disturbance force, and These are the stiffness coefficient and damping coefficient, respectively. This is the preset speed command.

[0036] Based on formulas (3) and (4), a strategy learning model based on a hybrid force-position control mechanism was constructed. Specifically, it is a hybrid force-position correction reward model that considers the effects of force and torque to optimize the robot's speed command and posture.

[0037] In some embodiments of the present invention, such as Figure 3 As shown, the training process of the preset motion control policy network is as follows: S301. Obtain the historical ontological perception observation sequence, environmental privilege information, active dynamic load characteristic information, and real disturbance force information of the quadruped robot in multiple training rounds.

[0038] In this embodiment of the invention, the training of the motion strategy employs a combination of supervised learning and reinforcement learning. Reinforcement learning utilizes the Proximal Policy Optimization (PPO) algorithm, combining the loss function of PPO with that of supervised learning to update the policy network parameters. Supervised learning helps the environmental encoder extract key information about the active dynamic load characteristics and predict potential environmental disturbances. Specifically, the active dynamic load problem of a quadruped robot on complex terrain is further modeled as a partially observable Markov decision process (POMDP), denoted as... ,in , ,O Representing state, action, and observation space respectively, in time At that time, the state changed from This indicates that the agent selects actions. transfer function Decide the next state The reward for this state and action is determined by... Given. However, the agent cannot obtain the complete environmental state, but instead obtains it through an observation function. Received observations The goal of an intelligent agent is to learn a policy. , to make the expected total discount reward Maximize, where This is the discount factor.

[0039] Preset observation space and preset state space: Robot body perception observations in the preset observation space Data can be collected through the body's sensing sensors, including: body angular velocity. Projected gravity Speed ​​command Joint angle Joint velocity and the previous step The states in the preset state space. Includes ontological observations Environmental privilege information Robotic arm load characteristics and load interference force The environmental privilege information includes the base linear velocity. Topographic elevation map and joint forces The robotic arm load information is derived from [robotic arm joint positions]. ,speed End effector position ,speed and total mass ] The system is composed of the aforementioned preset observation space and preset state space. This allows the acquisition of the quadruped robot's historical ontological perception observation sequences, environmental privilege information, active dynamic load characteristic information, and real disturbance force information across multiple training rounds.

[0040] S302. The historical ontology perception observation sequence is encoded using an environmental encoder network to generate environmental latent variables, base linear velocity estimates, and load disturbance force estimates at different times.

[0041] In this embodiment of the invention, the environmental encoder processes historical ontology-aware observation sequences. generate in Reconstruction is generated through the environment decoder. , 、 The system was trained separately to estimate the base linear velocity and load disturbance force, and the environmental latent variables, base linear velocity estimates, and load disturbance force estimates at different times were obtained.

[0042] S303. Using the proprioceptive observation value, environmental latent variables, base linear velocity estimate, and load disturbance force estimate in the historical proprioceptive observation sequence as the input to the preset motion control strategy network, the preset motion control strategy network is updated with parameters using a reinforcement learning algorithm to obtain the trained preset motion control strategy network.

[0043] In this embodiment of the invention, the proprioceptive observation value at the current moment, the environmental latent variable, the base linear velocity estimate, and the load disturbance force estimate in the historical proprioceptive observation sequence can be used as the input of the preset motion control strategy network to obtain the output action. Then, the parameters of the preset motion control strategy network are updated using a reinforcement learning algorithm to obtain the trained preset motion control strategy network.

[0044] In some embodiments of the present invention, such as Figure 4 As shown, step S301 includes: S401. The environmental encoder network is optimized by using the supervised learning loss function, the base linear velocity estimate, and the load disturbance force estimate to obtain the mean square error loss.

[0045] In this embodiment of the invention, the loss function of supervised learning includes two estimation loss terms and one reconstruction loss term. During the training process, mean squared error (MSE) loss is used for optimization. The supervised learning loss function is shown in formula (5): (5) In the formula, For mean square error loss, and These are the actual value and the estimated value of the base linear velocity, respectively. and These are the actual value and the estimated value of the load disturbance force, respectively. and They are respectively t Proprioceptive observations at time +1 and reconstructed values ​​from the environment decoder. To balance the weights of different loss terms, This is the expected operation.

[0046] S402. Input the proprioceptive observation value, environmental latent variable, base linear velocity estimate, and load disturbance force estimate of the current moment in the historical proprioceptive observation sequence into the preset motion control strategy network, and output the current action.

[0047] In this embodiment of the invention, a preset motion control strategy network is used based on the current proprioceptive observations. and environmental encoder output As input, output the current action. The current action in the action space. This represents the expected 12-dimensional joint torque applied to the actuator relative to the initial posture, and the current motion. As shown in formula (6): (6) S403. Input the current proprioceptive observation, environmental privilege information, active dynamic load characteristic information, and real disturbance force information into the preset evaluation network to obtain the current value.

[0048] In this embodiment of the invention, a preset evaluation network is used to... Privilege information Real robotic arm load characteristics information and real interference information As input, output the current value. As shown in formula (7): (7) S404. Calculate the current action and current value based on the preset reward function to obtain the current advantage estimate.

[0049] In this embodiment of the invention, a preset reward function can be set. The preset reward function used during training is shown in Table 1. Specifically, the reward items used for correction according to the hybrid force-potential control model include linear velocity and angular velocity tracking rewards, target height penalties, and horizontal attitude penalties.

[0050] Table 1. Reward Function and Weights

[0051] In some embodiments of the present invention, step S404 includes: The parameters of the current action are transformed according to the preset state space to obtain the current state value.

[0052] In this embodiment of the invention, the parameters of the current action are transformed according to the default values ​​in the preset state space, that is, the current action is determined to control the robotic arm to move, and then the current state value of the robotic arm is obtained, such as state. Includes ontological observations Environmental privilege information Robotic arm load characteristics and load interference force The environmental privilege information includes the base linear velocity. Topographic elevation map and joint forces The robotic arm load information is derived from [robotic arm joint positions]. ,speed End effector position ,speed and total mass ] composition.

[0053] Based on the preset observation space, the robot performs body perception observation and data collection on the current action to obtain the current observation value.

[0054] In this embodiment of the invention, sensors in a pre-defined observation space can be used to perform robot-body perception and observation data acquisition to obtain current observation values, such as body angular velocity. Projected gravity Speed ​​command Joint angle Joint velocity and the previous step .

[0055] The current reward value is obtained by calculating the current state value and the current observation value using a preset reward function.

[0056] In this embodiment of the invention, the preset reward function includes a force-position correction reward term constructed based on a preset hybrid force-position control model. The force-position correction reward term compensates for motion commands based on real disturbance force information to guide the preset motion control strategy network to learn anti-interference behavior. The preset reward function is shown in the formula in Table 1. Substituting the current state value and the current observation value into the formula for calculation yields the corresponding current reward value.

[0057] Based on the current reward value and the current value, the current advantage estimate is obtained.

[0058] In this embodiment of the invention, the current reward value and the current value can be calculated to obtain the current advantage estimate. .

[0059] S405. The parameters of the preset motion control strategy network are updated using a reinforcement learning algorithm and the current advantage estimate to obtain the reinforcement learning loss.

[0060] In this embodiment of the invention, a reinforcement learning algorithm and the current advantage estimate can be used to update the parameters of the preset motion control strategy network to obtain the reinforcement learning loss. The reinforcement learning loss is calculated as shown in formula (8): (8) In the formula, This represents the probability ratio between the old and new strategies. This represents the corresponding advantage estimate. This indicates the clipping threshold.

[0061] S406. Obtain the total policy training loss at the current moment based on the mean squared error loss and the reinforcement learning loss.

[0062] In this embodiment of the invention, the total loss of policy training is the sum of the supervised learning loss and the PPO loss, as shown in formula (9): (9) S407. Based on the gradient descent method, the total training loss of the policy at adjacent time points is judged to obtain the preset motion control policy network after training is completed.

[0063] In this embodiment of the invention, after training the preset motion control strategy network at the current time t, the preset motion control strategy network can be trained again using the data at the next time t+1 to obtain the total training loss of the strategy. Similarly, the total training loss of adjacent strategies is judged by the gradient descent method. When the total training loss of the strategy tends to a stable value, it is determined that the training of the preset motion control strategy network is complete.

[0064] Furthermore, all network modules are designed as multilayer perceptrons (MLPs) with exponential linear units (ELUs) as activation functions. Details are shown in Table 2. Table 2. Network Structure

[0065] This invention employs the physics simulation platform Isaac Sim and the reinforcement learning framework Isaac Lab for training and evaluating the policy network. The policy uses 4096 agents for parallel training, with a historical data size of H = 6, and runs on an Nvidia RTX 4090 GPU. The specific training hyperparameters are shown in Table 1. Table 3. Training Hyperparameters

[0066] A curriculum-based learning approach was employed to facilitate the robot's progressive learning of its active load-bearing capacity on complex terrain. The training curriculum comprised three parts: terrain training, speed training, and load-bearing training. The terrain training included: staircases ranging from [0.05 to 0.15] meters, and slopes ranging from [0...] meters... ° , 20 ° The maximum xy linear velocity is [-1.0, 1.0] m / s and the maximum z angular velocity is [-0.4, 0.4] radians / s, respectively, with learning ranges of [-0.1, 0.1] m / s and [-0.1, 0.1] radians / s. The maximum angular velocity is [-0.4, 0.4] radians / s, respectively, with learning ranges of [-0.1, 0.1] m / s and [-0.1, 0.1] radians / s, respectively, with the maximum angular velocity being [-0.0, 3.0] kg, and the maximum angular velocity is [-0.1, 0.1] kg, respectively. The maximum angular velocity is [-0.0, 3.0] kg, with learning ranges of [-0.1, 0.1] kg, with the maximum angular velocity being [-0.0, 0.0] m / s and the maximum angular velocity is [-0.4, 0.4] radians / s, respectively, with learning ranges of [-0.1, 0.1] kg.

[0067] To train the methods used in this study and narrow the gap between simulation results and actual conditions, a domain randomization method was employed. Specifically, as shown in Table 4, this included randomization of the general robot body and additional design features with active dynamic loads, including different actions and unknown robotic arm loads.

[0068] Table 4. Domain Randomization Settings

[0069] This invention proposes a deep reinforcement learning (DALFPDRL) framework based on active dynamic load modeling and hybrid force-position control. First, active dynamic load modeling parameterizes the active and dynamic characteristics of the load, enabling adaptive perception of load changes and real-time disturbances. Second, hybrid force-position control, based on the perceived load disturbance information, guides the robot to actively adjust its speed and posture through a reward mechanism, constructing an effective anti-interference compensation mechanism. Finally, the deep reinforcement learning framework learns control strategies and, using only the robot's historical perception information, predicts environmental features, load dynamics, and disturbances, thereby achieving adaptive adjustment of the robot's speed and posture and improving its balance control capability and anti-interference robustness in active dynamic load handling tasks in complex terrain.

[0070] To better implement the quadruped robot hybrid force-position control method in this embodiment of the invention, correspondingly, this embodiment of the invention also provides a quadruped robot hybrid force-position control device, such as... Figure 5 As shown, the quadruped robot hybrid force-position control device 500 includes: The perception acquisition module 501 is used to acquire the body perception sequence of the quadruped robot; The dynamic load quantization module 502 is used to extract features from the ontology perception sequence and quantify the key features of the active dynamic load carried by the quadruped robot through a preset active dynamic load model; the key features include the real-time interactive interference force generated by the active dynamic load on the quadruped robot. The active correction module 503 is used to actively correct the motion commands of the quadruped robot by using a preset hybrid force-position control model and real-time interactive interference force to generate anti-interference target commands. The motion execution module 504 is used to input motion commands into a preset motion control strategy network, and use the anti-interference target command as the guiding reference for the reward function in the preset motion control strategy network to generate actuator motion commands for driving the quadruped robot, and maintain the balance of the quadruped robot in the process of carrying active dynamic loads in complex terrain through the actuator motion commands.

[0071] The quadruped robot hybrid force position control device 500 provided in the above embodiments can realize the technical solutions described in the above quadruped robot hybrid force position control method embodiments. The specific implementation principles of each module or unit can be found in the corresponding content in the above quadruped robot hybrid force position control method embodiments, and will not be repeated here.

[0072] The above provides a detailed description of the quadruped robot hybrid force-position control method and device provided by the present invention. Specific examples have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of ​​the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A hybrid force-position control method for a quadruped robot, characterized in that, include: Obtain the proprioception sequence of the quadruped robot; The key features of the active dynamic load carried by the quadruped robot are extracted and quantified by a preset active dynamic load model. The key features include the real-time interactive interference force generated by the active dynamic load on the quadruped robot. The motion commands of the quadruped robot are actively corrected by a preset hybrid force-position control model and the real-time interactive interference force to generate anti-interference target commands. The motion command is input into a preset motion control strategy network, and the anti-interference target command is used as the guiding benchmark for the reward function in the preset motion control strategy network to generate actuator motion commands for driving the quadruped robot. The actuator motion commands are used to maintain the balance of the quadruped robot in the process of carrying the active dynamic load in complex terrain.

2. The hybrid force-position control method for a quadruped robot according to claim 1, characterized in that, The active dynamic load is a load system with active motion capability; the key characteristic quantities also include dynamic geometric features and active dynamic state, the dynamic geometric features include mass and center of mass position, and the active dynamic state includes joint position, joint velocity and joint force.

3. The hybrid force-position control method for a quadruped robot according to claim 1, characterized in that, The motion commands include velocity commands and pose commands; the anti-interference target commands include target pose commands and target velocity commands; the active correction of the quadruped robot's motion commands through a preset hybrid force-position control model and the real-time interactive interference force to generate anti-interference target commands includes: Based on the preset hybrid force-position control model and the torque component in the real-time disturbance force, the pose command is corrected to generate the target pose command. Based on the preset hybrid force-position control model and the force components in the real-time disturbance force, the speed command is corrected to generate the target speed command.

4. The hybrid force-position control method for a quadruped robot according to claim 1, characterized in that, The training process of the preset motion control strategy network is as follows: Acquire the historical ontological perception observation sequence, environmental privilege information, active dynamic load characteristic information, and real disturbance force information of the quadruped robot in multiple training rounds; The historical ontology perception observation sequence is encoded using an environmental encoder network to generate environmental latent variables, base linear velocity estimates, and load disturbance force estimates at different times. Using the proprioceptive observation value at the current moment in the historical proprioceptive observation sequence, the environmental latent variable, the base linear velocity estimate, and the load disturbance force estimate as inputs to the preset motion control strategy network, a reinforcement learning algorithm is used to update the parameters of the preset motion control strategy network to obtain the trained preset motion control strategy network.

5. The hybrid force-position control method for a quadruped robot according to claim 4, characterized in that, The process involves using the proprioceptive observation value at the current moment in the historical proprioceptive observation sequence, the environmental latent variables, the estimated base linear velocity, and the estimated load disturbance force as inputs to the preset motion control strategy network. A reinforcement learning algorithm is then used to update the parameters of the preset motion control strategy network to obtain a trained preset motion control strategy network, including: The environmental encoder network is optimized by using a supervised learning loss function, the base linear velocity estimate, and the load disturbance force estimate to obtain the mean square error loss. The proprioceptive observation value at the current moment in the historical proprioceptive observation sequence, the environmental latent variable, the base linear velocity estimate, and the load disturbance force estimate are input into the preset motion control strategy network, and the current action is output. The proprioceptive observation value at the current moment in the historical proprioceptive observation sequence, the environmental privilege information, the active dynamic load characteristic information, and the real interference force information are input into a preset evaluation network to obtain the current value; The current advantage estimate is obtained by calculating the current action and the current value based on a preset reward function; The parameters of the preset motion control strategy network are updated using a reinforcement learning algorithm and the current advantage estimate to obtain the reinforcement learning loss. The total policy training loss at the current moment is obtained based on the mean squared error loss and the reinforcement learning loss. The total training loss of the strategy at adjacent time points is determined by the gradient descent method to obtain the pre-set motion control strategy network after training.

6. The hybrid force-position control method for a quadruped robot according to claim 5, characterized in that, The step of calculating the current advantage estimate based on the current action and the current value using a preset reward function includes: The current action is parameterized according to a preset state space to obtain the current state value; Based on a preset observation space, the robot body perception observation and collection of the current action are performed to obtain the current observation value; The current reward value is obtained by calculating the current state value and the current observation value using the preset reward function; Based on the current reward value and the current value, the current advantage estimate is obtained.

7. The hybrid force-position control method for a quadruped robot according to claim 5, characterized in that, The preset reward function includes a force-position correction reward term constructed based on the preset hybrid force-position control model; the force-position correction reward term is used to compensate the motion command according to the real interference force information, so as to guide the preset motion control strategy network to learn anti-interference behavior.

8. The hybrid force-position control method for a quadruped robot according to claim 5, characterized in that, The supervised learning loss function is: In the formula, For mean square error loss, and These are the actual value and the estimated value of the base linear velocity, respectively. and These are the actual value and the estimated value of the load disturbance force, respectively. and They are respectively t Proprioceptive observations at time +1 and reconstructed values ​​from the environment decoder. To balance the weights of different loss terms, This is the expected operation.

9. The hybrid force-position control method for a quadruped robot according to claim 1, characterized in that, The correction formulas for the target pose command and the target velocity command are as follows: In the formula, For the target pose command, For the target speed command, X This is a preset constant value. The torque component in the real-time disturbance force. For the force component in real-time disturbance force, and These are the stiffness coefficient and damping coefficient, respectively. This is the preset speed command.

10. A hybrid force-position control device for a quadruped robot, characterized in that, include: The perception and acquisition module is used to acquire the body perception sequence of the quadruped robot; The dynamic load quantification module is used to extract features from the ontology perception sequence and quantify the key features of the active dynamic load carried by the quadruped robot through a preset active dynamic load model; the key features include the real-time interactive interference force generated by the active dynamic load on the quadruped robot. An active correction module is used to actively correct the motion commands of the quadruped robot by using a preset hybrid force-position control model and the real-time interactive interference force, and generate anti-interference target commands. The motion execution module is used to input the motion command into a preset motion control strategy network, and use the anti-interference target command as the guiding reference for the reward function in the preset motion control strategy network to generate actuator motion commands for driving the quadruped robot, and maintain the balance of the quadruped robot in the process of carrying the active dynamic load in complex terrain through the actuator motion commands.