Falling self-recovery control method and system for quadruped robot under complex terrain

Through deep reinforcement learning technology and teacher-student network model, combined with NP3O algorithm, the four-legged robot can independently recover and stand in complex terrain, solving the problem that traditional control methods rely on external intervention and achieving efficient and safe autonomous recovery capabilities.

CN120215301APending Publication Date: 2025-06-27SHANDONG YOUBAOTE INTELLIGENT ROBOTICS CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510341118.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Four-legged robots are prone to fall due to external interference, terrain unevenness or dynamic obstacles in complex terrain. Traditional control methods rely on external intervention, limit their application capabilities, and lack dynamic recovery strategies based on terrain characteristics.

Method used

Deep reinforcement learning technology is adopted, and through the teacher-student network model and NP3O algorithm, the Actor network, Critic network and Cost network are integrated into a unified optimization framework to realize the four-legged robot automatically restores to a standing posture after falling, and dynamically plan the optimal stand-up movement based on terrain characteristics.

Benefits of technology

Four-legged robots can independently recover and stand in complex terrain, significantly improving recovery efficiency, improving generalization capabilities, reducing dependence on external interventions, and improving the safety and robustness of the movement process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215301A_ABST
    Figure CN120215301A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of robot control, and provides a quadruped robot fall-down self-recovery control method and system under complex terrains. The method comprises the following steps: extracting a tumble self-recovery strategy feature from privilege information subjected to normalization processing by using a teacher network, so as to provide an optimization reference for a student network; learning strategy characteristics simulating a teacher network by using a student network so as to process the normalized historical observation information of the ontology state information; after the strategy converges, the Actor network is used for processing the output characteristics of the student network and the real-time observation information of the normalized body state information, the next state observation information, the reward signal, the cost feedback and the task end mark are obtained through interaction with the simulation environment, and the action of the quadruped robot is output; evaluating the value of adopting a specific action in the current state; and processing the privilege information subjected to normalization processing by using a Cost network, and evaluating the cost of the current action of the quadruped robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robot control, and particularly relates to a control method and system for a quadruped robot to self-recover from a fall in complex terrains. Background Art

[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.

[0003] In an unstructured environment (such as steep slopes, gravel, stairs, etc.), a quadruped robot is prone to imbalance and even fall due to external interference, uneven terrain, or dynamic obstacles. Traditional motion control methods usually rely on external intervention to restore the standing posture after the quadruped robot falls, so as to continue to execute tasks. This not only easily leads to task interruption or failure, but also severely limits the practical application ability of the quadruped robot.

[0004] Deep reinforcement learning technology mainly focuses on improving the motion planning and balance control capabilities of robots, but there is less research on autonomous recovery after falling, and there is a lack of dynamic recovery strategies based on terrain features. During the process of falling and recovering to a standing position, how to plan the optimal getting-up action according to terrain conditions while taking into account speed, energy consumption, and stability is a problem to be solved in the current technical field. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a control method and system for a quadruped robot to self-recover from a fall in complex terrains. By using deep reinforcement learning technology, the quadruped robot can not only autonomously recover to a standing posture after falling in complex terrains, but also dynamically plan the optimal getting-up action according to terrain features.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The first aspect of the present invention provides a control method for a quadruped robot to self-recover from a fall in complex terrains.

[0008] In one or more embodiments, a control method for a quadruped robot to self-recover from a fall in complex terrains is provided, including:

[0009] Obtain real-time observation information, historical observation information of the body state information, and privileged information of the quadruped robot body and perform normalization processing;

[0010] Use the teacher network to extract the fall self-recovery strategy features from the normalized privileged information to provide optimization reference for the student network; use the student network to learn the strategy features imitating the teacher network to process the normalized historical observation information of the body state information;

[0011] After the policy converges, the Actor network processes the real-time observation information of the output features of the student network and the normalized ontology state information. By interacting with the simulation environment, it obtains the next state observation information, reward signal, cost feedback, and task end flag, and outputs the actions of the quadruped robot.

[0012] The Critic network processes the normalized privileged information to evaluate the value of taking specific actions in the current state, so as to guide the update of the Actor network.

[0013] The Cost network processes the normalized privileged information to evaluate the cost of the current actions of the quadruped robot, which is used to constrain the update of the fall self-recovery strategy.

[0014] As an implementation manner, the reward signal is obtained by weighted summation of the weights corresponding to the key reward function, walking-related reward function, and safety constraint reward function during the process of falling and recovering to the standing posture.

[0015] As an implementation manner, the cost feedback is calculated by the constraint cost function of joint position, torque, and speed.

[0016] As an implementation manner, according to the actions of the quadruped robot output by the Actor network, after being superimposed with the initial positions of the current joints, the target joint positions are generated, and then the output torques are calculated, and finally the actions are executed.

[0017] As an implementation manner, the real-time observation information of the ontology state information includes the angular velocity of the body, the projected gravity information, the current desired three-dimensional motion command, the offset of the current joint position relative to the default joint position, the current speed of each joint, and the action value output by the Actor network.

[0018] As an implementation manner, the privileged information includes the observation of the robot ontology information, the linear velocity of the body, the difference information between the current root height and the measured terrain height, and the randomization parameters.

[0019] As an implementation manner, the NP3O algorithm is used to integrate the Actor network, Critic network, and Cost network into a unified optimization framework, combining the goals of policy optimization, value evaluation, and cost evaluation to ensure maximizing the cumulative reward under the condition of meeting safety constraints; the final optimization goal of NP3O is expressed as:

[0020]

[0021] Where:

[0022]

[0023] is the loss function of the Actor network, which is used to optimize the policy. κ is the penalty factor, which is used to control the penalty intensity of constraint violation; is the probability ratio of the old and new policies, is the normalized advantage function provided by the Critic network, is the normalized cost advantage function provided by the Cost network, is the expected cost return of the current policy, d i is the upper limit of the cost constraint, and are the mean and standard deviation of the cost advantage function respectively. ∈ is a hyperparameter, which is used to control the amplitude of policy update; represents under the policy π k according to the state distribution sampling the state s, and then sampling the action a according to the policy π k to obtain the expectation.

[0024] As an implementation, the update objectives of the Critic network and the Cost network also include: the Critic network predicts the state value V φ (s), minimizes the difference between the value function V φ (s) and the actual return R, so as to improve the evaluation ability of the state value; the Cost network predicts the state cost C ψ (s), minimizes the difference between the cost function C ψ (s) and the actual environment cost C actual to enhance the prediction ability of the cost; the objective function is expressed as:

[0025]

[0026] wherein, represents the expectation obtained by sampling the state s according to the state distribution D.

[0027] The second aspect of the present invention provides a four-legged robot fall self-recovery control system under complex terrain.

[0028] In one or more embodiments, a four-legged robot fall self-recovery control system under complex terrain includes:

[0029] An information processing module, which is used to obtain real-time observation information, historical observation information of the body state information of the four-legged robot, and privileged information and perform normalization processing;

[0030] A student network learning module, which is used to extract the falling self-recovery strategy features from the normalized privilege information by using a teacher network, so as to provide an optimization reference for the student network; and use the student network to learn the strategy features that imitate the teacher network, so as to process the normalized historical observation information of the ontology state information.

[0031] An action output module, which is used to, after the strategy converges, use the Actor network to process the output features of the student network and the real-time observation information of the normalized ontology state information, and obtain the next state observation information, reward signal, cost feedback and task end flag by interacting with the simulation environment, and output the quadruped robot actions.

[0032] An action value evaluation module, which is used to process the normalized privilege information by using the Critic network, evaluate the value of taking a specific action in the current state, so as to guide the update of the Actor network.

[0033] An action cost evaluation module, which is used to process the normalized privilege information by using the Cost network, evaluate the cost of the current action of the quadruped robot, and is used to constrain the update of the falling self-recovery strategy.

[0034] The third aspect of the present invention provides a computer-readable storage medium.

[0035] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it realizes the steps in the method for controlling the falling self-recovery of a quadruped robot in complex terrain as described above.

[0036] The fourth aspect of the present invention provides a quadruped robot.

[0037] A quadruped robot, including a robot body, on which a memory, a processor and a computer program stored on the memory and executable on the processor are provided, and when the processor executes the program, it realizes the steps in the method for controlling the falling self-recovery of a quadruped robot in complex terrain as described above.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] (1) By introducing a teacher-student network model, the present invention reduces the dependence on external intervention, enables the quadruped robot to autonomously learn and optimize the recovery strategy after falling, the student network learns efficient strategies from the teacher network, significantly improves the recovery efficiency, and at the same time enhances the generalization ability.

[0040] (2) The present invention adopts the NP3O (Normalized Penalized Proximal Policy Optimization) algorithm to integrate the Actor network, Critic network, and Cost network into a unified optimization framework. Combining the objectives of policy optimization, value evaluation, and cost evaluation, it ensures the maximization of cumulative rewards under the condition of meeting safety constraints, and optimizes the generation of policies. Compared with the traditional PPO (Proximal Policy Optimization) algorithm, the NP3O algorithm overcomes the challenges of complex constraint problems through the normalization process of the reward function, enabling the robot to stably generate high-quality actions in complex terrains.

[0041] (3) The design of the reward function of the present invention emphasizes the priority and constraints of the upright state, prompting the robot to optimize its posture in a stable and safe manner during fall recovery, reducing the number of falls and instabilities in complex terrains, and improving the safety of the movement process. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0043] Figure 1 is a schematic flow chart of the control method for a quadruped robot to self-recover from a fall under complex terrains in an embodiment of the present invention;

[0044] Figure 2 is a training framework diagram of the deep reinforcement learning algorithm in an embodiment of the present invention

[0045] Figure 3 is a deployment model of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0047] It should be noted that the following detailed descriptions are all illustrative and are intended to provide a further description of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0048] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0049] Figure 1 It is a schematic flow chart of a method for controlling the self - recovery of a quadruped robot under complex terrain in an embodiment of the present invention. As Figure 1 shown, the method for controlling the self - recovery of a quadruped robot under complex terrain in this embodiment may include:

[0050] Step 1: Obtain real - time observation information, historical observation information of the body state information, and privilege information of the quadruped robot and perform normalization processing;

[0051] Among them, the observation information includes real - time observation o t of the body state information, historical observation of the body state information, and privilege information s t . These observation information provide comprehensive context support and enhance the timeliness and predictability of decision - making.

[0052] The observation of the body information includes the angular velocity of the body, the projected gravity information, the current desired three - dimensional motion commands (such as linear velocity, angular velocity), the offset of the current joint position relative to the default joint position, the current speed of each joint, and the action value output by the policy network.

[0053] The historical observation of the body state information is formed by adding the observation information of the previous time steps, which is used to form a time series and enhance the timeliness and predictability of decision - making.

[0054] The observation of the privilege information includes the observation of the robot body information, the linear velocity of the body, the height difference between the current root height and the measured terrain height, etc. In addition, it also includes a plurality of randomized parameters, such as friction coefficient, recovery coefficient, joint control gain, motor strength, mass center of mass parameters, etc.

[0055] Step 2: Use the teacher network to extract the self - recovery strategy features for falling from the normalized privilege information to provide an optimization reference for the student network; use the student network to learn the strategy features imitating the teacher network to process the normalized historical observation information of the body state information;

[0056] Step 3: When the policy converges, the Actor network processes the real-time observation information of the output features of the student network and the normalized ontology state information. By interacting with the simulation environment, the next state observation information, reward signal, cost feedback, and task end flag are obtained, and the quadruped robot actions are output;

[0057] Among them, the reward signal is obtained by weighted summation of the weights corresponding to the key reward function during the process of falling and recovering to the standing posture, the walking-related reward function, and the safety constraint reward function.

[0058] The reward functions in the environment can be classified into three major parts: the key reward function during the process of falling and recovering to the standing posture, the walking-related reward function, and the safety constraint reward function. All reward functions will be calculated simultaneously and weighted and summed according to their corresponding weights.

[0059] The key reward function during the process of falling and recovering to the standing posture is shown in Table 1.

[0060] Table 1 Key Reward Function for Self-Recovery from Falling

[0061] Reward Name Equation Weight upward <![CDATA[1-g z > 0.7 stand_still <![CDATA[|θ - θ d |·(1 - g z )·(||u||2 < 0.15)]]> -1.0 base_height <![CDATA[(h target -h) 2 ·clip(-g z ,0,1)]]> -5.0

[0062] Among them, h target is the target body height; h represents the depth height; θ represents the joint angle; θ d represents the set joint angle; in the design of the reward function, the reward upward for the upright body state is particularly introduced, where g z is the z-axis component of the projected gravity information, aiming to prompt the robot to prioritize optimizing the ability to maintain upright during the process of falling and recovering. In addition, in order to enhance the robot's ability to recover to a more accurate standing posture after falling, the body height reward function base_height and the joint position reward function stand_still are also added. The body height reward function adds the constraint clip(-g z , 0, 1) for the upright body state. The expression of clip(-g z , 0, 1) is as shown in Equation 1:

[0063]

[0064] The joint position reward function adds the constraints 1 - g z and the command constraint ||u||2 < 0.15, where u represents the three-dimensional velocity command (the body linear velocity and body angular velocity in the xy direction) issued to the quadruped robot.

[0065] In order to make the robot more robust during movement, the walking-related reward function also introduces the constraint of the upright body state, as shown in Table 2.

[0066] Table 2 Walking-related Reward Function

[0067]

[0068] Wherein:

[0069] tracking_lin_vel: Linear velocity tracking error;

[0070] tracking_ang_vel: Angular velocity tracking error;

[0071] lin_vel_z: Z-axis linear velocity;

[0072] ang_vel_xy: XY-plane angular velocity;

[0073] Orientation: Attitude;

[0074] Collision: Collision;

[0075] trot_phase: Gait phase;

[0076] feet_air_time: Foot air time;

[0077] force_arrangement: Force distribution.

[0078] Wherein, represents the desired xy-direction linear velocity command of the robot; v xy represents the xy-direction linear velocity of the robot base; represents the desired yaw angular velocity; ω yaw represents the yaw angular velocity of the robot base; σ represents the standard deviation of the control tracking error; ν z represents the z-direction velocity of the robot base; ω x represents the x-direction angular velocity of the robot base; ω y represents the y-direction angular velocity of the robot base; g x represents the projection of gravity in the x-axis direction of the robot base; g y represents the projection of gravity in the y-axis direction of the robot base; f penalised represents the force that will be penalized for collision; f feet (0, 1, 2, 3) represents the contact states of the left front foot, right front foot, left hind foot, and right hind foot; f feet (1, 0.3, 2) represents the contact states of the right front foot, left front foot, right hind foot, and left hind foot; represents the exclusive OR operation; θ (FL,LR) represents the joint positions of the left front leg and right front leg of the robot; θ (RR,RL)Indicates the joint positions of the robot's right hind leg and left hind leg; Indicates the time the foot is in the air; c foot Indicates whether all the feet of the robot are fully in contact with the ground; F left_diag Indicates the resultant force of the left front leg and the right hind leg; F right_diag Indicates the resultant force of the left hind leg and the right front leg.

[0079] During the movement of the robot, it follows some physical limitations. To ensure the safety and stability of the robot's behavior, a safety constraint reward function is set, as shown in Table 3.

[0080] Table 3 Safety Constraint Reward Function

[0081]

[0082]

[0083] Among them, dof_vel: joint velocity;

[0084] torques: joint torque;

[0085] dof_acc: joint acceleration;

[0086] action_rate: action change rate;

[0087] action_smoothness: action smoothness;

[0088] dof_pos_limits: joint position limits;

[0089] dof_vel_limits: joint velocity limits;

[0090] torque_limits: joint torque limits.

[0091] τ represents the joint torque of the robot;

[0092] θ min Indicates the minimum value of the robot's joint position;

[0093] θ max Indicates the maximum value of the robot's joint position;

[0094] τ max Indicates the maximum of the robot's joint torque.

[0095] The cost feedback is calculated by the cost functions of the joint position, torque, and velocity limits.

[0096] The cost function in the environment sets the cost functions for joint position, torque, and speed to ensure policy safety and protect the hardware device. The cost function formula is expressed as follows:

[0097] cost_pos_limits = max(0, θ - θ min ) + max(0, θ max - θ) (2)

[0098]

[0099] cost_torque_limits = max(0, |τ| - τ max ) (4)

[0100] Step 4: Use the Critic network to process the normalized privileged information, evaluate the value of taking a specific action in the current state, and guide the update of the Actor network;

[0101] Step 5: Use the Cost network to process the normalized privileged information, evaluate the cost of the current action of the quadruped robot, and use it to constrain the update of the fall self-recovery strategy.

[0102] Among them, according to the action of the quadruped robot output by the Actor network, after being superimposed with the initial position of the current joint, the target joint position is generated, and then the output torque is calculated through a proportional-derivative (PD) controller, and finally the action is executed.

[0103] To enhance the diversity of training data, a random state reset mechanism is designed in the training environment. When the running time of the simulation environment exceeds the preset task time, a reset operation is triggered. By simulating the robot falling from a random posture in the simulation environment, it trains the robot to recover from any posture to the standing state, improving the diversity and robustness of the recovery ability.

[0104] The training model adopts an asymmetric Actor-Critic network architecture, combined with the design of the Teacher-Student Network, to improve the learning efficiency and policy generalization ability under complex terrains. At the same time, to meet the requirements of the NP3O algorithm, a Cost network for calculating costs is additionally introduced. The Actor-Critic network, the Cost network, and the Teacher-Student network are all constructed based on the Multi-Layer Perceptron.

[0105] As Figure 2 shown, the Actor network takes the normalized body state information o tTaking the feature information output by the teacher network or the student network as input, and through interacting with the simulation environment, obtaining the next state observation information, reward signal, cost feedback, and task end flag, and outputting the action a t The Critic network takes the normalized privileged information s t as input, evaluates the value of taking a specific action in the current state, and is used to guide the update of the Actor network. The Cost network takes the normalized privileged information as input, evaluates the cost of the current action, and is used to constrain the policy update to ensure that the generated policy maintains high efficiency and robustness while complying with safety restrictions.

[0106] Using the NP3O algorithm, the Actor network, Critic network, and Cost network are integrated into a unified optimization framework, combining the goals of policy optimization, value evaluation, and cost evaluation to ensure maximizing the cumulative reward under the condition of meeting safety constraints; the final optimization goal of NP3O is expressed as:

[0107]

[0108] where:

[0109]

[0110] is the loss function of the Actor network, used to optimize the policy, κ is the penalty factor, used to control the penalty intensity of constraint violation; is the probability ratio of the new and old policies, is the normalized advantage function provided by the Critic network, is the normalized cost advantage function provided by the Cost network, is the expected cost return of the current policy, d i is the upper limit of the cost constraint, and are the mean and standard deviation of the cost advantage function respectively, ∈ is a hyperparameter, used to control the amplitude of policy update; represents under the policy π k sampling the state s according to the state distribution and then sampling the action a according to the policy π k to obtain the expectation; clip represents the constraint of the body upright state.

[0111] The update objectives of the Critic network and the Cost network also include: The Critic network predicts the state value V φ (s), minimizes the difference between the value function V φ (s) and the actual return R, thereby improving the ability to evaluate the state value; The Cost network predicts the state cost Cψ (s) to minimize the cost function C ψ (s) and the actual environment cost C actual to reduce the difference therebetween, thereby enhancing the cost prediction ability; the objective function is expressed as:

[0112]

[0113] wherein, denotes the expectation obtained by sampling the state s according to the state distribution D.

[0114] The teacher network extracts high-performance features by inputting the normalized privileged information and provides optimization references for the student network. The student network takes the normalized historical observation information of the ontology state as the input and learns by imitating the policy features of the teacher network. The student network does not rely on privileged information and is more suitable for the deployment of actual robot systems. In the initial stage of training, the Actor network uses the feature information output by the teacher network as the input and guides the student network to gradually imitate the teacher network by minimizing the mean square error between the output features of the teacher network and the student network. When the policy converges, the output feature z t (2) replaces the teacher network feature z t (1) and becomes the direct input of the Actor network, realizing a smooth transfer from the teacher to the student.

[0115] Through this design, the framework realizes the close interaction among the Actor-Critic network, the Cost network and the environment, combines the optimization characteristics of the NP3O algorithm, and the synergistic effect of the teacher-student network, enabling the robot to quickly adapt to complex terrains and effectively learn to recover from any posture to the standing posture.

[0116] To narrow the gap between simulation and actual deployment, an environment randomization mechanism is introduced, including friction coefficient randomization, elastic recovery coefficient randomization, dynamic parameter randomization, motor parameter randomization and perturbation randomization.

[0117] The randomization of the friction coefficient and the restitution coefficient is attached to the rigid bodies of the robot model and is set through the rigid body properties of the physics engine; the randomization of the dynamic parameters refers to the randomization of the mass and inertia distribution of the robot; the randomization of the motor parameters refers to the randomization of the key parameters in the motor control system, usually including the proportional gain (Kp) and the derivative gain (Kd), as well as the randomization of the motor torque (i.e., the motor strength); the disturbance randomization is to simulate the external uncertain factors that the robot encounters in the actual environment, such as violent impacts, vibrations, uneven ground, etc. The disturbance randomization is added to the robot body in two forms of random disturbances, linear velocity disturbance and torque disturbance. The disturbance is usually applied intermittently rather than continuously, and is added at certain time intervals to simulate the randomness and irregularity of external disturbances in reality.

[0118] These randomized designs enhance the adaptability of the strategy to environmental changes and further improve the deployment reliability. The randomized content and its value range are shown in Table 4.

[0119] Table 4 Randomization Parameter Settings

[0120]

[0121] The course design draws on the progressive difficulty adjustment mechanism in games, including terrain courses and command courses.

[0122] Terrain Course: The robot starts training on simple terrains and gradually transitions to complex terrains. When the robot performs poorly on high-difficulty terrains, it will be switched back to low-difficulty terrains for review. After completing the highest-difficulty terrain, it will be randomly assigned to different terrains to avoid overfitting.

[0123] Command Course: Dynamically adjust the range of motion commands according to the robot's performance in the task. Gradually increase the command difficulty as the tracking reward improves, prompting the robot to maintain efficient performance in a larger action space.

[0124] Sim-to-real is an important link in realizing the transition of the robot system from simulation to actual deployment. We need to export the model for deployment from the entire training framework, as Figure 3 shown.

[0125] Export the model to be deployed into an ONNX format file, and optimize the model through ONNXRuntime to reduce inference latency and improve running efficiency. Deploy the optimized ONNX model to the robot hardware platform, ensuring that the deployment environment supports ONNX Runtime and preloading the required dynamic libraries. To ensure the consistency between training and actual deployment, it is particularly required that the update frequencies of the control signals in the simulation environment and the actual hardware are the same. This requires a sufficiently high motor response speed to ensure that the robot can achieve the same control accuracy on the actual hardware as in the simulation, thus achieving a smooth deployment effect.

[0126] The present invention is designed from aspects such as training environment setting, training model and optimization algorithm, environment randomization, training curriculum setting, model deployment, etc., comprehensively improving the adaptability, recovery ability and control performance of the robot in complex dynamic scenarios, ensuring that when the robot encounters emergencies such as falling, it can quickly adapt to terrain changes and autonomously recover, maintaining efficient and stable control performance.

[0127] The randomized state reset mechanism simulates various falling scenarios, enabling the robot to recover to a standing position from different initial states. This diverse training method significantly improves the adaptability of the robot in real complex terrains and has strong robustness in the face of unknown environments.

[0128] Through methods such as training environment randomization, curriculum design, and model optimization, ensure that the model can be successfully deployed to the physical robot and effectively reduce the differences between the simulation and the actual environment.

[0129] The present invention adopts a dynamic recovery strategy based on deep reinforcement learning, enabling the robot to autonomously identify and adapt to complex terrains after falling and complete autonomous recovery to a standing position. At the same time, the present invention significantly improves the robustness of the robot in unstructured environments such as rough terrains and steep slopes. It can not only efficiently recover from any posture to a standing posture, but also maintain stability and adaptability in complex environments, ensuring the continuity, stability and efficiency of the robot during task execution and avoiding the dependence on external intervention in traditional methods.

[0130] Figure 3 It is a schematic structural diagram of a four-legged robot's self-recovery control system for falling in complex terrains in an embodiment of the present invention. This embodiment corresponds to Figure 1 the four-legged robot's self-recovery control method for falling in complex terrains, as Figure 3 shown. The four-legged robot's self-recovery control system for falling in complex terrains in this embodiment may include:

[0131] An information processing module, which is used to obtain real-time observation information, historical observation information of the body state information, and privilege information of the four-legged robot's body state information and perform normalization processing;

[0132] A student network learning module, which is used to extract the falling self-recovery strategy features from the normalized privilege information by using the teacher network to provide optimization reference for the student network; and use the student network to learn the strategy features imitating the teacher network to process the normalized historical observation information of the ontology state information.

[0133] An action output module, which is used to process the output features of the student network and the real-time observation information of the normalized ontology state information by using the Actor network after the strategy converges, and obtain the next state observation information, reward signal, cost feedback and task end flag through interaction with the simulation environment, and output the quadruped robot action.

[0134] An action value evaluation module, which is used to process the normalized privilege information by using the Critic network to evaluate the value of taking a specific action in the current state to guide the update of the Actor network.

[0135] An action cost evaluation module, which is used to process the normalized privilege information by using the Cost network to evaluate the cost of the current action of the quadruped robot to constrain the update of the falling self-recovery strategy.

[0136] It should be noted here that Figure 3 Each module in the falling self-recovery control system of the quadruped robot in complex terrain corresponds one by one to each step in the above-mentioned falling self-recovery control method of the quadruped robot in complex terrain, and its specific implementation process is the same, which will not be repeated here.

[0137] A quadruped robot, including a robot body, on which a memory, a processor and a computer program stored on the memory and executable on the processor are provided. When the processor executes the program, it implements the steps in the above-mentioned falling self-recovery control method of the quadruped robot in complex terrain.

[0138] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0139] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A self-recovery control method for a quadruped robot falling down in complex terrain, characterized in that: include: Acquire the real-time observation information of the quadruped robot's body state information, the historical observation information of the body state information, and the privileged information and perform normalization processing; The teacher network is used to extract the fall self-recovery strategy features from the normalized privileged information to provide optimization reference for the student network; the student network is used to learn the strategy features that imitate the teacher network to process the normalized historical observation information of the ontology state information; When the strategy converges, the Actor network is used to process the output features of the student network and the real-time observation information of the normalized entity state information. By interacting with the simulation environment, the next state observation information, reward signal, cost feedback and task end mark are obtained, and the quadruped robot action is output; The normalized privileged information is processed using the Critic network to evaluate the value of taking a specific action in the current state to guide the update of the Actor network. The normalized privileged information is processed using the Cost network to evaluate the cost of the quadruped robot's current action, which is used to constrain the update of the fall self-recovery strategy.

2. The method for controlling a quadruped robot to recover from a fall in complex terrain as claimed in claim 1, characterized in that: The reward signal is obtained by weighted summing the corresponding weights of the key reward function, the walking-related reward function and the safety constraint reward function in the process of recovering from a fall to a standing posture.

3. The method for controlling a quadruped robot to recover from a fall in complex terrain as claimed in claim 1, characterized in that: The cost feedback is calculated by limiting the cost function of joint position, torque and velocity.

4. The method for controlling a quadruped robot to recover from a fall in complex terrain as claimed in claim 1, characterized in that: According to the quadruped robot action output by the Actor network, after superimposing it with the initial position of the current joint, the target joint position is generated, and then the output torque is calculated to finally realize the action execution.

5. The self-recovery control method for a quadruped robot falling down in complex terrain as claimed in claim 1, characterized in that: The real-time observation information of the body state information includes the angular velocity of the body, the projected gravity information, the current desired three-dimensional motion command, the offset of the current joint position relative to the default joint position, the current velocity of each joint, and the action value output by the Actor network; Or / and the privileged information includes observation of robot body information, linear speed of the body, current root height and measured terrain height difference information and randomization parameters.

6. The self-recovery control method for a quadruped robot falling down in complex terrain as claimed in claim 1, characterized in that: The NP3O algorithm is used to integrate the Actor network, Critic network, and Cost network into a unified optimization framework, combining the goals of strategy optimization, value evaluation, and cost evaluation to ensure that the cumulative reward is maximized while satisfying safety constraints; the final optimization goal of NP3O is expressed as: in: is the loss function of the Actor network, used to optimize the strategy, and κ is the penalty factor, used to control the penalty intensity of constraint violation; is the probability ratio of the new and old strategies, is the normalized advantage function provided by the Critic network, It is the normalized cost advantage function provided by the Cost network. is the expected cost-return of the current strategy, d i is the upper limit of the cost constraint, μ Ci and σ Ci are the mean and standard deviation of the cost advantage function, respectively, ∈ is a hyperparameter used to control the amplitude of the policy update; In the strategy π k Next, according to the state distribution Sampling state s, and then following strategy π k Sample action a and obtain the expected result; clip represents the constraint of the body's upright state.

7. The self-recovery control method for a quadruped robot falling down in complex terrain as claimed in claim 1, characterized in that: The update goals of the Critic network and the Cost network also include: The Critic network predicts the state value V φ (s), minimize the value function V φ (s) and the difference between the actual reward R, thereby improving the ability to evaluate the state value; the Cost network predicts the state cost C ψ (s), minimize the cost function C ψ (s) and the actual environmental cost C actual The difference between them can enhance the ability to predict the cost.

8. A self-recovery control system for a quadruped robot falling down in complex terrain, characterized in that: include: An information processing module, which is used to obtain real-time observation information of the quadruped robot's body state information, historical observation information of the body state information, and privileged information and perform normalization processing; The student network learning module is used to extract the fall self-recovery strategy features from the normalized privileged information using the teacher network to provide an optimization reference for the student network; the student network is used to learn the strategy features that imitate the teacher network to process the normalized historical observation information of the ontology state information; The action output module is used to process the output characteristics of the student network and the real-time observation information of the normalized entity state information using the Actor network after the strategy converges, and obtain the next state observation information, reward signal, cost feedback and task end mark by interacting with the simulation environment, and output the quadruped robot action; The action value evaluation module is used to process the normalized privileged information using the Critic network and evaluate the value of taking a specific action in the current state to guide the update of the Actor network; The action cost evaluation module is used to process the normalized privileged information using the Cost network, evaluate the cost of the quadruped robot's current action, and constrain the update of the fall self-recovery strategy.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the self-recovery control method for a quadruped robot falling down in complex terrain are implemented as described in any one of claims 1-7.

10. A quadruped robot, comprising a robot body, wherein the robot body is provided with a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps in the self-recovery control method for a quadruped robot falling under complex terrain are implemented as described in any one of claims 1-8.

Citation Information

Cited By

  • Robot falling-down self-recovery method and device based on deep reinforcement learning

    CN120508129A

  • Robot falling self-recovery method and device

    CN121492126A

  • Robot standing posture recovery method and device, equipment, storage medium and product

    CN121946529A

  • Quadruped robot falling recovery motion control method and device and medium

    CN122488528A