Humanoid robot falling recovery control method based on multi-stage course learning
Through multi-stage course learning and reinforced learning framework, the problem of insufficient fall recovery ability of humanoid robots is solved, rapid and stable fall recovery and environmental adaptation are achieved, hardware costs are reduced, and the robot's robustness and energy efficiency are improved.
Patent Information
- Application Number
- CN202510719962.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-08
AI Technical Summary
Existing humanoid robots have insufficient autonomous recovery capabilities after falling, especially in diverse fall postures and complex environments, with long recovery time and unstable hardware costs, and simulation strategies fail in real environments.
A multi-stage course learning method is adopted, combined with a reinforcement learning framework with mixed internal models and proximal strategy optimization, through staged learning and mirror loss functions and designing reward functions, the robot's recovery ability under different fall postures is gradually improved, and environmental adaptability and recovery efficiency are enhanced.
It realizes the robot's rapid and stable fall recovery in complex environments, shortens the recovery time to 1.2 seconds, improves the speed control accuracy, reduces joint oscillations, reduces hardware costs, enhances the conversion success rate from simulation to reality, and improves the robot's robustness and energy efficiency.
Smart Images

Figure CN120269573A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and specifically to a humanoid robot fall recovery control method based on multi-stage curriculum learning. Background Art
[0002] As an important branch of robot technology, humanoid robots aim to imitate human motor abilities and possess complex actions such as walking, standing, and fall recovery. Although humanoid robots have made remarkable progress in the field of motion control in recent years, their autonomous recovery ability after falling still faces significant challenges. Most traditional fall recovery methods rely on preset motion trajectories or rule-based control strategies, which are difficult to adapt to diverse fall postures and complex environmental interferences, and show low robustness and adaptability in practical applications.
[0003] Therefore, the following technical problems exist in the prior art:
[0004] Adaptability problem: Existing multi-stage curriculum learning methods usually optimize the action strategies of each stage independently, but there will be discontinuous and unstable situations during stage switching, resulting in recovery failure.
[0005] Long and unstable recovery time: Existing methods are difficult to meet both the recovery time and stability requirements, resulting in a failure rate in multiple recoveries.
[0006] Simulation-to-reality conversion problem: The strategy performs excellently in simulation, but often fails in the real environment, mainly due to environmental differences and unmodeled dynamic interferences.
[0007] Moreover, the following problems generally exist in humanoid robots in the prior art:
[0008] High requirements for motor performance: Existing recovery strategies have high requirements for the accuracy, response speed, and shock resistance of joint motors, resulting in increased hardware costs and difficulty in implementation on ordinary platforms. Summary of the Invention
[0009] To solve the above technical problems, the present invention proposes a humanoid robot fall recovery control method based on multi-stage curriculum learning. Through a multi-stage training strategy, it gradually learns the recovery path from different fall postures to stable standing, and introduces domain randomization and reward constraints in each stage to improve the environmental adaptability and recovery efficiency of the strategy. Experiments prove that this method can achieve efficient and stable fall recovery under various complex ground conditions.
[0010] To achieve the above object, the technical solution adopted by the present invention is:
[0011] A humanoid robot fall recovery control method based on multi-stage curriculum learning, characterized in that it includes the following steps:
[0012] S1. Construct a hybrid internal model that fuses regular observations and privileged observations, and use historical observation data to estimate the robot's body velocity and latent variables;
[0013] S2. Decompose the robot's falling and getting up actions into consecutive key frames and learn them frame by frame in stages;
[0014] S3. Introduce a reinforcement learning framework that combines hybrid internal optimization and proximal policy optimization. The reinforcement learning framework uses the regular observations and privileged observations stored in the hybrid internal model. The source encoder extracts feature representations from the regular observations, and the target encoder generates contrastive targets to guide representation learning. These latent representations are used by the actuator network to generate actions and are estimated for value by the evaluator. The actions are executed in the simulator, and the reward is calculated based on the similarity to the predefined key frame targets;
[0015] S4. Add a mirror loss function to measure the policy symmetry of the robot before and after mirror mapping, ensuring that the robot performs consistently in left-right symmetric or up-down symmetric motions;
[0016] S5. Design a reward function to guide the robot to learn the falling recovery action from different aspects.
[0017] As a preferred technical solution of the present invention: In step S1, the network of the hybrid internal model takes as input and outputs the context state vector and the estimated linear velocity ,
[0018] The network of the hybrid internal model is a feature that processes multi-frame observations to infer the environment and the robot's state. By combining current and historical observations, the estimator network can provide latent state estimates.
[0019] As a preferred technical solution of the present invention: In step S2, the staged learning is as follows:
[0020] The first stage requires the robot to learn from the first frame to the second frame. After successfully learning the second frame, enter the next learning, requiring it to learn from the third frame, and so on;
[0021] The second stage adds various randomizations when the robot can successfully stand in the simulation. When the robot can still stand after adding randomizations, add restricted training again. After completing the real machine verification, enter the third stage.
[0022] The third stage uses all fourteen-dimensional joint positions output by the hybrid internal model network, enabling the robot to recover from any supine posture to a flat position and finally stand up.
[0023] As a preferred technical solution of the present invention: in step S3, proximal policy optimization uses Proximal Policy Optimization as the optimization algorithm. By continuously interacting with the simulation environment, the robot can gradually master the ability to quickly and stably transition from one key frame to the next key frame.
[0024] 5. The humanoid robot fall recovery control method based on multi-stage curriculum learning according to claim 1, wherein: step S4 is specifically as follows:
[0025] When the robot executes the policy π(s) in state s, the state is transformed into a symmetric state M[s] through the mirror mapping M, and then a new action π(M[s]) is obtained through the calculation of the policy network. If the policy is perfectly symmetric, then after this action passes through the mirror mapping M[π(M[s])] again, it should be exactly the same as the policy output in the original state. The mirror loss is defined as the difference in the two-norm of the two, and the formula is:
[0026]
[0027] By minimizing this loss, it is ensured that the robot behaves consistently in left-right symmetric or up-down symmetric movements.
[0028] As a preferred technical solution of the present invention: in step S5, the reward function is specifically as follows:
[0029] reference joint pos is used to encourage the robot's current action to learn from the next key frame.
[0030] base height, height increase are used to encourage the robot to gradually increase its height.
[0031] orientation, body up are used to assist the robot in quickly learning the hand support state to the semi-standing state.
[0032] hand force increase is used for the first two frames of training to encourage the robot's hand to cooperate with the legs to exert force.
[0033] feet force increase is used to encourage the robot to land on both feet.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] Through a series of innovative technical means, the present invention has achieved remarkable technical effects in the imitation of human walking and running by humanoid robots, specifically as follows:
[0036] 1. Improve the recovery speed after falling and getting up:
[0037] In real robot tests, the average recovery time from lying down to standing up was only 1.2 seconds.
[0038] 2. Improve speed control accuracy:
[0039] The robot exhibits improved accuracy when executing velocity command tracking, reducing velocity control errors caused by terrain changes or speed increases.
[0040] The introduction of internal models effectively predicts future states, reduces control delays, and achieves smooth acceleration and deceleration.
[0041] 3. Reduce joint oscillation:
[0042] Smooth joint transitions reduce reliance on high-torque motors, reduce wear and increase system life.
[0043] 4. Successful simulation to reality conversion:
[0044] The robot is able to replicate in the real world the fall recovery actions learned in the simulation environment, significantly improving the success rate of deployment from simulation to reality.
[0045] 5. Optimized energy consumption and motor requirements:
[0046] Optimized motion planning and motion control reduce unnecessary energy consumption, improve the robot's energy efficiency when performing tasks, and reduce motor torque requirements.
[0047] 6. Enhanced robustness:
[0048] When faced with external disturbances and non-ideal conditions, the robot can quickly recover from falls, stand up and remain stable, demonstrating greater robustness.
[0049] 7. Expanded application scope:
[0050] The technical solution of the present invention improves the applicability of the robot in a variety of practical application scenarios. For example, in disaster scenarios such as earthquakes and fires where there is a lot of interference, the robot is prone to falling, and the factory floor is uneven. This method can quickly restore the robot to a standing state after falling, thus avoiding secondary damage to the robot.
[0051] 8. Promote technological progress:
[0052] The implementation of this invention will promote the development of humanoid robot technology, innovative fall recovery safety mechanism, and provide new possibilities for future research and innovation.
[0053] Through the above technical effects, the present invention optimizes the safety performance of human-machine interaction, enabling the robot to operate reliably in scenarios such as medical care and home services that require close human-robot collaboration, laying a technical foundation for building a more intelligent and safer next-generation service robot system, and is expected to drive the entire robot industry towards higher safety standards. Description of the Drawings
[0054] Figure 1 is the overall framework diagram of the present invention;
[0055] Figure 2 is the key frame for fall recovery state planning;
[0056] Figure 3 is the detailed table diagram of the reward function. Detailed Embodiments
[0057] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0058] As Figure 1 shown, the humanoid robot fall recovery control method based on multi-stage curriculum learning proposed by the present invention includes the following steps:
[0059] S1. Construct a hybrid internal model that fuses conventional observations and privileged observations, and use historical observation data to estimate the robot's body velocity and latent variables;
[0060] S2. Decompose the robot's fall and get-up actions into continuous key frames and learn them frame by frame in stages;
[0061] S3. Introduce a reinforcement learning framework that combines hybrid internal optimization and proximal policy optimization. The reinforcement learning framework uses the conventional observations and privileged observations stored in the hybrid internal model. The source encoder extracts feature representations from the conventional observations, the target encoder generates contrastive targets to guide representation learning, these latent representations are used by the actuator network to generate actions, and value estimation is performed by the evaluator. The actions are executed in the simulator, and rewards are calculated based on the similarity to the predefined key frame targets;
[0062] S4. Add a mirror loss function to measure the policy symmetry of the robot before and after mirror mapping to ensure that the robot performs consistently in left-right symmetric or up-down symmetric movements;
[0063] S5. Design a reward function to guide the robot to learn fall recovery actions from different aspects.
[0064] In step S1, the network of the hybrid internal model takes as the input and outputs the context state vector and the estimated linear velocity ,
[0065] The network of the hybrid internal model is designed as a feature that processes multi-frame observations to infer the environment and the robot's state. By combining current and historical observations, the estimator network can provide potential state estimates.
[0066] A reinforcement learning framework that combines Hybrid Internal Optimization (HIM) with Proximal Policy Optimization (PPO) is used for humanoid robot motion learning. This framework utilizes regular observations and privileged observations stored in a buffer. The source encoder extracts feature representations from regular observations, and the target encoder generates contrastive targets to guide representation learning. These latent representations are used by the actuator network to generate actions and are evaluated by a value estimator. The actions are executed in a simulator, and rewards are calculated based on the similarity to predefined key-frame targets.
[0067] Updated through policy gradients, thus facilitating the learning of complex humanoid motion behaviors. The network part of the hybrid internal model takes as input and outputs a context state vector and an estimated linear velocity .
[0068] The network is designed as a feature that processes multi-frame observations to infer the environment and the robot's state. By combining current and historical observations, our estimator network can provide more accurate and robust potential state estimates, ultimately improving the performance of the policy.
[0069] In step S2, the phased learning is as follows:
[0070] In the first phase, the robot is required to learn from the first frame to the second frame. After successfully learning the second frame, it proceeds to the next learning, which requires it to learn from the third frame, and so on.
[0071] In the second phase, after the robot can successfully stand in the simulation, various randomizations are added. When the robot can still stand after adding randomizations, restricted training is added again. After completing the real machine verification, it enters the third phase.
[0072] In the third phase, all fourteen-dimensional joint positions output by the hybrid internal model network are used, enabling the robot to recover from any supine posture to a flat position and ultimately achieve standing.
[0073] At the very beginning of learning and training, only ten of the fourteen joints are used (four arm joints and six leg joints, without using the yaw and roll joints of the legs). The key-frame training is as Figure 2As shown, in the first stage, the robot is required to learn from the first frame to the second frame. After successfully learning the second frame, it enters the next stage and is required to learn from the third frame, and so on. Such training can significantly reduce the difficulty of parameter tuning and accelerate the training speed. After the robot can successfully stand in the simulation, various randomizations are added, such as joint position and speed randomization, and robot orientation angle randomization. When the robot can still stand after adding randomizations, torque, speed, and other restrictions need to be added for training. After completing the real machine verification, it enters the final stage, expands the joints, and uses all fourteen-dimensional joint positions output by the network, so that the robot can recover from any supine posture to the flat position and finally stand up. This progressive learning method can significantly accelerate the robot's mastery of complex recovery actions.
[0074] In step S3, Proximal Policy Optimization is used as the optimization algorithm for proximal policy optimization. By continuously interacting with the simulation environment, the robot can gradually master the ability to quickly and stably transition from one key frame to the next.
[0075] Step S4 is specifically as follows:
[0076] When the robot executes the policy π(s) in state s, the state is transformed to the symmetric state M[s] through the mirror mapping M, and then a new action π(M[s]) is calculated through the policy network. If the policy is perfectly symmetric, then after this action passes through the mirror mapping M[π(M[s])] again, it should be exactly the same as the policy output in the original state. The mirror loss is defined as the difference in the two-norm of the two, and the formula is:
[0077]
[0078] By minimizing this loss, it is ensured that the robot behaves consistently in left-right symmetric or up-down symmetric movements, thereby improving the coordination and stability of the recovery actions.
[0079] In step S5, the reward function is specifically as follows:
[0080] reference joint pos is used to encourage the robot's current action to learn from the next key frame.
[0081] base height, height increase are used to encourage the robot to gradually increase its height.
[0082] orientation, body up are used to assist the robot in accelerating the learning of the hand support state to the semi-standing state.
[0083] hand force increase is used for the training of the first two frames to encourage the robot's hand to cooperate with the legs to exert force.
[0084] feet force increase is used to encourage the robot to land on both feet.
[0085] Such as Figure 3 As shown in the figure are the reward function and its corresponding formula. Among them, reference joint pos encourages the current action to learn from the next key frame and is the most important among all rewards. Since the height of the key frames gradually increases, base height and height increase are applicable in the learning process of all key frames, encouraging the robot to gradually increase its height. Orientation and body up help the hand support state to quickly learn to the semi-standing state. Hand force increase is mainly used for the training of the first two frames to encourage the hand to cooperate with the legs to exert force and achieve the effect of getting up quickly.
[0086] Due to the design differences of robots, some robots cannot land on their feet when lying flat. During the training process, feet force increase is used to encourage foot landing, which can help robots with non-landing soles to learn the target action faster. Feet orientation (ankle direction), stand on feet (whether to stand and contact the ground at this moment), standstill (static standing position), stand still vel (static standing speed), ankle vel limit (ankle speed limit), dof vel limit (joint speed limit), dof acc (joint acceleration), base acc (center of gravity acceleration), action smoothness (action smoothness), feet contact forces (foot contact force) make the robot stand stably, avoid jumping due to too fast speeds in the x and z directions, punish the sole for leaving the ground, and ensure the stability of the getting-up process. While torque limit (torque limit) and dof pos limits (joint position limit) are to minimize the damage to the actual machine motor. Hip roll yaw is used in the final stage of expanding the joint dimension to avoid excessive hip roll and yaw causing the legs to spread too wide.
[0087] The above description is only a preferred embodiment of the present invention and does not impose any other form of limitation on the present invention. Any modification or equivalent change made based on the technical essence of the present invention still falls within the scope of protection required by the present invention.
Claims
1. A humanoid robot fall recovery control method based on multi-stage curriculum learning, characterized in that: It includes the following steps: S1. Construct a hybrid internal model, which fuses regular observations and privileged observations, and uses historical observation data to estimate the robot's body velocity and latent variables; S2. Decompose the robot's falling and getting up actions into consecutive key frames and learn them frame by frame in stages; S3. Introduce a reinforcement learning framework that combines hybrid internal optimization and proximal policy optimization. The reinforcement learning framework uses the regular observations and privileged observations stored in the hybrid internal model. The source encoder extracts feature representations from the regular observations, and the target encoder generates contrastive targets to guide representation learning. These latent representations are used by the actuator network to generate actions, and value estimation is performed by the evaluator. The actions are executed in the simulator, and rewards are calculated based on the similarity to the predefined key frame targets; S4. Add a mirror loss function to measure the policy symmetry of the robot before and after mirror mapping, ensuring that the robot behaves consistently in left-right symmetric or up-down symmetric movements; S5. Design a reward function to guide the robot to learn the falling recovery action from different aspects.
2. The humanoid robot falling recovery control method based on multi-stage curriculum learning according to claim 1, wherein: In step S1, the network of the hybrid internal model takes as input and outputs a context state vector and an estimated linear velocity , The network of the hybrid internal model is a feature that processes multi-frame observations to infer the environment and the robot's state. By combining current and historical observations, the estimator network can provide latent state estimates.
3. The humanoid robot falling recovery control method based on multi-stage curriculum learning according to claim 1, wherein: In step S2, the staged learning is specifically as follows: In the first stage, that is, the robot is required to learn from the first frame to the second frame. After successfully learning the second frame, it enters the next learning, and is required to learn from the second frame to the third frame, and so on; In the second stage, after the robot can successfully stand in the simulation, various randomizations are added. When the robot can still stand after adding randomizations, restricted training is added again. After completing the real machine verification, it enters the third stage. In the third stage, all fourteen-dimensional joint positions output by the hybrid internal model network are used, enabling the robot to recover from any supine posture to a flat position and finally stand up.
4. The humanoid robot fall recovery control method based on multi-stage curriculum learning according to claim 1, characterized in that: In step S3, proximal policy optimization uses Proximal Policy Optimization as the optimization algorithm. By continuously interacting with the simulation environment, the robot can gradually master the ability to quickly and stably transition from one key frame to the next key frame.
5. The humanoid robot fall recovery control method based on multi-stage curriculum learning according to claim 1, wherein: Step S4 is specifically as follows: When the robot executes the policy π(s) in state s, the state is transformed to the symmetric state M[s] through the mirror mapping M, and then a new action π(M[s]) is calculated through the policy network. If the policy is perfectly symmetric, then after this action passes through the mirror mapping M[π(M[s])] again, it should be exactly the same as the policy output in the original state. The mirror loss is defined as the difference in the two-norm of the two, and the formula is: By minimizing this loss, it is ensured that the robot behaves consistently in left-right symmetric or up-down symmetric movements.
6. The humanoid robot fall recovery control method based on multi-stage curriculum learning according to claim 1, characterized in that: In step S5, the reward function is specifically as follows: reference joint pos is used to encourage the robot's current action to learn from the next key frame, base height, height increase are used to encourage the robot to gradually increase its height, orientation, body up are used to assist the robot in quickly learning the hand support state to the semi-standing state, hand force increase is used for the training of the first two frames to encourage the robot's hand to cooperate with the legs to exert force, feet force increase is used to encourage the robot to land on both feet.
Citation Information
Cited By
Robot motion control strategy network training method and device based on imitation learning
CN121447651A