Humanoid robot lower limb walking control method and device

By combining a Transformer-based policy network and an adversarial prior network, the problem of gait instability in humanoid robots in complex environments in existing technologies is solved, improving adaptability and robustness, and generating natural and stable gait control.

CN121069743AActive Publication Date: 2025-12-05CITIC HEAVY INDUSTRIES CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511620815.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2025-12-05
Estimated Expiration
2045-11-07

AI Technical Summary

Technical Problem

Existing humanoid robot walking control methods lack adaptability and robustness in complex or unknown environments. Traditional trajectory planning requires a lot of manual design, deep reinforcement learning is prone to generating asymmetric and unnatural gait with obvious motion jitter, and single-step prediction strategies have the problem of error accumulation.

Method used

A Transformer policy network based on a self-attention structure is used for multi-step action prediction. An adversarial prior network is combined to fuse style scoring and task rewards. The policy network is optimized through policy gradient reinforcement learning to construct smooth and coherent control instructions. Safety constraints and fall recovery mechanisms are added to the robot controller.

Benefits of technology

It achieves stability and continuity of the gait for the lower limbs of humanoid robots, improves adaptability and robustness in complex environments, and generates a natural and stable gait.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121069743A_ABST
    Figure CN121069743A_ABST
Patent Text Reader

Abstract

The invention relates to a humanoid robot lower limb walking control method and device, and belongs to the technical field of robot control, and the method comprises the following steps: obtaining state information and environment information of a robot; the state information and the environment information are input into a strategy network based on a self-attention structure, a joint action sequence of multiple steps in the future is output, candidate actions of the output overlapped action sequence at the same future moment are subjected to weighted fusion, and a smooth and coherent control instruction is generated; inputting the state transition segment of the robot into an adversarial prior network, outputting a style score and mapping the style score into a style reward; constructing a comprehensive reward function in combination with the task completion reward and the style reward, and optimizing the strategy network; and deploying the optimized strategy network to a robot controller to realize lower limb walking control. The walking naturalness, stability and environmental adaptability can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot control, in particular to a humanoid robot lower limb walking control method and device. BACKGROUND

[0002] The existing humanoid robot walking control method mainly has the following deficiencies: the traditional trajectory planning and control method needs a large amount of manual design of gait trajectory, and is difficult to cope with complex or unknown environment, and lacks self-adaptability and robustness.

[0003] The deep reinforcement learning method can realize walking through interaction learning, but if only relying on task reward, it is easy to generate asymmetric and unnatural gait, and the action jitter is obvious, and it lacks robustness to disturbance.

[0004] The single-step prediction strategy has the problem of error accumulation, which easily leads to unstable or invalid walking. SUMMARY

[0005] Therefore, the purpose of the present application is to provide a humanoid robot lower limb walking control method and device.

[0006] The technical scheme adopted by the present application to solve the above technical problems is: a humanoid robot lower limb walking control method, comprising the following steps: S1, obtaining state information and environment information of the robot; S2, inputting the state information and environment information into a policy network based on a self-attention structure, and outputting a joint action sequence of multiple future steps; S3, weighting and fusing candidate actions at the same future time of the overlapping action sequence output in step S2 to generate smooth and coherent control instructions; S4, inputting the state transition segment of the robot into an adversarial prior network, outputting a style score and mapping it into a style reward; S5, combining the task completion degree reward and the style reward to construct a comprehensive reward function, and optimizing the policy network; S6, deploying the optimized policy network to the robot controller to realize lower limb walking control.

[0007] Further, in step S2, the policy network adopts a Transformer structure based on a self-attention mechanism to encode the historical state sequence and output a multi-step action block sequence.

[0008] Further, in step S3, the weighting fusion adopts an exponential weighted average method of a time decay factor and a state adaptive factor.

[0009] Further, in step S4, the adversarial prior network is trained with high-quality reference motion data as positive samples and robot-generated trajectories as negative samples, adopts a least square adversarial loss, and introduces gradient regularization on expert samples.

[0010] Further, in step S5, the reward function is a weighted sum of task completion reward and style reward, the style reward is derived from the output score of the adversarial prior network, and the task completion reward includes speed, posture stability, energy consumption, and contact stability indicators.

[0011] Further, in step S5, the optimization of the policy network adopts a reinforcement learning algorithm based on policy gradient, including any one of PPO, SAC or TRPO.

[0012] The application also provides a device for implementing the above-mentioned humanoid robot lower limb walking control method, which comprises a thigh segment, a calf segment and a foot segment, the thigh segment is connected with the torso of the humanoid robot through a hip joint, the thigh segment and the calf segment are connected through a knee joint, and the calf segment and the foot segment are connected through an ankle joint, motor drivers and angle sensors are arranged in the hip joint, the knee joint and the ankle joint, a force sensor is arranged on the sole of the foot segment, and the device further comprises a robot controller connected with the motor drivers, the angle sensors and the force sensor.

[0013] The application has the following beneficial effects: 1. The application proposes a sequence action block prediction policy network based on Transformer, which expands the traditional single-step action output to multi-step trajectory block output, and adopts time decay and state adaptive exponential weighted fusion on overlapping blocks in the reasoning stage, thereby effectively suppressing long-time sequence error accumulation and action jitter while maintaining response real-time, and greatly improving the stability and continuity of gait.

[0014] 2. The application constructs an adversarial prior network, takes high-quality reference motion data as true samples and robot-generated trajectories as false samples, outputs a style score by discriminating the state transition segment, and uses the style score as a style reward and a task completion reward for joint optimization, thereby significantly improving the coordination, symmetry and humanoid nature of the generated gait.

[0015] 3. The application integrates the style reward and the task completion reward output by the adversarial prior network into a total reward signal, and uses an existing mature policy gradient type reinforcement learning algorithm (such as PPO, SAC, etc.) to optimize the policy network, thereby effectively improving the training sample efficiency while ensuring the stability of the algorithm. BRIEF DESCRIPTION OF DRAWINGS

[0016] Fig. 1 is the overall framework diagram of the application.

[0017] Fig. 2is the training deployment flowchart of the present application. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application is further described in detail below in combination with the drawings and examples. It should be understood that in the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more; the terms "upper", "lower", "left", "right", "inner", "outer", "front end", "rear end", "head", "tail" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the purpose of facilitating the description of the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" and the like are only for the purpose of description, and cannot be understood as indicating or implying relative importance.

[0019] Please refer to Figs. 1-2 A humanoid robot lower limb walking control method, comprising the following steps: Step one, sensing input: collecting robot state information and environment information, the robot state information including centroid position, torso posture, joint angle and angular velocity, foot end position and contact information. The environment information includes terrain height.

[0020] Step two, strategy network prediction based on self-attention structure: a strategy network based on Transformer structure is constructed, the historical state sequence is input, and the joint control action sequence of future steps is output, realizing global modeling of long-time sequence action.

[0021] Step three, overlapping action sequence weighted fusion: in the reasoning stage, the action sequence predicted by adjacent time steps produces multiple candidates at the same future time, and these candidates are exponentially weighted and fused according to time decay and state adaptive factor to generate smooth and coherent action instructions, effectively suppressing long-time sequence error accumulation and action jitter.

[0022] The action sequence weighted fusion adopts an exponential weighted average method: , In the formula, is the predicted action of the tth step after fusion, is the predicted action of the ith action block at the tth step, is a weight coefficient, used to weight and fuse the candidate actions of the overlapping block according to time decay, to suppress long-term prediction error, is a decay coefficient, the decay coefficient is a scalar, controlling the weight decay with the time sequence of the block, and i is the action block index.

[0023] Step four, against the prior network style constraint: build a discriminant network with high-quality reference motion as true sample and robot-generated trajectory as false sample, and output style score for state transition segment, and map it to style reward.

[0024] The input of the adversarial prior network is the local features of the state transition segment , including root velocity, joint local rotation and angular velocity, end effector relative position, contact state, etc.

[0025] The network training adopts least square adversarial loss, with high-quality reference motion segment as true sample and strategy generation segment as false sample. Gradient regularization is introduced on expert samples to improve training stability.

[0026] The adversarial prior network outputs style score , ranging from 0 to 1, and the closer to 1, the more similar to the reference motion. The style reward is obtained by linear mapping: In the formula, is the robot state of adjacent two frames, including root velocity, joint angle, angular velocity, end relative position, contact state; is the style reward, ranging from 0 to 1.

[0027] Step five, reward function comprehensive optimization: the total reward is formed by combining the task completion degree reward and the style reward output by the adversarial prior network, guiding the strategy network to generate actions that can complete the task and have natural style.

[0028] The comprehensive reward function is: In the formula, is the total reward at time t, is the task completion degree reward, is the style reward, and are weight coefficients for balancing task and style rewards. The weight is determined by experience setting and parameter search. In the early stage of training , , it can be appropriately fine-tuned according to the learning curve.

[0029] Step six, reinforcement learning strategy optimization: the strategy network is trained by using the reinforcement learning algorithm based on policy gradient, such as proximal policy optimization (PPO), soft actor-critic (SAC) or trust region policy optimization (TRPO), to maximize the expected comprehensive return.

[0030] Step seven, simulation training and migration deployment: course learning and domain randomization training are performed in a simulation environment, parallel training is performed on a simulation platform, course learning strategy is adopted, and gradually transition from simple terrain to complex terrain. Domain randomization includes kinetic parameters, friction coefficient, sensor noise and delay disturbance. Before real deployment, the system identifies and calibrates the real robot kinetic parameters, and performs strategy distillation and delay compensation to improve the real-time performance and stability of the strategy.

[0031] The application performs parallel training on a simulation platform, adopts course learning and domain randomization to improve robustness, performs system identification and parameter correction before real deployment, adopts strategy distillation and delay compensation to improve real-time performance, adds a safety constraint layer and a fall recovery mechanism in the robot controller, and further improves performance through online fine-tuning or residual learning when necessary.

[0032] Step eight, real deployment and safety layer: deploying the optimized strategy on a real robot, adding joint limit, torque saturation, contact hysteresis and fall recovery mechanism.

[0033] The application also provides a device for implementing the above-mentioned humanoid robot lower limb walking control method, which includes a thigh segment, a calf segment and a foot segment. The thigh segment is connected to the humanoid robot torso through a hip joint, the thigh segment and the calf segment are connected through a knee joint, and the calf segment and the foot segment are connected through an ankle joint. Motor drivers and angle sensors are arranged in the hip joint, the knee joint and the ankle joint. The sole of the foot segment is provided with a force sensor. The device also includes a robot controller connected to the motor drivers, angle sensors and force sensors. The foot structure is provided with a non-slip pad layer and a buffer mechanism to improve contact stability. The controller is an embedded microprocessor running a microkernel real-time operating system and supporting EtherCAT or CAN bus communication. The embodiment communicates with the motor drivers and sensors through the EtherCAT bus, executes the above-mentioned humanoid robot lower limb walking control method, outputs smooth joint commands, and realizes natural and stable lower limb walking.

[0034] The humanoid robot torso is provided with an inertial measurement unit (IMU) for obtaining center of mass and attitude information. Angle encoders are installed in the hip joint, the knee joint and the ankle joint for collecting angle and angular velocity. A six-axis force sensor is installed on the sole for detecting ground contact force and friction information. The hip joint and the knee joint are single-degree-of-freedom rotary joints, and the ankle joint is a two-degree-of-freedom joint allowing pitch and roll.

[0035] It should be noted that the above embodiments are only used to illustrate the application, but the application is not limited to the above embodiments. Any simple modification, equivalent change and modification made according to the technical essence of the application to the above embodiments also falls within the protection scope of the application.

Claims

1. A humanoid robot lower leg walking control method characterized by comprising: The method comprises the following steps: S1, obtaining state information and environment information of the robot; S2, inputting the state information and the environment information into a policy network based on a self-attention structure to output a joint action sequence in multiple future steps; S3, weighting and fusing candidate actions at the same future time in the overlapping action sequence output in step S2 to generate a smooth and coherent control instruction; S4, inputting a state transition segment of the robot into an adversarial prior network to output a style score and map it to a style reward; S5, combining a task completion degree reward and the style reward to construct a comprehensive reward function, and optimizing the policy network; S6, deploying the optimized policy network to a robot controller to realize lower limb walking control.

2. The walking control method of a lower leg of a humanoid robot according to claim 1, characterized by, In step S2, the policy network adopts a Transformer structure based on a self-attention mechanism to encode a historical state sequence and output a future multi-step action block sequence.

3. The walking control method of a lower leg of a humanoid robot according to claim 1, characterized by, In step S3, the weighted fusion adopts an exponential weighted average method of a time decay factor and a state adaptive factor.

4. The walking control method of a lower leg of a humanoid robot according to claim 1, characterized by, In step S4, the adversarial prior network is trained with high-quality reference motion data as positive samples and robot-generated trajectories as negative samples, adopts a least squares adversarial loss, and introduces gradient regularization on expert samples.

5. The walking control method of a lower leg of a humanoid robot according to claim 1, characterized by, In step S5, the reward function is a weighted sum of a task completion degree reward and a style reward, the style reward is derived from the output score of the adversarial prior network, and the task completion degree reward includes speed, posture stability, energy consumption, and contact stability indicators.

6. The walking control method of a lower leg of a humanoid robot according to claim 1, characterized by, In step S5, the optimization of the policy network adopts a reinforcement learning algorithm based on a policy gradient, including any one of PPO, SAC or TRPO.

7. An apparatus for implementing the walking control method of the lower limbs of a humanoid robot according to any one of claims 1 to 6, characterized in that, The robot comprises a thigh segment, a lower leg segment and a foot segment, the thigh segment is connected with a humanoid robot torso through a hip joint, the thigh segment and the lower leg segment are connected through a knee joint, the lower leg segment and the foot segment are connected through an ankle joint, motor drivers and angle sensors are arranged in the hip joint, the knee joint and the ankle joint, a force sensor is arranged on the sole of the foot segment, and the robot further comprises a robot controller connected with the motor drivers, the angle sensors and the force sensor.

Citation Information

Patent Citations

  • Humanoid robot gait imitation learning method combined with periodic reward

    CN118664586A

  • Multi-modal sensing humanoid robot action self-adaptive control method and multi-modal sensing humanoid robot action self-adaptive control system

    CN119610112A

  • Humanoid robot welding pose adjusting method and system

    CN120326636A

  • Man-machine hybrid autonomous navigation system in unknown dynamic environment

    CN120467325A

  • Arm-hand robot grabbing method based on deep reinforcement learning

    CN120533737A