Gait planning method for biped robot

By performing supervised stochastic divergent trial-and-error learning and reward-penalty condition optimization within the Humanoid-Gym framework, the robustness of motion control for bipedal robots in complex environments is addressed. Stable, low-energy gait planning is achieved, adapting to varied terrains and simplifying the research process.

CN120848247APending Publication Date: 2025-10-28SHANGHAI LUOBO PARTY TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510957796.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing motion control algorithms for bipedal robots lack robustness in complex environments, and may fail, especially when faced with more complex external environments. They also require a large amount of computation, making it difficult to achieve stable and low-energy gait planning.

Method used

A reinforcement learning training framework based on Humanoid-Gym is adopted. The robot is described as a rigid body serial kinematic chain through URDF files. Supervised random divergent trial and error learning is carried out in the simulation environment. Reward and punishment conditions are set, the reward function and noise addition are optimized, and a stable, low-energy and highly robust control strategy is achieved.

Benefits of technology

It achieves stable, low-energy, and highly robust gait planning for robots in complex environments, enabling them to walk autonomously in real-world environments and adapt to varying terrains, simplifying the research process and improving the level of systematization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848247A_ABST
    Figure CN120848247A_ABST
Patent Text Reader

Abstract

The invention relates to a biped robot gait planning method. The invention relates to the technical field of robot gait planning, and the method comprises the steps: taking a kinetic model of a robot as input, so as to achieve the neural network reasoning of simulation; describing the robot by using a URDF file, and enabling the robot to be equivalent to a kinematic chain formed by connecting rigid bodies in series; through accurate description, a robot consistent with that in reality is established in a simulation environment, the robot is enabled to carry out gait learning in a simulation physical environment, and supervised random divergence trial and error are carried out in a mode of setting initial conditions, termination conditions and reward and punishment conditions so as to realize gait learning. According to the invention, through repeated modification of parameters and structural adjustment of adaptability, a stable, low-energy-consumption and high-robustness control strategy is finally obtained through training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot gait planning technology, and specifically to a method for gait planning of bipedal robots. Background Technology

[0002] In recent years, legged robots have received widespread attention in the field of robotics. Compared to wheeled or tracked robots, legged robots can navigate and walk on more complex terrains, thus enabling their deployment in a wider range of applications. Bipedal robots, in particular, have advantages over quadrupedal and hexapedal robots due to their smaller size, making them easier to operate in narrow or complex environments, and they typically have higher energy efficiency. Furthermore, in the scientific research field, the motion control of bipedal humanoid robots is a hot area, and developing a low-cost bipedal robot provides an excellent research platform.

[0003] Based on these characteristics, bipedal robots can play an important role in fields such as general terrain exploration and disaster search and rescue. Their unique configuration provides flexibility and stability, enabling them to adapt to various complex and changing real-world environments, including walking on broken terrain and unpaved roads, or entering narrow, abandoned spaces that are normally difficult to access, and effectively performing their destination tasks. With the explosive development of reinforcement learning (RL) technology, many solutions have been found for the motion control problem of legged robots. Based on this technology, bipedal robots equipped with sensory input and well-trained neural networks can autonomously identify terrain and obstacles and walk, solving the fundamental problem of motion control for legged robots.

[0004] The BRAVER robot achieved unassisted, fully free, flat-surface walking in three-dimensional space with a certain degree of interference resistance. Its mechanical design aligns with industry mainstream practices; the linkage system, compared to the synchronous belt drive of the Little HERMES robot, can withstand and transmit greater torque while also boasting a longer service life. The high-rigidity linkage transmission facilitates the detection of joint feedback, and the output torque data derived from the current and voltage data at the motor ends more closely approximates the torque experienced at the actual joint position, thus reducing the number of sensors required and increasing control stability. Dynamic analysis also allows for the calculation of the force mapped to the sole of the foot, eliminating the need for a separate foot force sensor. However, like MPC, this control algorithm requires high precision in kinematic and dynamic modeling and involves significant computational load, necessitating a high-performance main controller for real-time calculation. Compared to RL, this control framework has lower robustness and may fail in more complex environments.

[0005] At the control algorithm level, the first-generation SUSTech point-foot bipedal robot uses the MPC control algorithm and constructs a two-stage method for map building. The first stage extracts and merges planar regions from images obtained by a depth camera to form a terrain map. The second stage approximates the extracted planar regions using convex polygons to simplify the map while maintaining its accuracy, thus simplifying the scene while building a terrain map with the required precision. Through this control method, the first-generation SUSTech point-foot bipedal robot has basically achieved terrain perception and autonomous balance maintenance. The parallel structure also conforms to biomimetic design principles.

[0006] Compared to traditional control algorithms such as MPC (Model Predictive Control), reinforcement learning control algorithms require less inference computation and can be deployed on most x86-64 computers and ARM architecture development boards. For robot locomotion control, it can also be divided into a control framework of "brain," "cerebellum," and "spinal cord."

[0007] Specifically, the complete chain requires a host computer with multimodal capabilities (usually with GPU graphics acceleration, NPU inference acceleration, and hardware encoding / decoding capabilities) as the "brain." This host computer executes mixed encoding of various types of information and inputs it into a medium-sized neural network for inference, enabling upper-level control such as VLA (Visual-Language-Action Model), robot path planning and navigation, lightweight GPT (Generative Pre-trained Transformer), LLaMA (Large Language Model Meta AI), and robot movement control strategy switching. A dedicated host computer for motion control uses a CPU to directly perform inference and data transmission, reception, encoding, and decoding with sensors and actuators. In this project, the host computer is mainly responsible for storing and executing neural network inference in the "flat ground walking" task, such as an Rk3588s embedded development board or a regular x86-64 computer—essentially the "cerebellum" of robot control. Regarding the chip selection and necessity for the "spinal cord" part, each joint motor module has an independent MCU responsible for executing the input information at a fixed frequency at high speed according to the given execution logic. Commonly used CAN and CANd chips... The format and content of the message information in the protocol robot integrated joint module are determined at the factory, which inevitably leads to redundancy in information transmission during use. However, the resources consumed by information transmission and reception in computer processing logic are enormous and cannot be optimized. Therefore, this project selects an STM32 processor for data transmission, reception, processing, and compression, which is equivalent to adding an external processing core to reduce the computing resource consumption of the host computer.

[0008] After providing a motion control method with terrain feedback, and combining it with higher-level functions such as an autonomous navigation system and a remote control transmission system, the bipedal robot can autonomously search for and locate target objects or survivors during missions. Simultaneously, it collects real-time on-site data, providing crucial environmental information for mission execution and aiding in the decision-making process. In conclusion, bipedal robots have irreplaceable significance in search and rescue, exploration, scientific research, and education. Developing a complete system design for this type of robot not only simplifies the research process but also improves the systematization and efficiency of its research. Summary of the Invention

[0009] This invention utilizes a control framework derived from the reinforcement learning training framework Humanoid-Gym. To achieve clear and natural gait in the simulation environments Issac Gym and MuoCo, this invention discloses a gait planning method for bipedal robots.

[0010] This invention provides the following technical solutions: A gait planning method for a bipedal robot, the method comprising the following steps: Step 1: Use the robot's dynamic model as input to achieve neural network inference for simulation; Step 2: Describe the robot using a URDF file, representing the robot as a kinematic chain composed of a series of rigid bodies; Step 3: Through precise description, a robot consistent with reality is built in the simulation environment, allowing it to learn gait in the simulated physical environment. Supervised random divergent trial and error is carried out by setting initial conditions, termination conditions, and reward and punishment conditions to achieve gait learning.

[0011] Preferably, each category describes the position, orientation, mass, moment of inertia, kinematic pair type, visual display, and collision element of the relative motion coordinate system of the active component.

[0012] Preferably, the high-precision URDF model uses the method of weighing the physical part and assigning the weighed mass data to the part in Solidworks to minimize the quality errors caused by machining and the quality differences that are difficult to assess.

[0013] Preferably, the robot needs to have a collision volume setting, where the leg collision volume is used to detect self-touching of the robot's limbs, so that the robot's gait will not collide with itself, and to achieve some gait correction functions. The collision volume of the body base is replaced by a cube larger than the body volume; The impactor of the foot is used to make contact with the ground, and the forces on the sole of the foot are solved.

[0014] Preferably, the perturbation function settings need to be reasonably adjusted according to the actual situation and robot pattern. Modify the training reward function Reward, as the setting of the reward function directly affects the robot's final behavioral performance. The reward function sets rewards or penalties for the following behavioral patterns: Velocity tracking: This feature encourages the robot to track specified linear and angular velocities using an error metric function. To calculate the error and give the reward, where e is the error value and w is the weighting coefficient; Pose tracking: Encourages the robot to maintain the desired pose, including base height and pose angle, also using an error metric function. To calculate attitude tracking error; Speed ​​imbalance: Penalizes the robot for speed deviations in the z, γ, and β directions; if the deviation is too large, it will be penalized. Contact pattern: Encourages the robot's foot contact pattern to be consistent with the preset contact mask Ip(t), reflecting the periodicity of the gait; Joint position tracking: Encourages the robot's joints to track the target position; Default joints: Encourage robot joints to remain in their default positions; Energy consumption: This penalty affects the robot's energy consumption. This item may train the robot to walk on straight legs, which poses a risk of near-double solution. Smooth motion: Punishes drastic changes in the robot's movements to ensure smooth motion; Large contact force: This is used to punish robots for excessive contact force, in order to ensure the stability of their movement.

[0015] Preferably, zero-sample transfer is achieved in a real environment. If the robot produces unreasonable behavior that exceeds the above-mentioned rewards and penalties during training, the weights of the reward and penalty function are modified or a reward and penalty function is added according to the specific circumstances of the error. The trained control strategy is then ported to MuJoCo. The simulation dt is 0.004s, and the control strategy takeover dt is 0.02s. Therefore, the hardware frequency is 250Hz, and the strategy takeover frequency is 50Hz.

[0016] Preferably, in the noise addition process, in addition to modifying the noise accumulation mechanism over time, noise related to gravitational acceleration is also added to achieve the filtering effect of the IMU; in the addition of randomization parameters, random phase is added to achieve random left and right foot lifting of the robot. During training, random parameters such as robot thrust, motor stiffness coefficient, motor damping coefficient, motor zero point position, joint resistance, and command noise were added to make the trained model more robust. Coefficients for bilateral tracking were added to the PPO algorithm, so that the symmetry parameters of the trained robot do not need to be set separately, but are automatically achieved by the fitting function of PPO.

[0017] A bipedal robot gait planning system, characterized in that: the system comprises: A data input module that takes the robot's dynamic model as input to achieve neural network inference in the simulation; The simulation module describes the robot using URDF files, equating the robot to a kinematic chain composed of a series of rigid bodies. The gait learning module establishes a robot in a simulation environment that is consistent with the real world through precise description, allowing the robot to learn gait in the simulated physical environment. By setting initial conditions, termination conditions, and reward and punishment conditions, it conducts supervised random divergent trial and error to achieve gait learning.

[0018] A computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement a bipedal robot gait planning method.

[0019] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a bipedal robot gait planning method.

[0020] The present invention has the following beneficial effects: This invention addresses the need for adaptation and modification of the Humanoid-Gym reinforcement learning training framework. It provides a complete porting approach and offers improvements to the training framework. Through repeated parameter modifications and structural adjustments for adaptability, a stable, low-energy, and highly robust control strategy is ultimately obtained after training. Attached Figure Description

[0021] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 The image shows the robot URDF export effect and the simplified collision body effect of the present invention. Figure 2 The robot's gait over a week is shown in IssacGym of this invention; Figure 3 The simulation environment of this invention is migrated to MuJoCo, showing the robot's gait within half a cycle; Figure 4 The flowchart shown is a process of the method of the present invention. Detailed Implementation

[0023] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] The present invention will be described in detail below with reference to specific embodiments. Specific Implementation Example 1: according to Figures 1 to 4 As shown, the specific optimized technical solution adopted by the present invention to solve the above-mentioned technical problems is: The present invention relates to a bipedal robot gait planning method.

[0026] This invention provides a gait planning method for a bipedal robot, the method comprising the following steps: Step 1: Use the robot's dynamic model as input to achieve neural network inference for simulation; Step 2: Describe the robot using a URDF file, representing the robot as a kinematic chain composed of a series of rigid bodies; Step 3: Through precise description, a robot consistent with reality is built in the simulation environment, allowing it to learn gait in the simulated physical environment. Supervised random divergent trial and error is carried out by setting initial conditions, termination conditions, and reward and punishment conditions to achieve gait learning.

[0027] This invention addresses the need for adaptation and modification of the Humanoid-Gym reinforcement learning training framework. It provides a complete porting approach and offers improvements to the training framework. Through repeated parameter modifications and structural adjustments for adaptability, a stable, low-energy, and highly robust control strategy is ultimately obtained after training. Specific Implementation Example 2: The only difference between Embodiment 2 and Embodiment 1 of this application is that: Each category describes the position, orientation, mass, moment of inertia, kinematic pair type, visual display, and collider elements of the moving component's relative motion coordinate system. Specific Implementation Example 3: The only difference between Embodiment 3 and Embodiment 2 of this application is that: The high-precision URDF model uses the method of weighing physical parts and assigns the weighed mass data to the part in Solidworks to minimize the quality errors caused by machining and the quality differences that are difficult to assess. Specific Implementation Example 4: The only difference between Embodiment 4 and Embodiment 3 of this application is that: The robot needs to have its collision volume set. The leg collision volume is used to detect self-touching of the robot's limbs, so that the robot's gait will not collide with itself, and to realize some gait correction functions. The collision volume of the body base is replaced by a cube larger than the body volume; The impactor of the foot is used to make contact with the ground, and the forces on the sole of the foot are solved. Specific Implementation Example 5: The difference between Embodiment 5 and Embodiment 4 of the present invention lies only in: The perturbation function settings need to be adjusted reasonably according to the actual situation and robot pattern. Modify the training reward function Reward, as the reward function settings directly affect the robot's final behavior. The reward function sets rewards or penalties for the following behavioral patterns: Velocity tracking: This feature encourages the robot to track specified linear and angular velocities using an error metric function. To calculate the error and give the reward, where e is the error value and w is the weighting coefficient; Pose tracking: Encourages the robot to maintain the desired pose, including base height and pose angle, also using an error metric function. To calculate attitude tracking error; Speed ​​imbalance: Penalizes the robot for speed deviations in the z, γ, and β directions; if the deviation is too large, it will be penalized. Contact pattern: Encourages the robot's foot contact pattern to be consistent with the preset contact mask Ip(t), reflecting the periodicity of the gait; Joint position tracking: Encourages the robot's joints to track the target position; Default joints: Encourage robot joints to remain in their default positions; Energy consumption: This penalty affects the robot's energy consumption. This item may train the robot to walk on straight legs, which poses a risk of near-double solution. Smooth motion: Punishes drastic changes in the robot's movements to ensure smooth motion; Large contact force: This is used to punish robots for excessive contact force, in order to ensure the stability of their movement. Specific Implementation Example Six: The difference between Embodiment Six and Embodiment Five of the present invention lies only in: To achieve zero-sample transfer in a real-world environment, if the robot exhibits unreasonable behavior exceeding the aforementioned rewards and penalties during training, the weights of the reward and penalty functions are modified or additional reward and penalty functions are added based on the specific circumstances of the error. The trained control strategy is then ported to MuJoCo. The simulation dt is 0.004s, and the control strategy takeover dt is 0.02s. Therefore, the hardware frequency is 250Hz, and the strategy takeover frequency is 50Hz. Specific Implementation Example 7: The difference between Embodiment Seven and Embodiment Six of the present invention lies only in: In the noise addition process, in addition to modifying the noise accumulation mechanism over time, noise related to gravitational acceleration was also added to achieve the filtering effect of the IMU; in the addition of randomization parameters, random phase was added to achieve random left and right foot lifting of the robot. During training, random parameters such as robot thrust, motor stiffness coefficient, motor damping coefficient, motor zero point position, joint resistance, and command noise were added to make the trained model more robust. Coefficients for bilateral tracking were added to the PPO algorithm, so that the symmetry parameters of the trained robot do not need to be set separately, but are automatically achieved by the fitting function of PPO. Specific Implementation Example 8: The difference between Embodiment 8 and Embodiment 7 of the present invention lies only in: A bipedal robot gait planning system, the system comprising: A data input module that takes the robot's dynamic model as input to achieve neural network inference in the simulation; The simulation module describes the robot using URDF files, equating the robot to a kinematic chain composed of a series of rigid bodies. The gait learning module establishes a robot in a simulation environment that is consistent with the real world through precise description, allowing the robot to learn gait in the simulated physical environment. By setting initial conditions, termination conditions, and reward and punishment conditions, it conducts supervised random divergent trial and error to achieve gait learning.

[0035] This invention addresses the need for adaptation and modification of the Humanoid-Gym reinforcement learning training framework. It provides a complete porting approach and offers improvements to the training framework. Through repeated parameter modifications and structural adjustments for adaptability, a stable, low-energy, and highly robust control strategy is ultimately obtained after training. Specific Implementation Example Nine: The difference between Embodiment Nine and Embodiment Eight of the present invention lies only in: The present invention provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement a bipedal robot gait planning method.

[0037] The method includes the following steps: Step 1: Use the robot's dynamic model as input to achieve neural network inference for simulation; Step 2: Describe the robot using a URDF file, representing the robot as a kinematic chain composed of a series of rigid bodies; Step 3: Through precise description, a robot consistent with reality is built in the simulation environment, allowing it to learn gait in the simulated physical environment. Supervised random divergent trial and error is carried out by setting initial conditions, termination conditions, and reward and punishment conditions to achieve gait learning.

[0038] This invention addresses the need for adaptation and modification of the Humanoid-Gym reinforcement learning training framework. It provides a complete porting approach and offers improvements to the training framework. Through repeated parameter modifications and structural adjustments for adaptability, a stable, low-energy, and highly robust control strategy is ultimately obtained after training. Specific Implementation Example 10: The only difference between Embodiment 10 and Embodiment 9 of the present invention is that: The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a bipedal robot gait planning method.

[0040] The method includes the following steps: Step 1: Use the robot's dynamic model as input to achieve neural network inference for simulation; Step 2: Describe the robot using a URDF file, representing the robot as a kinematic chain composed of a series of rigid bodies; Step 3: Through precise description, a robot consistent with reality is built in the simulation environment, allowing it to learn gait in the simulated physical environment. Supervised random divergent trial and error is carried out by setting initial conditions, termination conditions, and reward and punishment conditions to achieve gait learning.

[0041] This invention addresses the need for adaptation and modification of the Humanoid-Gym reinforcement learning training framework. It provides a complete porting approach and offers improvements to the training framework. Through repeated parameter modifications and structural adjustments for adaptability, a stable, low-energy, and highly robust control strategy is ultimately obtained after training. Specific Implementation Example Eleven: The only difference between Embodiment Eleven and Embodiment Ten of this invention is that: The method of the present invention includes the following steps: Step 1: Use the robot's dynamic model as input to achieve neural network inference for simulation; Step 2: Describe the robot using a URDF file, representing the robot as a kinematic chain composed of a series of rigid bodies; Step 3: Through precise description, a robot consistent with reality is built in the simulation environment, allowing it to learn gait in the simulated physical environment. Supervised random divergent trial and error is carried out by setting initial conditions, termination conditions, and reward and punishment conditions to achieve gait learning.

[0043] This invention addresses the need for adaptation and modification of the Humanoid-Gym reinforcement learning training framework. It provides a complete porting approach and offers improvements to the training framework. Through repeated parameter modifications and structural adjustments for adaptability, a stable, low-energy, and highly robust control strategy is ultimately obtained after training.

[0044] The input to reinforcement learning not only includes information from 15 frames of historical observations but also requires the robot's dynamic model as input to achieve neural network inference in the simulation. In reinforcement learning, robots are typically described using URDF (Unified Robot Description Format) files, which represent the robot as a kinematic chain composed of a series of rigid bodies. Each category describes the position, orientation, mass, moment of inertia, kinematic pair type, visual display, and collision elements of the moving components in their relative motion coordinate system. This precise description is equivalent to building a robot consistent with reality in the simulation environment, allowing it to learn gait within the simulated physical environment. Supervised random divergence (trial and error) can be easily achieved by setting initial conditions, termination conditions, and reward / penalty conditions. To build a high-precision URDF model, this project uses the method of weighing physical parts, assigning the weighed mass data to the part in Solidworks to minimize quality errors caused by manufacturing and difficult-to-assess quality differences, such as in wires, PCBs, and emergency stop switches. In addition, the robot needs to have its collision volume set. A reasonable setting simplifies calculations and can prevent many problems in advance. The leg collision volume is mainly used to detect self-collisions of the robot's limbs, preventing self-collisions during gait and enabling some gait correction. The body (Base) collision volume can be replaced by a cube exceeding the body volume. This reduces the total number of collision surfaces, lowering performance requirements, and allows training sessions to end quickly when the robot's posture is abnormal, improving training efficiency. The foot collision volume is used to determine contact with the ground and solve for the forces acting on the soles of the feet. The overall URDF and collision volumes are as follows: Figure 1 As shown.

[0045] After modeling and representing the robot, the environment function (Env) needs to be modified to match the training environment with the robot to be trained. Specifically, the simulation parameters of the perturbations need to be balanced (Confg) according to the robot's size, mass, and configuration. Increasing the perturbations will result in convergence of the training results, indicating greater robustness of the robot at that degree of freedom. However, unreasonably increasing external perturbations can cause the training results to diverge, making the robot unable to walk or even stand in the simulation environment. Therefore, the perturbation function settings need to be adjusted appropriately based on the actual situation and the robot's characteristics. Furthermore, the training reward function (Reward) also needs to be modified, as its setting directly affects the robot's final behavioral performance. Typically, the reward function sets rewards or penalties for the following behavioral patterns: 1) Velocity tracking: This feature encourages the robot to track specified linear and angular velocities. An error metric function φ(e,w) is used to calculate the error and provide a reward, where e is the error value and w is the weighting coefficient.

[0046] 2) Attitude Tracking: Encourages the robot to maintain the desired attitude, including base height and attitude angle. The attitude tracking error is calculated using the error metric φ(e,w). Velocity Imbalance: Penalizes the robot for velocity deviations in the z, γ, and β directions; excessive deviations are penalized.

[0047] 3) Speed ​​imbalance: Penalizes the robot for speed deviations in the Y, β directions. If the deviation is too large, it will be subject to penalty 3).

[0048] 4) Contact pattern: Encourage the robot's foot contact pattern to be consistent with the preset contact mask Ip(t), which reflects the periodicity of gait.

[0049] 5) Joint position tracking: Encourage robot joints to track the target position.

[0050] 6) Default joints: Encourage robot joints to remain in their default positions.

[0051] 7) Energy Consumption: Penalizes the robot's energy consumption. This may lead to training the robot to walk with straight legs, posing a near-double-break risk. Smooth Movement: Penalizes the robot for drastic changes in movement, ensuring smooth motion. High Contact Force: Penalizes the robot for excessive contact force, ensuring stability of movement.

[0052] In summary, these reward functions are designed to guide the robot to learn stable, smooth, and efficient gaits, achieving zero-shot transfer in real-world environments. Therefore, if the robot exhibits unreasonable behavior exceeding the aforementioned reward / penalty limits during training, the weights of the reward / penalty functions can be modified or additional functions can be added, depending on the specific circumstances of the error. Ultimately achieved as Figure 2 The trained control strategy was ported to MuJoCo, and the robot was found to still be able to walk, proving that the control strategy has transferability and robustness. The simulation dt was 0.004s, and the control strategy takeover dt was 0.02s; therefore, the hardware frequency was 250Hz, and the strategy takeover frequency was 50Hz. Figure 3 The robot is able to walk in the new simulation environment using a trained control strategy.

[0053] In noise addition, besides modifying the noise accumulation mechanism over time, noise related to gravitational acceleration was added to achieve the filtering effect of the IMU. In randomization parameter addition, random phase was added to enable random left and right foot lifting of the robot, making the observed experimental results more generalizable; for example, in the default settings, the robot always lifts its left foot first. When the robot's Sim-to-Real Gap is large, it may lead to the failure to observe general phenomena, instead repeatedly observing surface features of the same phenomenon. During training, randomization of the robot's thrust, motor stiffness coefficient, motor damping coefficient, motor zero-point position, joint resistance, and command noise was added to make the trained model more robust. In the PPO (Proximal Policy Optimization) algorithm, coefficients for bilateral tracking were added, so the symmetry parameters of the trained robot do not need to be set separately, as they are automatically implemented by the PPO fitting function.

[0054] Regarding the reward function settings, a function to determine the degree to which the feet are parallel to the ground has been added to make the gait more stable; a penalty function based on energy loss calculation has been added to make the robot's movements more energy-efficient; and a reward for machine speed tracking has been added to make the robot stay closer to the remote control target during remote control. The above reward functions, along with the original reward function and parameters, are shown in the figure.

[0055] The use of reinforcement learning training frameworks based on Humanoid-Gym requires adaptation and modification. This project provides a complete porting approach and offers improvements to the training framework. Through repeated modification of parameters and structural adjustments to adaptability, a stable, low-energy, and highly robust control strategy was finally obtained for physical deployment.

[0056] The above description is merely a preferred embodiment of a bipedal robot gait planning method. The scope of protection for this bipedal robot gait planning method is not limited to the above embodiments; all technical solutions falling within this conceptual framework are within the scope of protection of this invention. It should be noted that for those skilled in the art, any improvements and variations made without departing from the principles of this invention should also be considered within the scope of protection of this invention.

Claims

1. A gait planning method for a bipedal robot, characterized by: The method includes the following steps: Step 1: Use the robot's dynamic model as input to achieve neural network inference for simulation; Step 2: Describe the robot using a URDF file, representing the robot as a kinematic chain composed of a series of rigid bodies; Step 3: Through precise description, a robot consistent with reality is built in the simulation environment, allowing it to learn gait in the simulated physical environment. Supervised random divergent trial and error is carried out by setting initial conditions, termination conditions, and reward and punishment conditions to achieve gait learning.

2. The method according to claim 1, characterized in that: Each category describes the position, orientation, mass, moment of inertia, kinematic pair type, visual display, and collider elements of the moving component's relative motion coordinate system.

3. The method according to claim 2, characterized in that: The high-precision URDF model uses the method of weighing physical parts and assigns the weighed mass data to the part in Solidworks to minimize the quality errors caused by machining and the quality differences that are difficult to assess.

4. The method according to claim 3, characterized in that: The robot needs to have its collision volume set. The leg collision volume is used to detect self-touching of the robot's limbs, so that the robot's gait will not collide with itself, and to realize some gait correction functions. The collision volume of the body base is replaced by a cube larger than the body volume; The impactor of the foot is used to make contact with the ground, and the forces on the sole of the foot are solved.

5. The method according to claim 4, characterized in that: The perturbation function settings need to be adjusted reasonably according to the actual situation and robot pattern. Modify the training reward function Reward, as the reward function settings directly affect the robot's final behavior. The reward function sets rewards or penalties for the following behavioral patterns: Velocity tracking: This feature encourages the robot to track specified linear and angular velocities using an error metric function. To calculate the error and give the reward, where e is the error value and w is the weighting coefficient; Pose tracking: Encourages the robot to maintain the desired pose, including base height and pose angle, also using an error metric function. To calculate attitude tracking error; Speed ​​imbalance: Penalizes the robot for speed deviations in the z, γ, and β directions; if the deviation is too large, it will be penalized. Contact pattern: Encourages the robot's foot contact pattern to be consistent with the preset contact mask Ip(t), reflecting the periodicity of the gait; Joint position tracking: Encourages the robot's joints to track the target position; Default joints: Encourage robot joints to remain in their default positions; Energy consumption: This penalty affects the robot's energy consumption. This item may train the robot to walk on straight legs, which poses a risk of near-double solution. Smooth motion: Punishes drastic changes in the robot's movements to ensure smooth motion; Large contact force: This is used to punish robots for excessive contact force, in order to ensure the stability of their movement.

6. The method according to claim 5, characterized in that: To achieve zero-sample transfer in a real-world environment, if the robot exhibits unreasonable behavior exceeding the aforementioned rewards and penalties during training, the weights of the reward and penalty functions are modified or additional reward and penalty functions are added based on the specific circumstances of the error. The trained control strategy is then ported to MuJoCo. The simulation dt is 0.004s, and the control strategy takeover dt is 0.02s. Therefore, the hardware frequency is 250Hz, and the strategy takeover frequency is 50Hz.

7. The method according to claim 6, characterized in that: In the noise addition process, in addition to modifying the noise accumulation mechanism over time, noise related to gravitational acceleration was also added to achieve the filtering effect of the IMU; in the addition of randomization parameters, random phase was added to achieve random left and right foot lifting of the robot. During training, random parameters such as robot thrust, motor stiffness coefficient, motor damping coefficient, motor zero point position, joint resistance, and command noise were added to make the trained model more robust. Coefficients for bilateral tracking were added to the PPO algorithm, so that the symmetry parameters of the trained robot do not need to be set separately, but are automatically achieved by the fitting function of PPO.

8. A bipedal robot gait planning system, characterized in that: The system includes: A data input module that takes the robot's dynamic model as input to achieve neural network inference in the simulation; The simulation module describes the robot using URDF files, equating the robot to a kinematic chain composed of a series of rigid bodies. The gait learning module establishes a robot in a simulation environment that is consistent with the real world through precise description, allowing the robot to learn gait in the simulated physical environment. By setting initial conditions, termination conditions, and reward and punishment conditions, it conducts supervised random divergent trial and error to achieve gait learning.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method as claimed in claims 1-7.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the method of claims 1-7.

Citation Information

Cited By

  • Quadruped robot low-noise gait control method and system based on soft landing reward function

    CN121209388A