A humanoid robot obstacle-crossing reinforcement learning control method based on DCM constraint
By designing a reinforcement learning control method for humanoid robots to cross obstacles based on DCM constraints, and using an external force observer and reward function to optimize the robot's obstacle crossing, the instability problem caused by noise and obstacles in traditional methods is solved, and a more efficient and safer obstacle crossing capability is achieved.
Patent Information
- Application Number
- CN202610776907.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-25
AI Technical Summary
When bipedal robots walk/cross in complex environments, traditional methods rely on precise kinematic/dynamic modeling, which cannot effectively cope with environmental noise and obstacles, leading to instability and limitations.
An external force observer based on torque balance and zero-torque point stability criteria is designed. By combining nonlinear finite-time sliding mode terms and adaptive gain, obstacle crossing control is optimized through DCM model and reward function. By combining terrain sampling point grid and multi-dimensional reward function, the stability and safety of robot crossing obstacles are achieved.
It improves the stability and safety of robots crossing obstacles, reduces the fall rate, enhances adaptability to complex terrain, reduces energy consumption and reliance on precise modeling, and achieves end-to-end reinforcement learning control.
Smart Images

Figure CN122632837A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control technology, specifically to a reinforcement learning control method for humanoid robots to cross obstacles based on DCM constraints. Background Technology
[0002] Bipedal robots face various external disturbances when walking / crossing in complex environments. Traditional methods rely on accurate kinematic / dynamic modeling, but noise and obstacles in the environment can have significant adverse effects, thus limiting their effectiveness. Summary of the Invention
[0003] To address the problems existing in the prior art, the purpose of this invention is to provide a humanoid robot obstacle crossing reinforcement learning control method based on DCM constraints.
[0004] To solve the above problems, the present invention adopts the following technical solution.
[0005] A reinforcement learning control method for humanoid robots to cross obstacles based on DCM constraints, characterized by the following steps: An external force observer based on torque balance and zero-torque point stability criteria is designed. By correcting the observer update law and introducing a nonlinear finite-time sliding mode term, the transient response speed of external force estimation is improved and the accumulation of observer estimation error is suppressed.
[0006] Based on the difference between the estimated and calculated external force values, a hyperbolic tangent adaptive gain based on observation error is designed to improve the observer performance by dynamically adjusting the gain.
[0007] Based on the newly established motion divergence (DCM) model that considers the influence of external forces, the correction amount of disturbed motion divergence (DCM) is calculated, and a DCM tracking reward function is designed to ensure dynamic balance across obstacles.
[0008] The robot uses a terrain sampling point grid to detect obstacle height and designs a reward function for the height of the swing leg foot. By adjusting the robot's leg lifting height in real time, the height of the foot is ensured to be higher than the obstacle, enabling the robot to successfully cross the obstacle.
[0009] By combining the relative distance between the robot and obstacles and the structural parameters of its own feet, a foot-landing penalty function is designed to avoid the risk of stepping on obstacles during obstacle crossing.
[0010] To address the shaking issue that can easily occur during obstacle crossing, a motion stabilization reward function is set up to ensure that the robot's movements are smooth and its feet land parallel to the ground.
[0011] As a further improvement of the present invention, the external force estimation interference method in step (1) is based on the linear inverted pendulum dynamic model, torque balance and zero torque point stability criterion constraints, and the following external force observer is designed: in, This represents the estimated external force. Represents the current position of the center of mass. Indicates the estimated gain. This indicates the current position of the zero torque point. Let be the natural frequency of the linear inverted pendulum. A nonlinear finite-time sliding mode term is introduced, and the convergence speed is improved by designing a sliding mode term based on the hyperbolic tangent function. Sliding mode item: Complete formula for estimating external forces: in, For sliding mode gain, For smoothing parameters, For a finite-time power exponent, The difference between the calculated external force and the estimated external force: This external force estimation method adopts a two-layer structure of "linear correction + nonlinear sliding mode compensation" to solve the problem that the traditional linear correction term can only converge asymptotically and has a lag response. A nonlinear sliding mode term based on the hyperbolic tangent function is designed to improve the transient response speed of the observer under sudden impact scenarios, while effectively suppressing the steady-state estimation bias caused by model error and sensor noise.
[0012] As a further improvement of the present invention, the following adaptive variable gain is designed in step (2). : in, and For response parameters, The difference between the calculated value and the estimated value of the external force. This is the error amplification parameter. When encountering a sudden external impact and the estimation error increases significantly, this adaptive variable gain... The gain is rapidly increased to track sudden external forces; when the system tends to stabilize and the estimation error is small, the gain is automatically reduced to suppress observation noise and improve steady-state estimation accuracy. As a further improvement of the present invention, in step (3), the disturbed motion divergence (DCM) is derived and corrected based on the linear inverted pendulum model, and the dynamic balance performance of the robot is improved by designing a DCM tracking reward function. An external force disturbance correction term is designed so that when the robot is subjected to an external horizontal disturbance, the DCM will respond synchronously to the disturbance and adjust the dynamic trajectory to ensure the stability of the robot during walking. The estimated external force term, combined with the dynamic equation and the DCM equation, yields: Based on the calculated modified DCM considering external forces, design the DCM tracking reward function: in The custom error amplification parameter has a value of 0.02. Weights for the reward function As a further improvement to the present invention, in order to enhance the robot's accurate perception of the location, height, and range of obstacles, the following terrain height sampling values are designed: .
[0013] in Let be the transformation matrix. The three-dimensional position of the sampling point in the world coordinate system, including terrain height. Obstacle height information can be extracted from this. Obstacle height information collected based on sampled values. and the height of the foot of the current swinging leg Design a reward function to ensure that the robot lifts its leg above the obstacle.
[0014] in These are the height of the foot and the target leg lift height, respectively. The obstacle height is sensed in real time through terrain sampling grids to adapt to obstacles of different heights, such as protrusions.
[0015] As a further improvement of the present invention, in step (5), the following landing area reward function is designed. in and It is the first Obstacle location, For the reserved foot length threshold, This represents the upper and lower limits of the penalty, i.e. Maximum The minimum value is 0. This is the current landing position. When the distance between the foot and the obstacle is greater than the preset safety threshold, the penalty value is 0, and no additional restrictions are imposed on the strategy. When the relative distance is less than the safety threshold, the penalty value increases monotonically as the distance decreases, with the penalty being stronger the closer the distance. This avoids obstacle areas, effectively improves the safety of landing on obstacles, ensures smooth and continuous execution of obstacle-crossing actions, and avoids posture instability or obstacle-crossing failure caused by stepping on obstacles. As a further improvement of the present invention, in order to solve the problem of easy shaking in the obstacle crossing action involved in step (6), the following action anti-shake reward function is designed: in For real-time joint angles, For joint torque, These represent the current action and the previous action, respectively. By constraining the real-time joint angle change rate, joint torque fluctuation amplitude, and the difference between actions at adjacent moments, the robot is penalized when its movements are too large or its energy consumption is too high, ensuring that the robot's movements transition smoothly and its feet land parallel to each other.
[0016] Beneficial effects of the present invention Compared with the prior art, the advantages of this invention are: 1. The external force observer designed based on the torque balance and zero torque point stability criterion of this invention introduces hyperbolic tangent sliding mode for compensation, which improves the transient response speed of external force estimation and suppresses the accumulation of observer estimation error.
[0017] 2. By estimating external forces, a DCM considering external forces is derived, which can respond quickly to and recover from external disturbances. Compared with the traditional ZMP / CMP method, it has more direct physical meaning and response characteristics.
[0018] 3. Improve obstacle crossing stability and reliability. By combining DCM dynamic constraints and reward function design, the fall rate of the robot during obstacle crossing is effectively reduced, the tendency to tip over is suppressed, and dynamic balance is ensured.
[0019] 4. Enhance the ability to adapt to complex terrain. Relying on the terrain sampling grid to achieve global terrain perception, and with the gradient composite obstacle training environment, the robot can autonomously deal with protrusion-type obstacles with a height of 0.05m~0.3m and a maximum length of 0.2m, and adapt to different obstacle shapes.
[0020] 5. Achieve precise obstacle-crossing motion control. Through a multi-objective reward function system (foot height, landing penalty, anti-shake, etc.), guide the robot to adjust the leg lift height, stride and landing timing in real time to avoid stepping on or kicking obstacles, thereby improving obstacle-crossing safety.
[0021] 6. Reduces modeling dependence and enhances disturbance rejection capability. Employing an end-to-end reinforcement learning model, it eliminates the need for precise kinematic / dynamic modeling and can respond to external disturbances through self-sensing, overcoming the limitations of traditional control methods. Optimize obstacle crossing efficiency and hardware protection by incorporating penalty / reward mechanisms related to joint acceleration, torque, and energy loss to reduce robot energy consumption, minimize hardware wear caused by joint movement jumps, and extend equipment lifespan. Attached Figure Description
[0022] Figure 1 The overall control flowchart of the method of this invention is as follows: First, an external force observer with adaptive variable gain and hyperbolic tangent sliding mode compensation is designed to realize fast and smooth estimation of impact external force; then, an estimated external force correction DCM model is introduced and a tracking reward is set; finally, an obstacle crossing strategy is trained based on multi-dimensional reward, terrain sampling and PPO algorithm to realize stable obstacle crossing control under disturbance. Figure 2 The image shows the observation results of the external force observer, where the solid red line represents the given external force and the dashed blue line represents the external force observed by the observer. It can be seen that after a sudden change in the actual external force (red), the estimated external force (blue) can quickly track it, and the error rapidly converges to zero. The sliding mode term dynamically adjusts with the error, verifying the observer's fast response and chatter-free characteristics. Figure 3 This is a schematic diagram of the sampling points, where the yellow dots represent the sampling points laid out around the robot; Figure 4 To demonstrate the effectiveness of the proposed method in controlling a robot to traverse a raised obstacle 0.25m high and 0.1m wide, the ZMP curve trajectory was obtained. The green dashed lines represent the upper and lower boundaries, and the red vertical segments represent the robot's center of mass reaching the obstacle. It can be seen that the robot exhibits good stability throughout the entire walking process. Figure 5 The ZMP trajectory diagram obtained by controlling the robot to cross a 0.25m high and 0.1m wide boss-like obstacle without using the external force observer, DCM model considering the influence of external forces, and sampling point grid method based on terrain height proposed in this invention shows that the robot is very unstable during walking and falls down after passing the obstacle. Detailed Implementation
[0023] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0024] Please see Figures 1 to 5 A reinforcement learning control method for humanoid robots to cross obstacles based on DCM constraints includes the following steps: Step 1: Based on the system torque balance relationship and the steady-state constraint at the zero torque point, establish the observation update equation and design the following external force observer: Step 2: Introduce a nonlinear finite-time sliding mode term to improve the transient response speed of external force estimation, suppress the accumulation of estimation errors, and solve and output the horizontal external force disturbance value experienced by the robot: Sliding mode item: Complete formula for estimating external forces: The observer's estimation effect is determined by Figure 2 As can be seen, after the real external force in red suddenly changes, the estimated external force in blue can quickly track it, and the error quickly converges to zero. The sliding mode term is dynamically adjusted with the error, which verifies the fast response and chatter-free characteristics of the observer.
[0025] Step 3: Design Adaptive Variable Gain Dynamic adjustment To achieve a rapid response at the moment of impact: Step 4: Derive and correct the perturbed motion divergence (DCM), design the DCM tracking reward function to ensure dynamic equilibrium across obstacles, and apply the estimated external force terms. Combining the dynamic equations and the DCM equations, we can obtain: Based on the calculated modified DCM considering external forces, design the DCM tracking reward function: The error amplification parameter is set to 0.02. By adjusting the weights, the robot's dynamic balance during obstacle crossing is ensured, and the tendency to tip over is suppressed.
[0026] Step 5: Deploy a terrain sampling point grid to optimize terrain perception efficiency. To address the issues of slow training and high computational cost of visual sensors recognizing terrain in the simulation environment, deploy a terrain sampling point grid in the planned area in front of the robot, sampling the terrain height in front of it. .
[0027] By using the terrain height query interface of Isaac Gym, the terrain height data of sampling points can be obtained in real time, accurately perceiving the distance, height, and width information of obstacles. This replaces the traditional visual perception method, significantly reducing training time and improving the real-time performance and accuracy of complex terrain perception.
[0028] Step 6: Based on the real-time obstacle height perceived by the terrain sampling point grid, match the height of the robot's swinging leg foot and design a foot height reward function: Ensure the robot's leg lift height is higher than the obstacle height, enabling adaptive adjustment of the leg lift height to accommodate protrusion-type obstacles of different heights from 0.05m to 0.3m, ensuring successful crossing.
[0029] Step 7: Design a foot-landing penalty function to avoid the risk of stepping on obstacles. Based on the 12cm foot length parameter of the Unitree H1 robot, set a reserved foot length threshold and construct a foot-landing penalty area reward function: Design the following landing area reward function: Step 8: By constraining the real-time joint angle change rate, joint torque fluctuation amplitude, and the difference in motion between adjacent moments, the robot will be penalized when its movements are too large or its energy consumption is too high, ensuring that the robot's movements transition smoothly and its feet land parallel to the ground. Step nine employs the PPO Lagrange algorithm for training to achieve autonomous obstacle crossing. Fractal noise generation technology was introduced to construct diverse and complex training terrains, forcing the robot to adjust its stride, leg lift height, and landing timing. Reinforcement learning training was performed based on the PPO Lagrange algorithm to collaboratively optimize the total reward and constraint cost. Integrating dynamic constraints, gait stability, and energy efficiency, the robot ultimately achieved autonomous traversal of protrusion-like obstacles ranging from a maximum height of 0.3m (three-quarters of the lower leg length) to a maximum length of 0.2m. Figure 4 It is evident that the ZMP trajectory remains stable within the safety boundary during obstacle crossing, demonstrating excellent dynamic balance performance.
[0030] The above description is merely a preferred embodiment of the present invention; however, the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and its improved concepts, should be covered within the scope of protection of the present invention.
Claims
1. A reinforcement learning control method for humanoid robots to cross obstacles based on DCM constraints, characterized in that, Includes the following steps: 1) Design an external force observer based on torque balance and zero torque point stability criteria. Improve the transient response speed of external force estimation and suppress the accumulation of observer estimation error by correcting the observer update law and introducing a nonlinear finite-time sliding mode term. 2) Based on the difference between the estimated and calculated external force values, a hyperbolic tangent adaptive gain update law based on observation error is designed to improve the observer performance by dynamically adjusting the gain. 3) Based on the newly established motion divergence model considering the influence of external forces, the correction amount of disturbed motion divergence is calculated, and a motion divergence tracking reward function is designed to ensure dynamic balance across obstacles. 4) An obstacle height detection method based on terrain sampling point grid is adopted, and a reward function for the height of the swing leg foot is designed. By adjusting the robot's leg lifting height in real time, the height of the foot is ensured to be higher than the obstacle, so that the robot can successfully cross the obstacle. 5) Based on the relative distance between the robot and obstacles and its own foot structure parameters, design a foot-landing penalty function to avoid the risk of stepping on obstacles during obstacle crossing. To address the shaking issue that can easily occur during obstacle crossing, a motion stabilization reward function is set up to ensure that the robot's movements are smooth and its feet land parallel to the ground.
2. The humanoid robot obstacle-crossing reinforcement learning control method based on DCM constraints according to claim 1, characterized in that: The external force estimation disturbance method in step (1) is based on the linear inverted pendulum dynamic model, torque balance, and zero torque point stability criterion constraints. The following external force observer is designed: in, This represents the estimated external force. Represents the current position of the center of mass. Indicates the estimated gain. This indicates the current position of the zero torque point. To determine the natural frequency of the linear inverted pendulum, a nonlinear finite-time sliding mode term is introduced. By designing a hyperbolic tangent function sliding mode term, the convergence speed is improved. Sliding mode item: Complete formula for estimating external forces: in, For sliding mode gain, For smoothing parameters, For a finite-time power exponent, The difference between the calculated external force and the estimated external force: .
3. The humanoid robot obstacle-crossing reinforcement learning control method based on DCM constraints according to claim 2, characterized in that: In step (2), the following adaptive variable gain is designed. : in, and For response parameters, The difference between the calculated value and the estimated value of the external force. This is the error amplification parameter, used to estimate the error when encountering a sudden external impact. It will increase significantly, while adaptive variable gain Rapid improvement enables the external force observer to accurately track sudden external forces; as the system stabilizes, the estimation error decreases. When the gain is low, it automatically decreases to suppress observation noise and improve the accuracy of steady-state estimation.
4. The humanoid robot obstacle-crossing reinforcement learning control method based on DCM constraints according to claim 3, characterized in that: In step (3), the perturbed motion divergence is derived and corrected based on the linear inverted pendulum model. The dynamic balance performance of the robot is improved by designing a motion divergence tracking reward function. An external force disturbance correction term is designed. When the robot is subjected to an external horizontal disturbance, the DCM will respond to the disturbance synchronously and adjust the dynamic trajectory to ensure the stability of the robot when walking. The estimated external force term is used to correct the motion divergence. Combining the dynamic equations and the DCM equations, we can obtain: Based on the calculated modified DCM considering external forces, design the DCM tracking reward function: in The custom error amplification parameter has a value of 0.
02. This represents the weight of the reward function.
5. The humanoid robot obstacle-crossing reinforcement learning control method based on DCM constraints according to claim 4, characterized in that: To improve the robot's accurate perception of obstacle location, height, and range, the following terrain height sampling values are designed: in The transformation matrix is... The three-dimensional position of the sampling point in the world coordinate system, including terrain height. Obstacle height information can be extracted from it. Obstacle height information collected based on sampled values and the height of the foot of the current swinging leg Design a reward function to ensure the robot lifts its leg above the obstacle: in These are the height of the foot and the target leg lift height, respectively. The obstacle height is sensed in real time through terrain sampling grids to adapt to obstacles of different heights, such as protrusions.
6. The humanoid robot obstacle-crossing reinforcement learning control method based on DCM constraints according to claim 5, characterized in that: In step (5), the following landing area reward function is designed: in and It is the first Obstacle location, For the reserved foot length threshold, This represents the upper and lower limits of the penalty, i.e., the maximum is The minimum value is 0. This is the current landing position. When the distance between the foot and the obstacle is greater than the preset safety threshold, the penalty value is 0, and no additional restrictions are imposed on the strategy. When the relative distance is less than the safety threshold, the penalty value increases monotonically as the distance decreases. The closer the distance, the stronger the penalty. This avoids obstacle areas, effectively improves the safety of landing on obstacles, ensures smooth and continuous execution of obstacle crossing actions, and avoids posture instability or obstacle crossing failure caused by stepping on obstacles.
7. The humanoid robot obstacle-crossing reinforcement learning control method based on DCM constraints according to claim 6, characterized in that: To address the issue of jitter in the obstacle-crossing motion involved in step (6), the following motion stabilization reward function is designed: in For real-time joint angles, For joint torque, These are the current action and the previous action, respectively. By constraining the real-time joint angle change rate, joint torque fluctuation amplitude, and the difference between actions at adjacent moments, the robot will be penalized when its movements are too large or its energy consumption is too high, ensuring that the robot's movements transition smoothly and its feet land parallel to each other.