Self-adaptive anti-interference recovery method for humanoid robot based on hierarchical model predictive control enhanced by reinforcement learning

By employing a hierarchical model predictive control method enhanced by reinforcement learning, the multi-step footing sequence and contact wrench of the humanoid robot are optimized, solving the dynamic balance problem of the humanoid robot under complex terrain and external disturbances, and achieving stable and adaptive motion control.

CN121374584APending Publication Date: 2026-01-23BEIJING INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511660480.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing humanoid robots lack dynamic balance capabilities and have poor motion stability under complex terrain and external interference, making it difficult to achieve stable operation.

Method used

A hierarchical model predictive control method based on reinforcement learning is adopted. By combining the reinforcement learning decision module with the hierarchical model predictive control, the multi-step landing point sequence and contact wrench are optimized. Combined with dynamic constraints and event triggering signals, adaptive disturbance rejection and recovery are achieved.

Benefits of technology

It significantly improves the safety and robustness of humanoid robots under unknown disturbances, possesses online self-learning and adaptive capabilities, and can maintain stable movement in complex terrain and external disturbance environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121374584A_ABST
    Figure CN121374584A_ABST
Patent Text Reader

Abstract

The invention discloses a humanoid robot adaptive anti-interference recovery method based on reinforcement learning enhanced hierarchical model prediction control, and the method comprises the steps: receiving a user instruction, a robot state and a model prediction control parameter through a reinforcement learning decision module, and outputting a first-layer MPC parameter and a second-layer MPC parameter which are subjected to strategy enhancement after training; the first-layer MPC is combined with dynamic constraints, the multi-step foothold sequence is optimized and updated through the first-layer MPC parameters with the enhanced output strategy and event triggering signals, the second-layer MPC is combined with the first-layer model to predict and control the optimized robot state, the updated multi-step foothold sequence and the second-layer MPC parameters with the enhanced output strategy are optimized, and the multi-step foothold sequence is optimized. The invention relates to a contact wrench for optimizing a multi-step foothold sequence. According to the method, the walking disturbance recovery agility and the self-adaptive balance capability of the humanoid robot in complex terrains and external disturbance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of humanoid robots, and particularly relates to a humanoid robot adaptive anti-disturbance recovery method based on a reinforcement learning enhanced hierarchical model predictive control. BACKGROUND

[0002] As a kind of anthropomorphism physical carrier of artificial intelligence technology, humanoid robots can assist or replace human beings to complete single repetitive or high-risk work, and its morphological structure is highly similar to human beings. Compared with wheeled robots and tracked robots, humanoid robots have a similar bipedal movement mode to human beings, which makes them have higher terrain adaptation potential and environmental passing capacity when facing uneven terrains such as steps, slopes and rubble piles. Therefore, humanoid robots have broad application prospects in disaster rescue, dangerous operation and national defense and military fields. However, the stable operation of humanoid robots under the action of uneven terrain and external disturbance is one of the key prerequisites for their practical application. The movement of current humanoid robots in complex terrain is easily affected by terrain changes, external impact and control precision limitations, and there are problems such as insufficient dynamic balance ability and poor movement stability. In view of the above problems, it is urgent to study a stable passing control method with adaptive characteristics and anti-disturbance ability to improve the balance keeping and movement robustness of humanoid robots in complex terrain environment. SUMMARY

[0003] In view of the deficiencies in the prior art, the present application provides a humanoid robot adaptive anti-disturbance recovery method based on reinforcement learning enhanced hierarchical model predictive control.

[0004] The present application achieves the above technical purpose through the following technical means.

[0005] The humanoid robot motion adaptive anti-disturbance recovery method based on reinforcement learning enhanced hierarchical model predictive control comprises the following steps:

[0006] The reinforcement learning decision module receives user instructions, robot states and model predictive control parameters, and outputs policy enhanced first layer model predictive control parameters and second layer model predictive control parameters after training;

[0007] First layer model predictive control: combining the kinetic constraints, the output policy enhanced first layer model predictive control parameters and the event trigger signal are used to optimize and update the multi-step landing point sequence;

[0008] Second layer model predictive control: combining the first layer model predictive control optimized and updated multi-step landing point sequence, the robot state and the output policy enhanced second layer model predictive control parameters, the contact wrench of the multi-step landing point sequence is optimized.

[0009] Further, the reinforcement learning decision module adopts standing balance Standing anti-disturbance Walking balance Multi-stage training method for walking disturbance rejection.

[0010] Further, the first-layer model predictive control comprises constructing a target function as follows:

[0011]

[0012] wherein, is the target function of the first-layer model predictive control, is the relative footfall position sequence, is the slack variable, is the center of mass position, is the center of mass velocity, is the n-step footfall position, , is a positive definite matrix, is a non-negative constant.

[0013] Further, the constraints of the first-layer model predictive control comprise:

[0014] Constraint 1: , , ; wherein, is the center of mass state at the end of the current support phase, is the swing time remaining in the current support leg support phase, the swing phase time adjusted by reinforcement learning , is the expected swing phase time, represents the swing phase time residual provided by reinforcement learning, is the swing time of the current single-leg support phase, is the ZMP position relative to the floating base calculated by state estimation and forward kinematics, is the center of mass state estimation value of the robot at the current time, and are both discrete state equation matrices, , respectively represent the center of mass state of the first-step footfall position and the center of mass state of the n-step footfall position ;

[0015] Constraint 2: , and represent the upper and lower bounds of the relative distance between adjacent two footfalls;

[0016] Constraint 3: , is the optimized first-step footfall position, the position calculated for the heuristic foothold control strategy;

[0017] Constraint 4: , the complex adjustable heuristic parameters learned for the reinforcement learning.

[0018] Further, after optimizing the multi-step foothold sequence, the swing leg controller is composed of a trajectory planner and a tracking controller, which is used to accurately track the swing leg trajectory; the tracking controller includes gravity term compensation, reinforcement learning enhanced joint space feedforward and reinforcement learning enhanced task space virtual force feedback.

[0019] The gravity term compensation represents the torque required to be applied by the swing leg joint to maintain the current joint angle posture wherein, is the gravity term vector in the floating base multi-rigid body dynamics equation, is the number of driven joints of the humanoid robot, and q is the generalized position of the robot;

[0020] The reinforcement learning enhanced joint space feedforward represents the PD control of the error between the reference angle and the actual value of the swing leg joint, and the PD control of the error between the angular velocity and the actual value of the swing leg joint wherein, respectively represent the proportional gain and the differential gain in the joint space feedforward controller output by reinforcement learning, respectively represent the P and D coefficient diagonal matrices of the joint space feedforward controller, respectively represent the reference angle and the reference angular velocity of the swing leg joint solved by inverse kinematics, respectively represent the joint angle and the joint angular velocity of the state estimation;

[0021] The reinforcement learning enhanced task space virtual force feedback wherein, respectively represent the proportional gain and the differential gain in the task space virtual force feedback controller output by reinforcement learning, respectively represent the error between the swing foot reference pose and the actual pose, and the error between the swing foot reference angular velocity and linear velocity and the actual velocity, respectively represent the P and D coefficient diagonal matrices of the task space virtual force feedback controller.

[0022] Further, in the second layer model predictive control, the centroid angular velocity of the robot is approximated to be equal to the upper body posture angular velocity, and the posture angle is approximated to be equal to the upper body posture angle, and the centroid dynamics equation is converted into the following form:

[0023]

[0024] where, is the center of mass position, is the center of mass velocity, is the upper body posture angle, is the upper body posture angular velocity, g is the gravity vector, R i is the rotation matrix of the local coordinate system to the inertial coordinate system, f i is the contact force at the contact point P i represents the contact torque, is the total mass of the robot, Ф is the mapping of the derivative of the Euler angle to the angular velocity, is the optimal inertia configuration predicted by reinforcement learning according to the expected action, is the intrinsic identifier, represents the first whether the foot is in contact with the ground.

[0025] Further, the second layer model predictive control includes constructing the following objective function:

[0026]

[0027] where, W is the optimization variable, U is the control variable, , , is the semi-positive definite diagonal weight matrix reinforced by reinforcement learning, N is the number of sampling points, J T is the terminal state target function about the state variable at the last sampling point in the prediction domain, J R is the target function of the state variable and the control variable at each sampling point in the prediction domain, represents the reference value of the state variable.

[0028] Further, the constraints of the second layer model predictive control include:

[0029] Constraint 1: Discrete time domain state space equation where, and are discrete space matrices;

[0030] Constraint 2: represents the estimated center of mass state variable in the current control period;

[0031] Constraint 3: represents the restriction of the contact torque;

[0032] Constraint 4: represents the support leg joint torque restriction, where, ​​a torque constraint parameter, a first support foot contact wrench obtained by model predictive control optimization, upper and lower limits of a support foot joint torque, a Jacobian matrix of the support foot relative to the floating base, a rotation matrix of a local coordinate system of the support foot relative to a coordinate system of the floating base,

[0033] Constraint 5: an augmentation of the first support foot contact wrench, wherein, a floating base height clearance rate.

[0034] The present application has the following beneficial effects:

[0035] (1) The present application proposes a collaborative architecture integrating reinforcement learning and hierarchical model predictive control. The architecture outputs high-level motion parameters including step frequency and MPC weight through an upper-layer reinforcement learning policy. A lower-layer double-sequence MPC is responsible for multi-step foothold point sequence optimization and contact wrench prediction, respectively, realizing the separation and collaboration of strategic decision-making and tactical execution. Reinforcement learning guides parameterization at the strategic level without interfering with the real-time optimization of the MPC bottom layer, so that the system has both learning and adaptive capabilities and strictly guarantees the dynamics constraints, significantly improving safety and robustness under unknown disturbances.

[0036] (2) The present application designs a hierarchical prediction optimization mechanism based on a linear inverted pendulum model (LIPM) and a variable inertia centroid dynamics model. The first-layer MPC uses the efficiency of LIPM for multi-step foothold point planning, and the second-layer MPC integrates a real-time updated variable inertia model for accurate contact force optimization, improving the dynamics prediction accuracy and control performance while ensuring real-time performance.

[0037] (3) The present application establishes a reinforcement learning enhanced strategy based on global dynamics information and MPC performance feedback. Reinforcement learning observes multi-element information including full-order robot state, external disturbance, and lower-layer MPC performance indicators, dynamically adjusts the parameters and constraints of the lower-layer controller, and enables the system to have online self-learning and adaptive recovery capabilities for complex terrain and unknown disturbances. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 Fig. 1 is a schematic diagram of the hierarchical model predictive control process based on reinforcement learning enhancement according to the present application;

[0039] Figure 2 Fig. 2 is a simplified model diagram of a BHR humanoid robot with a floating base according to the present application;

[0040] Figure 3 Fig. 3 is a schematic diagram of a multi-step foothold point gait optimization generator based on a linear inverted pendulum according to the present application.

[0041] Figure 4 Fig. 1 is a schematic diagram of a variable inertia centroid dynamics model according to the present application;

[0042] Fig. 5(a) is a schematic diagram of centroid momentum of a humanoid robot according to the present application when not disturbed;

[0043] Fig. 5(b) is a schematic diagram of centroid momentum of a humanoid robot according to the present application when disturbed backward;

[0044] Fig. 5(c) is a schematic diagram of centroid momentum of a humanoid robot according to the present application when disturbed forward;

[0045] Figure 6 Fig. 6 is a schematic diagram of multi-step contact wrench prediction based on the variable inertia centroid dynamics model according to the present application;

[0046] Fig. 7(a) is a schematic diagram of motion of a humanoid robot according to the present application in the case of flat terrain and small disturbance;

[0047] Fig. 7(b) is a schematic diagram of motion of a humanoid robot according to the present application in the case of uneven terrain and larger disturbance;

[0048] Fig. 7(c) is a schematic diagram of full-body coordinated motion of a humanoid robot according to the present application in the case of large disturbance. DETAILED DESCRIPTION

[0049] The present application will be further described below in conjunction with the accompanying drawings and specific embodiments, but the scope of protection of the present application is not limited thereto.

[0050] The present application proposes a humanoid robot motion adaptive anti-disturbance recovery method based on a hierarchical model predictive control enhanced by reinforcement learning. The method adopts a hierarchical structure combining upper-layer reinforcement learning (RL) decision-making and lower-layer hierarchical model predictive control (MPC), the upper-layer RL incorporates full-order dynamics information and domain randomization, and outputs policy parameters by observing the relevant parameters of the lower-layer MPC and the state of the robot; the lower layer decomposes the nonlinear MPC problem into two layers of linear MPC problems, the first layer of MPC combines with the dynamics constraints, receives the output policy parameters of the RL upper layer and the event trigger signal to optimize and update the multi-step footfall sequence, so as to accurately track the state of the robot, the second layer of MPC combines with the optimized robot state and footfall sequence after the optimization of the first layer of MPC, and combines with the output policy parameters of the RL upper layer, and is responsible for optimizing the contact wrench of the multi-step footfall sequence, so that the support leg controller has the self-learning and rapid recovery ability to the complex terrain and strong disturbance environment.

[0051] As Figure 1As shown, the whole system is divided into two parts: RL decision module and policy-strengthened MPC module. The policy-strengthened MPC module contains: the first layer MPC, which is used for multi-step foothold planning, generates gait sequence, and distinguishes the supporting leg and swing leg through the finite state machine, the swing leg controller is composed of trajectory planner and tracking controller, which is used to accurately track the swing leg trajectory; the supporting leg controller, i.e. the second layer MPC, is used to optimize the supporting leg contact wrench. The swing leg and the supporting leg input the expected torque to the robot simulator, and the robot simulator inputs the sensor data to the state estimator to provide state information for the whole system. The RL decision module receives user instructions, robot state and MPC related parameters, and outputs policy-strengthened MPC related parameters after training.

[0052] I. Modeling of humanoid robot

[0053] Take the BHR humanoid robot of Beijing University of Technology as an example. The whole robot has 20 degrees of freedom, including 12 degrees of freedom in the lower limbs and 8 degrees of freedom in the upper limbs. The torso of the humanoid robot is not actually fixedly connected with the inertial system, but a typical under-actuated floating base system. The floating base (usually the pelvis or torso) is the “root node” connecting all limbs (legs, arms). Its 6-dimensional state (3-dimensional position + 3-dimensional attitude) defines the global pose of the robot in the real world. Without this reference, the positions of all limbs are only relative, meaningless local coordinates, and cannot be associated with the world coordinate system.

[0054] The torso of the robot is connected to the inertial system through 6 virtual joints to establish a unified global reference system for the simplified model of the BHR humanoid robot, as shown in Figure 2

[0055] The generalized position of the humanoid robot is defined as , and the generalized velocity is , is the number of driven joints of the humanoid robot; the generalized position and the generalized velocity of the BHR humanoid robot are:

[0056]

[0057] where x, y, z represent the translational distances of the 3 virtual joints of the robot floating base, respectively, z , y , x represent the rotation angles of the 3 virtual joints of the robot floating base, respectively, represents the vector composed of the rotation angles of each driven joint; the generalized velocity is the rate of change of each element in the generalized position.

[0058] ​The application is based on the data structure of the floating base humanoid robot model constructed by the space vector method, so as to deduce the kinematics and dynamics, and provide a model basis for subsequent motion optimization and stable control algorithm.

[0059] II. First layer MPC

[0060] When the humanoid robot walks on uneven terrain, the selection of the foot placement position plays an important role in maintaining the dynamic stability of the robot. In order to improve the adaptive ability and the rapid recovery ability after suffering external disturbance of the humanoid robot, the first layer MPC: multi-step foot placement gait optimization generator is designed.

[0061] (1) Establish the linear inverted pendulum model discrete time domain dynamics equation:

[0062] The humanoid robot is simplified into a linear inverted pendulum model (LIPM) with constant centroid height, as shown in Figure 3 , and its dynamics equation is expressed as: , wherein the ZMP position , the centroid position , the superscripts x, y, z represent the components in the corresponding coordinate axes of the corresponding coordinate system, , and g represents the gravitational acceleration.

[0063] The centroid position and the centroid velocity of the humanoid robot are taken as the state variables x, and the ZMP position is taken as the control variable u, to establish the continuous time domain state space equation of the LIPM:

[0064] (1)

[0065] , wherein the state variable is the centroid position and the centroid velocity, the foot placement position is the control variable, , and w is the oscillation frequency of the linear inverted pendulum model.

[0066] The established LIPM state space equation is discretized to obtain the following equation:

[0067] (2)

[0068] , wherein is the initial centroid state, is the ZMP position , the centroid state after time, , and are the discrete state equation matrices, which are respectively as follows:

[0069] (3)

[0070] (4)

[0071] LIPM uses the lowest complexity to capture the core of maintaining dynamic balance, and "tailors" the prediction model for MPC, achieving the best balance between speed and accuracy, allowing the first layer MPC to focus on the strategic task of "foot placement planning" without being dragged down by the details of the underlying execution.

[0072] (2) First layer MPC problem construction

[0073] The first layer MPC is a multi-step foot placement gait optimizer. It is assumed that the humanoid robot only has a single leg support phase during movement, and does not include a double leg support phase, that is, the two legs are judged by the gait signal to be a support phase or a swing phase. The presence of a double leg support phase is used to train the stability in the standing balance state, and the transition from a double leg support phase to a single leg support phase is determined by a state signal.

[0074] In the single leg support phase, the geometric center position of the support leg is approximated as the relative position of the ZMP to the position of the robot floating base, providing a stability criterion to ensure that all plans are physically feasible. Given the desired swing phase time and the current single leg support phase swing time , the center of mass state at the end of the current single leg support phase is calculated according to equation (2):

[0075] (5)

[0076] where is the remaining swing time of the current support leg support phase, used to predict the center of mass state at the end of the current support phase , is the swing phase time adjusted by reinforcement learning, which is used to adjust the step frequency adjustment after suffering external disturbances, △T sw represents the swing phase time residual provided by reinforcement learning, and the symbol ~ indicates the reinforcement learning enhancement; is the ZMP relative to the floating base position calculated by state estimation and forward kinematics, is the current time estimate of the robot's center of mass state.

[0077] In order to enable the humanoid robot to maintain the desired motion speed when moving on uneven terrain by adjusting the foot placement, the present application optimizes the relative foot placement positions of future multiple steps through the design of the first layer MPC. Specifically, the construction of this optimization problem is to approximate the relative foot placement positions of future multiple steps as the positions of the ZMP relative floating base of future multiple steps, and take it as the optimization variable, thereby bringing the foot placement planning problem into the predictive control framework. Formula (5) represents the optimization variable as a discrete value. According to formula (2), the center of mass state at the end of the swing phase of future N steps is calculated as:

[0078] (6)

[0079] Among them, 、 respectively represent the center of mass state of the first step foot placement position and the n th step foot placement position . Considering the prediction domain length of the model predictive controller of the subsequent support leg, the present application sets N to 4.

[0080] The MPC problem of the present application for optimizing the relative foot placement positions of future steps is as follows:

[0081] (7)

[0082] Among them, is the objective function of the first layer MPC, and the optimization variables include the relative foot placement position sequence and the slack variable , which is used to calculate the relative position error between the first foot placement of the gait optimization generator and the foot placement position calculated by using the heuristic strategy.

[0083] The objective function of the MPC optimization problem proposed by the present application consists of three parts: first, the error penalty of the center of mass velocity at the i th step foot placement and the expected center of mass velocity , second, the relative distance penalty between adjacent foot placements , and finally the penalty of the slack variable . Among them, 、 are positive definite matrices, is a non-negative constant.

[0084] The constraints of the first layer MPC optimization problem in the present application include:

[0085] Constraint 1: The above-mentioned LIPM dynamics equation constraint, corresponding to (5) and (6), constraint 1 ensures that the optimized foot placement sequence is dynamically feasible, so that the foot placement sequence can make the center of mass motion remain stable.​

[0086] Constraint 2: to represent the kinematic accessibility of humanoid robot, and to represent the upper and lower bounds of the relative distance between two adjacent footholds, which together with the relative distance penalty in the objective function above, optimize the multi-step foothold sequence to be within the reachable set.

[0087] Constraint 3: to limit the position of the first-step foothold optimized to be close to the position calculated by the heuristic foothold control strategy .

[0088] Constraint 4: for the calculation of the heuristic foothold position, based on the error between the desired centroid velocity and the actual centroid velocity, to achieve acceleration and deceleration by controlling the relative position of the foothold, to strengthen the learning of complex adjustable heuristic parameters, to provide a good initial guess for the lower MPC, so that the MPC does not need to search for solutions in a huge space, but only needs to fine-tune in a space reduced and optimized by reinforcement learning, which greatly improves the success rate and efficiency, and is more forward-looking than giving fixed parameters manually.

[0089] Since the constraints are linear, the optimal control problem in equation (7) can be converted into a quadratic programming (QP) problem, and since the problem only contains 5 optimization variables, it can be solved quickly using a QP solver. The MPC in the multi-step foothold gait optimization generator designed in the present application uses the Eigen-QuadProg solver to complete the solution.

[0090] (3) Swing leg trajectory planning and tracking control

[0091] Reinforcement learning can learn to use what gait and step frequency to ensure the stability of the robot according to the different disturbances and external forces generated by different terrains in the training environment, which is difficult to manually design based on rule-based controllers. The swing leg trajectory planner and tracking controller described in the present application form a closed-loop control link of planning-execution:

[0092] The trajectory planner generates a continuous and smooth trajectory of the foot end in the task space in real time based on the first sequence of MPC-optimized foothold sequences, the current gait phase and the robot state, including position, velocity and acceleration profiles, through high-order polynomials;

[0093] The tracking controller, as the lower-level executor, receives the output of the trajectory planner and realizes accurate trajectory tracking through a parallel double-loop control architecture.

[0094] The trajectory planner of the swing leg in the application is to plan the swing leg by using a 5th order polynomial The axis and The axis position, the swing leg is planned by using a 6th order polynomial The axis direction position, are all relative to the reference position of the robot floating base. The height of the swing leg is , including the expected swing leg height, the height residual provided by the reinforcement learning according to different terrains When the robot moves on flat ground, it does not need to lift too high to ensure the stability of movement, and when it is subjected to external disturbance, it may need to lift the height of the swing leg to ensure the stability of the robot, thereby ensuring the effective use of energy.

[0095] The above polynomial equation is expressed as follows:

[0096] , (8)

[0097] , (9)

[0098] Among them, And are polynomial coefficients, the size of which changes with the swing phase time and the expected next foot landing position, and the dynamically changing polynomial coefficients are as follows:

[0099] (10)

[0100] (11)

[0101] Among them, is the initial position of the swing leg relative to the robot floating base at the beginning of the current swing phase, and the value thereof is calculated by state estimation and forward kinematics, is the first foot landing position obtained by the first layer MPC optimization at the current swing phase time, is the expected lifting height of the swing leg.

[0102] The tracking controller of the swing leg in the application is composed of three parts, which are gravity term compensation, reinforcement learning enhanced joint space feedforward and virtual force feedback in task space. That is, the joint torque of the swing leg is obtained by adding the torque vectors of the above three parts:

[0103] (12)

[0104] The gravity term compensation represents the torque required to be applied by the swing leg joint to maintain the current joint angle posture, which is obtained by calculating the multi-rigid body dynamics equation:

[0105] (13)

[0106] where, is the gravity term vector in the floating-base multi-body dynamics equation.

[0107] The reinforcement learning enhanced joint space feedforward controller represents the PD control of the error between the reference angle and angular velocity of the swing leg joints and the actual values, denoted as:

[0108] (14)

[0109] where, and represent the proportional gain and the derivative gain in the joint space feedforward controller output by the reinforcement learning, respectively, and represent the P and D coefficient diagonal matrices of the controller, respectively, and represent the reference angle and the reference angular velocity of the swing leg joints solved by inverse kinematics, respectively, and represent the joint angle and the joint angular velocity of the state estimation, respectively.

[0110] The reinforcement learning enhanced task space virtual force feedback controller is represented as:

[0111] (15)

[0112] where, and represent the proportional gain and the derivative gain in the task space virtual force feedback controller output by the reinforcement learning, respectively, and represent the error between the reference pose and the actual pose of the swing foot, and the error between the reference angular velocity and linear velocity and the actual velocity of the swing foot, respectively, and represent the P and D coefficient diagonal matrices of the task space virtual force feedback controller, respectively.

[0113] III. Second layer MPC

[0114] After the multi-step footfall position sequence obtained by the first layer MPC enhanced based on reinforcement learning, the contact wrench (i.e. force and torque) interacting with the environment is optimized, which is very important for the humanoid robot to maintain stability and track the required speed. Therefore, the second layer MPC based on the variable inertia center of mass dynamics model is proposed, which is expressed as a convex MPC problem and can be solved in real time.

[0115] (1) Variable inertia center of mass dynamics model

[0116] The center of mass dynamics model of the humanoid robot is often applied in the field of motion planning and control. By regarding the robot as a variable inertia ellipsoid model moving in space, such as Figure 4The effects of the contact force on the robot's center of mass position and center of mass momentum are considered when the feet positions and contacts are taken into account. The transformations of the center of mass momentum of the robot when it is disturbed in different degrees are shown in Figures 5(a), (b), and (c). The center of mass dynamics model considers more dimensions than the simplified model, especially the introduced center of mass angular momentum, which makes the motion planning and control based on the model more accurate.

[0117] The dynamics equation of the center of mass dynamics model is as follows:

[0118] (16)

[0119] wherein, represents the center of mass linear momentum and the angular momentum around the center of mass of the robot, respectively, is the total mass of the robot, is the gravity vector, represents the i-th only the position of the contact foot in the world coordinate system, represents the contact force at the contact point . represents the contact torque.

[0120] The center of mass momentum is the projection and of the linear momentum and the angular momentum of all the robot links in the center of mass coordinate system, and is expressed as:

[0121] (17)

[0122] wherein, represents the moment of inertia of the ellipsoid in the center of mass coordinate system, represents the angular velocity of the ellipsoid. In the known robot rigid body model and the generalized position , the moment of inertia is calculated using the spatial vector dynamics.

[0123] Considering that the change of the angular velocity of the ellipsoid closely follows the change of the upper body, the center of mass angular velocity of the robot is approximated to be equal to the upper body attitude angular velocity, and the attitude angle is also approximated to be equal to the upper body attitude angle, i.e.:

[0124] , (18)

[0125] wherein, is the upper body attitude angular velocity of the robot, θ G is the attitude angle, and θ fb is the upper body attitude angle of the robot.

[0126] Based on the above approximation, the center of mass dynamics equation is converted into the following form: ​

[0127] (19)

[0128] Among them, the state variables include the position of the centroid. Center of mass velocity Upper body posture and angle and upper body angular velocity ; Indicates the first Whether the foot makes contact with the ground, and when it does, ,otherwise, ; It is the rotation matrix for transforming from the local coordinate system to the inertial coordinate system. It is the optimal inertia configuration predicted by reinforcement learning based on the expected action. It is updated in each MPC cycle and is considered a constant within a single cycle. It is By convention, the matrix mapping the derivative of Euler angles to angular velocity is obtained from... The result is obtained, where "ZYX" refers to a specific rotation sequence convention for Euler angles, also commonly known as the "roll-pitch-yaw" sequence. In control systems, As inherent identifiers, they respectively refer to the actual physical positions of the left and right feet. The first sequence of MPC optimization generates the landing point sequence. It is the expected future foot placement position. Through a dynamic mapping relationship based on gait phase, the optimized foot placement position is assigned in real time to the physical foot position in the swing phase as the target command for its position servoing.

[0129] Gravitational constant Also used as components of state variables, these are augmented to construct standard linear state equations:

[0130] (20)

[0131] Based on (19) and (20), the continuous-time state-space equation is obtained:

[0132] (twenty one)

[0133] in, To control the amount, including the contact wrench applied to both feet, subscript and Representing the right and left feet. Matrices A and B are derived from (19):

[0134] (twenty two)

[0135] (23)

[0136] where, .

[0137] Discretize the matrix A to get the discrete time domain state space equation:

[0138] (24)

[0139] where, and are discrete space matrices, which can be obtained by forward Euler method:

[0140] (25)

[0141] (26)

[0142] where, is the prediction domain sampling time interval.

[0143] Equation (23) contains the mass center position of the state variable , where the mass center trajectory in the prediction domain is combined with the reaction formula multi-step foot point MPC optimization to obtain the foot point sequence, which is brought into to obtain the matrix at each sampling point in the prediction domain.

[0144] (2) Construction of the second layer MPC problem

[0145] After discretizing the state equation of the variable inertia mass center dynamics model, the multi-step foot point contact wrench optimization is constructed into an MPC problem as follows:

[0146] (27)

[0147] where, the optimization variable is , which represents the sequence of double foot contact wrenches in the prediction domain, , , is a semi-positive definite diagonal weight matrix strengthened by reinforcement learning, N is the number of sampling points, is the terminal state target function about the state variable at the last sampling point in the prediction domain, is the target function of the state variable and the control variable at each sampling point in the prediction domain, represents the reference value of the state variable, which is represented as: , , , , is the mass center reference position, mass center reference velocity, expected Euler angle and expected angular velocity in the current control period.

[0148] The constraints of the second layer MPC optimization problem in the present application include:

[0149] Constraint 1: The constraint of the variable inertia centroid dynamics equation mentioned above, corresponding to (24).

[0150] Constraint 2: , represents the estimated centroid state quantity in the current control period.

[0151] Constraint 3: , represents the contact wrench constraint.

[0152] Constraint 4: , represents the support leg joint torque constraint, torque constraint parameter is given by reinforcement learning.

[0153] Constraint 5: is the augmented contact wrench of the first support foot, the contact wrench of the first support foot obtained by MPC optimization and the floating base height clearance rate are superimposed to obtain , represents the upper and lower limits of the support foot joint torque, is the Jacobian matrix of the support foot relative to the floating base, is the rotation matrix of the local coordinate system of the support foot relative to the floating base coordinate system.

[0154] As shown in Figure 6 , it is the multi-step contact wrench prediction diagram based on the variable inertia centroid dynamics model proposed in the present application. As with the above-mentioned first layer MPC problem, in order to improve the solving speed, the present application converts the second layer sequence MPC problem into a QP problem, and uses the Eigen-QuadProg solver to solve the converted QP problem.

[0155] Four, reinforcement learning upper strategy

[0156] The goal of the reinforcement learning strategy is to estimate the action parameters required by the MPC layer. The problem of finding these action parameters is regarded as a Markov Decision Process (MDP). In the training process, at each time step , the Agent receives an observation from the observation space, performs an action according to its current policy , and obtains a nominal reward . The ultimate goal of the strategy is to maximize the reward discount return maximization: , is the discount coefficient, represents the expected value based on the policy .

[0157] The state space, action space and reward space designed by the application are as follows:

[0158] (1) State space

[0159] Complete state space It is difficult to completely measure the state space by the sensor system in the humanoid robot motion task, but it can be approximated by the compression and learning of the environment observation space in the process of environment state transition Information level purpose.

[0160] In the application, the RL upper layer receives the robot state observed from the state space , which is an observation sequence with high-dimensional hidden historical information . The robot upper body posture angle that can be observed by the robot sensor data , the robot upper body posture angular velocity , the robot upper body posture velocity , the robot upper body posture angular acceleration , the generalized position , the generalized velocity , the end contact wrench , the first layer MPC related parameters , the second layer MPC related parameters , and the action of the previous time step are directly observed parts , and the trunk motion estimation information and other non-ontological sensor information are auxiliary observation parts , and finally combined with the user instruction to form a complete environment state space .

[0161] (2) Action space

[0162] The action space of the application does not contain the related parameters of the joint space, and it contains the heuristic parameters output to the first layer MPC , the swing phase time , the swing leg height residual error , the positive definite matrix , the non-negative constant , the slack variable , the PD gain parameters related to the swing leg , the semi-positive diagonal weight matrix of the second layer MPC , the moment of inertia , and the torque constraint parameter .

[0163] (3) Reward space

[0164] reward is the core measure of continuous evaluation of reinforcement learning training effect, is the direct basis of humanoid robot adjusting behavior preference between environment and user. Unlike end-to-end RL, this kind of comprehensive utilization of MPC solution as prior knowledge provides sample efficient and reliable strategy learning.

[0165] In the present application, task reward and stage reward are designed, task reward encourages robot to walk according to instruction, including linear velocity tracking, vertical linear velocity penalty, angular velocity penalty, trunk posture penalty, trunk height penalty, energy loss, joint acceleration penalty, foot position penalty, wherein the addition of energy loss reward will minimize the energy consumption of the robot, which is reflected in the fast / slow step frequency and torque limit parameter adjustment; stage reward is divided into standing and walking, wherein the corresponding weight and disturbance setting are different.

[0166] (4) training details

[0167] The present application uses proximal policy optimization (PPO) algorithm to optimize the upper policy of RL, and is matched with asymmetric actor-critic architecture (i.e. strategy network and performance evaluation network architecture, critic uses more state information to improve stability), the strategy network and the performance evaluation network architecture both adopt single three-layer MLP, the size is (512, 256, 128). In the training, the present application adopts domain randomization strategy to randomize the following key parameters:

[0168] environment interaction parameters: randomize the dynamic / static friction coefficient between the foot and the ground, and the height field, slope and surface stiffness of the terrain;

[0169] robot body parameters: randomize the total mass and mass distribution, actual position of the center of mass, and joint arm inertia of the robot within the preset range;

[0170] external disturbance parameters: in the process of robot movement, external thrust and torque disturbance of different size, direction and duration is applied to the random position on the body.

[0171] The present application adopts standing balance standing anti-interference walking balance The multi-stage training mode of walking anti-interference is completed, and the motion self-adaptation and anti-interference recovery training of the humanoid robot in the unknown terrain environment are completed. As shown in FIG. 7 (a) and (b), the motion conditions reflect that after being strengthened by reinforcement learning, the robot adapts to the terrain and external disturbance changes with different foot lifting heights, and the motion condition of FIG. 7 (c) reflects that when subjected to external large disturbance, the upper strategy guides the inertia configuration of each joint of the robot to achieve the compensation of the center of mass momentum, and then realize the anti-interference recovery at a fast speed.

[0172] The embodiments are preferred embodiments of the present application, but the present application is not limited to the above embodiments, and any obvious improvements, replacements or modifications made by those skilled in the art without departing from the essential content of the present application shall fall within the protection scope of the present application.

Claims

1. A method for humanoid robot motion adaptive anti-disturbance recovery based on reinforcement learning enhanced hierarchical model predictive control, characterized in that: A reinforcement learning decision module receives user instructions, robot states and model predictive control parameters, and outputs policy enhanced first layer model predictive control parameters and second layer model predictive control parameters after training; First layer model predictive control: combining with the dynamics constraint, the output policy enhanced first layer model predictive control parameters and the event trigger signal are used to optimize and update the multi-step foothold sequence; Second layer model predictive control: combining the first layer model predictive control optimized and updated multi-step foothold sequence and robot state, the output policy enhanced second layer model predictive control parameters, the contact wrench of the multi-step foothold sequence is optimized.

2. The humanoid robot motion self-adaptation and disturbance recovery method according to claim 1, characterized in that, Reinforcement learning decision module employs standing balance Standing disturbance rejection Walking balance Multi-stage training approach for walking disturbance rejection.

3. The humanoid robot motion self-adaptation and disturbance recovery method according to claim 1, characterized in that, The first layer model predictive control comprises constructing a target function as follows: wherein, is the target function of the first layer model predictive control, is a relative landing point position sequence, is a slack variable, is a centroid position, is a centroid velocity, is an n-step landing point position, , is a positive definite matrix, is a non-negative constant.

4. The humanoid robot motion self-adaptive and disturbance recovery method according to claim 3, characterized in that, The constraints of the first layer model predictive control include: Constraint 1: , , ; wherein, is the center of mass state at the end of the current support phase, is the swing time remaining in the current support leg support phase, swing phase time adjusted by reinforcement learning , is the desired swing phase time, ΔT sw represents the swing phase time residual provided by reinforcement learning, is the swing time of the current single leg support phase, is the position of the ZMP relative to the floating base calculated by state estimation and forward kinematics, is the center of mass state estimate of the robot at the current time, and are both discrete state equation matrices, , respectively represent the center of mass state of the first step landing point position and the center of mass state of the n-step landing point position . Constraint 2: , and denote the upper and lower bounds of the relative distance between two adjacent footprints. Constraint 3: , the position of the first step of the optimization foot point, the position calculated by the heuristic foot point control strategy; Constraint 4: , Complex tunable heuristic parameters learned for reinforcement learning.

5. The humanoid robot motion self-adaptation and disturbance recovery method according to claim 1, characterized in that, After optimizing and updating the multi-step foothold sequence, the finite state machine is used to distinguish the supporting leg and the swing leg, the swing leg controller is composed of a trajectory planner and a tracking controller, and is used to accurately track the swing leg trajectory; the tracking controller includes a gravity term compensation, a reinforcement learning enhanced joint space feedforward and a reinforcement learning enhanced task space virtual force feedback; The gravity term compensation represents a moment of force required to be exerted by the swing leg joint to maintain the current joint angle posture wherein, is a gravity term vector in the floating-base multi-body dynamics equation, is the number of driven joints of the humanoid robot, and q is the generalized position of the robot. PD control of an error of a reference angle of a swing leg joint from an actual value, PD control of an error of an angular velocity of the swing leg joint from an actual value wherein, respectively represent a proportional gain, a differential gain in the joint space feedforward controller output by reinforcement learning, respectively represent P, D coefficient diagonal matrices of the joint space feedforward controller, respectively represent a reference angle, a reference angular velocity of the swing leg joint solved by inverse kinematics, respectively represent a joint angle, a joint angular velocity of the state estimation; The reinforcement learning enhanced task space virtual force feedback wherein, respectively represent proportional gain, derivative gain in the reinforcement learning output task space virtual force feedback controller, respectively represent error of swing leg reference pose and actual pose, error of swing leg reference angular velocity and linear velocity and actual velocity, respectively represent task space virtual force feedback controller P, D coefficient diagonal matrix.

6. The humanoid robot motion self-adaptation and disturbance recovery method according to claim 1, characterized in that, In the second layer model predictive control, the centroid angular velocity of the robot is approximately equal to the upper body posture angular velocity, and the posture angle is approximately equal to the upper body posture angle, and the centroid dynamics equation is transformed into the following form: wherein, is the center of mass position, is the center of mass velocity, is the upper body pose angle, is the robot upper body pose angular velocity, g is the gravity vector, R i is the rotation matrix of the local coordinate frame transformation to the inertial coordinate frame, f i is the contact force at the contact point P i denotes the contact torque, is the total mass of the robot, Ф is the force-torque vector, is the matrix that maps the derivative of the Euler angles to the angular velocity, is the optimal inertia configuration predicted by the reinforcement learning from the expected action, is the intrinsic identifier, denotes the whether the sole is in contact with the ground.​ 7. The humanoid robot motion self-adaptation and disturbance recovery method according to claim 6, characterized in that, The second layer model predictive control comprises constructing a target function as follows: wherein W is an optimization variable, U is a control variable, , , is a semi-positive definite diagonal weight matrix reinforced by reinforcement learning, N is a number of sampling points, J T is a terminal state target function about the state variable at the last sampling point in a prediction domain, J R is a target function of the state variable and the control variable at each sampling point in the prediction domain, denotes a reference value of the state variable.

8. The humanoid robot motion self-adaptation and disturbance recovery method according to claim 6, characterized in that, The constraints of the second layer model predictive control include: Constraint 1: Discrete-time state-space equation where, and is a discrete-space matrix; Constraint 2: xkdenotes the estimated state of the center of mass at the current control period; Constraint 3: , indicating the limit of contact with the wrench; Constraint 4: represents the support leg joint torque limit, where, is the torque constraint parameter, is the first support foot contact wrench obtained by model predictive control optimization, represents the upper and lower limits of the support foot joint torque, is the Jacobian matrix of the support foot relative to the floating base, is the rotation matrix of the local coordinate system of the support foot relative to the coordinate system of the floating base; Constraint 5: For the first support foot to contact the wrench augmentation, where, is the floating base height vacancy rate.

Citation Information

Cited By

  • Long-time-domain robot autonomous operation method based on hierarchical reinforcement learning

    CN121696986A