A state feedback model predictive control method and system for a legged robot

Through the state feedback model predictive control method, the closed-loop state observer and model predictive control strategy MPC are combined with constraint relaxation WBC to optimize the joint torque, which solves the stability and reliability problems of the leg-foot robot under uncertainty and external disturbances and achieves more efficient control performance.

CN119937611BActive Publication Date: 2025-10-14SHANDONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510101628.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-10-14
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing control methods for legged robots lack stability and reliability when faced with uncertainty and external disturbances. In particular, model predictive control methods have difficulty maintaining effective control performance when dealing with model uncertainty and external disturbances.

Method used

The state feedback model predictive control method is adopted to track the dynamic disturbance information of the leg-foot robot through a closed-loop state observer, and the model predictive control strategy MPC is used to predict the motion state at future moments. Combined with constraint relaxation WBC, the MPC solution is optimized secondary to calculate the joint torque and enhance stability and reliability.

Benefits of technology

The control performance of the legged robot is improved, its stability and reliability in the face of uncertainty and external disturbances are enhanced, and it can more accurately capture dynamic changes and reduce static errors in state tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937611B_ABST
    Figure CN119937611B_ABST
Patent Text Reader

Abstract

The application provides a state feedback model predictive control method and system for a leg-foot robot, and belongs to the technical field of robot control; firstly, disturbance information in a single rigid body dynamics model of the leg-foot robot is dynamically tracked by using a closed-loop state observer; subsequently, the motion state of the leg-foot robot at a future time is predicted by using a model predictive control strategy MPC according to a set expected motion state of the leg-foot robot and a disturbance prediction value, and the ground reaction force meeting the MPC constraint condition is solved; finally, the MPC solving result is secondarily optimized by using a constraint relaxation WBC, and the joint torque of the leg-foot robot is calculated. The application can improve the control performance of the leg-foot robot, and enhance the stability and reliability of the leg-foot robot when facing uncertainty and external disturbance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robot control, and in particular relates to a state feedback model predictive control method and system for a leg-foot robot. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Virtual model controllers (VMCs) are widely used in the motion control of legged robots, attracting attention for their low computational complexity and fast response speed. VMCs construct appropriate virtual components at each degree of freedom to be controlled to generate appropriate virtual forces. These virtual forces are transformed through the action of actuators. For some control problems, it is necessary to map forces or torques in the workspace into joint torques in the joint space. However, VMC methods have significant limitations in long-term motion planning, as they generally lack the ability to fully predict and plan the robot's future motion states.

[0004] In contrast, model predictive control (MPC) technology has gained popularity in the field of legged robot control due to its ability to plan long-term and handle multiple constraints. MPC builds a dynamic model of the system and uses this model to predict the system's future behavior at each control moment. Based on these predictions, it generates an optimized control sequence. MPC involves the basic steps of system modeling, prediction, and optimization. It solves an optimization problem to compute the optimal control input sequence to maximize or minimize a performance metric.

[0005] However, despite its significant theoretical advantages, MPC faces challenges in practical deployment, such as dealing with model uncertainties, such as modeling errors and external disturbances. These uncertainties can lead to degraded control performance and even compromise system stability. Effectively handling model uncertainties and external disturbances to achieve stable walking has become a major technical challenge in the control of legged robots.

[0006] For example, patent application number 202211050051.6 is a method for bipedal robot motion control based on deep reinforcement learning. This method relies on a large amount of training data to learn the control strategy. Although this method may perform well in a simulation environment, its generalization ability may be limited when faced with complex and changeable situations in the real world; in particular, when the external disturbance exceeds the range of the training data, the algorithm may not be able to cope effectively. In addition, external disturbances have a significant impact on the optimal control problem of nonlinear systems. Disturbances may cause changes in the state equations and control equations of the system, thereby changing the optimal control strategy; at the same time, disturbances may also introduce uncertainty factors, making the solution of the optimal control problem more difficult. Therefore, the stability and reliability of the legged robot under this method in the face of uncertainty and external disturbances are still not ideal. Summary of the Invention

[0007] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a state feedback model predictive control method and system for a leg-leg robot, which can enhance the stability and reliability of the leg-leg robot in the face of uncertainty and external disturbances on the basis of improving the control performance of the leg-leg robot.

[0008] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0009] A first aspect of the present invention provides a state feedback model predictive control method for a legged robot.

[0010] A state feedback model predictive control method for a legged robot, comprising:

[0011] Obtaining the state information of the leg-foot robot; deriving the single rigid body dynamics model of the leg-foot robot based on the obtained state information and discretizing it;

[0012] The closed-loop state observer is used to dynamically track the disturbance information in the single rigid body dynamics model of the legged robot.

[0013] According to the expected motion state and disturbance prediction value set by the leg-foot robot, the model predictive control strategy MPC is used to predict the motion state of the leg-foot robot at the future moment and solve the ground reaction force that meets the MPC constraints.

[0014] Constraint relaxation WBC is used to perform secondary optimization on the MPC solution and calculate the joint torque of the leg-foot robot.

[0015] Furthermore, based on the obtained state information, the single rigid body dynamics model of the legged robot is derived and discretized, including: mathematically representing the model predictive control strategy MPC by solving a constrained optimization problem, and differentially defining the generalized coordinate vector in the mathematical representation; mapping the mathematical representation in the model predictive control strategy MPC optimization problem to a coordinate system, and determining the single rigid body dynamics model of the legged robot; and predicting the motion state of the legged robot at future moments based on the single rigid body dynamics model.

[0016] Furthermore, before predicting the motion state of the legged robot at a future moment based on the single rigid body dynamics model, the single rigid body dynamics model is first converted from a nonlinear model to a linear model containing time-varying disturbances.

[0017] Furthermore, the closed-loop state observer is provided with state variables related to tracking the motion state of the leg-foot robot, and the state variables include: the actual state of the leg-foot robot, the reconstructed state of the closed-loop state observer, and the predicted state of the model predictive control strategy MPC for the future moment.

[0018] Furthermore, it also includes quadratic programming based on input variation, specifically: defining the state tracking error, converting the cost function involved in the model predictive control strategy MPC into a discrete domain representation and defining the input variation; writing the input variation as an optimization variable into the defined state tracking error to achieve quadratic programming based on the input variation.

[0019] Furthermore, the model predictive control strategy MPC takes the closed-loop state observer equation as the prediction benchmark, uses the attenuation matrix to approximate the time domain evolution in the prediction domain, and then makes a preliminary prediction of the motion state of the legged robot at the next moment.

[0020] A second aspect of the present invention provides a state feedback model predictive control system for a legged robot.

[0021] A state feedback model predictive control system for a legged robot, comprising:

[0022] The state information acquisition module is configured to: obtain state information of the leg-foot robot; derive a single rigid body dynamics model of the leg-foot robot based on the obtained state information and discretize the model;

[0023] The closed-loop state observation module is configured to: dynamically track disturbance information in a single rigid body dynamics model of the legged robot using a closed-loop state observer;

[0024] The state feedback MPC module is configured to: predict the future motion state of the leg-foot robot using the model predictive control strategy MPC based on the desired motion state and disturbance prediction value set by the leg-foot robot and solve the ground reaction force that meets the MPC constraints;

[0025] The constraint relaxation WBC module is configured to: perform secondary optimization on the MPC solution using the constraint relaxation WBC and calculate the joint torque of the legged robot;

[0026] The state estimation module is used to feed back the posture information of the leg-foot robot to the state feedback MPC module and the constraint relaxation WBC module.

[0027] Furthermore, the closed-loop state observation module is integrated into the state feedback MPC module.

[0028] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a state feedback model predictive control method for a legged robot as described in the first aspect of the present invention.

[0029] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and runnable on the processor. When the processor executes the program, it implements the steps in the state feedback model predictive control method for a leg-foot robot as described in the first aspect of the present invention.

[0030] One or more of the above technical solutions have the following beneficial effects:

[0031] The present invention first uses a closed-loop state observer to dynamically track the disturbance information in the single rigid body dynamics model of the leg-legged robot; then uses the model predictive control strategy MPC to predict the motion state of the leg-legged robot at future moments and solve the ground reaction force that meets the MPC constraint conditions; finally, the constraint relaxation WBC is used to perform secondary optimization on the MPC solution results and calculate the joint torque of the leg-legged robot. By converting the single rigid body dynamics model (nonlinear floating basis dynamics model) containing uncertainty into an augmented linear model containing time-varying disturbance terms, and using the closed-loop state observer equation as the benchmark for MPC state prediction, it is possible to accurately capture the dynamics of the leg-legged robot; at the same time, by constructing an optimization problem based on the input variation, the static error in the MPC state tracking is reduced. Therefore, the present invention can enhance the stability and reliability of the leg-legged robot in the face of uncertainty and external disturbances on the basis of improving the control performance of the leg-legged robot.

[0032] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0034] Figure 1 This is a flow chart of a state feedback model predictive control method for a legged robot in Example 1 of the present invention.

[0035] Figure 2 This is a schematic diagram of mapping the mathematical representation of the MPC optimization problem to a coordinate system in the first embodiment of the present invention.

[0036] Figure 3 Schematic diagram of the impact resistance simulation test process in Example 1 of the present invention.

[0037] Figure 4 Schematic diagram of the impact resistance simulation test results in Example 1 of the present invention.

[0038] Figure 5 Schematic diagram of the comparison results of state estimation errors of two different MPCs in Example 1 of the present invention.

[0039] Figure 6 Schematic diagram of the simulation test results of OSF-MPC in Example 1 of the present invention.

[0040] Figure 7 Schematic diagram of the simulation test results of RF-MPC in Example 1 of the present invention.

[0041] Figure 8 Schematic diagram of objective function values ​​and ground reaction forces (GRF) under different optimization schemes in Example 1 of the present invention.

[0042] Figure 9 This is a structural diagram of a state feedback model predictive control system for a legged robot in Example 2 of the present invention. DETAILED DESCRIPTION

[0043] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0044] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.

[0045] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0046] Example 1

[0047] This embodiment discloses a state feedback model predictive control method for a leg-foot robot.

[0048] like Figure 1 As shown, a state feedback model predictive control method for a legged robot includes:

[0049] Step S1, obtaining state information of the leg-foot robot; deriving a single rigid body dynamics model of the leg-foot robot based on the obtained state information and discretizing it;

[0050] Step S2: using a closed-loop state observer to dynamically track disturbance information in the single rigid body dynamics model of the legged robot;

[0051] Step S3: Based on the desired motion state and disturbance prediction value set for the leg-foot robot, the model predictive control strategy MPC is used to predict the motion state of the leg-foot robot at future moments and solve the ground reaction force that satisfies the MPC constraint conditions;

[0052] Step S4: Use constraint relaxation WBC to perform secondary optimization on the MPC solution and calculate the joint torque of the leg-foot robot.

[0053] Based on the above process, the present invention can improve the control performance of the legged robot while enhancing its stability and reliability in the face of uncertainty and external disturbances. To facilitate understanding of the technical solution of the present invention, the specific implementation steps of the technical solution of the present invention are further explained and illustrated below.

[0054] Step S1, obtaining state information of the leg-foot robot; deriving a single rigid body dynamics model of the leg-foot robot based on the obtained state information and discretizing it.

[0055] Based on the obtained state information, a single rigid body dynamic model of the legged robot is derived and discretized, including: mathematically representing the model predictive control strategy (MPC) by solving a constrained optimization problem, and differentially defining the generalized coordinate vector in the mathematical representation; mapping the mathematical representation in the model predictive control strategy (MPC) optimization problem to a coordinate system, and determining the single rigid body dynamic model of the legged robot; and making a preliminary prediction of the motion state of the legged robot at the next moment based on the single rigid body dynamic model, specifically:

[0056] Step S1-1: mathematically represent the model predictive control strategy MPC by solving a constrained optimization problem:

[0057]

[0058]

[0059] Among them, l T () is the terminal cost function, l() is the stage cost function, x(t) and u(t) represent the state vector and input vector at time t respectively, T N The goal of the optimization problem is to find an optimal input sequence that minimizes the above cost function while satisfying the system dynamic equation State constraint x(t)∈X f , and the input constraint u(t)∈R u .

[0060] Step S1-2: Differentiate and define the generalized coordinate vector in the mathematical representation, namely:

[0061]

[0062] Among them, p c ∈R 3 represents the center of mass position of the leg-foot robot's body {O} in the world coordinate system, represents the Euler angle of the fuselage, where θ and φ represent the roll, pitch and yaw angles respectively. c ∈R 3 and ω c ∈R 3 They represent the linear velocity and angular velocity of the center of mass of the fuselage in the world system respectively. j ∈R 3 and Represents joint position and joint velocity. All foot end contact forces u i ∈R 3 Collection of For the convenience of representation, let all subscripts of the legs in the support phase be included in the set C, and the identity matrix is ​​represented by I.

[0063] Step S1-3: Map the mathematical representation of the MPC optimization problem to the coordinate system, as follows: Figure 2 shown.

[0064] First, in order to determine the dynamic model of the leg-foot robot's trunk, the whole-body dynamic model of the leg-foot robot is considered, namely:

[0065]

[0066]

[0067] Among them, M B Represents the inertia matrix of the floating base part, M Bjis the matrix encoding the inertial coupling between the legs and the floating base, M Jb It is with M Bj The inertial coupling matrix, M, has a similar effect. J is the inertia matrix of the leg joints, is the nonlinear term representing the Coriolis force and gravity, τ dist Represents the external disturbance force and disturbance torque received by the robot. S represents the drive part selection matrix in the dynamic model, τ represents the joint drive torque; M(q) represents the inertia matrix, represents the joint acceleration, S T Indicates the force generated by the driving part, Represents the transpose of the Jacobian matrix of joint i.

[0068] The main focus of motion planning in legged (multi-jointed) robot systems is the 6-dimensional state information of the floating base, that is, only the first 6 rows of the underactuation are relevant. This concept is called single rigid body dynamics. Therefore, extracting the first 6 rows of the above dynamic equations can determine the single rigid body dynamic model of the legged robot. Based on the single rigid body dynamic model, a preliminary prediction of the legged robot's motion state at the next moment can be made, namely:

[0069]

[0070] in, represents the angular acceleration of the center of mass, represents the acceleration of the center of mass, h B Represents the nonlinear terms of Coriolis force and gravity. All terms with subscript B correspond to the first 6 rows of the relevant terms in the dynamic equation. Therefore, the MPC state equation in the continuous time domain in the mathematical representation of the above model predictive control strategy MPC is It can be expressed as:

[0071]

[0072] in, Represents the rotation matrix, which is used to transform the angular velocity from the joint coordinate system to the center of mass coordinate system; G B (x,u) represents the function of the center of mass acceleration, Represents the total force and torque of the system. Furthermore, before making a preliminary prediction of the motion state of the legged robot at the next moment based on the single rigid body dynamics model, it is necessary to first convert the single rigid body dynamics model from a nonlinear model to a linear model containing time-varying disturbances. Specifically, in order to ensure that the optimization problem is strictly convex, the nonlinear part needs to be linearized at the MPC update point. is the state space vector of the system. B Performing a first-order Taylor expansion on (x,u) yields:

[0073]

[0074] where variable t belongs to a neighborhood of t0, x0and u0are the state and control input at time t0, respectively. A ∈ R 13×13 and B ∈ R 13×13 are the state transition matrix and input matrix obtained after discretization of G B (x, u), respectively, and O(x, u) represents the high-order terms left after the first-order Taylor expansion. Thus, the centroid dynamics of the legged robot can be expressed in the complete state space equation, i.e.,

[0075]

[0076] where D(t) = [d1, d2,..., d 12 , 0] T is a time-varying term used to represent the linearization error and disturbance; y represents the system output. Here, only the case of bounded disturbance is considered, because the disturbance resistance of the robot is limited due to the constraints of the hardware such as actuators. For the legged robot system, the 6-dimensional state information of the floating base is measurable, and thus, C is a unit matrix. It is worth noting that when the state variable definition in this paper is adopted, the matrix A is always in the upper triangular form regardless of the linearization method chosen, i.e.,

[0077]

[0078] It is worth noting that in the prior art, most of the MPC algorithms for legged robots ignore the existence of D(t) and directly establish an open-loop state equation to estimate the dynamic behavior of the system in the prediction horizon, i.e.,

[0079]

[0080] where represents the estimation of the true state x, and is forced to be equal to x(t) at each MPC update time t, ignoring the bias in the state equation, i.e., not considering any deviation of the actual system output from the theoretical system output. Ignoring the existence of D(t) means that the system is assumed to run under ideal conditions without modeling errors and external disturbances. When the external disturbance is small, the influence of D(t) on the system dynamic behavior is small, and the method can also exhibit good control effect. When the external disturbance is large, D(t) has a significant influence on the evolution of the state variable x(t), and at this time, the above method cannot completely capture the dynamic behavior of the actual system, resulting in a large deviation between the model prediction trajectory and the actual system trajectory.

[0081] Step S2, the disturbance information in the single rigid body dynamics model of the leg-foot robot is dynamically tracked by using a closed-loop state observer.

[0082] The motion state of the leg-foot robot is tracked by the closed-loop state observer, and the closed-loop state observer equation is represented as:

[0083]

[0084] Wherein, L∈R 13×13 represents the feedback gain matrix, and ξ represents the reconstructed state of the observer to x; A represents the state transition matrix of the system, B represents the input matrix of the system, and u represents the control input of the system, represents the estimated output, and C represents the output matrix. Define as the state estimation error of the closed-loop state observer.

[0085] Further, the closed-loop state observer is provided with state variables related to tracking the motion state of the leg-foot robot, and the state variables include: the real state x(t) of the leg-foot robot, the reconstructed state ξ of the closed-loop state observer, and the predicted state of the next time

[0086] Wherein, the state estimation error The state transition equation in the continuous time domain of the state estimation error

[0087]

[0088] The above closed-loop state observer is discretized to obtain a discrete time model of the closed-loop state observer, that is:

[0089]

[0090]

[0091] Wherein, T represents the control period of the leg-foot robot, and y(t) represents the measured value of x at time t. The leg-foot robot usually maintains a high control bandwidth, so T is very small. For convenience of calculation, an approximate discretization method is used here, that is:

[0092]

[0093] Further, the characteristic roots of the closed-loop state observer can be calculated by the following formula, that is:

[0094]

[0095] Wherein, det() is the determinant, represents the i-th diagonal element of the matrix Z. To make the state estimation error converge, all eigenvalues of the equation should be located in the unit circle, i.e., |Z i |<1, we have

[0096] Further, in order to reduce the steady-state error in MPC state tracking, the embodiment further proposes an optimization problem construction method based on input variation on the basis of the above scheme, in particular: defining a state tracking error, converting a cost function involved in the model predictive control strategy MPC into a representation in a discrete domain and defining an input variation; writing the input variation as an optimization variable into the definition of the state tracking error to realize quadratic programming based on the input variation, i.e.,

[0097] First, a state tracking error is defined. In the embodiment, the cost function is constructed as a quadratic function, which punishes the deviation from the reference trajectory and the system input. Therefore, the state tracking error can be defined as:

[0098]

[0099] Subsequently, the cost function involved in the model predictive control strategy MPC is converted into a representation in a discrete domain and an input variation is defined. The cost function involved in the model predictive control strategy MPC is converted into a representation in a discrete domain, in particular:

[0100]

[0101] where Q t and Q are weight matrices of terminal state and stage state respectively, and R is a weight matrix of input; J represents the cost function, e t+i|t represents the state error at time t+i predicted at time t, e t+N|t represents the terminal state (i.e., at time t+N) error predicted at time t, and u t+i|t represents the input at time t+i predicted at time t. For a leg-foot robot, the order of magnitude of the input u is much larger than that of the state vector error e, which leads to u T Ru accounts for a large proportion in the cost function J, which will cause a large control steady-state error in the system. Therefore, a scheme of using the input variation as the decision variable of the optimization problem is proposed, and the input variation is defined as:

[0102]

[0103] where u t-1 represents the control input at the last time of the MPC prediction update point, and thus the representation form of the cost function based on the input variation is:

[0104]

[0105] Using this method, the input change amount Δu can be directly penalized to control the fluctuation amplitude of the foot force, while reducing the proportion of the foot force in the objective function, and reducing the state tracking static error. Since the input change amount is used as the optimization variable, the feasible region R u The constraint condition becomes:

[0106]

[0107] where C fri is a block diagonal matrix used to encode the boundary of the friction cone of the supporting leg, whose lower and upper bounds are represented by d and Finally, the input change amount is written as an optimization variable to define the state tracking error to achieve quadratic programming based on the input change amount.

[0108] Further, in order to verify the performance of the closed-loop state observer, the embodiment also carries out mathematical demonstration on the dynamic tracking performance of the proposed closed-loop state observer for time-varying disturbance systems. The proof will be assisted by the following lemma:

[0109] Lemma 1: For any matrix and vector, if holds, where γ and β are both greater than 0, then the following inequality holds:

[0110]

[0111] The state transition equation of the state estimation error in the continuous time domain can be written as:

[0112]

[0113] The dynamics of the leg-foot robot can be divided into two independent parts, namely linear dynamics and rotational dynamics. Among them, the nonlinear rotational dynamics can be equivalent to the form of LTV (linear time-varying) model plus time-varying disturbance. According to the previous analysis, it can be concluded that by selecting the gain matrix L, (A-L) can be configured as a Hurwitz matrix, that is, there exists a positive definite matrix P such that the following formula is true:

[0114] (A-L) T P+P(A-L)=-Φ;

[0115] where Φ is a positive definite matrix. Taking the estimation error as the state variable, the energy function is defined as follows:

[0116]

[0117] Differentiate the above formula, and combine the above state transition equation to obtain:

[0118]

[0119] According to the positive definite property of the quadratic form, for any α>0, the following inequality holds:

[0120]

[0121] Expanding the above equation, we have:

[0122]

[0123] Thus, we have:

[0124]

[0125] where ξ max (-Φ) and ξ max (P) represent the maximum eigenvalues of -Φ and P, respectively. It can be further deduced that ξ max (-Φ)<0. Therefore, there must exist a positive number such that the inequality ξ max (-Φ)+α<0 holds. If holds, then In this case, the observation error is globally asymptotically stable. If the above equation does not hold, let γ=(ξ max (-Φ)+α)、 Applying Lemma 1 to the above equation, we have:

[0126]

[0127] From the above equation, it can be seen that is uniformly bounded (UUB) and will converge to the residual set

[0128]

[0129] Therefore, by adjusting the feedback gain matrix L, the norm of the residual set can be configured to the desired range. In contrast, when using the standard MPC, Since all eigenvalues of the matrix A are located on the imaginary axis, it cannot be ensured that A T P+PA is negative definite or is bounded, in which case may exhibit a divergent state.

[0130] Further, in order to verify the stability of the MPC, the following stability analysis is also made in this embodiment. Specifically, the iterative process of the MPC in the prediction domain can be written as:

[0131]

[0132] where, denotes the predicted state vector at time t+1 at time t, denotes the predicted state feedback gain term at time t+i at time t. At each MPC update point, the value of is predetermined and used to approximate the external disturbance term D(t) in the actual system model. For convenience of representation, define The above equation can be converted to:

[0133]

[0134] where, can be regarded as another generalized force field similar to gravity, which exerts additional forces on the legged robot. Writing the above equation in the augmented form yields:

[0135]

[0136] where, Λ dec denotes the decay matrix. Define The above equation can be converted to the standard state-space form:

[0137] Z t+i+1|t = A i *Z t+i|t + B i Δu t+i|t ;

[0138]

[0139] where, A i denotes the state transition matrix, B i denotes the input matrix. Then, the MPC problem can be reformulated as a QP problem, i.e.,

[0140]

[0141]

[0142] where, Z t+i|t ∈ Z and u t+i|t ∈ U are the state constraints and input constraints of the MPC problem, respectively, Z t+N|t ∈ Z f is the terminal state constraint. The convergence of the MPC control law depends on the setting of its cost function and terminal constraint, while the state transition matrix A i and the input matrix B iThere are no special requirements for the form of . The MPC strategy proposed in this embodiment can be mathematically transformed into a form consistent with the existing general formula, and the convergence proof method proposed therein still applies to the MPC in the above formula. Therefore, the MPC with correction term proposed in this embodiment does not affect the convergence of the original MPC.

[0143] Step S3: Based on the expected motion state and disturbance prediction value set for the leg-foot robot, the model predictive control strategy MPC is used to predict the motion state of the leg-foot robot at future moments and solve the ground reaction force that meets the MPC constraint conditions.

[0144] First, the MPC optimization problem in step S2 can be converted into a standard QP problem for solution. Since the closed-loop state observer has the ability to dynamically track the real model of the legged robot, this embodiment uses the closed-loop state observer equation as the prediction benchmark for MPC. The closed-loop state observer still contains noise caused by the sensor equipment or measurement method. To improve this problem, this embodiment introduces a coefficient matrix F, namely:

[0145]

[0146] For convenience, define When the system performs MPC update at time t, the state recursive formula in the prediction domain can be written as:

[0147]

[0148] in, represents the state vector at time t+1 predicted at time t, Δu t+i|t Represents the change in input at time t+i relative to time t. The initial state of the state recursion in MPC Represents the current true state x(t), that is This method does not change the state transfer matrix A and input matrix B in the original system equations, but corrects the system model by adding feedback correction terms. Through the above derivation, the nonlinear model is linearized near the MPC update point, and a locally effective linear time-varying (LTV) system is obtained. Legged robots usually maintain a high control bandwidth. Therefore, in this embodiment, the external disturbance term D(t) is regarded as a constant in the MPC prediction domain. The eigenvalue has a negative real part, and the state estimation error It shows a decay trend within the MPC range. In this embodiment, an attenuation matrix Λ is used dec To approximate the time domain evolution of x in the prediction domain, that is:

[0149]

[0150]

[0151] in, represents the attenuation matrix Λ dec Λ dec The value of depends on the actual application scenario of the legged robot and the feedback gain matrix L, and needs to be adjusted according to actual considerations. All dynamic iterative equations in the prediction domain are written into a compact state space form:

[0152]

[0153] in, and They represent the augmented vectors of the robot state and input changes in the next N control cycles respectively. is the vector of all feedback items in the prediction domain, A c,t and B v,t are the augmented state transfer matrix and input matrix respectively.

[0154] Therefore, the augmented form of the MPC objective function is:

[0155]

[0156] in, Represents the augmented vector consisting of all expected states in the prediction domain. a Represents the augmented state weight matrix after integrating the stage vector weight and terminal vector weight matrix. R a is the augmented input weight matrix after integrating all stage input weight matrices. Converting the above formula into the standard QP form yields:

[0157]

[0158] stΔu t+i|t ∈R u ;

[0159] st

[0160]

[0161]

[0162] Define the optimal solution to the above optimization problem for:

[0163]

[0164] use The first element in with ut-1 The sum is taken as the optimal plantar force before the next MPC update, that is:

[0165] Step S4: Use constraint relaxation WBC to perform secondary optimization on the MPC solution and calculate the joint torque of the leg-foot robot.

[0166] Due to the large amount of computation, MPC is difficult to maintain a high update frequency on the airborne computing system. Therefore, the constraint relaxation WBC is selected as the compensation controller of MPC to perform secondary optimization on the MPC calculation results to compensate for the adverse effects of the low control frequency of MPC. The secondary optimization result of WBC on the optimal plantar force output by MPC is denoted as u * ,Right now:

[0167] u * =u t +δu;

[0168] Where δu is the plantar force correction value output by WBC. * Substitute this as input into the closed-loop state observer equation for iteration. The WBC incorporates a joint-level PD feedback term into the final control torque, but this is primarily intended to overcome the swing phase trajectory tracking lag caused by model errors and disturbances. During the stance phase, the joint position and velocity do not change significantly, so the impact of the stance phase PD on the plantar force can be ignored.

[0169] Furthermore, in order to evaluate the MPC controller designed in this embodiment, a series of comparative experiments were carried out in the simulation environment Webots. By comparing and analyzing the experimental results of the three MPC controllers deployed on the leg-foot robot simulation platform, it is evaluated which scheme has more advantages in dealing with model uncertainty, specifically including: the standard MPC (Standard MPC) deployed on MIT-Humanoid, the nonlinear MPC (RF-MPC) and the state feedback MPC (OSF-MPC) proposed in this embodiment. In the simulation, the other control submodules used in conjunction with MPC are kept consistent, and the bipedal module (i.e., the leg-foot robot) is set to walk at a step frequency of 2.5Hz, the MPC update frequency is 125Hz, and the prediction step size is 10. For the convenience of calculation, OSF-MPC is configured to use the same approximate linearization method as the standard MPC. The above three MPC controllers used for comparative testing all have the same control parameters. As a specific implementation method, in this embodiment: MPC weight matrix Q = Q t =diag(130,160,150,20,20,600,1,1,20,5,5,20), weight matrix R=10 -6 *I 6*6 The parameters of the joint PD controller are in, represents the proportional gain matrix, Represents the differential gain matrix. The related gain matrix of OSF-MPC proposed in this embodiment is set as follows: The diagonal feedback gain matrix Coefficient matrix F = diag(0.1*I 6*6 ,I 6*6 ,0), attenuation matrix Λ dec =diag(0.01*I 6*6 ,0.3*I 6*6 ,0). The dynamic parameters of the bipedal module are shown in Table 1:

[0170] Table 1 Dynamic parameters of the bipedal module

[0171]

[0172]

[0173] 1. Conduct an impact test. Specifically: One of the features of the MPC framework proposed in this embodiment is the realization of dynamic correction of the model. To verify its effectiveness, an impact test was conducted compared with the standard MPC. In the simulation, the leg-foot robot was set to walk in the +x direction at a constant speed. A step external disturbance force with an amplitude of 15N was applied to a point on the right side of the leg-foot robot's head in the -y direction, causing it to deviate from the reference trajectory. The experimental process is as follows: Figure 3 When facing the same external disturbance force, the robustness of the two MPC controllers is evaluated by quantifying which MPC controller can make the robot's desired trajectory deviate less and recover faster.

[0174] Figure 4 The curves showing the changes in the roll angle and lateral displacement of the bipedal module during the experiment are shown. The shaded area represents the area affected by external disturbances. Compared with the standard MPC, the method proposed in this embodiment can control the robot to recover to the desired state more quickly when subjected to external disturbances, showing significant improvement in system robustness. The state estimation error changes in the two sets of experiments are shown in Figure 2. Figure 5 As shown in FIG, compared with the standard MPC, the method proposed in this embodiment shows a smaller state estimation error in rotational dynamics, and improves the model prediction accuracy under disturbance conditions.

[0175] 2. Conducting a weighted walking test: To further demonstrate the disturbance resistance advantages of the proposed method, a weighted walking test was conducted comparing the proposed MPC controller and RF-MPC. The test setup was as follows: At the beginning of the simulation, a 5 kg weight was placed on the head of the bipedal module. After a period of operation, the simulator increased the weight of the load to 8 kg, representing approximately 88% of the total mass of the bipedal module.

[0176] Figure 6 and Figure 7 The test results of this simulation experiment are presented. Initially, both controllers achieved stable walking and roughly achieved the desired body height and posture. However, the pitch angle of the legged robot using RF-MPC deviated slightly from the desired state, and its average body height was lower than that of the OSF-MPC. After adding a payload, the pitch angle of the robot using RF-MPC diverged, and it eventually collapsed after 3 seconds. In contrast, the OSF-MPC controller demonstrated robustness to external disturbances, allowing the legged robot to maintain stable walking.

[0177] 3. Conduct a comparative test of the optimization schemes. Specifically: To verify the effectiveness of the optimization based on the input variation in this embodiment, a comparative experiment was conducted between the optimization solution based on the input u (Scheme 1) and the optimization solution based on the input variation Δu (Scheme 2). In this test, two different MPC optimization schemes (Scheme 1 and Scheme 2) were deployed on the simulation prototype, and the legged robot was set to walk at a constant speed. The advantages and disadvantages of the two schemes were evaluated by recording and comparing the MPC optimization objective function values ​​of the two schemes and the smoothness of the calculated optimal input u. The experimental results are as follows: Figure 8 As shown, Figure 8 The vertical ground reaction force on the right side of the middle image is the data of the right leg during the movement of the leg-footed robot. Figure 8 At t = 10s, 10.2s, 10.4s, and 10.6s, the transition from the swing phase to the support phase occurs, and large mutations in the plantar force are unavoidable. However, since the optimization based on Δu in Scheme 2 has a smaller state tracking error e(k) during the non-commutation period, the MPC can maintain the legged robot near the desired state with a smaller input during the support / swing transition. Therefore, Scheme 2 can effectively alleviate the sudden spike in contact force when the robot transitions from the swing phase to the support phase. At other times, Scheme 2 has a more obvious advantage than Scheme 1, significantly reducing the tracking error e(k). This confirms that optimization based on the input change Δu can effectively alleviate the problem of the input penalty term in the objective function being too large relative to the state penalty term.

[0178] Example 2

[0179] This embodiment discloses a state feedback model predictive control system for a legged robot.

[0180] As shown in Figure 9 A state feedback model predictive control system for a legged robot, comprising:

[0181] A state information acquisition module configured to: acquire state information of the legged robot; and derive a single-rigid-body dynamics model of the legged robot and discretize the model according to the acquired state information;

[0182] A closed-loop state observation module configured to: utilize a closed-loop state observer to dynamically track disturbance information in the single-rigid-body dynamics model of the legged robot;

[0183] A state feedback MPC module configured to: utilize a model predictive control strategy (MPC) to predict a motion state of the legged robot at a future time and solve a ground reaction force satisfying MPC constraint conditions, according to a desired motion state set for the legged robot and a disturbance prediction value;

[0184] A constraint relaxation WBC module configured to: utilize a constraint relaxation WBC to secondarily optimize a solution of the MPC and calculate a joint torque of the legged robot.

[0185] A state estimation module configured to feed back pose information of the legged robot to the state feedback MPC module and the constraint relaxation WBC module.

[0186] Further, the closed-loop state observation module is integrated in the state feedback MPC module. The state feedback model predictive control system provided by the present application simultaneously integrates MPC and WBC, aiming to provide a robust control strategy for the legged robot, as shown in Figure 9 wherein represents a desired velocity of a body of the legged robot, u mpc represents an optimal foot force calculated by the MPC, τ cmd and τ * are a joint torque output by the WBC and a final joint torque acting on the robot, respectively. Figure 9 The blue numerical labels in the figure represent a running frequency of each control submodule, i.e., the state feedback MPC module runs at a frequency of 125 Hz, the constraint relaxation WBC module runs at a frequency of 500 Hz, and the state estimation module runs at a frequency of 500 Hz.

[0187] Embodiment Three

[0188] An object of the present embodiment is to provide a computer-readable storage medium.

[0189] A computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of the state feedback model predictive control method for a legged robot according to embodiment one of the present disclosure.

[0190] Embodiment four

[0191] An object of the present embodiment is to provide an electronic device.

[0192] An electronic device comprising a memory, a processor, and a program stored on the memory and executable on the processor, wherein the processor implements the steps of the state feedback model predictive control method for a legged robot according to embodiment one of the present disclosure when executing the program.

[0193] The steps involved in the apparatuses of embodiments two, three and four above correspond to the method of embodiment one, and the detailed description can be found in the relevant description section of embodiment one. The term "computer readable storage medium" should be understood to include a single medium or multiple media that store one or more sets of instructions; it should also be understood that any medium that is capable of storing, encoding or carrying a set of instructions for execution by a processor and that causes the processor to perform any of the processes of the present disclosure is within the scope of the term "computer readable storage medium".

[0194] Those skilled in the art should understand that the modules or steps of the present disclosure described above can be implemented by a general computer device, and alternatively, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by a computing device, or they can be made into individual integrated circuit modules, or a plurality of modules or steps among them can be made into a single integrated circuit module. The present disclosure is not limited to any specific combination of hardware and software.

[0195] The specific embodiments of the present disclosure described above in conjunction with the accompanying drawings are not intended to limit the scope of protection of the present disclosure, and those skilled in the art should understand that various modifications or variations made by those skilled in the art on the basis of the technical solutions of the present disclosure without inventive labor are still within the scope of protection of the present disclosure.

Claims

1. A state feedback model predictive control method for a legged robot, characterized in that: include: Obtaining the state information of the leg-foot robot; deriving the single rigid body dynamics model of the leg-foot robot based on the obtained state information and discretizing it; The closed-loop state observer is used to dynamically track the disturbance information in the single rigid body dynamics model of the legged robot. According to the expected motion state and disturbance prediction value set by the leg-foot robot, the model predictive control strategy MPC is used to predict the motion state of the leg-foot robot at the future moment and solve the ground reaction force that meets the MPC constraints. Constraint relaxation WBC is used to perform secondary optimization on the MPC solution and calculate the joint torque of the leg-foot robot. The closed-loop state observer is provided with state variables related to tracking the motion state of the leg-foot robot, and the state variables include: the real state of the leg-foot robot, the reconstructed state of the closed-loop state observer, and the predicted state of the model predictive control strategy MPC at the future moment; All dynamic iterative equations in the prediction domain are written into a compact state space form: ; in, and Represent the augmented vectors of the robot state and input changes in the next N control cycles, is the vector of all feedback items in the prediction domain, and are the augmented state transfer matrix and input matrix respectively, Represents the control input at the moment before the model predictive control strategy MPC prediction update point.

2. A state feedback model predictive control method for a legged robot according to claim 1, characterized in that: Based on the obtained state information, the single rigid body dynamic model of the legged robot is derived and discretized, including: The model predictive control strategy MPC is mathematically represented by solving a constrained optimization problem, and the generalized coordinate vector in the mathematical representation is differentially defined; the mathematical representation in the model predictive control strategy MPC optimization problem is mapped to a coordinate system, and the single rigid body dynamics model of the leg-legged robot is determined; based on the single rigid body dynamics model, the motion state of the leg-legged robot at future moments is predicted.

3. A state feedback model predictive control method for a legged robot according to claim 2, characterized in that: Before predicting the motion state of the legged robot at a future moment based on the single rigid body dynamics model, the single rigid body dynamics model is first converted from a nonlinear model to a linear model containing time-varying disturbances.

4. A state feedback model predictive control method for a legged robot according to claim 1, characterized in that: It also includes quadratic programming based on input variation, specifically: defining the state tracking error, converting the cost function involved in the model predictive control strategy MPC into a discrete domain representation and defining the input variation; writing the input variation as an optimization variable into the defined state tracking error to achieve quadratic programming based on the input variation.

5. The state feedback model predictive control method for a legged robot according to claim 1, characterized in that: The model predictive control strategy MPC takes the closed-loop state observer equation as the prediction benchmark and uses the attenuation matrix to approximate the time domain evolution in the prediction domain, thereby making a preliminary prediction of the motion state of the legged robot at the next moment.

6. A state feedback model predictive control system for a legged robot, applying the state feedback model predictive control method for a legged robot according to any one of claims 1 to 5, characterized in that: include: The state information acquisition module is configured to: obtain state information of the leg-foot robot; derive a single rigid body dynamics model of the leg-foot robot based on the obtained state information and discretize the model; The closed-loop state observation module is configured to: dynamically track disturbance information in a single rigid body dynamics model of the legged robot using a closed-loop state observer; The state feedback MPC module is configured to: predict the future motion state of the leg-foot robot using the model predictive control strategy MPC based on the desired motion state and disturbance prediction value set by the leg-foot robot and solve the ground reaction force that meets the MPC constraints; The constraint relaxation WBC module is configured to perform secondary optimization on the MPC solution using the constraint relaxation WBC and calculate the joint torque of the leg-foot robot.

7. A state feedback model predictive control system for a legged robot according to claim 6, characterized in that: It also includes: a state estimation module, which is used to feed back the posture information of the leg-foot robot to the state feedback MPC module and the constraint relaxation WBC module; at the same time, the closed-loop state observation module is integrated into the state feedback MPC module.

8. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps of the state feedback model predictive control method for a legged robot as described in any one of claims 1 to 5 are implemented.

9. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the state feedback model predictive control method for a legged robot as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • A method and system for motion control of bipedal robots based on deep reinforcement learning

    CN115128960B

  • Biped robot motion control method and system based on deep reinforcement learning

    CN115128960A

  • Humanoid robot high-dynamic jumping motion control method based on online centroid trajectory optimization

    CN118682750A