State feedback model prediction control method and system for legged robot

Through the state feedback model prediction control method, the closed-loop state observer and model prediction control strategy are used to solve the stability and reliability problems of leg foot robots in the face of uncertainty and external disturbances, achieving higher control performance and dynamic capture accuracy.

CN119937611AActive Publication Date: 2025-05-06SHANDONG UNIV

Patent Information

Application Number
CN202510101628.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with model uncertainty and external perturbations in leg foot robot motion control, resulting in degradation of control performance and impaired system stability.

Method used

The state feedback model prediction control method is adopted, by obtaining the state information of the leg foot robot, the single rigid body dynamic model is derived and discretized, and the disturbance information is dynamically tracked by the closed-loop state observer, the future motion state is predicted based on the model prediction control strategy, and the second optimization is performed through the constraint relaxation WBC to calculate the joint torque.

Benefits of technology

The control performance of the leg foot robot is improved, the stability and reliability in the face of uncertainty and external disturbances are enhanced, and the precise capture of dynamics and state tracking is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937611A_ABST
    Figure CN119937611A_ABST
Patent Text Reader

Abstract

The invention provides a state feedback model prediction control method and system for a legged robot, and belongs to the technical field of robot control. Firstly, a closed-loop state observer is used for dynamically tracking disturbance information in a single rigid body dynamic model of the legged robot; then, according to the expected motion state set by the legged robot and a disturbance prediction value, a model prediction control strategy MPC is adopted to predict the motion state of the legged robot at the future moment, and ground reaction force meeting MPC constraint conditions is solved; and finally, constraint relaxation WBC is adopted to carry out secondary optimization on the MPC solving result, and the joint torque of the leg-foot robot is calculated. According to the invention, on the basis of improving the control performance of the leg-foot robot, the stability and reliability of the leg-foot robot facing uncertainty and external disturbance can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robot control, and in particular relates to a state feedback model predictive control method and system for a leg-foot robot. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Virtual Model Controller (VMC) is widely used in legged robot motion control. This method has attracted attention for its small computational complexity and fast response speed. VMC generates appropriate virtual forces by constructing appropriate virtual components on each degree of freedom that needs to be controlled. These virtual forces are converted by the action of the actuators. For some control problems, it is necessary to map the force or torque in the workspace into the joint torque in the joint space. However, the VMC method has obvious limitations in long-term motion planning. It usually lacks the ability to fully predict and plan the future motion state of the robot.

[0004] In contrast, model predictive control (MPC) technology is favored in the field of legged robot control because of its long-term planning capabilities and ability to handle multiple constraints. MPC builds a dynamic model of the system and uses this model to predict the future behavior of the system at each control moment. Based on these predictions, it can generate an optimized control sequence. MPC includes basic steps such as system modeling, prediction, and optimization. It calculates the optimal control input sequence by solving an optimization problem to maximize or minimize a performance indicator.

[0005] However, despite its significant theoretical advantages, MPC faces challenges in handling model uncertainties, such as modeling errors and external disturbances, when deployed in practice. These uncertainties may lead to degraded control performance and even compromised system stability. How to effectively handle model uncertainties and external disturbances to achieve stable walking has become a technical challenge in the field of legged robot control.

[0006] For example, the patent application number is 202211050051.6, which is a bipedal robot motion control method based on deep reinforcement learning. This method relies on a large amount of training data to learn control strategies. Although this method may perform well in a simulation environment, its generalization ability may be limited when faced with complex and changeable situations in the real world; in particular, when the external disturbance exceeds the range of the training data, the algorithm may not be able to respond effectively. In addition, external disturbances have a significant impact on the optimal control problem of nonlinear systems. Disturbance may cause changes in the state equations and control equations of the system, thereby changing the optimal control strategy; at the same time, disturbances may also introduce uncertainty factors, making the solution of the optimal control problem more difficult. Therefore, the stability and reliability of the legged robot under this method in the face of uncertainty and external disturbances are still not ideal. Summary of the invention

[0007] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a state feedback model predictive control method and system for a leg-leg robot, which can enhance the stability and reliability of the leg-leg robot when facing uncertainty and external disturbances on the basis of improving the control performance of the leg-leg robot.

[0008] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0009] A first aspect of the present invention provides a state feedback model predictive control method for a legged robot.

[0010] A state feedback model predictive control method for a legged robot, comprising:

[0011] Acquire the state information of the leg-foot robot; derive the single rigid body dynamics model of the leg-foot robot and discretize it according to the acquired state information;

[0012] The closed-loop state observer is used to dynamically track the disturbance information in the single rigid body dynamics model of the legged robot.

[0013] According to the expected motion state and disturbance prediction value set by the leg-foot robot, the model predictive control strategy MPC is used to predict the motion state of the leg-foot robot at future moments and solve the ground reaction force that meets the MPC constraint conditions.

[0014] Constraint relaxation WBC is used to perform secondary optimization on the MPC solution and calculate the joint torque of the leg-foot robot.

[0015] Furthermore, based on the obtained state information, the single rigid body dynamics model of the legged robot is derived and discretized, including: mathematically representing the model predictive control strategy MPC by solving a constrained optimization problem, and differentially defining the generalized coordinate vector in the mathematical representation; mapping the mathematical representation in the model predictive control strategy MPC optimization problem to the coordinate system, and determining the single rigid body dynamics model of the legged robot; and predicting the motion state of the legged robot at future moments based on the single rigid body dynamics model.

[0016] Furthermore, before predicting the motion state of the legged robot at a future moment based on the single rigid body dynamics model, the single rigid body dynamics model is first converted from a nonlinear model to a linear model containing time-varying disturbances.

[0017] Furthermore, the closed-loop state observer is provided with state variables related to tracking the motion state of the leg-foot robot, and the state variables include: the real state of the leg-foot robot, the reconstructed state of the closed-loop state observer, and the predicted state of the model predictive control strategy MPC for the future moment.

[0018] Furthermore, it also includes quadratic programming based on input changes, specifically: defining a state tracking error, converting the cost function involved in the model predictive control strategy MPC into a discrete domain representation and defining an input change; writing the input change as an optimization variable into the defined state tracking error to achieve quadratic programming based on the input change.

[0019] Furthermore, the model predictive control strategy MPC uses the closed-loop state observer equation as a prediction benchmark, and uses the attenuation matrix to approximate the time domain evolution in the prediction domain, thereby making a preliminary prediction of the motion state of the legged robot at the next moment.

[0020] A second aspect of the present invention provides a state feedback model predictive control system for a legged robot.

[0021] A state feedback model predictive control system for a legged robot, comprising:

[0022] The state information acquisition module is configured to: obtain the state information of the leg-foot robot; derive the single rigid body dynamics model of the leg-foot robot and discretize it according to the obtained state information;

[0023] The closed-loop state observation module is configured to: dynamically track disturbance information in a single rigid body dynamics model of the legged robot using a closed-loop state observer;

[0024] The state feedback MPC module is configured to: predict the motion state of the leg-foot robot at the future moment and solve the ground reaction force that meets the MPC constraint conditions according to the expected motion state and disturbance prediction value set by the leg-foot robot using the model predictive control strategy MPC;

[0025] The constraint relaxation WBC module is configured to: perform secondary optimization on the MPC solution using the constraint relaxation WBC and calculate the joint torque of the leg-foot robot;

[0026] The state estimation module is used to feed back the posture information of the leg-foot robot to the state feedback MPC module and the constraint relaxation WBC module.

[0027] Furthermore, the closed-loop state observation module is integrated into the state feedback MPC module.

[0028] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a state feedback model predictive control method for a legged robot as described in the first aspect of the present invention.

[0029] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in a state feedback model predictive control method for a legged robot as described in the first aspect of the present invention are implemented.

[0030] One or more of the above technical solutions have the following beneficial effects:

[0031] The present invention first uses a closed-loop state observer to dynamically track the disturbance information in the single rigid body dynamics model of the leg-foot robot; then uses the model predictive control strategy MPC to predict the motion state of the leg-foot robot at future moments and solve the ground reaction force that satisfies the MPC constraint conditions; finally, the constraint relaxation WBC is used to perform secondary optimization on the MPC solution results, and the joint torque of the leg-foot robot is calculated. By converting the single rigid body dynamics model (nonlinear floating basis dynamics model) containing uncertainty into an augmented linear model containing time-varying disturbance terms, and using the closed-loop state observer equation as the benchmark for MPC state prediction, accurate capture of the dynamics of the leg-foot robot can be achieved; at the same time, by constructing an optimization problem based on the input variation, the static error in the MPC state tracking is reduced. Therefore, the present invention can enhance the stability and reliability of the leg-foot robot in the face of uncertainty and external disturbances on the basis of improving the control performance of the leg-foot robot.

[0032] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0034] Figure 1 This is a flow chart of a state feedback model predictive control method for a legged robot in Embodiment 1 of the present invention.

[0035] Figure 2 It is a schematic diagram of mapping the mathematical representation in the MPC optimization problem to a coordinate system in the first embodiment of the present invention.

[0036] Figure 3 Schematic diagram of the anti-impact simulation test process in Example 1 of the present invention.

[0037] Figure 4 Schematic diagram of the impact resistance simulation test results in Example 1 of the present invention.

[0038] Figure 5 Schematic diagram of comparison results of state estimation errors of two different MPCs in Example 1 of the present invention.

[0039] Figure 6 Schematic diagram of the simulation test results of OSF-MPC in Example 1 of the present invention.

[0040] Figure 7 Schematic diagram of the simulation test results of RF-MPC in Example 1 of the present invention.

[0041] Figure 8 Schematic diagram of objective function values ​​and ground reaction forces (GRF) under different optimization schemes in Example 1 of the present invention.

[0042] Fig. 9 This is a structural diagram of a state feedback model predictive control system for a legged robot in Embodiment 2 of the present invention. DETAILED DESCRIPTION

[0043] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0044] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0045] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0046] Embodiment 1

[0047] This embodiment discloses a state feedback model predictive control method for a legged robot.

[0048] like Figure 1 As shown, a state feedback model predictive control method for a legged robot comprises:

[0049] Step S1, obtaining state information of the leg-foot robot; deriving a single rigid body dynamics model of the leg-foot robot based on the obtained state information and discretizing it;

[0050] Step S2, using a closed-loop state observer to dynamically track disturbance information in a single rigid body dynamics model of the leg-foot robot;

[0051] Step S3, according to the expected motion state and disturbance prediction value set by the leg-foot robot, the model predictive control strategy MPC is used to predict the motion state of the leg-foot robot at the future moment and solve the ground reaction force that meets the MPC constraint condition;

[0052] Step S4: Use constraint relaxation WBC to perform secondary optimization on the MPC solution and calculate the joint torque of the leg-foot robot.

[0053] Based on the above process, the present invention can enhance the stability and reliability of the legged-legged robot in the face of uncertainty and external disturbances on the basis of improving the control performance of the legged-legged robot. In order to facilitate the understanding of the technical solution of the present invention, the specific implementation steps in the technical solution of the present invention are further explained and illustrated below.

[0054] Step S1, obtaining the state information of the leg-foot robot; according to the obtained state information, deriving the single rigid body dynamics model of the leg-foot robot and discretizing it.

[0055] According to the obtained state information, the single rigid body dynamics model of the legged robot is derived and discretized, including: mathematically representing the model predictive control strategy MPC by solving a constrained optimization problem, and differentially defining the generalized coordinate vector in the mathematical representation; mapping the mathematical representation in the model predictive control strategy MPC optimization problem to a coordinate system, and determining the single rigid body dynamics model of the legged robot; and preliminarily predicting the motion state of the legged robot at the next moment based on the single rigid body dynamics model, specifically:

[0056] Step S1-1, mathematically represent the model predictive control strategy MPC by solving a constrained optimization problem:

[0057]

[0058]

[0059] Among them, l T () is the terminal cost function, l() is the stage cost function, x(t) and u(t) represent the state vector and input vector at time t respectively, T N Represents the length of the prediction period. The goal of the optimization problem is to find an optimal input sequence that minimizes the above cost function and satisfies the system dynamic equation State constraint x(t)∈X f , and the input constraint u(t)∈R u .

[0060] Step S1-2, differentially define the generalized coordinate vector in the mathematical representation, namely:

[0061]

[0062] Among them, p c ∈R 3 represents the center of mass position of the legged robot's body {O} in the world coordinate system, represents the Euler angle of the fuselage, where θ and φ represent the roll, pitch and yaw angles respectively. c ∈R 3 and ω c ∈R 3 They represent the linear velocity and angular velocity of the center of mass of the fuselage in the world system. j ∈R 3 and represents the joint position and joint velocity. All foot contact forces u i ∈R 3 The collection of For the convenience of representation, let all the subscripts of the legs in the support phase be included in the set C, and the identity matrix is ​​represented by I.

[0063] Step S1-3, mapping the mathematical representation in the MPC optimization problem of the model predictive control strategy to the coordinate system, as follows Figure 2 shown.

[0064] First, in order to determine the dynamic model of the leg-foot robot trunk, the whole body dynamic model of the leg-foot robot is considered first, namely:

[0065]

[0066]

[0067] Among them, M B Represents the inertia matrix of the floating base part, M Bjis the matrix encoding the inertial coupling between the legs and the floating base, M Jb It is with M Bj The inertial coupling matrix, M, has a similar effect. J is the inertia matrix of the leg joints, is the nonlinear term representing the Coriolis force and gravity, τ dist represents the external disturbance force and disturbance torque on the robot. S represents the drive part selection matrix in the dynamic model, τ represents the joint drive torque; M(q) represents the inertia matrix, represents the joint acceleration, S T Indicates the force generated by the drive part, Represents the transpose of the Jacobian matrix of joint i.

[0068] The main focus of motion planning in legged (multi-jointed legged) robot systems is the 6-dimensional state information of the floating base, that is, only the first 6 rows of underactuation are relevant. This concept is called single rigid body dynamics. Therefore, extracting the first 6 rows in the above dynamic equations can determine the single rigid body dynamics model of the legged robot, and based on the single rigid body dynamics model, a preliminary prediction of the motion state of the legged robot at the next moment is made, namely:

[0069]

[0070] in, represents the angular acceleration of the center of mass, represents the acceleration of the center of mass, h B represents the nonlinear terms of Coriolis force and gravity. All terms with subscript B correspond to the first 6 rows of relevant terms in the dynamic equation. Therefore, the MPC state equation in the continuous time domain in the mathematical representation of the above model predictive control strategy MPC is It can be expressed as:

[0071]

[0072] in, represents the rotation matrix, which is used to transform the angular velocity from the joint coordinate system to the center of mass coordinate system; G B (x,u) is the function of the center of mass acceleration, Represents the total force and torque of the system. Furthermore, before making a preliminary prediction of the motion state of the legged robot at the next moment based on the single rigid body dynamics model, it is necessary to first convert the single rigid body dynamics model from a nonlinear model to a linear model containing time-varying disturbances. Specifically, in order to ensure that the optimization problem is strictly convex, the nonlinear part needs to be linearized at the MPC update point. is the state space vector of the system. 0 The neighborhood of G B The first-order Taylor expansion of (x,u) yields:

[0073]

[0074] Among them, the variable t belongs to t 0 Neighborhood of x 0 and u 0 t 0 The state and control input at the moment. A∈R 13×13 and B∈R 13×13 For G B The state transfer matrix and input matrix obtained after discretization of (x,u), O(x,u) represents the remaining high-order terms after the first-order Taylor expansion. Therefore, the center of mass dynamics of the legged robot can be expressed by the complete state space equation, namely:

[0075]

[0076] Where D(t) = [d 1 ,d 2 ,...,d 12 ,0] T is a time-varying term used to represent linearization error and disturbance; y represents the system output. Here, only the case where the external disturbance is bounded is considered, because the robot's disturbance resistance is limited by the constraints of hardware such as actuators. For the leg-foot robot system, the 6-dimensional state information of the floating base is measurable, so C is a unit matrix. It is worth noting that when the state variable definition method in this article is adopted, no matter which linearization method is selected, the matrix A is always in upper triangular form, that is:

[0077]

[0078] It should be noted that in the prior art, most legged robot MPC algorithms ignore the existence of D(t) and directly establish an open-loop state equation to estimate the dynamic behavior of the system in the prediction time domain, namely:

[0079]

[0080] in, represents the estimate of the true state x, and At each MPC update time t is forced to be equal to x(t), and the deviation in the state equation is ignored, that is, any deviation between the actual system output and the theoretical system output is not considered. Ignoring the existence of D(t) means that the system is assumed to operate under ideal conditions without modeling errors and external disturbances. When the external disturbance suffered by the system is very small, the impact of D(t) on the dynamic behavior of the system is small, and this method can also show good control effect. When the external disturbance is large, D(t) has a significant impact on the evolution of the state variable x(t). At this time, the above method cannot fully capture the dynamic behavior of the actual system, resulting in a large deviation between the model prediction trajectory and the actual system change trajectory.

[0081] Step S2: Use a closed-loop state observer to dynamically track disturbance information in the single rigid body dynamics model of the legged robot.

[0082] The motion state of the legged robot is tracked by a closed-loop state observer; the closed-loop state observer equation is expressed as:

[0083]

[0084] Among them, L∈R 13×13 represents the feedback gain matrix, ξ represents the observer's reconstructed state of x; A represents the system's state transfer matrix, B represents the system's input matrix, u represents the system's control input, Denotes the estimated output, and C denotes the output matrix. Definition is the state estimation error of the closed-loop state observer.

[0085] Furthermore, the closed-loop state observer is provided with state variables related to tracking the motion state of the leg-foot robot, which include: the real state x(t) of the leg-foot robot, the reconstructed state ξ of the closed-loop state observer, and the predicted state of the model predictive control strategy MPC at the next moment

[0086] Among them, the state estimation error The state transfer equation in the continuous time domain can be written as:

[0087]

[0088] By discretizing the above closed-loop state observer, we can obtain the discrete time model of the closed-loop state observer, namely:

[0089]

[0090]

[0091] Where T represents the control period of the legged robot, and y(t) represents the measured value of x sampled at time t. Legged robots usually maintain a high control bandwidth, so T is very small. For the convenience of calculation, an approximate discretization method is used here, namely:

[0092]

[0093] Furthermore, the characteristic root of the closed-loop state observer can be calculated by the following formula:

[0094]

[0095] Among them, det() is the determinant, express To make the state estimation error converge, let all eigenvalues ​​of the equation be within the unit circle, that is, |Z i |<1, we can get

[0096] Furthermore, in order to reduce the static error in MPC state tracking, this embodiment, based on the above scheme, also proposes a method for constructing an optimization problem based on input variation. Specifically, the state tracking error is defined, the cost function involved in the model predictive control strategy MPC is converted into a discrete domain representation and the input variation is defined; the input variation is written as an optimization variable into the defined state tracking error to realize the quadratic programming based on the input variation, that is:

[0097] First, define the state tracking error. In this embodiment, the cost function is constructed as a quadratic function that penalizes deviations from the reference trajectory and system input. Therefore, the state tracking error can be defined as:

[0098]

[0099] Subsequently, the cost function involved in the model predictive control strategy MPC is converted into a discrete domain representation and the input variation is defined. The cost function involved in the model predictive control strategy MPC is converted into a discrete domain representation, specifically:

[0100]

[0101] Among them, Q t and Q are the weight matrices of the terminal state and stage state respectively, R is the input weight matrix; J represents the cost function, e t+i|t represents the state error at time t+i predicted at time t, e t+N|t represents the error of the terminal state predicted at time t (i.e., time t+N), u t+i|trepresents the input at time t+i predicted at time t. For legged robots, the order of magnitude of the input u is much larger than the state vector error e, resulting in u T Ru accounts for a large proportion in the cost function J, which will cause a large control static error in the system. Therefore, a solution is proposed to use the input change as the decision variable of the optimization problem, and the input change is defined as:

[0102]

[0103] Among them, u t-1 represents the control input at the last moment of the MPC prediction update point. Therefore, the cost function based on the input change is expressed as:

[0104]

[0105] Using this method, the input change Δu can be directly penalized to control the fluctuation amplitude of the plantar force, while reducing the weight of the plantar force in the objective function and reducing the state tracking static error. Since the input change is used as the optimization variable, the feasible domain R of the input u is u The constraints become:

[0106]

[0107] Among them, C fri is a block diagonal matrix encoding the boundaries of the friction cone of the supporting leg, whose lower and upper limits are given by d and Finally, the input variation is written as the optimization variable to define the state tracking error to achieve the quadratic programming based on the input variation.

[0108] Furthermore, in order to verify the performance of the closed-loop state observer, this embodiment also provides a mathematical demonstration of the dynamic tracking performance of the proposed closed-loop state observer for the time-varying disturbance system. The proof is based on the following lemma:

[0109] Lemma 1: For any matrix and vector, if holds true, where both γ and β are greater than 0, then the following inequality holds true:

[0110]

[0111] The state transfer equation in the continuous time domain of state estimation error can be written as:

[0112]

[0113] The dynamics of the legged robot can be decomposed into two independent parts, namely: linear dynamics and rotational dynamics. Among them, the nonlinear rotational dynamics can be equivalently converted into the form of LTV (linear time-varying) model plus time-varying perturbations. According to the previous analysis, it can be concluded that by selecting the gain matrix L, (AL) can be configured as a Hurwitz matrix, that is, there is a positive definite matrix P such that the following formula holds:

[0114] (AL) T P+P(AL)=-Φ;

[0115] Where Φ is a positive definite matrix. is the state variable, and the energy function is defined as follows:

[0116]

[0117] Differentiate the above formula and combine it with the above state transfer equation to obtain:

[0118]

[0119] According to the positive definite property of quadratic forms, for any α>0, the following inequality holds:

[0120]

[0121] Expand the above formula to get:

[0122]

[0123] Therefore, we can get:

[0124]

[0125] Among them, ξ max (-Φ) and ξ max (P) represent the maximum eigenvalues ​​of -Φ and P respectively. It can be inferred that ξ max (-Φ) < 0. Therefore, there must be a positive number such that the inequality ξ max (-Φ)+α<0 holds. If established, At this time, the observation error Global asymptotic stability. If the above equation is not satisfied, let γ=(ξ max (-Φ)+α)、 Applying Lemma 1 to the above equation yields:

[0126]

[0127] From the above formula, we can see that is uniformly bounded (UUB) and will converge to the residual set at an exponential rate

[0128]

[0129] Therefore, by adjusting the feedback gain matrix L, the residual set The norm of is configured to the desired range. In contrast, when using standard MPC, Since all eigenvalues ​​of matrix A lie on the imaginary axis, there is no way to ensure that A T P+PA is negative or is bounded, at this time It may appear diffuse.

[0130] Furthermore, in order to verify the stability of MPC, this embodiment also performs the following stability analysis. Specifically, the iterative process of MPC in the prediction domain can be written as:

[0131]

[0132] in, represents the state vector at time t+1 predicted at time t, represents the state feedback gain term at time t+i predicted at time t. At each MPC update point, The value of is predetermined and is used to approximate the external disturbance term D(t) in the actual system model. For convenience of representation, define The above equation can be converted to:

[0133]

[0134] in, It can be regarded as another generalized force field similar to gravity, exerting additional force on the legged robot. Writing the above formula in an augmented form yields:

[0135]

[0136] Among them, Λ dec represents the attenuation matrix. Definition The above equation can be converted to the standard state space form:

[0137] Z t+i+1|t =A i *Z t+i|t +B i Δu t+i|t ;

[0138]

[0139] Among them, Ai represents the state transfer matrix, B i represents the input matrix. Then, the MPC problem can be reformulated as a QP problem, namely:

[0140]

[0141]

[0142] Among them, Z t+i|t ∈Z and u t+i|t ∈U are the state constraints and input constraints of the MPC problem, respectively, and Z t+N|t ∈Z f is the terminal state constraint. The convergence of the MPC control law depends on the setting of its cost function and terminal constraints. i and the input matrix B i There is no special requirement for the form of . The MPC strategy proposed in this embodiment can be mathematically transformed into a form consistent with the existing general formula, and the convergence proof method proposed in it is still applicable to the MPC in the above formula. Therefore, the MPC with correction term proposed in this embodiment will not affect the convergence of the original MPC.

[0143] Step S3: According to the expected motion state and disturbance prediction value set by the leg-foot robot, the model predictive control strategy MPC is used to predict the motion state of the leg-foot robot at future moments and solve the ground reaction force that meets the MPC constraint conditions.

[0144] First, the MPC optimization problem in step S2 can be converted into a standard QP problem for solution. Since the closed-loop state observer has the ability to dynamically track the real model of the legged robot, this embodiment uses the closed-loop state observer equation as the prediction benchmark of the MPC. The closed-loop state observer still contains noise caused by the sensor device or measurement method. To improve this problem, this embodiment introduces a coefficient matrix F, namely:

[0145]

[0146] For convenience, define When the system performs an MPC update at time t, the state recursive formula in the prediction domain can be written as:

[0147]

[0148] in, represents the state vector at time t+1 predicted at time t, Δu t+i|t Represents the change in input at time t+i relative to time t. The initial state of the state recursion in MPC Represents the current true state x(t), that is This method does not change the state transfer matrix A and input matrix B in the original system equation, but corrects the system model by adding feedback correction terms. Through the above derivation, the nonlinear model is linearized near the MPC update point, and a locally effective linear time-varying (LTV) system is obtained. Legged robots usually maintain a high control bandwidth. Therefore, in this embodiment, the external disturbance term D(t) is regarded as a constant in the MPC prediction domain. The eigenvalue of has a negative real part, and the state estimation error It shows a decay trend within the MPC range. In this embodiment, an attenuation matrix Λ is used dec To approximate the time domain evolution of x in the prediction domain, that is:

[0149]

[0150]

[0151] in, Denotes the attenuation matrix Λ dec The ith power of Λ dec The value of depends on the actual application scenario of the legged robot and the feedback gain matrix L, and needs to be adjusted according to practical considerations. All dynamics iteration equations in the prediction domain are written in a compact state space form:

[0152]

[0153] in, and They represent the augmented vectors of the robot state and input changes in the next N control cycles respectively. is the vector of all feedback items in the prediction domain, A c,t and B v,t are the augmented state transfer matrix and input matrix respectively.

[0154] Therefore, the augmented form of the MPC objective function is:

[0155]

[0156] in, Represents the augmented vector consisting of all expected states in the prediction domain. a represents the augmented state weight matrix after integrating the stage vector weight and the terminal vector weight matrix. R a is the augmented input weight matrix after integrating all the input weight matrices of the stages. Converting the above formula into the standard QP form yields:

[0157]

[0158] stΔu t+i|t ∈R u ;

[0159] st

[0160]

[0161]

[0162] Define the optimal solution to the above optimization problem for:

[0163]

[0164] use The first element in with u t-1 The sum is taken as the optimal plantar force before the next MPC update, that is:

[0165] Step S4: Use constraint relaxation WBC to perform secondary optimization on the MPC solution and calculate the joint torque of the leg-foot robot.

[0166] Due to the huge amount of calculation, MPC is difficult to maintain a high update frequency on the airborne computing system. Therefore, the constraint relaxation WBC is selected as the compensation controller of MPC to perform secondary optimization on the MPC calculation results to compensate for the adverse effects of the low control frequency of MPC. The secondary optimization result of WBC on the optimal plantar force output by MPC is denoted as u * ,Right now:

[0167] u * =u t +δu;

[0168] Where δu is the plantar force correction value output by WBC. * Substitute it as an input into the closed-loop state observer equation for iteration. WBC adds the PD feedback term at the joint level to the final control torque, but this is mainly aimed at overcoming the lag problem of swing phase trajectory tracking caused by model errors and disturbances. For the support phase, the joint position speed does not change too much, so the influence of the support phase PD on the plantar force can be ignored.

[0169] Furthermore, in order to evaluate the MPC controller designed in this embodiment, a series of comparative experiments were carried out in the simulation environment Webots. By comparing and analyzing the experimental results of the three MPC controllers deployed on the leg-foot robot simulation platform, it is evaluated which scheme has more advantages in dealing with model uncertainty, including: the standard MPC (Standard MPC) deployed on MIT-Humanoid, the nonlinear MPC (RF-MPC) and the state feedback MPC (OSF-MPC) proposed in this embodiment. In the simulation, the other control submodules used in conjunction with the MPC are consistent. The bipedal module (i.e., the leg-foot robot) is set to walk at a step frequency of 2.5Hz, the MPC update frequency is 125Hz, and the prediction step length is 10. For the convenience of calculation, OSF-MPC is configured to use the same approximate linearization method as the standard MPC. The above three MPC controllers used for comparative testing all have the same control parameters. As a specific implementation method, in this embodiment: MPC weight matrix Q = Q t =diag(130,160,150,20,20,600,1,1,20,5,5,20), weight matrix R=10 -6 *I 6*6 The parameters of the joint PD controller are in, represents the proportional gain matrix, Denotes the differential gain matrix. The relevant gain matrix of the OSF-MPC proposed in this embodiment is set as follows: The diagonal feedback gain matrix Coefficient matrix F = diag(0.1*I 6*6 ,I 6*6 ,0), attenuation matrix Λ dec =diag(0.01*I 6*6 ,0.3*I 6*6 ,0). The dynamic parameters of the bipedal module are shown in Table 1:

[0170] Table 1 Dynamic parameters of the bipedal module

[0171]

[0172]

[0173] 1. Conduct an impact test. Specifically: One of the features of the MPC framework proposed in this embodiment is the realization of dynamic correction of the model. To verify its effectiveness, an impact test was conducted in comparison with the standard MPC. In the simulation, the legged robot was set to walk along the +x direction at a constant speed. A step external disturbance force with an amplitude of 15N was applied to a point on the right side of the legged robot's head along the -y direction, causing it to deviate from the reference trajectory. The experimental process is as follows: Figure 3 As shown in Figure 2, facing the same external disturbance force, the robustness of the two is evaluated by quantifying which MPC controller can make the robot's desired trajectory deviate less and recover faster.

[0174] Figure 4 The curves of the roll angle and lateral displacement of the biped module during the experiment are given, where the shaded area represents the area of ​​external disturbance. Compared with the standard MPC, the method proposed in this embodiment can control the robot to return to the desired state faster when subjected to external disturbances, showing significant improvement in system robustness. The state estimation error change process in the two groups of experiments is shown in Figure 2. Figure 5 As shown, compared with the standard MPC, the method proposed in this embodiment presents a smaller state estimation error in rotational dynamics and improves the model prediction accuracy under disturbance conditions.

[0175] 2. Carry out a load-bearing walking test. Specifically: To further test the advantages of the method proposed in this embodiment in terms of disturbance resistance, a load-bearing walking comparison test was carried out on the proposed MPC controller and RF-MPC. The test setting is as follows: At the beginning of the simulation, a 5kg weight is placed on the head of the bipedal module. After running for a period of time, the weight of the load is increased to 8kg through the simulator, which accounts for about 88% of the total mass of the bipedal module.

[0176] Figure 6 and Figure 7 The test results of the simulation experiment are given. In the initial stage, both controllers can achieve stable walking and roughly reach the desired body height and posture. However, the legged robot using RF-MPC deviates slightly from the preset desired state in pitch angle, and the average body height is lower than that of OSF-MPC. After adding the load weight, the pitch angle of the robot using RF-MPC diverges and eventually falls down after 3s. In contrast, OSF-MPC is robust to external disturbances, and the legged robot can still maintain stable walking.

[0177] 3. Conduct a comparative test of the optimization schemes. Specifically: To verify the effectiveness of the optimization based on the input variation in this embodiment, a comparative experiment based on the input u (Scheme 1) and the optimization solution based on the input variation Δu (Scheme 2) was conducted. In this test, two different MPC optimization schemes (Scheme 1 and Scheme 2) were deployed on the simulation prototype, and the legged robot was set to walk at a constant speed. The advantages and disadvantages of the two schemes were evaluated by recording and comparing the MPC optimization objective function values ​​of the two schemes and the smoothness of the calculated optimal input u. The experimental results are as follows: Figure 8 As shown, Figure 8 The vertical ground reaction force on the middle right side is the data of the right leg during the movement of the leg-footed robot. Figure 8At t=10s, 10.2s, 10.4s and 10.6s, the switching moment from the swing phase to the support phase is reached. At this time, the plantar force inevitably has a large mutation. However, since the optimization based on Δu in Scheme 2 has a smaller state tracking error e(k) during the non-switching period, the MPC can maintain the legged robot near the desired state with a smaller input at the support / swing switching moment. Therefore, Scheme 2 can effectively alleviate the sudden contact force spike when the robot enters the support phase from the swing phase. At other moments, compared with Scheme 1, the advantage of Scheme 2 is more obvious, which significantly reduces the tracking error e(k). This confirms that the optimization based on the input change Δu can effectively alleviate the problem that the input penalty term in the objective function is too large relative to the state penalty term.

[0178] Embodiment 2

[0179] This embodiment discloses a state feedback model predictive control system for a legged robot.

[0180] like Fig. 9 As shown, a state feedback model predictive control system for a legged robot comprises:

[0181] The state information acquisition module is configured to: obtain the state information of the leg-foot robot; derive the single rigid body dynamics model of the leg-foot robot and discretize it according to the obtained state information;

[0182] The closed-loop state observation module is configured to: dynamically track disturbance information in a single rigid body dynamics model of the legged robot using a closed-loop state observer;

[0183] The state feedback MPC module is configured to: predict the motion state of the leg-foot robot at the future moment and solve the ground reaction force that meets the MPC constraint conditions according to the expected motion state and disturbance prediction value set by the leg-foot robot using the model predictive control strategy MPC;

[0184] The constraint relaxation WBC module is configured to: use the constraint relaxation WBC to perform secondary optimization on the MPC solution results and calculate the joint torque of the leg-foot robot.

[0185] The state estimation module is used to feed back the posture information of the leg-foot robot to the state feedback MPC module and the constraint relaxation WBC module.

[0186] Furthermore, the closed-loop state observation module is integrated into the state feedback MPC module. The state feedback model predictive control system provided by the present invention integrates both MPC and WBC, aiming to provide a robust control strategy for the legged robot, such as Fig. 9 As shown, represents the expected velocity of the legged robot body, u mpcrepresents the optimal plantar force calculated by MPC, τ cmd and τ * They are the joint torque output by WBC and the joint torque finally acting on the robot; Fig. 9 The blue digital marks in the figure indicate the operating frequency of each control submodule, that is, the operating frequency of the state feedback MPC module is 125 Hz, the operating frequency of the constraint relaxation WBC module is 500 Hz, and the operating frequency of the state estimation module is 500 Hz.

[0187] Embodiment 3

[0188] The purpose of this embodiment is to provide a computer-readable storage medium.

[0189] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in a state feedback model predictive control method for a legged robot as described in the first embodiment of the present disclosure.

[0190] Embodiment 4

[0191] The purpose of this embodiment is to provide an electronic device.

[0192] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in a state feedback model predictive control method for a legged robot as described in the first embodiment of the present disclosure are implemented.

[0193] The steps involved in the apparatuses of the above embodiments 2, 3 and 4 correspond to the method embodiment 1, and the specific implementation methods can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0194] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0195] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A state feedback model predictive control method for a legged robot, characterized in that: include: Acquire the state information of the leg-foot robot; derive the single rigid body dynamics model of the leg-foot robot and discretize it according to the acquired state information; The closed-loop state observer is used to dynamically track the disturbance information in the single rigid body dynamics model of the legged robot. According to the expected motion state and disturbance prediction value set by the leg-foot robot, the model predictive control strategy MPC is used to predict the motion state of the leg-foot robot at future moments and solve the ground reaction force that meets the MPC constraint conditions. Constraint relaxation WBC is used to perform secondary optimization on the MPC solution and calculate the joint torque of the leg-foot robot.

2. A state feedback model predictive control method for a legged robot as claimed in claim 1, characterized in that: According to the obtained state information, the single rigid body dynamics model of the legged robot is derived and discretized, including: The model predictive control strategy MPC is mathematically represented by solving a constrained optimization problem, and the generalized coordinate vector in the mathematical representation is differentially defined; the mathematical representation in the model predictive control strategy MPC optimization problem is mapped to a coordinate system, and the single rigid body dynamics model of the leg-legged robot is determined; based on the single rigid body dynamics model, the motion state of the leg-legged robot at future moments is predicted.

3. A state feedback model predictive control method for a legged robot as claimed in claim 2, characterized in that: Before predicting the motion state of the legged robot at a future moment based on the single rigid body dynamics model, the single rigid body dynamics model is first converted from a nonlinear model to a linear model containing time-varying disturbances.

4. A state feedback model predictive control method for a legged robot as claimed in claim 1, characterized in that: The closed-loop state observer is provided with state variables related to tracking the motion state of the leg-foot robot, and the state variables include: the real state of the leg-foot robot, the reconstructed state of the closed-loop state observer, and the predicted state of the model predictive control strategy MPC at future moments.

5. A state feedback model predictive control method for a legged robot as claimed in claim 1, characterized in that: It also includes quadratic programming based on input changes, specifically: defining a state tracking error, converting the cost function involved in the model predictive control strategy MPC into a discrete domain representation and defining an input change; writing the input change as an optimization variable into the defined state tracking error to achieve quadratic programming based on the input change.

6. A state feedback model predictive control method for a legged robot as claimed in claim 1, characterized in that: The model predictive control strategy MPC takes the closed-loop state observer equation as the prediction benchmark, uses the attenuation matrix to approximate the time domain evolution in the prediction domain, and then makes a preliminary prediction of the motion state of the legged robot at the next moment.

7. A state feedback model predictive control system for a legged robot, characterized in that: include: The state information acquisition module is configured to: obtain the state information of the leg-foot robot; derive the single rigid body dynamics model of the leg-foot robot and discretize it according to the obtained state information; The closed-loop state observation module is configured to: dynamically track disturbance information in a single rigid body dynamics model of the legged robot using a closed-loop state observer; The state feedback MPC module is configured to: predict the motion state of the leg-foot robot at the future moment and solve the ground reaction force that meets the MPC constraint conditions according to the expected motion state and disturbance prediction value set by the leg-foot robot using the model predictive control strategy MPC; The constraint relaxation WBC module is configured to: use the constraint relaxation WBC to perform secondary optimization on the MPC solution results and calculate the joint torque of the leg-foot robot.

8. A state feedback model predictive control system for a legged robot as claimed in claim 7, characterized in that: It also includes: a state estimation module, which is used to feed back the posture information of the leg-foot robot to the state feedback MPC module and the constraint relaxation WBC module; at the same time, the closed-loop state observation module is integrated into the state feedback MPC module.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps of a state feedback model predictive control method for a legged robot as described in any one of claims 1 to 6 are implemented.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the state feedback model predictive control method for a legged robot as described in any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • A method and system for motion control of bipedal robots based on deep reinforcement learning

    CN115128960B

  • Robust control method of dynamic biped walking robot

    CN109164705A

  • Robot balance control method and device, readable storage medium and robot

    CN113927585A

  • Biped robot motion control method and system based on deep reinforcement learning

    CN115128960A

  • Method and system for controlling stable movement of quadruped robot

    CN115755594A

Cited By

  • Robot dynamic balance control method and system

    CN121143417A

  • A robot dynamic balance control method and system

    CN121143417B