Control method and device for stable walking of exoskeleton robot
By constructing the state space expression of the augmented system and the on-policy data-driven Q learning algorithm based on the state feedback PI strategy, the problem of stable walking of the crutchless self-balancing lower limb exoskeleton robot under unknown user parameters is solved, and autonomous and stable walking under different patient loads is achieved, which improves the applicability and performance of the exoskeleton robot.
Patent Information
- Application Number
- CN202510643232.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-26
AI Technical Summary
The stability of the existing crutchless self-balancing lower limb exoskeleton robot walking control algorithm depends on the user's physical parameters, resulting in poor versatility and adaptability, making it difficult to maintain stable walking under unknown user parameters.
The on-policy data-driven Q learning algorithm based on the state feedback PI strategy is adopted to construct the state space expression of the augmented system, and the cost function and control law iteratively update the cost function and control law using the system state information to solve the optimal control problem in real time, and maintain the stable walking of the exoskeleton robot under the unknown dynamic information.
The crutchless self-balancing lower limb exoskeleton robot is realized to walk independently and stably under different patient loads, which improves the scope of application and overall performance, avoids dependence on system dynamic parameters, and enhances versatility and adaptability.
Smart Images

Figure CN120533692A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of exoskeleton robot control technology, and in particular to a control method, device, terminal equipment and computer-readable storage medium for an exoskeleton robot to stably walk. Background Art
[0002] Globally, millions of healthy individuals have lost their mobility due to conditions such as stroke, polio, multiple sclerosis, spinal cord injury, and cerebral palsy. Paralysis not only severely impacts patients' daily lives but also significantly increases the risk of secondary complications such as pressure ulcers, muscle atrophy, and osteoporosis. Consequently, self-balancing lower-limb exoskeleton robotic technology, capable of restoring the ability of paraplegics to walk, is gaining increasing attention in the field of motion assistance and rehabilitation.
[0003] Currently, existing self-balancing lower-limb exoskeleton robots primarily utilize two technical approaches: crutch-assisted exoskeleton robots and crutch-free self-balancing lower-limb exoskeleton robots. Due to the high-dimensional dynamics, kinematic redundancy, hardware constraints, and model uncertainty involved in exoskeleton robots, achieving stable motion control presents significant challenges. To reduce the difficulty of exoskeleton development, some researchers in this field have developed crutch-assisted exoskeleton robots. These exoskeleton robots rely on external assistive devices such as crutches or cranes to maintain dynamic balance during standing and walking. However, this approach increases costs and reduces its scope of application. Therefore, the development and control algorithm research of crutch-free self-balancing lower-limb exoskeleton robots are more suitable for current realities.
[0004] Stability issues are common in walking control algorithms for crutch-free, self-balancing lower-limb exoskeleton robots. These issues are highly dependent on the physical parameters of both the exoskeleton and the user (e.g., mass, inertia, height, etc.). Researchers at home and abroad have developed a variety of control algorithms for crutch-free, self-balancing exoskeleton robots capable of achieving autonomous, stable walking. However, most of these algorithms are based on the coupled dynamics between the exoskeleton and the user, resulting in walking stability being affected by the user's individual physical parameters. Due to the significant differences in physical parameters between different users, this dependency limits the versatility and adaptability of crutch-free, self-balancing exoskeleton robots. Summary of the Invention
[0005] In order to address the deficiencies of the above-mentioned prior art, the present invention provides a control method, device, terminal device and computer-readable storage medium for the stable walking of an exoskeleton robot, so that the exoskeleton robot can still maintain stable walking even when the user's physical parameters are unknown, thereby improving the scope of application and practical application value of the crutch-free self-balancing lower limb exoskeleton robot.
[0006] The first object of the present invention is to provide a control method for an exoskeleton robot to walk stably.
[0007] The second object of the present invention is to provide a control device for an exoskeleton robot to walk stably.
[0008] The third object of the present invention is to provide a terminal device.
[0009] A fourth object of the present invention is to provide a computer-readable storage medium.
[0010] The first object of the present invention can be achieved by adopting the following technical solutions:
[0011] A method for controlling an exoskeleton robot to stably walk, the method comprising:
[0012] Based on the crutch-free self-balancing lower limb exoskeleton robot model, the state space expression of the augmented system is constructed;
[0013] Based on the state space expression of the augmented system, an on-policy data-driven Q-learning algorithm based on a state feedback PI strategy is used to solve the control problem. The on-policy data-driven Q-learning algorithm based on a state feedback PI strategy utilizes system state information measured along the system trajectory, and approximates the cost function and control law in the dynamic programming method through continuous iterative updates, thereby solving the optimal control problem online and in real time.
[0014] Furthermore, the state space expression of the augmented system is:
[0015]
[0016] The state quantity of the augmented system is
[0017] Where k is the discrete time sampling moment of the system, X k+1 、X k are the states of the augmented system at time k+1 and k respectively, and All are constant matrices; are the discrete state variables and control inputs of the crutch-free self-balancing lower limb exoskeleton robot model at time k; is the reference signal system status;
[0018] In order to avoid V, it is necessary to satisfy Hurwitz's condition, which leads to the i-k The performance indicator function is:
[0019]
[0020] Among them, r(e i ,u i) is the utility function at the i-th moment, e i 、u i are the tracking error and control input at the i-th moment respectively; Q = Q T ≥0 and R=R T >0 are all positive definite matrices.
[0021] Furthermore, the process of constructing the state space expression of the augmented system includes:
[0022] The self-balancing lower limb exoskeleton robot is simplified into a linear inverted pendulum model, and the linear inverted pendulum model is represented in the same way in the x-axis direction and the y-axis direction;
[0023] The expression of the linear inverted pendulum model in the x-axis direction is:
[0024]
[0025] Among them, [x CoM ,z CoM ] T 、[x zmp ,z zmp ] T , m and g are the center of mass position, zero moment point position, mass and gravitational acceleration of the exoskeleton robot respectively; are the accelerations of the center of mass of the exoskeleton robot in the z-axis and x-axis directions respectively;
[0026] If the ground is flat and the height of the center of mass is fixed, the linear inverted pendulum model is expressed as:
[0027]
[0028] Select x CoM 、 and As system state variables is the velocity of the center of mass of the exoskeleton robot in the x-axis direction; zmp As system output The state space expression of the self-balancing lower limb exoskeleton robot model is:
[0029]
[0030] in,
[0031] Discretize the state space expression of the self-balancing lower limb exoskeleton robot model and obtain:
[0032] x k+1 =Ax k +Bu k
[0033] y k =Cx k
[0034] in, is the discrete state variable at time k+1, is the discrete output at time k, is a constant matrix;
[0035] The state space equation of the reference signal generator is:
[0036]
[0037] in, is the reference signal output, is a constant matrix;
[0038] Define tracking error:
[0039] e k =y k -r k
[0040] For the linear quadratic tracking problem of a balanced lower limb exoskeleton robot system, the cost function is a quadratic function consisting of the tracking error and the control input, and its performance index is defined as follows:
[0041]
[0042] Based on the self-balancing lower limb exoskeleton robot system and the reference signal generator system, the state space expression of the augmented system is obtained.
[0043] Furthermore, the state space expression based on the augmented system uses an on-policy data-driven Q learning algorithm based on a state feedback PI strategy to solve the control problem, including:
[0044] Based on the performance index function, the linear quadratic tracking Bellman equation of the augmented system is expressed as follows:
[0045]
[0046] The linear quadratic pursuit Bellman equation of the augmented system is rewritten as k And the quadratic form represented by the positive definite matrix U:
[0047]
[0048] Among them, U=U T >0;
[0049] From the linear quadratic pursuit Bellman equation and quadratic form of the augmented system, we can get:
[0050]
[0051] in,
[0052] Define the Hamiltonian function:
[0053]
[0054] According to the necessary conditions for optimality, let We can get:
[0055] u k =-KX k =-(R+γG T UG) -1 γG T UFX k (4-3)
[0056] Substituting Equation (4-3) into Equation (4-2) and combining the state space expression of the augmented system, the algebraic Riccati equation of the augmented system can be obtained as:
[0057] Ω-U+γF T UF-γ 2 F T UG(R+γG T UG) -1 G T UF=0
[0058] For augmented systems, in order to solve the problem of performance index optimization, we first solve U in the algebraic Riccati equation; the optimal cost and optimal control law can be obtained by equations (4-1) and (4-3), respectively;
[0059] For equation (4-2), define the Q function of the discrete augmented system:
[0060]
[0061] Substituting the state space expression of the augmented system into equation (4-4), we can obtain:
[0062]
[0063] According to the necessary conditions for optimality, let equation (4-5) be k Taking partial derivatives, we can get:
[0064]
[0065] From formula (4-2) and formula (4-4), we can know that J(X k )=Q(X k ,uk ), so the Q function Bellman equation is expressed as:
[0066]
[0067] Defining state variables Then from formula (4-5) and formula (4-7), we can get:
[0068]
[0069] Based on equations (4-6) and (4-8), the on-policy data-driven Q learning algorithm based on the state feedback PI strategy is described as:
[0070] Step 41, initialization: j = 0, with a stable control law Start iterating;
[0071] Step 42: Strategy evaluation: Solve using formula (4-8) Right now:
[0072]
[0073] Step 43: Strategy Improvement: Update Control Law
[0074] Step 44: Judgment Is it true? If so, stop the iteration and use Go to control system, l represents the algorithm learning accuracy; otherwise, let j=j+1, and return to step 42 to continue to perform subsequent operations.
[0075] Furthermore, the method further comprises:
[0076] Based on the state space expression of the augmented system, an on-policy data-driven Q learning algorithm based on an output feedback PI strategy is used to solve the control problem; the on-policy data-driven Q learning algorithm based on an output feedback PI strategy utilizes input and output sequences to reconstruct the system state, and based on the reconstructed system state, the on-policy data-driven Q learning algorithm based on a state feedback PI strategy is used to solve the control problem.
[0077] Furthermore, the state space expression based on the augmented system uses an on-policy data-driven Q learning algorithm based on an output feedback PI strategy to solve the control problem, including:
[0078] The state X of the augmented system k The input sequence consists of [kN,k-1] time periods Output sequence And the reference trajectory sequence Indicates that:
[0079]
[0080] Where N is the order of the system, and the input sequence is Output sequence
[0081] U N =[G AG A 2 G…A N-1 G];
[0082] W N =[(CA N-1 ) T (CA N-2 ) T … (CA) T C T ] T ;
[0083]
[0084] Based on the state X of the augmented system k And the on-policy data-driven Q learning algorithm based on state feedback PI strategy, using the state reconstruction method, derives the on-policy data-driven Q learning algorithm based on output feedback PI strategy to solve the control problem.
[0085] Furthermore, the state X based on the augmented system k And the on-policy data-driven Q learning algorithm based on the state feedback PI strategy, using the state reconstruction method, derives the on-policy data-driven Q learning algorithm based on the output feedback PI strategy, including:
[0086] First, the state X of the augmented system is k Substituting into equation (4-7), we obtain the Q function based on input / output measurement data:
[0087]
[0088] in,
[0089] And M is divided into in:
[0090]
[0091] u for the Q function based on input / output measurement data k Find the partial derivative and get the optimal controller u k for:
[0092]
[0093] in,
[0094] Substituting Equation (4-6) into the Q function based on input / output measurement data, we obtain the Bellman equation of the Q function consisting of the input, output, and reference trajectory sequence:
[0095]
[0096] in,
[0097] At the same time, u k+1 It can be calculated by the following formula:
[0098]
[0099] In order to obtain the components of the matrix M in the Q function based on the input / output measurement data, the parameters in the Q function based on the input / output measurement data are linearized as follows:
[0100]
[0101] in,
[0102] Where m ij Represents the element represented by the i-th row and j-th column in the M matrix, i, j = 1, 2, ..., l, l = aN + 2 * pN + a; vector Obtained through Kronecker product operation:
[0103] The Bellman equation of the Q function consisting of the input, output and reference trajectory sequence is rewritten as:
[0104]
[0105] Based on the optimal controller u k And the rewritten Q function Bellman equation, the on-policy data-driven Q learning algorithm based on the output feedback PI strategy is described as:
[0106] Step 71, initialization: j = 0, with a stable control law Start iterating;
[0107] Step 72: Strategy evaluation: Solve the Bellman equation using the rewritten Q function
[0108]
[0109] Step 73: Strategy Improvement: Update Control Law for:
[0110]
[0111] Step 74: Judgment Is it true? If so, stop the iteration and use Go to control the system; otherwise, let j=j+1 and return to step 72 to continue execution.
[0112] The second object of the present invention can be achieved by adopting the following technical solutions:
[0113] A control device for an exoskeleton robot to stably walk, the device comprising:
[0114] A building block for constructing the state space representation of the augmented system based on a crutch-free self-balancing lower limb exoskeleton robot model;
[0115] The first solving module is used to solve the control problem based on the state space expression of the augmented system using an on-policy data-driven Q learning algorithm based on a state feedback PI strategy; the on-policy data-driven Q learning algorithm based on a state feedback PI strategy utilizes system state information measured along the system trajectory, and approximates the cost function and control law in the dynamic programming method through continuous iterative updates, thereby solving the optimal control problem online and in real time.
[0116] Furthermore, the device further comprises:
[0117] The second solving module is used to solve the control problem based on the state space expression of the augmented system using an on-policy data-driven Q learning algorithm based on an output feedback PI strategy; the on-policy data-driven Q learning algorithm based on an output feedback PI strategy uses input and output sequences to reconstruct the system state, and then uses the on-policy data-driven Q learning algorithm based on a state feedback PI strategy to solve the control problem based on the reconstructed system state.
[0118] The third object of the present invention can be achieved by adopting the following technical solutions:
[0119] A terminal device includes a processor and a memory for storing a program executable by the processor. When the processor executes the program stored in the memory, the above-mentioned control method for stable walking of the exoskeleton robot is implemented.
[0120] The fourth object of the present invention can be achieved by adopting the following technical solutions:
[0121] A computer-readable storage medium stores a program, which, when executed by a processor, implements the above-mentioned control method for stable walking of an exoskeleton robot.
[0122] The present invention has the following beneficial effects compared to the prior art:
[0123] The present invention provides a universal control method for a crutch-free self-balancing lower limb exoskeleton robot that can maintain autonomous and stable walking when carrying different patient loads. This method is different from traditional control methods. For example, the linear quadratic tracking (LQT) of optimal control requires knowing all the dynamic information of the system or identifying the system. The method provided by the present invention can ensure that the exoskeleton robot system tracks the preset gait and operates stably when the internal dynamic parameter information of the system is unknown, thereby improving the scope of application and overall performance of the crutch-free self-balancing lower limb exoskeleton robot. BRIEF DESCRIPTION OF THE DRAWINGS
[0124] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0125] Figure 1 This is a flow chart of a method for controlling stable walking of an exoskeleton robot according to Example 1 of the present invention;
[0126] Figure 2 This is a simplified model LIPM structure diagram of a crutch-free self-balancing lower limb exoskeleton robot according to Example 1 of the present invention;
[0127] Figure 3 This is a graph showing the x-axis target trajectory of the crutch-free self-balancing lower limb exoskeleton robot according to Example 1 of the present invention;
[0128] Figure 4 A graph showing the y-axis target trajectory of the crutch-free self-balancing lower limb exoskeleton robot according to Example 1 of the present invention;
[0129] Figure 5 This is a structural block diagram of a control device for stable walking of an exoskeleton robot according to Example 2 of the present invention;
[0130] Figure 6 This is a structural block diagram of the terminal device of embodiment 3 of the present invention. DETAILED DESCRIPTION
[0131] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. It should be understood that the specific embodiments described are only used to explain this application and are not used to limit this application.
[0132] In the description of the embodiments of this application, the technical terms "first," "second," etc. are used only to distinguish different objects and should not be understood to indicate or imply relative importance or to implicitly indicate the quantity, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "plurality" is more than two, unless otherwise specifically defined.
[0133] Example 1:
[0134] like Figure 1 As shown, this embodiment provides a control method for an exoskeleton robot to stably walk, comprising the following steps:
[0135] S101. Based on the crutch-free self-balancing lower limb exoskeleton robot model, construct the state space expression of the augmented system.
[0136] like Figure 2 As shown in Figure 1, the self-balancing lower limb exoskeleton robot is simplified into a linear inverted pendulum model (LIPM), with the x-axis and y-axis represented in the same way. Below, using the x-axis as an example, this simplified model is used to generate the walking trajectory of a crutch-free self-balancing lower limb exoskeleton robot (SBLLE). A stable walking feedback controller is designed to maintain the stability of the crutch-free self-balancing lower limb exoskeleton robot during walking. The linear inverted pendulum model (LIPM) can be expressed as:
[0137]
[0138] Among them, [x CoM ,z CoM ] T 、[x zmp ,z zmp ] T , m and g are the center of mass (CoM) position, zero moment point (ZMP) position, mass and gravitational acceleration of the exoskeleton robot respectively; x CoM 、z CoM are the x-axis coordinates and z-axis coordinates of the center of mass of the exoskeleton robot; zmp 、z zmp are the x-axis coordinates and z-axis coordinates of the zero moment point respectively; are the accelerations of the center of mass of the exoskeleton robot in the z-axis and x-axis directions, respectively.
[0139] The center of mass of an exoskeleton robot is the average location of the mass distribution of the entire system (robot body + user). The zero moment point of an exoskeleton robot is the point in the robot's contact area with the ground where the net moment exerted by the ground on the robot is zero horizontally. In short, if the ZMP falls within the support surface, the robot is dynamically stable.
[0140] If the ground is considered to be flat and the height of the center of mass (CoM) is a fixed value and does not change, the inverted pendulum of the crutch-free self-balancing lower limb exoskeleton robot degenerates into a linear inverted pendulum, as shown in Figure 2 As shown. The linear inverted pendulum model (LIPM) in this case can be expressed as:
[0141]
[0142] Select x CoM 、 and As system state variables is the velocity of the center of mass of the exoskeleton robot in the x-axis direction; zmp As system output The state space expression of the self-balancing lower limb exoskeleton robot model is:
[0143]
[0144] in,
[0145] Discretize the model formula (3) and describe the discretized state space expression as follows:
[0146]
[0147] Where k is the discrete time sampling moment of the system, and is the discrete state change between time k+1 and time k, and are the discrete input and output quantities at time k, are all constant matrices.
[0148] Specifically, A, B, and C are matrices corresponding to the self-balancing lower limb exoskeleton robot model in the discrete sampling period T:
[0149]
[0150] It should be noted that T in this article represents the transpose of a matrix, unless otherwise specified.
[0151] The state space equation of the reference signal generator is:
[0152]
[0153] in, is the reference signal system state, is the reference signal output, are all constant matrices. The tracking error is defined as follows:
[0154] e k =y k -r k (6)
[0155] For the linear quadratic tracking (LQT) problem of a balanced lower limb exoskeleton robot system, the cost function is a quadratic function consisting of the tracking error and the control input, and its performance index is defined as follows:
[0156]
[0157] Among them, r(e i ,u i ) is the utility function at the i-th moment, Q = Q T ≥0 and R=R T >0 are all positive definite matrices.
[0158] Based on the self-balancing lower limb exoskeleton robot system and the reference signal generator system, the state space equation expression of the augmented system can be obtained as follows:
[0159]
[0160] Among them, the state quantity of the augmented system is
[0161] In order to avoid V, it is necessary to satisfy Hurwitz's condition, which leads to the i-k The performance indicator function is:
[0162]
[0163] S102. Based on the state space expression of the augmented system, an on-policy data-driven Q learning algorithm based on the state feedback PI strategy is used to solve the control problem.
[0164] In this embodiment, when the dynamic information is not completely known but the system state can be measured, an on-policy data-driven Q learning algorithm based on a state feedback PI strategy is used to obtain the optimal controller.
[0165] Based on Equation (9), the linear quadratic tracking (LQT) Bellman equation of the augmented system can be expressed as follows:
[0166]
[0167] The performance index function (10) can be rewritten as k And the quadratic form represented by the positive definite matrix U:
[0168]
[0169] Among them, U=U T >0.
[0170] Therefore, from formula (10) and formula (11), we can get:
[0171]
[0172] in,
[0173] The Hamiltonian function is defined as follows:
[0174]
[0175] According to the necessary conditions for optimality, let equation (13) be k Find the partial derivative, that is We can get:
[0176] u k =-KX k =-(R+γG T UG) -1 γG T UFX k (14)
[0177] Substituting Equation (14) into Equation (12) and combining it with Equation (8), we can obtain the algebraic Riccati equation (ARE) of the augmented system:
[0178] Ω-U+γF T UF-γ 2 F T UG(R+γG T UG) -1 G T UF=0 (15)
[0179] For the augmented system (Equation (8)), in order to solve the problem of optimizing the performance index (Equation (9)), we first need to solve the algebraic Riccati equation U in Equation (15); then, the optimal cost and optimal control law can be obtained by Equations (11) and (14), respectively.
[0180] For the Bellman equation (12) represented by the augmented system state and the positive definite matrix U, the Q function of the discrete augmented system is defined as:
[0181]
[0182] Substituting the augmented system state space equation (8) into (16), we can obtain:
[0183]
[0184] According to the necessary conditions for optimality, let equation (17) be k Find the partial derivative, that is We can get:
[0185]
[0186] From Bellman equation (12) and discrete augmented system Q function (16), we know that J(X k )=Q(X k ,u k ), so the Q function Bellman equation can be expressed as:
[0187]
[0188] Defining state variables Then from formula (17) and formula (19), we can get:
[0189]
[0190] Therefore, using Bellman equation (20), the on-policy data-driven Q learning algorithm based on the state feedback PI strategy can be described as:
[0191] Step 1: Initialize j = 0, with a stable control law Start iterating.
[0192] Step 2: Strategy evaluation, using the Q function Bellman equation (20) to solve
[0193]
[0194] Step 3: Strategy improvement, update control law for:
[0195]
[0196] Step 4: Judgement (l represents the learning accuracy of the algorithm) is established, if so, stop the iteration and use Go to control the system; otherwise, let j = j + 1 and return to step 2 to continue execution.
[0197] The solution process of the classical optimal control algorithm is offline. To execute these classical algorithms, it is necessary to know the complete information of the system dynamics in advance. Therefore, the traditional optimal control algorithm is generally not applicable to situations where the parameters of the controlled system change.
[0198] The on-policy, data-driven Q-learning algorithm based on a state-feedback PI strategy, provided in this embodiment, leverages system state information measured along the system trajectory and, through continuous iterative updates, approximates the cost function and control law used in traditional dynamic programming methods (such as linear quadratic programming (LQR)), thereby solving optimal control problems online and in real time. This algorithm can adapt to parameter changes in the controlled system in real time without prior knowledge of the system's dynamics. As the controlled system parameters change, an optimal controller can still be obtained to control the system output.
[0199] S103. Based on the state space expression of the augmented system, an on-policy data-driven Q learning algorithm based on the output feedback PI strategy is used to solve the control problem.
[0200] In practical engineering applications, system state information is not always easy to measure, and the measurement cost of some system state information is also high. The system state of the augmented system can be reconstructed using measurement data such as input data, output data, and reference trajectory data.
[0201] In this embodiment, when dynamic information is incomplete and the system state is difficult to measure, an on-policy data-driven Q-learning algorithm based on an output-feedback PI strategy is used to obtain the optimal controller. This on-policy data-driven Q-learning algorithm based on an output-feedback PI strategy is an improvement on the on-policy data-driven Q-learning algorithm based on a state-feedback PI strategy. Because some system states are difficult to measure, the system state is reconstructed using input and output data. Based on this reconstructed state, the on-policy data-driven Q-learning algorithm based on a state-feedback PI strategy is then used to solve the control problem.
[0202] When (A, C) is observable, the state X of the augmented system k The input sequence can be [kN, k-1] time period Output sequence And the reference trajectory sequence Indicates that:
[0203]
[0204] Where N represents the order of the system, and the input sequence Output sequence Reference trajectory sequence
[0205]
[0206] U N =[G AG A 2 G … A N-1 G]
[0207] W N =[(CA N-1 ) T (CA N-2 ) T … (CA) T C T ] T
[0208]
[0209] Using the state reconstruction method, an on-policy data-driven Q learning algorithm based on the output feedback PI strategy is derived. First, substitute Equation (23) into the Q function Equation (20) to obtain the Q function based on input / output measurement data:
[0210]
[0211] in, And M is divided into in:
[0212]
[0213] According to the optimality principle, the optimal controller u k Should meet For u in formula (24) k Find the partial derivative to get the optimal controller u k for:
[0214]
[0215] in,
[0216] Substituting Equation (19) into Equation (24), we can obtain a Q-function Bellman equation consisting of the input, output, and reference trajectory sequences:
[0217]
[0218] in,
[0219] At the same time, u k+1It can be calculated by the following formula:
[0220]
[0221] In order to obtain the components of the matrix M in equation (24), the parameters of equation (24) can be linearized as follows:
[0222]
[0223] in:
[0224]
[0225] In the above formula, m ij (i,j=1,2,…,l,l=aN+2*pN+a) represents the element represented by the i-th row and j-th column in the M matrix. The vector It can be obtained by the following Kronecker product operation:
[0226]
[0227] Right now
[0228] The Bellman equation (27) based on the measured data can be rewritten as:
[0229]
[0230] Therefore, based on Equations (26) and (29), we can obtain the on-policy data-driven Q learning algorithm based on the output feedback PI strategy:
[0231] Step 1: Initialize j = 0, with a stable control law Start iterating.
[0232] Step 2: Strategy evaluation, using Bellman equation (29) to solve
[0233]
[0234] Step 3: Strategy improvement, update the control law according to the following formula
[0235]
[0236] Step 4: Judgement (l represents the learning accuracy of the algorithm) is established, if so, stop the iteration and use Go to control the system; otherwise, let j = j + 1 and return to step 2 to continue execution.
[0237] This embodiment provides an on-policy data-driven Q learning algorithm based on the output feedback PI strategy. The algorithm uses the input data sequence, output data sequence and reference signal data sequence to solve the performance index optimization problem of the augmented system. In the strategy evaluation formula (30), the least squares method is used to solve the unknown vector Since the vector There are l(l+1) / 2 independent vectors, so the number of collected data samples is L≥l(l+1) / 2, and they are stored in the following data matrix:
[0238]
[0239] in, Therefore, the least squares solution of formula (30) can be calculated as:
[0240]
[0241] However, from Equation (26), we can see that due to the control input u k Depend on and Linear representation, so Φ j Dissatisfaction with rank, leading to (Φ j ) T Φ j The matrix is not invertible. To solve this problem, the control input u k Add detection noise w k , ensuring that Equation (33) has a unique solution. The purpose of adding detection noise is to make the matrix Φ j The full rank condition is satisfied, namely:
[0242] rank(Φ j )=l(l+1) / 2 (34)
[0243] In the strategy improvement step, the improved controller is solved Used to generate data during the j+1th iteration.
[0244] The effectiveness of the controller of the crutch-free self-balancing lower limb exoskeleton robot using the on-policy data-driven Q learning algorithm based on the output feedback PI strategy is verified through MATLAB simulation.
[0245] The control time is selected as T = 0.001, and the center of mass (CoM) height is selected as z CoM =0.883, the robot weight is selected as m=84.766. The tracking trajectory is selected as F y(t) is a square wave with a period of 4 and an amplitude of 0.3. The initial conditions of the external system are set to x1(0)=0, x2(0)=0, and x3(0)=0.
[0246] According to the selected parameters, the simulation results can be obtained by referring to Figure 3 、 4 . Figure 3 The curve showing the actual x-axis displacement and target trajectory of the crutch-free self-balancing exoskeleton robot over time. Figure 4 The graph shows the actual y-axis displacement and target trajectory of the crutch-free self-balancing exoskeleton robot over time. The simulation graph shows that the output of the self-balancing lower-limb exoskeleton robot tracks the target trajectory. This simulation result demonstrates that the feedback controller designed in this embodiment can effectively control and regulate the output of the crutch-free self-balancing lower-limb exoskeleton robot, enabling the robot to maintain stable walking despite varying user loads.
[0247] Those skilled in the art will appreciate that all or part of the steps in the method for implementing the above embodiments may be completed by instructing related hardware through a program, and the corresponding program may be stored in a computer-readable storage medium.
[0248] It should be noted that although the method operations of the above embodiments are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all of the illustrated operations must be performed to achieve the desired results. Rather, the depicted steps may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into a single step, and / or a single step may be broken down into multiple steps.
[0249] Example 2:
[0250] like Figure 5 As shown, this embodiment provides a control device for stable walking of an exoskeleton robot, which includes a construction module 501, a first solution module 502 and a second solution module 503, wherein:
[0251] A construction module 501 is used to construct a state space expression of an augmented system based on a crutch-free self-balancing lower limb exoskeleton robot model;
[0252] The first solution module 502 is used to solve the control problem based on the state space expression of the augmented system using an on-policy data-driven Q learning algorithm based on a state feedback PI strategy; the on-policy data-driven Q learning algorithm based on a state feedback PI strategy utilizes system state information measured along the system trajectory, and continuously iteratively updates the cost function and control law in the approximate dynamic programming method, thereby solving the optimal control problem online and in real time.
[0253] Furthermore, the device further comprises:
[0254] The second solution module 503 is used to solve the control problem based on the state space expression of the augmented system using an on-policy data-driven Q learning algorithm based on an output feedback PI strategy; the on-policy data-driven Q learning algorithm based on an output feedback PI strategy uses input and output sequences to reconstruct the system state, and then uses the on-policy data-driven Q learning algorithm based on a state feedback PI strategy to solve the control problem based on the reconstructed system state.
[0255] The specific implementation of each module in this embodiment can be found in the above-mentioned embodiment 1, and will not be described one by one here; it should be noted that the system provided in this embodiment is only illustrated by the division of the above-mentioned functional modules. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above.
[0256] Example 3:
[0257] This embodiment provides a terminal device, which can be a computer, such as Figure 6 As shown, it comprises a processor 602, a memory, an input device 603, a display 604 and a network interface 605 connected via a system bus 601. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 606 and an internal memory 607. The non-volatile storage medium 606 stores an operating system, a computer program and a database. The internal memory 607 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor 602 executes the computer program stored in the memory, the control method for the stable walking of the exoskeleton robot of the above-mentioned embodiment 1 is implemented.
[0258] Example 4:
[0259] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method for controlling the stable walking of the exoskeleton robot of the above-mentioned embodiment 1 is implemented.
[0260] It should be noted that the computer-readable storage medium of the present embodiment may be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0261] The above is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes based on the technical solution and inventive concept of the present invention within the scope disclosed by the present invention, which falls within the scope of protection of the present invention.
Claims
1. A control method for stable walking of an exoskeleton robot, characterized in that: The method comprises: Based on the crutch-free self-balancing lower limb exoskeleton robot model, the state space expression of the augmented system is constructed; Based on the state space expression of the augmented system, an on-policy data-driven Q-learning algorithm based on a state feedback PI strategy is used to solve the control problem. The on-policy data-driven Q-learning algorithm based on a state feedback PI strategy utilizes system state information measured along the system trajectory, and approximates the cost function and control law in the dynamic programming method through continuous iterative updates, thereby solving the optimal control problem online and in real time.
2. The control method according to claim 1, characterized in that: The state space expression of the augmented system is: The state quantity of the augmented system is Among them, X k+1 、X k are the states of the augmented system at time k+1 and k respectively, and All are constant matrices; are the discrete state variables and control inputs of the crutch-free self-balancing lower limb exoskeleton robot model at time k; is the reference signal system status; In order to avoid V, it is necessary to satisfy Hurwitz's condition, which leads to the i-k The performance indicator function is: Among them, r(e i ,u i ) is the utility function at the i-th moment, e i 、u i are the tracking error and control input at the i-th moment respectively; Q = Q T ≥0 and R=R T >0 are all positive definite matrices.
3. The control method according to claim 2, characterized in that: The process of constructing the state space expression of the augmented system includes: The self-balancing lower limb exoskeleton robot is simplified into a linear inverted pendulum model, and the linear inverted pendulum model is represented in the same way in the x-axis direction and the y-axis direction; The expression of the linear inverted pendulum model in the x-axis direction is: Among them, [x CoM ,z CoM ] T 、[x zmp ,z zmp ] T , m and g are the center of mass position, zero moment point position, mass and gravitational acceleration of the exoskeleton robot respectively; are the accelerations of the center of mass of the exoskeleton robot in the z-axis and x-axis directions respectively; If the ground is flat and the height of the center of mass is fixed, the linear inverted pendulum model is expressed as: Select x CoM 、 and As system state variables is the velocity of the center of mass of the exoskeleton robot in the x-axis direction; zmp As system output The state space expression of the self-balancing lower limb exoskeleton robot model is: in, Discretize the state space expression of the self-balancing lower limb exoskeleton robot model and obtain: x k+1 =Ax k +Bu k y k =Cx k in, is the discrete state variable at time k+1, is the discrete output at time k, is a constant matrix; The state space equation of the reference signal generator is: in, is the reference signal output, is a constant matrix; Define tracking error: e k =y k -r k For the linear quadratic tracking problem of a balanced lower limb exoskeleton robot system, the cost function is a quadratic function consisting of the tracking error and the control input, and its performance index is defined as follows: Based on the self-balancing lower limb exoskeleton robot system and the reference signal generator system, the state space expression of the augmented system is obtained.
4. The control method according to any one of claims 2 and 3, characterized in that: The state space expression based on the augmented system is used to solve the control problem using an on-policy data-driven Q learning algorithm based on a state feedback PI strategy, including: Based on the performance index function, the linear quadratic tracking Bellman equation of the augmented system is expressed as follows: The linear quadratic pursuit Bellman equation of the augmented system is rewritten as k And the quadratic form represented by the positive definite matrix U: Where, U = U T >0; From the linear quadratic pursuit Bellman equation and quadratic form of the augmented system, we can get: in, Define the Hamiltonian function: According to the necessary conditions for optimality, let We can get: u k =-KX k =-(R+γG T UG) -1 γG T UFX k (4-3) Substituting Equation (4-3) into Equation (4-2) and combining the state space expression of the augmented system, the algebraic Riccati equation of the augmented system can be obtained as: Ω-U+γF T UF-γ 2 F T UG(R+γG T (UG) -1 G T UF=0 For augmented systems, in order to solve the problem of performance index optimization, we first solve U in the algebraic Riccati equation; the optimal cost and optimal control law can be obtained by equations (4-1) and (4-3), respectively; For equation (4-2), define the Q function of the discrete augmented system: Substituting the state space expression of the augmented system into equation (4-4), we can obtain: According to the necessary conditions for optimality, let equation (4-5) be k Taking partial derivatives, we can get: From formula (4-2) and formula (4-4), we can know that J(X k )=Q(X k ,u k ), so the Q function Bellman equation is expressed as: Defining state variables Then from formula (4-5) and formula (4-7), we can get: Based on equations (4-6) and (4-8), the on-policy data-driven Q learning algorithm based on the state feedback PI strategy is described as: Step 41, initialization: j = 0, with a stable control law Start iterating; Step 42: Strategy evaluation: Solve using formula (4-8) Right now: Step 43: Strategy Improvement: Update Control Law Step 44: Judgment Is it true? If so, stop the iteration and use Go to control system, l represents the algorithm learning accuracy; otherwise, let j=j+1, and return to step 42 to continue to perform subsequent operations.
5. The control method according to claim 4, characterized in that: The method further comprises: Based on the state space expression of the augmented system, an on-policy data-driven Q learning algorithm based on an output feedback PI strategy is used to solve the control problem; the on-policy data-driven Q learning algorithm based on an output feedback PI strategy utilizes input and output sequences to reconstruct the system state, and based on the reconstructed system state, the on-policy data-driven Q learning algorithm based on a state feedback PI strategy is used to solve the control problem.
6. The control method according to claim 5, characterized in that: The state space expression based on the augmented system is used to solve the control problem using an on-policy data-driven Q learning algorithm based on an output feedback PI strategy, including: The state X of the augmented system k The input sequence consists of [kN,k-1] time periods Output sequence And the reference trajectory sequence Indicates that: Where N is the order of the system, and the input sequence is Output sequence U N =[G AG A 2 G…A N-1 G]; W N =[(CA N-1 ) T (THAT N-2 ) T …(THAT) T C T ] T ; Based on the state X of the augmented system k And the on-policy data-driven Q learning algorithm based on state feedback PI strategy, using the state reconstruction method, derives the on-policy data-driven Q learning algorithm based on output feedback PI strategy to solve the control problem.
7. The control method according to claim 6, characterized in that: The state X based on the augmented system k And the on-policy data-driven Q learning algorithm based on the state feedback PI strategy, using the state reconstruction method, derives the on-policy data-driven Q learning algorithm based on the output feedback PI strategy, including: First, the state X of the augmented system is k Substituting into equation (4-7), we obtain the Q function based on input / output measurement data: in, And M is divided into in: u for the Q function based on input / output measurement data k Find the partial derivative and get the optimal controller u k for: in, Substituting Equation (4-6) into the Q function based on input / output measurement data, we obtain the Bellman equation of the Q function consisting of the input, output, and reference trajectory sequence: in, At the same time, u k+1 It can be calculated by the following formula: In order to obtain the components of the matrix M in the Q function based on the input / output measurement data, the parameters in the Q function based on the input / output measurement data are linearized as follows: in, Where m ij Represents the element represented by the i-th row and j-th column in the M matrix, i, j = 1, 2, ..., l, l = aN + 2 * pN + a; vector Obtained through Kronecker product operation: The Bellman equation of the Q function consisting of the input, output and reference trajectory sequence is rewritten as: Based on the optimal controller u k And the rewritten Q function Bellman equation, the on-policy data-driven Q learning algorithm based on the output feedback PI strategy is described as: Step 71, initialization: j = 0, with a stable control law Start iterating; Step 72: Strategy evaluation: Solve the Bellman equation using the rewritten Q function Step 73: Strategy Improvement: Update Control Law for: Step 74: Judgment Is it true? If so, stop the iteration and use Go to control the system; otherwise, let j=j+1 and return to step 72 to continue execution.
8. A control device for stable walking of an exoskeleton robot, characterized in that: The device comprises: A building block for constructing the state space representation of the augmented system based on a crutch-free self-balancing lower limb exoskeleton robot model; The first solving module is used to solve the control problem based on the state space expression of the augmented system using an on-policy data-driven Q learning algorithm based on a state feedback PI strategy; the on-policy data-driven Q learning algorithm based on a state feedback PI strategy utilizes system state information measured along the system trajectory, and approximates the cost function and control law in the dynamic programming method through continuous iterative updates, thereby solving the optimal control problem online and in real time.
9. The control device according to claim 8, characterized in that The device further comprises: The second solving module is used to solve the control problem based on the state space expression of the augmented system using an on-policy data-driven Q learning algorithm based on an output feedback PI strategy; the on-policy data-driven Q learning algorithm based on an output feedback PI strategy uses input and output sequences to reconstruct the system state, and then uses the on-policy data-driven Q learning algorithm based on a state feedback PI strategy to solve the control problem based on the reconstructed system state.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the control method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Mobile robot tracking control method based on multi-target point information fusion
CN118131628A
Control method for off-policy output feedback data driving Q learning based on VI strategy
CN118709559A