Under-actuated system hierarchical sliding mode optimization control method based on adaptive dynamic programming
Patent Information
- Application Number
- CN202410022633.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-01-08
AI Technical Summary
然而,随着工业过程的不断发展,传统的单层滑膜变结构控制已无法满足实际要求,因而具有分层结构的滑膜控制近年来备受关注
[0086]针对欠驱动系统,本发明提出一种基于自适应动态规划的分层滑膜优化控制方法。通过引入包含全部状态信息的分层滑膜面,使得系统鲁棒性被提高。利用Bellman最优性原理构造基于分层滑膜面的代价函数和HJB方程并通过梯度下降法求解非显式优化控制策略。采用神经网络对非显式优化控制策略进行有效近似,构建执行网络对控制策略进行调谐并构建评价网络对每一次调谐后的控制策略进行评估后再反馈给执行网络以实现对欠驱动系统的优化控制目标。
Smart Images

Figure QLYQS_1 
Figure QLYQS_2 
Figure QLYQS_3
Abstract
Description
Technical Field
[0001] This invention belongs to the field of underactuated system optimization control technology, specifically relating to a hierarchical sliding film optimization control method for underactuated systems based on adaptive dynamic programming. Background Technology
[0002] Underactuated systems are a typical type of unstable dynamic system. Due to their flexible operating space and characteristics such as underactuation, nonlinearity, and open-loop instability, they have gradually become a hot research topic in the field of control and have been widely used in aerospace and robotic control systems. In recent years, controllers designed based on various classical control theories and modern control theory methods have enabled the dynamic characteristics of underactuated systems to be well reflected.
[0003] As a special type of nonlinear control method, sliding membrane variable structure control has been widely studied in recent years. Its unique feature lies in the fact that the structure of the sliding surface is not fixed during construction, but rather can force the controlled system to move according to a predetermined sliding mode, and in this process, it is insensitive to parameter perturbations and external disturbances. Therefore, sliding membrane variable structure control possesses good robustness and other characteristics. However, with the continuous development of industrial processes, traditional single-layer sliding membrane variable structure control can no longer meet practical requirements, thus, layered sliding membrane control has attracted much attention in recent years. Each sub-sliding surface of the layered sliding surface is constructed according to the state of each subsystem, and the top layer sliding surface is composed of all the sub-sliding surfaces, meaning that the top layer sliding surface contains all the system state information. For underactuated systems, each sub-sliding surface can appropriately correspond to each subspace of the underactuated system; therefore, layered sliding membrane control can effectively solve control problems related to underactuated systems.
[0004] Adaptive dynamic programming (ADMP) provides an effective technique for handling optimal decision-making and control problems under conditions of uncertainty and stochasticity, without linear assumptions. It employs approximate structures, such as neural networks, to generate online approximations of optimal solutions through recursive numerical methods. Specifically, ADMP combines reinforcement learning with traditional dynamic programming to approximate solutions to the Hamilton-Jacobi-Bellman (HJB) equations. For linear time-varying cases, the HJB equations can evolve into algebraic Riccati equations, and the well-known iterative solution strategy is proposed by transforming the algebraic Riccati equations into a series of linear Lyapunov equations. Along this direction, the iterative solution strategy of ADMP has been gradually extended and applied to nonlinear time-varying cases. Summary of the Invention
[0005] To address the optimization control problem of underactuated systems, this invention proposes a layered sliding surface optimization control method for underactuated systems based on adaptive dynamic programming. By introducing layered sliding surfaces, the system response rate is improved. To enhance the dynamic performance of the underactuated system, an optimized controller design is employed, using an adaptive dynamic programming method to ensure the optimal performance of the underactuated system.
[0006] A hierarchical sliding film optimization control method for underactuated systems based on adaptive dynamic programming specifically includes:
[0007] S1. Establish a dynamic model of the underactuated system in nominal form;
[0008] Define an augmented state variable X = [x1, x2, x3, x4] for an underactuated system. T Let the state components be x1, x2, x3, and x4, then the dynamic model of the underactuated system can be formulated as follows:
[0009]
[0010] Where p1(X) and p2(X) are system dynamic functions; q1(X) and q2(X) are control gain functions; and u is the control input.
[0011] S2. Based on the dynamic model of the underactuated system, construct a layered sliding surface containing state information;
[0012] S21. The sub-slip surface vector s is constructed as follows:
[0013] s = [s1, 0, s2, 0] T =HX o +MX e (2)
[0014] Where s i =h i x 2i-1 +x 2i Let i = 1, 2, representing the i-th sub-slip surface; H = diag(h1, 1, h2, 1) is a diagonal matrix, where h1 and h2 are elements on the diagonal; vector X o and X e Defined as X o =[x1,0,x3,0] T and X e =[0,x2,0,x4] T ; Matrix M is:
[0015]
[0016] S22. Based on equation (2), the layered synovial surface vector S is constructed as follows:
[0017] S = [S1, 0, S2, 0] T =s+Z S (3)
[0018] Where S i =s i +z i-1 S i-1 , i = 1, 2, representing the i-th layer of the sliding membrane surface; Z = diag(z0, 1, z1, 1) is a diagonal matrix, where z0 and z1 are elements on the diagonal; vector S Defined as S =[S0,0,S1,0] T S0 and S1 are both vectors S Element;
[0019] S23, based on equations (2) and (3) and obtained by iterative calculation, the i-th layer of the layered slip surface S is obtained. i As shown below:
[0020]
[0021] in, like Then the parameter otherwise,
[0022] S24. Based on equations (2), (3), and (4), the layered synovial surface vector S is re-expressed as follows:
[0023]
[0024] Among them, matrix for:
[0025]
[0026] S25. Taking the derivative of equation (5) yields the following result:
[0027]
[0028] Where, vector P(X) and U(t) are respectively P(X) = [0, p1(X), 0, p2(X)] T and U=[0,0,0,u] T The vector Q(X) is:
[0029]
[0030] S26. Based on the hierarchical structure, the control input u is designed as follows:
[0031]
[0032] Where i ≤ i; u eq,i and u sw,i These represent the equivalent control law and the switching control law for the i-th layer of the sliding surface, respectively.
[0033] S27. Based on equation (7), the vector U is re-formulated as:
[0034] U = Ξ(u eq +u sw (8)
[0035] Where, vector u eq and u sw They are respectively and Matrix Ξ is:
[0036]
[0037] S3. Based on step S2, construct the cost function and HJB equation based on the layered sliding surface and solve the optimization control strategy by gradient descent method;
[0038] S31. Using the Bellman optimality principle, construct a cost function based on a layered sliding surface. as follows:
[0039]
[0040] Where A and B are both positive definite matrices; α is a positive constant; ||X|| represents the norm of the augmented state variable X; F(u sw Let be a positive definite function and be defined as:
[0041]
[0042] Where β is a positive constant; ω i Represents the integral variable; It is a monotonically odd function;
[0043] S32. Based on equation (9), the Hamiltonian function based on the layered sliding surface is constructed as follows:
[0044]
[0045] in, for The gradient; represent Transpose of;
[0046] S33, The HJB equation of constructive formula (11) is shown below:
[0047]
[0048] in, The cost function for optimization; yes The gradient; represent Transpose of; and These are the optimized equivalent control law vector and the switching control law vector, respectively.
[0049] S34, based on equations (11) and (12), and by solving partial differential equations and have to:
[0050]
[0051]
[0052] S4. Construct an evaluation network and an execution network to evaluate the optimization control strategy for each execution in order to achieve the optimization control objective for the underactuated system;
[0053] S41. By using the general approximation performance of radial basis function neural networks, the optimized cost function is... It is approximated to the following form:
[0054]
[0055] Among them, W * For bounded weights, (W) * ) T For W * Transpose of; Represents the activation function; To approximate the error;
[0056] S42. Based on equation (15), the optimized cost function gradient Formulated as:
[0057]
[0058] in, for The gradient; represent Transpose of; for The gradient;
[0059] S43, Based on equation (16), and It is rephrased as follows:
[0060]
[0061]
[0062] S44, due to the bounded weight W * It is unknown, therefore, The following estimates are made:
[0063]
[0064] in, for The estimate; For W * The estimate; for Transpose of;
[0065] S45, based on equation (19), The gradient is formulated as:
[0066]
[0067] in, for The gradient;
[0068] S46, Based on formula (20), and The estimation form is:
[0069]
[0070]
[0071] in, and They are respectively and The estimate; To implement network weights;
[0072] S47. Based on equations (21) and (22), the estimated form of the HJB equation is:
[0073]
[0074] in, To evaluate the weights of the network;
[0075] S48. Based on equations (12) and (23), the Bellman residual is defined as follows:
[0076]
[0077] S49. To ensure the minimum Bellman residual, the objective function e is defined as follows:
[0078]
[0079] S50. Based on equation (25) and using the standardization principle and gradient descent method, the evaluation network update law is designed as follows:
[0080]
[0081] Where, ξ c To evaluate network update rate; represent Transpose of;
[0082] S51. To ensure system stability, the network update law is designed as follows:
[0083]
[0084] Where ε and δ are both positive constants; J T Represents the transpose of J; γ a To implement network update rate.
[0085] Beneficial technical effects of the present invention:
[0086] For underactuated systems, this invention proposes a hierarchical sliding surface optimization control method based on adaptive dynamic programming. By introducing a hierarchical sliding surface containing all state information, the robustness of the system is improved. A cost function and HJB equation based on the hierarchical sliding surface are constructed using the Bellman optimality principle, and the non-explicit optimization control strategy is solved using the gradient descent method. A neural network is used to effectively approximate the non-explicit optimization control strategy. An execution network is constructed to tune the control strategy, and an evaluation network is constructed to evaluate the control strategy after each tuning before feeding it back to the execution network to achieve the optimization control objective for the underactuated system. Attached Figure Description
[0087] Figure 1 This is a flowchart of the hierarchical sliding film optimization control method for underactuated systems based on adaptive dynamic programming, as described in this invention.
[0088] Figure 2 This is a simulation response curve of the state variables in an embodiment of the present invention;
[0089] Figure 3 This is a simulation curve of the layered sliding membrane variable response in an embodiment of the present invention;
[0090] Figure 4This is a simulation graph of the network weight update performed in an embodiment of the present invention.
[0091] Figure 5 This is a simulation diagram of the evaluation network weight update in an embodiment of the present invention. Detailed Implementation
[0092] The present invention will be further described below with reference to the accompanying drawings and embodiments;
[0093] A hierarchical sliding film optimization control method for underactuated systems based on adaptive dynamic programming is shown in the appendix. Figure 1 As shown, it specifically includes:
[0094] S1. Establish a dynamic model of the underactuated system in nominal form;
[0095] Define an augmented state variable X = [x1, x2, x3, x4] for an underactuated system. T Let the state components be x1, x2, x3, and x4, then the dynamic model of the underactuated system can be formulated as follows:
[0096]
[0097] Where p1(X) and p2(X) are system dynamic functions; q1(X) and q2(X) are control gain functions; and u is the control input.
[0098] S2. Based on the dynamic model of the underactuated system, construct a layered sliding surface containing state information;
[0099] S21. The sub-slip surface vector s is constructed as follows:
[0100] s = [s1, 0, s2, 0] T =HX o +MX e (2)
[0101] Where s i =h i x 2i-1 +x 2i Let i = 1, 2, representing the i-th sub-slip surface; H = diag(h1, 1, h2, 1) is a diagonal matrix, where h1 and h2 are elements on the diagonal; vector X o and X e Defined as X o =[x1,0,x3,0] T and X e =[0,x2,0,x4] T ; Matrix M is:
[0102]
[0103] S22. Based on equation (2), the layered synovial surface vector S is constructed as follows:
[0104] S = [S1, 0, S2, 0] T =s+Z S (3)
[0105] Where S i =s i +z i-1 S i-1 , i = 1, 2, representing the i-th layer of the sliding membrane surface; Z = diag(z0, 1, z1, 1) is a diagonal matrix, where z0 and z1 are elements on the diagonal; vector S Defined as S =[S0,0,S1,0] T S0 and S1 are both vectors S Element;
[0106] S23, based on equations (2) and (3) and obtained by iterative calculation, the i-th layer of the layered slip surface S is obtained. i As shown below:
[0107]
[0108] in, like Then the parameter otherwise,
[0109] S24. Based on equations (2), (3), and (4), the layered synovial surface vector S is re-expressed as follows:
[0110]
[0111] Among them, matrix for:
[0112]
[0113] S25. Taking the derivative of equation (5) yields the following result:
[0114]
[0115] Where, vector P(X) and U(t) are respectively P(X) = [0, p1(X), 0, p2(X)] T and U=[0,0,0,u] T The vector Q(X) is:
[0116]
[0117] S26. Based on the hierarchical structure, the control input u is designed as follows:
[0118]
[0119] Where i ≤ i; u eq,i and u sw,i These represent the equivalent control law and the switching control law for the i-th layer of the sliding surface, respectively.
[0120] S27. Based on equation (7), the vector U is re-formulated as:
[0121] U = Ξ(u eq +u sw (8)
[0122] Where, vector u eq and u sw They are respectively and Matrix Ξ is:
[0123]
[0124] S3. Based on step S2, construct the cost function and HJB equation based on the layered sliding surface and solve the optimization control strategy by gradient descent method;
[0125] S31. Using the Bellman optimality principle, construct a cost function based on a layered sliding surface. as follows:
[0126]
[0127] Where A and B are both positive definite matrices; α is a positive constant; ||X|| represents the norm of the augmented state variable X; F(u sw Let be a positive definite function and be defined as:
[0128]
[0129] Where β is a positive constant; ω i Represents the integral variable; It is a monotonically odd function;
[0130] S32. Based on equation (9), the Hamiltonian function based on the layered sliding surface is constructed as follows:
[0131]
[0132] in, for The gradient; represent Transpose of;
[0133] S33, The HJB equation of constructive formula (11) is shown below:
[0134]
[0135] in, The cost function for optimization; yes The gradient; represent Transpose of; and These are the optimized equivalent control law vector and the switching control law vector, respectively.
[0136] S34, based on equations (11) and (12), and by solving partial differential equations and have to:
[0137]
[0138]
[0139] S4. Construct an evaluation network and an execution network to evaluate the optimization control strategy for each execution in order to achieve the optimization control objective for the underactuated system;
[0140] S41. By using the general approximation performance of radial basis function neural networks, the optimized cost function is... It is approximated to the following form:
[0141]
[0142] Among them, W * For bounded weights, (W) * ) T For W * Transpose of; Represents the activation function; To approximate the error;
[0143] S42. Based on equation (15), the optimized cost function gradient Formulated as:
[0144]
[0145] in, for The gradient; represent Transpose of; for The gradient;
[0146] S43, Based on equation (16), and It is rephrased as follows:
[0147]
[0148]
[0149] S44, due to the bounded weight W * It is unknown, therefore, The following estimates are made:
[0150]
[0151] in, for The estimate; For W * The estimate; for Transpose of;
[0152] S45, based on equation (19), The gradient is formulated as:
[0153]
[0154] in, for The gradient;
[0155] S46, Based on formula (20), and The estimation form is:
[0156]
[0157]
[0158] in, and They are respectively and The estimate; To implement network weights;
[0159] S47. Based on equations (21) and (22), the estimated form of the HJB equation is:
[0160]
[0161] in, To evaluate the weights of the network;
[0162] S48. Based on equations (12) and (23), the Bellman residual is defined as follows:
[0163]
[0164] S49. To ensure the minimum Bellman residual, the objective function e is defined as follows:
[0165]
[0166] S50. Based on equation (25) and using the standardization principle and gradient descent method, the evaluation network update law is designed as follows:
[0167]
[0168] Where, ξ c To evaluate network update rate; represent Transpose of;
[0169] S51. To ensure system stability, the network update law is designed as follows:
[0170]
[0171] Where ε and δ are both positive constants; J T Represents the transpose of J; γ a To implement network update rate.
[0172] To verify the effectiveness of the control method of this invention, numerical simulation was performed on an underactuated system. The initial values of the state variables were given as X(0) = [x1(0), x2(0), x3(0), x4(0)]. T =[0.03,1,0.002,0.24] T The values of parameters μ1 and μ2 are set to μ1 = 2.1 and μ2 = 4.9. Cost function The parameters were set to α = 0.3 and β = 2. The correlation matrix is given as follows:
[0173]
[0174] The weights of the execution network and the evaluation network are defined as follows: and Its initial values were respectively and Activation function Selected as a Gaussian type function and Its center and width are c, respectively. NN,l ={-2,-1,0,1,2} and w NN,l={1,2,1,2,1}. The parameters of the weight update law for the execution network and the evaluation network are given as ξ. c =0.95, γ a =1, ε=0.6, δ=1.5.
[0175] Simulation results are attached. Figure 2-5 As shown. Figure 2 The response curves for system states x1, x2, x3, and x4 show that the underactuated system reaches a stable state when the system states gradually converge and reach an equilibrium point. Figure 3 The response curves of the layered synovial variables S1 and S2 show that the trajectories of S1 and S2 initially oscillate significantly, then stabilize and gradually converge to the equilibrium point. Figure 4 and Figure 5 The figures shown are the weight update diagrams of the execution network and the evaluation network in the embodiment. It can be seen that the weights of the execution network and the evaluation network converge and reach a stable state after training in a short period of time.
Claims
1. A hierarchical sliding film optimization control method for underactuated systems based on adaptive dynamic programming, characterized in that, Specifically, it includes: S1. Establish a dynamic model of the underactuated system in nominal form; S2. Based on the dynamic model of the underactuated system, construct a layered sliding surface containing state information; S3. Based on step S2, construct the cost function and HJB equation based on the layered sliding surface and solve the optimization control strategy by gradient descent method; Step S3 is as follows: S31. Using the Bellman optimality principle, construct a cost function based on a layered sliding surface. ; S32. Based on equation (9), the Hamiltonian function based on the layered sliding surface is constructed as follows: (11) in, Represents the layered synovial surface vector. for The gradient; represent The transpose of; where, and All are positive definite matrices; It is a positive constant; Represents augmented state variables The norm; It is a positive definite function; S33, HJB equation of construction formula (11); S34, based on equations (11) and (12), and by solving partial differential equations and have to: (13) (14) in, The cost function representing optimization gradient, This represents the optimized equivalent control law vector. This represents the optimized switching control law vector. It is a positive constant; S4. Construct an evaluation network and an execution network to evaluate the optimization control strategy for each execution in order to achieve the optimization control objective for the underactuated system.
2. The hierarchical sliding film optimization control method for underactuated systems based on adaptive dynamic programming according to claim 1, characterized in that, Step S1 is as follows: Define an augmented state variable for an underactuated system And let the state components be respectively , , and The dynamic model of the underactuated system can then be formulated as follows: (1) in, and For system dynamic functions; and For control gain function; For controlling input.
3. The hierarchical sliding film optimization control method for underactuated systems based on adaptive dynamic programming according to claim 1, characterized in that, Step S2 is as follows: S21, Constructing the subsliposome surface vector as follows: (2) in , , indicating the first Layered slip film surface; It is a diagonal matrix. and All are elements on the diagonal; vector and Defined respectively and ;matrix for: S22. Based on equation (2), construct the layered sliding surface vector. as follows: (3) in , , indicating the first Layered slip film surface; It is a diagonal matrix. and All are elements on the diagonal; vector Defined as , and Both are vectors Element; S23, based on equations (2) and (3) and obtained through iterative calculation, is the first... Layered slip film surface As shown below: (4) in, , ;like Then the parameter ;otherwise, ; S24. Based on equations (2), (3), and (4), the layered sliding surface vector It is rephrased as follows: (5) Among them, matrix for: S25. Taking the derivative of equation (5) yields the following result: (6) Where, vector , and They are respectively , and ;matrix for: ; S26. Based on a hierarchical structure, control input Designed as follows: (7) in, ; and Representing the first Equivalent control law and switching control law for layered sliding surfaces; S27, Based on equation (7), vector The formula has been reformulated as follows: (8) Where, vector and They are respectively and ;matrix for: 。 4. The hierarchical sliding film optimization control method for underactuated systems based on adaptive dynamic programming according to claim 1, characterized in that, The cost function based on the layered sliding surface Specifically: (9) in, and All are positive definite matrices; It is a positive constant; Represents augmented state variables The norm; It is a positive definite function and is defined as: (10) in, It is a positive constant; Represents the integral variable; It is a monotonically odd function.
5. The hierarchical sliding film optimization control method for underactuated systems based on adaptive dynamic programming according to claim 1, characterized in that, The HJB equation is specifically as follows: (12) in, The cost function for optimization; yes The gradient; represent transpose; and These are the optimized equivalent control law vector and the switching control law vector, respectively.
6. The hierarchical sliding film optimization control method for underactuated systems based on adaptive dynamic programming according to claim 1, characterized in that, Step S4 is as follows: S4-1. By using the general approximation performance of radial basis function neural networks, the cost function is optimized. It is approximated to the following form: (15) in, For bounded weights, for transpose; Represents the activation function; To approximate the error; S4-2, Based on equation (15), the optimized cost function gradient Formulated as: (16) in, for The gradient; represent transpose; for The gradient; S4-3, Based on equation (16), and It is rephrased as follows: (17) (18) S4-4, Due to bounded weights It is unknown, therefore, The following estimates are made: (19) in, for The estimate; for The estimate; for transpose; S4-5, Based on equation (19), The gradient is formulated as: (20) in, for The gradient; S4-6, Based on equation (20), and The estimation form is: (21) (22) in, and They are respectively and The estimate; To implement network weights; S4-7. Based on equations (21) and (22), the estimated form of the HJB equation is: (23) in, ; To evaluate the weights of the network; S4-8. Based on equations (12) and (23), the Bellman residual is defined as follows: (24) S4-9. To ensure the minimum Bellman residual, define the objective function. As shown below: (25) S5-10. Based on equation (25) and using the standardization principle and gradient descent method, the evaluation network update law is designed as follows: (26) in, To evaluate network update rate; represent transpose; S5-11. To ensure system stability, the network update law is designed as follows: (27) in, and All are positive constants; ; represent transpose; To implement network update rate.
Citation Information
Patent Citations
Underactuated mechanical arm layering sliding mode control method based on fuzzy optimization
CN108972560A
Self-adaptive optimal tracking control method for linear system
CN112445131A