A self-triggered adaptive dynamic programming method based on model prediction

CN120578053BActive Publication Date: 2026-09-01BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510452909.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2026-09-01
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

例如,模型预测控制虽然在处理未知非线性系统的优化问题方面有显著优势,但是在线计算需求量大和解决非线性MPC问题的复杂性阻碍了其发展

Benefits of technology

[0114]与现有技术相比,本发明存在以下技术效果:针对未知非线性离散时间系统,开发了一种基于迭代模型预测过程的自触发自适应动态规划(MSTADP)方法。将自触发机制与MPC融合,通过动态触发间隔减少预测时域内的优化问题求解次数,降低在线计算负载。然后提出一种估计批评结构来实现基于迭代模型预测过程的自触发自适应动态规划算法。最后,提供仿真结果来验证所提出算法的有效性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578053B_ABST
    Figure CN120578053B_ABST
Patent Text Reader

Abstract

This invention discloses a self-triggering adaptive dynamic programming method based on model prediction, belonging to the field of optimal control. It includes: developing a self-triggering adaptive dynamic programming algorithm based on an iterative model prediction process for unknown nonlinear systems; designing trigger condition functions through Lyapunov stability analysis to achieve intermittent control that depends only on the previous trigger state, reducing communication overhead; designing a three-network architecture consisting of a model network, an evaluation network, and an execution network; proposing a rolling optimization mechanism to decompose infinite-time domain optimization into finite-time domain subproblems, reducing computational time; and finally, providing simulation results to verify the effectiveness of the proposed method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optimal control, and in particular to a self-triggered adaptive dynamic programming method based on model prediction. Background Technology

[0002] Optimal control of systems is a key research direction in modern control theory, playing a vital role in various practical applications such as autonomous vehicles, industrial production processes, and aircraft control. Optimizing control strategies can significantly improve system efficiency, reduce costs, and enhance system stability and robustness. Adaptive dynamic programming (ADP) is an important method for solving optimization and control problems in such complex systems, and various ADP methods have been developed in recent years to achieve optimal control. However, due to limitations in computational power, finding feasible solutions to the Hamilton-Jacobi-Bellman (HJB) equations remains a significant challenge.

[0003] To overcome the challenges of optimal control, researchers have proposed various solutions. Model predictive control (MPC) has shown significant advantages in handling optimization problems of unknown nonlinear systems. MPC establishes predictive models to solve practical industrial problems by implementing a rolling time (RH) mechanism. Researchers have proposed a functional MPC method using the ADP algorithm, which can handle constraints and disturbances to achieve optimal control. Furthermore, the self-triggering mechanism (STM) method can significantly reduce computational load, communication overhead, and energy consumption, while adapting to the characteristics of dynamic nonlinear systems. These methods enhance the robustness and practicality of controllers to varying degrees.

[0004] Despite progress in existing research, several unresolved challenges remain in the field of optimal control. For example, while model predictive control (MMC) offers significant advantages in handling optimization problems of unknown nonlinear systems, its development is hampered by the high computational demands and complexity of solving nonlinear MPC problems. Although event-triggered mechanisms and MMC have received theoretical attention, effectively integrating these mechanisms in practical applications to design more robust and adaptive control strategies remains an open question. Therefore, research in optimal control needs to continue exploring more effective methods to address these challenges. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and solve the optimal control problem of a class of nonlinear systems with terminal state constraints.

[0006] To achieve the above objectives, this invention proposes a model-predictive self-triggering adaptive dynamic programming method (MSTADP), comprising the following steps:

[0007] S1. Construct a general nonlinear system to establish an accurate mathematical model for subsequent control algorithm design and describe the dynamic characteristics of the controlled object.

[0008] First, consider the nonlinear system as follows:

[0009] x k+1 =f(x) k )+g(x k )u k

[0010] In the formula, k is the time step index, indicating that the system is at the k-th step, x k ∈R n Let u be the state of the system at step k. k ∈R m f(x) is the control input of the system at step k. k Let g(x) be the state transition function of the system. k Let f(x) be the control input matrix of the system. Assume x = 0 is the equilibrium point of the system. Furthermore, f(x) k )+g(x k )u(x k In R n ×R m →R n The expression is continuously differentiable, and f and g are considered unknown. Define u(x) as... k ) is a feedback control law.

[0011] S2. This invention designs a novel self-triggering mechanism for real-time optimization. This mechanism can reduce the frequency of online computation and reduce resource consumption through dynamic triggering of control updates. Traditional event-triggered mechanisms may lead to honor calculations, while the self-triggering mechanism only updates control when the state deviates from a threshold, reducing the computational burden. This invention combines the rolling optimization characteristics of MPC to balance real-time performance and optimality within a finite time domain. The main method is as follows:

[0012] First, define a strictly monotonically increasing sequence. This serves as the self-sampling time. The controller only triggers and updates at the self-sampling time. Next, the sampling interval is defined as:

[0013]

[0014] in, This represents the system's self-sampled state. Since the controller only updates at the self-sampled time, we will use the feedback control law u(x) k ) is defined as v(x) k The cost function is defined as follows:

[0015]

[0016] Where p is the time step index, indicating that the system is at step p, and x p ξ represents the state of the system at step p. p V(x) represents the sampling interval of the system at step p. p +ξ p ) indicates that the system at x p +ξ p Control law and utility function in the state Here, Q is a positive semi-definite matrix, and R is a positive definite symmetric matrix. Assumption 1: The system is controllable, and its state is observable. Furthermore, u represents any admissible control strategy of the system.

[0017] Combining model predictive control mechanisms, let H P H represents the prediction time domain. C Indicates the control time domain, and H C ≤H P To reduce computational requirements, parameter N is set to make H... C =H P =N, here This is the terminal state. Therefore, J(x) k This can be further expressed as:

[0018]

[0019] in, This represents the terminal constraint within the finite prediction time domain. It can be expressed as:

[0020]

[0021] in, For the system in The transpose of the state matrix at time step P is a positive definite weight matrix. According to Bellman's optimality theorem, the optimal cost function J... * (x k )satisfy:

[0022]

[0023]

[0024] in, Indicates that the system is in The utility function of time, J * (x k+1 ) for the system in x k+1 The optimal cost function at that time, This is the transpose of the system state. For the system in Time control law, for transpose,

[0025] To determine the optimal strategy that minimizes the cost function and ensures system stability Its expression is:

[0026]

[0027] Note 1: The self-triggering mechanism STM designed in this invention can effectively reduce the computational burden, but it also leads to a decrease in system performance. Furthermore, since the design of STM includes the current sampling time k... j Offline algorithms cannot achieve real-time control of the system. Therefore, we introduce the MPC mechanism to accelerate system convergence. This paper proposes a self-triggered adaptive dynamic programming method based on MPC, which essentially incorporates STM into the MPC control process to reduce the number of predictions, thereby reducing the online computational burden while maintaining system performance.

[0028] S3. Design a self-triggering mechanism (STM) to calculate the trigger interval in real time and calculate the trigger condition threshold to ensure system stability within the trigger interval. Assumption 2: Assume that for all x... k ∈Ω k Optimal control strategy u * (x k If ) is Lipschitz continuous, that is, there exists a Lipschitz constant Υ that satisfies:

[0029]

[0030] Theorem 1: If both Assumption 1 and Assumption 2 are true, then the trigger threshold can be derived:

[0031]

[0032] Where η∈(0,1), κ is a constant greater than 0, and λ min (·) and λ max (·) represent the minimum and maximum eigenvalues ​​of the matrix, respectively. When Ξ(ξ) k When )≤0, the system is always eventually bounded (UUB).

[0033] Proof: Consider the change in the Lyapunov function:

[0034]

[0035]

[0036] in, The squared term representing the upper bound of the model uncertainty.

[0037] Further decompose the quadratic terms:

[0038]

[0039] Combining the two equations above, we get:

[0040]

[0041] When Ξ(ξ) k When )≤0, in order to ensure ΔJ * (x k If )≤0, the following conditions must be met:

[0042]

[0043] That is, the sampling state must satisfy:

[0044]

[0045] Assumption 3: For a closed-loop system, there exists a parameter σ > 0.5 such that the sampling error ξ k The sampled state satisfies the inequality:

[0046] ||x k+1 || = σ||x k ‖+σ‖ξ k ||

[0047] Lemma 1: If the inequality in Assumption 3 holds, then for a nonlinear system, the measurement error ξ k With sampling state satisfy:

[0048]

[0049] Theorem 3: If Assumptions 2 and 3 hold, then the next triggering time k j+1 It can be directly derived from the current sampling state. It is determined and satisfies the stability condition. Specifically, it takes the following form:

[0050]

[0051] in It can be obtained from the following formula:

[0052]

[0053] Proof: According to Theorem 1 and Theorem 3, we have:

[0054]

[0055] To simplify the calculation, let the right side of the expression equal to Right now We will discuss two scenarios:

[0056] Case 1: When k = kj +1 and At that time, there were:

[0057]

[0058] Since 0 < η < 1 and σ > 0.5, we can obtain therefore:

[0059]

[0060] Depend on achievable therefore The trigger interval is:

[0061]

[0062] in, This indicates rounding down. Therefore, in this case:

[0063]

[0064] Therefore, Δk j =1, the trigger condition is met.

[0065] Case 2: When k = k j +1 and At that time, there were:

[0066]

[0067] because The expression increases monotonically with increasing k; therefore, it is assumed that ε≥k exists. j +1 makes:

[0068]

[0069] achievable because The next trigger interval k can be solved. j =1+(ε), where

[0070] In summary, the self-triggered mechanism can be designed as follows:

[0071]

[0072] Q.E.D.

[0073] This step mainly analyzes the system energy change based on the Lyapunov function, combines the Lipschitz continuity assumption, quantifies the error limit of the control strategy, and then designs the trigger threshold.

[0074] S4. Neural networks possess powerful fitting capabilities. This step obtains the inherent dynamic characteristics of the controlled system by establishing a data-driven model. Let the neural network fitting the nonlinear system be a model neural network. In this framework, the system state x... k and control input u k The measurement is performed within the current prediction time domain, and then the neural network model can infer the state x at the next time step. k+1 For discrete-time systems, The state needs to be estimated by training a neural network model before performing the learning process. The neural network fits the state expression as follows:

[0075]

[0076] in, Let be the input vector, and θ(·) be the bounded activation function, which is tanh(·) in this invention. m1 and w m2 The weight matrix is ​​randomly initialized. The recognition error of the neural network model is defined as... Therefore, the performance metrics of the model neural network are as follows:

[0077]

[0078] Note 2: After the neural network model is trained, its weights w m1 and w m2 It will be fixed and used in the online prediction phase.

[0079] S5. Construct a comment neural network. In the current prediction time domain, the terminal penalty is fitted using a comment neural network, which improves the system performance. The cost function is calculated as follows:

[0080]

[0081] in, and Let l1 be the target weight matrix, l1 be the number of hidden layer neurons, and τ be the weight matrix. ck This represents the bounded estimation error. The fitted cost function is expressed as:

[0082]

[0083] Among them, w c1 and w c2 It is randomly initialized. Combined with the MPC mechanism, the comment neural network needs to fit two objectives simultaneously: the cost function J(x) in the time domain [k, k+N-1]. k Terminal status Penalty items The specific implementation is as follows:

[0084]

[0085] According to the Bellman equation, the ideal cost function should satisfy:

[0086]

[0087] in, Represents the estimated control law, This represents the terminal state of the neural network fit. A relaxation factor is introduced. To balance the convergence rate, the error function is rewritten as:

[0088]

[0089] Wherein, Δθ(w c1 x k )=θ(w c1 x k+1 )-θ(w c1 x k ), Δτ ck =τ ck+1 -τ ck The total residual error is:

[0090]

[0091] in, Δτ c =Δτ ck +Δτ ck+N For simplification, let Δθ c =Δθ(w c1 x k )+Δθ(w c1 x k+N Therefore, the performance metrics are:

[0092]

[0093] The weight update rules are as follows:

[0094]

[0095] Where, α c ∈(0,1) is the learning rate.

[0096] S6. Design an action neural network to fit the optimal control strategy. The target strategy is:

[0097]

[0098] in, and l1 represents the target weight, and l2 represents the number of neurons in the hidden layer. This represents the bounded error after fitting the neural network. The estimated value is:

[0099]

[0100] Where, ω a1 and ω a2 This is the corresponding weight matrix. From this formula, we can obtain:

[0101]

[0102] Similarly, the corresponding control input error is:

[0103]

[0104] in, make For τ a Then the objective error function is:

[0105]

[0106] The weight update rules are as follows:

[0107]

[0108] Where, α a ∈(0,1) represents the learning rate. This invention uses a designed action neural network to fit the HJB equation. It is worth noting that using a neural network to fit the HJB equation to determine the control law differs from the method of determining the control law based on MPC.

[0109] S7. Train the neural network model and update the neural network weights using gradient descent:

[0110]

[0111] Where, α m ∈(0,1) is the learning rate.

[0112] S8. Perform online predictions for the system.

[0113] Furthermore, it was proved that the characteristics of nonlinear systems under optimal control and the weight error of neural networks are eventually uniformly bounded under certain given conditions.

[0114] Compared with existing technologies, this invention offers the following technical advantages: For unknown nonlinear discrete-time systems, a self-triggered adaptive dynamic programming (MSTADP) method based on an iterative model prediction process is developed. The self-triggered mechanism is integrated with MPC (Multi-Process Control), reducing the number of optimization problems solved in the prediction time domain through dynamic triggering intervals, thus lowering the online computational load. Then, an estimation critique structure is proposed to implement the self-triggered adaptive dynamic programming algorithm based on the iterative model prediction process. Finally, simulation results are provided to verify the effectiveness of the proposed algorithm. Attached Figure Description

[0115] Figure 1 This is a flowchart of a self-triggered adaptive dynamic programming method based on model prediction according to the present invention.

[0116] Figure 2 Example 1 of the present invention is a spring-mass-damper system.

[0117] Figure 3 This refers to the training error of the neural network model in Example 1 of this invention.

[0118] Figure 4 This is the system state under the MSTADP and HDP algorithms in Example 1 of this invention.

[0119] Figure 5 This is the control input generated by the MSTADP and HDP algorithms in Example 1 of this invention.

[0120] Figure 6 Example 1 of this invention: Triggering time interval Δk j . Detailed Implementation

[0121] To further illustrate the features of the present invention, please refer to the following detailed description and appendix. Figure 1-6 The accompanying drawings are for reference and illustration only and are not intended to limit the scope of protection of this invention.

[0122] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings:

[0123] like Figure 1 As shown, this embodiment discloses a self-triggered adaptive dynamic programming method based on model prediction, including the following steps:

[0124] S1. The present invention will be implemented through, as follows: Figure 2 The spring-mass-damper system shown verifies the effectiveness of the MSTADP algorithm. The system is as follows:

[0125]

[0126] Where x1 and x2 represent state vectors, representing horizontal displacement and velocity, respectively. F(u) is the control input, representing the applied force, and M is the mass of the object. Furthermore, K is the spring stiffness constant, and P is the damping coefficient of the damper. Table 1 shows the relevant system parameters.

[0127] Table 1

[0128] I2 4I 0.5I2 5 10 0.7 0.6 0.98

[0129] The dynamic formula is as follows:

[0130]

[0131] S2. This invention designs a novel self-triggering mechanism for real-time optimization. This mechanism can reduce the frequency of online computation and reduce resource consumption through dynamic triggering of control updates. Traditional event-triggered mechanisms may lead to honor calculations, while the self-triggering mechanism only updates control when the state deviates from a threshold, reducing the computational burden. This invention combines the rolling optimization characteristics of MPC to balance real-time performance and optimality within a finite time domain. The main method is as follows:

[0132] First, define a strictly monotonically increasing sequence. As the self-sampling time. According to the MPC principle, the continuous-time system is discretized, and the sampling interval is set to 0.5s. The system is discretized as follows:

[0133]

[0134] The controller triggers and updates only at the self-sampling time. Next, the sampling interval is defined as:

[0135]

[0136] in, This represents the system's self-sampled state. Since the controller only updates at the self-sampled time, we will use the feedback control law u(x) k ) is defined as v(x) k The cost function is defined as follows:

[0137]

[0138] Where p is the time step index, indicating that the system is at step p, and x p ξ represents the state of the system at step p. p V(x) represents the sampling interval of the system at step p. p +ξ p ) indicates that the system at x p +ξ p Control law and utility function in the state Here, Q is a positive semi-definite matrix, and R is a positive definite symmetric matrix. Assumption 1: The system is controllable, and its state is observable. Furthermore, u represents any admissible control strategy of the system.

[0139] Combining model predictive control mechanisms, let H P H represents the prediction time domain. C Indicates the control time domain, and H C ≤H P To reduce computational requirements, parameter N is set to make H... C =H P =N, here This is the terminal state. Therefore, J(x) k This can be further expressed as:

[0140]

[0141] in, This represents the terminal constraint within the finite prediction time domain. It can be expressed as:

[0142]

[0143] in, For the system in The transpose of the state matrix at time step P is a positive definite weight matrix. According to Bellman's optimality theorem, the optimal cost function J... * (x k )satisfy:

[0144]

[0145] in, Indicates that the system is in The utility function of time, J * (x k+1 ) for the system in x k+1 The optimal cost function at that time, This is the transpose of the system state. For the system in Time control law, for transpose,

[0146] To determine the optimal strategy that minimizes the cost function and ensures system stability Its expression is:

[0147]

[0148] Note 1: The self-triggering mechanism STM designed in this invention can effectively reduce the computational burden, but it also leads to a decrease in system performance. Furthermore, since the design of STM includes the current sampling time k... jOffline algorithms cannot achieve real-time control of the system. Therefore, we introduce the MPC mechanism to accelerate system convergence. This paper proposes a self-triggered adaptive dynamic programming method based on MPC, which essentially incorporates STM into the MPC control process to reduce the number of predictions, thereby reducing the online computational burden while maintaining system performance.

[0149] S3. Design a self-triggering mechanism (STM) to calculate the trigger interval in real time and calculate the trigger condition threshold to ensure system stability within the trigger interval. Assumption 2: Assume that for all x... k ∈Ω k Optimal control strategy u * (x k If ) is Lipschitz continuous, that is, there exists a Lipschitz constant Υ that satisfies:

[0150]

[0151] Theorem 1: If both Assumption 1 and Assumption 2 are true, then the trigger threshold can be derived:

[0152]

[0153] Where η∈(0,1), κ is a constant greater than 0, and λ min (·) and λ max (·) represent the minimum and maximum eigenvalues ​​of the matrix, respectively. When Ξ(ξ) k When )≤0, the system is always eventually bounded (UUB).

[0154] Proof: Consider the change in the Lyapunov function:

[0155]

[0156] in, The squared term representing the upper bound of the model uncertainty.

[0157] Further decompose the quadratic terms:

[0158]

[0159] Combining the two equations above, we get:

[0160]

[0161] When Ξ(ξ) k When )≤0, in order to ensure ΔJ * (x k If )≤0, the following conditions must be met:

[0162]

[0163] That is, the sampling state must satisfy:

[0164]

[0165] Assumption 3: For a closed-loop system, there exists a parameter σ > 0.5 such that the sampling error ξ k The sampled state satisfies the inequality:

[0166] ||x k+1 || = σ||x k ‖+σ‖ξ k ||

[0167] Lemma 1: If the inequality in Assumption 3 holds, then for a nonlinear system, the measurement error ξ k With sampling state satisfy:

[0168]

[0169] Theorem 3: If Assumptions 2 and 3 hold, then the next triggering time k j+1 It can be directly derived from the current sampling state. It is determined and satisfies the stability condition. Specifically, it takes the following form:

[0170]

[0171] in It is a dynamic trigger threshold function, which is obtained by the following formula:

[0172]

[0173] Proof: According to Theorem 1 and Theorem 3, we have:

[0174]

[0175] To simplify the calculation, let the right side of the expression equal to Right now We will discuss two scenarios:

[0176] Case 1: When k = k j +1 and At that time, there were:

[0177]

[0178] Since 0 < η < 1 and σ > 0.5, we can obtain therefore:

[0179]

[0180] Depend on achievable therefore The trigger interval is:

[0181]

[0182] in, This indicates rounding down. Therefore, in this case:

[0183]

[0184] Therefore, Δk j =1, the trigger condition is met.

[0185] Case 2: When k = k j +1 and At that time, there were:

[0186]

[0187] because The expression increases monotonically with increasing k; therefore, it is assumed that ε≥k exists. j +1 makes:

[0188]

[0189] achievable because The next trigger interval k can be solved. j =1+(ε), where

[0190] In summary, the self-triggered mechanism can be designed as follows:

[0191]

[0192] Q.E.D.

[0193] This step mainly analyzes the system energy change based on the Lyapunov function, combines the Lipschitz continuity assumption, quantifies the error limit of the control strategy, and then designs the trigger threshold.

[0194] S4. Neural networks possess powerful fitting capabilities. This step obtains the inherent dynamic characteristics of the controlled system by establishing a data-driven model. Let the neural network fitting the nonlinear system be a model neural network, with a 3-8-2 architecture. In this framework, the system state x... k and control input u k The measurement is performed within the current prediction time domain, and then the neural network model can infer the state x at the next time step. k+1 For discrete-time systems, The state needs to be estimated by training a neural network model before performing the learning process. The neural network fits the state expression as follows:

[0195]

[0196] in, Let be the input vector, and θ(·) be the bounded activation function, which is tanh(·) in this invention. m1 and w m2 The weight matrix is ​​randomly initialized. The recognition error of the neural network model is defined as... Therefore, the performance metrics of the model neural network are as follows:

[0197]

[0198] Note 2: After the neural network model is trained, its weights w m1 and w m2 It will be fixed and used in the online prediction phase.

[0199] S5. Construct a comment neural network with a 2-8-1 configuration. In the current prediction time domain, the terminal penalty is fitted using the comment neural network, which improves the system performance. The cost function is calculated as follows:

[0200]

[0201] in, and Let l1 be the target weight matrix, l1 be the number of hidden layer neurons, and τ be the weight matrix. ck This represents the bounded estimation error. The fitted cost function is expressed as:

[0202]

[0203] Among them, w c1 and w c2 It is randomly initialized. Combined with the MPC mechanism, the comment neural network needs to fit two objectives simultaneously: the cost function J(x) in the time domain [k, k+N-1]. k Terminal status Penalty items The specific implementation is as follows:

[0204]

[0205] According to the Bellman equation, the ideal cost function should satisfy:

[0206]

[0207] in, Represents the estimated control law, This represents the terminal state of the neural network fit. A relaxation factor is introduced. To balance the convergence rate, the error function is rewritten as:

[0208]

[0209] Wherein, Δθ(w c1 x k )=θ(w c1 x k+1 )-θ(w c1 x k ), Δτ ck =τ ck+1 -τ ck The total residual error is:

[0210]

[0211] in, Δτ c =Δτ ck +Δτ ck+N For simplification, let Δθ c =Δθ(w c1 x k )+Δθ(w c1 x k+N Therefore, the performance metrics are:

[0212]

[0213] The weight update rules are as follows:

[0214]

[0215] Where, α c =0.7 is the learning rate.

[0216] S6. Design an action neural network to fit the optimal control policy. The network architecture is configured as 2-8-1. The target policy is:

[0217]

[0218] in, and l1 represents the target weight, and l2 represents the number of neurons in the hidden layer. This represents the bounded error after fitting the neural network. The estimated value is:

[0219]

[0220] Where, ω a1 and ω a2 This is the corresponding weight matrix. From this formula, we can obtain:

[0221]

[0222] Similarly, the corresponding control input error is:

[0223]

[0224] in, make For τ a Then the objective error function is:

[0225]

[0226] The weight update rules are as follows:

[0227]

[0228] Where, α a =0.2 represents the learning rate. This invention uses an action neural network of this design to fit the HJB equation. It is worth noting that using a neural network to fit the HJB equation to determine the control law differs from the method of determining the control law based on MPC.

[0229] S7. Train the neural network model and update the neural network weights using gradient descent:

[0230]

[0231] Where, α m =0.1 is the learning rate.

[0232] For the location system, 500 datasets were randomly selected within the range [0.01, 0.01] for model network training, and another 500 datasets were used to test accuracy. Each dataset underwent 500 training iterations. To verify the accuracy of the neural network model's approximation system, the training error was set to a minimum. Figure 3 shown in 10 -3 The accuracy of the model network demonstrates its reliability and accuracy. After testing the model network, the final w is obtained. m1 and w m2 The parameter configuration is shown in Table 2. The initial state is set as x0 = [1.2, -1.2]. T .

[0233] Table 2

[0234] 1kg 5N / s 0.5 Ns / m 0.5s

[0235] S8. Perform online predictions for the system.

[0236] The effectiveness of the MSTADP algorithm is demonstrated by comparing it with the HDP algorithm. (Selection of relaxation factor) exist Figure 4In the middle, all system state trajectories are in Figure 5 All of them converge to zero under the influence of the control input shown. Figure 6 The trigger interval is shown. The control input was updated 41 times in 150 time steps. The results show that the STM method significantly reduces resource utilization. The observation results show that the convergence speed of all curves is significantly faster after adopting the MSTADP algorithm, indicating that the control effect of the MSTADP algorithm is better than that of the HDP algorithm.

[0237] Next, we will further analyze the convergence and stability of the neural network. Hypothesis 2: [The following is a separate, unrelated statement:] and These are used as target weights for the comment neural network and the action neural network, respectively. They satisfy... in, and It is a positive constant. Also, assume... β m ≤‖τ c ||≤θ cM γ m ≤‖τ a ||≤γ aM .

[0238] Theorem 2: The weight estimation error w of a comment neural network and an action neural network is determined if the following conditions are met. CM and w AM It is uniformly eventually bounded (UUB):

[0239]

[0240] Proof: Construct the Lyapunov function:

[0241]

[0242] in, Its first difference is divided into:

[0243]

[0244] We can obtain:

[0245]

[0246] From the Cauchy-Schwartz inequality, we can obtain:

[0247]

[0248] It can be represented as:

[0249]

[0250] make We can obtain:

[0251]

[0252] The objective of this invention is to ensure that Δv ≤ 0, therefore the following inequalities must be satisfied:

[0253]

[0254] That is, it must satisfy:

[0255]

[0256] Q.E.D.

[0257] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as ROM, RAM, magnetic disk, or optical disk.

[0258] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A self-triggered adaptive dynamic programming method based on model prediction, characterized in that, Includes the following steps: S1. Construct a nonlinear system to establish an accurate mathematical model for subsequent control algorithm design and describe the dynamic characteristics of the controlled object. S2. Design a self-triggering mechanism for real-time optimization. This self-triggering mechanism reduces the frequency of online calculations and reduces resource consumption by dynamically triggering and controlling updates. The self-triggering mechanism updates control only when the state deviates from the threshold, reducing computational burden; S3. Design a self-triggering mechanism (STM) to realize real-time calculation of the trigger interval and calculate the trigger condition threshold to ensure that the system remains stable within the trigger interval. S4. Obtain the inherent dynamic characteristics by establishing a data-driven model of the controlled system; Let the neural network fitting the nonlinear system be a model neural network; the system state and control input The data is measured within the current prediction time domain, and then the neural network model infers the state at the next time step. For discrete-time systems, It is necessary to estimate by training a neural network model before performing the learning process; S5. Construct a comment neural network. In the current prediction time domain, the terminal penalty is fitted using the comment neural network to improve the system performance. S6. Design an action neural network to fit the optimal control strategy; S7. Train the neural network model and update the neural network weights using gradient descent: S8. Perform online prediction of the system; the characteristics of the nonlinear system under optimal control and the weight error of the neural network are ultimately uniformly bounded under given conditions; In step S3, there are two cases: Scenario 1: When and At that time, there were: ; because and ,have to , ,therefore: ; Depend on ,have to ,therefore The trigger interval is: ; in, This indicates rounding down; therefore, in case 1: ; Therefore The triggering condition is met; Scenario 2: When and At that time, there were: ; because Follow The increase is monotonically increasing, therefore, suppose there exists Make: ; have to ;because Solve for the next trigger interval ,in, ; In summary, the self-triggered mechanism is designed as follows: 。 2. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S1, consider the nonlinear system as follows: ; In the formula, The time step index indicates that the system is at the _ . step, For the system in the first The state of the step, For the system in the first Step control input, Let be the state transition function of the system. Let be the control input matrix of the system; It is the equilibrium point of the system; in addition, exist The above is continuously differentiable, and and Considered unknown; defined This is a feedback control law.

3. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S2, a strictly monotonically increasing sequence is defined. As the self-sampling time; the controller only triggers and updates at the self-sampling time; then, the sampling interval is defined as: ; in, This represents the system's self-sampled state; since the controller only updates at the self-sampled time, the feedback control law will be used. Defined as The cost function is defined as: ; Where p is the time step index, indicating that the system is at the _ ... step, Indicates the system at the 1st The state of the step, This represents the sampling interval of the system at step p. Indicates that the system is in Control law and utility function in the state ;here It is a positive semi-definite matrix. It is a positive definite symmetric matrix.

4. The self-triggered adaptive dynamic programming method based on model prediction according to claim 3, characterized in that: Assume the system is controllable and its state is observable; furthermore, This represents any permissible control strategy of the system; Combining model predictive control mechanisms, let Indicates the prediction time domain, Indicates control over the time domain, and To reduce computational requirements, parameters are set. make , here It is in terminal state; therefore, To further express: ; in, The terminal constraint within the finite prediction time domain is represented as: ; in, For the system in Transpose of the state matrix at time step The weight matrix is ​​positive definite; according to Bellman's optimality theorem, the optimal cost function is... satisfy: ; ; in, Indicates that the system is in The utility function at time For the system in The optimal cost function at that time, This is the transpose of the system state. For the system in Time control law, for transpose, To determine the optimal strategy that minimizes the cost function and ensures system stability Its expression is: 。 5. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S4, the neural network fits the state expression as follows: ; in, For the input vector, For bounded activation functions, choose and The weight matrix is ​​randomly initialized; the identification error of the neural network model is defined as... Therefore, the performance metrics of the model neural network are as follows: ; After the neural network model is trained, its weights and It will be fixed and used in the online prediction phase.

6. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S5, the cost function is calculated as follows: ; in, It is a bounded activation function. and For the target weight matrix, This represents the number of neurons in the hidden layer. The cost function expression is: (The cost function expression is missing from the original text.) ; in, and Randomly initialized; combined with the MPC mechanism, the comment neural network needs to fit two targets simultaneously: the time domain and the time domain. Cost function within Terminal status Penalty items The specific implementation is as follows: ; According to the Bellman equation, the ideal cost function should satisfy: ; in, Represents the estimated control law, This represents the terminal state of the neural network fit.

7. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: Introducing relaxation factors To balance the convergence rate, the error function is rewritten as: ; in, , , The total residual error is: ; in, , For simplicity, let Therefore, the performance indicators are: ; The weight update rules are as follows: ; in, This is the learning rate.

8. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S6, the strategy is: ; in, It is a bounded activation function. and For the target weight, The number of neurons in the hidden layer. The bounded error after fitting the neural network; The estimated value is: ; in, and For the corresponding weight matrix, It is a positive definite symmetric matrix; From this formula, we can obtain: ; The corresponding control input error is: ; in, , ;make for Then the objective error function is: ; The weight update rules are as follows: ; in, This is the learning rate.

9. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S7, the gradient descent method is used to update the weights of the model's neural network: ; ; in, This is the learning rate.