Self-triggering adaptive dynamic planning method based on model prediction

By combining the self-trigger mechanism and model prediction control MSTADP method, the optimal control computing complexity problem of unknown nonlinear systems is solved, and an efficient and robust control strategy is realized, which reduces the online computing load and improves system performance.

CN120578053AActive Publication Date: 2025-09-02BEIHANG UNIV

Patent Information

Application Number
CN202510452909.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-09-02
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The prior art, when dealing with optimal control of unknown nonlinear systems, has high computational requirements and high complexity, making it difficult to effectively integrate event triggering mechanisms and model prediction controls to achieve robust and adaptive control strategies.

Method used

A self-triggered adaptive dynamic programming method (MSTADP) based on model prediction is proposed. Combining the self-trigger mechanism and model prediction control, the number of optimization problems is reduced through dynamic trigger intervals, the online calculation load is reduced, and the dynamic characteristics of the system are fitted through the neural network, and the self-trigger mechanism and comment neural network are designed to improve system performance.

Benefits of technology

It effectively reduces the online computing burden, while ensuring system performance, realizes efficient and optimal control of unknown nonlinear systems, and improves the robustness and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578053A_ABST
    Figure CN120578053A_ABST
Patent Text Reader

Abstract

The invention discloses a self-triggering self-adaptive dynamic programming method based on model prediction, which belongs to the field of optimal control, and comprises the step of developing a self-triggering self-adaptive dynamic programming algorithm based on an iterative model prediction process for an unknown nonlinear system. A trigger condition function is designed through Lyapunov stability analysis, intermittent control only depending on the previous trigger state is achieved, and communication traffic is reduced. A three-network architecture composed of a model network, an evaluation network and an execution network is designed. Then, a rolling optimization mechanism is provided, infinite time domain optimization is decomposed into finite time domain sub-problems, and calculation time consumption is reduced; and finally, providing a simulation result to verify the effectiveness of the proposed method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of optimal control, and in particular to a self-triggering adaptive dynamic programming method based on model prediction. Background Art

[0002] Optimal control of systems is a key research topic in modern control theory. It plays a vital role in numerous practical applications, including autonomous vehicles, industrial production processes, and aircraft control. By optimizing control strategies, system efficiency can be significantly improved, costs can be reduced, and stability and robustness can be enhanced. Adaptive dynamic programming (ADP) is an important approach for solving optimization and control problems in such complex systems. In recent years, various ADP methods have been developed to achieve optimal control. However, due to computational limitations, finding feasible solutions to the Hamilton-Jacobi-Bellman (HJB) equations remains a significant challenge.

[0003] To overcome the challenges in optimal control, researchers have proposed a variety of solutions. Model predictive control (MPC) has shown significant advantages in dealing with optimization problems of unknown nonlinear systems. MPC solves practical industrial problems by implementing a rolling horizon (RH) mechanism and establishing a predictive model. Researchers have proposed a functional MPC method using the ADP algorithm, which can handle constraints and disturbances to achieve optimal control. In addition, the self-triggering mechanism (STM) method can significantly reduce the computational load, communication overhead, and energy consumption, while adapting to the characteristics of dynamic nonlinear systems. These methods have enhanced the robustness and practicality of the controller to varying degrees.

[0004] Despite recent progress, several unresolved challenges remain in the field of optimal control. For example, while model predictive control (MPC) offers significant advantages in optimizing unknown nonlinear systems, its development is hampered by the high computational requirements and complexity of solving nonlinear MPC problems. While event-triggered mechanisms and model predictive control (MPC) have received theoretical attention, effectively integrating these mechanisms in practical applications to design more robust and adaptive control strategies remains an open question. Therefore, optimal control research needs to continue exploring more effective approaches to address these challenges. Summary of the Invention

[0005] The purpose of the present invention is to overcome the defects of the prior art and solve the optimal control problem of a class of nonlinear systems with terminal state constraints.

[0006] To achieve the above objectives, the present invention proposes a self-triggered adaptive dynamic programming method based on model prediction (MSTADP), which includes the following steps:

[0007] S1. Construct a general nonlinear system, establish an accurate mathematical model for subsequent control algorithm design, and describe the dynamic characteristics of the controlled object.

[0008] First consider the nonlinear system as:

[0009] x k+1 =f(x k )+g(x k )u k

[0010] Where k is the time step index, indicating that the system is at the kth step, x k ∈R n is the state of the system at step k, u k ∈R m is the control input of the system at step k, f(x k ) is the state transfer function of the system, g(x k ) is the control input matrix of the system. Assume that x=0 is the equilibrium point of the system. In addition, f(x k )+g(x k )u(x k ) in R n ×R m →R n is continuously differentiable, and f and g are considered unknown. Define u(x k ) is the feedback control law.

[0011] S2. This invention designs a novel self-triggering mechanism for real-time optimization. This triggering mechanism can reduce the frequency of online computations and reduce resource consumption through dynamic triggering control updates. Traditional event-triggered mechanisms can lead to credit calculations, while self-triggering mechanisms only update control when the state deviates from a threshold, reducing the computational burden. This invention combines the rolling optimization characteristics of MPC to balance real-time performance and optimality within a limited time domain. The main methods are:

[0012] First, define a strictly monotonically increasing sequence As the self-sampling time. The controller is triggered and updated only at the self-sampling moment. Next, define the sampling interval as:

[0013]

[0014] in, is the self-sampling state of the system. Since the controller is only updated at the self-sampling moment, we will feedback the control law u(x k ) is defined as v(x k ), the cost function is defined as:

[0015]

[0016] Among them, p is the time step index, indicating that the system is in the pth step, x p represents the state of the system at step p, ξ p represents the sampling interval of the system at step p, v(x p +ξ p ) indicates that the system is at x p +ξ p Control law in state, utility function Here Q is a positive semidefinite matrix and R is a positive definite symmetric matrix. Assumption 1: The system is controllable and its state is observable. In addition, u represents any admissible control policy for the system.

[0017] Combined with the model predictive control mechanism, let H P represents the prediction time domain, H C represents the control time domain, and H C ≤H P To reduce the computational requirements, set the parameter N so that H C =H P =N, where is a terminal state. Therefore, J(x k ) can be further expressed as:

[0018]

[0019] in, Represents the terminal constraint within the finite prediction horizon. It can be expressed as:

[0020]

[0021] in, For the system The transpose of the state matrix at the moment, P is the positive definite weight matrix. According to Bellman's optimality theorem, the optimal cost function J * (x k )satisfy:

[0022]

[0023]

[0024] in, Indicates that the system is The utility function when J * (x k+1 ) is the system at x k+1 The optimal cost function when is the transpose of the system state, For the system The control law when for The transpose of

[0025] To determine the optimal strategy to minimize the cost function and ensure system stability Its expression is:

[0026]

[0027] Note 1: The self-triggering mechanism STM designed in this invention can effectively reduce the computational burden, but it will also lead to a decrease in system performance. In addition, since the design of STM includes the current sampling time k j ,Offline algorithms cannot achieve real-time control of the system. To this end, we introduce the MPC mechanism to accelerate the convergence of the system. Therefore, this paper proposes a self-triggered adaptive dynamic programming method based on MPC. Its essence is to add STM to the MPC control process to reduce the number of predictions, thereby reducing the online computing burden while ensuring system performance.

[0028] S3. Design a self-triggering mechanism STM to realize real-time calculation of the trigger interval and calculate the trigger condition threshold to ensure that the system remains stable within the trigger interval. Assumption 2: Assume that for all x k ∈Ω k , optimal control strategy u * (x k ) is Lipschitz continuous, that is, there exists a Lipschitz constant Υ satisfying:

[0029]

[0030] Theorem 1: If both Assumptions 1 and 2 hold, the trigger threshold can be derived:

[0031]

[0032] Among them, η∈(0,1), κ is a constant greater than 0, and λ min (·) and λ max (·) represent the minimum and maximum eigenvalues ​​of the matrix respectively. k )≤0, the system is always ultimately bounded (UUB).

[0033] Proof: Consider the variation of the Lyapunov function:

[0034]

[0035]

[0036] in, The square term representing the upper bound of the uncertainty of the model

[0037] Further decomposition of the quadratic term:

[0038]

[0039] Combining the above two equations, we can get:

[0040]

[0041] When Ξ(ξ k )≤0, to ensure ΔJ * (x k )≤0, must satisfy:

[0042]

[0043] That is, the sampling status must meet the following requirements:

[0044]

[0045] Assumption 3: For a closed-loop system, there exists a parameter σ > 0.5 such that the sampling error ξ k The sampling state satisfies the inequality:

[0046] ‖x k+1 ‖=σ‖x k ‖+σ‖ξ k ‖

[0047] Lemma 1: If the inequality in Assumption 3 holds, then for a nonlinear system, the measurement error ξ k and sampling status satisfy:

[0048]

[0049] Theorem 3: If Assumptions 2 and 3 hold, then the next trigger time k j+1 The current sampling status can be directly Determine and meet the stability conditions. The specific form is:

[0050]

[0051] in It is obtained by the following formula:

[0052]

[0053] Proof: According to Theorem 1 and Theorem 3, we can get:

[0054]

[0055] To simplify the calculation, let the right side of the formula be equal to Right now There are two cases to discuss:

[0056] Case 1: When k = kj +1 and When:

[0057]

[0058] Since 0<η<1 and σ>0.5, we can get therefore:

[0059]

[0060] Depend on Available therefore The trigger intervals are:

[0061]

[0062] in, means round down. So, in this case:

[0063]

[0064] Therefore, Δk j =1, the trigger condition is met.

[0065] Case 2: When k = k j +1 and When:

[0066]

[0067] because It increases monotonically with k, so, assuming there exists ε ≥ k j +1 makes this:

[0068]

[0069] Available because The next trigger interval k can be solved j =1+(ε), where

[0070] In summary, the trigger mechanism design can be obtained as follows:

[0071]

[0072] The proof is complete.

[0073] This step mainly analyzes the energy changes of the system based on the Lyapunov function, combines the Lipschitz continuity assumption, quantifies the error limit of the control strategy, and thus designs the trigger threshold.

[0074] S4. Neural networks have strong fitting capabilities. This step obtains the inherent dynamic characteristics of the controlled system by establishing a data-driven model. Let the neural network that fits the nonlinear system be the model neural network. In this framework, the system state x k and control input u k is measured in the current prediction time domain, and then the neural network model can infer the state x at the next moment k+1 For discrete-time systems, It needs to be estimated by training the neural network model before executing the learning process. The neural network fitting state expression is:

[0075]

[0076] in, is the input vector, θ(·) is the bounded activation function, and tanh(·) is selected in the present invention. m1 and w m2 is a randomly initialized weight matrix. The identification error of the model neural network is defined as Therefore, the performance indicators of the model neural network are as follows:

[0077]

[0078] Note 2: After the model neural network training is completed, its weight w m1 and w m2 will be fixed and used in the online prediction phase.

[0079] S5. Construct a review neural network. In the current prediction time domain, the terminal penalty is fitted using the review neural network, which improves the performance of the system. The cost function is calculated as:

[0080]

[0081] in, and is the target weight matrix, l1 is the number of hidden layer neurons, τ ck is the bounded estimation error. The cost function expression of the fitting is:

[0082]

[0083] Among them, w c1 and w c2 is randomly initialized. Combined with the MPC mechanism, the comment neural network needs to fit two objectives at the same time: the cost function J(x k ), terminal state Penalty items The specific implementation is:

[0084]

[0085] According to the Bellman equation, the ideal cost function should satisfy:

[0086]

[0087] in, represents the estimated control law, Indicates the terminal state of the neural network fitting. Introducing the relaxation factor To balance the convergence speed, the error function is rewritten as:

[0088]

[0089] Among them, Δθ(w c1 x k )=θ(w c1 x k+1 )-θ(w c1 x k ), Δτ ck =τ ck+1 -τ ck The total residual error is:

[0090]

[0091] in, Δτ c =Δτ ck +Δτ ck+N , for simplicity, let Δθ c =Δθ(w c1 x k )+Δθ(w c1 x k+N ). Therefore, the performance indicators are:

[0092]

[0093] The weight update rule is:

[0094]

[0095] Among them, α c ∈(0,1) is the learning rate.

[0096] S6. Design an action neural network to fit the optimal control strategy. The target strategy is:

[0097]

[0098] in, and is the target weight, l2 is the number of hidden layer neurons, is the bounded error after neural network fitting. The estimated value of is:

[0099]

[0100] Among them, ω a1 and ω a2 is the corresponding weight matrix. From this formula we can get:

[0101]

[0102] Similarly, the corresponding control input error is:

[0103]

[0104] in, make is τ a , then the objective error function is:

[0105]

[0106] The weight update rule is:

[0107]

[0108] Among them, α a ∈(0,1) is the learning rate. The present invention uses the designed action neural network to fit the HJB equation. It is worth noting that the method of using the neural network to fit the HJB equation to determine the control law is different from the method of determining the control law according to MPC.

[0109] S7. Train the model neural network and use the gradient descent method to update the model neural network weights:

[0110]

[0111] Among them, α m ∈(0,1) is the learning rate.

[0112] S8. Perform online prediction on the system.

[0113] Furthermore, it is proved that the properties of nonlinear systems under optimal control and the weight errors of neural networks are ultimately uniformly bounded under certain given conditions.

[0114] Compared with existing technologies, this invention achieves the following technical advantages: A self-triggered adaptive dynamic programming (MSTADP) method based on an iterative model prediction process is developed for unknown nonlinear discrete-time systems. By integrating the self-triggering mechanism with MPC, the dynamic triggering interval reduces the number of optimization problem solutions within the prediction time domain, thereby reducing the online computational load. An estimation-criticism structure is then proposed to implement the self-triggered adaptive dynamic programming algorithm based on the iterative model prediction process. Finally, simulation results are provided to verify the effectiveness of the proposed algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0115] Figure 1 This is a flow chart of a self-triggered adaptive dynamic programming method based on model prediction according to the present invention.

[0116] Figure 2 This is a spring-mass-damper system according to the present invention.

[0117] Figure 3 This is the training error of the neural network model of Example 1 of the present invention.

[0118] Figure 4 This is the system status under the MSTADP and HDP algorithms of Example 1 of the present invention.

[0119] Figure 5 This is the control input generated under the MSTADP and HDP algorithms in Example 1 of the present invention.

[0120] Figure 6 The trigger time interval Δk of the present invention is shown in Example 1 j . DETAILED DESCRIPTION

[0121] In order to further illustrate the features of the present invention, please refer to the following detailed description of the present invention and the attached Figure 1-6 The accompanying drawings are for reference and illustration purposes only and are not intended to limit the scope of protection of the present invention.

[0122] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings:

[0123] like Figure 1 As shown, this embodiment discloses a self-triggered adaptive dynamic programming method based on model prediction, comprising the following steps:

[0124] S1. The present invention will be Figure 2 The spring-mass-damper system shown here verifies the effectiveness of the MSTADP algorithm. The system is as follows:

[0125]

[0126] Here, x1 and x2 represent the state vectors, representing the horizontal displacement and velocity, respectively. F(u) is the control input, representing the applied force, and M is the mass of the object. Furthermore, K is the spring stiffness constant, and P is the damping coefficient of the damper. Table 1 lists the relevant system parameters.

[0127] Table 1

[0128] Q R P N Υ σ κ η I2 4I 0.5I2 5 10 0.7 0.6 0.98

[0129] The dynamic formula is as follows:

[0130]

[0131] S2. This invention designs a novel self-triggering mechanism for real-time optimization. This triggering mechanism can reduce the frequency of online computations and reduce resource consumption through dynamic triggering control updates. Traditional event-triggered mechanisms can lead to credit calculations, while self-triggering mechanisms only update control when the state deviates from a threshold, reducing the computational burden. This invention combines the rolling optimization characteristics of MPC to balance real-time performance and optimality within a limited time domain. The main methods are:

[0132] First, define a strictly monotonically increasing sequence As the self-sampling time. According to the MPC principle, the continuous time system is discretized and the sampling interval is set to 0.5s. The system is discretized as follows:

[0133]

[0134] The controller is only triggered and updated at the self-sampling moment. Next, define the sampling interval as:

[0135]

[0136] in, is the self-sampling state of the system. Since the controller is only updated at the self-sampling moment, we will feedback the control law u(x k ) is defined as v(x k ), the cost function is defined as:

[0137]

[0138] Among them, p is the time step index, indicating that the system is in the pth step, x p represents the state of the system at step p, ξ p represents the sampling interval of the system at step p, v(x p +ξ p ) indicates that the system is at x p +ξ p Control law in state, utility function Here Q is a positive semidefinite matrix and R is a positive definite symmetric matrix. Assumption 1: The system is controllable and its state is observable. In addition, u represents any admissible control policy for the system.

[0139] Combined with the model predictive control mechanism, let H P represents the prediction time domain, H C represents the control time domain, and H C ≤H P To reduce the computational requirements, set the parameter N so that H C =H P =N, where is a terminal state. Therefore, J(x k ) can be further expressed as:

[0140]

[0141] in, Represents the terminal constraint within the finite prediction horizon. It can be expressed as:

[0142]

[0143] in, For the system The transpose of the state matrix at the moment, P is the positive definite weight matrix. According to Bellman's optimality theorem, the optimal cost function J * (x k )satisfy:

[0144]

[0145] in, Indicates that the system is The utility function when J * (x k+1 ) is the system at x k+1 The optimal cost function when is the transpose of the system state, For the system The control law when for The transpose of

[0146] To determine the optimal strategy to minimize the cost function and ensure system stability Its expression is:

[0147]

[0148] Note 1: The self-triggering mechanism STM designed in this invention can effectively reduce the computational burden, but it will also lead to a decrease in system performance. In addition, since the design of STM includes the current sampling time k j,Offline algorithms cannot achieve real-time control of the system. To this end, we introduce the MPC mechanism to accelerate the convergence of the system. Therefore, this paper proposes a self-triggered adaptive dynamic programming method based on MPC. Its essence is to add STM to the MPC control process to reduce the number of predictions, thereby reducing the online computing burden while ensuring system performance.

[0149] S3. Design a self-triggering mechanism STM to realize real-time calculation of the trigger interval and calculate the trigger condition threshold to ensure that the system remains stable within the trigger interval. Assumption 2: Assume that for all x k ∈Ω k , optimal control strategy u * (x k ) is Lipschitz continuous, that is, there exists a Lipschitz constant Υ satisfying:

[0150]

[0151] Theorem 1: If both Assumptions 1 and 2 hold, the trigger threshold can be derived:

[0152]

[0153] Among them, η∈(0,1), κ is a constant greater than 0, and λ min (·) and λ max (·) represent the minimum and maximum eigenvalues ​​of the matrix respectively. k )≤0, the system is always ultimately bounded (UUB).

[0154] Proof: Consider the variation of the Lyapunov function:

[0155]

[0156] in, The square term representing the upper bound of the uncertainty of the model

[0157] Further decomposition of the quadratic term:

[0158]

[0159] Combining the above two equations, we can get:

[0160]

[0161] When Ξ(ξ k )≤0, to ensure ΔJ * (x k )≤0, must satisfy:

[0162]

[0163] That is, the sampling status must meet the following requirements:

[0164]

[0165] Assumption 3: For a closed-loop system, there exists a parameter σ > 0.5 such that the sampling error ξ k The sampling state satisfies the inequality:

[0166] ‖x k+1 ‖=σ‖x k ‖+σ‖ξ k ‖

[0167] Lemma 1: If the inequality in Assumption 3 holds, then for a nonlinear system, the measurement error ξ k and sampling status satisfy:

[0168]

[0169] Theorem 3: If Assumptions 2 and 3 hold, then the next trigger time k j+1 The current sampling status can be directly Determine and meet the stability conditions. The specific form is:

[0170]

[0171] in is a dynamic trigger threshold function, which is obtained by the following formula:

[0172]

[0173] Proof: According to Theorem 1 and Theorem 3, we can get:

[0174]

[0175] To simplify the calculation, let the right side of the formula be equal to Right now There are two cases to discuss:

[0176] Case 1: When k = k j +1 and When:

[0177]

[0178] Since 0<η<1 and σ>0.5, we can get therefore:

[0179]

[0180] Depend on Available therefore The trigger intervals are:

[0181]

[0182] in, means round down. So, in this case:

[0183]

[0184] Therefore, Δk j =1, the trigger condition is met.

[0185] Case 2: When k = k j +1 and When:

[0186]

[0187] because It increases monotonically with k, so, assuming there exists ε ≥ k j +1 makes this:

[0188]

[0189] Available because The next trigger interval k can be solved j =1+(ε), where

[0190] In summary, the trigger mechanism design can be obtained as follows:

[0191]

[0192] The proof is complete.

[0193] This step mainly analyzes the energy changes of the system based on the Lyapunov function, combines the Lipschitz continuity assumption, quantifies the error limit of the control strategy, and thus designs the trigger threshold.

[0194] S4. Neural networks have strong fitting capabilities. This step obtains the inherent dynamic characteristics of the controlled system by establishing a data-driven model. Let the neural network that fits the nonlinear system be the model neural network, and the configuration of the model neural network architecture is 3-8-2. In this framework, the system state x k and control input u k is measured in the current prediction time domain, and then the neural network model can infer the state x at the next moment k+1 For discrete-time systems, It needs to be estimated by training the neural network model before executing the learning process. The neural network fitting state expression is:

[0195]

[0196] in, is the input vector, θ(·) is the bounded activation function, and tanh(·) is selected in the present invention. m1 and w m2 is a randomly initialized weight matrix. The identification error of the model neural network is defined as Therefore, the performance indicators of the model neural network are as follows:

[0197]

[0198] Note 2: After the model neural network training is completed, its weight w m1 and w m2 will be fixed and used in the online prediction phase.

[0199] S5. Construct a review neural network with a configuration of 2-8-1. In the current prediction time domain, the terminal penalty is fitted using the review neural network, which improves the performance of the system. The cost function is calculated as:

[0200]

[0201] in, and is the target weight matrix, l1 is the number of hidden layer neurons, τ ck is the bounded estimation error. The cost function expression of the fitting is:

[0202]

[0203] Among them, w c1 and w c2 is randomly initialized. Combined with the MPC mechanism, the comment neural network needs to fit two objectives at the same time: the cost function J(x k ), terminal state Penalty items The specific implementation is:

[0204]

[0205] According to the Bellman equation, the ideal cost function should satisfy:

[0206]

[0207] in, represents the estimated control law, Indicates the terminal state of the neural network fitting. Introducing the relaxation factor To balance the convergence speed, the error function is rewritten as:

[0208]

[0209] Among them, Δθ(w c1 x k )=θ(w c1 x k+1 )-θ(w c1 x k ), Δτ ck =τ ck+1 -τ ck The total residual error is:

[0210]

[0211] in, Δτ c =Δτ ck +Δτ ck+N , for simplicity, let Δθ c =Δθ(w c1 x k )+Δθ(w c1 x k+N ). Therefore, the performance indicators are:

[0212]

[0213] The weight update rule is:

[0214]

[0215] Among them, α c =0.7 is the learning rate.

[0216] S6. Design an action neural network to fit the optimal control strategy. The network architecture is configured as 2-8-1. The target strategy is:

[0217]

[0218] in, and is the target weight, l2 is the number of hidden layer neurons, is the bounded error after neural network fitting. The estimated value of is:

[0219]

[0220] Among them, ω a1 and ω a2 is the corresponding weight matrix. From this formula we can get:

[0221]

[0222] Similarly, the corresponding control input error is:

[0223]

[0224] in, make is τ a , then the objective error function is:

[0225]

[0226] The weight update rule is:

[0227]

[0228] Among them, α a =0.2 is the learning rate. The present invention uses the designed action neural network to fit the HJB equation. It is worth noting that the method of using the neural network to fit the HJB equation to determine the control law is different from the method of determining the control law according to MPC.

[0229] S7. Train the model neural network and use the gradient descent method to update the model neural network weights:

[0230]

[0231] Among them, α m =0.1 is the learning rate.

[0232] For the position system, 500 data sets were randomly selected in the range of [0.01, 0.01] for model network training, and another 500 data sets were used to test accuracy. Each data set was trained 500 times. To verify the accuracy of the neural network model approximation system, the training error reached Figure 3 shown in 10 -3 The accuracy of the model network proves the reliability and accuracy of the model network. After testing the model network, the final w m1 and w m2 The parameter configuration is shown in Table 2. The initial state is set to x0 = [1.2, -1.2]. T .

[0233] Table 2

[0234] M K C Ts 1kg 5N / s 0.5Ns / m 0.5s

[0235] S8. Perform online prediction on the system.

[0236] By comparing with the HDP algorithm, the effectiveness of the MSTADP algorithm is proved. exist Figure 4In the , all system state trajectories are Figure 5 As shown, both converge to zero under the influence of the control input. Figure 6 The trigger interval is shown. The control input is updated 41 times in 150 time steps. The results show that the STM method significantly reduces resource utilization. From the observation results, it can be seen that the convergence speed of all curves is significantly accelerated after the MSTADP algorithm is adopted, indicating that the control effect of the MSTADP algorithm is better than that of the HDP algorithm.

[0237] Next, we will further analyze the convergence and stability of the neural network. Assumption 2: and As the target weights of the comment neural network and action neural network respectively. in, and is a positive constant. At the same time, assuming β m ≤‖τ c ‖≤θ cM , γ m ≤‖τ a ‖≤γ aM .

[0238] Theorem 2: If the following conditions are met, the weight estimation error w of the comment neural network and the action neural network CM and w AM is uniformly eventually bounded (UUB):

[0239]

[0240] Proof: Construct the Lyapunov function:

[0241]

[0242] in, Its first-order difference is:

[0243]

[0244] We can get:

[0245]

[0246] According to the Cauchy-Schwartz inequality, we can get:

[0247]

[0248] It can be expressed as:

[0249]

[0250] make We can get:

[0251]

[0252] The purpose of the present invention is to make Δv≤0, so the following inequality must be satisfied:

[0253]

[0254] That is, it must meet:

[0255]

[0256] The proof is complete.

[0257] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk, etc. Various media that can store program codes.

[0258] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A self-triggered adaptive dynamic programming method based on model prediction, characterized in that: The steps include: S1. Construct nonlinear systems, establish accurate mathematical models for subsequent control algorithm design, and describe the dynamic characteristics of the controlled object; S2. Design a self-triggering mechanism for real-time optimization. This self-triggering mechanism reduces the frequency of online calculations and reduces resource consumption through dynamic trigger control updates. The self-triggering mechanism updates the control only when the state deviates from the threshold, reducing the computational burden; S3. Design a self-triggering mechanism (STM) to achieve real-time calculation of the trigger interval and the trigger condition threshold to ensure that the system remains stable within the trigger interval. S4. Obtain the inherent dynamic characteristics by establishing a data-driven model of the controlled system; let the neural network fitting the nonlinear system be the model neural network; the system state x k and control input u k is measured in the current prediction time domain, and then the neural network model infers the state x at the next moment k+1 ; For discrete-time systems, It needs to be estimated by training a neural network model before performing the learning process; S5. Construct a review neural network. In the current prediction time domain, the terminal penalty is fitted using the review neural network to improve the performance of the system. S6. Design an action neural network to fit the optimal control strategy; S7. Train the model neural network and use the gradient descent method to update the model neural network weights: S8. Perform online prediction on the system; the characteristics of the nonlinear system under optimal control and the weight error of the neural network are ultimately uniformly bounded under given conditions.

2. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S1, the nonlinear system is considered as: x k+1 =f(x k )+g(x k )u k Where k is the time step index, indicating that the system is at the kth step, x k ∈R n is the state of the system at step k, u k ∈R m is the control input of the system at step k, f(x k ) is the state transfer function of the system, g(x k ) is the control input matrix of the system; let x = 0 be the equilibrium point of the system; in addition, f(x k )+g(x k )u(x k ) in R n ×R m →R n is continuously differentiable, and f and g are considered unknown; define u(x k ) is the feedback control law.

3. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S2, define a strictly monotonically increasing sequence As the self-sampling time; the controller is triggered and updated only at the self-sampling time; then, define the sampling interval as: in, is the self-sampling state of the system; since the controller is only updated at the self-sampling moment, the feedback control law u(x k ) is defined as v(x k ), the cost function is defined as: Among them, p is the time step index, indicating that the system is in the pth step, x p represents the state of the system at step p, ξ p represents the sampling interval of the system at step p, v(x p +ξ p ) indicates that the system is at x p +ξ p Control law in state, utility function Here Q is a positive semidefinite matrix and R is a positive definite symmetric matrix.

4. The self-triggered adaptive dynamic programming method based on model prediction according to claim 3, characterized in that: Assume that the system is controllable and its state is observable; in addition, u represents any admissible control strategy of the system; Combined with the model predictive control mechanism, let H P represents the prediction time domain, H C represents the control time domain, and H C ≤H P ; To reduce the computational requirements, set the parameter N so that H C =H P =N, where is a terminal state; therefore, J(x k ) is further expressed as: in, represents the terminal constraint within the finite prediction horizon; it is expressed as: in, For the system The transpose of the state matrix at the moment, P is the positive definite weight matrix; according to Bellman's optimality theorem, the optimal cost function J * (x k )satisfy: in, Indicates that the system is The utility function when J * (x k+1 ) is the system at x k+1 The optimal cost function when is the transpose of the system state, For the system The control law when for The transpose of To determine the optimal strategy to minimize the cost function and ensure system stability Its expression is:

5. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S3, there are two cases: Case 1: When k = k j +1 and When: Since 0<η<1 and σ>0.5, we have therefore: Depend on have to therefore The trigger intervals are: in, means round down; therefore, in this case 1: Therefore, Δk j =1, the trigger condition is met; Case 2: When k = k j +1 and When: because It increases monotonically with the increase of k, so, let there exist ε≥k j +1 makes this: have to because Solve for the next trigger interval k j =1+(ε), where In summary, the self-triggering mechanism design is:

6. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S4, the neural network fitting state expression is: in, is the input vector, θ(·) is the bounded activation function, tanh(·) is selected; w m1 and w m2 is a randomly initialized weight matrix; the identification error of the model neural network is defined as Therefore, the performance indicators of the model neural network are as follows: After the model neural network training is completed, its weight w m1 and w m2 will be fixed and used in the online prediction phase.

7. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S5, the cost function is calculated as: in, and is the target weight matrix, l1 is the number of hidden layer neurons, τ ck is a bounded estimation error; the fitting cost function expression is: Among them, w c1 and w c2 is randomly initialized; combined with the MPC mechanism, the comment neural network needs to fit two objectives at the same time: the cost function J(x k ), terminal state Penalty items The specific implementation is: According to the Bellman equation, the ideal cost function should satisfy: in, represents the estimated control law, Represents the terminal state of the neural network fit.

8. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: Introducing relaxation factors To balance the convergence speed, the error function is rewritten as: Among them, Δθ(w c1 x k )=θ(w c1 x k+1 )-θ(w c1 x k ), Δτ ck =τ ck+1 -τ ck The total residual error is: in, Δτ c =Δτ ck +Δτ ck+N , for simplicity, let Δθ c =Δθ(w c1 x k )+Δθ(w c1 x k+N ); therefore, the performance indicators are: The weight update rule is: Among them, α c ∈(0,1) is the learning rate.

9. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S6, the strategy is: in, and is the target weight, l2 is the number of hidden layer neurons, is the bounded error after neural network fitting; The estimated value of is: Among them, ω a1 and ω a2 is the corresponding weight matrix; From this formula we can get: The corresponding control input error is: in, make is τ a , then the objective error function is: The weight update rule is: Among them, α a ∈(0,1) is the learning rate.

10. The self-triggered adaptive dynamic programming method based on model prediction according to claim 1, characterized in that: In step S7, the gradient descent method is used to update the model neural network weights: Among them, α m ∈(0,1) is the learning rate.

Citation Information

Patent Citations

  • Optimization control system and method based on control strategy update time prediction

    CN116382070A

  • Self-adaptive dynamic programming control method for specified time

    CN118192224A

  • Self-learning differential game collaborative guidance method based on consensus initiative

    CN118550322A

  • Three-phase grid-connected inverter model prediction control strategy optimization method

    CN118868217A

  • Unmanned ship trajectory tracking control method based on self-adaptive dynamic event triggering

    CN119396141A

Cited By

  • Dynamic correction method for digital twin model parameters of line equipment of MPC

    CN120871633A