An Adaptive Dynamic Programming Control Method with Specified Time
By using the adaptive dynamic programming control method within the specified time, the problem that the stability time is not affected by the initial state of the system is solved, and the optimality and stability of the system within the specified time is achieved, which is better than the traditional control method.
Patent Information
- Application Number
- CN202410197075.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-02-22
AI Technical Summary
The prior art is difficult to accurately specify the stability time within the specified time to ensure that the stability time is not affected by the initial state of the system, while meeting the optimality and specified time stability requirements.
A control method for adaptive dynamic programming (ADP) is proposed to establish a nonlinear time-varying delay system model, design a prescribed time cost function and Hamiltonian equation, approximate unknown terms by judging neural networks, and solve the optimal control strategy through the neural network weight update law.
It realizes the simultaneous optimization and stability within the specified time, proves the stability of the nonlinear time-varying system in the specified time, and is better than the traditional prescribed/fixed time control method.
Smart Images

Figure CN118192224B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of finite-time optimal control, and particularly to a finite-time adaptive dynamic programming control method. Background Art
[0002] Finite-time control is an important topic in the fields of missile guidance, multi-agent cooperation, and smart home management. Generally speaking, finite-time control pays more attention to achieving precise and efficient control compared to classical finite / fixed-time control schemes. Its advantages lie in strong adaptability, good robustness, and the ability to better handle nonlinear systems and uncertainties. In practical engineering, the efficient and stable control performance is restricted by three factors: the initial conditions of the system or external disturbances; dependence on specific model assumptions; relatively accurate system models.
[0003] To solve the finite-time control problem, current research mainly includes three aspects: 1) Selection of time-varying functions. Current methods include data-driven methods, which learn the patterns of time-varying functions from historical data through machine learning algorithms. 2) Convergence techniques. Many schemes use techniques such as numerical optimization methods and introduction of adaptive learning rate strategies to improve the convergence speed and stability. 3) Robustness enhancement and observer design. Researchers tend to combine sliding mode observers and deep learning techniques to improve the accuracy and robustness of system states. In practical engineering, to ensure the steady-state performance of nonlinear time-varying delay systems, complex partial differential equations need to be solved. To address this problem, many scholars have studied the adaptive dynamic programming (ADP) technique.
[0004] Currently, many schemes have been proposed for finite-time control, but the following problems still remain: (i) How to accurately specify the settling time and ensure that the settling time is not affected by the initial state of the system. (ii) The weight update rule needs to understand the dynamic conditions and at the same time ensure compliance with the terminal constraint conditions. (iii) How to make the system achieve both optimality and finite-time stability. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects existing in the prior art and solve the problem of solving the nonlinear Hamilton-Jacobi-Bellman (HJB) equation within a finite time range.
[0006] To achieve the above purpose, the present invention proposes a finite-time adaptive dynamic programming (ADP) control method, which includes the following steps:
[0007] S1. Construct a nonlinear time-varying delay system model, and the specific form is:
[0008]
[0009] Wherein, They are the state vector, the control vector, and the uncertain parameter vector respectively; t 0 represents the initial time; the function f(t, x(t), x τ (t), θ) is locally Lipschitz with respect to the state vector x and piecewise continuous with respect to time t. For any t ≥ 0, f(t, 0, x τ (t), θ) = 0; x τ (t) = x(t - τ(t)) represents the time-varying delay term, τ(t) represents the unknown time-varying delay, and its derivative satisfies v(s) represents a continuous function.
[0010] The model needs to satisfy Assumption 1:
[0011] ||f(t, x(t), x τ (t), θ)|| ≤ Φ(θ)(γ(x(t))||x(t)|| + η(x(t))λ(x τ (t))||x τ (t)||),
[0012] where ||·|| represents the vector norm or the matrix 2-norm, γ(x(t)), η(x(t)), and λ(x τ (t)) are continuous positive semi-definite functions, and Φ(θ) represents a function related to θ.
[0013] S2. Construct the prescribed-time cost function; establish the cost function on the interval [t 0 , t 0 + T p ) The specific process is as follows: Design the prescribed-time cost function as:
[0014]
[0015] where the function ψ(·) is the terminal constraint condition, T p is the preset time parameter, and the stage cost is:
[0016]
[0017] where is a positive definite symmetric matrix.
[0018] S3. Construct the HJB equation and calculate the expression of the optimal solution. The specific process is as follows: Establish the Hamiltonian equation as:
[0019]
[0020] where u = u(t), g = g(x), f = f(t, x(t), x τ (t), θ), x = x(t),
[0021] Define the cost function for the optimal specified time as Solve the HJB equation The optimal solution expression can be obtained as That is:
[0022]
[0023] Among them, To obtain the optimal control, it is necessary to solve H(x, u * ). However, H(x, u * ) is a non - linear equation, and it is challenging to obtain an analytical solution. Given that the HJB equation is a non - linear partial differential equation, it is even more difficult to find its solution. The present invention approximates the unknown term by designing an appropriate neural network structure.
[0024] S4. Construct the evaluation neural network and the Hamiltonian equation. The specific process is as follows: Design the evaluation neural network as to approximately estimate the specified - time cost function;
[0025] Among them, is the weight vector of the evaluation neural network, is the activation function, and ε(x, t) is the neural network approximation error. The terminal constraint is:
[0026]
[0027] To accurately estimate and approximate V x and V t , design the evaluation neural network as:
[0028]
[0029]
[0030] Among them,
[0031] Substitute V x into , and we can get:
[0032]
[0033] Among them, W a represents the weight vector of the execution neural network, represents the weight vector of the execution neural network under the optimal strategy, To solve the optimal control strategy, substitute the neural - network - approximated V x and V tSubstitute into the HJB equation \(H(x, u * )\), we can get:
[0034]
[0035] where, is the residual of the HJB equation.
[0036] S5. Design the approximate weight vector to approximate the time-varying cost function and calculate the approximate expression of the HJB equation. The specific process is as follows: Since the specified time cost function fluctuates with time, design the approximate weight vector to estimate the ideal weight and construct the approximate evaluation neural network The approximate terminal constraint condition is where represents the estimated value of the state. \(V x and \(V t The approximate expressions are:
[0037]
[0038]
[0039] Design the weight estimation error of the evaluation neural network as where, is the weight vector of the evaluation neural network under the optimal strategy, is the approximate weight vector of the evaluation neural network. Since is unknown, \(u * can be approximated as Combined with the weight estimation error, the HJB equation can be transformed into By transforming it into the form to solve the optimal strategy.
[0040] S6. Solve the extreme value of the approximate HJB equation, minimize the residual of the HJB equation and the terminal error. The specific process is as follows: To solve the equation Design \(P(t)=0\) is equivalent to Then the solution of the extreme value of the HJB equation is transformed into solving \(P(t)=0\).
[0041] Construct the terminal constraint estimation error as:
[0042]
[0043] where, To minimize the terminal constraint error, define the squared constraint error as
[0044] S7. Update the weights of the execution neural network and the evaluation neural network. The specific form is as follows: The weight update rule of the execution neural network is where γ 1 represents the learning rate of the execution neural network.
[0045] The weight update rule of the evaluation neural network is:
[0046]
[0047] where γ 2 and γ 3 are positive parameters, representing the learning rate of the evaluation neural network.
[0048] S8. End, and obtain the optimal control law as:
[0049]
[0050] where the continuous function μ(t) is a specified time adjustment (T p -PTA) function. That is, it satisfies the following conditions:
[0051]
[0052]
[0053] Furthermore, it is proved that the nonlinear time-varying system using this control method is stable within the specified time.
[0054] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: Aiming at the problem that the specified / fixed-time control method can only achieve stability within the specified time range, a specified-time adaptive dynamic programming control method is proposed to ensure both optimality and stability at the specified time; A kind of execution evaluation neural network with a time-varying activation function is constructed, and a new weight update rule is derived according to the terminal error of the system and the approximation error of the HJB equation; According to the proposed specified-time stability criterion, it is proved that this control scheme satisfies the specified-time stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a flowchart of the specified-time adaptive dynamic programming control method of the present invention.
[0056] Figure 2 is a comparison graph of the convergence performance of the system state and the controller under the ADPPTC and APCPTC schemes.
[0057] Figure 3 is a cost function curve graph and a terminal constraint error graph under different strategies.
[0058] Figure 4 Convergence curves of x under different Ts p and strategies.
[0059] Figure 5 For different Ts p under u and convergence curves.
[0060] Figure 6 Convergence curves of x under different deltas
[0061] Figure 7 For different deltas under u and convergence curves.
[0062] Figure 8 For different Ts p convergence curves of x.
[0063] Figure 9 For different Ts p under x 1 and x 2 convergence curves.
[0064] Figure 10 For different Ts p under u 1 and u 2 versus and convergence curves.
[0065] Figure 11 For different deltas under x 1 and x 2 convergence curves.
[0066] Figure 12 For different deltas under u 1 and u 2 versus and convergence curves. Detailed implementation manners
[0067] To further illustrate the features of the present invention, please refer to the following detailed description of the present invention and the accompanying drawings. The accompanying drawings are for reference and illustration only and are not used to limit the scope of protection of the present invention. The detailed implementation manners of the present invention are described in detail below with reference to the accompanying drawings:
[0068] As Figure 1 shown, this embodiment discloses a specified-time adaptive dynamic programming (ADP) control method, including the following steps:
[0069] S1. Construct a nonlinear time-varying delay system model
[0070] In a missile guidance system, it is first necessary to achieve the tracking and positioning of the target. The position and velocity information of the target are obtained by using sensors, and then, taking them as inputs and combining with the current state information of the missile, a nonlinear time-varying delay system model is constructed. The cost function can reflect various constraint conditions and performance indicators during the missile flight. According to the actual industrial process, a nonlinear time-varying delay system model is constructed, and its specific form is:
[0071]
[0072] where, are the state vector, control vector, and uncertain parameter vector respectively; t 0 represents the initial time; the function f(t, x(t), x τ (t), θ) is locally Lipschitz with respect to the state vector x, piecewise continuous with respect to time t, and for any t ≥ 0, f(t, 0, x τ (t), θ) = 0; x τ (t) = x(t - τ(t)) represents the time-varying delay term, τ(t) represents the unknown time-varying delay, and its derivative satisfies v(s) represents a continuous function. The model needs to satisfy Assumption 1:
[0073] ||f(t, x(t), x τ (t), θ)|| ≤ Φ(θ)(γ(x(t))||x(t)|| + η(x(t))λ(x τ (t))||x τ (t)||),
[0074] where, ||·|| represents the norm of a vector or the two-norm of a matrix, γ(x(t)), η(x(t)), and λ(x τ (t)) are continuous semi-positive definite functions, and Φ(θ) represents a function related to θ.
[0075] According to the established system model, a cost function on the interval [t 0 , t 0 + T p ) is established, and the specified-time cost function is designed as:
[0076]
[0077] where, the function ψ(·) is the terminal constraint condition, T p is the preset time parameter, and the stage cost is:
[0078]
[0079] where, is a positive definite symmetric matrix. To solve the optimal control strategy of the nonlinear time-varying delay system, the Hamiltonian equation is defined as:
[0080]
[0081] where,
[0082] S2. Construct an approximate evaluation neural network
[0083] Design an approximate evaluation neural network to estimate the weights and the approximate cost function to better approximate the optimal control strategy. This network can help adjust the control input of the missile guidance system in real time to adapt to the uncertain environment and target changes. And according to the characteristics of the missile guidance system, establish the HJB equation. This equation describes the optimal control problem, that is, how to minimize the cost function and satisfy the system constraints. Design an approximate weight vector to approximate the time-varying cost function and calculate the approximate expression form of the HJB equation. The specific process is as follows: Since the specified time cost function fluctuates with time, design an approximate weight vector to estimate the ideal weight and construct an approximate evaluation neural network The approximate terminal constraint condition is Approximate V x and V t as:
[0084]
[0085]
[0086] Design the weight estimation error of the evaluation neural network as where, is the weight vector of the evaluation neural network under the optimal strategy, is the approximate weight vector of the evaluation neural network. Since is unknown, u * can be approximated as Combined with the weight estimation error, the HJB equation can be transformed into:
[0087]
[0088] By transforming it into form to achieve the solution of the optimal strategy.
[0089] S3. Obtain the specified time optimal control strategy and the neural network update rate
[0090] Using an approximate evaluation neural network, solve the approximate HJB equation and minimize the residual sum and constraint error. This process can help optimize the control strategy and improve the performance and robustness of the missile guidance system. Update the weights and parameters of the execution neural network and the evaluation neural network according to the system feedback and error conditions. This process is an iterative optimization process that can continuously improve the performance and adaptability of the missile guidance system. The optimal control strategy for the nonlinear time-varying delay system designed based on the neural network is as follows:
[0091]
[0092] where The continuous function μ(t) is a specified-time adjustment (T p -PTA) function. That is, it satisfies the following conditions:
[0093]
[0094]
[0095] The optimal control strategy for the nonlinear time-varying delay system is the solution of the HJB equation. The constructed neural network approximates the specified-time cost function by minimizing the HJB equation residual and the terminal error. Solve the HJB equation for the cost function approximated by the neural network to obtain the optimal control strategy of the system. Define P(t) and the squared constraint error. The specific process is as follows: To solve the equation Design P(t) = 0 is equivalent to Then, the extremum solution of the HJB equation is transformed into solving P(t) = 0.
[0096] Solving P(t) = 0 can obtain the weight update rule of the execution neural network as where γ 1 represents the learning rate of the execution neural network.
[0097] Construct the terminal constraint estimation error as:
[0098]
[0099] where To minimize the terminal constraint error, define the squared constraint error as:
[0100]
[0101] The weight update rule of the evaluation neural network is:
[0102]
[0103] where γ 2and γ 3 are positive parameters representing the learning rate of the evaluation neural network.
[0104] Through the above steps, precise control and adaptive adjustment can be achieved in the missile guidance system, thereby improving the hit accuracy and combat effectiveness of the missile. This method can adapt to various missile types and combat scenarios and provides new ideas and methods for the further development of missile guidance technology.
[0105] Furthermore, it is proved that the tracking error of the nonlinear time-varying delay system using this control method can be stabilized within a specified time. The specific process is as follows:
[0106] For the nonlinear time-varying delay system, design the optimal control law and the uncertain parameter adaptive law as follows:
[0107]
[0108]
[0109] and are bounded; the state variable x is globally stable within a specified time. Construct the cost function:
[0110]
[0111] where and are non-negative differentiable functions; γ 3 is a positive constant; is the constant upper bound of the uncertain parameter and the unknown time-varying delay function; is its approximation error.
[0112] It can be inferred that L and V 1 are non-negative differentiable functions, and V 1 ≤ L. If holds, then L(t) is bounded when t ∈ [t 0 , t 0 + T p ) and under the condition. Define the function ρ(t). When t ∈ [t 0 , t 0 + T p ), ρ(t) = μ(t); when t ∈ [t 0 + T p , +∞), ρ(t) = κ > 0.
[0113] To prove that and are bounded, select the Lyapunov function as:
[0114]
[0115] Hypothesis 2: The ideal neural network weights, activation functions, function reconstruction errors, and their derivatives are constrained by known normal numbers. Calculate the derivative of V ac :
[0116]
[0117] Considering the definition of matrix norm, it can be deduced that:
[0118]
[0119] Furthermore, it can be deduced that:
[0120]
[0121] Substitute inequality (7) into (5):
[0122]
[0123] where
[0124]
[0125] If then π > 0. Therefore, if the condition is satisfied, it can be concluded that: and are bounded. The stability analysis consists of two parts: t ∈ [t 0 , t 0 + T p ) and t ∈ [t 0 + T p , +∞). When t ∈ [t 0 , t 0 + T p ), from the definition of V 1 in the system, it can be deduced that:
[0126]
[0127] According to the definition of θ and Hypothesis 1, the following conclusion can be drawn:
[0128]
[0129] Substitute equation (10) into equation (9), and we can get:
[0130]
[0131] According to the definition of V 5 in (1), the inequality can be verified:
[0132]
[0133] According to (1), (11) and (12), it can be calculated that:
[0134]
[0135] Substituting the weight update rate and (2) - (13) into it, we can get:
[0136]
[0137] According to the proposed control law and Young's inequality, we can get:
[0138]
[0139] According to inequalities (6) and (7), it can be deduced that:
[0140]
[0141] Considering Hypothesis 2, we can get ε N is a positive constant; If γ 1 ≥ 1 and then we can get:
[0142]
[0143] Inequality (17) is equal to satisfying Therefore, there is and it can be deduced that When t ∈ [t 0 + T p , +∞), the control input is designed as:
[0144]
[0145] It can be deduced that:
[0146]
[0147] Due to the existence and continuity of the solution, we can get x(t 0 + T p ) = 0. According to x(t 0 + T p ) = 0 and the controller (18), we can get u(t 0 + T p ) = 0. In addition, it can be proved that f(0, x τ , θ) = 0, x(t 0 ) = 0, Therefore, the system is stable within a specified time, and the proof is complete.
[0148] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0149] The following is a simulation example.
[0150] To prove the effectiveness of the specified-time optimal control (ADPPTC) scheme based on ADP, the present invention will be compared with the adaptive specified-time control (APCPTC) scheme. The simulation results show that under different specified-time conditions, the scheme proposed by the present invention is superior to other control schemes in terms of system performance and control cost reduction. The non-linear time-varying time-delay system is selected as:
[0151]
[0152] where, θ a = 1.5, θ b = 0.5. Other functions involved are Φ(θ)=|θ 1 |+|θ 2 |, γ(x(t))=|x(t)|, η(x(t))=|x(t)|, λ(x τ (t))=|x τ (t)| and κ = 10.
[0153] τ(t) and v(s) are selected as:
[0154] τ(t)=0.4(1 + sin(t)) ≤ 0.8;
[0155]
[0156] Let ψ(x(t 0 + T p ), t 0 + T p ) = x T (t 0 + T p )x(t 0 + T p ).
[0157] To determine the optimal control strategy, a time-varying activation function is selected when the number of neurons L = 3, as follows:
[0158]
[0159] In this case, the parameter in the weight update law is γ 1 = 3, γ 2 = 0.5, γ 3 = 0.5. The initial value of the weight is W a (0) = W c (0) = [2, 2, 2] T .
[0160] Finally, the convergence performance of the system state and the controller under the ADPPTC and APCPTC schemes is as Figure 2 shown, and the cost function and the terminal constraint error under different strategies are as Figure 3 shown. It can be seen that under the action of the designed controller, the ADPPTC scheme achieves faster initial convergence. At the same time, it also reaches a lower stable value at the preset time.
[0161] The present invention is simulated under scenarios where different specified times T p = 1, 2, 3, 4 s, and δ = 2. The general scale and logarithmic scale of the system state at different specified times are as Figure 4 shown. The trajectories and adaptive parameters of the controller at different specified times are as Figure 5 shown. It can be seen that under any T p condition, the controller can converge to the origin within T p , and the convergence effect of the ADPPTC scheme is always better than that of the APCPTC scheme.
[0162] The present invention is simulated under scenarios where different δ = 2, 4, 6 and T p = 4 s, and the changes of the controller and the adaptive parameters are as Figure 6 and Figure 7 shown. u(t) converges to the origin within T p . In addition, as δ increases, the convergence speed of x(t) also increases. The system states when δ = 2 and T p = 0.2, 0.4, 0.6, 0.8 s are as Figure 8 shown. Under any T p condition, the convergence speed of the system state of the ADPPTC scheme is always faster than that of the APCPTC scheme.
[0163] The present invention is simulated under scenarios where different specified times T p = 1, 2, 3, 4 s, and δ = 2. The change curves of each system state at different specified times are as Figure 9 shown, and the trajectories and adaptive parameters of each controller at different specified times are as Figure 10 shown. As T pAs it increases, the convergence speed of each system state also increases, u 1 and u 2 The convergence speed also increases, and both can converge to the origin within T p seconds.
[0164] The present invention is simulated under scenarios with different δ = 1, 3, 5 and T p = 4 s. The change curves of each system state under different δ are as shown in Figure 11 Figure [specific figure number], and the trajectories and adaptive parameters of each controller under different δ are as shown in Figure 12 Figure [specific figure number]. As δ increases, the convergence speed of each system state also increases, u 1 and u 2 The convergence speed also increases, and both can converge to the origin within T p seconds.
[0165] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention. It should be noted that the specific figure numbers in the translation are placeholders and need to be filled in according to the actual figures in the original text.
Claims
1. A time-limited adaptive dynamic programming control method, characterized in that: The steps include: S1. Construct a nonlinear time-varying delay system model; S2, constructing a prescribed time cost function; S3, construct the HJB equation and calculate the optimal solution expression; S4, construct the judgment neural network and Hamiltonian equation; S5. Design approximate weight vector Approximate the time-varying cost function and calculate the approximate expression of the HJB equation; S6, solve the extreme value of the approximate HJB equation and minimize the residual and terminal error of the HJB equation; S7, updating the weights of the execution neural network and the evaluation neural network; S8, end, and obtain the optimal control law; In step S4, the specific process is: design the judgment neural network as To approximate the prescribed time cost function; in, is the weight vector of the judging neural network, is the activation function, ε(x,t) is the neural network approximation error; the terminal constraint is: In order to accurately estimate and approximate V x and V t , the design judgment neural network is: in, V x Substitution In, we get: Among them, W a represents the weight vector for executing the neural network, represents the weight vector of the execution neural network under the optimal strategy, In order to solve the optimal control strategy, the neural network approximation V x and V t Substitute into the HJB equation H(x,u * ), we get: in, is the residual of the HJB equation; In step S8, the optimal control law is: in, The continuous function μ(t) is the time adjustment (T p -PTA) function; that is, the following conditions are met: The stability of the nonlinear time-varying system using this control method within the specified time is proved.
2. The adaptive dynamic programming control method for a specified time according to claim 1, characterized in that: In step S1, the specific form is: in, are the state vector, control vector and uncertain parameter vector respectively; t0 represents the initial time; function f(t,x(t),x τ (t),θ) is locally Lipschitz for the state vector x, piecewise continuous for time t, and for any t ≥ 0, we have f(t,0,x τ (t),θ)=0;x τ (t) = x(t-τ(t)) represents the time-varying delay term, τ(t) represents the unknown time-varying delay, and its derivative satisfies v(s) represents a continuous function.
3. The adaptive dynamic programming control method for a specified time according to claim 2, characterized in that: The model must meet the following requirements: ∥f(t,x(t),x τ (t),θ)∥≤Φ(θ)(γ(x(t))∥x(t)∥+η(x(t))λ(x τ (t))∥x τ (t)∥, Among them, ∥·∥ represents the norm of the vector or the bi-norm of the matrix, γ(x(t)), η(x(t)) and λ(x τ (t)) is a continuous positive semidefinite function, and Φ(θ) represents a function related to θ.
4. The adaptive dynamic programming control method for a specified time according to claim 1, characterized in that: In step S2, establish the interval [t0, t0+T p The specific process of the cost function on ) is as follows: The design time cost function is: Among them, the function ψ(·) is the terminal constraint, T p is the preset time parameter, and the stage cost is: in, is a positive definite symmetric matrix.
5. The adaptive dynamic programming control method of a specified time according to claim 1, characterized in that: In step S3, the specific process is: establish the Hamiltonian equation as: Among them, u=u(t), g=g(x), f=f(t,x(t),x τ (t),θ), x=x(t), The cost function that defines the optimal stipulated time is: Solving the HJB equation The optimal solution expression is Right now: in, 6. The adaptive dynamic programming control method of a specified time according to claim 1, characterized in that: In step S5, the specific process is as follows: Since the time cost function will fluctuate over time, an approximate weight vector is designed To estimate the ideal weights and construct an approximate judgment neural network The approximate terminal constraints are: in Represents the estimated value of the state; V x and V t The approximate expression of is: The weight estimation error of the design judgment neural network is in, is the weight vector of the neural network under the optimal strategy, is the approximate weight vector of the judging neural network; since Unknown, u * Approximately Combining the weight estimation error, the HJB equation is transformed into By converting into The optimal strategy is solved in the form of .
7. The adaptive dynamic programming control method of a specified time according to claim 1, characterized in that: In step S6, the specific process is: in order to solve the equation design P(t)=0 is equivalent to Then the extreme value solution of the HJB equation is transformed into solving P(t) = 0; The estimated error of the constructed terminal constraint is: in, In order to minimize the terminal constraint error, the square constraint error is defined as 8. The method of adaptive dynamic programming control within a specified time according to claim 7, characterized in that: In step S7, the specific form is: executing the neural network weight update rule is: Where γ1 represents the learning rate of the execution neural network; The weight update rule for judging the neural network is: in, γ2 and γ3 are positive parameters, representing the learning rate of the neural network.
Citation Information
Patent Citations
Reconfigurable robot decentralized neural optimal control method based on judgment identification structure
CN109581868A
Finite time optimal tracking control method for uncertain nonlinear system
CN109976161A