Four-rotor unmanned aerial vehicle attitude optimization control method based on execution-evaluation network
By proposing an attitude optimization control method for quadrotor UAVs based on an execution-evaluation network, the problem of uncertain convergence time in the fixed-time tracking controller of quadrotor UAVs is solved. This method enables accurate tracking and optimized control under external disturbances, improves system performance, and reduces communication costs.
Patent Information
- Application Number
- CN202510892470.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-11
AI Technical Summary
In the existing design of tracking controllers for quadcopter drones, the upper bound of the theoretical convergence time for fixed-time tracking is uncertain, which causes the tracking error to fail to converge quickly and affects tracking performance.
A quadrotor UAV attitude optimization control method based on execution-evaluation network is adopted. By constructing a filtered compensation signal and a virtual controller, and combining reinforcement learning and Lyapunov stability theory, a Hamilton-Jacobi-Bellman equation for a predetermined time is designed. The attitude subsystem is decomposed and a Lyapunov function is constructed to achieve predetermined time stability and optimization control of the system.
It enables quadcopter UAVs to accurately track reference trajectories under external disturbances, improves transient and steady-state performance, reduces the impact of filtering errors, and reaches stability within a predetermined time, thereby reducing communication costs.
Smart Images

Figure CN120928839A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of quadcopter unmanned aerial vehicle (UAV) technology, specifically to a quadcopter UAV attitude optimization control method based on an execution-evaluation network. Background Technology
[0002] Unmanned aerial vehicles (UAVs) are short for unmanned aerial vehicles, which can be manually controlled by a remote controller or fully autonomously. They are mainly classified by their practical use into military UAVs, commercial UAVs, and civilian UAVs. Military UAVs are mainly used for weapon delivery, battlefield reconnaissance, electronic communication, and attack missions; commercial UAVs are mainly used for aerial photography, gaming, and formation performances; and civilian UAVs are mainly used for fire safety, agricultural protection, power maintenance, and environmental monitoring. Quadrotor UAVs can achieve tracking control faster by designing their settling time to ensure the tracking error meets the expected convergence time. Furthermore, the transient performance of quadrotor UAVs is also an important performance indicator in trajectory tracking control problems; small overshoot, fast convergence time, and small steady-state error enable quadrotor UAVs to quickly and smoothly track the reference trajectory. The invention patent application with application number CN 116954067B discloses a design method for a tracking controller for a quadcopter drone, which achieves fixed-time tracking error convergence constraints. However, the upper bound of the theoretical convergence time for the fixed time cannot be determined, which may prevent the tracking error of the quadcopter drone from converging quickly, thus affecting the inability to quickly achieve better tracking performance. Summary of the Invention
[0003] To address the aforementioned technical problems, this invention provides a quadrotor UAV attitude optimization control method based on an execution-evaluation network. An ideal trajectory tracking controller is designed to improve the transient and steady-state performance of the quadrotor UAV system, and it can still accurately track the reference trajectory even in the presence of external disturbances.
[0004] To achieve the above technical objectives, the adopted technical solution is: a quadcopter UAV attitude optimization control method based on an execution-evaluation network, comprising the following steps:
[0005] S1. Establish an attitude dynamics model of a quadrotor UAV containing external disturbances, and obtain the state-space equation of the attitude system of the quadrotor UAV based on the attitude dynamics model of the quadrotor UAV.
[0006] S2. Based on the state-space equation and adaptive command filtering back-propagation control method of the quadrotor UAV attitude system, construct a filter compensation signal x based on the defined state-space equation and adaptive command filtering back-propagation control method. i,1 and x i,2 The coordinate transformation equations are used to decompose the attitude subsystem of the quadrotor UAV into a two-level attitude subsystem.
[0007] S3. Construct a first-level optimal performance function for the first-level attitude subsystem, and based on the first-level optimal performance function, construct the first-level Hamilton-Jacobi-Bellman equations, introducing the cost function c. i,1 and intermediate cost function Using radial basis function neural networks to approximate the intermediate cost function By approximating the first-level Hamilton-Jacobi-Bellman equations using an execution-evaluation neural network, the execution-evaluation adaptive law is obtained. and And combine reinforcement learning to design a virtual controller α i ;
[0008] S4. Based on the first-level attitude subsystem and Lyapunov stability theory, construct the first Lyapunov function V. i,1 and V i,1 Differentiation yields By designing the virtual controller α i Execution-evaluation adaptive law and Substitution This makes the first-level attitude subsystem tend to be time-stable;
[0009] S5. Construct a second-level optimal performance function for the second-level attitude subsystem, and based on the second-level optimal performance function, construct the second-level Hamilton-Jacobi-Bellman equations, introducing the cost function c. i,2 and intermediate cost function Using radial basis function neural networks to approximate the intermediate cost function By approximating the second-level Hamilton-Jacobi-Bellman equation through an identification-execution-evaluation neural network, the learning law of the identification-execution-evaluation weights is obtained. And combine reinforcement learning to design virtual controllers
[0010] S6. Based on the second-level attitude subsystem and Lyapunov stability theory, a self-triggered control strategy is introduced to construct the second Lyapunov function V. i,2 and V i,2 Differentiation yields Design of the virtual controller Identification-Execution-Evaluation Weighted Learning Law Substitute V i,2 And by simplifying the inequalities, the second-level attitude subsystem tends to stabilize at a predetermined time.
[0011] Furthermore, the coordinate transformation equation is as follows:
[0012]
[0013] where i=1,2,3,e i,1 and e i,2 Let z be the tracking error and dummy error variables, respectively. i,1 and z i,2 Representing the error variable, x i,1 and x i,2 Represented as a filtered compensation signal, v i,d Represented as the desired tracking signal, (v1,v2,v3)=(φ,θ,ψ), where φ, θ, and ψ represent Euler angles;
[0014] Define the filtering error s i for:
[0015]
[0016] The filtered compensation signal x i,1 and x i,2 Defined as:
[0017]
[0018] Where η = η1 / η2, η1 is a positive even number, and η2 is a positive odd number, satisfying η1 < η2; and l i k i,1 k i,2 >0, for x i,1 (0)=x i,2 (0) = 0 and i = 1, 2, 3, T d Given a predetermined time and a positive design parameter, execute the following nonlinear instruction filter:
[0019]
[0020] Where r i >0 and H i ≥1 / 2 is a design parameter.
[0021] Furthermore, the first-order Hamilton-Jacobi-Bellman equation is:
[0022]
[0023] use The optimal controller can be obtained. The cost function is The intermediate cost function is introduced as follows:
[0024]
[0025] Among them, z i,1 As the first-level attitude subsystem, and l i r i k i,1 k i,2 >0, η = η1 / η2, where η1 is a positive even number and η2 is a positive odd number, satisfying η1 < η2; T d It is a predetermined time and a positive design parameter.
[0026] Furthermore, the virtual controller α i for
[0027]
[0028] in, express The estimated value, These are basis functions. v i,d This is represented as the desired tracking signal.
[0029] Furthermore, the execution-evaluation adaptive law and for
[0030]
[0031] in, and express and The estimated value, These are basis functions. v i,d Represented as the desired tracking signal, κ ai1 and κ ci1 The gain of the controller to be designed.
[0032] Furthermore, the first Lyapunov function V... i,1
[0033]
[0034] in, and It is an estimation error. and express and The estimated value.
[0035] Furthermore, the second-order Hamilton-Jacobi-Bellman equations...
[0036]
[0037] This allows for the optimization of virtual controllers. Furthermore, the cost function is: To achieve the optimal control objective, the following intermediate cost function is constructed:
[0038]
[0039] Among them, z i,2 This is the second-level attitude subsystem. and l i r i k i,1 k i,2 >0, η = η1 / η2, where η1 is a positive even number and η2 is a positive odd number, satisfying η1 < η2; i = 1, 2, 3. J xx J yy J zz This represents the moment of inertia of a quadcopter drone. It is the ideal weight vector, Φ i It is a basis function; δ i This indicates the approximation error.
[0040] Furthermore, the virtual controller for
[0041]
[0042] in, and l i r i k i,1 k i,2 >0, η = η1 / η2, where η1 is a positive even number and η2 is a positive odd number, satisfying η1 < η2; i = 1, 2, 3. J xx J yy J zz This represents the moment of inertia of a quadcopter drone.
[0043] Furthermore, the identification-execution-evaluation adaptive law and They are respectively
[0044]
[0045] Where, p i κ ai2 and κ ci2 The gain of the controller to be designed.
[0046] The beneficial effects of this invention are:
[0047] 1. This method adds a filtered compensation signal x i,1 and x i,2 By compensating for tracking error and virtual error variables, the impact of filtering error on control performance is eliminated, while also avoiding a series of problems such as "complexity explosion." Furthermore, by designing the filtering compensation signal, it is also made stable within a predetermined time.
[0048] 2. Design a predetermined time controller through steps S2-S5, wherein the settling time can be determined by a defined filtered compensation signal x. i,1 and x i,2 The pre-adjustment parameters η and T in d The size of the settling time determines the settling time, which differs from existing finite-time and fixed-time controller design methods. The settling time is no longer limited by the known initial state conditions, and can be calculated through design, making attitude control of quadcopter UAVs more precise.
[0049] 3. This invention achieves a balance between control cost and tracking performance by constructing HJB equations and designing an optimized control method. The proposed optimal control with a predetermined time can better achieve a balance between control performance and cost efficiency. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of the steps of the attitude optimization control method for a quadrotor UAV based on an execution-evaluation network according to the present invention;
[0051] Figure 2 This is an attitude tracking trajectory diagram of the quadcopter unmanned aerial vehicle system of the present invention;
[0052] Figure 3 This is a comparison chart of the attitude tracking errors of the quadcopter unmanned aerial vehicle system of this invention;
[0053] Figure 4 This is a graph showing the weights of the attitude execution-evaluation neural network for the first-level subsystem of a quadcopter UAV.
[0054] Figure 5 This is a graph showing the weights of the attitude execution-evaluation neural network for the second-level subsystem of a quadcopter UAV.
[0055] Figure 6 This is a comparison chart of attitude control input curves for quadcopter UAV systems;
[0056] Figure 7 This is a diagram showing the attitude self-trigger interval of the control input for a quadcopter UAV system.
[0057] Figure 8 This is a diagram showing the attitude event trigger intervals for the control input of a quadcopter UAV system.
[0058] Figure 9 This is a cost function comparison curve of the first-level subsystem of a quadcopter UAV;
[0059] Figure 10 This is a cost function comparison curve of the second-level subsystem of a quadcopter UAV. Detailed Implementation
[0060] The preferred embodiments of the invention are given below with reference to the accompanying drawings to illustrate the technical solution of the invention in detail. The corresponding drawings will be provided for detailed explanation of the invention. It should be particularly noted that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit or restrict the invention.
[0061] like Figure 1 As shown, the attitude optimization control method for a quadcopter UAV based on an execution-evaluation network includes the following steps:
[0062] Step 1: Establish the dynamic model of the quadcopter UAV containing external disturbances as follows:
[0063]
[0064] Where φ, θ, and ψ represent Euler angles; J xx J yy J zz G represents the moment of inertia of a quadcopter drone. φ G θ G ψ d represents the aerodynamic damping coefficient. φ d θ d ψ Represents an externally bounded disturbance, satisfying u φ u θ u ψ This represents the control torque.
[0065] Furthermore, based on the attitude dynamics model of the quadrotor UAV, the state-space equation of the quadrotor UAV attitude system is obtained:
[0066] The state-space equations are expressed as follows:
[0067]
[0068] Among them (v1, v2, v3) = (φ, θ, ψ), (u1, u2, u3) = (u φ ,u θ ,u ψ ), d1=d φ d2=d θ d3=d ψ ,
[0069] Step 2: Based on the state-space equations of the quadrotor UAV attitude system and the adaptive command filtering back-propagation control method, construct a filter compensation signal x based on the defined parameters. i,1 and x i,2 The coordinate transformation equations are used to decompose the attitude subsystem of the quadrotor UAV into a two-level attitude subsystem.
[0070] The expression for the coordinate transformation equation is as follows:
[0071]
[0072] where i=1,2,3,e i,1 and e i,2 Let z be the tracking error and dummy error variables, respectively. i,1 and z i,2 The variables represent error variables, representing the first-level attitude subsystem and the second-level attitude subsystem, respectively. i,1 and x i,2 Represented as a filtered compensation signal, v i,d The desired tracking signal is assumed to have a smooth and bounded derivative. The filtered output signal can be obtained from the design with a virtual controller α. i It is obtained from the filter.
[0073] Define the filtering error s i for:
[0074]
[0075] According to the aforementioned attitude optimization control method for quadrotor UAVs based on execution-evaluation networks, the adaptive command filtering back-propagation control method requires the addition of a filter compensation signal to eliminate the impact of filtering errors on system performance. The filter compensation signal x... i,1 and x i,2 Defined as:
[0076]
[0077] Where η = η1 / η2, η1 is a positive even number, and η2 is a positive odd number, satisfying η1 < η2; T d This is called the scheduled time, which is a positive design parameter. and l i k i,1 k i,2 >0, for x i,1 (0)=x i,2 (0) = 0 and i = 1, 2, 3, execute the following nonlinear instruction filter:
[0078]
[0079] Where r i >0 and H i ≥1 / 2 is a design parameter.
[0080] Step 3: Construct the first-level optimal performance function for the first-level attitude subsystem, and build the first-level Hamilton-Jacobi-Bellman equation based on the first-level optimal performance function, introducing the cost function c. i,1 and intermediate cost function Using radial basis function neural networks to approximate the intermediate cost function By approximating the first-level Hamilton-Jacobi-Bellman equations using an execution-evaluation neural network, the execution-evaluation adaptive law is obtained. and And combine reinforcement learning to design a virtual controller α i .
[0081] The first-level subsystem constructs the first-level optimal performance function as follows:
[0082]
[0083] Where Ω represents the set of all allowed control inputs, The ideal and optimal virtual controller.
[0084] The Hamilton-Jacobi-Bellman equation is further as follows:
[0085]
[0086] use The optimal controller can be obtained. The cost function is The intermediate cost function is introduced as follows:
[0087]
[0088] Then you can get and
[0089]
[0090] Next, we use a radial basis function neural network to approximate the first-level intermediate cost function:
[0091]
[0092] in It is the ideal weight vector. It is a basis function; the input vector is It is an approximation error.
[0093] Therefore, by using the radial basis function neural network to approximate the cost function to obtain the optimized virtual controller, we can get:
[0094]
[0095]
[0096] In the formula and express and The estimated value, of which and It is the estimation error.
[0097] The first-level HJB equation is solved by approximating it using an execution-evaluation neural network, resulting in the execution-evaluation adaptive law. and for:
[0098]
[0099] In the formula κ ai1 and κ ci1 The gain of the controller to be designed.
[0100] Step 4: Based on the first-level attitude subsystem and Lyapunov stability theory, construct the first Lyapunov function V. i,1 and V i,1 Differentiation yields By designing the virtual controller α i Execution-evaluation adaptive law and Substitution The simplified approach makes the first-level attitude subsystem tend to be time-stable.
[0101] Constructing Lyapunov function V for the first-level subsystem of a quadcopter unmanned aerial vehicle system model i,1 :
[0102]
[0103] Furthermore, the Lyapunov function V i,1 Taking the derivative with respect to time, we get:
[0104]
[0105] The optimal virtual controller α designed i Execution-evaluation adaptive law and Substituting into the above equation and simplifying, we get:
[0106]
[0107] Step 5: Construct the second-level optimal performance function for the second-level attitude subsystem, and based on the second-level optimal performance function, construct the second-level Hamilton-Jacobi-Bellman equation, introducing the cost function c. i,2 and intermediate cost function Using radial basis function neural networks to approximate the intermediate cost function By approximating the second-level Hamilton-Jacobi-Bellman equation through an identification-execution-evaluation neural network, the learning law of the identification-execution-evaluation weights is obtained. And combine reinforcement learning to design virtual controllers
[0108] Based on the coordinate transformation equation in step S2, we can obtain
[0109]
[0110] Based on an attitude control method for a quadrotor UAV using an execution-evaluation network, in step S5, the second-level attitude subsystem constructs the second-level optimal performance function as follows:
[0111]
[0112] The second-order Hamilton-Jacobi-Bellman equation is further defined as follows:
[0113]
[0114] This allows for the optimization of virtual controllers. Furthermore, the cost function is: To achieve the optimal control objective, the following intermediate cost function is constructed:
[0115]
[0116] In this process, fuzzy logic systems are used for approximation. in It is the ideal weight vector, Φ i It is a basis function; where δ i Describes the approximate error and satisfies
[0117] because There are unattainable ideal parameter values in the function, so we use neural networks to approximate the unknown function. in It is the ideal weight vector. These are basis functions. The approximation error is given; and the input vector is...
[0118] To achieve the optimal performance constraint control objective within a predetermined time, and Represented as
[0119]
[0120] Simultaneously, by utilizing an execution-evaluation network to obtain an optimized virtual controller, it is possible to achieve...
[0121]
[0122] In the formula and They are respectively and The estimation. Identification-Execution-Evaluation Adaptive Law and They are respectively:
[0123]
[0124] Where p i κ ai2 and κ ci2 The gain of the controller to be designed.
[0125] Step 6: Based on the second-level attitude subsystem and Lyapunov stability theory, a self-triggered control strategy is introduced to construct the second Lyapunov function V. i,2 and V i,2 Differentiation yields Design of the virtual controller Identification-Execution-Evaluation Weighted Learning Law Substitute V i,2 And by simplifying the equations using inequalities, the second-level attitude subsystem is brought to a stable state at a predetermined time.
[0126] To achieve the optimized control objective, a self-triggering control strategy is introduced into the controller design, and its triggering rule is expressed as follows:
[0127]
[0128] In the formula t k ,t k+1 ∈R + 0 <M i <1 and I i ,P i >0. M i It can be used to adjust the correlation weight between trigger time and controller, while I i The adjustment has a direct impact on the trigger time. From t=t k At the beginning, It will directly affect the system. Next trigger time t k+1 It can be calculated according to the above formula. During this time interval, the control signal u... i (t) always remains at t k The value of time. From the above formula, we can deduce...
[0129] u i (t k+1 )-u i (t)≤M i |u i (t)|+I i .
[0130] Assume that the time variable is a continuous function that satisfies and t∈[t k ,t k+1 Then we have:
[0131]
[0132] According to the self-triggering theory, the control rate The structure is as follows:
[0133]
[0134] in To be designed later, R i , ι i,1 For the gain of the controller to be designed, and R i >0
[0135] Construct a second Lyapunov function V that guarantees system stability i,2 :
[0136]
[0137] Calculate V i,2 The derivative is:
[0138]
[0139] The following inequalities hold:
[0140]
[0141] Using Yang's inequality, we get:
[0142]
[0143] Substituting Young's inequality and the above inequality into... Zhongde:
[0144]
[0145]
[0146] Step 7: Construct the global Lyapunov function and prove that the compensated signal tends to be time-stable. The global Lyapunov function is:
[0147]
[0148] The following inequalities hold:
[0149]
[0150] Differentiating the Lyapunov function of the whole system using the above inequality, we get:
[0151]
[0152] By definition and For matrix and Find the minimum eigenvalue and determine p. i The parameter corresponding to / 2>1, that is (2κ ci1 -1) / 4>0, (2κ ci2 -1) / 4>0 yields:
[0153]
[0154]
[0155] in and
[0156] Therefore, it can be inferred that all signals in the closed-loop system are actually predefined and time-bounded, and the variable e i,1 ,e i,2 ,Θ i ,Θ ai1 ,Θ ci1 ,Θ ai2 ,Θ ci2 In the following areas Within, in the steady-state time T d Inside.
[0157] Based on the proof that the compensation signal tends to be stable over a predefined time, in step S6, the Lyapunov function for constructing the compensation signal is:
[0158]
[0159] By differentiating the Lyapunov function of the compensation signal, we obtain:
[0160]
[0161] definition And design appropriate parameters l i Satisfy L i <l i L i Given a known parameter, we can obtain:
[0162]
[0163] in Similarly to the discussion above, therefore x i,1 x i,2 It tends to reach predefined time stability.
[0164] To illustrate the control effect of the method of the present invention in detail, a simulation experiment will be conducted in MATLAB, with the reference trajectory set as v. 1,d =1 / 4sin(πt / 8), v 2,d =1 / 4cos(πt / 7), v 3,d =1 / 4cos(πt / 9). For the initial state of the attitude system [φ(0),θ(0),ψ(0)]=[0.2,0.7,0.6], the external disturbance is defined as d φ =d θ =d ψ =0.01sin(πt / 25). The designed parameter is defined as: κ ci1 =κ ci2 =11,κ ai1 =κ ai2 =5(i=1,2,3); p1=p2=0.6, p3=0.3;k 1,1 =k 2,1 =k 3,1 =5,k 1,2 =0.01, k 2,2 =k 3,2 =10; H1=H2=H3=10; r1=40, r2=r3=10; R1=R2=R3=0.8; M1=M2=M3=0.05; η=2 / 25; T d =4. The model parameters of the quadcopter UAV are shown in Table 1.
[0165] Table 1: Model Parameters of Quadrotor UAVs
[0166] variable value unit variable value unit <![CDATA[J xx ]]> 0.0256 <![CDATA[Kg.m 2 ]]> <![CDATA[G φ ]]> 0.01 Kg / rad <![CDATA[J yy ]]> 0.0256 <![CDATA[Kg.m 2 ]]> <![CDATA[G θ ]]> 0.01 Kg / rad <![CDATA[J zz ]]> 0.0489 <![CDATA[Kg.m 2 ]]> <![CDATA[G ψ ]]> 0.01 Kg / rad
[0167] The following results were obtained through simulation experiments using MATLAB. Figure 2 This is an attitude tracking trajectory diagram of the quadcopter unmanned aerial vehicle system of the present invention; Figure 3 This is a comparison chart of the attitude tracking errors of the quadcopter unmanned aerial vehicle system of this invention; Figure 3 Comparison 1 and Comparison 2 represent finite-time and fixed-time controllers, respectively. The difference lies in the fact that, compared to the predetermined-time controller designed in this patent, the convergence speed and settling time of the predetermined-time controller can be adjusted by modifying η and T. d The size is used to control it. The result is as follows: Figure 2 and Figure 3 As shown, the proposed predetermined time control mechanism is known to be in place at a predetermined time T. d It provides system stability. Figure 4 This is a graph showing the weights of the attitude execution-evaluation neural network for the first-level subsystem of a quadcopter UAV. Figure 5 This is a graph showing the weights of the attitude execution-evaluation neural network for the second-level subsystem of a quadcopter UAV. Figure 6 This is a comparison chart of the attitude control input curves of a quadcopter UAV system; it can be seen that the input curve converges faster than the other two conventional methods. Figure 7 This is a diagram showing the attitude self-trigger interval of the control input for a quadcopter UAV system.
[0168] Figure 8 This is a diagram showing the attitude event trigger intervals for the control input of a quadcopter UAV system. A comparison of the trigger counts and intervals shows that the designed control method effectively avoids continuous updates of the control signal, reducing the communication burden. Through simulation experiments, the self-triggering method, based on reinforcement learning, proposed in this invention, ensures that all signals are bounded and the control system remains stable even with limited communication bandwidth. Figure 9 and Figure 10 The figures show a comparison of the cost functions of the first and second stage systems of the quadcopter UAV. Compared with finite-time and fixed-time control methods, the figures show that the lower cost function of this patent application results in lower communication costs and achieves a balance between communication costs and energy consumption, indicating that this control strategy is superior to traditional control methods.
[0169] The above are merely preferred embodiments of the present invention and are not intended to limit or restrict the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection declared by the present invention.
Claims
1. A method for attitude optimization control of a quadrotor UAV based on an execution-evaluation network, characterized in that, Includes the following steps: S1. Establish an attitude dynamics model of a quadrotor UAV containing external disturbances, and obtain the state-space equation of the attitude system of the quadrotor UAV based on the attitude dynamics model of the quadrotor UAV. S2. Based on the state-space equation and adaptive command filtering back-propagation control method of the quadrotor UAV attitude system, construct a filter compensation signal x based on the defined state-space equation and adaptive command filtering back-propagation control method. i,1 and x i,2 The coordinate transformation equations are used to decompose the attitude subsystem of the quadrotor UAV into a two-level attitude subsystem. S3. Construct a first-level optimal performance function for the first-level attitude subsystem, and based on the first-level optimal performance function, construct the first-level Hamilton-Jacobi-Bellman equations, introducing the cost function c. i,1 and intermediate cost function Using radial basis function neural networks to approximate the intermediate cost function By approximating the first-level Hamilton-Jacobi-Bellman equations using an execution-evaluation neural network, the execution-evaluation adaptive law is obtained. and And combine reinforcement learning to design a virtual controller α i ; S4. Based on the first-level attitude subsystem and Lyapunov stability theory, construct the first Lyapunov function V. i,1 and V i,1 Differentiation yields By designing the virtual controller α i Execution-evaluation adaptive law and Substitution This makes the first-level attitude subsystem tend to be time-stable; S5. Construct a second-level optimal performance function for the second-level attitude subsystem, and based on the second-level optimal performance function, construct the second-level Hamilton-Jacobi-Bellman equations, introducing the cost function c. i,2 and intermediate cost function Using radial basis function neural networks to approximate the intermediate cost function By approximating the second-level Hamilton-Jacobi-Bellman equation through an identification-execution-evaluation neural network, the learning law of the identification-execution-evaluation weights is obtained. And combine reinforcement learning to design virtual controllers S6. Based on the second-level attitude subsystem and Lyapunov stability theory, a self-triggered control strategy is introduced to construct the second Lyapunov function V. i,2 and V i,2 Differentiation yields Design of the virtual controller Identification-Execution-Evaluation Weighted Learning Law Substitute V i,2 And by simplifying the equations using inequalities, the second-level attitude subsystem is brought to a stable state at a predetermined time.
2. The attitude optimization control method for a quadrotor UAV based on an execution-evaluation network as described in claim 1, characterized in that: The coordinate transformation equation is as follows: where i=1,2,3,e i,1 and e i,2 Let z be the tracking error and dummy error variables, respectively. i,1 and z i,2 Representing the error variable, x i,1 and x i,2 Represented as a filtered compensation signal, v i,d Represented as the desired tracking signal, (v1,v2,v3)=(φ,θ,ψ), where φ, θ, and ψ represent Euler angles; Define the filtering error s i for: The filtered compensation signal x i,1 and x i,2 Defined as: Where η = η1 / η2, η1 is a positive even number, and η2 is a positive odd number, satisfying η1 < η2; and l i k i,1 k i,2 >0, for x i,1 (0)=x i,2 (0) = 0 and i = 1, 2, 3, T d Given a predetermined time and a positive design parameter, execute the following nonlinear instruction filter: Where r i >0 and H i ≥1 / 2 is a design parameter.
3. The attitude optimization control method for a quadrotor UAV based on an execution-evaluation network as described in claim 1, characterized in that: The first-order Hamilton-Jacobi-Bellman equation is: use The optimal controller can be obtained. The cost function is The intermediate cost function is introduced as follows: Among them, z i,1 As the first-level attitude subsystem, and l i r i k i,1 k i,2 >0, η = η1 / η2, where η1 is a positive even number and η2 is a positive odd number, satisfying η1 < η2; T d It is a predetermined time and a positive design parameter.
4. The attitude optimization control method for a quadrotor UAV based on an execution-evaluation network as described in claim 1, characterized in that: The virtual controller α i for in, express The estimated value, These are basis functions. v i,d This is represented as the desired tracking signal.
5. The attitude optimization control method for a quadrotor UAV based on an execution-evaluation network as described in claim 1, characterized in that: Execution-evaluation adaptive law and for in, and express and The estimated value, These are basis functions. v i,d Represented as the desired tracking signal, κ ai1 and κ ci1 The gain of the controller to be designed.
6. The attitude optimization control method for a quadrotor UAV based on an execution-evaluation network as described in claim 1, characterized in that: The first Lyapunov function V i,1 in, and It is an estimation error. and express and The estimated value.
7. The attitude optimization control method for a quadrotor UAV based on an execution-evaluation network as described in claim 1, characterized in that: The second-order Hamilton-Jacobi-Bellman equations This allows for the optimization of virtual controllers. Furthermore, the cost function is To achieve the optimal control objective, the following intermediate cost function is constructed: Among them, z i,2 This is the second-level attitude subsystem. and l i r i k i,1 k i,2 >0, η = η1 / η2, where η1 is a positive even number and η2 is a positive odd number, satisfying η1 < η2; i = 1, 2, 3. J xx J yy J zz This represents the moment of inertia of a quadcopter drone. It is the ideal weight vector, Φ i It is a basis function; δ i This indicates the approximation error.
8. The attitude optimization control method for a quadrotor UAV based on an execution-evaluation network as described in claim 1, characterized in that: The virtual controller for in, and l i r i k i,1 k i,2 >0, η = η1 / η2, where η1 is a positive even number and η2 is a positive odd number, satisfying η1 < η2; i = 1, 2, 3. J xx J yy J zz This represents the moment of inertia of a quadcopter drone.
9. The attitude optimization control method for a quadrotor UAV based on an execution-evaluation network as described in claim 1, characterized in that: Identification-Execution-Evaluation Adaptive Law and They are respectively Where, p i , κ ai2 and κ ci2 The gain of the controller to be designed.
Citation Information
Patent Citations
A tracking controller design method for a quadrotor drone
CN116954067B