A Q-learning H-infinity anti-interference control method for fixed-wing UAV
By combining H∞ theory and Q learning algorithm, a model-free anti-interference control law is designed, which solves the problem of insufficient anti-interference capability of the Q learning algorithm in the face of external disturbances, and achieves higher anti-interference capability and system performance of the drone.
Patent Information
- Application Number
- CN202211023187.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-25
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-08-25
AI Technical Summary
When existing Q-learning algorithms face external disturbances, their anti-interference ability is weak, resulting in a degradation of UAV control performance.
Combined with H∞ theory, a model-free anti-interference control law based on input, output and perturbation input data is designed. Through iterative solution of the output feedback Q function and the optimal kernel matrix, the control input is optimized to suppress perturbation.
It improves the anti-interference capability and system performance of the drone, and enhances the control stability and response performance under external disturbance conditions.
Smart Images

Figure CN115407654B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of unmanned aerial vehicle flight control, and in particular relates to a Q-learning H-infinity anti-interference control method for fixed-wing unmanned aerial vehicle. Background Art
[0002] At present, there are many kinds of UAV flight control algorithms, such as the classic PID algorithm with simple control structure and principle, model predictive control algorithm and sliding film control algorithm, which are all model-based control methods. In response to the model-free problem, some intelligent control algorithms have also been applied to UAV flight control systems, such as the Q learning algorithm. The Q learning algorithm can use the system input and output data to complete the adaptive adjustment of the control law. It is an intelligent control algorithm, but in the actual flight process, external disturbances are inevitable. Under some disturbances, the adjustment ability of the optimal controller designed by the Q learning algorithm is reduced, and the UAV control performance is relatively worse than when there is no disturbance input.
[0003] In view of the increasingly diverse flight missions and changing flight environments of drones, drones need to improve their anti-interference capabilities in order to successfully complete flight missions. Drones need to be able to adaptively adjust flight control laws according to changes in external disturbance inputs, thereby demonstrating certain intelligent characteristics. Summary of the invention
[0004] Purpose of the invention: The technical problem to be solved by the present invention is to provide a Q-learning fixed-wing UAV H infinite anti-interference control method in view of the shortcomings of the prior art. ∞ A model-free anti-interference control law based on input, output and disturbance input data is theoretically designed. This method not only continues the data-driven advantage of the Q-learning algorithm, but also has a suppressive effect on disturbances that meet the disturbance attenuation conditions, thereby improving the anti-interference ability and system performance of the UAV.
[0005] The method of the present invention comprises the following steps:
[0006] Step 1: Model the UAV and establish a UAV longitudinal state space model with disturbance terms;
[0007] Step 2: Under the above-mentioned observable conditions of the UAV longitudinal system state, reconstruct the system state using the system control input, observation output and disturbance input data;
[0008] Step 3: Based on the reconstructed state information, combined with H ∞ The output feedback Q function is theoretically constructed. The Q function is differentiated with respect to the disturbance term and the control input to obtain the optimal control input and the worst disturbance expression under disturbance.
[0009] Step 4: Parameterize the output feedback Q function, collect system input, output and disturbance data, and iterate online to solve the optimal kernel matrix;
[0010] Step 5: Use the optimal control input and the worst-case disturbance expression, combined with the optimal kernel matrix, to calculate the optimal control input and the worst-case disturbance input.
[0011] Step 6: Using the disturbance attenuation condition, select the disturbance input that meets the disturbance condition, add the disturbance input to the system, and verify the effective suppression of the learned anti-disturbance controller on the disturbance input under the disturbance attenuation condition;
[0012] Preferably, the nonlinear differential equation of the longitudinal motion of the UAV is:
[0013]
[0014]
[0015]
[0016]
[0017] The state variables V, α, θ, and q represent the velocity, angle of attack, pitch angle, and pitch rate, respectively. The control input δ e represents the elevator deflection angle; and t represents time; I y is the moment of inertia around the y-axis of the body coordinate system, T is the engine thrust, g is the acceleration of gravity, M is the pitch moment, and m is the mass of the drone.
[0018] Preferably, in step 1, after obtaining the longitudinal nonlinear differential equation of the UAV, the longitudinal dynamics model is written using the Matlab / simulink tool using a matlab script file, the trim() function is used to find the UAV working point for balancing, the linmod() function is used for linearization to obtain the state space expression, and the c2d() function is used for discretization of the continuous system, and finally the following form is obtained:
[0019] x k+1 =Ax k +Bu k
[0020] y k =Cx k ,
[0021] in n1, n, and p are the control input dimension, state variable dimension, real number set and system output dimension respectively; A, B and C are the system matrix, control matrix and observation matrix respectively; xk Contains four state variables of the drone, namely V, α, θ and q; subscript k represents the current moment, subscript k+1 represents the next moment; y k is the output value of the UAV system, and then the disturbance term is introduced into the state space model, and the state space expression becomes:
[0022] x k+1 =Ax k +Bu k +Ew k
[0023] y k =Cx k ,
[0024] in is the disturbance input term added to the system, n2 is the disturbance input dimension, and E is the disturbance input matrix.
[0025] Preferably, step 2 includes: constructing state variables through information of input, output and disturbance input:
[0026]
[0027] Among them, M u 、M w and M y are the control input reconstruction matrix, disturbance input reconstruction matrix and system output reconstruction matrix respectively, is the control input vector composed of control inputs at different times, is the output vector composed of the disturbance input at different times, is the output vector consisting of the system output at different times, and
[0028]
[0029]
[0030] and denote the control input, disturbance input and system output at time kN respectively; the superscript T denotes matrix transposition; D N1 and D N2 are the control input coupling matrix and the disturbance input coupling matrix respectively; U N and V N Respectively represent the controllability and observability matrices of the system; N is the step size of the past moment, and N satisfies N≥n; each symbol is expressed as follows:
[0031]
[0032] U N =[B AB … AN-1 B],V N =[(CA N-1 ) T … (CA) T C T ] T ,
[0033]
[0034] Preferably, in step 3, H ∞ theory, design output feedback Q function Q(x k ,w k ,u k )for:
[0035]
[0036] Where Q1 and R are weight matrices, γ d is the disturbance attenuation factor, γ is the discount factor, 0<γ≤1; P is the solution of the generalized Riccati equation, z k is a vector composed of control input, disturbance input and system output in the past N moments, and the output feedback Q function Q(x k ,w k ,u k ) is further written as:
[0037]
[0038] Where H is the kernel matrix of the Q function and satisfies l=n1N+n2N+pN+n1+n2,H (△)(□) The performance function matrix representing △ and □, such as for and The coupling performance function matrix is as follows:
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045] H uu =R+γB TPB, H uw =γB T PE,
[0046]
[0047] Preferably, in step 4, the state space expression of the UAV longitudinal model is used to generate the system input, output and disturbance input data within the required [kN, k-1] time step. k-1 ,u k-2 ,...,u k-N ,w k-1 ,w k-2 ,...,w k-N ,y k-1 ,y k-2 ,...,y k-N , based on the output feedback Q function H ∞ The policy iteration algorithm learns the optimal kernel matrix H.
[0048] Preferably, the output feedback Q learning H ∞ Policy iteration algorithm, including:
[0049] Step 4-1, initialization: H 0 , Q y , R, γ, γ d Initialization assignment, H 0 is the initialization value of H, j is initialized to 1, Q y is the weight matrix;
[0050] Step 4-2, strategy evaluation: collect the input, output and disturbance input data generated by the system and solve the output feedback Q function H ∞ Bellman equation, collection interval step size L and Bellman equation are as follows:
[0051] L≥(n1N+n2N+pN+n1+n2)×(n1N+n2N+pN+n1+n2+1) / 2:
[0052]
[0053] Among them, H j is the H matrix of the jth iteration;
[0054] Step 4-3, strategy update: Update the control input and disturbance input in a data-driven manner:
[0055]
[0056]
[0057] and are the control input and disturbance input at the j+1th iteration of the kth step, respectively;
[0058] Step 4-4, termination condition: when the convergence condition || H is met j -H j-1 The iteration is stopped when ||≤ε, ε is the convergence condition constant, and H at the termination of the iteration is the optimal kernel matrix; if the convergence condition is not met, j=j+1 is executed and the iteration continues.
[0059] Preferably, in step 5, the elements in the optimal kernel matrix H are combined into the form of the optimal control law and the worst disturbance, and the optimal control input and the worst disturbance input are expressed as:
[0060]
[0061]
[0062] Above Represents the submatrix contained in the optimal kernel matrix H corresponding to the termination condition.
[0063] Preferably, the disturbance attenuation condition is used to select the disturbance input that meets the disturbance condition. The disturbance attenuation condition is expressed as
[0064]
[0065] After the disturbance selection is completed, the disturbance input is added to the system to verify the effective suppression of the disturbance input under the disturbance attenuation condition by the learned anti-disturbance controller.
[0066] Beneficial effects:
[0067] 1. The present invention relates to a Q-learning fixed-wing drone H ∞ The anti-interference control method is to control the ∞ The invention is an improvement of the Q learning algorithm. Generally, the Q learning algorithm designs the optimal control law online without considering external disturbances. Although Q learning is a model-free data-driven method that can adaptively adjust the control gain online, the control effect is often not ideal for some large disturbances, and even causes system oscillation and divergence. The invention combines Q learning with H ∞ Combining theory with practice, the disturbance input is also taken into account when designing the control law, which greatly reduces the impact of the disturbance on the system and improves the system response performance.
[0068] 2. The invention relates to a Q-learning fixed-wing drone H ∞The anti-interference control method, like Q-learning, is a model-free method, that is, in the process of control law design, it is necessary to collect data on system input, output and disturbance input, and use these data to design the anti-interference optimal controller online. When the external disturbance changes, the control law will also be adaptively adjusted. This method makes the UAV flight control system more intelligent.
[0069] 3. The present invention uses the system and disturbance generation data to learn the anti-interference control law, which can solve the flight problem of the UAV when there is external disturbance, such as wind disturbance. In addition, the model-free characteristic of Q learning itself can effectively solve the complex, uncertain and even unknown UAV model. The simulation comparison between the present invention and the Q learning method proves that Q learning is more effective in combining H ∞ The ability to suppress disturbances is better than Q learning. The verification on nonlinear models also shows the effectiveness and feasibility of the present invention.
[0070] 4. The anti-interference control method described in the present invention is essentially to solve the linear quadratic problem (LQR) of the discrete system. The kernel matrix of the Q function Bellman equation is simple and fast to solve. Therefore, this method is suitable for aircraft with weak computing power and small storage space. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 It is a flow chart for realizing the method of the present invention;
[0072] Figure 2 Yes Q Learning H ∞ The angle of attack variation curve of the algorithm;
[0073] Figure 3 Yes Q Learning H ∞ The pitch angle variation curve of the algorithm;
[0074] Figure 4 Yes Q Learning H ∞ The pitch angle rate change curve of the algorithm;
[0075] Figure 5 Yes Q Learning H ∞ The elevator deflection angle variation curve of the algorithm;
[0076] Figure 6 Yes Q Learning H ∞ Schematic diagram of the convergence of the elements in the algorithm H matrix related to the control input;
[0077] Figure 7 Yes Q Learning H ∞ Schematic diagram of the convergence of the elements in the algorithm H matrix related to the disturbance input;
[0078] Figure 8 It is to learn H matrix and H online * Schematic diagram of norm change;
[0079] Fig. 9 is the disturbance attenuation change curve of the selected disturbance;
[0080] Fig.10 is to select Q learning H under disturbance ∞ Angle of attack change curve compared with the algorithm and Q learning algorithm;
[0081] Fig.11 is to select Q learning H under disturbance ∞ Pitch angle change curve compared with the algorithm and Q learning algorithm;
[0082] Fig.12 is to select Q learning H under disturbance ∞ Pitch angle rate change curve compared with the algorithm and Q learning algorithm;
[0083] Fig.13 is to select Q learning H under disturbance ∞ The elevator deflection angle change curve under the comparison of the algorithm and the Q learning algorithm;
[0084] Fig.14 It is the verification of the anti-disturbance control law learned under disturbance on the nonlinear model (angle of attack response comparison);
[0085] Fig.15 It is the verification of the anti-disturbance control law learned under disturbance on the nonlinear model (pitch angle response comparison); Fig.16 It is the verification of the anti-disturbance control law learned under disturbance on the nonlinear model (pitch angle rate response comparison). DETAILED DESCRIPTION
[0086] The present invention will be further explained below in conjunction with the accompanying drawings.
[0087] like Figure 1 As shown, the present invention provides a Q-learning fixed-wing drone H ∞ Anti-interference control methods include:
[0088] Step 1: Model the UAV and establish a UAV longitudinal state space model with disturbance terms;
[0089] The nonlinear differential equation of the longitudinal motion of the UAV is:
[0090]
[0091]
[0092]
[0093]
[0094] The state variables V, α, θ, and q represent the velocity, angle of attack, pitch angle, and pitch rate, respectively. The control input δ e represents the elevator deflection angle; and t represents time; I y is the moment of inertia around the y-axis of the body coordinate system, T is the engine thrust, g is the acceleration of gravity, M is the pitch moment, and m is the mass of the drone.
[0095] After the drone is balanced, the balance state is V trim =53.6448 m / s, α trim =0.566°, T trim =1354.3 Newtons, keeping the speed V = 53.6448 m / s as a constant value, the discrete state space of the UAV after linearizing the working point and adding the disturbance term is expressed as
[0096]
[0097] where x k =[α θ q] T ,u k =δ e .
[0098] Assuming that the system sensor can only obtain information about the pitch angle, the anti-interference controller is designed through the data of the pitch angle, elevator, and disturbance input. The system output is as follows:
[0099] y k =[0 1 0]x k
[0100] Step 2: The present invention is H for output feedback Q learning ∞ Anti-interference control method, assuming that the system state information cannot be fully obtained, uses the system input and output and disturbance data to reconstruct the state information. In this example, N is 3:
[0101]
[0102] and
[0103]
[0104]
[0105] The symbols are expressed as follows:
[0106]
[0107] U N =[B AB A 2 B],V N =[(CA2 ) T (CA) T C T ] T ,
[0108]
[0109] Step 3: Combine H ∞ Theory, design an output feedback Q function:
[0110]
[0111] remember P is the solution of the generalized Riccati equation, and the output feedback Q function Q(x k ,w k ,u k ) is further written as:
[0112]
[0113] It is expressed as follows:
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120] H uu =R+γB T PB, H uw =γB T PE,
[0121]
[0122] Design parameter Q y =1, R=1, γ d =1, γ=1, calculate Q1 and the generalized Riccati equation:
[0123]
[0124] After calculating the optimal kernel matrix, the submatrices of the optimal control input and the worst disturbance input are:
[0125]
[0126]
[0127]
[0128]
[0129] Step 4: Use the state space expression of the UAV longitudinal model to generate the system input, output, and disturbance input data within the required [kN, k-1] time step. k-1 ,u k-2 ,...,u k-N ,w k-1 ,w k-2 ,...,w k-N ,y k-1 ,y k-2 ,...,y k-N , based on the output feedback Q function H ∞ The policy iteration algorithm learns and obtains the optimal kernel matrix H. When the termination condition is met, the iteration stops and the optimal kernel matrix H is obtained. Since N = 3, L≥66, here L is taken as 70; H 0 =0.01*I, I is the unit matrix. ;
[0130] Step 5: The elements in the optimal kernel matrix H are combined into the form of optimal control input and worst disturbance input. The optimal control input and worst disturbance input are expressed as:
[0131]
[0132]
[0133] The steps of the policy iteration algorithm based on output feedback Q learning are as follows:
[0134] 1) Initialization: H 0 , Q y , R, γ, γ d Initialize the value, j=1, and select the stable initial control strategy u0;
[0135] 2) Strategy evaluation: Collect the input, output, and disturbance input data generated by the system and solve the output feedback Q function H ∞ Bellman equation:
[0136]
[0137] 3) Strategy update: Update control input and disturbance input in a data-driven manner:
[0138]
[0139]
[0140] 4) Termination condition: When the convergence condition || H is met j -H j-1 The iteration stops when ||≤ε, where ε is the convergence condition constant. If the convergence condition is not met, the iteration continues after j=j+1.
[0141] The partial matrix of the optimal kernel matrix H obtained by online learning is:
[0142]
[0143]
[0144]
[0145]
[0146] In step 6, the disturbance attenuation condition is used to select the disturbance input that meets the disturbance condition. The disturbance attenuation condition is expressed as
[0147]
[0148] Choose the disturbance as w k = sin(k)e -0.005k In steady state, the value of the above formula is 0.1, which is less than γ d The setting of 1 satisfies the condition, indicating that the designed anti-disturbance control law can suppress the disturbance. After the disturbance selection is completed, the disturbance input is added to the system to verify the effective suppression of the learned anti-disturbance controller on the disturbance input under the disturbance attenuation condition.
[0149] The Q-learning of this invention ∞ The anti-interference control method is an improvement on the Q learning algorithm. Generally, the Q learning algorithm designs the optimal control law online without considering external disturbances, and its anti-interference ability is poor. ∞ After combining the theory, this method can evaluate the external disturbance input online and adaptively adjust the control input to make the system response optimal. In addition, relying on the model-free advantage of the Q learning algorithm, this method learns the control law in a data-driven environment. The system and disturbance input only need to obtain the data of their input, output and disturbance input. Figure 1 Design a flow chart for the present invention; Figure 2 Schematic diagram of angle of attack change of the method of the present invention; Figure 3 Schematic diagram of pitch angle variation of the method of the present invention; Figure 4It is a schematic diagram of the pitch angle rate change of the method of the present invention; the three states of the UAV longitudinal model are all stable at the 50th step, and the adjustment effect is good. Figure 5 is the schematic diagram of the elevator deflection angle, the deflection range is in the interval [-7.6°, 0]; Figure 6 and Figure 7 Learn H for Q ∞ Schematic diagram of the convergence of the control input-related items and disturbance input-related items in the algorithm's H matrix; Figure 8 It is a schematic diagram of the norm convergence of the kernel matrix H; Fig. 9 To select the disturbance attenuation of the disturbance, the disturbance attenuation is 0.1 in the steady state, which is less than γ d The setting of 1 can suppress this disturbance; Fig.10 , Fig.11 , Fig.12 and Fig.13 The method in this paper is different from that of unbound H ∞ The state response curves of the Q-learning algorithm are compared. It can be seen from the figure that the three state fluctuations of this paper are smaller than those of the Q-learning method, and the curve of the elevator deflection angle is obviously flatter, so it has a stronger disturbance suppression ability; Fig.14 , Fig.15 , Fig.16 In order to apply the anti-interference control law learned by the method of the present invention to the nonlinear system of the UAV and compare it with the linear system, it can be seen from the figure that under the selected disturbance input, the response curves of the linear system and the nonlinear system are very consistent, and the nonlinear system has an anti-interference performance that is very close to that of the linear system, which shows that the anti-interference controller designed by the method of this paper is also effective for the actual nonlinear system of the UAV.
[0150] The present invention provides a Q-learning fixed-wing unmanned aerial vehicle H ∞ Anti-interference control method, there are many methods and ways to implement the technical solution. The above is only the preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies. The above is only the preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A Q-learning H-infinity anti-interference control method for fixed-wing UAV, characterized in that: The following steps are involved: 1) Establishing a spatial model of the longitudinal system state of the UAV, wherein the spatial model of the longitudinal system state of the UAV includes a disturbance term; 2) Under the condition that the longitudinal system state of the UAV is observable, the longitudinal system state of the UAV is reconstructed through control input, observation output and disturbance input data; 3) Combined with H ∞ Theoretical construction of output feedback Q function; 4) Parameterize the output feedback Q function, collect the input, output and disturbance data of the UAV longitudinal system, and solve the optimal kernel matrix through online iteration; 5) Using the optimal control input and the worst-case disturbance expression, combined with the optimal kernel matrix, the optimal control input and the worst-case disturbance input are calculated; 6) Using the disturbance attenuation condition, select the disturbance input that meets the disturbance condition, add the disturbance input to the longitudinal system of the UAV, and verify the effective suppression of the disturbance input under the disturbance attenuation condition by the learned anti-interference controller; In step 1), the spatial model of the longitudinal system state of the UAV excluding the disturbance term is first constructed, and its expression is: x k+1 =Ax k +Bu k y k =Cx k , Among them: the state quantity at time k UAV system output value n1, n, and p are the control input dimension, state variable dimension, real number set and UAV longitudinal system output dimension respectively; A, B and C are the state transfer matrix, control input matrix and observation matrix respectively; x k Contains four state variables of the drone, namely speed V, angle of attack α, pitch angle θ and pitch angle rate q; The four state variables of the drone are: in t represents time; I y is the moment of inertia around the y-axis of the body coordinate system, T t is the engine thrust of the drone, D is the drag, L is the lift, g is the acceleration of gravity, M is the pitch moment, and m is the mass of the drone; Then the spatial model of the UAV longitudinal system state including the disturbance term is: x k+1 =Ax k +Bu k +Ew k y k =Cx k , in is the disturbance input term added to the UAV longitudinal system, n2 is the disturbance input dimension, and E is the disturbance input matrix; In step 2), the longitudinal system state of the UAV is reconstructed as: Among them, M u 、M w and M y They are the control input reconstruction matrix, the disturbance input reconstruction matrix and the UAV longitudinal system output reconstruction matrix, is the control input vector composed of control inputs at different times, is the output vector composed of the disturbance input at different times, is the output vector composed of the longitudinal system output of the UAV at different times, and: U N =[B AB … A N-1 B],V N =[(CA N-1 ) T … (CA) T C T ] T , Among them: U N and V N Denote the controllability and observability matrices of the UAV longitudinal system respectively; D N1 and D N2 are the control input coupling matrix and the disturbance input coupling matrix respectively; N is the step size at the past moment, N satisfies N≥n, and n is the dimension of the state variable; and They represent the control input, disturbance input and UAV longitudinal system output at time kN respectively; T represents the transpose of the matrix, A N-1 , A N-2 , A N-3 Represents the N-1, N-2, N-3 power of the state transfer matrix A; In step 3): the output feedback Q function is: Where Q1 and R are weight matrices, γ d is the disturbance attenuation factor, γ is the discount factor, 0<γ≤1; P is the solution of the generalized Riccati equation, z k is a vector consisting of control input, disturbance input and UAV longitudinal system output at the past N moments with time k as the base point. The Q function is rewritten as: Where: H is the kernel matrix of the output feedback Q function and satisfies Parameter l = n1N + n2N + pN + n1 + n2; n1 is the control input dimension, n2 is the disturbance input dimension, and p is the output dimension; for The performance coupling matrix, for and The performance coupling matrix, for and The performance coupling matrix, for and u k The performance coupling matrix, for and w k Performance coupling matrix; for and The performance coupling matrix, for The performance coupling matrix, for and The performance coupling matrix, for and u k The performance coupling matrix, for and w k Performance coupling matrix; for and The performance coupling matrix, for and The performance coupling matrix, for The performance coupling matrix, for and u k The performance coupling matrix, for and w k Performance coupling matrix; for u k and The performance coupling matrix, for u k and The performance coupling matrix, for u k and The performance coupling matrix, H uu for u k The performance coupling matrix, H uw for u k and w k Performance coupling matrix; w k and The performance coupling matrix, w k and The performance coupling matrix, w k and The performance coupling matrix, H wu w k and u k The performance coupling matrix, H ww w k The performance coupling matrix of ; I is the unit matrix; In step 4), the spatial model expression of the longitudinal system state of the UAV is used to generate the system input, output and disturbance input data within the required [kN, k-1] time step, and the obtained: u k-1 ,u k-2 ,...,u k-N ,w k-1 ,w k-2 ,...,w k-N ,y k-1 ,y k-2 ,...,y k-N , Based on the output feedback Q function H ∞ Strategy iteration algorithm, learn the optimal kernel matrix H 最优 ; Based on the output feedback Q function H ∞ Strategy iteration algorithm, learn the optimal kernel matrix H 最优 The specific implementation process is: Step 4-1), initialization: respectively 0 , Q y , R, γ, γ d Initialization assignment, where: H 0 is the initialization value of the kernel matrix H, Q y is the weight matrix; Step 4-2), strategy evaluation: collect the input, output and disturbance input data generated by the UAV longitudinal system, and solve the output feedback Q function H ∞ Bellman equation, obtain the collection interval step length L, the value of L should satisfy L ≥ l × (l + 1) / 2, so: L≥(n1N+n2N+pN+n1+n2)×(n1N+n2N+pN+n1+n2+1) / 2, Output feedback Q function H ∞ Bellman equation: Among them, H j is the H matrix of the jth iteration, z k+1 is a vector consisting of control input, disturbance input and UAV longitudinal system output in the past N moments with k+1 as the base point; Step 4-3), strategy update: Update the control input and disturbance input in a data-driven manner: and are the control input and disturbance input at the j+1th iteration of the kth step, respectively; Step 4-4), termination condition: when the convergence condition ||H is met j -H j-1 The iteration stops when ||≤ε, where ε is the convergence condition constant. When the iteration is terminated, the optimal kernel matrix H is obtained. 最优 ; If the convergence condition is not met, execute j=j+1 and continue iterating; In step 5), the optimal kernel matrix H 最优 The elements in form the optimal control law and the worst disturbance form. The optimal control input and the worst disturbance input are expressed as: in: Indicates the optimal kernel matrix H corresponding to the termination condition 最优 The sub-matrices contained.
2. According to claim 1, a Q-learning fixed-wing UAV H-infinity anti-interference control method is characterized in that: In step 6), the disturbance attenuation condition is used to select the disturbance input that meets the disturbance condition. The disturbance attenuation condition is expressed as: After the disturbance selection is completed, the disturbance input is added to the longitudinal system of the UAV to verify the effective suppression of the disturbance input under the disturbance attenuation condition.
Citation Information
Patent Citations
Deep Q neural network anti-interference model and intelligent anti-interference algorithm
CN108777872A
Unmanned aerial vehicle active-disturbance-rejection controller design method based on FOPSO algorithm
CN113253603A