High-speed aircraft formation control method based on optimizing terminal time

Through the high-speed aircraft formation control method based on optimized terminal time, the problem of rapid response of high-speed aircraft is solved, the system performance and stability are improved, and the formation control optimization is achieved.

CN120103864BActive Publication Date: 2025-08-29NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510586006.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-29
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing optimal control methods cannot meet the requirements of high-speed aircraft for rapid response, resulting in performance degradation, overload loss, trajectory deviation and other problems, affecting flight safety and stability.

Method used

The high-speed aircraft formation control method based on optimized terminal time is adopted, and the terminal response time is optimized to achieve fast response by establishing a nonlinear mathematical model, building an error control system, designing an optimal inversion controller, and using the Critic neural network approximation controller in reinforcement learning.

Benefits of technology

It improves the performance, stability and efficiency of the aircraft formation system, and ensures the safe and stable operation of high-speed aircraft in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103864B_ABST
    Figure CN120103864B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-speed aircraft formation control method based on optimizing terminal time. The method establishes a nonlinear mathematical model for a single high-speed aircraft. Based on the nonlinear mathematical model of the single high-speed aircraft, an error control system for multiple high-speed aircraft is constructed, taking into account the relative positions, speeds, and total distances between the high-speed aircraft. For the error control system of multiple high-speed aircraft, an optimal inversion controller is designed to optimize the terminal time of the high-speed aircraft. A critic neural network based on reinforcement learning is used to approximate the designed optimal inversion controller, and high-speed aircraft formation control is achieved through the optimal inversion controller. By designing an optimal inversion controller that optimizes the terminal response time of the high-speed aircraft, the method addresses the high requirements of the high-speed aircraft formation system for rapid response, optimizes the terminal time, and thereby achieves optimal control of the aircraft formation, improving the system's performance, stability, and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-agent system formation flight control of multiple high-speed aircraft, and in particular to a high-speed aircraft formation control method based on optimizing terminal time. Background Art

[0002] High-speed aircraft refers to aircraft with a flight Mach number greater than 5. Research and development of advanced control technologies for high-speed aircraft, especially the optimal control of formation systems, has become a key frontier and hot research direction in the modern aerospace field.

[0003] Optimal control of high-speed aircraft formation systems can significantly improve flight performance, safety, and mission efficiency. For example, optimal control methods can help aircraft optimize speed, maneuverability, and stability through efficient resource allocation during different flight phases (such as acceleration, cruising, and braking), ensuring optimal performance in complex environments. These methods are an indispensable core technology in the design and operation of modern high-speed aircraft.

[0004] However, current optimal control methods still cannot meet the rapid response requirements of high-speed aircraft, leading to a series of potential problems. Specifically, a sluggish control system response can lead to adverse consequences such as degraded aircraft performance, overload loss of control, and trajectory deviation. This not only increases energy waste and fuel consumption, but also increases safety risks and even causes flight accidents. Furthermore, a sluggish control system response can cause overcompensation or oscillation, compromising flight stability and even causing control system failure or crash. Overall, a rapidly responsive control system is crucial to ensuring the safe, stable, and efficient operation of high-speed aircraft during flight. Summary of the Invention

[0005] Purpose of the invention: In view of the above shortcomings, the present invention provides a high-speed aircraft formation control method based on optimizing terminal time and enabling rapid response of a high-speed aircraft formation system.

[0006] Technical solution: To solve the above problems, the present invention adopts a high-speed aircraft formation control method based on optimizing terminal time, comprising the following steps:

[0007] (1) Establish a nonlinear mathematical model for a single high-speed aircraft;

[0008] (2) Based on the nonlinear mathematical model of a single high-speed aircraft, the error control system for multiple high-speed aircraft is constructed by considering the relative position, speed, and total distance between the high-speed aircraft;

[0009] (3) Design an optimal backstepping controller for the error control system of multiple high-speed aircraft to optimize the terminal response time of high-speed aircraft;

[0010] (4) The critic neural network in reinforcement learning is used to approximate the designed optimal inversion controller, and the high-speed aircraft formation control is realized through the optimal inversion controller.

[0011] Furthermore, the nonlinear mathematical model of a single high-speed aircraft in step (1) is:

[0012] ;

[0013] in, For the The position status of a high-speed aircraft, For Find the first-order derivative, For the The speed state of a high-speed aircraft, For Find the first-order derivative, , and are all known nonlinear functions, For the Control input for a high-speed aircraft.

[0014] Furthermore, the error control system model of the multi-high-speed aircraft is:

[0015] ;

[0016] ;

[0017] in, , , For the The desired position of a high-speed aircraft, For the The expected speed of a high-speed aircraft; , For the The position status of a high-speed aircraft, For the The desired position of a high-speed aircraft; , is the position status of the leader high-speed aircraft, The desired position of the leader's high-speed aircraft; For Find the first-order derivative, For Find the first-order derivative, For Find the first-order derivative; is the global formation position error, For the A high-speed aircraft and The communication relationship between high-speed aircraft, For the A communication relationship between a high-speed aircraft and a leader.

[0018] Furthermore, the Desired position of a high-speed aircraft The calculation formula is:

[0019] ;

[0020] ;

[0021] ;

[0022] ;

[0023] ;

[0024] in, For the The desired relative position between the follower high-speed aircraft and the leader high-speed aircraft, 、 、 They are the follower's first The relative position of the high-speed aircraft to the leader high-speed aircraft is expected. For the The expected straight-line distance between the follower high-speed aircraft and the leader high-speed aircraft, For the The connection line between the follower high-speed aircraft and the leader high-speed aircraft The angle between the planes, For the The connection line between the follower high-speed aircraft and the leader high-speed aircraft is The speeds of the follower high-speed aircraft are The angle of projection on the plane, is the track inclination angle of the leader high-speed aircraft, is the track azimuth of the leader high-speed aircraft;

[0025] No. Expected speed of a high-speed aircraft The calculation formula is: .

[0026] Furthermore, the virtual controller of the optimal inversion controller is:

[0027] ;

[0028] Optimal terminal time satisfy: ;

[0029] in, is a positive definite symmetric matrix, For the communication relationship between high-speed aircraft, is the optimal performance indicator function Error variable Seek the derivative, It's about terminal time function, is the Hamiltonian function.

[0030] Furthermore, the actual controller of the optimal inversion controller is:

[0031] ;

[0032] Optimal terminal time satisfy: ;

[0033] in, is a positive definite symmetric matrix, is the optimal performance indicator function Error variable Find the derivative.

[0034] Furthermore, the specific steps of designing the optimal inversion controller include:

[0035] Defining the error variable and :

[0036] ;

[0037] in, It is Estimation of virtual control for a high-speed aircraft;

[0038] According to the global formation position error The calculation formula for the error variable To find the first-order derivative, we have:

[0039] ;

[0040] in,

[0041] ;

[0042] ;

[0043] Then we have:

[0044] ;

[0045] ;

[0046] For the error variable :

[0047] Setting the cost function :

[0048] ;

[0049] in, and is a positive definite symmetric matrix, It is Virtual control of a high-speed aircraft;

[0050] Setting performance indicator functions :

[0051] ;

[0052] in, It's about terminal time function, represents the time variable;

[0053] The performance indicator function converges, and there is an optimal performance indicator function:

[0054] ;

[0055] in, for The admissible control subset on , express A compact subset of express Euclidean space of dimensional real vectors, The optimal terminal time function, is the discount factor;

[0056] Substituting the Hamiltonian function into the Hamiltonian-Jacobi-Bellman equation is:

[0057] ;

[0058] in, is the optimal performance indicator function right Seek derivation, is the transpose of the matrix, For the error variable The derivative of

[0059] Solution Optimal virtual control ;

[0060] For the error variable , similarly, we get the optimal actual control .

[0061] Furthermore, the specific steps of using the Critic neural network in reinforcement learning to approximate the designed optimal inversion controller are as follows:

[0062] Will Decomposing, considering the solution of the Hamilton-Jacobi-Bellman equation and the first neural network formula, and estimating the unknown terms, the optimal virtual controller is:

[0063] ;

[0064] in, is the normal number to be designed, Represents the activation function of the first neural network right Find the partial derivative, is the estimated value of the ideal weight of the first neural network;

[0065] The weight update law is:

[0066] ;

[0067] in, is the learning rate parameter, is the training error, is the partial derivative of the training error with respect to the estimate of the ideal weights of the first neural network;

[0068] Will Decomposition, considering the solution of the Hamilton-Jacobi-Bellman equation and the second neural network formula, and estimating the unknown terms, the optimal actual controller is:

[0069] ;

[0070] in, is the normal number to be designed, Represents the activation function of the second neural network right Find the partial derivative, is the estimated value of the ideal weight of the second neural network;

[0071] The weight update law is:

[0072] ;

[0073] in, is the learning rate parameter, is the training error, is the partial derivative of the training error with respect to the estimate of the ideal weights of the second neural network.

[0074] The present invention also adopts a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0075] The present invention also adopts a computer-readable storage medium having a computer program stored thereon, and the computer program implements the steps of the above method when executed by a processor.

[0076] Beneficial effect: Compared with the existing technology, the significant advantage of the present invention is that it meets the high requirements of the high-speed aircraft formation system for rapid response by designing an optimal inversion controller that optimizes the terminal response time of high-speed aircraft, optimizes the terminal time, thereby achieving optimal control of the aircraft formation and improving the performance, stability and efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 Schematic diagram of the flow of the high-speed aircraft formation control method of the present invention. DETAILED DESCRIPTION

[0078] like Figure 1 As shown, a high-speed aircraft formation control method based on optimizing terminal time in this embodiment of the present invention includes the following steps:

[0079] Step 1: Build a system model

[0080] High-speed aircraft is a typical nonlinear model with a wide range of flight altitude and Mach number. Generally, the flight altitude is above 30,000 meters and the flight Mach number is greater than 5. The flight environment is also complex and changeable, which brings great challenges to the control law design. The dynamic equations and kinematic equations of a high-speed aircraft are:

[0081] (1);

[0082] (2);

[0083] (3);

[0084] (4);

[0085] (5);

[0086] (6);

[0087] in, Indicates the The position status of a high-speed aircraft, Indicates the The speed of a high-speed aircraft, and , (Simplified to )and (Simplified to ) represent the The track inclination angle and track azimuth of a high-speed aircraft, For the The angle of attack of a high-speed aircraft; Indicates the The mass of a high-speed aircraft, is the acceleration due to gravity; 、 、 For the Thrust vectoring for a high-speed aircraft Component in the body coordinate system; 、 、 They represent drag, side force and lift respectively.

[0088] Specify the mass of all high-speed aircraft in the high-speed aircraft formation system and angle of attack Similarly, the aerodynamic control surfaces used by each high-speed aircraft during flight include elevators , rudder and ailerons , where the aerodynamic expressions of each high-speed aircraft are:

[0089] (7);

[0090] (8);

[0091] (9);

[0092] in, , , .

[0093] In order to more conveniently design the control law for the formation system, the following expression is defined:

[0094] (10);

[0095] (11);

[0096] (12);

[0097] The above The expression of a high-speed aircraft can be transformed into:

[0098] (13);

[0099] (14);

[0100] (15);

[0101] (16);

[0102] (17);

[0103] (18);

[0104] definition is the control input of the aircraft. Derivatives of Equations (13)-(15) are taken and Equations (16)-(18) are substituted into the equations to obtain:

[0105] (19);

[0106] in,

[0107] (20);

[0108] (twenty one);

[0109] Defining new variables , ,in, and Respectively represent The speed state and control input of a high-speed aircraft, then The dynamic model of a high-speed aircraft can be rewritten as:

[0110] (twenty two);

[0111] in, , express Euclidean space of dimensional real vectors, express A compact subset of .

[0112] Step 2: Considering the relative position, speed, and total distance between high-speed aircraft, construct the collaborative formation error of multiple high-speed aircraft.

[0113] The multi-speed vehicle system includes a leader and followers, the relative position expected by the followers can be described as , and has the following expression:

[0114] (twenty three);

[0115] (twenty four);

[0116] (25);

[0117] in, Indicates the The expected straight-line distance between the follower and the leader, and , It is the safe distance between high-speed aircraft; Indicates the Build a connection between followers and leaders The angle between planes; Represents a straight line and speed respectively The angle of the projection on the plane; 、 Denote the leader’s track inclination angle and track azimuth respectively. According to the above description and definition, the expected position of the follower can be obtained , .

[0118] definition, ,in 、 Respectively represent The expected position and velocity of a high-speed aircraft. Then the error system is:

[0119] (26);

[0120] Define the global formation position error:

[0121] (27);

[0122] in, 、 Respectively represent A high-speed aircraft and Communication relationship between a high-speed aircraft and a leader.

[0123] Step 3: Design of optimal inverse controller for high-speed aircraft formation system.

[0124] Considering the high speed requirements of high-speed aircraft formation systems, an inverse optimal controller for high-speed aircraft formation systems based on optimized terminal time is designed. First, a coordinate transformation is defined, and then the optimal controller is obtained using inversion techniques, resulting in the optimized terminal time. The specific steps are as follows:

[0125] First define the coordinate transformation:

[0126] (28);

[0127] in, is the global position error, It is The estimated value of virtual control of a high-speed aircraft, and there is hour .

[0128] According to formula (27), we have:

[0129] (29);

[0130] definition have to:

[0131] (30);

[0132] in, , , , and assuming , Is a positive number.

[0133] Setting the cost function ,in is a positive definite symmetric matrix, It's about terminal time function, let For virtual control, set the performance index function to:

[0134] (31);

[0135] in, is the discount factor, for the admissible control law , express The admissible control subset on , and the cost function is continuously differentiable.

[0136] Abbreviated as or , if Equation (31) is continuously differentiable, by introducing Dividing the integral into two parts, we get:

[0137] (32).

[0138] In equation (32), let the first term on the right side be Taylor expansion, and let the second term on the right side be:

[0139] ;

[0140] Then we have:

[0141] ;

[0142] in 、 Respectively right and pair Derivative, this article abbreviates it as 、 . eliminate , and remove , and then have to:

[0143] (33).

[0144] Set the Hamiltonian function to:

[0145] ;

[0146] Since the performance indicator function converges, there exists the following optimal performance indicator function:

[0147] (34);

[0148] Indicates that the control changes with the state, and its essence is still changing with time, and The representation is actually the same. Here, in order to highlight its relevance to the state, it is written as .

[0149] Substituting the Hamiltonian function into the Hamiltonian-Jacobi-Bellman equation is:

[0150] (35);

[0151] By solving The optimal virtual control can be obtained:

[0152] (36);

[0153] And the optimal terminal time Satisfy the following form:

[0154] (37).

[0155] According to the coordinate transformation of formula (28), we have:

[0156] (38);

[0157] (39);

[0158] Assume that , set the cost function ,in 、 is a positive definite symmetric matrix, and the performance index function is:

[0159] (40).

[0160] Substituting the optimal performance index function into the Hamilton-Jacobi-Bellman equation yields:

[0161] (41);

[0162] By solving The optimal actual control is:

[0163] (42);

[0164] Optimal terminal time Satisfies the following formula:

[0165] (43).

[0166] Step 4: Design of adaptive optimal controller and proof of stability.

[0167] Next, we use the Critic neural network from reinforcement learning to approximate the designed optimal backstepping controller to optimize controller performance and improve accuracy. Finally, using Lyapunov stability theory, we prove that the designed formation control system has global asymptotic stability, thus ensuring good convergence performance in a dynamically changing environment. The specific steps are as follows:

[0168] In order to achieve the control objectives, Breaks down to:

[0169] (44);

[0170] in, is a positive constant to be designed, .

[0171] Using the characteristics of neural networks, It can be expressed as:

[0172] (45);

[0173] in, is the ideal weight of the neural network, is the number of hidden neurons, is the activation function, is the approximation error.

[0174] because is an unknown matrix, so we use an estimated value Instead, design as follows:

[0175] (46);

[0176] The gradient vector for estimating the optimal cost is:

[0177] (47);

[0178] in, Respectively 、 right Find the partial derivative.

[0179] Considering the solution of the Hamilton-Jacobi-Bellman equation and the neural network formulation, and estimating the unknown terms, the optimal controller can be written as:

[0180] (48);

[0181] According to equations (47) and (48), the training error is corrected to obtain the Hamilton-Jacobi-Bellman equation , then the partial derivative of the error with respect to the estimated weight is:

[0182] (49).

[0183] Next, train the Critic neural network and design , using the gradient descent method to make the objective function At least, the update law of Critic NN is established as follows:

[0184] (50);

[0185] in, is the learning rate parameter, using Normalization is performed to improve the convergence speed of the gradient descent method.

[0186] After coordinate transformation, repeat the above steps and The same decomposition can be done as follows:

[0187] (51);

[0188] in, is a positive constant to be designed, .

[0189] Also using the characteristics of neural networks, It can be expressed as:

[0190] (52);

[0191] in, is the ideal weight of the neural network, is the activation function, is the approximation error.

[0192] because is an unknown matrix, so we use an estimated value Instead, design as follows:

[0193] (53);

[0194] Then we have:

[0195] (54);

[0196] in, Respectively 、 right Find the partial derivative.

[0197] Considering the Hamilton-Jacobi-Bellman solution and the neural network formulation and estimating the unknowns, the optimal controller can be written as:

[0198] (55);

[0199] Substituting Equation (55) into the Hamilton-Jacobi-Bellman equation, we get the training error , the partial derivative of the error with respect to the estimated weight is:

[0200] (56);

[0201] Next, train the critic neural network and design , using the gradient descent method to make the objective function Minimum, establish the following evaluation network update law:

[0202] (57);

[0203] in, is the learning rate parameter, using Perform normalization.

[0204] Theorem 1: For a high-speed aircraft formation system with high requirements for speed (such as Equation (22)), under the action of the designed virtual controller (such as Equation (48)), the actual controller (such as Equation (55)), and the adaptive law (such as Equations (50) and (57)), the formation system can achieve optimal control within the optimal terminal time (such as Equations (37) and (43)), and the formation closed-loop system is stable within the terminal time.

[0205] Proof: Step 1: Take Taking a high-speed aircraft as an example, the error between the ideal weight and the estimated weight is recorded as , so there is To facilitate analysis, the following definitions are made:

[0206] ;

[0207] And guarantee , we can get:

[0208] (58);

[0209] in is the residual of the neural network approximation. Assume , , ;in 、 、 are all positive real numbers.

[0210] For the For a high-speed aircraft, the Lyapunov function is designed as follows:

[0211] (59);

[0212] Taking the derivative of equation (59) and substituting it into equations (32) and (51), we can obtain:

[0213] (60);

[0214] According to Young's inequality, we have the following relationship:

[0215] (61);

[0216] (62);

[0217] (63);

[0218] (64);

[0219] Substituting equations (61)-(64) into equation (60), we have:

[0220] (65);

[0221] Further scaling formula (65) to:

[0222] (66);

[0223] make, , , there exists a , making , and define , then formula (66) can be written as:

[0224] (67);

[0225] Step 2: Select The Lyapunov function of a high-speed aircraft system is as follows:

[0226] (68);

[0227] in,

[0228] (69);

[0229] in , so there is .

[0230] For the convenience of analysis, the following definitions are made:

[0231] ;

[0232] And guarantee get:

[0233] (70);

[0234] in To approximate the residual, assume Bounded, exists , ,in All are normal numbers.

[0235] Taking the derivative of formula (71), we can get:

[0236] (71);

[0237] According to Young's inequality, we have:

[0238] (72);

[0239] (73);

[0240] (74);

[0241] (75);

[0242] Substituting equations (73)-(75) into equation (72), we can obtain:

[0243] (76);

[0244] make , there exists a , making .

[0245] Substituting into formula (76), we have:

[0246] (77);

[0247] Taking the derivative of equation (68) and according to equations (67) and (77), we can obtain:

[0248] (78);

[0249] make , , then formula (78) can be finally scaled to:

[0250] (79);

[0251] in, Therefore, we have:

[0252] (80);

[0253] Rule No. The high-speed aircraft system is stable.

[0254] Step 3: Select the Lyapunov function:

[0255] (81);

[0256] According to formula (79), we can get:

[0257] (82);

[0258] in, . Further, we have:

[0259] (83);

[0260] This completes the proof.

Claims

1. A high-speed aircraft formation control method based on optimizing terminal time, characterized in that: The following steps are involved: (1) Establish a nonlinear mathematical model for a single high-speed aircraft; (2) Based on the nonlinear mathematical model of a single high-speed aircraft, the error control system for multiple high-speed aircraft is constructed by considering the relative position, speed, and total distance between the high-speed aircraft; (3) Design an optimal inversion controller to optimize the terminal time of high-speed aircraft for the error control system of multiple high-speed aircraft; The virtual controller of the optimal inversion controller is: ; Optimal terminal time satisfy: ; in, is a positive definite symmetric matrix, For the communication relationship between high-speed aircraft, is the optimal performance indicator function Error variable Seek derivation, It's about terminal time function, is the Hamiltonian function; The performance indicator function is: ; in, is the discount factor; The cost function is: ; in, and is a positive definite symmetric matrix, It is Virtual control of a high-speed aircraft; The actual controller of the optimal inversion controller is: ; Optimal terminal time satisfy: ; in, is a positive definite symmetric matrix, is the optimal performance indicator function Error variable Derivative; The performance indicator function is: ; The cost function is: ; in, 、 is a positive definite symmetric matrix, It is Practical control of a high-speed aircraft; (4) The critic neural network in reinforcement learning is used to approximate the designed optimal inversion controller, and the high-speed aircraft formation control is realized through the optimal inversion controller.

2. The high-speed aircraft formation control method according to claim 1, characterized in that: The nonlinear mathematical model of a single high-speed aircraft in step (1) is: ; in, For the The position status of a high-speed aircraft, For Find the first-order derivative, For the The speed state of a high-speed aircraft, For Find the first-order derivative, , and are all known nonlinear functions, For the Control input for a high-speed aircraft.

3. The high-speed aircraft formation control method according to claim 2, characterized in that: The model of the error control system of the multi-high-speed aircraft is: ; ; in, , , For the The desired position of a high-speed aircraft, For the The expected speed of a high-speed aircraft; , For the The position status of a high-speed aircraft, For the The desired position of a high-speed aircraft; , is the position status of the leader high-speed aircraft, The desired position of the leader's high-speed aircraft; For Find the first-order derivative, For Find the first-order derivative, For Find the first-order derivative; is the global formation position error, For the A high-speed aircraft and The communication relationship between high-speed aircraft, For the A communication relationship between a high-speed aircraft and a leader.

4. The high-speed aircraft formation control method according to claim 3, characterized in that: No. Desired position of a high-speed aircraft The calculation formula is: ; ; ; ; ; in, For the The desired relative position between the follower high-speed aircraft and the leader high-speed aircraft, 、 、 They are the follower's first The relative position of the high-speed aircraft to the leader high-speed aircraft is expected. For the The expected straight-line distance between the follower high-speed aircraft and the leader high-speed aircraft, For the The connection line between the follower high-speed aircraft and the leader high-speed aircraft The angle between the planes, For the The connection line between the follower high-speed aircraft and the leader high-speed aircraft is The speeds of the follower high-speed aircraft are The angle of projection on the plane, is the track inclination angle of the leader high-speed aircraft, is the track azimuth of the leader high-speed aircraft; No. Expected speed of a high-speed aircraft The calculation formula is: 。 5. The high-speed aircraft formation control method according to claim 4, characterized in that: The specific steps of designing the optimal inversion controller include: Defining the error variable and : ; in, It is Estimation of virtual control for a high-speed aircraft; According to the global formation position error The calculation formula for the error variable To find the first-order derivative, we have: ; in, ; ; Then we have: ; ; For the error variable : Setting the cost function : ; in, and is a positive definite symmetric matrix, It is Virtual control of a high-speed aircraft; Setting performance indicator functions : ; in, It's about terminal time function, represents the time variable; The performance indicator function converges, and there is an optimal performance indicator function: ; in, for The admissible control subset on , express A compact subset of express Euclidean space of dimensional real vectors, The optimal terminal time function, is the discount factor; Substituting the Hamiltonian function into the Hamiltonian-Jacobi-Bellman equation is: ; in, is the optimal performance indicator function right Seek derivation, is the transpose of the matrix, For the error variable The derivative of Solution Optimal virtual control ; For the error variable , similarly, we get the optimal actual control .

6. The high-speed aircraft formation control method according to claim 5, characterized in that: The specific steps of using the Critic neural network in reinforcement learning to approximate the designed optimal backtracking controller are as follows: Will Decomposing, considering the solution of the Hamilton-Jacobi-Bellman equation and the first neural network formula, and estimating the unknown terms, the optimal virtual controller is: ; in, is the normal number to be designed, Represents the activation function of the first neural network right Find the partial derivative, is the estimated value of the ideal weight of the first neural network; The weight update law is: ; in, is the learning rate parameter, is the training error, is the partial derivative of the training error with respect to the estimate of the ideal weights of the first neural network; Will Decomposition, considering the solution of the Hamilton-Jacobi-Bellman equation and the second neural network formula, and estimating the unknown terms, the optimal actual controller is: ; in, is the normal number to be designed, Represents the activation function of the second neural network right Find the partial derivative, is the estimated value of the ideal weight of the second neural network; The weight update law is: ; in, is the learning rate parameter, is the training error, is the partial derivative of the training error with respect to the estimate of the ideal weights of the second neural network.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Unmanned air vehicle attitude robust fault tolerance control method based on neural network observer

    CN104049640A

  • Four-rotor formation fault-tolerant control method based on adaptive neural network

    CN111948944A