A safety optimization control method for UAV formation based on leader-follower
By introducing the integrator model and adaptive dynamic programming algorithm into the leader-follower formation control structure, the problems of speed limitation and command tracking error in the UAV formation are solved, the safe optimization control of the UAV formation is achieved, and the stability and safety of the formation are improved.
Patent Information
- Application Number
- CN202410828872.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-06-25
AI Technical Summary
Existing leader-follower-based UAV formation control methods fail to effectively consider the speed limit of the UAV system and the underlying unknown command tracking error, resulting in unstable formations and prone to safety accidents.
Under the leader-follower formation control structure, an integrator model is constructed, and the speed command saturation function and uncertainty terms are introduced. The optimal control problem is established through the Hamilton-Jacobi-Bellman equation, and an adaptive dynamic programming algorithm is used for iterative learning to adaptively compensate for speed limitations and command tracking errors.
The safety and stability of UAV formations are improved, the formation tracking error is reduced, and the safety of UAV formation flight and compliance with dynamic constraints are ensured.
Smart Images

Figure CN118915815B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of UAV control technology, and in particular to a safety optimization control method for UAV formation based on leader-follower. Background Art
[0002] With the rapid development of drone technology, drone formation control has become a research hotspot in the field of drone applications. In specific applications, drone formations can efficiently cover large areas and coordinately execute precise target missions. In other applications, drone formations can significantly improve the efficiency of tasks such as emergency rescue, power inspections, forest fire prevention, remote sensing mapping, and agricultural plant protection. However, in practical applications, the safety of drone formations is often crucial. Drone systems are susceptible to speed limitations and underlying unknown command tracking errors, which can lead to unstable formation control and even crashes. This can lead to mission failures and potentially cause casualties and property damage. Therefore, it is necessary to model the speed limitations and underlying unknown command tracking errors of drones and then design a formation control algorithm that uses a hybrid model- and data-driven adaptive learning approach to ensure the safety of drone formation flight.
[0003] Among existing UAV formation control methods, the leader-follower approach offers numerous advantages, including simple structure, flexible scalability, and ease of implementation. It is the most commonly used formation control method in practical applications. In this approach, the leader UAV is typically responsible for tracking a flight path planned online or offline, while the follower UAVs make corresponding control adjustments based on the leader UAV's status information to ensure formation stability and safety. However, existing leader-follower formation control methods fail to account for the speed limitations and underlying unknown tracking errors faced by real-world UAV systems. This can lead to the leader UAV being unable to accurately track the flight path and, on the other hand, causing instability in the follower UAVs during formation control adjustments, leading to potential collisions or crashes. Summary of the Invention
[0004] The technical problem to be solved by the present invention is: in response to the technical problems existing in the prior art, the present invention provides a leader-follower based UAV formation safety optimization control method with simple principle, wide application range and high safety performance.
[0005] In order to solve the above technical problems, the technical solution proposed by the present invention is:
[0006] A leader-follower-based UAV formation safety optimization control method, comprising:
[0007] Step S01: Under the leader-follower formation control structure, a UAV formation is formed; wherein the leader UAV flies according to a preset or real-time planned flight route, and the follower UAVs adjust their flight attitude based on the status information of the leader UAV to maintain their relative position within the formation;
[0008] Step S02: The formation tracking dynamics model of the follower UAV is transformed into an integrator model. A velocity command saturation function is introduced into the integrator model, and the underlying unknown command tracking error is expressed as an uncertainty term in the model.
[0009] Step S03: Establish an optimal control problem, consider the actual speed limit of the UAV, introduce the Hamiltonian and the Hamilton-Jacobi-Bellman equation, and solve the optimal control strategy considering the speed constraint;
[0010] Step S04: Applying an adaptive dynamic programming algorithm, based on the UAV simulation and actual flight data, through a neural network training process, iteratively learns and updates the control law parameters, and adaptively compensates for the speed limit and command tracking error.
[0011] As a further improvement of the method of the present invention: in step S02, the speed limit of the actual UAV and the underlying unknown instruction tracking error are considered to establish the following equation:
[0012]
[0013] in, is the actual calculated speed instruction, Represents the underlying unknown command tracking error, which is about the speed command Unknown function.
[0014] As a further improvement of the method of the present invention: in step S04, the process of the adaptive dynamic programming algorithm includes:
[0015] Step S401: Initially set the neural network parameters and use the simple P control law to perform preliminary data collection;
[0016] Step S402: Based on the collected flight data, the neural network weights are updated using the least squares method to approach the optimal control strategy;
[0017] Step S403: Repeat the above data collection and weight updating process until the predetermined learning convergence condition is reached.
[0018] As a further improvement of the method of the present invention: the adaptive dynamic programming algorithm solution process:
[0019] Step S411: Two neural networks, namely the evaluation network and the execution network, are used to approximate the optimal value function , and optimal virtual speed control instructions as follows:
[0020]
[0021] in , , and They are and Wiki function;
[0022] Step S412: Collect the UAV's flight data for neural network training. A simple P control law is used during the collection:
[0023]
[0024] Step S413: Initialization , ;
[0025] Step S414: Run the loop to iterate and solve the following equation until , is the decision threshold, usually set to :
[0026]
[0027] Step S415: When ,make , , exit the loop.
[0028] As a further improvement of the method of the present invention, step S05 is also included: establishing a data link between the follower UAV and the leader UAV through the wireless communication module, so that the follower UAV can receive the leader's position and heading information in real time, and make flight control adjustments according to the updated control law to make the formation stable and safe.
[0029] As a further improvement of the method of the present invention, step S06 is also included: using the training results of the adaptive dynamic programming algorithm, that is, the iteratively optimized control instructions, to guide the actual UAV formation to perform tasks, ensuring that the formation control instructions comply with the UAV dynamic constraints under the conditions of speed restrictions and unknown instruction errors.
[0030] As a further improvement to the method of the present invention: the flight data is collected from the simulated flight or the actual flight test of the UAV.
[0031] Compared with the prior art, the advantages of the present invention are:
[0032] 1. The leader-follower-based UAV formation safety optimization control method of the present invention has a simple principle, a wide range of applications, and high safety performance. The present invention can adaptively learn and adjust the optimal formation speed control instructions based on the speed limit of actual UAVs and the underlying unknown instruction tracking error, thereby avoiding the problem of safety accidents caused by not considering the speed limit and the underlying unknown instruction tracking error in actual UAV formation flight.
[0033] 2. The leader-follower-based UAV formation safety optimization control method of the present invention models the speed limit of the UAV system and the underlying unknown command tracking error, and designs a model and data hybrid driven adaptive learning formation control algorithm to further ensure the safety of UAV formation flight.
[0034] 3. The leader-follower-based UAV formation safety optimization control method of the present invention models speed constraints and underlying unknown command tracking errors. It also employs an adaptive dynamic programming approach to adaptively learn based on UAV simulation and actual flight data. Therefore, compared to formation control algorithms that do not consider actual UAV speed constraints and underlying unknown command tracking errors, the optimal formation control algorithm after iterative learning achieves a smaller formation tracking error. Furthermore, because the present invention considers speed constraints and underlying unknown command tracking errors, the UAV formation control commands calculated by the optimal formation control algorithm better conform to the actual UAV dynamic constraints, further ensuring the flight safety of the UAV formation.
[0035] 4. The leader-follower-based UAV formation safety optimization control method of the present invention, through the construction of a formation control system structure, effectively manages the communication topology between multiple UAVs within the formation, ensures accurate information transmission and response, and enhances the overall collaborative operation capability of the formation. By implementing this method, formation tracking error can be significantly reduced, flight safety can be improved, and the instability caused by ignoring speed limits and command tracking errors in traditional formation control can be effectively resolved. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic diagram of a leader-follower UAV formation in a specific application example of the present invention.
[0037] Figure 2 It is a schematic diagram of relative position parameters corresponding to the "I"-shaped and "V"-shaped formations used in a specific application example of the present invention.
[0038] Figure 3 Schematic diagram of the communication topology relationship between multiple follower drones and a leader drone in a specific application example of the present invention.
[0039] Figure 4 It is a flow chart of the adaptive dynamic programming algorithm of the present invention in a specific application example.
[0040] Figure 5 It is a schematic diagram of the formation control system structure of follower UAVs in a specific application example of the present invention.
[0041] Figure 6 It is a flowchart of flight data collection in a specific application example of the present invention.
[0042] Figure 7 It is a schematic diagram of the process of the method of the present invention in a specific application example. DETAILED DESCRIPTION
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0044] In actual UAV formation flight, the leader UAV is usually responsible for tracking the flight path, while the formation maintenance task is mainly completed by the follower UAVs. However, existing formation control methods often fail to consider the speed limitations and underlying unknown command tracking errors faced by actual UAV systems, resulting in unstable formations and prone to safety accidents. The purpose of the present invention is to ensure the stability and safety of actual formation flight by modeling the speed limitations and underlying unknown command tracking errors of actual UAV systems, establishing an optimal control problem, and using an adaptive dynamic programming algorithm to solve the optimal formation speed control instructions. That is, within a leader-follower formation control framework, by modeling the speed limitations and underlying unknown command tracking errors of UAVs, a hybrid model- and data-driven adaptive learning method is used to achieve safe and optimized control of UAV formations.
[0045] In practical UAV formation systems, despite employing a simple and reliable leader-follower formation control architecture, the failure to account for speed limits and underlying unknown command tracking errors can easily lead to formation control system failures and even to safety incidents such as UAV collisions and crashes. Therefore, this paper models the UAV system's speed limits and underlying unknown command tracking errors, designing a formation control algorithm driven by a hybrid model and data-driven adaptive learning approach to further ensure the safety of UAV formation flight.
[0046] To this end, the innovative principle of this invention is to model the formation tracking dynamics of follower UAVs as an integrator model within a leader-follower formation control framework. Secondly, building on the integrator model, a saturation function for the speed command is introduced, and the underlying unknown command tracking error is modeled as an uncertain term. Finally, an iterative optimization control algorithm based on adaptive dynamic programming is designed. Using data collected from simulations or actual flights, adaptive learning and training are performed to iteratively learn the optimal formation control algorithm that accounts for speed constraints and the underlying unknown command tracking error.
[0047] like Figure 1 As shown in FIG, a schematic diagram of a UAV formation based on a leader-follower system in a specific application of the present invention; Figure 7 As shown, the present invention provides a safety optimization control method for a UAV formation based on a leader-follower system, which includes:
[0048] Step S01: Under the leader-follower formation control structure, a UAV formation is formed; wherein the leader UAV flies according to a preset or real-time planned flight route, and the follower UAVs adjust their flight attitude based on the status information of the leader UAV to maintain their relative position within the formation;
[0049] Step S02: The formation tracking dynamics model of the follower UAV is transformed into an integrator model. A velocity command saturation function is introduced into the integrator model, and the underlying unknown command tracking error is expressed as an uncertainty term in the model.
[0050] Step S03: Establish an optimal control problem, consider the actual speed limit of the UAV, introduce the Hamiltonian and the Hamilton-Jacobi-Bellman equation, and solve the optimal control strategy considering the speed constraint;
[0051] Step S04: Applying an adaptive dynamic programming algorithm, based on the UAV simulation and actual flight data, through a neural network training process, iteratively learns and updates the control law parameters, and adaptively compensates for the speed limit and command tracking error.
[0052] In a specific application example, in step S02, the following equation is established by considering the actual speed limit of the UAV and the underlying unknown command tracking error:
[0053]
[0054] in, is the actual calculated speed instruction, Represents the underlying unknown command tracking error, which is about the speed command Unknown function.
[0055] In a specific application example, in step S04, the process of the adaptive dynamic programming algorithm includes:
[0056] Step S401: Initially set the neural network parameters and use the simple P control law to perform preliminary data collection;
[0057] Step S402: Based on the collected flight data, the neural network weights are updated using the least squares method to approach the optimal control strategy;
[0058] Step S403: Repeat the above data collection and weight updating process until the predetermined learning convergence condition is reached.
[0059] As a further improvement of the method of the present invention: the adaptive dynamic programming algorithm solution process:
[0060] Step S411: Two neural networks, namely the evaluation network and the execution network, are used to approximate the optimal value function , and optimal virtual speed control instructions as follows:
[0061]
[0062] in , , and They are and Wiki function;
[0063] Step S412: Collect the UAV's flight data for neural network training. A simple P control law is used during the collection:
[0064]
[0065] Step S413: Initialization , ;
[0066] Step S414: Run the loop to iterate and solve the following equation until , is the decision threshold, usually set to :
[0067]
[0068] Step S415: When ,make , , exit the loop.
[0069] In a specific application example, step S05 is also included: establishing a data link between the follower UAV and the leader UAV through the wireless communication module, so that the follower UAV can receive the leader's position and heading information in real time, and make flight control adjustments according to the updated control law to make the formation stable and safe.
[0070] In a specific application example, step S06 is also included: using the training results of the adaptive dynamic programming algorithm, that is, the iteratively optimized control instructions, to guide the actual UAV formation to perform tasks, ensuring that the formation control instructions comply with the UAV dynamic constraints under the conditions of speed restrictions and unknown instruction errors.
[0071] In specific application examples, flight data is collected from simulated flights or actual flight tests of drones.
[0072] As can be seen from the above, the formation control method proposed in this embodiment is based on a leader-follower formation control structure, where the leader UAV tracks the flight route planned online or offline, and the follower UAV tracks the leader UAV and maintains a given relative position relationship with it.
[0073] In a specific application example, a given relative position relationship is usually expressed as a set of three-dimensional relative position parameters. Indicates, for example Indicates that the follower UAV is 10 meters away from the leader UAV in the north direction and 20 meters away in the east direction in the inertial coordinate system, and the flight altitude is 5 meters higher. According to different formation configuration requirements, the corresponding relative position parameters can be adjusted. The relative position parameters corresponding to the "I" formation and the "V" formation are as follows: Figure 2 shown.
[0074] The follower drone establishes a communication link with the leader drone through the wireless communication module and receives information from the leader drone, including the GPS position and heading angle of the leader drone, and the communication topology relationship between multiple follower drones and the leader drone, such as Figure 3 shown.
[0075] The drone is modeled as an integrator as follows:
[0076]
[0077] in, and are the position and velocity of the follower UAV in the inertial coordinate system.
[0078] Considering the actual speed limit of the drone and the underlying unknown command tracking error, the following equation is established:
[0079]
[0080] in, is the actual calculated speed instruction, Represents the underlying unknown command tracking error, which is about the speed command Unknown function.
[0081] To realize UAV formation, we first need to obtain the formation error equation of the follower UAV tracking the leader UAV, and let the tracking error be , which is a three-dimensional vector, namely Substituting the UAV model into the equation, we get the formation error equation as follows:
[0082]
[0083] Among them, the speed command must meet the constraints .
[0084] Next, we design the following performance index function (for the convenience of describing the method, we use , , the former represents the state of the system, and the latter represents the control input of the system):
[0085]
[0086] Formulate the optimal control problem:
[0087]
[0088] The optimal control law is designed as follows:
[0089]
[0090] in, It is a stable reference speed control instruction, which can be calculated by simple PID control. It is the virtual speed control instruction to be optimized.
[0091] Therefore, the optimal control problem can be rewritten as:
[0092]
[0093] in, , , .
[0094] Define the optimal value function for:
[0095]
[0096] in, is the set of all feasible control laws.
[0097] Next, for the above optimal control problem, the Hamiltonian of the system is defined as:
[0098]
[0099] in, for right The partial derivative of .
[0100] Then the optimal value function Satisfies the Hamilton-Jacobi-Bellman equation:
[0101]
[0102] By deriving the above formula, the optimal control law of the system is obtained as:
[0103]
[0104] Substituting the above formula into the Hamilton-Jacobi-Bellman equation, we get:
[0105]
[0106] Solving the above Hamilton-Jacobi-Bellman equation can obtain the optimal virtual speed command , but the equation is highly nonlinear and contains unknown terms , so it is difficult to solve it directly analytically. The following uses the adaptive dynamic programming algorithm to solve it.
[0107] Adaptive dynamic programming algorithm solution process:
[0108] 1) Two neural networks, the evaluation network and the execution network, are used to approximate the optimal value function , and optimal virtual speed control instructions as follows:
[0109]
[0110] in , , and They are and Wiki function.
[0111] 2) Collect UAV flight data for neural network training. A simple P control law can be used during the collection:
[0112]
[0113] 3) Initialization , .
[0114] 4) Run the loop to iteratively solve the following equation until , is the decision threshold, usually set to :
[0115]
[0116] 5) When ,make , , exit the loop.
[0117] The specific derivation of the neural network weight update in the above adaptive dynamic programming algorithm loop is as follows:
[0118] 1) The neural network Substitute into the equation have:
[0119]
[0120] 2) Order , ,in is the order of the input vector, then:
[0121] 、
[0122] Right now
[0123] 3) Rewrite the above formula as ;
[0124] in , ,
[0125] .
[0126] 4) Due to and is linearly related, so it can be obtained by sampling Group and , the least squares method is used to update the neural network weights.
[0127] 5) Specifically, the least squares method is used to update the neural network weights as follows:
[0128]
[0129] in, , .
[0130] The flow chart of the adaptive dynamic programming algorithm is as follows Figure 4 shown.
[0131] Based on the above method of the present invention, in a specific example, the formation control system structure of the follower UAV is as follows: Figure 5 shown.
[0132] Specifically, the data used for training the adaptive dynamic programming algorithm can be either simulated flight data or real flight data. The flow chart of simulated and real flight data collection is as follows: Figure 6 shown.
[0133] Those skilled in the art will appreciate that the above-mentioned embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0134] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed above with reference to the preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent variations, and modifications to the above embodiment that do not depart from the technical solution of the present invention and are based on the technical essence of the present invention shall fall within the scope of protection of the technical solution of the present invention.
Claims
1. A safety optimization control method for UAV formation based on leader-follower, characterized by: include: Step S01: Building a UAV formation under the leader-follower formation control structure; The leader drone flies according to a preset or real-time planned flight route, and the follower drones adjust their flight attitude based on the status information of the leader drone to maintain their relative position within the formation; Step S02: The formation tracking dynamics model of the follower UAV is transformed into an integrator model. A velocity command saturation function is introduced into the integrator model, and the underlying unknown command tracking error is expressed as an uncertainty term in the model. Step S03: Establish an optimal control problem, consider the actual speed limit of the UAV, introduce the Hamiltonian and the Hamilton-Jacobi-Bellman equation, and solve the optimal control strategy considering the speed constraint; The optimal control problem is: in, , , ; Define the optimal value function for: in, is the set of all feasible control laws; Next, for the above optimal control problem, the Hamiltonian of the system is defined as: in, for right The partial derivative of Then the optimal value function Satisfies the Hamilton-Jacobi-Bellman equation: By deriving the above formula, the optimal control law of the system is obtained as: Substituting the above formula into the Hamilton-Jacobi-Bellman equation, we get: Solving the above Hamilton-Jacobi-Bellman equation can obtain the optimal virtual speed command , but the equation is highly nonlinear and contains unknown terms , so it is difficult to solve it directly analytically, so the adaptive dynamic programming algorithm is used to solve it; Step S04: Applying an adaptive dynamic programming algorithm, based on the UAV simulation and actual flight data, through a neural network training process, iteratively learns and updates the control law parameters, and adaptively compensates for the speed limit and command tracking error.
2. The leader-follower based UAV formation safety optimization control method according to claim 1 is characterized in that: In step S02, the following equation is established by considering the actual speed limit of the UAV and the underlying unknown command tracking error: in, is the actual calculated speed instruction, Represents the underlying unknown command tracking error, which is about the speed command Unknown function.
3. The leader-follower based UAV formation safety optimization control method according to claim 1, characterized in that: In step S04, the process of the adaptive dynamic programming algorithm includes: Step S401: Initially set the neural network parameters and use the simple P control law to perform preliminary data collection; Step S402: Based on the collected flight data, the neural network weights are updated using the least squares method to approach the optimal control strategy; Step S403: Repeat the above data collection and weight updating process until the predetermined learning convergence condition is reached.
4. The leader-follower based UAV formation safety optimization control method according to claim 3 is characterized in that: The adaptive dynamic programming algorithm solution process: Step S411: Two neural networks, namely the evaluation network and the execution network, are used to approximate the optimal value function , and optimal virtual speed control instruction as follows: in , , and They are and Wiki function; Step S412: Collect the UAV's flight data for neural network training. A simple P control law is used during the collection: Step S413: Initialization , ; Step S414: Run the loop to iterate and solve the following equation until , is the decision threshold: Step S415: When ,make , , exit the loop.
5. The leader-follower-based UAV formation safety optimization control method according to any one of claims 1 to 4, characterized in that: It also includes step S05: establishing a data link between the follower UAV and the leader UAV through the wireless communication module, so that the follower UAV can receive the leader's position and heading information in real time, and make flight control adjustments according to the updated control law to make the formation stable and safe.
6. The leader-follower-based UAV formation safety optimization control method according to any one of claims 1 to 4, characterized in that: The method further includes step S06: utilizing the training results of the adaptive dynamic programming algorithm, i.e., the iteratively optimized control instructions, to guide the actual UAV formation to perform tasks, ensuring that the formation control instructions comply with the UAV dynamic constraints under conditions of speed limits and unknown instruction errors.
7. The leader-follower-based UAV formation safety optimization control method according to any one of claims 1 to 4, characterized in that: The collection of flight data comes from the UAV's simulated flight or actual flight test.
Citation Information
Patent Citations
Unmanned ship formation path tracking method based on deep reinforcement learning
CN111694365A
Unmanned aerial vehicle-unmanned ship heterogeneous cooperative obstacle avoidance formation control method based on preset performance control
CN118151527A