Online optimization decision-making method for aircraft swarm game confrontation based on pseudo-spectral method
By combining the pseudo-spectral method of rolling time domain ideas, the optimal control problem is established, and the online optimization problem of the drone cluster in complex environments is solved, efficient coordinated obstacle avoidance and approximation of the drone cluster is achieved, and the efficiency of independent decision-making and control accuracy are improved.
Patent Information
- Application Number
- CN202411688472.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2044-11-25
AI Technical Summary
The existing technology is difficult to achieve rapid online optimization in the collaborative flight of drone clusters, especially in complex environments, which cannot effectively deal with multiple constraints such as angle and time, resulting in poor results in the drone cluster approaching the target.
The pseudo-spectral method combined with the rolling time domain idea is adopted to establish the optimal control problem, and nonlinear constraints are handled through the precise penalty function, and discrete interpolation and Lagrangian polynomial interpolation are used to realize the online optimization decision of the aircraft cluster, and comprehensively consider multiple constraints such as angle and time.
It significantly improves the autonomous decision-making efficiency and coordinated obstacle avoidance capabilities of the drone cluster, improves the control accuracy and dynamic response capabilities of the aircraft cluster, and enhances the task adaptability and safety in complex environments.
Smart Images

Figure CN119576004B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cooperative autonomous flight of unmanned aerial vehicle (UAV) swarms, and particularly proposes an online autonomous decision-making method for UAV swarm cooperative obstacle avoidance and approaching a target by combining a pseudospectral method with the idea of receding horizon. Background Technique
[0002] The technology of aircraft has developed rapidly. Its characteristics of being driverless, flexible, and cost-effective have enabled it to be widely used in many fields such as military, agriculture, and logistics. With the continuous progress of intelligent and automated technologies, the remote control and autonomous control systems of aircraft have become increasingly perfect, further broadening their application scenarios and value. However, with the improvement of control technology, the flight environment faced by aircraft has become increasingly complex. In order to better adapt to task execution in diverse scenarios, the research on aircraft autonomous control flight technology has become the focus of global attention.
[0003] With the continuous expansion of the application scenarios of aircraft, its flight environment has become increasingly complex, posing higher requirements for the task execution ability of aircraft. Simply improving the autonomous decision-making ability of a single aircraft can no longer cope with the current increasingly severe task challenges. In this context, cooperative flight of aircraft demonstrates significant advantages and can complete tasks more effectively, so this research direction has received increasing attention. However, the optimization problems involved in multi-aircraft cooperative flight are no longer limited to the states of individual aircraft, but become more complex. The optimization process takes a long time and is vulnerable to factors such as dynamic environment, communication limitations, computational complexity, and security.
[0004] In order to effectively solve the online optimization problem with complex constraints of fast autonomous decision-making during cooperative flight of UAV swarms, improve the cooperation among UAV swarms, and ensure the task completion rate for specific scenarios. Therefore, researching an autonomous decision-making method for UAV swarm cooperative obstacle avoidance and approaching a target based on a rolling optimization method of the pseudospectral method will provide strong support for the research on UAV swarm problems.
[0005] The current methods include:
[0006] 1) UAV Swarm Flight Path Planning Based on Pseudospectral Method [J]. Li Zheng, Chen Jianwei, Peng Bo. Aerospace Defense, 2021, 4(01): 52-59. This study proposed a UAV swarm path planning method based on the Radau pseudospectral method, aiming to solve the path planning problem in multi-UAV cooperative flight. First, the dynamic model of the UAV was established, and the constraint conditions during the swarm flight were considered and transformed into nonlinear constraints. Then, the performance index of path planning was determined, and the optimal control model for UAV swarm path planning was established. Finally, the Radau pseudospectral method was used to solve the optimal control problem, and the feasibility and effectiveness of this method were verified through simulation. The simulation results show that this method can obtain flight trajectories that meet state constraints, path constraints, and control constraints, meet the requirements of multi-UAV cooperative flight, and have strong practical value. However, its disadvantages are: (1) This method optimizes the RPM for the overall movement route and cannot achieve the effect of online optimization. (2) More constraint conditions, such as angle constraints and time constraints, are not established, and the effect of the UAV swarm approaching the target is not good.
[0007] 2) Multi-UAV Cooperative Trajectory Planning Method Based on Q-Learning [J]. Yin Yiyi, Wang Xiaofang, Zhou Jian. Acta Armamentarii, 2023, 44(02): 484-495. This study proposed a solution based on Q-learning for the trajectory planning problem of multi-UAVs arriving at the target simultaneously. By establishing a battlefield environment model and a Markov decision model for single-UAV trajectory planning, the Q-learning algorithm was used to find the optimal trajectory with the shortest flight distance. Then, the experience matrix obtained by the Q-learning algorithm was used to quickly calculate the shortest trajectories of each UAV, and by adjusting the action selection strategy of the detouring UAVs, a trajectory group that meets time coordination was obtained. To solve the collision problem between multi-UAVs, this method designed a retreat parameter to determine the local replanning area, and based on the deep Q-learning theory, a neural network was used to replace the Q-table to replan the local trajectory, avoiding the curse of dimensionality problem. In addition, this method also designed an obstacle Q matrix by referring to the idea of the artificial potential field method and superimposed it on the original Q matrix to achieve collision avoidance of UAVs for unprobed obstacles. The simulation results show that this method can obtain cooperative trajectories that meet time coordination and collision avoidance and avoid unprobed obstacles during environmental modeling. (1) When the number of UAVs is large, the training time and space complexity of the Q-learning algorithm are still high. (2) The collision detection module and the local replanning method are relatively simple and cannot handle complex collision scenarios and dynamic obstacles. (3) More constraint conditions, such as angle constraints and time constraints, are not established, and the effect of the UAV swarm approaching the target is not good. Summary of the Invention
[0008] In the field of cooperative autonomous flight technology for aircraft clusters, the present invention proposes an online autonomous decision-making method for cooperative obstacle avoidance and approaching targets of aircraft clusters by combining the pseudospectral method with the idea of rolling horizon. This method establishes an optimal control problem optimized based on the pseudospectral method and designs a corresponding objective function, comprehensively considering factors such as angle constraints and time constraints. Through rolling optimization, online optimization of the aircraft formation is achieved, thus significantly improving the efficiency of autonomous decision-making for cooperative obstacle avoidance and approaching targets of aircraft clusters. The specific steps are as follows:
[0009] S1. Obtain relevant parameters required for the aircraft cluster to execute the target task, and define the state update expression of the aircraft cluster as:
[0010] In the initial stage, determine information such as the number of aircraft clusters executing the task, the number of enemy interceptors, the number, positions, and radii of threat areas in the environment, the position of the target point, the initial positions of the aircraft clusters, the initial states, and the angle-of-attack settings for the aircraft clusters to approach the target cooperatively. Then define the aircraft cluster state update expression as follows.
[0011] Select the ground coordinate system Arbitrarily select a spatial point on the ground , and choose an arbitrary direction in the horizontal plane as the axis direction, The axis direction is perpendicular to the horizontal plane upward, Perpendicular to the plane, and its direction is determined according to the right-hand rule.
[0012] Then consider fixed-wing aircraft groups. The kinematic model of the th fixed-wing aircraft can be expressed in the ground coordinate system as:
[0013]
[0014] Wherein, is the position coordinate of the UAV in the ground coordinate system, represents the overload coefficient received by the UAV, represents the normal overload, is the gravitational acceleration, is the track deviation angle, is the track inclination angle, is the speed roll angle.
[0015] Based on this kinematic model, the state update equation of the aircraft cluster can be represented. From the above model, it can be seen that for the th aircraft, its state quantity vector can be expressed as follows:
[0016]
[0017] The input control quantity vector is as follows:
[0018]
[0019] Therefore, the state update expression of the UAV swarm at any time can be expressed as the following formula:
[0020]
[0021] S2. Establish the optimal control problem of the UAV swarm as follows:
[0022]
[0023] Wherein, is the augmented objective function after processing the relevant constraints using the exact penalty function; is the state update expression of the UAV swarm; is the initial state constraint of the initial swarm; and are the boundary constraints of the state quantity and the control quantity respectively; is the range constraint of the slack decision variable; is the time domain range of the optimization solution. The derivation of these quantities will be specifically introduced below.
[0024] Since in the cooperative trajectory planning problem of the UAV swarm, it contains complex non-linear constraints, the computational dimension of using conventional solution methods is very large, and the solution speed for this problem is very slow. In order to improve the solution efficiency of this trajectory planning problem, the present invention introduces the precise penalty function method to process the non-linear constraints of the UAV swarm trajectory planning problem. Specifically, the threat zone constraint, distance constraint and evasion interceptor constraint of the UAVs are augmented into the objective function, that is, the term in the above, and its specific expression is as follows:
[0025]
[0026]
[0027] In the formula, is a constant, is the penalty coefficient. Obviously, when the degree of non-satisfaction of the state constraint and the time interval are larger, the value of is larger. is the minimum safety distance constraint between any two UAVs in the UAV swarm, is the maximum communication distance constraint between any two UAVs in the UAV swarm, is the minimum safety distance between two UAVs, is the maximum communication distance between two UAVs, Indicates the distance between the current aircraft and the th aircraft. Is the threat area constraint, Indicates the coordinates of the center of the th threat area, Is the radius of the planar circle of the threat area, Is the number of threat areas. Is the interceptor constraint, Indicates the position of the current th interceptor, Indicates the position of the current aircraft, Indicates the minimum safe distance between the aircraft and the interceptor, Is the number of aircraft in the aircraft cluster, Is the number of interceptors.
[0028] It can be seen that in the augmented objective function, when the constraints are fully satisfied, the optimal solution of the augmented objective function after being processed by the exact penalty function is the optimal solution of the original problem . The original problem contains three optimization terms: approaching the target, time coordination, and angle coordination, which are expressed as follows:
[0029]
[0030] Among them, Is the optimization term for approaching the target, Represents the position coordinates of the expected target point, Represents the Euclidean distance between two points. Is the optimization term for time coordination, , , respectively represent the position, opportunistic position, and target point position of the current aircraft. Is the optimization term for angle coordination, Indicates the angle between the vector from the th aircraft to the target and the plane, Indicates the angle between the projection of the vector from the th aircraft to the target on the plane and the axis, Indicates the expected angle of the angle between the vector from the th aircraft to the target and the plane, Indicates the expected angle of the angle between the projection of the vector from the th aircraft to the target on the plane and the axis. , , respectively 、 、 weight coefficients
[0031] The differential equation in the established optimal control problem has been defined and described in S1, which is the state quantity of the initial cluster given in S1. Regarding and the boundary constraints of these two state quantities and the control quantity are specifically expressed as follows:
[0032]
[0033] where 、 、 、 respectively represent the boundary ranges of the flight altitude, flight speed, track deviation angle, and track inclination angle of the aircraft, 、 、 respectively represent the boundary ranges of the tangential overload, normal overload, and roll angle of the aircraft.
[0034] S3. Solve the optimization problem based on the pseudospectral method combined with the moving horizon approach, specifically as follows:
[0035] (1) Select the number of nodes for pseudospectral method discrete interpolation, the initial time , the initial state , and apply the Radau pseudospectral method to calculate the open-loop optimal control within 4s time domain ;
[0036] (2) Within the interval , apply the optimal control quantity to the system, and denote the state measurement value at as , and let ;
[0037] (3) Within the interval , apply the open-loop optimal control to the system; meanwhile, based on the state measurement value at , calculate the open-loop optimal control , and let be the optimization calculation time, then ;
[0038] (4) Let , and return to (3).
[0039] The aircraft is continuously controlled by the above-mentioned rolling horizon update method until the mission is completed or it dies. During the above update process, the Radau pseudospectral method is used to solve the optimal control problem established previously. The solution process of the Radau pseudospectral method is introduced below:
[0040] First, for the state within the time interval , it is mapped to the interval using the following formula for solving and optimizing the problem. The mapping formula is as follows:
[0041]
[0042] Then, the control variables and state variables are discretized using LGR points, and Lagrange polynomial interpolation is used to obtain approximate functions of the state variables and control variables. The basis functions of the Lagrange polynomial interpolation are:
[0043]
[0044] The approximate expressions of the state variables and control variables are:
[0045]
[0046] where and are respectively and After the above time transformation, a series of points discretized using LGR points on , and there is .
[0047] Taking the derivative of the state variable gives
[0048]
[0049] Taking the derivative of the state variable , through the differential equation established in S1, can also be expressed in another form:
[0050]
[0051] After the above processing method, the state differential equation in the S2 optimal control problem can be expressed as the following equality constraint:
[0052]
[0053] Therefore, the pseudo-spectral method, through this discrete method, converts the optimal control problem of S2 into a non-linear programming problem. There are many mature methods for solving non-linear programming problems. By using the SNOPT package for solving, the solution to the optimal control problem proposed by S2, that is, the optimal control quantity, can be obtained. The obtained optimal control quantity is applied to the aircraft system. The non-linear programming problem transformed by the above method can be described by the following formula:
[0054]
[0055] This continuous update process of open-loop optimal control ensures that the algorithm has the ability of real-time and online application, and realizes the online optimization of the aircraft group by using rolling optimization. The update process is as Figure 1 shown.
[0056] S4. State perception and update:
[0057] In S3, the solution algorithm and update iteration framework for the optimal control problem are introduced to complete the given task. It can be seen that after solving the optimal control problem proposed by S2, the solution to this problem, that is, the optimal control quantity, is obtained. Applying it to the aircraft system and solving by rolling time-domain update until the task is completed. Therefore, in the algorithm, it is necessary to add a judgment and detection on whether the task is completed. If the task has been completed, the algorithm ends; if the task is not completed in the current optimization update, the current state of the aircraft is detected, and the opportunistic information and the perception and update of the environmental information are obtained. These information are used as the initial settings of S2 to re-solve the optimal control problem. The overall flow chart of the solution is as Figure 2 shown.
[0058] Advantages of the present invention:
[0059] Based on the rolling time-domain pseudo-spectral method, the present invention realizes the online optimization of cooperative obstacle avoidance and approaching the target of the aircraft cluster, significantly improving the efficiency of autonomous decision-making. By optimizing the establishment of the optimal control problem through the pseudo-spectral method and comprehensively considering multiple constraints such as angle and time, the control accuracy and dynamic response ability of the aircraft cluster are enhanced, thereby improving the adaptability and safety of cooperative flight. Description of the Drawings
[0060] Figure 1 is the continuous application process diagram of open-loop optimal control.
[0061] Figure 2 is the flow chart of the pseudo-spectral method aircraft cluster trajectory optimization based on the rolling time-domain idea.
[0062] Figure 3 is the process of the aircraft cluster approaching.
[0063] Figure 4 It is the overload curve of the aircraft.
[0064] Figure 5 It is the speed curve of the aircraft. Specific implementation manner
[0065] The present invention utilizes the aircraft motion model mentioned in S1 to verify the pseudospectral method under the rolling time domain mentioned above. Taking the example of five aircrafts cooperating to approach a static target point and avoiding the interceptors of each other during the approaching process, this algorithm is used to perform online planning for the approaching trajectory and complete relevant cooperative actions. As Figures 3-5 shown.
Claims
1. An online optimization decision-making method for the game confrontation of aircraft clusters based on the pseudospectral method, characterized in that Including the following steps: S1. Obtain the relevant parameters required for the aircraft group to perform the target task, and define the state update expression of the aircraft group at any time as: , Among them, is the state vector of the aircraft group, is the control vector of the aircraft group; S2. Establish the optimal control problem of the aircraft cluster as follows: , Among them, is the augmented objective function for processing relevant constraints using the exact penalty function; is the initial state constraint of the initial cluster; and are the boundary constraints of the state variable and the control variable respectively; is the range constraint of the slack decision variable; is the time domain range of the optimization solution; The specific expression is: , Wherein, , , , , , , , Among them, is a constant, is the penalty coefficient, is the minimum safety distance constraint between any two aircraft in the aircraft cluster, is the maximum communication distance constraint between any two aircraft in the aircraft cluster, is the minimum safety distance between two aircraft, is the maximum communication distance between two aircraft, represents the distance between the current aircraft and the th aircraft, is the threat area constraint, represents the th threat area center coordinates, is the radius of the threat area plane circle, is the number of threat areas, is the interceptor constraint, represents the position of the th interceptor, represents the position of the current aircraft, represents the minimum safety distance between the aircraft and the interceptor, is the number of aircraft in the aircraft cluster, is the number of interceptors; It includes three optimization items: approaching the target, time, and angle coordination, which are expressed as: , Among them, is an optimization term for approaching the target, represents the position coordinates of the expected target point, represents the Euclidean distance between two points; is an optimization term for time coordination, 、 and respectively represent the position, opportunistic position and target point position of the current aircraft; is an optimization term for angle coordination, represents the angle between the vector from the th aircraft to the target and the plane, represents the angle between the projection of the vector from the th aircraft to the target on the plane and the axis, represents the expected angle of the angle between the vector from the th aircraft to the target and the plane, represents the expected angle of the angle between the projection of the vector from the th aircraft to the target on the plane and the axis; 、 、 are respectively 、 、 's weight coefficients; S3. Solve the optimal control problem based on the pseudospectral method combined with the receding horizon idea, specifically: (1)Determine the number of nodes for discrete interpolation by the pseudospectral method and the initial time . The initial state . Apply the Radau pseudospectral method to calculate the open-loop optimal control within 4 s in the time domain ; (2) In the interval apply the optimal control quantity to the system. Denote the state measurement value at as , and let ; (3) In the interval apply the open-loop optimal control to the system; meanwhile, according to the state measurement value at perform the calculation of the open-loop optimal control Let be the optimization calculation time, then ; (4) Let , and return to (3); Control the aircraft continuously through the above receding horizon update method until the task is completed or the aircraft dies; in the above update process, use the Radau pseudospectral method to solve the established optimal control problem, and the specific solution process is: First, for the state within the time interval , it is mapped to the interval using the following formula, and then the problem is solved and optimized. The mapping formula is as follows: , Then, discretize the control quantity and the state quantity using the LGR points, and use Lagrange polynomial interpolation to obtain the approximate functions of the state quantity and the control quantity. The basis function of the Lagrange polynomial interpolation is: , The approximate expressions of the state quantity and the control quantity are: , Among them, and are respectively and After the above-mentioned time transformation, a series of discrete points using LGR points on , and there is ; Take the derivative: , Derivation of the state quantity Through the differential equation established in S1, it can also be expressed in another form: , The above two derivatives of the state quantity are approximately equivalent, so the following equality constraint can be established: , Through this discrete method, the pseudospectral method converts the optimal control problem of S2 into a nonlinear programming problem. By using the SNOPT package for solution, the solution to the optimal control problem proposed by S2, that is, the optimal control quantity, can be obtained. ; Apply the obtained optimal control quantity to the aircraft system; The nonlinear programming problem is expressed as: ; S4. State perception and update: After obtaining the solution of the optimal control problem through S3, judge whether the target task of the aircraft cluster is completed. If so, end; otherwise, detect the current state of the aircraft, and obtain the opportunistic information and sense and update the environmental information, and use these information as the initial settings to re-solve the optimal control problem.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle formation recombination trajectory planning method
CN113050687A
Hypersonic aircraft double-layer trajectory planning method based on pseudo-spectral method
CN117008634A