Aircraft deployment method based on dynamic airspace network and reinforcement learning

Through dynamic airspace network and reinforcement learning methods, the limitations of aircraft allocation in complex flight simulator training scenarios are solved, precise flight path planning and conflict management are achieved, aviation control efficiency and airspace resource utilization efficiency are improved, and air transportation needs of multiple flights and multiple aircraft models are met.

CN120412339BActive Publication Date: 2025-09-05NAVAL AVIATION UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510918845.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-09-05
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

The existing aircraft deployment methods have limitations in complex and changeable flight simulator training scenarios, which are difficult to meet the growing training needs, and are unable to achieve accurate flight path planning and effective management of conflict risks.

Method used

Using a method based on dynamic airspace network and reinforcement learning, the airspace units are dynamically divided, the safe area is calculated, the flight path collection is generated, and the action instructions for height adjustment, heading correction and speed changes are provided to achieve accurate flight allocation.

Benefits of technology

It improves aviation control efficiency, reduces conflict risks, meets the air transportation needs of multiple flights and multiple aircraft models, and improves the utilization efficiency of airspace resources and flight operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412339B_ABST
    Figure CN120412339B_ABST
Patent Text Reader

Abstract

The present application discloses an aircraft deployment method based on a dynamic airspace network and reinforcement learning, which relates to the field of air traffic management. The method includes obtaining flight information and task priority weight information of each flight in the historical time period and the current time period; dynamically dividing the airspace into independent sub-airspaces and shared airspaces using the flight information in the historical time period; calculating the safe area of ​​each flight through a spatiotemporal buffer zone model and determining the flight path accordingly; generating a set of release action instructions including altitude adjustment, heading correction and speed change based on the flight information, airspace division results and flight path, and deploying flights in real time. The technical effect of the present application is that it can respond to airspace changes in real time, optimize flight paths, reduce potential conflicts, and improve airspace utilization, thereby improving aviation control efficiency and flight safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of air traffic management, and in particular to an aircraft deployment method based on dynamic airspace networks and reinforcement learning. Background Art

[0002] The air traffic control system plays a vital role in ensuring the safe and orderly operation of flights and the rational allocation of airspace resources. It not only ensures passenger safety but also reduces costs by improving efficiency. As air traffic continues to grow, the efficiency and safety of air traffic control become increasingly critical.

[0003] Flight simulator training plays an irreplaceable role as a core means of training pilots and enhancing the capabilities of air traffic controllers. The highly simulated environment created by flight simulators allows for the simulation of extremely complex and diverse flight scenarios. Among these, making scientific and rational aircraft deployment decisions based on the current situation to assist controllers in their decision-making is both a key and challenging aspect of training and is directly related to flight safety and efficiency. However, existing aircraft deployment methods have certain limitations when dealing with complex and ever-changing flight simulator training scenarios, making it difficult to meet the growing demand for training.

[0004] Therefore, in-depth research on aircraft deployment methods based on current situation analysis is of great practical significance for improving the quality of flight simulator training, strengthening the ability of air traffic controllers to deal with complex situations, and thus ensuring the safety and smoothness of actual air transportation. Summary of the Invention

[0005] The purpose of this application is to provide an aircraft deployment method based on dynamic airspace networks and reinforcement learning, which can improve air traffic control efficiency, reduce conflict risks, and meet the air transportation needs of multiple flights and multiple aircraft types.

[0006] To achieve the above objectives, this application provides the following solutions:

[0007] This application provides an aircraft deployment method based on dynamic airspace network and reinforcement learning, including:

[0008] Obtain flight information and task priority weight information for each flight in the historical time period and the current time period;

[0009] Based on the flight information of each flight in the historical time period, the airspace unit is divided through the dynamic sub-airspace network to obtain an airspace division result; the airspace division result includes independent sub-airspace and shared airspace;

[0010] Inputting flight information and mission priority weight information of each flight in the current time period into a spatiotemporal buffer zone model to obtain a current flight path set; the spatiotemporal buffer zone model is used to calculate the safe zone area of ​​each flight based on the flight information, and determine the starting point, waypoint, destination, and flight schedule of each flight based on the safe zone area and mission priority weight information of each flight to obtain the current flight path set;

[0011] Obtaining a current disengagement action instruction set based on flight information, airspace division results, and a current flight path set for each flight in a current time period; the disengagement action instruction set includes an altitude adjustment instruction, a heading correction instruction, and a speed change instruction;

[0012] Each flight in the current time period is deployed according to the current release action instruction set.

[0013] According to the specific embodiments provided in this application, this application has the following technical effects:

[0014] The present application provides an aircraft deployment method based on dynamic airspace network and reinforcement learning. By obtaining the flight information and mission priority weight information of each flight in the historical time period and the current time period, the real-time status and flight conditions of each flight can be accurately grasped, which solves the problem of delayed information update in traditional air traffic control and realizes real-time monitoring and analysis of flight dynamics. Based on the flight information of each flight in the historical time period, the airspace is divided through a dynamic sub-airspace network to obtain the airspace division result including independent sub-airspace and shared airspace, which solves the problem of uneven distribution and low utilization efficiency of airspace resources and realizes the optimization of airspace resources. Dynamic adjustment and optimal configuration of flight resources; by inputting flight information and task priority weight information into the spatiotemporal buffer zone model, the safe area of ​​each flight is calculated, and the flight path set is determined accordingly, which solves the problem of unscientific and inaccurate flight path planning of flights, and realizes effective path planning based on safety and priority; by generating a set of release action instructions including altitude adjustment, heading correction and speed change according to flight status information, airspace division results and flight path set, and deploying each flight, it solves the problem of potential conflicts and inflexible deployment between flights, and realizes accurate deployment of flights and safe flight.

[0015] To sum up, this application not only improves the efficiency and safety of air traffic control, but also enhances the utilization efficiency of airspace resources, meets the air transportation needs of multiple flights and multiple aircraft types, and has significant practical application value and broad development prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 This is a diagram of the application environment of an aircraft deployment method based on dynamic airspace network and reinforcement learning in one embodiment of the present application.

[0018] Figure 2 A flowchart of an aircraft deployment method based on a dynamic airspace network and reinforcement learning is provided in accordance with an embodiment of the present application.

[0019] Figure 3 A flowchart of an aircraft deployment method based on a dynamic airspace network and reinforcement learning is provided in accordance with another embodiment of the present application.

[0020] Figure 4 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0023] The aircraft deployment method based on dynamic airspace network and reinforcement learning provided in the embodiment of the present application can be applied to Figure 1In the application environment shown, the terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. Terminal 102 may send flight information and mission priority weight information for each flight in the historical time period and the current time period to server 104. After receiving the flight information and mission priority weight information for each flight in the historical time period and the current time period, server 104 divides airspace units using a dynamic sub-airspace network based on the flight information of each flight in the historical time period to obtain an airspace division result. The airspace division result includes independent sub-airspaces and shared airspaces. The flight information and mission priority weight information for each flight in the current time period are input into a spatiotemporal buffer zone model to obtain a current flight path set. The spatiotemporal buffer zone model is used to calculate the safe zone area of ​​each flight based on the flight information, and determine the starting point, waypoint, destination, and flight schedule of each flight based on the safe zone area and mission priority weight information of each flight to obtain a current flight path set. A current escape action instruction set is obtained based on the flight information of each flight in the current time period, the airspace division result, and the current flight path set. The escape action instruction set includes an altitude adjustment instruction, a heading correction instruction, and a speed change instruction. Each flight in the current time period is deployed based on the current escape action instruction set. The server 104 may feed back the obtained current release action instruction set to the terminal 102. Furthermore, in some embodiments, the aircraft deployment method based on the dynamic airspace network and reinforcement learning may also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 may directly perform flight deployment processing based on the flight information and task priority weight information of each flight in the historical time period and the current time period. Alternatively, the server 104 may obtain the flight information and task priority weight information of each flight in the historical time period and the current time period from a data storage system and perform flight deployment processing based on the flight information and task priority weight information of each flight in the historical time period and the current time period.

[0024] The terminal 102 may be, but is not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. Portable wearable devices may include smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as a standalone server or a server cluster consisting of multiple servers, or may be a cloud server.

[0025] In an exemplary embodiment, Figure 2As shown, a method for aircraft deployment based on dynamic airspace network and reinforcement learning is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in FIG. 1 is taken as an example to illustrate the method, which includes the following steps 201 to 205. Among them:

[0026] Step 201: Obtain flight information and task priority weight information for each flight in the historical time period and the current time period. The flight information includes flight position information, flight speed information, heading angle information, aircraft model parameter information, and meteorological data information.

[0027] Step 202: Based on the flight information of each flight in the historical time period, the airspace units are divided through the dynamic sub-airspace network to obtain an airspace division result; the airspace division result includes independent sub-airspaces and shared airspaces.

[0028] In step 203, the flight information and mission priority weight information of each flight in the current time period are input into the spatiotemporal buffer zone model to obtain the current flight path set. The spatiotemporal buffer zone model is used to calculate the safe zone area of ​​each flight based on the flight information, and determine the starting point, transit point, destination, and flight schedule of each flight based on the safe zone area and mission priority weight information of each flight to obtain the current flight path set.

[0029] Step 204 , obtaining a current release action instruction set based on the flight information of each flight in the current time period, the airspace division result and the current flight path set; the release action instruction set includes an altitude adjustment instruction, a heading correction instruction and a speed change instruction.

[0030] Step 205: Allocate each flight in the current time period according to the current release action instruction set.

[0031] By implementing the above steps 201 to 205, the present application can achieve efficient, safe and automated management of air traffic, optimize airspace resource allocation, improve flight operation efficiency, and enhance flight safety. It is particularly suitable for complex air transportation needs with multiple flights and multiple aircraft types.

[0032] In another exemplary embodiment of the present application, step 202 specifically includes:

[0033] Divide the airspace into multiple cubic grids.

[0034] According to the flight information of each flight in the historical time period, the flight density between adjacent cube grids is determined, and the modularity of the airspace partition is calculated based on the flight density through the modularity function.

[0035] Through the greedy algorithm, with the goal of maximizing the increase in the change of the spatial division modularity, the adjacent cube grids are merged to obtain the initial spatial unit division result.

[0036] Calculate the flight density between adjacent airspace units in the initial airspace unit division result, and return to the step of "using a greedy algorithm to maximize the increase in the change in the airspace division modularity, merge adjacent cube grids, and obtain the initial airspace unit division result" until the airspace is divided into a state containing only two airspace units, and obtain the airspace division result; among them, a high-flow airspace unit with high flight density is used as an independent sub-airspace; the other low-flow airspace unit with low flight density is used as a shared airspace.

[0037] As an optional implementation, the modularity function is:

[0038] .

[0039] in, represents the modularity of the spatial division at time t; represents the flight density from the i-th cube grid to the j-th cube grid at time t; represents the total number of flights flowing into and out of the i-th cube grid at time t; represents the total number of flights flowing into and out of the j-th cube grid at time t; represents the total number of flights flowing into and out of the airspace at time t; Indicates the airspace unit to which the i-th cube grid belongs at time t; Indicates the airspace unit to which the j-th cube grid belongs at time t; Represents the first indicator function.

[0040] In another exemplary embodiment of the present application, step 203 specifically includes:

[0041] The spatiotemporal buffer zone model includes a safety area calculation module and a time window insertion module which are connected in sequence.

[0042] The safety zone calculation module uses the following formula to calculate the safety zone area of ​​each flight in the current time period:

[0043] .

[0044] in, Indicates the time t The safe area of ​​each flight; No. The mission priority weight of each flight; Indicates the Minimum turning radius of aircraft for each flight; is the control response time; Indicates the The maximum speed of the aircraft for each flight.

[0045] The time window insertion module is used to:

[0046] Traverse the current flight path set. If the time interval between any two flights arriving at the current path intersection is less than the preset time interval, insert a preset time window to delay the flight with the lower priority weight according to the task priority weight information until all flights at the current path intersection meet the preset stop condition, and then determine the next path intersection.

[0047] As an optional implementation, the mission priority weight information specifically includes: commercial mission priority weight, cargo mission priority weight and military mission priority weight; wherein, cargo mission priority weight < commercial mission priority weight < military mission priority weight.

[0048] As an optional implementation, the calculation formula for the preset time interval is:

[0049] .

[0050] in, Indicates any Flights and The preset time interval for each flight to arrive at the intersection of the current path; Indicates the time t The safe area of ​​each flight; Indicates the time t The safe area of ​​each flight; Indicates the time t The speed of each flight; Indicates the time t The speed of each flight; Indicates the preset minimum time interval.

[0051] In another exemplary embodiment of the present application, step 204 specifically includes:

[0052] The flight information and airspace division results of each flight in the current time period are used as the input status information of each flight.

[0053] Based on the input state information of each flight, the DQN algorithm is used to adjust the action of each flight to obtain the current local release action set.

[0054] According to the current local release action set, the current release action instruction set is obtained through genetic algorithm.

[0055] In another exemplary embodiment of the present application, based on the input state information of each flight, the DQN algorithm is used to adjust the action of each flight to obtain the current local release action set, which specifically includes:

[0056] The input state information of each flight is fed into the trained deep Q network model, and the Q value of all actions is calculated using the following formula:

[0057] .

[0058] in, Indicates the The Q value of the flight taking the ath action; Indicates the The mission priority weight of each flight; 、 and represent the first weight coefficient, the second weight coefficient and the third weight coefficient respectively; Indicates the preset minimum separation distance; Indicates the The change in energy consumption after a flight takes the ath action.

[0059] By comparing the Q values ​​of all actions in each flight and selecting the action corresponding to the largest Q value as the optimal action for each flight, the local release action set at the current moment is obtained.

[0060] In another exemplary embodiment of the present application, according to the current local release action set, a current release action instruction set is obtained by a genetic algorithm, specifically including:

[0061] Based on the local release action set at the current moment, the following formula is used to calculate the Flights to Flight impact weight:

[0062] .

[0063] in, Indicates the Flights to Flight impact weight; and Respectively represent Flights and The feature vector corresponding to the optimal action of each flight; Indicates that the Flights and The feature vectors corresponding to the optimal actions of each flight are concatenated; represents the activation function, Represents the normalization function.

[0064] Based on the Flights to The flight impact weights are calculated and the action combination that minimizes the number of global conflicts is selected through a genetic algorithm to obtain the current set of release action instructions.

[0065] The calculation formula for the number of global conflicts is:

[0066] .

[0067] .

[0068] in, Indicates the number of global conflicts; represents the second indicator function, Indicates the Flights and The distance between flights, Indicates the preset safety distance.

[0069] In another exemplary embodiment of the present application, an aircraft deployment method based on dynamic airspace network and reinforcement learning is provided, such as Figure 3 As shown, specifically including:

[0070] Step 1: Dynamic sub-airspace network design.

[0071] In actual aviation traffic, flight traffic varies significantly across time periods and regions. The static airspace division methods used in related technologies cannot be flexibly adjusted based on real-time traffic flow, which can easily lead to congestion in some airspace and idle airspace resources in others, reducing overall airspace resource utilization. Therefore, a method is needed to dynamically divide airspace units based on real-time traffic flow to optimize the allocation and use of airspace resources. To this end, a dynamic sub-airspace network is designed that can dynamically divide airspace units based on real-time traffic flow and optimize airspace resource utilization.

[0072] Specifically, the dynamic sub-airspace network input is the real-time flight location (latitude, longitude, altitude), and aircraft parameters (speed, climb rate) and meteorological data (wind speed, turbulence area) can also be added for reference.

[0073] Among them, real-time flight position can determine the specific location and movement trajectory of the flight in the airspace; aircraft parameters affect the flight performance and behavior of the aircraft; and meteorological data will affect the flight safety and efficiency of the flight.

[0074] The output of the dynamic sub-spatial network is the dynamically adjusted spatial division result (topological structure).

[0075] The specific process is as follows:

[0076] Step 1.1: Airspace map modeling.

[0077] The airspace is divided into a 1km×1km×300m cubic grid. Each grid is a graph node, and the edge weight is the flight density between adjacent grids. Therefore, each grid has a maximum of eight edges connecting to eight surrounding grids. This grid size was chosen based on a comprehensive consideration of airspace management accuracy and computational complexity. Smaller grids can more accurately represent flight locations and flow distribution, but increase computational complexity. Larger grids, while requiring less computation, may not accurately reflect flight distribution. The 1km×1km×300m grid size strikes a good balance between these two requirements. By treating each grid as a graph node and setting the edge weight to the flight density between adjacent grids, the relationship between flight flows across different grids can be intuitively represented, providing a foundation for subsequent airspace partitioning.

[0078] Flight density can be calculated as follows: Use real-time flight location information to count the flight flow between each grid and the adjacent grid at each time within a period of time (such as the past 12 months). Specifically, it can be calculated by counting the grids that each flight passes through: For a certain flight T, the specific location and movement trajectory of the flight in the airspace can be obtained based on its historical real-time flight location. Then, when it flies out of the airspace grid node i and enters the adjacent grid node j, the weight of i→j is increased by 1, and finally the number of flights between adjacent grids is obtained. Dividing it by 12 months is the flight density, that is, the flow between grids. .

[0079] Step 1.2: Calculate dynamic modularity.

[0080] Define the modularity function:

[0081] .

[0082] in, represents the modularity of the spatial division at time t; It represents the flight density from the i-th cube grid to the j-th cube grid at time t, reflecting the closeness of the connection between the two grids; It represents the total number of flights flowing into and out of the ith cube grid at time t, that is, the sum of the number of flights flowing into and out of the grid i, reflecting the importance of the ith cube grid in the entire airspace; represents the total number of flights flowing into and out of the j-th cube grid at time t; represents the total number of flights flowing into and out of the airspace at time t; Indicates the airspace unit to which the i-th cube grid belongs at time t; Indicates the airspace unit to which the j-th cube grid belongs at time t; Represents the first indicator function.

[0083] The modularity function measures the rationality of the current airspace division by calculating the traffic differences between different grids. The larger the modularity Q value, the more reasonable the airspace division. Indicates whether the i-th cube grid and the j-th cube grid belong to the same spatial unit. hour, ,otherwise That is, when the two belong to the same airspace unit, they are counted into the modularity, otherwise they are not counted.

[0084] Step 1.3: Real-time optimization.

[0085] Merging high-traffic areas into independent sub-airspaces can reduce interference between different flights and improve flight efficiency; merging low-traffic areas into shared airspace can avoid waste of resources and improve airspace utilization.

[0086] A greedy algorithm maximizes the Q value, merging high-traffic areas into independent sub-airspaces and low-traffic areas into shared airspaces. A greedy algorithm takes the best or optimal (i.e., most favorable) option at each step, hoping to achieve a globally optimal result. In this scenario, the optimal airspace partitioning is gradually found by continuously selecting operations that increase the modularity Q value (i.e., merging or splitting airspace units).

[0087] Specifically, the grid cells divided in step 1.1 are used as the initial state of the spatial cells, and the Q value is calculated according to step 1.2. Then, all spatial cells are traversed to perform merging operations, the change in Q value under each operation is calculated, and the operation that increases the Q value with the largest increase is selected for execution. After the operation is executed, the spatial cell topology structure is updated and the Q value is recalculated, and the above operation is repeated. The algorithm terminates when no operation that increases the Q value is found. At this time, the spatial division method is the result that maximizes the Q value.

[0088] In this implementation, the results can be updated at regular intervals. For example, the airspace division can be performed every hour based on the flight status within the next hour. Alternatively, when a conflict occurs, the results can be updated based on the new route status after adjustment and resolution.

[0089] Step 2: Classify and schedule multiple models.

[0090] Different aircraft models have significant differences in performance, such as speed and maneuverability. Failure to account for these differences in air traffic control and the adoption of a unified scheduling strategy can easily lead to scheduling conflicts and flight delays. Therefore, it is necessary to design a hierarchical priority queue based on aircraft performance differences to more rationally arrange flight paths and reduce scheduling conflicts.

[0091] To this end, a multi-type aircraft scheduling method (MTAS) was designed based on the performance differences of various aircraft types. The spatio-temporal buffer zone model was designed to reduce scheduling conflicts.

[0092] The input for step 2 is the aircraft type parameters (maximum speed , minimum turning radius ), mission urgency (commercial, cargo, and military, with priority weights set to 1, 0.9, and 1.1, respectively; for commercial aircraft, this application divides aircraft into three categories based on speed and maneuverability, with weights of 1.05, 1, and 0.95, respectively. See step 2.1 for details). Among them, the aircraft model parameters directly affect the flight capability and behavior of the aircraft, the maximum speed determines the upper limit of the aircraft's flight speed, and the minimum turning radius reflects the aircraft's maneuverability. Mission urgency reflects the importance and urgency of the flight mission. This information can be used to determine the priority and scheduling strategy of the flight. The output is a set of flight paths sorted by priority.

[0093] A spatiotemporal buffer zone model was proposed, combining queuing theory with kinematic constraints to optimize flight conflicts. Queuing theory can be used to analyze the temporal relationship between flights waiting and flying, while kinematic constraints consider the physical characteristics of aircraft and flight safety requirements. By combining these two, the designed "spatiotemporal buffer zone" model can more accurately handle conflicts between flights and ensure flight safety.

[0094] The specific process is as follows:

[0095] Step 2.1: Classify the models.

[0096] For commercial aircraft, this application divides aircraft into three categories based on speed and maneuverability, with weights of 1.05, 1, and 0.95 respectively.

[0097] Class A (large passenger aircraft): Large passenger aircraft typically carry more passengers and their missions are more important, so they are given higher priority. However, due to their larger size, they have relatively poor maneuverability and require more space and time to operate during flight.

[0098] Class B (mid-size): , medium priority. Medium aircraft are intermediate in speed and maneuverability, and their priority is set accordingly.

[0099] Class C (small aircraft / drone): Small aircraft and drones typically perform auxiliary or specialized tasks that are relatively unimportant, so their priority is set to low. However, they are highly maneuverable and can more flexibly adjust their flight paths during flight.

[0100] Step 2.2: Generate spatiotemporal buffer zones.

[0101] Generate a safety zone for each flight. The calculation formula for the safety zone area is:

[0102] .

[0103] in, Indicates the time t The safe area of ​​each flight; No. The mission priority weight of each flight; Indicates the Minimum turning radius of aircraft for each flight; is the control response time; Indicates the The meaning of this formula is that, combined with its priority weight, the maximum speed of the aircraft in the first flight Minimum turning radius of aircraft on flights Plus the regulatory response time The aircraft is at maximum speed The flight distance is the safety zone radius R, that is: , generating a circular safety zone. This safety zone ensures that the aircraft has enough space to respond to emergencies during flight and avoid collisions with other aircraft.

[0104] Insert time window at path intersection: If the time difference between two flights is expected to enter the same area minutes, then the low-priority flight is delayed according to the priority weight, satisfying the constraints:

[0105] .

[0106] in, Indicates any Flights and The preset time interval for each flight to arrive at the intersection of the current path; Indicates the time t The safe area of ​​each flight; Indicates the time t The safe area of ​​each flight; Indicates the time t The speed of each flight; Indicates the time t The speed of each flight; Indicates the preset minimum time interval. The process first determines the time difference between the two flights expected to enter the same area. Is it less than 2 minutes? minutes, indicating that the time interval is sufficient and no special treatment is required; if Minutes, further judgment is required. Further calculation of the new time interval Then determine whether the time difference between the two flights expected to enter the same area is less than the new time interval If it is less than , the low priority flight will be delayed. 2 minutes is a preset initial safety time interval threshold. This step is performed first to quickly screen out flights with relatively sufficient time intervals and no conflict risk in the initial stage, reducing unnecessary complex calculations. For cases where the time difference is less than 2 minutes, the new time interval is accurately calculated. To further determine whether the flight needs to be delayed, balancing computing efficiency and safety.

[0107] The purpose of this constraint is to ensure that there is enough time between two flights when they enter the same area to avoid collision. The time interval is calculated by dividing the sum of the safe areas of the two flights by the sum of their speeds. The larger of the two minutes is used as the new time interval. If the time difference between the two flights entering the same area is less than the new time interval , low-priority flights are delayed, and finally a set of flight paths sorted by priority is obtained to meet safety requirements. This flight path set is a combination of flight paths after obtaining today's flight missions, and flights of different aircraft models and different mission urgency are sorted according to priority E. In this set, the flight paths of high-priority flights are determined first and are more secure, and the flight paths of low-priority flights are adjusted according to high-priority flights and safety constraints. The set is represented as a list, where each element represents the flight path information of a flight, including the starting point, transit points, end point, and time schedule during the flight, indicating the flight order and path planning of different flights.

[0108] In addition, the time window size for inserting delayed low-priority flights is not fixed. First, a time value is calculated based on the sum of the safety area and the speed of the two flights, and then compared with the preset minimum time interval. Take the larger value to get the new time interval , which determines the size of the time window to insert when delaying low-priority flights.

[0109] Step 3: Conflict prediction and resolution based on multi-agent reinforcement learning.

[0110] During the execution of today's flight plan generated in Step 2, a conflict prediction and resolution method based on multi-agent reinforcement learning is designed to monitor aircraft flight status and avoid possible conflicts such as dangerous approaches. This method monitors aircraft status in real time. In the complex aviation traffic environment, conflicts between flights can occur at any time. Traditional conflict prediction and resolution methods often rely on fixed rules and models, lacking real-time and flexibility. Multi-agent reinforcement learning allows each flight to act as an agent, making autonomous decisions based on its own state and that of its surroundings, thereby more effectively predicting and resolving conflicts. To this end, a MARL-CDR module is constructed to use distributed agents to predict conflicts and generate resolution strategies in real time.

[0111] Specifically, its input flight status is (position (x, y, z), speed v, heading angle ), and whether the location is a shared area (determined by the DSAN partitioning results designed in step 1). Flight status information is the basis for the agent's decision-making, used to understand the position, speed, and heading angle of itself and surrounding flights, supporting the agent's prediction of potential conflicts. The DSAN partitioning results provide information about the current airspace structure, allowing the agent to adjust its decisions based on the airspace partitioning.

[0112] Its output is a set of release instructions (height adjustment , Course Correction , speed change ).

[0113] A "local-global" dual-loop reinforcement learning architecture is designed to integrate individual decision-making with global coordination. This architecture considers the local decisions of each flight agent to quickly respond to potential conflicts, while also comprehensively evaluating and adjusting the decisions of all agents through a global coordinator to ensure overall flight safety and efficiency.

[0114] The specific process is as follows:

[0115] Step 3.1: Build a local agent: Each flight corresponds to an agent, and the DQN algorithm (Deep Q-Network) is used to optimize individual actions. The agent includes the following two aspects:

[0116] The state space consists of the state of the current flight and its neighboring flights (within 10 km) as input. The 10 km range was chosen based on a comprehensive consideration of the impact of flights on each other and computational complexity. Flights within this range have a greater impact on potential conflicts with the current flight, while also minimizing the computational burden of considering too many flight information.

[0117] The output action space metrics include climb 100 meters, descend 100 meters, turn left 5 degrees, turn right 5 degrees, accelerate 5%, and decelerate 5%, which can be represented by a vector of length 6. These actions are common ways for aircraft to adjust their flight status. By selecting appropriate actions, the agent can avoid conflicts with other flights.

[0118] Through the reward function, the agent can learn the action strategy that minimizes energy consumption while ensuring safety. The specific formula of the reward Q function is:

[0119] .

[0120] in, Indicates the The Q value of the flight taking the ath action; Indicates the The mission priority weight of each flight; 、 and represent the first weight coefficient, the second weight coefficient and the third weight coefficient respectively; Represents the preset minimum separation distance, which is obtained from the state space calculation. That is, for the state space of an agent, the Euclidean distance of the aircraft closest to the agent is calculated according to its three-dimensional coordinates as ; Indicates the The change in energy consumption after a flight takes the ath action is related to the size and power of the specific aircraft model, and can be obtained based on the performance indicators of each aircraft model. and Used to balance the effects of changes in minimum separation distance and energy consumption on rewards. Represents the impact of the generated adjustment action strategy on the airspace distribution. When the output action causes the aircraft to enter the shared airspace from an independent sub-airspace, =1.5; when the output action causes the aircraft to enter the independent sub-airspace from the shared airspace, =0.7; when the output action makes the aircraft still in the independent sub-airspace, =0.9; when the output action makes the aircraft still in the shared airspace, =1.

[0121] In this way, the model is forced to use the shared sub-airspace with low traffic volume to resolve conflicts. The larger the value, the safer the flight, the higher the reward, and the energy consumption changes. The smaller it is, the more efficient the flight is and the higher the reward is.

[0122] Specifically, the DQN and target network are built, and the agent's state and actions are combined into input-output data pairs, from which random sampling is performed for training. The reward value Q is calculated and the DQN parameters are updated based on Q. After training, a state space is input, and a specific action selected from the action space is output.

[0123] According to step 2.2 Whether it complies with safety standards. The agent within the safe zone radius R inputs its state space into the DQN, which then outputs the action for the aircraft agent. This can be represented by a 1×6 vector, where a 1 in a bit in the vector indicates the selection of a specific action space.

[0124] Step 3.2, Global Coordinator:

[0125] Assuming there are N flights, based on the actions for each flight agent obtained in step 3.1, we can obtain N×6 action combinations, which means that each flight has 6 possible actions. Based on this action combination matrix, we evaluate the conflict risk and obtain the final action combination for all N flights today. This involves the following two steps:

[0126] ① Calculate the impact weight of flight i on flight j:

[0127] .

[0128] in, Indicates the Flights to Flight impact weight; and Respectively represent Flights and The feature vector corresponding to the optimal action of each flight; Indicates that the Flights and The feature vectors corresponding to the optimal actions of each flight are concatenated; represents the activation function, Represents the normalization function. In this way, the first Flights to The influence weight of each flight. The larger the weight, the greater the impact of the flight. Some characteristics of the first flight (such as position, heading, etc.) are different from those of the second flight. There is a strong correlation or potential conflict trend between the flights, which means that Flights to The greater the impact of each flight.

[0129] ② Output final action: Select the number of global conflicts The smallest action combination. Among them, Indicates the number of global conflicts; represents the second indicator function, ; Indicates the Flights and The distance between flights, Indicates the preset safety distance calculated in step 2.2.

[0130] A genetic algorithm selects the action combination that minimizes the number of global conflicts, thereby ensuring flight safety across the entire airspace from a global perspective. Specifically, each flight's action combination is treated as a chromosome of length N. An initial population of action combinations is randomly generated, with each chromosome representing a selected action combination. For each action combination in the population, the number of global conflicts is calculated and used as its fitness. Next, a selection operation is performed, selecting the action combination with the lowest number of conflicts to advance to the next generation. A crossover operation then randomly selects two action combinations by exchanging some of their actions to generate new combinations. Simultaneously, a mutation operation is performed on certain action combinations with a certain probability, randomly changing individual actions. These steps are repeated, allowing the population to gradually evolve, ultimately approaching the optimal action combination—the one with the lowest number of conflicts. Executing the actions corresponding to this combination achieves the greatest possible resolution of conflicts.

[0131] During the process of air traffic controllers deploying aircraft on flight days, the module designed in this step can read the real-time status information of each flight, then analyze and output the optimal action combination, and provide it to air traffic controllers as deployment suggestions to avoid potential conflicts in the actual deployment process.

[0132] In summary, this application has the following technical effects:

[0133] Through step 1, a dynamic sub-airspace network (DSAN) is designed, which can dynamically adjust the airspace unit division based on real-time flight positions, aircraft parameters, and meteorological data, thereby optimizing airspace resource utilization. Specifically, DSAN divides the airspace into fine-grained cubic grids and calculates the modularity value based on the flight density between adjacent grids. This allows high-traffic areas to be merged into independent sub-airspaces to reduce interference, while low-traffic areas are merged into shared airspace to avoid resource waste. This approach not only improves flight efficiency but also reduces the risk of congestion and delays.

[0134] In step 2, a multi-type scheduling method (MTAS) was designed. It categorizes flights based on their performance and mission urgency, generating spatiotemporal buffer zones, thereby reducing scheduling conflicts and improving flight safety. MTAS first divides aircraft into three categories based on speed and maneuverability. It then utilizes queuing theory and kinematic constraints to generate a safe zone for each flight, inserting time windows at route intersections to ensure sufficient spacing and prevent potential conflicts. Consequently, flights of varying priorities can be rationally scheduled based on their characteristics, ensuring the orderly operation of each flight.

[0135] Through step 3, a Multi-Agent Reinforcement Learning for Conflict Detection and Resolution (MARL-CDR) module is constructed. This module enables each flight to act as an agent, making autonomous decisions based on its own state and that of its surrounding environment, thereby achieving more efficient conflict prediction and resolution strategies. This module adopts a "local-global" dual-loop reinforcement learning architecture, in which local agents use the DQN algorithm to optimize individual actions, while the global coordinator uses a graph attention network to evaluate the conflict risk of all local action proposals and select the optimal action combination to minimize the number of global conflicts. This design not only enhances the system's real-time responsiveness but also ensures flight safety and efficiency throughout the entire airspace.

[0136] The present application also provides an application scenario, which applies the above-mentioned aircraft allocation method based on dynamic airspace network and reinforcement learning. Specifically: the aircraft allocation method based on dynamic airspace network and reinforcement learning provided in this embodiment can be applied in the airspace management scenario around the airport. The airspace management scenario around the airport includes the take-off phase, the cruising phase and the landing phase; the flight enters the cruising phase from the take-off phase, obtains the optimal flight path through the allocation of the dynamic airspace network, and enters the landing phase. The aircraft allocation method based on dynamic airspace network and reinforcement learning provided in this embodiment belongs to the flight path optimization phase in the cruising phase. Specifically, the method monitors and plans the flight path in real time through the dynamic sub-airspace network division and the spatiotemporal buffer zone model to ensure the flight safety and efficiency of the flight during the cruising phase, while reducing the risk of flight delays and conflicts, and improving the automation and intelligence level of the overall air traffic control.

[0137] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 4As shown. The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, memory, and I / O interface are connected via a system bus, and the communication interface is connected to the system bus via the I / O interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store flight information processing data. The I / O interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an aircraft deployment method based on a dynamic airspace network and reinforcement learning.

[0138] Those skilled in the art will understand that Figure 4 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above-mentioned method embodiments when executing the computer program.

[0139] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0140] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0141] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0142] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0143] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0144] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0145] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. An aircraft deployment method based on dynamic airspace network and reinforcement learning, characterized in that: The aircraft deployment method based on dynamic airspace network and reinforcement learning includes: Obtain flight information and task priority weight information for each flight in the historical time period and the current time period; According to the flight information of each flight in the historical time period, the airspace unit is divided by a dynamic sub-airspace network to obtain an airspace division result, specifically including: dividing the airspace into multiple cubic grids; according to the flight information of each flight in the historical time period, determining the flight density between adjacent cubic grids, and calculating the airspace division modularity through a modularity function based on the flight density; merging adjacent cubic grids by a greedy algorithm with the goal of maximizing the change increase of the airspace division modularity to obtain an initial airspace unit division result; calculating the flight density between adjacent airspace units in the initial airspace unit division result, returning to the step of "merging adjacent cubic grids by a greedy algorithm with the goal of maximizing the change increase of the airspace division modularity to obtain an initial airspace unit division result" until the airspace is divided into a state containing only two airspace units, thereby obtaining an airspace division result; the airspace division result includes independent sub-airspaces and shared airspaces; wherein, a high-flow airspace unit with high flight density is used as an independent sub-airspace; and another low-flow airspace unit with low flight density is used as a shared airspace; The flight information and mission priority weight information of each flight in the current time period are input into the spatiotemporal buffer zone model to obtain a current flight path set; the spatiotemporal buffer zone model is used to calculate the safe area of ​​each flight based on the flight information, and determine the starting point, transit point, destination, and flight schedule of each flight based on the safe area and mission priority weight information of each flight to obtain the current flight path set, which specifically includes: The spatiotemporal buffer zone model includes a safety zone calculation module and a time window insertion module connected in sequence. The safety zone calculation module calculates the safety zone area of ​​each flight in the current time period using the following formula: ; in, Indicates the time t The safe area of ​​each flight; No. The mission priority weight of each flight; Indicates the Minimum turning radius of aircraft for each flight; is the control response time; Indicates the Maximum aircraft speed for each flight; The time window insertion module is used to: traverse the current flight path set, and if the time interval between any two flights arriving at the current path intersection is less than a preset time interval, then insert a preset time window to delay the flight with a lower priority weight according to the task priority weight information, until all flights at the current path intersection meet the preset stop condition, and then determine the next path intersection; Obtaining a current set of disengagement action instructions based on the flight information, airspace division results, and current flight path set of each flight in the current time period, specifically comprising: using the flight information and airspace division results of each flight in the current time period as input state information of each flight; adjusting the action of each flight using a DQN algorithm based on the input state information of each flight to obtain a current set of local disengagement actions; obtaining a current set of disengagement action instructions based on the current set of local disengagement actions using a genetic algorithm; the set of disengagement action instructions including an altitude adjustment instruction, a heading correction instruction, and a speed change instruction; Each flight in the current time period is deployed according to the current release action instruction set.

2. The aircraft deployment method based on dynamic airspace network and reinforcement learning according to claim 1 is characterized in that: The modularity function is: ; in, represents the modularity of the spatial division at time t; represents the flight density from the i-th cube grid to the j-th cube grid at time t; represents the total number of flights flowing into and out of the i-th cube grid at time t; represents the total number of flights flowing into and out of the j-th cube grid at time t; represents the total number of flights flowing into and out of the airspace at time t; Indicates the airspace unit to which the i-th cube grid belongs at time t; Indicates the airspace unit to which the j-th cube grid belongs at time t; Represents the first indicator function.

3. The aircraft deployment method based on dynamic airspace network and reinforcement learning according to claim 1, characterized in that: Mission priority weight information specifically includes: commercial mission priority weight, cargo mission priority weight, and military mission priority weight; among which, cargo mission priority weight < commercial mission priority weight < military mission priority weight.

4. The aircraft deployment method based on dynamic airspace network and reinforcement learning according to claim 1, characterized in that: The calculation formula for the preset time interval is: ; in, Indicates any Flights and The preset time interval for each flight to arrive at the intersection of the current path; Indicates the time t The safe area of ​​each flight; Indicates the time t The safe area of ​​each flight; Indicates the time t The speed of each flight; Indicates the time t The speed of each flight; Indicates the preset minimum time interval.

5. The aircraft deployment method based on dynamic airspace network and reinforcement learning according to claim 1, characterized in that: Based on the input state information of each flight, the DQN algorithm is used to adjust the action of each flight to obtain the current local release action set, which includes: The input state information of each flight is fed into the trained deep Q network model, and the Q value of all actions is calculated using the following formula: ; in, Indicates the The Q value of the flight taking the ath action; Indicates the The mission priority weight of each flight; 、 and represent the first weight coefficient, the second weight coefficient and the third weight coefficient respectively; Indicates the preset minimum separation distance; Indicates the The change in energy consumption after a flight takes the ath action; By comparing the Q values ​​of all actions in each flight, and selecting the action corresponding to the largest Q value as the optimal action for each flight, the current local release action set is obtained.

6. The aircraft deployment method based on dynamic airspace network and reinforcement learning according to claim 1, characterized in that: According to the current local release action set, the current release action instruction set is obtained through the genetic algorithm, which specifically includes: Based on the current local release action set, the first Flights to Flight impact weight: ; in, Indicates the Flights to Flight impact weight; and Respectively represent Flights and The feature vector corresponding to the optimal action of each flight; Indicates that the Flights and The feature vectors corresponding to the optimal actions of each flight are concatenated; represents the activation function, represents the normalization function; Based on the Flights to The flight impact weights are calculated and the action combination that minimizes the number of global conflicts is selected through a genetic algorithm to obtain the current set of release action instructions.

7. The aircraft deployment method based on dynamic airspace network and reinforcement learning according to claim 6, characterized in that: The calculation formula for the number of global conflicts is: ; ; in, Indicates the number of global conflicts; represents the second indicator function, Indicates the Flights and The distance between flights, Indicates the preset safety distance.

Citation Information

Patent Citations

  • Conflict minimization flight path collaborative planning method considering high-altitude wind time variation

    CN115938162A

  • Multi-level low-altitude air route network construction method in complex urban environment

    CN119516846A