A large-scale low-altitude flight plan intelligent collaborative scheduling method based on reinforcement learning
Patent Information
- Application Number
- CN202610798637.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-15
AI Technical Summary
基于规则的方法求解速度快、可求解问题规模大,但求解性能一般;基于优化的方法求解性能优异,但受问题规模和求解速度限制严重
[0122] This invention provides a large-scale intelligent collaborative scheduling method for low-altitude flight plans based on reinforcement learning. Addressing the needs of large-scale integrated operation in differentiated low-altitude scenarios, it achieves conflict-free intelligent scheduling of low-altitude flight plans under arbitrary scenario scales based on reinforcement learning and partially observable Markov decision processes. Compared with existing rule-based conflict scheduling methods, it achieves a higher number of conflict-free executions and stronger performance. Compared with optimization-based conflict scheduling methods, it has a faster solution speed and larger scale, possessing 10... 6 With a solution capacity of / hour, it can meet the needs of urban-level low-altitude flight plan management and control, has strong engineering application potential, and can effectively improve the service and operation support capabilities of low-altitude airspace.
Smart Images

Figure CN122759628A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of aviation technology, specifically relating to a method for intelligent collaborative scheduling of large-scale low-altitude flight plans based on reinforcement learning. Background Technology
[0002] Low-altitude flight, with its advantages in utilizing three-dimensional space, has become a crucial engine for economic and social development. In 2025, my country's drone flight hours exceeded 45 million, a growth rate of nearly 70%, covering scenarios such as cultural tourism, consumer experiences, logistics, and urban management. Large-scale, multi-scenario low-altitude flight has become an inevitable trend. In high-density operational scenarios, relying solely on real-time conflict perception and collision avoidance strategies faces performance bottlenecks. It is necessary to pre-allocate and resolve conflicts through pre-flight planning management. Traditional flight planning methods can be mainly divided into rule-based methods and optimization-based methods. Rule-based methods offer fast solution speeds and can solve large-scale problems, but their performance is generally average; optimization-based methods offer excellent solution performance, but are severely limited by problem size and solution speed. To improve the control capabilities of low-altitude flight, there is an urgent need to propose an intelligent, scalable, and highly robust large-scale low-altitude flight planning collaborative scheduling technology. This technology should integrate conflict resolution, mission value, and customer satisfaction goals to ultimately achieve a low-altitude flight planning scheduling scheme that balances scale, speed, and performance, providing core technical support for the safe and efficient large-scale operation of low-altitude flights. Summary of the Invention
[0003] To address this, the present invention provides a large-scale intelligent collaborative scheduling method for low-altitude flight plans based on reinforcement learning, which has scalability for any scenario scale and can realize the scheduling of low-altitude flight plan conflicts for city-level needs, thereby solving the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides the following technical solution:
[0005] A method for intelligent collaborative scheduling of large-scale low-altitude flight plans based on reinforcement learning includes the following steps:
[0006] Step 1. Based on the mission type (line operation and area operation), start time, end time, acceptable adjustment time, maximum flight speed, and aircraft endurance of the low-altitude flight plan, define the flight plan allocation constraints, such as the adjustment range of start time, end time, and operation duration for each flight plan.
[0007] Step 2. Based on the airspace usage patterns of route operations and area operations, establish a flight plan conflict detection model that integrates altitude conflict detection, horizontal conflict detection, and temporal conflict detection.
[0008] Step 3. Establish a hierarchical decision-making model of "height and execution decision - start and end time decision", design the observation and action space required for each level of decision, define the training rewards for each agent, and establish a reinforcement learning model based on a partially observable Markov decision process.
[0009] Step 4. Develop decision-making functions for the corresponding altitude allocation agent, route operation time allocation agent, and area operation time allocation agent, and design the agent structure based on graph neural networks and self-attention mechanisms.
[0010] Step 5. Based on the characteristics of dense action space and non-independent discrete adjacent actions in the allocation model, the PPO (Proximal Policy Optimization) algorithm is modified by adding an adjacent action advantage diffusion mechanism, and a more efficient agent training is achieved by combining independent training and joint training.
[0011] Step 6. Utilize the trained agent and combine it with a deterministic spatiotemporal occupancy masking mechanism to achieve intelligent collaborative scheduling of conflict-free low-altitude flight plans of any scale.
[0012] Furthermore, the flight plan allocation constraint construction in step 1 includes:
[0013] Step 1.1: Establish a low-altitude flight plan model for a given low-altitude flight plan. It can be represented as:
[0014]
[0015] in, and They are respectively The type of plan and the nature of the task. Airspace usage type (line operations and area operations). For the coordinates of the airspace point, For use of height, and For departure and arrival times, For speed, Operating modes (one-way and round-trip). The status is planned (not allocated, allowed to execute, canceled). To determine whether there are flight altitude requirements, and For the permitted maximum and minimum airspace altitudes, The maximum allowable time adjustment amount, For the maximum permissible flight time, The maximum permissible flight speed.
[0016] Step 1.2: Calculate the allocation constraints, specifically including:
[0017] (1) Start time adjustment range constraints:
[0018]
[0019] in, This refers to the rescheduled departure time.
[0020] (2) End time adjustment range constraints:
[0021]
[0022] in, This refers to the arrival time after the relocation.
[0023] (3) Constraints on the range of adjustment of route operation time:
[0024]
[0025]
[0026] in, The flight distance for the route.
[0027] (4) Constraints on the range of regional operation time adjustment:
[0028]
[0029] Furthermore, the flight plan conflict detection model in step 2 is designed for flight plans. and Its conflict resolution process includes:
[0030] Step 2.1: Execution status check. If and If any plan is in the "cancelled" state, it is considered that there is no conflict; otherwise, proceed to step 2.2.
[0031] Step 2.2: High Conflict Judgment. When and If an intersection exists, proceed to step 2.3; otherwise, it is assumed there is no conflict. For If its status is "unallocated", then and Take respectively and If its status is "Execution Allowed", then and All .
[0032] Step 2.3: Horizontal Conflict Judgment. When and The minimum distance between projections on the same plane is less than the horizontal safety interval. If the conflict is not present, proceed to step 2.4; otherwise, it is assumed that there is no conflict.
[0033] Step 2.4: Time Conflict Judgment. For and Flight segments / areas with a horizontal safety interval smaller than the standard distance. Entry and exit times are and ,like For line operation, then and Based on the location and departure time of their conflicting flight segments Arrival time If it is a regional operation, then the departure time should be extracted separately. and arrival time .when and If there is an intersection, it is considered a conflict; otherwise, it is considered a non-conflict. Wherein... If the status is "unallocated", then ,like If the status is "Execution Allowed", then ; For safety time intervals.
[0034] Step 2.5: Conflict Type Differentiation. If and The status is "Execution Allowed". If there is a conflict between the two, it is "Confirmed Conflict"; otherwise, it is "Potential Conflict".
[0035] Furthermore, the reinforcement learning model based on a partially observable Markov decision process in step 3 is constructed through the following steps:
[0036] Step 3.1: Establish a hierarchical decision-making model of "altitude and execution decision-start and end time decision-making": First, the altitude allocation agent decides whether to allow the plan to be executed based on the determination of multiple altitude layers and potential spatiotemporal occupancy information. If execution is allowed, the altitude layer to be used for the plan is selected and handed over to the time adjustment agent for further decision-making. According to different airspace usage types, the corresponding time adjustment agent makes decisions on the start and end times of the plan to complete the allocation of the plan.
[0037] Step 3.2: Establish an environmental state model. The environmental state can be represented as follows: ,in For flight plan The state at time step t. The agent makes decisions in ascending order of the expected departure times of each flight plan, deciding on one flight plan and updating the environmental state at each time step.
[0038] Step 3.3: Establish a local observation model. The local observations input to different agents include:
[0039] (1) Highly coordinated local observation of intelligent agents
[0040] Because there are complex potential conflicts between the plans, and the number of potential conflicting objects varies among different plans, this conflict relationship is represented by a graph structure, and key features are pieced together as local observations of a highly coordinated agent, corresponding to time step t. Local observations during allocation are represented as follows:
[0041]
[0042] in, For node information, for As feature information of nodes To and Plans with potential conflict relationships are used as feature information for nodes. To and A set of plans with potentially conflicting relationships; for Edge characteristics between plans that have potentially conflicting relationships; For the determined spatiotemporal resource occupancy at each altitude, for The proportion of spatiotemporal resources that are confirmed to be occupied and unusable at height layer k. This is the highest level within the environment; This is altitude preference information, when there is a flight altitude requirement. The one-hot encoding is used; otherwise, it is an all-zero matrix. for The planned quantity was not allocated.
[0043]
[0044] in, for The execution value is determined by its program type, mission type, and flight service distance / area.
[0045]
[0046] in, and for In, it and The conflict segments are arranged according to their starting and ending proportions within the overall route (if...). For regional operation tasks =0、 =1); for and The overall angle of the conflict segment (if at least one area operation task exists within it) =180°); to for and The conflict is highly repetitive and occupies the one-heat variable. If it is "not allocated", then it is and The intersection height, if If "Allow execution" is enabled, then... The unique heat variable.
[0047] (2) Local observation of the intelligent agent for route operation time allocation
[0048] The flight path operation time allocation agent is activated after the altitude allocation agent selects the flight altitude; therefore, it only focuses on the potential and certain spatiotemporal resource occupancy at that altitude. For flight path operation tasks... The proportion of flights to the designated routes The spatiotemporal resource occupancy requirement can be expressed as:
[0049]
[0050] Therefore, the goal of the route operation time allocation agent is to find a suitable... and Make Maintain an appropriate time interval from other spatiotemporal resource occupancy. (The nominal...) Expand upwards and downwards This will allow us to obtain information about operations that affect that flight path. The parallelogram-shaped decision space. Similarly, with The time occupancy of potential conflicts can also be represented in a similar form. These time occupancy values are filled into the parallelogram decision space described above, and then proportionally transformed into a square to obtain the flight path operation. The spatiotemporal conflict diagram. Local observations by the flight path operation time allocation agent can be represented as:
[0051]
[0052] in, , and To discretize the spatiotemporal conflict map into a two-dimensional grid, the grid occupancy status of the number of unallocated plans, the grid occupancy status of the value-weighted unallocated plans, and the grid occupancy status of the number of allocated plans; To determine the occupancy characteristics of critical allocated plans in the spatiotemporal conflict map, the top l UFPs with the largest spatiotemporal resource occupancy area are selected as critical allocated plans. Compared to The spatiotemporal occupancy characteristics include the coordinates of the four points it occupies, the coordinates of the center point, the length in the x-direction, the length in the y-direction, and the area it occupies.
[0053] (3) Local observation of the intelligent agent for regional operation time allocation
[0054] The regional operation time allocation agent is activated after the altitude allocation agent selects the flight altitude; therefore, it only focuses on the potential and certain spatiotemporal resource occupancy at that altitude. For regional operation tasks... Its resource consumption is to The decision space is the entire interval of . to Expand upwards and downwards The one-dimensional space. Similarly, filling the spatiotemporal resource occupancy of other UFPs into the decision space yields the regional operation tasks. The spatiotemporal conflict diagram. Therefore, the local observation of the regional operation allocation agent can be represented as:
[0055]
[0056] in, , and To discretize the spatiotemporal conflict map into a one-dimensional grid, we need to determine the grid occupancy of the number of unallocated plans, the grid occupancy weighted by the value of the unallocated plans, and the grid occupancy of the number of allocated plans.
[0057] Step 3.4: Establish the action space, where the actions of each agent include:
[0058] (1) Highly adaptable agent action space. Using a one-dimensional discrete action space, it can be represented as:
[0059]
[0060] (2) Action space of the agent for route operation / regional time allocation. To ensure the feasibility of the output allocation plan, it is necessary to ensure that the adjustment granularity is not too fine. Therefore, a two-dimensional discrete action space is established with 15 seconds as the minimum adjustment force. To match the fixed-dimensional output requirements of the neural network, a uniform adjustment space of ±15 minutes is set. The unselectable parts can be masked during execution using a masking mechanism. The action space of the flight path operation time allocation agent can be represented as:
[0061]
[0062] Step 3.5: Establish state transition rules:
[0063] according to The allocation decisions for each UFP are made in ascending order. If the allocation agent at a high altitude chooses to "cancel", the environmental state is updated directly and the decision for the next plan is made at the next time step. If the plan is executed at a certain altitude, the corresponding time decision agent is selected according to the airspace type of the plan to allocate the time. After the allocation is completed, the environmental state is updated and the decision for the next plan is made at the next time step.
[0064] Step 3.6: Design and allocate the comprehensive reward function:
[0065] (1) Conflict penalty. During the allocation process, the flight plan currently being decided is determined according to the decision-making order. Need to avoid to The time and space resources already allocated for these completed decision-making plans must then be selected for subsequent... to The combination of start and end times with the least impact is used, so the conflict penalty consists of two parts: active conflict penalty (conflict between the current assignment and the already decided UFP) and passive conflict penalty (conflict between the current assignment and the future decision UFP).
[0066] ①Proactive conflict punishment:
[0067]
[0068] in, For the spatiotemporal conflict diagram and The actual overlapping area for The total area of the spatiotemporal conflict map, This represents the superlinear penalty coefficient.
[0069] ② Passive conflict penalty:
[0070]
[0071] (2) Proximity penalty. In addition to direct conflict, dangerous proximity with a short time interval is also penalized, which can be expressed as:
[0072]
[0073]
[0074] in, In the absence of conflict and The minimum time interval.
[0075] (3) Potential conflict avoidance reward. If, after the agent makes a decision, the UFP occupies less space and time when traversing / covering an undecided UFP compared to before the decision, a corresponding reward is applied to that action, which can be expressed as:
[0076]
[0077] in and These represent the number of grid cells occupied in the spatiotemporal space of the undecided UFP before / after allocation, and when crossing / covering it. and These are weighted values of spatiotemporal occupancy grids for undecided UFPs before / after allocation and for crossing / covering undecided UFPs. To prevent division by 0 for a decimal.
[0078] (4) Penalties for deviation from intent. During the allocation process, it is desirable to avoid conflicts while adhering as closely as possible to the flight intent stated in the plan application. Therefore, penalties are imposed for deviations from intent, including:
[0079] ① Time Intent Deviation Penalty. For regional operations, the time intent deviation penalty consists of the center time advantage of the selected time slot and the overall length deviation of the time slot; for route operations, the time intent penalty consists of the departure time deviation and the arrival time advantage, which can be expressed as:
[0080]
[0081]
[0082] in, This represents the superlinear penalty coefficient.
[0083] ② Penalty for High Deviance from Intent. For plans with high requirements, the penalty for high deviation can be expressed as:
[0084]
[0085] in, The selected height after adjustment.
[0086] (5) Resource waste penalty. Since highly selective agents employ a "cancel" strategy, a resource waste penalty is established to punish agents for avoiding resource waste by canceling large-scale manipulations. This penalty can be expressed as:
[0087]
[0088] in, After all plans are allocated Is there still room for adjustment in the allocation?
[0089] In summary, the reward function for a highly coordinated agent can be expressed as:
[0090]
[0091] The reward function for a time-managed agent can be expressed as:
[0092]
[0093] It should be noted that in the sequential decision-making process, since the UFP of the decision at step t and the UFP of the decision at step t+1 may only be close in time but far apart in spatial location, there are problems in directly estimating the state-action rewards using the traditional difference method. Therefore, the discount rate is set to 0, and the expected value of each state action is estimated by reward backfilling through passive conflict penalty, resource waste penalty and other methods.
[0094] Furthermore, step 4 includes the following steps:
[0095] Step 4.1: Design the structure of the highly adapted agent. Based on the local observations and action space of the highly adapted agent, a graph attention network based on fused edge features is used to process the graph structure data, and other feature data is processed using fully connected layers. After merging, the data is then passed through multiple fully connected layers to output discrete action probabilities.
[0096] Step 4.2: Design the agent structure for flight route operation time allocation. Based on the local observation and action space of the agent, a two-dimensional convolutional network is used to process the spatiotemporal conflict raster map. A self-attention mechanism is used to handle the occupation of key spatiotemporal resources. Fully connected layers are used to map and process other feature data. After merging, the data is passed through multiple fully connected layers to output the discrete action probability of departure time adjustment a0. After selecting the departure time adjustment action, the action is encoded through an embedding layer and concatenated to the original processed features to output the discrete action probability of arrival time adjustment a1.
[0097] Step 4.3: Design the structure of the regional task time allocation agent. Based on the local observation and action space of the regional task time allocation agent, a one-dimensional convolutional network is used to process the spatiotemporal conflict map. Fully connected layers are used to map and process other feature data. After merging, the data is passed through multiple fully connected layers to output the discrete action probability of departure time adjustment. After selecting the departure time adjustment action, the action is encoded through an embedding layer and concatenated to the original processed features to output the discrete action probability of arrival time adjustment.
[0098] Furthermore, step 5 includes the following steps:
[0099] Step 5.1: In the traditional PPO algorithm, its loss can be expressed as:
[0100]
[0101] in, This represents the probability distribution of actions under the current policy. The action probability distribution for the old strategy. For the advantage of the action, This is a truncation function. These are the hyperparameters used for PPO truncation;
[0102] Based on the partially observable Markov decision model for this problem, a combination of adjacent departure and arrival time adjustments may lead to the same conflict outcome. In a dense decision space, the rewards for each state action are not completely isolated. Therefore, a dominant neighborhood diffusion mechanism is added during training. Positive dominant diffusion is activated if both the active conflict penalty and the proximity penalty are 0; negative dominant diffusion is activated if the active conflict penalty is negative; otherwise, dominant diffusion is not performed.
[0103] Step 5.2: Construct the diffusion neighborhood distribution. The positive dominance diffusion neighborhood has a peak-shaped distribution. The negative-dominance diffusion neighborhood is a trough-shaped neighborhood distribution. , u represents the neighborhood action.
[0104] Step 5.3: Calculate the positive and negative neighborhood cross-entropy of the two-dimensional action.
[0105] ① Positive dominant neighborhood cross-entropy:
[0106]
[0107] in, Let be the probability of neighborhood action u in the policy distribution of action i;
[0108] ② Negative dominant neighborhood cross-entropy:
[0109]
[0110] Step 5.4: Calculate the neighborhood diffusion loss, which can be expressed as:
[0111]
[0112]
[0113] in, and These are separate decisions regarding whether to open the positive and negative diffusion entrances;
[0114] The final policy network loss function can be expressed as:
[0115]
[0116] in, and These are the participation coefficients for positive and negative diffusion, respectively.
[0117] Step 5.5: During the training process, since the final allocation result is affected by both the high-allocation agent and the time-allocation agent, direct training is prone to instability. Therefore, a first-come, first-served strategy is first adopted for high-allocation, and the time-allocation agent is pre-trained. Then, the high-allocation agent and the time-allocation agent are jointly trained to output the trained agent.
[0118] Step 6 includes the following steps:
[0119] Step 6.1: During the test, perform height and time allocation sequentially. When performing time allocation, select... After that, in the selected Conflicts are detected synchronously. If no conflict exists, the allocation is completed; otherwise, the allocation is blocked. Continue by selecting the highest value after shading. .
[0120] Step 6.2: If If all motion space is obscured, cancel the plan and ensure that the final output scheme has no conflicts.
[0121] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0122] This invention provides a large-scale intelligent collaborative scheduling method for low-altitude flight plans based on reinforcement learning. Addressing the needs of large-scale integrated operation in differentiated low-altitude scenarios, it achieves conflict-free intelligent scheduling of low-altitude flight plans under arbitrary scenario scales based on reinforcement learning and partially observable Markov decision processes. Compared with existing rule-based conflict scheduling methods, it achieves a higher number of conflict-free executions and stronger performance. Compared with optimization-based conflict scheduling methods, it has a faster solution speed and larger scale, possessing 10... 6 With a solution capacity of / hour, it can meet the needs of urban-level low-altitude flight plan management and control, has strong engineering application potential, and can effectively improve the service and operation support capabilities of low-altitude airspace. Attached Figure Description
[0123] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein:
[0124] Figure 1 This is a flowchart illustrating a large-scale low-altitude flight planning intelligent collaborative scheduling method based on reinforcement learning, according to the present invention.
[0125] Figure 2 This is a schematic diagram of the horizontal interval discrimination of the present invention;
[0126] Figure 3 This is a schematic diagram of the time and space occupation of the flight route operation according to the present invention;
[0127] Figure 4 This is a schematic diagram of the spatial and temporal occupancy of the operation area in this invention;
[0128] Figure 5 This is a schematic diagram of the highly coordinated intelligent agent structure of the present invention;
[0129] Figure 6 This is a schematic diagram of the intelligent agent structure for route operation time allocation in this invention;
[0130] Figure 7 This is a schematic diagram of the intelligent agent structure for regional operation time allocation in this invention;
[0131] Figure 8 This is an agent training curve provided in one embodiment of the present invention; (a) is the pre-training reward, and (b) is the joint training reward.
[0132] Figure 9 This is a diagram showing the effect of flight plan conflict resolution according to an embodiment of the present invention; (a) is before conflict resolution, and (b) is after conflict resolution. Detailed Implementation
[0133] To make the objectives, uses, and technical solutions of this invention easier to understand, and to provide a clearer and more complete description of this invention, the following detailed description is provided in conjunction with embodiments. Obviously, the described embodiments are only some embodiments of this invention, and it should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the scope of protection of this invention.
[0134] The application principle of the present invention will be described in detail below with reference to the accompanying drawings.
[0135] like Figure 1 As shown in Example 1: This example provides a method for intelligent collaborative scheduling of large-scale low-altitude flight plans based on reinforcement learning, including:
[0136] Step 1. Define the flight plan allocation constraints, including the adjustment range of start time, end time, and operation duration for each flight plan.
[0137] Step 1.1: Establish a low-altitude flight plan model for a given low-altitude flight plan. It can be represented as:
[0138]
[0139] in, and They are respectively The type of plan and the nature of the task. Airspace usage type (line operations and area operations). For the coordinates of the airspace point, For use of height, and For departure and arrival times, For speed, Operating modes (one-way and round-trip). The status is planned (not allocated, allowed to execute, canceled). To determine whether there are flight altitude requirements, and For the permitted maximum and minimum airspace altitudes, The maximum allowable time adjustment amount, For the maximum permissible flight time, The maximum permissible flight speed.
[0140] Step 1.2: Calculate the allocation constraints, specifically including:
[0141] (1) Start time adjustment range constraints:
[0142]
[0143] in, This refers to the rescheduled departure time.
[0144] (2) End time adjustment range constraints:
[0145]
[0146] in, This refers to the arrival time after the relocation.
[0147] (3) Constraints on the range of adjustment of route operation time:
[0148]
[0149]
[0150] in, The flight distance for the route.
[0151] (4) Constraints on the range of regional operation time adjustment:
[0152]
[0153] Step 2. Based on the airspace usage patterns of route operations and area operations, establish a flight plan conflict detection model that integrates altitude conflict detection, horizontal conflict detection, and temporal conflict detection.
[0154] Step 2.1: Execution status check. If and If any plan is in the "cancelled" state, it is considered that there is no conflict; otherwise, proceed to step 2.2.
[0155] Step 2.2: High Conflict Judgment. When and If an intersection exists, proceed to step 2.3; otherwise, it is assumed there is no conflict. For If its status is "unallocated", then and Take respectively and If its status is "Execution Allowed", then and All .
[0156] Step 2.3: Horizontal Conflict Judgment. When and The minimum distance between projections on the same plane is less than the horizontal safety interval. If the condition is met, proceed to step 2.4; otherwise, it is considered that there is no conflict. Cases where the horizontal spacing requirement is not met include... Figure 2 .
[0157] Step 2.4: Time Conflict Judgment. For and Flight segments / areas with a horizontal safety interval smaller than the standard distance. Entry and exit times are and ,like For line operation, then and Based on the location and departure time of their conflicting flight segments Arrival time If it is a regional operation, then the departure time should be extracted separately. and arrival time .when and If there is an intersection, it is considered a conflict; otherwise, it is considered a non-conflict. Wherein... If the status is "unallocated", then ,like If the status is "Execution Allowed", then ; For safety time intervals.
[0158] Step 2.5: Conflict Type Differentiation. If and The status is "Execution Allowed". If there is a conflict between the two, it is "Confirmed Conflict"; otherwise, it is "Potential Conflict".
[0159] Step 3. Establish a hierarchical decision-making model of "height and execution decision - start and end time decision", design the observation and action space required for each level of decision, define the training rewards for each agent, and establish a reinforcement learning model based on a partially observable Markov decision process.
[0160] Step 3.1: Establish a hierarchical decision-making model of "altitude and execution decision-start and end time decision-making": First, the altitude allocation agent decides whether to allow the plan to be executed based on the determination of multiple altitude layers and potential spatiotemporal occupancy information. If execution is allowed, the altitude layer to be used for the plan is selected and handed over to the time adjustment agent for further decision-making. According to different airspace usage types, the corresponding time adjustment agent makes decisions on the start and end times of the plan to complete the allocation of the plan.
[0161] Step 3.2: Establish an environmental state model. The environmental state can be represented as follows: ,in For flight plan The state at time step t. The agent makes decisions in ascending order of the expected departure times of each flight plan, deciding on one flight plan and updating the environmental state at each time step.
[0162] Step 3.3: Establish a local observation model. The local observations input to different agents include:
[0163] (1) Highly coordinated local observation of intelligent agents
[0164] Because there are complex potential conflicts between the plans, and the number of potential conflicting objects varies among different plans, this conflict relationship is represented by a graph structure, and key features are pieced together as local observations of a highly coordinated agent, corresponding to time step t. Local observations during allocation are represented as follows:
[0165]
[0166] in, For node information, for As feature information of nodes To and Plans with potential conflict relationships are used as feature information for nodes. To and A set of plans with potentially conflicting relationships; for Edge characteristics between plans that have potentially conflicting relationships; For the determined spatiotemporal resource occupancy at each altitude, for The proportion of spatiotemporal resources that are confirmed to be occupied and unusable at height layer k. The highest layer in the environment is 5, which is taken as 5 in this embodiment; This is altitude preference information, when there is a flight altitude requirement. The one-hot encoding is used; otherwise, it is an all-zero matrix. for The planned quantity was not allocated.
[0167]
[0168] in, for The execution value is determined by its program type, mission type, and flight service distance / area.
[0169]
[0170] in, and for In, it and The conflict segments are arranged according to their starting and ending proportions within the overall route (if...). For regional operation tasks =0、 =1); for and The overall angle of the conflict segment (if at least one area operation task exists within it) =180°); to for and The conflict is highly repetitive and occupies the one-heat variable. If it is "not allocated", then it is and The intersection height, if If "Allow execution" is enabled, then... The unique heat variable.
[0171] (2) Local observation of the intelligent agent for route operation time allocation
[0172] The flight path operation time allocation agent is activated after the altitude allocation agent selects the flight altitude; therefore, it only focuses on the potential and certain spatiotemporal resource occupancy at that altitude. For flight path operation tasks... The proportion of flights to the designated routes The spatiotemporal resource occupancy requirement can be expressed as:
[0173]
[0174] Therefore, the goal of the route operation time allocation agent is to find a suitable... and Make Maintain an appropriate time interval from other spatiotemporal resource occupancy. (The nominal...) Expand upwards and downwards This will allow us to obtain information about operations that affect that flight path. The parallelogram-shaped decision space. Similarly, with The time occupancy of potential conflicts can also be represented in a similar form. These time occupancy values are filled into the parallelogram decision space described above, and then proportionally transformed into a square to obtain the flight path operation. Spatiotemporal conflict diagram, such as Figure 3 As shown. The local observation of the flight path operation time allocation agent can be represented as:
[0175]
[0176] in, , and To discretize the spatiotemporal conflict map into a two-dimensional grid, the grid occupancy status of the number of unallocated plans, the grid occupancy status of the value-weighted unallocated plans, and the grid occupancy status of the number of allocated plans; To determine the occupancy characteristics of critical allocated plans in the spatiotemporal conflict map, the top l UFPs with the largest spatiotemporal resource occupancy area are selected as critical allocated plans. Compared to The spatiotemporal occupancy characteristics include the coordinates of the four points it occupies, the coordinates of the center point, the length in the x-direction, the length in the y-direction, and the area it occupies.
[0177] (3) Local observation of the intelligent agent for regional operation time allocation
[0178] The regional operation time allocation agent is activated after the altitude allocation agent selects the flight altitude; therefore, it only focuses on the potential and certain spatiotemporal resource occupancy at that altitude. For regional operation tasks... Its resource consumption is to The decision space is the entire interval of . to Expand upwards and downwards The one-dimensional space. Similarly, filling the spatiotemporal resource occupancy of other UFPs into the decision space yields the regional operation tasks. Spatiotemporal conflict diagram, such as Figure 4 As shown. Therefore, the local observation of the regional operation dispatching agent can be represented as:
[0179]
[0180] in, , and To discretize the spatiotemporal conflict map into a one-dimensional grid, we need to determine the grid occupancy of the number of unallocated plans, the grid occupancy weighted by the value of the unallocated plans, and the grid occupancy of the number of allocated plans.
[0181] Step 3.4: Establish the action space, where the actions of each agent include:
[0182] (1) Highly adaptable agent action space. Using a one-dimensional discrete action space, it can be represented as:
[0183]
[0184] (2) Action space of the agent for route operation / regional time allocation. To ensure the feasibility of the output allocation plan, it is necessary to ensure that the adjustment granularity is not too fine. Therefore, a two-dimensional discrete action space is established with 15 seconds as the minimum adjustment force. To match the fixed-dimensional output requirements of the neural network, a uniform adjustment space of ±15 minutes is set. The unselectable parts can be masked during execution using a masking mechanism. The action space of the flight path operation time allocation agent can be represented as:
[0185]
[0186] Step 3.5: Establish state transition rules:
[0187] according to The allocation decisions for each UFP are made in ascending order. If the allocation agent at a high altitude chooses to "cancel", the environmental state is updated directly and the decision for the next plan is made at the next time step. If the plan is executed at a certain altitude, the corresponding time decision agent is selected according to the airspace type of the plan to allocate the time. After the allocation is completed, the environmental state is updated and the decision for the next plan is made at the next time step.
[0188] Step 3.6: Design and allocate the comprehensive reward function:
[0189] (1) Conflict penalty. During the allocation process, the flight plan currently being decided is determined according to the decision-making order. Need to avoid to The time and space resources already allocated for these completed decision-making plans must then be selected for subsequent... to The combination of start and end times with the least impact is used, so the conflict penalty consists of two parts: active conflict penalty (conflict between the current assignment and the already decided UFP) and passive conflict penalty (conflict between the current assignment and the future decision UFP).
[0190] ①Proactive conflict punishment:
[0191]
[0192] in, For the spatiotemporal conflict diagram and The actual overlapping area for The total area of the spatiotemporal conflict map, In this embodiment, the superlinear penalty coefficient is set to 0.3.
[0193] ② Passive conflict penalty:
[0194]
[0195] (2) Proximity penalty. In addition to direct conflict, dangerous proximity with a short time interval is also penalized, which can be expressed as:
[0196]
[0197]
[0198] in, In the absence of conflict and The minimum time interval.
[0199] (3) Potential conflict avoidance reward. If, after the agent makes a decision, the UFP occupies less space and time when traversing / covering an undecided UFP compared to before the decision, a corresponding reward is applied to that action, which can be expressed as:
[0200]
[0201] in and These represent the number of grid cells occupied in the spatiotemporal space of the undecided UFP before / after allocation, and when crossing / covering it. and These are weighted values of spatiotemporal occupancy grids for undecided UFPs before / after allocation and for crossing / covering undecided UFPs. To prevent division by zero, the value is set to 10 in this embodiment. -6 .
[0202] (4) Penalties for deviation from intent. During the allocation process, it is desirable to avoid conflicts while adhering as closely as possible to the flight intent stated in the plan application. Therefore, penalties are imposed for deviations from intent, including:
[0203] ① Time Intent Deviation Penalty. For regional operations, the time intent deviation penalty consists of the center time advantage of the selected time slot and the overall length deviation of the time slot; for route operations, the time intent penalty consists of the departure time deviation and the arrival time advantage, which can be expressed as:
[0204]
[0205]
[0206] in, The superlinear penalty coefficient is set to 0.6 in this embodiment.
[0207] ② Penalty for High Deviance from Intent. For plans with high requirements, the penalty for high deviation can be expressed as:
[0208]
[0209] in, The selected height after adjustment.
[0210] (5) Resource waste penalty. Since highly selective agents employ a "cancel" strategy, a resource waste penalty is established to punish agents for avoiding resource waste by canceling large-scale manipulations. This penalty can be expressed as:
[0211]
[0212] in, After all plans are allocated Is there still room for adjustment in the allocation?
[0213] In summary, the reward function for a highly coordinated agent can be expressed as:
[0214]
[0215] The reward function for a time-managed agent can be expressed as:
[0216]
[0217] It should be noted that in the sequential decision-making process, since the UFP of the decision at step t and the UFP of the decision at step t+1 may only be close in time but far apart in spatial location, there are problems in directly estimating the state-action rewards using the traditional difference method. Therefore, the discount rate is set to 0, and the expected value of each state action is estimated by reward backfilling through passive conflict penalty, resource waste penalty and other methods.
[0218] Step 4. Develop decision-making functions for the corresponding altitude allocation agent, route operation time allocation agent, and area operation time allocation agent, and design the agent structure based on graph neural networks and self-attention mechanisms.
[0219] Step 4.1: Design the highly adapted agent structure. Based on the local observations and action space of the highly adapted agent, a graph attention network based on fused edge features is used to process the graph structure data. Other feature data is then mapped and processed using fully connected layers. After merging, the data is passed through multiple fully connected layers to output discrete action probabilities, such as... Figure 5 As shown.
[0220] Step 4.2: Design the agent structure for flight route operation time allocation. Based on the local observation and action space of the agent, a two-dimensional convolutional network is used to process the spatiotemporal conflict raster map. A self-attention mechanism is used to handle the occupation of key spatiotemporal resources. Fully connected layers are used to map and process other feature data. After merging, the data is passed through multiple fully connected layers to output the discrete action probability of departure time adjustment. After selecting the departure time adjustment action, the action is encoded through an embedding layer and concatenated to the original processed features to output the discrete action probability of arrival time adjustment. Figure 6 As shown.
[0221] Step 4.3: Design the structure of the regional task time allocation agent. Based on the local observations and action space of the regional task time allocation agent, a one-dimensional convolutional network is used to process the spatiotemporal conflict map. Fully connected layers are used to map and process other feature data. After merging, the data is passed through multiple fully connected layers to output the discrete action probability for departure time adjustment. After selecting the departure time adjustment action, this action is encoded through an embedding layer and concatenated to the original processed features to output the discrete action probability for arrival time adjustment, such as... Figure 7 As shown.
[0222] Step 5. Based on the characteristics of dense action space and non-independent discrete adjacent actions in the allocation model, the PPO (Proximal Policy Optimization) algorithm is modified by adding an adjacent action advantage diffusion mechanism, and a more efficient agent training is achieved by combining independent training and joint training.
[0223] Step 5.1: In the traditional PPO algorithm, its loss can be expressed as:
[0224]
[0225] in, This represents the probability distribution of actions under the current policy. The action probability distribution for the old strategy. For the advantage of the action, This is a truncation function. In this embodiment, the hyperparameter used for PPO truncation is set to 0.2;
[0226] Based on the partially observable Markov decision model for this problem, a combination of adjacent departure and arrival time adjustments may lead to the same conflict outcome. In a dense decision space, the rewards for each state action are not completely isolated. Therefore, a dominant neighborhood diffusion mechanism is added during training. Positive dominant diffusion is activated if both the active conflict penalty and the proximity penalty are 0; negative dominant diffusion is activated if the active conflict penalty is negative; otherwise, dominant diffusion is not performed.
[0227] Step 5.2: Construct the diffusion neighborhood distribution. The positive dominance diffusion neighborhood has a peak-shaped distribution. The negative-dominance diffusion neighborhood is a trough-shaped neighborhood distribution. u represents the neighborhood action, and the neighborhood radius is set to 2 in this embodiment. , .
[0228] Step 5.3: Calculate the positive and negative neighborhood cross-entropy of the two-dimensional action.
[0229] ① Positive dominant neighborhood cross-entropy:
[0230]
[0231] in, Let be the probability of neighborhood action u in the policy distribution of action i;
[0232] ② Negative dominant neighborhood cross-entropy:
[0233]
[0234] Step 5.4: Calculate the neighborhood diffusion loss, which can be expressed as:
[0235]
[0236]
[0237] in, and These are separate decisions regarding whether to open the positive and negative diffusion entrances;
[0238] The final policy network loss function can be expressed as:
[0239]
[0240] in, and These are the participation coefficients for positive and negative diffusion, respectively, both set to 0.1 in this embodiment;
[0241] Step 5.5: During training, since the final allocation result is influenced by both the high-level allocation agent and the time-based allocation agent, direct training is prone to instability. Therefore, a first-come, first-served strategy is first used for high-level allocation, and the time-based allocation agent is pre-trained. Then, the high-level allocation agent and the time-based allocation agent are jointly trained to output the trained agent. The agent training curve is shown below. Figure 8 As shown, compared to other methods, the advantage diffusion PPO converged to a higher reward value in this problem and achieved a better allocation effect.
[0242] Step 6. Utilize the trained agent and combine it with a deterministic spatiotemporal occupancy masking mechanism to achieve intelligent collaborative scheduling of conflict-free low-altitude flight plans of any scale.
[0243] Step 6.1: During the test, perform height and time allocation sequentially. When performing time allocation, select... After that, in the selected Conflicts are detected synchronously. If no conflict exists, the allocation is completed; otherwise, the allocation is blocked. Continue by selecting the highest value after shading. .
[0244] Step 6.2: If If all motion space is obscured, cancel the plan. Ensure the final output scheme has no conflicts. Figure 9 As shown.
Claims
1. A method for intelligent collaborative scheduling of large-scale low-altitude flight plans based on reinforcement learning, characterized in that, Includes the following steps: Step 1. Based on the mission type, start time, end time, acceptable adjustment time, maximum flight speed, and aircraft endurance of the low-altitude flight plan, define the flight plan allocation constraints, including the adjustment range of start time, end time, and operation duration for each flight plan. Step 2. Combine the airspace usage patterns of route operations and area operations to establish a flight plan conflict detection model that integrates altitude conflict detection, horizontal conflict detection, and temporal conflict detection; Step 3. Establish a hierarchical decision-making model of "height and execution decision-start and end time decision", design the observation and action space required for each level of decision, define the training reward for each agent, and establish a reinforcement learning model based on a partially observable Markov decision process. Step 4. Develop decision-making functions for the corresponding altitude allocation agent, route operation time allocation agent, and area operation time allocation agent, and design the agent structure based on graph neural networks and self-attention mechanisms; Step 5. Based on the characteristics of dense action space and non-independent discrete adjacent actions in the allocation model, the PPO algorithm is modified by adding an adjacent action advantage diffusion mechanism, and a more efficient agent training is achieved by combining independent training and joint training. Step 6. Utilize the trained agent and combine it with a deterministic spatiotemporal occupancy masking mechanism to achieve intelligent collaborative scheduling of conflict-free low-altitude flight plans of any scale.
2. The method of claim 1, wherein, The flight plan allocation constraint construction in step 1 includes: Step 1.1: Establishing a low-altitude flight plan model, a certain low-altitude flight plan is represented as: in, and They are respectively The type of plan and the nature of the task. For airspace usage type, For the coordinates of the airspace point, For use of height, and For departure and arrival times, For speed, For operation mode, In the planned state, To determine whether there are flight altitude requirements, and For the permitted maximum and minimum airspace altitudes, The maximum allowable time adjustment amount, For the maximum permissible flight time, The maximum permissible flight speed; Step 1.2: Calculate the allocation constraints, specifically including: (1) Start time adjustment range constraints: in, The adjusted departure time; (2) End time adjustment range constraints: in, The arrival time after rescheduling; (3) Constraints on the range of adjustment of route operation time: in, For the flight distance of the route; (4) Constraints on the range of regional operation time adjustment: 。 3. The intelligent collaborative scheduling method for large-scale low-altitude flight planning based on reinforcement learning according to claim 2, characterized in that, The flight plan conflict detection model in step 2 is for flight plans. and Its conflict resolution process includes: Step 2.1: Execute status check; if and If any plan is in the "cancelled" state, it is considered to be without conflict; otherwise, proceed to step 2.
2. Step 2.2: High Conflict Judgment; when and If an intersection exists, proceed to step 2.3; otherwise, it is assumed that there is no conflict. If its status is "unallocated", then and Take respectively and If its status is "Execution Allowed", then and All ; Step 2.3: Horizontal conflict judgment; when and The minimum distance between projections on the same plane is less than the horizontal safety interval. If the condition is met, proceed to step 2.4; otherwise, it is assumed that there is no conflict. Step 2.4: Time Conflict Judgment; For and Flight segments / areas with a horizontal safety interval smaller than the standard distance. Entry and exit times are and ,like For line operation, then and Based on the location and departure time of their conflicting flight segments Arrival time If it is a regional operation, then the departure time should be extracted separately. and arrival time ;when and If there is an intersection, it is considered a conflict; otherwise, it is considered a non-conflict. If the status is "unallocated", then ,like If the status is "Execution Allowed", then ; For safe time intervals; Step 2.5: Distinguish between conflict types; if and The status is "Execution Allowed". If there is a conflict between the two, it is "Confirmed Conflict"; otherwise, it is "Potential Conflict".
4. The intelligent collaborative scheduling method for large-scale low-altitude flight planning based on reinforcement learning according to claim 1, characterized in that, The construction process of the reinforcement learning model based on partially observable Markov decision processes in step 3 includes: Step 3.1: Establish a hierarchical decision-making model of "altitude and execution decision - start and end time decision": First, the altitude allocation agent decides whether to allow the plan to be executed based on the determination of multiple altitude layers and potential spatiotemporal occupancy information. If execution is allowed, the altitude layer to be used for the plan is selected and handed over to the time adjustment agent for further decision-making. According to different airspace usage types, the corresponding time adjustment agent makes decisions on the start and end times of the plan to complete the allocation of the plan. Step 3.2: Establish an environmental state model, where the environmental state is represented as follows: ,in For flight plan The state at time step t; the agent makes decisions in ascending order of the expected departure time of each flight plan, making a decision on one flight plan and updating the environmental state at each time step; Step 3.3: Establish a local observation model. The local observations input to different agents include: (1) Highly coordinate local observations of the intelligent agent; at time step t, corresponding to Local observations during allocation are represented as follows: in, For node information, for As feature information of nodes To and Plans with potential conflict relationships are used as feature information for nodes. To and A set of plans with potentially conflicting relationships; for Edge characteristics between plans that have potentially conflicting relationships; For the determined spatiotemporal resource occupancy at each altitude, for The proportion of spatiotemporal resources that are confirmed to be occupied and unusable at height layer k. This is the highest level within the environment; This is altitude preference information, when there is a flight altitude requirement. The one-hot encoding is used; otherwise, it is an all-zero matrix. for The planned quantity was not allocated in the middle; in, for The execution value is determined by its program type, mission type, and flight service distance / area; in, and for In, it and The conflict segment is determined by the proportion of its start and end points within the overall route, if... For regional operation tasks =0、 =1; for and The overall angle of the conflict segment, if there is at least one area operation mission within it. =180°; to for and The conflict is highly repetitive and occupies the one-heat variable. If it is "not allocated", then it is and The intersection height, if If "Allow execution" is enabled, then... The unique heat variable; (2) Local observation of the intelligent agent for route operation time allocation; for route operation tasks The proportion of flights to the designated routes The spatiotemporal resource occupancy requirement is expressed as: The nominal Expand upwards and downwards This will allow us to obtain information about operations that affect that flight path. The parallelogram decision space; and The time occupancy of potential conflicts can also be represented in a similar form. These time occupancy values are filled into the parallelogram decision space described above, and then proportionally transformed into a square to obtain the flight path operation. The spatiotemporal conflict diagram; the local observations of the flight path operation time allocation agent are represented as: in, , and To discretize the spatiotemporal conflict map into a two-dimensional grid, the grid occupancy status of the number of unallocated plans, the grid occupancy status of the value-weighted unallocated plans, and the grid occupancy status of the number of allocated plans; To determine the occupancy characteristics of critical allocated plans in the spatiotemporal conflict map, the top l UFPs with the largest spatiotemporal resource occupancy area are selected as critical allocated plans. Compared to The spatiotemporal occupancy characteristics include the coordinates of the four points it occupies, the coordinates of the center point, the length in the x-direction, the length in the y-direction, and the area it occupies; (3) Local observation of the regional operation time allocation agent; for regional operation tasks Its resource consumption is to The decision space is the entire interval of . to Expand upwards and downwards The one-dimensional space; filling the spatiotemporal resource occupancy of other UFPs into the decision space to obtain the regional operation task. Spatiotemporal conflict diagram; local observations of the regional operation allocation agent can be represented as: in, , and To discretize the spatiotemporal conflict map into a one-dimensional grid, the grid occupancy status of the number of unallocated plans, the grid occupancy status of the value-weighted unallocated plans, and the grid occupancy status of the number of allocated plans; Step 3.4: Establish the action space, where the actions of each agent include: (1) Highly allocate the agent's action space; adopt a one-dimensional discrete action space, represented as: (2) Action space of the intelligent agent for route operation / area time allocation; a two-dimensional discrete action space is established with 15 seconds as the minimum adjustment force, and a uniform adjustment space is set at ±15 minutes. The unselectable parts can be masked during execution using a masking mechanism; the action space of the flight path operation time allocation agent is represented as: Step 3.5: Establish state transition rules: according to The allocation decisions for each UFP are made in ascending order. If the allocation agent selects the "cancel" action, the environmental state is updated directly and the decision on the next plan is made in the next time step. If the plan is executed at a certain altitude, the corresponding time decision agent is selected according to the airspace type of the plan to allocate the time. After the allocation is completed, the environmental state is updated and the decision on the next plan is made in the next time step. Step 3.6: Design and allocate the comprehensive reward function: (1) Conflict penalty; The conflict penalty consists of two parts, including active conflict penalty and passive conflict penalty. Active conflict penalty is the conflict between the current allocation and the already decided UFP, and passive conflict penalty is the conflict between the current allocation and the future decision UFP. ①Proactive conflict punishment: in, For the spatiotemporal conflict diagram and The actual overlapping area for The total area of the spatiotemporal conflict map, The superlinear penalty coefficient; ② Passive conflict penalty: (2) Proximity penalty; In addition to direct conflict, dangerous proximity with a short time interval is also penalized, as shown below: in, In the absence of conflict and The minimum time interval; (3) Potential conflict avoidance reward; if the UFP occupies less space-time when crossing / covering the undecided UFP after the agent makes a decision compared to before the allocation, a corresponding reward is applied to the action, which can be expressed as: in and These represent the number of grid cells occupied in the spatiotemporal space of the undecided UFP before / after allocation, and when crossing / covering it. and These are weighted values of spatiotemporal occupancy grids for undecided UFPs before / after allocation and for crossing / covering undecided UFPs. To prevent division by 0 for a decimal; (4) Penalties for deviation from intent; During the allocation process, it is desirable to avoid conflicts while adhering as much as possible to the flight intent stated in the plan application. Therefore, penalties are imposed for deviations from intent, including: ① Penalty for deviation from intended time: in, The superlinear penalty coefficient; ② Punishment for highly intentional deviation: in, The selected height after adjustment; (5) Resource waste penalty; To penalize agents for avoiding resource waste by massively canceling manipulation, a resource waste penalty is established, represented as: in, After all plans are allocated Is there still room for adjustment within the system? In summary, the reward function for a highly coordinated agent is expressed as: The reward function for the time-managed agent is expressed as: The discount rate is set to 0, and the expected value of each state action is estimated by rewarding passive conflict penalties and resource waste penalties.
5. The intelligent collaborative scheduling method for large-scale low-altitude flight planning based on reinforcement learning according to claim 1, characterized in that, The agent structure in step 4: Step 4.1: Design a highly tuned agent structure; process graph structure data based on a graph attention network that integrates edge features, use fully connected layers to map and process other feature data, merge them, and then output discrete action probabilities through multiple fully connected layers; Step 4.2: Design the intelligent agent structure for flight route operation time allocation; process the spatiotemporal conflict grid map based on a two-dimensional convolutional network, use a self-attention mechanism to handle the occupation of key spatiotemporal resources, use fully connected layers to map and process other feature data, merge them, and then output the discrete action probability of departure time adjustment a0 after passing through multiple fully connected layers; after selecting the departure time adjustment action, encode the action through an embedding layer and concatenate it to the original processed features to output the discrete action probability of arrival time adjustment a1. Step 4.3: Design the intelligent agent structure for regional operation time allocation; process the spatiotemporal conflict map based on a one-dimensional convolutional network, use fully connected layers to map and process other feature data, merge them, and then output the discrete action probability of departure time adjustment a0 after passing through multiple fully connected layers; after selecting the departure time adjustment action, encode the action through the embedding layer and concatenate it to the original processed features to output the discrete action probability of arrival time adjustment a1.
6. The intelligent collaborative scheduling method for large-scale low-altitude flight planning based on reinforcement learning according to claim 1, characterized in that, Step 5 includes the following steps: Step 5.1: Determine the diffusion type; in the traditional PPO algorithm, the loss is expressed as: in, This represents the probability distribution of actions under the current policy. The action probability distribution for the old strategy. For the advantage of the action, This is a truncation function. These are the hyperparameters used for PPO truncation; In conjunction with the partially observable Markov decision model for this problem, a dominant neighborhood diffusion mechanism was added during training; if both the active conflict penalty and the proximity penalty are 0, positive dominant diffusion is activated; if the active conflict penalty is negative, negative dominant diffusion is activated; otherwise, dominant diffusion is not performed. Step 5.2: Construct the diffusion neighborhood distribution; the positive dominance diffusion neighborhood is a peak-shaped neighborhood distribution. The negative-dominance diffusion neighborhood is a trough-shaped neighborhood distribution. u represents the neighborhood action; Step 5.3: Calculate the positive and negative advantage neighborhood cross-entropy for the two-dimensional action; ① Positive dominant neighborhood cross-entropy: in, Let be the probability of neighborhood action u in the policy distribution of action i; ② Negative dominant neighborhood cross-entropy: Step 5.4: Calculate the neighborhood diffusion loss, expressed as: in, and These are separate decisions regarding whether to open the positive and negative diffusion entrances; The final policy network loss function is expressed as: in, and These are the participation coefficients for positive and negative diffusion, respectively. Step 5.5: Use a first-come, first-served strategy for high-allocation, pre-train the time-allocation agent, then jointly train the high-allocation agent and the time-allocation agent, and output the trained agent.
7. The intelligent collaborative scheduling method for large-scale low-altitude flight planning based on reinforcement learning according to claim 1, characterized in that, The spatiotemporal occupancy occupancy mechanism in step 6; during the test, height and time are adjusted sequentially; when adjusting the time, select... After that, in the selected Conflicts are detected synchronously. If no conflict exists, the allocation is completed; otherwise, the allocation is blocked. Continue by selecting the highest value after shading. ;like If all motion space is obscured, cancel the plan and ensure that the final output scheme has no conflicts.