City low-altitude multi-aircraft four-dimensional flight plan collaborative optimization method based on preference perception
By improving the MADDPG algorithm and AMMP mechanism, non-compliant adjustment actions are filtered out using task preference attributes, and the HPER mechanism is combined to optimize the generation of urban low-altitude multi-aircraft four-dimensional flight plans. This solves the problems of dynamic task requirements and heterogeneous task preferences in urban low-altitude airspace management, and realizes rapid and compliant flight plan generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies are unable to quickly adapt to the dynamic and diverse task requirements in urban low-altitude airspace management, and fail to effectively meet the preference constraints of heterogeneous tasks, resulting in high computing costs, poor flexibility, and non-compliant operation.
By employing the improved MADDPG algorithm and AMMP mechanism, each UAV is modeled as an intelligent agent. By discretizing urban airspace, identifying spatiotemporal conflicts, and using task preference attributes to filter non-compliant adjustment actions, the multi-UAV four-dimensional flight plan is optimized and generated in conjunction with the HPER mechanism.
It significantly shortens the flight plan submission time, improves the flexibility and response speed of the UAM system, ensures that the generated plans meet the specific adjustment constraints of heterogeneous tasks, achieves on-demand compliance, and improves the learning efficiency and convergence speed of the algorithm.
Smart Images

Figure CN121789516A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of urban air traffic management technology, specifically relating to a preference-aware collaborative optimization method for four-dimensional flight plans of multiple aircraft in urban low-altitude airspace. Background Technology
[0002] With the evolution of intelligent transportation systems, Urban Air Mobility (UAM) has become an innovative solution reshaping the future urban transportation landscape. UAM systems utilize advanced vehicles such as drones to provide diverse on-demand and pre-booked air transport services, including logistics delivery, emergency response, and security patrols. Future urban low-altitude airspace operations will exhibit complex characteristics such as large scale, high density, heterogeneous missions, and dynamically changing operating environments. Against this backdrop, coordinating diverse operational needs with limited airspace resources to ensure the safe, efficient, and collaborative management of urban low-altitude airspace has become a core challenge for the industry.
[0003] In traditional civil aviation air traffic management, air traffic controllers typically employ centralized control to strategically optimize aircraft 4D flight plans to achieve effective conflict management. This offers some reference for urban low-altitude airspace management, where strategic-level global collaborative management of 4DT flight plans can optimize the allocation of limited spatial and temporal resources. However, unlike civil aviation, future UAM operations must adapt to dynamic flight requests, heterogeneous task priorities, and frequent plan revisions. In other words, even well-planned global flight plans may face newly submitted time-sensitive tasks, necessitating rapid reconfiguration of entirely new global plans. Therefore, urban low-altitude flight requires strategic-level collaborative management of flight plans from the management's perspective. The goal is to deduce and predict flight conflicts between multiple UAVs from a global perspective before takeoff, and to generate globally optimized, conflict-free, and preference-compliant multi-aircraft 4D flight plans before flight mission execution.
[0004] Existing research typically models the above problem as a large-scale combinatorial optimization problem, using traditional or improved heuristic optimization algorithms to solve it. A generally accepted approach is to generate conflict-free plans through a combination of three strategies: adjusting takeoff time, adjusting flight speed, and adjusting local paths. However, existing models and methods still face key challenges in practical applications in complex, dynamic, and heterogeneous mission scenarios: 1) Heuristic optimization methods are essentially "one-time solutions" for the global flight plan. When mission requirements change dynamically, this method needs to iterate and solve for a completely new global plan online, resulting in relatively poor scalability and extremely high computational costs, making it difficult to meet the needs of rapid adjustments. 2) Existing research focuses on planning operational risks and time costs when generating conflict-free plans, often neglecting the heterogeneous mission preferences that will inevitably exist in future urban low-altitude operations. Although some studies have attempted to introduce mission priorities or specific mission constraints, these are usually considered "soft weights" in the optimization objective, rather than "hard constraints" in actual operation.
[0005] Therefore, there is an urgent need for a preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight planning to solve the problems existing in the current technology. Summary of the Invention
[0006] In view of this, the present invention provides a preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight plans, which is used to solve the above-mentioned problems existing in the prior art.
[0007] To achieve the above objectives, this invention provides a preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight planning, comprising the following steps: The urban airspace is discretized to obtain several three-dimensional airspace grids, and the multidimensional attribute vector of each three-dimensional airspace grid is determined. The initial flight plan information of several UAVs is obtained and modeled to obtain an initial flight plan set; The initial flight plan sets all UAV flight plan tasks are mapped to urban airspace, and spatiotemporal conflicts are identified to obtain conflict data; the conflict data includes: the three-dimensional airspace grid of the conflict point, time, conflicting UAV, and the UAV's flight plan; Based on the improved MADDPG algorithm, each UAV is treated as an intelligent agent, and the flight plan set is iteratively optimized to obtain the final executable multi-UAV four-dimensional flight plan set. The improved MADDPG algorithm includes an AMMP mechanism, which is used to filter non-compliant adjustment action components of the intelligent agent according to the task preference attributes.
[0008] As an embodiment of the present invention, the urban airspace is discretized to obtain several three-dimensional airspace grids, and the multi-dimensional attribute vector of each three-dimensional airspace grid is determined, including: A Cartesian coordinate system is constructed within the urban airspace, and the continuous three-dimensional urban operational airspace is discretized into several three-dimensional airspace grids along the three coordinate axes of the Cartesian coordinate system, as shown below: In the formula, Indicates urban airspace, Represents a three-dimensional spatial raster. Represents three-dimensional spatial grid coordinates. This indicates the extent of urban airspace on the X-axis. This indicates the extent of urban airspace on the Y-axis. This indicates the extent of urban airspace on the Z-axis. This indicates the floor function. This represents the size of a three-dimensional spatial raster. The multidimensional attribute vector for each 3D spatial raster is determined as follows: In the formula, Representing a three-dimensional spatial raster A multidimensional attribute vector, Representing a three-dimensional spatial raster The reachability of the airspace, Representing a three-dimensional spatial raster The attributes of infrastructure service capabilities Representing a three-dimensional spatial raster Operational risk attributes; in, As shown below: In the formula, Representing a three-dimensional spatial raster Horizontal area Representing a three-dimensional spatial raster The reachable airspace area; in, As shown below: In the formula, This represents the initial ground risk value. This indicates the reduction or mitigation of ground risks.
[0009] As one embodiment of the present invention, initial flight plan information of several UAVs is obtained and modeled to obtain an initial flight plan set, including: Obtain initial flight plan information for several drones; Based on the initial flight plan information of several UAVs, a model is created to obtain the initial flight plan set, as shown below: In the formula, Represents the initial flight plan set. Indicates the first The flight plan mission of the drone This indicates the total number of drones. Indicates the first The mission number of the drone. Indicates the first The four-dimensional waypoint sequence information of the initial flight path of the drone. Indicates the first The range of values for the flight speed of the drone. Indicates the first The level of onboard CNS service capabilities possessed by the drone itself. Indicates the first The level of airborne safety capabilities possessed by the drone itself. Indicates the first The mission preferences of drones; in, In the formula, Indicates the first A drone passes through three-dimensional waypoints The estimated arrival time, Indicates the first The number of waypoints for each drone; As shown below: In the formula, This indicates that the strategy is not allowed. or , This indicates that the strategy is not allowed. , This indicates that the strategy is not allowed. , This indicates that all strategies are allowed. and , This indicates an adjustment to the departure time. This indicates an adjustment to the flight speed. This indicates that the local path is being adjusted.
[0010] As an embodiment of the present invention, all flight plan tasks in the initial flight plan set are mapped to urban airspace, and spatiotemporal conflicts are identified to obtain conflict data, including: The initial flight plan will map all flight plan missions to urban airspace. When decision step Entering the three-dimensional spatial grid Number of drones inside More than three-dimensional spatial grid When the airspace capacity is reached, a flight conflict is determined, and the corresponding conflict data is identified. When the task number The drone enters the three-dimensional airspace grid Within the time frame, determine the task number. Do the drones meet the requirements of a three-dimensional airspace grid? The set airborne CNS service capability level or airborne safety capability level is determined; if it is not met, it is determined that a flight conflict has occurred, and the corresponding conflict data is determined. As shown below: In the formula, This represents the Boolean condition for determining when a flight conflict occurs. Representing a three-dimensional spatial raster airspace capacity, Representing a three-dimensional spatial raster volume, This represents the safe envelope volume required for the operation of a single drone. and Representing three-dimensional spatial grids The minimum airborne CNS service capability level and airborne safety capability level are set for entry.
[0011] As an embodiment of the present invention, based on the improved MADDPG algorithm, each UAV is treated as an intelligent agent, and the flight plan set is iteratively optimized to obtain the final executable multi-UAV four-dimensional flight plan set, including: Each drone is treated as an agent. Before each training session, an initial Main Actor network and Main Critic network are constructed for each agent, and a Target Actor network and Target Critic network are determined for each of the initial Main Actor network and Main Critic network. Determine each agent's role in the decision-making step. The local observation state at that time is used to obtain the decision-making state of all agents. The global state at that time; Based on the MADDPG algorithm and AMMP mechanism, according to the Main Actor network and decision steps of each agent... The state at that time determines the action; Generate joint actions based on the deterministic actions of each agent, as shown below: In the formula, Indicate decision steps Joint actions at the time Indicates the first An intelligent agent in the decision-making step Deterministic actions at that time; Perform joint actions and obtain global rewards and decision steps based on the reward functions of all agents. The global state at that time, combined with joint actions and decision steps. The global state at that time is obtained as an empirical tuple. ;in, Indicate decision steps The global state at that time, Indicates joint action, This represents the reward function value. Indicate decision steps The global state at that time; Based on the HPER mechanism, experience tuples are stored in the experience replay pool, and the target value is calculated by extracting experience tuples from the experience replay pool according to preset rules. The Main Critic network and Main Actor network are updated based on the target value, and the Target Actor network and Target Critic network are updated based on the updated Main Actor network and Main Critic network. Repeat the above steps to iteratively optimize the flight plan set and obtain a multi-UAV four-dimensional flight plan set.
[0012] As an embodiment of the present invention, each agent in the decision-making step The local observation state at that time is shown below: In the formula, Indicates the first An intelligent agent in the decision-making step The local observation status, Indicates the first An intelligent agent in the decision-making step Its own predicted state Indicates the first An intelligent agent in the decision-making step Task preference state, Indicates the first An intelligent agent in the decision-making step The state of conflict perception Indicates the first An intelligent agent in the decision-making step Local perception state, Let represent a feature tensor, where This refers to the number of feature channels. For each relative grid cell in the tensor, its corresponding eigenvector contains multiple feature channel parameters, which need to be processed. Dimensionality reduction; Indicates the first An intelligent agent in the decision-making step Predicting three-dimensional spatial raster coordinates, Indicates the first The final three-dimensional airspace grid coordinates of the initial flight plan of an agent Indicates the first An intelligent agent in the decision-making step The planned flight speed, Indicate decision steps The expected physical time, Indicates the first An intelligent agent in the decision-making step The original scheduled arrival time This indicates the presence of an indicator of future conflict. Indicates the time of conflict warning. Indicates the spatial distance for conflict early warning. Indicates the severity factor of the conflict; The reward functions for all agents are as follows: In the formula, Represents the reward function, , , , and Each represents a non-negative weighting coefficient for each reward component. Indicates the reward for conflict penalty items. This indicates a reward for risk-based penalties. This indicates a reward for delayed cost items. This indicates a reward for operational efficiency costs. Indicates a preference for rewards that violate penalties. This indicates the total number of spatiotemporal conflict points in the next state. This represents the total operational risk of all agents' waypoints. This represents the total delay time for all flight plans. and They represent the first The intelligent agent executes the new flight plan, specifying the takeoff and arrival times. and They represent the first The initial flight plan, including takeoff and arrival times, for each agent. This represents the total flight time for all planned flight missions. This indicates the overall degree of mission preference violation across all planned flight missions. and Both represent defined indicator functions. and All represent positive penalty constants. and All of these represent parameter sensitivity thresholds.
[0013] As an embodiment of the present invention, based on the MADDPG algorithm and AMMP mechanism, according to the MainActor network and decision steps of each agent... The state at that time determines a deterministic action, including: Through the Main Actor network of each agent, based on each agent's decision-making step... The state output at that time is the original action intent, as shown below: In the formula, Indicates the first An intelligent agent in the decision-making step The original intention of the action at that time Indicates the first An intelligent agent in the decision-making step The local path component of the original action intention at that time. Indicates the first An intelligent agent in the decision-making step The time component of the original action intention at that time Indicates the first An intelligent agent in the decision-making step The velocity component of the original intention of the action at that time; Based on the AMMP mechanism, the original action intent is masked as follows: In the formula, Indicates the first An intelligent agent in the decision-making step The intended action after being masked. This represents the Hadamard product operation; A linear scaling mapping is applied to the masked action intent to obtain a deterministic action, as shown below: In the formula, Indicates the first An intelligent agent in the decision-making step Deterministic actions at that time Indicates the first An intelligent agent in the decision-making step The time component of a deterministic action at a given time. Indicates the first An intelligent agent in the decision-making step The velocity component of the deterministic motion at time, Indicates the first An intelligent agent in the decision-making step Local path components of deterministic actions at time.
[0014] As an embodiment of the present invention, based on the HPER mechanism, experience tuples are stored in an experience replay pool, and the target value is calculated by extracting experience tuples from the experience replay pool according to preset rules, including: Construct an experience replay pool; the experience replay pool includes: a high-priority experience replay layer, a medium-priority experience replay layer, and a low-priority experience replay layer; The experience tuples are assigned to the experience replay pool based on the preset Boolean conditions and the components of the reward function. Among them, the preset Boolean conditions include: high-priority Boolean conditions, medium-priority Boolean conditions, and low-priority Boolean conditions; High-priority Boolean conditions As shown below: Medium-priority Boolean conditions are as follows: In the formula, , , and Both represent Boolean conditions. This represents the acceptable threshold for the reward of the risk penalty term in the reward function. This represents the acceptable threshold for the delayed cost term reward in the reward function. This represents the acceptable threshold for the reward function's efficiency cost component reward. Indicating in the decision-making step Risk penalty and reward for a single agent at any given time. Indicating in the decision-making step The latency cost of a single agent's reward at any given time Indicating in the decision-making step The operational efficiency cost of a single intelligent agent at any given time is a reward. This represents the acceptable individual threshold for the reward of a risk penalty term for a single agent. This represents the acceptable single-agent threshold for the delay cost reward of a single agent. This represents the acceptable individual threshold for the cost-benefit of the operational efficiency of a single agent. Low-priority Boolean conditions are shown below: According to preset rules, several experience tuples are extracted from the experience replay pool to obtain the target experience tuple set. The target value of each experience tuple in the target experience tuple set is then calculated, as shown below: In the formula, Represents the first set of target empirical tuples The target value of an empirical tuple. Indicates the first The reward function value of each experience tuple. Indicates the discount factor. Indicates the first The joint action is executed in each empirical tuple. The global state afterwards.
[0015] As an embodiment of the present invention, the preset rules include: inter-layer priority sampling and intra-layer priority experience replay sampling; According to preset rules, several experience tuples are extracted from the experience replay pool to obtain the target experience tuple set, including: The sampling quantity for each of the high-priority, medium-priority, and low-priority experience replay layers in the experience replay pool is determined based on inter-layer priority sampling, as shown below: In the formula, This represents the set of sample counts for the empirical tuples in each empirical replay layer. This indicates the number of experience tuples extracted from the high-priority experience replay layer. This indicates the number of experience tuples extracted from the middle-priority experience replay layer. This indicates the number of experience tuples extracted from the low-priority experience replay layer. This represents the sampling weight of each experience playback layer. This indicates the number of empirical tuples in the target empirical tuple set. This represents the set of number of experience tuples in each experience replay layer. This represents the total number of experience tuples in the experience replay pool. This represents the set of priority scaling factors for each experience playback layer. This represents the scaling factor for the high-priority experience playback layer. This indicates the scaling factor for the medium-priority experience playback layer. This represents the scaling factor for low-priority experience playback layers. , This indicates the number of experience tuples in the high-priority experience replay layer. This indicates the number of experience tuples in the medium-priority experience replay layer. This indicates the number of experience tuples in the low-priority experience replay layer; The sampling probability of each empirical tuple is calculated based on the priority empirical replay sampling within the layer, and sampling is performed according to the sampling probability of each empirical tuple to obtain the target empirical tuple set, as shown below: In the formula, Representing empirical tuples TD Error Representing empirical tuples The target value, Representing empirical tuples The global state, Representing empirical tuples The sampling probability, Hyperparameters that indicate the control priority This represents a smooth term with a probability of 0.
[0016] As an embodiment of the present invention, the Main Critic network and the Main Actor network are updated according to the target value, and the Target Actor network and the Target Critic network are updated using the updated Main Actor network and the Main Critic network, including: Based on the target value of each empirical tuple in the target empirical tuple set, the network parameters in the Main Critic network are... The network parameters of the Main Actor network are updated using the updated Main Critic network. Update as follows: In the formula, Represents the first set of target empirical tuples The sampling probability of an empirical tuple Indicates the degree of control deviation correction. Representing empirical tuples Importance sampling weights; According to the updated network parameters and Update the parameters of the Target network as follows: In the formula, This indicates the soft update parameter.
[0017] The beneficial effects of this invention are as follows: First, by learning a generalizable global collaborative optimization strategy, the traditional high-computation-cost iterative search is transformed into rapid inference by neural networks. This can significantly shorten the lead time window required for flight plan submission, making it more adaptable to temporary dynamic flight mission requests in large-scale, high-density operational scenarios, thereby improving the on-demand flexibility and efficient response speed of future UAM system operations. Second, by utilizing the task preference-aware action masking (AMMP) mechanism, non-compliant adjustment action components are forcibly filtered at the decision-making execution end based on preference constraints, ensuring that the generated flight plan can 100% meet the specific adjustment constraints of heterogeneous tasks, thus achieving true on-demand operational compliance. Finally, through the hierarchical priority experience replay (HPER) mechanism, hierarchical management and priority sampling of key and high-value experience samples are achieved, which can significantly improve the convergence speed and learning efficiency of the algorithm in multi-objective reward environments.
[0018] Other advantages, objectives, and features of the invention will be set forth in the following description and will be apparent to those skilled in the art in some respects, or may be learned by practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0019] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration: Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation
[0020] like Figure 1 As shown, this invention provides a preference-aware collaborative optimization method for multi-aircraft four-dimensional flight plans in urban low-altitude environments, comprising the following steps: The urban airspace is discretized to obtain several three-dimensional airspace grids, and the multidimensional attribute vector of each three-dimensional airspace grid is determined. The initial flight plan information of several UAVs is obtained and modeled to obtain an initial flight plan set; The initial flight plan sets all UAV flight plan tasks are mapped to urban airspace, and spatiotemporal conflicts are identified to obtain conflict data; the conflict data includes: the three-dimensional airspace grid of the conflict point, time, conflicting UAV, and the UAV's flight plan; Based on the improved MADDPG algorithm, each UAV is treated as an intelligent agent, and the flight plan set is iteratively optimized to obtain the final executable multi-UAV four-dimensional flight plan set. The improved MADDPG algorithm includes an AMMP mechanism, which is used to filter non-compliant adjustment action components of the intelligent agent according to the task preference attributes.
[0021] The working principle of the above technical solution is as follows: A digital urban low-altitude operating environment is constructed, and initial spatiotemporal tasks and conflict states are extracted to provide environmental constraints and input benchmarks for subsequent optimization. Specifically, this includes: characterizing the multidimensional static attributes of the airspace grid, including airspace reachability, infrastructure service capability, and operational risk attributes, to perform three-dimensional raster modeling of the complex urban low-altitude operating environment; formally defining and modeling the initial flight plans received from multiple UAVs from different operators, including task number, 4DT (4-dimensional time-of-flight) information, permissible flight speed range, onboard capability level, and task preference attributes. The initial 4DT information represents the individual optimal flight trajectory of a single UAV, calculated based on a path cost function with the optimization objective of minimizing operational risk and efficiency; the task preference attribute explicitly defines the operational preferences for each task, determining the operator's acceptance of takeoff time, flight speed, and local path adjustments. All initial flight plans of UAVs are mapped to a low-altitude grid environment. By integrating the matching degree between airspace grid attributes and the initial flight missions of the UAVs, a 4DT (4D flight conflict) judgment criterion for urban low-altitude strategic flight is established from two dimensions: airspace capacity constraints and operational operability matching. Through global spatiotemporal analysis of all initial flight plans, all potential 4DT spatiotemporal conflicts are identified.
[0022] Then, global collaborative optimization is achieved through an improved multi-agent reinforcement learning mechanism. Specifically, this includes: Each UAV is modeled as an independent agent. By designing and defining reasonable local observation states, action spaces, and global reward functions, the global conflict optimization problem is modeled as a multi-agent Markov decision process. An optimization method based on an improved Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is proposed. The core mechanism modules include Action Masking for Mission Preference (AMMP) and Hierarchical Prioritized Experience Replay (HPER). AMMP ensures the hard constraint execution of heterogeneous preferences by forcibly filtering non-compliant action components based on mission preference attributes at the action output. HPER improves the learning efficiency and policy performance during algorithm training by constructing three independent experience pools (high, medium, and low) and using inter-layer priority sampling and intra-layer priority experience replay sampling. The algorithm is iteratively trained to obtain the trained global collaborative optimization policy parameters. During the algorithm execution phase, relying on the optimization strategy learned through training, each agent adjusts its actions in parallel based on local observations, resolves conflicts and optimizes the initial flight plan, and finally generates a conflict-free and preference-aware strategic-layer multi-UAV four-dimensional flight plan.
[0023] The beneficial effects of the above technical solution are as follows: By learning a generalizable global collaborative optimization strategy, the traditional high-computation-cost iterative search is transformed into rapid inference by neural networks. This significantly shortens the lead time window required for flight plan submission, making it more adaptable to temporary dynamic flight mission requests in large-scale, high-density operational scenarios, thereby improving the on-demand flexibility and efficient response speed of the future UAM system. Secondly, by utilizing the task preference-aware action masking (AMMP) mechanism, non-compliant adjustment action components are forcibly filtered at the decision-making and execution end based on preference constraints, ensuring that the generated flight plan can 100% meet the specific adjustment constraints of heterogeneous tasks, thus achieving true on-demand operational compliance. Finally, through the hierarchical priority experience replay (HPER) mechanism, hierarchical management and priority sampling of key and high-value experience samples are achieved, which can significantly improve the convergence speed and learning efficiency of the algorithm in multi-objective reward environments.
[0024] In one implementation, the urban airspace is discretized to obtain several three-dimensional airspace grids, and the multidimensional attribute vector of each three-dimensional airspace grid is determined, including: A Cartesian coordinate system is constructed within the urban airspace, and the continuous three-dimensional urban operational airspace is discretized into several three-dimensional airspace grids along the three coordinate axes of the Cartesian coordinate system, as shown below: In the formula, Indicates urban airspace, Represents a three-dimensional spatial raster. Represents three-dimensional spatial grid coordinates. This indicates the extent of urban airspace on the X-axis. This indicates the extent of urban airspace on the Y-axis. This indicates the extent of urban airspace on the Z-axis. This indicates the floor function. This represents the size of a three-dimensional spatial raster. The multidimensional attribute vector for each 3D spatial raster is determined as follows: In the formula, Representing a three-dimensional spatial raster A multidimensional attribute vector, Representing a three-dimensional spatial raster The reachability of the airspace, Representing a three-dimensional spatial raster The attributes of infrastructure service capabilities Representing a three-dimensional spatial raster Operational risk attributes; in, As shown below: In the formula, Representing a three-dimensional spatial raster Horizontal area Representing a three-dimensional spatial raster The reachable airspace area; in, As shown below: In the formula, This represents the initial ground risk value. This indicates the reduction or mitigation of ground risks.
[0025] The working principle and beneficial effects of the above technical solution are as follows: To achieve urban airspace management, a multi-dimensional attribute parameter vector is assigned to each airspace grid to describe the inherent operational attribute conditions of that airspace grid; among which, This is an airspace reachability attribute, which is derived from the airspace structure itself and the distribution data of static buildings in the city. It represents the occupancy of static obstacles within the airspace grid and defines the physical space in which UAVs can safely and effectively perform flight missions. This is an infrastructure service capability attribute, which characterizes the service level of infrastructure such as communication, navigation, and surveillance that an airspace grid can provide. It reflects the grid's support and guarantee capabilities for UAV operations, and its values can be abstracted as follows: Any constant between, when When it indicates that the raster is completely unable to provide CNS services, conversely when This indicates that the grid possesses ideal CNS service capabilities; This is an operational risk attribute, characterizing the ground collision risk in urban operations, specifically the safety impact on people on the ground should an unmanned aerial vehicle (UAV) system fail in the air and crash. It integrates multiple risk factors, including ground population density and urban environmental shielding. The initial ground risk value refers to the level of casualties caused by a drone crash when the drone has gone out of control and crashed, and the ground population is completely exposed to the drone's impact area without any ground risk mitigation measures being taken. This is an inherent attribute of drone operation and is only related to the drone's own characteristic parameters and the ground population density in the crash area. The value represents the reduction in ground risk, indicating the combined effect of ground environmental shielding and the risk mitigation provided by the UAV parachute system.
[0026] In one embodiment, initial flight plan information of several UAVs is obtained for modeling to obtain an initial flight plan set, including: Obtain initial flight plan information for several drones; Based on the initial flight plan information of several UAVs, a model is created to obtain the initial flight plan set, as shown below: In the formula, Represents the initial flight plan set. Indicates the first The flight plan mission of the drone This indicates the total number of drones. Indicates the first The mission number of the drone. Indicates the first The four-dimensional waypoint sequence information of the initial flight path of the drone. Indicates the first The range of values for the flight speed of the drone. Indicates the first The level of onboard CNS service capabilities possessed by the drone itself. Indicates the first The level of airborne safety capabilities possessed by the drone itself. Indicates the first The mission preferences of drones; in, In the formula, Indicates the first A drone passes through three-dimensional waypoints The estimated arrival time, Indicates the first The number of waypoints for each drone; As shown below: In the formula, This indicates that the strategy is not allowed. or , This indicates that the strategy is not allowed. , This indicates that the strategy is not allowed. , This indicates that all strategies are allowed. and , This indicates an adjustment to the departure time. This indicates an adjustment to the flight speed. This indicates that the local path is being adjusted.
[0027] The working principle and beneficial effects of the above technical solution: It centralizes the data received by the management from the operator. The initial flight plan information of the drone was used to create a model; among which... By introducing a time dimension into the traditional three-dimensional spatial trajectory, this path is the [missing information - likely a specific path or feature]. The optimal trajectory of an individual drone is calculated by a path cost function that optimizes both operational risk and efficiency. Indicates the first The permissible flight speed range of a drone is determined by the physical characteristics of the drone; this invention only considers multi-rotor drones. This reflects the specific needs and preferences of operators for drone operation tasks. By integrating these into collaborative optimization goals and constraints, the acceptability of subsequent unified scheduling optimization schemes by management can be effectively improved, making them more in line with the actual operability of future UAM operations. Each type of task preference aims to reflect typical application scenarios, such as the attributes of drones performing urban law enforcement tasks. When performing inspection tasks, the attribute is When carrying out scheduled logistics delivery tasks, the attribute is Entertainment tasks have lower priority, fewer restrictions, and lower attributes. Furthermore, since the UAV's flight speed is acceptable as long as it remains within its permissible range regardless of the mission being performed during the strategic phase, this invention does not impose constraints on flight speed adjustment strategies in the mission preferences. Local path adjustment strategies must not be applied to the path's starting and ending points.
[0028] In one embodiment, all flight plan tasks in the initial flight plan set are mapped to urban airspace, and spatiotemporal conflicts are identified to obtain conflict data, including: The initial flight plan will map all flight plan missions to urban airspace. When decision step Entering the three-dimensional spatial grid Number of drones inside More than three-dimensional spatial grid When the airspace capacity is reached, a flight conflict is determined, and the corresponding conflict data is identified. When the task number The drone enters the three-dimensional airspace grid Within the time frame, determine the task number. Do the drones meet the requirements of a three-dimensional airspace grid? The set airborne CNS service capability level or airborne safety capability level is determined; if it is not met, it is determined that a flight conflict has occurred, and the corresponding conflict data is determined. As shown below: In the formula, This represents the Boolean condition for determining when a flight conflict occurs. Representing a three-dimensional spatial raster airspace capacity, Representing a three-dimensional spatial raster volume, This represents the safe envelope volume required for the operation of a single drone. and Representing three-dimensional spatial grids The minimum airborne CNS service capability level and airborne safety capability level are set for entry.
[0029] In one embodiment, based on the improved MADDPG algorithm, each UAV is treated as an intelligent agent, and the flight plan set is iteratively optimized to obtain the final executable multi-UAV four-dimensional flight plan set, including: Each drone is treated as an agent. Before each training session, an initial Main Actor network and Main Critic network are constructed for each agent, and a Target Actor network and Target Critic network are determined for each of the initial Main Actor network and Main Critic network. Determine each agent's role in the decision-making step. The local observation state at that time is used to obtain the decision-making state of all agents. The global state at that time; Based on the MADDPG algorithm and AMMP mechanism, according to the Main Actor network and decision steps of each agent... The state at that time determines the action; Generate joint actions based on the deterministic actions of each agent, as shown below: In the formula, Indicate decision steps Joint actions at the time Indicates the first An intelligent agent in the decision-making step Deterministic actions at that time; Perform joint actions and obtain global rewards and decision steps based on the reward functions of all agents. The global state at that time, combined with joint actions and decision steps. The global state at that time is obtained as an empirical tuple. ;in, Indicate decision steps The global state at that time, Indicates joint action, This represents the reward function value. Indicate decision steps The global state at that time; Based on the HPER mechanism, experience tuples are stored in the experience replay pool, and the target value is calculated by extracting experience tuples from the experience replay pool according to preset rules. The Main Critic network and Main Actor network are updated based on the target value, and the Target Actor network and Target Critic network are updated based on the updated Main Actor network and Main Critic network. Repeat the above steps to iteratively optimize the flight plan set, and obtain the final executable multi-UAV four-dimensional flight plan set.
[0030] The working principle and beneficial effects of the above technical solution are as follows: The AMMP-HPER-MADDPG algorithm is constructed, and the overall training process is described below: First, the Actor policy network, Critic value network, and their corresponding target network parameters for each agent are initialized. An independent experience replay pool containing high, medium, and low priorities is also initialized. During the centralized training phase, at the start of each training episode, the initial environment and state are obtained. Each UAV agent outputs its original action intent in parallel based on its local observation state. Subsequently, through the AMMP mechanism designed in this invention, an action mask vector is generated based on task preference attributes. This vector is applied to the original action intent to dynamically filter non-compliant action components and linearly scale it to the legal physical action space, generating the final executable action. Multi-agent joint actions are executed, and a global state transition is performed to obtain the next state and global reward. The generated experience tuples are then classified and stored in the corresponding priority levels according to the judgment conditions defined in the HPER mechanism designed in this invention. Subsequently, experience sampling is performed through inter-layer priority sampling and intra-layer priority experience replay sampling methods defined by the HPER mechanism to construct high-value training mini-batch samples. The parameters of the centralized Critic network are updated using global state information and joint action information from the samples, and their gradients guide the independent updates of each Actor network. Furthermore, the target network parameters of each agent are softly updated to stabilize the learning process. Finally, when the algorithm model reaches the preset end-training conditions (maximum number of rounds or reward value reaches the target) and meets the convergence requirements, the optimal policy network parameters are derived for subsequent distributed collaborative optimization scheduling of flight plans. The trained algorithm model is deployed to the centralized decision-making unit of the urban air traffic management system, where it makes decisions. During the algorithm execution phase, the trained global optimal collaborative optimization strategy is used for rapid distributed simulation. Each agent outputs conflict resolution and adjustment actions in parallel. These actions are strictly constrained by the AMMP mechanism to ensure that each action strictly conforms to the hard constraints of task preferences. Based on the compliant actions output by each agent, the management performs conflict resolution and optimization on the initial flight plan. All revised plans are aggregated to generate a conflict-free and preference-aware strategic-layer multi-UAV four-dimensional flight plan set. After verification, this set is fed back to the operators for each UAV to execute flight missions.
[0031] In one embodiment, each agent at the decision step The local observation state at that time is shown below: In the formula, Indicates the first An intelligent agent in the decision-making step The local observation status, Indicates the first An intelligent agent in the decision-making step Its own predicted state Indicates the first An intelligent agent in the decision-making step Task preference state, Indicates the first An intelligent agent in the decision-making step The state of conflict perception Indicates the first An intelligent agent in the decision-making step Local perception state, Let represent a feature tensor, where This refers to the number of feature channels. For each relative grid cell in the tensor, its corresponding eigenvector contains multiple feature channel parameters, which need to be processed. Dimensionality reduction; Indicates the first The initial flight plan of an agent predicts three-dimensional airspace grid coordinates. Indicates the first An intelligent agent in the decision-making step The endpoint three-dimensional spatial raster coordinates, Indicates the first An intelligent agent in the decision-making step The planned flight speed, Indicate decision steps The expected physical time, Indicates the first An intelligent agent in the decision-making step The original scheduled arrival time This indicates the presence of an indicator of future conflict. Indicates the time of conflict warning. Indicates the spatial distance for conflict early warning. Indicates the severity factor of the conflict; The reward functions for all agents are as follows: In the formula, Represents the reward function, , , , and Each represents a non-negative weighting coefficient for each reward component. Indicates the reward for conflict penalty items. This indicates a reward for risk-based penalties. This indicates a reward for delayed cost items. This indicates a reward for operational efficiency costs. Indicates a preference for rewards that violate penalties. This indicates the total number of spatiotemporal conflict points in the next state. This represents the total operational risk of all agents' waypoints. This represents the total delay time for all flight plans. and They represent the first The intelligent agent executes the new flight plan, specifying the takeoff and arrival times. and They represent the first The initial flight plan, including takeoff and arrival times, for each agent. This represents the total flight time for all planned flight missions. This indicates the overall degree of mission preference violation across all planned flight missions. and Both represent defined indicator functions. and All represent positive penalty constants. and All of these represent parameter sensitivity thresholds.
[0032] Based on the MADDPG algorithm and AMMP mechanism, according to the Main Actor network and decision steps of each agent... The state at that time determines a deterministic action, including: Through the Main Actor network of each agent, based on each agent's decision-making step... The state output at that time is the original action intent, as shown below: In the formula, Indicates the first An intelligent agent in the decision-making step The original intention of the action at that time Indicates the first An intelligent agent in the decision-making step The local path component of the original action intention at that time. Indicates the first An intelligent agent in the decision-making step The time component of the original action intention at that time Indicates the first An intelligent agent in the decision-making step The velocity component of the original intention of the action at that time; Based on the AMMP mechanism, the original action intent is masked as follows: In the formula, Indicates the first An intelligent agent in the decision-making step The intended action after being masked. This represents the Hadamard product operation; A linear scaling mapping is applied to the masked action intent to obtain a deterministic action, as shown below: In the formula, Indicates the first An intelligent agent in the decision-making step Deterministic actions at that time Indicates the first An intelligent agent in the decision-making step The time component of a deterministic action at a given time. Indicates the first An intelligent agent in the decision-making step The velocity component of the deterministic motion at time, Indicates the first An intelligent agent in the decision-making step Local path components of deterministic actions at time.
[0033] The working principle and beneficial effects of the above technical solution are as follows: The state and actions taken by each intelligence at different decision steps are displayed in its corresponding local observation state and action space, as detailed below: State space: specifically includes: First, self-predicted state This state component describes the agent. In the decision step t, the planned operation parameters are as follows: The predicted three-dimensional spatial grid coordinates of the agent at the current decision step t are derived from the currently effective flight plan and represent the agent's predicted physical moment at the current time. The grid corresponding to the time; The three-dimensional airspace grid coordinates representing the endpoint of the intelligent agent's execution of the initial flight plan, i.e., the target point coordinates. The planned flight speed at the current moment, which is within the specified speed range. Within. This represents the current grid. The originally scheduled arrival time. The intelligent agent compares... and First, it can internally perceive the current cumulative time deviation; second, it can detect the task preference state. This state component describes the agent. The hard constraints of task preference attributes that must be followed are the driving feature of the subsequent action masking mechanism; thirdly, the conflict perception state. This state component is obtained based on the strategic 4DT flight conflict determination and describes the attributes of the first spatiotemporal conflict point that occurs on the current planned simulation path. As an indicator of future conflicts, based on the flight plan that takes effect in the current decision step, if the agent's future projected path contains at least one spatiotemporal conflict, then... Conversely, if the remaining paths are entirely free of conflicts, then... This indicator bit serves as a conflict alarm signal, directly triggering the agent's policy adjustment mechanism. The conflict warning time is quantified as the time difference between the current moment and the earliest predicted time of conflict occurrence, reflecting the urgency of conflict handling and helping the agent weigh whether to make "minor adjustments" or "drastic interventions". Representing the spatial distance of the conflict, it quantifies the Euclidean distance between the predicted position of the agent at the current decision step and the center point of the conflict grid. This metric serves as an aid to time-based early warning and describes the proximity of conflicts in the spatial dimension. The conflict severity factor quantifies the degree to which predicted demand exceeds grid capacity at a specific spatiotemporal point of conflict. This factor helps the agent predict the potential costs of conflict resolution, thereby guiding it to find a better spatiotemporal trade-off among multiple compliance strategies. Fourth, local perception state. : Construct a grid based on the agent's current grid A fixed-size local perception window centered on the agent, covering the area within the agent's field of vision. The directly adjacent grids along the three axes form a spatial dimension of 3, including the center and 26 surrounding neighbors. 3 3. Three-dimensional spatial structure. Therefore, intelligent agents In the decision-making step The perceptible local environmental state information is defined as a 4-dimensional feature tensor. .in, This refers to the number of feature channels. For each relative grid cell in the tensor, its corresponding eigenvector contains multiple feature channel parameters, as shown below: Action Space: Based on three core optimization scheduling strategies, the agent is... In the decision-making step action Defined as a multidimensional continuous vector related to the adjustment policy. In the algorithm implementation, each action component output by the output layer of the Main Actor network is activated by an activation function. Mapped to Within the interval, the original action intent vector is obtained. The masked action intent vector is obtained through the AMMP mechanism module. Then, it is linearly scaled and mapped to the actual legal action space in the environment, specifically including: adjusting the takeoff time action component: this action component represents the initial 4DT flight plan. The overall time shift change, The maximum allowable variation in takeoff time. When This indicates a delay in takeoff. This indicates an advance takeoff maneuver. This indicates a constant takeoff maneuver. The flight speed adjustment component represents the change in speed applied to the current cruise speed. For the agent, its flight speed must strictly adhere to speed constraints. Therefore, it is necessary to first calculate the allowable speed adjustment boundary based on the current speed. Represents a decision step size The maximum allowable acceleration is usually a positive value. This represents the maximum permissible deceleration, typically a negative value. It is then dynamically mapped to a legal adjustment range, ensuring the agent's flight speed value in the next decision step. It can satisfy physical characteristics. Adjust local path motion components: specify the adjustment action within a decision time step. within, within The maximum adjustment of the path in each Cartesian direction of the axis cannot exceed the size of one grid cell. Therefore, this motion component can be defined as a three-dimensional continuous displacement vector, which is the normalized original three-dimensional intention vector of the local path adjustment motion output by the Main Actor network, with each component multiplied by the grid cell size of the corresponding dimension. To perform scaling.
[0034] Reward function: This flight plan should achieve flight planning and scheduling of multiple UAVs, including safety costs, while satisfying all constraints. Efficiency and cost and preference costs Overall comprehensive cost Minimize, as shown below: All N agents in the decision-making step Sharing the same global reward function It comprises five reward components. This reward function comprehensively considers factors such as flight safety, operational efficiency, and mission preferences from a global perspective, and includes a conflict penalty reward. and risk penalty items reward Together they represent security costs Delayed cost item reward and operational efficiency cost item rewards Together they characterize efficiency and cost And preferences for violating penalties and rewards This corresponds to the preference cost. Specifically, this includes: Risk Penalty Reward: This aims to encourage all agents to avoid high-risk areas on the ground as much as possible to improve overall safety. This reward accumulates the ground collision risk values of all airspace grids traversed by new flight plans; Delay Cost Reward: This aims to minimize deviations caused by flight plan adjustments, i.e., to encourage flight plans to adhere to the original mission time window as much as possible. Therefore, when a new flight plan deviates from the original takeoff and arrival times, a penalty is imposed; Operational Efficiency Cost Reward: This aims to encourage agents to adopt more efficient flight paths by minimizing total flight time; Preference Violation Penalty Reward: This aims to encourage agents to generate mission preference attributes consistent with the prescribed mission preferences. Maintain consistent original actions Define three indicator functions. , and For each preference attribute type, if the agent is not allowed to use strategies such as adjusting takeoff time, adjusting flight speed, or adjusting local path, the function takes a value of 1 respectively; otherwise, it takes a value of 0. In this invention, all types of preference attributes are allowed to adjust flight speed; therefore, It is always 0.
[0035] In one embodiment, based on the HPER mechanism, experience tuples are stored in an experience replay pool, and the target value is calculated by extracting experience tuples from the experience replay pool according to preset rules, including: Construct an experience replay pool; the experience replay pool includes: a high-priority experience replay layer, a medium-priority experience replay layer, and a low-priority experience replay layer; The experience tuples are assigned to the experience replay pool based on the preset Boolean conditions and the components of the reward function. Among them, the preset Boolean conditions include: high-priority Boolean conditions, medium-priority Boolean conditions, and low-priority Boolean conditions; High-priority Boolean conditions As shown below: Medium-priority Boolean conditions are as follows: In the formula, , , and Both represent Boolean conditions. This represents the acceptable threshold for the reward of the risk penalty term in the reward function. This represents the acceptable threshold for the delayed cost term reward in the reward function. This represents the acceptable threshold for the reward function's efficiency cost component reward. Indicating in the decision-making step Risk penalty and reward for a single agent at any given time. Indicating in the decision-making step The latency cost of a single agent's reward at any given time Indicating in the decision-making step The operational efficiency cost of a single intelligent agent at any given time is a reward. This represents the acceptable individual threshold for the reward of a risk penalty term for a single agent. This represents the acceptable single-agent threshold for the delay cost reward of a single agent. This represents the acceptable individual threshold for the cost-benefit of the operational efficiency of a single agent. Low-priority Boolean conditions are shown below: According to preset rules, several experience tuples are extracted from the experience replay pool to obtain the target experience tuple set. The target value of each experience tuple in the target experience tuple set is then calculated, as shown below: In the formula, Represents the first set of target empirical tuples The target value of an empirical tuple. Indicates the first The reward function value of each experience tuple. Indicates the discount factor. Indicates the first The joint action is executed in each empirical tuple. The global state afterwards.
[0036] The working principle and beneficial effects of the above technical solution are as follows: The core of this mechanism is to classify and store experience tuples into three independent experience pools with different priority levels based on the importance of different experiences to policy optimization. Samples are then taken from each level according to a standard priority experience replay method to obtain small batches of training samples, aiming to improve learning efficiency and policy performance during training. Three independent experience pools with high, medium, and low priorities are constructed. Specifically, these include: a high-priority experience replay layer: This layer stores the most critical experiences, including tuples that cause spatiotemporal conflicts or violate task preferences, aiming to force the algorithm to learn frequently "how to avoid catastrophic failure"; a medium-priority experience replay layer: This layer stores high-value heuristic experiences, including "successful endpoint" events and tuples that provide "warning" signals or cause significant efficiency drops, aiming to enable the algorithm to efficiently learn "how to succeed" and "avoid near-failure"; , , and Both represent Boolean conditions. This indicates that the global optimization target state has been reached, meaning it is triggered when there are no conflicts and no violation of task preferences. Indicates a new state There exists a critical state on the verge of a spatiotemporal conflict, namely, the existence of any four-dimensional spatiotemporal grid where the number of drones entering the grid is exactly equal to its airspace capacity. This indicates that the global cost defined in the reward function is too high. This condition is triggered whenever any one of these costs falls below its corresponding acceptable global threshold. The first layer represents extreme individual events diluted by the "global summation," triggered when the individual cost of at least one agent exceeds its corresponding acceptable individual threshold. The second layer, a low-priority experience replay layer, stores regular experience tuples from all other ordinary cases, providing foundational samples for the algorithm to learn its basic policy and optimize soft-constraint objectives. These three independent experience pools have fixed storage capacities; when storage within a layer reaches its limit, older experiences are deleted using a first-in, first-out (FIFO) strategy. This ensures that more critical experiences are retained for training, while less important experiences are discarded over time, thus improving the efficiency of the training process.
[0037] In one embodiment, the preset rules include: inter-layer priority sampling and intra-layer priority experience replay sampling; According to preset rules, several experience tuples are extracted from the experience replay pool to obtain the target experience tuple set, including: The sampling quantity for each of the high-priority, medium-priority, and low-priority experience replay layers in the experience replay pool is determined based on inter-layer priority sampling, as shown below: In the formula, This represents the set of sample counts for the empirical tuples in each empirical replay layer. This indicates the number of experience tuples extracted from the high-priority experience replay layer. This indicates the number of experience tuples extracted from the middle-priority experience replay layer. This indicates the number of experience tuples extracted from the low-priority experience replay layer. This represents the sampling weight of each experience playback layer. This indicates the number of empirical tuples in the target empirical tuple set. This represents the set of number of experience tuples in each experience replay layer. This represents the total number of experience tuples in the experience replay pool. This represents the set of priority scaling factors for each experience playback layer. This represents the scaling factor for the high-priority experience playback layer. This indicates the scaling factor for the medium-priority experience playback layer. This represents the scaling factor for low-priority experience playback layers. , This indicates the number of experience tuples in the high-priority experience replay layer. This indicates the number of experience tuples in the medium-priority experience replay layer. This indicates the number of experience tuples in the low-priority experience replay layer; The sampling probability of each empirical tuple is calculated based on the priority empirical replay sampling within the layer, and sampling is performed according to the sampling probability of each empirical tuple to obtain the target empirical tuple set, as shown below: In the formula, Representing empirical tuples TD Error Representing empirical tuples The target value, Representing empirical tuples The global state, Representing empirical tuples The sampling probability, Hyperparameters that indicate the control priority This represents a smooth term with a probability of 0.
[0038] In one embodiment, the Main Critic network and the Main Actor network are updated according to the target value, and the Target Actor network and the Target Critic network are updated using the updated Main Actor network and the Main Critic network, including: Based on the target value of each empirical tuple in the target empirical tuple set, the network parameters in the Main Critic network are... The network parameters of the Main Actor network are updated using the updated Main Critic network. Update as follows: In the formula, Represents the first set of target empirical tuples The sampling probability of an empirical tuple Indicates the degree of control deviation correction. Representing empirical tuples The importance sampling weights are typically normalized during the calculation process to maintain stability, resulting in... ; According to the updated network parameters and Update the parameters of the Target network as follows: In the formula, This indicates the soft update parameter.
[0039] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.
Claims
1. A preference-aware collaborative optimization method for multi-aircraft four-dimensional flight planning in urban low-altitude environments, characterized in that, Includes the following steps: The urban airspace is discretized to obtain several three-dimensional airspace grids, and the multidimensional attribute vector of each three-dimensional airspace grid is determined. The initial flight plan information of several UAVs is obtained and modeled to obtain an initial flight plan set; The initial flight plan sets all UAV flight plan tasks are mapped to urban airspace, and spatiotemporal conflicts are identified to obtain conflict data; the conflict data includes: the three-dimensional airspace grid of the conflict point, time, conflicting UAV, and the UAV's flight plan; Based on the improved MADDPG algorithm, each UAV is treated as an intelligent agent, and the flight plan set is iteratively optimized to obtain the final executable multi-UAV four-dimensional flight plan set. The improved MADDPG algorithm includes an AMMP mechanism, which is used to filter non-compliant adjustment action components of the intelligent agent according to the task preference attributes.
2. The preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight planning according to claim 1, characterized in that, The urban airspace is discretized to obtain several three-dimensional airspace grids, and the multi-dimensional attribute vector of each three-dimensional airspace grid is determined, including: A Cartesian coordinate system is constructed within the urban airspace, and the continuous three-dimensional urban operational airspace is discretized into several three-dimensional airspace grids along the three coordinate axes of the Cartesian coordinate system, as shown below: In the formula, Indicates urban airspace, Represents a three-dimensional spatial raster. Represents three-dimensional spatial grid coordinates. This indicates the extent of urban airspace on the X-axis. This indicates the extent of urban airspace on the Y-axis. This indicates the extent of urban airspace on the Z-axis. This indicates the floor function. This represents the size of a three-dimensional spatial grid. The multidimensional attribute vector for each 3D spatial raster is determined as follows: In the formula, Representing a three-dimensional spatial raster A multidimensional attribute vector, Representing a three-dimensional spatial raster The reachability of the airspace, Representing a three-dimensional spatial raster The attributes of infrastructure service capabilities Representing a three-dimensional spatial raster Operational risk attributes; in, As shown below: In the formula, Representing a three-dimensional spatial raster Horizontal area Representing a three-dimensional spatial raster The reachable airspace area; in, As shown below: In the formula, This represents the initial ground risk value. This indicates the reduction or mitigation of ground risks.
3. The preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight planning according to claim 1, characterized in that, Initial flight plan information from several UAVs is obtained and modeled to obtain an initial flight plan set, including: Obtain initial flight plan information for several drones; Based on the initial flight plan information of several UAVs, a model is created to obtain the initial flight plan set, as shown below: In the formula, Represents the initial flight plan set. Indicates the first The flight plan mission of the drone This indicates the total number of drones. Indicates the first The mission number of the drone. Indicates the first The four-dimensional waypoint sequence information of the initial flight path of the drone. Indicates the first The range of values for the flight speed of the drone. Indicates the first The level of onboard CNS service capabilities possessed by the drone itself. Indicates the first The level of airborne safety capabilities possessed by the drone itself. Indicates the first The mission preferences of drones; in, In the formula, Indicates the first A drone passes through three-dimensional waypoints The estimated arrival time, Indicates the first The number of waypoints for each drone; As shown below: In the formula, This indicates that the strategy is not allowed. or , This indicates that the strategy is not allowed. , This indicates that the strategy is not allowed. , This indicates that all strategies are allowed. and , This indicates an adjustment to the departure time. This indicates an adjustment to the flight speed. This indicates that the local path is being adjusted.
4. The preference-aware urban low-altitude multi-aircraft four-dimensional flight plan collaborative optimization method according to claim 2, characterized in that, The initial flight plan set maps all flight plan missions to urban airspace, and identifies spatiotemporal conflicts to obtain conflict data, including: The initial flight plan will map all flight plan missions to urban airspace. When decision step Entering the three-dimensional spatial grid Number of drones inside More than three-dimensional spatial grid When the airspace capacity is reached, a flight conflict is determined, and the corresponding conflict data is identified. When the task number The drone enters the three-dimensional airspace grid Within the time frame, determine the task number. Do the drones meet the requirements of a three-dimensional airspace grid? The set airborne CNS service capability level or airborne safety capability level is determined; if it is not met, it is determined that a flight conflict has occurred, and the corresponding conflict data is determined. As shown below: In the formula, This represents the Boolean condition for determining when a flight conflict occurs. Representing a three-dimensional spatial raster airspace capacity, Representing a three-dimensional spatial raster volume, This represents the safe envelope volume required for the operation of a single drone. and Representing three-dimensional spatial grids The minimum airborne CNS service capability level and airborne safety capability level are set for entry.
5. The preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight planning according to claim 1, characterized in that, Based on the improved MADDPG algorithm, each UAV is treated as an intelligent agent, and the flight plan set is iteratively optimized to obtain the final executable multi-UAV four-dimensional flight plan set, including: Each drone is treated as an agent. Before each training session, an initial Main Actor network and Main Critic network are constructed for each agent, and a Target Actor network and Target Critic network are determined for each of the initial Main Actor network and Main Critic network. Determine each agent's role in the decision-making step. The local observation state at that time is used to obtain the decision-making state of all agents. The global state at that time; Based on the MADDPG algorithm and AMMP mechanism, according to the Main Actor network and decision steps of each agent... The state at that time determines the action; Generate joint actions based on the deterministic actions of each agent, as shown below: In the formula, Indicate decision steps Joint actions at the time Indicates the first An intelligent agent in the decision-making step Deterministic actions at that time; Perform joint actions and obtain global rewards and decision steps based on the reward functions of all agents. The global state at that time, combined with joint actions and decision steps. The global state at that time is obtained as an empirical tuple. ;in, Indicate decision steps The global state at that time, Indicates joint action, This represents the reward function value. Indicate decision steps The global state at that time; Based on the HPER mechanism, experience tuples are stored in the experience replay pool, and the target value is calculated by extracting experience tuples from the experience replay pool according to preset rules. The Main Critic network and Main Actor network are updated based on the target value, and the Target Actor network and Target Critic network are updated based on the updated Main Actor network and Main Critic network. Repeat the above steps to iteratively optimize the flight plan set, and obtain the final executable multi-UAV four-dimensional flight plan set.
6. The preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight planning according to claim 5, characterized in that, Each agent in the decision-making step The local observation state at that time is shown below: In the formula, Indicates the first An intelligent agent in the decision-making step The local observation status, Indicates the first An intelligent agent in the decision-making step Its own predicted state Indicates the first An intelligent agent in the decision-making step Task preference state Indicates the first An intelligent agent in the decision-making step The state of conflict perception Indicates the first An intelligent agent in the decision-making step Local perception state, Let represent a feature tensor, where This refers to the number of feature channels. For each relative grid cell in the tensor, its corresponding eigenvector contains multiple feature channel parameters, which need to be processed. Dimensionality reduction; Indicates the first An intelligent agent in the decision-making step Predicting three-dimensional spatial raster coordinates, Indicates the first The final three-dimensional airspace grid coordinates of the initial flight plan of an agent Indicates the first An intelligent agent in the decision-making step The planned flight speed, Indicate decision steps The expected physical time, Indicates the first An intelligent agent in the decision-making step The original scheduled arrival time This indicates the presence of an indicator of future conflict. Indicates the time of conflict warning. Indicates the spatial distance for conflict early warning. Indicates the severity factor of the conflict; The reward functions for all agents are as follows: In the formula, Represents the reward function, , , , and Each represents a non-negative weighting coefficient for each reward component. Indicates the reward for conflict penalty items. This indicates a reward for risk-based penalties. This indicates a reward for delayed cost items. This indicates a reward for operational efficiency costs. Indicates a preference for rewards that violate penalties. This indicates the total number of spatiotemporal conflict points in the next state. This represents the total operational risk of all agents' waypoints. This represents the total delay time for all flight plans. and They represent the first The intelligent agent executes the new flight plan, specifying the takeoff and arrival times. and They represent the first The initial flight plan, including takeoff and arrival times, for each agent. This represents the total flight time for all planned flight missions. This indicates the overall degree of mission preference violation across all planned flight missions. and Both represent defined indicator functions. and All represent positive penalty constants. and All of these represent parameter sensitivity thresholds.
7. The preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight planning according to claim 5, characterized in that, Based on the MADDPG algorithm and AMMP mechanism, according to the Main Actor network and decision steps of each agent... The state at that time determines a deterministic action, including: Through the Main Actor network of each agent, based on each agent's decision-making step... The state output at that time is the original action intent, as shown below: In the formula, Indicates the first An intelligent agent in the decision-making step The original intention of the action at that time Indicates the first An intelligent agent in the decision-making step The local path component of the original action intention at that time. Indicates the first An intelligent agent in the decision-making step The time component of the original action intention at that time Indicates the first An intelligent agent in the decision-making step The velocity component of the original intention of the action at that time; Based on the AMMP mechanism, the original action intent is masked as follows: In the formula, Indicates the first An intelligent agent in the decision-making step The intended action after being masked. This represents the Hadamard product operation; A linear scaling mapping is applied to the masked action intent to obtain a deterministic action, as shown below: In the formula, Indicates the first An intelligent agent in the decision-making step Deterministic actions at that time Indicates the first An intelligent agent in the decision-making step The time component of a deterministic action at a given time. Indicates the first An intelligent agent in the decision-making step The velocity component of the deterministic motion at time, Indicates the first An intelligent agent in the decision-making step Local path components of deterministic actions at time.
8. The preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight planning according to claim 5, characterized in that, Based on the HPER mechanism, experience tuples are stored in the experience replay pool, and the target value is calculated by extracting experience tuples from the experience replay pool according to preset rules, including: Construct an experience replay pool; the experience replay pool includes: a high-priority experience replay layer, a medium-priority experience replay layer, and a low-priority experience replay layer; The experience tuples are assigned to the experience replay pool based on the preset Boolean conditions and the components of the reward function. Among them, the preset Boolean conditions include: high-priority Boolean conditions, medium-priority Boolean conditions, and low-priority Boolean conditions; High-priority Boolean conditions As shown below: Medium-priority Boolean conditions are as follows: In the formula, , , and Both represent Boolean conditions. This represents the acceptable threshold for the reward of the risk penalty term in the reward function. This represents the acceptable threshold for the delayed cost term reward in the reward function. This represents the acceptable threshold for the reward function's efficiency cost component reward. Indicating in the decision-making step Risk penalty and reward for a single agent at any given time. Indicating in the decision-making step The latency cost of a single agent's reward at any given time Indicating in the decision-making step The operational efficiency cost of a single intelligent agent at any given time is a reward. This represents the acceptable individual threshold for the reward of a risk penalty term for a single agent. This represents the acceptable single-agent threshold for the delay cost reward of a single agent. This represents the acceptable individual threshold for the cost-benefit of the operational efficiency of a single agent. Low-priority Boolean conditions are shown below: According to preset rules, several experience tuples are extracted from the experience replay pool to obtain the target experience tuple set. The target value of each experience tuple in the target experience tuple set is then calculated, as shown below: In the formula, Represents the first set of target empirical tuples The target value of an empirical tuple. Indicates the first The reward function value of each experience tuple. Indicates the discount factor. Indicates the first The joint action is executed in each empirical tuple. The global state afterwards.
9. The preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight planning according to claim 8, characterized in that, The preset rules include: inter-layer priority sampling and intra-layer priority experience replay sampling; According to preset rules, several experience tuples are extracted from the experience replay pool to obtain the target experience tuple set, including: The sampling quantity for each of the high-priority, medium-priority, and low-priority experience replay layers in the experience replay pool is determined based on inter-layer priority sampling, as shown below: In the formula, This represents the set of sample counts for the empirical tuples in each empirical replay layer. This indicates the number of experience tuples extracted from the high-priority experience replay layer. This indicates the number of experience tuples extracted from the middle-priority experience replay layer. This indicates the number of experience tuples extracted from the low-priority experience replay layer. This represents the sampling weight of each experience playback layer. This indicates the number of empirical tuples in the target empirical tuple set. This represents the set of number of experience tuples in each experience replay layer. This represents the total number of experience tuples in the experience replay pool. This represents the set of priority scaling factors for each experience playback layer. This represents the scaling factor for the high-priority experience playback layer. This indicates the scaling factor for the medium-priority experience playback layer. This represents the scaling factor for low-priority experience playback layers. , This indicates the number of experience tuples in the high-priority experience replay layer. This indicates the number of experience tuples in the medium-priority experience replay layer. This indicates the number of experience tuples in the low-priority experience replay layer; The sampling probability of each empirical tuple is calculated based on the priority empirical replay sampling within the layer, and sampling is performed according to the sampling probability of each empirical tuple to obtain the target empirical tuple set, as shown below: In the formula, Representing empirical tuples TD Error Representing empirical tuples The target value, Representing empirical tuples The global state, Representing empirical tuples The sampling probability, Hyperparameters that indicate the control priority This represents a smooth term with a probability of 0.
10. The preference-aware collaborative optimization method for urban low-altitude multi-aircraft four-dimensional flight planning according to claim 5, characterized in that, The Main Critic network and Main Actor network are updated based on the target value, and the Target Actor network and Target Critic network are then updated using the updated Main Actor network and Main Critic network, including: Based on the target value of each empirical tuple in the target empirical tuple set, the network parameters in the Main Critic network are... The network parameters of the Main Actor network are updated using the updated Main Critic network. Update as follows: In the formula, Represents the first set of target empirical tuples The sampling probability of an empirical tuple Indicates the degree of control deviation correction. Representing empirical tuples Importance sampling weights; According to the updated network parameters and Update the parameters of the Target network as follows: In the formula, This indicates the soft update parameter.