An airport flight delay collaborative mitigation and dynamic scheduling method and system
Patent Information
- Application Number
- CN202610799016.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-08
AI Technical Summary
[0006]本发明所要解决的技术问题是提供一种机场航班延误协同缓解与动态调度方法及系统,其旨在解决现有机场航班延误协同缓解技术中存在的以下技术问题:
[0073] 1. Achieve multi-agent game equilibrium and improve global collaborative efficiency: Through distributed Dec-POMDP modeling and Nash equilibrium solution, each agent negotiates autonomously under goal conflict, and the total system delay time is reduced compared with the traditional CDM mechanism, thereby improving the utilization rate of ground resources.
Smart Images

Figure CN122713622A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent decision-making and smart air traffic control technology, specifically relating to a collaborative method and system for mitigating airport flight delays based on constrained Markov decision processes and multi-agent game theory. Addressing the challenge of collaborative scheduling among multiple stakeholders—control tower, ground services, and gate allocation—under conditions of information asymmetry and conflicting objectives, this invention constructs a distributed, partially observable decision framework. Under the constraints of safety intervals, priorities, and resource capacity, it achieves multi-agent game equilibrium and delay coordination optimization, providing highly secure and interpretable intelligent decision support for the construction of smart civil aviation. Background Technology
[0002] Mitigating airport flight delays primarily relies on single-center scheduling optimization models and rule-driven collaborative decision-making mechanisms, but significant bottlenecks remain in practical applications. Traditional methods, such as collaborative decision-making (CDM) mechanisms, while enabling information sharing among multiple departments through fixed protocols, cannot effectively handle the conflicting objectives and dynamic game relationships among various decision-making entities (control tower, ground services, gate allocation, and traffic management), leading to scheduling strategies often resulting in local optimization but global deterioration.
[0003] While multi-agent reinforcement learning has partially improved collaborative decision-making capabilities in traffic scheduling, it still has significant limitations. On one hand, existing methods typically model each agent independently, neglecting the complex information dependencies and inter-agent game dynamics, making it difficult for strategies to converge to game equilibrium. Air traffic control prioritizes safe intervals, while ground services aim to maximize resource utilization; when these goals conflict, there is a lack of effective negotiation mechanisms. On the other hand, traditional reinforcement learning cannot embed aviation safety constraints (such as intervals, priorities, and resource capacity) during the exploration process, often generating scheduling actions that violate safety rules, such as consecutive takeoff intervals of less than 2 minutes or high-priority flights being overtaken by low-priority flights, directly preventing the strategy from being effectively applied.
[0004] Existing methods lack explicit modeling of the causal mechanisms of delays. Policy networks only learn data correlations rather than causal interventions, resulting in severe decision-making failures in off-distribution scenarios such as sudden weather events and traffic control. The dynamic scheduling of ground resources (parking spaces, jet bridges, shuttle buses) and the lack of robustness to uncertainties such as weather and traffic control significantly reduce the adaptability of policies in complex operating environments.
[0005] To address the aforementioned problems, this invention proposes an innovative solution capable of achieving multi-agent game equilibrium solving, safety constraint embedding and real-time correction, causal-guided decision optimization, and dynamic resource scheduling and uncertainty-robust training. This invention systematically solves existing technical bottlenecks through a distributed partially observable decision framework, Lagrange-constrained reinforcement learning, causal regularization mechanisms, and distributed robust optimization strategies, providing a safe, efficient, and interpretable intelligent decision-making solution for collaborative mitigation of airport flight delays. Summary of the Invention
[0006] The technical problem to be solved by this invention is to provide a method and system for coordinated mitigation and dynamic scheduling of airport flight delays, which aims to solve the following technical problems existing in the current airport flight delay coordination and mitigation technologies:
[0007] 1. Multi-agent goal conflict and game imbalance problem: Existing single-center scheduling models and collaborative decision-making mechanisms cannot effectively handle information asymmetry and interest conflicts among entities such as control tower, ground service, station allocation, and traffic management, resulting in local optimization and global deterioration of scheduling strategies, and failing to converge to multi-agent game equilibrium.
[0008] 2. Problems with the inability to embed safety constraints and violations of rules: Traditional multi-agent reinforcement learning cannot embed aviation safety constraints into the policy network during the exploration process, often generating scheduling actions that violate the rules, such as insufficient intervals or high-priority flights being cut in line by low-priority flights, which makes the policy unworkable in actual air traffic control.
[0009] 3. The problem of failure due to lack of causal mechanism: The existing policy network only learns the correlation of data and lacks explicit modeling of the causal structure of delay. In scenarios where the training is out of distribution, such as sudden weather and traffic control, the decision-making is seriously ineffective and the generalization ability is weak.
[0010] 4. Insufficient robustness to uncertain disturbances: The dynamic scheduling of ground resources has poor adaptability to uncertain factors such as weather changes, flow control commands, and equipment failures, and the performance of the strategy deteriorates significantly in complex operating environments.
[0011] This invention systematically solves the above problems by constructing a distributed partially observable decision framework, Lagrange-constrained reinforcement learning, causal regularization mechanisms, and a distributed bar optimization strategy.
[0012] First, this invention provides a method for collaborative mitigation and dynamic scheduling of airport flight delays, comprising the following steps:
[0013] Step 1: By modeling the distributed partially observable Markov decision process of multiple decision-making agents in airport surface operations, a multi-agent local observation and global joint decision-making framework is obtained.
[0014] Step 2: Based on the multi-agent local observation and global joint decision-making framework obtained in Step 1, aviation safety hard constraints are embedded through Lagrange constraint reinforcement learning to obtain candidate actions that satisfy the constraints and an adaptive penalty mechanism.
[0015] Step 3: Based on the multi-agent local observation and global joint decision-making framework obtained in Step 1, and the candidate actions and adaptive penalty mechanism that satisfy the constraints obtained in Step 2, a multi-agent Nash equilibrium collaborative strategy is obtained through centralized training, decentralized execution and counterfactual baseline optimization.
[0016] Step 4: Based on the multi-agent Nash equilibrium cooperative strategy obtained in Step 3, a causal decision model resistant to external perturbations is obtained through delay propagation causal graph modeling and causal regularization constraints.
[0017] Step 5: Based on the causal decision-making model against external disturbances obtained in Step 4, an uncertain adaptive robust scheduling strategy is obtained through sub-Bruker optimization and meta-learning adaptation.
[0018] Step 6: Based on the decision framework, constraint mechanism, equilibrium strategy, causal model, and robust strategy obtained in Steps 1–5, output a real-time visualized flight scheduling scheme through experience replay and multi-loss joint update.
[0019] In some possible implementations, step 1 specifically involves:
[0020] Step 1.1: Model the control tower, ground services, gate allocation, and traffic management entities as independent intelligent agents, each of which can only observe local state information. This constitutes part of the observable environment;
[0021] Step 1.2, Define the global state space Joint Action Space State transition probability and reward function :
[0022]
[0023] in The amount of reduction in total system delay. For the degree of achievement of each intelligent agent's own goals, Penalties for violations of safety intervals and prioritization conflicts. , , Preset weighting coefficients;
[0024] Step 1.3: Gated cyclic units are used to encode historical observation sequences and generate hidden states as input to the policy network to alleviate the lack of observation information.
[0025] Step 1 uses the distributed partially observable Markov decision process (Dec-POMDP) modeling module to characterize the information asymmetry and goal conflict among multiple agents, realize the transformation from local observation to global collaboration, and solve the problems of imbalance in multi-agent games and scheduling local optimality.
[0026] In some possible implementations, step 2 specifically involves:
[0027] Step 2.1: Express the three types of hard constraints—interval, priority, and capacity—as inequality constraint functions. :
[0028] (1) Spacing constraint: ;
[0029] (2) Priority constraints: ;
[0030] (3) Capacity constraints: ;
[0031] Step 2.2, construct the constrained Lagrangian optimization objective:
[0032]
[0033] And the Lagrange multipliers are adaptively updated according to gradient ascent:
[0034]
[0035] in The learning rate of the multiplier is set to non-negative to ensure that the constraint penalty does not weaken in the opposite direction.
[0036] Step 2.3: Project the violation action to the nearest legal action space via the security projection layer:
[0037]
[0038] in .
[0039] Step 2 transforms safety constraints into learnable penalty items, enforces compliance in the safety projection layer, eliminates illegal scheduling from the source, and ensures that the strategy meets air traffic control safety standards and can be directly implemented.
[0040] In some possible implementations, step 3 specifically involves:
[0041] Step 3.1: A centralized training and decentralized execution framework is adopted, with global information used for training and local observations used for execution;
[0042] Step 3.2 introduces a counterfactual baseline to evaluate the agent's marginal contribution and reduce the gradient estimation variance:
[0043]
[0044] Step 3.3: Drive convergence to Nash equilibrium using the joint strategy optimization objective:
[0045]
[0046] in For intelligent agents Cumulative rewards This is a strategy for reference.
[0047] Step 3: Counterfactual baseline reduces training fluctuations, joint optimization promotes game equilibrium among multiple agents, significantly improves global collaborative efficiency, and reduces total delay significantly compared to the traditional CDM mechanism.
[0048] In some possible implementations, step 4 specifically involves:
[0049] Step 4.1: Construct a cause-effect graph of delayed propagation structure. The nodes include meteorological data, flow control data, queue length, resource utilization, and delay duration.
[0050] Step 4.2: Randomly intervene in the cause nodes of the cause-effect graph to generate counterfactual samples;
[0051] Step 4.3: Calculate the KL divergence between factual and counterfactual features as the causal regularization loss:
[0052]
[0053] Step 4 explicitly models the delay causal mechanism to avoid learning spurious correlations and significantly improve the decision-making effectiveness in unseen scenarios such as sudden weather and traffic control.
[0054] In some possible implementations, step 5 specifically involves:
[0055] Step 5.1: Construct a set of environmental uncertainties using the Wasserstein sphere. It covers meteorological, flow control, and equipment malfunction disturbances;
[0056] Step 5.2, perform worst-case split optimization:
[0057]
[0058] Step 5.3, Learning the dynamic model of the environment Enhance exploration in areas with large prediction errors;
[0059] Step 5.4: Employ a meta-learning mechanism to achieve rapid adaptation to the new environment.
[0060] Step 5: Adversarial training optimizes performance in worst-case scenarios, while meta-learning enhances dynamic adaptability, keeping performance degradation caused by uncertainty to a minimum and exhibiting robustness far exceeding traditional methods.
[0061] In some possible implementations, step 6 specifically involves:
[0062] Step 6.1: Store the state, joint action, reward, and constraint violation into the replay buffer;
[0063] Step 6.2: The total loss is the weighted sum of the multi-agent proximal policy optimization loss, the causal regularization loss, and the constraint loss.
[0064] Step 6.3: After the model converges, it outputs a flight sorting Gantt chart, a parking stand / taxiway occupancy map, and a ground vehicle routing map, and supports interactive correction.
[0065] Step 6, decentralized execution, reduces single-step inference latency, meets the high-concurrency real-time scheduling needs of airports, and provides interpretable decision-making basis, thereby enhancing air traffic controllers' trust.
[0066] The present invention also provides an airport flight delay collaborative mitigation and dynamic scheduling system, used to apply the above-mentioned airport flight delay collaborative mitigation and dynamic scheduling method, including:
[0067] The distributed partially observable Markov decision process modeling module models multiple decision-making entities such as control tower, ground services, gate allocation, and traffic management in airport operations as independent intelligent agents, constructs a distributed collaborative decision-making framework in a partially observable environment, realizes unified modeling and joint decision-making of multiple entities under conditions of information asymmetry, and provides basic model support for subsequent constraint embedding, game coordination, and causal optimization.
[0068] The safety constraint embedding module based on Lagrange constraint reinforcement learning embeds three types of aviation safety hard constraints—interval, priority, and capacity—into the reinforcement learning policy network. It adaptively adjusts the constraint penalty intensity through Lagrange multipliers and corrects violations online through a safety projection layer, ensuring that all scheduling actions fully comply with air traffic control safety rules and achieving the integrated fusion of safety constraints and decision optimization.
[0069] The multi-agent game coordination and Nash equilibrium solution module, under the centralized training and decentralized execution architecture, realizes game interaction, contribution evaluation and strategy coordination among multiple agents, prompting each agent to negotiate autonomously under goal conflict, and finally converge to Nash equilibrium, realizing the leap from local optimization to global optimum, and improving the overall operational efficiency and coordination level of the airport.
[0070] The causal-guided decision optimization module explicitly models the causal propagation mechanism of flight delays. By using causal graphs and counterfactual interventions to guide the strategy network to learn real causal relationships rather than spurious data correlations, it significantly improves the generalization ability of the strategy in out-of-distribution scenarios such as sudden weather and air traffic control, and solves the out-of-distribution failure problem of traditional models.
[0071] The module for robust optimization and uncertainty adaptation models uncertainties such as weather forecast errors, flow control fluctuations, and equipment failures as a set of uncertainties. It optimizes the strategy performance in the worst case through adversarial training and combines meta-learning and model-driven exploration to improve the system's adaptability and robustness to dynamic disturbances.
[0072] The beneficial effects of this invention are:
[0073] 1. Achieve multi-agent game equilibrium and improve global collaborative efficiency: Through distributed Dec-POMDP modeling and Nash equilibrium solution, each agent negotiates autonomously under goal conflict, and the total system delay time is reduced compared with the traditional CDM mechanism, thereby improving the utilization rate of ground resources.
[0074] 2. Ensure that safety constraints are fully met and strategies can be directly deployed: Lagrange constraint reinforcement learning and safety projection layer ensure that all scheduling actions do not violate hard constraints such as interval, priority, and capacity. The probability of generating illegal actions is effectively reduced from a high level in traditional methods to zero, meeting air traffic control safety standards.
[0075] 3. Enhance out-of-distribution generalization ability to cope with sudden scenarios: The causal regularization mechanism significantly improves the decision-making effectiveness of the strategy in unseen scenarios such as sudden weather and traffic control compared to pure reinforcement learning, and the increase in delay is significantly reduced.
[0076] 4. Improve robustness to uncertain disturbances: The bibliometric optimization and meta-learning mechanism keep the performance degradation of the strategy within a small range under conditions such as large weather forecast errors and random disturbances in flow control commands, while the traditional method suffers significant performance degradation.
[0077] 5. Provide interpretable decision-making basis: Cause-effect diagrams and counterfactual intervention analysis provide diagnostic information such as delays being mainly caused by flow control, enhancing air traffic controllers' trust in decision-making and significantly improving the acceptance rate of collaborative decision-making.
[0078] 6. Computational efficiency meets real-time scheduling requirements: Decentralized execution reduces single-step inference latency, supports rescheduling in shorter cycles, and meets the real-time requirements of high-concurrency airport scenarios. Attached Figure Description
[0079] Figure 1 This is a diagram showing the overall architecture of the system described in this invention.
[0080] Figure 2 The structure of the Dec-POMDP model diagram;
[0081] Figure 3 This is a schematic diagram of the security projection layer;
[0082] Figure 4 An evolution diagram of multi-agent game interaction and joint actions;
[0083] Figure 5 Here is an example of a cause-effect graph structure;
[0084] Figure 6 This is a schematic diagram of the Lagrange multiplier adaptive update mechanism;
[0085] Figure 7 Training flowchart;
[0086] Figure 8 This is a schematic diagram of airport surface operations scheduling. Detailed Implementation
[0087] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0088] This invention first provides a method for collaborative mitigation and dynamic scheduling of airport flight delays, comprising the following steps:
[0089] Step 1: Multi-agent cooperative decision-making modeling based on Dec-POMDP
[0090] This step details the specific implementation of the distributed partially observable Markov decision process (Dec-POMDP) modeling module in this invention, and its model structure is as follows: Figure 2 As shown.
[0091] Step 1.1: Agent Definition and Environment Modeling
[0092] Taking a typical hub airport as an application scenario, the four key decision-making entities in the airport's operations are modeled as independent intelligent agents:
[0093] Tower agent: responsible for runway resource scheduling and deciding on the order and interval of takeoffs / landings;
[0094] Ground service intelligent agent: responsible for the coordination and scheduling of ground resources such as boarding bridges, shuttle buses, refueling trucks, and baggage carts;
[0095] Machine station allocation agent: responsible for the dynamic allocation and release of parking spaces;
[0096] Traffic Management Agent: Responsible for responding to air traffic control instructions and adjusting flight schedules.
[0097] Each agent at time Only local observations can be obtained. This includes flight queue lengths, resource occupancy status, and communication information of neighboring agents within the region, constituting a portion of the observable environment. Global State It contains complete information about all agents, but is not visible to the individual agents.
[0098] Step 1.2: Design of State Space, Action Space, and Reward Function
[0099] Global state space Scheduled times, actual times, and delay durations for all flights; occupancy status and queue lengths for each runway; occupancy status of each parking stand; available number and location of ground vehicles; flow control command level.
[0100] Joint Action Space :
[0101] in (Tower): Select the next flight to depart / arrive, specify the runway and time; (Ground): Allocating ground service resources to specific flights; (Aircraft parking positions): Assigning inbound flights to available parking positions; (Flow): Adjust flight schedules in response to flow control instructions.
[0102] reward function The design is a weighted sum of collaborative rewards, individual rewards, and conflict penalties.
[0103]
[0104] in The amount of reduction in total system delay. For each intelligent agent's own goal achievement (such as interval compliance reward for tower intelligent agents, and resource utilization reward for ground service intelligent agents). Penalties for conflicting behaviors such as violation of safety intervals and reversal of priorities. , , The preset weighting coefficients can be adjusted according to actual business needs.
[0105] Step 1.3: Partial Observability Processing
[0106] Gated recurrent units are used to encode historical observation sequences and generate the hidden state at the current moment as input to the policy network to alleviate the problem of insufficient information caused by missing observations.
[0107] Step 2: Secure Constraint Embedding Based on Lagrange Constraint Reinforcement Learning
[0108] This step details the specific implementation of the security constraint embedding and correction module, and its core mechanism is as follows: Figure 3 As shown, the Lagrange multiplier adaptive update is as follows: Figure 6 As shown.
[0109] Step 2.1: Formal Definition of Aviation Safety Constraints
[0110] The three types of hard constraints are expressed as inequality constraint functions. :
[0111] Interval constraint: Time difference between consecutive takeoffs or landings It must not be less than the minimum safety interval , Choose 90 seconds or 120 seconds depending on the model.
[0112]
[0113] Priority constraint: The position of higher-priority flights in the sorting sequence. The priority level must not exceed the position of the lower priority flight. .
[0114]
[0115] Capacity constraints: The number of runways, taxiways, or parking positions occupied simultaneously. It must not exceed its capacity limit. .
[0116]
[0117] Step 2.2: Adaptive Update of Lagrange Multipliers
[0118] Constructing the Lagrangian function incorporates constraints into the optimization objective:
[0119]
[0120] Lagrange multipliers Adaptive updates using gradient ascent:
[0121]
[0122] in The learning rate is set to non-negative to ensure that the constraint penalty does not weaken in the opposite direction. When the constraint violation... hour, Increase the penalty for violating the constraint; conversely... Gradually decrease.
[0123] Step 2.3: Security Projection Layer Correction
[0124] Output candidate actions in the policy network Then, proceed to the secure projection layer:
[0125] like Then output directly. As a safety measure; otherwise, solve the following convex optimization problem, Projected onto the nearest movable space:
[0126]
[0127] in The alternating direction multiplier method is used for online solution. This algorithm can pre-compute some steps offline to accelerate the process and meet the needs of real-time decision-making.
[0128] Step 3: Multi-agent game coordination and Nash equilibrium solution
[0129] This step details the specific implementation of the multi-agent game coordination and Nash equilibrium solution module, including the game interaction and joint action evolution process among multiple agents, as follows: Figure 4 As shown.
[0130] Step 3.1: Centralized Training and Decentralized Execution Framework
[0131] A centralized training and decentralized execution framework is adopted. During the training phase, global state is used. and joint actions Guides the policy updates of each agent; during the execution phase, each agent relies only on its own local observations. Output action.
[0132] Step 3.2: Counterfactual Baseline
[0133] Introducing a counterfactual baseline to evaluate each agent's marginal contribution to the team reward reduces gradient estimation variance:
[0134]
[0135] That is, by fixing the actions of other agents, the average Q-value of the current agent when taking different actions is calculated as the baseline. The expected value is approximated by Monte Carlo sampling, with the number of samplings preset to 10.
[0136] Step 3.3: Joint Policy Optimization and Nash Equilibrium
[0137] Design a joint policy optimization objective to encourage the policy to converge to Nash equilibrium:
[0138]
[0139] in For intelligent agents Cumulative rewards This serves as a reference strategy. An iterative update algorithm using a multi-agent proximal policy optimization method ensures monotonic improvement and convergence to an equilibrium solution.
[0140] Step 4: Causal-guided decision optimization
[0141] This step details the specific implementation of the causal-guided decision optimization module, and its causal graph structure is as follows: Figure 5 As shown.
[0142] Step 4.1: Constructing the Cause-and-Effect Graph of Delayed Propagation
[0143] Based on expert knowledge and historical data in the aviation field, construct such as Figure 5 Cause-and-effect diagram shown The nodes include: meteorological conditions ( ), flow control instructions ( ), takeoff queue length ( ), Ground resource occupancy rate ( ), actual delay time ( ).
[0144] Causality edge: , , , , The Peter-Clark algorithm is used to learn causal structures from historical data, and confidence thresholds are set for pruning.
[0145] Step 4.2: Counterfactual Intervention and Feature Coding
[0146] During training, the cause nodes in the cause-effect graph are randomly intervened. For example, flow control commands are forcibly applied. The current observations are intervened at a specific level to generate counterfactual samples. The factual and counterfactual samples are then input into a weighted causal encoder to obtain factual features. Counterfactual features The encoder structure uses a three-layer multilayer perceptron (MLP), with 256 neurons in each layer and ReLU activation function.
[0147] Step 4.3: Calculation of Causal Regularization Loss
[0148] Calculate the KL divergence between factual and counterfactual features as the causal regularization loss:
[0149]
[0150] The feature representations learned by the loss-constrained policy network are insensitive to intervention operations in the causal graph, i.e., they satisfy counterfactual invariance, thus learning the true causal structure rather than spurious correlations. The weighting coefficient is set to 0.1.
[0151] Step 5: Optimization of Blue Bars and Overall Training Process
[0152] This step details the Bruker optimization and uncertainty adaptive module, as well as the overall training process of the system. The training process is as follows: Figure 7 As shown.
[0153] Step 5.1: Construction of the Uncertainty Set
[0154] Modeling environmental dynamics as a set of uncertainties The set of uncertainties includes the following possible distributions: preset weather forecast error range (e.g., ±30%); preset flow control command level offset range (e.g., ±1 level); downtime follows a Poisson distribution; and a preset failure rate parameter is used. A Wasserstein sphere is used to construct the uncertainty set.
[0155] Step 5.2: Worst-case optimization and adversarial training
[0156] The optimization objective is to maximize the cumulative reward in the worst-case scenario:
[0157]
[0158] In each training round, the opponent starts from... The system selects the environment distribution that results in the worst-case performance for the current policy, and the policy network updates accordingly. A meta-learning mechanism is employed to accelerate adaptation to new environments, with a preset inner loop learning rate. and outer loop learning rate .
[0159] Step 5.3: Model-based exploration enhancement
[0160] Learning Environment Dynamic Model In areas with large prediction errors, exploration weights are increased. The environment model employs an ensemble neural network, and prediction uncertainty is estimated through the variance of the ensemble members' outputs. Exploration reward. Add to the total reward to guide the strategy to access areas of high uncertainty.
[0161] Step 5.4: Overall Training Cycle
[0162] First, initialization is performed, loading airport operation data, including flight schedules, weather forecasts, and flow control instructions. At the same time, the policy network parameters, Lagrange multipliers, causal graph structures, and encoder parameters of each agent are initialized, and the replay buffer is cleared.
[0163] At each decision step (e.g., every 30 seconds), each agent outputs candidate actions through the policy network based on its observed local information. Subsequently, the safety projection layer checks the candidate actions and corrects any actions that violate constraints. After all agents' actions are executed jointly, the environment state is updated, and the reward obtained in this round of scheduling and the degree of constraint violation are calculated. The Lagrange multipliers are adaptively adjusted based on the constraint violation situation. Finally, the current state, joint actions, reward, next-time state, and constraint violation amount are stored in the replay buffer.
[0164] During the training and update phase, a batch of historical experiences is randomly sampled from the replay buffer, and the multi-agent proximal policy optimization loss, causal regularization loss, and constraint loss are calculated sequentially. These three losses are weighted and summed to obtain the total loss, with the weights of the causal regularization loss and constraint loss set to preset values. The optimizer updates the parameters of the policy network, value network, and causal encoder based on the total loss, with the learning rate set to a preset value.
[0165] Step 6: Output and Visualization of Scheduling Instructions
[0166] This step details the specific implementation of the system output module, and its output format is as follows: Figure 8 As shown. After training and convergence, the system can output scheduling plans in real time. The convergence condition can be set as the total system delay reduction falling below a preset threshold for several consecutive steps, or the change in policy network parameters being less than a set tolerance. The output scheduling plan mainly includes three visualization parts. The first part is a flight sequence Gantt chart, with the time axis as the horizontal axis, showing the takeoff or landing time, priority marking, and interval compliance of each flight, making it easy for air traffic controllers to intuitively grasp the timing arrangement of flight sequences. The second part is a taxiway and parking stand occupancy time sequence diagram, which uses color coding to show the occupancy status of each key resource at different time periods and marks possible conflict areas to help identify resource bottlenecks. The third part is a ground vehicle path map, which dynamically displays the travel paths and current positions of ground vehicles such as shuttle buses, refueling trucks, and baggage trucks on the airport map. The above visualization interface supports interactive adjustments by air traffic controllers, such as manually modifying flight priorities or temporarily inserting emergency flights. The system will re-call the safety projection layer based on the adjusted status to correct actions and output decisions, ensuring the real-time performance and safety of the scheduling plan.
[0167] like Figure 1 As shown, the present invention also provides an airport flight delay collaborative mitigation and dynamic scheduling system, used to apply the above-mentioned airport flight delay collaborative mitigation and dynamic scheduling method, including:
[0168] (1) Distributed Partially Observable Markov Decision Process (Dec-POMDP) Modeling Module
[0169] The various decision-making entities in airport operations are modeled as intelligent agents, including: a tower agent (responsible for takeoff / landing sequencing), a ground service agent (responsible for scheduling jet bridges, shuttle buses, and refueling trucks), a parking stand allocation agent (responsible for dynamic allocation of parking stands), and a traffic management agent. These agents can only observe local state information (such as local flight queues and resource occupancy), constituting a partially observable environment. Their model structure is as follows: Figure 2 As shown.
[0170] Define the global state space Joint Action Space State transition probability and reward function The reward function is designed as follows:
[0171] Collaborative reward: Reduction in total system delay time (negative values indicate penalties);
[0172] Individual reward items: the degree of achievement of each intelligent agent's own goals (e.g., tower pursues interval compliance reward +1, ground service pursues resource utilization rate +0.5).
[0173] Conflict penalties: Violations of safety intervals, reversals of priorities, etc., will result in a negative reward of -10.
[0174] (2) Safety constraint embedding module based on Lagrange constraint reinforcement learning
[0175] To meet the hard constraints of aviation safety, the Lagrange multiplier method is used to transform the constraints into penalty terms in the objective function. The constraint function is defined as follows: ,include:
[0176] Interval constraint: Time difference between consecutive takeoffs / landings ;
[0177] Priority constraint: The ranking position of a high-priority flight (international, VIP) cannot be surpassed by a low-priority flight;
[0178] Capacity constraints: Number of parking spaces, taxiways, and runways simultaneously occupied Capacity limit;
[0179] Constructing the Lagrange function:
[0180] ;
[0181] in The Lagrange multipliers are used for adaptive updates via gradient ascent. A safety projection layer (e.g., ...) is added after the policy network outputs the action. Figure 3 (As shown): If an action violates any hard constraints, it is projected into the nearest legal action space.
[0182] (3) Multi-agent game coordination and Nash equilibrium solution module
[0183] A centralized training and decentralized execution framework is adopted. During training, global information guides the policy updates of each agent, while the execution phase relies only on local observations. A counterfactual baseline is introduced to evaluate the marginal contribution of each agent to the team reward, reducing gradient estimation variance. The game interaction and joint action evolution process among multiple agents is as follows: Figure 4 As shown. To ensure the policy converges to Nash equilibrium, a joint policy optimization objective is designed:
[0184] ;
[0185] in For intelligent agents Cumulative rewards The reference strategy is used. Iterative updates are performed using the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm, ensuring monotonic improvement and convergence to an equilibrium solution.
[0186] (4) Causal-guided decision optimization module
[0187] To enhance the generalization ability of the strategy in out-of-distribution scenarios, a structural causal model is introduced to explicitly model the delay propagation mechanism. A causal graph is constructed. The nodes include: weather conditions, air traffic control instructions, takeoff queue length, ground resource occupancy rate, and actual delay time. An example of its structure is shown below. Figure 5 As shown. The features learned by the causal regularization constraint policy network are consistent with the intervention distribution in the causal graph:
[0188] ;
[0189] in Indicates the causal variable The feature distribution after intervention. During training, randomly intervening in certain nodes of the causal graph requires the policy network output to be insensitive to the intervention, thereby learning the true causal structure rather than spurious correlations.
[0190] (5) Optimization of Blue Bars and Adaptive Module for Uncertainty
[0191] To address uncertainties such as weather and flow control, the Distributed Bar Optimization (DRO) method is employed to model the environmental dynamics as a set of uncertainties. The worst-case distribution within the range. The optimization objective is:
[0192] ;
[0193] Worst-case transition probabilities are generated through adversarial training, and a meta-learning mechanism is used to quickly adapt to new environments. A model-based multi-agent reinforcement learning approach is employed to learn a dynamic model of the environment. Increase the exploration weight in areas with large prediction errors.
[0194] The overall system workflow is as follows:
[0195] First, airport operation data is loaded, and the policy networks, Lagrange multipliers, and causal graph structures of each agent are initialized. Then, the cyclic scheduling phase begins. At each decision step, each agent, based on its local observations… Output candidate actions Subsequently, the safety projection layer checks the candidate actions and corrects any violations; after the joint action is executed, the environment state is updated, and rewards and constraint violations are calculated simultaneously. The Lagrange multipliers are adaptively adjusted based on the constraint violations (the mechanism is as follows). Figure 6 As shown in the diagram), the experience is stored in the replay buffer. The entire training process is as follows: Figure 7 As shown. During the training and update phase, samples are taken from the replay buffer, and the MAPPO loss, causal regularization loss, and constraint loss are calculated. Based on these, the policy network, value network, and causal encoder are updated. The final output is a scheduling instruction (such as...). Figure 8 As shown in the figure): After training and convergence, the system can output scheduling schemes such as flight sequencing, gate allocation, and ground vehicle routes in real time.
[0196] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for collaborative mitigation and dynamic scheduling of airport flight delays, characterized in that, The steps include the following: Step 1: By modeling the distributed partially observable Markov decision process of multiple decision-making agents in airport surface operations, a multi-agent local observation and global joint decision-making framework is obtained. Step 2: Based on the multi-agent local observation and global joint decision-making framework obtained in Step 1, aviation safety hard constraints are embedded through Lagrange constraint reinforcement learning to obtain candidate actions that satisfy the constraints and an adaptive penalty mechanism. Step 3: Based on the multi-agent local observation and global joint decision-making framework obtained in Step 1, and the candidate actions and adaptive penalty mechanism that satisfy the constraints obtained in Step 2, a multi-agent Nash equilibrium collaborative strategy is obtained through centralized training, decentralized execution and counterfactual baseline optimization. Step 4: Based on the multi-agent Nash equilibrium cooperative strategy obtained in Step 3, a causal decision model resistant to external perturbations is obtained through delay propagation causal graph modeling and causal regularization constraints. Step 5: Based on the causal decision-making model against external disturbances obtained in Step 4, an uncertain adaptive robust scheduling strategy is obtained through sub-Bruker optimization and meta-learning adaptation. Step 6: Based on the decision framework, constraint mechanism, equilibrium strategy, causal model, and robust strategy obtained in Steps 1–5, output a real-time visualized flight scheduling scheme through experience replay and multi-loss joint update.
2. The airport flight delay collaborative mitigation and dynamic scheduling method according to claim 1, characterized in that, Step 1 is as follows: Step 1.1: Model the control tower, ground services, gate allocation, and traffic management entities as independent intelligent agents, each of which can only observe local state information. This constitutes part of the observable environment; Step 1.2, Define the global state space Joint Action Space State transition probability and reward function : in The amount of reduction in total system delay. For the degree of achievement of each intelligent agent's own goals, Penalties for violations of safety intervals and prioritization conflicts. , , Preset weighting coefficients; Step 1.3: Gated cyclic units are used to encode historical observation sequences and generate hidden states as input to the policy network to alleviate the lack of observation information.
3. The airport flight delay collaborative mitigation and dynamic scheduling method according to claim 1, characterized in that, Step 2 is as follows: Step 2.1: Express the three types of hard constraints—interval, priority, and capacity—as inequality constraint functions. : (1) Spacing constraint: ; (2) Priority constraints: ; (3) Capacity constraints: ; Step 2.2, construct the constrained Lagrangian optimization objective: And the Lagrange multipliers are adaptively updated according to gradient ascent: in The learning rate of the multiplier is set to non-negative to ensure that the constraint penalty does not weaken in the opposite direction. Step 2.3: Project the violation action to the nearest legal action space via the security projection layer: in .
4. The airport flight delay collaborative mitigation and dynamic scheduling method according to claim 1, characterized in that, Step 3 specifically involves: Step 3.1: A centralized training and decentralized execution framework is adopted, with global information used for training and local observations used for execution; Step 3.2 introduces a counterfactual baseline to evaluate the agent's marginal contribution and reduce the gradient estimation variance: Step 3.3: Drive convergence to Nash equilibrium using the joint strategy optimization objective: in For intelligent agents Cumulative rewards This is a strategy for reference.
5. The airport flight delay collaborative mitigation and dynamic scheduling method according to claim 1, characterized in that, Step 4 specifically involves: Step 4.1: Construct a cause-effect graph of delayed propagation structure. The nodes include meteorological data, flow control data, queue length, resource utilization, and delay duration. Step 4.2: Randomly intervene in the cause nodes of the cause-effect graph to generate counterfactual samples; Step 4.3: Calculate the KL divergence between factual and counterfactual features as the causal regularization loss:
6. The airport flight delay collaborative mitigation and dynamic scheduling method according to claim 1, characterized in that, Step 5 specifically involves: Step 5.1: Construct a set of environmental uncertainties using the Wasserstein sphere. It covers meteorological, flow control, and equipment malfunction disturbances; Step 5.2, perform worst-case split optimization: Step 5.3, Learning the dynamic model of the environment Enhance exploration in areas with large prediction errors; Step 5.4: Employ a meta-learning mechanism to achieve rapid adaptation to the new environment.
7. The airport flight delay collaborative mitigation and dynamic scheduling method according to claim 1, characterized in that, Step 6 specifically involves: Step 6.1: Store the state, joint action, reward, and constraint violation into the replay buffer; Step 6.2: The total loss is the weighted sum of the multi-agent proximal policy optimization loss, the causal regularization loss, and the constraint loss. Step 6.3: After the model converges, it outputs a flight sorting Gantt chart, a parking stand / taxiway occupancy map, and a ground vehicle routing map, and supports interactive correction.
8. An airport flight delay collaborative mitigation and dynamic scheduling system, used to apply the airport flight delay collaborative mitigation and dynamic scheduling method as described in any one of claims 1-7, characterized in that, include: The distributed partially observable Markov decision process modeling module models multiple decision-making entities such as control tower, ground services, gate allocation, and traffic management in airport operations as independent intelligent agents, constructs a distributed collaborative decision-making framework in a partially observable environment, realizes unified modeling and joint decision-making of multiple entities under conditions of information asymmetry, and provides basic model support for subsequent constraint embedding, game coordination, and causal optimization. The safety constraint embedding module based on Lagrange constraint reinforcement learning embeds three types of aviation safety hard constraints—interval, priority, and capacity—into the reinforcement learning policy network. It adaptively adjusts the constraint penalty intensity through Lagrange multipliers and corrects violations online through a safety projection layer, ensuring that all scheduling actions fully comply with air traffic control safety rules and achieving the integrated fusion of safety constraints and decision optimization. The multi-agent game coordination and Nash equilibrium solution module, under the centralized training and decentralized execution architecture, realizes game interaction, contribution evaluation and strategy coordination among multiple agents, prompting each agent to negotiate autonomously under goal conflict, and finally converge to Nash equilibrium, realizing the leap from local optimization to global optimum, and improving the overall operational efficiency and coordination level of the airport. The causal-guided decision optimization module explicitly models the causal propagation mechanism of flight delays. By using causal graphs and counterfactual interventions to guide the strategy network to learn real causal relationships rather than spurious data correlations, it significantly improves the generalization ability of the strategy in out-of-distribution scenarios such as sudden weather and air traffic control, and solves the out-of-distribution failure problem of traditional models. The module for robust optimization and uncertainty adaptation models uncertainties such as weather forecast errors, flow control fluctuations, and equipment failures as a set of uncertainties. It optimizes the strategy performance in the worst case through adversarial training and combines meta-learning and model-driven exploration to improve the system's adaptability and robustness to dynamic disturbances.