Multi-stage complex operation layered collaborative planning method
Through a three-layer solution framework based on layered reinforcement learning, multi-cycle aircraft transportation operations are divided into process layer, apron layer and resource layer, and information exchange and decision-making coordination are carried out through inter-layer collaboration mechanisms, which solves the problems of local optimization and high computing costs of multi-cycle aircraft logistics transportation operations planning in the existing technology, achieving more efficient and global optimization planning effects.
Patent Information
- Application Number
- CN202510297034.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems such as local optimization, high computational cost, insufficient diversity of solutions, dependence on initial conditions and lack of global perspective in multi-cycle aircraft logistics and transportation operation planning.
A three-layer solution framework based on layered reinforcement learning is adopted to divide multi-cycle aircraft transportation operations into process layer, apron layer and resource layer, and information exchange and decision-making coordination are carried out through inter-layer collaboration mechanisms to achieve global optimization.
It improves the efficiency and effectiveness of multi-cycle aircraft transportation operation planning, overcomes the problems of local optimization and high computing costs, provides more diverse solutions, and enhances the adaptability and information coordination capabilities of the system.
Smart Images

Figure CN120218524A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of civil aviation aircraft logistics transportation operation planning, and particularly to a multi-stage complex operation hierarchical collaborative planning method. Background Art
[0002] The civil aviation airport aircraft logistics transportation task is a key component to ensure the punctual and safe operation of flights, representing the most large-scale and complex operation system in airport ground support. Its main task is to achieve the rapid transfer of aircraft, the rapid loading and unloading of goods and passengers through an efficient logistics support system, so as to improve the operation efficiency of the airport and make it a multi-purpose platform in air transportation. A civil aviation airport equipped with complete support equipment and a logistics team can perform efficient aircraft transportation tasks and expand the service capabilities of airlines. Therefore, in terms of improving flight punctuality and enhancing airport competitiveness, the logistics transportation system of civil aviation airports is crucial. And the flexibility and rapid response ability of the airport ground support team are the key factors to effectively improve transportation capacity and ensure aviation safety. In other words, the ground support ability of aircraft directly affects the overall operation efficiency of the airport. Similar to high-intensity training, one of the main goals of civil aviation airports and their ground support systems is to improve the departure and arrival efficiency. Among the many factors affecting the ground support ability of aircraft, the efficiency of logistics transportation operations is particularly important. Multi-cycle logistics transportation operations are usually divided into three stages: the arrival stage, the apron stage, and the departure stage. Therefore, the planning of multi-cycle aircraft logistics transportation operations can be divided into three parts: 1) arrival planning, 2) apron planning, and 3) departure planning.
[0003] Existing research on multi-cycle aircraft logistics transportation operation planning mainly focuses on two aspects: model construction and algorithm design. In terms of model construction, existing research has established scheduling models for material transfer, interference handling, ground support team deployment, apron arrangement, aircraft arrival and departure, and considered process, apron, and resource constraints. However, these studies usually only consider part of the operation process and lack a comprehensive modeling of the entire logistics transportation process. In actual aircraft scheduling and ground logistics transportation operations, the decisions at each stage are coupled and interact with each other and cannot be simply separated. In terms of algorithm design, aircraft logistics transportation operation planning essentially belongs to a multi-stage resource-constrained complex operation planning problem. Current research on solving this problem is mainly divided into three categories: optimization-based methods, heuristic-based methods, and learning-based methods.
[0004] The optimization-based method can obtain the optimal planning solution, but its high computational cost makes it difficult to be practically applied in complex multi-period aircraft logistics transportation operation planning scenarios. Although heuristic-based methods have significantly improved the efficiency in solving complex aircraft logistics transportation operation planning problems, these methods inevitably encounter problems such as local optimal solutions and insufficient diversity of solutions during the search process. In addition, the results of heuristic methods highly depend on the initial values and search strategies, usually requiring a large number of adjustments, resulting in poor interpretability of the solutions. With the development of machine learning technology, some studies have explored learning-based methods in the field of aircraft transportation operation planning and proved that they are superior to traditional methods in terms of effectiveness and efficiency.
[0005] Although existing studies have solved the planning problem of multi-period aircraft transportation operations to a certain extent, there are still the following main deficiencies:
[0006] 1. Local optimization: Existing optimization methods and heuristic methods usually can only solve local problems and are difficult to achieve global optimization. In multi-period aircraft transportation operations, the decisions at each stage are interdependent, and it is often difficult to achieve global optimality by optimizing a single stage alone;
[0007] 2. High computational cost: Although optimization-based methods can obtain the optimal solution, their computational complexity is relatively high. Especially when facing complex scenarios with multiple periods and multiple stages, the computational cost is unacceptable;
[0008] 3. Insufficient diversity of solutions: Although heuristic methods have improved the solution efficiency, they are prone to falling into local optimal solutions during the search process, lacking sufficient diversity of solutions and being difficult to find the global optimal solution;
[0009] 4. Dependence on initial conditions: The effectiveness of heuristic methods often depends on the initial conditions and search strategies, requiring a large number of parameter adjustments, increasing the application difficulty and the non-interpretability of the solutions;
[0010] 5. Lack of a global perspective: Most existing studies focus on the optimization of specific operation processes and lack a mechanism for overall planning of the entire transportation operation process from a global and long-term perspective.
[0011] Through in-depth analysis of the above deficiencies, the present invention proposes a three-layer solution framework based on hierarchical reinforcement learning, aiming to solve the planning problem of multi-period aircraft transportation operations. This framework optimizes the processes, apron, and resource allocation of multi-period aircraft transportation operations from a global and long-term perspective, overcomes the limitations of the prior art, and improves the efficiency and effectiveness of planning. Summary of the Invention
[0012] The object of the present invention is to address the above problems and provide a hierarchical collaborative planning method for multi-stage complex operations. The multi-stage complex operations are divided into a process layer, an apron layer, and a resource layer, and a three-layer solution framework based on hierarchical reinforcement learning is adopted to solve the planning problem of multi-cycle aircraft transportation operations. This framework optimizes the processes, apron, and resource allocation of multi-cycle aircraft transportation operations from a global and long-term perspective, overcomes the limitations of the prior art, and improves the efficiency and effectiveness of planning.
[0013] To achieve the above object, the technical solution of the present invention is as follows:
[0014] A hierarchical collaborative planning method for multi-stage complex operations, comprising the following steps:
[0015] Step S101: Divide the multi-stage complex operations into a process layer, an apron layer, and a resource layer, and model the planning processes of the process layer, apron layer, and resource layer as a decentralized partially observable Markov decision process (Dec-POMDP); the decisions of the process layer, apron layer, and resource layer coordinate the decision-making information among the process layer, apron layer, and resource layer through an inter-layer collaboration mechanism based on local information;
[0016] Step S102: Process layer planning, using a reinforcement learning algorithm to plan a reasonable flight schedule according to the given aircraft flight patterns and flight cycles;
[0017] Step S103: Apron layer planning, using a reinforcement learning algorithm to allocate a suitable apron for each aircraft according to the flight schedule and apron distribution;
[0018] Step S104: Resource layer planning: Using a reinforcement learning algorithm to allocate resources (such as taxiways, refueling trucks, boarding bridges, shuttle buses, etc.) for each apron according to the resource service scope;
[0019] Step S105: Output a feasible allocation plan for multi-cycle aircraft transportation operations.
[0020] As an improvement to the above technical solution, in step S101, the nine-tuple of the decentralized partially observable Markov decision process is represented as M D =<F, S, U, P, R, Ω, O, n, γ>;
[0021] where F represents a set of aircraft (i.e., agents), F = {f1, f2,..., f n}), where f represents an aircraft and n is the number of aircraft; S represents the global state space of the airport environment; U represents the action space, and each aircraft f ∈ F selects a corresponding action u f ∈ Ω at each time slice according to its local observation o f∈U; P(s′|s, u): S×U×S→[0, 1] represents the state transition function, where s∈S is the current state, s′∈S is the next state, and u∈U is the joint action of all aircraft; represents the shared reward function, which is shared by all aircraft; Ω represents the observation space; O(s, f): S×F→Ω represents the observation function, which determines the local observation o of each aircraft f in state s f ; n represents the number of aircraft; γ∈[0, 1] represents the discount factor, which is used to balance long-term and short-term rewards;
[0022] It is defined that each aircraft f has a process layer policy π p,f , apron layer policy π v,f and resource layer policy π z,f respectively in the process planning stage, apron planning stage and resource planning stage of the multi-stage complex operation; The joint policy of each layer is: Process layer joint policy π p ={π p,1 , π p,2 , …, π p,n}; Apron layer joint policy π v ={π v,1 , π v,2 , …, π v,n}; Resource layer joint policy π z ={π z,1 , π z,2 , …, π z,n};
[0023] It is defined that the time scales of the process layer, apron layer and resource layer are TP, TV and t respectively to control the decision-making frequency of each layer's policy; Then the inter-layer coordination mechanism is manifested as: After observing the initial state s0, the process layer joint policy π p selects a joint process action u p every TP = k·TV; Subsequently, the apron layer joint policy π v selects a joint apron action u v every TP = h·t; In each time slice t, the resource layer joint policy π z selects a joint resource action u z .
[0024] As an improvement to the above technical solution, all aircraft in each layer share the reward function given by the airport environment; The rewards of each layer's policy π are respectively denoted as r p , r v , r z ; The replay buffer of each layer is expressed as D p , D v and D z , and the corresponding stored experience is expressed as: Process layer experience: Apron layer experience: Resource layer experience:
[0025] As an improvement to the above technical solution, the inter-layer communication mechanism for each layer is as follows:
[0026] 1) Information exchange: At each time scale, the process layer, apron layer, and resource layer share their respective decision-making and status information to ensure the consistency and coordination of the decision-making process; 2) Reward function synchronization: Each layer shares the global reward function provided by the airport environment to ensure that the decisions of each layer are directed towards the global optimum; the reward functions r p 、r v 、r z are used by the policies of each layer and are synchronously updated at different time scales; 3) Experience replay: The experience data of each layer's policy is stored in the corresponding replay buffer for policy training.
[0027] As an improvement to the above technical solution, in the step S102, using the reinforcement learning algorithm to plan a reasonable flight plan according to the given flight mode and flight schedule of the aircraft means:
[0028] 1) State, the state s p ∈S p is represented as a quadruple (I f ,Γ e ,i w ,t), where: I f is the set of aircraft indices participating in flight plan scheduling, Γ e is the set of transportation operation types participating in flight plan scheduling, i w is the current cycle index, and t is the current time slice;
[0029] 2) Observation, the observation of each aircraft f is o p =(i f ,δ e ,i w ,t), where i f is the index of the aircraft, δ e is the transportation operation type of the aircraft to be sorted currently, i w is the current cycle index, and t is the current time slice;
[0030] 3) Action, at each time slice TP, the action u p of each aircraft = {0, 1}, where 1 means the aircraft is sorted, and 0 means the aircraft is not sorted;
[0031] 4) Reward Function. After all the aircraft in a certain period are scheduled, it obtains a process reward. where V f is the apron set selected by the aircraft according to the process. After executing the joint action u p , the total reward r p of this process policy π p is the sum of the process rewards of all aircraft, and the calculation formula is:
[0032] As an improvement to the above technical solution, the reinforcement learning algorithm described in step S102 includes a set of state value calculation and evaluation networks (Value Agent Net and Value Mixing Net) and a set of process action value calculation and evaluation networks (ProcessAgent Net and Process Mixing Net); they are learned respectively through the following two loss functions and :
[0033]
[0034] where and θ p are network parameters, are target network parameters.
[0035] As an improvement to the above technical solution, since the aircraft indices of each flight crew must be adjacent, action reorganization is performed after all the aircraft are scheduled, that is, the aircraft of the same flight crew are adjusted to the position where the first aircraft of the flight crew appears.
[0036] As an improvement to the above technical solution, in step S103, using the reinforcement learning algorithm to assign an apron to each aircraft according to the flight plan and apron distribution means:
[0037] 1) State. The state s v ∈S v is represented as a quadruple s v =(I f , Γ F , I v , t), where I f is the set of aircraft indices participating in job assignment, Γ F is the set of unfinished transportation jobs of the aircraft participating in job assignment, I v is the current apron index of the aircraft participating in job assignment (one-hot encoding), and t is the current time slice;
[0038] 2) Observation, the observation of each aircraft f is denoted as o v =(i f , Γ f , i v , t), where i f is the index of the aircraft, Γ f is the set of unfinished transportation operations of this aircraft, i v is the current apron index of this aircraft, and t is the current time slice;
[0039] 3) Action, at each time slice TV, the action of each aircraft is denoted as u v ={v1, v2,..., v n}, where v n represents the apron index selected by the aircraft;
[0040] 4) Reward Function; when aircraft f is assigned to an apron, it can obtain an apron reward where is the set of transportation operations that the aircraft can complete on the selected apron. After executing the joint action u v , the total reward r v of this apron policy π v is the sum of the apron rewards of all aircraft, and the calculation formula is: The design of this reward function can prompt the aircraft to complete as many operations as possible on an apron, thereby reducing the overall transportation operation time of the current cycle.
[0041] As an improvement to the above technical solution, the reinforcement learning algorithm described in step S103 includes a group of apron action value calculation and evaluation networks (Station Agent Net and Station Mixing Net), and learns through the following loss function :
[0042]
[0043] where, θ v is the network parameter, is the target network parameter.
[0044] As an improvement to the above technical solution, in step S104, allocating resources to each apron according to the resource service scope through the reinforcement learning algorithm means;
[0045] 1) State, the state s z ∈S z is represented as a five-tuple s z =(If , I v , δ F , I z , t), where I f is the set of aircraft indices participating in job allocation, I v is the current apron index of the aircraft participating in job allocation, δ F is the set of support resources currently in need of resources, I z is the set of support resource indices that can provide services for transportation operation δ F , and t is the current time slice;
[0046] 2) Observation, the observation of each aircraft f is represented as o z =(i f , i v , δ f , I z , t), where i f is the index of the aircraft, i v is the current apron index of the aircraft, δ f is the current operation in need of resources, I z is the set of support resource indices that can provide services for transportation operation δ f , and t is the current time slice;
[0047] 3) Action, at each time slice t, the action of each aircraft is represented as u z ={z1, z2,..., z n}, where z n represents the support resource index selected by the aircraft for the current support apron;
[0048] 4) Reward Function, when aircraft f allocates resources to the support apron where transportation operation δ is located, it can obtain a resource reward where t wait and t move are the waiting time and moving time from the end of the previous transportation operation to the start of the current transportation operation respectively;
[0049] After executing the joint action u z , the total reward r z of this resource strategy π z is the sum of the resource rewards of all aircraft, and the calculation formula is:
[0050] As an improvement to the above technical solution, the reinforcement learning algorithm package in step S104 includes a group of resource action value calculation and evaluation networks (Resource Agent Net and Resource Mixing Net). Learning can be performed through the following loss function as follows:
[0051]
[0052] where θ z is the network parameter, and is the target network parameter.
[0053] As an improvement to the above technical solution, in step S105, outputting a feasible multi-period aircraft transportation operation allocation plan means that:
[0054] For each aircraft in each period, first use the process layer network to select a suitable process layer action according to the observed global state and local observation;
[0055] If the action of the aircraft is to perform choreography (i.e., ), then use the apron layer network and the process layer network to allocate suitable aprons and resources for each of its transportation operations;
[0056] Finally, find a set of feasible aircraft-apron-resource matching pairs and add them to the allocation plan; the specific process is as follows:
[0057] Repeat until the time slice t ends:
[0058] For each flight cycle w, for each aircraft f in each flight cycle w:
[0059] Phase 1: Process planning
[0060] 1. Obtain the current process step
[0061] 2. Obtain the global state and the local observation
[0062] 3. Select the process action according to the ε-greedy process layer policy π p Select the process action
[0063] Obtain the reward The next state and the next local observation
[0064] 4. Store the process layer experience in the experience pool D p ;
[0065] 5. If then perform the following operations:
[0066] 6. Obtain the current cycle index w.i;
[0067] Phase 2: Apron Planning
[0068] For each transportation operation δ of aircraft f:
[0069] 1. Obtain the current apron step
[0070] 2. Obtain the global state and local observation
[0071] 3. Select an apron action according to the ε-greedy apron layer policy π v
[0072] 4. Obtain the reward next state and next local observation
[0073] 5. Store the experience in the experience pool D v ;
[0074] Phase 3: Resource Planning
[0075] 6. Obtain the global state and local observation
[0076] 7. Select a resource action according to the ε-greedy resource layer policy π z
[0077] 8. Obtain the reward next state and next local observation
[0078] 9. Store the experience in the experience pool D z ;
[0079] 10. Obtain the start time t of transportation operation δ s ←t;;
[0080] 11. Obtain the end time t of transportation operation δ e ←t + δ.t;
[0081] 12. Obtain the match p←(w.i, f, δ, j, u v , u z , ts , t e );
[0082] 13. Add the matching p to the allocation plan P;
[0083] Return the allocation plan P.
[0084] Compared with the prior art, the present invention includes but is not limited to the following advantages and positive effects:
[0085] 1. Enhanced local optimization ability. Through the hierarchical reinforcement learning method, the complex multi-cycle aircraft transportation operation is decomposed into three sub-problems: the process layer, the apron layer, and the resource layer, and optimized separately. This hierarchical design not only considers the interdependence between layers but also realizes efficient information exchange and global optimization through the inter-layer communication mechanism, improving the overall planning effect.
[0086] 2. Improved computational efficiency. Although the existing optimization methods can obtain the optimal solution, their computational complexity is high, making it difficult to meet real-time requirements in practical applications. In contrast, the present invention adopts an algorithm based on reinforcement learning, which not only ensures efficient solution but also significantly reduces the computational cost, and can quickly obtain high-quality solutions in complex multi-cycle aircraft planning scenarios.
[0087] 3. Enhanced diversity and robustness of solutions. The heuristic method is prone to falling into local optimal solutions during the solution process, while the present invention can better explore the solution space through the reinforcement learning algorithm, provide more diverse solutions, and avoid the local optimal problem to a certain extent. The introduction of multi-agent reinforcement learning also enhances the adaptability and robustness of the system to environmental changes.
[0088] 4. Improved scalability and adaptability of the system. The hierarchical framework and the reinforcement learning-based method of the present invention have good scalability and can adapt to different combat tasks and scenarios. By adjusting the planning algorithms at each layer, it can flexibly cope with changes such as different aircraft flight patterns, aviation support cycles, and resource constraints, and has strong adaptability.
[0089] 5. Enhanced information coordination and conflict resolution ability. By designing the inter-layer communication mechanism, the present invention can achieve efficient information exchange and decision coordination between layers, avoiding planning conflicts caused by information asymmetry. This mechanism helps to optimize the planning results from a global and long-term perspective and further improves the overall performance of the system.
[0090] In summary, through the innovative hierarchical reinforcement learning method and efficient inter-layer communication mechanism, the present invention overcomes the defects in the prior art such as local optimization, high computational cost, insufficient solution diversity, dependence on initial conditions, and lack of global perspective, significantly improves the planning efficiency and effect of multi-period aircraft transportation operations, and has important practical application value and promotion prospects. Brief Description of the Drawings
[0091] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0092] Figure 1 is the overall architecture diagram of the present invention;
[0093] Figure 2 is the framework diagram of the inter-layer cooperation mechanism of the present invention;
[0094] Figure 3 is the process layer network structure of the rack. Detailed Embodiments
[0095] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts, any modifications, equivalent replacements, improvements, etc., shall be included in the protection scope of the present invention.
[0096] Figure 1 shows the overall architecture diagram of the multi-period aircraft transportation operation planning based on hierarchical reinforcement learning proposed by the present invention. As Figure 1 can be seen, the multi-stage complex operation hierarchical cooperation planning method of the present invention includes three levels: the process layer, the apron layer, and the resource layer. Each level conducts planning and decision-making respectively, and coordinates information through the inter-layer cooperation mechanism, so as to realize the global optimization of the multi-period aircraft logistics transportation operation scheduling plan.
[0097] Includes the following steps:
[0098] Step S101: Decentralized Partially Observable Markov Decision Process (Dec-POMDP) Modeling: Model the above three-level planning process as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP). In this model, the decision-making at each level is based on local information, but the decision-making information between different levels is coordinated through an inter-level cooperation mechanism to achieve global optimization;
[0099] Step S102: Process Layer Planning: Use a reinforcement learning algorithm to plan a reasonable flight schedule according to the given aircraft flight patterns and flight cycles;
[0100] Step S103: Apron Layer Planning: Use a reinforcement learning algorithm to allocate a suitable apron for each aircraft according to the flight schedule and apron distribution;
[0101] Step S104: Resource Layer Planning: Allocate resources (such as taxiways, refueling trucks, boarding bridges, shuttle buses, etc.) to each apron according to the resource service scope through a reinforcement learning algorithm.
[0102] Step S105: Output a feasible multi-period aircraft transportation operation allocation plan.
[0103] In Step S101, the multi-period aircraft transportation operation planning process is modeled as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP). The specific modeling is as follows: A Dec-POMDP can be represented as a nine-tuple M D = <F, S, U, P, R, Ω, o, n, γ>, where:
[0104] F represents a set of aircraft (i.e., agents), F = {f1, f2,..., f n}, where f represents an aircraft and n is the number of aircraft.
[0105] S represents the global state space of the airport environment.
[0106] U represents the action space. Each aircraft f ∈ F selects the corresponding action u f ∈ Ω at each time slice according to its local observation o f ∈ U.
[0107] P(s′|s, u): S × U × S → [0, 1] represents the state transition function, where s ∈ S is the current state, s′ ∈ S is the next state, and u ∈ U is the joint action of all aircraft.
[0108] Represents the shared reward function, which is shared by all aircraft.
[0109] Ω represents the observation space.
[0110] O(s, f): S×F→Ω represents the observation function, which determines the local observation o of each aircraft f at state s f 。
[0111] n represents the number of aircraft.
[0112] γ∈[0, 1] represents the discount factor, which is used to balance long-term and short-term rewards.
[0113] The present invention abstracts the complex multi-period aircraft transportation operation planning into a process layer, an apron layer, and a resource layer according to the aircraft support business process, and designs a three-layer solution framework based on Hierarchical Reinforcement Learning (HRL) to solve the Dec-POMDP problem. Specifically, each aircraft f has a process layer policy π p,f 、apron layer policy π v,f and resource layer policy π z,f respectively in the process planning stage, apron planning stage, and resource planning stage of the multi-period transportation operation planning problem. When a new flight task arrives, a flight plan is first formulated, which describes the arrangement of aircraft by cycle and gives the expected start / end times for aircraft to perform outbound, apron support, and return operations in each cycle. The specific steps are as follows: 1) Determine the flight mode (i.e., short / long-haul flight, day / night flight, crew structure and number) and cycle pattern (i.e., number of takeoffs and landings and interval time), and use the process layer policy π p,f to generate a "multi-crew - multi-period" flight plan. 2) According to the cycle start / end times in the flight plan, use the apron layer policy π v,f and resource layer policy π z,f to allocate apron and resources for the aircraft, and then complete the transportation operation, thus solving the multi-period aircraft transportation operation planning problem.
[0114] The inter-layer cooperation mechanism framework diagram is as shown in Figure 2 . Specifically, the joint policies of each layer are expressed as follows: Process layer joint policy: π p ={π p,1 , π p,2 , …, π p,n}; Apron layer joint policy: π v ={π v,1 , π v,2 , …, π v,n}; Resource layer joint policy: π z ={π z,1 , π z,2, …, π z,n}。
[0115] In the three - layer solution framework, the present invention designs three time scales to control the decision - making frequency of each layer's strategy. The time scales of the process layer, apron layer, and resource layer are defined as TP, TV, and t respectively. The specific operations are as follows:
[0116] After observing the initial state s0, the process - layer joint strategy π p selects a joint process action u every TP = k·TV p 。
[0117] Subsequently, the apron - layer joint strategy π v selects a joint apron action u every TP = h·t v 。
[0118] In each time slice t, the resource - layer joint strategy π z selects a joint resource action u z 。
[0119] In the Dec - POMDP model of the present invention, all the aircraft in each layer share the reward function given by the airport environment. The rewards of each layer's strategy π are denoted as r p 、r v 、r z 。In addition, the present invention represents the replay buffers of the three - layer strategies as D p 、D v and D z respectively. Their corresponding stored experiences can be expressed as:
[0120] Process - layer experience:
[0121] Apron - layer experience:
[0122] Resource - layer experience:
[0123] To coordinate the decision - making information between different layers, the present invention designs an inter - layer communication mechanism to further optimize the global and long - term planning results. The specific steps are as follows:
[0124] 1) Information exchange:
[0125] At each time scale, the process layer, apron layer, and resource layer share their respective decision - making and state information.
[0126] Each layer exchanges necessary information through communication to ensure the consistency and coordination of the decision - making process.
[0127] 2) Reward - function synchronization:
[0128] Each level shares the global reward function provided by the airport environment to ensure that the decisions of each level are directed towards the global optimum.
[0129] Reward function r p 、r v 、r z is used by the policies of each level and is synchronously updated on different time scales.
[0130] 3) Experience replay:
[0131] The experience data of each level's policy is stored in the corresponding replay buffer.
[0132] The experience data in the replay buffer is used for policy training to ensure that the policy can effectively learn and improve at different levels.
[0133] Through the above design, the present invention realizes the efficient hierarchical collaborative planning of multi-cycle aircraft transportation operations, improving the decision-making efficiency and effectiveness of the overall system.
[0134] In step S102, a reasonable flight plan is planned according to the flight mode and flight cycle of the given aircraft by using a reinforcement learning algorithm. Specifically:
[0135] 1) State
[0136] State s p ∈S p is represented as a quadruple (I f ,Γ e ,i w ,t), where:
[0137] I f is the set of aircraft indices participating in flight plan scheduling,
[0138] Γ e is the set of transportation operation types participating in flight plan scheduling,
[0139] i w is the current cycle index,
[0140] t is the current time slice.
[0141] 2) Observation
[0142] The observation of each aircraft f is o p =(i f ,δ e ,i w ,t), where:
[0143] i f is the index of the aircraft,
[0144] δ e is the transportation operation type of the aircraft to be sorted currently,
[0145] i w is the current cycle index,
[0146] t is the current time slice.
[0147] 3) Action
[0148] At each time slice TP, the action u of each aircraft p = {0, 1}, where:
[0149] 1 means the aircraft is sorted,
[0150] 0 means the aircraft is not sorted.
[0151] Note: The aircraft indices of each crew are usually adjacent. Therefore, after all aircraft are arranged, action reorganization needs to be performed, that is, the aircraft of the same crew are adjusted to the position where the first aircraft of the crew appears. For example, the joint action The joint action after action reorganization is where represents the j-th aircraft of the i-th crew.
[0152] 4) Reward Function
[0153] When all aircraft in a certain cycle are arranged, it can obtain a process reward where V f is the apron set selected by the aircraft according to the process. After executing the joint action u p the total reward r of this process policy π p is the sum of the process rewards of all aircraft, and the calculation formula is: p In summary, through defining the state, observation, action and reward function, the process layer planning realizes the reasonable arrangement of multi-cycle aircraft transportation operations. The process layer network structure is as shown in
[0154] S301 of. It includes a group of state value calculation and evaluation networks (ValueAgentNet and ValueMixingNet) and a group of process action value calculation and evaluation networks (ProcessAgentNet and ProcessMixingNet). It can be learned respectively through the following two loss functions Figure 3 and and as follows:
[0155]
[0156] Among them, and θ p are network parameters, and are target network parameters.
[0157] In step S103, a reinforcement learning algorithm is used to allocate aprons for each aircraft according to the flight plan and apron distribution. Specifically:
[0158] State
[0159] State s v ∈ S v is represented as a quadruple s v =(I f , Γ F , I v , t), where:
[0160] I f is the set of aircraft indices participating in job allocation,
[0161] Γ F is the set of unfinished transportation jobs of the aircraft participating in job allocation,
[0162] I v is the current apron index (one-hot encoded) of the aircraft participating in job allocation,
[0163] t is the current time slice.
[0164] Observation
[0165] The observation of each aircraft f is represented as o v =(i f , Γ f , i v , t), where:
[0166] i f is the index of the aircraft,
[0167] Γ f is the set of unfinished transportation jobs of this aircraft,
[0168] i v is the current apron index of this aircraft,
[0169] t is the current time slice.
[0170] Action
[0171] At each time slice TV, the action of each aircraft is represented as u v ={v1, v2,..., v n}, where v nIndicates the guaranteed apron index selected by the aircraft.
[0172] Reward Function
[0173] When aircraft f is assigned to an apron, it can obtain an apron reward where is the set of transportation operations that can be completed by the aircraft on the selected apron. After executing the joint action u v the total reward r v of this apron policy π v is the sum of all aircraft apron rewards, and the calculation formula is: The design of this reward function can prompt the aircraft to complete as many operations as possible on an apron, thereby reducing the overall transportation operation time of the current cycle. The apron layer network structure is as shown in Figure 3 S302. It contains a set of apron action value calculation and evaluation networks (StationAgentNet and StationMixingNet). It can be learned through the following loss function as follows:
[0174]
[0175] where θ v is the network parameter, and
[0176] is the target network parameter. In step S104, resources are allocated to each apron according to the resource service scope through a reinforcement learning algorithm. Specifically:
[0177] State
[0178] State s z ∈S z is represented as a five-tuple s z =(I f , I v , δ F , I z , t), where:
[0179] I f is the set of aircraft indices participating in job assignment,
[0180] I v is the current apron index of the aircraft participating in job assignment,
[0181] δ F is the set of guaranteed resources currently in need of resources,
[0182] I z is the set of resources that can provide transportation operations for δ FIndex set of guarantee resources for providing services
[0183] t is the current time slice
[0184] Observation
[0185] The observation of each aircraft f is denoted as o z =(i f , i v , δ f , I z , t), where
[0186] i f is the index of the aircraft
[0187] i v is the current apron index of the aircraft
[0188] δ f is the operation that currently requires resources
[0189] I z is the index set of guarantee resources that can provide services for the transportation operation δ f t is the current time slice
[0190] t is the current time slice
[0191] Action
[0192] At each time slice t, the action of each aircraft is denoted as u z ={z1, z2,..., z n}, where z n represents the index of the guarantee resource selected by the aircraft for the current guarantee apron
[0193] Reward Function
[0194] When aircraft f allocates resources to the guarantee apron where the transportation operation δ is located, it can obtain a resource reward where t wait and t move are the waiting time and moving time from the end of the previous transportation operation to the start of the current transportation operation respectively. After executing the joint action u z , the total reward r z of this resource strategy π z is the sum of the resource rewards of all aircraft, and the calculation formula is The design of this reward function can prompt the aircraft to complete resource allocation as soon as possible, thereby reducing the waiting and moving time and improving the overall guarantee efficiency of the current cycle. The resource layer network structure is as Figure 3As shown in S303. It includes a set of resource action value calculation and evaluation networks (ResourceAgentNet and ResourceMixingNet). It can be learned through the following loss function as follows:
[0195]
[0196] where θ z is the network parameter, and is the target network parameter.
[0197] In step S105, a feasible multi-period aircraft transportation operation assignment plan is output. The following algorithm describes the main idea of using a three-layer planning network for transportation operation assignment. For each aircraft in each period, first, the process layer network selects a suitable process layer action according to the observed global state and local observation. If the action of the aircraft is to perform choreography (i.e., ), then the apron layer network and the resource layer network are used to allocate suitable aprons and resources for each of its transportation operations. Finally, a set of feasible aircraft-apron-resource matching pairs are found and added to the assignment plan. The specific process is as follows:
[0198] Repeat until the end of time slice t:
[0199] For each flight cycle w:
[0200] For each aircraft f in the cycle:
[0201] Phase 1: Process planning
[0202] Obtain the current process step
[0203] Obtain the global state and local observation
[0204] According to the ε-greedy process layer policy π p Select the process action
[0205] Obtain the reward Next state and next local observation
[0206] Store the process layer experience in the experience pool D p ;
[0207] If then perform the following operations:
[0208] Obtain the current cycle index w.i;
[0209] Phase 2: Apron Planning
[0210] For each transportation operation δ of aircraft f:
[0211] Obtain the current apron step
[0212] Obtain the global state and the local observation
[0213] According to the ε-greedy apron layer policy π v Select an apron action
[0214] Obtain the reward Next state and the next local observation
[0215] Store the experience in the experience pool D v ;
[0216] Phase 3: Resource Planning
[0217] Obtain the global state and the local observation
[0218] According to the ε-greedy resource layer policy π z Select a resource action
[0219] Obtain the reward Next state and the next local observation
[0220] Store the experience in the experience pool D z ;
[0221] Obtain the start time t of transportation operation δ s ←t;
[0222] Obtain the end time t of transportation operation δ e ←t + δ.t;
[0223] Obtain the match p ← (w.i, f, δ, j, u v ,u z ,t s ,t e );
[0224] Add the match p to the allocation plan P.
[0225] Return the allocation plan P.
[0226] Through the innovative hierarchical reinforcement learning method and efficient inter-layer communication mechanism, the present invention overcomes the defects in the prior art such as local optimization, high computational cost, insufficient diversity of solutions, dependence on initial conditions, and lack of global perspective, significantly improves the planning efficiency and effect of multi-cycle aircraft transportation operations, and has important practical application value and popularization prospects.
Claims
1. A multi-stage complex operation hierarchical collaborative planning method, characterized by: The steps include: Step S101: Divide the multi-stage complex operation into a process layer, an apron layer and a resource layer, and model the planning process of the process layer, the apron layer and the resource layer as a decentralized partially observable Markov decision process; the decision of the process layer, the apron layer and the resource layer coordinates the decision information between the process layer, the apron layer and the resource layer through an inter-layer coordination mechanism based on local information; Step S102: Process-level planning, using reinforcement learning algorithms to plan a reasonable flight plan based on a given aircraft flight mode and flight cycle; Step S103: apron layer planning, using reinforcement learning algorithm to allocate appropriate aprons for each aircraft according to flight plans and apron distribution; Step S104: Resource layer planning: Allocate resources to each apron according to the resource service range through a reinforcement learning algorithm; Step S105: Output a feasible multi-period aircraft transportation operation allocation plan.
2. The multi-stage complex operation hierarchical collaborative planning method according to claim 1, characterized in that: In step S101, the nine-tuple of the decentralized partially observable Markov decision process is represented by M D =<F,S,U,P,R,Ω,O,n,γ> ; Where F represents a group of aircraft, F = {f1, f2, ..., f n }, where f represents an aircraft, n is the number of aircraft; S represents the global state space of the airport environment; U represents the action space, and each aircraft f∈F moves according to its local observation o in each time slice. f ∈Ω selects the corresponding action u f ∈U; P(s′|s,u): S×U×S→[0,1] represents the state transfer function, where s∈S is the current state, s′∈S is the next state, and u∈U is the joint action of all aircraft; represents a shared reward function, which is shared by all aircraft; Ω represents the observation space; O(s, f): S×F→Ω represents the observation function, which determines the local observation o of each aircraft f in state s. f ; n represents the number of aircraft; γ∈[0,1] represents the discount factor, which is used to balance long-term and short-term rewards; Define each aircraft f to have a process-level strategy π in the process planning stage, apron planning stage, and resource planning stage of a multi-stage complex operation p,f 、Ramp layer strategy π v,f and resource layer strategy π z,f ; The joint strategies of each layer are: Process layer joint strategy π p ={π p,1 , π p,2 , …, π p,n }; Aerodrome layer joint strategy π v ={π v,1 , π v,2 , …, π v,n }; Resource layer joint strategy π z ={π z,1 , π z,2 , …, π z,n }; Define the time scales of the process layer, apron layer and resource layer as TP, TV and t respectively to control the decision frequency of each layer strategy; then the inter-layer coordination mechanism is as follows: after the initial state s0 is observed, the process layer joint strategy π p Each TP = k·TV selects a joint process action u p ; Then, the ramp layer joint strategy π v Select a joint ramp action u for each TP = h·t v ; At each time slice t, the resource layer joint strategy π z Select a joint resource action u z ; All aircraft in each layer share the reward function given by the airport environment; the rewards of each layer strategy π are recorded as r p 、r v 、r z ; The playback buffer of each layer is represented by D p , D v and D z , the corresponding storage experience is expressed as: Process layer experience: Ramp Level Experience: Resource layer experience:
3. The multi-stage complex operation hierarchical collaborative planning method according to claim 2, characterized in that: The inter-layer communication mechanism of each layer is: 1) Information exchange: At each time scale, the process layer, apron layer, and resource layer share their respective decision and status information to ensure consistency and coordination of the decision-making process; 2) Reward function synchronization: Each level shares the global reward function provided by the airport environment to ensure that the decision-making of each level is made in the direction of global optimization; the reward function r p 、r v 、r z Used by strategies at all levels and updated synchronously at different time scales; 3) Experience replay: The experience data of each layer of strategy is stored in the corresponding replay buffer for strategy training.
4. The multi-stage complex operation hierarchical collaborative planning method according to claim 2, characterized in that: In step S102, using the reinforcement learning algorithm to plan a reasonable flight plan according to the given flight mode and flight cycle of the aircraft means: 1) State, state s p ∈S p Represented as a four-tuple (I f , Γ e ,i w , t), where: I f is the set of aircraft indexes participating in flight planning, Γ e is the set of transport operation types involved in flight planning, i w is the current cycle index, and t is the current time slice; 2) Observation, the observation of each aircraft f is o p =(i f , δ e ,i w , t), where i f is the index of the aircraft, δ e is the transport operation type of the current aircraft to be sorted, i w is the current cycle index, and t is the current time slice; 3) Action: At each time slice TP, the action u of each aircraft p ={0,1}, where 1 means the aircraft are sorted and 0 means the aircraft are not sorted; 4) Reward function: When all aircraft in a cycle are arranged, it receives a process reward Where V f It is the set of aprons selected by the aircraft according to the process; In performing the joint action u p After that, the process strategy π p The total reward r p It is the sum of all aircraft process rewards, calculated as:
5. The multi-stage complex operation hierarchical collaborative planning method according to claim 4, characterized in that: The reinforcement learning algorithm in step S102 includes a set of state value calculation evaluation networks and a set of process action value calculation evaluation networks; respectively, through the following two loss functions and To learn: in, and θ p are network parameters, are the target network parameters.
6. The multi-stage complex operation hierarchical collaborative planning method according to claim 2, characterized in that: In step S103, using a reinforcement learning algorithm to allocate an apron to each aircraft according to the flight plan and apron distribution means: 1) State, state s v ∈S v Represented as a four-tuple s v =(I f , Γ F , I v , t), where I f is the set of aircraft indexes participating in the job allocation, Γ F is the set of unfinished transport operations of the aircraft participating in the operation allocation, I v is the current apron index of the aircraft participating in the job allocation, and t is the current time slice; 2) Observation, the observation of each aircraft f is represented by o v =(i f , Γ f ,i v , t), where i f is the index of the aircraft, Γ f is the set of unfinished transport operations of the aircraft, i v is the current ramp index of the aircraft, and t is the current time slice; 3) Action, in each time slice TV, the action of each aircraft is represented by u v = {v1, v2, ..., v n }, where v n Indicates the support apron index selected by the aircraft; 4) Reward function: When aircraft f is assigned to a ramp, it receives a ramp reward in It is the collection of transport operations completed by the aircraft on the selected apron; when performing the joint action u v After that, the apron strategy π v The total reward r v It is the sum of all aircraft ramp rewards, calculated as:
7. The multi-stage complex operation hierarchical collaborative planning method according to claim 6, characterized in that: The reinforcement learning algorithm in step S103 includes a set of apron action value calculation and evaluation networks, which are implemented through the following loss function To learn: Among them, θ v are network parameters, are the target network parameters.
8. The multi-stage complex operation hierarchical collaborative planning method according to claim 2, characterized in that: The step S104, allocating resources to each apron according to the resource service range by using a reinforcement learning algorithm, refers to; 1) State, state s z ∈S z Represented as a five-tuple s z =(I f , I v , δ F , I z , t), where I f is the index set of aircraft participating in the job allocation, I v is the current apron index of the aircraft participating in the job allocation, δ F is the set of guaranteed resources that currently require resources, I z Is capable of transport operationsδ F The index set of guaranteed resources providing services, t is the current time slice; 2) Observation, the observation of each aircraft f is represented by o z =(i f ,i v , δ f , I z , t), where i f is the index of the aircraft, i v is the current ramp index of the aircraft, δ f is the operation that currently requires resources, I z Is capable of transport operationsδ f The index set of guaranteed resources providing services, t is the current time slice; 3) Action, at each time slice t, the action of each aircraft is represented by u z ={z1, z2, ..., z n }, where z n Indicates the support resource index selected by the aircraft for the current support apron; 4) Reward function: When aircraft f allocates resources to the support apron where the transport operation δ is located, it obtains a resource reward where t wait and t move They are the waiting time and the moving time from the end of the previous transport operation to the start of the current transport operation; In performing the joint action u z After that, the resource strategy π z The total reward r z It is the sum of all aircraft resource rewards, calculated as:
9. The multi-stage complex operation hierarchical collaborative planning method according to claim 8, characterized in that: Step S104: The reinforcement learning algorithm includes a set of resource action value calculation evaluation networks; through the following loss function To learn: Among them, θ z are network parameters, is the target network parameter.
10. The multi-stage complex operation hierarchical collaborative planning method according to claim 2, characterized in that: The step S105, outputting a feasible multi-period aircraft transport operation allocation plan, refers to: For each aircraft in each cycle, a process-layer network is first used to select a suitable process-layer action based on the observed global state and local observations; If the aircraft's action is to perform a choreography, The apron layer network and process layer network are used to allocate appropriate aprons and resources to each transportation operation; Finally, a set of feasible aircraft-airport-resource matching pairs is found and added to the allocation plan; the specific process is: Repeat until time slice t ends: For each flight cycle w, for each aircraft f in each flight cycle w: Phase 1: Process Planning 1. Get the current process step 2. Get the global status and local observation 3. According to the ε-greedy process layer strategy π p Select process action Get rewards Next state and the next local observation 4. Process-level experience Stored in experience pool D p middle; 5. If Then do the following:
6. Get the current cycle index wi; Phase 2: Apron Planning For each transport operation δ of aircraft f:
1. Get the current apron step 2. Get the global status and local observation 3. According to the ε-greedy apron layer strategy π v Select the apron action 4. Get rewards Next state and the next local observation 5. Experience Stored in experience pool D v middle; Phase 3: Resource Planning 6. Get the global status and local observation 7. According to the ε-greedy resource layer strategy π z Select Resource Action 8. Get rewards Next state and the next local observation 9. Experience Stored in experience pool D z middle; 10. Get the start time t of the transport operation δ s ←t; 11. Get the end time t of the transport operation δ e ←t+δ.t; 12. Get the matching p←(wi,f,δ,j,u v ,u z , t s , t e ); 13. Add match p to allocation plan P; Returns the allocation plan P.