Hybrid Control Method for Taxiing of Scene Aircraft Combining Sorting and Path Planning
By constructing a reinforced learning environment with an undirected graph structure and a multi-layer noise adaptation layer unit neural network, combined with course learning and training strategies, the aircraft's path planning and conflict management are optimized, and the problem of heavy workload of tower controllers is solved, and more efficient and safe aircraft scheduling is achieved.
Patent Information
- Application Number
- CN202411553762.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-11-03
AI Technical Summary
In the prior art, tower controllers have too heavy workloads, traditional scheduling algorithms are inefficient, deep reinforcement learning aircraft scheduling methods are difficult to train effective strategies when large action space and movement distribution are uneven, and the existing observations do not cover the overall environment, resulting in insufficient optimization of aircraft scheduling.
A hybrid control method for scene aircraft taxiing combined with sorting and path planning is constructed. By analyzing scene geographic information into an undirected graph structure, a reinforcement learning environment is constructed, and a neural network of multi-layer noise adaptation layer units is used for path planning and conflict detection. Combined with course learning and training strategies, the path selection and conflict management of aircraft are optimized.
It improves the safety and efficiency of aircraft scheduling, enhances the accuracy and adaptability of decision-making in complex environments, optimizes path planning and conflict management, and reduces the workload of tower controllers.
Smart Images

Figure CN119418558B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent air traffic control, and particularly relates to a method for mixed control of aircraft taxiing on the apron by combining sorting and path planning. Background Art
[0002] The traditional dispatching method for airport aircraft is that the tower controller guides the aircraft to take off and land safely, manages the taxiing of the aircraft on the taxiway, coordinates the intervals between flights, and handles emergencies to ensure the safety and smoothness of airport operations. However, with the continuous growth of aviation demand but the large time span of airport facility expansion, the workload intensity of tower controllers has increased, and thus the dispatching decisions of controllers for aircraft often reduce a certain efficiency.
[0003] In terms of path planning problems, traditional dispatching algorithms such as Dijkstra and heuristic algorithms have been well studied, but they often lack generality, and the problems studied are only applicable to specific scenarios. For the path planning problem of aircraft conflict resolution, the modeling is difficult and the application scenarios are highly restrictive. Deep reinforcement learning is a common machine learning method widely used to handle decision-making problems in complex states. For the existing apron aircraft auxiliary decision-making based on deep reinforcement learning, if the observation value does not cover the overall environment, due to the complexity of the design of the dispatching and path planning reward functions, the dispatching of aircraft is often not the optimal result; if the observation value covers the overall environment, it is difficult to train an effective strategy when the action space is large and the action distribution is uneven. Summary of the Invention
[0004] The present invention provides a method for mixed control of aircraft taxiing on the apron by combining sorting and path planning to assist the tower controller in decision-making in view of the problems existing in the prior art.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A method for mixed control of aircraft taxiing on the apron by combining sorting and path planning, comprising the following steps:
[0007] Step 1: Parse the apron geographical information into an undirected graph structure, and pre-calculate the apron structure information and the initial action mask set in combination with the undirected graph and the apron activity prohibition specification to construct a reinforcement learning environment;
[0008] Step 2: Obtain the real-time position information of the aircraft and construct the input observation value of the method for mixed control of aircraft taxiing on the apron in combination with the apron structure information;
[0009] Step 3: Construct a model for mixed control of aircraft taxiing on the apron;
[0010] The model for mixed control of aircraft taxiing on the apron includes a path planning module, a conflict detection module, and a conflict scheduling module;
[0011] The path planning module is implemented based on a neural network, and its structure is a neural network composed of multiple noise adaptation layer units. The path planning module is used to take the observation values obtained in step 2 as inputs, and after being processed by the neural network, obtain the action probability vector of the aircraft, and transmit the action probability vector to the conflict scheduling module.
[0012] The conflict detection module constructs the protection areas of taxiway intersections and runways as conflict domains, and calculates conflict information based on the real-time positions of aircraft to determine whether an aircraft has a conflict. The conflict detection module is used to judge whether an aircraft has a conflict and transmit the conflict information to the conflict scheduling module.
[0013] The conflict scheduling module obtains action masks from the initial action mask set according to the real-time position information of the aircraft, sorts the aircraft in conflict and updates the action masks, and finally calculates the next action selection of the aircraft from the action masks and the action probability vector. The conflict scheduling module is used to integrate the information submitted by the path planning module and the conflict detection module and calculate the next action selection of the aircraft.
[0014] Step 4: Use the training strategy of curriculum learning to train the surface aircraft taxiing hybrid control model;
[0015] Step 5: Input the observation values into the surface aircraft taxiing hybrid control model to obtain the action selection of the aircraft, and finally form the path planning trajectory of the aircraft.
[0016] Furthermore, the construction of the reinforcement learning environment in step 1 includes the following processes:
[0017] S11: Abstract the surface into a topological graph based on the taxiway and runway features of the surface environment. The intersections between taxiways and between taxiways and runways are used as the nodes of the topological graph, and the taxiways and runways are used as the edges of the topological graph. The runways and taxiways are two-way accessible, and this topological graph is an undirected graph structure;
[0018] S12: Divide the environmental action space into 9 actions, namely east, south, west, north, southeast, northeast, southwest, northwest, and stop;
[0019] S13: Generate an initial action mask set according to the surface geographical information. The action masks are represented by one-hot encoded vectors and initialized as [1,1,1,1,1,1,1,1,1]. Each position represents actions such as east, south, west, north, southeast, northeast, southwest, northwest, and stop in sequence. The 1 in the action mask indicates an available action, and 0 indicates an unavailable action. The process of generating the initial action mask set is as follows:
[0020] The initial action mask set includes action masks for edges and nodes. The executable actions for an edge include moving along a certain direction of the edge, the reverse of that direction, and stopping. Taking an edge towards the northeast as an example, the action mask is [0,0,0,0,0,1,1,0,1]. The executable actions for a node include directions to other nodes and stopping. Taking a node from which other nodes can be reached eastward and northward as an example, the action mask is [1,0,0,1,0,0,0,0,1];
[0021] S14: Calculate the scene structure information according to the scene geographical information, including: runway and taxiway length information; relative positions from node to node, position information of aircraft parking positions and both ends of the runway; takeoff or landing time information while waiting on the runway.
[0022] Furthermore, the construction of the input observation values in step 2 includes the following process:
[0023] Construct the aircraft position observation, including [x,y,x e ,y e ,x b ,y b ,t,v]. Taking the center point of the scene as the origin, the eastward direction as the positive x-axis, and the northward direction as the positive y-axis. (x,y) represents the horizontal and vertical coordinates of the aircraft, (x e ,y e ) represents the horizontal and vertical coordinates of the aircraft's end point, (x b ,y b ) represents the horizontal and vertical coordinates of the aircraft's target node, t represents the number of steps the aircraft has executed, and v represents the speed of the aircraft.
[0024] Construct the path planning observation, including the concatenation of 8 path vectors, which respectively represent the paths of the aircraft's target node in the eight directions of east, south, west, north, southeast, northeast, southwest, and northwest. Each path vector is represented as [l,s,n,x d ,y d , representing the path information adjacent to the aircraft's target node. l represents the path length, s represents the shortest distance from this path to the nearest other aircraft to the aircraft's target node, n represents the number of aircraft on this path, x d and y d represent the horizontal and vertical coordinates of another node adjacent to the aircraft's target node on this path.
[0025] Construct the conflict observation, including [t r ,n r ,r,s r . t r represents the takeoff or landing time of the aircraft while waiting on the runway, n rThe number of aircraft waiting in the same conflict as the aircraft is represented by, r represents the radius of the conflict domain, and represents the shortest distance from other waiting aircraft to the conflict node.
[0026] Finally, the aircraft position observation with dimension 8, the path planning observation spliced by 8 path vectors with dimension 5, and the conflict observation with dimension 4 are spliced to construct an input observation value with dimension 52.
[0027] Furthermore, the path planning module, conflict detection module, and conflict scheduling module in step 3 are as follows:
[0028] The path planning module is implemented based on a neural network, including the optimal path value estimation network Q1 and the optimal path auxiliary network Q2. The optimal path value estimation network Q1 outputs the aircraft action, and the optimal path auxiliary network Q2 is used to assist the training of the optimal path value estimation network Q1. Their structures and initial parameters are the same, and the structure is a neural network composed of multiple noise adaptation layer units. The input of the first layer of the Q1 and Q2 networks is the splicing of the aircraft position observation, path planning observation, and conflict observation. The output of each layer of the subsequent noise adaptation layer units is jump-connected to the aircraft position observation and used as the input of the next layer of noise adaptation layer units to further extract features. The network finally outputs an action probability vector and passes the action probability vector to the conflict scheduling module.
[0029] The conflict detection module constructs the protected areas of taxiway intersections and runways as conflict domains. The conflict domain is a circle centered on the conflict node. There are two types of aircraft conflicts: aircraft waiting for takeoff or landing on the same runway are determined to be in conflict; different aircraft whose Euclidean distance between the real-time position and the conflict node position is less than the radius of the conflict domain are determined to be in conflict. The conflict detection module is used to determine whether the aircraft is in conflict and transfer the conflict information to the conflict scheduling module.
[0030] The conflict scheduling module obtains the action mask from the initial action mask set according to the real-time position information of the aircraft. For the aircraft in conflict, the conflict scheduling module adopts the shortest time first strategy. Based on the distance to the conflict node, conflict resolution time, speed of the aircraft, and runway waiting time, with the shortest total conflict resolution time as the optimization goal, the aircraft queue at each conflict is obtained. For the aircraft in the queue, the action mask of the leading aircraft remains unchanged, and the action masks of the remaining aircraft are replaced with zero vectors. Finally, the Hadamard product of the action mask and the action probability vector is calculated to obtain the next action selection of the aircraft. The conflict scheduling module is used to calculate the next action selection of the aircraft according to the action probability vector and conflict information.
[0031] Furthermore, the training strategy of using curriculum learning to train the parameters of the path planning module in step 4;
[0032] Due to the large scope of the airport, uneven distribution of actions, and various rule constraints, the scheduling of aircraft is very complex. The course study solves the problem of scheduling in the path planning module under complex scene environments by gradually increasing the aircraft density and adding rule constraints. The course study is divided into the following two courses:
[0033] Course 1: In the first stage (1-1) of Course 1, without using the conflict detection module and the scheduling module while shielding the aircraft on adjacent taxiways, the landing aircraft is trained to reach the parking position; in the second stage (1-2) of Course 1, without using the conflict detection module and the scheduling module while shielding the aircraft on adjacent taxiways, the take-off aircraft is trained to reach the runway and take off smoothly; in the third stage (1-3) of Course 1, while shielding the conflict detection module and the scheduling module, a penalty for the safe distance between aircraft is added, and the take-off and landing aircraft are trained to reach the end point safely, and the aircraft density is gradually increased until the density upper bound.
[0034] Course 2: The conflict detection module and the conflict scheduling module are added, rules such as runway waiting information and intersection conflict areas are increased, and a penalty is added for the situation where there are other aircraft about to land or take off and enter the runway in the approach direction of the runway. The aircraft in Course 1 are trained to reach the end point safely under rule constraints, and the aircraft density is gradually increased until the density upper bound.
[0035] Furthermore, the training methods for each stage of Claim 4 are as follows:
[0036] Step 3 The path planning module is trained using the D3QN reinforcement learning method, and the training process is as follows:
[0037] Construct the optimal path value estimation network Q1 and the optimal path auxiliary network Q2; Q1 and Q2 have the same network structure. At each training step, Q1 outputs the state-action value function of the current step, and Q2 outputs the state-action value function of the next step.
[0038] The decision actions include: east, south, west, north, southeast, northeast, southwest, northwest, stop, 9 optional actions;
[0039] In Stage 1-1 and 1-2, rewards are given for the distance of the current aircraft to the end point and the current number of execution steps; in Stage 1-3, rewards are given for the distance of the current aircraft to the end point, the current number of execution steps, and the relative distance to other aircraft; in the training of Course 2, rewards are given for the distance of the current aircraft to the end point, the current number of execution steps, and the rules defined by the scene activity prohibition specifications;
[0040] Calculate the loss function based on the state-action value functions of the current step and the next step and the given rewards, and perform backpropagation of the gradient of the network Q1 to update the parameters of Q1.
[0041] Every other cycle, Q2 copies the parameters of Q1 to increase the stability of training.
[0042] Train networks Q1 and Q2 until the networks converge.
[0043] Furthermore, the neural network of the path planning module in step 3 improves the calculation of the action value function under the mask condition;
[0044] The conflict scheduling module obtains the action mask from the initial action mask set according to the real-time position information of the aircraft. For the aircraft in conflict, the conflict scheduling module adopts the shortest time first strategy. Based on the distance to the conflict node, the conflict resolution time, the speed of the aircraft, and the runway waiting time, with the goal of minimizing the total conflict resolution time, the aircraft queue at each conflict is obtained. For the aircraft in the queue, the action mask of the leading aircraft remains unchanged, and the action masks of the remaining aircraft are replaced with zero vectors. Finally, the Hadamard product of the action mask and the action probability vector is calculated to obtain the next action selection of the aircraft. The conflict scheduling module is used to calculate the next action selection of the aircraft according to the action probability vector and the conflict information. The calculation method is as follows:
[0045] A(s,a;θ)=Q π (s,a;θ)e m
[0046]
[0047] Where Q(s,a;θ) is the current state-action value function, m represents the action mask, A(s,a;θ) is the Hadamard product of the two, and the argmax function returns the index value of the maximum value in a vector.
[0048] Furthermore, the reward functions for each stage in claim 4 are as follows:
[0049]
[0050] In the formula: reward step represents the aircraft step penalty, reward away represents the aircraft moving away from the end penalty, reward completion represents the aircraft reaching the end reward, reward distance represents the aircraft distance from the end penalty, reward rule represents the surface movement prohibition norm penalty, reward air represents the relative distance penalty from other aircraft, t is the current execution step, s and sl are the Euclidean distances of the aircraft from the end at the previous step and the current step respectively, d ijIt is the sum of the Euclidean distances from aircraft i to the target node and from the nearest aircraft j on the adjacent taxiway. The ground movement prohibition norms and penalties include an aircraft about to land entering the runway in the approach direction of the runway, an aircraft on the runway taxiing entering the runway, and not maintaining a safe distance from other aircraft, etc. ω1, ω2, ω3, ω4, ω5, ω6 are all weight coefficients, b is a constant, and min is a function for obtaining the smaller value of the two.
[0051] In Course 1, the rewards for Stages 1-1 and 1-2 are reward step 、reward away 、reward completion and reward distance 。The reward for Stage 1-3 is reward step 、reward away 、reward completion 、reward distance and reward air 。In Course 2, the rewards are reward step 、reward away 、reward completion 、reward distance 、reward air and reward rule 。
[0052] A decision-making device for a ground aircraft taxiing hybrid control method combining sorting and path planning, comprising:
[0053] An information collection module: obtaining ground structure information and an initial action mask set based on ground geographical information;
[0054] An information processing module: preprocessing the ground structure information and aircraft position information;
[0055] A decision-making module: making a ground scheduling decision based on the processed information.
[0056] The beneficial effects of the present invention are:
[0057] (1) The present invention constructs a ground aircraft taxiing hybrid control model, including a path planning module, a conflict detection module, and a conflict scheduling module. The combined operation of these three modules optimizes path planning and conflict management. The path planning module dynamically optimizes the flight route through reinforcement learning, enabling the aircraft to select the best path according to real-time situations. The conflict detection module accurately identifies potential conflicts to ensure timely response of the system. The conflict scheduling module optimizes the instructions based on the conflict information to minimize the conflict resolution time and improve flight safety and efficiency. This integrated solution overall improves the safety, efficiency, and decision-making accuracy of the aviation system.
[0058] (2) The present invention trains the surface aircraft taxiing hybrid control model through a curriculum learning training strategy, gradually introducing complex rules and constraint conditions, enabling the curriculum learning method to effectively train the model to gradually improve the path planning ability in a complex aircraft scheduling environment, and enhancing its comprehensive adaptability and reliability in conflict detection, scheduling optimization, and safety distance management during actual operation.
[0059] (3) The present invention calculates the state-value function of the action mask according to the surface environment reinforcement learning architecture. By eliminating invalid actions in the mask during the calculation of the state-value function and action advantage function, the accuracy of path planning under the mask condition is improved, enabling the model to effectively exclude the interference of invalid actions on the calculation when facing them, thereby optimizing the estimation of the state-action value function. This improvement enables the model to more accurately evaluate the value of each action when dealing with complex scheduling and conflict management, thus achieving more efficient and safe aircraft taxiing control. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a schematic flowchart of the method of the present invention.
[0061] Figure 2 It is a schematic flowchart of the path planning module structure of the present invention.
[0062] Figure 3 It is a schematic flowchart of the conflict detection module structure of the present invention.
[0063] Figure 4 It is a schematic flowchart of the conflict scheduling model of the present invention.
[0064] Figure 5 It is a schematic diagram of the model training strategy of the present invention.
[0065] Figure 6 It is a schematic diagram of the structure of the decision-making device of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0066] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0067] As Figure 1 shown, a surface aircraft taxiing hybrid control method combining sorting and path planning includes the following steps:
[0068] Step 1: Parse the surface geographical information into an undirected graph structure, build a reinforcement learning environment based on the undirected graph and surface activity prohibition specifications, and pre-calculate the surface structure information and the initial action mask set;
[0069] The process of building the reinforcement learning environment and pre-calculating the surface structure information and the initial action mask set is as follows:
[0070] S11: Abstract the scene as a topological graph based on the characteristics of taxiways and runways in the scene environment. The intersections between taxiways and taxiways, and between taxiways and runways are used as the nodes of the topological graph, and the taxiways and runways are used as the edges of the topological graph. The roads of the runway and taxiway are two-way reachable, and this topological graph is an undirected graph structure;
[0071] S12: Divide the environmental action space into 9 actions, namely east, south, west, north, southeast, northeast, southwest, northwest, and stop;
[0072] S13: Generate an initial action mask set according to the scene geographical information. The action mask is represented by a one-hot encoded vector and initialized as [1,1,1,1,1,1,1,1,1]. Each position represents actions such as east, south, west, north, southeast, northeast, southwest, northwest, and stop in sequence. The 1 in the action mask indicates an available action, and 0 indicates an unavailable action. The generation process of the initial action mask set is as follows:
[0073] The initial action mask set includes action masks for edges and nodes. The executable actions for edges include along a certain direction of the edge and the reverse of that direction, as well as stop. Taking the northeast edge as an example, the action mask is [0,0,0,0,0,1,1,0,1]. The executable actions for nodes include the directions to other nodes and stop. Taking the node that can reach other nodes in the east and north directions as an example, the action mask is [1,0,0,1,0,0,0,0,1];
[0074] S14: Calculate the scene structure information according to the scene geographical information, including: the length information of the runway and taxiway; the relative position of node to node, the position information of the aircraft's parking position and the two endpoints of the runway; the takeoff or landing time information waiting on the runway.
[0075] Step 2: Obtain the real-time position information of the aircraft and construct the input observation value of the scene aircraft taxiing hybrid control method in combination with the scene structure information;
[0076] The process of constructing the input observation value of the scene aircraft taxiing hybrid control method is as follows:
[0077] Construct the aircraft position observation, including [x,y,x e ,y e ,x b ,y b ,t,v]. Taking the center point of the scene as the origin, the east direction as the positive x-axis, and the north direction as the positive y-axis. (x,y) represents the horizontal and vertical coordinates of the aircraft, (x e ,y e ) represents the horizontal and vertical coordinates of the aircraft's end point, (x b ,y b) represents the horizontal and vertical coordinates of the aircraft target node, t represents the number of steps executed by the aircraft, and v represents the aircraft speed.
[0078] Construct path planning observations, including the splicing of 8 path vectors, which respectively represent the paths in the eight directions of east, south, west, north, southeast, northeast, southwest, and northwest of the aircraft target node. Each path vector is represented as [l, s, n, x d , y d , representing the path information adjacent to the aircraft target node. l represents the path length, s represents the shortest distance from this path to the nearest other aircraft from the aircraft target node, n represents the number of aircraft existing on this path, x d and y d represent the horizontal and vertical coordinates of another node adjacent to the aircraft target node on this path.
[0079] Construct conflict observations, including [t r , n r , r, s r . t r represents the takeoff or landing time of the aircraft waiting on the runway, n r represents the number of aircraft waiting in the same conflict as the aircraft, r represents the conflict domain radius, and represents the shortest distance from other waiting aircraft to the conflict node.
[0080] Finally, splice the above aircraft position observations with dimension 8, the path planning observations spliced by 8 path vectors with dimension 5, and the conflict observations with dimension 4, so as to construct an input observation value with dimension 52.
[0081] Step 3: Construct a surface aircraft taxiing hybrid control model;
[0082] The surface aircraft taxiing intelligent path planning model includes a path planning module, a conflict detection module, and a conflict scheduling module;
[0083] The path planning module is implemented based on a neural network. First, construct an optimal path value estimation network Q1 and an optimal path auxiliary network Q2. Their structures and initial parameters are the same. The structure is a neural network composed of multiple noise adaptation layer units, as Figure 2 shown. The input of the first layer of the network is the splicing of the aircraft position observation, the path planning observation, and the conflict observation. The output of each subsequent noise adaptation layer unit is jump-connected to the aircraft position observation and used as the input of the next noise adaptation layer unit for further feature extraction. The networks Q1 and Q2 of the path planning module finally output two branches, the value function V π (s) of the current state of the aircraft, and the action value advantage function A π (s, a) corresponding to the current state of the aircraft. Calculate the advantage function A output by the neural network π(s, a) and the Hadamard product of the submission mask of the conflict scheduling module, and subtract the mean of the result vector from the result vector to obtain the effective advantage function vector The effective advantage function vector Then, sum it with the value function V of the current state π (s) to obtain the state-action value function Q(s, a; θ). The calculation method is as follows:
[0084] V π (s) = E a: π (s) [Q π (s, a)]
[0085]
[0086] In the formula, A π (s, a) represents the advantage value function, NEA represents the number of effective actions, represents the masked advantage value function, θ represents the network parameters, Q(s, a; θ) is the current state-action value function, and m represents the action mask. Mask the action set and estimate the action value advantage function A π (s, a). The neural network output eliminates the influence of invalid actions on the calculation of the action value function Q(s, a; θ), making the estimation of the current state-action value function Q(s, a; θ) more accurate. Finally, the network calculates the action probability vector and passes the action probability vector to the conflict scheduling module.
[0087] The conflict detection module's judgment process is as Figure 3 shown. It constructs the protected areas of the taxiway intersections and runways as the conflict domains, and the conflict domains are circles centered on the conflict nodes. There are two types of conflict categories for aircraft: aircraft waiting for takeoff or landing on the same runway are determined to have conflicts; different aircraft with the Euclidean distance between their real-time positions and the conflict node position less than the radius of the conflict domain are determined to have conflicts. The conflict detection module is used to judge whether an aircraft has a conflict and pass the conflict information to the conflict scheduling module.
[0088] As Figure 4As shown in the figure, the conflict scheduling module obtains the action mask from the initial action mask set according to the real-time position information of the aircraft. For the aircraft in conflict, the conflict scheduling module adopts the shortest time first strategy. Based on the distance to the conflict node, conflict resolution time, speed of the aircraft, and runway waiting time, with the goal of minimizing the total conflict resolution time, it obtains the queue of each conflicting aircraft. For the aircraft in the queue, the action mask of the leading aircraft remains unchanged, and the action masks of the remaining aircraft are replaced with zero vectors. Finally, the Hadamard product of the action mask and the action probability vector is calculated to obtain the next instruction of the aircraft. The conflict scheduling module is used to calculate the next instruction of the aircraft according to the action probability vector and conflict information. The calculation method is as follows:
[0089] A(s,a;θ)=Q π (s,a;θ)e m
[0090]
[0091] Among them, Q(s,a;θ) is the current state-action value function, m represents the action mask, A(s,a;θ) is the Hadamard product of the two, and the argmax function returns the index value of the maximum value in a vector.
[0092] Step 4: Train the surface aircraft taxiing hybrid control model using the curriculum learning training strategy;
[0093] As Figure 5 shown, the curriculum training process is as follows:
[0094] Due to the large airport area, uneven action distribution, and various rule constraints, the scheduling of aircraft is very complex. Curriculum learning solves the problem of scheduling in the complex surface environment by the path planning module, which is achieved by gradually increasing the aircraft density and adding rule constraints. Curriculum learning is divided into the following two courses:
[0095] Course 1: In the first stage (1-1) of Course 1, when shielding the aircraft on adjacent taxiways and not using the conflict detection module and scheduling module, train the landing aircraft to reach the parking position; in the second stage (1-2) of Course 1, when shielding the aircraft on adjacent taxiways and not using the conflict detection module and scheduling module, train the takeoff aircraft to reach the runway and take off smoothly; in the third stage (1-3) of Course 1, when shielding the conflict detection module and scheduling module, add the aircraft safety distance penalty, train the takeoff and landing aircraft to reach the end safely, and gradually increase the aircraft density until the density upper bound;
[0096] Course 2: Add a conflict detection module and a conflict scheduling module, add rules such as runway waiting information and intersection conflict areas, increase the penalty for an aircraft when there is another aircraft about to land or take off and enter the runway in the runway approach direction, train the aircraft in Course 1 to safely reach the end point under the rules, and gradually increase the aircraft density until the density upper bound.
[0097] The specific process is as follows:
[0098] Construct an optimal path value estimation network Q1 and an optimal path auxiliary network Q2; Q1 and Q2 have the same network structure. At each training step, Q2 outputs an action, and Q1 outputs the state value function value of the selected action.
[0099] Adjust the reward function according to the training objectives in different courses and different training stages.
[0100]
[0101] In the formula: reward step represents the aircraft step penalty, reward away represents the aircraft distance from the end point penalty, reward completion represents the aircraft reaching the end point reward, reward distance represents the aircraft distance from the end point penalty, reward rule represents the surface movement prohibition specification penalty, reward air represents the relative distance penalty from other aircraft, t is the current execution step, s and sl are the Euclidean distances of the aircraft from the end point in the previous step and the current step respectively, and d ij is the sum of the Euclidean distances from aircraft i to the nearest aircraft j on the adjacent taxiway to the target node. The surface movement prohibition specification penalty includes that there is another aircraft about to land entering the runway in the runway approach direction, another aircraft on the runway taxiing entering the runway, and not maintaining a safe distance from other aircraft, etc. ω1, ω2, ω3, ω4, ω5, ω6 are all weight coefficients, b is a constant, and min is a function to find the smaller value of the two.
[0102] In Course 1, the rewards in Stages 1-1 and 1-2 are reward step , reward away , reward completion and reward distance . The rewards in Stage 1-3 are reward step , reward away , reward completion , reward distance and reward air . In Course 2, the reward is reward step, reward away , reward completion , reward distance , reward air and reward rule .
[0103] The decision-making actions include: east, south, west, north, southeast, northeast, southwest, northwest, stop, 9 optional actions;
[0104] In each stage, the training processes of networks Q1 and Q2 are as follows:
[0105] Let the network parameters of Q1 be φ and those of Q2 be θ. The initial network parameters of Q1 and Q2 are the same. The action-value function Q(s, a, θ) of D3QN is expressed as:
[0106]
[0107] The optimization objective of the Q2 network is expressed as:
[0108]
[0109] y = r + γQ2(s', a max ; θ)
[0110] In the formula, r represents the reward for executing the current action, a max represents the optimal action in state s, γ is the discount factor, and s' represents the next state of state s.
[0111] When the D3QN network performs the i-th iteration, randomly and uniformly sample from the experience pool U(D) to update the following loss function:
[0112]
[0113] In the formula, Q1(s, a max ; φ) is the maximum value of the current state-action value function of the Q1 network, and y2 is the optimization objective of the Q2 network; s t and s t+1 are respectively the fused representations of the aircraft's own characteristics, the characteristics of the target node's adjacent taxiway, the characteristics of the aircraft on the target node's adjacent taxiway, and the conflict characteristics at the current moment t and t + 1; r is the reward obtained at time t, and a represents the decision-making action taken. After calculating the loss function, use the gradient descent method to minimize the loss function to update the network parameters, so that the advantage value function A mask (s, a) continuously approaches the optimal advantage function State value function continuously approaches the optimal state value function
[0114] Calculate the loss function according to the state-action value functions of the current step and the next step and the given reward, and perform backpropagation on the gradient of network Q1 to update the parameters of Q1.
[0115] Every other period, Q1 copies the parameters of Q2 to increase the stability of training.
[0116] Train networks Q1 and Q2 until the networks converge.
[0117] The decision-making method of the present invention constructs a topologized reinforcement learning environment based on the scene geographical information and uses a mask to constrain the action set, calculates the position observation, the path planning observation and the conflict observation, and splices the observations as the input. The input is processed by the joint work of multiple modules such as the path planning module, the conflict detection module, and the conflict scheduling module. This integrated solution optimizes path planning and conflict management as a whole. The observation value of the scene aircraft assisted decision-making based on deep reinforcement learning only covers part of the environment, making the training fast and effective; the aircraft at the conflict nodes are sorted based on the shortest time first, making the scheduling efficient and safe. The present invention can effectively assist the tower controller to schedule the aircraft on the scene and relieve the scheduling pressure.
Claims
1. A taxiing hybrid control method for scene aircraft combining sorting and path planning, characterized in that including the following steps: Step 1: Parse the scene geographical information into an undirected graph structure, and pre-calculate the scene structure information and the initial action mask set by combining the undirected graph and the scene activity prohibition specification to construct a reinforcement learning environment; Step 2: Obtain the real-time position information of the aircraft and construct the input observation value of the scene aircraft taxiing hybrid control method in combination with the scene structure information; Step 3: Construct a scene aircraft taxiing hybrid control model; The scene aircraft taxiing hybrid control model includes a path planning module, a conflict detection module, and a conflict scheduling module; The path planning module is implemented based on a neural network, and its structure is a neural network composed of multi-layer noise adaptation layer units; the path planning module is used to take the observation value obtained in Step 2 as input, and after being processed by the neural network, obtain the action probability vector of the aircraft, and transfer the action probability vector to the conflict scheduling module; The conflict detection module constructs the protection areas of taxiway intersections and runways as conflict domains, and calculates conflict information based on the real-time position of the aircraft to determine whether the aircraft has a conflict; the conflict detection module is used to judge whether the aircraft has a conflict and transfer the conflict information to the conflict scheduling module; The conflict scheduling module obtains the action mask from the initial action mask set according to the real-time position information of the aircraft, sorts the aircraft in conflict and updates the action mask, and finally calculates the next action selection of the aircraft from the action mask and the action probability vector; the conflict scheduling module is used to integrate the information submitted by the path planning module and the conflict detection module and calculate the next action selection of the aircraft; Step 4: Train the scene aircraft taxiing hybrid control model using the training strategy of curriculum learning; Step 5: Input the observation value into the scene aircraft taxiing hybrid control model to obtain the action selection of the aircraft, and finally form the path planning trajectory of the aircraft.
2. The method for hybrid control of taxiing of a scene aircraft combining sorting and path planning according to claim 1, wherein The construction of the reinforcement learning environment in Step 1 includes the following process: S11: Abstract the scene as a topological graph based on the taxiway and runway features of the scene environment, where the intersections between taxiways and taxiways, and between taxiways and runways are used as the nodes of the topological graph, and the taxiways and runways are used as the edges of the topological graph; the runway and the taxiway are two-way reachable, and this topological graph is an undirected graph structure; S12: Divide the environmental action space into 9 actions, namely east, south, west, north, southeast, northeast, southwest, northwest, and stop; S13: Generate an initial action mask set according to the scene geographical information. The action mask is represented by a one-hot encoded vector, initialized as [1,1,1,1,1,1,1,1,1], and each position represents the east, south, west, north, southeast, northeast, southwest, northwest, and stop actions in turn; the 1 in the action mask indicates that the action can be selected, and 0 indicates the action that cannot be selected; the generation process of the initial action mask set is as follows: The initial action mask set includes action masks for edges and nodes; the executable actions for an edge include moving along a certain direction of the edge and the reverse of that direction, as well as stopping; taking the edge towards the northeast as an example, the action mask is [0,0,0,0,0,1,1,0,1]; the executable actions for a node include directions to other nodes and stopping; taking the node that can reach other nodes eastward and northward as an example, the action mask is [1,0,0,1,0,0,0,0,1]; S14: Calculate the scene structure information according to the scene geographical information, including: runway and taxiway length information; relative positions from node to node, position information of aircraft parking positions and both ends of the runway; takeoff or landing waiting time information on the runway.
3. The hybrid control method for the taxiing of a scene aircraft combining sorting and path planning according to claim 1, wherein The input observation value construction in step 2 includes the following process: Construct aircraft position observation, including [x, y, x e , y e , x b , y b , t, v]; with the center point of the scene as the origin, the eastward direction as the positive direction of the x-axis, and the northward direction as the positive direction of the y-axis; (x, y) represents the horizontal and vertical coordinates of the aircraft, (x e , y e ) represents the horizontal and vertical coordinates of the end point of the aircraft, (x b , y b ) represents the horizontal and vertical coordinates of the target node of the aircraft, t represents the number of steps executed by the aircraft, and v represents the speed of the aircraft; Construct path planning observations, including the splicing of 8 path vectors, which respectively represent the paths in the eight directions of east, south, west, north, southeast, northeast, southwest and northwest of the aircraft target node; each path vector is expressed as [l, s, n, x d , y d , representing the path information adjacent to the aircraft target node; l represents the path length, s represents the nearest distance from this path to the aircraft target node to other aircraft, n represents the number of aircraft existing on this path, x d and y d represent the abscissa and ordinate of another node adjacent to the aircraft target node on this path; Conflict situation observation, including [t r , n r , r, s r ; t r represents the take-off or landing time of the aircraft waiting on the runway, n r represents the number of aircraft waiting in the same conflict as the aircraft, r represents the conflict domain radius, and s represents the shortest distance from other waiting aircraft to the conflict node; Finally, splice the aircraft position observation with a dimension of 8, the path planning observation spliced by 8 path vectors with a dimension of 5, and the conflict observation with a dimension of 4, so as to construct an input observation value with a dimension of 52.
4. A method for hybrid control of the taxiing of a scene aircraft combining sorting and path planning according to claim 1, characterized in that, The path planning module, conflict detection module, and conflict scheduling module in step 3 are as follows: The path planning module is implemented based on a neural network, including the optimal path value estimation network Q1 and the optimal path auxiliary network Q2; the optimal path value estimation network Q1 outputs the actions of the aircraft, and the optimal path auxiliary network Q2 is used to assist the training of the optimal path value estimation network Q1; their structures and initial parameters are the same, and the structures are both neural networks composed of multiple noise adaptation layer units; the input of the first layer of the Q1 and Q2 networks is the splicing of the aircraft position observation, path planning observation, and conflict observation, and the output of each layer of the subsequent noise adaptation layer units is jump-connected to the aircraft position observation and used as the input of the next layer of noise adaptation layer units to further extract features; the network finally outputs an action probability vector and passes the action probability vector to the conflict scheduling module; The conflict detection module constructs the protection areas of taxiway intersections and runways as conflict domains, and the conflict domain is a circle centered on the conflict node; there are two conflict categories for aircraft: aircraft waiting for takeoff or landing on the same runway are determined to have a conflict; different aircraft with the Euclidean distance between the real-time position of the aircraft and the position of the conflict node less than the radius of the conflict domain are determined to have a conflict; the conflict detection module is used to judge whether an aircraft has a conflict and transfer the conflict information to the conflict scheduling module; The conflict scheduling module obtains the action mask from the initial action mask set according to the real-time position information of the aircraft; For the aircraft in conflict, the conflict scheduling module adopts the shortest time first strategy. Based on the distance to the conflict node, conflict resolution time, speed of the aircraft, and runway waiting time, with the goal of minimizing the total conflict resolution time, an aircraft queue at each conflict is obtained; for the aircraft in the queue, the action mask of the head aircraft remains unchanged, and the action masks of the remaining aircraft are replaced with zero vectors; finally, the Hadamard product of the action mask and the action probability vector is calculated to obtain the next action selection of the aircraft; the conflict scheduling module is used to calculate the next action selection of the aircraft according to the action probability vector and conflict information; the calculation method is as follows: A(s,a;θ)=Q π (s,a;θ)e m Among them, Q(s,a;θ) is the current state-action value function, m represents the action mask, A(s,a;θ) is the Hadamard product of the two, and the argmax function returns the index value of the maximum value in a vector.
5. The method for taxiing hybrid control of a scene aircraft combining sorting and path planning according to claim 1, characterized in that, The training strategy for using curriculum learning to train the parameters of the path planning module in step 4: Curriculum learning solves the problem of scheduling the path planning module in complex scene environments by gradually increasing the aircraft density and adding rule constraints; Curriculum learning is divided into the following two curriculums: Curriculum 1: In the first stage (1-1) of Curriculum 1, when shielding the aircraft on adjacent taxiways and not using the conflict detection module and the scheduling module, train the landing aircraft to reach the parking position; In the second stage (1-2) of Curriculum 1, when shielding the aircraft on adjacent taxiways and not using the conflict detection module and the scheduling module, train the takeoff aircraft to reach the runway and take off smoothly; The third stage (1-3) of Curriculum 1 is to add a penalty for the safe distance of the aircraft when shielding the conflict detection module and the scheduling module, and train the takeoff and landing aircraft to reach the end safely, while gradually increasing the aircraft density until the density upper bound; Curriculum 2: Add a conflict detection module, a conflict scheduling module, increase rules such as runway waiting information and intersection conflict areas, add a penalty for other aircraft about to land or take off and enter the runway in the runway approach direction, and train the aircraft in Curriculum 1 to reach the end safely under the rule constraints, while gradually increasing the aircraft density until the density upper bound.
6. The hybrid control method for taxiing of a scene aircraft combining sorting and path planning according to claim 1, characterized in that, The training methods for each stage in claim 4 are as follows: In step 3, the path planning module is trained using the D3QN reinforcement learning method, and the training process is as follows: Construct an optimal path value estimation network Q1 and an optimal path auxiliary network Q2; Q1 and Q2 have the same network structure. At each training step, Q1 outputs the state-action value function of the current step, and Q2 outputs the state-action value function of the next step; The decision actions include: east, south, west, north, southeast, northeast, southwest, northwest, stop, 9 optional actions; In stages 1-1 and 1-2, the distance from the current aircraft to the end and the current number of execution steps are used as rewards; In stage 1-3, the distance from the current aircraft to the end, the current number of execution steps, and the relative distance from other aircraft are used as rewards; The training in Curriculum 2 combines the distance from the current aircraft to the end, the current number of execution steps, and the rules defined by the scene activity prohibition specification as rewards; Calculate the loss function according to the state-action value functions of the current step and the next step and the given rewards, and perform backpropagation of the gradient of network Q1 to update the parameters of Q1; Every other cycle, Q2 copies the parameters of Q1 to increase the stability of training; Train networks Q1 and Q2 until the networks converge.
7. A method for hybrid control of surface aircraft taxiing combining sorting and path planning according to claim 1, characterized in that, In the neural network of the path planning module in step 3, the calculation of the action value function under the mask condition is improved; The networks Q1 and Q2 of the path planning module finally output two branches: the value function V of the current state of the aircraft π (s), and the action value advantage function A of the current state of the aircraft π (s,a); calculate the Hadamard product of the advantage function A π (s,a) output by the neural network and the mask submitted by the conflict scheduling module, and subtract the mean of the result vector from the result vector to obtain the effective advantage function vector Take the effective advantage function vector And then sum it with the value function V of the current state π (s) to obtain the state-action value function Q(s,a;θ), and the calculation method is as follows: V π (s) = E a:π(s) [Q π (s,a)] Where, A π (s,a) represents the advantage value function, NEA represents the number of valid actions, represents the masked advantage value function, θ represents the network parameters, Q(s,a;θ) is the current state-action value function, m represents the action mask; masking the action space and estimating the action value advantage function A π (s,a), the neural network output eliminates the influence of invalid actions on the calculation of the action value function Q(s,a;θ), so that the current state-action value function Q(s,a; The estimation of θ is more accurate.
8. A method for hybrid control of surface aircraft taxiing combining sorting and path planning according to claim 5, characterized in that, The reward functions for each stage in claim 4 are as follows: where: reward step represents the aircraft step penalty, reward away represents the penalty for the aircraft being far from the end point, reward completion represents the reward for the aircraft reaching the end point, reward distance represents the penalty for the aircraft's distance from the end point, reward rule represents the penalty for the surface movement prohibition specification, reward air represents the penalty for the relative distance from other aircraft. t is the current execution step, s and sl are the Euclidean distances of the aircraft from the end point in the previous step and the current step respectively, and d ij is the sum of the Euclidean distances from aircraft i to the nearest aircraft j on the adjacent taxiway to the target node; the surface movement prohibition specification penalty includes that when there is an aircraft about to land in the runway approach direction and an aircraft enters the runway, when another aircraft is taxiing on the runway and an aircraft enters the runway, and when not maintaining a safe distance from other aircraft, etc. ω1, ω2, ω3, ω4, ω5, ω6 are all weight coefficients, b is a constant, and min is a function to obtain the smaller value of the two; in Course 1, the rewards for Stages 1-1 and 1-2 are reward step 、reward away 、reward completion and reward distance ; the rewards for Stage 1-3 are reward step 、reward away 、reward completion 、reward distance and reward air ; in Course 2, the rewards are reward step 、reward away 、reward completion 、reward distance 、reward air and reward rule .
9. The method for taxiing hybrid control of a scene aircraft combining sorting and path planning according to claim 5, characterized in that, The process of training networks Q1 and Q2 is as follows: Let the network parameters of Q1 be φ and the network parameters of Q2 be θ. The initial network parameters of Q1 and Q2 are the same; The action value function Q(s,a,θ) of DDQN is expressed as: The Q2 network optimization objective is expressed as: y = r + γQ2(s', a max ; θ) where r represents the reward for performing the current action, a max represents the optimal action in state s, γ is the discount factor, and s' represents the next state of state s; When the D3QN network performs the i-th iteration, randomly and uniformly sample from the experience pool U(D) to update the following loss function: where Q1(s,a max ; φ) is the maximum value of the Q1 network's current state-action value function, and y2 are the optimization objectives of the Q2 network respectively; s t and s t+1 are the aircraft's own characteristics, the characteristics of the target node's adjacent taxiways, the characteristics of the aircraft on the target node's adjacent taxiways, and the conflict characteristics at the current time t and t + 1 respectively Integrated representation; r is the reward obtained at time t, and a represents the decision-making action taken; after calculating the loss function, the gradient descent method is used to minimize the loss function to update the network parameters, so that the advantage value function A of the neural network mask (s,a) continuously approaches the optimal advantage function State value function Continuously approaches the optimal state value function 10. The decision-making device adopting a scene aircraft taxiing hybrid control method combining sorting and path planning as described in claim 1, characterized in that, Including: Information collection module: Obtain the scene structure information and the initial action mask set based on the scene geographical information; Information processing module: Preprocess the scene structure information and the aircraft position information; Decision-making module: Make a scene scheduling decision based on the processed information.
Citation Information
Patent Citations
Multi-Agent airport surface taxiing path planning method based on historical data analysis
CN113610271A
Intelligent airport sliding scheduling method based on multi-agent reinforcement learning
CN116402273A