Aviation trajectory semi-supervised clustering method based on Transform noise reduction auto-encoder

Through the semi-supervised clustering method of aviation trajectory based on Transformer noise reduction autoencoder, the problem of track clustering in the airport terminal area is solved, efficient and accurate clustering of tracks is achieved, and operation efficiency and safety are improved.

CN120256992APending Publication Date: 2025-07-04QINGDAO CIVIL AVIATION AIR TRAFFIC CONTROL IND DEV CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510328287.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing technology lacks unified quantitative evaluation standards and performance benchmarks, which makes it difficult to achieve cluster analysis of aircraft entry and exit tracks in the airport terminal area, affecting operational efficiency and safety.

Method used

The semi-supervised clustering method of aviation trajectory based on Transformer noise reduction autoencoder is adopted. By constructing key indicators for quantitative evaluation of exit track clustering, the noise reduction autoencoder and multi-layer perceptron model of Transformer architecture are used, and combined with the fine-grained semi-supervised clustering algorithm, the track features are extracted and clustered.

Benefits of technology

It realizes efficient and accurate clustering of the flight tracks, improves the operating efficiency and safety of the airport terminal area, is widely applicable, does not require additional data assistance, and has fast calculations and high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256992A_ABST
    Figure CN120256992A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of civil aviation data analysis, and particularly relates to an aviation trajectory semi-supervised clustering method based on a Transform noise reduction auto-encoder. Comprising the following steps: S100, preprocessing a flight arrival and departure trajectory data set; s200, data set feature extraction; s300, constructing a model and training the model; s400, carrying out a semi-supervised clustering algorithm; and S500, training the model constructed in the step S300 to obtain a new test set, extracting the data features of the entry and departure and runway numbers according to the step S200, inputting the track data into the model, and clustering the output noise reduction track by a semi-supervised clustering algorithm to obtain a clustering result. According to the method, data mining is carried out on the flight path data of the airspace of the terminal area, a flight path clustering model is constructed, the mode category to which the flight path belongs is identified, and then reference is provided for fine safe operation of the airport terminal area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of civil aviation data analysis, and particularly relates to an aviation trajectory semi-supervised clustering method based on a Transformer denoising autoencoder. Background Art

[0002] The airport terminal area is an air traffic control area established over one or several adjacent busy airports. It is a key area for the safe, efficient, and orderly operation of aircraft during arrival and departure. It is also the air traffic service airspace with the highest flight density, the most complex airspace environment, and the greatest difficulty in air traffic control and command. In the airport terminal area, the arrival and departure routes are densely intertwined, and there are potential flight conflict risks. The takeoff and landing flights of aircraft are easily affected by complex meteorological and other airspace environments, resulting in difficulties for aircraft to fly according to the specified standard arrival and departure flight procedures, and flight track deviations during arrival and departure often occur. To improve the operation efficiency and safety level of the airport terminal area, the air traffic control system must monitor in real time and accurately identify the flight track deviations and true flight intentions of each aircraft in the terminal area.

[0003] Due to the lack of a unified quantitative evaluation standard and performance comparison benchmark, the in-depth development and application of relevant research have been restricted to a certain extent. To improve the operation efficiency and safety level of the airport terminal area, it is necessary to focus on the clustering analysis of aircraft arrival and departure tracks in the airport terminal area, that is, to conduct data mining on the track data in the terminal area airspace, construct a track clustering model, identify the pattern categories to which the tracks belong, and thus provide a reference for the refined and safe operation of the airport terminal area.

[0004] Therefore, the above-mentioned existing problems need to be solved urgently. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technology, the present invention provides an aviation trajectory semi-supervised clustering method based on a Transformer denoising autoencoder. By providing a unified airport terminal area track simulation data set and manually labeled clustering labels, key indicators for quantitatively evaluating the arrival and departure track clustering are constructed. The present invention innovatively designs a track clustering method for actual application requirements. Through quantitative evaluation and performance comparison, it can automatically and accurately identify and output track category labels based on the given original track information input.

[0006] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0007] An aviation trajectory semi-supervised clustering method based on a Transformer denoising autoencoder includes the following steps:

[0008] S100. Preprocessing of the flight arrival and departure trajectory data set: S110. Cleaning and interpolating and filling in the trajectory data;

[0009] S200. Feature extraction of the dataset: S210. Extract the features of the arrival trajectory and the departure trajectory from the preprocessed trajectory dataset; S220. Extract the features of the runway number to which the track belongs from the trajectory dataset.

[0010] S300. Model construction and model training: S310. Construct a trajectory noise reduction model for each flight arrival and departure trajectory data to reduce the influence of noise points in the trajectory data; S320. Divide the flight arrival and departure trajectory dataset into a training set and a validation set for model training.

[0011] S400. Semi-supervised clustering algorithm: S410. Cluster the flight arrival and departure trajectory data through a fine-grained semi-supervised clustering method.

[0012] After the model constructed in step S300 is trained well, a new test set is obtained. First, extract the data features of arrival / departure and runway number according to step S200, then input the track data into the model, and then submit the output denoised trajectory to the semi-supervised clustering algorithm for clustering to obtain the clustering result.

[0013] Furthermore, in S110, clean and interpolate the trajectory data to fill in the missing values:

[0014] In the flight arrival and departure trajectory dataset, the total number of tracks is n, and each track data contains the following fields: track index flightID, sampling timestamp t, longitude x, latitude y, and altitude height. Clean and interpolate each track data to fill in the missing values.

[0015] Furthermore, in S210, extract the features of the arrival trajectory and the departure trajectory from the preprocessed trajectory dataset:

[0016] Select the altitude height information of several timestamps t before and after a track index flightID, and judge whether the track belongs to an arrival track or a departure track according to the change of the height value.

[0017] In S220, extract the features of the runway number to which the track belongs from the trajectory dataset:

[0018] It includes obtaining the number of runways, obtaining the number of runway numbers, determining the runway position, and dividing the runway numbers.

[0019] Furthermore, in S310, construct a trajectory noise reduction model for each flight arrival and departure trajectory data:

[0020] For the trajectory flight set Traj n, with the autoencoder of AutoEncoder as the core, a denoising autoencoder based on the Transformer architecture, Transformer-based Denoised AutoEncoder, T-DAE; the encoder Encoder is built based on the Transformer model, and the decoder Decoder is built based on the multi-layer perceptron MLP model to construct a trajectory denoising model to reduce the influence of noise points in the trajectory data;

[0021] Further, in S320, the flight arrival and departure trajectory dataset is divided into a training set and a validation set for model training:

[0022] The flight trajectory dataset is divided into a training set and a validation set. The AdamW optimizer is used, the batch size is 64, the initial learning rate is set to 10e-3, and the Early Stopping early stopping mechanism is used, with the Early Stopping value set to 5;

[0023] Taking RMSE as the error index, the parameters of T-DAE are optimized through the backpropagation BP algorithm;

[0024] Further, in S410, the flight arrival and departure trajectory data is clustered by a fine-grained semi-supervised clustering method:

[0025] For the set of denoised trajectory flights Traj obtained from the T-DAE model nr , first according to TOL and runway k (k = 0, 1... 2l - 1) label values, all flight trajectories are divided into 4l large clusters C i , i ∈ [1, 4l], i ∈ N to supervise the subsequent fine-grained clustering.

[0026] Further, in S210, for the arrival flight trajectory, the height height values of several time stamps t at the end of the track are close to 0. For the departure flight trajectory, the height height values of several time stamps t at the start of the track are close to 0; by judging the arrival and departure of flights, we add a new label TOL (Takeoff or Land) in the training set and the validation set to characterize the category of track takeoff or landing; among them, when TOL = 0, it represents the departure track, that is, the takeoff category, and when TOL = 1, it represents the arrival track, that is, the landing category;

[0027] Further, in S220, the steps for obtaining the number of runways, obtaining the number of runway numbers, determining the runway position, and dividing the runway numbers are as follows:

[0028] S221. There is no relevant information about the airport runway situation in the flight arrival and departure trajectory dataset in step S100. First, through data mining, the number of airport runways is extracted. Secondly, the runway numbers are divided according to the runway directions; the runway number mining determines the number of airport runways by visualizing the trajectories (x, y) of the takeoff flights at the m time stamps t before takeoff and the trajectories (x, y) of the landing flights at the n time stamps t before landing.

[0029] S222. For an east-west runway, there are two major types of flights taking off from this runway: taking off from east to west and taking off from west to east. Similarly, there are two major types of flights landing on this runway: landing from the east and landing from the west; for one runway, one runway number should be assigned to each end of the runway to identify the runway type to which the flight belongs, that is, if the number of runways at an airport is l, the number of runway numbers at this airport is 2l.

[0030] S223. Perform Kmeans clustering on the trajectories [(x1,y1)(x2,y2)...(x m ,y m )] of the takeoff flights at the m time stamps t before takeoff and the trajectories [(x1,y1)(x2,y2)...(x n ,y n )] of the landing flights at the n time stamps t before landing respectively to obtain the trajectory clustering centers. The trajectory clustering centers are the parts where the trajectories of the takeoff flights at the m time stamps t before takeoff or the landing flights at the n time stamps t before landing highly overlap. Use the obtained trajectory clustering centers as the approximation of the airport runways.

[0031] The trajectory clustering centers obtained for the takeoff category flights are:

[0032] {runway1[(x 11 ,y 11 )(x 12 ,y 12 )…(x 1m ,y 1m )]…runway 2l [(x 2l1 ,y 2l1 )(x 2l2 ,y 2l2 )...(x 2lm ,y 2lm )]}

[0033] The trajectory clustering centers obtained for the landing category flights are:

[0034] {runway1[(x 11 ,y 11 )(x 12 ,y 12 )…(x1n,y 1n)]…runway 2l [(x 2l1 ,y 2l1 )(x 2l2 ,y 2l2 )...(x 2ln ,y 2ln )]}

[0035] S224. Add a new column of labels RunwayNumber to represent the runway number of the flight track. Select the Euclidean distance to calculate the similarity between the flight track and the runway number, so as to determine the runway number to which a certain flight track belongs;

[0036] For departing flights, first extract the track [(x1,y1)(x2,y2)...(x m ,y m )] of the m time stamps t before takeoff of the track, and calculate the distance between it and the clustering center track runway k (k = 0,1...2l - 1):

[0037]

[0038] Find the minimum S k And output the corresponding runway number. Define the sum of the minimum distances as S min , and find the corresponding runway number:

[0039]

[0040] The runway number label of this track is RunwayNumber = k min ;

[0041] For landing type flights, extract the track [(x1,y1)(x2,y2)...(x n ,y n )] of the n time stamps t before landing of the track, and calculate the distance between it and the clustering center track runway k (k = 0,1...2l - 1):

[0042]

[0043] Find the minimum S k And output the corresponding runway number. Define the sum of the minimum distances as S min , and find the corresponding runway number: Through formula 2, the runway number label of this track is RunwayNumber = k min ;

[0044] In the above S310, the trajectory noise reduction model takes the trajectory flight trajn All the trajectory points are used as input, and the output is the denoised trajectory flight traj nr The set of denoised trajectory flights is Traj nr ;

[0045] The number of heads of the multi-head attention in the encoder Transformer Block is 8, the number of Transformer layers is 4, and the decoder MLP is set with 3 linear layers, and the ReLU activation function is used to activate between the linear layers;

[0046] Further, in the S410, the Kmeans clustering algorithm is selected to perform clustering on each large cluster C i The specific steps are as follows:

[0047] S411. Use the number of clusters hyperparameter ncluster = n to train the Kmeans model K for the trajectories in C i on the training set, where ncluster ∈ [2, 20] and ncluster ∈ N; in

[0048] S412. Use K i to calculate the clustering score Score for the trajectories in C in on the validation set in ; Score is a user-defined clustering metric function, and its formula is as follows:

[0049]

[0050] This metric is calculated by weighting two internal metrics, the silhouette coefficient SC and the Davies-Bouldin index DBI, and two external metrics, the adjusted Rand index ARI and the normalized mutual information NMI based on the validation set labels. The calculation formulas for each metric are as follows:

[0051]

[0052] S413. For the clustering score set Score i of the C i cluster, find the number of clusters hyperparameter BestNcluster i with the highest score = argmax(Score i ), and the Kmeans model under this hyperparameter is the optimal model;

[0053] S414. Return the clustering results of the optimal Kmeans models for the 4l large clusters.

[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:​

[0055] 1. The present invention extracts the statistical features of the dataset after preprocessing the dataset by mining the capture time sequence relationship between the airport track and the runway number. Subsequently, the relationship between the track category and the takeoff and landing situation and the runway is mined, and the efficient clustering of the track is realized by using the mined takeoff and landing and runway number category features.

[0056] 2. By first dividing the base clusters and then performing fine-grained clustering, the similarity of the same type of trajectories is effectively increased; compared with directly clustering all the data, the semi-supervised clustering method using the base clusters significantly improves the clustering accuracy. At the same time, the present invention proposes a model based on the T-DAE (Transformer-based Denoising AutoEncoder) structure for auto-encoding analysis of the input trajectory data; the present invention can efficiently cluster the input track data clusters without external factor variables for the cluster track classes.

[0057] 3. The present invention conducts research based on the existing aviation trajectory big data, has wide applicability, and does not require other additional data such as weather, wind direction, etc. for auxiliary clustering; at the same time, the steps of the algorithm design of the present invention are easy to understand, fast in calculation, small in overhead, and high in accuracy. The designed Transformer-based denoising autoencoder model is easy to train and can effectively reveal the flight operation rules and airspace usage characteristics, thereby improving the efficiency and safety of air traffic management. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a schematic flow chart of the method of the present invention;

[0059] Figure 2 is a schematic visualization diagram of the landing trajectory of the landing flight of the present invention;

[0060] Figure 3 is a schematic diagram of the clustering runway numbers of the landing flights of the present invention;

[0061] Figure 4 is a schematic diagram of the runway numbers of the takeoff flights of the present invention;

[0062] Figure 5 is a schematic diagram of the T-DAE model architecture of the present invention;

[0063] Figure 6 is a schematic visualization diagram of the clustering index results of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0065] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention.

[0066] All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0067] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it is not necessary to further define and explain it in subsequent figures.

[0068] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0069] An aviation trajectory semi-supervised clustering method based on a Transformer denoising autoencoder includes the following steps:

[0070] S100. Preprocessing of the flight arrival and departure trajectory dataset: S110. Cleaning and interpolating the trajectory data to complete it;

[0071] S200. Feature extraction of the dataset: S210. Extracting the features of arrival trajectories and departure trajectories from the preprocessed trajectory dataset, and S220. Extracting the features of the runway numbers to which the flight tracks belong from the trajectory dataset;

[0072] S300. Model construction and model training: S310. Constructing a trajectory denoising model for each flight arrival and departure trajectory data to reduce the influence of noise points in the trajectory data, and S320. Dividing the flight arrival and departure trajectory dataset into a training set and a validation set for model training;

[0073] S400, Semi-supervised clustering algorithm: S410, Cluster the flight arrival and departure trajectory data through a fine-grained semi-supervised clustering method;

[0074] S500, After training the model constructed in step S300, obtain a new test set. First, extract the data features of arrival, departure, and runway numbers according to step S200, then input the track data into the model, and hand over the output denoised trajectory to the semi-supervised clustering algorithm for clustering to obtain the clustering result.

[0075] Furthermore, for the semi-supervised clustering method of aviation trajectories based on the Transformer denoising autoencoder, S110, Clean and interpolate the trajectory data:

[0076] In the flight arrival and departure trajectory dataset, the total number of trajectories is n, and each trajectory data contains the following fields: trajectory index flightID, sampling timestamp t, longitude x, latitude y, altitude height. Clean and interpolate each trajectory data;

[0077] S210, Extract the features of arrival trajectories and departure trajectories from the preprocessed trajectory dataset:

[0078] Select the altitude height information of several timestamps t before and after a trajectory index flightID, and judge whether this trajectory belongs to an arrival trajectory or a departure trajectory based on the change of the height value;

[0079] S220, Extract the features of the runway number to which the trajectory belongs from the trajectory dataset:

[0080] Including obtaining the number of runways, obtaining the number of runway numbers, determining the runway position, and dividing the runway numbers;

[0081] S310, Construct a trajectory denoising model for each flight arrival and departure trajectory data:

[0082] For the trajectory flight set Traj n , with the autoencoder of AutoEncoder as the core, a denoising autoencoder based on the Transformer architecture, Transformer-based Denoised AutoEncoder, T-DAE; the encoder Encoder is built based on the Transformer model, and the decoder Decoder is built based on the multi-layer perceptron MLP model to construct a trajectory denoising model to reduce the influence of noise points in the trajectory data;

[0083] S320, Divide the flight arrival and departure trajectory dataset into a training set and a validation set for model training:

[0084] The flight trajectory dataset is divided into a training set and a validation set. The AdamW optimizer is used, with a batch size of 64 and an initial learning rate set to 10e-3. The Early Stopping mechanism is used, and the Early Stopping value is set to 5;

[0085] Taking RMSE as the error metric, the parameters of T-DAE are optimized through the backpropagation BP algorithm;

[0086] S410. Cluster the flight arrival and departure trajectory data through fine-grained semi-supervised clustering:

[0087] For the denoised trajectory flight set Traj obtained from the T-DAE model nr , first, according to the TOL and runway k (k = 0, 1... 2l - 1) label values, all flight trajectories are divided into 4l large clusters C i , i ∈ [1, 4l], i ∈ N to supervise the subsequent fine-grained clustering.

[0088] Furthermore, a semi-supervised clustering method for aviation trajectories based on the Transformer denoising autoencoder

[0089] In S210, for the arrival flight trajectory, the height values of several time stamps t at the end of the trajectory are close to 0. For the departure flight trajectory, the height values of several time stamps t at the start of the trajectory are close to 0. By judging the arrival and departure of flights, a new label TOL (Takeoff or Land) is added to the training set and the validation set to represent the category of the trajectory takeoff or landing. Among them, when TOL = 0, it represents the departure trajectory, that is, the takeoff category. When TOL = 1, it represents the arrival trajectory, that is, the landing category;

[0090] In S220, obtain the number of runways, obtain the number of runway numbers, determine the runway position, and divide the runway numbers. The specific steps are as follows:

[0091] S221. There is no relevant information about the airport runway situation in the flight arrival and departure trajectory dataset in step S100. First, extract the number of airport runways through data mining, and then divide the runway numbers according to the direction of the runway. The number of runways is mined by visualizing the trajectories (x, y) of m time stamps t before takeoff of takeoff flights and the trajectories (x, y) of n time stamps t before landing of landing flights to determine the number of airport runways;

[0092] S222. An east-west runway. The flights taking off from this runway can be divided into two major types: taking off from east to west and taking off from west to east. Similarly, the flights landing on this runway can also be divided into two major types: landing from the east and landing from the west. For a runway, a runway number should be assigned to each end of the runway to identify the runway type to which the flight belongs. That is, if the number of runways at an airport is l, the number of runway numbers at this airport is 2l;

[0093] S223. For the trajectories [(x1,y1)(x2,y2)...(x m ,y m )] of the takeoff flights at m time stamps t before takeoff and the trajectories [(x1,y1)(x2,y2)...(x n ,y n )] of the landing flights at n time stamps t before landing, perform Kmeans clustering respectively to obtain the trajectory clustering centers. The trajectory clustering centers are the parts where the trajectories of the takeoff flights at m time stamps t before takeoff or the landing flights at n time stamps t before landing highly overlap. Use the obtained trajectory clustering centers as the approximation of the airport runway;

[0094] The trajectory clustering centers obtained for the takeoff category flights are:

[0095] {runway1[(x 11 ,y 11 )(x 12 ,y 12 )…(x 1m ,y 1m )]…runway 2l [(x 2l1 ,y 2l1 )(x 2l2 ,y 2l2 )...(x 2lm ,y 2lm )]}

[0096] The trajectory clustering centers obtained for the landing category flights are:

[0097] {runway1[(x 11 ,y 11 )(x 12 ,y 12 )…(x 1n ,y 1n )]…runway 2l [(x 2l1 ,y 2l1 )(x 2l2 ,y 2l2 )...(x 2ln ,y 2ln )]}

[0098] S224. Add a new column of labels, RunwayNumber, to represent the runway number of the flight track. Select the Euclidean distance to calculate the similarity between the flight track and the runway number, so as to determine the runway number to which a certain flight track belongs.

[0099] For departing flights, first extract the track [(x1,y1)(x2,y2)...(x m ,y m )] of the m time stamps t before takeoff of the track, and calculate the distance between it and the clustering center track runway k (k = 0,1...2l - 1):

[0100]

[0101] Find the minimum S k And output the corresponding runway number. Define the sum of the minimum distances as S min , and find the corresponding runway number:

[0102]

[0103] The runway number label of this track is RunwayNumber = k min ;

[0104] For landing type flights, extract the track [(x1,y1)(x2,y2)...(x n ,y n )] of the n time stamps t before landing of the track, and calculate the distance between it and the clustering center track runway k (k = 0,1...2l - 1):

[0105]

[0106] Find the minimum S k And output the corresponding runway number. Define the sum of the minimum distances as S min , and find the corresponding runway number: Through formula 2, the runway number label of this track is RunwayNumber = k min ;

[0107] In S310, the trajectory noise reduction model takes all the trajectory points of the trajectory flight traj n as input, and the output is the trajectory flight traj nr after noise reduction. The set of trajectory flights after noise reduction is Traj nr ;

[0108] The number of heads in the multi-head attention of the encoder Transformer block is 8, the number of Transformer layers is 4, and the decoder MLP is set with 3 linear layers, which are activated by the ReLU activation function between the linear layers;

[0109] In S410, the Kmeans clustering algorithm is selected to cluster each large cluster C i specifically, the steps are as follows:

[0110] S411. For the trajectories in C i the Kmeans model K is trained on the training set using the hyperparameter ncluster=n of the number of clusters, in where ncluster∈[2,20] and ncluster∈N;

[0111] S412. For the trajectories in C i the clustering score Score is calculated using K in on the validation set; in ;

[0112] Score is a custom clustering metric function, and its formula is as follows:

[0113]

[0114] This metric is calculated by weighted combination of two internal metrics, the silhouette coefficient SC and the Davies-Bouldin index DBI, and two external metrics, the adjusted Rand index ARI and the normalized mutual information NMI based on the validation set labels. The calculation formulas for each metric are as follows:

[0115]

[0116] S413. For the set of clustering scores Score i of the C i cluster, find the hyperparameter BestNcluster of the number of ncluster clusters with the highest score i = argmax(Score i ), and the Kmeans model under this hyperparameter is the optimal model;

[0117] S414. Return the clustering results of the optimal Kmeans models for the 4l large clusters.

[0118] The present invention will be further explained below in conjunction with the attached Figure 1-6 drawings and Embodiment 1:

[0119] The aircraft arrival and departure flight track data used in this embodiment are data under a certain simulation scenario. The original data has been normalized and consists of 8,573 arrival tracks, 8,329 departure tracks, and some track category labels. Each track contains information such as flight index flightID, timestamp t, longitude x, latitude y, altitude height, etc.; the preprocessing includes abnormal data filtering and interpolation to fill in the blanks. The process is shown in Figure 1 as follows:

[0120] S100. Preprocessing of the flight arrival and departure track dataset:

[0121] Filter the track records with excessive curvature in a section before takeoff for the takeoff type tracks; by statistically analyzing the sampling time difference between consecutive track records of each aircraft, the sampling rate of this dataset is determined to be 4 seconds. At the same time, by statistically analyzing the number of track points included in consecutive tracks of each aircraft, it is determined that the number of track points included in each track is within 500 track points. In this example 1, cubic spline interpolation is used to interpolate each track data to 550 data points.

[0122] S200. Feature extraction of the dataset:

[0123] S210. Extract the features of arrival tracks and departure tracks from the preprocessed track dataset:

[0124] By selecting the altitude height information at several timestamps t before and after a track index flightID, and judging whether this track belongs to an arrival track or a departure track based on the change of the height value; for the arrival flight track, the altitude height values at the last 30 timestamps t of its track are close to 0; for the departure flight track, the altitude height values at the first 30 timestamps t of its track are close to 0; through the judgment of flight arrival and departure, we add a new label TOL, that is, Takeoff or Land, to the training set and the validation set to characterize the category of track takeoff or landing. Among them, when TOL = 0, it represents a departure track, that is, the takeoff category, and when TOL = 1, it represents an arrival track, that is, the landing category.

[0125] S220. Extract the features of the runway number to which the track belongs from the track dataset:

[0126] The process of extracting the runway number will be divided into four aspects: obtaining the number of runways, obtaining the number of runway numbers, determining the runway position, and dividing the runway numbers:

[0127] Since there is no relevant information about the airport runway situation in the original dataset, we first extract the number of airport runways through data mining, and then divide the runway numbers according to the runway directions. Runway number mining method: Determine the number of airport runways by visualizing the trajectories (x, y) of departing flights in the 6 time stamps t before takeoff and the trajectories (x, y) of arriving flights in the 30 time stamps t before landing. Figure 2 It is a visualization schematic diagram of the landing trajectory of the landing flight in Embodiment 1, and it can be determined that the airport has two runways.

[0128] One east-west runway. The flights taking off from this runway have two major types: taking off from east to west and taking off from west to east. Similarly, the flights landing on this runway also have two major types: landing from the east and landing from the west. Refer to Figure 2 Visible, Figure 2 The label in is the true label of the flight track. Different labels represent different categories of aircraft flight tracks. For a runway, the labels of the flight tracks on both sides of the runway are not the same, that is, for a runway, the flight tracks landing from different directions of the runway belong to different categories. A similar conclusion can also be obtained for takeoff tracks, that is, for a runway, the flight tracks taking off from different directions of the runway belong to different categories. Therefore, for a runway, a runway number should be assigned to each end of the runway to identify the runway type to which the flight belongs. In Embodiment 1, the number of runways in the airport terminal area is 2, so the number of runway numbers in this airport is 4.

[0129] By performing Kmeans clustering on the trajectories [(x1,y1)(x2,y2)...(x m ,y m )] of departing flights in the m time stamps t before takeoff and the trajectories [(x1,y1)(x2,y2)...(x n ,y n )] of arriving flights in the n time stamps t before landing respectively, the trajectory clustering centers are obtained. The trajectory clustering centers are the parts where the trajectories of the departing flights in the 6 time stamps t before takeoff or the arriving flights in the 30 time stamps t before landing highly coincide. Use the obtained trajectory clustering centers as the approximation of the airport runway.

[0130] The trajectory clustering centers obtained for the departing flight category are:

[0131] {runway1[(x 11 ,y 11 )(x 12 ,y 12 )…(x 16 ,y 16 )]…runway4[(x 41 ,y 41 )(x 42 ,y 42)...(x 46 ,y 46 )]}

[0132] The trajectory clustering centers obtained for landing category flights are:

[0133] {runway1[(x 11 ,y 11 )(x 12 ,y 12 )…(x 130 ,y 130 )]…runway4[(x 41 ,y 41 )(x 42 ,y 42 )...(x 430 ,y 430 )]}

[0134] Figure 3 and Figure 4 are schematic diagrams of the runway numbers for clustering landing and takeoff flights. The cross marks pointed by the runway number arrows in Figure 3 and 4 are the obtained trajectory clustering centers.

[0135] Runway number division:

[0136] A new column label RunwayNumber is added to the data, and its values 0 - 3 represent runways 1 - 4 respectively. The Euclidean distance is selected to calculate the similarity between the flight trajectory and the runway number, so as to determine the runway number to which a certain flight trajectory belongs. For example, for takeoff flights, first extract the trajectory [(x1,y1)(x2,y2)...(x6,y6)] of the 6 time stamps t before takeoff of the flight track, and let it be compared with the trajectory of each takeoff runway clustering center

[0137] runway k (k = 0,1...3) to calculate the distance:

[0138]

[0139] Finally, we need to find the minimum S k and output the corresponding runway number. Define the sum of the minimum distances as S min , and finally find the corresponding runway number:

[0140]

[0141] Then the runway number label of this flight track is RunwayNumber = k min . For landing category flights, use the same method to extract the trajectory of the 30 time stamps t before landing of the flight track

[0142] [(x1,y1)(x2,y2)...(x 30 ,y 30 )], let it calculate the distance from each landing runway clustering center trajectory runway k (k = 0, 1...3):

[0143]

[0144] And find the minimum S k (Formula (2)) and output the corresponding runway number.

[0145] S300, Model Construction and Model Training:

[0146] Model Construction:

[0147] For the trajectory flight set Traj n , the present invention takes the auto - encoder structure of AutoEncoder as the core, and researches and designs a denoising auto - encoder based on the Transformer architecture (Transformer - based DenoisedAutoEncoder, T - DAE). The architecture diagram of the denoising auto - encoder is shown in Figure 5 . The encoder Encoder is built based on the Transformer model, and the decoder Decoder is built based on the multi - layer perceptron MLP model. For the input track data, the auto - encoder first maps it to the latent space to extract effective features, and then processes these features through multiple layers of Transformer encoders to capture complex patterns and relationships in the input data. In the decoder part, the model reconstructs the latent representation generated by the encoder back to the input space, finally obtaining the denoised trajectory data, which is used as the input for the subsequent K - means model. The purpose of constructing the trajectory denoising model is to reduce the influence of noise points in the trajectory data.

[0148] The trajectory denoising model takes all the trajectory points of the trajectory flight traj n as the input, and the output is the denoised trajectory flight traj nr , and the set of denoised trajectory flights is Traj nr .

[0149] The number of heads of the multi - head attention in the encoder Transformer Block is 8, and the number of Transformer layers is 4. The decoder MLP sets 3 linear layers, and the ReLU activation function is used for activation between the linear layers.

[0150] Model Training:

[0151] The flight trajectory dataset is divided into a training set and a validation set. The AdamW optimizer is used, with a batch size of 64. The initial learning rate is set to 10e-3, and the Early Stopping mechanism is used, with the Early Stopping value set to 5.

[0152] Finally, with RMSE as the error metric, the parameters of the T-DAE are optimized through the backpropagation BP algorithm.

[0153] S400, semi-supervised clustering algorithm:

[0154] Through the fine-grained semi-supervised clustering idea of "divide and conquer", the arrival and departure trajectory data of flights are clustered efficiently and accurately:

[0155] For the set of denoised trajectory flights Traj obtained from the T-DAE model nr , first, according to the TOL and runway k (=0,1...3) label values, all flight trajectories are divided into 8 large clusters c i , i ∈ [1,8], i ∈ N to supervise the subsequent fine-grained clustering.

[0156] Select the Kmeans clustering algorithm to cluster each large cluster C G specifically, the steps are as follows:

[0157] For the trajectories in C i , use the number of clusters hyperparameter ncluster = n to train the Kmeans model K on the training set in , ncluster ∈ [2,20] ncluster ∈ N;

[0158] For the trajectories in C i , use K in to calculate the clustering score Score in ;

[0159] Score is a custom clustering metric function, and its formula is as follows:

[0160]

[0161] This metric is calculated by weighting two internal metrics, the silhouette coefficient SC and the Davies-Bouldin index DBI, and two external metrics, the adjusted Rand index ARI and the normalized mutual information NMI based on the validation set labels. The calculation formulas for each metric are as follows:

[0162]

[0163] For C iThe set of clustering scores of clusters, Score i , find the hyperparameter BestNclusteri of the number of ncluster clusters with the highest score, BestNclusteri = argmax(Score i ). The Kmeans model under this hyperparameter is the optimal model;

[0164] Return the clustering results of the optimal Kmeans model for 8 large clusters.

[0165] S500, Clustering of arrival and departure track trajectories

[0166] After training the model constructed in step S300, when a new test set with labels is obtained, first extract the data features of arrival and departure and runway numbers according to step S200, and then input the track data Traj n to the model, and input the output denoised trajectory Traj nr to the semi-supervised clustering algorithm for clustering to obtain the clustering results. Table 1 shows the evaluation index scores of each category in the clustering results of Example 1:

[0167] Table 1 Evaluation index scores of each category for clustering

[0168]

[0169] Category(x, y), where x represents the TOL value and y represents the RunwayNum value

[0170] Clustering results:

[0171] In this Example 1 for 4 types of departing flights, i.e., TOL = 0, both the adjusted Rand index ARI and the normalized mutual information NMI are 1.0, which indicates that the model can completely consistently identify the classification method of the samples, that is, among all samples, samples of the same class are correctly classified into the same class in both clustering results, while samples of different classes are also correctly separated. For 4 types of arriving flights, i.e., TOL = 1, the ARI and NMI scores on RunwayNum = 0 and RunwayNum = 2 are relatively high, indicating that the model has excellent classification ability; while the correct classification ability on RunwayNum = 1 and RunwayNum = 3 is good. The results are shown in Figure 6 , the horizontal axis represents eight large clusters, where Category(x, y), x represents the TOL value and y represents the RunwayNum value, the vertical axis is the clustering index score, and the columns of different colors represent different evaluation indexes. This result achieves efficient clustering of cluster track trajectories.

[0172] The present invention has been described exemplarily in conjunction with the accompanying drawings. Obviously, the specific implementation of the present invention is not limited by the above embodiments. Those skilled in the art can make various modifications or variations to the present invention without departing from the technical concept of the present invention, and of course, these modifications or variations also fall within the protection scope of the present invention.

Claims

1. A semi-supervised clustering method for aviation trajectories based on a Transformer denoising autoencoder, characterized in that: It includes the following steps: S100. Preprocessing of the flight arrival and departure trajectory dataset: S110. Cleaning and interpolating the trajectory data to fill in the gaps; S200. Feature extraction of the dataset: S210. Extracting the features of the arrival and departure trajectories from the preprocessed trajectory dataset, S220. Extracting the features of the runway numbers to which the flight tracks belong from the trajectory dataset; S300. Model construction and model training: S310. Constructing a trajectory noise reduction model for each flight arrival and departure trajectory data to reduce the influence of noise points in the trajectory data, S320. Dividing the flight arrival and departure trajectory dataset into a training set and a validation set for model training; S400. Semi-supervised clustering algorithm: S410. Clustering the flight arrival and departure trajectory data through a fine-grained semi-supervised clustering method; After training the model constructed in step S300, a new test set is obtained. First, extract the data features of arrival / departure and runway numbers according to step S200, then input the flight track data into the model, and then submit the output denoised trajectory to the semi-supervised clustering algorithm for clustering to obtain the clustering result.

2. The semi-supervised clustering method for aviation trajectories based on the Transformer denoising autoencoder according to claim 1, characterized in that: In step S110, cleaning and interpolating the trajectory data to fill in the gaps: In the flight arrival and departure trajectory dataset, the total number of flight tracks is n, and each trajectory data contains the following fields: flight track index flightID, sampling timestamp t, longitude x, latitude y, and altitude height. Clean and interpolate each trajectory data to fill in the gaps; In step S210, extracting the features of the arrival and departure trajectories from the preprocessed trajectory dataset: Select the altitude height information of a certain number of timestamps t before and after a flight track index flightID, and judge whether the flight track belongs to an arrival track or a departure track according to the change of the height value; In step S220, extracting the features of the runway numbers to which the flight tracks belong from the trajectory dataset: It includes obtaining the number of runways, obtaining the number of runway numbers, determining the runway positions, and dividing the runway numbers; In step S310, constructing a trajectory noise reduction model for each flight arrival and departure trajectory data: For the trajectory flight set Traj n , with the autoencoder of the AutoEncoder as the core, a denoising autoencoder based on the Transformer architecture, the Transformer-based Denoised AutoEncoder (T-DAE); the encoder Encoder is built based on the Transformer model, and the decoder Decoder is built based on the multi-layer perceptron MLP model to construct a trajectory denoising model to reduce the impact of noise points in the trajectory data; In step S320, dividing the flight arrival and departure trajectory dataset into a training set and a validation set for model training: Divide the flight trajectory dataset into a training set and a validation set, use the AdamW optimizer, the batch size is 64, the initial learning rate is set to 10e-3, use the Early Stopping early stopping mechanism, and the Early Stopping value is set to 5; Taking RMSE as the error index, optimize the parameters of the T-DAE through the backpropagation BP algorithm; In step S410, clustering the flight arrival and departure trajectory data through a fine-grained semi-supervised clustering method: For the denoised trajectory flight set Traj obtained from the T-DAE model nr , first, according to the TOL and runway k (k = 0, 1... 2l - 1) label values, all flight trajectories are divided into 4l large clusters C i , i ∈ [1, 4l], i ∈ N to supervise the subsequent fine-grained clustering.

3. The semi-supervised clustering method for aviation trajectories based on the Transformer denoising autoencoder according to claim 2, characterized in that: In S210, for the approach flight trajectory, the height values of several time stamps t at the end of the trajectory are close to 0. For the departure flight trajectory, the height values of several time stamps t at the start of the trajectory are close to 0. By judging the approach and departure of flights, we add a new label TOL (Takeoff or Land) to the training set and the validation set to represent the category of takeoff or landing of the trajectory. Among them, when TOL = 0, it represents a departure trajectory, that is, the takeoff category. When TOL = 1, it represents an approach trajectory, that is, the landing category. In S220, obtaining the number of runways, obtaining the number of runway numbers, determining the runway position, and dividing the runway numbers, the specific steps are as follows: S221. There is no relevant information about the airport runway situation in the flight approach and departure trajectory dataset in step S100. First, through data mining, the number of airport runways is extracted. Secondly, the runway numbers are divided according to the direction of the runway. The number of runways is mined by visualizing the trajectories (x, y) of m time stamps t before takeoff of takeoff flights and the trajectories (x, y) of n time stamps t before landing of landing flights to determine the number of airport runways. S222. For an east-west runway, there are two major types of flights taking off from this runway: taking off from east to west and taking off from west to east. Similarly, there are two major types of flights landing on this runway: landing from the east and landing from the west. For one runway, each end of the runway should be assigned a runway number to identify the runway type to which the flight belongs. That is, if the number of runways at an airport is l, the number of runway numbers at this airport is 2l. S223. Perform Kmeans clustering on the trajectories [(x1, y1)(x2, y2)...(x m , y m )] of the takeoff flight at the m timestamps t before takeoff and the trajectories [(x1, y1)(x2, y2)...(x n , y n )] of the landing flight at the n timestamps t before landing respectively to obtain the trajectory clustering centers. The trajectory clustering centers are the parts where the trajectories of the takeoff flight at the m timestamps t before takeoff or the landing flight at the n timestamps t before landing highly overlap. Use the obtained trajectory clustering centers as the approximation of the airport runway; The trajectory clustering center obtained by the takeoff category flight is: {runway1[(x 11 ,y 11 )(x 12 ,y 12 )…(x 1m ,y 1m )]… runway 2l [(x 2l1 ,y 2l1 )(x 2l2 ,y 2l2 )...(x 2lm ,y 2lm )]} The trajectory clustering center obtained by the landing category flight is: {runway1[(x 11 ,y 11 )(x 12 ,y 12 )…(x 1n ,y 1n )]… runway 2l [(x 2l1 ,y 2l1 )(x 2l2 ,y 2l2 )...(x 2ln ,y 2ln )]} S224. Add a new column of labels RunwayNumber to represent the runway number of the trajectory. Select the Euclidean distance to calculate the similarity between the flight trajectory and the runway number, so as to judge the runway number to which a certain flight trajectory belongs. For the departing flight, first extract the track [(x1,y1)(x2,y2)...(x m ,y m )] of the m timestamps t before takeoff of the track, and calculate the distance between it and the track runway k (k = 0, 1... 2l - 1) for each takeoff runway clustering center track: Find the minimum S k and output the corresponding runway number. Define the minimum distance sum as S min and find the corresponding runway number: The runway number label of this track is RunwayNumber = k min ; For landing-type flights, extract the trajectories [(x1,y1)(x2,y2)...(x n ,y n )] of the first n timestamps t before landing of the flight path, and calculate the distance between them and the trajectory runwary k (k = 0, 1... 2l-1) of the clustering center of each landing runway: Find the minimum S k And output the corresponding runway number. Define the minimum distance sum as S min , and find the corresponding runway number: Through formula 2, the runway number label of this track is RunwayNumber = k min ; In the S310, the trajectory noise reduction model takes all the trajectory points of the trajectory flight traj n as input, and the output is the trajectory flight traj nr after noise reduction. The set of trajectory flights after noise reduction is Traj nr ; The number of heads of the multi-head attention (Multi-Head Attention) of the encoder Transformer Block is 8, the number of Transformer layers is 4, and the decoder MLP is set with 3 linear layers, and the ReLU activation function is used for activation between the linear layers. In S410, the Kmeans clustering algorithm is selected to perform clustering on each large cluster C i respectively, and the specific steps are as follows: S411. For C i Use the Kmeans model K with the number of clusters hyperparameter ncluster=n to train the trajectories in the training set respectively in , where ncluster ∈ [2, 20] and ncluster ∈ N; S412. For C i Use K for the trajectory in the validation set in Calculate the clustering score Score in ; Score is a custom clustering metric function, and its formula is as follows: This metric is obtained by weighted calculation of two internal metrics, the silhouette coefficient SC and the Davies-Bouldin index DBI, and two external metrics, the adjusted Rand index ARI and the normalized mutual information NMI based on the validation set labels. The calculation formulas of each metric are as follows: S413. For C i the clustering score set Score of the cluster i , find the hyperparameter BestNcluster of the number of ncluster clusters with the highest score i = argmax(Score i ), and the Kmeans model under this hyperparameter is the optimal model; S414. Return the clustering results of the optimal Kmeans model for 4l large clusters.

Citation Information

Cited By

  • Terminal area airspace traffic flow spatial pattern mining method considering abnormal trajectory

    CN121093299A

  • A terminal area airspace traffic flow spatial pattern mining method considering abnormal trajectories

    CN121093299B