Hybrid network 4D flight path prediction method based on LSTM and Transformer

Through the hybrid network of LSTM and Transformer, the time dependence relationship of flight trajectory data is dynamically captured, which solves the problem that dynamic characteristics in the existing methods are not fully considered, and achieves high-precision flight trajectory prediction and calculation efficiency improvement.

CN120471313APending Publication Date: 2025-08-12KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510186424.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing flight trajectory prediction methods do not fully consider the dynamic characteristics of the input sequence, resulting in insufficient dynamic capture capabilities, unable to effectively adapt to the requirements of flight trajectory prediction tasks, and high computing resource consumption.

Method used

Using a hybrid network of LSTM and Transformer, the time dependence and correlation of flight trajectory data are dynamically captured through the embedding layer, LSTM network, multi-head attention mechanism and mask multi-head attention mechanism to generate high-precision flight trajectory prediction.

Benefits of technology

It significantly improves the modeling ability of dynamic sequences, improves the accuracy and efficiency of flight trajectory prediction, and reduces the demand for computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471313A_ABST
    Figure CN120471313A_ABST
Patent Text Reader

Abstract

The invention relates to a hybrid network 4D flight path prediction method based on LSTM and Transform, and belongs to the technical field of flight path prediction. Comprising the following steps: inputting preprocessed flight path data into an embedding layer, mapping the preprocessed flight path data into high-dimensional feature representation, and forming an input sequence; inputting the input sequence into the LSTM network to generate an output representation with a time-dependent feature; performing global feature extraction on the output representation of the LSTM network by using a multi-head attention mechanism so as to capture a dependency relationship and relevance between different trajectory data points; performing normalization processing on global features output by the multi-head attention mechanism, and further enhancing feature expression ability through a feedforward neural network; a mask multi-head attention mechanism is adopted in a decoder, and prediction features are output; and mapping the prediction features output by the decoder to a target space through a linear layer, and generating a final result of flight path prediction. The method is excellent in performance in complex flight path prediction tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a 4D flight trajectory prediction method based on a hybrid network of LSTM and Transformer, belonging to the technical field of flight trajectory prediction. Background Art

[0002] With rapid economic development, demand for air transport continues to grow rapidly, and the contradiction between aviation demand and airspace capacity is becoming increasingly prominent. Airspace complexity is increasing. Currently, the operating model of fixed airspace segments and routes suffers from structural rigidity, cascading failures, and limited capacity. This not only limits the scope for aviation communication optimization but also fails to support future trajectory-based and performance-based airspace operating models. Data in the ADS-B system is expected to grow explosively in the future. Time series analysis of ADS-B can help address the challenges posed by this growing aviation demand.

[0003] The flight trajectory prediction (FTP) mission is attracting increasing research attention from academia and industry worldwide due to its support capabilities for future trajectory-based operations (TBO), including the Single European Sky ATM Research (SESAR) and the Next Generation Air Transportation System (NextGen). The core concept of TBO is that allowing traffic participants to share future flight trajectory predictions can enhance the interconnection between the air and the ground, and achieve safe and efficient air traffic control (ATC). Accurate prediction of aircraft four-dimensional (4D) trajectories is a basic technology for improving the predictability of TBO air traffic and is of great significance for completing downstream tasks such as estimated arrival time, conflict detection and air traffic flow prediction. The main goal of the FTP mission is to predict the motion attributes that describe the discrete trajectory points of the aircraft, such as longitude, latitude, altitude, and speed, which are referred to as 4D.

[0004] Although existing fusion models using the EMD-LSTM method can improve the model's predictive ability to a certain extent, the inherent problems of the EMD method, such as mode mixing and boundary effects, will affect the model's predictive performance. Therefore, it is necessary to develop a model with better decomposition capabilities to extract the hidden features of time series.

[0005] Current research still faces challenges with insufficient prediction performance and uneven prediction accuracy. For example, clustering-based methods have limited predictive performance due to input information limitations. Regression statistical models require extensive trajectory data to model each flight. Because BP prediction models consider the two-dimensional location of the aircraft, namely its latitude and longitude, their prediction dimensionality is insufficient. Furthermore, conventional machine learning methods, such as the K-nearest neighbor algorithm, support vector machines (SVMs), decision trees, and random forests, require extensive feature engineering to ensure model fit and evaluation, preserving temporal features in the data. Deep learning methods, however, have previously been primarily based on LSTM networks, which have a relatively simple architecture, due to their excellent performance for time series data prediction. Furthermore, to further exploit the spatial characteristics of trajectory datasets, hybrid models combining convolutional neural networks with LSTMs have seen increasing research in recent years. However, these models, due to their large network size, require high computing power.

[0006] The technology of this invention comes from the Yunnan Provincial Major Science and Technology Special Project (202302AD080002); the Yunnan Provincial Key Laboratory of Computer Technology Application Open Fund (CB22144S073A); and the "Xingdian Talent Support Plan" Industrial Innovation Talent Project (Yunnan Development and Reform Commission Personnel

[2019] No. 1096). Summary of the Invention

[0007] The technical problem solved by the present invention is: the present invention provides a hybrid network 4D flight trajectory prediction method based on LSTM and Transformer, which is used to solve the problems that the current technology does not fully consider the dynamic characteristics of the input sequence, does not optimize the specific dynamic characteristics in the flight trajectory data, and sacrifices a certain degree of dynamic capture ability. The present invention makes up for the shortcomings of the existing methods in dynamic time series tasks, can adapt to the needs of flight trajectory prediction tasks, and significantly improves the modeling ability of dynamic sequences.

[0008] The technical solution of the present invention is: a hybrid network 4D flight trajectory prediction method based on LSTM and Transformer, the method comprising:

[0009] Step 1: Collect the original flight trajectory data and pre-process the original flight trajectory data;

[0010] Step 2: Input the pre-processed flight trajectory data into the embedding layer and map it into a high-dimensional feature representation to form an input sequence;

[0011] Step 3: Input the input sequence into the LSTM network to extract the context information of the time series and generate an output representation with time-dependent features;

[0012] Step 4: Use the multi-head attention mechanism to extract global features from the output representation of the LSTM network to capture the dependencies and correlations between different trajectory data points;

[0013] Step 5: Normalize the global features output by the multi-head attention mechanism and further enhance the feature expression capability through a feedforward neural network;

[0014] Step 6: Use the masked multi-head attention mechanism in the decoder to output the predicted features;

[0015] Step 7: Map the predicted features output by the decoder to the target space through a linear layer to generate the final result of the flight trajectory prediction.

[0016] Furthermore, the Step 1 includes:

[0017] Extracting original flight trajectory data from flight data records, the original flight trajectory data includes time-space sequence information, and the time-space sequence information includes the aircraft's timestamp, latitude and longitude, altitude, speed, and heading angle;

[0018] Then perform data cleaning, normalization, interpolation and denoising.

[0019] Furthermore, in Step 2, the embedding layer adopts a fully connected layer or a feature encoding function, and the feature encoding function performs feature encoding including: temporal feature encoding and spatial feature encoding;

[0020] The time feature encoding converts the timestamp into a periodic feature;

[0021] The spatial feature encoding maps the latitude and longitude information to a fixed-dimensional feature vector through MLP;

[0022] The output of the embedding layer is an embedding feature matrix with a shape of (T, D), where T is the time step and D is the embedding dimension.

[0023] Furthermore, in Step 3, the key structure of the LSTM network includes an input gate, a forget gate, and an output gate; the LSTM network dynamically stores important timing information and filters irrelevant information through its memory unit.

[0024] Furthermore, the Step 3 includes:

[0025] For the input sequence X={x1,x2,...,x T}Calculate the hidden state h through the LSTM network tand cell state c t , the LSTM network outputs a context representation of a time series H = {h1,h2,...,h T}, as a high-dimensional representation of temporal features; hidden state h t and cell state c t The calculation process is as follows:

[0026]

[0027] Among them, f t Represents the output vector of the forget gate, with a value of 0 to 1, which determines how much of the cell state C at the previous moment is retained t-1 ;W f Represents the weight matrix of the forget gate, controlling h t-1 and x t Effect on the degree of forgetting; h t-1 Indicates the hidden state of the previous moment, carrying historical information; x t represents the input vector at time t; [h t-1 ,x t ] means h t-1 and x t The concatenated vector; b f Represents the bias vector of the forget gate, which adjusts the threshold of the activation function; σ represents the Sigmoid function, which compresses the output to between 0 and 1; i t Represents the output vector of the input gate, ranging from 0 to 1, which determines how many candidate states are updated To the current cell state; W i 、W C The weight matrix representing the input gate and candidate state; Represents the candidate cell state, new information generated by nonlinear transformation; b i 、b C Represents the bias vector of the input gate and the candidate state; tanh represents the hyperbolic tangent function, which compresses the candidate state to between -1 and 1; C t Indicates the cell state at the current moment; C t-1 Indicates the cell state at the previous moment; · indicates element-by-element multiplication; f t ·C t-1 The forget gate determines how much old information to keep; Indicates that the input gate decides how much new information to add; o t Represents the output vector of the output gate, which takes a value from 0 to 1 and determines how much of the cell state is exposed to the hidden state; W o Represents the weight matrix of the output gate; h t Represents the hidden state of the current moment, short-term memory, as output or passed to the next moment; tanh(C t) means compressing the cell state to between -1 and 1; o t tanh(C t ) indicates that the output gate controls the amount of information exposed.

[0028] Furthermore, the Step 4 includes:

[0029] (1) Generate query, key and value vectors for the data at each time step. The calculation formula for query, key and value vectors is expressed as:

[0030]

[0031] Among them, W i Q 、W i K 、W i V Q represents the learnable weight matrix of the i-th attention head, which is used to generate query, key, and value respectively; i , K i , V i represents the query, key, and value matrix corresponding to the i-th attention head;

[0032] (2) Calculate the similarity score between the query and the key. The calculation formula of the similarity score is expressed as:

[0033]

[0034] Among them, d k is the dimension of the key vector, used to scale the dot product result to stabilize the gradient; Represents the scaling factor to prevent the dot product result from being too large and causing the softmax gradient to disappear; QK T Represents the dot product of the query and the key, calculates the similarity, V represents the value; softmax represents the probability distribution normalized by row, and the weight represents the degree of attention to each position; Attention(Q,K,V) represents the value after weighted summation;

[0035] (3) Use the multi-head attention mechanism to fuse the features of different subspaces; multi-head attention calculates multiple sets of attention values in parallel. The calculation formula of multi-head attention is expressed as:

[0036] MultiHead(Q,K,V)=Concat(head1...,head i ,...head h )W O ;

[0037] Where h is the number of attention heads; headi Represents the output of the i-th attention head, namely Attention(Q i ,K i ,V i ); Concat means concatenating the outputs of multiple heads along the feature dimension; W O is the output projection matrix, which represents the learnable output weight matrix and maps the concatenated result to the final dimension; MultiHead(Q,K,V) represents the result of multi-head attention;

[0038] The output of the multi-head attention mechanism serves as the global feature representation of the encoder.

[0039] Furthermore, the Step 5 includes:

[0040] Normalize the output of the multi-head attention mechanism: retain the input features through residual connections to avoid the gradient disappearance of the deep network; the normalization operation process is expressed as:

[0041] Output=LayerNorm(X+AttentionOutput);

[0042] LayerNorm() represents layer normalization, making the feature distribution of each sample have a mean of 0 and a variance of 1; AttentionOutput represents the result of the multi-head attention obtained in Step 4; Output represents the result obtained after the normalization operation;

[0043] The normalized results are input into the feedforward neural network to perform nonlinear transformation on the features and enhance the feature expression capability. The feedforward network consists of two fully connected layers with an activation function introduced in the middle. The processing process of the input feedforward neural network is expressed as follows:

[0044] FFN(Output)=ReLU(OutputW1+b1)W2+b2;

[0045] FFN(x) represents the output of the feedforward neural network; ReLU() represents the ReLU activation function, which sets all negative values to 0 and keeps positive values unchanged. The role of ReLU is to introduce nonlinearity; W1 is the weight matrix of the first linear transformation, with a dimension of d model ×d f , where d f is the hidden layer dimension of the feedforward neural network; b1 is the bias vector of the first linear transformation, with dimension d f ; W2 is the weight matrix of the second linear transformation, dimension d model ×d f , which maps the output of the hidden layer back to the original input dimension d model; b2 is the bias vector of the second linear transformation, dimension d model .

[0046] Furthermore, the Step 6 includes:

[0047] The embedding representation of the historical trajectory data is calculated by Step 5; the embedding representation of the target trajectory data is obtained from the output of Step 3. The embedding representations of the historical trajectory data and the target trajectory data are input to the decoder: the historical trajectory is captured by the ordinary multi-head attention mechanism. The embedding representation of the target trajectory data is then processed by the masked multi-head attention mechanism to ensure that the current time step only depends on the prediction value of the previous time step.

[0048] Implementation of the mask operation: The weights of future time steps are set to negative infinity, making their attention scores close to zero:

[0049]

[0050] Where M is the mask matrix.

[0051] Furthermore, the Step 7 includes:

[0052] The predicted features output by the decoder are projected through a linear layer to map the high-dimensional features back to the target space. The output prediction value is the flight trajectory at the next moment, and the final output is a complete flight trajectory prediction sequence. The projection matrix W and bias b are defined as:

[0053] Prediction=DecoderOutputW+b;

[0054] Decoderoutput is the hidden state from the last layer of the decoder, that is, the final output of the masked multi-head attention mechanism in Step 6; W represents the weight matrix, which is the parameter learned during the model training process; b represents the bias term, which is the parameter learned during the model training process; Prediction represents the final prediction value, with dimension (T, O), where T represents the time step, which is the number of predicted time points, and O is the number of output features.

[0055] The present invention also provides a hybrid network 4D flight trajectory prediction system based on LSTM and Transformer, which includes: a module for executing the hybrid network 4D flight trajectory prediction method based on LSTM and Transformer.

[0056] This paper replaces the Transformer's Positional Encoding with LSTM. The predefined positional encoding part of the Transformer network is insufficient to capture dynamic context changes during the generation process and cannot be dynamically adjusted to adapt to the temporal dependencies in the trajectory sequence. In 4D flight trajectory prediction, the time intervals of historical trajectory points may be uneven and highly dynamic. Compared with positional encoding, LSTM can more effectively capture the long-range dependencies of time series and adapt to irregular time distributions.

[0057] The beneficial effects of the present invention are:

[0058] 1. Design improvements for time series tasks: Although methods such as Informer, Autoformer, and FEDformer are optimized for time series tasks, their positional encoding still uses traditional static or simple dynamic positional representations, which does not fully consider the dynamic characteristics of the input sequence. In contrast, the LSTM-Transformer architecture of this invention overcomes the shortcomings of existing methods for dynamic time series tasks by dynamically replacing positional encoding.

[0059] 2. Optimization for flight trajectory prediction: This paper combines the dynamic modeling capabilities of LSTM with the global modeling capabilities of Transformer for the first time, adapting to the needs of flight trajectory prediction tasks and improving the modeling capabilities of self-attention. The self-attention mechanism can capture long-range dependencies by weightedly aggregating the global information of the input sequence.

[0060] 3. Model efficiency and performance trade-off: The LSTM network of the present invention significantly improves the ability to model dynamic sequences. When combined with LSTM, the model not only enhances the modeling of local dynamics of sequences but also further improves the ability to capture overall dependencies through the attention mechanism. This combination enables the model to have both the parallelization capabilities of the Transformer and the sequence memory advantages of the LSTM, and excels in complex flight trajectory modeling tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is a model flow chart of the present invention;

[0062] Figure 2 The feature distribution structure of the data set in the present invention;

[0063] Figure 3 This is the LSTM structure diagram in the present invention;

[0064] Figure 4 This is the multi-head attention structure diagram in the present invention;

[0065] Figure 5A bar chart comparing the prediction results of the method of the present invention and other existing methods on the dataset 3U8287Datasets;

[0066] Figure 6 A bar chart comparing the prediction results of the method of the present invention and other existing methods on the dataset MU5576Dataset. DETAILED DESCRIPTION

[0067] Example 1: All experiments of the present invention were carried out on a device with an Intel I5-13500H CPU, an NVIDIA GeForce GTX 4060 GPU, 8G of GPU memory, 64-bit Windows 11 OS, and 16G of RAM; the present invention uses the Anaconda platform to conduct experiments under the Pytorch framework.

[0068] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0069] like Figures 1-6 As shown, a hybrid network 4D flight trajectory prediction method based on LSTM and Transformer, the method includes:

[0070] Step 1: Collect the original flight trajectory data and pre-process the original flight trajectory data;

[0071] Furthermore, the Step 1 includes:

[0072] Extracting original flight trajectory data from flight data records, the original flight trajectory data includes time-space sequence information, and the time-space sequence information includes the aircraft's timestamp, latitude and longitude, altitude, speed, and heading angle;

[0073] Then perform data cleaning, normalization, interpolation, denoising, etc. to eliminate noise and outliers, improve data quality, and ensure that the numerical range of the data is suitable for model training (such as normalization to [0,1] or standardization to zero mean unit variance).

[0074] Specifically, the data used in this paper is all real flight data, sourced from https: / / flightadsb.variflight.com / , a website that records real aircraft flight trajectories every 5 seconds. To demonstrate the effectiveness of the method presented in this paper, we compiled and created flight datasets for two flights: Sichuan Airlines flight 3U8287 and China Eastern Airlines flight MU5576. These data are labeled Data3U8287 and DataMU5576. Both flights represent real flight trajectories from Kunming, China to Zhengzhou, China, but the detailed flight trajectories of the two flights are different. Data3U8287 covers flight trajectory data for 12 months, from April 2023 to April 2024; DataMU5576 covers flight trajectory data from April 2023 to July 2024. The following describes the process of constructing the dataset: (1) First, daily flight trajectory data for each flight is obtained from the website, and the present invention merges the data; (2) missing values and unreasonable data in the data are filled and replaced using the moving average method; (3) the present invention deletes irrelevant features and only retains the timestamp, altitude, speed, heading angle, longitude, and latitude. The structure of the flight dataset of the two flights is shown in Table 1:

[0075] Table 1 shows the data set structure.

[0076]

[0077] Figure 2 This is the experimental data sample structure diagram of the present invention. After the data set is constructed, the present invention analyzes the characteristic structures of the data set:

[0078] from Figure 2 As can be seen from the data, both datasets exhibit strong overall periodicity. The flight altitude ranges from approximately 0 to 10,000 meters; the longitude ranges from approximately 102.5 to 112.5 degrees; and the latitude ranges from approximately 26 to 34 degrees north. Based on the above analysis, the aircraft's altitude, longitude, latitude, and speed are the most important factors influencing the future flight trajectory.

[0079] Step 2: Input the preprocessed flight trajectory data into the embedding layer and map it into a high-dimensional feature representation to form an input sequence; the embedding layer maps the input features into a high-dimensional feature space, allowing the model to capture complex nonlinear relationships.

[0080] Furthermore, in Step 2, the embedding layer adopts a fully connected layer or a feature encoding function, and the feature encoding function performs feature encoding including: temporal feature encoding and spatial feature encoding;

[0081] The time feature encoding converts the timestamp into a periodic feature (such as a time embedding in the form of sine / cosine);

[0082] The spatial feature encoding maps the latitude and longitude information to a fixed-dimensional feature vector through MLP;

[0083] The output of the embedding layer is an embedding feature matrix with a shape of (T, D), where T is the time step and D is the embedding dimension.

[0084] Step 3: Input the input sequence into the LSTM network to extract the context information of the time series, generate an output representation with time-dependent features, and enhance the model's ability to model time series features;

[0085] Furthermore, in Step 3, the key structure of the LSTM network includes an input gate, a forget gate, and an output gate, which can effectively capture the temporal dependency of trajectory data. The LSTM network dynamically stores important timing information through its cell state and filters irrelevant information.

[0086] Furthermore, the Step 3 includes:

[0087] For the input sequence X={x1,x2,...,x T}Calculate the hidden state h through the LSTM network t and cell state c t , the LSTM network outputs a context representation of a time series H = {h1,h2,...,h T}, as a high-dimensional representation of temporal features; hidden state h t and cell state c t The calculation process is as follows:

[0088]

[0089] Among them, f t Represents the output vector of the forget gate, with a value of 0 to 1, which determines how much of the cell state C at the previous moment is retained t-1 ;W f Represents the weight matrix of the forget gate, controlling h t-1 and x t Effect on the degree of forgetting; h t-1 Indicates the hidden state of the previous moment, carrying historical information; x t represents the input vector at time t; [h t-1 ,x t ] means h t-1 and x t The concatenated vector; b fRepresents the bias vector of the forget gate, which adjusts the threshold of the activation function; σ represents the Sigmoid function, which compresses the output to between 0 and 1; i t Represents the output vector of the input gate, ranging from 0 to 1, which determines how many candidate states are updated To the current cell state; W i 、W C The weight matrix representing the input gate and candidate state; Represents the candidate cell state, new information generated by nonlinear transformation; b i 、b C Represents the bias vector of the input gate and the candidate state; tanh represents the hyperbolic tangent function, which compresses the candidate state to between -1 and 1; C t Indicates the current cell state (long-term memory); C t-1 represents the cell state at the previous moment; represents element-wise multiplication (Hadamard product); f t ·C t-1 The forget gate determines how much old information to keep; Indicates that the input gate decides how much new information to add; o t Represents the output vector of the output gate, which takes a value from 0 to 1 and determines how much of the cell state is exposed to the hidden state; W o represents the weight matrix of the output gate; h t Represents the hidden state of the current moment, short-term memory, as output or passed to the next moment; tanh(C t ) means compressing the cell state to between -1 and 1; o t tanh(C t ) indicates that the output gate controls the amount of information exposed.

[0090] Step 4: In the encoder, a multi-head attention mechanism is used to extract global features from the output representation of the LSTM network to capture the dependencies and correlations between different trajectory data points.

[0091] Furthermore, the Step 4 includes:

[0092] (1) Generate query, key and value vectors for the data at each time step. The calculation formula for query, key and value vectors is expressed as:

[0093]

[0094] Among them, W i Q 、W i K 、W i VQ represents the learnable weight matrix of the i-th attention head, which is used to generate query, key, and value respectively; i , K i , V i represents the query, key, and value matrix corresponding to the i-th attention head;

[0095] (2) Calculate the similarity score between the query and the key (ScaledDot-ProductAttention). The calculation formula of the similarity score is expressed as:

[0096]

[0097] Among them, d k is the dimension of the key vector (i.e. K i The size of the last dimension of ), used to scale the dot product result to stabilize the gradient; Represents the scaling factor to prevent the dot product result from being too large and causing the softmax gradient to disappear; QK T Represents the dot product of the query and the key, calculates the similarity (shape is [sequence length, sequence length]), V represents the value; softmax represents the probability distribution normalized by row, and the weight represents the degree of attention to each position; Attention(Q,K,V) represents the value after weighted summation (shape is the same as V);

[0098] (3) Use the multi-head attention mechanism to fuse the features of different subspaces; multi-head attention calculates multiple sets of attention values in parallel. The calculation formula of multi-head attention is expressed as:

[0099] MultiHead(Q,K,V)=Concat(head1...,head i ,...head h )W O ;

[0100] where h is the number of attention heads (e.g. 8 or 16); head i Represents the output of the i-th attention head, namely Attention(Q i ,K i ,V i ); Concat means concatenating the outputs of multiple heads along the feature dimension (the shape becomes [sequence length, h*d_v]); W O is the output projection matrix, which represents the learnable output weight matrix, mapping the concatenated result to the final dimension (shape is [h*d_v, output dimension]); MultiHead(Q,K,V) represents the result of multi-head attention (shape is [sequence length, output dimension]);

[0101] The output of the multi-head attention mechanism is used as the global feature representation of the encoder; among them, LSTM outputs a context representation of the time series H = {h1,h2,...,h T}, as a high-dimensional representation of temporal features.

[0102] The context representation H output by LSTM is input into the multi-head attention mechanism to capture the global dependencies between trajectory points:

[0103] Step 5: Normalize the global features output by the multi-head attention mechanism and further enhance the feature expression capability through a feedforward neural network. This step includes residual connections and layer normalization to ensure model stability and accelerate training convergence.

[0104] Furthermore, the Step 5 includes:

[0105] Normalize the output of the multi-head attention mechanism: Residual connection is used to retain the input features to avoid the gradient disappearance of the deep network; the normalization operation process is expressed as:

[0106] Output=LayerNorm(X+AttentionOutput);

[0107] LayerNorm() represents layer normalization, making the feature distribution of each sample have a mean of 0 and a variance of 1; AttentionOutput represents the result of the multi-head attention obtained in Step 4; Output represents the result obtained after the normalization operation;

[0108] The normalized results are input into the feedforward neural network to perform nonlinear transformation on the features and enhance the feature expression capability. The feedforward network consists of two fully connected layers, and an activation function (such as ReLU) is introduced in the middle. The processing process of the input feedforward neural network is expressed as follows:

[0109] FFN(Output)=ReLU(OutputW1+b1)W2+b2;

[0110] FFN(x) represents the output of the Feed-Forward Neural Network (FFN); ReLU() represents the ReLU (Rectified Linear Unit) activation function, which sets all negative values to 0 and keeps positive values unchanged. The role of ReLU is to introduce nonlinearity; W1 is the weight matrix of the first linear transformation, with dimension d model ×d f , where d fis the hidden layer dimension of the feedforward neural network; b1 is the bias vector of the first linear transformation, with dimension d f ; W2 is the weight matrix of the second linear transformation, dimension d model ×d f , which maps the output of the hidden layer back to the original input dimension d model ; b2 is the bias vector of the second linear transformation, dimension d model .

[0111] Step 6: Use the Masked Multi-Head Attention mechanism in the decoder to output the predicted features to ensure that only the trajectory data of the current time step and before is used for modeling during trajectory prediction to avoid information leakage.

[0112] Furthermore, the Step 6 includes:

[0113] The embedding representation of the historical trajectory data is calculated by Step 5; the embedding representation of the target trajectory data is obtained from the output of Step 3. The embedding representations of the historical trajectory data and the target trajectory data are input to the decoder: the historical trajectory is captured by the ordinary multi-head attention mechanism. The embedding representation of the target trajectory data is then processed by the masked multi-head attention mechanism to ensure that the current time step only depends on the prediction value of the previous time step.

[0114] Implementation of the mask operation: The weights of future time steps are set to negative infinity, making their attention scores close to zero:

[0115]

[0116] Where M is the mask matrix.

[0117] Step 7: Map the predicted features output by the decoder to the target space through a linear layer to generate the final result of the flight trajectory prediction.

[0118] Furthermore, the Step 7 includes:

[0119] The predicted features output by the decoder are projected through a linear layer to map the high-dimensional features back to the target space. The output prediction value is the flight trajectory at the next moment (such as position, speed, or heading angle), and the final output is a complete flight trajectory prediction sequence. The projection matrix W and bias b are defined as:

[0120] Prediction=DecoderOutputW+b;

[0121] Decoderoutput is the hidden state (feature representation) from the last layer of the decoder, that is, the final output of the masked multi-head attention mechanism in Step 6; W represents the weight matrix, which is the parameter learned during model training; b represents the bias term, which is the parameter learned during model training; Prediction represents the final prediction value, with dimension (T, O), where T represents the time step, which is the number of predicted time points, and O is the number of output features, such as the predicted trajectory coordinates (x, y, z) and speed information.

[0122] The present invention also provides a hybrid network 4D flight trajectory prediction system based on LSTM and Transformer, the system comprising:

[0123] A preprocessing module is used to collect and preprocess the original flight trajectory data;

[0124] The input sequence generation module is used to input the pre-processed flight trajectory data into the embedding layer and map it into a high-dimensional feature representation to form an input sequence;

[0125] The output representation generation module is used to input the input sequence into the LSTM network to extract the context information of the time series and generate an output representation with time-dependent features;

[0126] The global feature extraction module is used to extract global features from the output representation of the LSTM network using a multi-head attention mechanism to capture the dependencies and correlations between different trajectory data points;

[0127] Normalization module, which is used to normalize the global features output by the multi-head attention mechanism and further enhance the feature expression capability through the feedforward neural network;

[0128] The predicted feature output module is used to output the predicted features using the masked multi-head attention mechanism in the decoder;

[0129] The flight trajectory prediction module is used to map the prediction features output by the decoder to the target space through a linear layer to generate the final result of the flight trajectory prediction.

[0130] The present invention replaces the Positional Encoding module of Transformer with LSTM. The LSTM module can dynamically capture the long-range dependency characteristics of the time series through the gating mechanism (input gate, forget gate and output gate). LSTM can adjust its state at each time step and dynamically update the memory unit according to the changes in the input sequence to capture the dynamic changes in the flight trajectory (such as speed, direction, and environmental impact). Explicit time dependency modeling: The inherent recursive characteristics of LSTM enable it to effectively model complex temporal relationships in time series, avoiding the limitations brought by the static nature of positional encoding. Compared with static positional encoding, LSTM provides a more flexible way to capture the position and context information of the input sequence.

[0131] To demonstrate the effectiveness of our method, we compared it with the Autoformer, FEDformer, Informer, and Transformer (Non-Positional Encoding). For each method, we set the input model sequence to 96 and evaluated the predictions for 96 steps, with a time difference of 5 seconds between each step, equivalent to predicting the flight trajectory for the next 480 seconds. We also set the number of attention heads to 8, the sliding window to 24, the learning rate to 0.0001, and used the Mean Sequential Error (MSE) as the loss function.

[0132] Historical flight trajectory data is used for model training, and evaluation indicators such as MSE (mean square error) and MAE (mean absolute error) are used to evaluate the model to verify its prediction accuracy and generalization performance.

[0133] Evaluation indicators:

[0134] The present invention uses MAE and MSE as evaluation indicators of the experiment. Mean square error (MSE) and mean absolute error (MAE) are the most commonly used evaluation indicators for regression problems.

[0135] MAE is the mean of the absolute errors between the predicted and observed values:

[0136]

[0137] MSE is the average of the squared differences between the predicted results and the actual targets:

[0138]

[0139] The smaller the two indicators, the better.

[0140] The present invention conducted experiments on two data sets, obtaining the values of the two evaluation metrics, MAE and MSE, based on the predicted and actual trajectories. The present invention conducted a statistical analysis of the errors in altitude, longitude, and latitude in flight trajectory prediction, with the results shown in Tables 2 and 3. The error metrics in the tables show that the present invention (LS-TF) achieved better prediction results than other models for both altitude and longitude and latitude in the 96-step flight trajectory prediction. This demonstrates that the present invention (LS-TF) achieves superior prediction accuracy compared to other models in the flight trajectory prediction task.

[0141] Table 2 shows the experimental results of 3U8287Datasets

[0142]

[0143] Table 3 shows the experimental results of MU5576Datasets

[0144]

[0145] Figure 5 A bar chart comparing the prediction results of the method of the present invention and the Informer, Autoformer, FEDformer, and Transformer (Non Positional Encoding-Transformer) methods of the ablation experiment on the dataset 3U8287Datasets;

[0146] Figure 6 A bar chart comparing the prediction results of the method of the present invention and the Informer, Autoformer, FEDformer, and Transformer (Non Positional Encoding-Transformer) methods of the ablation experiment on the MU5576Dataset;

[0147] The method (LS-TF) shown in the present invention achieves better prediction accuracy than other methods in terms of altitude, longitude, and latitude features. Because the aircraft's altitude change rate is relatively small except during takeoff and landing, each method performs worse than longitude and latitude in terms of altitude. In the experimental results on the 3U8287 dataset, the present invention performs best in terms of MSE, but the MAE is similar to the ablation experiment results due to some outliers and outliers in the dataset. The present invention performs well in all indicators on the MU5576 dataset. Overall, the present invention has a certain effect on improving prediction accuracy in the field of flight trajectory prediction.

[0148] The specific embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.

Claims

1. A 4D flight trajectory prediction method based on a hybrid network of LSTM and Transformer, characterized by: The method comprises: Step 1: Collect the original flight trajectory data and pre-process the original flight trajectory data; Step 2: Input the pre-processed flight trajectory data into the embedding layer and map it into a high-dimensional feature representation to form an input sequence; Step 3: Input the input sequence into the LSTM network to extract the context information of the time series and generate an output representation with time-dependent features; Step 4: Use the multi-head attention mechanism to extract global features from the output representation of the LSTM network to capture the dependencies and correlations between different trajectory data points; Step 5: Normalize the global features output by the multi-head attention mechanism and further enhance the feature expression capability through a feedforward neural network; Step 6: Use the masked multi-head attention mechanism in the decoder to output the predicted features; Step 7: Map the predicted features output by the decoder to the target space through a linear layer to generate the final result of the flight trajectory prediction.

2. The 4D flight trajectory prediction method based on a hybrid network of LSTM and Transformer according to claim 1, characterized in that: Step 1 includes: Extracting original flight trajectory data from flight data records, the original flight trajectory data includes time-space sequence information, and the time-space sequence information includes the aircraft's timestamp, latitude and longitude, altitude, speed, and heading angle; Then perform data cleaning, normalization, interpolation and denoising.

3. The 4D flight trajectory prediction method based on a hybrid network of LSTM and Transformer according to claim 1, characterized in that: In Step 2, the embedding layer uses a fully connected layer or a feature encoding function, and the feature encoding function performs feature encoding including: temporal feature encoding and spatial feature encoding; The time feature encoding converts the timestamp into a periodic feature; The spatial feature encoding maps the latitude and longitude information to a fixed-dimensional feature vector through MLP; The output of the embedding layer is an embedding feature matrix with a shape of (T, D), where T is the time step and D is the embedding dimension.

4. The 4D flight trajectory prediction method based on a hybrid network of LSTM and Transformer according to claim 1, characterized in that: In Step 3, the key structure of the LSTM network includes an input gate, a forget gate, and an output gate; the LSTM network dynamically stores important timing information and filters irrelevant information through its memory unit.

5. The 4D flight trajectory prediction method based on a hybrid network of LSTM and Transformer according to claim 1, characterized in that: Step 3 includes: For the input sequence X={x1,x2,...,x T }Calculate the hidden state h through the LSTM network t and cell state c t , the LSTM network outputs a context representation of a time series H = {h1,h2,...,h T }, as a high-dimensional representation of temporal features; hidden state h t and cell state c t The calculation process is as follows: Among them, f t Represents the output vector of the forget gate, with a value of 0 to 1, which determines how much of the cell state C at the previous moment is retained t-1 ;W f Represents the weight matrix of the forget gate, controlling h t-1 and x t Effect on the degree of forgetting; h t-1 Indicates the hidden state of the previous moment, carrying historical information; x t represents the input vector at time t; [h t-1 ,x t ] means h t-1 and x t The concatenated vector; b f Represents the bias vector of the forget gate, which adjusts the threshold of the activation function; σ represents the Sigmoid function, which compresses the output to between 0 and 1; i t Represents the output vector of the input gate, ranging from 0 to 1, which determines how many candidate states are updated To the current cell state; W i 、W C The weight matrix representing the input gate and candidate state; Represents the candidate cell state, new information generated by nonlinear transformation; b i 、b C Represents the bias vector of the input gate and the candidate state; tanh represents the hyperbolic tangent function, which compresses the candidate state to between -1 and 1; C t Indicates the cell state at the current moment; C t-1 Indicates the cell state at the previous moment; · indicates element-by-element multiplication; f t ·C t-1 The forget gate determines how much old information to keep; Indicates that the input gate decides how much new information to add; o t Represents the output vector of the output gate, which takes a value from 0 to 1 and determines how much of the cell state is exposed to the hidden state; W o Represents the weight matrix of the output gate; h t Represents the hidden state of the current moment, short-term memory, as output or passed to the next moment; tanh(C t ) means compressing the cell state to between -1 and 1; o t tanh(C t ) indicates that the output gate controls the amount of information exposed.

6. The 4D flight trajectory prediction method based on a hybrid network of LSTM and Transformer according to claim 1, characterized in that: Step 4 includes: (1) Generate query, key and value vectors for the data at each time step. The calculation formula for query, key and value vectors is expressed as: Among them, W i Q 、W i K 、W i V Q represents the learnable weight matrix of the i-th attention head, which is used to generate query, key, and value respectively; i , K i , V i represents the query, key, and value matrix corresponding to the i-th attention head; (2) Calculate the similarity score between the query and the key. The calculation formula of the similarity score is expressed as: Among them, d k is the dimension of the key vector, used to scale the dot product result to stabilize the gradient; Represents the scaling factor to prevent the dot product result from being too large and causing the softmax gradient to disappear; QK T Represents the dot product of the query and the key, calculates the similarity, V represents the value; softmax represents the probability distribution normalized by row, and the weight represents the degree of attention to each position; Attention(Q,K,V) represents the value after weighted summation; (3) Use the multi-head attention mechanism to fuse the features of different subspaces; multi-head attention calculates multiple sets of attention values in parallel. The calculation formula of multi-head attention is expressed as: MultiHead(Q,K,V)=Concat(head1...,head i ,...head h )W O ; Where h is the number of attention heads; head i Represents the output of the i-th attention head, namely Attention(Q i ,K i ,V i ); Concat means concatenating the outputs of multiple heads along the feature dimension; W O is the output projection matrix, which represents the learnable output weight matrix and maps the concatenated result to the final dimension; MultiHead(Q,K,V) represents the result of multi-head attention; The output of the multi-head attention mechanism serves as the global feature representation of the encoder.

7. The 4D flight trajectory prediction method based on a hybrid network of LSTM and Transformer according to claim 1, characterized in that: Step 5 includes: Normalize the output of the multi-head attention mechanism: retain the input features through residual connections to avoid the gradient disappearance of the deep network; the normalization operation process is expressed as: Output=LayerNorm(X+AttentionOutput); LayerNorm() represents layer normalization, making the feature distribution of each sample have a mean of 0 and a variance of 1; AttentionOutput represents the result of the multi-head attention obtained in Step 4; Output represents the result obtained after the normalization operation; The normalized results are input into the feedforward neural network to perform nonlinear transformation on the features and enhance the feature expression capability. The feedforward network consists of two fully connected layers with an activation function introduced in the middle. The processing process of the input feedforward neural network is expressed as follows: FFN(Output)=ReLU(OutputW1+b1)W2+b2; FFN(x) represents the output of the feedforward neural network; ReLU() represents the ReLU activation function, which sets all negative values to 0 and keeps positive values unchanged. The role of ReLU is to introduce nonlinearity; W1 is the weight matrix of the first linear transformation, with a dimension of d model ×d f , where d f is the hidden layer dimension of the feedforward neural network; b1 is the bias vector of the first linear transformation, with dimension d f ; W2 is the weight matrix of the second linear transformation, dimension d model ×d f , which maps the output of the hidden layer back to the original input dimension d model ; b2 is the bias vector of the second linear transformation, dimension d model .

8. The 4D flight trajectory prediction method based on a hybrid network of LSTM and Transformer according to claim 1, characterized in that: Step 6 includes: The embedding representation of the historical trajectory data is calculated by Step 5; the embedding representation of the target trajectory data is obtained from the output of Step 3. The embedding representations of the historical trajectory data and the target trajectory data are input to the decoder: the historical trajectory is captured by the ordinary multi-head attention mechanism. The embedding representation of the target trajectory data is then processed by the masked multi-head attention mechanism to ensure that the current time step only depends on the prediction value of the previous time step. Implementation of the mask operation: The weights of future time steps are set to negative infinity, making their attention scores close to zero: Where M is the mask matrix.

9. The 4D flight trajectory prediction method based on a hybrid network of LSTM and Transformer according to claim 1, characterized in that: Step 7 includes: The predicted features output by the decoder are projected through a linear layer to map the high-dimensional features back to the target space. The output prediction value is the flight trajectory at the next moment, and the final output is a complete flight trajectory prediction sequence. The projection matrix W and bias b are defined as: Prediction=DecoderOutputW+b; Decoderoutput is the hidden state from the last layer of the decoder, that is, the final output of the masked multi-head attention mechanism in Step 6; W represents the weight matrix, which is the parameter learned during the model training process; b represents the bias term, which is the parameter learned during the model training process; Prediction represents the final prediction value, with dimension (T, O), where T represents the time step, which is the number of predicted time points, and O is the number of output features.

10. A 4D flight trajectory prediction system based on a hybrid network of LSTM and Transformer, characterized by: The system includes: a module for executing the LSTM and Transformer-based hybrid network 4D flight trajectory prediction method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Indoor environment data prediction method and system based on LSTM-iTransform

    CN120744481A

  • Indoor environment data prediction method and system based on LSTM-iTransformer

    CN120744481B

  • Multi-target track collaborative planning method and system based on hierarchical reinforcement learning

    CN120869165A

  • Transform-based radar trajectory correlation method

    CN121049848A

  • Expressway interleaving area lane change accident risk early warning method and system based on vehicle trajectory data

    CN121305922A