Airport flight delay prediction system and method
Through the dual attention cascade architecture and multi-scale temporal enhancement mechanism, the problem of insufficient spatiotemporal coupling modeling in airport flight delay prediction is solved, high-precision, low-latency intelligent prediction is achieved, and the adaptability and real-time performance of the model are improved.
Patent Information
- Application Number
- CN202510791645.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies have difficulty effectively modeling the dynamic spatiotemporal coupling characteristics of multi-source heterogeneous data in airport flight delay prediction, resulting in insufficient ability to analyze complex delay patterns. Traditional methods are unable to adaptively enhance the local state weights of key areas, ignore the dynamic interaction between time and space characteristics, and lack a business rule embedding mechanism, resulting in insufficient adaptability of prediction results to air traffic control dynamic scheduling strategies.
A dual attention cascade architecture, multi-scale temporal enhancement mechanism and dynamic path discarding strategy are adopted. The data preprocessing module is used to perform spatiotemporal decoupling, and the spatial feature module and temporal feature module are used to process spatial heterogeneity and temporal dependence features respectively. The feature fusion and output module are used for efficient feature fusion, and finally hierarchical attention fusion and lightweight output module are used for prediction.
It significantly improves the accuracy and real-time performance of airport flight delay predictions, reduces the number of model parameters, optimizes inference latency, improves adaptability to multi-source heterogeneous data and prediction accuracy, and meets the needs of airport real-time decision-making.
Smart Images

Figure CN120745899A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of air traffic management and monitoring, and in particular to an airport flight delay prediction system and method. Background Art
[0002] Currently, airport inbound flight anomaly diagnosis and delay warning primarily rely on traditional time series models and basic deep learning techniques, but significant bottlenecks remain in practical application. Traditional methods, such as ARIMA and linear regression models, can fit simple time series patterns using historical data, but they are unable to effectively model the dynamic spatiotemporal coupling characteristics of heterogeneous multi-source data (such as meteorological data, airspace traffic flow, and parking congestion index).
[0003] In recent years, the application of deep learning models (such as CNNs and LSTMs) has partially improved predictive capabilities, but they still have significant limitations. On the one hand, existing methods typically treat temporal and spatial features independently, ignoring the impact of their dynamic interaction ( ), resulting in insufficient analysis of complex delay patterns. On the other hand, traditional CNNs use globally equal-weighted convolution kernels, which cannot adaptively enhance the local state weights of key areas (such as runways and highly congested parking stands). Furthermore, pooling operations can easily blur detailed information (such as gradient mutations at congestion boundaries), resulting in a 15% to 20% increase in spatial heterogeneity modeling errors. Furthermore, in time series modeling, dilated convolutions or unidirectional LSTMs with a single dilation rate struggle to account for both short-term sudden fluctuations (such as 7-minute takeoff queue accumulations) and long-term regularities (hourly traffic accumulation). Furthermore, they lack a mechanism for embedding business rules, resulting in insufficient compatibility between prediction results and dynamic air traffic control scheduling strategies. Summary of the Invention
[0004] The purpose of the present invention is to provide an airport flight delay prediction system and method. Through a dual attention cascade architecture, a multi-scale temporal enhancement mechanism, and a dynamic path discarding strategy, the system systematically addresses the bottlenecks of existing technologies and provides a high-precision, low-latency intelligent solution for abnormal diagnosis and early warning of airport inbound flights.
[0005] To achieve the above-mentioned object, the present invention provides an airport flight delay prediction system, comprising a data preprocessing module, a spatial feature module, a temporal feature module, and a feature fusion and output module;
[0006] The input data is reduced in dimension and screened for features by the data preprocessing module, and a multimodal feature matrix with spatiotemporal decoupling is output. The matrix is then passed through the spatial feature module and the temporal feature module to reduce the interference of redundant features. The spatiotemporal features are then fused by the feature fusion and output module to reduce the number of model parameters and obtain prediction results.
[0007] The input data includes flight dynamic time series data, meteorological time series, airspace flow control signals and parking space congestion index; the data preprocessing module constructs a flight delay propagation chain in the time dimension, extracts time series dependency features, and converts timestamps into minute-level continuous variables; and grid-encodes geographic information in the spatial dimension to generate a spatial feature tensor.
[0008] Among them, the spatial feature module is collaboratively composed of a dual attention mechanism and a dynamic random deep residual network. The dual attention mechanism includes channel attention and spatial attention. The residual module in the dynamic random deep residual network is embedded in the spatial feature module, and adopts a three-stage design of "compression-feature extraction-restoration". Specifically, the number of input channels is halved through convolution to reduce computational complexity; the gradient flow is optimized in combination with the pre-activation structure; and the original channel dimension is restored to ensure information integrity.
[0009] The temporal feature module is dynamically enhanced through multi-expansion rate dilated convolution and bidirectional LSTM priority gating, where the forward LSTM models historical dependencies and the reverse LSTM captures future constraints.
[0010] Among them, the feature fusion and output module uses hierarchical attention fusion to dynamically adjust the feature importance through the spatiotemporal feature weighting formula; the lightweight output replaces the fully connected layer with LayerNorm.
[0011] The present invention also proposes an airport flight delay prediction method, which uses the airport flight delay prediction system and includes the following steps:
[0012] Step 1: Input data and decouple data preprocessing based on spatiotemporal features;
[0013] Step 2: A spatial feature module with a dual attention mechanism is used to address the problem of insufficient spatial heterogeneity;
[0014] Step 3: Perform multi-scale temporal feature extraction and dynamic enhancement;
[0015] Step 4: Use hierarchical attention-guided spatiotemporal feature fusion and affine transformation to output the prediction results.
[0016] The present invention provides an airport flight delay prediction system and method. The system includes a data preprocessing module, a spatial feature module, a temporal feature module, and a feature fusion and output module. First, data preprocessing is performed based on the decoupling of spatial and temporal features to provide normalized input for subsequent modules. Then, the spatial feature module with a dual attention mechanism is used to address the problem of insufficient spatial heterogeneity. The temporal feature module uses a parallel dilated convolution and dynamic pooling strategy to take into account both short-term sudden fluctuations and long-term regularities. Finally, hierarchical attention-guided spatiotemporal fusion is used to improve the recognition accuracy of the model through multi-scale feature splicing of dilated convolution and bidirectional LSTM. The lightweight output driven by affine transformation replaces the fully connected layer with LayerNorm, thereby optimizing the inference delay while ensuring the feature expression capability and completing the delay prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 It is a schematic diagram of the principle architecture of an airport flight delay prediction system of the present invention.
[0019] Figure 2 It is a schematic diagram of the overall combination of spatial feature modules of an airport flight delay prediction system of the present invention.
[0020] Figure 3 It is a schematic diagram of the overall structure of the channel attention of the airport flight delay prediction system of the present invention.
[0021] Figure 4 It is a schematic diagram of the overall structure of the spatial attention of an airport flight delay prediction system of the present invention.
[0022] Figure 5 It is a schematic diagram of the residual module structure of an airport flight delay prediction system of the present invention.
[0023] Figure 6 It is a schematic diagram of the overall framework construction of a time feature module of an airport flight delay prediction system of the present invention.
[0024] Figure 7 It is a schematic diagram of the overall structure of data preprocessing of an embodiment of an airport flight delay prediction method of the present invention.
[0025] Figure 8 It is a schematic diagram of the overall structure of a multi-expansion rate dilated convolution of an airport flight delay prediction method of the present invention.
[0026] Figure 9 This is a schematic diagram of the bidirectional LSTM model structure in an airport flight delay prediction method of the present invention. DETAILED DESCRIPTION
[0027] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.
[0028] See also Figure 1 ,The present invention provides an airport flight delay prediction system, including a data preprocessing module, a spatial feature module, a temporal feature module, and a feature fusion and output module;
[0029] The input data is reduced in dimension and screened for features by the data preprocessing module, and a multimodal feature matrix with spatiotemporal decoupling is output. The matrix is then passed through the spatial feature module and the temporal feature module to reduce the interference of redundant features. The spatiotemporal features are then fused by the feature fusion and output module to reduce the number of model parameters and obtain prediction results.
[0030] The following is combined with Figures 2 to 6 Detailed description of each module:
[0031] 1. Data preprocessing and multimodal feature decoupling
[0032] The input data includes flight dynamic time series data (take-off and landing times, delay records), meteorological time series, airspace flow control signals, and parking stand congestion index. Through spatiotemporal decoupling, the flight delay propagation chain is constructed in the time dimension, time-series dependent features (such as the cumulative effect of delays) are extracted, and the timestamps are converted into minute-level continuous variables. In the spatial dimension, geographic information such as parking stands and runways are grid-encoded to generate spatial feature tensors. Categorical features (such as parking stand numbers) are sparsely one-hot encoded, and numerical features (such as congestion index) are standardized and missing values are interpolated. Finally, a spatiotemporal decoupled multimodal feature matrix is output to provide standardized input for subsequent modules.
[0033] 2. Dual attention cascade and residual combination of spatial feature modules
[0034] The spatial feature module is composed of a dual attention mechanism (CAM→SAM) and a dynamic random depth residual network. Figure 2 shown.
[0035] Channel Attention (CAM): Generate channel descriptors by compressing spatial dimensions through global average pooling, use bottleneck layers (fully connected layers + SiLU activation) to model nonlinear relationships between channels, generate dynamic weights to suppress noise channels (such as secondary meteorological indicators), and enhance key features (such as runway congestion index). The overall structure of channel attention is as follows Figure 3 shown.
[0036] Spatial Attention Detection (SAM): It uses multi-scale convolution kernels to capture local congestion gradients (such as sudden changes in parking space boundaries) and regional anomaly associations (such as the nonlinear coupling of flight density and capacity). It also uses Sobolev space projection to uniformly map discrete events (mechanical failures) and continuous fields (congestion index) into a high-dimensional functional space to solve the dimensional incompatibility problem. The continuous activation function generates a spatial mask to dynamically enhance high-delay risk areas (such as parking spaces with a capacity-occupancy ratio > 1). The overall structure of spatial attention is as follows: Figure 4 shown.
[0037] The residual module is embedded in the spatial feature module and adopts a three-stage design of "compression-feature extraction-restoration". The number of input channels is halved through convolution to reduce computational complexity; the pre-activation structure (Conv-BN-SiLU) is combined to optimize gradient flow; the original channel dimension is restored to ensure information integrity. The residual block calculation path is randomly skipped with a probability of 0.8. The structure of the residual module is as follows Figure 5 shown.
[0038] 3. Multi-scale dynamic enhancement of temporal feature modules
[0039] The temporal feature module achieves dynamic enhancement through multi-expansion rate dilated convolution and bidirectional LSTM priority gating: multi-expansion rate dilated convolution, bidirectional LSTM, and priority gating. The forward LSTM models historical dependencies, while the backward LSTM captures future constraints. Flight priority vectors are embedded in the input and output gates, memory strength is dynamically adjusted, and redundant states are removed through a memory compression bottleneck layer, reducing the number of parameters by 33%. The basic framework of the temporal feature module is as follows: Figure 6 shown.
[0040] 4. Feature fusion and lightweight output
[0041] Hierarchical attention fusion dynamically adjusts feature importance through a spatiotemporal feature weighting formula, increasing attention weight during peak hours. In terms of lightweight output design, affine transformations are used instead of traditional fully connected layers to improve parameter efficiency.
[0042] Furthermore, the present invention also proposes an airport flight delay prediction method, which uses the airport flight delay prediction system and includes the following steps:
[0043] S1: input data, data preprocessing based on decoupling of spatiotemporal features;
[0044] S2: A spatial feature module with dual attention mechanism is used to deal with the problem of insufficient spatial heterogeneity;
[0045] S3: Perform multi-scale temporal feature extraction and dynamic enhancement;
[0046] S4: Use hierarchical attention-guided spatiotemporal feature fusion and affine transformation to output prediction results.
[0047] The following is a further explanation of the execution steps compared with the traditional model:
[0048] The data preprocessing method based on the decoupling of spatiotemporal features in step S1;
[0049] The traditional model has problems such as dimension redundancy, inconsistent dimensions, and ordinal deviation of category features when processing multi-source heterogeneous data (such as flight dynamics, weather, and airspace traffic), which leads to poor model generalization. Figure 7 shown.
[0050] Technical solutions and differences:
[0051] Structure and steps: Data dimensionality reduction removes redundant label columns (such as "anomaly type"), reducing the feature dimension from |F| to |F|-1. For feature differentiation, categorical features (such as parking slots) are sparsely one-hot encoded, and numerical features (such as flight distance) are normalized to construct a bimodal feature space Φcat⊕Φnum. For high-dimensional mapping, categorical features are expanded into a d-dimensional binary space using one-hot encoding, and sparse matrix compression techniques are used to control dimensionality explosion.
[0052] Principle: Eliminate dimensional differences and ordinal deviations through orthogonal feature space, and improve the model's adaptability to multi-source data.
[0053] The spatial feature module of the dual attention mechanism in step S2;
[0054] This invention involves a spatial feature module with a dual-attention mechanism, designed to enhance convolutional neural networks' ability to model airport spatial heterogeneity. This technical solution aims to address the inadequacy of traditional convolutional networks in handling spatial heterogeneity, improving prediction accuracy while reducing the interference of redundant features.
[0055] Here’s how it works:
[0056] CAM suppresses minor meteorological channels (such as humidity) and enhances the runway congestion index. SAM captures apron boundary gradients through multi-scale convolutions and, combined with L2,1 regularization, improves the accuracy of weight allocation in key areas. The random depth mechanism of the residual network reduces training computation by 20%, and the SiLU activation function reduces the decay rate of the gradient norm in deep layers.
[0057] Among them, the channel attention mechanism algorithm is explained in detail as follows:
[0058] (1) Step 2.1.1: Channel statistical information aggregation, whose operation goal is to compress the spatial information of each channel and generate channel-level statistics.
[0059] The mean is calculated for all spatial positions (H×W) of each channel to obtain the global average pooling result:
[0060]
[0061] Physical meaning: Global average pooling can capture the global context information of the channel and suppress redundant spatial details.
[0062] (2) Step 2.1.2: Channel interaction learning, learn the nonlinear relationship between channels through the fully connected layer and generate attention weights. This step is further divided into several sub-steps.
[0063] ① Bottleneck layer: Use a fully connected layer to compress the channel dimension from C to C / r (r is the reduction ratio)
[0064] s=W1·z+b1,W1∈R C×(C / r) ,b1∈R C / r
[0065] ② Activation function (SiLU): Enhances nonlinearity, combines linearity with gating mechanisms, and improves the model's expressiveness.
[0066]
[0067] ③Recovery layer: restore the channel dimension to C through the fully connected layer
[0068] s'=W2·SiLU(s)+b2,W2∈R (C / r)×C ,b2∈R C
[0069] ④Attention gate generation: Use the Sigmoid function to generate channel attention weight g:
[0070] g=σ(s'),g∈[0,1] C
[0071] Physical Implications: This approach establishes high-order dependencies between channels, going beyond simple linear associations; dynamically assigns feature importance to achieve "on-demand enhancement"; and improves the model's robustness to noise and task variations. This process enables the model to autonomously focus on task-relevant features, ultimately improving the performance of downstream tasks.
[0072] (3) Step 2.1.3: Feature recalibration, dynamically adjust the original feature map using attention weights.
[0073] Expand g to the same dimension as X (B×C×H×W) and multiply it channel by channel with the original features.
[0074]
[0075] Physical significance: Enhance the response of important channels, suppress irrelevant channels, and achieve adaptive feature optimization.
[0076] The spatial attention mechanism algorithm is explained in detail as follows:
[0077] (1) Step 2.2.1: Channel statistical information aggregation, whose operation goal is to compress the spatial information of each channel and generate channel-level statistics.
[0078] The mean is calculated for all spatial positions (H×W) of each channel to obtain the global average pooling result:
[0079]
[0080] Physical meaning: Global average pooling can capture the global context information of the channel and suppress redundant spatial details.
[0081] (2) Step 2.2.2: Channel interaction learning, learn the nonlinear relationship between channels through the fully connected layer and generate attention weights. This step is further divided into several sub-steps.
[0082] ① Bottleneck layer: Use a fully connected layer to compress the channel dimension from C to C / r (r is the reduction ratio)
[0083] s=W1·z+b1,W1∈R C×(C / r) ,b1∈R C / r
[0084] ② Activation function (SiLU): Enhances nonlinearity, combines linearity with gating mechanisms, and improves the model's expressiveness.
[0085]
[0086] ③Recovery layer: restore the channel dimension to C through the fully connected layer
[0087] s'=W2·SiLU(s)+b2,W2∈R (C / r)×C,b2∈R C
[0088] ④Attention gate generation: Use the Sigmoid function to generate channel attention weight g:
[0089] g=σ(s'),g∈[0,1] C
[0090] Physical Implications: This approach establishes high-order dependencies between channels, going beyond simple linear associations; dynamically assigns feature importance to achieve "on-demand enhancement"; and improves the model's robustness to noise and task variations. This process enables the model to autonomously focus on task-relevant features, ultimately improving the performance of downstream tasks.
[0091] (3) Step 2.2.3: Feature recalibration: Use attention weights to dynamically adjust the original feature map. Expand g to the same dimensions as X (B×C×H×W) and multiply it with the original features channel by channel.
[0092]
[0093] Physical significance: Enhance the response of important channels, suppress irrelevant channels, and achieve adaptive feature optimization.
[0094] Detailed explanation of the spatial attention mechanism algorithm:
[0095] Step 2.3.1: Multi-scale feature extraction, extract multi-scale context information from the input features.
[0096] ① Micro-feature extraction: Use group convolution (groups = 4) with a 3×3 convolution kernel to extract local detail features.
[0097]
[0098] Physical meaning: Grouped convolution divides the input channels into 4 groups, and each group independently learns local patterns (such as edges and corners)
[0099] ② Macro feature extraction: Use grouped convolution (groups = 2) with a 3×3 convolution kernel to capture global context features.
[0100]
[0101] Physical meaning: Larger groups (groups = 2) allow the convolution kernel to cover a wider receptive field and capture target shape or region-level semantics.
[0102] Step 2.3.2: Sobolev space projection, mapping multi-scale features to Sobolev space to enhance the smoothness and stability of features.
[0103] ① Feature stitching: stitching microscopic and macroscopic features along the channel dimension
[0104] F fused ←Concat(F micro ,F macro )
[0105] Physical meaning: Fusion of local details and global context to provide more comprehensive spatial information.
[0106] ② Projection operation: Apply group normalization (GNorm), SiLU activation function and 3×3 convolution in sequence
[0107] F sobolev =Conv2D 3×3 (SiLU(GNorm(F fused )))
[0108] Spatial significance: The projected features satisfy the properties of the Sobolev space, suppressing high-frequency noise and improving feature continuity.
[0109] Step 2.3.3: Holder continuous activation, generate spatial attention map, and enhance the robustness of activation through Holder continuity.
[0110] ① Spatial averaging: average the projected features in the channel dimension and compress them into a single channel.
[0111]
[0112] Physical meaning: Aggregate multi-channel information and generate preliminary spatial attention distribution.
[0113] ②Holder-Sigmoid operator: combines Sigmoid and custom nonlinearity.
[0114]
[0115] β: controls the decay rate of the activation value. The larger the value, the more focused the attention area.
[0116] α: Adjusts the slope of the hyperbolic tangent (implicit parameter).
[0117] Physical meaning: Holder continuity ensures that the activation function transitions smoothly in space and avoids mutation noise.
[0118] Step 2.3.4: Regularize the system to improve the sparsity and spatial consistency of the attention map through regularization constraints.
[0119] ①Sobolev smoothness regularization: Frobenius norm of constraint weights:
[0120]
[0121] Physical meaning: Suppress drastic changes in the weight matrix to ensure that the model is robust to input perturbations.
[0122] ② Spatial coherence regularization: combining L1 regularization and spatial gradient penalty:
[0123]
[0124] γ1: Encourages the attention map to be sparse and focus on key areas.
[0125] γ2: Force attention to transition smoothly in space to avoid fragmentation.
[0126] Step 2.3.5: Feature modulation, using attention maps to dynamically adjust the original features.
[0127]
[0128] Physical meaning: The features of high attention value areas (A≈1) are enhanced, and the features of low value areas (A≈0) are suppressed.
[0129] The spatial attention mechanism achieves adaptive enhancement of the spatial position of feature maps through multi-scale feature fusion, Sobolev space projection, Holder continuous activation, and regularization constraints. It balances local details with global semantics, suppresses noise, improves model stability, focuses on key task areas, and enhances downstream task performance.
[0130] The technical background of the multi-scale temporal feature extraction and dynamic enhancement module in step S3 is as follows:
[0131] In many time series forecasting tasks, traditional time series feature extraction methods are insufficient in capturing both short-term fluctuations and long-term patterns, and lack the ability to effectively integrate business rules. To address this technical challenge, this paper provides a multi-scale time series feature extraction and dynamic enhancement module that effectively accounts for both short-term fluctuations and long-term patterns, while integrating business rules to improve forecast accuracy and reliability.
[0132] Multi-expansion rate dilated convolution: Parallel dilated convolutions with dilation rates d = 1, 2, and 4 can cover a time window of 7 minutes to 31 minutes. Its function is to capture the instantaneous accumulation of the takeoff queue (d = 1) and the cumulative effect of the morning rush hour (d = 4) through convolution operations with different dilation rates, thereby extracting multi-scale temporal features. The overall structure of the multi-expansion rate dilated convolution is as follows: Figure 8 shown.
[0133] Dynamic pooling: Pooling kernel size adjustment: adaptively adjust the pooling kernel size according to the sequence length L. Through dynamic pooling, the scale of the pooling operation is automatically adjusted according to the length of the input sequence, thereby better adapting to time series data of different lengths and ensuring the effectiveness and accuracy of feature extraction.
[0134] Bidirectional LSTM: Priority code P is embedded in the bidirectional LSTM to modify the input gate and output gate logic. Input gate and output gate formula: i t =σ(W i ·[h t-1 ,x t ]+U i ·P). Where i t is the input gate activation value, σ is the activation function, W i and U i are weight matrices, h t-1 is the hidden state at the previous moment, x t is the current input, and P is the priority code. Priority coding increases the input gate activation value by 42%, thereby strengthening the key state memory and ensuring that the model prioritizes important business rules. The bidirectional LSTM model structure is as follows Figure 9 shown.
[0135] Threshold residual connection: h t =o t ⊙tanh(c' t )+λh t-1 , where ht is the hidden state at the current moment, ot is the output gate activation value, ⊙ represents element-wise multiplication, tanh is the hyperbolic tangent activation function, and c' t is the corrected cell state, and λ is the threshold parameter of the residual connection. The thresholded residual connection alleviates the vanishing gradient problem in long sequences, reducing the gradient norm decay rate from 0.78 to 0.32, improving the model's stability and performance in long sequence modeling. The overall architecture of the time series model is shown in Algorithm 3.
[0136] Working Principle (1) Multi-expansion convolution captures short-term and long-term features: Multi-expansion convolution uses convolution kernels with different expansion rates in parallel to capture the instantaneous accumulation phenomenon of the takeoff queue (d = 1) and the cumulative effect of the morning peak (d = 4), thereby taking into account both short-term fluctuations and long-term regularities. (2) Priority embedding: Priority encoding embedding increases the activation value of the input gate, allowing the model to more effectively remember key state information, follow business rules, and prioritize important events. (3) Residual connection: Gradient vanishing mitigation: The gated residual connection alleviates the gradient vanishing problem of long sequences by transferring hidden states across time steps, making the model more stable and performant in long sequence modeling, and significantly reducing the gradient norm decay rate.
[0137] Detailed explanation of the timing modeling architecture algorithm:
[0138] (1) Step 3.1: Micro-scale feature extraction, capturing local temporal patterns through multi-scale dilated convolution.
[0139] ① Parallel dilated convolution: Use three parallel branches with dilation rate r∈{1,2,4}, and apply one-dimensional convolution to each branch
[0140] Dilation rate r: controls the spacing of the convolution kernel, r = 1 captures the relationship between adjacent time points, and r = 4 captures long-range dependencies.
[0141] Physical significance: Multi-scale convolution jointly models short-term fluctuations and long-term trends.
[0142] ②GBLU activation function: Apply gated bilinear unit (GatedBilinearUnit) to enhance nonlinearity.
[0143] C r =GBLU(C r )=σ(W g C r )⊙(W l C r )
[0144] Application significance: Combine gating mechanism with linear transformation to screen key features and suppress noise.
[0145] ③Conditional maximum pooling: Dynamically adjust the pooling area according to the priority index P
[0146] C r =ConditionalMaxPool(kernel=2)(C r ,P)
[0147] Physical meaning: more details are retained near high-priority time steps, and low-priority areas are downsampled.
[0148] ④ Feature splicing: splice the outputs of the three branches along the feature dimension
[0149]
[0150] Physical significance: Fusion of multi-scale temporal features provides rich contextual information.
[0151] (2) Step 3.2: Priority encoding, mapping the priority index into a learnable embedding vector for dynamic adjustment by LSTM. Mapping the discrete priority index P into a continuous vector.
[0152]
[0153] Embedding dimension dp = 16: compresses priority information into a low-dimensional space to avoid overfitting.
[0154] Physical meaning: The embedding vectors of high-priority time steps encode more significant features.
[0155] (3) Step 3.3: Enhance LSTM, model temporal dependencies through bidirectional LSTM, and integrate priority information.
[0156] ① Bidirectional processing: Forward LSTM: processes the sequence from t = 1 to t = T, capturing historical dependencies; Reverse LSTM: processes the sequence from t = T to t = 1, capturing future dependencies:
[0157] Input Gate:
[0158] Forget Gate:
[0159] Candidate memory:
[0160] Outputs:
[0161] Memory update: c t =f t ⊙c t-1 +i t ⊙g t
[0162] Hide status update:h t =o t ⊙tanh(c t )+h t-1
[0163] ② Layer stacking and bottleneck compression:
[0164] Bidirectional hidden state concatenation:
[0165] Bottleneck layer: compress the dimension to db=64 through MLP
[0166]
[0167] Physical meaning: Reduce redundant information and retain the core features of time series modeling.
[0168] (4) Step 3.4: Attention pooling, dynamically aggregate time-step features, and generate a global context vector.
[0169] ①Context projection: Use the GELU activation function to map the hidden state to a low-dimensional space, which is similar to ReLU but smoother and suitable for probabilistic attention weights.
[0170] ②Attention mechanism: Generate time step weights through Softmax. t Indicates the contribution of the t-th time step to the final prediction.
[0171] α=Softmax(W a Z)∈R B×T
[0172] ③ Time series aggregation: Weighted summation is used to obtain the global context vector, focusing on key time steps (such as stock turning points and voice keywords) and suppressing irrelevant noise.
[0173]
[0174] (5) Step 3.5: Output normalization to stabilize the training process and avoid gradient explosion or gradient disappearance.
[0175] The time series modeling architecture achieves efficient modeling of time series data through multi-scale convolution, priority encoding, bidirectional LSTM, and attention mechanisms. The core physical implications include: jointly capturing local details and global trends, enhancing key time steps based on task requirements, and focusing on core events to improve prediction robustness.
[0176] Hierarchical attention-guided spatiotemporal feature fusion in step S4
[0177] Traditional methods struggle to balance local mutations with global regularity during spatiotemporal feature fusion. Furthermore, the number of parameters in fully connected layers is excessive, leading to high model complexity and low inference efficiency. To address this issue, a hierarchical attention-guided spatiotemporal feature fusion module was proposed. This module aims to efficiently fuse spatiotemporal features, reduce model parameters, improve inference efficiency, and meet real-time requirements.
[0178] Composition structure: (1) Feature splicing: Splice the features extracted by multi-scale dilated convolution, obtain feature information of different scales through dilated convolution, and then splice and integrate them to provide a rich multi-scale feature foundation for subsequent feature fusion. (2) Temporal attention: Calculate the temporal attention weight, and use the obtained temporal attention weight to perform weighted aggregation on the LSTM output, so that the model can focus on key periods in the time series, such as peak periods, to enhance attention to important time information and improve the ability to capture key information in the time dimension during feature fusion. (3) Affine transformation output: Replacing the traditional fully connected layer with an affine transformation can reduce the number of parameters while retaining fine-grained feature information, avoiding information loss, and making the output result more accurately reflect the fused spatiotemporal features.
[0179] Working Principle: (1) Multi-scale features are extracted through dilated convolution, capturing feature information at different spatial and temporal scales, providing a comprehensive foundation for feature fusion. (2) The temporal attention mechanism focuses on key time intervals such as peak hours, with a weight of up to 68%, allowing the model to fully consider the important information of the time dimension during the feature fusion process and achieve effective fusion of spatiotemporal features. (3) The affine transformation retains fine-grained information while reducing the number of parameters, ensuring the accuracy and precision of the output results, and meeting the model's requirements for the fusion effect of spatiotemporal features.
[0180] This solution addresses the core pain points of traditional models in airport delay prediction (such as insufficient spatiotemporal coupling modeling and high computational redundancy) through decoupling spatiotemporal features, enhancing dynamic attention, and utilizing a lightweight output design. Experimental data and principle analysis demonstrate that the synergistic effect of these modules significantly improves prediction accuracy and real-time performance, providing reliable technical support for intelligent airport systems.
[0181] In summary, the present invention has the following beneficial effects:
[0182] 1. Data preprocessing methods and feature optimization
[0183] Through data dimensionality reduction and feature screening, non-critical feature columns (such as redundant labels) are removed, and the feature dimension is reduced from the original |F| to |F|-1, which significantly reduces model input noise, reduces computational complexity, and improves training efficiency by approximately 18%.
[0184] Feature type differentiation and high-dimensional mapping construct an orthogonal feature space through one-hot encoding of categorical features and standardization of numerical features, eliminating dimensional differences and ordinal bias, thereby improving the model's adaptability to multi-source heterogeneous data by 22% and reducing the prediction error (MAE) by 8.5%.
[0185] 2. Spatial Feature Module
[0186] The channel attention mechanism (CAM) dynamically assigns feature channel weights, suppresses interference from redundant features (such as secondary meteorological indicators), and focuses on key factors (such as runway congestion index). Experiments show that its sensitivity to emergencies is increased by 37% and prediction accuracy is improved by 2.5%.
[0187] Spatial Convolutional Attention (SAM) uses multi-scale convolution kernels and Sobolev regularization to capture local congestion gradients (such as sudden changes in apron boundaries). Combined with mixed norm constraints (L2,1 sparsity), it improves the weight allocation accuracy of key areas (such as runways) by 15% and optimizes inference latency by 22%.
[0188] The residual network module introduces a random depth mechanism (survival probability 0.8), which reduces the computational complexity in the training phase by 20%. The SiLU activation function and pre-activation structure design alleviate the gradient vanishing problem, increasing the convergence speed of the deep network by 30% and reducing the test set error by 12.3%.
[0189] 3. Temporal Feature Module
[0190] Multi-scale time series feature extraction uses parallel dilated convolution (d=1, 2, 4) and dynamic pooling strategies, covering a time window of 7 to 31 minutes, taking into account both short-term sudden fluctuations and long-term regularities. The model's MAE in long-sequence prediction tasks is reduced by 14.7%.
[0191] The bidirectional LSTM and priority embedding mechanism explicitly integrate flight priority rules, strengthening the memory strength of high-priority flight status. The prediction error (MAE = 8.2 minutes) is reduced by 19% compared with the unidirectional LSTM, and the rationality of decision-making in business scenarios is improved by 28%.
[0192] Thresholded residual connections and memory compression alleviate information loss in long-term dependencies through cross-timestep residual paths (λ=1) and bottleneck layer dimensionality reduction. The gradient norm decay rate is reduced from 0.78 to 0.32, and the accuracy of attention weight allocation during critical periods (such as take-off peak) is improved by 23%.
[0193] 4. Feature Fusion and Lightweight Output
[0194] Hierarchical attention-guided spatiotemporal fusion combines dilated convolution with multi-scale feature splicing of bidirectional LSTM, combined with Softmax weighting in the time dimension, with a feature redundancy increase of only 9.7%, and the model's recognition accuracy for complex anomalies is improved by 16%.
[0195] The lightweight output driven by affine transformation replaces the fully connected layer with LayerNorm, reducing the number of parameters by 76%. While ensuring feature expression capabilities, the inference latency is optimized to 12.3ms, meeting the real-time requirements of airports and improving deployment efficiency by 58%.
[0196] The above disclosure is merely one or more preferred embodiments of the present invention, and certainly cannot be used to limit the scope of the present invention. A person skilled in the art can understand that all or part of the processes of the above embodiments and equivalent changes made in accordance with the claims of the present invention still fall within the scope of the invention.
Claims
1. An airport flight delay prediction system, characterized in that: Including data preprocessing module, spatial feature module, temporal feature module, feature fusion and output module; The input data is reduced in dimension and screened for features by the data preprocessing module, and a multimodal feature matrix with spatiotemporal decoupling is output. The matrix is then passed through the spatial feature module and the temporal feature module to reduce the interference of redundant features. The spatiotemporal features are then fused by the feature fusion and output module to reduce the number of model parameters and obtain prediction results.
2. The airport flight delay prediction system according to claim 1, wherein: The input data includes flight dynamic time series data, meteorological time series, airspace flow control signals and parking stand congestion index; the data preprocessing module constructs a flight delay propagation chain in the time dimension, extracts time series dependency features, and converts timestamps into minute-level continuous variables; in the spatial dimension, the geographic information is grid-encoded to generate a spatial feature tensor.
3. The airport flight delay prediction system according to claim 2, wherein: The spatial feature module is collaboratively constructed by a dual attention mechanism and a dynamic random deep residual network. The dual attention mechanism includes channel attention and spatial attention. The residual module in the dynamic random deep residual network is embedded in the spatial feature module and adopts a three-stage design of "compression-feature extraction-restoration". Specifically, the number of input channels is halved through convolution to reduce computational complexity; the gradient flow is optimized by combining a pre-activation structure; and the original channel dimension is restored to ensure information integrity.
4. The airport flight delay prediction system according to claim 3, wherein: The temporal feature module is dynamically enhanced through multi-expansion rate dilated convolution and bidirectional LSTM priority gating, where the forward LSTM models historical dependencies and the reverse LSTM captures future constraints.
5. The airport flight delay prediction system according to claim 4, characterized in that: In the feature fusion and output module, hierarchical attention fusion is used to dynamically adjust the feature importance through a spatiotemporal feature weighting formula; the lightweight output replaces the fully connected layer with LayerNorm.
6. A method for predicting airport flight delays, using the airport flight delay prediction system according to any one of claims 1 to 5, characterized in that: The following steps are involved: Step 1: Input data and decouple data preprocessing based on spatiotemporal features; Step 2: A spatial feature module with a dual attention mechanism is used to address the problem of insufficient spatial heterogeneity; Step 3: Perform multi-scale temporal feature extraction and dynamic enhancement; Step 4: Use hierarchical attention-guided spatiotemporal feature fusion and affine transformation to output the prediction results.
Citation Information
Cited By
Flight guarantee time prediction method and device, medium and equipment
CN121279556A