Flight delay prediction method and device based on space-time multi-mode fusion and medium

By combining a layer-by-layer graph convolution model with multimodal feature fusion, the problem of existing technologies failing to fully consider the dynamic changes of airport networks and the influence of multiple factors is solved, thus achieving higher accuracy in flight delay prediction.

CN120849871AActive Publication Date: 2025-10-28CIVIL AVIATION UNIV OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511351832.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-10-28
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing flight delay prediction technologies fail to fully consider the dynamic changes in the mutual influence between airports in the airport network and the combined effects of multiple factors, resulting in insufficient prediction accuracy.

Method used

A spatiotemporal prediction model based on layer-by-layer graph convolution is adopted. Through a temporal feature extraction module, a layer-by-layer graph learning module, a multimodal feature fusion module, and a delay prediction module, multimodal features are fused to predict flight delays by combining airport association network, physical connection matrix, temporal association matrix, and spatial interaction matrix.

Benefits of technology

It improves the accuracy of flight delay prediction, effectively captures the spatiotemporal propagation patterns of flight delays in airport networks, and enhances the precision of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849871A_ABST
    Figure CN120849871A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer technology application, and provides a flight delay prediction method and device based on space-time multi-modal fusion and a medium, and the method comprises the steps: firstly obtaining flight delay associated data and airport basic information of multiple airports in a preset time window, and then generating a node embedding matrix; and three spatial dependence matrixes are constructed based on route physical connection, a historical cooperative delay rate and a time-space relationship. And carrying out time feature mining on the preprocessed flight data through a time feature extraction module, and fusing node embedding and a multi-source spatial dependency matrix by utilizing a layer-by-layer graph learning module to realize hierarchical learning of spatial features. And finally, integrating time and space features through a multi-modal feature fusion module, and inputting the time and space features into a delay prediction module to obtain a flight delay time prediction result. According to the method, through joint modeling of multi-modal spatial-temporal characteristics, the spatial-temporal propagation rule of flight delay in an airport network can be effectively captured, and the flight delay prediction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology applications, and in particular to a method, device and medium for predicting flight delays based on spatiotemporal multimodal fusion. Background Technology

[0002] In the aviation operations system, flight delays have always been a key factor affecting the quality of air service and operational efficiency. Airport networks, as complex systems, exhibit high correlation in both time and space, and the generation and propagation of flight delays are heavily influenced by these spatiotemporal correlations. To address the challenge of flight delay prediction, related research has led to the development of complex spatiotemporal graph convolutional network models, such as Spatiotemporal Graph Convolutional Networks (STGCN) and STSGCN. These models attempt to utilize spatiotemporal graph convolutional networks to uncover the patterns in the spatiotemporal distribution of flight delays in order to predict delay situations. However, existing graph convolutional models have significant shortcomings. On the one hand, they primarily focus on the static correlations between airport nodes, failing to fully consider the dynamic changes in the degree of mutual influence between different airports within the airport network over time. In actual aviation operations, the intensity and manner of mutual influence between airports dynamically change due to factors such as peak flight arrivals and departures, weather changes, and airspace control adjustments at different times, making it difficult for static correlation assumptions to accurately characterize such complex changes. On the other hand, the mutual influence between airports is the result of the combined effects of multiple factors. While existing models involve some factors, they do not comprehensively consider them. The complexity of airport networks means that their mutual influence is not only constrained by geographical distance—airports geographically close by may experience more direct impacts due to flight flow interactions and airspace overlap—but also involves historical delay times. Airports with frequent and interconnected historical delays are prone to cascading effects on each other in subsequent operations. Flight traffic is also a key factor; flight scheduling and resource competition among high-traffic airports significantly alter the degree of mutual influence. Existing models lack the ability to fully explore and apply the synergistic effects of these multiple factors. Therefore, existing flight delay prediction technologies are insufficient in characterizing the spatiotemporal dynamics of airport networks and the comprehensive impact of multiple factors, making it difficult to meet the practical needs of accurately predicting the spatiotemporal distribution of flight delays. Innovative technical solutions are urgently needed to overcome these bottlenecks. Summary of the Invention

[0003] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows: This invention provides a flight delay prediction method based on spatiotemporal multimodal fusion. The method is implemented based on a pre-trained spatiotemporal prediction model using layer-by-layer graph convolution. This spatiotemporal prediction model includes a temporal feature extraction module, a layer-by-layer graph learning module, a multimodal feature fusion module, and a delay prediction module. The method includes the following steps: S100, obtain the prediction dataset and airport basic information; the prediction dataset includes flight delay correlation data for multiple airports within the next w time windows.

[0004] S200, based on the airport basic information, obtain the airport association network, and obtain three spatial dependency matrices based on physical routes, historical collaborative delay rates and spatial relationships, denoted as physical connection matrix, temporal association matrix and spatial interaction matrix respectively.

[0005] S300, the predicted dataset is preprocessed to obtain a preprocessed dataset, which is then input into the time feature extraction module to extract time features and sent to the multimodal feature fusion module.

[0006] S400, the node embedding matrix, physical connection matrix, temporal correlation matrix and spatial interaction matrix of the airport association network are input into the layer-by-layer graph learning module to obtain spatial features through layer-by-layer graph feature learning, and then sent to the multimodal feature fusion module.

[0007] S500, the multimodal feature fusion module is used to fuse the temporal features and the spatial features to obtain multimodal fused features, which are then sent to the delay prediction module to obtain the prediction target corresponding to the prediction dataset, wherein the prediction target is the flight delay time.

[0008] The present invention has at least the following beneficial effects: This invention provides a flight delay prediction method based on spatiotemporal multimodal fusion, implemented using a pre-trained layered graph convolution spatiotemporal prediction model. This model includes a temporal feature extraction module, a layered graph learning module, a multimodal feature fusion module, and a delay prediction module. The method first acquires flight delay correlation data and basic airport information from multiple airports within a preset time window, then generates a node embedding matrix and constructs three spatial dependency matrices based on route physical connections, historical collaborative delay rates, and spatiotemporal relationships. The temporal feature extraction module mines temporal features from the preprocessed flight data, while the layered graph learning module fuses node embeddings and multi-source spatial dependency matrices to achieve hierarchical learning of spatial features. Finally, the multimodal feature fusion module integrates temporal and spatial features and inputs them into the delay prediction module to obtain the flight delay prediction result. This method, through joint modeling of multimodal spatiotemporal features, can effectively capture the spatiotemporal propagation patterns of flight delays in airport networks, improving the accuracy of flight delay prediction.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart of a flight delay prediction method based on spatiotemporal multimodal fusion provided for an embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0014] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of these steps can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the steps can be rearranged. A process can be terminated when its operation is complete, but it may also have additional steps not included in the figures. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.

[0015] This invention provides a flight delay prediction method based on spatiotemporal multimodal fusion. The method is implemented based on a trained spatiotemporal prediction model of layer-by-layer graph convolution. The spatiotemporal prediction model based on layer-by-layer graph convolution includes a time feature extraction module, a layer-by-layer graph learning module, a multimodal feature fusion module, and a delay prediction module.

[0016] In this embodiment of the invention, the temporal feature extraction module includes a temporal convolutional network and an encoder connected in sequence.

[0017] In this embodiment of the invention, a temporal convolutional network is used to extract temporal features from input sequence data and capture long-short-term dependencies. The temporal convolutional network may include a first convolutional module, a second convolutional module, a one-dimensional convolutional layer, and a concatenation module, wherein the first and second convolutional modules are sequentially connected and have identical structures. Each convolutional block includes a sequentially connected dilated causal convolutional layer, a weight normalization layer, a batch normalization layer, an activation function layer, and a dropout layer. The dropout layer of the first convolutional block is connected to the dilated causal convolutional layer of the second convolutional block; that is, the output of the dropout layer of the first convolutional block serves as the input to the dilated causal convolutional layer of the second convolutional block. The one-dimensional convolutional layer and the dropout layer of the second convolutional block are connected to the concatenation module.

[0018] In this embodiment of the invention, the activation function layer uses the ReLU activation function. The concatenation module is used to concatenate the outputs of the one-dimensional convolutional layer and the dropout layer of the second convolutional block. The sequence data input to the temporal convolutional network is divided into two paths: one path expands the receptive field to capture long-term dependencies by expanding the causal convolutional layer, performs weight normalization and batch normalization, and then reduces overfitting through the ReLU activation function and dropout layer; the other path performs residual connection operations. Finally, the outputs of the two paths are added together to obtain the final output of the temporal convolutional network.

[0019] In this embodiment of the invention, the encoder is a transform architecture used to further model complex dependencies in the temporal dimension through a self-attention mechanism. The encoder may include a position encoding module, a multi-head attention mechanism layer, a first combination module, a feedforward neural network layer, and a second combination module connected in sequence.

[0020] The positional encoding module is used to positionally encode the output features of the temporal convolutional network to clarify the relative positional relationship of each element in the sequence. In this embodiment of the invention, sine and cosine positional encoding can be used to positionally encode the output features of the temporal convolutional network. This encoding provides continuous temporal information input while maintaining sensitivity to different time scales, making it an efficient method for processing complex flight data.

[0021] In this embodiment of the invention, the multi-head attention mechanism layer achieves information capture from different semantic subspaces by mapping input features to multiple subspaces (heads) and computing attention in parallel, then concatenating the output results. The first and second combination modules have the same structure, both being a combination of residual connections and layer normalization. The first combination module is connected to the position encoding module and the multi-head attention mechanism layer, and is used to perform layer normalization after residual connection between the outputs of the position encoding module and the multi-head attention mechanism layer. The second combination module is connected to the first combination module and the feedforward neural network, and is used to perform layer normalization after residual connection between the outputs of the first combination module and the feedforward neural network.

[0022] Those skilled in the art should understand that the working principles of multi-head attention mechanisms, residual connections and layer normalization combination (Add&Norm), and feedforward neural network layers are already within the scope of existing technology.

[0023] In this embodiment of the invention, the layer-by-layer graph learning module may include a first layer-by-layer learning network and a second layer-by-layer learning network. The first layer-by-layer learning network includes a first heterogeneous graph convolutional network, a second heterogeneous graph convolutional network, and a third heterogeneous graph convolutional network. The second layer-by-layer learning network includes a first heterogeneous graph feature enhancement network, a second heterogeneous graph feature enhancement network, and a third heterogeneous graph feature enhancement network.

[0024] Furthermore, the delay prediction module may include a bidirectional long short-term memory (LSTM) network and a fully connected layer connected in sequence. The bidirectional LSM network consists of two independent LSM network layers: a forward LSM network layer and a backward LSM network layer. The forward LSM network layer processes the forward sequence to generate forward hidden states, and the backward LSM network layer processes the reverse sequence to generate backward hidden states. The output of the bidirectional LSM network is a concatenated feature obtained by concatenating the forward and backward hidden states.

[0025] Furthermore, such as Figure 1 As shown, the method may include the following steps: S100, obtain the prediction dataset and airport basic information; the prediction dataset includes flight delay correlation data for multiple airports within the next w time windows.

[0026] In this embodiment of the invention, flight delay-related data may include basic flight attributes, time-related features, spatial-related features, operational status, and environmental conditions. The basic attributes include flight number, aircraft tail number, origin airport, destination airport, etc. Time-related features include estimated departure time, departure delay time, aircraft taxiing time, landing gear closure time, landing gear opening time, aircraft taxiing in time, estimated arrival time, estimated flight time, actual flight time, and in-flight time, etc. Spatial-related features include controlled flight distance. Environmental conditions include meteorological parameters of the origin airport and meteorological parameters of the destination airport.

[0027] Those skilled in the art will understand that flight delay data within the next w time windows can be obtained from the flight schedule. Each time window can be one day in length, and the next w time windows constitute the next w days. w can be set based on actual needs and can be an integer greater than or equal to 1; for example, w can be set to 1, 3, or 7, etc.

[0028] In this embodiment of the invention, the airport's basic information may include the airport's geographical location and flight route information, etc.

[0029] S200, based on the airport basic information, obtain the airport association network, and obtain three spatial dependency matrices based on physical routes, historical collaborative delay rates and spatial relationships, denoted as physical connection matrix, temporal association matrix and spatial interaction matrix respectively.

[0030] In this embodiment of the invention, a node in the airport association network represents an airport, and the number of nodes is equal to the number of airports. If there is a flight route between two airports, then there is an edge between the corresponding two nodes. The characteristics of each node may include the airport's average takeoff delay time, average arrival delay time, average taxiing time, average taxiing time, average landing gear opening time, average landing gear closing time, average temperature, average humidity, average wind speed, precipitation, visibility, and weather phenomena over the next w time windows.

[0031] In this embodiment of the invention, weather phenomena can be historical weather phenomena, such as fog, heavy rain, moderate rain, light rain, heavy snow, light snow, tornadoes, funnel clouds, etc. Weather phenomena are represented in vector form, with the vector length equal to the number of weather phenomena. If a certain weather phenomenon occurs, the corresponding position is 1; otherwise, it is 0. The features corresponding to weather phenomena can be a weighted sum of all weather phenomena, and the weight of each weather phenomenon can be determined based on the actual situation.

[0032] Furthermore, the node embedding matrix of the airport association network satisfies the following condition: X = {X1, X2, ..., X} r , ..., X n}

[0033] Where X is the node embedding matrix, X r Let X be the feature vector of the r-th airport node in the airport association network, where r ranges from 1 to n, and n is the number of airports; r ={X r1 X r2 , ..., X rj , ..., X rm}, X rj For X r The j-th feature, X rj =∑ w k=1 w rj (k)•x rj (k), x rj (k) represents the value of the j-th feature in the k-th time window within w future time windows, where j ranges from 1 to m, and m is the number of node features. k ranges from 1 to w, where w... rj (k) is x rj The weight corresponding to (k).

[0034] Furthermore, wrj (k) = x rj (k) / ∑ w k=1 x rp (k). x rp (k) is X r The value of the p-th feature in the k-th time window within w future time windows, where p ranges from 1 to m.

[0035] In this embodiment of the invention, the physical connectivity matrix captures the propagation of flight delays by representing the route connectivity between airports.

[0036] Furthermore, the physical connection matrix satisfies the following condition: Ag = ; Where Ag is the physical connection matrix, α ig Let α be the route identifier between airport i and airport g. If there is a route between airport i and airport g, then α ig =1, otherwise, α ig =0, i takes values ​​from 1 to n, g takes values ​​from 1 to n, and n is the number of airports.

[0037] Historical delay times at airports reflect the propagation characteristics of flight delays between airports. If there are historically large delay times on routes between two airports, then the probability of flight delays propagating between these two airports is also high. Based on this, a time-series correlation matrix based on historical collaborative delay rates is constructed.

[0038] Furthermore, the temporal correlation matrix satisfies the following condition: At= ; Where At is the temporal correlation matrix, β ig Let β represent the impact of airport i on flight delays at airport g over a future time window of w. ig =(F ig -minF ig ) / (maxF ig -minF ig ), F ig Let F be the average of w average arrival delay times for flights from airport i to airport g over w future time windows. ig Let maxF be the minimum of the w average arrival delay times for flights from airport i to airport g over the next w time windows. ig Let $\frac{i}{g}$ be the maximum of the $w average arrival delay times for flights from airport $i$ to airport $g$ over the next $w$ time windows.

[0039] Flight delays rarely occur in isolation; they can spread from one airport to other connected airports, especially when flight traffic is high or the airports are close together, accelerating the spread of delays. By combining traffic flow and distance relationships, a spatial interaction matrix based on spatial relationships can be constructed to comprehensively consider multiple influencing factors between airports.

[0040] Furthermore, the spatial interaction matrix satisfies the following condition: Af = ; Where Af is the spatial interaction matrix, c ig Let c be the normalized flight traffic from airport i to airport g within w time windows in the future. ig =(c ig -minc) / (maxc-minc), where minc is the minimum value in matrix B, maxc is the maximum value in matrix B, and B= , Let d represent the actual flight traffic from airport i to airport g within w time windows in the future. ig Let d be the normalized flight path distance between airport i and airport g. ig =(d ig -mind) / (maxd-mind), where mind is the minimum value in matrix C, maxd is the maximum value in matrix C, and C= , Let be the actual flight path distance between the i-th airport and the g-th airport.

[0041] In this embodiment of the invention, flight traffic refers to the number of flights. Those skilled in the art should understand that if there are no routes between two airports, the corresponding normalized flight traffic is 0. If there are multiple routes between two airports, the route distance between the two airports can be the average of the route distances of the multiple routes. S300, the predicted dataset is preprocessed to obtain a preprocessed dataset, which is then input into the time feature extraction module to extract time features and sent to the multimodal feature fusion module.

[0042] In this embodiment of the invention, preprocessing may include operations such as data cleaning, encoding conversion, and normalization to convert the prediction dataset into a feature vector that can be recognized by a computer.

[0043] Those skilled in the art should understand that the method of using temporal convolutional networks and encoders to extract features from input temporal data falls within the scope of existing technology.

[0044] S400, the node embedding matrix, physical connection matrix, temporal correlation matrix and spatial interaction matrix of the airport association network are input into the layer-by-layer graph learning module to obtain spatial features through layer-by-layer graph feature learning, and then sent to the multimodal feature fusion module.

[0045] In this embodiment of the invention, the layer-by-layer graph learning module combines Graph Convolutional Networks (GCNs) and Graph Attention (GAT) to mine the topological relationships and dynamic influences between airports in an alternating layer-by-layer manner. Specifically, the GCNs capture the spatial dependencies in the airport network, while the GAT dynamically learns the interactions between different airports, optimizing feature learning.

[0046] Furthermore, the S400 specifically includes: S401, the node embedding matrix and the physical connection matrix are input into the first heterogeneous graph convolutional network to obtain the first graph convolutional features, and the node embedding matrix and the physical connection matrix are input into the first heterogeneous graph feature enhancement network to obtain the first graph attention features.

[0047] S402, the first graph convolutional features and the first graph attention features are fused to obtain a first fused feature, and the first fused feature and the temporal correlation matrix are input into a second heterogeneous graph convolutional network to obtain a second graph convolutional feature, and the first fused feature and the temporal correlation matrix are input into a second heterogeneous graph feature enhancement network to obtain a second graph attention feature.

[0048] S403, the second graph convolutional features and the second graph attention features are fused to obtain the second fused features, and the second fused features and the spatial interaction matrix are input into the third heterogeneous graph convolutional network to obtain the third graph convolutional features, and the second fused features and the spatial interaction matrix are input into the third heterogeneous graph feature enhancement network to obtain the third graph attention features; S404, the convolutional features of the third graph and the attention features of the third graph are fused to obtain the third fused feature, which is used as the spatial feature.

[0049] In this embodiment of the invention, the output feature HG1 of the first heterogeneous graph convolutional network is: HG1 = σ(Ag•X•WG1). Wherein, WG1 is the weight matrix of the first heterogeneous graph convolutional network, σ() is the ReLU activation function, and • represents dot product. The output feature HG2 of the second heterogeneous graph convolutional network is: HG2 = σ(At•HG1•WG2). Wherein, WG2 is the weight matrix of the second heterogeneous graph convolutional network. The output feature HG3 of the third heterogeneous graph convolutional network is: HG3 = σ(Af•HG2•WG3). Wherein, WG3 is the weight matrix of the third heterogeneous graph convolutional network.

[0050] In this embodiment of the invention, the output feature HT1 of the first heterogeneous graph feature enhancement network is: HT1 = (HT11, HT12, ..., HT1...). p , ..., HT1 n ),HT1 p The feature of the p-th airport node obtained through the first heterogeneous graph feature enhancement network, where p takes values ​​from 1 to n, is HT1. p =σ(∑ N1(p) q1=1 f1 pq1 •WT1•h1 q1 ), f1 pq1 Let be the attention coefficient between the p-th airport node and its q1-th neighbor node obtained from the first heterogeneous graph feature enhancement network, where q1 ranges from 1 to N1(p), and N1(p) is the number of neighbor nodes of the p-th airport node obtained based on the physical connection matrix. In the physical connection matrix, if α... pe =1 indicates that the e-th airport node is a neighbor of the p-th airport node; otherwise, it is not a neighbor of the p-th airport node. The value of e ranges from 1 to n. Where exp() is the natural exponential function, LeakyReLU() is the LeakyReLU activation function, and f1 T Here, WT1 is the attention vector of the first heterogeneous graph feature enhancement network, and h1 is the weight matrix of the first heterogeneous graph feature enhancement network. p h1 represents the features of the p-th airport node obtained based on X. q1 h1 represents the features of the q1th airport node obtained based on X. u1 Let u1 be the feature of the u1th airport node obtained based on X, where u1 takes values ​​from 1 to N1 (p), and || represents concatenation.

[0051] In this embodiment of the invention, the output feature HT2 of the second heterogeneous graph feature enhancement network is: HT2 = (HT21, HT22, ..., HT2...). p , ..., HT2 n ), HT2 p This refers to the features of the p-th airport node obtained through the second heterogeneous graph feature enhancement network. HT2 p =σ(∑ N2(p) q2=1 f2 pq2 •WT2•h2 q2 f2 pq2Let q2 be the attention coefficient between the p-th airport node and its q2-th neighbor node obtained from the second heterogeneous graph feature enhancement network. The value of q2 ranges from 1 to N2(p), where N2(p) is the number of neighbor nodes of the p-th airport node obtained based on the temporal correlation matrix. In the temporal correlation matrix, if β... pe If ≠0, it means that the e-th airport node is a neighbor of the p-th airport node; otherwise, it is not a neighbor of the p-th airport node.

[0052] in, , where f2 T WT2 is the attention vector of the second heterogeneous graph feature enhancement network, and h2 is the weight matrix of the second heterogeneous graph feature enhancement network. p h2 represents the features of the p-th airport node obtained based on HT1. q2 Based on the features of the q2th airport node obtained from HT1, h2 u2 The u2th airport node is a feature obtained based on HT1, where u2 takes values ​​from 1 to N2 (p).

[0053] In this embodiment of the invention, the output feature HT3 of the third heterogeneous graph feature enhancement network is: HT3 = (HT31, HT32, ..., HT3...). p , ..., HT3 n ), HT3 p This refers to the features of the p-th airport node obtained through the third heterogeneous graph feature enhancement network. HT3 p =σ(∑ N3(p) q3=1 f3 pq3 •WT3•h3 q3 f3 pq3 Let q3 be the attention coefficient between the p-th airport node and its q3th neighbor node obtained from the third heterogeneous graph feature enhancement network, where q3 ranges from 1 to N3(p), and N3(p) is the number of neighbor nodes of the p-th airport node obtained based on the spatial interaction matrix. In the spatial interaction matrix, if c... pe •d pe If ≠0, it means that the e-th airport node is a neighbor of the p-th airport node; otherwise, it is not a neighbor of the p-th airport node.

[0054] in, Among them, f3 T h3 is the attention vector of the third heterogeneous graph feature enhancement network, WT3 is the weight matrix of the third heterogeneous graph feature enhancement network, and h3 is the weight vector of the third heterogeneous graph feature enhancement network. p h3 represents the features of the p-th airport node obtained based on HT2. q3 For the features of the q3th airport node obtained based on HT2, h3u3 The u3th airport node is a feature obtained based on HT2, where u3 takes values ​​from 1 to N3 (p).

[0055] In this embodiment of the invention, the output features of each heterogeneous graph convolutional network and each heterogeneous graph feature enhancement network can be fused using a fusion module. The fusion module may include a fusion layer, a fully connected layer, and a ReLU activation function layer. The fused feature obtained after fusing the output features of each heterogeneous graph convolutional network and each heterogeneous graph feature enhancement network by the fusion module satisfies the following condition: HM = ReLU(Wf•contact(HG,HT) + bf).

[0056] Where HM is the fusion feature, HG is the output feature of each heterogeneous graph convolutional network, HT is the output feature of each heterogeneous graph feature enhancement network, Wf is the weight matrix of the fully connected layer, bf is the bias vector of the fully connected layer, and contact() represents the concatenation operation.

[0057] S500, the multimodal feature fusion module is used to fuse the temporal features and the spatial features to obtain multimodal fused features, which are then sent to the delay prediction module to obtain the prediction target corresponding to the prediction dataset. The prediction target is the flight delay time, specifically the flight arrival delay time.

[0058] In this embodiment of the invention, the multimodal fused feature after the temporal feature and the spatial feature are fused by the multimodal feature fusion module satisfies the following condition: HMS = (Ws•contact(HM, HS) + bs). HMS is the multimodal fused feature, HS is the temporal feature, Ws is the weight matrix of the multimodal feature fusion module, and bs is the bias vector of the multimodal feature fusion module.

[0059] Those skilled in the art should understand that the method of using bidirectional long short-term memory networks and fully connected layers to process input multimodal fusion features to obtain flight delay times falls within the scope of existing technology.

[0060] In summary, the flight delay prediction method based on spatiotemporal multimodal fusion provided in this invention significantly enhances the ability to capture time-series features in flight delay prediction by combining temporal convolutional networks and the Transformer mechanism. Furthermore, the model introduces heterogeneous graph convolutional networks and graph attention networks. By constructing multi-type airport graph structures, it can better simulate the interrelationships and influences between airports. A specific aggregation method is used for node embedding, deeply mining the topological relationships and dynamic influences between airports in a layer-by-layer alternating manner. Through this layer-by-layer alternating mechanism, the model can capture more fine-grained dynamic change information at different levels. Finally, by fusing the extracted spatiotemporal features into a unified high-dimensional feature representation and inputting it into a bidirectional long short-term memory network for flight delay prediction, the model can improve the accuracy of flight delay prediction.

[0061] In this embodiment of the invention, the trained spatiotemporal prediction model of layer-by-layer graph convolution can be obtained by training the initial spatiotemporal prediction model using a sample dataset. In this embodiment, the sample dataset can be a publicly available dataset. The flight-related data is flight punctuality data provided by the U.S. Bureau of Transportation Statistics, with the selected data being historical flight data for the entire year of 2023, totaling over 6 million flight records. Historical weather data is sourced from the U.S. National Oceanic and Atmospheric Administration (NOAA). Because the dataset itself contains a large amount of data with missing features and outliers, rows with missing weather data are filled with the average of the three days before and after the missing weather data; if there is no data for the three days before and after, the row is deleted. For flight data, rows containing outliers are directly deleted.

[0062] During training, the input time series window length can be set to 7 days, meaning that each batch of sample data consists of 7 days of data. The prediction target time window length can be set to 1 day, 3 days, or 5 days, etc.

[0063] Those skilled in the art should understand that the method of training an initial layer-by-layer graph convolutional spatiotemporal prediction model using a sample dataset falls within the scope of existing technology.

[0064] The flight delay prediction method based on spatiotemporal multimodal fusion provided in this invention, taking a sample dataset as an example, reduces the prediction error by approximately 3% compared to existing prediction models.

[0065] This invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being configured to perform the method described in this invention.

[0066] This invention also provides a computer-readable storage medium storing computer-executable instructions for performing the methods described in this invention.

[0067] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0068] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A flight delay prediction method based on spatiotemporal multimodal fusion, characterized in that, The method is based on a pre-trained spatiotemporal prediction model using layer-by-layer graph convolution. This model includes a temporal feature extraction module, a layer-by-layer graph learning module, a multimodal feature fusion module, and a delay prediction module. The method comprises the following steps: S100, Obtain the prediction dataset and basic airport information; the prediction dataset includes flight delay correlation data for multiple airports within the next w time windows; S200, based on the airport basic information, obtain the airport association network, and obtain three spatial dependency matrices based on physical routes, historical collaborative delay rates and spatial relationships, which are respectively denoted as physical connection matrix, temporal association matrix and spatial interaction matrix; S300, the prediction dataset is preprocessed to obtain a preprocessed dataset, which is then input into the time feature extraction module to extract time features and sent to the multimodal feature fusion module; S400, the node embedding matrix, physical connection matrix, temporal correlation matrix and spatial interaction matrix of the airport association network are input into the layer-by-layer graph learning module to obtain spatial features through layer-by-layer graph feature learning, and then sent to the multimodal feature fusion module; S500, the multimodal feature fusion module is used to fuse the temporal features and the spatial features to obtain multimodal fused features, which are then sent to the delay prediction module to obtain the prediction target corresponding to the prediction dataset, wherein the prediction target is the flight delay time.

2. The method according to claim 1, characterized in that, The layer-by-layer graph learning module includes a first layer-by-layer learning network and a second layer-by-layer learning network. The first layer-by-layer learning network includes a first heterogeneous graph convolutional network, a second heterogeneous graph convolutional network, and a third heterogeneous graph convolutional network. The second layer-by-layer learning network includes a first heterogeneous graph feature enhancement network, a second heterogeneous graph feature enhancement network, and a third heterogeneous graph feature enhancement network.

3. The method according to claim 2, characterized in that, The S400 specifically includes: S401, the node embedding matrix and the physical connection matrix are input into the first heterogeneous graph convolutional network to obtain the first graph convolutional features, and the node embedding matrix and the physical connection matrix are input into the first heterogeneous graph feature enhancement network to obtain the first graph attention features; S402, the first graph convolutional features and the first graph attention features are fused to obtain the first fused features, and the first fused features and the temporal correlation matrix are input into the second heterogeneous graph convolutional network to obtain the second graph convolutional features, and the first fused features and the temporal correlation matrix are input into the second heterogeneous graph feature enhancement network to obtain the second graph attention features; S403, the second graph convolutional features and the second graph attention features are fused to obtain the second fused features, and the second fused features and the spatial interaction matrix are input into the third heterogeneous graph convolutional network to obtain the third graph convolutional features, and the second fused features and the spatial interaction matrix are input into the third heterogeneous graph feature enhancement network to obtain the third graph attention features; S404, the convolutional features of the third graph and the attention features of the third graph are fused to obtain the third fused feature, which is used as the spatial feature.

4. The method according to claim 1, characterized in that, The temporal feature extraction module includes a temporal convolutional network and an encoder connected in sequence.

5. The method according to claim 1, characterized in that, The delay prediction module includes a bidirectional long short-term memory network and a fully connected layer connected in sequence.

6. An electronic device, characterized in that, Including processor and memory; The processor executes the steps of the method as described in any one of claims 1 to 5 by invoking programs or instructions stored in the memory.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a program or instructions that cause a computer to perform the steps of the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Flight operation network delay prediction method based on deep learning combination model

    CN116205120A

  • Multi-airport flight delay prediction method based on space-time diagram convolutional neural network

    CN116862061A

  • Air route network flow prediction method based on spatio-temporal feature fusion

    CN117058927A