A train delay prediction method based on multiple time scales
Through multi-time-scale data processing methods, combined with graph attention and channel attention, the accuracy problem of train delay prediction in complex networks with multiple stations in the existing technology is solved, and higher prediction accuracy and real-time performance are achieved.
Patent Information
- Application Number
- CN202510193540.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Existing train delay prediction methods have difficulty in effectively capturing spatiotemporal correlation and nonlinear complexity when dealing with multi-station and complex railway networks, resulting in low prediction accuracy.
A multi-time-scale approach is adopted to obtain spatiotemporal embedding representation through multi-source data fusion, and fast Fourier transform is used to identify periodicity. Graph attention and multi-time-scale channel attention are combined to capture spatiotemporal correlation and perform decomposition and aggregation to obtain delay prediction results.
It improves the accuracy of train delay prediction under multi-station and complex railway network conditions, can better consider the influence of factors at different time scales, and improves the scientificity and real-time nature of the prediction.
Smart Images

Figure CN120124794B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of transportation technology, and in particular to a train delay prediction method based on multiple time scales. Background Art
[0002] By predicting train delays at multiple stations in advance, train dispatchers can be provided with a basis for pre-planning and train scheduling, formulating more scientific operation plans, and thus improving train punctuality. Currently, train delay prediction methods can be divided into three categories: statistical prediction methods, traditional machine learning prediction methods, and neural network prediction methods.
[0003] Statistical train delay prediction uses statistical models to analyze historical train operation data, uncovering delay patterns and patterns, and thus predicting future train operation status. Statistical train delay prediction methods primarily include multivariate logistic regression, linear regression, lognormal distribution models, probability distribution models, and analytical stochastic models. Statistical methods offer significant advantages in train delay prediction. By establishing mathematical relationships between variables, they provide efficient and highly interpretable models that clearly demonstrate the key factors and mechanisms influencing delays. Statistical models offer a simple structure and low computational cost, making them suitable for simple train operation scenarios with relatively small data sets. However, these simplicity makes it difficult to model and analyze nonlinearities and complex dependencies within the data. Generally, the severity of train delays can only be analyzed based on the distribution presented in the data. Statistical methods also lack real-time and adaptability, hindering timely updates in dynamic and complex train delay scenarios, and their generalizability is limited.
[0004] Traditional machine learning methods are the most popular approach in train delay prediction. Traditional machine learning models primarily include tree models, Markov models, and Bayesian networks. Tree models, primarily random forests, employ an ensemble of multiple decision trees to capture nonlinear relationships between independent and dependent variables. They are robust against high-dimensional and diverse train operation data and offer excellent results in predicting train delays on specific routes. However, the large number of trees in the ensemble increases model complexity and training time. Markov models use a state transition matrix to describe the transition patterns between train operating states and have a simple structure. They can model the relationships between train events (such as travel time and dwell time) and effectively capture the dependencies between stations and within stations in train operation time series. By defining a dynamic evolution matrix between stations, they are often used to predict the delay status of a train at the next station. Because Markov models assume that the current state is only related to the previous state, they ignore the long-term dependencies and delay evolution patterns in historical train operation data, failing to effectively capture the cumulative effects of delays during train operation. In addition, the basic Markov model only uses train schedule data to construct the state transition matrix and does not introduce external variables. Therefore, it has poor adaptability for train delay prediction in complex scenarios. This shortcoming needs to be compensated by combining the hidden Markov model or other extension methods. The Bayesian network represents the possibility of prediction results through probability distribution and can be used to predict train delays in real time. The network can effectively represent and calculate complex probabilistic reasoning between random variables, clarify the causal relationship between variables, and help analyze the key factors affecting train delays. However, when dealing with high-dimensional data or large-scale railway network systems (such as multi-station train networks), the Bayesian network faces greater difficulties in model structure design and parameter learning.
[0005] Because train operation time series data exhibits temporal dependence and spatial correlation, statistical and traditional machine learning methods struggle to capture the temporal and spatial correlations in the data. Compared to traditional methods, neural networks can automatically learn complex nonlinear relationships, making them suitable for multi-factor train delay prediction and effectively mining underlying patterns and regularities in complex traffic data. The extreme learning machine (ELM) neural network architecture is adaptable to diverse datasets, making it effective across diverse railway networks and conditions. It can be easily extended to deeper networks with more hidden layers, improving model accuracy and capturing more complex features. Long short-term memory (LSTM) networks, through input, forget, and output gates, selectively memorize event information in time series, facilitating the capture of long-term dependencies within the sequence. Another advantage of neural networks is their flexible integration of different architectures to form hybrid models, which effectively handle heterogeneous and multi-attribute data in dynamic railway systems. Combining LSTM with convolutional neural networks (CNNs) can capture not only long-term dependencies in train operation data but also spatial correlations between stations. Combining LSTM with fully connected neural networks (FCNNs) can capture interactions between train operation and weather data. Heterogeneous models combined with graph convolutional neural networks (GNNs) overcome the limitations of CNNs in capturing non-Euclidean static data, improving the accuracy of train delay prediction. While these neural network models can capture temporal and spatial correlations within sequences, they only predict train delays at a single time scale. This makes it difficult to extract the underlying periodicity and complex correlations between features in the data, making them unsuitable for predicting train delays at multiple stations. Summary of the Invention
[0006] The technical problem addressed by this invention is how to improve the accuracy of train delay predictions in multi-station, complex railway networks using a multi-time-scale analysis method, while taking into account the spatiotemporal correlation of train events and the nonlinear complexity of influencing factors. Compared to single-time-scale prediction methods, a multi-time-scale train delay prediction method can address train delay issues by considering factors at different scales, such as short-term, medium-term, and long-term cyclical influences. By using multiple time scales for prediction and integrating information from multiple time scales, the model can improve the accuracy of multi-station delay predictions.
[0007] To achieve the above object, the present invention provides the following solutions:
[0008] A train delay prediction method based on multiple time scales, comprising:
[0009] Perform multi-source data fusion on train operation data to obtain a spatiotemporal embedding representation of the data tensor; wherein the spatiotemporal embedding representation includes: a time embedding representation, a space embedding representation, and a feature embedding representation;
[0010] Embedding the space-time representation, identifying periodicity in the time series through fast Fourier transform, and dividing information at multiple time scales;
[0011] Capturing the spatiotemporal correlation of the spatiotemporal embedding representation and generating a spatiotemporal correlation feature representation;
[0012] Based on the information of the multiple time scales, the spatiotemporal correlation feature representation is decomposed into multiple time scales, and channel attention learning is performed on the decomposed information to obtain aggregated multi-time scale sequence information;
[0013] Obtain delay prediction results based on the aggregated multi-time scale sequence information.
[0014] Optionally, the train operation data includes:
[0015] Train operation related information: timestamp, train number, planned running time, planned stop time, actual running time, actual stop time, running buffer time, stop buffer time, train arrival delay, train departure delay;
[0016] Site related information: site number;
[0017] Environmental related information: temperature, weather.
[0018] Optionally, obtaining a spatiotemporal embedding representation of a data tensor includes:
[0019] Perform convolution operations on the train running data to model the features and project the tensors into a high-dimensional space to obtain feature embedding representations;
[0020] The trains in the train operation data are sorted according to the order of arrival. The position encoding operation is used to assign position information to the tensor, thereby identifying the relative position relationship between the train and the station and obtaining a spatial embedding representation.
[0021] The timestamp information in the train operation data is divided into four granularities according to hour, week, month, and year. The four timestamp information are encoded separately to obtain time embedding representation;
[0022] The temporal embedding representation is integrated with the spatial embedding representation and the feature embedding representation to obtain the spatiotemporal embedding representation of the data tensor.
[0023] Optionally, dividing into multiple time scales includes:
[0024] The train operation data is converted from the time domain to the frequency domain using fast Fourier transform. The average amplitude of the data is calculated from the frequency domain analysis to obtain the top K most prominent frequency values. Based on the relationship between frequency and period, the K most significant time scales and the corresponding average amplitude values of the time scales are obtained.
[0025] Optionally, capturing the spatiotemporal correlation of the spatiotemporal embedding representation to generate a spatiotemporal correlation feature representation includes:
[0026] Performing multi-head linear attention processing on the spatiotemporal embedding representation to obtain an enhanced feature representation;
[0027] Adaptive graph convolution is performed on the enhanced feature representation to obtain feature representation containing spatiotemporal correlations.
[0028] Optionally, performing multi-head linear attention processing on the spatiotemporal embedding representation includes:
[0029] Performing convolution operations on the spatiotemporal embedding representation to capture local temporal dependencies in the train run sequence;
[0030] Perform linear transformation on the feature representation after convolution operation;
[0031] Perform multi-head linear attention processing on the linearly transformed feature representation to capture the correlation between trains;
[0032] The multi-layer perceptron is used to learn the nonlinear relationship of the feature representation after attention processing to obtain the enhanced feature representation.
[0033] Optionally, performing adaptive graph convolution processing on the enhanced feature representation includes:
[0034] Based on the enhanced feature representation, an adaptive graph matrix is generated according to the number of features of the site, and the negative values in the graph matrix are set to 0 using the ReLU activation function;
[0035] Use the SoftMax function to normalize the weights between the nodes of the graph matrix to generate an adaptive adjacency matrix;
[0036] The adaptive adjacency matrix and the enhanced feature representation are multiplied using the Einstein summation convention to obtain the multiplied information;
[0037] The multiplied information is subjected to multi-layer graph convolution operations to aggregate information from multiple sites and obtain feature representations that include spatiotemporal correlations.
[0038] Optionally, based on the information of the multiple time scales, decomposing the spatiotemporal correlation feature representation into multiple time scales, and learning channel attention on the decomposed information includes:
[0039] Based on the information of the multiple time scales, the spatiotemporal correlation feature representation is divided into K time scales, each different time scale contains different periodic information, the data of the K time scales are converted into the frequency domain using DCT, and the channel attention weights are solved for the sequences of different time scales;
[0040] A SoftMax operation is performed on the average amplitude value of the channel attention weight to solve the weight value of each time scale, and the sequence information of multiple time scales is aggregated according to different weight information.
[0041] The beneficial effects of the present invention are:
[0042] The present invention first fuses multi-source train operation data to obtain a spatiotemporal embedding representation of the data tensor. Secondly, using the spatiotemporal embedding representation, Fast Fourier Transform (FFT) is used to identify periodicity in the time series and partition information across multiple time scales. The spatiotemporal embedding representation is then used to capture spatiotemporal correlations and generate spatiotemporal correlation feature representations. Finally, based on the information across multiple time scales, the spatiotemporal correlation feature representations are aggregated to obtain aggregated multi-time scale sequence information. Finally, based on this aggregated multi-time scale sequence information, delay prediction results are obtained. This invention can accurately predict train delays in complex railway networks with multiple stations. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0044] Figure 1 The overall network structure of a train delay prediction method based on multiple time scales according to an embodiment of the present invention;
[0045] Figure 2 This is the GA module structure of an embodiment of the present invention. DETAILED DESCRIPTION
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0047] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0048] This embodiment proposes a train delay prediction method based on multiple time scales, including:
[0049] The train operation data is fused with multiple sources to obtain the spatiotemporal embedding representation of the data tensor; the spatiotemporal embedding representation includes: time embedding representation, space embedding representation, and feature embedding representation;
[0050] Embedding time and space into representations, identifying periodicity in time series through fast Fourier transform, and dividing information at multiple time scales;
[0051] For spatiotemporal embedding representation, capture spatiotemporal correlation and generate spatiotemporal correlation feature representation;
[0052] Based on information from multiple time scales, the spatiotemporal feature representation is decomposed into multiple time scales, and channel attention is learned on the decomposed information to obtain aggregated multi-time scale sequence information.
[0053] Obtain delay prediction results based on the aggregated multi-time scale sequence information.
[0054] Specifically, the purpose of this embodiment is to realize multi-station delay prediction of trains by capturing the potential correlation between multiple complex factors through the proposed train delay prediction method based on multiple time scales. Specifically: a heterogeneous model combining a graph attention (GA) module and a multi-time-scale channel attention (MTSCA) module is proposed. The train operation data is enhanced by fusing train schedule information, equipment and train interaction information and external environment information to reflect the relationship between train delays and internal and external factors. Predictions are made using multiple time scales, taking into account factors of different scales, such as short-term, medium-term and long-term cyclical effects. By fusing multi-scale information, the problems of difficulty in capturing the potential correlation between multiple complex factors and low accuracy of multi-station delay prediction are solved.
[0055] Furthermore, the train operation data includes:
[0056] Train operation related information: timestamp, train number, planned running time, planned stop time, actual running time, actual stop time, running buffer time, stop buffer time, train arrival delay, train departure delay;
[0057] Site related information: site number;
[0058] Environmental related information: temperature, weather.
[0059] Furthermore, obtaining the spatiotemporal embedding representation of the data tensor includes:
[0060] Perform convolution operations on the train running data to model the features and project the tensors into a high-dimensional space to obtain feature embedding representations;
[0061] The trains in the train operation data are sorted according to the order of arrival. The position encoding operation is used to assign position information to the tensor, thereby identifying the relative position relationship between the train and the station and obtaining a spatial embedding representation.
[0062] The timestamp information in the train operation data is divided into four granularities according to hour, week, month, and year. The four timestamp information are encoded separately to obtain time embedding representation;
[0063] The temporal embedding representation is integrated with the spatial embedding representation and the feature embedding representation to obtain the spatiotemporal embedding representation of the data tensor.
[0064] Furthermore, the division into multiple time scales includes:
[0065] The train operation data is converted from the time domain to the frequency domain using fast Fourier transform. The average amplitude of the data is calculated from the frequency domain analysis to obtain the top K most prominent frequency values. Based on the relationship between frequency and period, the K most significant time scales and the corresponding average amplitude values of the time scales are obtained.
[0066] Furthermore, the spatiotemporal embedding representation captures the spatiotemporal correlation and generates spatiotemporal correlation feature representations including:
[0067] Perform multi-head linear attention processing on the spatiotemporal embedding representation to obtain enhanced feature representation;
[0068] Adaptive graph convolution is performed on the enhanced feature representation to obtain feature representation containing spatiotemporal correlations.
[0069] Furthermore, multi-head linear attention processing is performed on the spatiotemporal embedding representation, including:
[0070] Perform convolution operations on the spatiotemporal embedding representation to capture local temporal dependencies in the train run sequence;
[0071] Perform linear transformation on the feature representation after convolution operation;
[0072] Perform multi-head linear attention processing on the linearly transformed feature representation to capture the correlation between trains;
[0073] The multi-layer perceptron is used to learn the nonlinear relationship of the feature representation after attention processing to obtain the enhanced feature representation.
[0074] Furthermore, performing adaptive graph convolution processing on the enhanced feature representation includes:
[0075] Based on the enhanced feature representation, an adaptive graph matrix is generated according to the number of features of the site, and the negative values in the graph matrix are set to 0 using the ReLU activation function;
[0076] Use the SoftMax function to normalize the weights between the nodes of the graph matrix to generate an adaptive adjacency matrix;
[0077] The adaptive adjacency matrix and the enhanced feature representation are multiplied using the Einstein summation convention to obtain the multiplied information;
[0078] The multiplied information is subjected to multi-layer graph convolution operations to aggregate information from multiple sites and obtain feature representations that include spatiotemporal correlations.
[0079] Furthermore, based on information at multiple time scales, the spatiotemporal feature representation is decomposed at multiple time scales, and channel attention learning is performed on the decomposed information, including:
[0080] Based on information from multiple time scales, the spatiotemporal feature representation is divided into K time scales. Each different time scale contains different periodic information. The data of K time scales are transformed into the frequency domain using DCT, and the channel attention weights are solved for sequences of different time scales.
[0081] A SoftMax operation is performed on the average amplitude value of the channel attention weight to solve the weight value of each time scale, and the sequence information of multiple time scales is aggregated according to different weight information.
[0082] More specifically, if Figure 1 As shown, the train delay prediction method based on multiple time scales proposed in this embodiment mainly consists of an Embedding module, an FFT for Patch module, a GA module and an MTSCA module.
[0083] The Embedding module consists of three main parts: Convolution (Conv) operations, Position Embedding (P-Embed), and Time Embedding (T-Embed). Convolution is a one-dimensional convolution operation with a kernel size of 3 and a stride of 1. It primarily increases the dimensionality of the data while capturing the temporal dependencies between various factors affecting train operation. Position Embedding uses position encoding to model the relative positions of trains. Time Embedding is a global, learnable timestamp embedding that divides train time features into four granularities: hour, week, month, and year. It then performs Embedding operations on different timestamps and fuses them together. Feature embedding is combined with spatial and temporal embedding to create a spatiotemporal embedding. This effectively extracts spatiotemporal features and correlations between variables in train operation data.
[0084] The FFT for Patch module is a multi-time-scale identification module for train operation data. It uses the Fast Fourier Transform (FFT) to detect K prominent periodicities as time scales, thereby scaling the train operation time series. The Fast Fourier Transform (FFT) is used to convert train operation data from the time domain to the frequency domain. The average amplitude of the data is calculated from the frequency domain to determine the top K most prominent frequency values. Based on the relationship between frequency and period, the K most significant time scales and their corresponding average amplitudes are then determined. The FFT for Patch module can adaptively identify the K most significant time scales in historical data, avoiding the subjective division of data periods that could disrupt the data's periodicity.
[0085] The GA module is a graph attention module, which consists of a linear attention module MLLA similar to Mamba and an adaptive graph convolution AGCN. Its structure is as follows Figure 2As shown in the figure, this module primarily consists of three convolutional modules with a kernel size of 3 and a stride of 1, five linearly varying linear modules, one multi-head linear attention module, one multi-layer perceptron (MLP) module, and one AGCN module. First, a convolution operation is applied to capture local temporal dependencies in the train run sequence, enabling the model to better understand the temporal evolution of the train state. A linear transformation is then used to map the input feature dimensions to the same dimension. This process transforms the input features in a specific direction, helping the model learn more effective representations. The linear attention mechanism reduces the computational complexity of traditional attention mechanisms by avoiding large-scale matrix multiplications. Attention information is stored and updated in the state space matrix (SSM) at each step. The updated SSM is directly added to the original matrix, enabling efficient attention computation. The MLP module learns nonlinear relationships in the data. The AGCN module generates an adaptive graph matrix based on the number of features at the station. The ReLU activation function sets all negative values in the resulting matrix to zero, enhancing sparsity and retaining only positive relationships. The SoftMax function normalizes the weights between nodes to generate an adaptive adjacency matrix. The adaptive adjacency matrix and the MLP output are efficiently multiplied using the Einstein summation convention to enable information flow between nodes (stations). Information from multiple stations is aggregated through multi-layer graph convolution operations.
[0086] The MTSCA module is a multi-timescale channel attention module. Based on the K most significant timescales and the corresponding average amplitude values obtained by the FFT for Patch module, the MTSCA module divides the train operation data into K timescales. The MTSCA module consists of a Reshape operation, a Discrete Cosine Transform (DCT) module, a Fusion module, an FC module, and a Sigmod function. The Reshape operation transforms the data information at multiple timescales into specific dimensions. For each timescale tensor, the tensor is split along the feature dimension to obtain independent channels. Next, a Discrete Cosine Transform (DCT) is applied to each channel to convert the tensor from the time domain to the frequency domain. The frequency values of each channel are then aggregated through the Fusion module, and the tensor dimension is transformed using a fully connected operation. A channel attention mechanism is then used to calculate weights for each channel. These weights are multiplied with the original timescale tensor, enhancing the model's ability to capture delay-related features and capturing inter-sequence correlations. The SoftMax function assigns weights to the average amplitude values corresponding to the time scale. By multiplying each time scale series by the corresponding weight value and adding them together, information from multiple time scale series is aggregated. Finally, the time series is passed through the Output Layer to output the delay prediction results.
[0087] Combine Figure 1 The overall network structure of the model shown in the figure is as follows. The workflow of this embodiment consists of the following steps. First, the processed train operation history data is input into the model in a two-dimensional format. The train operation history data is first standardized and normalized to improve the data stationarity, resulting in the standardized and normalized data. Then, a one-dimensional convolution operation with a kernel size of 3 and a stride of 1 is used to model the features and project the tensor into a high-dimensional space. A parameter balance factor is used to determine the importance of the convolution features in the overall embedding. The spatial embedding operation is used to represent the relative position relationship between the train and its adjacent related trains. The trains are sorted according to the order of arrival, which is treated as an ordered sequence. The position encoding operation is used to assign position information to the tensor, thereby identifying the relative position relationship between the train and the station. The temporal embedding operation divides the timestamp information in the train data into four granularities: hour, week, month, and year. The four timestamp information is encoded separately. The temporal embedding is then combined with the spatial embedding and feature embedding to obtain a spatiotemporal embedding representation of the data tensor. This enhances the model's ability to capture spatiotemporal features. The resulting high-dimensional spatiotemporal embedding tensor is transformed into the frequency domain using the FFT forPatch module. The data's periodic characteristics are analyzed, and the top K most significant timescale values and the corresponding average amplitude values of the timescales are determined. The General Asynchronous Array (GA) module then captures spatiotemporal correlations in the high-dimensional spatiotemporal embedding tensor, while the MLLA module uses a Mamba-like multi-head linear attention mechanism to capture inter-train correlations, allowing the model to focus on important train information. The adaptive graph convolution module then constructs a matrix tensor of the same size based on the station feature dimensions. This matrix tensor is activated using the ReLU function, which sets all negative values in the resulting matrix to zero, enhancing sparsity and retaining only positive relationships. The SoftMax function is then used to normalize the weights between nodes to generate a station-adaptive adjacency matrix. The resulting adaptive adjacency matrix is then efficiently multiplied with the spatiotemporal embedding tensor using the Einstein summation convention, thus ensuring information flow between nodes (stations). Multiple layers of graph convolution then aggregate the feature information of multiple adjacent stations. Based on the information required at multiple time scales, the MTSCA module divides train operation time series data into K time scales. Each time scale contains different periodic information. The data at each K time scale is transformed into the frequency domain using the DCT. Channel attention weights are calculated for the sequences at different time scales, thereby increasing the channel information at different time scales and capturing nonlinear correlations between sequences. A SoftMax operation is then performed on the average amplitude values at different time scales to determine the weight value for each time scale. Based on the different weight information, the sequence information at multiple time scales is aggregated, and a linear layer is used to output the multi-station train delay prediction results.
[0088] This embodiment improves the accuracy of train delay prediction in multi-station and complex railway network conditions by applying a combined model of the Embedding module, the FFT for Patch module, the GA module, and the MTSCA module.
[0089] The Embedding module consists of T-Embed, P-Embed, and Conv modules. The encoding of temporal and spatial features effectively captures the spatiotemporal correlations of train operations. Train operation data is inherently time-dependent, and a train's operating status changes dynamically over time. When a train is delayed, the delay may recover, worsen, or remain constant over time. Spatially, delays propagate both vertically and horizontally, causing delays for other trains running in the same time period. This interdependence between trains is key to delay prediction. The model uses time and position encoding to represent these features. Furthermore, considering the complexity of the railway network and potential external disturbances (such as severe weather like strong winds or heavy snow), the model utilizes convolution operations to encode these environmental factors that significantly impact train operations. The FFT for Patch module uses a fast Fourier transform (FFT) module to identify the periodicity of train operation data and decompose it into multiple time scales. This multi-scale approach helps capture potential periodic features, thereby improving prediction accuracy. The General Asynchronous Module (GA) module combines two submodules: the MLLA module and the AGCN module. The MLLA module uses linear attention to capture the interactions between trains. The AGCN module learns the dynamic correlations between train events at different stations by constructing an adaptive graph matrix. Unlike traditional graph convolutional networks (GCNs), which are limited to aggregating information from directly connected nodes, the adaptive adjacency matrix can aggregate information from distant stations over time. This design enables the model to more comprehensively capture complex spatiotemporal relationships. The MTSCA module processes periodic features at multiple scales. Train operation data is influenced by multiple factors, including latent variables such as delay distribution. Train delays exhibit a long-tail distribution, with short and medium delays occurring frequently, while long delays are relatively rare. This characteristic makes it challenging to extract explicit feature information in the time domain. The MTSCA module uses DCT to convert time domain data into multi-scale frequency domain data, which is then processed using a channel-wise attention mechanism. This approach enables the model to capture correlations between sequences and assign higher attention weights to factors most relevant to train delays, resulting in more accurate analysis and prediction of delays.
[0090] In summary, the combination of the Embedding module, FFT for Patch module, GA module, and MTSCA module can achieve accurate train delay prediction under multi-station and complex railway network conditions.
[0091] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A train delay prediction method based on multiple time scales, characterized in that: include: Perform multi-source data fusion on train operation data to obtain a spatiotemporal embedding representation of the data tensor; wherein the spatiotemporal embedding representation includes: a time embedding representation, a space embedding representation, and a feature embedding representation; Embedding the space-time representation, identifying periodicity in the time series through fast Fourier transform, and dividing information at multiple time scales; Capturing the spatiotemporal correlation of the spatiotemporal embedding representation and generating a spatiotemporal correlation feature representation; including: Performing multi-head linear attention processing on the spatiotemporal embedding representation to obtain an enhanced feature representation; Perform adaptive graph convolution on the enhanced feature representation to obtain feature representation containing spatiotemporal correlations; Performing multi-head linear attention processing on the spatiotemporal embedding representation includes: Performing convolution operations on the spatiotemporal embedding representation to capture local temporal dependencies in the train run sequence; Perform linear transformation on the feature representation after convolution operation; Perform multi-head linear attention processing on the linearly transformed feature representation to capture the correlation between trains; Use a multi-layer perceptron to learn the nonlinear relationship of feature representation after attention processing to obtain enhanced feature representation; Adaptive graph convolution processing of the enhanced feature representation includes: Based on the enhanced feature representation, an adaptive graph matrix is generated according to the number of features of the site, and the negative values in the graph matrix are set to 0 using the ReLU activation function; Use the SoftMax function to normalize the weights between the nodes of the graph matrix to generate an adaptive adjacency matrix; The adaptive adjacency matrix and the enhanced feature representation are multiplied using the Einstein summation convention to obtain the multiplied information; The multiplied information is subjected to multi-layer graph convolution operations to aggregate information from multiple sites and obtain feature representations that include spatiotemporal correlations. Based on the information of the multiple time scales, the spatiotemporal correlation feature representation is decomposed into multiple time scales, and channel attention learning is performed on the decomposed information to obtain aggregated multi-time scale sequence information; Obtain delay prediction results based on the aggregated multi-time scale sequence information.
2. The train delay prediction method based on multiple time scales according to claim 1, characterized in that: The train operation data includes: Train operation related information: timestamp, train number, planned running time, planned stop time, actual running time, actual stop time, running buffer time, stop buffer time, train arrival delay, train departure delay; Site related information: site number; Environmental related information: temperature, weather.
3. The train delay prediction method based on multiple time scales according to claim 1, characterized in that: Obtaining the spatiotemporal embedding representation of the data tensor includes: Perform convolution operations on the train running data to model the features and project the tensors into a high-dimensional space to obtain feature embedding representations; The trains in the train operation data are sorted according to the order of arrival. The position encoding operation is used to assign position information to the tensor, thereby identifying the relative position relationship between the train and the station and obtaining a spatial embedding representation. The timestamp information in the train operation data is divided into four granularities according to hour, week, month, and year. The four timestamp information are encoded separately to obtain time embedding representation; The temporal embedding representation is integrated with the spatial embedding representation and the feature embedding representation to obtain the spatiotemporal embedding representation of the data tensor.
4. The train delay prediction method based on multiple time scales according to claim 1, characterized in that: The division into multiple time scales includes: The train operation data is converted from the time domain to the frequency domain using fast Fourier transform. The average amplitude of the data is calculated from the frequency domain analysis to obtain the top K most prominent frequency values. Based on the relationship between frequency and period, the K most significant time scales and the corresponding average amplitude values of the time scales are obtained.
5. The train delay prediction method based on multiple time scales according to claim 1, characterized in that: Decomposing the spatiotemporal correlation feature representation at multiple time scales based on the information at the multiple time scales, and learning channel attention on the decomposed information includes: Based on the information of the multiple time scales, the spatiotemporal correlation feature representation is divided into K time scales, each different time scale contains different periodic information, the data of the K time scales are converted into the frequency domain using DCT, and the channel attention weights are solved for the sequences of different time scales; A SoftMax operation is performed on the average amplitude value of the channel attention weight to solve the weight value of each time scale, and the sequence information of multiple time scales is aggregated according to different weight information.
Citation Information
Patent Citations
High-speed railway train delay prediction method and device based on graph convolutional neural network, and storage medium
CN119106765A