Traffic flow prediction method based on space-time masking and graph convolution
By adopting a two-dimensional masking reconstruction spatiotemporal feature extraction method and an optimized GNN structure in traffic flow prediction, the problem that existing methods are difficult to capture long-term and long-space dependencies is solved, and higher prediction accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510328674.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
AI Technical Summary
Existing traffic flow prediction methods are difficult to effectively capture long-term and long-space dependencies, resulting in low prediction accuracy in complex traffic environments.
The spatial and temporal feature extraction method based on two-dimensional masking reconstruction is adopted, combined with attention mechanism, adaptive graph learning and efficient spatiotemporal feature fusion strategy, the GNN structure is optimized to improve the model's ability to capture long-term and long-space dependencies.
It significantly improves the accuracy and stability of traffic flow prediction, can predict complex traffic flow fluctuations more accurately, and improves the intelligence level of traffic management.
Smart Images

Figure CN120183187A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a traffic flow prediction model based on spatio-temporal masking and graph convolution and a method for constructing the same. Background Art
[0002] With the acceleration of the urbanization process, modern urban traffic faces increasingly complex challenges. Problems such as traffic congestion, frequent accidents, and low utilization efficiency of limited road resources have seriously affected the travel efficiency and quality of life of urban residents. In order to relieve traffic pressure and improve road capacity, traffic management departments urgently need efficient traffic optimization and management strategies. As an important tool for traffic management, traffic flow prediction can analyze historical data and real-time traffic information to predict future traffic conditions in advance, providing a scientific basis for traffic signal optimization, road resource scheduling, intelligent navigation, etc., thereby effectively reducing congestion, improving road utilization efficiency, and reducing safety hazards caused by sudden traffic conditions.
[0003] Traditional traffic flow prediction methods mainly rely on technologies such as statistical analysis, time series modeling, and machine learning, such as ARIMA (Autoregressive Integrated Moving Average Model), LSTM (Long Short-Term Memory Network), etc. These methods can, to a certain extent, model the temporal characteristics of traffic flow, but often have difficulty fully capturing complex spatial dependence relationships, especially in the case of abnormal fluctuations in traffic flow during emergencies and holidays, resulting in low prediction accuracy. In recent years, with the rapid development of deep learning technology, methods based on graph neural networks (GNNs) have demonstrated powerful data modeling capabilities in the field of traffic flow prediction. GNNs can process data with complex spatial topological structures, making them particularly suitable for traffic flow modeling in road networks. By combining time series features and spatial dependence information, GNNs can effectively improve the accuracy of traffic flow prediction.
[0004] In existing research, Lan et al. proposed the Dynamic Spatial-Temporal Aware Graph Neural Network (DSTAGNN), which can dynamically model the spatio-temporal dependence of traffic flow and effectively capture the flow change patterns during peak hours and emergencies. Li and Zhu developed the Spatial-Temporal Fusion Graph Neural Networks (STFGNN), which deeply integrates spatial and temporal features, significantly enhancing the model's ability to capture the spatio-temporal dependence relationship of traffic flow and thus improving the prediction accuracy. Guo et al. proposed the method of Learning Dynamics and Heterogeneity of Spatial-Temporal Graph Data (ASTGNN), which combines the attention mechanism and can adaptively weight important spatio-temporal information, enabling the model to more accurately predict complex traffic flow fluctuations. However, these methods still have certain limitations in extracting long-time and long-space features and are difficult to effectively model long-range dependence relationships, resulting in a decline in performance in long-term prediction tasks.
[0005] To address the above problems, the present invention proposes an improved traffic flow prediction method. By constructing a more efficient spatio-temporal feature extraction model and optimizing the GNN structure, the model's ability to capture long-time and long-space dependence relationships is enhanced. This method combines the attention mechanism, adaptive graph learning, and an efficient spatio-temporal feature fusion strategy, enabling more accurate prediction of complex traffic flow fluctuations, improving the intelligent level of traffic management, and thus providing strong support for the development of intelligent transportation. Summary of the Invention
[0006] The present invention relates to the field of intelligent transportation. Specifically, it relates to a traffic flow prediction method based on spatio-temporal feature extraction, which can efficiently capture long-time and long-space dependence relationships, improve the accuracy and stability of traffic flow prediction, and is applicable to application scenarios such as intelligent traffic management, navigation optimization, and road resource scheduling.
[0007] With the accelerating urbanization process, modern urban transportation faces increasingly complex challenges. Problems such as traffic congestion, frequent accidents, and limited utilization efficiency of road resources have severely affected the travel efficiency and quality of life of urban residents. As a core task in the Intelligent Transportation System (ITS), traffic flow prediction aims to predict the traffic flow changes in the traffic network in the future period based on historical traffic data, providing a scientific basis for traffic signal optimization, road resource scheduling, intelligent navigation, etc.
[0008] The traffic flow prediction problem can essentially be formulated as a spatio-temporal sequence modeling task. In a given traffic network, let the traffic flow data be represented by the traffic flow tensor X ∈ R T×N×C where:
[0009] T represents the number of time steps, that is, the length of the time window of historical observations;
[0010] N represents the number of spatial nodes in the traffic network (such as road intersections, sensor stations, etc.);
[0011] C represents the number of information channels, such as features like traffic flow, speed, occupancy, etc.
[0012] The goal of traffic flow prediction is to predict the traffic flow in the next T' time steps by learning the mapping function f, that is:
[0013]
[0014] where, is the traffic flow data for the next T' time steps predicted by the model.
[0015] Existing methods have certain limitations in spatio-temporal modeling. Especially, the input length T of most traffic flow prediction models is usually 12 (corresponding to 1 hour of data), which limits the model's ability to capture long-term spatio-temporal dependencies. In a complex traffic environment, traffic flow is affected not only by short-term influences but also by dynamic effects over longer time and larger spatial scales. Therefore, how to efficiently model long-time and long-space dependencies has become a key challenge in improving the accuracy of traffic flow prediction.
[0016] To solve the above problems, the present invention proposes an improved spatio-temporal feature extraction method for efficiently capturing long-time and long-space dependencies. The core idea is to design a feature extraction mechanism based on two-dimensional masked reconstruction, which performs masked reconstruction tasks in the time dimension and the spatial dimension respectively, so as to achieve efficient information extraction.
[0017] Details are as follows: Through long-term and long-space feature extraction, dual-dimensional masked reconstruction tasks, and spatio-temporal feature fusion enhanced by a Graph Neural Network (GNN), efficient and accurate traffic flow prediction is achieved. First, a long input sequence T>>12 is set to cover traffic flow information over a longer time range, and temporal masking and spatial masking are respectively adopted to enable the model to separately learn temporal and spatial dependency patterns. Let M t and M s represent the temporal and spatial masking matrices respectively, where M t ∈{0,1} T*N , controls the masking of the temporal dimension, and M s ∈{0,1} T*N controls the masking of the temporal dimension.
[0018] Secondly, in the dual-dimensional masked reconstruction task, a temporal decoder and a spatial decoder are respectively used for the reconstruction task, and their calculation formulas are as follows:
[0019]
[0020] Among them, f t and f s are the temporal decoder and spatial decoder based on the Transformer structure respectively. During the training process, masked mean squared error (MaskedMSE) loss is used for optimization. Brief Description of the Drawings
[0021] Figure 1 is a flowchart of a method for constructing a traffic flow prediction model based on spatio-temporal masking and graph convolution provided by the present invention;
[0022] Figure 2 is a flowchart of spatio-temporal masking of a traffic flow prediction model based on spatio-temporal masking and graph convolution provided by the present invention;
[0023] Figure 3 is a schematic diagram of the mamba structure of a traffic flow prediction model based on spatio-temporal masking and graph convolution provided by the present invention; Detailed Embodiments
[0024] Specifically, given an input spatio-temporal time series X∈R T*N*C , we propose the following masking strategy: (1) Spatial masking (S-Mask) randomly masks the time series of N×p sensors, where p is the masking rate between 0 and 1. This produces a spatially masked input, Temporal masking (T-Mask) randomly masks T×p time steps of a time series. This produces a temporally masked input Two masking strateg
[0025]
[0026] The masked data is set to 0, and the model needs to learn global and local spatio-temporal patterns from the unmasked data to reconstruct the masked data.
[0027] In the reconstruction task of the invention, we choose the mamba model. The mamba model combines the State Space Model (SSM) and local convolutional characteristics, can efficiently capture long-term and local dependencies, and has high computational efficiency. For the masked data masked_x, it will pass through a fully connected layer to map the feature dimension from input_dim (original feature dimension) to embed_dim (embedding dimension)
[0028] X embed = W embed *X + b embed
[0029] This can improve the expressive ability of features, facilitate the subsequent mamba encoder to capture spatio-temporal patterns, then model the input sequence on the ssm module in the mamba model, and then capture local context through a one-dimensional convolution. The output obtained after encoder processing is a high-dimensional spatio-temporal feature representation. For this output, passing through a fully connected layer to restore to the original feature dimension can reconstruct the complete spatio-temporal data. The Mamba model is trained to optimize its ability to reconstruct masked data, and the loss function is selected as the Mean Absolute Error (MAE).
[0030] The loss function of the model is as follows
[0031]
[0032] In the traffic flow prediction task, the spatial structure information of the traffic network is the core part. Due to its advantage in dealing with graph-structured data, graph neural networks are widely used to model the spatial dependencies between traffic nodes. The traffic network can be naturally represented as a graph structure G=(V, E, W), where V represents the set of traffic nodes (such as sensor locations), E represents the set of edges between nodes (such as road connection relationships), and W∈RN×N is the weight matrix, representing the similarity or connection strength between nodes. To capture the complex relationships between nodes, we use the Graph Laplacian matrix for modeling: L = D - W, where D is the degree matrix. For the normalized graph Laplacian matrix, its form is where \(I\) is the identity matrix. This representation method can effectively model the global dependencies between traffic nodes and provide a basis for subsequent calculations of the graph neural network.
[0033] After that, the eigenvalue decomposition of the graph Laplacian matrix is used, \(L = U\Lambda U^T\) T , where \(U\) is the eigenvector matrix and \(\Lambda\) is the diagonal matrix of eigenvalues. The GFT transforms the signal \(X\) in the node domain into a spectral domain representation: In the spectral domain, frequency-domain convolution is performed on the signal, in the form of: where \(\Phi\) is the frequency-domain filter, and then the result in the spectral domain is transformed back to the node domain
[0034] Based on the frequency-domain convolution, multi-order Chebyshev polynomial approximation is further added:
[0035]
[0036] where is the Chebyshev polynomial of the normalized graph Laplacian matrix, \(\theta\) k are the parameters to be learned. To further enhance the expression ability of spatial features, a node self-attention mechanism is introduced to dynamically adjust the weight distribution between nodes
[0037]
[0038] where \(Q\) and \(K\) are the query and key-value vectors of the node features
[0039] In STMFNet, the present invention designs a novel feature fusion mechanism that can be flexibly embedded into the existing prediction framework. This mechanism significantly improves the prediction ability of the model by fusing the spatial features and temporal features generated by the spatio-temporal masking module with the intermediate representation of the predictor. The specific implementation is as follows:
[0040] First, we input the long-time-span input sequence \(X\) long (containing \(L\) time steps) into the pre-trained spatio-temporal masking module to generate the spatial feature matrix \(F\) S \(\in\mathbb{R}\) N×L×C and the temporal feature matrix \(F\) T \(\in\)
[0041] \(\mathbb{R}\) N×L×C , where \(N\) represents the number of spatial nodes and \(C\) represents the feature dimension. Next, we use a downstream spatio-temporal predictor (with parameters ) to process the short-term input \(X\) short =\(\{X\) t-T+1 , \(X\) t-T+2 , \(\cdots\), \(X\)t Process it to obtain its hidden representation H P ∈R N×C' The formula is as follows
[0042]
[0043] where C’ is the dimension of the predictor hidden representation. To align the spatial feature F S and the temporal feature F T with H P of the present invention, truncation and reshaping operations are performed on F S and F T to obtain F' S ∈R N×T'×C and F' T ∈R N×T'×C , where T’ is the number of time steps of the short-term input. Then, F' S and F′ T are projected onto the C’ dimension through two multi-layer perceptrons (MLPs) to obtain H S and H T
[0044] H S = MLP(F' S ) H T = MLP(F' T )
[0045] Finally, the present invention adds the hidden representation H P of the predictor to the projected spatial feature H S and the temporal feature H T to obtain the enhanced representation H enhanced = H P + H S + H T .
[0046] In this way, H enhanced not only contains the short-term spatio-temporal features extracted by the predictor itself, but also integrates the long-range spatial and temporal features generated by the spatio-temporal masking module, thus significantly improving the prediction performance of the model.
[0047] The above embodiments only illustrate the technical solutions of the present invention in a specific implementation manner. Any equivalent replacement, modification or partial replacement of the present invention without departing from the spirit and scope of the present invention shall be covered by the scope of protection of the claims of the present invention.
Claims
1. A traffic flow prediction method based on spatiotemporal masking and feature fusion, characterized in that: The following steps are involved: (1) Data preprocessing: constructing the traffic flow data tensor X∈R T*N*C , where T represents the time step, N represents the number of sensor nodes in the traffic network, and C represents the number of input data channels (including flow, speed, occupancy, etc.). (2) Spatiotemporal masking mechanism: The temporal masking (T-Mamba) and spatial masking (S-Mamba) modules are used to randomly mask the input data and reconstruct the masked data through self-supervised learning to enhance the model’s ability to learn long-term and long-spatial dependencies. (3) Time series feature extraction: Traffic flow data is transformed into the frequency domain using fast Fourier transform (FFT) to extract periodic patterns. (4) Spatial feature extraction: A graph convolutional network (GCN) is used to model the topological relationship between sensor nodes and obtain spatial correlation. (5) Spatiotemporal feature fusion: The STGrpBlock structure is designed to combine convolution operations with the self-attention mechanism to simultaneously capture the complex dependencies in both temporal and spatial dimensions. (6) Prediction output: Based on the extracted spatiotemporal features, the traffic flow prediction results for future time steps are generated through the regression layer.
2. The traffic flow prediction method according to claim 1, characterized in that: The spatiotemporal masking mechanism adopts a temporal decoder and a spatial decoder based on a mamba structure to reconstruct the masked information respectively, so as to improve the expression capability of spatiotemporal features.
3. The traffic flow prediction method according to claim 1, characterized in that: The time series feature extraction module uses fast Fourier transform (FFT) to perform frequency domain analysis and combines it with an adaptive periodic modeling mechanism to capture the periodic changes of holidays and working days.
4. The traffic flow prediction method according to claim 1, characterized in that: The graph convolutional network (GCN) uses an adaptive adjacency matrix to model the dynamic topological relationship of the traffic network to enhance the model's ability to model traffic changes under emergencies.
5. The traffic flow prediction method according to claim 1, characterized in that: The STGrpBlock adopts a multi-head self-attention mechanism to strengthen the information interaction between different time steps and sensor nodes.
6. The traffic flow prediction method according to claim 1, characterized in that: The prediction output layer uses a multi-layer perceptron (MLP) for nonlinear mapping to improve the final prediction accuracy.
7. A traffic flow prediction system based on spatiotemporal masking and feature fusion, comprising: Data acquisition module, which collects traffic flow data from multiple sensors; Data preprocessing module, used to construct traffic flow data tensor; A spatiotemporal feature extraction module that performs spatiotemporal masking, FFT frequency domain conversion, GCN calculation, and STGrpBlock calculation to extract spatiotemporal correlations; The prediction module is used to generate future traffic flow prediction results based on the extracted spatiotemporal features; Storage module, used to store historical traffic data and trained prediction models.
8. The traffic flow prediction system according to claim 8, characterized in that: The data acquisition module includes a plurality of sensor nodes deployed at different road locations to obtain multi-source heterogeneous traffic flow information.
9. The traffic flow prediction system according to claim 8, characterized in that: The prediction module is based on the Mamba structure and combines the self-attention mechanism with the graph convolutional network (GCN) for spatiotemporal modeling to improve the prediction accuracy.
Citation Information
Cited By
Traffic flow prediction method based on time-frequency domain joint modeling
CN121354346A