Pre-training enhanced space-time Transform network traffic flow prediction method
Through pre-training enhanced spatiotemporal Transformer traffic flow prediction method, long-term historical traffic flow data and graph structure learners generate spatiotemporal sparse graphs, solving the shortcomings of existing models in capturing long-term spatiotemporal dependencies and achieving more accurate traffic flow prediction effects.
Patent Information
- Application Number
- CN202510304498.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing traffic flow prediction model has shortcomings in capturing long-term spatial and temporal dependencies, and ordinary attention calculation methods cannot effectively capture global dependencies, resulting in the difficulty of the model to make accurate predictions of different future trends based on limited historical data.
The pre-training enhanced spatiotemporal Transformer traffic flow prediction method is used to pre-train the encoder through long-term historical traffic flow data to extract the hidden state of the long-term mode, and a graph structure learner is used to generate spatiotemporal sparse maps to replace the predefined road network structure maps. In the prediction stage, the time feature extraction module and the spatial feature extraction module are used to model the correlation between time and space respectively, and long sequence features are fused through adaptive strategies.
It realizes more accurately capturing complex spatio-temporal relationships and global information in traffic data, can more accurately predict medium- and long-term traffic flow under complex traffic conditions, and improves the accuracy and efficiency of traffic flow prediction.
Smart Images

Figure CN120148236A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of intelligent transportation and deep learning, and specifically relates to a spatio-temporal Transformer traffic flow prediction method enhanced by pre-training. This method is applicable to traffic flow prediction applications in intelligent transportation systems, such as urban traffic management, public transportation scheduling, and traffic optimization. Background Art
[0002] With the continuous development of intelligent transportation systems along with urbanization and population growth, traffic congestion has become one of the important problems faced by modern cities. Effectively predicting traffic flow conditions is crucial for improving the operation and management of the traffic system. By understanding the changing trends of traffic flow in advance, traffic managers can formulate corresponding measures to reduce congestion, improve travel efficiency, and provide a more convenient travel experience. However, due to the complexity and randomness of traffic flow, accurate traffic flow prediction is not an easy task, and how to accurately extract effective feature information remains a key issue. Secondly, existing traffic flow prediction models have deficiencies in capturing long-term spatio-temporal dependence relationships, and ordinary attention calculation methods cannot effectively capture global dependence relationships. The model may have difficulty distinguishing short-term time series under different backgrounds. Therefore, based on limited historical data, it is difficult for the model to accurately predict different future trends, and the patterns of traffic flow sequences and the spatio-temporal dependence relationships between them need to be analyzed based on long-term historical sequence data. Modeling the characteristics of spatio-temporal polymorphism and complex correlations of traffic flow data forces the transformation and upgrading of traffic flow prediction methods in the big data era. Summary of the Invention
[0003] Object of the Invention: The present invention aims to solve the above existing technical problems, and provides a spatio-temporal Transformer traffic flow prediction method enhanced by pre-training.
[0004] Technical Solution: The present invention mainly includes five parts: (1) Selection and processing of the dataset; (2) Pre-training stage; (3) Graph structure learning stage; (4) Prediction stage; (5) Verification of method effectiveness.
[0005] Step 1: Define the traffic flow prediction problem and clarify the representation of traffic flow sequence data.
[0006] The present invention sets the original traffic flow data of N sensors in the traffic network at T time slices Divide the historical traffic flow data into segments in hours, sample once every 5 minutes, and each segment has 12 sampling data. Use the traffic flow information of each week as an input sequence for pre-training, and each input sequence has 168 segments. The present invention focuses on using traffic flow data for prediction, establishing a model using historical traffic flow data and using this model to predict the traffic flow on the same road section for a period of time in the future. The calculation expression is as follows:
[0007] Among them, is the historical traffic flow data sequence,
[0008] is the future traffic flow data sequence;
[0009] Step 2: Use the long-term historical traffic flow data to pre-train the encoder through the masked auto-encoding strategy, extract long sequence features and periodic information, and restore the original sequence through the Transformer network with an encoder-decoder architecture to create a self-supervised task: Specifically: The original traffic flow sequence X i from node i is segmented into p sequence segments of length T, and some sequence segments are randomly masked, an identifier flag is set, and the masking rate is set to 75%, and then input into the encoder.
[0010] The encoder consists of three parts: input embedding, positional encoding, and time-aware attention block: The input embedding layer is a linear projection that can map the space of the unmasked sequence to the latent space, and the conversion formula is: Among them, is the traffic data of segment j of node i in the input sequence, and are learnable parameters, The positional encoding layer is used to append sequential information, and the sine-cosine encoding in the Transformer is used to learn the positional embedding; then, by stacking multiple time-aware attentions, the long-term dependencies and periodic features between the traffic flow and the nodes are obtained. The time-aware attention calculation formula is:
[0011]
[0012] is the time series matrix, are respectively and weight matrices of, d q 、d k and d v are the lengths of the query, key, and value respectively. Further, the latent representations of all unmasked sequence segments are obtained through encoding. This encoder part only extracts features from the unmasked sequence segments.
[0013] The decoder is designed independently of the encoder and consists of a Transformer block and a prediction layer; the latent representation of the input segment is used to reconstruct the sequence The prediction layer uses a multi - layer perceptron (MLP) for prediction, mainly to restore the masked part of the sequence segment, and trains with this as the goal. This decoder part calculates for all sequences, and its operations are parallelly calculated for all time series.
[0014] Furthermore, through the above - mentioned method, a pre - trained encoder can be obtained. It can encode the information of historical long - term sequences, extract the hidden states and periodic features of long - sequence information, obtain an overall understanding of the traffic flow sequence, and apply it in the prediction stage. The pre - trained model can effectively learn time patterns from very long - term historical time series and generate segment representations containing rich context information.
[0015] Step 3: Adopt a data - driven method to dynamically update the weights, adaptively capture the spatial relationship between traffic nodes from time - series data, and generate a spatio - temporal sparse graph A. PEG , and the calculation formula is:
[0016] where M 1 =tanh(αm 1 Θ 1 ), M 2 =tanh(αm 2 Θ 2 ), m 1 , m 2 are the initialized node embeddings, Θ 1 , Θ 2 are the model parameters, α is the saturation rate of the activation function. To make the matrix reach a sparse effect and reduce the computational cost of subsequent graph convolution, the following formula is used to reduce the number of neighbor nodes:
[0017]
[0018] where argtopk(·) returns the indices of the k maximum values in the vector, retains the weights of the k most relevant neighbor nodes, and sets the weights of non - connected nodes to 0. The finally generated graph A PEG is a sparse graph that can automatically learn the relationship between each node in the graph, replacing the pre - defined graph structure in advance.
[0019] Step 4: Further, take the last segment of the original historical sequence and the spatio - temporal sparse graph A PEG as the input of the spatio - temporal Transformer network model, and at the same time fix the pre - trained encoder parameters to obtain long - sequence features, as an enhancement of the long - term historical sequence information. Furthermore, without increasing the computational amount, the introduction of longer historical cycle information is realized to capture long - term time - dependent relationships and dynamic spatial - dependent relationships.
[0020] In the further prediction stage, the spatio-temporal Transformer network consists of multiple temporal feature extraction modules and spatial feature extraction modules, which respectively use the improved multi-head self-attention mechanism to model the temporal and spatial correlations and then fuse them through an adaptive strategy. Among them, the temporal attention calculation formula is:
[0021] MTA = concat(TAtt 1 , TAtt 2 ,..., TAtt h )W t ,
[0022] where are the weight matrices of Q (TA) , K (TA) , and V (TA) respectively;
[0023] The spatial graph convolution calculation formula is:
[0024]
[0025] is the normalized Laplacian matrix,
[0026] The spatial attention calculation formula is:
[0027]
[0028] MSA = concat(SAtt 1 , SAtt 2 ,..., SAtt h )W s ,
[0029] Q (SA) , K (SA) , and V (SA) are the spatial attention query, key, and value matrices generated by graph convolution respectively, A PEG is the spatio-temporal sparse graph matrix generated by the graph learning layer, and W m and W s are learnable parameter matrices.
[0030] In the further fusion stage, the temporal features and spatial features are concatenated and then input into a fully connected network, and the spatio-temporal features are aggregated through skip connections to obtain the output of the spatio-temporal feature extraction module. The calculation formula is: O i = FC(concat(MTA, MSA)). Further, an adaptive strategy is used to fuse the long sequence features and spatio-temporal features, and the long sequence features H i in the pre-training stage are adaptively fused through a weight matrix β. The calculation formula is:
[0031]
[0032] Among them and are learnable parameter matrices, is the future traffic flow obtained through the output of the fully connected network.
[0033] Step 5: Train and optimize the present invention, and verify the effectiveness of the method. The present invention uses the Adam optimizer to optimize the model, and selects the mean absolute error (MAE), mean absolute percentage error (MAPE), and root mean square error (RMSE) as evaluation indicators. The specific calculation formulas are as follows:
[0034]
[0035] The advantages and beneficial effects of the present invention are as follows: The present invention uses long-term historical traffic flow data to pre-train the encoder, extracts the hidden states of long-term patterns, and fully explores the non-linear time-dependent relationship and dynamic spatial-dependent relationship before the long sequence. Then, a graph structure learner is used to generate a spatio-temporal sparse graph to replace the predefined road network structure diagram. In the prediction stage, a time feature extraction module and a spatial feature extraction module are used to model the time and spatial correlations respectively, and then the long sequence features are adaptively fused to better capture the dynamic changes of the data. The present invention can effectively extract the information of the long-term historical traffic flow sequence, and can more accurately predict the medium- and long-term traffic flow under complex traffic conditions by capturing the complex spatio-temporal relationships and global information in the traffic data. In addition, the hidden layer features extracted by the pre-trained model are adapted to other downstream tasks of traffic flow, which can further enhance the significance and generality of the model. Description of the Drawings
[0036] Figure 1 is a schematic diagram of the traffic flow prediction process of the spatio-temporal Transformer based on pre-training enhancement of the present invention;
[0037] Figure 2 is Figure 1 a schematic diagram of the process of the pre-training stage in step S1 in;
[0038] Figure 3 is a schematic diagram of the overall architecture of the spatio-temporal Transformer network model based on pre-training enhancement of the present application;
[0039] Figure 4 is Figure 3 a calculation flow diagram of the graph structure learning stage in;
[0040] Figure 5 is Figure 3 a calculation flow diagram of the time feature extraction module in;
[0041] Figure 6 is Figure 3 the flow chart of the spatial feature extraction module in
[0042] Figure 7 is the comparison chart of the simulation results and the true values of the pre-trained enhanced spatio-temporal Transformer traffic flow prediction in this application;
[0043] Figure 8 is the flow chart of the present invention. Detailed implementation manners
[0044] To make the objectives, technical solutions and advantages of the present invention clearer, the following further elaborates the present invention in detail with reference to specific embodiments. The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings of the specification.
[0045] Refer to Figure 1 , Figure 1 the flow chart of the pre-trained enhanced spatio-temporal Transformer traffic flow prediction, including the following steps: Step 1) Data preprocessing; Step 2) Pre-training; Step 3) Constructing a spatio-temporal Transformer network; Step 4) Training the spatio-temporal Transformer network constructed in Step 3); Step 5) Using the trained model to predict the traffic flow at the next moment.
[0046] Specifically, for data preprocessing, the historical traffic flow data is divided into segments in hours, sampled every 5 minutes, so each segment has a total of 12 sampling data. If the traffic flow information of each week is used as an input sequence for pre-training, and a segment is used as the prediction input for overall training.
[0047] Specifically, for the pre-training stage, refer to Figure 2 , Figure 2 is the schematic diagram of the pre-training stage process of Step 2), including random masking, input embedding, positional encoding, time-aware attention calculation, the pre-training task of constructing a reconstruction sequence, and subsequent fine-tuning.
[0048] Specifically, for Step 3), refer to Figure 3 , Figure 3 is the overall architecture schematic diagram of the pre-trained enhanced spatio-temporal Transformer network model of this application, mainly including a pre-trained encoder, a sparse spatio-temporal graph adaptively generated in a data-driven manner, and a spatio-temporal Transformer network, where the spatio-temporal Transformer network includes a time feature extraction module composed of improved time attention, a spatial feature extraction module composed of graph convolution and improved spatial attention, an adaptive fusion module, and an output layer.
[0049] Specifically, refer to Figure 4, Figure 4 is the computational flow chart of the graph structure learning stage in step 3), including calculating the generation graph matrix and the sparse graph matrix;
[0050] Refer to Figure 5 , Figure 5 is the computational flow chart of the time feature extraction module in step 3), including the specific calculation of time attention and the splicing of multi-head attention;
[0051] Refer to Figure 5 , Figure 5 is the computational flow chart of the spatial feature extraction module in step 3), including performing graph convolution on the spatio-temporal sparse graph, the specific calculation of spatial attention, and the splicing of multi-head attention;
[0052] Specifically, the training process of step 4) can be summarized as follows:
[0053] Step1: First, perform stable normalization processing on the traffic flow speed time series data, and divide the data set into three parts: training set, test set, and validation set; perform masking operation on the long historical traffic flow sequence and input it into the pre-trained model for self-supervised training. Continuously update the parameters during the model training until the termination condition is reached, and save the parameters and status of the model for subsequent prediction tasks.
[0054] Step2: Fix the parameters of the pre-trained encoder, use the last segment of the long sequence and its long sequence features as the input of the spatio-temporal Transformer network, and train the model. Randomly initialize the network parameters, perform forward propagation to obtain the predicted output of the model, and then calculate the gradient of the loss function with respect to the network parameters through backpropagation, and continuously update the parameters until the termination condition is reached. If the training does not meet the requirements, update the training samples and continue to train the model in the next round;
[0055] Step3: After the training is completed, save the model training results, including the network system used in the training and the weight relationship between each node after training. After successfully saving the model, testing can be performed. Process the test data in the same way, and then input it into the model for prediction. Compare the final result with the actual data to obtain the prediction accuracy of the model. Comprehensive evaluation of the performance and efficiency of the model can be achieved by considering indicators such as training time.
[0056] Refer to Figure 7 , Figure 7 is the comparison chart of the pre-training enhanced spatio-temporal Transformer traffic flow prediction simulation results and the true values of this application, which intuitively shows the model prediction performance.
[0057] The present invention processes and analyzes traffic flow data, learns complex features and spatio-temporal dependence relationships in long sequence data, and realizes more precise and accurate traffic flow prediction. The above description is only for the purpose of illustrating the present invention and not for limiting the protection scope of the present invention. Any equivalent modifications and other modified changes made by those of ordinary skill in the art according to the disclosure of the present invention shall be included in the protection scope recorded in the claims.
Claims
1. A spatiotemporal Transformer traffic flow prediction method based on pre-training enhancement, characterized in that: The following steps are involved: Step 1: Obtain historical traffic flow data and preprocess it, divide the time series into segments, and use weekly traffic flow as the pre-training input sequence; Step 2: Use the masked auto-encoding strategy to pre-train the encoder, extract long-term temporal dependencies and periodic features by masking some sequence segments and reconstructing the original data; Step 3: Build a spatiotemporal Transformer network, including a pre-trained encoder, a graph structure learner, a temporal feature extraction module, a spatial feature extraction module, and a fusion module; Step 4: Fix the pre-trained encoder parameters, generate a spatiotemporal sparse graph and train the spatiotemporal Transformer network; Step 5: Predict future traffic flow based on the trained model and verify the prediction accuracy using the test set.
2. The method according to claim 1, characterized in that In step 1, the preprocessing includes: expressing the historical traffic flow data as Where T represents the length of the time series, C is the number of features for each node, and the samples are sampled every 5 minutes. Each segment contains 12 sampled data, and the input sequence is divided into hours.
3. The method according to claim 1, characterized in that In step 2, the masked automatic encoding strategy specifically includes: dividing the traffic flow sequence of node i into p segments, randomly masking 75% of the sequence segments, and inputting them into the encoder for reconstruction training; the encoder consists of an input embedding layer, a position encoding layer, and a time-aware attention block, and maps the unmasked sequence to the latent space through linear projection, and attaches weekly and daily period embedding information; the decoder restores the masked sequence segments through a multi-layer perceptron, and the loss function is the mean square error of the masked part.
4. The method according to claim 1, characterized in that: In step 3, the spatiotemporal sparse graph A PEG The generation method is: dynamically learn the spatial dependency relationship between nodes based on time series data, and the calculation formula is: Where M1 = tanh(αm1Θ1), M2 = tanh(αm2Θ2), m1 and m2 are initialization node embeddings, Θ1 and Θ2 are model parameters, and α is the saturation rate of the activation function; The generated graph A PEG It is a sparse graph that can automatically learn the relationship between nodes in the graph, replacing the pre-defined graph structure.
5. The method according to claim 1, characterized in that In step 3, the temporal feature extraction module models the temporal correlation through an improved multi-head self-attention mechanism, and the calculation formula is: in Q (TA) , K (TA) 、V (TA) The weight matrix of The spatial feature extraction module is based on the spatiotemporal sparse graph A PEG Perform graph convolution and spatial attention calculation, the formula is: where σ(·) is the Sigmoid activation function, is the activation matrix of layer l, h (0) Take the initial features of the last segment of the traffic flow length sequence feature, W (l) is the learnable parameter matrix, is the normalized Laplace matrix, and Where Q (SA) , K (SA) 、V (SA) are the spatial attention query, key and value matrices generated by graph convolution, A PEG is the spatiotemporal sparse graph matrix generated by the graph learning layer, W m is a learnable parameter matrix; The fusion module fuses the long sequence features with the spatiotemporal features through the adaptive weight matrix β and outputs the prediction results.
6. The method according to claim 1, characterized in that In step 5, Huber Loss is used as the loss function, the Adam optimizer is used to train the model, and the mean absolute error (MAE), mean absolute percentage error (MAPE) and root mean square error (RMSE) are used as evaluation indicators.
7. The method according to claim 1, characterized in that The time-aware attention mechanism dynamically adjusts the attention score by embedding weekly and daily cycle information. The calculation formula is: The time series feature matrix C t Generated by linear transformation, it captures dependencies between different time periods; in They are and The weight matrix, d q d k and d v are the lengths of query, key and value respectively, Z t is the output of the embedding layer with additional positional encoding.
8. The method according to claim 1, characterized in that The output long sequence features of the pre-trained encoder and the spatiotemporal sparse graph are used together as the input of the spatiotemporal Transformer network, introducing long-term historical information without increasing the amount of computation.
Citation Information
Cited By
Traffic flow prediction method and device
CN120452209A
Model training method and device, container preheating method and device and electronic equipment
CN120850052A
Space-time frequency multi-step indoor temperature prediction method and device
CN121434650A
Adaptive space-time coding mask pre-training traffic flow prediction method and device
CN121982899A
State-aware dynamic asymmetric space-time Transform long-term traffic speed prediction method and system
CN122157505A