Traffic flow prediction method based on space-time convolution self-attention network
By introducing the self-attention mechanism of time convolution network and graph convolution network into the traffic flow prediction method, the shortcomings of the existing methods in modeling the local trend of traffic flow sequences and the spatial heterogeneity of road networks are solved, and more efficient traffic flow prediction is achieved.
Patent Information
- Application Number
- CN202510304790.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120124809A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of intelligent transportation and deep learning, and specifically relates to a traffic flow prediction method based on a spatio-temporal convolutional self-attention network. Background Art
[0002] Traffic flow prediction is a hot research field in intelligent transportation systems. This technology constructs a traffic flow prediction network based on deep learning and uses a large amount of historical traffic flow data to train the network to achieve accurate and real-time traffic flow prediction, aiming to improve traffic management efficiency, achieve precise allocation of traffic resources, and provide high-quality traffic services.
[0003] In recent years, a large number of studies have applied the Transformer network to traffic flow prediction and proposed many improved methods based on the Transformer network architecture. Most of these methods directly use the standard self-attention mechanism to model the temporal and spatial correlations of traffic flow data without making appropriate improvements to the traffic flow data. Therefore, there are the following problems: First, the standard self-attention mechanism treats each time step as a discrete data point, ignoring the inherent local trend information of sequence data, resulting in inaccurate correlation matching. Second, when using the self-attention mechanism to calculate the spatial correlation of nodes, the existing methods only consider the similarity of traffic flow values between nodes, while ignoring the impact of spatial heterogeneity attributes such as road network topology, road type, and road width on node correlation. In addition, the existing methods use the global self-attention mechanism to model the spatial correlation of nodes, and there is a serious problem of attention dispersion when facing a complex road network with a large number of nodes, making it difficult to focus on the key nodes that play a leading role. Summary of the Invention
[0004] In view of the above problems, the present invention proposes a traffic flow prediction method based on a spatio-temporal convolutional self-attention network. The network structure mainly consists of a mask matrix generation module, a data embedding module, a temporal convolutional self-attention module, and a spatial convolutional self-attention module. The mask matrix generation module filters out nodes with high correlation degrees according to the short-term and long-term correlations of node traffic flows and the adjacency relationship between nodes. The data embedding module is used to map the original data into a high-dimensional space and embed an adaptive time period vector. The temporal convolutional self-attention module introduces a temporal convolutional network into the standard self-attention mechanism to consider the local trend information in the traffic flow sequence and improve the accuracy of temporal correlation modeling. The spatial convolutional self-attention module embeds an additional node heterogeneity feature vector for each node and, based on the mask matrix, uses a self-attention mechanism that integrates a graph convolutional network to model the spatial correlation between nodes. The above improvements effectively improve the accuracy of traffic flow prediction.
[0005] The traffic flow prediction method based on the spatio-temporal convolutional self-attention network, its construction and application include the following steps:
[0006] Step 1) Collect road network information to obtain the original data, normalize the original data through preprocessing, and divide the data set into a training set, a validation set, and a test set according to the ratio of 6:2:2;
[0007] Step 2) Construct a mask matrix generation module, screen out nodes with high correlation according to the short-term correlation, long-term correlation of node traffic flow, and the adjacency relationship between nodes, and generate a mask matrix;
[0008] Step 3) Construct a data embedding module, map the original data to a high-dimensional space, and embed an adaptive periodic vector to enhance the expression ability of the data;
[0009] Step 4) Construct a temporal convolutional self-attention module to capture the local change trend of the traffic flow sequence and model the temporal correlation between each time step;
[0010] Step 5) Construct a spatial convolutional self-attention module to add additional heterogeneous feature vectors to the nodes, mine the local traffic patterns of the road network nodes, and model the spatial correlation between each node;
[0011] Step 6) Concatenate the mask matrix generation module, the data embedding module, the temporal convolutional self-attention module, and the spatial convolutional self-attention module, add a residual connection to stabilize the training gradient and accelerate network convergence, add a fully connected layer to fuse the spatio-temporal features, and use it to output the prediction result, finally completing the network construction;
[0012] Step 7) Use the traffic flow data set collected from the road network to train the network and evaluate the network prediction accuracy.
[0013] Step 8) Collect the real-time traffic flow data of the road network, input it into the trained spatio-temporal convolutional self-attention network, and obtain the traffic flow prediction result.
[0014] Furthermore, in the above step 1), in the urban road network, given the road network structure G=(S, E, A), where S is the set of traffic nodes in the road network, E is the set of roads connecting traffic nodes, is the adjacency matrix of the road network nodes, reflecting the connection relationship between nodes. The traffic flow data is collected by sensors arranged on the road network nodes. The traffic flow data recorded by N sensors at time step t is expressed as This method learns a function f according to the given road network structure G and the traffic flow data of the previous T time steps to predict the traffic flow data of the next τ time steps, that is, {X t+1 ,…,X t+τ} = f(X t+1 ,…,X t+τ; G).
[0015] Further, in step 2), a mask matrix generation module is constructed to screen key nodes. The module screens nodes with high correlation degrees according to three methods: short-term correlation, long-term correlation of node traffic flow, and adjacency relationship between nodes. The specific generation method of the mask matrix is as follows:
[0016] Step 2-1): Initialize the mask matrix The initial value is set to 0;
[0017] Step 2-2): Input the complete historical traffic flow data, calculate the similarity of the long-term traffic flow sequences between each pair of nodes through the Pearson correlation coefficient, select the k nodes with the highest similarity for each node, and set the corresponding positions in the mask matrix to 1;
[0018] Step 2-3): Input the short-term historical traffic flow data of the adjacent T time steps, calculate the similarity of the short-term traffic flow sequences between each pair of nodes through the Pearson correlation coefficient, select the k nodes with the highest similarity for each node, and set the corresponding positions in the mask matrix to 1;
[0019] Step 2-4): According to the topological structure of the traffic road network, select the reachable nodes within five hops of each node in the traffic road network as key nodes, and set the corresponding positions in the mask matrix to 1.
[0020] Further, in step 3), a data embedding module is constructed to map the original data to a high-dimensional space, which can enhance the expression ability of the data. At the same time, an adaptive periodic embedding is introduced to enhance the periodic characteristics of the data, thereby improving the time correlation modeling ability of the network; the specific construction method is as follows:
[0021] Step 3-1): Use a fully connected network to project the original input into a high-dimensional space d is the feature dimension after projection;
[0022] Step 3-2): Initialize a learnable weekly periodic feature matrix This matrix assigns an independent d-dimensional feature vector to each day of the week, and the data collected on the same day has the same weekly periodic feature vector;
[0023] Step 3-3): Initialize a learnable daily periodic feature matrix In the dataset used in this method, traffic flow data is recorded every five minutes. Therefore, a day is divided into 288 time steps. The date feature matrix assigns an independent d-dimensional feature vector to each time step, and the data collected at the same time step has the same daily periodic feature vector;
[0024] Step 3-4): For the input which contains traffic flow data for a total of T time steps, select the corresponding vector from the period matrix according to the data collection time to obtain the weekly period embedding and the daily period embedding Then add the two to to obtain the output of the data embedding module
[0025] Furthermore, in the said step 4), a temporal convolutional self-attention module is constructed to extract temporal features from the output X (0) of the data embedding module. The specific construction method is as follows:
[0026] Step 4-1): Define a Temporal Convolutional Network (TCN). The TCN can effectively extract the local change trend in the traffic flow sequence, so as to assign higher attention weights to time steps with similar local change trends during the attention calculation process, improving the accuracy of temporal correlation modeling. For the traffic flow history sequence of node n, the calculation formula of the TCN is:
[0027]
[0028] where i is the time step index, m is the input dimension index, k is the convolutional kernel size, is the convolutional kernel.
[0029] Step 4-2): Perform temporal convolutional self-attention calculation. Through the TCN, the output X (0) of the data embedding module is converted into a query matrix Q T and a key matrix K T , and X (0) is converted into a value matrix V T through a linear transformation. Then, perform scaled dot-product attention calculation on the query matrix Q T and the key matrix K T to obtain the temporal correlation matrix A T . Then multiply A T by the value matrix V T to obtain the output of a single attention head; finally, concatenate the outputs of each attention head to obtain the output of the temporal convolutional self-attention module. The calculation process is as follows:
[0030] Q T = TCN Q (X (0) ), K T = TCN K (X (0) ), VT = X (0) W T
[0031]
[0032] head i = A T V T
[0033]
[0034] Among them, is a learnable parameter matrix. The Softmax function represents the normalization operation on the matrix, and the Concat function represents matrix concatenation. is the output of the time convolutional self-attention module.
[0035] Furthermore, in step 5), a spatial convolutional self-attention module is constructed. First, additional heterogeneous feature embeddings are added to the nodes, and spatial features are extracted from the output X of the data embedding module based on the mask matrix. The specific construction method is as follows: (0) in
[0036] Step 5-1): Randomly initialize the learnable node heterogeneous feature matrix and concatenate it with the output X of the data embedding module (0) to obtain The node heterogeneous feature embedding assigns a node heterogeneous feature vector to each node in the traffic road network. This vector can be continuously learned during model training and is used to characterize the influence of spatial attributes such as road type, width, and local topological structure of the road network on node correlation.
[0037] Step 5-2): Define a Graph Convolutional Network (GCN). The GCN can capture the local spatial features of nodes according to the adjacency matrix of the traffic road network, so as to assign higher attention weights to nodes with similar local traffic patterns during the attention calculation process, improving the accuracy of spatial correlation modeling. The calculation formula of the GCN is:
[0038]
[0039] where ReLU is the activation function, is the adjacency matrix of the traffic road network nodes with self-loops added, I is the identity matrix, is the corresponding degree matrix, is a learnable parameter matrix.
[0040] Step 5-3): Perform spatial convolutional self-attention calculation. Convert X through GCN (0) ' into query matrix Q S and key matrix K S , convert X through linear transformation (0) into value matrix V S , then perform scaled dot-product attention calculation on query matrix Q S and key matrix K S , and perform Hadamard product with the output of the mask matrix generation module to obtain spatial correlation matrix A S , then multiply A S with value matrix V S to get the output of a single attention head; finally, concatenate the outputs of each attention head to obtain the output of the spatial convolutional self-attention module The calculation process is as follows:
[0041] Q S = GCN Q (X (0)′ ), K S = GCN K (X (0)′ ), V S = X (0) W S
[0042]
[0043] head i = A S V S
[0044]
[0045] wherein, is a learnable parameter matrix, ⊙ is the Hadamard product, is the output of the spatial convolutional self-attention module.
[0046] Furthermore, in step 6), concatenate the mask matrix generation module, data embedding module, temporal convolutional self-attention module, and spatial convolutional self-attention module, and add residual connections and fully connected layers to complete the network construction. The input of the network is The output of the network is The specific process is as follows:
[0047] Step 6-1): Input X first undergoes feature embedding through the data embedding module to obtain
[0048] Step 6-2): X (0)The time - convolutional self - attention module and the space - convolutional self - attention module are respectively input for spatio - temporal feature extraction to obtain and A fully - connected network is used to fuse the spatio - temporal features, and the calculation formula is:
[0049]
[0050] where is a learnable parameter matrix.
[0051] Step 6 - 3): Repeat the spatio - temporal feature extraction process of Step 6 - 2) a total of L times, and finally obtain the spatio - temporal feature X (l) of the traffic flow, and input it into a two - layer fully - connected network to obtain the final prediction result Y. The calculation formula is as follows:
[0052] Y = ReLU(ReLU(X (l) W O1 )W O2 )
[0053] where ReLU is the activation function, is a learnable parameter matrix.
[0054] Furthermore, in the said Step 7), the traffic flow data set collected by the road network is used to train the network, and the prediction accuracy of the network is evaluated. The specific process is as follows:
[0055] Step 7 - 1): Initialize parameters such as the number of heads of the multi - head attention, the learning rate, the batch size, the maximum number of iterations, and the feature dimension;
[0056] Step 7 - 2): Divide the training set, the validation set, and the test set in a ratio of 6:2:2, and use the mean absolute error as the loss function;
[0057] Step 7 - 3): Use the training set to train the network. Stop training when the maximum number of iterations is reached or the error no longer decreases, and use the test set to test the prediction accuracy of the network;
[0058] The beneficial effects of the present invention:
[0059] In view of the fact that most existing methods ignore the important influence of the local change trend and spatial heterogeneity of the sequence when using the self-attention mechanism to model spatio-temporal correlation, and there is also the problem of scattered spatial attention in complex road networks, a traffic flow prediction method based on a spatio-temporal convolutional self-attention network is proposed. This method uses a convolutional self-attention mechanism to extract spatio-temporal features from traffic flow data. Among them, the temporal convolutional self-attention introduces a temporal convolutional network into the standard self-attention mechanism to consider the local trend information in the traffic flow sequence and improve the accuracy of temporal correlation modeling; the spatial convolutional self-attention embeds an additional node heterogeneity feature vector for each node, and filters key nodes based on a mask matrix, and uses a self-attention mechanism integrating a graph convolutional network to model the spatial correlation between nodes and improve the accuracy of spatial correlation modeling. Experimental results on the real traffic dataset PeMS08 show that the prediction accuracy of this method is better than that of existing traffic flow prediction methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a schematic diagram of the steps of a traffic flow prediction method based on a spatio-temporal convolutional self-attention network of the present invention;
[0061] Figure 2 It is a model structure diagram of a traffic flow prediction method based on a spatio-temporal convolutional self-attention network of the present invention; (a) is the overall network structure; (b) is the structure of the temporal convolutional self-attention module; (c) is the structure of the spatial convolutional self-attention module;
[0062] Figure 3 It is a comparison diagram of the predicted values and the true values of a traffic flow prediction method based on a spatio-temporal convolutional self-attention network of the present invention on the PeMS08 dataset. DETAILED DESCRIPTION OF THE INVENTION
[0063] The technical method of the present invention will be further described in detail below with reference to the accompanying drawings of the specification.
[0064] As Figure 1 described, a traffic flow prediction method based on a spatio-temporal convolutional self-attention network includes the following steps:
[0065] Step 1) Collect traffic flow data of the road network. After preprocessing operations, the dataset is divided into a training set, a validation set, and a test set according to a ratio of 6:2:2;
[0066] In the above step 1), in the urban road network, a given road network structure G=(S, E, A) is provided, where S is the set of traffic nodes in the road network, E is the set of roads connecting traffic nodes, is the adjacency matrix of the road network nodes, reflecting the connection relationship between nodes. The traffic flow data is collected by sensors arranged on the road network nodes. The traffic flow data recorded by N sensors at the t time step is expressed as This method learns a function f to predict traffic flow data for the next τ time steps based on a given road network structure G and traffic flow data for the previous T time steps, i.e., {X t+1 , …, X t+τ} = f(X t+1 , …, X t+τ ; G).
[0067] Step 2): Construct a mask matrix generation module for screening key nodes. The module screens nodes with high correlation based on the short-term and long-term correlations of node traffic flows and the adjacency relationships between nodes. The specific method for generating the mask matrix is as follows:
[0068] Step 2-1): Initialize the mask matrix The initial value is set to 0;
[0069] Step 2-2): Input the complete traffic flow historical data, calculate the similarity of the long-term traffic flow sequences between each pair of nodes through the Pearson correlation coefficient, select the k nodes with the highest similarity for each node, and set the corresponding positions in the mask matrix to 1;
[0070] Step 2-3): Input the short-term traffic flow historical data for the adjacent T time steps, calculate the similarity of the short-term traffic flow sequences between each pair of nodes through the Pearson correlation coefficient, select the k nodes with the highest similarity for each node, and set the corresponding positions in the mask matrix to 1;
[0071] Step 2-4): According to the topological structure of the traffic road network, select the reachable nodes within five hops of each node in the traffic road network as key nodes, and set the corresponding positions in the mask matrix to 1.
[0072] Step 3): Construct a data embedding module to map the original data into a high-dimensional space, which can enhance the expression ability of the data and introduce adaptive periodic feature embedding to improve the time correlation modeling ability of the network; the specific construction method is as follows:
[0073] Step 3-1): Use a fully connected network to project the original input into a high-dimensional space d is the feature dimension after projection;
[0074] Step 3-2): Initialize a learnable weekly periodic feature matrix This matrix assigns an independent d-dimensional feature vector to each day of the week, and data collected on the same day has the same weekly periodic feature vector;
[0075] Step 3-3): Initialize a learnable daily periodic feature matrix The traffic flow data in the dataset used in this method is recorded every five minutes. Therefore, a day is divided into 288 time steps. The date feature matrix assigns an independent d-dimensional feature vector to each time step, and the data collected at the same time step has the same daily cycle feature vector;
[0076] Step 3-4): For the input which contains traffic flow data of a total of T time steps, select the corresponding vector from the cycle matrix according to the data collection time to obtain the weekly cycle embedding and the daily cycle embedding Then add the two to to obtain the output of the data embedding module
[0077] Step 4), as Figure 2 (b) shows, construct a temporal convolutional self-attention module to extract temporal features from the output X (0) of the data embedding module. The specific construction method is as follows:
[0078] Step 4-1): Define a Temporal Convolutional Network (TCN). The TCN can effectively extract the local change trend in the traffic flow sequence, so as to assign higher attention weights to time steps with similar local change trends during the attention calculation process, improving the accuracy of temporal correlation modeling. For the traffic flow historical sequence of node n, the calculation formula of the TCN is:
[0079]
[0080] where i is the time step index, m is the input dimension index, k is the convolution kernel size, is the convolution kernel.
[0081] Step 4-2): Perform temporal convolutional self-attention calculation. Convert the output X (0) of the data embedding module into a query matrix Q T and a key matrix K T through the TCN, convert X (0) into a value matrix V T through a linear transformation, then perform scaled dot-product attention calculation on the query matrix Q T and the key matrix K T to obtain the temporal correlation matrix A T , then multiply A T with the value matrix V T to obtain the output of a single attention head; finally, concatenate the outputs of each attention head to obtain the output of the temporal convolutional self-attention module. The calculation process is as follows:
[0082] Q T = TCN Q (X (0) ), K T = TCN K (X (0) ), V T = X (0) W T
[0083]
[0084] head i = A T V T
[0085]
[0086] Among them, is a learnable parameter matrix, the Softmax function represents the normalization operation on the matrix, and the Concat function represents matrix concatenation, is the output of the time convolutional self-attention module.
[0087] Step 5), as Figure 2 (c) shows, construct a spatial convolutional self-attention module to add additional heterogeneous feature vectors to the nodes, and extract spatial features from the output X of the data embedding module based on the mask matrix (0) The specific construction method is as follows:
[0088] Step 5-1): Randomly initialize the learnable node heterogeneity feature matrix and concatenate it with the output X of the data embedding module (0) to obtain The node heterogeneity feature embedding assigns a node heterogeneity feature vector to each node in the traffic road network. This vector can be continuously learned during model training and is used to characterize the influence of spatial attributes such as road type, width, and local topology of the road network on node correlation;
[0089] Step 5-2): Define a Graph Convolutional Network (GCN). The GCN can capture the local spatial features of nodes according to the adjacency matrix of the traffic road network, so as to assign higher attention weights to nodes with similar local traffic patterns during the attention calculation process, improving the accuracy of spatial correlation modeling. The calculation formula of the GCN is:
[0090]
[0091] where ReLU is the activation function, is the adjacency matrix with self-loops added to the nodes of the traffic road network, and I is the identity matrix. is the corresponding degree matrix. is the learnable parameter matrix.
[0092] Step 5-3): Perform spatial convolutional self-attention calculation. Convert X (0) ' into the query matrix Q S and the key matrix K S through GCN, and convert X (0) into the value matrix V S through linear transformation. Then perform scaled dot-product attention calculation on the query matrix Q S and the key matrix K S , and perform Hadamard product with the output of the mask matrix generation module to obtain the spatial correlation matrix A S . Then multiply A S with the value matrix V S to get the output of a single attention head. Finally, concatenate the outputs of each attention head to obtain the output of the spatial convolutional self-attention module The calculation process is as follows:
[0093] Q S = GCN Q (X (0)′ ), K S = GCN K (X (0)′ ), V S = X (0) W S
[0094]
[0095] head i = A S V S
[0096]
[0097] where is the learnable parameter matrix, ⊙ is the Hadamard product, is the output of the spatial convolutional self-attention module.
[0098] Step 6), as shown in Figure 2 (a), concatenate the mask matrix generation module, data embedding module, temporal convolutional self-attention module, and spatial convolutional self-attention module, and add residual connections and fully connected layers to complete the network construction. The input of the network is The output of the network is The specific process is as follows:
[0099] Step 6-1): The input X first undergoes data embedding through the data embedding module to obtain
[0100] Step 6-2): Input X (0) into the temporal convolutional self-attention module and the spatial convolutional self-attention module respectively for spatio-temporal feature extraction to obtain and Adopt a fully connected network to fuse the spatio-temporal features. The calculation formula is:
[0101]
[0102] where is a learnable parameter matrix.
[0103] Step 6-3): Repeat the spatio-temporal feature extraction process in Step 6-2 a total of L times to finally obtain the spatio-temporal feature X (l) of the traffic flow, and input it into a two-layer fully connected network to obtain the final prediction result Y. The calculation formula is as follows:
[0104] Y = ReLU(ReLU(X (l) W O1 )W O2 )
[0105] where ReLU is the activation function, is a learnable parameter matrix.
[0106] Step 7), use the traffic flow dataset collected from the road network to train the network and evaluate the prediction accuracy of the network. The specific process is as follows:
[0107] Step 7-1): Initialize parameters such as the number of multi-head attention heads, learning rate, batch size, maximum number of iterations, and feature dimension;
[0108] Step 7-2): Divide the training set, validation set, and test set in a ratio of 6:2:2, and use the mean absolute error as the loss function;
[0109] Step 7-3): Use the training set to train the network, stop training when the maximum number of iterations is reached or the error no longer decreases, and use the test set to test the prediction accuracy of the network;
[0110] Step 8), load the optimal weight parameters, collect and input real-time data for traffic flow prediction.
[0111] To verify the prediction performance of the spatio-temporal convolutional self-attention network described in the present invention, based on the PeMS08 dataset, using the historical traffic flow observations in the previous hour (T = 12) to predict the traffic flow in the next hour (τ = 12), the mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE) are used as evaluation metrics. And the following models are selected for comparison:
[0112] T-GCN: A traffic prediction model that combines graph convolution and gated recurrent units.
[0113] GWNET: A traffic prediction model that combines diffusion convolution and gated temporal convolution layers.
[0114] STSGCN: Adopts a spatio-temporal synchronous graph convolutional network to simultaneously extract the spatio-temporal correlations between traffic nodes.
[0115] GMAN: A traffic flow prediction model based on the encoder-decoder architecture with a pure self-attention mechanism.
[0116] ASTGNN: A traffic flow prediction model that combines a temporal self-attention mechanism and dynamic graph convolution.
[0117] The experimental results on the PeMS08 dataset are shown in Table 1:
[0118] Table 1 Experimental results of each model on the PeMS08 dataset
[0119]
[0120] It can be seen from the experimental results that on the PeMS08 dataset, the error metrics of the method described in the present invention are all better than those of all comparison models. Compared with the ASTGNN with the highest prediction accuracy, the MAE, MAPE, and RMSE are reduced by 9.7%, 5.7%, and 7.1% respectively.
[0121] The above is only a preferred implementation manner of the present invention under the highway network dataset. The protection scope of the present invention is not limited to the above implementation manner. Any equivalent modifications and other modified variations made by those of ordinary skill in the art according to the content disclosed in the present invention shall be included in the protection scope recorded in the claims.
Claims
1. A traffic flow prediction method based on spatiotemporal convolutional self-attention network, characterized in that: The following steps are involved: Collect road network traffic flow data, perform preprocessing and normalization operations, and divide the data set into training set, validation set and test set; The mask matrix generation module, data embedding module, temporal convolution self-attention module and spatial convolution self-attention module are constructed respectively; the mask matrix generation module selects nodes with high correlation according to the short-term correlation, long-term correlation and adjacency between nodes of node traffic flow; the data embedding module maps the original data to a high-dimensional space and embeds an adaptive time period vector; the temporal convolution self-attention module adds one-dimensional convolution to the standard self-attention mechanism to capture the temporal correlation of traffic flow data; the spatial convolution self-attention module embeds additional node heterogeneity feature vectors for each node, selects key nodes based on the mask matrix, and adopts the self-attention mechanism of the fused graph convolution network to model the spatial correlation between nodes; The above modules are spliced together, and residual connections are added to stabilize the training gradient and accelerate network convergence. A fully connected layer is added to fuse the spatiotemporal features and output the prediction results to complete the network construction. The network is trained using the training set, with the mean absolute error as the loss function. Training is stopped when the maximum number of iterations is reached or the error no longer decreases, and the prediction accuracy of the network is tested using the test set. Collect real-time traffic flow data from the road network, input it into the trained network, and obtain traffic flow prediction results.
2. According to claim 1, a traffic flow prediction method based on spatiotemporal convolutional self-attention network is characterized in that: The method for generating a mask matrix by the mask matrix generating module includes: Initialize the mask matrix The initial value is set to 0; Input the complete historical traffic flow data, calculate the long-term traffic flow sequence similarity between nodes through the Pearson correlation coefficient, select the k nodes with the highest correlation for each node, and set the corresponding position in the mask matrix to 1; Input the short-term traffic flow history data of the adjacent T time steps, calculate the short-term traffic flow sequence similarity between nodes through the Pearson correlation coefficient, select the k nodes with the highest correlation for each node, and set the corresponding position in the mask matrix to 1; According to the topological structure of the traffic network, the reachable nodes within five hops of each node in the traffic network are selected as key nodes, and the corresponding positions are set to 1 in the mask matrix.
3. The traffic flow prediction method based on spatiotemporal convolutional self-attention network according to claim 1 is characterized in that: The operations of the data embedding module include: The original input is transformed into Projection to high-dimensional space Initialize the learnable periodic feature matrix Assign an independent d-dimensional feature vector to each day of the week; initialize a learnable daily cycle feature matrix Assign a separate d-dimensional feature vector to each time step in a day; According to the input X FC The acquisition time of each data in is selected, and the corresponding vector is obtained from the period matrix to obtain the period embedding and daily cycle embedding Add the two together to get the output of the data embedding module 4. The traffic flow prediction method based on spatiotemporal convolutional self-attention network according to claim 1 is characterized in that: The calculation process of the temporal convolution self-attention module includes: The data is embedded into the output X of the module through the temporal convolutional network (0) Convert to query matrix Q T and key matrix K T , through linear transformation, X (0) Convert to value matrix V T , where Q T =TCN Q (X (0) ), K T =TCN K (X (0) ), V T =X (0) W T , is a learnable parameter matrix; For the query matrix Q T and key matrix K T Perform scaled dot product attention calculation to get the temporal correlation matrix A T and value matrix V T Multiply them together to get the output head of a single attention head i =A T V T ; The outputs of each attention head are concatenated to obtain the output of the temporal convolutional self-attention module. The Concat function represents matrix concatenation.
5. The traffic flow prediction method based on spatiotemporal convolutional self-attention network according to claim 1, characterized in that: The spatial convolution self-attention module includes: Node heterogeneity feature embedding: The specific operation is to randomly initialize the node heterogeneity feature matrix and combine it with the output X of the data embedding module (0) Splice, get Spatial convolution self-attention calculation: X is converted into (0) 'Convert to query matrix Q S and key matrix K S , through linear transformation, X (0) Convert to value matrix V S , where Q S =GCN Q (X (0) ′),K S =GCN K (X (0) ′),V S =X (0) W S , is a learnable parameter matrix; For the query matrix Q S and key matrix K S Perform the scaled dot product attention calculation and generate the output of the module with the mask matrix Perform the Hadamard product to get the spatial correlation matrix A S and value matrix V S Multiply them together to get the output head of a single attention head i =A S V S ; The outputs of each attention head are concatenated to obtain the output of the spatial convolution self-attention module.
6. The traffic flow prediction method based on spatiotemporal convolutional self-attention network according to claim 1 is characterized in that: In the network construction step, the data is embedded into the output X of the module (0) Input the temporal convolution self-attention module and the spatial convolution self-attention module respectively to extract spatiotemporal features, and obtain and A fully connected network is used to fuse spatiotemporal features, and the calculation formula is: in, is a learnable parameter matrix; repeat the above spatiotemporal feature extraction process for L times, and finally obtain the spatiotemporal feature X of traffic flow (l) Input a two-layer fully connected network to get the final prediction result Y, the calculation formula is Y = ReLU (ReLU (X (l) W O1 )W O2 ),in, is the learnable parameter matrix.
Citation Information
Cited By
Multi-mode space-time traffic flow modeling method supporting large-scale road network real-time prediction
CN120337795A
Transform architecture-based traffic flow prediction method and system
CN120823711A