Transform-based traffic flow prediction method and device capable of sensing local-global time-space relationship

By combining the advantages of Transformer, TCN and GNN, a traffic flow prediction method based on Transformer, LGSTfromer is proposed, which solves the problem that existing methods are difficult to capture the spatial and temporal dependence of traffic data, and realizes accurate prediction of traffic flow data and effective capture of spatial and temporal dependence.

CN119990191AActive Publication Date: 2025-05-13ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510050036.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Existing traffic flow prediction methods are difficult to effectively capture the local-global space-time dependence in traffic data, especially when learning long-term and global spatial characteristics, they face problems such as gradient vanishing and error accumulation.

Method used

Combining the advantages of Transformer, TCN and GNN, a traffic flow prediction method based on Transformer is proposed. Through the spatial and temporal information embedding layer, local-global time dependency module and local-global space dependency module, short- and long-term time dependencies and local and global space dependencies in traffic data are captured.

Benefits of technology

Accurate prediction of traffic flow data is achieved, which can perform better overall prediction performance than the baseline model in short-medium-long-term, and can effectively capture the space-time dependence of complex time periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990191A_ABST
    Figure CN119990191A_ABST
Patent Text Reader

Abstract

The invention discloses a Transform-based traffic flow prediction method and device capable of sensing a local-global time-space relationship, and the method comprises the steps: firstly obtaining the historical traffic data of all to-be-predicted road sections, the historical data comprising the traffic flow feature data of the predicted road sections, the static adjacent matrix data between road network nodes, and the sampling timestamp data of the historical data; secondly, a spatio-temporal information embedding layer is constructed to provide multiple types of embedding input for a model trunk, the learning ability of the model is enhanced, and three different types of embedding are respectively historical data information embedding, time information embedding and space node self-adaptive embedding; then, constructing a local-global time dependence extraction module, respectively learning short-time and long-time time dependence relationships in the data by using a multi-scale TCN and a self-attention mechanism in a time dimension, and meanwhile, introducing a double-path self-adaptive information gating fusion technology to realize effective fusion of time features of different hierarchies; then constructing a local-global spatial dependency extraction module, respectively learning local and global spatial dependency relationships in the data by using a dynamic-static graph convolutional network and a self-attention mechanism in spatial dimension, and realizing effective fusion of spatial features of different levels based on a two-way adaptive information gating fusion technology; and finally, mapping the potential spatial-temporal feature representation into a prediction result through a full connection layer network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of data mining and artificial intelligence, and in particular to a method for predicting traffic flow data, in particular to a method and device for predicting traffic flow based on a Transformer that can perceive a local-global spatiotemporal relationship. Background Art

[0002] Accurate traffic flow data prediction is a key technology for data-driven intelligent transportation systems. Traffic flow prediction is a type of multivariate time series data prediction task. It requires predicting future traffic data based on historical traffic data recorded by traffic sensors, including vehicle flow prediction, road speed prediction, etc. It has an important impact on optimizing urban traffic management, travel efficiency, and traffic experience.

[0003] In recent years, with the increase in the number of private cars, the contradiction between people and the limited road network has become more and more serious. Alleviating this contradiction through certain means has also become a research hotspot. In recent years, scholars have proposed a large number of traffic flow prediction methods. In the early development of this field, due to technical limitations, only the time dependency of traffic data was considered. Some simple time series models such as recurrent neural network (RNN) and its variants long short-term memory network (LSTM) and gated recurrent unit (GRU) are used to model the time dependency of traffic data. However, the spatial dependency between different road network nodes is ignored. The existing view is generally that in addition to capturing temporal dynamic patterns such as periodicity and seasonality, the spatial dependency under the constraints of road network structure is more important for traffic flow prediction tasks. With the further development of deep learning, inspired by spectral graph neural networks, graph neural networks (GNNs) have been proposed for processing non-Euclidean irregular road network data. This method builds spatial connection relationships between nodes in different road sections based on the static topological structure of the road network, and realizes the flow of traffic information between nodes through the message propagation mechanism. This has become a classic and widely used model for traffic prediction tasks. For example, spatiotemporal graph convolutional networks (STGCNs), spatiotemporal fusion graph convolutional networks (STFGCNs), and dynamic frequency domain graph convolutional networks (DFDGCNs) either consider the impact of spatial relationships on the time dimension or try to capture spatial relationships and time relationships simultaneously. On the other hand, with the success of Transformer in many different fields (such as natural language processing, computer vision, data mining, etc.), traffic flow data prediction methods based on Transformer have also been widely studied, which is essentially different from GNNs. Thanks to its core self-attention mechanism, Transformer can effectively capture the spatiotemporal dependencies in traffic data by calculating the weight coefficients between different nodes in different dimensions (Self-attention, SA). This has become a powerful spatiotemporal relationship modeling tool for traffic flow prediction tasks.For example, there are traffic transformers (Traffic Transformer), spatiotemporal transformer network (STTN), etc. There are also methods that consider embedding spatiotemporal information in the transformer, embedding the time points of historical data with the road network structure or nodes, and inputting the same data into the transformer structure, such as spatiotemporal adaptive embedding transformer (STAEformer), propagation delay-aware transformer (PDformer), etc.

[0004] The essence of spatiotemporal dependency modeling lies in how to describe the transmission of traffic information between different nodes. However, road network traffic behavior usually has huge uncertainty, and the transmission of traffic information is not only static at the local level, but also dynamic at the global level. The above mainstream traffic flow prediction methods are effective, but they are challenging in learning the local and global spatiotemporal dependencies of traffic flow. 1) For RNN and temporal convolutional network (TCN), their advantage is to learn short-term time features; when learning long-term time features, they often face problems such as gradient vanishing and error accumulation. 2) For GNN, its advantage is that it can capture local spatial features well based on the static adjacency relationship of the road network; while learning global spatial features often depends on the quality of the learned global dynamic adjacency relationship matrix. 3) For Transformer, relying on its internal self-attention mechanism SA, it is good at capturing global spatiotemporal features; however, the impact on local spatiotemporal features is often not well modeled. Summary of the invention

[0005] Aiming at complex traffic flow data, the present invention overcomes the above-mentioned shortcomings of the prior art by combining the advantages of TCN and GNN in local spatiotemporal relationship mining and the advantages of Transformer in global spatiotemporal relationship mining, and provides a Transformer-based traffic flow prediction method and device that can perceive local-global spatiotemporal relationship, so as to achieve accurate prediction of future traffic flow data.

[0006] The first aspect of the present invention relates to a Transformer-based traffic flow prediction method capable of perceiving local-global spatiotemporal relationships (LGSTformer), comprising the following steps:

[0007] S1: Obtain historical traffic data of all road sections to be predicted, including traffic flow feature data of the predicted road section, static adjacency matrix data between road network nodes, and sampling timestamp data of historical data. The above data are merged to obtain a historical sample data set, and the historical sample data set is further divided into a training data set, a validation set, and a test set.

[0008] S2: Based on the training sample data set, a spatiotemporal information embedding layer is constructed to provide multiple types of embedding inputs for the model backbone, including historical data feature embedding, time periodicity embedding, and spatial node adaptive embedding, so as to enhance the spatiotemporal dependency learning ability of the subsequent spatiotemporal module;

[0009] S3: Based on the embedded data obtained by the embedding layer, learning time features by constructing a local-global time dependency module, wherein the time features include short-term and long-term time dependencies in the traffic flow data;

[0010] S4: Based on the time feature data learned by the time module, spatial feature learning is performed by constructing a local-global spatial dependency module, wherein the spatial feature includes local and global spatial dependencies in the traffic flow data;

[0011] S5: Mapping the future traffic flow data result of the final road section to be predicted through a fully connected layer network based on the spatiotemporal features;

[0012] S6: Experimental verification, by conducting experimental comparison with multi-type traffic flow prediction methods on various types of traffic flow data sets, the effectiveness of the present invention is demonstrated.

[0013] The third aspect of the present invention relates to a Transformer-based traffic flow prediction device capable of perceiving local-global spatiotemporal relationships, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the Transformer-based traffic flow prediction method capable of perceiving local-global spatiotemporal relationships of the present invention.

[0014] The third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the Transformer-based traffic flow prediction method of the present invention that is capable of perceiving local-global spatiotemporal relationships.

[0015] The beneficial effects of the present invention are as follows:

[0016] 1) A novel Transformer-based traffic data prediction model, LGSTfromer, is proposed, which can accurately capture the local-global spatiotemporal dependencies in historical traffic data.

[0017] 2) Two new spatiotemporal modules LGTM and LGSM are designed to capture spatiotemporal dependencies of different ranges. The former combines TCN and temporal self-attention mechanism to learn short-term and long-term temporal dependencies, while the latter combines DSGCN and spatial self-attention mechanism to learn local-global spatial dependencies.

[0018] 3) Performance comparison experiments were conducted on four real-world data sets with high prediction complexity. The results showed that the method of the present invention can achieve better overall prediction performance in the short, medium and long term than the baseline model. The visualization results on different data sets show that the present invention can fit the changing trend of the data well and can accurately capture the spatiotemporal dependencies of complex time periods, which can further verify the effectiveness of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a structural diagram of the present invention.

[0020] Figure 2 This is a visualization result diagram of the present invention on the node Index=40 in the data set METR-LA from 2012 / 6 / 8 to 2012 / 6 / 9.

[0021] Figure 3 This is a visualization result diagram of the present invention on the node Index=120 in the data set METR-LA from 2012 / 6 / 8 to 2012 / 6 / 9.

[0022] Figure 4 This is a visualization result diagram of the present invention on 2017 / 4 / 26-2017 / 4 / 27 at the node Index=18 in the data set PEMS-BAY.

[0023] Figure 5 This is a visualization result diagram of the present invention on 2017 / 4 / 26-2017 / 4 / 27 at the node Index=149 in the data set PEMS-BAY.

[0024] Figure 6 This is a visualization result diagram of the present invention on 2018 / 2 / 18-2018 / 2 / 19 at the node Index=25 in the data set PEMS04.

[0025] Figure 7 This is a visualization result diagram of the present invention on 2018 / 2 / 18-2018 / 2 / 19 at the node Index=35 in the data set PEMS04.

[0026] Figure 8 This is a visualization result diagram of the present invention on 2016 / 8 / 20-2016 / 8 / 21 at the node Index=15 in the data set PEMS08.

[0027] Fig. 9 This is a visualization result diagram of the present invention on 2016 / 8 / 20-2016 / 8 / 21 at the node Index=35 in the data set PEMS08. DETAILED DESCRIPTION

[0028] In order to make the objectives, technical solutions and advantages of the present invention more clear, specific embodiments of the present invention are further described in detail below with reference to the accompanying drawings.

[0029] Example 1

[0030] This embodiment provides a Transformer-based traffic flow prediction method that can perceive the local-global spatiotemporal relationship. The model structure is as follows: Figure 1 As shown, the method includes:

[0031] S1: Obtain historical traffic data for all road sections to be predicted. The steps are as follows:

[0032] S1.1: The road segments in the road network to be predicted are regarded as basic node units, which are represented as a node set V = {v1, v2, …, v N}∈R N , where N represents the number of road segments to be predicted. The physical topology of the road can be expressed as G = {V, P, A s}, where P is the edge set that determines the node connection relationship, A s is the static neighbor matrix of the traffic graph structure. T}∈R T The historical traffic flow data is represented as X = {X1, X2, ..., X N}∈R T×N×D , where D represents the number of features. Furthermore, the data to be predicted in the future can be expressed as

[0033] S1.2: Divide the above historical data set into a training data set, a validation data set, and a test data set according to a preset ratio. Based on the preset loss function of the training data set, establish and train the traffic flow data prediction model. The preset loss function on the traffic dormitory data set is HuberLoss, and the preset loss function on the traffic flow data set is MAELoss.

[0034] S2: Construct a spatiotemporal information embedding layer to provide various types of embedding inputs for the model backbone. The steps are as follows:

[0035] S2.1: Perform time periodic embedding on the sampling timestamp data. This embedding mainly includes embedding the timestamp within a day. dayand the embedding E of the days of the week week The former is related to the sampling frequency of the data. If the sampling frequency of the data set is 5 minutes in a day, then a day will be divided into 288 timestamps; the latter is the number of days in a week, that is, 7 days. First, two learnable parameter matrices W are defined 288×E and W 7×E To represent the embedded dictionary, where E is the dimension size of the embedding. Then use the input daily timestamp data X day ∈R T and weekly data X week ∈R T As the index, get E day ∈R T×N×E and E week ∈R T×N×E . Where T and N represent the length of historical data time and the number of nodes respectively.

[0036] S2.2: Adaptive embedding of spatial nodes for spatial feature data. Traffic flow has dynamic spatiotemporal correlations, and it is insufficient to extract spatial relationships based only on static road network adjacency graphs. In order to better capture the dynamic changes in time and space of traffic flow, a spatial node adaptive embedding matrix e with trainable parameters is introduced. spa ∈R N×ESA To handle the interaction between different nodes, the input is then aligned through the broadcast mechanism to obtain E spa ∈R T×N×ESA , where ESA represents the dimension size of spatial node embedding.

[0037] S2.3: High-dimensional feature embedding of historical data information of road segment nodes. In order to maintain the original information of the original data and maintain dimension alignment with the spatiotemporal embedding, only one fully connected layer FC() is used to project the original data into a high-dimensional latent space to obtain a new feature representation H X =FC(X 1:T )∈R T×N×E , where X 1:T Represents historical input data of length T.

[0038] S2.4: Finally, by concatenating these four comprehensive embeddings, we obtain the spatiotemporal representation H of the traffic input data. In =E day ||E week ||E spa ||H X ∈R T×N×Ein , where E in is the final sum of the embedding dimensions.

[0039] S3: Local-global time dependency module LGTM, which realizes the extraction of time dependency relationships in different ranges. Based on the embedded data obtained by the embedding layer, the local-global time dependency module is constructed to learn time features, and the time features include short-term and long-term time dependencies in traffic flow data. The steps are as follows:

[0040] S3.1: Short-term temporal dependency extraction. As analyzed before, the mainstream tools for modeling short-term temporal dependencies include recursive neural networks (RNNs) and temporal convolutional networks (TCNs). The sequential computing method of RNNs faces these two problems: the gradient vanishing problem and the inability to perform parallel computing. Compared with RNNs, TCNs are more computationally efficient. It uses convolutional networks (causal convolutions and dilated convolutions) to process data, which enables it to perform parallel computing, and can adjust the perception range of local temporal features by changing the size of the receptive field, which is more flexible. Therefore, the present invention introduces a multi-scale convolution module (Multi-scale Gated Tanh Unit, M-GTU) to capture the short-term temporal dynamic information of traffic flow data. It is mainly composed of three gated Tanh units (GTUs) with different receptive fields. The mathematical expression of the calculation process of a single GTU is as follows:

[0041]

[0042] in is the input of LGTM, Γ K is a convolution kernel with a kernel size of 1×K, and *τ is a gated convolution operator. φ() and σ() are the nonlinear activation Tanh function and Sigmoid function, respectively, which are used as a gating mechanism to control the output of the data flow, which has been proven to be very effective for TCN. Convolution kernel Γ K Will The channel characteristic number E in Doubled to 2E in , and then the channel features are divided in a compromise to obtain F and B, so as to ensure that the output features are consistent in the channel dimension. Further, three different convolution kernels can be used to obtain the mathematical expression of M-GTU:

[0043]

[0044] where Γ K1 , Γ K2 and Γ K3 They are convolution kernels with kernel sizes of 1×K1, 1×K2, and 1×K3 respectively. || represents the connection operation, which connects the outputs of three GTUs of different scales in the time dimension. The fully connected layer FC() is responsible for restoring the time dimension of the connection output to T. It is the short-term time dependency learned by M-GTU.

[0045] S3.2: Extraction of long-term temporal dependencies. M-GTU learns short-term sequence changes, but cannot further learn the long-term dependencies of the road network. In order to better capture long-term temporal dependencies, the present invention uses the scaled dot product self-attention mechanism commonly used in the classic Transformer model to calculate the importance of each timestamp position in the entire input sequence in the time dimension, and fuses the temporal information of the entire input sequence through the importance coefficient. According to the formal definition, the query matrix Q required for the self-attention mechanism calculation can be derived through three different linear mapping layers: temp , key matrix K temp Sum value matrix V temp The mathematical expression of the attention calculation matrix can be expressed as follows:

[0046]

[0047] in are different trainable parameters. Then, we can temp and K temp To calculate the self-attention score Attn between different timestamps in the time dimension temp , thereby realizing the extraction of global time dependency. Specifically, by temp and K temp T (K temp The transpose of ) is used for matrix multiplication, and the result is normalized to [0,1] through the Softmax function. The mathematical expression of the whole process can be expressed as follows:

[0048]

[0049] So, Attn temp (Q temp ,K temp ) can adaptively learn the global temporal dependencies between different timestamps. temp It is possible to obtain the representation of long-term temporal dependency features Its mathematical expression can be expressed as follows:

[0050]

[0051] Among them, in order to be able to Keep the dimensions consistent. Furthermore, a multi-head attention mechanism (MHA) is introduced to capture more diverse temporal features. Specifically, when calculating the temporal self-attention score Attn tempPreviously, the query matrix Q was usually temp , key matrix K temp Sum value matrix V temp Divide into some finer-grained block structures and assign different trainable parameters. This method is more effective than the non-multi-head attention mechanism. Its mathematical expression can be expressed as follows:

[0052]

[0053] head i =Attn temp,i (Q temp,i ,K temp,i )V temp,i (7)

[0054] Among them, Q temp,i , K temp,i and V temp,i Represents Q temp , K temp and V temp The ith partition block of . temp Then represents the weight matrix used to combine the outputs of all attention heads.

[0055] S3.3: Obtaining short-term time dependence and long-term time dependence Finally, we need to consider how to effectively fuse them. On the one hand, we need to consider the importance of different time dependencies, and on the other hand, we also need to consider the ratio of information fusion between the two. To achieve this goal, the present invention introduces a dual-path adaptive information gating fusion mechanism, which is mainly composed of two gating mechanisms, namely, a hold gate and a reset gate. The hold gate controls the information flow on a single path, and the reset gate controls the fusion ratio of information on different paths. and After two gating mechanisms, the output fusion obtains the time-dependent features The mathematical expression of the whole process can be expressed as follows:

[0056]

[0057] Among them, φ() and σ() are the nonlinear activation Tanh function and Sigmoid function respectively.

[0058] S3.4: Consistent with the original Transformer structure design, LGSTformer also retains the original residual connection layer (Add), layer normalization (Norm) and position-level feedforward layer (FFN) in the LGTM module. Existing experience shows that Add can ensure the flow of input information, avoid gradient vanishing and network degradation; Norm stabilizes the model training process and accelerates convergence; FFN performs nonlinear transformation on the representation of each position to enhance the expressiveness of the model. Time-dependent features After the above three operations, the input of LGSM is output from the LGTM submodule. The mathematical expression of the whole process can be expressed as follows:

[0059]

[0060] Among them, LN and Dropout are LayerNorm and Dropout method functions respectively. FFN consists of two linear layers with ReLU nonlinear operators.

[0061] S4: Local-global time dependency module LGSM, to extract spatial dependencies in different ranges. Based on the time feature data learned by the time module, the spatial feature learning is performed by constructing a local-global spatial dependency module, and the spatial feature includes the local and global spatial dependencies in the traffic flow data. The steps are as follows:

[0062] S4.1: Local spatial dependency extraction. In recent years, the use of graph convolutional neural networks (GCNs) based on adjacency matrices to model local spatial dependencies of traffic data is one of the most commonly used modeling methods. GCNs can handle direct connections between nodes and model local dependencies by propagating information between adjacent nodes. Therefore, in the same way, the present invention also uses this method to process local dependencies, and its mathematical expression is as follows:

[0063]

[0064] Among them, GCN() represents the graph convolution operation, represents the input of the Lth layer LGSM, A s The static adjacency matrix of the road network nodes. On the other hand, the local spatial structure may be affected by short-term external factors such as bad weather and traffic accidents. This may cause changes in the dependencies between local nodes. Considering this influence is also beneficial for modeling traffic data, the spatial node adaptive embedding is also used to model the unconventional local changes in spatial relationships as a reference for A s In addition, GCN() is used to propagate node information. Its mathematical expression can be expressed as follows:

[0065]

[0066] A d =Softmax(relu(φ(ΦΦ T )))∈R N×N (16)

[0067]

[0068] Among them, φ(), relu() and softmax() represent Tanh, ReLU and Softmax method functions respectively. is a learnable parameter. Subsequently, a simple fully connected layer FC() is used to extract the local spatial information of different representations. Its mathematical expression can be expressed as follows:

[0069]

[0070] in, Represents the learned local spatial dependency features.

[0071] S4.2: Global spatial dependency extraction. Another characteristic of traffic networks is the highly dynamic changes in traffic flow, which widely affects multiple regions. Some traffic patterns (such as overall traffic flow, the impact of major traffic events, etc.) involve the interaction of all nodes or multiple regions in the network and they may be indirectly dependent on each other. Similar to the global temporal dependency learning method, we use a scaled dot product self-attention mechanism to better capture the global spatial dependency. Given an input First, three different linear layers are used to map it into different feature spaces to obtain the query matrix Q spa , key matrix K spa Sum value matrix V spa The mathematical expression of the attention calculation matrix can be expressed as follows:

[0072]

[0073] in are different trainable parameters. Then, we can spa and K spa To calculate the self-attention score Attn between different nodes in the spatial dimension spa , thereby realizing the extraction of global spatial dependencies. The mathematical expression of the whole process can be expressed as follows:

[0074]

[0075] So, Attn spa (Q spa ,K spa) can adaptively learn the global spatial dependencies between different nodes. Then, combined with V spa A representation of the global spatial dependence feature can be obtained Its mathematical expression can be expressed as follows:

[0076]

[0077] Furthermore, as in LGTM, a multi-head attention mechanism MHA is introduced to capture richer spatial information from different representation subspaces. Its mathematical expression can be expressed as follows:

[0078]

[0079] head i =Attn spa,i (Q spa,i ,K spa,i )V spa,i (26)

[0080] Among them, Q spa,i , K spa,i and V spa,i Represents Q spa , K spa and V spa The ith partition block of . spa Then represents the weight matrix used to combine the outputs of all attention heads.

[0081] S4.3: Obtaining Local Spatial Dependence and global space dependencies Finally, we need to consider how to effectively fuse them. The same method is used in LGTM to fuse short-term temporal dependencies and long-term temporal dependencies. and The fusion also adopts a dual-path adaptive information gating fusion method. and After two gating mechanisms, the spatial dependency features obtained by fusion are output The mathematical expression of the whole process can be expressed as follows:

[0082]

[0083] Among them, φ() and σ() are the nonlinear activation Tanh function and Sigmoid function respectively.

[0084] S4.4: As done in LGTM, a residual connection layer (Add), layer normalization (Norm), and position-level feed-forward layer (FFN) are introduced to stabilize the training of the entire module. Spatial Dependent Features After the above three operations, the output from the LGSM submodule is The mathematical expression of the whole process can be expressed as follows:

[0085]

[0086] Among them, LN and Dropout are LayerNorm and Dropout method functions respectively. FFN consists of two linear layers with ReLU nonlinear operators.

[0087] S5: Output layer. After the two submodules LGTM and LGSM capture the short-term and long-term temporal dependencies and the local-global spatial dependencies, a fully connected layer network is used to map the potential spatiotemporal feature representation to the prediction results of future data. The mathematical expression of the whole process can be expressed as follows:

[0088]

[0089] Among them, W out is a learnable parameter, b out It's a deviation.

[0090] S6: Experimental verification. The present invention will evaluate the traffic prediction capabilities of the output results of multiple types of traffic prediction models (including traditional methods, STGNN-based methods, spatiotemporal regularization-based methods, linear-based methods, and Transformer-based methods) on different traffic flow data sets (including flow data sets and speed data sets) to demonstrate the effectiveness of the present invention.

[0091] S6.1: Dataset. The present invention conducted experiments on four real-world datasets, PEMS-BAY, METR-LA, PEMS04 and PEMS08, which are all benchmark test sets for traffic prediction. It is worth noting that short (15min)-medium (30min)-long (60min) predictions were performed on PEMS-BAY and METR-LA, while long (60min) predictions were performed on PEMS04 and PEMS08. PEMS-BAY, PEMS04 and PEMS08 were all collected by the US Performance Measurement System (PeMS), but at different collection locations. METR-LA is a Los Angeles highway dataset. The basic information of the four datasets is shown in the following table:

[0092]

[0093] S6.2: Evaluation indicators. In order to evaluate the prediction data x of different models pred and the true value x trueIn order to quantify the differences between the two, this paper uses three commonly used evaluation indicators: mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE) to quantify all experimental results. The mathematical expression formulas of these indicators are as follows:

[0094]

[0095]

[0096] S6.3: Comparison model. In order to demonstrate the performance of the present invention in predicting different types of traffic data, the present invention will be compared with five different types of traffic prediction models. 1) The traditional prediction method HA (Historical Average), which uses the historical average as the prediction value. 2) STGNN-based methods include GWNet, DCRNN, AGCRN, GTS, SGODE, STFGCN, DFDGCN, etc. The characteristics of this type of method are to use the graph structure as spatial information and combine it with time information to extract spatiotemporal dependencies. 3) STNorm, a method based on spatiotemporal regularization, focuses on using normalization techniques to balance the different characteristics of the spatiotemporal dimensions in traffic flow data. 4) STID, a linear-based method, attempts to use a multilayer perceptron (MLP) to design a simple and effective model that can compete with STGNN. 5) Transformer-based methods include PDformer, STAEformer, MTESformer, etc. This type of method uses the spatiotemporal self-attention mechanism to learn the spatiotemporal dependencies in traffic data.

[0097] S6.4: Experimental results. The present invention evaluates the performance of the three prediction dimensions (horizon) of short-term, medium-term and long-term on the two velocity datasets of METR-LA and PEMS-BAY. It is worth noting that the best result is bolded, the second best result is underlined, and the third best result is marked with an asterisk. The comparison results are as follows:

[0098]

[0099]

[0100] According to the experimental results table on the speed dataset, it can be found that: 1) For the traditional HA method, the prediction performance in both METR-LA and PEMS-BAY is the worst. This is because it only averages the historical data and cannot effectively extract the spatial and temporal dependencies in non-stationary data. 2) For the STGNN-based method, SGODE achieved the best results in short-term and medium-term length prediction of METR-LA, and DFDGCN was second only to LGSTfromer in short-term and medium-term length prediction of PEMS-BAY. Other methods performed generally. 3) For methods based on spatiotemporal regularization, STNorm outperformed METR-LA on PEMS-BAY. In addition, STNorm's performance on METR-LA is lower than that of the STGNN-based method, while its performance on PEMS-BAY is better than that of the STGNN-based method. This is because METR-LA has more complex changes and spatiotemporal relationships, which makes it more difficult to balance spatiotemporal features. 4) For linear-based methods, STID outperforms STNorm on PEMS-BAY, thanks to STID's spatiotemporal embedding technology. 5) For Transformer-based methods, the overall performance of the LGSTformer method proposed in this invention is the best among all models, and it also performs best in long-term prediction. Compared with STAEformer, in the long-term prediction of METR-LA and PEMS-BAY, the RMSE index of LGSTformer is reduced by 0.43% and 2.3%, the MAE index is relatively reduced by 0.3% and 2.13%, and the MAPE index is relatively reduced by -0.40% and 3.17%. This proves that considering local dependencies is effective. Compared with MTESformer, the RMSE index of LGSTformer is reduced by 2.10% and 3.20%, and the MAE index is relatively reduced by 1.19% and 1.60%.

[0101]

[0102]

[0103] The MAPE index is relatively reduced by -1.25% and 1.84%. This proves the effectiveness of the introduced dual-path adaptive information fusion method for local-global information fusion. Then, according to the experimental results table analysis on the traffic dataset, it can be found that: 1) For the traditional method HA, the prediction performance is similar to that of the speed dataset, and the results on PEMS04 and PEMS08 are the worst. 2) For the STGNN-based method, DFDGCN performs best on PEMS04, while STFGCN performs better on PEMS08. Compared with the earlier GWNet, DCRNN, and AGCRN methods, STFGCN and DFDGCN use more complex dynamic graph construction methods to model the dynamic spatial dependencies of traffic data. 3) For methods based on spatiotemporal regularization, STNorm performs better than GWNet, DCRNN, and AGCRN on PEMS04, but worse than DFDGCN and STFGCN; while it is slightly lower than the STGNN method on the PEMS08 dataset. 4) For linear-based methods, STID outperforms STGNN-based methods and data decomposition-based methods on the PEMS04 dataset; and its performance on the PEMS08 dataset is similar to that of DFDGCN. This further proves that good results can be achieved by embedding temporal and spatial information in the model. 5) For Transformer-based methods, this type of method performs better than other methods on the PEMS04 and PEMS08 datasets. This shows that the self-attention mechanism of the Transformer model can better capture spatiotemporal dependencies. The LGSTfromer proposed in this invention performs best on PEMS04. Compared with STAEformer, the RMSE indicator of LGSTformer is reduced by 0.70%, the MAE indicator is relatively reduced by -0.05%, and the MAPE indicator is relatively reduced by 0.33%.

[0104]

[0105]

[0106] S6.5: Visualization. To further demonstrate the prediction performance of the present invention, the prediction results are visualized on four data sets. Figure 2-8 As shown. Data from two nodes are selected for display in each data set, the prediction length is Horizon 12 (60 minutes), and the visualized data time range is 2 days. It can be seen that the LGSTformer model can fit the changing trend of the data well, whether it is the changing trend on the speed data set or the flow data set, which shows that it can accurately capture the spatiotemporal dependencies of complex time periods.

[0107] In summary, the implementation application cases show that the Transformer-based traffic flow prediction method capable of perceiving local-global spatiotemporal relationships proposed in the present invention is effective. Compared with other design methods, the present invention fully considers the spatiotemporal dynamics of traffic data. In view of the temporal dependencies of traffic data, multi-scale temporal convolution and self-attention mechanism in the time dimension are used to learn short-term and long-term temporal dependencies respectively; in view of the dynamic spatial dependencies of traffic data, dynamic and static graph convolution and self-attention mechanism in the spatial dimension are used to learn local and global spatial dependencies respectively. In order to more effectively fuse dependencies at different levels, a two-way adaptive information gating fusion mechanism is introduced to achieve this. The experiment used multiple real traffic flow data, and the results fully demonstrated the feasibility and superiority of the model.

[0108] Example 2

[0109] The third aspect of the present invention relates to a Transformer-based traffic flow prediction device capable of perceiving local-global spatiotemporal relationships, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the Transformer-based traffic flow prediction method capable of perceiving local-global spatiotemporal relationships of Example 1.

[0110] Example 3

[0111] This embodiment relates to a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the traffic flow prediction method based on Transformer and capable of perceiving local-global spatiotemporal relationship of Embodiment 1 is implemented.

[0112] The above description is a specific embodiment of the present invention and the technical principles used. If the changes made according to the concept of the present invention do not exceed the spirit covered by the description and drawings, they should still fall within the scope of protection of the present invention.

Claims

1. A Transformer-based traffic flow prediction method capable of perceiving local-global spatiotemporal relationships, comprising the following steps: S1: Obtain historical traffic data of all road sections to be predicted, wherein the historical data includes traffic flow characteristic data of the predicted road section, static adjacency matrix data between road network nodes, and sampling timestamp data of historical data; fuse the above data to obtain a historical sample data set, and further divide the historical sample data set into a training data set, a validation set, and a test set; S2: Based on the training sample data set, a spatiotemporal information embedding layer is constructed to provide multiple types of embedding inputs for the model backbone, including historical data feature embedding, time periodicity embedding, and spatial node adaptive embedding, so as to enhance the spatiotemporal dependency learning ability of the subsequent spatiotemporal module; S3: Based on the embedded data obtained by the embedding layer, learning time features by constructing a local-global time dependency module, wherein the time features include short-term and long-term time dependencies in the traffic flow data; S4: Based on the time feature data learned by the time module, spatial feature learning is performed by constructing a local-global spatial dependency module, wherein the spatial feature includes local and global spatial dependencies in the traffic flow data; S5: Based on the spatiotemporal features, a fully connected layer network is used to map out the future traffic flow data results of the final road section to be predicted.

2. The Transformer-based traffic flow prediction method capable of perceiving local-global spatiotemporal relationships as claimed in claim 1, characterized in that: Step S1 includes the following steps: S1.1: The road segments in the road network to be predicted are regarded as basic node units, which are represented as a node set V = {v1, v2, …, v N }∈R N , where N represents the number of road segments to be predicted; the physical topological structure of the road can be expressed as G = {V, P, A s }, where P is the edge set that determines the node connection relationship, A s is the static neighbor matrix of the traffic graph structure; the road section is in a specific time period T = {t1, t2, …, t T }∈R T The historical traffic flow data is represented as X = {X1, X2, ..., X N }∈R T×N×D , where D represents the number of features; further, the data to be predicted in the future can be expressed as S1.2: Divide the above historical data set into a training data set, a verification data set and a test data set according to a preset ratio; Based on the preset loss function of the training data set, a traffic flow data prediction model is established and trained; the preset loss function on the traffic dormitory data set is HuberLoss, and the preset loss function on the traffic flow data set is MAELoss.

3. The Transformer-based traffic flow prediction method capable of perceiving local-global spatiotemporal relationships as claimed in claim 1, characterized in that: Step S2 comprises the following steps: S2.1: Time periodic embedding of sampling timestamp data, mainly including embedding of timestamps within a day day and the embedding E of the days of the week week ; Define two learnable parameter matrices W 288×E and W 7×E To represent the embedded dictionary; then use the input daily timestamp data X day ∈R T and weekly data X week ∈R T As the index, get E day ∈R T×N×E and E week ∈R T×N×E ; S2.2: Perform spatial node adaptive embedding on spatial feature data and introduce a spatial node adaptive embedding matrix e with trainable parameters spa ∈R N×ESA To handle the interaction between different nodes, the input is then aligned through the broadcast mechanism to obtain E spa ∈R T×N×ESA ; S2.3: Perform high-dimensional feature embedding on the historical data information of the road segment nodes, and use a fully connected layer FC() to project the original data into a high-dimensional latent space to obtain a new feature representation H X =FC(X 1:T )∈R T×N×E ; S2.4: Concatenate the above four embeddings to get the comprehensive embedding H In , the obtained spatiotemporal representation of the input data is The comprehensive embedding is then input into the spatiotemporal dependency extraction module.

4. The Transformer-based traffic flow prediction method capable of perceiving local-global spatiotemporal relationships as claimed in claim 1, characterized in that: Step S3 includes the following steps: S3.1: Based on the comprehensive embedding H In To extract short-term temporal dependencies, a multi-scale convolution module MGTU is introduced to capture the short-term temporal dynamic information of traffic flow data; the mathematical expression of the calculation process of a single GTU is as follows: in is the input of LGTM, Γ K is a convolution kernel with a kernel size of 1×K, *τ is a gated convolution operator; φ() and σ() are the nonlinear activation Tanh function and Sigmoid function respectively; further, three different convolution kernels can be used to obtain the mathematical expression of M-GTU: where Γ K1 , Γ K2 and Γ K3 They are convolution kernels with kernel sizes of 1×K1, 1×K2, and 1×K3 respectively; || represents the connection operation, which connects the outputs of three GTUs of different scales in the time dimension; the fully connected layer FC() is responsible for restoring the time dimension of the connection output to T; S3.2: Based on the comprehensive embedding H In Extract long-term temporal dependencies and use the scaled dot product self-attention mechanism commonly used in the classic Transformer model to calculate the importance of each timestamp position in the entire input sequence in the time dimension, and fuse the temporal information of the entire input sequence through the importance coefficient; The query matrix Q required for the self-attention mechanism calculation is derived through three different linear mapping layers temp , key matrix K temp Sum value matrix V temp ; Afterwards, based on Q temp and K temp To calculate the self-attention score Attn between different timestamps in the time dimension temp , thereby realizing the extraction of global time dependence; then, combined with V temp It can obtain the representation of long-term temporal dependency features The mathematical expression of the whole process can be expressed as follows: in, are all different trainable parameters; further, the Multi-head Attention (MHA) mechanism is introduced to capture more diverse temporal features; its mathematical expression can be expressed as follows: head i =Attn temp,i (Q temp,i ,K temp,i )V temp,i (7) Among them, Q temp,i , K temp,i and V temp,i Represents Q temp , K temp and V temp The i-th partition block of temp represents the weight matrix used to combine the outputs of all attention heads; S3.3: Dependence on the short-term timing described and long-term temporal dependencies The time-dependent features are fused and a dual-path adaptive information gating fusion mechanism is introduced. The structure mainly consists of two gating mechanisms, namely the hold gate and the reset gate. The hold gate controls the information flow on a single path, and the reset gate controls the fusion ratio of information on different paths. and After two gating mechanisms, the output fusion obtains the time-dependent features The mathematical expression of the whole process can be expressed as follows: Among them, φ() and σ() are the nonlinear activation Tanh function and Sigmoid function respectively; S3.4: Consistent with the original Transformer structure design, LGSTformer also retains the original residual connection layer (Add), layer normalization (Norm) and position-level feedforward layer (FFN) in the LGTM module; time-dependent features After the above three operations, the input of LGSM is output from the LGTM submodule.

5. The Transformer-based traffic flow prediction method capable of perceiving local-global spatiotemporal relationships as claimed in claim 1, characterized in that: Step S4 comprises the following steps: S4.1: Based on the temporal dependency features, local spatial dependency is extracted and GCN based on static graph structure is used to process local dependency. The mathematical expression is as follows: Among them, GCN() represents the graph convolution operation, represents the input of the Lth layer LGSM, A s The static adjacency matrix of the road network nodes is used to model the spatial relationship of the unconventional local changes based on spatial node adaptive embedding, as a s In addition, GCN() is used for node information propagation; its mathematical expression can be expressed as follows: A d =Softmax(relu(φ(ΦΦ T )))∈R N×N (16) Among them, φ(), relu() and softmax() represent Tanh, ReLU and Softmax method functions respectively; is a learnable parameter; then, a simple fully connected layer FC() is used to extract the local spatial information of different representations; its mathematical expression can be expressed as follows: in, Represents the learned local spatial dependency features; S4.2: Extract global spatial dependency based on the time-dependent features, use a scaled dot product self-attention mechanism to better capture the global spatial dependency, and calculate the correlation between nodes in the entire road network in the spatial dimension; the overall calculation process is similar to S3.2; S4.3: Local spatial dependence on the and global space dependencies The spatial dependency features are obtained by fusion, and the dual-path adaptive information gating fusion mechanism in S3.3 is introduced. and After two gating mechanisms, the spatial dependency features obtained by fusion are output S4.4: Consistent with the LGTM submodule design, the original residual connection layer (Add), layer normalization (Norm) and position-level feedforward layer (FFN) are also retained in the LGTM module; spatial dependency features After the above three operations, the output from the LGSM submodule is 6. The Transformer-based traffic flow prediction method capable of perceiving local-global spatiotemporal relationships as claimed in claim 1, characterized in that: The step S5 specifically includes: After the two submodules LGTM and LGSM capture the short-term and long-term temporal dependencies and the local-global spatial dependencies, a fully connected layer network is used to map the potential spatiotemporal feature representation to the prediction results of future data. The mathematical expression of the whole process can be expressed as follows: Among them, W out is a learnable parameter, b out It's a deviation.

7. A Transformer-based traffic flow prediction device capable of perceiving local-global spatiotemporal relationships, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the traffic flow prediction method based on Transformer and capable of perceiving local-global spatiotemporal relationship as described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the traffic flow prediction method based on Transformer and capable of perceiving local-global spatiotemporal relationship described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Trajectory prediction method for multi-dimensional spatio-temporal feature fusion for automatic driving

    CN118296090A

  • Space-time adaptive dynamic graph convolutional network traffic flow prediction method

    CN118629226A

  • Transform traffic flow prediction method and device based on space-time fusion embedding, computer readable storage medium and product

    CN118886526A

Cited By

  • Self-adaptive space-time diagram neural network missing data completion method

    CN120256845A

  • An Adaptive Spatiotemporal Graph Neural Network Missing Data Completion Method

    CN120256845B

  • Multi-scale space-time fusion traffic flow prediction method based on coupling graph

    CN120319030A

  • Traffic flow prediction method and device

    CN120452209A

  • Road traffic flow prediction method based on space-time mixed attention network

    CN120526608A