A traffic flow prediction method, system, device and medium

CN120564411BActive Publication Date: 2026-09-08INNER MONGOLIA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510658379.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2026-09-08
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

[0004]然而,基于图神经网络的方法还存在以下问题:图构建方面的信息不充分,很多现有方法在构建图时仅利用给定的空间邻近性等简单信息,忽略了节点间在时间上的相似性等其他潜在关系

Benefits of technology

本发明通过TCN能够在不增加过多参数的情况下扩大感受野,挖掘交通流量数据在时间序列上的局部动态变化,从而捕捉时间序列中的长短期依赖关系,提取出时间特征;通过GAT能够计算交通网络邻接矩阵中每个节点与邻居节点的注意力系数,通过注意力机制能够学习节点之间的重要性;通过注意力系数对邻居节点进行聚合,得到交通网络的空间特征,能够得到该节点更新后的特征表示,从而能够自适应地捕捉节点之间的空间依赖关系;本发明能够更好地融合时间和空间特征,提高交通流的预测精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564411B_ABST
    Figure CN120564411B_ABST
Patent Text Reader

Abstract

The application provides a traffic flow prediction method, system, device and medium, and belongs to the technical field of traffic flow prediction. The method comprises the following steps: acquiring traffic flow data and a traffic network adjacency matrix of different time scales; combining a time convolution network (TCN) and a graph attention network (GAT) to obtain a traffic flow prediction model GAT-TCN; inputting the traffic flow data of different time scales into the GAT-TCN, using the TCN to mine the local dynamic change of the traffic flow data on a time sequence, and extracting time features; inputting the time features and the traffic network adjacency matrix into the GAT, using the GAT to calculate the attention coefficients of each node and neighbor nodes in the traffic network adjacency matrix, aggregating the neighbor nodes according to the attention coefficients, and obtaining spatial features of the traffic network; and fusing the time features and the spatial features, and predicting the traffic flow according to the fused data. The application can improve the prediction accuracy of the traffic flow.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of traffic flow prediction technology, specifically relating to a traffic flow prediction method, system, device, and medium. Background Technology

[0002] The rapid urbanization process in my country has brought about problems such as traffic congestion. Accurate traffic flow forecasting can provide valuable suggestions for urban traffic management and vehicle scheduling, thereby reducing traffic accidents.

[0003] Currently, traffic flow prediction methods can be broadly categorized into statistical methods, machine learning methods, and deep learning methods. Statistical methods are only applicable to short-term predictions in simple environments. While machine learning methods still hold an important position, they can only handle medium-scale data and traffic flow patterns with low complexity. Deep learning has become the mainstream trend, demonstrating significant advantages such as automatic feature extraction, complex data structure processing, large-scale data adaptability, and strong model generalization ability. Deep learning methods based on graph neural networks have shown remarkable advancement, with core advantages manifested in several aspects: First, graph neural networks possess powerful spatiotemporal relationship modeling capabilities, accurately representing complex spatial connections in traffic networks. Second, they exhibit excellent adaptability to the complex and dynamically changing structure of traffic networks, susceptible to factors such as road construction, accidents, and special events, handling traffic networks of varying sizes and complexities. Third, they enable efficient information propagation and aggregation within graph structures, fully utilizing global structural information for learning and reasoning, thereby uncovering more valuable features for traffic flow prediction. Fourth, they are highly compatible with other deep learning methods, further enhancing traffic flow prediction performance through combinations with recurrent neural networks and convolutional neural networks.

[0004] However, graph neural network-based methods still suffer from the following problems: insufficient information in graph construction. Many existing methods only utilize simple information such as spatial proximity when constructing graphs, ignoring other potential relationships such as temporal similarity between nodes. For example, some spatially distant nodes may have high correlation during specific time periods (such as morning and evening rush hours), but graphs built based on simple spatial proximity cannot capture this relationship, which affects the model's understanding and prediction of complex spatiotemporal data such as traffic flow. Furthermore, the constructed graph structures are usually static, making it difficult to adapt to the dynamically changing spatiotemporal dependencies in data such as traffic flow. In reality, the state of traffic networks is constantly changing, and the connections and influence between nodes also change over time. Static graph structures cannot accurately reflect this dynamism, limiting the model's ability to learn and predict dynamic spatiotemporal features and affecting prediction accuracy. Summary of the Invention

[0005] To overcome the shortcomings of the existing technology, the present invention provides a traffic flow prediction method, comprising the following steps: Obtain traffic flow data and traffic network adjacency matrix at different time scales; By combining the Temporal Convolutional Network (TCN) and the Graph Attention Network (GAT), a traffic flow prediction model, GAT-TCN, is obtained. Traffic flow data at different time scales are input into GAT-TCN. TCN is used to mine the local dynamic changes of traffic flow data in the time series and extract time features. The time features and traffic network adjacency matrix are input into GAT. GAT is used to calculate the attention coefficient between each node in the traffic network adjacency matrix and its neighboring nodes. The neighboring nodes are aggregated based on the attention coefficient to obtain the spatial features of the traffic network. By fusing temporal and spatial features, traffic flow can be predicted based on the fused data.

[0006] Preferably, the step of using TCN to mine local dynamic changes in traffic flow data over time and extract time features specifically involves: using TCN to transform the channel dimension of the input data, using temporal convolution to slide the transformed data through a convolution kernel with specific parameters over time to mine local dynamic changes in the time series, and outputting time features through temporal convolution.

[0007] Preferably, the step of fusing temporal and spatial features and predicting traffic flow based on the fused data includes the following steps: GAT's temporal attention mechanism adjusts the weights of each time step in the temporal features based on spatial features, and fuses the spatial and temporal features according to the weights. GAT's residual connection then fuses the traffic flow data with the obtained fused features again to generate prediction results.

[0008] Preferably, the traffic flow data includes weekly traffic flow data, daily traffic flow data, and recent traffic flow data.

[0009] Preferably, before inputting traffic flow data at different time scales into GAT-TCN, the method further includes training GAT-TCN using a Generative Adversarial Network (GAN). The loss function used for training includes auxiliary loss terms related to specific attributes of traffic flow. These auxiliary loss terms include a spatiotemporal correlation loss term, a road attribute matching loss term, and a special scene feature loss term. The spatiotemporal correlation loss term is used to measure the correlation between generated data and real data in the time dimension. The road attribute matching loss term is used to calculate the error between the generated traffic flow data and the real traffic flow features under the same road attribute. The special scene feature loss term is used to calculate the difference between the generated special scene traffic flow data and the real special scene features.

[0010] The spatiotemporal correlation loss term is determined by the following formula: ; In the formula, The mean absolute error of spatial characteristics, The mean square error represents the time characteristic. for The weight parameters, for Weight parameters; The road attribute matching loss term is determined by the following formula: ; In the formula, These are the weighting coefficients for speed features in the calculation of road attribute matching loss. The mean square error of the velocity characteristic. These are the weighting coefficients for traffic flow features in the calculation of road attribute matching loss. The mean square error of the flow characteristics. The density feature is the weighting coefficient in the calculation of road attribute matching loss. The mean square error of the density feature; The feature loss term for special scenarios is determined by the following formula: ; In the formula, For special scene feature loss, To highlight the differences in characteristics under extreme weather scenarios, To differentiate features in high-traffic fluctuation scenarios. To highlight the unique characteristics of high-traffic scenarios during major events, To highlight the unique characteristics of specific flow scenarios during major events, , ν , ξ , η These are the weight coefficients for the feature differences corresponding to specific scenarios.

[0011] The present invention also provides a traffic flow prediction system, comprising: The data acquisition module is used to acquire traffic flow data and traffic network adjacency matrix at different time scales; The model building module is used to combine the Temporal Convolutional Network (TCN) and the Graph Attention Network (GAT) to obtain the traffic flow prediction model GAT-TCN. The feature extraction module is used to input traffic flow data at different time scales into GAT-TCN, use TCN to mine the local dynamic changes of traffic flow data in the time series, and extract time features; input the time features and traffic network adjacency matrix into GAT, use GAT to calculate the attention coefficient between each node in the traffic network adjacency matrix and its neighboring nodes, and aggregate the neighboring nodes according to the attention coefficient to obtain the spatial features of the traffic network. The traffic flow prediction module is used to fuse temporal and spatial features and predict traffic flow based on the fused data.

[0012] The present invention also provides a computer device including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to perform the traffic flow prediction method.

[0013] The present invention also provides a computer-readable storage medium storing a computer program adapted for loading by a processor to execute the traffic flow prediction method.

[0014] The traffic flow prediction method provided by this invention has the following beneficial effects: This invention utilizes TCN to expand the receptive field without adding too many parameters, thereby uncovering local dynamic changes in traffic flow data over time series and capturing long-term and short-term dependencies to extract temporal features. GAT (Generative Attention Scale) calculates the attention coefficient between each node and its neighbors in the traffic network adjacency matrix, enabling the learning of node importance through the attention mechanism. Aggregating neighboring nodes using these attention coefficients yields the spatial features of the traffic network, providing an updated feature representation for each node and allowing for adaptive capture of spatial dependencies between nodes. This invention better integrates temporal and spatial features, improving the accuracy of traffic flow prediction. Attached Figure Description

[0015] To more clearly illustrate the embodiments and design schemes of the present invention, the accompanying drawings required for this embodiment will be briefly described below. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart of the traffic flow prediction method according to an embodiment of the present invention; Figure 2 A comparison chart of prediction performance before and after enhancement; Figure 3 For the GAT model; Figure 4This is a multi-head attention mechanism; Figure 5 For TCN model; Figure 6 The effects of different parameters on the experiment; among them, Figure 6 (a) represents the influence of feature dimensions; Figure 6 (b) represents the influence of head count on the multi-head attention mechanism; Figure 6 (c) represents the effect of kernel size; Figure 6 (d) represents the effect of the expansion coefficient; Figure 7 Ablation experiments were conducted on PEMS04; among which, Figure 7 (a) represents the mean absolute error of each model's prediction in PEMS04; Figure 7 (b) represents the root mean square error of each model's prediction in PEMS04; Figure 7 (c) represents the mean absolute percentage error of each model in PEMS04 prediction.

[0017] Figure 8 Ablation experiments were conducted on PEMS08; among which, Figure 8 (a) represents the mean absolute error of each model's prediction in PEMS08; Figure 8 (b) represents the root mean square error of each model's prediction in PEMS08; Figure 8 (c) represents the mean absolute percentage error of each model in PEMS08 predictions.

[0018] Figure 9 The fitted curves of the predicted and actual values ​​of the GAT-TCN model on PEMS04; Figure 10 This is the fitted curve of the predicted and actual values ​​of the GAT-TCN model on PEMS08. Detailed Implementation

[0019] To enable those skilled in the art to better understand and implement the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be construed as limiting the scope of protection of the present invention.

[0020] Example This invention provides a traffic flow prediction method, specifically as follows: Figure 1 As shown, it includes the following steps: Step 1: Obtain traffic flow data and traffic network adjacency matrix at different time scales.

[0021] Step 2: Combine the Temporal Convolutional Network (TCN) and the Graph Attention Network (GAT) to obtain the traffic flow prediction model GAT-TCN.

[0022] (1) Generative Adversarial Network (GAN).

[0023] To enhance the model's adaptability to special environments, this invention uses a Generative Adversarial Network (GAN) to augment the PEMS08 dataset to train the GAT-TCN model, thereby improving its robustness. The GAN mainly consists of a generator and a discriminator. The generator takes random noise as input and undergoes complex transformations through a multi-layered neural network architecture and non-linear activation functions to learn the distribution patterns of real traffic flow data, generating traffic flow samples similar to the real data to expand the dataset. The discriminator receives both real traffic flow data and fake data generated by the generator. It is also composed of a multi-layered neural network, learning the features of the input data and using the output of a single neuron to determine the authenticity of the data. In adversarial training, the generator and discriminator are trained alternately. The generator makes the generated data more difficult for the discriminator to identify as fake, while the discriminator continuously improves its ability to distinguish between real and fake data. Their objective functions are, respectively, minimizing the probability of the generator identifying the generated data as fake, and maximizing the probability of correctly distinguishing real and fake data. Through this continuous game-theoretic approach, the GAN model can effectively generate high-quality new data with a distribution similar to real traffic flow data.

[0024] The objective function of the generator is as follows: ; In the formula, For discriminator, For generator, The expectation operator (representing the expectation of a probability distribution) random noise vector (Seek expectation) It is a random noise vector. for The probability distribution function.

[0025] The objective function of the discriminator is as follows: ; In the formula, For expected operators, For the true data probability distribution, It is a discriminator.

[0026] For the generator, the parameters are updated during training based on the gradient of the objective function. Let the parameters of the generator be... The update formula (taking stochastic gradient descent as an example) is: ; in, It is the learning rate of the generator. It is the objective function Regarding generator parameters The gradient.

[0027] For the discriminator, let its parameters be... The update formula (again, using stochastic gradient descent as an example) is: ; in It is the learning rate of the discriminator. It is the objective function Regarding discriminator parameters The gradient.

[0028] During training, in addition to traditional adversarial losses, auxiliary loss terms related to specific traffic flow attributes are added, enabling the generator to produce data that more closely resembles actual traffic flow patterns. The generated data exhibits stronger periodicity (such as morning and evening rush hour patterns) and spatial correlation (the mutual influence of traffic flow on adjacent road segments). This allows the GAT-TCN model to utilize this more realistic data during training, better capturing the spatiotemporal characteristics of traffic flow and thus reducing errors. The generated data contains these latent feature information, from which the GAT-TCN model can learn deeper feature representations.

[0029] The auxiliary loss items mainly include the following three types of losses: The spatiotemporal correlation loss term considers the spatiotemporal characteristics of traffic flow data and designs a loss term based on time series prediction error. By calculating the correlation between generated data and real data in the time dimension, as well as their consistency in different spatial locations, it is ensured that the data conforms to the spatiotemporal variation patterns of traffic flow.

[0030] To measure the correlation between generated and real data over time, a time series prediction model is first constructed. This invention uses GAT-TCN to train an LSTM by inputting real traffic flow data (PEMS08) into the LSTM in time series, allowing it to learn the patterns of traffic flow changes over time. The trained time series prediction model is then used to predict the real data, obtaining a sequence of predicted values. For generated data, it is similarly input into the model in chronological order to obtain corresponding predicted values. The error between the predicted values ​​of the real data and the predicted values ​​of the generated data is then calculated. The mean squared error (MSE) is used as the metric, as shown in the following formula:

[0031] ; In the formula, Mean square error, For the sample size, For the first i The true predicted value of each sample For the first i Generate predicted values ​​for each sample.

[0032] Spatial Dimension Consistency Calculation: The transportation network is divided into different regions, each of which can be considered a node. For each node, its spatial neighborhood is defined, which is the set of adjacent nodes. For example, in an urban road network, adjacent intersections or road segments can constitute the spatial neighborhood of a node.

[0033] Spatial consistency measure: Calculates the consistency between generated data and real data within their spatial neighborhood. Node j and its neighboring nodes k For both real and generated data, calculate the traffic flow difference between these two nodes. and The consistency of the differences is then measured using the mean absolute error (MAE), as shown in the following formula:

[0034] ; In the formula, It is a measure of spatial consistency (mean absolute error of spatial features). This represents the number of samples or the total number of nodes in the node set. K The number of neighboring nodes considered for each node. N J For nodes J The set of neighboring nodes.

[0035] We obtain the spatiotemporal correlation loss by weighted summing of the errors in the time dimension and the spatial dimension: ; In the formula, for The weight parameters, for The weight parameters.

[0036] The road attribute matching loss term is designed based on different road attributes (such as road type, number of lanes, etc.). For different types of roads, the generated data is required to have corresponding traffic flow characteristics to ensure that the data conforms to the actual road operation. For the generated traffic flow data, the error between its corresponding road attribute and the actual traffic flow characteristics under that road attribute is calculated. For example, for traffic flow data generated from a road assumed to be an urban arterial road, the difference between its average speed and the historical average speed of urban arterial roads, as well as the difference between its flow rate and the historical average flow rate, are calculated. The mean squared error is used to calculate these feature matching errors. For speed features, the formula is as follows:

[0037] ; In the formula, The mean square error of the velocity characteristic. n It is the number of samples generated. It is the historical average speed of the city's main roads. Is the data generated in the first... i The speed in each sample is weighted and summed to obtain the road attribute matching loss. Considering three features—speed, flow, and density—the formula for the road attribute matching loss is as follows: ; In the formula, These are the weighting coefficients for speed features in the calculation of road attribute matching loss. These are the weighting coefficients for traffic flow features in the calculation of road attribute matching loss. The mean square error of the flow characteristics. The density feature is the weighting coefficient in the calculation of road attribute matching loss. The mean square error of the density feature.

[0038] The special scenario feature loss term extracts specific traffic flow features for extreme scenarios such as extreme weather and major events, focusing on low vehicle speeds during extreme weather and high traffic flow fluctuations during major events to ensure the data retains these key features. For the generated traffic flow data in these special scenarios, the difference between the data and the features of the actual special scenarios is calculated. Taking low vehicle speed features under extreme weather as an example, the difference between the average vehicle speed of the generated data and the historical average vehicle speed under extreme weather conditions is calculated, using absolute error as the metric. The formula is as follows:

[0039] ; In the formula, Indicates absolute error (used to measure the difference between the average vehicle speed in the generated data and the average vehicle speed in historical extreme weather conditions). This represents the average vehicle speed under actual extreme weather conditions in history. This represents the average vehicle speed generated in the data.

[0040] The differences in features across different specific scenarios are weighted and summed to obtain the loss of features specific to those scenarios. , Considering the four characteristics of low vehicle speed and high traffic flow fluctuations under extreme weather conditions, as well as high traffic flow and specific flow directions under major events, the formula for the feature loss of special scenarios is as follows: ; In the formula, For special scene feature loss, To highlight the differences in characteristics under extreme weather scenarios, To differentiate features in high-traffic fluctuation scenarios. To highlight the unique characteristics of high-traffic scenarios during major events, To highlight the unique characteristics of specific flow scenarios during major events, , ν , ξ , η These are the weight coefficients for the feature differences corresponding to specific scenarios.

[0041] Experimental Design: This invention trains the GAT-TCN model using both the original dataset before and the dataset after enhancement. The same training parameters are set: learning rate of 0.001, 50 iterations, and batch size of 32. Mean absolute error (MAE) is used as the evaluation metric to predict traffic flow over the next 12 time steps. The model's performance in specific scenarios is calculated, and the prediction performance of the models trained on the pre- and post-enhancement datasets is compared across different scenarios. Experimental results are as follows: Figure 2 As shown, from Figure 2 As can be seen, the performance metrics of the augmented model are significantly better than those before augmentation, indicating that data augmentation effectively improves the model's adaptability and predictive performance in special scenarios.

[0042] (2) Graph Attention Network (GAT).

[0043] like Figure 3 As shown, the Graph Attention Network (GAT) mainly consists of node feature input, attention mechanism calculation, and feature update output. This module is based on graph-structured data, where nodes represent different elements in the traffic network (such as road segments and intersections), and edges reflect the relationships between them. For each node, GAT first transforms its original features into a more computationally suitable form through a linear transformation. It then calculates the attention coefficients between the node and its neighbors using the attention mechanism. This process, based on factors such as the similarity of node features, uses learnable parameters to determine the importance of each neighbor node to the current node, adaptively capturing spatial correlations in the traffic network. Finally, it performs a weighted summation of the neighbor node features based on the calculated attention coefficients to obtain the updated node features. This updated feature contains both its own information and incorporates important feature information from neighbor nodes, enabling GAT to effectively extract spatial features from graph data and adapt to the complex topology and data characteristics of traffic networks. The core formula is as follows:

[0044] ; In the formula, The nodes are the output after calculation. i eigenvectors; K The number of attention heads indicates the number of attention heads running in parallel; σ The activation function is used to introduce a nonlinear transformation; common examples include ReLU. In the first k In each attention head, the nodei With nodes j The attention weights between nodes reflect the attention weights between them. j For nodes i The degree of importance under this attention; W k : This is the weight matrix corresponding to the k-th attention head, used to weight the input features. x j Perform a linear transformation; x j For input nodes j eigenvectors; b k For the first k The bias vector corresponding to each attention head is used for adjustment after linear transformation; N i For nodes i The set of neighboring nodes, i.e., the nodes participating in the computation. i The set of nodes related to attention weights.

[0045] like Figure 4 As shown, multi-head attention is a feature extraction and information fusion method that processes data by running multiple attention heads in parallel. Each attention head has its own independent parameters, enabling it to focus on different parts or features of the input data from different perspectives. During computation, for a given input (such as node features in traffic flow data), it is first mapped to a query, key, and value vector space through a linear transformation. Each attention head calculates the similarity between the query vector and the key vector, obtaining attention weights. These weights determine the importance of different elements in the value vector to the current computation. By weighted summing of the value vectors, the output features of each attention head are obtained. Finally, the output features of multiple attention heads are concatenated or linearly combined to obtain the final multi-head attention output. This mechanism can more comprehensively capture complex relationships and patterns in the data, enhancing the model's ability to focus on different feature dimensions and semantic information.

[0046] (3) Temporal Convolutional Network (TCN).

[0047] Temporal Convolutional Networks (TCNs) are primarily used for processing time-series data. The key internal detail lies in the temporal convolution operation. This module consists of a series of convolutional layers, where the convolutional kernel in each layer slides along the temporal dimension to process the input traffic flow time-series data. The convolutional kernel has a specific size and stride; the size determines the time range covered by each convolutional operation, and the stride controls the interval at which the convolutional kernel moves across the time series. To effectively capture dependencies over long time periods, TCNs employ dilated causal convolution. This convolutional method can extract correlation features between distant time points by increasing the receptive field without relying on future information.

[0048] Meanwhile, to avoid the vanishing gradient problem common in deep networks and improve training efficiency, TCN uses residual connections to add the output of the previous layer to the output of the current layer after convolution and other operations. After each convolutional layer, an activation function is added to increase non-linearity, and batch normalization is used to accelerate convergence and stabilize the training process, effectively extracting temporal features from traffic flow time-series data. Specifically, as shown below... Figure 5 As shown:

[0049] (a) Causal convolution.

[0050] This module primarily utilizes causal convolution to ensure no information leakage occurs when processing time series data. Causal convolution stipulates that the convolution output at any given time depends only on the input values ​​from the past and current times. Let the input time series be... The filter is ,in k Let be the filter length, then the causal convolution at time is t Output It can be represented as:

[0051] ; (b) Dilated convolution.

[0052] To handle longer time series, TCN introduces dilated convolution. Dilated convolution increases the receptive field by inserting zeros between filter elements. For the dilation factor... d Dilated convolution at time 1 t The output formula can be expressed as:

[0053] ; In this way, TCN can expand the receptive field of convolution without significantly increasing computation, thereby effectively capturing long-range temporal dependencies. Furthermore, TCN employs residual connections, assuming the input is... The output is obtained after passing through multiple convolutional layers. The output after residual connection is This helps alleviate the vanishing gradient problem, enabling the model to train deeper network structures and further improve its ability to model time series data.

[0054] Regarding the global spatiotemporal dependency module, in scenarios such as traffic flow prediction, it can simultaneously consider long-distance dependencies in both temporal and spatial dimensions, effectively integrating global information. The global spatiotemporal dependency module aims to model temporal and spatial information in a unified manner. In the spatial dimension, TCN uses a graph structure to represent the connections between different traffic nodes (such as intersections and road segments). This spatial relationship is described using an adjacency matrix A, where... Represents a node i and nodes j The connection weights between nodes. For the feature vector of each node... This involves updating node features through spatial aggregation operations, enabling each node to access information from its neighboring nodes. In the temporal dimension, a mechanism similar to TCN is employed, using causal convolution and dilated convolution to process node features at different time steps. This is achieved by using causal convolution and dilated convolution to process node features at different time steps. Assuming the node feature matrix varies across time steps... t for After processing by the spatiotemporal module, a new feature matrix is ​​obtained. Specifically, in the spatiotemporal interaction process, the fusion operation adds the features obtained after temporal convolution to the features obtained after spatial aggregation. The formula can be expressed as:

[0055] ; in, and These represent time and space processing functions, respectively. This approach allows the module to capture global spatiotemporal dependencies, meaning that the characteristics of a node at the current time are affected by the characteristics of other nodes at different times.

[0056] The model employs residual connections to combine the initial features with features that have undergone a series of complex operations (Dropout - ReLU - WeightNorm - Causal Convol - Dropout - ReLU - WeightNorm - Causal Convol). In traffic flow prediction, the initial features may contain basic temporal and spatial characteristics of the original traffic flow data, such as initial flow magnitude, simple time period features, and basic relationships between adjacent nodes. After operations such as temporal convolution, graph attention, and temporal attention, the features become more complex and abstract, containing deeper spatiotemporal correlation information.

[0057] Residual connections combine these two types of features, ensuring that the final features used for prediction retain both the basic characteristics of the original data and the complex information from deep mining. This approach effectively prevents information loss caused by excessive layers and complex operations in deep neural networks. Compared to models without residual connections, it better utilizes all the information in the input data, improving prediction accuracy.

[0058] The traffic flow prediction model GAT-TCN is an architecture that integrates Graph Attention Network (GAT), Temporal Convolutional Network (TCN), and Generative Adversarial Network (GAN). It is specifically designed for processing and predicting traffic flow data, which possesses both spatiotemporal characteristics. In the generator, weekly, daily, and recent traffic flow data are integrated with the traffic network adjacency matrix as input. Temporal features are extracted using the temporal convolutional operation of TCN. Then, GAT aggregates neighbor node features based on the attention coefficients between nodes to incorporate spatial structure information. Subsequent temporal attention operations accurately capture the importance of each time step and perform weighted processing. Residual connections fuse the initial and processed features to preserve the original information. Finally, normalization and activation functions generate the prediction result. The discriminator consists of multiple fully connected layers that receive the generator's prediction results and real data. These are progressively transformed through fully connected layers and activation functions, and finally, a sigmoid activation function outputs a numerical value representing the probability of the data being real. During model training, the discriminator uses BCELoss (loss function) to distinguish between real and fake data and optimizes its parameters, while the generator uses MSELoss (loss function) to measure the difference from the real data and uses adversarial BCELoss to make the generated data fool the discriminator. The two train against each other through an adversarial training mechanism, which improves the generator's predictive ability and makes the entire model exhibit high accuracy and generalization in spatiotemporal data processing. When processing temporal and spatial features, the model does not simply concatenate or add them, but deeply integrates them through temporal and spatial attention mechanisms. After using GAT to capture spatial attention, temporal attention operations are further used to explore the relative importance between different time steps. This sequential operation allows the model to adjust the weights of different time steps based on the incorporated spatial information.

[0059] When predicting traffic flow at an intersection, if spatial attention reveals traffic congestion at adjacent intersections due to road construction, the temporal attention stage will focus more on the time steps following the occurrence of congestion. Features from these time steps are weighted, allowing the model to more flexibly respond to changes in traffic conditions—a refined feature fusion and processing approach not found in previous models. Previous traffic flow prediction models focused solely on time-series features or spatial structure information. This model combines Temporal Convolutional Networks (TCNs) and Graph Attention Networks (GATs). TCNs effectively capture the dynamic changes in traffic flow data over time, such as periodicity and trends. It can analyze the patterns of traffic flow changes at different times of the day (morning peak, evening peak, etc.) and the differences in traffic flow between weekdays and weekends within a week using convolutional kernels at different time scales.

[0060] GAT focuses on uncovering the spatial dependencies between nodes (such as intersections and road segment monitoring points) in traffic networks. In the complex graph structure of a traffic network, the traffic flow of each node is influenced by its surrounding nodes. GAT, through its attention mechanism, can adaptively learn the degree of correlation between different nodes. For example, the traffic flow at an intersection on a main road significantly impacts the traffic flow at adjacent side road intersections. GAT effectively captures this relationship and learns spatial relationships from multiple perspectives through multi-head attention mechanisms, enhancing the richness of spatial feature representation. This fusion of TCN and GAT allows the model to simultaneously consider the spatiotemporal characteristics of traffic flow, integrating temporal variations and spatial interactions when predicting traffic flow. Compared to previous models that only considered a single dimension, it provides a more comprehensive prediction.

[0061] Step 3: Input traffic flow data at different time scales into GAT-TCN, use TCN to mine the local dynamic changes of traffic flow data in the time series and extract time features; input the time features and traffic network adjacency matrix into GAT, use GAT to calculate the attention coefficient between each node in the traffic network adjacency matrix and its neighboring nodes, aggregate the neighboring nodes according to the attention coefficients to obtain the spatial features of the traffic network; fuse the time features and spatial features, and predict traffic flow based on the fused data.

[0062] (4) Combining GAT and TCN.

[0063] The GAT-TCN model of this invention is a spatiotemporal architecture combining GAT, TCN, and GAN. It first integrates weekly, daily, and recent traffic flow data with the traffic network adjacency matrix as input. Temporal features are extracted through the temporal convolution operation of TCN. Then, the attention coefficients between nodes in GAT aggregate neighbor node features to incorporate spatial structure information. Subsequent temporal attention operations capture the importance of each time step and perform weighted processing. Residual connections fuse the initial and processed features to retain the original information. Finally, a fully connected layer generates the prediction result. After capturing spatial attention using GAT, temporal attention operations are further used to explore the relative importance between different time steps. This sequential operation allows the model to adjust the weights of different time steps based on the incorporated spatial information. When considering traffic flow prediction at an intersection, if spatial attention reveals traffic congestion at an adjacent intersection due to road construction, the temporal attention stage will focus more on the time steps after the congestion occurs, weighting the features of these time steps to enable the model to respond more flexibly to changes in traffic conditions. In the adversarial training process, the generator and discriminator of a generative adversarial network continuously analyze and learn from the data, which can uncover the potential features of traffic flow and improve the representation ability of the features.

[0064] like Figure 1 As shown, the model first receives traffic flow data at different time scales as input. After entering the model, this data undergoes initial operations in the TCN (Traffic Channel Network) section to normalize the data. Perform the initial convolution operation according to the convolution formula: ; in, n Indicates the sample index. t Representing the time step index, we get , This represents the weight parameters in the initial feature aggregation stage, where, i This indicates that, in the time dimension, it is relative to the current time step. t The offset is used to aggregate information from adjacent time steps in the time dimension. j Represents a spatial location index. c This represents the feature channel index, used to operate on features from different channels. Different channels can capture different types of feature information.

[0065] Then, a temporal convolution operation is performed, considering information from the current time step and its adjacent time steps, to extract local dynamic changes in the time series. Feature_ 1. Perform temporal convolution operations: ; In the formula, This represents the weight parameters in the feature extraction stage.

[0066] This operation considers the information of the current time step and its adjacent time steps with an interval of 2, and mines the local dynamic changes in the time series to obtain... .

[0067] Spatial structural information is incorporated through GAT. For nodes i and j Its eigenvectors First, through the weight matrix Perform a linear transformation on the node features to obtain , Attention coefficient The calculation is as follows:

[0068] ; in, It is the attention weight vector, used softmax Function normalization: ; in, It is a node i The set of neighboring nodes is used to measure the importance of neighboring nodes to the current node. Then, the features of neighboring nodes are aggregated based on the attention coefficient to update the node's features. An activation function that measures the element i and elements j The degree of correlation between them.

[0069] For the m The node is updated by aggregating neighbor node features based on the attention coefficient. i Features : ; in, It is the activation function. After three layers of graph convolution operations, each layer repeats the above process of calculating attention coefficients and feature aggregation. The attention mechanism in the GAT module captures spatial relationship information from different perspectives, and finally the feature representations obtained from multi-head attention are merged to obtain... Feature_ 3:

[0070] ; In the formula, m This represents the head index in a multi-head attention mechanism. i Typically represents a node index, Indicates the first m The first index is calculated in the graph. i The characteristic representation of a node index.

[0071] This integrates GAT into the model, combining spatial information with the initial temporal features extracted by TCN. The result is a model incorporating spatial information. Feature_ After step 3, it undergoes a subsequent time-attention operation, which, based on the spatial information already incorporated, further explores the relative importance between different time steps. Feature_ 3. Perform linear transformation:

[0072] ; in, This is a linear transformation matrix. The temporal features extracted by TCN and the spatial features incorporated by GAT interact at this stage. When calculating the temporal attention coefficients, the node features that have already incorporated spatial information are taken into account, allowing the temporal attention to better adjust the weights of different time steps according to spatial relationships. Calculate the temporal attention coefficients:

[0073] ; The weights of the fully connected layer are: , bias is The activation function is According to the time attention coefficient We obtain the weighted summation of features at different time steps. Feature_ 4: ; In the formula, represents the characteristic tensor Z Feature information from all channels of all samples at time step t is selected and used for temporal attention coefficients. a time After multiplying, a weighted sum is performed to obtain the final Feature_4.

[0074] This fused feature representation, Feature_4, is combined with the original Feature_1 via a residual connection: ; Finally, the prediction results are output through a fully connected layer. Through this fusion approach, the model can fully utilize the spatiotemporal characteristics of traffic flow data, taking into account both the changing patterns over time and spatial structural information when predicting traffic flow.

[0075] Model output results: Generator output The data enters the discriminator and is calculated via the fully connected layer. The calculation formula is as follows: ; in, This indicates the discriminator calculation process. Represents computation for different fully connected layers. This is the final predicted traffic flow value.

[0076] Example 2 (1) Experimental setup.

[0077] Experimental environment: The operating system was Windows 11, and the hardware environment consisted of a desktop computer equipped with an NVIDIA GeForce RTX 3060 graphics card. The PyTorch deep learning library was used, as detailed in Table 1. Table 1 Experimental Operating Environment This invention uses the PEMS04 and PeMS08 datasets to experimentally validate the traffic flow prediction model. Both datasets originate from the PeMS system of a local transportation bureau. The PEMS04 dataset contains data from 3848 detectors on 29 highways in the area from January 1, 2018 to February 28, 2018. These detectors collected data every 5 minutes, recording the number of vehicles passing by each of the 3848 sensors every 5 minutes. The PEMS08 dataset contains data from 1979 detectors on 8 highways in San Bernardino, Southern California, from July 1, 2016 to August 31, 2016. These sensors collected data every 5 minutes, covering the number of vehicles passing by each of the 1979 sensors every 5 minutes. (See Table 2 for details.)

[0078] To maintain consistency with baseline methods and ensure fair comparison, the two datasets used in the experiment were divided chronologically into training (60%), validation (20%), and test (20%) sets, and different models were used to predict traffic flow for the next hour.

[0079] Table 2 Dataset Parameters To test the accuracy of this algorithm, this invention uses Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Mean Absolute Percentage Error (MAPE) to measure model performance. These three metrics represent prediction errors, and smaller values ​​are better. Their calculation methods are as follows: (1) Mean Absolute Error: It is defined as the average of the absolute errors between the predicted and the actual values. Its magnitude directly reflects the average error between the predicted and the actual values.

[0080] (2) Root mean square error: It is defined as the square root of the average of the sum of squares of the errors between the predicted and actual values. Its advantage is that it is a relative indicator, which is not affected by the dimensions of the data and can intuitively reflect the relative error between the predicted and actual values.

[0081] (3) Mean Absolute Percentage Error: It is defined as an evaluation index that measures the relative error between the predicted value and the actual value. Its value reflects the average deviation of the predicted value from the actual value, and the unit is percentage.

[0082] See the formula below for details: ; ; ; In the formula, MAE The mean absolute error, RMSE The root mean square error, MAPE The mean absolute percentage error, n For sample size, y i For the first i The true value of each sample For the first i The predicted value for each sample.

[0083] Experimental parameter settings: The ratio of training set to test set was set to 8:2. Training and test sets were divided into training and test sets. Training parameters: On the PEMS04 dataset, the number of training epochs was 50, the batch size was 16, and the initial learning rate was 0.0005. The learning rate was multiplied by 0.92 every 10 epochs. On the PEMS08 dataset, the number of training epochs was 150, the batch size was 64, the initial learning rate was 0.001, and the learning rate was multiplied by 0.9 every 10 epochs.

[0084] GAT parameters: node feature dimension is 8, number of graph convolutional layers is 3, number of multi-head attention mechanisms is 8, and attention coefficient learning rate is 0.0005.

[0085] TCN parameters: kernel size is 3, dilation factor is 2, hidden layer channels are 32, and residual connection layers are 3.

[0086] (2) Parameter sensitivity experiment.

[0087] Using the PEMS04 dataset, we analyzed the fluctuations in model prediction results when different parameter values ​​changed, in order to examine the impact of each parameter on model performance.

[0088] 1) To study the sensitivity of GAT node feature dimensions, the node feature dimensions were set to 6, 8, and 10, respectively. When the feature dimension is 8, a good balance can be achieved between fully representing data features and avoiding excessive redundancy. It can cover the key spatiotemporal information in traffic flow data without making the model too complex.

[0089] 2) To investigate the sensitivity of the number of heads in the GAT multi-head attention mechanism, the multi-head attention mechanism was set to 6, 8, and 10 heads respectively. When the number of heads is 8, a good balance can be achieved between effectively capturing the relationships between nodes and controlling the complexity of the model, enabling the model to better handle the complex spatial relationships in traffic flow data.

[0090] 3) To study the sensitivity of TCN convolution kernel size, the kernel size was set to 1, 3, and 5. When the kernel size is 3, a good balance can be achieved between capturing short-term and medium-term temporal features. It can effectively extract the temporal correlation in traffic flow data without introducing too much noise, thus enabling the model to show good performance.

[0091] 4) To study the sensitivity of the TCN expansion coefficient, the expansion coefficient was set to 1, 2, and 4. An expansion coefficient of 2 strikes a balance between reasonably expanding the receptive field and ensuring effective information transmission, enabling the model to extract long-term and short-term temporal features from traffic flow data more effectively. The results are as follows: Figure 6 As shown, where, where, Figure 6 (a) represents the influence of feature dimensions; Figure 6 (b) represents the effect of head count on the multi-head attention mechanism; Figure 6 (c) represents the effect of kernel size; Figure 6 (d) represents the effect of the expansion coefficient.

[0092] (3) Ablation experiment.

[0093] By removing specific modules, components, or features from the model one by one, and comparing the performance of the original model and the simplified model, the degree of contribution of each part to the overall performance of the model can be determined.

[0094] like Figure 7 , Figure 8 The figures shown are ablation experiment diagrams for PEMS04 and PEMS08, respectively. Figure 7 (a) represents the mean absolute error of each model's prediction in PEMS04; Figure 7 (b) represents the root mean square error of each model's prediction in PEMS04; Figure 7 (c) represents the mean absolute percentage error of each model in PEMS04 prediction; Figure 8 (a) represents the mean absolute error of each model's prediction in PEMS08; Figure 8(b) represents the root mean square error of each model's prediction in PEMS08; Figure 8 (c) represents the mean absolute percentage error of each model in PEMS08 predictions.

[0095] 1) The light blue squares represent the most basic version of the model, lacking all modules. Only the basic Temporal Processing (TCN) module is used. This highlights the importance of the Spatial Processing (GAT) module, the Deep Fusion module, and the Adversarial Training component.

[0096] 2) The dark blue squares represent models that lack the deep fusion module and adversarial training module in the GAN model generator. This highlights the importance of the spatiotemporal fusion module and adversarial training components.

[0097] 3) The light green squares represent the absence of the adversarial training module, highlighting its importance.

[0098] 4) The dark green square represents the complete form of this model.

[0099] Ablation experiments clearly demonstrate that in this traffic flow prediction model, the spatial structure fusion module (based on GAT), the refined spatiotemporal fusion mechanism (including spatiotemporal embedding), and the adversarial training optimization module have the greatest impact on the model's prediction accuracy and overall performance. These modules play crucial roles in utilizing the spatial structure of the traffic network, fusing spatiotemporal features, and optimizing the matching degree between the generated results and real data, respectively. The absence of these important modules will lead to varying degrees of performance degradation in predicting traffic flow. These modules help improve the model's predictive ability at different prediction times.

[0100] This ablation experiment aims to explore the impact of each component of the traffic flow prediction model on performance. By comparing the simplified model with the original model after removing specific modules, we found that: the spatial processing module (GAT) can capture the spatial structure correlation of the traffic network, and its absence will lead to a significant performance drop; the deep fusion module is of great significance for integrating multi-source information and mining feature relationships, and its absence will reduce model performance; the adversarial training module can optimize the matching degree between the generated results and real data, and its absence will reduce prediction accuracy; although the spatiotemporal fusion module was not tested separately, its indispensable role in processing the spatiotemporal variation patterns of traffic flow and improving the prediction ability at different prediction times can be seen from the performance decline of other simplified models.

[0101] (4) Comparison of model prediction performance.

[0102] This invention compares the GAT-TCN model with 12 traffic flow prediction models.

[0103] 1) High Availability (HA): This method collects historical traffic flow data for the same time period and calculates its average value to predict future traffic flow for the same time period. However, because it relies solely on historical average data, it struggles to capture sudden changes and anomalies in traffic flow.

[0104] 2) ARIMA: This model integrates autoregressive, differencing, and moving average elements. Based on the periodicity and correlation of traffic flow, it uses historical data and, after differencing, comprehensively considers various factors to predict future traffic flow. This model requires a large amount of data and is based on the assumption of linear relationships; its accuracy is limited when facing nonlinear and complex situations such as sudden events.

[0105] 3) SVR: Based on the SVM regression algorithm, it uses historical traffic flow data as input features to construct a regression function to predict traffic flow. However, it has high computational complexity and takes a long time to train when processing large-scale data.

[0106] 4) DCRNN: A diffusing convolutional recurrent neural network model that replaces the fully connected layers in the GRU with diffusing convolution, forming a new diffusing convolutional gated recurrent unit. This model integrates diffusing convolution into the architecture of the GRU. Diffusing convolution is a graph-based convolution method that can simulate the diffusion process of information in a graph. In DCRNN, the state updates of nodes are calculated through diffusing convolution, allowing the model to update node states based on the diffusion relationships of nodes in the graph and information in the time series. For example, in traffic flow data, information (such as traffic congestion) diffuses from one node to adjacent nodes, and DCRNN can utilize this diffusion characteristic to better predict the dynamic changes in traffic flow. As a recurrent neural network, DCRNN still retains the ability to model time series. By cyclically processing sequential data and combining the spatial information processing capabilities of diffusing convolution, it can effectively capture long-term dependencies in time series and spatial dependencies in graph structures.

[0107] 5) STGCN: Spatiotemporal Graph Convolutional Network combines spectral domain GCN and one-dimensional convolution to capture spatiotemporal correlations. Spatially, STGCN utilizes spectral domain GCN to process graph-structured data. Spectral domain GCN defines convolution operations based on the eigenvalue decomposition of the graph's Laplacian matrix, effectively handling spatial relationships between nodes. By performing eigenvalue decomposition on the graph's Laplacian matrix and projecting node features onto the spectral domain for convolution, spatial correlation features between nodes can be extracted. Temporally, one-dimensional convolution is used to process time-series data. One-dimensional convolution automatically extracts local features from time series data and captures long-term dependencies by stacking convolutional layers. By combining the spatial features processed by spectral domain GCN with the temporal features processed by one-dimensional convolution, STGCN can simultaneously consider the temporal and spatial correlations in spatiotemporal data.

[0108] 6) VAR: Based on the interrelationships of multiple time series variables, it uses historical data of each variable to form a vector sequence to predict future values.

[0109] 7) Graph WaveNet: Combining GCN, it proposes an adaptive adjacency matrix to learn the dynamic correlations between nodes. This model constructs the adjacency matrix through a learnable process. Specifically, it utilizes node feature information and calculates dynamic connection weights between nodes through a series of neural network layers (such as multilayer perceptrons) to form an adaptive adjacency matrix. This approach can dynamically adjust the connection relationships between nodes based on the features in the data, unlike traditional methods that rely on a predefined fixed adjacency matrix. For processing time series data, Graph WaveNet adopts an architecture similar to WaveNet. This model processes time series data by stacking multiple convolutional layers with gating mechanisms, effectively capturing long-term dependencies in the time series. In Graph WaveNet, these time series processing layers are combined with the adaptive adjacency matrix, enabling the model to simultaneously consider the spatial dynamic correlations between nodes and the changes in the time series.

[0110] 8) ResNet-CNN: Combining CNN and ResNet structures, CNN extracts spatial features of traffic flow data, and ResNet residual structure solves the training problem of deep networks, but it has shortcomings such as complex structure and long training time.

[0111] 9) ASTGCN: It can consider information in both time and space dimensions simultaneously. It models spatial correlations through graph convolutional neural networks (GCN), treating the traffic network as a graph structure, where nodes represent road monitoring points and edges represent the connections between roads, effectively capturing the spatial dependencies between different monitoring points.

[0112] 10) STSGCN: Spatiotemporal Synchronous Graph Convolutional Network Model. This model constructs multiple local spatiotemporal graphs to simultaneously capture spatiotemporal dependencies. It divides the entire spatiotemporal data into multiple local regions and constructs a spatiotemporal graph within each local region. These spatiotemporal graphs can simultaneously consider the relationships between spatially adjacent nodes and the temporal sequence relationships. For example, in traffic flow data, a local spatiotemporal graph can represent the traffic conditions of a block over a period of time, where nodes represent intersections, edges represent connections between intersections, and traffic changes over time. Convolutional operations are performed on the constructed local spatiotemporal graphs. Graph convolution is used to handle spatial relationships, and temporal convolution is used to handle time-series relationships. These convolutional operations can effectively extract features from the local spatiotemporal graphs, and because they are performed on multiple local spatiotemporal graphs, they can comprehensively consider the spatiotemporal features of different local regions, avoiding information loss or over-smoothing that may occur when performing a single convolutional operation on the entire spatiotemporal data.

[0113] 11) CNN-LSTM: Combining CNN and LSTM structures, CNN extracts spatial features of traffic flow, while LSTM processes temporal characteristics. The combination of the two can accurately capture road associations and temporal dependencies.

[0114] 12) STFGNN: A neural network architecture for spatiotemporal data prediction, primarily used to process data with spatiotemporal characteristics such as traffic flow. Its core idea is to effectively fuse spatial and temporal information to improve prediction accuracy. This architecture consists of multiple modules, including a spatial feature extraction module, a temporal feature extraction module, and a fusion module.

[0115] 13) DDSTCRN: It decouples two hidden time series signals through gating and residual decomposition mechanisms, captures the spatiotemporal correlation of traffic data using dynamic cyclic graph convolution, and then captures the dynamically changing spatial dependencies through a dynamic graph builder.

[0116] Representative traditional models and traditional spatiotemporal models were selected from Table 3 for analysis. On the PEMS04 and PEMS08 datasets, the average MAE of this model was reduced by 48.61%, 41.17%, and 43.56% for HA, 52.61%, 41.79%, and 50.57% for ARIMA and SVR, respectively, and the average RMSE was reduced by 43.54%, 54.54%, and 41.78% for HA, 44.24%, 43.29%, and 43.72%, respectively. This demonstrates that traditional models are insufficient in capturing complex spatiotemporal features, cannot consider the spatial relationships between traffic network nodes, and cannot effectively mine long-term dependencies in time series data.

[0117] Table 3 Performance comparison of GAT-TCN and different models on PEMS04 and PEMS08 datasets. On the PEMS04 and PEMS08 datasets, this model reduced the average MAE by 35.26% and 11.81% and the average RMSE by 31.92% and 22.58% respectively compared to traditional spatiotemporal models DCRNN and STGCN, and reduced the average RMSE by 25.31% and 13.76% and the average RMSE by 17.54% and 10.48% respectively. This demonstrates that traditional spatiotemporal models are insufficient in spatial feature extraction, fail to recognize the differences in importance between different nodes, and are only suitable for short-term prediction tasks, making them unsuitable for processing long-term series data.

[0118] On the PEMS04 and PEMS08 datasets, this model reduces MAE, RMSE, and MAPE by 0.11%, 1.45%, and 2.48% and 6.48%, 4.96%, and 5.75% respectively compared to the best-performing model DDSTCRN. This demonstrates the superiority and rationality of the GAT-TCN model's network architecture, indicating that it can effectively capture long-term and short-term dependencies in time series data and model traffic flow data with complex temporal dynamics. The GAT-TCN model achieves good performance on different datasets (PEMS03 and PEMS07), demonstrating its good generalization ability and adaptability to traffic flow prediction tasks in various scenarios.

[0119] (5) Visual display of results.

[0120] exist Figure 9 , Figure 10 In the figure, the vertical axis represents traffic flow data per unit time, with a sampling interval of 5 minutes. The figure visualizes the comparison between the predicted and actual values ​​of the GAT-TCN model for a portion of the time period at node 96 on the PeMS04 and PeMS08 datasets. The blue line represents the actual traffic flow data, showing how the actual traffic flow changes over time. The orange line represents the traffic flow data predicted by the GAT-TCN model, reflecting the model's prediction of traffic flow changes over time. The figure shows a high degree of fit between the model's predictions and the actual values, indicating that the model accurately captures the characteristics of traffic flow data and demonstrating the effectiveness of the model construction.

[0121] In summary, the present invention has the following advantages: (1) To address the issues of insufficient data and lack of diversity in the original traffic flow dataset under specific scenarios, this project employs GAN data augmentation. Random noise is introduced to generate new traffic flow data samples, while spatiotemporal correlation loss, road attribute matching loss, and special scenario feature loss are added. This ensures that the generated data possesses features consistent with actual road attributes and includes characteristics specific to the scenario, effectively expanding the data's diversity and feature dimensions, and providing more diverse data for model training. Finally, predictions are performed on the test set of this dataset. Experimental results show that the average MAE of the model using the augmented dataset during training decreased by 3.5% compared to the unaugmented dataset. This enhances the model's robustness, enabling it to learn more useful features and better cope with complex traffic flow conditions.

[0122] (2) Traffic flow data has complex spatiotemporal characteristics, exhibiting periodicity, trends, and sudden changes in the time dimension. In the spatial dimension, traffic nodes in different geographical locations influence each other, forming a complex network structure. Traditional traffic flow prediction models struggle to deeply explore the complex dependencies between time and space dimensions. To address these issues, a traffic flow fusion prediction model based on graph neural networks and deep generative models (GAT-TCN) is proposed. In terms of temporal features, TCN extracts temporal features using causal convolution and dilated convolution, which can expand the receptive field without adding too many parameters, thereby capturing the long-term and short-term dependencies in the time series. In terms of spatial features, GAT constructs a graph structure to capture the spatial characteristics of traffic flow. It learns the importance between nodes in the graph through an attention mechanism. For each node in the graph, GAT calculates the attention coefficient between it and its neighboring nodes, and then weights and sums the features of the neighboring nodes based on these coefficients to obtain the updated feature representation of the node. This method can adaptively capture the spatial dependencies between nodes. The data at each time step is passed through a TCN layer to extract temporal features, and then these features are fed into the GAT network for graph structure modeling. Finally, a generative adversarial network (GAT-TCN) is introduced to further optimize the model. Through continuous adversarial training, the generator can better integrate temporal and spatial features, improving the model's prediction accuracy and achieving effective fusion of temporal and spatial features. Experimental results show that GAT-TCN outperforms other models in prediction.

[0123] The present invention also provides a traffic flow prediction system, comprising: The data acquisition module is used to acquire traffic flow data and traffic network adjacency matrix at different time scales; The model building module is used to combine the Temporal Convolutional Network (TCN) and the Graph Attention Network (GAT) to obtain the traffic flow prediction model GAT-TCN. The feature extraction module is used to input traffic flow data at different time scales into GAT-TCN, use TCN to mine the local dynamic changes of traffic flow data in the time series, and extract time features; input the time features and traffic network adjacency matrix into GAT, use GAT to calculate the attention coefficient between each node in the traffic network adjacency matrix and its neighboring nodes, and aggregate the neighboring nodes according to the attention coefficient to obtain the spatial features of the traffic network. The traffic flow prediction module is used to fuse temporal and spatial features and predict traffic flow based on the fused data.

[0124] The present invention also provides a computer device including a memory and a processor; the memory stores a computer program, and the processor runs the computer program in the memory to perform a traffic flow prediction method.

[0125] The present invention also provides a computer-readable storage medium storing a computer program adapted for loading by a processor to execute a traffic flow prediction method.

[0126] The above-described embodiments are merely preferred embodiments of the present invention, and the scope of protection of the present invention is not limited thereto. Any simple changes or equivalent substitutions of the technical solutions that can be obviously obtained by those skilled in the art within the scope of the technology disclosed in the present invention shall fall within the scope of protection of the present invention.

Claims

1. A traffic flow prediction method, characterized in that, Includes the following steps: Obtain traffic flow data and traffic network adjacency matrix at different time scales; By combining a Temporal Convolutional Network (TCN) and a Graph Attention Network (GAT), a traffic flow prediction model, GAT-TCN, is obtained. A Generative Adversarial Network (GAN) is used to train GAT-TCN. The loss function used for training includes auxiliary loss terms related to specific traffic flow attributes. These auxiliary loss terms include a spatiotemporal correlation loss term, a road attribute matching loss term, and a special scene feature loss term. The spatiotemporal correlation loss term measures the correlation between generated data and real data in the time dimension. The road attribute matching loss term calculates the error between the generated traffic flow data and the real traffic flow features under different road attributes. The special scene feature loss term calculates the difference between the generated traffic flow data in special scenarios and the real special scene features. The spatiotemporal correlation loss term is determined by the following formula: ; In the formula, The mean absolute error of spatial characteristics, The mean square error represents the time characteristic. for The weight parameters, for Weight parameters; The road attribute matching loss term is determined by the following formula: ; In the formula, The weighting coefficients for speed features in the calculation of road attribute matching loss are: The mean square error of the velocity characteristic. These are the weighting coefficients for traffic flow features in the calculation of road attribute matching loss. The mean square error of the flow characteristics. The density feature is the weighting coefficient in the calculation of road attribute matching loss. The mean square error of the density feature; The feature loss term for special scenarios is determined by the following formula: ; In the formula, For special scene feature loss, To highlight the differences in characteristics under extreme weather scenarios, To differentiate features in high-traffic fluctuation scenarios. To highlight the unique characteristics of high-traffic scenarios during major events, To highlight the unique characteristics of specific flow scenarios during major events, , ν , ξ , η These are the weight coefficients for the feature differences corresponding to specific scenarios; Traffic flow data at different time scales are input into GAT-TCN. TCN is used to mine the local dynamic changes of traffic flow data in the time series and extract time features. The time features and traffic network adjacency matrix are input into GAT. GAT is used to calculate the attention coefficient between each node in the traffic network adjacency matrix and its neighboring nodes. The neighboring nodes are aggregated based on the attention coefficient to obtain the spatial features of the traffic network. The GAT time attention mechanism adjusts the weights of each time step in the time features based on the spatial features, and then fuses the spatial and temporal features according to the weights. The GAT residual connection fuses the traffic flow data with the obtained fused features again to generate prediction results.

2. The traffic flow prediction method according to claim 1, characterized in that, The method of using TCN to mine local dynamic changes in traffic flow data over time and extract time features involves: using TCN to transform the channel dimension of the input data, using temporal convolution to slide the transformed data through a convolution kernel with specific parameters over time to mine local dynamic changes in the time series, and outputting time features through temporal convolution.

3. The traffic flow prediction method according to claim 1, characterized in that, The traffic flow data includes weekly traffic flow data, daily traffic flow data, and recent traffic flow data.

4. A traffic flow prediction system, characterized in that, include: The data acquisition module is used to acquire traffic flow data and traffic network adjacency matrix at different time scales; The model building module combines the Temporal Convolutional Network (TCN) and the Graph Attention Network (GAT) to obtain the traffic flow prediction model GAT-TCN. A Generative Adversarial Network (GAN) is used to train GAT-TCN. The loss function used for training includes auxiliary loss terms related to specific traffic flow attributes. These auxiliary loss terms include a spatiotemporal correlation loss term, a road attribute matching loss term, and a special scene feature loss term. The spatiotemporal correlation loss term measures the correlation between generated data and real data in the time dimension. The road attribute matching loss term calculates the error between the generated traffic flow data and the real traffic flow features under different road attributes. The special scene feature loss term calculates the difference between the generated special scene traffic flow data and the real special scene features. The spatiotemporal correlation loss term is determined by the following formula: ; In the formula, The mean absolute error of spatial characteristics, The mean square error represents the time characteristic. for The weight parameters, for Weight parameters; The road attribute matching loss term is determined by the following formula: ; In the formula, The weighting coefficients for speed features in the calculation of road attribute matching loss are: The mean square error of the velocity characteristic. These are the weighting coefficients for traffic flow features in the calculation of road attribute matching loss. The mean square error of the flow characteristics. The density feature is the weighting coefficient in the calculation of road attribute matching loss. The mean square error of the density feature; The feature loss term for special scenarios is determined by the following formula: ; In the formula, For special scene feature loss, To highlight the differences in characteristics under extreme weather scenarios, To differentiate features in high-traffic fluctuation scenarios. To highlight the unique characteristics of high-traffic scenarios during major events, To highlight the unique characteristics of specific flow scenarios during major events, , ν , ξ , η These are the weight coefficients for the feature differences corresponding to specific scenarios; The feature extraction module is used to input traffic flow data at different time scales into GAT-TCN, use TCN to mine the local dynamic changes of traffic flow data in the time series, and extract time features; input the time features and traffic network adjacency matrix into GAT, use GAT to calculate the attention coefficient between each node in the traffic network adjacency matrix and its neighboring nodes, and aggregate the neighboring nodes according to the attention coefficient to obtain the spatial features of the traffic network. The traffic flow prediction module is used to adjust the weights of each time step in the temporal features based on the spatial features using the time attention mechanism of GAT. The spatial and temporal features are then fused according to the weights. The residual connection of GAT further fuses the traffic flow data with the obtained fused features to generate the prediction result.

5. A computer device, characterized in that, It includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to perform the traffic flow prediction method according to any one of claims 1-3.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to execute the traffic flow prediction method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Traffic flow prediction method and device based on synthetic data

    CN113570861A

  • Traffic flow prediction method based on fusion of space-time adaptive graph learning and dynamic graph convolution

    CN117392846A