A traffic flow prediction method for realizing co-mining of traffic point semantics and network topology

By combining the STIGNN method of ChebNet and GAT models, the problem that existing traffic flow prediction methods are difficult to mine traffic network topology and semantic information is solved, and a higher precision traffic flow prediction is achieved.

CN119443156BActive Publication Date: 2025-06-17CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510030960.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-17
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Existing traffic flow prediction methods are difficult to fully explore the topological information and node semantic information of the traffic network, resulting in limited prediction accuracy.

Method used

A STIGNN method combining ChebNet model and GAT model is proposed. The spatial correlation and node semantic information of the traffic network are simultaneously mined through the ChebGAT module, and the timing characteristics are captured through the GateTCN module to realize traffic flow prediction.

Benefits of technology

By combining spatial and temporal information, STIGNN can predict traffic flow more accurately, significantly improving prediction accuracy, especially in complex traffic networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119443156B_ABST
    Figure CN119443156B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of transportation technologies, and particularly relates to a traffic flow prediction model that realizes co-mining of traffic point semantics and network topology. The name of the model is STIGNN. In terms of spatial feature aggregation, STIGNN focuses on the coupling of traffic point features and the topological relationship between traffic points, comprehensively considering potential and existing spatial dependencies. STIGNN can process long time series through gated temporal convolutional layers. By introducing a dual attention mechanism for node features and topological relationships in the spatio-temporal framework, this model can be directly applied to inductive learning tasks and can be generalized to any network with complex entities and relationships. Experimental results on two public real traffic network datasets, METR-LA and PEMS-BAY, show that STIGNN outperforms advanced baseline models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of transportation, and particularly relates to a traffic flow prediction method for realizing co-mining of traffic point semantics and network topology. Background Art

[0002] Traffic flow prediction is a key task in urban traffic control and scheduling in the field of transportation, with high non-linearity and complexity. The traffic system is an important infrastructure of modern cities, supporting the daily communication and travel of millions of people. With the rapid development of urbanization and population, the traffic system has become increasingly complex, and safety problems such as frequent traffic jams and traffic accidents have emerged. The intelligent transportation system has emerged and developed rapidly, aiming to achieve the efficient coordination of elements such as people, vehicles, and roads by making full use of cutting-edge technologies such as artificial intelligence, the Internet of Things, and mobile communication. Traffic flow prediction is a core component of the intelligent transportation system, and its task is to predict the future traffic flow speed using the given historical speed data and the basic road network, effectively avoiding traffic congestion and improving the operation efficiency of the traffic system. However, traffic flow prediction is a complex spatio-temporal prediction problem. The traffic state in a specific area is not only affected by the distribution of surrounding traffic routes, but also changes dynamically according to different time periods. The dependencies in space and time pose great challenges to traffic flow prediction.

[0003] Researchers have proposed many algorithmic models to solve the traffic flow prediction problem. Traditional statistical models such as the historical average (HA) model, the auto-regressive integrated moving average (ARIMA) model, the Kalman filter model, etc. are limited by the dependence on the stationarity assumption and cannot capture the complex dynamic relationships between traffic data, so their prediction ability is very limited. Traditional machine learning algorithms such as support vector regression (SVR), linear regression (LR), etc. have been further improved on the basis of traditional statistical models through the mining of historical traffic data, but they still cannot break through the complexity of the model in space and time. Inspired by the great success of deep learning techniques in the fields of computer vision and natural language processing, researchers introduced the recurrent neural network (RNN) model that can mine time dependencies and the convolutional neural network (CNN) that can capture spatial dependencies into traffic prediction, and unprecedented success has been achieved in the traffic flow prediction task. During this period, excellent traffic flow prediction models such as LSTMNN and GBM emerged. However, CNN and RNN can only be used for Euclidean data (i.e., images, texts, and videos), and the spatial dependence of the road network is represented by a non-Euclidean graph structure. Serializing the graph structure will inevitably result in information loss, which is not conducive to the mining of spatial topological information.

[0004] In recent years, graph neural networks (GNNs) have become the forefront of deep learning research, showing state-of-the-art performance in various applications. It can capture the complex relationships between objects and make inferences based on data described by graphs. GNNs have been proven to be effective in various node-level, edge-level, and graph-level prediction tasks. In fact, GNNs are very suitable for traffic prediction problems. The road network is essentially a graph, the sensors on the roads are nodes, and the connections between sensors are edges. Taking the graph as the input, benefiting from the complement of traffic network topology information, the GNN-based model will show superior performance in the road traffic flow prediction task compared with previous methods. The STGCN model first applied the GNN network to traffic network prediction. By using the Graph Convolutional Network (GCN) to mine traffic network topology information and using the Gated temporal convolutional layer (Gated TCN) module based on the gating mechanism to replace the RNN model to capture the time period characteristics, this model achieved excellent prediction results in the METR-LA and PEMS-BAY datasets. Subsequently, researchers focused on the migration of graph neural network algorithms to traffic prediction, and a large number of traffic prediction models based on the combination of GNN and RNN emerged. Such as the DCRNN model using Diffusion convolutional neural networks (DCNNs) and Gated recurrent unit (GRU), the STGAT model using Graph attation neural network (GAT) and Gated TCN module, the GC-LSTM model using GCN and Long short-term memory (LSTM), etc. It can be clearly observed that the current traffic flow prediction models based on spatio-temporal graph structures focus on the optimization of graph neural network and recurrent neural network modules, and different combination methods can always bring surprises to traffic flow prediction. Of course, many strategies have been introduced during this period, such as the learnable adjacency matrix introduced in the GraphWaveNet model. The road structure information mined by researchers according to the map is restricted by the manual processing granularity and still ignores some key network information. By allowing the network to adaptively learn graph structure information and make joint predictions with the known graph structure, the problem of missing graph structure can be effectively solved. Graph WaveNet is also based on the GCN and Gated TCN modules, but it achieves higher prediction accuracy than DCRNN. The learnable matrix strategy also appears in the STGAT model, and the experiment once again proves the feasibility of this strategy.

[0005] By considering the underlying graph structure, significant progress has been made in GNNs. However, there are still some challenges; for example, although the left and right lanes as nodes are very close, their relevance is not necessarily stronger than that of lanes in the same direction at a greater distance. Another example is that traffic sections with similar infrastructure, even if they are far apart, can still be of reference value. There is a hope to further couple the information of the road itself with the topological information of the road network, but there are still certain gaps in the joint mining of topological information and node semantic information by graph neural networks. Whether it is the GCN model, GAT model, DCNN model, etc., they are all node-level information aggregation methods. The target node can only obtain the second-order neighbor information through the first-order neighbors, and cannot fully mine the network topological information. In the SGC model and ChebNet model in the model, although they can simultaneously mine multi-order topological information, their mining of the semantic associations between nodes is not in place, and they only formulate the graph information aggregation strategy according to the graph structure. And only a small number of models have learnable information aggregation rules, and the model cannot select by itself which information is important and which is not. In the models with learnable aggregation rules, GAT adjusts the information aggregation method between nodes by calculating the dynamic attention between nodes, but it is only limited to the first-order graph structure, and the dependence mining of information is more based on the node's own semantic information, and the relationship between different-order neighbors in the network cannot be well mined. ChebNet applies static attention to the node sets of different-order neighbors and learns to judge which set of neighbors of each node is more important, but the learning parameters are set for the network topology, and the node semantic associations cannot be fully extracted. The fusion of the two models may help to find a traffic flow prediction scheme that can simultaneously mine the topological information of the traffic network and the semantic information of traffic nodes.

[0006] In order to better mine the spatio-temporal relationship of the traffic network and improve the accuracy of traffic flow prediction. The present invention proposes a spatio-temporal graph neural network that combines traffic node semantic information and network topological information. Summary of the Invention

[0007] The present invention proposes a traffic flow prediction method that realizes the co-mining of traffic point semantics and network topology. This model solves the problem in the existing methods that the importance of the coupling of traffic point features and traffic network information for traffic flow prediction is ignored, and realizes an effective traffic flow prediction task.

[0008] The technical solution of the present invention is realized as follows:

[0009] A traffic flow prediction method for realizing co - mining of traffic point semantics and network topology. The name of this method is STIGNN. STIGNN includes two ChebGAT modules with parallel computing and three GateTCN modules connected in series. Among them, the ChebGAT module couples ChebNet static attention and GAT dynamic attention to fully mine the spatial correlation of the traffic network and realize the simultaneous mining of topological features and semantic information.

[0010] The GateTCN module integrates a time - series model with a gating mechanism. After mining the temporal correlation of road information detected at different times, it can obtain road features that fuse temporal information. Then, the road features that fuse time series and the sensor network are input into the ChebGAT module to deeply mine the spatio - temporal features of the road network for final traffic flow prediction.

[0011] Optionally, the process of the ChebNet model for aggregating features is as follows:

[0012]

[0013] Among them, σ(·) represents a non - linear activation function, and θ i is the learnable static weight of the i - th order neighbor. When i = 0, it represents the importance degree of its own information. d i represents the degree of node i, W is a learnable weight matrix, and x i represents the traffic point information of the first - order neighbor of node i, x j represents the traffic point information of the second - order neighbor of node j, x k represents the traffic point information of the third - order neighbor of node k, d j represents the degree of node j, and d k represents the degree of node k.

[0014] Optionally, the process of GAT for aggregating features is as follows:

[0015]

[0016] Among them, α ij represents the semantic association degree between node i and j;

[0017]

[0018] Among them, || represents the concatenation operation, is a learnable weight vector, and N i represents the set of first - order neighbors of node i.

[0019] Optionally, the neighborhood perception range of the ChebGAT module is fixed to 3 - order, and the process of its aggregating features is as follows:

[0020]

[0021] Optionally, the STIGNN transposed sensor feature matrix X ∈ R M×N×F → X ∈ R N×F×M , and at the same time, set the dilated causal convolution kernel f ∈ R K , where K represents the size of the convolution kernel, and the process of dilated causal convolution can be expressed as:

[0022]

[0023] where d is the dilation factor of the dilated causal convolution, m represents the time series length of the intercepted traffic data, s is the value traversing K, and X ★ m f represents the convolution of X using the dilated causal convolution kernel f, and f(s) represents the s-th parameter of the convolution kernel f.

[0024] Optionally, the data feature processing process of the GateTCN module is as follows:

[0025]

[0026] where z l represents the update gate of the l-th layer of GateTCN, which is used to control the information ratio of the longer time series and the shorter time series, and r l represents the reset gate, C l is the obtained long time series information, f z , f r and f c are both learnable convolution kernels, b z , b r and b c are all learnable feature biases, represents the dot product.

[0027] Optionally, when the ChebGAT module captures spatial features, an adaptively learned traffic graph G ′ =(V, E ′ ) is introduced. In G ′ , the edges and edge weights are learned end-to-end through the stochastic gradient descent algorithm. Denote the adjacency matrix of G ′ as A adp , and use the randomly initialized matrices E1, E2 ∈ R N×c to learn A adp :

[0028] A adp = Softmax(Relu(E1E2 T ))#

[0029] where E1 is named the source node embedding and E2 is named the target node embedding.

[0030] After adopting the above technical solution, the beneficial effects of the present invention are as follows:

[0031] The traffic flow prediction method in the present invention innovatively combines the key elements of the ChebNet model and the GAT model, enabling the graph neural network to simultaneously calculate the feature correlations between traffic points and the dependencies between adjacent traffic points of different orders. The traffic flow prediction method in the present invention can mine the spatial information of the traffic network more fully. At the same time, the addition of the Gated TCN module enables the model to have the ability to mine the temporal characteristics between traffic flows. The coupling of node semantic information and the traffic network topology structure allows the traffic point to be predicted to fully refer to the states of nearby traffic points and the traffic network structure, making the traffic flow prediction method in the present invention mine the spatial information of the traffic network more fully. In view of the dynamic temporality of the traffic flow network, the present invention designs a Gated TCN module, which learns to control the traffic information dependencies at different time periods by adding a gating function, enabling the present invention to have the ability to predict the traffic flow changes over time. Through the bidirectional mining and combination of spatial and temporal information, the present invention can achieve more accurate traffic flow prediction compared with existing methods. Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 is the process of aggregating neighbor information by the GAT model and the ChebNet module;

[0034] Figure 2 is the structure of the RNN model;

[0035] Figure 3 is the dilated causal convolution with a convolution size of 2;

[0036] Figure 4 is the structural diagram of the GateTCN module;

[0037] Figure 5 is the overall architecture diagram of the STIGNN. Detailed Embodiments

[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] The embodiment of the present application discloses a traffic flow prediction method that realizes co-mining of traffic point semantics and network topology.

[0040] Example

[0041] according to Figures 1 to 5 As shown, a traffic flow prediction method that realizes co-mining of traffic point semantics and network topology is presented.

[0042] 1. Related Work

[0043] 1.1 Problem Modeling

[0044] The goal of traffic prediction is to predict future traffic flow based on the traffic flow observed by N sensors preset on the road network. Sensors are regarded as nodes, and sensors with close distance or upstream and downstream relationships are regarded as edges. A traffic network can be defined as an undirected graph G = (V, E), where V represents the node set and E represents the edge set. Let A∈R N×N represents the adjacency matrix of graph G, then we have a ij ∈A>0 indicates that there is a connection between sensors i and j, otherwise a ij = 0. Using the same sampling frequency, the time step of each interval is t, and the traffic data is sampled M times. In each step, the graph G has a sensor information matrix. The sensor information matrices recorded each time are concatenated and recorded as the feature matrix X∈R M×N×F , where F represents the dimension of the information recorded by the sensor, including traffic flow, vehicle speed, road occupancy, etc. Let X m represents the information recorded by the sensor in the traffic data collected for the mth time, Y m represents the traffic flow value in the traffic data collected for the mth time. The traffic flow prediction problem aims to learn a function h(·), which combines the graph information and the sensor status and predicted traffic flow at the historical moment to infer the traffic flow at the next moment. The process is as follows:

[0045]

[0046] In fact, traffic flow prediction can be divided into several types according to the different time steps t, among which t<30min is short-term prediction, 30min≤t<60min is medium-term prediction, and t≥60min is long-term prediction.

[0047] 1.2 Graph Neural Network

[0048] To mine the topological information of graphs, Scarselli et al. proposed the concept of graph neural networks. Based on Banach's fixed-point theory, they combined recurrent neural networks with random walk models to achieve the fusion of node features and topological information. Based on this, models such as GraphSAGE, DCNNs, and GAT have emerged in the field of graph neural networks. These models intuitively set the dependence of target nodes on the features of their neighbor nodes in the problem space, so they are called spatial domain models. Among them, the GAT model first introduced the attention mechanism into graph neural networks. By calculating the dependence between nodes through the features of adjacent nodes, it fully incorporates node semantic information into the aggregation of network information. Its aggregation process is as shown in Figure 1 shown on the left. The black lines represent edges, and the arrows represent the direction of information aggregation. α ij represents the attention score between nodes i and j, and the calculation process is as follows:

[0049]

[0050] where, || represents the concatenation operation, is the learnable weight vector, W is the learnable weight matrix, and N i represents the first-order neighbor set of node i. It can be clearly seen that node x0 can only control the information aggregation within the first-order neighbors, and at the same time, the information aggregation weight is controlled by node semantic information.

[0051] Convolutional neural networks have achieved great success in the field of images. However, due to the irregularity of graph structures, convolutional neural networks cannot be migrated to the field of graphs. Bruna et al. combined graph signal processing and proposed the first-generation graph convolutional neural network, Spectral CNN. By orthogonally diagonalizing the graph Laplacian matrix, the Fourier basis and frequencies of the graph can be obtained and a complete graph Fourier domain (also known as the graph spectral domain) can be formed. Mapping the graph signal from the spatial domain to the spectral domain, the mapping of the graph signal in the graph Fourier basis is serialized data, which can be effectively captured by convolution. Based on this, models such as ChebNet, GCN, and GraphHeat have emerged in the field of graph neural networks. These models all achieve graph tasks by adjusting graph frequencies in the frequency domain and amplifying effective graph signals. Among them, ChebNet first proposed that the convolutional kernel can be fitted by a polynomial of frequency, reducing the resource consumption of convolutional kernel learning. There is a direct association between the frequency polynomial function in the frequency domain and the polynomial function of the adjacency matrix in the spatial domain, which means that while adjusting the order of the frequency, it is also adjusting the weights of different-order neighbors. Its aggregation process is as shown in Figure 1 shown on the right. The black lines represent edges, and the arrows represent the direction of information aggregation. θ iDenote the weight of the $i$-th neighbor node, which is a learnable static parameter.

[0052] The GAT model focuses on the semantics of nodes and adjusts the contribution of first-order neighbors to the target node through semantic information between nodes. ChebNet pays attention to the order of neighbor nodes and sets different learnable parameters for neighbors of different orders. Both models have their own focuses and have achieved excellent results.

[0053] 1.3 Temporal Networks

[0054] The architecture of the RNN model is as Figure 2 shown. It stores the hidden features of traffic information at historical moments through its unique memory unit and uses them for prediction at the next moment. RNN has outstanding performance in dealing with time series modeling problems. To further optimize the RNN algorithm, Graves et al. proposed LSTM. LSTM adds gate structures on the basis of RNN. The forget gate determines which traffic information to discard, the input gate determines which newly input information to pass to the next step to update the information stored in the old memory unit, and the output gate predicts the traffic flow at that moment. Chung et al. further proposed the GRU model on the basis of LSTM. GRU only uses the update gate and the reset gate, is simpler than LSTM, and is more time-saving to train.

[0055] However, the extended models based on RNN still face the problems of requiring a large amount of resources for training and being unable to handle longer sequence predictions. Wu et al. adopted dilated causal convolution as the temporal convolution layer (TCN) to capture the temporal characteristics of nodes. Dilated causal convolution obtains an exponentially growing receptive field by increasing the depth of the network. Different from the RNN-based method, dilated causal convolution can correctly process long time series in a non-recursive manner, which helps with parallel computing and alleviates the problem of gradient explosion. In the present invention, zeros are padded in the dilated causal convolution to maintain the causal order of time, so the prediction of the current time step only involves historical information. The convolution process of dilated causal convolution is as Figure 3 shown.

[0056] 2. STIGNN

[0057] 2.1 Spatial Correlation Mining

[0058] Transportation networks are usually represented by graph structures. To fully exploit the spatial correlation of transportation networks, the present invention simultaneously considers the distances between transportation points and the differences in road conditions. Among them, topology refers to the transportation network structure, and semantics refers to the information of transportation points. The present invention simultaneously mines both the transportation network structure and the information of transportation points. Graph neural networks have spawned many algorithms to mine node information or topological structure information in graph data, but there are still deficiencies in the simultaneous mining of topological features and semantic information. To solve this problem, the present invention proposes a ChebGAT module that couples ChebNet static attention and GAT dynamic attention to fully exploit the spatial correlation of transportation networks. The process of ChebNet aggregating features is as follows:

[0059]

[0060] Among them, σ(·) represents a non-linear activation function, and θ i is the learnable static weight of the i-th order neighbor. When i = 0, it represents the importance of its own information. d i represents the degree of node i, which is used for diagonal normalization to ensure that the node features of different order neighbors are of the same scale. W is a learnable weight matrix, which, combined with the non-linear activation function σ(·), performs a non-linear transformation on the nodes after fusing the network topological information to further extract features. x i represents the transportation point information of the first-order neighbor of node i, x j represents the transportation point information of the second-order neighbor of node j, x k represents the transportation point information of the third-order neighbor of node k, d j represents the degree of node j, and d k represents the degree of node k. The process of the GAT model aggregating features is as follows:

[0061]

[0062] Among them, α ij is calculated from Equation (2) and represents the semantic association degree between nodes i and j. Combining Equation (3) and Equation (4) gives the ChebGAT module. Considering that the sensors are too far away after the third-order neighbor and cannot form a strong correlation, the present invention fixes the neighborhood perception range of the ChebGAT module to 3 orders, and its process of aggregating features is as follows:

[0063]

[0064] When the distance between two traffic sensors is very close but the monitoring data is quite different, or the detected data is similar but the distance is very far, the two sensors cannot provide much effective information for each other. The coupling of node semantic information and the traffic network topology structure enables the model to select the most suitable nodes for reference and fully capture the effective information in the traffic network.

[0065] The combination of formula (2) and formula (4) will increase a certain time complexity. In the traditional GAT model, the feature mapping of all vertices is first calculated as \(\vec{F}\to F\) ′ , and then the attention coefficients are calculated. In the subsequent feature aggregation process, mainly weighted summation operations are involved, and no high-complexity multiplication operations are involved anymore. Therefore, the time complexity of GAT is \(O(|V|\times F\times F\) ′ ) + \(O(|F|\times F\) ′ ). While the ChebGAT module needs to calculate the attention scores of the target node with its first-order neighbors, second-order neighbors, and third-order neighbors. In the most complex case, the attention coefficients need to be calculated between all nodes. Therefore, the time complexity of the ChebGAT module is \(O(|V|\times F\times F\) ′ ) + \(O(|V|\times|V|\times F\) ′ ), and the complexity at this time mainly depends on the number of road sensors and the dimension of node features. However, benefiting from the fact that there are not many main road sections in a region, the time complexity brought by the number of sensors does not significantly affect the training time of the model. By comparing the prediction results of the ChebGAT module and the GAT model on the datasets METR-LA and PeMS-BAY, the present invention confirms the powerful ability of the ChebGAT module in mining spatial correlation.

[0066] 2.2 Temporal Correlation Mining

[0067] Wu et al. have demonstrated the powerful performance of dilated causal convolution in mining temporal correlation. The present invention decides to further develop a temporal feature mining model based on this model. To facilitate the study of the temporal correlation of the traffic network, the present invention transposes the sensor feature matrix \(X\in R\) M×N×F \(\to X\in R\) N×F×M , and at the same time sets the dilated causal convolution kernel \(f\in R\) K , where \(K\) represents the size of the convolution kernel. The process of dilated causal convolution can be expressed as:

[0068] (6)

[0069] where \(d\) is the dilation factor of the dilated causal convolution, \(m\) represents the time series length of the intercepted traffic data, \(s\) is the value traversing \(K\), and \(X^*\) mf represents convolving X using the dilated causal convolution kernel f, and f(s) represents the s-th parameter of the convolution kernel f. The dilation factor increases with the stacking order of the dilated causal convolution layers, which enables the extended causal convolution network to capture longer sequences with fewer layers, thus saving computational resources.

[0070] In RNNs, the gating mechanism is crucial. Chung et al. have experimentally demonstrated the powerful ability of the gating mechanism in controlling the flow of temporal information. The present invention aims to introduce an update gate and a reset gate in the dilated causal convolution layer, enabling the model to further establish long-term dependencies of temporal correlations in an exponentially growing temporal receptive field. The present invention names the time series model incorporating the gating mechanism as the gated temporal convolution layer, denoted as the GateTCN module. The structure of the GateTCN module is as Figure 4 shown, and its data feature processing process is as follows:

[0071] (7)

[0072] where z l represents the update gate of the l-th layer of the GateTCN, which is used to control the information ratio between longer time series and shorter time series. r l represents the reset gate, which determines which information of the shorter time series to retain for the calculation of the long sequence information. C l is the obtained long time series information. f z , f r and f c are all learnable convolution kernels. b z , b r and b c are all learnable feature biases. represents the dot product.

[0073] 2.3 STIGNN Based on Spatiotemporal Feature Mining

[0074] The present invention presents the framework of STIGNN in Figure 5 . It includes three serially connected GateTCN modules and two parallel computing ChebGAT modules. The model input data is the road information detected by road sensors at different time periods and the sensor distribution map. Among them, after the road sensors pass the road information detected at different times through the three-layer GateTCN module to mine the temporal correlation, the road features integrating the temporal sequence information can be obtained. Then, the road features integrating the temporal sequence and the sensor network are input into the ChebGAT module to deeply mine the spatiotemporal features of the road network for the final traffic flow prediction.

[0075] When the ChebGAT module captures spatial features, the present invention introduces an adaptively learned traffic graph G ′ =(V, E′ ),G ′ The edge and edge weight are learned end-to-end through the stochastic gradient descent algorithm. By doing so, the present invention enables the model to discover hidden spatial dependencies. Denote the adjacency matrix of G ′ as A adp , and use randomly initialized matrices E1, E2 ∈ R N×c to learn A adp :

[0076] A adp = Softmax(Relu(E1E2 T )) (8)

[0078] Name E1 as the source node embedding and E2 as the target node embedding. By multiplying E1 and E2, the spatial dependence weight between the source node and the target node can be obtained. The present invention uses the ReLU activation function to eliminate weak connections and the SoftMax function to normalize the adaptive adjacency matrix.

[0079] Through two ChebGAT modules, the present invention respectively obtains the spatio-temporal feature vectors captured by the realistic sensor network and the sensor network based on adaptive learning, denoted as X real and X adp . The feature fusion process of the two is as follows:

[0080] X fusion = βX real + (1 - β)X adp (9)

[0082] where β is a hyperparameter adjusted manually. It is found through experiments that when β = 0.6, the model can achieve the best prediction results.

[0083] 3. Experiments and Analysis

[0084] 3.1 Dataset and Preprocessing

[0085] The present invention conducts experiments on two public transportation network datasets METR-LA and PEMS-BAY released by Li et al. Among them, METR-LA includes traffic speed data of 207 sensors collected on the highways in Los Angeles County from March 1, 2012 to June 30, 2012, and PEMS-BA is collected by the California Department of Transportation (CalTrans) Performance Evaluation System (PEMS), collecting 325 sensor data collected in the Bay Area from January 1, 2017 to May 31, 2017. Table 1 provides detailed dataset information.

[0086] Table 1 Dataset Details

[0087] Dataset Number of nodes Number of edges Time step Time interval METR-LA 207 1515 34272 5 minutes PEMS-BAY 325 2369 52116 5 minutes

[0088] The original data contains the longitude and latitude coordinates of the sensors, the road map, and the data measured by the sensors. In order to determine which sensors should be connected to form a sensor network, the present invention introduces a threshold Gaussian kernel function as follows:

[0089]

[0090] where a ij ∈A represents the spatial distance weight between nodes i and j, and dist(v i , v j ) is the shortest distance between sensors calculated by the Floyd algorithm. δ is the standard deviation of the sensor distances, and k is a predefined threshold. Only when the shortest distance between sensors is less than k, it is adopted as an edge of the sensor network and the weight is calculated.

[0091] 3.2 Experimental Setup

[0092] All experiments were conducted on a single Tesla v100s GPU with 384GB of RAM, and the comparison models were reproduced using the Pytorch deep learning framework. The road information detected by the sensors and the sensor network in the first 60 minutes were used as inputs to detect the traffic flow in the next 15, 30, and 60 minutes, where 15, 30, and 60 are in the short-term, medium-term, and long-term prediction intervals respectively. Both the METR-LA and PEMS-BAY datasets involved in model evaluation were randomly divided into three parts, with 70% as the training set, 10% as the validation set, and the remaining 20% as the test set. During the training process, STIGNN used the Adam algorithm to optimize the gradient descent process, with the initial learning rate set to 0.0003, dropout set to 0.6, and epoch set to 200.

[0093] The present invention conducted comparative tests with the following baselines: ARIMA: Autoregressive Integrated Moving Average model using a Kalman filter. FC-LSTM: A recurrent neural network traffic flow prediction model with fully connected LSTM hidden units. DCRNN: A traffic flow prediction model that captures spatial dependencies using bidirectional random walks on a traffic network and captures temporal dependencies using a pre-sampled encoder-decoder structure. STGCN: Directly constructs a graph convolutional model in a time-series based multi-layer traffic network, replacing the conventional graph neural network + recurrent neural network architecture, effectively reducing the number of parameters and achieving a faster training speed. Graph WaveNet: Uses a learnable adaptive dependency matrix to capture hidden spatial dependencies in the data and processes long time series through dilated one-dimensional convolutions with different receptive fields. ST-MetaNet: Adopts a sequence-to-sequence architecture consisting of an encoder that learns historical information and a decoder that makes predictions step by step to mine the spatio-temporal features of a traffic network. STGRAT: Proposes an attention-based encoder-decoder architecture that uses node attention to capture spatial correlations between roads and uses temporal attention to capture the temporal dynamics of long sequences. STGAT: Processes long time series by stacking gated time convolutional layers and uses a graph attention network to mine the spatial features of a traffic network.

[0094] To verify the effectiveness of STIGNN, the present invention uses the following three metrics to test the performance of each model. Including Root Mean Square Error (RMSE):

[0095]

[0096] Mean Absolute Error (MAE):

[0097]

[0098] Mean Absolute Percentage Error (MAPE):

[0099]

[0100] Among them, Z represents the number of samples, represents the traffic flow prediction value of the i-th sample, y i represents the actual traffic flow of the i-th sample.

[0101] 3.3 Experimental Results

[0102] Table 2 shows the overall performance of STIGNN and all comparison models on the PEMS - BAY and METR - LA datasets. STIGNN has achieved excellent results in the three evaluation metrics (MAE, RMSE, and MAPE) on these two datasets. The results indicate that STIGNN in the present invention can be applied to short - term, medium - term, and long - term traffic flow prediction scenarios.

[0103] The models that combine graph neural networks and time - series networks have lower prediction errors in all evaluation metrics on the two datasets compared to the traditional time - series analysis models ARIMA and FC - LSTM. The traffic flow prediction ability that ignores the spatial topological information of the traffic network due to the interconnection between roads is limited and cannot handle complex traffic data well. At the same time, models that combine spatio - temporal information will also show different characteristics due to different spatial feature extraction components and temporal feature mining components. For example, the model STGRAT with dual spatial and temporal attention has strong fitting ability in short - term prediction and has achieved the best short - term prediction results in both datasets. However, the existence of attention will weaken its generalization ability, which makes its performance slightly inferior to STIGNN in long - term prediction. Other models such as Graph WaveNet, by developing a new adaptive correlation matrix and learning through node embedding, can accurately capture long - term dependencies in the data, making the model also show high accuracy in long - term prediction. STIGNN combines the neighbor order and the relevance of node features in spatial feature mining and uses dilated causal convolution with fusion gates in the time series, enabling the model to achieve the best prediction performance in both datasets.

[0104] Due to the extremely complex traffic network in Los Angeles, the performance of the model in the METR - LA dataset is lower than that in the PEMS - BAY dataset. However, STIGNN significantly reduces the prediction error in the METR - LA dataset, which means that compared with other baseline methods, STIGNN can better handle complex network situations.

[0105] Table 2 Performance comparison of STIGNN and other models on the PEMS - BAY and METR - LA datasets

[0106]

[0107]

[0108] The present invention proposes a spatio-temporal deep learning method for traffic flow prediction, STIGNN. STIGNN is a general model for processing spatio-temporal prediction tasks. In order to comprehensively consider the differences in distances between traffic points and road conditions, the present invention uses dual attention of road node features and inter-road distances, and designs a ChebGAT module to mine potential and existing spatial dependencies. In order to effectively extract the time series information of the traffic network, STIGNN introduces a gated unit in the TCN module, named the GateTCN module, to achieve the mining and application of long time series features. Finally, STIGNN has achieved good results on two public real traffic network datasets.

[0109] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the technical solutions of the present invention shall be included within the protection scope of the present invention.

Claims

1. A traffic flow prediction method for realizing co-mining of traffic point semantics and network topology, characterized in that: The method is called STIGNN. STIGNN consists of two parallel computing ChebGAT modules and three-layer serially connected GateTCN modules. The ChebGAT module couples the ChebNet static attention and the GAT dynamic attention to fully explore the spatial correlation of the transportation network and realize the simultaneous mining of topological features and semantic information. The GateTCN module integrates the time series model of the gating mechanism to mine the time correlation of road information detected at different times, and then obtains the road features of the fused time series information. The road features of the fused time series and the sensor network are then input into the ChebGAT module to deeply mine the spatiotemporal characteristics of the road network for the final traffic flow prediction; The process of aggregating features in the ChebNet model is as follows: Among them, σ(·) represents the nonlinear activation function, θ i is the learnable static weight of the i-th order neighbor. When i=0, it indicates the importance of its own information. i represents the degree of node i, W is the learnable weight matrix, x i represents the traffic point information of the first-order neighbors of node i, x j represents the traffic point information of the second-order neighbors of node j, x k represents the traffic point information of the third-order neighbors of node k, d j represents the degree of node j, d k represents the degree of node k; The neighborhood perception range of the ChebGAT module is fixed to 3rd order, and the process of aggregating features is as follows: STIGNN transposes the sensor feature matrix X∈R M×N×F →X∈R N×F×M , and set the dilated causal convolution kernel f∈R K , where K represents the size of the convolution kernel. The process of dilated causal convolution can be expressed as: Among them, d is the dilation factor of the dilated causal convolution, m represents the length of the time series of the intercepted traffic data, s is the value of traversal K, X★ m f represents the convolution of X using the dilated causal convolution kernel f, and f(s) represents the sth parameter of the convolution kernel f; The GateTCN module data feature processing process is as follows: Among them, zl represents the update gate of the lth layer of GateTCN, which is used to control the information ratio of longer time series and short time series, r l Represents the reset gate, C l To obtain long time series information, f z , f r With f c are all learnable convolution kernels, b z , b r and b c are all learnable feature biases, Represents the dot product.

2. The traffic flow prediction method for realizing co-mining of traffic point semantics and network topology according to claim 1 is characterized in that: The GAT aggregation feature process is as follows: Among them, α ij Indicates the degree of semantic association between nodes i and j; Among them, || represents the splicing operation, is the learnable weight vector, N i represents the first-order neighbor set of node i.

3. The traffic flow prediction method for realizing co-mining of traffic point semantics and network topology according to claim 1 is characterized in that: When capturing spatial features, the ChebGAT module introduces an adaptively learned traffic graph G′=(V, E′). The edges and edge weights in G′ are learned end-to-end using a stochastic gradient descent algorithm. The adjacency matrix of G′ is denoted by A adp , using randomly initialized matrices E1, E2∈R N×c Study A adp : YOUR adp =Softmax(Relu(E1E2 T )) Among them, E1 is named as the source node embedding, and E2 is named as the target node embedding.

Citation Information

Patent Citations

  • Space-time diagram node attribute prediction method fusing adaptive graph diffusion convolutional network

    CN115828990A

  • Space-time adaptive dynamic graph convolutional network traffic flow prediction method

    CN118629226A