Urban-level traffic flow prediction method fusing graph attention network and Transform
By fusing graph attention networks and Transformers, a city-level traffic flow prediction method is constructed, which solves the shortcomings of existing models in terms of topological and dynamic dependencies in urban road networks, achieves high-precision traffic flow prediction, and improves the efficiency and robustness of urban traffic management.
Patent Information
- Application Number
- CN202511470694.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-03-17
AI Technical Summary
Existing traffic flow prediction models lack sufficient exploration of road network topology and dynamic dependencies in urban road networks, making it difficult to effectively capture the diverse patterns and non-stationary characteristics in traffic sequences, and their prediction performance is poor on complex urban road networks.
By integrating graph attention networks and Transformers, a city-level traffic flow prediction method is constructed. This method dynamically learns the spatial dependencies between nodes by constructing a road network adjacency graph and a spatiotemporal tensor sequence, combined with node representations based on deep temporal context semantics, and models the long- and short-term cycle patterns of traffic flow through a self-attention mechanism.
It significantly improves the accuracy and robustness of traffic flow forecasting, enabling more accurate prediction of urban traffic flow and providing data support for signal control optimization and congestion management.
Smart Images

Figure CN121686801A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of traffic flow prediction, specifically relating to a city-level traffic flow prediction method that integrates graph attention networks and Transformers. Background Technology
[0002] With the continuous growth of motor vehicle ownership and the rapid increase in road traffic demand in Chinese cities, urban road network structures are becoming increasingly complex, and traffic congestion is becoming more and more prominent. High-precision traffic flow prediction, as a core supporting technology of intelligent transportation systems, is of great significance for improving urban traffic efficiency, optimizing signal control strategies, and alleviating congestion. However, existing research mostly uses graphical models and time-series models to independently model spatiotemporal relationships. Time modeling is difficult to effectively capture the diverse patterns and non-stationary characteristics in traffic sequences, and it does not pay enough attention to short-term and long-term cycles; spatial modeling is prone to gradient problems when dealing with long-distance sequences, and the independent design of spatiotemporal modules makes it difficult to fully explore road network topology and dynamic dependencies. At the same time, current traffic flow prediction models are mostly validated on highway data with regular structure and less interference. Summary of the Invention
[0003] This invention addresses the common shortcomings of existing traffic flow prediction models, which generally lack sufficient exploration of road network topology and dynamic dependencies, and are mostly based on the structurally regular highway scenario, with relatively insufficient research on the complex and heterogeneous urban road networks. It provides a city-level traffic flow prediction method that integrates graph attention networks and Transformers, thereby improving the accuracy and robustness of traffic flow prediction. This invention will provide data support for urban signal control optimization and congestion management, guiding intelligent urban traffic planning and management.
[0004] To address the above technical problems, this invention provides the following technical solution: a city-level traffic flow prediction method integrating graph attention networks and Transformers, comprising the following steps:
[0005] S1. Based on actual license plate recognition probe data, construct a refined urban traffic flow dataset and build a road network adjacency graph;
[0006] S2. Construct a graph attention neural network layer that takes the road network adjacency graph as input and a spatiotemporal tensor sequence containing spatial information as output; construct a linear projection layer that takes traffic flow data as input and high-dimensional traffic flow features as output.
[0007] A Transformer layer is constructed by taking a spatiotemporal tensor sequence containing spatial information and a weighted fusion feature of high-dimensional traffic flow features as input and a node representation of deep temporal context semantics as output.
[0008] Based on graph attention neural network layers, linear projection layers, and transformer layers, combined with the output projection layer, a spatiotemporal traffic flow prediction network is constructed.
[0009] S3. Train a spatiotemporal traffic flow pre-network using a traffic flow dataset to obtain a traffic flow pre-model, which can be used to predict urban traffic flow.
[0010] Furthermore, the aforementioned step S1 includes the following sub-steps:
[0011] S1.1. Based on the density of point distribution and road accessibility, select collection points, match spatial coordinates with road base maps, and construct a time series data structure;
[0012] S1.2, Complete the missing data and perform normalization processing;
[0013] S1.3 On the loaded high-precision urban road network base map, the direct road connection relationships between the collection points are identified and drawn one by one by the manual interpretation and manual digitization methods to construct the road network adjacency map.
[0014] Furthermore, the aforementioned step S1.2 includes the following sub-steps:
[0015] S1.2.1 For single-point short-term missing data, a time-dimensional linear interpolation method is used, employing the effective flow values of adjacent times before and after the missing point for interpolation calculation, as shown in the following formula:
[0016]
[0017] in, and These represent traffic flow at times before and after the missing data points;
[0018] S1.2.2 For multiple consecutive missing points, based on the principle of spatial correlation, weighted interpolation is performed according to the flow rate change trend of adjacent collection points in the same time period, as shown in the following formula:
[0019] in, It represents the traffic flow at the known point preceding the missing time period. It is the traffic flow at the first known point after the missing time period. It is any missing point in time between these two known points;
[0020] S1.2.3. The min-max normalization algorithm is used to linearly map the completed dataset to the [0,1] interval, eliminating the dimensional differences between different collection nodes and generating structured matrix data, as shown in the following formula:
[0021]
[0022] In the formula, This represents the maximum value of this feature in the dataset. This represents the minimum value of this feature in the dataset. This represents the raw traffic flow data.
[0023] Furthermore, in step S1.3 above, when identifying and drawing the direct road connections between collection points one by one, the existence of a valid path between each pair of adjacent collection points is determined based on the actual road traffic direction and actual accessibility to ensure that the connection relationship has physical rationality. The attribute information of the start and end points and edges is then labeled, and an adjacency matrix A is constructed based on the connection pairs.
[0024] .
[0025] Furthermore, in step S2 described above, the linear projection layer represents the input traffic data as... In the form of, For batch size, This refers to the number of nodes. Given the length of the input sequence, the original input sequence is mapped to a high-dimensional feature space:
[0026]
[0027] in, belong gather, belong gather, This represents the dimension of the hidden layer features.
[0028] Furthermore, in step S2 above, the graph attention neural network layer inputs the feature vectors of each node into the graph attention layer at each time step, performs weighted interaction with the information of neighboring nodes, dynamically learns the spatial dependencies between nodes in the traffic network, and defines the node features as... The formula for updating the node features of the graph attention neural network layer is as follows:
[0029]
[0030] in, It is an activation function used to implement nonlinear transformations. The weight matrix is adjustable through learning. What it represents is a node. The neighborhood group, The attention coefficient is represented by the following formula:
[0031]
[0032] in, These are attention parameters that can be used for learning, where T is the transpose. This represents a vector concatenation operation. Represents a node eigenvectors, Indicates adjacent nodes eigenvectors, Representing neighboring nodes eigenvectors, This represents a linear rectified function with leakage. This represents the linear rectification function.
[0033] Furthermore, the aforementioned graph attention neural network layer uses multiple independent attention heads to perform the above calculations in parallel, and concatenates their output features to output:
[0034]
[0035] in, This is to increase the number of heads that receive attention.
[0036] Furthermore, in step S2 above, the spatial features extracted by the graph attention neural network layer and the high-dimensional traffic flow features output by the linear projection layer are weighted and fused to output a weighted fused feature, calculated as follows:
[0037]
[0038] in, The fusion coefficient is... Indicates the characteristics after fusion. This represents the spatial features extracted layer by layer by the graph attention neural network. This represents the time characteristics after linear projection.
[0039] Furthermore, the aforementioned Transformer layer reorganizes the spatiotemporal tensor sequence output by the graph attention module into an input sequence along the time dimension, and adds position encoding for each time step to explicitly preserve the positional information of the temporal order in the sequence; the sequence is input into a Transformer encoder structure composed of multiple stacked sub-layers, each layer containing a multi-head self-attention module and a feedforward neural network;
[0040] The self-attention mechanism transforms the input feature vector into three sets of representations: query, key, and value. By calculating the attention score between each time step, it performs global dependency modeling on the entire sequence and automatically identifies the influence of important historical time points on the current prediction time. The calculation is as follows:
[0041]
[0042] in, , and These correspond to the query, key, and value, respectively.
[0043] Multi-head attention mechanisms utilize multiple self-attention heads working in parallel to capture different patterns in the data, as shown in the following equation:
[0044] ,
[0045] in, , These represent the linear transformation matrices for query, key, value, and output, respectively.
[0046] After the self-attention output, a non-linear transformation is performed on the representation through a feedforward sub-layer, and residual connections and layer normalization are used to improve stability and training efficiency. This sub-layer independently applies a two-layer neural network to each time step position in the self-attention output:
[0047] ,
[0048] Here, FFN is a two-layer linear mapping plus a nonlinear activation function. The weight matrix representing the first-level linear transformation. Indicates the bias term of the first layer, The weight matrix representing the second-level linear transformation. This represents the bias term of the second layer.
[0049] After stacking multiple layers of encoders in the Transformer layer, the output node representation with deep temporal context semantics is obtained. This representation integrates long-term temporal dependencies and local change trends, and is passed as input to the prediction module to complete the regression prediction of traffic flow or other indicators in future time steps.
[0050] ,
[0051] in Belonging to the real number field OK Column matrix, i.e. , This refers to the prediction step size. This represents the temporal features extracted by the Transformer encoder.
[0052] Furthermore, the aforementioned step S3 includes the following sub-steps:
[0053] S3.1 Input the training data into the model. The model calculates the predicted output through forward propagation. The mean squared error (MSE) is used as the loss function to measure the error between the predicted result and the true label, as shown in the following formula:
[0054] ,
[0055] In the formula, , These represent the model's actual value and predicted value, respectively. This represents the total number of samples.
[0056] S3.2. Based on the error, the gradient information of the parameters of each layer of the model is calculated by the backpropagation algorithm, and the Adam optimizer is used to update the model parameters;
[0057] S3.3. Repeat steps S3.1 to S3.2 to gradually optimize the model parameters and finally train a high-precision traffic flow prediction model that can capture complex spatiotemporal dependencies.
[0058] Compared with the prior art, the beneficial technical effects of the present invention using the above technical solution are as follows:
[0059] This invention innovatively proposes a collaborative learning framework integrating graph attention networks and Transformers, overcoming the aforementioned limitations based on high-precision digital road network representation. By constructing a realistic topological adjacency matrix containing arterial roads, secondary roads, and complex intersections, and utilizing the GAT layer to dynamically parse the spatial dependency weights between nodes, it accurately characterizes heterogeneous effects such as ramp diversion and intersection propagation. Combined with the Transformer encoder to synchronously model the long- and short-term periodic patterns and non-stationary characteristics of traffic flow, it effectively overcomes the problem of long-distance sequence gradient decay. Furthermore, through a deep spatiotemporal feature coupling mechanism and adaptive computation optimization, this invention significantly improves the efficiency and robustness of the model in micro-level road segment prediction, providing core technical support for intelligent transportation decision-making such as dynamic signal timing optimization and precise public transport capacity scheduling. Attached Figure Description
[0060] Figure 1 This is a schematic diagram of the method flow of the present invention.
[0061] Figure 2 This is a schematic diagram of the road network topology structure of the present invention in ArcGIS.
[0062] Figure 3 This is a structural diagram of the spatiotemporal traffic flow prediction model that integrates GAT and Transformer models according to the present invention.
[0063] Figure 4(a) shows the performance of the traffic flow prediction method based on graph neural network of the present invention on the training set.
[0064] Figure 4(b) shows the performance of the traffic flow prediction method based on graph neural networks in this invention on the test set. Detailed Implementation
[0065] To better understand the technical content of the present invention, specific embodiments are described below in conjunction with the accompanying drawings.
[0066] In this invention, various aspects of the invention are described with reference to the accompanying drawings, in which numerous illustrative embodiments are shown. Embodiments of the invention are not limited to those depicted in the drawings. It should be understood that the invention is implemented through any of the various concepts and embodiments described above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed herein are not limited to any particular implementation. Furthermore, some aspects of the invention disclosed may be used alone or in any suitable combination with other aspects of the invention disclosed.
[0067] like Figure 1 As shown, this invention provides a city-level traffic flow prediction method that integrates graph attention networks and Transformers, characterized by the following steps:
[0068] S1. Based on actual license plate recognition probe data, a refined urban traffic flow dataset is constructed, and a road network adjacency graph is built. In this embodiment, taking the main urban area of a city as an example, a road network covering 81 square kilometers with more than 300 actual license plate recognition (LPR) probe data collection points is selected as the data source. After screening, 252 representative collection points are retained, forming 7776 time slices with 5-minute intervals, with a total data volume of approximately 2 million records, which can fully reflect the traffic fluctuation characteristics throughout the day.
[0069] Step S1 includes the following sub-steps:
[0070] S1.1. Based on the density of point distribution and road accessibility, collection points are selected, and their spatial coordinates are matched with the road base map to construct a time-series data structure. Specifically, the original LPR collection data is imported into the Gaode Cloud Map system. Based on the geographical distribution density, traffic accessibility, and importance of road nodes, 252 representative target collection points are selected. The latitude and longitude coordinates of the selected points are exported and imported into the ArcGIS geographic information system. By overlaying them with the road network base map, spatial matching and visual confirmation of the collection points and road topology are achieved. Pandas and NumPy are used to clean, sort, and reconstruct the data, organizing the traffic flow data into NumPy arrays with defined dimensions.
[0071] S1.2, Complete the missing data and perform normalization processing, which includes the following sub-steps:
[0072] S1.2.1. Regarding the few missing data points in LPR traffic data, assuming that traffic flow between two immediately preceding and following known data points exhibits a linear trend, a time-axis-based linear interpolation method is used to fill in the single-point short-term missing data in the traffic data. For valid single-point missing data at adjacent time points... interpolation Calculate as follows:
[0073]
[0074] in, and These represent traffic flow before and after the data gap point.
[0075] S1.2.2 For data missing within a short continuous time period, a weighted interpolation estimation is performed by combining the traffic flow trend of surrounding points during the same time period, calculated as follows:
[0076] As shown in the following formula:
[0077] in, It represents the traffic flow at the known point preceding the missing time period. It is the traffic flow at the first known point after the missing time period. It is any missing point in time between these two known points;
[0078] S1.2.3. The min-max normalization algorithm is used to linearly map the completed dataset to the [0,1] interval, eliminating the dimensional differences between different collection nodes and generating structured matrix data, as shown in the following formula:
[0079]
[0080] In the formula, This represents the maximum value of this feature in the dataset. This represents the minimum value of this feature in the dataset. This represents the raw traffic flow data.
[0081] S1.3. Using manual interpretation and cross-validation methods, the connection relationship between point pairs is marked according to the actual road traffic rules. An adjacency matrix is established based on the spatial coordinates of the collected points. If two points are directly connected by a road, the corresponding matrix position is assigned a value of 1; otherwise, it is assigned a value of 0. This constructs a road network adjacency graph.
[0082] like Figure 2As shown in the embodiment, the latitude and longitude coordinates of the LPR collection points are imported into the ArcGIS system. On the loaded high-precision urban road network base map, direct road connections between the collection points are identified and drawn one by one using a combination of manual interpretation and manual digitization. During the connection process, based on the actual road traffic direction and accessibility, it is determined whether there is a valid path between each pair of adjacent collection points to ensure the physical rationality of the connection relationship. The origin, destination, and edge attribute information are then labeled, thereby constructing a 252×252 adjacency matrix A based on the connection pairs.
[0083]
[0084] S2, such as Figure 3 As shown, a graph-annotation neural network layer is constructed with a road network adjacency graph as input and a spatiotemporal tensor sequence containing spatial information as output. A linear projection layer is constructed with traffic flow data as input and high-dimensional traffic flow features as output.
[0085] A Transformer layer is constructed by taking a spatiotemporal tensor sequence containing spatial information and a weighted fusion feature of high-dimensional traffic flow features as input and a node representation of deep temporal context semantics as output.
[0086] Based on graph attention neural network layers, linear projection layers, and transformer layers, combined with the output projection layer, a spatiotemporal traffic flow prediction network is constructed.
[0087] The road network structure is modeled using a Graph Attention Network (GAT). The input traffic data is represented as... In the form of, For batch size, This refers to the number of nodes. Given the length of the input sequence, the original input sequence is mapped to a high-dimensional feature space:
[0088]
[0089] in, belong gather, belong gather, This represents the dimension of the hidden layer features.
[0090] By leveraging the GAP layer to process the input traffic time-series data, such as historical traffic flow, at each time step, a graph with the road network as its topology is constructed. The feature vectors of each node at that time are input into the graph attention layer, where they are weighted and interact with information from neighboring nodes. This dynamically learns the spatial dependencies between nodes in the traffic network and defines the node features as... The node feature update formula for the GAT layer is as follows:
[0091]
[0092] in, It is an activation function used to implement nonlinear transformations. The weight matrix is adjustable through learning. What it represents is a node. The neighborhood group, The attention coefficient is represented by the following formula:
[0093]
[0094] in, These are attention parameters that can be used for learning, where T is the transpose. This represents a vector concatenation operation. Represents a node eigenvectors, Indicates adjacent nodes eigenvectors, Representing neighboring nodes eigenvectors, This represents a linear rectified function with leakage. This represents the linear rectification function.
[0095] To improve training stability and model expressive power, multiple independent attention heads are used to perform the above calculations in parallel, and their output features are concatenated to output:
[0096]
[0097] in, This is to increase the number of heads that receive attention.
[0098] After passing through the GAT layer, the model can adaptively adjust neighbor weights based on node feature similarity, emphasizing important connections and weakening irrelevant connections, thereby generating node embeddings containing rich spatial context information. To further enhance the spatial feature representation capability, this invention designs a feature fusion mechanism, which weights and fuses the spatial features extracted by GAT with the features obtained by initial linear projection of the input data to improve the ability to express structural information in the final representation. The calculation is as follows:
[0099]
[0100] in, The fusion coefficient is... Indicates the characteristics after fusion. This represents the spatial features extracted by the GAT layer. This represents the time characteristics after linear projection.
[0101] The Transformer architecture is used to perform deep modeling of the dynamic characteristics of traffic state evolution at each node over time. First, the efficient sparse graph computation capabilities provided by toolkits such as PyTorch Geometric are utilized to process large-scale road networks. The spatiotemporal tensor sequence output by the graph attention module is reorganized into the input sequence along the time dimension, and a positional encoding is added to each time step to explicitly preserve the temporal order of the positional information in the sequence. The sequence is then input into a Transformer encoder structure consisting of multiple stacked sub-layers, each layer containing a multi-head self-attention module and a feedforward neural network.
[0102] The self-attention mechanism first transforms the input feature vector into three sets of representations: query, key, and value. By calculating the attention score between each time step, it performs global dependency modeling on the entire sequence, automatically identifying the influence of important historical time points on the current prediction time. Specifically, the calculation is as follows:
[0103]
[0104] in, , and These correspond to the query, key, and value, respectively.
[0105] Multi-head attention mechanisms utilize multiple self-attention heads working in parallel to capture different patterns in the data, as shown in the following equation:
[0106]
[0107] in, , These represent the linear transformation matrices for query, key, value, and output, respectively.
[0108] After the self-attention output, a non-linear transformation is performed on the representation through a feedforward sub-layer, and residual connections and layer normalization are used to improve stability and training efficiency. This sub-layer independently applies a two-layer neural network to each time step position in the self-attention output:
[0109]
[0110] Here, FFN is a two-layer linear mapping plus a nonlinear activation function (SiLU). The weight matrix representing the first-level linear transformation. Indicates the bias term of the first layer, The weight matrix representing the second-level linear transformation. This represents the bias term of the second layer.
[0111] After stacking multiple layers, the Transformer encoder outputs a node representation with deep temporal context semantics. This representation integrates long-term temporal dependencies and local change trends, and is then passed as input to the prediction module to complete regression predictions of traffic flow or other indicators in future time steps.
[0112]
[0113] in Belonging to the real number field OK Column matrix, i.e. , This represents the temporal features extracted by the encoder. This refers to the prediction step size, which is set to 1 in this experiment.
[0114] S3. Train a spatiotemporal traffic flow pre-network using a traffic flow dataset to obtain a traffic flow pre-model, which can be used to predict urban traffic flow.
[0115] The processed traffic dataset was divided into training and testing sets in an 8:2 ratio to ensure sufficient model training and objective evaluation. During training, mean squared error (MSE) was used as the loss function to calculate the error between the predicted results and the true labels.
[0116]
[0117] The gradient information of each layer parameter is calculated using the backpropagation algorithm based on the error. The Adam optimizer is used to update the model parameters. Adam automatically adjusts the learning rate by combining the first and second moment estimates of the gradient, achieving efficient and stable parameter optimization. The above steps constitute one training iteration. The entire training process is repeated through multiple iterations to gradually optimize the model parameters. Finally, the fully trained model can effectively extract deep spatiotemporal dependencies and achieve high-precision traffic flow prediction. After training, the mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE) are used to measure the prediction accuracy from different dimensions such as absolute deviation, variance deviation, and relative error, respectively, to comprehensively evaluate the model performance. The performance of the traffic flow prediction method based on graph neural networks in this invention on the training set and test set are shown in Figure 4(a) and Figure 4(b), respectively. The evaluation index results of this experiment are shown in the table:
[0118] Table 1. Prediction Model Error Indicators
[0119]
[0120] While the present invention has been described above with reference to preferred embodiments, it is not intended to limit the invention. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the invention. Therefore, the scope of protection of the present invention shall be determined by the claims.
Claims
1. A method of urban traffic flow prediction by fusing graph attention network and Transformer, characterized in that, The method comprises the following steps: S1, based on the actual license plate recognition probe data, constructing a city fine feature traffic flow data set, and constructing a road network adjacency graph; S2, constructing a graph attention neural network layer taking the road network adjacency graph as input and a space information containing spatio-temporal tensor sequence as output, constructing a linear projection layer taking the traffic flow data as input and the traffic flow high-dimensional feature as output, taking the space information containing spatio-temporal tensor sequence and the traffic flow high-dimensional feature weighted fusion feature as input, and the node representation of deep time context semantics as output, constructing a Transformer layer; Based on the graph attention neural network layer, the linear projection layer, the transformer layer, and the output projection layer, a spatio-temporal traffic flow prediction network is constructed; S3, training the spatio-temporal traffic flow prediction network by using the traffic flow data set to obtain a traffic flow prediction model, which is used to realize the prediction of urban traffic flow.
2. The urban traffic flow prediction method of claim 1, wherein, Step S1 includes the following sub-steps: S1.1, according to the point distribution density and the road accessibility, screening the collection points, matching the spatial coordinates with the road base map, and constructing a time series data structure; S1.2, filling in the missing data and performing normalization processing; S1.3, on the loaded high-precision city road network base map, the direct road connection relationship between the collection points is identified and drawn one by one by using artificial interpretation and manual digitization method, and the road network adjacency graph is constructed.
3. The urban traffic flow prediction method of claim 1, wherein, Step S1.2 includes the following sub-steps: S1.2.1, for single-point short-time missing, a time dimension linear interpolation method is used to interpolate and calculate the effective flow value of the missing point before and after the adjacent time, as follows: , wherein, and respectively are traffic flow at time instants before and after the data missing point. S1.2.2, for multi-point continuous missing, combined with the spatial correlation principle, based on the flow change trend of adjacent collection points in the same time period, weighted interpolation is carried out, as follows: , wherein, is the traffic flow at the known point preceding the missing time period, is the traffic flow at the known point following the missing time period, is any missing time point between these two known points; S1.2.3, the completed data set is linearly mapped to the [0, 1] interval by using the minimum-maximum normalization algorithm, the dimension difference between different collection nodes is eliminated, and a structured matrix data is generated, as follows: , wherein, denotes the maximum value of this feature in the data set, denotes the minimum value of this feature in the data set, denotes the original traffic flow data.
4. The urban traffic flow prediction method of claim 1, wherein, In step S1.3, when identifying and drawing the direct road connection relationship between the collection points one by one, according to the real road traffic direction and the actual accessibility, whether there is an effective path between each pair of adjacent collection points is determined, the connection relationship is ensured to have physical rationality, and the attribute information of the start and end points and the edge is marked, and the adjacency matrix A is constructed based on the connection pair: 。 5. The urban traffic flow prediction method of claim 1, wherein, In step S2, the linear projection layer represents the input traffic data as where is the batch size, is the number of nodes, is the input sequence length, and maps the original input sequence to a high-dimensional feature space: , wherein, belongs to the set, belongs to the set, represents the dimension of the hidden layer features.
6. The urban traffic flow prediction method of claim 1, wherein, In step S2, the graph attention neural network layer inputs the feature vector of each node to the graph attention layer at each time step, performs weighted interaction with the information of the neighboring nodes, dynamically learns the spatial dependency between the nodes in the traffic network, and defines the node feature as The node feature update formula of the graph attention neural network layer is as follows: , wherein, is an activation function for implementing a non-linear transformation, is a weight matrix that can be adjusted by learning, represents a neighbor set of the node , denotes an attention coefficient, which is calculated as follows: , in, These are attention parameters that can be used for learning, where T is the transpose. This represents a vector concatenation operation. Represents a node eigenvectors, Indicates adjacent nodes eigenvectors, Representing neighboring nodes eigenvectors, This represents a linear rectified function with leakage. This represents the linear rectifier function.
7. The urban traffic flow prediction method of claim 6, wherein, The graph attention neural network layer uses multiple independent attention heads to perform the above calculation in parallel, and outputs the output features by splicing them: , wherein, is the number of attention heads.
8. The city-level traffic flow prediction method of fusing graph attention network and Transformer according to claim 7, in step S2, the spatial features extracted by the graph attention neural network layer and the traffic flow high-dimensional features output by the linear projection layer are weighted and fused to output the weighted fusion features, which are calculated as follows: , wherein, is a fusion coefficient, denotes the fused feature, denotes the spatial feature extracted by the graph attention neural network layer, denotes the time feature after linear projection.
9. The urban traffic flow prediction method of claim 1, wherein, The Transformer layer reorganizes the spatio-temporal tensor sequence output by the graph attention module in the time dimension into an input sequence, and adds position encoding to each time step to explicitly retain the position information of the time sequence. The sequence is input into the Transformer layer encoder structure composed of multiple stacked sub-layer groups, each of which contains a multi-head self-attention module and a feedforward neural network. The self-attention mechanism transforms the input feature vector into three groups of query, key and value representations, models the global dependence of the entire sequence by calculating the attention score between time steps, and automatically identifies the influence degree of important historical time points on the current prediction time. The calculation is as follows: , wherein, , and correspond to a query, a key and a value, respectively; The multi-head attention mechanism uses multiple self-attention heads to work in parallel to capture different patterns in the data, as follows: , wherein, , respectively represent linear transformation matrices for query, key, value and output. After the self-attention output, a feedforward sub-layer is used to perform nonlinear transformation on the representation, and a residual connection and layer normalization are used to improve stability and training efficiency. The sub-layer independently applies a two-layer neural network to each time step position in the self-attention output: , where FFN is a two-layer linear mapping plus a non-linear activation function, denotes a weight matrix of the first layer linear transformation, denotes a bias term of the first layer, denotes a weight matrix of the second layer linear transformation, denotes a bias term of the second layer; After the encoder stack of the Transformer layer, the node representation with deep temporal context semantics is output, which integrates long-term temporal dependence and local change trend, and is transmitted as input to the prediction module to complete the regression prediction of future time step traffic flow or other indicators: , wherein belongs to the row column matrix, i.e. , refers to the prediction step, denotes the time feature extracted by the Transformer encoder.
10. The urban traffic flow prediction method of claim 1, wherein, Step S3 includes the following sub-steps: S3.1, input the training data into the model, and calculate the prediction output by forward propagation. The mean square error (MSE) is used as the loss function to measure the error between the prediction result and the true label, as follows: , wherein, , denote the true and predicted values of the model, respectively, denotes the total number of samples; S3.2, based on the error, calculate the gradient information of each layer parameter of the model by the back propagation algorithm, and update the model parameters using the Adam optimizer; S3.3, steps S3.1 to S3.2 are executed in a loop to gradually optimize the model parameters, and finally a high-precision traffic flow prediction model capable of capturing complex spatio-temporal dependence relationships is trained.
Citation Information
Cited By
Charging station waiting time prediction method based on graph neural network
CN122335245A