Traffic flow prediction method based on enhanced space-time diagram convolutional network
By enhancing the granularity adjustment and feature extraction methods of the spatiotemporal graph convolution network, the problems of traffic flow dynamics and non-stationarity are solved, and high-precision prediction of the correlation of traffic flow in long-distance areas is achieved.
Patent Information
- Application Number
- CN202510567067.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
AI Technical Summary
Existing traffic flow prediction models are difficult to capture dynamics and non-stationarity caused by emergencies or weather factors, and fail to effectively capture traffic flow correlations between long-distance areas, especially spatial correlations between main roads and branches.
Using a method based on an enhanced spatiotemporal graph convolution network, the traffic flow characteristics of the future time step are acquired and predicted through the granularity adjustment module, the SLA gated causal convolution module, the adaptive graph augmentation module and the improved spatiotemporal encoder combined with a multi-layer perceptron module integrating efficient local attention.
It improves the accuracy and robustness of traffic flow prediction, can effectively capture the traffic flow correlation between long-distance areas, especially the spatial correlation between main roads and branches, and has strong model adaptability.
Smart Images

Figure CN120472664A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a traffic flow prediction method based on an enhanced spatiotemporal graph convolutional network. Background Art
[0002] While existing inventions have made significant progress in modeling spatiotemporal correlations, some limitations remain. For one thing, most models fail to fully account for the dynamic and non-stationary nature of traffic flow, often influenced by factors like emergencies and weather. Furthermore, these models often struggle to capture the correlations between traffic flows across distant regions, particularly the spatial correlations between arterial roads and their branch roads. Summary of the Invention
[0003] The present invention discloses a traffic flow prediction method based on an enhanced spatiotemporal graph convolutional network to overcome the above technical problems.
[0004] In order to achieve the above object, the technical solution of the present invention is:
[0005] A traffic flow prediction method based on enhanced spatiotemporal graph convolutional network includes the following steps:
[0006] S1: Obtain the historical traffic flow information of the road section to be analyzed for T time steps through traffic sensor information;
[0007] S2: Based on the granularity adjustment module, according to the preset traffic network diagram and grid division parameters, grid cells in the traffic network diagram after granularity adjustment are obtained, so as to obtain the granularity-adjusted historical traffic flow data based on the traffic flow information of the historical T time steps;
[0008] S3: Based on the SLA gated causal convolution module, the historical traffic flow data after granularity adjustment is obtained to extract the temporal characteristics of the traffic flow;
[0009] S4: Based on the adaptive graph augmentation module, the augmented traffic graph structure and augmented traffic flow data are obtained according to the traffic network graph after granularity adjustment and the historical traffic flow data with the time characteristics of traffic flow extracted;
[0010] S5: Based on the improved spatiotemporal encoder, the traffic flow characteristics of the future time step are obtained according to the augmented traffic map structure and augmented traffic flow data;
[0011] S6: Based on the multi-layer perceptron module with integrated efficient local attention, the traffic flow characteristics of the future time step are obtained by integrating the time and regional attention enhancement mechanism according to the traffic flow characteristics of the future time step to predict the traffic flow in the future time step.
[0012] Beneficial effects: The traffic flow prediction method based on the enhanced spatiotemporal graph convolutional network of the present invention obtains granularity-adjusted historical traffic flow data based on the traffic flow information of the historical T time steps, and obtains the historical traffic flow data for extracting the time characteristics of the traffic flow based on the SLA gated causal convolution module; then the historical traffic flow data for extracting the time characteristics of the traffic flow are augmented; then the traffic flow characteristics of the future time steps are obtained based on the improved spatiotemporal encoder; finally, the traffic flow characteristics of the future time steps that integrate the time and regional attention enhancement mechanism are obtained by integrating the multi-layer perceptron module with efficient local attention, so as to predict the traffic flow of the future time steps. The present invention solves the problem of the dynamic and non-stationary nature of traffic flow due to factors such as emergencies or weather by integrating the multi-layer perceptron module with efficient local attention. At the same time, through the improved spatiotemporal encoder, the prediction model of the present invention can capture the correlation of traffic flows between distant regions, especially the spatial correlation between main roads and branches, and the prediction results are highly accurate, and the prediction model has strong adaptability and higher robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0014] Figure 1 This is a flow chart of the traffic flow prediction method based on enhanced spatiotemporal graph convolutional network of the present invention;
[0015] Figure 2 A road topology diagram and a traffic flow diagram in an embodiment of the present invention;
[0016] Figure 3 Schematic diagram of the overall traffic prediction method model in an embodiment of the present invention;
[0017] Figure 4a Schematic diagram of the original traffic network diagram and historical traffic flow data in an embodiment of the present invention;
[0018] Figure 4b A schematic diagram of the traffic network diagram and historical traffic flow data after granularity adjustment in an embodiment of the present invention;
[0019] Figure 5a Schematic diagram of the overall structure of the SLA gated causal convolution module in an embodiment of the present invention;
[0020] Figure 5bSchematic diagram of the SLA module structure in an embodiment of the present invention;
[0021] Figure 5c Schematic diagram of the structure of the DWC module in the SLA module in an embodiment of the present invention;
[0022] Figure 6 Schematic diagram of an adaptive graph augmentation module in an embodiment of the present invention;
[0023] Figure 7 1 is a diagram showing the overall architecture of a space-time encoder in an embodiment of the present invention;
[0024] Figure 8 Schematic diagram of the ELA module structure in an embodiment of the present invention.
[0025] Figure 9 This is a flow chart of traffic flow prediction of the traffic prediction method model in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0027] This embodiment introduces a traffic flow prediction method based on enhanced spatiotemporal graph convolutional network. Figure 1 、 Figure 3 and Figure 9 As shown, the following steps are included:
[0028] S1: Obtain the historical traffic flow information of the road section to be analyzed for T time steps through traffic sensor information;
[0029] Specifically, this embodiment first gives a traffic network topology G = (V, E, A) and the city-wide traffic flow data X containing T time steps before time t. t-T:t ={X t-T ,…,X t}, learn a traffic flow prediction function to predict the traffic flow of all locations at time t+1, that is, X t+1 ∈R |V| , where R |V| A real number vector representing |V| nodes; |V| represents the total number of nodes in the set of nodes;
[0030] Among them, the traffic flow prediction problem is expressed as follows:
[0031]
[0032] Where G represents the traffic network topology; V represents the set of nodes; E represents the set of edges between two adjacent nodes in the set of nodes; A represents the adjacency matrix of the traffic network topology; t represents the time; T represents the total number of time steps; X t-T:t represents the city's traffic flow data T time steps before time t; X t-T represents the traffic flow data of the entire city at time tT; X t represents the traffic flow data of the entire city at time t; F(·) represents the method used for traffic flow prediction, and θ is a learnable parameter vector; represents the traffic flow prediction function;
[0033] Figure 2 Shows the topological structure of a road graph, Table 1 is the corresponding Figure 2 Traffic flow information of each node at different time points. Specifically, the traffic flow data of each node at time t1 and t2 are recorded in the table, which are {3, 5, 2, 1, 4, 7} and {2, 2, 3, 6, 8, 4} respectively. Based on these historical traffic data, combined with Figure 2 The topological structure of the graph can be used to predict traffic flow at the future time t3, resulting in a predicted traffic flow value of {1, 2, 5, 6, 7, 7} for each node. This process demonstrates that the use of graph structures and time series data can effectively capture the spatial and temporal dependencies of traffic flow, providing a reliable basis for accurate traffic flow prediction.
[0034] Table 1 Traffic flow information of traffic road nodes at different time points
[0035]
[0036] S2: Based on the granularity adjustment module, according to the preset traffic network diagram and grid division parameters, grid cells in the traffic network diagram after granularity adjustment are obtained, so as to obtain the granularity-adjusted historical traffic flow data based on the traffic flow information of the historical T time steps;
[0037] Specifically, in this embodiment, a traffic prediction method model is established to obtain traffic flow information in future time steps based on the traffic flow information in the historical T time steps; the traffic prediction method model includes a granularity adjustment module, an SLA gated causal convolution module, an adaptive graph enhancement module, an improved spatiotemporal encoder, and a multi-layer perceptron module with integrated efficient local attention.
[0038] S2: Based on the granularity adjustment module, according to the preset traffic network diagram and grid division parameters, grid cells in the traffic network diagram after granularity adjustment are obtained, so as to obtain the granularity-adjusted historical traffic flow data based on the traffic flow information of the historical T time steps;
[0039] Specifically, the granularity adjustment module is used to divide the traffic network area into multiple grid units according to the grid division parameters α and β, so as to obtain the grid units in the traffic network graph after granularity adjustment, and to obtain the historical traffic flow data X′ after granularity adjustment according to the preset traffic network graph G and the traffic flow information of the historical T time steps. t-T:t And the traffic network graph g′=(V′, E′, A′) after granularity adjustment; the SLA gated causal convolution module is used to calculate the historical traffic flow data X′ after granularity adjustment. t-T:t , obtain the historical traffic flow data of the time characteristics of the extracted traffic flow, that is, the historical traffic flow data output by the SLA gated causal convolution module The adaptive graph enhancement module is used to extract the time characteristics of traffic flow and effectively decompose its characteristic information based on the traffic network graph after granularity adjustment and the historical traffic flow data of the time characteristics of traffic flow. Obtain the augmented traffic graph structure and augmented traffic flow data The improved spatiotemporal encoder is used to and augmented traffic flow data Extract spatiotemporal features to obtain traffic flow characteristics at the future time step (t+1 moment) The multi-layer perceptron module with integrated efficient local attention is used to obtain the traffic flow characteristics of the future time step by integrating the time and regional attention enhancement mechanism according to the traffic flow characteristics of the future time step;
[0040] Specifically, the granularity adjustment module in this embodiment is an important component used for regional division and feature extraction in the traffic flow prediction model. In the core business districts of large cities, the transportation department needs to understand the overall traffic distribution in the region, rather than the traffic at a single intersection. For example, a business district in Shanghai can use the granularity adjustment module to divide the area into grids, each grid covering facilities such as shopping malls, office buildings, or subway stations. This division method effectively summarizes the overall traffic within the grid, solves the problem of scattered single-point data, facilitates congestion monitoring, optimizes resources, and provides a multi-level perspective for the model.
[0041] Specifically, the preset traffic network diagram G = (V, E, A), grid division parameters α, β and historical traffic flow information X of T time steps are input into the granularity adjustment module. t-T:t ={X t-T ,…,X t}, adjust the granularity of the input data to obtain the traffic network graph G′=(V′,E′,A′) after granularity adjustment and the historical traffic flow data X′ after granularity adjustment t-T:t ={X′ t-T ,…,X′ t This allows the traffic prediction model to flexibly adjust the level of refinement of regional divisions based on demand, effectively capturing the spatial distribution characteristics of traffic flow and providing a multi-level, multi-granular perspective for modeling. Here, G′ represents the granularity-adjusted traffic network graph; V′, E′, and A′ represent the set of nodes after granularity adjustment, the set of edges between two adjacent nodes in the set of nodes after granularity adjustment, and the adjacency matrix of the granularity-adjusted traffic network topology graph, respectively.
[0042] Specifically, first, the longitude and latitude range is [L lng ,U lng ],[L lat ,U lat ] is mapped into a two-dimensional space, and divided into N = α × β grid cells according to the specific characteristics of the area of the traffic section to be predicted and the granularity parameters α and β. Each grid cell represents an independent area, and all grids form a spatial area set, that is, {r1,…,r N}, where r N Represents the Nth grid unit; N is the total number of grids in the traffic network diagram after granularity adjustment; L lng ,U lng are the upper and lower limits of the longitude range respectively; L lat ,U lat By aggregating the connection edges between grids and internal traffic flow data, the set of edges between two adjacent nodes in the set of nodes after granularity adjustment, the adjacency matrix A' of the traffic network topology after granularity adjustment, and the historical traffic flow data X' after granularity adjustment are generated. t-T:t .
[0043] The calculation process of the granularity adjustment module in this embodiment is as follows:
[0044] exist Figure 4a In the process, the original traffic network graph G(V,E,A) and the historical traffic flow X are combined through the granularity adjustment module. t-T:tAccording to the set granularity parameters α = 2, β = 3, it is divided into 6 grid unit areas {r1,…,r6}. Taking the third grid unit r3 as an example, the node set it contains is {v2,v4,v5}. If there is an edge (v1,v2)∈E and v1,v2 belong to different grids r3,r5 respectively, then the corresponding edge (r3,r5)∈E′ is added to the traffic network graph after granularity adjustment, and the adjacency matrix is updated to [3][5]=1. Similarly, V′, E′ and A′ of the traffic flow graph after granularity adjustment can be derived.
[0045] In Table 1, the traffic flow of nodes v2, v4, and v5 at time t1 is 5, 1, and 4 respectively. The sum of them gives the traffic flow of region r3 at time t1 and t2 is 10, and so on. Figure 4b As shown, the historical traffic flow data X′ after granularity adjustment is generated t-T:t , as shown in formula (2):
[0046]
[0047] N=α×β
[0048]
[0049] X′ t-T:t ={X′ t-T ,X′ t-T+1 ,…,X′ t}
[0050] Where: Represents the nth grid unit r in the traffic network graph after granularity adjustment n The total traffic flow at time t; n is the index of the grid in the traffic network graph after granularity adjustment; t represents the current time step; is the traffic flow of the node in the original traffic network graph at time t; v is the index of the node in the original traffic network graph; S n represents the set of nodes in the nth grid unit in the traffic network graph after granularity adjustment; N is the total number of grid units in the traffic network graph after granularity adjustment; X′ t represents the total traffic flow data at time t in the traffic network graph after granularity adjustment; X′ t-T:t represents the historical traffic flow data after granularity adjustment; X′ t-T represents the city-wide traffic flow data at time tT after granularity adjustment; T represents the total number of historical time steps; α and β are both grid division parameters.
[0051] S3: Based on the SLA gated causal convolution module, the historical traffic flow data after granularity adjustment is obtained to extract the temporal characteristics of the traffic flow.
[0052] Specifically, in the SLA gated causal convolution module (Simplified Linear Attention), the input granularity-adjusted historical traffic flow data X′ t-T:t ={X′ t-T ,…,X′ t}∈R T×N , using the characteristic of the causal convolution module that “cannot see the future” to process the input data, and integrating the SLA module to efficiently capture the global dependency of the input data, we get in, represents the historical traffic flow data output by the SLA gated causal convolution module; D represents the feature dimension after processing by the SLA gated causal convolution module. This feature helps capture the temporal correlation of traffic flow and improves the model's ability to model time series patterns. The SLA gated causal convolution module feeds input data into the SLA module and the causal convolution module, respectively. It then performs a Hadamard product operation on the outputs of the two modules to obtain the result processed by the gating mechanism. This effectively combines global and local features, improving feature expression and prediction accuracy. The SLA gated causal convolution module can extract the temporal characteristics of traffic flow and effectively decompose its feature information.
[0053] Specifically, the SLA gated causal convolution module can accurately extract the temporal patterns of historical traffic flow data and provide the model with real time-dependent features.
[0054] The SLA gated causal convolution module includes an SLA module, a causal convolution module and a gating module;
[0055] The SLA module is used to obtain the embedded vector after SLA layer processing based on the historical traffic flow data after granularity adjustment.
[0056] The causal convolution module is used to obtain feature dimensions extracted from the convolved feature tensor based on the historical traffic flow data after granularity adjustment;
[0057] The gating module is used to embed the vector after the SLA layer processing and the feature dimensions extracted from the convolved feature tensor to obtain historical traffic flow data for extracting the temporal characteristics of traffic flow
[0058] Specifically, the SLA gated causal convolution module structure of this embodiment is as follows: Figure 5a As shown, first input X′ t-T:t Enter the SLA module and get X' t-T:t _in, input X′ at the same time t-T:tEnter the causal convolution module to obtain P, Q, and finally according to X' t-T:t _in, P, Q and formula (16) are processed using the gating mechanism to obtain the result X′ t-T:t _in represents the historical traffic flow data output by the SLA module; P represents the first D feature dimensions extracted from the convolutional feature tensor; Q represents the last D feature dimensions extracted from the convolutional feature tensor;
[0059] Preferably, the embedding vector after SLA layer processing is obtained Here’s how:
[0060] S31: Obtain the traffic flow data tensor with added location coding information. The formula used is as follows:
[0061] Specifically, the SLA module structure is as follows Figure 5b As shown, this module is mainly used to extract global features and capture long-term dependencies. t-T:t ∈R T×N , first undergoes linear transformation and adds position encoding, the calculation formula is as follows:
[0062] B1=Linear(X′ t-T:t ) (3)
[0063] B2=B1+Positional_Encoding (4)
[0064] Where: X′ t-T:t represents the historical traffic flow data after granularity adjustment; Linear(·) represents the fully connected layer; B1 represents the traffic flow data tensor output by the fully connected layer, with the feature dimension expanded from 1 to D; B2 represents the traffic flow data tensor with position encoding information added; Positional_Encoding represents the self-learned position information encoding;
[0065] Specifically, B1∈R T×N×D To expand the feature dimension of the input data, we then add self-learning position information encoding Positional_Encoding∈R to B1. T×N×D , to capture the order information and local feature dependencies to obtain B2∈R T ×N×D .
[0066] S32: Rearrange the traffic flow data tensor with the added position coding information to obtain a rearranged input feature tensor;
[0067] Specifically, in this embodiment, in order to adapt the attention mechanism, the traffic flow data tensor B2 with added position coding information is rearranged using a conventional rearrangement operation (reArrange operation) to obtain B′2∈R TN×D (TN = T × N) to calculate Query, Key, Value. TN represents the first dimension of the rearranged feature vector, where the second dimension is D;
[0068] S33: Obtain a query vector, a key vector, and a value vector according to the rearranged input feature tensor to obtain a query vector processed by the multi-head attention mechanism, a key vector processed by the multi-head attention mechanism, and a value vector processed by the multi-head attention mechanism;
[0069] First, get the query vector, key vector, and value vector:
[0070]
[0071] Where: Query represents the query vector; Key represents the key vector; Value represents the value vector; B′2 represents the rearranged input feature tensor, and W represents the flattened time-node feature sequence; Q Represents the projection weight matrix of the query vector; W K Represents the projection weight matrix of the key vector; W V represents the projection weight matrix of the value vector; b q represents the bias term of the query vector; b k represents the bias term of the key vector; b v represents the bias term of the value vector;
[0072] Among them, Query,Key,Value∈R TN×D , respectively using W Q ,W K ,W V ∈R D×D The query, key, and value are obtained. TN represents the total number of elements after flattening, that is, the total length after flattening the time step and node dimensions.
[0073] Next, we get the query vector and key vector after ReLU activation:
[0074] Specifically, in order to eliminate the complexity of Softmax when calculating attention, ReLU activation operations are performed on Query and key respectively. The calculation formula is as follows:
[0075]
[0076] Where: Represents the query vector after being processed by the ReLU activation function; represents the key vector after ReLU activation function processing; δ(·) represents the ReLU activation operation.
[0077] in, For more non-linearly expressive queries and keys.
[0078] Finally, the The value is decomposed into multiple heads, and the query vector processed by the multi-head attention mechanism, the key vector processed by the multi-head attention mechanism, and the value vector processed by the multi-head attention mechanism are taken: each head focuses on different feature patterns, and the calculation formula is as follows:
[0079]
[0080] Where: Represents the query vector after multi-head attention mechanism processing; represents the key vector after multi-head attention mechanism processing; Value′ represents the value vector after multi-head attention mechanism processing, heads is the number of heads of multi-head attention; Multi_Head represents the operation of the multi-head attention mechanism; R represents a real number.
[0081] S34: According to the query vector processed by the multi-head attention mechanism, the key vector processed by the multi-head attention mechanism, and the value vector processed by the multi-head attention mechanism, obtain the output matrix calculated by the efficient attention mechanism and the feature matrix processed by the depthwise separable convolution;
[0082] The formula used to obtain the output matrix calculated by the efficient attention mechanism is as follows:
[0083]
[0084] Where: B3 represents the output matrix calculated by the efficient attention mechanism, which is the output attention feature; Efficient_Att represents the efficient attention mechanism; i represents the index in the query vector; represents the i-th query vector after multi-head attention processing; represents the j-th key vector after multi-head attention processing; Trans(·) represents the matrix transposition operation; Value′ j represents the j-th value vector after multi-head attention processing; j represents the index in the key vector / value vector; N represents the total length after flattening the time step and node dimensions; N is the total number of grid cells in the traffic network graph after granularity adjustment; T is the total number of historical time steps;
[0085] Specifically, is the output attention feature.
[0086] The formula used to obtain the feature matrix after depth-wise separable convolution is as follows:
[0087] Specifically, the DWC module inside the SLA module is used to perform DWC convolution on Value′ to capture local features. The calculation formula is as follows:
[0088] B4=DWC(Value′) (10)
[0089] Where: B4 represents the feature matrix after depthwise convolution; DWC(·) represents depthwise separable convolution;
[0090] in, is a richer feature representation after DWC operation. DWC operation is as follows Figure 5c As shown in the figure, the Value′ is first subjected to deep convolution to obtain the local features of each channel, and then point-by-point convolution is performed to achieve information fusion between channels and generate a richer feature representation. Through DWC convolution, features are extracted and efficiency is significantly improved.
[0091] S35: Based on the output matrix calculated by the efficient attention mechanism and the feature matrix after depth-wise separable convolution, the embedding vector after SLA layer processing is obtained. The formula used is as follows:
[0092] Specifically, a linear layer is used to rearrange and integrate features and normalize the output. The calculation formula is as follows:
[0093]
[0094] Where: SoftMax represents the normalization function; X′ t-T:t _in∈R T×N×D represents the embedded vector after processing by the SLA layer, that is, the historical traffic flow data output by the SLA module;
[0095] Preferably, the feature dimension extracted from the convolved feature tensor is obtained, that is, the calculation process of the causal convolution module is as follows:
[0096] Specifically, for the input data X′ t-T:t ∈R T×N , first through zero padding and causal convolution, the calculation formula is as follows:
[0097] X′ t-T:t _zero=ZeroFill(X′ t-T:t ) (12)
[0098] X′ t-T:t_conv=2Dconv(X′ t-T:t ) (13)
[0099] Where: represents the embedding vector after zero padding; where K t Represents the size of the convolution kernel of the causal convolution module; N represents the total number of grids in the traffic network graph after granularity adjustment; D represents the feature dimension after processing by the SLA gated causal convolution module; Represents the shape of a three-dimensional tensor; X′ t-T:t _conv∈R T×N×2D represents the embedding vector after causal convolution processing; 2Dconv(·) represents a two-dimensional convolution operation, and the size of the convolution kernel is K t , used to extract local features; ZeroFill(·) represents the zero filling operation, that is, inserting (K t -1) zeros ensure that the time step dimension remains unchanged after convolution. This design makes the receptive field of the convolution operation only include the current and previous time steps, while keeping the time dimension of the output consistent with the input, ensuring causality in time series processing.
[0100] The embedding vector X′ after causal convolution t-T:t _conv is split along the feature dimension, and the calculation formula is as follows:
[0101] P=X′ t-T:t _conv[:,:,:D] (14)
[0102] Q=X′ t-T:t _conv[:,:,D:] (15)
[0103] Where: P represents the first D feature dimensions extracted from the feature tensor after convolution; Q represents the last D feature dimensions extracted from the feature tensor after convolution; X′ t-T:t _conv[:,:,:D] represents the first D feature dimensions of the convolution output tensor; X′ t-T:t _conv[:,:,D:] represents the last D feature dimensions of the convolution output tensor;
[0104] Specifically, {P, Q}∈R T×N×D Represents X′ t-T:t _conv has two parts in the feature dimension. The first part P is used to fuse with the SLA output features to make up for the deficiency of causal convolution in capturing only local features. The second part Q is used to generate the gating signal.
[0105] Specifically, the convolution operation in the causal convolution module uses a zero-padding method to keep the time step dimension unchanged. Since the traffic flow at time t+1 needs to be inferred, the convolution operation does not use zero-padding, but directly performs a one-dimensional convolution on the time step dimension. Specifically, 2C0 convolution kernels are used. Input along the time dimension Perform one-dimensional convolution to obtain Then, the output is calculated using formula (16) Where C0 represents the total number of convolution kernels; Γ represents the convolution kernel; [P|Q] represents the feature tensor after convolution; l1 represents the first layer in the spatiotemporal convolution block of the improved spatiotemporal encoder; represents the output of the first layer in the spatiotemporal convolutional block of the improved spatiotemporal encoder; K t Represents the size of the convolution kernel of the causal convolution module;
[0106] Preferably, the formula used to obtain historical traffic flow data for extracting the time characteristics of traffic flow through the gate control module is as follows:
[0107] Specifically, the gating mechanism aims to dynamically adjust the importance of features through the Sigmoid nonlinear activation function, control the flow of causal convolution and SLA fusion results, and reduce the interference of invalid information. The specific calculation formula for splitting and fusion is as follows:
[0108]
[0109] Where: σ is the Sigmoid activation function, which acts as a gating unit to dynamically control the flow of features, adjust the importance of information, and avoid interference from useless information; ⊙ is the Hadamard product operation, which effectively combines global and local features to improve feature expression ability and prediction accuracy; To extract the historical traffic flow data of the temporal characteristics of traffic flow, that is, the historical traffic flow data output by the SLA gated causal convolution module.
[0110] Specifically, the SLA (Simplified Linear Attention) gated causal convolution module in this embodiment is a key component in capturing time series features in traffic flow prediction models. Its advantage lies in its strict adherence to causality, ensuring that the model relies solely on historical data without leaking future information. This feature is particularly critical in real-world scenarios. For example, during rush hour, causal convolution can accurately extract past temporal patterns, providing the model with realistic time-dependent features, thereby improving understanding of complex traffic flow changes and supporting more accurate traffic flow prediction and congestion relief.
[0111] The SLA gated causal convolution module introduces simplified linear attention (SLA), improving its ability to capture both global and local features. Compared to traditional linear attention, SLA offers higher computational efficiency and provides a multi-head attention mechanism. This not only enables efficient modeling of long-range temporal dependencies, but also enhances local feature extraction by integrating depthwise separable convolution (DWC). This design enables the model to account for both local patterns in traffic flow and global trends over long time spans, providing more comprehensive and detailed feature support for complex traffic scenarios.
[0112] S4: Based on the adaptive graph augmentation module, the traffic network graph after granularity adjustment and the historical traffic flow data with extracted temporal characteristics of traffic flow are used. Obtaining the augmented traffic graph structure and augmented traffic flow data;
[0113] The adaptive graph augmentation module of this embodiment calculates the regional traffic pattern and aims to calculate the correlation between the regional embedding vector and the overall regional traffic pattern at each moment, mask the vectors with low correlation, remove the connections with low correlation to reduce noise, and add strongly dependent but unconnected areas, dynamically optimize the regional graph structure and enhance the useful feature information in the traffic flow, thereby improving the model's ability to model complex traffic patterns. The input of the adaptive graph augmentation module is the traffic network graph G′=(V′,E′,A′) after granularity adjustment and the historical traffic flow data for extracting the temporal characteristics of traffic flow. By adjusting the edge set E', adjacency matrix A' and traffic flow data of the traffic flow graph G' after granularity adjustment Perform adaptive augmentation and output the augmented traffic map structure and augmented traffic flow data This process can effectively eliminate the temporal and spatial heterogeneity of the input data and optimize the representation ability of the model. Represents the structure of the augmented traffic graph; represents the augmented edge set; represents the augmented adjacency matrix; represents the augmented traffic flow data; represents the traffic flow data of the entire city at time tT after augmentation; Represents the traffic flow data of the entire city at time t after augmentation; Specifically, the adaptive graph augmentation module adaptively enhances the traffic flow data and graph topology data based on the learned heterogeneous regional dependency pairs, thereby removing noise and enhancing useful information. The calculation process of the adaptive graph augmentation module is as follows: Figure 6 shown.
[0114] Preferably, the step S4 includes: obtaining the nth grid unit rn The probability that the embedding vector is masked to obtain the augmented traffic flow data; where, at any time τnth grid unit r n Transportation mode The calculation formula is as follows:
[0115]
[0116] in,
[0117]
[0118] Where: Represents the nth grid cell r n Traffic pattern at time τ; is the nth grid cell r n The embedding vector of The nth grid cell r in n ;ω represents the learnable parameter vector;τ represents the specific moment in the time window, τ∈(tT:t); represents the historical traffic flow data at time τ for extracting the temporal characteristics of traffic flow; Trans(·) represents the matrix transposition operation; is the traffic flow of the node in the original traffic network graph at time τ; v is the index of the node in the original traffic network graph; S n represents the set of nodes in the nth grid cell in the transportation network graph after granularity adjustment; is the Nth grid cell r N Embedding vector of
[0119] in, is the traffic flow information of N grid cells at any time, that is, The flow data of the τth time slice in . With the learnable parameter vector ω∈R D Multiplying them gives the nth grid unit r at time τ n Transportation mode The larger the value, the larger the nth grid unit r n The traffic pattern at time τ is more correlated with the overall regional traffic pattern.
[0120] like Figure 6 As shown, calculated Then, the historical traffic flow data output by the SLA gated causal convolution module is Perform augmentation to obtain the nth grid unit r n The probability that traffic flow data is masked at time τ as follows:
[0121]
[0122] Where: Represents the nth grid cell r at time τ n The probability that the embedding vector of is masked; Nern(·) represents the Bernoulli distribution;
[0123] Among them, the higher Represents the nth grid cell r n In the nth grid cell r n Embedding vector of The lower the correlation with the overall traffic flow pattern.
[0124] Specifically, this embodiment uses the augmentation parameter γ As the probability of the roulette algorithm, the traffic flow data with a ratio of γ is selected Set to zero to generate augmented traffic data in Represents the n1th grid cell Traffic flow at time τ1; τ1 is the time in the traffic flow data when the γ ratio is set to zero.
[0125] Preferably, the step S4 further includes: obtaining the mth grid unit r m and the nth grid cell r n The overall traffic pattern similarity is used to obtain the augmented traffic graph structure; m and n are the indexes of the grid cells.
[0126] Preferably, get the mth grid unit r m and the nth grid cell r n The formula used for the overall traffic mode similarity is as follows:
[0127] First, get the nth grid cell r n The overall traffic pattern u at time (tT:t) n ∈R D The calculation formula is as follows:
[0128]
[0129] Among them, u n Represents the nth grid cell r n The overall traffic pattern in the time period (tT:t); τ represents the specific moment in the time window; t represents the current time step; Represents the nth grid cell r n Traffic pattern at time τ; is the nth grid cell r nThe embedding vector of The nth grid cell r in n Traffic characteristic vector; T represents the total number of historical time steps;
[0130] Specifically, and embedding Multiplying and summing can give the area r n The overall traffic pattern u at time (tT:t) n ∈R D .
[0131] like Figure 6 As shown, calculate u n After that, the traffic network graph G′=(V′, E′, A′) after granularity adjustment is augmented. The process is as follows:
[0132] The formula used to calculate the similarity between two regions is as follows:
[0133]
[0134] Where: q m,n Represents the mth grid cell r m and the nth grid cell r n The similarity of traffic flow patterns between m Represents the mth grid cell r m The overall traffic pattern in the time period (tT:t); ||·|| represents the L2 norm;
[0135] where q m,n The larger the ∈R, the greater the area r m and r n The higher the similarity of traffic flow patterns between regions, the lower the degree of heterogeneity. m and r n The probability that the connecting edge is masked, that is, the probability that the adjacency matrix A′[m][n] is set to 0, is calculated as follows:
[0136] ρ m,n = Bern(1-q m,n ) (twenty one)
[0137] Where: ρ m,n Represents the mth grid cell r m and the nth grid cell r n The overall traffic pattern similarity of
[0138] Specifically, higher ρ m,n ∈R represents region r m and r n The overall traffic patterns are more similar.
[0139] In this embodiment, based on the augmentation parameter γ, ρ is used m,n As the probability of the roulette algorithm, select γ proportions of edges from the edge set E′ to perform insertion and deletion operations. For the spatially adjacent region r m and r n , if the flow regularity is low, then the edge in the edge set E′ (r m ,r n ) will be masked and its adjacency matrix A′[m][n]=0 will be updated. For non-adjacent regions, if their heterogeneity degree q m,n is lower, then a new edge (r m ,r n ), and update its adjacency matrix A′[m][n]=1. Generate augmented graph topology data
[0140] S5: Based on the improved spatiotemporal encoder, the traffic flow characteristics of the future time step are obtained according to the augmented traffic map structure and augmented traffic flow data.
[0141] Specifically, this embodiment uses an improved space-time encoder, such as Figure 7 Part (a) of the augmented traffic graph structure and augmented traffic flow data Extract spatiotemporal features to obtain To model complex temporal and spatial dependencies. At the same time, the newly added channel attention module (such as Figure 7 The SENet module in part (c) further enhances the network’s expressive power and helps solve the problem of weakened correlation between traffic flows in long-distance regions.
[0142] The improved spatiotemporal encoder consists of two spatiotemporal convolutional blocks and a causal convolutional layer. Figure 7 As shown in part (b), the spatiotemporal convolution block, which is the core of the network structure, consists of two causal convolution modules and a spatial graph convolution layer in the middle. On this basis, this embodiment introduces the SENet network to strengthen the fusion of global and local features, enhance the bottom of the "sandwich" structure, and effectively improve the feature extraction and expression capabilities.
[0143] Specifically, the input of the improved spatiotemporal encoder of this embodiment is the augmented traffic graph structure and augmented traffic flow data Through spatiotemporal modeling, the module outputs the traffic flow characteristics at time t+1 Combining the causal convolution module with the graph convolutional network, this module can capture the dynamic change patterns of time series and the geographical correlations between regions, enhancing the comprehensive expression ability of traffic characteristics.
[0144] The spatial graph convolutional network is performed on the graph at each time step, for the flow input at time τ Chebyshev graph convolution (ChebGCN) is used for processing and the output The formula is as follows:
[0145]
[0146] Where: represents the output of the second layer in the spatiotemporal convolution block of the improved spatiotemporal encoder; ChebGCN(·) represents Chebyshev graph convolution; represents the output of the first layer in the spatiotemporal convolution block of the improved spatiotemporal encoder; i′ represents the index number of the order of the Chebyshev polynomial; K represents the order of the Chebyshev polynomial, which can flexibly capture the features of different order neighborhoods; θ i′ represents the convolution kernel, θ i ∈R K×D ; represents Chebyshev polynomials; represents the optimized normalized Laplace matrix, where λ max is the maximum eigenvalue of matrix L; L is the normalized Laplace matrix, I N x represents the N-th order identity matrix; A is the adjacency matrix of the traffic network topology, which is used to improve numerical stability; and D′ is the degree matrix. This formula approximates the graph filter in an efficient polynomial manner, improving computational efficiency and capturing local and global dependencies between graph nodes.
[0147] Chebyshev polynomials Defined as:
[0148]
[0149] For full time step traffic flow data, the input is Use spatial graph convolution to integrate spatial information and build a more accurate and reliable prediction model to obtain the output represents the output of the first layer in the spatiotemporal convolutional block of the improved spatiotemporal encoder; represents the output of the second layer in the spatiotemporal convolution block of the improved spatiotemporal encoder; K t Represents the size of the convolution kernel of the causal convolution module; C0 represents the total number of convolution kernels;
[0150] The SENet structure is as follows Figure 7 As shown in part (c) of the calculation process, the calculation process is as follows:
[0151] enter The output of the third layer in the spatiotemporal convolution block of the improved spatiotemporal encoder is first compressed by the pooling operation. Z represents the pooled vector of the third layer output in the spatiotemporal convolution block of the improved spatiotemporal encoder to extract the global information of the feature channel. Then, the ReLU and Sigmoid activation functions are used to perform nonlinear activation on Z to generate channel weights. To determine the importance of each feature channel. Finally, the expanded channel weight S is used to Re-evaluate the channel weights to get the output The input features are recalibrated, key channel features are strengthened and redundant information is suppressed. The result after weighted adjustment of the input is shown.
[0152] The inference layer of this embodiment adopts a causal convolution module to extract features and make predictions from the output data generated by the spatiotemporal convolution block.
[0153] Its input is the features generated after being processed by two spatiotemporal convolution blocks Through the convolution kernel Perform a one-dimensional convolution operation on the time dimension, where the number of convolution kernels is D, so as to extract the features in the time series and finally obtain the output result Represents the prediction of the features at time t+1.
[0154] S6: Based on the multi-layer perceptron module with integrated efficient local attention, the traffic flow characteristics of the future time step are obtained by integrating the time and regional attention enhancement mechanism according to the traffic flow characteristics of the future time step to predict the traffic flow in the future time step.
[0155] Specifically, in the ELA (Efficient Local Attention) + MLP module, the local attention mechanism is used to Process and extract detailed information to obtain prediction results This module focuses on short-term features, preventing the model from neglecting detailed information, thereby improving the accuracy of model predictions. This module uses the ELA module and the MLP network in series, using the ELA module's output as the input to the MLP module. This enhances the model's utilization of spatial information, improves its generalization capabilities, and helps the model mitigate the dynamic and non-stationary nature of traffic flow.
[0156] The multi-layer perceptron (MLP) module with integrated efficient local attention (ELA) in this embodiment integrates efficient local attention (ELA) with a multi-layer perceptron (MLP) module to improve the accuracy and efficiency of traffic flow prediction. Traffic flow data often exhibits significant local variations, such as surges during peak hours or sudden traffic flows within specific areas.
[0157] The ELA module focuses on capturing these local dynamic characteristics and broader global patterns. Through its efficient attention mechanism, it can quickly locate and process key information points, while effectively reducing computational complexity and reducing the model's sensitivity to data fluctuations.
[0158] At the same time, the MLP module, as one of the basic architectures of deep learning, plays an indispensable role in this integration. Through multi-level nonlinear transformations, it enhances the learning ability of input features. It not only further refines the local and global features extracted by the ELA module, but also maps these features into higher-level abstract representations.
[0159] The combination of ELA and MLP achieves a perfect fusion of local feature capture and global pattern understanding, improving the model's adaptability and robustness while ensuring efficient and stable computational performance. This design leverages the strengths of both modules, ensuring attention to local details while enhancing understanding of overall patterns, providing a more accurate and reliable solution for traffic flow forecasting.
[0160] The input of ELA is the traffic flow characteristics of the future time step (t+1 moment) output by the improved spatiotemporal encoder Through local attention processing, the module outputs the final predicted traffic The calculation process of this module is as follows Figure 8 As shown. First, the input vector Average pooling is performed on the time dimension and the region dimension to obtain E1∈R 1×N×D and E2∈R T×1×D , which extracts a global feature summary, capturing the overall trend and local pattern in the input feature map in the time dimension and the regional dimension, respectively. Convolution result in time dimension; E2 represents the convolution of Convolution results in the region dimension;
[0161] Then perform one-dimensional convolution to learn the time series pattern and spatial correlation, and obtain E′1∈R 1×N×D and E′2∈R T×1×D, which is used to help the model understand the changing patterns at different time points or between regions. E1′ represents the result of one-dimensional convolution on E1′; E2′ represents the result of one-dimensional convolution on E2′.
[0162] Then use GroupNorm and Sigmoid activation function to map the features to the range of [0,1], and get E″1∈R 1 ×N×D and E″2∈R T×1×D , used to measure the importance of each position. Among them, GroupNorm represents the group normalization operation; finally, the input The product of E″1 and E″2 is passed through the multilayer perceptron to obtain the output. The calculation formula is as follows:
[0163]
[0164] Where: MLP is a multi-layer perceptron; E″2 represents the importance weight at the time step level, which is used to measure the contribution of each time step; E1″ represents the importance weight at the regional level, which is used to measure the importance of each regional unit; The traffic flow features of future time steps represented by the fused temporal and regional attention enhancement mechanism are combined with the information of the input itself so that the final result can preserve local and global patterns at the same time.
[0165] This embodiment proposes a method for solving the problem of traffic time series prediction. Through experimental verification of four public data sets, the results show that the evaluation indicators of the proposed model on key data sets are reduced by an average of 1.62%-2.78% compared with the current optimal model. The grid division method is used to summarize the overall traffic flow in the region, solve the problem of data fragmentation, and significantly improve the efficiency of traffic monitoring. The SLA (gated causal convolution) module proposed in this embodiment combines the multi-head attention mechanism to efficiently model long-distance time dependencies, and at the same time enhances the ability to extract local features through deep separable convolution, thereby taking into account the expression capabilities of global and local time patterns.
[0166] The improved spatiotemporal encoder proposed in this embodiment enhances the bottom of the "sandwich structure" in the traditional spatiotemporal convolutional layer, dynamically adjusts the feature weights, strengthens the expression of key features, and effectively suppresses redundant information. This design further improves the expressiveness and prediction accuracy of the model. The ELA local attention module of this embodiment reduces computational complexity and sensitivity to data fluctuations while capturing local and global features, further improving the model's ability to handle dynamic and non-stationary problems. The model proposed in this embodiment demonstrates excellent robustness and effectiveness in traffic flow prediction tasks.
[0167] The traffic flow prediction method based on the enhanced spatio-temporal graph convolutional network in this embodiment establishes a traffic flow prediction model - Enhanced Spatio-Temporal Graph Convolutional Network (ESTGCN). By introducing a granularity adjustment module, ESTGCN enables the model to flexibly adjust the degree of refinement of regional division according to demand to adapt to application requirements in different scenarios; at the same time, a simplified linear attention mechanism is integrated into the causal convolution module, taking into account the local pattern of traffic flow and the global trend of long-term span, thereby improving the modeling ability of complex traffic flow dynamic changes; using the traffic flow and graph topology levels to adaptively augment the traffic flow graph data, dynamically optimize the traffic graph structure, and enhance the model's understanding of the spatial correlation of traffic flow; combined with the enhanced spatio-temporal encoder, it can comprehensively infer the complex changing trends of traffic flow and improve the prediction accuracy; the integrated local attention module effectively alleviates the modeling difficulties caused by the dynamic and non-stationary nature of traffic flow, and enhances the robustness and generalization ability of the model. The ESTGCN model of this embodiment aims to solve the problems existing in existing technologies, such as the dynamics and non-stationarity of traffic flow and the difficulty in capturing correlations between long-distance regions, through its unique architectural design. It is deepened and expanded based on the spatiotemporal encoder and uses known road network structure information to predict traffic flow.
[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A traffic flow prediction method based on enhanced spatiotemporal graph convolutional network, characterized by: The steps include: S1: Obtain the historical traffic flow information of the road section to be analyzed for T time steps through traffic sensor information; S2: Based on the granularity adjustment module, according to the preset traffic network diagram and grid division parameters, grid cells in the traffic network diagram after granularity adjustment are obtained, so as to obtain the granularity-adjusted historical traffic flow data based on the traffic flow information of the historical T time steps; S3: Based on the SLA gated causal convolution module, the historical traffic flow data after granularity adjustment is obtained to extract the temporal characteristics of the traffic flow; S4: Based on the adaptive graph augmentation module, the augmented traffic graph structure and augmented traffic flow data are obtained according to the traffic network graph after granularity adjustment and the historical traffic flow data with the time characteristics of traffic flow extracted; S5: Based on the improved spatiotemporal encoder, the traffic flow characteristics of the future time step are obtained according to the augmented traffic map structure and augmented traffic flow data; S6: Based on the multi-layer perceptron module with integrated efficient local attention, the traffic flow characteristics of the future time step are obtained by integrating the time and regional attention enhancement mechanism according to the traffic flow characteristics of the future time step to predict the traffic flow in the future time step.
2. The traffic flow prediction method based on enhanced spatiotemporal graph convolutional network according to claim 1 is characterized in that: In S2, the formula used to obtain the historical traffic flow data after granularity adjustment is as follows: N=α×β X′ t-T:t ={X′ t-T ,X′ t-T+1 ,…,X′ t } Where: Represents the nth grid unit r in the traffic network graph after granularity adjustment n The total traffic flow at time t; n is the index of the grid in the traffic network graph after granularity adjustment; t represents the current time step; is the traffic flow of the node in the original traffic network graph at time t; v is the index of the node in the original traffic network graph; S n represents the set of nodes in the nth grid unit in the traffic network graph after granularity adjustment; N is the total number of grid units in the traffic network graph after granularity adjustment; X′ t represents the total traffic flow data at time t in the traffic network graph after granularity adjustment; X′ t-T:t represents the historical traffic flow data after granularity adjustment; X′ t-T represents the city-wide traffic flow data at time tT after granularity adjustment; T represents the total number of historical time steps; α and β are both grid division parameters.
3. The traffic flow prediction method based on enhanced spatiotemporal graph convolutional network according to claim 2 is characterized in that: In S3, the SLA gated causal convolution module includes an SLA module, a causal convolution module and a gating module; The SLA module is used to obtain the embedded vector processed by the SLA layer based on the historical traffic flow data after granularity adjustment; The causal convolution module is used to obtain feature dimensions extracted from the convolved feature tensor based on the historical traffic flow data after granularity adjustment; The gating module is used to obtain historical traffic flow data for extracting the temporal characteristics of traffic flow based on the embedding vector processed by the SLA layer and the feature dimensions extracted from the feature tensor after convolution.
4. The traffic flow prediction method based on enhanced spatiotemporal graph convolutional network according to claim 3 is characterized in that: The method to obtain the embedded vector after SLA layer processing is as follows: S31: Obtain the traffic flow data tensor with added location coding information. The formula used is as follows: B1=Linear(X′ t-T:t ) B2=B1+Positional_Encoding Where: X′ t-T:t represents the historical traffic flow data after granularity adjustment; Linear(·) represents the fully connected layer; B1 represents the traffic flow data tensor output by the fully connected layer; B2 represents the traffic flow data tensor with position encoding information added; Positional_Encoding represents the self-learned position information encoding; S32: Rearrange the traffic flow data tensor with the added position coding information to obtain a rearranged input feature tensor; S33: Obtain a query vector, a key vector, and a value vector according to the rearranged input feature tensor to obtain a query vector processed by the multi-head attention mechanism, a key vector processed by the multi-head attention mechanism, and a value vector processed by the multi-head attention mechanism; First, get the query vector, key vector, and value vector: Where: Query represents the query vector; Key represents the key vector; Value represents the value vector; B′2 represents the rearranged input feature tensor, and represents the flattened time-node feature sequence; W Q Represents the projection weight matrix of the query vector; W K Represents the projection weight matrix of the key vector; W V represents the projection weight matrix of the value vector; b q represents the bias term of the query vector; b k represents the bias term of the key vector; b v represents the bias term of the value vector; B′2 represents the rearranged input feature tensor; Next, we get the query vector and key vector after ReLU activation: Where: Represents the query vector after being processed by the ReLU activation function; represents the key vector after ReLU activation function processing; δ(·) represents the ReLU activation operation; Finally, get the query vector processed by the multi-head attention mechanism, the key vector processed by the multi-head attention mechanism, and the value vector processed by the multi-head attention mechanism: Where: Represents the query vector after multi-head attention mechanism processing; represents the key vector after multi-head attention mechanism processing; Value′ represents the value vector after multi-head attention mechanism processing; Multi_Head represents the multi-head attention mechanism operation; S34: According to the query vector processed by the multi-head attention mechanism, the key vector processed by the multi-head attention mechanism, and the value vector processed by the multi-head attention mechanism, obtain the output matrix calculated by the efficient attention mechanism and the feature matrix processed by the depthwise separable convolution; The formula used to obtain the output matrix calculated by the efficient attention mechanism is as follows: Where: B3 represents the output matrix calculated by the efficient attention mechanism, which is the output attention feature; Efficient_Att represents the efficient attention mechanism; i represents the index in the query vector; represents the i-th query vector after multi-head attention processing; represents the j-th key vector after multi-head attention processing; Trans(·) represents the matrix transposition operation; Value′ j represents the j-th value vector after multi-head attention processing; j represents the index in the value vector / key vector; TN represents the total length after flattening the time step and node dimensions; N is the total number of grid cells in the traffic network graph after granularity adjustment; T is the total number of historical time steps; The formula used to obtain the feature matrix after depth-wise separable convolution is as follows: B4=DWC(Value′) Where: B4 represents the feature matrix after depth-wise separable convolution; DWC(·) represents depth-wise separable convolution; S35: Based on the output matrix calculated by the efficient attention mechanism and the feature matrix after depth-wise separable convolution, the embedding vector after SLA layer processing is obtained. The formula used is as follows: Where: SoftMax represents the normalization function; Represents the embedded vector after processing by the SLA layer, that is, the historical traffic flow data output by the SLA module.
5. The traffic flow prediction method based on enhanced spatiotemporal graph convolutional network according to claim 3 is characterized in that: Get the feature dimensions extracted from the convolved feature tensor as follows; X′ t-T:t _zero=ZeroFill(X′ t-T:t ) X′ t-T:t _conv=2Dconv(X′ t-T:t ) Where: represents the embedding vector after zero padding; where K t Represents the size of the convolution kernel of the causal convolution module; N represents the total number of grids in the traffic network graph after granularity adjustment; D represents the feature dimension after processing by the SLA gated causal convolution module; Represents the shape of a three-dimensional tensor; X′ t-T:t _conv∈R T ×N×2D represents the embedding vector after causal convolution processing; 2Dconv(·) represents a two-dimensional convolution operation, and the size of the convolution kernel is K t , used to extract local features; ZeroFill(·) represents the zero filling operation; The embedding vector X′ after causal convolution t-T:t _conv is split to get: P=X′ t-T:t _conv[:,:,:D] Q=X′ t-T:t _conv[:,:,D:] Where: P represents the first D feature dimensions extracted from the feature tensor after convolution; Q represents the last D feature dimensions extracted from the feature tensor after convolution; X′ t-T:t _conv[:,:,:D] represents the first D feature dimensions of the convolution output tensor; X′ t-T:t _conv[:,:,D:] represents the last D feature dimensions of the convolution output tensor.
6. The traffic flow prediction method based on enhanced spatiotemporal graph convolutional network according to claim 5 is characterized in that: The formula used to obtain historical traffic flow data for extracting the time characteristics of traffic flow is as follows: Where: σ is the Sigmoid activation function; ⊙ is the Hadamard product operation; To extract the historical traffic flow data of the temporal characteristics of traffic flow, that is, the historical traffic flow data output by the SLA gated causal convolution module.
7. The traffic flow prediction method based on enhanced spatiotemporal graph convolutional network according to claim 6 is characterized in that: Said S4 includes: Get the nth grid cell r n The probability that the embedding vector of is masked to obtain the augmented traffic flow data; Get the mth grid cell r m and the nth grid cell r n The overall traffic pattern similarity is used to obtain the augmented traffic graph structure; m and n are the indexes of the grid cells.
8. The traffic flow prediction method based on enhanced spatiotemporal graph convolutional network according to claim 7 is characterized in that: Get the nth grid cell r n The probability that the embedding vector of a grid cell is masked is given by the following formula: in, Where: Represents the nth grid cell r n Traffic pattern at time τ; is the nth grid cell r n The embedding vector of The nth grid cell r in n ;ω represents the learnable parameter vector;τ represents the specific moment in the time window, τ∈(tT:t); represents the historical traffic flow data at time τ for extracting the temporal characteristics of traffic flow; Trans(·) represents the matrix transposition operation; is the traffic flow of the node in the original traffic network graph at time τ; v is the index of the node in the original traffic network graph; S n represents the set of nodes in the nth grid cell in the transportation network graph after granularity adjustment; is the Nth grid cell r N Embedding vector of Then we get the nth grid unit r n The probability that traffic flow data is masked at time τ as follows: Where: Represents the nth grid cell r at time τ n The probability that the embedding vector of is masked; Bern(·) represents the Bernoulli distribution.
9. The traffic flow prediction method based on enhanced spatiotemporal graph convolutional network according to claim 7 is characterized in that: Get the mth grid cell r m and the nth grid cell r n The formula used for the overall traffic mode similarity is as follows: Among them, u n Represents the nth grid cell r n The overall traffic pattern in the time period (tT:t); τ represents the specific moment in the time window; t represents the current time step; Represents the nth grid cell r n Traffic pattern at time τ; is the nth grid cell r n The embedding vector of The nth grid cell r in n Traffic characteristic vector; T represents the total number of historical time steps; Where: q m,n Represents the mth grid cell r m and the nth grid cell r n The similarity of traffic flow patterns between m Represents the mth grid cell r m The overall traffic pattern during the period (tT:t); u n Represents the nth grid cell r n The overall traffic pattern in the time period (tT:t); ||·|| represents the L2 norm; ρ m,n =Bern(1-q m,n ) Where: ρ m,n Represents the mth grid cell r m and the nth grid cell r n The similarity of the overall traffic pattern of ; Bern(·) represents the Bernoulli distribution.
10. The traffic flow prediction method based on enhanced spatiotemporal graph convolutional network according to claim 9 is characterized in that: In S6, the traffic flow characteristics of the future time step of the attention enhancement mechanism of the fusion time and region are obtained, Where: MLP is a multi-layer perceptron; E″2 represents the importance weight of the time step; E″1 represents the importance weight at the regional level; Traffic flow characteristics at future time steps representing the fusion of temporal and regional attention enhancement mechanisms.