Traffic prediction method based on expansion convolution strategy and diffusion diagram attention

By adopting the expansion convolution strategy and diffusion graph attention mechanism in the traffic prediction model, the space-time dependence in traffic data is solved, and the problem of information loss when handling abnormal points is achieved is achieved higher prediction accuracy.

CN120163277APending Publication Date: 2025-06-17ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510169393.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing traffic prediction models are difficult to effectively capture the complex spatiotemporal dependencies in traffic data, especially when anomalies appear, which may lead to information loss and reduced prediction accuracy.

Method used

The traffic prediction method based on expansion convolution strategy and diffusion graph attention is adopted, and the traffic data is mapped into high-dimensional spatiotemporal features through the spatiotemporal perception framework. The time sampling convolution module and diffusion graph attention module are used to capture the dependence between time and space respectively to reduce the impact of abnormal points on global information.

Benefits of technology

Improve the accuracy of traffic prediction, and achieve state-of-the-art performance on multiple public data sets by reducing the impact of anomalies and effectively modeling spatiotemporal correlations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163277A_ABST
    Figure CN120163277A_ABST
Patent Text Reader

Abstract

The invention relates to the field of traffic prediction, in particular to a traffic prediction method based on an expansion convolution strategy and diffusion diagram attention, and the method comprises the steps: S1, providing a space-time perception framework based on an expansion convolution idea; s2, constructing a space-time perception model used for capturing a space-time dependency relationship in the framework, wherein the space-time perception model comprises a data embedding module used for mapping original traffic data into high-dimensional space-time characteristics; the time sampling convolution module is used for modeling time correlation in the traffic data; a diffusion map attention module for modeling spatial correlation in the traffic sequence; and S3, inputting the spatio-temporal characteristics output by the spatio-temporal perception model into a decoder, and outputting a prediction result. The method has the beneficial effects that the input sequence is divided into a plurality of subsequences by utilizing the idea of expansion convolution, so that the influence of abnormal points on global information is reduced; and the time sampling convolution module and the diffusion diagram attention module are respectively used for capturing a dependency relationship between time and space and a spatial correlation relationship in a modeling sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of traffic prediction, and particularly to a traffic prediction method based on an extended convolution strategy and diffusion graph attention. Background Art

[0002] With the rapid development of intelligent transportation systems, how to accurately predict traffic has become a key issue in this field. The main challenge lies in effectively capturing the complex spatio-temporal dependencies in traffic data. In recent years, various neural network models have been designed to solve this problem, and most models rely on the dependencies between time steps to learn the hidden features in traffic sequences. However, when there are abnormal changes in traffic data at a certain time step, the specific mechanisms of these models will cause the abnormal time step to have a negative impact on the capture of global information. The Temporal Convolutional Network (TCN) can expand the receptive field through dilated sampling to capture long-term dependencies, but this may lead to information sparsity and information loss.

[0003] Existing research focuses on the following two issues: (i) how to reduce the impact of abnormal points in traffic sequences on global information; (ii) how to model the spatio-temporal correlations in sequences. Summary of the Invention

[0004] In order to overcome the above deficiencies, the purpose of the present invention is to provide a traffic prediction method based on an extended convolution strategy and diffusion graph attention, which can reduce the impact of abnormal points on global information while modeling the spatio-temporal correlations in traffic sequences to improve the accuracy of traffic prediction.

[0005] The present invention achieves the above purpose through the following solutions: A traffic prediction method based on an extended convolution strategy and diffusion graph attention, comprising the following steps:

[0006] S1: Propose a spatio-temporal perception framework based on the idea of extended convolution;

[0007] S2: Build a spatio-temporal perception model for capturing spatio-temporal dependencies within the framework, including: a data embedding module for mapping original traffic data into high-dimensional spatio-temporal features;

[0008] a temporal sampling convolution module for modeling the temporal correlations in traffic data;

[0009] a diffusion graph attention module for modeling the spatial correlations in traffic sequences;

[0010] S3: Input the spatio-temporal features output by the spatio-temporal perception model into a decoder to output a prediction result.

[0011] Preferably, the step S1 specifically includes the following steps:

[0012] S1-1: Set the dilation step to 4, and divide the input sequence into subsequences with a time length of 3 by interval sampling;

[0013] S1-2: In each layer, select subsequences by halving the dilation step and reassemble them according to the time steps.

[0014] Preferably, the specific steps of the data embedding module in S2 include:

[0015] S2-11: Extract the original input features of the original traffic data, and the original input features are processed through a fully connected layer to obtain high-dimensional features where N is the number of nodes, d1 is the number of feature channels, T is the number of time steps, and R is the set of real number matrices with the shape of (d1, N, T);

[0016] S2-12: Select two time periods, daily and weekly, and generate corresponding specific embedding vectors T day ∈R N ×T ,T week ∈R N×T ; S2-13: To ensure that the embedding vector can be dynamically updated according to real-time traffic data, two learnable time embeddings are defined: W day ∈R N×f ,W week ∈R N×7 , where f represents the number of time stamps in a day;

[0017] S2-14: Extract the corresponding weekly embedding and daily timestamp embedding from the traffic data, where d2 is the number of feature channels; by concatenating and broadcasting them, obtain the periodic embedding of the traffic time series: E t =E d +E w ;

[0018] S2-15: Concatenate the high-dimensional features and the periodic embedding to obtain the output of the hidden spatio-temporal representation: Z = E data ||E t .

[0019] Preferably, f = 288.

[0020] Preferably, the specific implementation steps of the time convolutional sampling module in step S2 include: The data embedding module inputs the output of the hidden spatio-temporal representation into the time convolutional sampling module, and through the convolutional operations of two groups of 2DCNNs, obtain the dynamic time dependence H' within the subsequence t= Conv2D(Conv2D(Padding(Z))), where padding(Z) refers to the padding applied to Z, and Conv2D represents the 2D convolution operation applied to the padded result.

[0021] Preferably, the specific implementation steps of the diffusion graph attention module in step S2 include:

[0022] S2-21: Introduce the adaptive spatial embedding E s , identify the potential relationships between nodes and model the nodes: H f = Sum(H′ t , dim = -1), where exp(·) represents the exponential function;

[0023] S2-22: In traffic data, since traffic patterns change over time; to capture these dynamic relationships, establish a dynamic relationship matrix: S2-23: Use the graph attention mechanism to enhance the information transmission between nodes; at the same time, to adjust the attention scores in the diffusion graph attention module, introduce a spatio-temporal relationship graph with dynamic correlation in the graph attention:

[0024] where W f ∈R d×d is used to map the features, FC(·) represents the fully connected layer, W att ∈R N×T and W dyn ∈R N×T are learnable parameters, ⊙ is the Hadamard product, LeakyReLU(·) is the activation function, and exp(·) represents the exponential function;

[0025] S2-24: Construct a mask by indexing the most relevant neighbors according to the maximum correlation of the nodes: where Γ represents the maximum number of neighbors of the nodes;

[0026] S2-25: Capture spatial correlation through the diffusion graph attention:

[0027]

[0028] where θ k represent the attention output and learnable parameters of the k-th head respectively, and ReLU(·) is the activation function.

[0029] Preferably, the specific implementation steps of S3 include:

[0030] S3-1: Take the spatio-temporal feature H′ output by step S2 gInput into the gated linear unit to obtain the hidden feature: H hid = (W1H g + b1) ⊙ σ(W2H g + b2),

[0031] where W1 and W2 are weight matrices, b1 and b2 are bias vectors, and σ is a non - linear activation function;

[0032] S3 - 2: Input the hidden feature into the prediction layer to obtain the final output representation: Yout = ReLUU((EC(ReLU(FC(Hhid)))) where ReLU(·) is an activation function.

[0033] A traffic prediction structure based on the dilated convolution strategy and diffusion graph attention for implementing the above - mentioned method, including an encoder and a decoder. The encoder includes a data embedding module and a spatio - temporal attention module. The spatio - temporal attention module includes a temporal sampling convolution module and a diffusion graph attention module. The encoder includes a gated linear layer and a prediction layer; the original traffic data is dilated convolved to obtain several subsequences, and after input into the data embedding module, a hidden spatio - temporal representation is output. The hidden spatio - temporal representation output passes through the temporal sampling convolution module and the diffusion graph attention module in sequence, and the spatio - temporal features are output; the spatio - temporal features are input into the decoder, and through the gated linear layer and the prediction layer, the traffic prediction result is obtained.

[0034] The beneficial effects of the present invention are as follows: Using the idea of dilated convolution to divide the input sequence into multiple subsequences, thereby reducing the impact of outliers on global information; at the same time, in order to model the spatial correlation relationship in the sequence, the present invention designs a temporal sampling convolution module and a diffusion graph attention module to capture temporal and spatial dependencies respectively. In the actual application process, this method achieves state - of - the - art performance in multiple public datasets, verifying its effectiveness and accuracy in traffic prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is the schematic diagram of the step - by - step process of the method of the present invention;

[0036] Figure 2 is the schematic diagram of the structure for implementing the method of the present invention;

[0037] Figure 3 is the schematic diagram of the present invention for modeling the spatio - temporal dependencies of subsequences; where (a) is the schematic diagram of the spatio - temporal feature extraction framework; (b) is the block diagram of the diffusion graph attention module. DETAILED DESCRIPTION OF THE INVENTION

[0038] The present invention will be further described below in conjunction with specific implementation examples, but the protection scope of the present invention is not limited thereto:

[0039] Example: AsFigure 1 As shown in the figure, a traffic prediction method based on the dilated convolution strategy and diffusion graph attention includes the following steps:

[0040] S1: A spatio-temporal perception framework based on the idea of dilated convolution is proposed;

[0041] S2: A spatio-temporal perception model for capturing spatio-temporal dependencies is constructed within the framework, including: a data embedding module for mapping original traffic data into high-dimensional spatio-temporal features;

[0042] a time sampling convolution module for modeling the temporal correlation in traffic data;

[0043] a diffusion graph attention module for modeling the spatial correlation in traffic sequences;

[0044] S3: The spatio-temporal features output by the spatio-temporal perception model are input into the decoder to output the prediction result.

[0045] As Figure 2 shown in the figure, a traffic prediction structure based on the dilated convolution strategy and diffusion graph attention for implementing the above method is realized through an encoder and a decoder. The encoder includes a data embedding module and a spatio-temporal attention module. The spatio-temporal attention module includes a time sampling convolution module and a diffusion graph attention module. The encoder includes a gated linear layer and a prediction layer. The original traffic data is dilated convolved to obtain several subsequences, and after being input into the data embedding module, a hidden spatio-temporal representation is output. The hidden spatio-temporal representation output passes through the time sampling convolution module and the diffusion graph attention module in sequence to output spatio-temporal features. The spatio-temporal features are input into the decoder, and through the gated linear layer and the prediction layer, the traffic prediction result is obtained.

[0046] The three indicators (RMSE, MAE, MAPE (%)) of this method on the PEMS04, PEMS07, PEMS08, PEMS - BAY, and METR - LA datasets reach 29.77 / 18.02 / 11.88, 32.95 / 19.15 / 8.01, 23.16 / 13.32 / 8.74, 3.57 / 1.54 / 3.46,

[0047] 6.11 / 2.96 / 8.21 respectively.

[0048] The following is a more detailed description of each step:

[0049] S1: A spatio-temporal perception framework based on the idea of dilated convolution is proposed. The traffic sequence is divided into multiple subsequences by interval sampling, and the traffic sequence is processed layer by layer using the idea of extended convolution, which enables it to integrate information of different time scales and avoid overfitting to noise or anomalies in consecutive time steps.

[0050] Specifically, it includes the following steps:

[0051] S1-1: Set the dilation step to 4, and divide the input sequence into subsequences with a time length of 3 by interval sampling;

[0052] S1-2: In each layer, select subsequences by halving the dilation step and reassemble them according to the time steps.

[0053] S2: Build a spatio-temporal awareness model for capturing spatio-temporal dependencies within the framework, including:

[0054] (1) A data embedding module for mapping raw traffic data into high-dimensional spatio-temporal features. Specifically:

[0055] S2-11: Extract the original input features of the raw traffic data. The original input features are processed through a fully connected layer to obtain high-dimensional features where N is the number of nodes, d1 is the number of feature channels, T is the number of time steps, and R is the set of real number matrices with the shape of (d1, N, T);

[0056] S2-12: Generate specific embedding vectors for each time period according to different time cycles. In this embodiment, two time cycles of daily and weekly are selected, and corresponding specific embedding vectors T day ∈R N×T ,T week ∈R N×T ;

[0057] S2-13: To ensure that the embedding vectors can be dynamically updated according to real-time traffic data, two learnable time embeddings are defined: W day ∈R N×f ,W week ∈R N×7 where f represents the number of time stamps in a day. Since observations are recorded every 5 minutes, f = 288; S2-14: Extract the corresponding weekly embedding from the traffic data and the daily time stamp embedding where d2 is the number of feature channels; By concatenating and broadcasting them, the periodic embedding of the traffic time series is obtained: E t =E d +E w ;

[0058] S2-15: Concatenate the high-dimensional features and the periodic embedding to obtain the output of the hidden spatio-temporal representation: Z = E data ||E t .

[0059] Build a spatio-temporal graph attention module that can capture spatio-temporal dependencies, and improve the accuracy of traffic prediction by modeling spatio-temporal correlations in traffic sequences. As Figure 3 As shown in (a), the specific steps are as follows: Input the features output by the data embedding module into the spatio-temporal attention module. First, use the temporal sampling convolution module to capture the temporal correlations of traffic subsequences; Next, the features are used to model spatial correlations through the diffusion graph attention module.

[0060] (2) Temporal sampling convolution module for modeling temporal correlations in traffic data: Each temporal sampling convolution module consists of a set of rich convolution filters, which are mainly composed of two groups of 2D CNNs for capturing dynamic temporal dependencies within subsequences: H′ t = Conv2D(Conv2D(Padding(Z))), where Padding(Z) refers to the padding applied to Z, and Conv2D represents the 2D convolution operation applied to the padded result.

[0061] (3) Diffusion graph attention module for modeling spatial correlations in traffic sequences, as Figure 3 shown in (b):

[0062] S2-31: To preserve the traffic patterns of each node, we use adaptive spatial embedding (E s ), which can identify and model the potential relationships between nodes: H f = Sum(H′ t , dim = -1), S2-32: In traffic data, since traffic patterns change over time; To capture these dynamic relationships, a dynamic relationship matrix is established: S2-33: Use the graph attention mechanism to enhance information transmission between nodes; At the same time, to adjust the attention scores in the diffusion graph attention module, a spatio-temporal relationship graph with dynamic correlations is introduced into the graph attention:

[0063] where W f ∈R d×d is used to map the features, FC(·) represents the fully connected layer, W att ∈R N×T and W dyn ∈R N×T are learnable parameters, ⊙ is the Hadamard product, LeakyReLU(·) is the activation function, and exp(·) represents the exponential function;

[0064] S2-34: Construct a mask by indexing the most relevant neighbors according to the maximum correlation of the nodes: where Γ represents the maximum number of neighbors of the nodes;

[0065] S2-35: Finally, spatial correlations are captured through diffusion graph attention to obtain spatio-temporal features

[0066] wherein θ k represent the attention output and learnable parameters of the k-th head respectively.

[0067] S3: The spatio-temporal features output by the spatio-temporal perception model are input into the decoder to output the prediction result. Specifically

[0068] S3-1: The spatio-temporal feature H′ output in step S2 g is input into the gated linear unit. The gated linear unit combines linear transformation with the gating mechanism to enhance the flexibility and expressiveness of the neural network, and obtains the hidden feature: H hid =(W1H g +b1)☉σ(W2H g +b2), where W1 and W2 are weight matrices, b1 and b2 are bias vectors, and σ is a non-linear activation function;

[0069] S3-2: The hidden feature is input into the prediction layer to obtain the final output representation: Y out =ReLU((EC(ReLU(FC(H hid )))) where ReLU(·) is an activation function.

[0070] The above are the specific embodiments of the present invention and the technical principles applied. If changes made according to the concept of the present invention do not exceed the spirit covered by the specification and drawings, they should still fall within the protection scope of the present invention.

Claims

1. Traffic prediction method based on dilated convolution strategy and diffusion graph attention, characterized by The following steps are involved: S1: A spatiotemporal perception framework based on the idea of ​​dilated convolution is proposed; S2: Building a spatiotemporal-aware model within the framework to capture spatiotemporal dependencies, including: A data embedding module for mapping raw traffic data into high-dimensional spatiotemporal features; A temporally sampled convolutional module for modeling temporal correlations in traffic data; Diffusion map attention module for modeling spatial correlations in traffic sequences; S3: Input the spatiotemporal features output by the spatiotemporal perception model into the decoder and output the prediction results.

2. The traffic prediction method based on dilated convolution strategy and diffusion graph attention according to claim 1, characterized in that: The step S1 specifically includes the following steps: S1-1: Set the expansion step to 4, and divide the input sequence into subsequences of time length 3 by interval sampling; S1-2: In each layer, the number of expansion steps is halved to select subsequences, and then reassembled according to the number of time steps.

3. The traffic prediction method based on dilated convolution strategy and diffusion graph attention according to claim 1, characterized in that: The specific steps of the data embedding module in S2 include: S2-11: Extract the original input features of the original traffic data. The original input features are processed by the fully connected layer to obtain high-dimensional features. Where N is the number of nodes, d1 is the number of feature channels, T is the number of time steps, and R is a set of real matrices with shape (d1, N, T); S2-12: Select two time periods, daily and weekly, and generate corresponding specific embedding vectors T day ∈R N×T , T week ∈R N×T ; S2-13: To ensure that the embedding vector can be dynamically updated according to real-time traffic data, two learnable temporal embeddings are defined: W dav ∈R N×f , W week ∈R N×7 , where f represents the number of timestamps in a day; S2-14: Extract corresponding weekly embeddings from traffic data and daily timestamp embed Where d2 is the number of feature channels; by connecting and broadcasting them, we get the periodic embedding of the traffic time series: E t =E d +E w ; S2-15: Connect high-dimensional features and periodic embedding to obtain hidden spatiotemporal representation output: Z = E data ||E t .

4. The traffic prediction method based on dilated convolution strategy and diffusion graph attention according to claim 3 is characterized in that f=288。 5. The traffic prediction method based on dilated convolution strategy and diffusion graph attention according to claim 1, characterized in that: The specific implementation steps of the temporal convolution sampling module in step S2 include: the data embedding module inputs the hidden spatiotemporal representation output into the temporal convolution sampling module, and obtains the dynamic temporal dependency H′ in the subsequence through two groups of 2DCNN convolution operations. t =Conv2D(Conv2D(Padding(Z))), where Padding(Z) refers to the padding applied to Z and Conv2D represents the 2D convolution operation applied to the padded result.

6. The traffic prediction method based on dilated convolution strategy and diffusion graph attention according to claim 1, characterized in that: The specific implementation steps of the diffusion map attention module in step S2 include: S2-21: Introducing Adaptive Spatial Embedding E s , identify nodes and model potential relationships between nodes: H f =Sum(H′ t , dim=-1), Where exp(·) represents the exponential function; S2-22: In traffic data, since traffic patterns change over time, in order to capture these dynamic relationships, a dynamic relationship matrix is ​​established: S2-23: Use the graph attention mechanism to enhance information transmission between nodes; at the same time, in order to adjust the attention score in the diffusion graph attention module, a spatiotemporal relationship graph with dynamic correlation is introduced in the graph attention: Where W f ∈R d×d Used to map features, FC(·) represents the fully connected layer, W att ∈R N×T and W dyn ∈R N×T is a learnable parameter, ⊙ is the Hadamard product, LeakyReLU(·) is the activation function, and exp(·) represents the exponential function; S2-24: Build a mask by indexing the most relevant neighbors based on the maximum relevance of a node: Where Γ represents the maximum number of neighbors of a node; S2-25: Capturing spatial correlations via diffuse map attention: in θ k denote the attention output and learnable parameters of the k-th head respectively, and ReLU(·) is the activation function.

7. The traffic prediction method based on dilated convolution strategy and diffusion graph attention according to claim 1, characterized in that: The specific implementation steps of S3 include: S3-1: The spatiotemporal features H output from step S2 g Input to the gated linear unit to obtain the hidden feature: H hid =(W1H′ g +b1)⊙ο(W2H′ g +b2), Among them, W1 and W2 are weight matrices, b1 and b2 are bias vectors, and σ is a nonlinear activation function; S3-2: Input the hidden features into the prediction layer to obtain the final output representation: Y out =ReLU(FC(ReLU(FC(H hid ))))), where ReLU(·) is the activation function.

8. A traffic prediction structure based on dilated convolution strategy and diffuse graph attention for implementing the above method, comprising an encoder and a decoder, characterized in that: The encoder includes a data embedding module and a spatiotemporal attention module, the spatiotemporal attention module includes a time sampling convolution module and a diffusion map attention module, and the encoder includes a gated linear layer and a prediction layer; the original traffic data is subjected to dilated convolution to obtain a plurality of subsequences, which are input into the data embedding module to obtain a hidden spatiotemporal representation output, which is sequentially passed through the time sampling convolution module and the diffusion map attention module to output spatiotemporal features, and the spatiotemporal features are input into a decoder, and the traffic prediction results are obtained through the gated linear layer and the prediction layer.