Traffic prediction method based on double-dynamic graph attention network

By adopting dual dynamic graph attention network and dual transformation technology in traffic flow prediction, the problem that existing methods are difficult to capture non-paired relationships and long-term spatial and temporal correlations is solved, and more accurate and flexible traffic signal prediction is achieved.

CN120069271APending Publication Date: 2025-05-30DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411915151.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing traffic flow prediction methods are difficult to effectively capture dynamic non-paired relationships and long-term spatial and temporal correlations in road networks, and traditional graph convolutional networks have poor results in spatial dependency learning.

Method used

Using a traffic prediction method based on a dual dynamic graph attention network, the road network structure is dynamically modeled without relying on prior knowledge, and the traffic flow graph is transformed into a hypergraph using dual transformation, the spatial correlation between nodes and edges is captured, and the temporal correlation is extracted through the gated attention linear unit stack.

Benefits of technology

Effectively capture dynamic spatiotemporal characteristics in traffic signals, improve the accuracy of traffic prediction and the memory ability of the model, and better describe the mutual influence between nodes and the dynamic changes of the road network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069271A_ABST
    Figure CN120069271A_ABST
Patent Text Reader

Abstract

The invention provides a traffic prediction method based on a double-dynamic graph attention network, and belongs to the technical field of traffic prediction methods. The method comprises the following steps: adding spatio-temporal information to an input traffic signal, generating a dynamic graph adjacency matrix through linear transformation, and performing dual transformation on a dynamic graph to generate a dual dynamic hypergraph adjacency matrix; inputting a traffic signal to be predicted and the double-dynamic graph adjacency matrix into a spatial feature extraction module, and capturing and integrating spatial correlation features; a time feature extraction module is used to capture time-related features on different time scales through a plurality of stacked gating attention linear units, and time-space related features captured by the current time-space feature extraction module are obtained; and the output module carries out linear processing and residual decomposition on the extracted spatio-temporal correlation features to obtain a prediction result of the current module and signal input of the next block, and integrates output of all spatio-temporal feature extraction modules to obtain a final prediction value. According to the method, a space-time convolutional network architecture is adopted to learn dynamic characteristics in traffic signals, and the traffic prediction precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic prediction methods, and in particular, to a traffic prediction method based on a dual dynamic graph attention network. Background Art

[0002] The rapid growth of urban population has led to the continuous expansion of urban scale, road congestion and frequent traffic accidents, which has increased the pressure on urban traffic management. Accurate traffic forecasting plays a key role in intelligent transportation systems (ITS), providing sufficient scientific basis for time-sensitive traffic decisions such as intelligent traffic control and navigation path planning. Traffic signals are often represented by geographically located multivariate time series (MTS), which are indistinguishable in time and space. How to fully extract the spatiotemporal characteristics of traffic data is a key issue that needs to be solved in traffic forecasting.

[0003] Early traffic flow prediction methods are usually regarded as time series mining tasks, such as autoregressive integrated moving average (ARIMA) and Kalman filtering to model the time series characteristics of traffic signals. However, these models are based on the assumption of smooth time series, which are suitable for scenarios with sufficient data volume and small fluctuation range, ignoring the spatial dimension characteristics contained in the data.

[0004] Traditional machine learning methods (such as K nearest neighbor and support vector regression) can achieve higher prediction accuracy and more complex data modeling, but the structure of such models is relatively simple and cannot meet the needs of accurately predicting traffic conditions. With the development of deep learning technology, convolutional neural networks (CNN), recurrent neural networks (RNN) and their variants (LSTM, GRU), long short-term memory networks (LSTM) and gated recurrent units (GRU) have been tried for traffic flow prediction tasks. In early studies, the road network was divided into grids of equal size so that CNN could be used to learn the spatial dependencies between grids, and RNN and its variants were used to capture temporal features. These methods can capture nonlinear temporal features in the data, but it is difficult to describe the mutual influence between nodes for the prediction of a single node, and as a non-Euclidean spatial structure, the road network has a poor learning effect on spatial dependencies using CNN.

[0005] Graph Convolutional Networks (GCNs) are widely used for modeling pairwise relationships due to their excellent spatial feature extraction capabilities and integrated with Recurrent Neural Networks (RNNs) and their variants or Temporal Convolutional Networks (TCNs) as spatio-temporal graph neural networks (such as STGCN, MTGNN, DCRNN, etc.) to provide accurate traffic predictions while exploring spatio-temporal information. The traffic flow prediction method based on GCN focuses on constructing a graph or adjacency matrix, which represents road segments as nodes and the relationships between road segments as edges. In early studies, researchers usually constructed predefined graph structures using prior knowledge (such as road segment distances, POI similarities, etc.), but these static graph structures cannot well reflect the dynamic changes in road network traffic flow. Therefore, some methods attempt to construct dynamic graphs or adjacency matrices using traffic data, and compared with static graph structures, the GCN methods based on dynamic graphs have achieved quite significant improvements.

[0006] However, despite the good results achieved by these studies, there are still some problems in existing traffic flow prediction methods that have not been solved. The first problem is that traditional graph convolutional networks cannot capture dynamic non-pairwise relationships in the road network. Graph convolution updates the representation of the current node by weighted aggregation of the features of the current node and its neighbor nodes, enabling the representation of the node to capture local graph structure information. However, this method can only capture pairwise relationships, but there may be complex non-pairwise relationships in the traffic network. For example, in many scenarios with similar historical data, the future results are completely opposite. Therefore, modeling based on pairwise relationships may not be able to fully utilize global topological structure information, limiting the representational ability of graph convolutional networks for the overall features of the graph. The second problem is the inability to fully explore long-term spatio-temporal correlations. Currently, the GCN methods based on dynamic graphs model the dynamic characteristics of traffic signals, use GCN to extract spatial features at each time step, and then use RNNs and their variants or TCNs to aggregate each time step to extract temporal features. In fact, such methods are vulnerable to long-term dependencies, and the limited memory capacity makes it difficult for the model to learn long-distance spatio-temporal dependencies. Summary of the Invention

[0007] According to the technical problems mentioned in the above background art, a traffic prediction method based on a dual dynamic graph attention network is provided. The present invention dynamically models the road network structure without relying on prior knowledge, and converts the traffic flow graph into a hypergraph using dual transformation to synchronously capture the spatial correlations of nodes and edges in the traffic flow graph. By designing a gated attention linear unit, it better captures the potential long-term dynamic spatio-temporal features in traffic signals; by stacking several gated linear units, it extracts temporal correlations at different time scales, achieving the capture of dynamic spatio-temporal correlations in the road network from both local and global perspectives.

[0008] The technical means adopted by the present invention are as follows:

[0009] A traffic prediction method based on a dual dynamic graph attention network, characterized by comprising the following steps:

[0010] S101: Without relying on prior knowledge, define the traffic prediction problem as a multivariate time series prediction problem, attach spatio-temporal information to the input traffic signals and generate a dynamic graph through linear transformation, and at the same time generate a dual dynamic hypergraph by dual transformation of the dynamic graph;

[0011] S102: Input the traffic signal to be predicted and the dual graph structure into the spatial feature extraction module to capture and integrate spatially relevant features;

[0012] S103: Use the time feature extraction module to capture time-related features at different time scales through multiple stacked gated attention linear units, and obtain the spatio-temporal related features captured by the current spatio-temporal feature extraction module;

[0013] S104: The output module performs linear processing and residual decomposition on the extracted spatio-temporal related features to obtain the prediction result of the current module and the signal input of the next block, and integrates the outputs of all spatio-temporal feature extraction modules to obtain the final prediction value.

[0014] Further, define the traffic prediction problem as a multivariate time series prediction problem, and define the road network structure as G=(V, E, A), where V={v 1 , v 2 ,..., v N}∈R N represents the set of nodes, N represents the number of nodes, E={e 1 , e 2 ,..., e M}∈R M is the set of edges, M represents the number of edges, A∈R N×N represents the adjacency matrix, and its elements represent the connectivity between nodes;

[0015] Represent the dual hypergraph structure through the adjacency matrix H. When the node v i ∈V h is associated with the edge e j ∈E h of the dual hypergraph structure, define:

[0016]

[0017] Adopt the dual transformation method for the graph G=(V, E, A) to obtain the dual hypergraph G h =(V h , E h , H); where, V h represents the node set, E hDenote the hyperedge set, Denote the adjacency matrix.

[0018] Use the feature matrix to describe the traffic signals of the road network at time t, where represents the observed value of node i at time t, and F represents the dimension of the node attribute features of the traffic road network;

[0019] That is, for a given traffic flow graph G and historical traffic signals χ P =[X t-P+1 ,...,X t-1 ,X t ∈R P×N×F , learn the mapping function f to predict the traffic data for the next Q time steps. The expression of the final predicted value is:

[0020] f([X t-P+1 ,...,X t-1 ,X t ; G)=[X t+1 ,X t+2 ,...,X t+Q .

[0021] Furthermore, considering the dynamic spatial correlation existing in the road network at different times, introduce a spatio-temporal embedding generator of spatio-temporal information to generate the dynamic graph;

[0022] According to the given traffic signal χ P find the corresponding daily time embedding T D and weekly time embedding T W , and then perform the Hadamard product operation with the node spatial embedding S to obtain a new spatio-temporal embedding Denoted as:

[0023]

[0024] where, ⊙ represents the Hadamard product operation;

[0025] Pass the input signal χ P through an MLP layer to extract the dynamic features, and perform the Hadamard product operation on the dynamic features and the corresponding spatio-temporal embedding to generate the dynamic graph embedding Denoted as:

[0026]

[0027] where, tanh(.) represents the activation function; the adjacency matrix of the dynamic graph is obtained by multiplying it by its own transpose, that is

[0028] Furthermore, the method of using dual transformation is adopted to convert the dynamic graph into a dynamic hypergraph, that is, the edges and nodes of the given graph are mutually mapped, the nodes of the graph are mapped to the target edges, and the edges of the graph are mapped to the target nodes, which is expressed as:

[0029]

[0030] where V h = E, E h = V, H ∈ R M×N ; H represents the structural information of the dynamic graph and also represents the corresponding dual dynamic hypergraph of the dynamic graph;

[0031] H is decomposed into two parts, that is:

[0032] H = H src + H dst ;

[0033] where H src and H dst respectively represent non-zero element matrices formed by the starting nodes and ending nodes of the directed edges in the dynamic graph;

[0034] The traffic signal χ P is subjected to feature transformation to generate node features suitable for input hypergraph convolution

[0035]

[0036] where W dis is the road network distance matrix, W 1 , W 2 are learnable parameters, and || represents the concatenation operation.

[0037] Furthermore, the spatial feature extraction module includes: a spatial collaborative learning layer, a dual dynamic graph convolution layer, and a temporal self-attention feature fusion layer;

[0038] The spatial collaborative learning layer is used to generate a hypergraph adjacency matrix based on graph attention and generate an edge representation matrix based on dynamic hypergraph attention; the dual dynamic graph convolution layer extracts the spatial correlation in the traffic signal through the dynamically generated graph structure and the dynamic graph edge representation matrix of the spatial collaborative learning layer to extract the spatial correlation in the traffic signal; the temporal self-attention fusion layer performs feature fusion on the spatial features extracted by the dynamic graph convolution and the hypergraph convolution network.

[0039] Furthermore, the temporal feature extraction module includes: stacked gated attention linear units GALU; the temporal feature extraction module uses a skip connection method to achieve information transfer between different gated attention linear units GALU;

[0040] The gated attention linear unit GALU includes: an dilated causal convolutional layer based on GLU and an dilated causal convolutional layer based on temporal attention.

[0041] Further, in step S104, in order to achieve residual decomposition and multi-step prediction, two output results are obtained through two fully connected layers and

[0042]

[0043] Among them, represents the prediction result of the spatio-temporal feature extraction module of block j ∈ [0, J-1], represents the backward prediction result of this block for the input , W j and W B are the learnable weights of the fully connected layer, b j and b B are the biases of the fully connected layer; the input of the next block, i.e., the spatio-temporal feature extraction module of block j+1, is expressed as:

[0044]

[0045] The final output y P of the model is the sum of the prediction results of the J-block dynamic graph convolutional recurrent module:

[0046]

[0047] Compared with the prior art, the present invention has the following advantages:

[0048] The present invention proposes a traffic prediction method based on a dual dynamic graph attention network. This method adopts a spatio-temporal convolutional network architecture, and uses a spatial feature extraction module based on a dual dynamic graph convolutional network and a temporal feature extraction module based on a temporal convolutional network to learn the dynamic characteristics in traffic signals.

[0049] The present invention uses a spatial feature extraction module to learn the spatial correlation of traffic signals. This module generates a dynamic graph through time-varying traffic signals, constructs a dual hypergraph using a dual transformation method, and captures the dynamic characteristics of nodes and edges in the traffic flow graph through a dual dynamic graph convolutional network, effectively modeling the spatial cooperation effect of traffic flow.

[0050] The present invention proposes a temporal feature extraction module that stacks multiple gated attention linear unit (GALU) blocks to model the temporal correlation of traffic signals. The GALU block flexibly learns the temporal dependence of traffic flow through a temporal gating mechanism, improving the prediction ability of the model in the temporal dimension.

[0051] The present invention has conducted a large number of comparative experiments on four real-world traffic datasets. The experimental results show that, compared with 18 baseline methods, the method of the present invention achieves more accurate prediction accuracy on different datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0053] Figure 1 It is a schematic flowchart of a traffic prediction method based on a dual dynamic graph attention network of the present invention;

[0054] Figure 2 It is a schematic diagram of the framework of a traffic prediction model based on a dual dynamic graph attention network of the present invention;

[0055] Figure 3 It is a schematic diagram of the dual transformation from a dynamic graph to a dynamic hypergraph of the present invention;

[0056] Figure 4 It is a schematic diagram of the structure of a gated attention linear unit (GALU) of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0058] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0059] As shown Figure 1 in the figure, a traffic prediction method based on a dual-dynamic graph attention network according to the present invention includes the following steps:

[0060] S101: Without relying on prior knowledge, define the traffic prediction problem as a multivariate time series prediction problem, attach spatio-temporal information to the input traffic signals and generate a dynamic graph through a linear transformation. At the same time, generate a dual dynamic hypergraph through a dual transformation of the dynamic graph;

[0061] S102: Input the input traffic signals and the dual-graph structure into the spatial feature extraction module to capture and integrate spatial correlations;

[0062] S103: Input the output of the spatial feature extraction module into the temporal feature extraction module, extract temporal correlations at different time scales through a plurality of stacked gated attention linear units, and obtain the spatio-temporal correlation features captured by the current spatio-temporal feature extraction module;

[0063] S104: The output module performs linear processing and residual decomposition on the extracted spatio-temporal correlation features to obtain the prediction result of the current module and the signal input of the next block, and integrates the outputs of all spatio-temporal feature extraction modules to obtain the final prediction value. In order to implement residual decomposition and multi-step prediction, two output results are obtained through two fully connected layers and

[0064]

[0065] Among them, represents the prediction result of the spatio-temporal feature extraction module of block j∈[0,J-1], represents the reverse prediction result of this block for the input , W j and W B are the learnable weights of the fully connected layer, b j and b B are the biases of the fully connected layer; the input of the next block, i.e., the spatio-temporal feature extraction module of block j+1, is expressed as:

[0066]

[0067] The final output y P of the model is the sum of the prediction results of the J-block dynamic graph convolutional recurrent modules as:

[0068]

[0069] As Figure 2As shown in the figure, in the spatial dimension, considering that traditional graph convolution cannot capture the dynamic non-pairwise relationships in the road network, the constructed dynamic graph is transformed into a dual hypergraph through dual transformation, and the hypergraph is used to capture the high-order complex information contained in the road network. The graph attention and hypergraph attention networks are embedded into the spatial collaborative learning module to propagate features between the dual graph structures, enabling the dual graph convolutional network to capture more complex relationships in the road network. The temporal self-attention feature fusion module is used to fuse the spatial features extracted by the dual graph convolutional network. In the temporal dimension, the temporal self-attention mechanism is used, and the position encoding is added to make up for the deficiency of the temporal feature extraction module in modeling long-term temporal dependencies to some extent. The gated attention linear unit is proposed, and the temporal correlation is extracted in different temporal dimensions by stacking multiple unit blocks. The present invention realizes the efficient prediction of traffic signals by dynamically extracting spatial and temporal patterns.

[0070] As a preferred embodiment, in the present application, the traffic prediction problem is defined as a multivariate time series prediction problem, and the road network structure is defined as G=(V, E, A), where V={v 1 , v 2 ,..., v N}∈R N represents the set of nodes, N represents the number of nodes, E={e 1 , e 2 ,..., e M}∈R M is the set of edges, M represents the number of edges, A∈R N×N represents the adjacency matrix, and its elements represent the connectivity between nodes;

[0071] The dual hypergraph structure is represented by the adjacency matrix H. When the node v i ∈V h is associated with the edge e j ∈E h of the dual hypergraph structure, it is defined as:

[0072]

[0073] The dual transformation method is used for the graph G=(V, E, A) to obtain the dual hypergraph G h =(V h , E h , H); where, V h represents the node set, E h represents the hyperedge set, represents the adjacency matrix.

[0074] The feature matrix is used to describe the traffic signals of the road network at time t, where Denote the observed value of node $i$ at time $t$, and $F$ represents the dimension of the traffic road network node attribute features;

[0075] That is, for a given traffic flow graph $G$ and historical traffic signals $\chi$ P $= [X$ t-P+1 , \ldots, X$ t-1 , X$ t \in \mathbb{R}$ P×N×F , learn the mapping function $f$ to predict the traffic data for the next $Q$ time steps. The final prediction value expression is:

[0076] $f([X$ t-P+1 , \ldots, X$ t-1 , X$ t ; G) = [X$ t+1 , X$ t+2 , \ldots, X$ t+Q .

[0077] Preferably, considering the dynamic spatial correlation existing in the road network at different times, a spatio-temporal embedding generator introducing spatio-temporal information is used to generate a dynamic graph;

[0078] According to the given traffic signal $X$ P find the corresponding daily time embedding $T$ D and weekly time embedding $T$ W , and then perform a Hadamard product operation with the node spatial embedding $S$ to obtain a new spatio-temporal embedding which is expressed as:

[0079]

[0080] where $\odot$ represents the Hadamard product operation;

[0081] Pass the input signal $X$ P through an MLP layer to extract dynamic features, and perform a Hadamard product operation on the dynamic features and the corresponding spatio-temporal embedding to generate a dynamic graph embedding which is expressed as:

[0082]

[0083] where $\tanh(.)$ represents the activation function; the dynamic graph adjacency matrix is obtained by multiplying it by its own transpose, that is

[0084] Adopt the method of dual transformation to convert the dynamic graph into a dynamic hypergraph, that is, map the edges and nodes of the given graph to each other, map the nodes of the graph to the edges of the target, and map the edges of the graph to the nodes of the target, which is expressed as:

[0085]

[0086] Among them, V h = E, E h = V, H ∈ R M×N ; H represents the structural information of the dynamic graph and also represents the dual dynamic hypergraph corresponding to the dynamic graph;

[0087] Decompose H into two parts, namely:

[0088] H = H src + H dst ;

[0089] Among them, H src and H dst respectively represent non-zero element matrices formed by the starting nodes and ending nodes of the directed edges in the dynamic graph;

[0090] Perform feature transformation on the traffic signal χ P to generate node features suitable for input to the hypergraph convolution

[0091]

[0092] Among them, W dis is the road network distance matrix, W 1 , W 2 are learnable parameters, and || represents the concatenation operation.

[0093] In this application, as a preferred implementation, the spatial feature extraction module includes: a spatial collaborative learning layer, a dual dynamic graph convolution layer, and a temporal self-attention feature fusion layer;

[0094] The spatial collaborative learning layer is used to generate a hypergraph adjacency matrix based on graph attention and generate an edge representation matrix based on dynamic hypergraph attention; the dual dynamic graph convolution layer extracts the spatial correlation in the traffic signal through the dynamically generated graph structure and the dynamic graph edge representation matrix of the spatial collaborative learning layer to extract the spatial correlation in the traffic signal; the temporal self-attention fusion layer performs feature fusion on the spatial features extracted by the dynamic graph convolution and the hypergraph convolution network.

[0095] The generation of the hypergraph adjacency matrix based on graph attention is to generate the dynamic hyperedge features of the hypergraph by using the dual relationship between the dynamic graph and its hypergraph, that is, to use the dynamic graph attention network DGAT to extract node features on the dynamic graph as the hyperedge features for updating the hypergraph. DGAT is defined as:

[0096]

[0097] Among them, represents the spatial features extracted by the graph convolution network, Denote the graph attention coefficient at time t, AdpAvgPool(.) represents the adaptive average pooling layer, RELU(.) represents the activation function, and Softmax(.) represents the normalization operation;

[0098] Use the graph attention mechanism to dynamically allocate different weights according to the feature similarity between the target node and adjacent nodes, and adaptively adjust the adjacency relationship; perform a linear transformation and diagonalization operation on the hyperedge features to obtain the weighted diagonal matrix of the hypergraph

[0099]

[0100] Among them, respectively represent the weight matrix and bias of the linear transformation, and diag(.) represents the diagonalization operation on the tensor; the adjacency matrix of the dynamic hypergraph convolutional network is expressed as:

[0101]

[0102] The adjacency matrix of the dynamic hypergraph convolutional network is the input of DHGCN in the double dynamic graph convolutional layer;

[0103] The generation of the edge representation matrix based on dynamic hypergraph attention is to use the dynamic hypergraph attention network DHGAT to extract the node features in the hypergraph to generate the edge features of the dynamic graph; first, extract the correlation between the start node and the end node in the graph as the input of the hypergraph attention network

[0104]

[0105] Among them, W 3 , W 4 represent the learnable parameters of the input linear transformation, (.)[ind src , :] and (.)[ind dst , :] respectively represent the tensors sliced by the ind src and ind dst corresponding to the start node and end node selection of the directed edge in the traffic flow graph, and || represents the concatenation operation;

[0106] Use DHGAT to implement the generation of dynamic edge features:

[0107]

[0108] Among them, represents the weighted diagonal matrix of the hypergraph for learning hyperedge features, and L adp represents the weight tensor for adaptively learning hyperedge features, and the hypergraph adjacency matrix of the dynamic graph attention network:

[0109]

[0110] Generate the edge representation matrix of the dynamic graph by changing the output shape of the edge features

[0111]

[0112] Among them, Reshape(.) represents the operation used to transform the shape of the tensor, represents the edge representation matrix of the dynamic graph and will be used as the input of DGCN in the double dynamic graph convolutional layer.

[0113] The double dynamic graph convolutional layer extracts the spatial correlation in traffic signals through the dynamically generated graph structure and the dynamic graph edge representation matrix of the spatial collaborative learning layer and is expressed as:

[0114]

[0115] For the high-order complex information contained in the traffic road network, the dual transformation method is used to convert the dynamic graph into a dynamic hypergraph, and the dynamic hypergraph convolutional DHGCN is introduced to capture the hidden non-pairwise relationships in traffic signals. The hypergraph node features of the dynamic hypergraph convolutional network are used as the input, and the hypergraph adjacency matrix generated by the spatial collaborative learning layer is used as the graph structure to extract the spatial correlation contained in the hypergraph structure, and the formula is:

[0116]

[0117] The temporal self-attention fusion layer is used to fuse the spatial features extracted by the dynamic graph convolution and the hypergraph convolution network. The temporal self-attention network mainly consists of a positional encoding, a temporal self-attention mechanism, and a residual connection. The temporal self-attention mechanism enables the model to observe the long-term time trend and focus on highly relevant information, thus making up for the deficiency of the temporal feature extraction module in modeling long-term dependencies to a certain extent.

[0118] In the temporal self-attention network, the positional encoding is added to the feature data at each time step, and the formula is as follows:

[0119]

[0120] Among them, is the feature after adding the positional encoding, i ∈ [0, N - 1], p ∈ [0, P - 1], e p ∈ R F represents the positional encoding, and its definition is as follows,

[0121]

[0122] Among them, Sin(.) and Cos(.) represent the sine function and cosine function respectively. The temporal self-attention network generates the context representation of the sequence by embedding the relationships between positions in the sequence. First, the input sequence is linearly transformed to generate Query, Key, and Value matrices:

[0123]

[0124] where, W Q 、W K 、W V represent the weight matrices learned by the linear transformation. Then, Q and K T are matrix-multiplied and normalized to obtain the attention distribution at each time step. The attention matrix is weighted-summed with V to generate the final temporal context corresponding to the dynamic graph convolution The formula is as follows:

[0125]

[0126] The features extracted by the dynamic hypergraph convolution are consistent with those of the dynamic graph convolution during the processing of the temporal self-attention network, and the corresponding temporal context is generated The dual transformation is used to transform the features of the hypergraph into the feature dimension of the graph, facilitating the feature fusion of the graph and hypergraph. The node features extracted by the dynamic hypergraph convolution are mapped to the node features of the dynamic graph

[0127]

[0128] where, W 5 ∈R M×N is a learnable parameter used to capture the correlation between edges and nodes in the dual hypergraph. Feature fusion is the summation operation on the temporal features of the dynamic graph convolution and the temporal features of the dynamic hypergraph convolution transformed into dynamic graph features . The formula is as follows,

[0129]

[0130] where, represents the output of the spatial feature extraction module. A dimension rearrangement operation is performed on to improve the efficiency of the temporal feature extraction module in processing time series features. Dropout(.) represents the regularization operation.

[0131] Preferably, the time feature extraction module includes: stacked gated attention linear units GALU; the time feature extraction module uses skip connections to implement information transfer between different gated attention linear units GALU;

[0132] The gated attention linear unit GALU includes: an dilated causal convolutional layer based on GLU and an dilated causal convolutional layer based on temporal attention.

[0133] Given the input feature sequence of the spatial feature extraction module Filter K is the size of the time filter, and the default value of K is 2; And The dilated causal convolution operation at step t is expressed as:

[0134]

[0135] Where Represents the convolution operation using the filter , d represents the dilation factor, and P′ represents the new sequence length after the convolution operation;

[0136] The dilated causal convolutional layer based on GLU uses a linear transformation to transform the input Into Capture the temporal dependence of traffic signals through dilated causal convolution, divide the output result into two parts along the feature dimension And input it into GLU; then the formula for the convolutional layer based on GLU is:

[0137]

[0138] Where Represents the output sequence of the GLU dilated causal convolutional layer, and σ(.) represents the Sigmoid activation function;

[0139] Finally, Is input into the residual network to avoid the problem of gradient disappearance during training.

[0140] Considering the problem that the GLU layer has insufficient ability to capture the dynamic features of traffic signals, the dilated causal convolutional layer based on temporal attention models the dynamic change law of temporal features and calculates the temporal attention score

[0141]

[0142] Where, W 6 , W 7 , W 8 , W 9They are all learnable parameters. Through the weighted fusion of the temporal attention scores and the temporal features extracted by the dilated causal convolution, the attention of the model to the key time periods is further enhanced:

[0143]

[0144] Among them, represents the output of the dilated causal convolution layer based on temporal attention. In summary, the output of the i-th GALU is the sum of the GLU dilated causal convolution layer and the dilated causal convolution layer of temporal attention. A residual connection is adopted between the input and output of each GALU block to avoid the problem of gradient disappearance in model training. Therefore, the output of the i-th GALU block is defined as:

[0145]

[0146] The temporal feature extraction module is composed of L GALU blocks combined through skip connections, and the final output considers the inputs of all GALU blocks and the output of the last GALU block They represent the hidden states at each temporal level. Adding skip connections to each hidden state is essentially a standard convolution with a convolution kernel of 1×P i+1 , and the final output feature of the connection is defined as,

[0147]

[0148] Among them, represents the final output feature of the temporal feature extraction module, P L = 1, are the learnable parameters of these standard convolutions.

[0149] Example 1

[0150] To better evaluate the effectiveness of the method of the present invention, it is compared with more than a dozen other traffic predictions, and the method of the present invention is verified based on the PEMS03, PEMS04, PEMS07, and PEMS08 datasets:

[0151] The PEMS03, PEMS04, PEMS07, and PEMS08 datasets are all collected by the loop sensors of the California Department of Transportation on highways. The data is collected every 30 seconds and summarized into five-minute time steps to record the current traffic conditions of the road section, and Z-Score is used for standardization. The datasets are divided in chronological order, with 60% for training, 20% for validation, and 20% for testing. The detailed information of the datasets is shown in Table 1.

[0152] Table 1 Datasets used in the experiment

[0153]

[0154] In Table 1, Dataset represents the dataset, Nodes represents the nodes, Edges represents the edges, Time steps represents the time steps, and Time Range represents the time range. The present invention uses the Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Mean Absolute Percentage Error (MAPE) as evaluation metrics.

[0155] The design historical step P = 12 and the prediction time step Q = 12, that is, the data for the next hour is predicted based on the observations of the previous hour. This model is implemented using the PyTorch deep learning framework, and all experiments are trained on an NVIDIA Tesla A800 GPU. The Adam optimizer with an initial learning rate of 0.001 is used for training. The early stopping strategy is used on the validation dataset, waiting patiently for 15 iterations to prevent overfitting.

[0156] Table 2 Comparison results of the present invention with other models

[0157]

[0158] In Table 2, Models represents the models. Table 2 shows the error metrics of the prediction of the present invention and existing methods on the test dataset. Traditional statistics-based models (such as HA, ARIMA, etc.) are based on the assumption of stationary data and are difficult to handle non-linear and non-stationary traffic data, showing the worst performance. Non-graph-based deep learning models (such as FC-LSTM, TCN, etc.) use deep learning methods to capture the temporal features in traffic data and achieve good results. However, these models only consider temporal correlations and ignore spatial correlations, resulting in poorer performance compared to graph structure-based models. Graph neural network models can capture temporal and spatial correlations. Among them, models such as DCRNN, STGCN, and GWNet use predefined graph structures to capture spatial correlations, and models such as AGCRN and STG-NCDE learn the spatial structure through an adaptive matrix. These models model spatial relationships through static structures and do not consider the dynamic change trend of traffic signals. Although models such as DSTAGNN use the dynamic characteristics in historical data to construct the graph structure, due to the fact that graph convolution can only capture pairwise relationships, there are still deficiencies in capturing spatio-temporal correlations. Compared with other models, the method of the present invention not only considers capturing non-pairwise relationships in traffic data but also mines the traffic fluctuation relationships of long and short distances from the perspective of temporal attention.

[0159] To evaluate the impact of the key components of the traffic prediction network of the present invention on the model prediction effect, three variants based on the present invention were designed for ablation experiments, and the method of the present invention was compared with these three variants on the PEMS04 and PEMS08 datasets to verify the effectiveness of these components for feature extraction. The differences among the four variants are as follows:

[0160] w / o DG: The dynamic graph is removed, and only the traditional predefined adjacency matrix is used.

[0161] w / o DHGCN: The dual dynamic hypergraph is removed, and only graph convolution is used to model spatially related features.

[0162] w / o GALU: The time feature extraction module is removed, and only the time-related feature extraction model is used to model traffic signals.

[0163] Table 3 reports the evaluation metrics of the three ablation comparison models and the model of the present invention on the test set. The results show that the model of the present invention has the best performance compared with the experimental results of the three variants. Specifically, the network model of the present invention is better than w / o DG, indicating that using the dynamic graph brings benefits to modeling the dynamic characteristics of traffic data, better than w / o DHGCN, indicating that hypergraph convolution plays a positive role in spatial feature extraction, and better than w / o GALU, indicating the necessity of the time feature extraction module to extract the global time correlation of traffic flow. From the above results, on the one hand, it shows the importance of the key components in the traffic prediction network of the present invention, and on the other hand, it also shows the effectiveness of the method of the present invention in extracting the spatio-temporal correlation of traffic flow sequences.

[0164] Table 3 Ablation Experiment Research on PEMS04 and PEMS08

[0165]

[0166] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0167] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0168] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in electrical or other forms.

[0169] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0170] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0171] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0172] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of each embodiment of the present invention.

Claims

1. A traffic prediction method based on dual dynamic graph attention network, characterized in that: The following steps are involved: S101: Without relying on prior knowledge, the traffic prediction problem is defined as a multivariate time series prediction problem. The temporal and spatial information of the input traffic signals is added and a dynamic graph adjacency matrix is ​​generated through linear transformation. Meanwhile, the dynamic graph is dually transformed to generate a dual dynamic hypergraph adjacency matrix. S102: inputting the traffic signal to be predicted and the dual dynamic graph adjacency matrix into a spatial feature extraction module to capture and integrate spatial related features; S103: Using the temporal feature extraction module, the temporal related features are captured at different time scales by stacking multiple gated attention linear units, and the temporal and spatial related features captured by the current temporal and spatial feature extraction module are obtained; S104: The output module performs linear processing and residual decomposition on the extracted spatiotemporal related features to obtain the prediction result of the current module and the signal input of the next block, and integrates the outputs of all spatiotemporal feature extraction modules to obtain the final prediction value.

2. A traffic prediction method based on dual dynamic graph attention network according to claim 1, characterized in that: The traffic prediction problem is defined as a multivariate time series prediction problem, and the road network structure is defined as G = (V, E, A), where V = {v1, v2, ..., v N }∈R N represents the set of nodes, N represents the number of nodes, E={e1,e2,...,e M }∈R M is a set of edges, M represents the number of edges, A∈R N×N represents the adjacency matrix, whose elements represent the connectivity between nodes; The dual hypergraph structure is represented by the adjacency matrix H. i ∈V h With the dual hypergraph structure edge e j ∈E h When associated, define: For graph G = (V, E, A), the dual transformation method is used to obtain the dual hypergraph G of graph G. h =(V h ,E h ,H); Among them, V h represents a node set, E h represents a hyperedge set, Represents the adjacency matrix. Using the feature matrix To describe the traffic signal of the road network at time t, represents the observed value of node i at time t, and F represents the dimension of the attribute characteristics of the traffic network node; That is, given a traffic flow graph G and historical traffic signals χ P =[X t-P+1 ,...,X t-1 ,X t ]∈R P×N×F , learn the mapping function f to predict the traffic data of the next Q time steps, and the final prediction value expression is: f([X t-P+1 ,...,X t-1 ,X t ];G)=[X t+1 ,X t+2 ,...,X t+Q ]。 3. The traffic prediction method based on dual dynamic graph attention network according to claim 1 is characterized in that: Taking into account the dynamic spatial associations existing in the road network at different times, a spatiotemporal embedding generator that introduces spatiotemporal information generates the dynamic graph; According to the given traffic signal χ P Find the corresponding daily time embedding T D and weekly time embedding T W , and then perform Hadamard product operation with the node space embedding S to obtain the new spatiotemporal embedding It is expressed as: Among them, ⊙ represents the Hadamard product operation; The input signal χ P After an MLP layer, dynamic features are extracted and embedded with the corresponding spatiotemporal embedding Perform Hadamard product operation to generate dynamic graph embedding It is expressed as: Among them, tanh(.) represents the activation function; the dynamic graph adjacency matrix is ​​obtained by Multiply it by its own transpose to get 4. A traffic prediction method based on dual dynamic graph attention network according to claim 1, characterized in that: The dual transformation method is used to convert the dynamic graph into a dynamic hypergraph, that is, the edges and nodes of a given graph are mapped to each other, the nodes of the graph are mapped to the edges of the target, and the edges of the graph are mapped to the nodes of the target, which is expressed as: Among them, V h =E,E h =V,H∈R M×N ; H represents the dynamic graph and also represents the structural information of the dual dynamic hypergraph corresponding to the dynamic graph; Decompose H into two parts, namely: H=H src +H dst ; Among them, H src and H dst They represent the zero-element matrices composed of the start and end nodes of the directed edges in the dynamic graph respectively; Traffic signal χ P Perform feature conversion to generate node features suitable for input hypergraph convolution Among them, W dis is the road network distance matrix, W1, W2 are learnable parameters, and || represents the concatenation operation.

5. The traffic prediction method based on dual dynamic graph attention network according to claim 1 is characterized in that: The spatial feature extraction module includes: a spatial collaborative learning layer, a dual dynamic graph convolution layer, and a temporal self-attention feature fusion layer; The spatial collaborative learning layer is used to generate a hypergraph adjacency matrix based on graph attention and an edge representation matrix based on dynamic hypergraph attention; the dual dynamic graph convolution layer is used to generate a dynamic graph edge representation matrix based on the dynamically generated graph structure and the spatial collaborative learning layer. To extract the spatial correlation in traffic signals; the temporal self-attention fusion layer performs feature fusion on the spatial features extracted by the dynamic graph convolution and hypergraph convolution networks.

6. A traffic prediction method based on dual dynamic graph attention network according to claim 1, characterized in that: The temporal feature extraction module comprises: a stacked gated attention linear unit GALU; the temporal feature extraction module uses a skip connection method to realize information transmission between different gated attention linear units GALU; The gated attention linear unit GALU includes: a GLU-based dilated causal convolutional layer and a temporal attention-based dilated causal convolutional layer.

7. A traffic prediction method based on dual dynamic graph attention network according to claim 1, characterized in that: In step S104, in order to achieve residual decomposition and multi-step prediction, two output results are obtained through two fully connected layers: and in, represents the prediction result of the spatiotemporal feature extraction module of block j∈[0,J-1], Indicates that the block is input The reverse prediction result, W j and W B is the learnable weight of the fully connected layer, b j and b B is the bias of the fully connected layer; the next block is the input of the j+1 spatiotemporal feature extraction module It is expressed as: The final output of the model is P The prediction results of the J-block dynamic graph convolution loop module are added together as follows:

Citation Information

Cited By

  • Green wave control method based on traffic flow prediction driving

    CN121459605A