Traffic flow prediction method of attention-based long and short time-space diagram neural network
Through the attention-based long and short space-time graph neural network, combined with the graph attention network and the time convolution network, the problem of insufficient space-time dependence in traffic flow prediction is solved, and high-precision and high-rootability traffic flow prediction is achieved.
Patent Information
- Application Number
- CN202510092271.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The existing traffic flow prediction methods have shortcomings in capturing the space-time dependence relationship of traffic flow, making it difficult to achieve high accuracy and high robustness, which limits its application in actual scenarios.
The attention-based long and short space-time graph neural network is adopted to construct an adjacency matrix and a multi-head attention mechanism, combined with the graph attention network (GAT) and the time convolution network (TCN), the spatiotemporal characteristics of traffic data are extracted, and the LSTM is used for prediction, integrating long and short-term traffic flow.
It improves the accuracy and generalization performance of traffic flow prediction, can effectively capture the complex spatial dependence and time-dependent characteristics of the traffic network, and adapt to the dynamic changes of traffic flow.
Smart Images

Figure CN120012828A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of traffic management and intelligent transportation, and relates to a traffic flow prediction method based on an attention-based long-short spatiotemporal graph neural network. Background Art
[0002] Traffic flow prediction has an indispensable and important application value in modern urban traffic management, intelligent transportation system construction and travel planning. Through accurate traffic flow prediction, traffic management departments can optimize the timing control of traffic lights and reasonably allocate traffic resources, thereby effectively alleviating traffic congestion and improving the overall traffic efficiency of the road network. In addition, accurate traffic flow prediction can also provide high-quality real-time route planning suggestions for individual travelers, helping them avoid peak hours or congested sections and achieve a more efficient and comfortable travel experience. This not only improves the operating efficiency of the urban transportation system, but also creates favorable conditions for the sustainable development of the social economy. Traditional traffic flow prediction methods have achieved certain results in early research and application. However, due to the significant limitations of these methods in capturing the complex dynamic characteristics of traffic flow changing with time and space, their prediction accuracy is often difficult to meet the high requirements of modern transportation systems. This is mainly because traffic flow is affected by a variety of factors, and its spatiotemporal distribution has highly nonlinear and dynamic characteristics. It is difficult to accurately characterize these complex characteristics by relying solely on traditional methods. In recent years, with the rapid development of deep learning technology, more and more studies have introduced it into the field of traffic flow prediction. Deep learning methods, with their advantages in processing complex nonlinear and high-dimensional data, have made up for the shortcomings of traditional methods to a certain extent and significantly improved the accuracy of traffic flow prediction. However, even so, existing deep learning models still have obvious deficiencies in modeling the spatiotemporal dependencies of traffic flow. Many models only focus on the single dependency characteristics of the time dimension or space dimension, or perform simple linear modeling of the spatiotemporal dependencies, and cannot fully explore the deep correlation of traffic flow in the spatiotemporal dimensions. These models still find it difficult to achieve high accuracy and high robustness when predicting traffic flow in future time periods, which limits their widespread application in actual scenarios.
[0003] Therefore, a prediction method for attention-based long-short spatiotemporal graph neural networks is needed. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a traffic flow prediction method based on attention-based long-short spatiotemporal graph neural network.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A traffic flow prediction method based on attention-based long-short spatiotemporal graph neural network, the core steps of which include:
[0007] S1: Obtain historical traffic data, including flow, speed and other characteristics, build an adjacency matrix based on the road segments or node relationships in the traffic network to represent the spatial dependency between nodes, divide the historical data into time windows, and normalize the data;
[0008] S2: Input the traffic history data and adjacency matrix into the model;
[0009] S3: Extract temporal and spatial features respectively through the spatiotemporal module, and improve the generalization ability through the feedforward neural network;
[0010] S4: Through the spatial attention module, a three-layer spatial attention block is used to realize the differential extraction of important features in the shared layer. The attention mechanism weighs the output according to the attention weight, and the output result is merged into the LSTM, and the information of the dense layer is processed by the LSTM to perform the task prediction, and finally the long and short flow predictions are integrated to obtain the final prediction result;
[0011] In S1, time series data and adjacency matrix data are obtained, and the dimensions are constructed according to the received data as (V, E, A), where V represents the number of nodes, represents the road sections included in the traffic network, E represents the edge set connecting the nodes, and A represents the adjacency matrix;
[0012] In S2, it is assumed that we have N nodes (i.e., road sections in the traffic network), and the feature of each node is a time series, which represents the traffic flow characteristics of the node at different time steps;
[0013] The adjacency matrix is an N×N matrix that represents the spatial dependencies between nodes in the transportation network;
[0014] In the model, time series data and adjacency matrix are used as input, and the temporal and spatial dependencies are processed by the spatiotemporal module and the spatial attention module;
[0015] Furthermore, S3 specifically includes:
[0016] S31. Construct a spatial GAT block, including a multi-head attention mechanism, to extract spatial features;
[0017] S32. Construct a TCN block, which includes two convolution kernels, a long TCN block and a short TCN block, to respectively simulate long-term and short-term dependencies;
[0018] Furthermore, in said S31,
[0019] GAT-TCN (Graph Attention Network-Time Domain Convolutional Network) is used to extract spatiotemporal features respectively. In the spatial GAT (Graph Attention Network) block, a multi-head attention mechanism is adopted to enable the model to jointly learn spatial dependencies through multiple independent attention blocks;
[0020] In the TCN (time domain convolutional network) block, two different convolution kernels are used to extract time series features;
[0021] The final feature extraction result is output to the spatial attention module and the feedforward neural network. The self-attention mechanism is introduced to propagate GAT. For a given node feature set R represents a set of real numbers, F is the number of features of each node, and the attention coefficient e between any two nodes ij It is expressed as:
[0022]
[0023] The softmax function is used to normalize the attention coefficient into a form that is easy to compare:
[0024] a ij =softmax(LeakyReLU(e ij ))
[0025] The Softmax function is an activation function; the LeakyReLU function is an activation function used in neural networks, a uses a shared self-attention mechanism for all nodes, and e ij is the attention coefficient between any two nodes;
[0026] Then use the GCN convolution rule to update the model features of these coefficients:
[0027]
[0028] Where h is the node feature set, the activation function (e.g. ReLU), N is the number of nodes, a is the shared self-attention mechanism used by all nodes; W is the learned weight matrix; the input data is passed to the TCN layer, multiple 1D convolutional layers are applied, and the 1D deconvolution layer is used to restore it to the original size and output the predicted value;
[0029]
[0030]
[0031] Where y(t) represents the predicted value, x(t) represents the input value, w(k) represents the convolution kernel, and K represents the number of convolution kernels. is a special form of the Kronecker product representing the tensor product, d k Represents delay.
[0032] In S4, the spatial attention module is used to extract the spatial correlation of traffic data. The Attention block in the module uses a score function to determine the size of the attention weight and obtains three matrices Q, K, and V;
[0033] The attention mechanism scales the value of V based on the relationship between the key K and the query Q as follows:
[0034] Att(Q,K,V)=f A (Q,K)V
[0035] Where Att represents the attention mechanism, f A () is a score function, Q, K, V are three matrices;
[0036] To improve learning ability, Q, K, and V can be linearly projected into different subspaces. By calculating the value of a single attention head, multiple attention heads can be combined to obtain the input Y.
[0037] MultiHead(Q,K,V)=[Head1·Head2·...·Head0]W0
[0038]
[0039]
[0040] in are the weight matrices of Q, K, and V respectively, W0 is the output of the linear combination, Softmax is the activation function, MultiHead represents the multi-head attention mechanism, and Head represents the attention head;
[0041] The final output result is merged into LSTM, and the results of temporal convolution and spatial attention are fused through the gating mechanism;
[0042] LSTM uses a cell state to store long-term memory, and then cooperates with the gate mechanism to filter information, thereby achieving control over long-term memory. Specifically:
[0043] Input gate f t The role of is to determine the unit status that needs to be updated:
[0044] f t =σ g (W f x t +U f h t-1 +b f )
[0045] The neural layer of the forget gate is responsible for filtering and processing the information in the joint state at the previous moment, as shown in the following two formulas;
[0046] Specifically, the forget gate determines whether the information in the state at the previous moment is retained or discarded, and then updates the unit state;
[0047] i t =σ g (W i x t +U t h t-1 +b i )
[0048] s t =tanh(W c ·[W c x t +U c h t-1 ]+b c
[0049] On this basis, the functions of the input gate and the forget gate are further integrated to jointly affect the updating process of the unit state;
[0050] c t =f t ·c t-1 +i t ·s t
[0051] Finally, the function of the output value layer is to determine how much processed information to pass to the input threshold layer of the next time step and output the result;
[0052] o t =σ g (W o x t +U o h t-1 +b o )
[0053] The information update of the hidden state is shown below, which is achieved through a series of nonlinear transformations;
[0054] h t =o t ·tanh(c t )
[0055] Where W f , b f , W i , b i , W c , b c , W oand b o are the weights and offsets of each threshold layer, f t is the input gate, i t It's the forget gate, t is the output gate, c t is the memory state, h t is the hidden state, σ and tanh are the activation functions.
[0056] Further, in said S32,
[0057] Use GAT-TCN (Graph Attention Network-Time Domain Convolutional Network) to extract spatiotemporal features respectively;
[0058] In the TCN (time domain convolutional network) block, in order to effectively capture short-term and long-term dependencies, the present invention uses two different sizes of convolution kernels when designing the model, thereby achieving different modeling of these two types of time series features to separate short-term and long-term dependencies;
[0059] In order to capture short-term and long-term temporal dependencies more accurately, convolution kernels of different sizes are used in the design of the Short-Term Convolutional Network (ShortTCN) and the Long-Term Convolutional Network (Long TCN);
[0060] In the short TCN, we selected convolution kernels of size 1×1, 1×2, and 1×3 to simulate dependencies in shorter time scales;
[0061] In contrast, in the design of long TCN, in order to accommodate a wider range of temporal dependency modeling needs;
[0062] We used 1×1, 1×5 and 1×6 convolution kernels to effectively extract long-term time-dependent features, and then merged the results of the three convolutions to output the predicted value;
[0063] In S3, in order to improve the generalization performance of a single task, a feature capture method based on a multi-layer feedforward neural network (FNN) was used;
[0064] This method processes the single-task output of the spatiotemporal module through multiple fully connected hidden layers. Specifically, the FNN structure of the framework consists of three fully connected hidden layers;
[0065] The feedforward neural network obtains the final output a of the network through layer-by-layer information transmission. (l) ;
[0066] The entire network can be viewed as a composite function, taking vector x as the input a of the first layer (0) , the output a of the Lth layer(l) As the output of the entire function;
[0067] Let a (l) represents the activation vector of layer l, θ (l) Represents a matrix that maps the weights of layer l to layer l+1 Represents the weight from neuron K in layer l to neuron j in layer l+1;
[0068] The neural network that FNN has through three fully connected hidden layers can be constructed with the following matrix:
[0069]
[0070] Where σ is the activation function, and the value of σ ranges from 0 to 1;
[0071] The final prediction result of S4 is mapped back to the real space to obtain the predicted traffic flow data;
[0072] The loss between the actual value and the predicted value is calculated using mean absolute error (MAE), mean absolute percentage error (MAPE), and root mean square error (RMSE);
[0073] Mean absolute error:
[0074]
[0075] Root Mean Square Error:
[0076]
[0077] Mean absolute percentage error:
[0078]
[0079] where x i,j and Represent the actual value and predicted value respectively, N and Q are the number of samples.
[0080] According to an embodiment of the present disclosure, the step of obtaining historical traffic flow data of traffic at a prediction point includes: obtaining historical traffic flow data of traffic, including features such as flow rate and speed, constructing an adjacency matrix based on the relationship between road sections or nodes in the traffic network to represent the spatial dependency between nodes, and segmenting the historical data into time windows;
[0081] According to an embodiment of the present disclosure, the steps of preprocessing and feature reading the acquired traffic flow data include: acquiring time series data and adjacency matrix data, constructing dimensions (V, E, A) based on the received data, dividing the historical data by time windows, and normalizing the data; the feature of each node is a time series, which represents the traffic flow features of the node at different time steps; the adjacency matrix is an N×N matrix, which represents the spatial dependency between nodes in the traffic network; then the time series data and the adjacency matrix are used as input, and the time dependency and spatial dependency are processed by the spatiotemporal module and the spatial attention module; then GAT-TCN is used to extract spatiotemporal features respectively. In the spatial GAT block, a multi-head attention mechanism is used to enable the model to jointly learn spatial dependencies through multiple independent attention blocks; in the TCN block, two different convolution kernels are used to extract time series features. The final feature extraction results are output to the spatial attention module and the feedforward neural network;
[0082] According to the embodiment of the present disclosure, a feature capture method based on a multi-layer feedforward neural network (FNN) is used for the output feature extraction result; this method processes the single-task output of the spatiotemporal module through multiple fully connected hidden layers. Specifically, the FNN structure of the framework consists of three fully connected hidden layers, and the feedforward neural network obtains the final output of the network a through layer-by-layer information transmission. (l) ; For the output feature extraction results, the spatial attention module is used to extract the spatial correlation of traffic data. The Attention block in the module uses a score function to determine the size of the attention weight and obtains three matrices Q, K, and V. The attention mechanism scales the value of V based on the relationship between the key K and the query Q;
[0083] A computer-readable storage medium is provided, on which a computer program or code is stored, which, when loaded and executed by a processor, can implement the steps of the traffic flow prediction method.
[0084] The beneficial effects of the present invention are as follows: the present invention realizes accurate modeling of the spatiotemporal dependency of traffic flow by combining a graph neural network (GAT) and a temporal convolutional network (TCN), thereby effectively improving the accuracy of traffic flow prediction; utilizing a spatial attention module and a multi-head attention mechanism, not only can the complex spatial dependency in the traffic network be captured, but also short-term and long-term temporal dependency features can be effectively extracted, adapting to the dynamic change characteristics of traffic flow; introducing a multi-layer feedforward neural network (FNN) helps to improve the model's adaptability to different traffic scenarios and enhances the generalization performance.
[0085] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0087] Figure 1 A diagram of the overall structure of the invention;
[0088] Figure 2 is an overall flow chart of the invention;
[0089] Figure 3 Note the network structure diagram for the invented graph;
[0090] Figure 4 A structural diagram of the time series feature reading module of the invention;
[0091] Figure 5 Structural diagram of the invented long short-term memory network. DETAILED DESCRIPTION
[0092] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0093] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0094] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0095] See also Figure 1 to Figure 5 The present invention provides the following technical solution: a traffic flow prediction method based on attention-based long-short spatiotemporal graph neural network, characterized in that the steps of the method are as follows:
[0096] S1: Obtain historical traffic data, including flow, speed and other characteristics, build an adjacency matrix based on the road segments or node relationships in the traffic network to represent the spatial dependency between nodes, divide the historical data into time windows, and normalize the data;
[0097] S2: Input the traffic history data and adjacency matrix into the model;
[0098] S3: Extract temporal and spatial features respectively through the spatiotemporal module, and capture the similarities between related tasks through the feedforward neural network to improve generalization ability;
[0099] S4: Through the spatial attention module, a three-layer spatial attention block is used to realize the differential extraction of important features in the shared layer. The attention mechanism weighs the output according to the attention weight, and the output result is merged into the LSTM, and the information of the dense layer is processed by the LSTM to perform the task prediction, and finally the long and short flow predictions are integrated to obtain the final prediction result;
[0100] S1 specifically includes: obtaining time series data and adjacency matrix data, and constructing dimensions (V, E, A) based on the received data, where V represents the number of nodes, represents the road sections included in the transportation network, E represents the edge set connecting the nodes, and A represents the adjacency matrix.
[0101] S2 specifically includes: assuming that we have N nodes (i.e., road sections in the traffic network), the feature of each node is a time series, which represents the traffic flow characteristics of the node at different time steps;
[0102] The adjacency matrix is an N×N matrix that represents the spatial dependencies between nodes in the transportation network;
[0103] In the model, time series data and adjacency matrix are taken as input, and the temporal and spatial dependencies are processed by the spatiotemporal module and the spatial attention module.
[0104] Furthermore, S3 specifically includes:
[0105] S31. Construct a spatial GAT block, including a multi-head attention mechanism, to extract spatial features;
[0106] S32. Construct a TCN block, which includes two convolution kernels, a long TCN block and a short TCN block, to respectively simulate long-term and short-term dependencies;
[0107] Furthermore, in said S31,
[0108] GAT-TCN (Graph Attention Network-Time Domain Convolutional Network) is used to extract spatial and temporal features respectively. In the spatial GAT (Graph Attention Network) block, a multi-head attention mechanism is adopted to enable the model to jointly learn spatial dependencies through multiple independent attention blocks;
[0109] In TCN (time domain convolutional network), two different convolution kernels are used to extract time series features. The final feature extraction results are output to the spatial attention module and the feedforward neural network. The self-attention mechanism is introduced to propagate GAT. For a given node feature set R represents a set of real numbers, F is the number of features of each node, and the attention coefficient e between any two nodes ij It is expressed as:
[0110]
[0111] The softmax function is used to normalize the attention coefficient into a form that is easy to compare:
[0112] a ij =softmax(LeakyReLU(e ij ))
[0113] The Softmax function is an activation function, the LeakyReLU function is an activation function used in neural networks, a uses a shared self-attention mechanism for all nodes, and e ij is the attention coefficient between any two nodes;
[0114] Then use the GCN convolution rule to update the model features of these coefficients:
[0115]
[0116] Where h is the node feature set, σ is the activation function (such as ReLU), N is the number of nodes, a is the shared self-attention mechanism used by all nodes; W is the learned weight matrix;
[0117] Pass the input data to the TCN layer, apply multiple 1D convolutional layers, restore it to the original size using a 1D deconvolutional layer, and output the predicted value;
[0118]
[0119]
[0120] Where y(t) represents the predicted value, x(t) represents the input value, w(k) represents the convolution kernel, and K represents the number of convolution kernels. is a special form of the Kronecker product representing the tensor product, d k Represents delay.
[0121] S4 specifically includes: the spatial attention module is used to extract the spatial correlation of traffic data. The Attention block in the module uses a score function to determine the size of the attention weight and obtain three matrices Q, K, and V;
[0122] The attention mechanism scales the value of V based on the relationship between the key K and the query Q as follows:
[0123] Att(Q,K,V)=f A (Q,K)V
[0124] Where Att represents the attention mechanism, f A () is a score function, Q, K, V are three matrices;
[0125] To improve learning ability, Q, K, and V can be linearly projected into different subspaces. By calculating the value of a single attention head, multiple attention heads can be combined to obtain the input Y.
[0126] MultiHead(Q,K,V)=[Head1·Head2·...·Head0]W0
[0127]
[0128]
[0129] in are the weight matrices of Q, K, and V respectively, W0 is the output of the linear combination, Softmax is the activation function, MultiHead represents the multi-head attention mechanism, and Head represents the attention head;
[0130] The final output result is merged into LSTM, and the results of temporal convolution and spatial attention are fused through the gating mechanism. LSTM uses a cell state to store long-term memory, and then cooperates with the gating mechanism to filter information, thereby achieving control of long-term memory. Specifically:
[0131] Input gate f t The role of is to determine the unit status that needs to be updated:
[0132] f t =σ g (W f x t +U f h t-1 +b f )
[0133] The neural layer of the forget gate is responsible for filtering and processing the information in the joint state at the previous moment, as shown in the following two formulas. Specifically, the forget gate determines whether the information in the state at the previous moment is retained or discarded, and then updates the unit state;
[0134] i t =σ g (W i x t +U t h t-1 +b i )
[0135] s t =tanh(W c ·[W c x t +U c h t-1 ]+b c
[0136] On this basis, the functions of the input gate and the forget gate are further integrated to jointly affect the updating process of the unit state;
[0137] c t =f t ·c t-1 +i t ·s t
[0138] Finally, the function of the output value layer is to determine how much processed information to pass to the input threshold layer of the next time step and output the result;
[0139] o t =σ g (W o x t +U o ht-1 +b o )
[0140] The information update of the hidden state is shown below, which is achieved through a series of nonlinear transformations;
[0141] h t =o t ·tanh(c t )
[0142] Where W f , b f , W i , b i , W c , b c , W o and b o are the weights and offsets of each threshold layer, f t is the input gate, i t It's the forget gate, t is the output gate, c t is the memory state, h t is the hidden state, σ and tanh are the activation functions.
[0143] Further, in said S32,
[0144] Use GAT-TCN to extract temporal and spatial features respectively;
[0145] In the TCN block, in order to effectively capture short-term and long-term dependencies, we use two different sizes of convolution kernels when designing the model, thereby achieving different modeling of these two types of temporal features to separate short-term and long-term dependencies;
[0146] In order to capture short-term and long-term temporal dependencies more accurately, convolution kernels of different sizes are used in the design of the Short-Term Convolutional Network (ShortTCN) and the Long-Term Convolutional Network (Long TCN);
[0147] In the short TCN, we selected convolution kernels of size 1×1, 1×2, and 1×3 to simulate dependencies in shorter time scales;
[0148] In contrast, in the design of long TCN, in order to adapt to a wider range of time-dependent modeling needs, we use 1×1, 1×5 and 1×6 convolution kernels, so that we can effectively extract long-term time-dependent features, and then merge the results of the three convolutions to output the predicted value;
[0149] In S3, in order to improve the generalization performance of a single task, a feature capture method based on a multi-layer feedforward neural network (FNN) was used;
[0150] This method processes the single-task output of the spatiotemporal module through multiple fully connected hidden layers. Specifically, the FNN structure of the framework consists of three fully connected hidden layers. The feedforward neural network obtains the final output of the network through layer-by-layer information transmission. (l) ; The entire network can be viewed as a composite function, taking vector x as the input a of the first layer (0) , the output a of the Lth layer (l) As the output of the entire function;
[0151] Let a (l) represents the activation vector of layer l, θ (l) Represents a matrix that maps the weights of layer l to layer l+1 Represents the weight from neuron K in layer l to neuron j in layer l+1;
[0152] The neural network that FNN has through three fully connected hidden layers can be constructed with the following matrix:
[0153]
[0154] Where σ is the activation function, the value of σ ranges from 0 to 1, a (l) is the output, θ (l) is the matrix of corresponding weights.
[0155] To evaluate the model performance, the loss between the actual and predicted values was calculated using mean absolute error (MAE), mean absolute percentage error (MAPE), and root mean square error (RMSE);
[0156] Mean absolute error:
[0157]
[0158] Root Mean Square Error:
[0159]
[0160] Mean absolute percentage error:
[0161]
[0162] where x i,j and Represent the actual value and predicted value respectively, N and Q are the number of samples;
[0163] According to an embodiment of the present disclosure, the step of obtaining historical traffic flow data of traffic at a prediction point includes: obtaining historical traffic flow data of traffic, including features such as flow rate and speed, constructing an adjacency matrix based on the relationship between road sections or nodes in the traffic network to represent the spatial dependency between nodes, and segmenting the historical data into time windows;
[0164] According to an embodiment of the present disclosure, the steps of preprocessing and feature reading the acquired traffic flow data include: acquiring time series data and adjacency matrix data, constructing dimensions (V, E, A) based on the received data, dividing the historical data by time windows, and normalizing the data; the feature of each node is a time series, which represents the traffic flow features of the node at different time steps; the adjacency matrix is an N×N matrix, which represents the spatial dependency between nodes in the traffic network; then the time series data and the adjacency matrix are used as input, and the time dependency and spatial dependency are processed by the spatiotemporal module and the spatial attention module; then GAT-TCN is used to extract spatiotemporal features respectively. In the spatial GAT block, a multi-head attention mechanism is used to enable the model to jointly learn spatial dependencies through multiple independent attention blocks; in the TCN block, two different convolution kernels are used to extract time series features. The final feature extraction results are output to the spatial attention module and the feedforward neural network;
[0165] According to the embodiment of the present disclosure, a feature capture method based on a multi-layer feedforward neural network (FNN) is used for the output feature extraction result; this method processes the single-task output of the spatiotemporal module through multiple fully connected hidden layers. Specifically, the FNN structure of the framework consists of three fully connected hidden layers, and the feedforward neural network obtains the final output of the network a through layer-by-layer information transmission. (l) ; For the output feature extraction results, the spatial attention module is used to extract the spatial correlation of traffic data. The Attention block in the module uses a score function to determine the size of the attention weight and obtains three matrices Q, K, and V. The attention mechanism scales the value of V based on the relationship between the key K and the query Q;
[0166] According to the embodiments of the present disclosure, if the learning ability needs to be improved, Q, K, and V can be linearly projected into different subspaces, and the value of a single attention head can be calculated and multiple attention heads can be combined to obtain the input Y; the final output result is merged into LSTM, and the results of temporal convolution and spatial attention are fused through a gating mechanism; LSTM uses a cell state to store long-term memory, and then cooperates with a gating mechanism to filter information, thereby achieving control of long-term and short-term memory;
[0167] The above-mentioned embodiment is based on the traffic flow prediction method of the attention-based long-short spatiotemporal graph neural network, and only uses the division of functional modules as an example to illustrate. In practical applications, the implementation method of the function can be flexibly adjusted according to specific needs, that is, the relevant modules or steps can be further decomposed or recombined. For example, the functional modules mentioned in the article can be integrated into a single module or split into multiple sub-modules to realize all or part of the functions of the method; in addition, the naming of modules and steps in this article is only for the convenience of functional distinction and should not be regarded as a restrictive description of the content of the method;
[0168] According to the embodiments of the present disclosure, a storage device can store a set of designed program codes; when these codes are loaded and run by the corresponding processing unit or processor, an efficient traffic flow prediction method can be implemented; the method is centered on long-term and short-term time predictions and aims to improve the accuracy and real-time response capability of traffic flow prediction; these program codes are designed
[0169] The complexity and diversity of traffic data are fully considered, so that the processing unit can perform multiple key steps, including but not limited to the following: first, complete the collection and preprocessing of traffic data to ensure the integrity and consistency of input data; second, take long-term and short-term time prediction as the core, and fully integrate long-term and short-term traffic flow prediction; finally, use the constructed model to accurately and efficiently predict traffic flow, providing reliable technical support for traffic management and scheduling; in addition, these program codes adopt a modular design and have good scalability and adaptability. In practical applications, the code can be adjusted according to different needs to meet the needs of traffic flow prediction in specific scenarios. For example, short-term traffic flow prediction has the following disadvantages: in the event of an accident or abnormal situation (such as bad weather, traffic accidents), the prediction results may become invalid; the prediction time span is limited and cannot provide effective support for long-term planning; in order to improve real-time performance and accuracy, it may be necessary to introduce complex algorithms, which have high requirements for computing resources; long-term traffic flow prediction has the following disadvantages: as the time span increases, the cumulative error and uncertainty of the model will increase significantly; it is impossible to effectively capture short-term dynamic changes and fluctuations in traffic flow; long-term predictions often involve more influencing factors and complex multivariate models, which require high modeling and computing capabilities; so we integrate long-term and short-term traffic flow predictions, combining the advantages of the two predictions and making up for their shortcomings; this model not only enhances the practicality and adaptability of the system, but also provides a solid technical support and theoretical basis for the further development of the field of smart transportation.
[0170] The computer-readable storage medium of this embodiment stores a computer program, and when the processor executes the program, the traffic flow prediction method described in the embodiment can be implemented. Specifically, the storage medium may include an internal storage unit of the terminal, such as a hard disk or a memory; the hard disk is used to store a large amount of historical traffic data, prediction models and program files, while the memory is responsible for storing the currently executed program and temporary data to ensure efficient data processing; in addition, the storage medium may also be an external storage device, such as an external hard disk, a memory card, a secure digital card, etc. These devices provide a large storage space and are suitable for storing large-scale traffic data, long-term accumulated prediction models and other necessary files; in addition, the computer-readable storage medium of this embodiment may also be used in combination with an internal storage unit and an external storage device to provide higher storage capacity and flexible data processing capabilities; through the combination of internal storage and external storage, the system can meet the large data storage requirements while ensuring data reading and writing efficiency.
[0171] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A traffic flow prediction method based on attention-based long-short spatiotemporal graph neural network, characterized by: The method comprises the following steps: S1: Obtain historical traffic data, including flow, speed and other characteristics, build an adjacency matrix based on the road segments or node relationships in the traffic network to represent the spatial dependency between nodes, divide the historical data into time windows, and normalize the data; S2: Input the traffic history data and adjacency matrix into the model; S3: Extract temporal and spatial features through the spatiotemporal module, and improve generalization ability through the feedforward neural network; S4: Through the spatial attention module, a three-layer spatial attention block is used to realize the differentiated extraction of important features in the shared layer; the attention mechanism weighs the output according to the attention weight, and the output result is merged into the LSTM, and the information of the dense layer is processed by the LSTM to perform the task prediction, and finally the long and short flow predictions are integrated to obtain the final prediction result.
2. The traffic flow prediction method based on attention-based long-short spatiotemporal graph neural network according to claim 1 is characterized by: In S1, time series data and adjacency matrix data are obtained, and the dimensions are constructed according to the received data as (V, E, A), where V represents the number of nodes, represents the road sections included in the transportation network, E represents the edge set connecting the nodes, and A represents the adjacency matrix.
3. The traffic flow prediction method based on attention-based long-short spatiotemporal graph neural network according to claim 1 is characterized by: In S2, there are N nodes, i.e., road sections in the traffic network, and the feature of each node is a time series, which represents the traffic flow features of the node at different time steps; The adjacency matrix is an N×N matrix that represents the spatial dependencies between nodes in the transportation network; In the model, time series data and adjacency matrix are taken as input, and the temporal and spatial dependencies are processed by the spatiotemporal module and the spatial attention module.
4. The traffic flow prediction method based on attention length and short-term spatiotemporal graph neural network according to claim 1 is characterized by: The S3 specifically includes: S31. Construct a spatial GAT block, including a multi-head attention mechanism, to extract spatial features; S32. Construct a TCN block containing two types of convolution kernels, a long TCN block and a short TCN block, which are used to model long-term and short-term dependencies respectively.
5. The traffic flow prediction method based on attention length and short-term spatiotemporal graph neural network according to claim 4 is characterized by: In S31, a graph attention network-time domain convolutional network GAT-TCN is used to extract spatiotemporal features respectively, and a multi-head attention mechanism is adopted in the spatial graph attention network GAT block, so that the model can jointly learn spatial dependencies through multiple independent attention blocks; In the time domain convolutional network TCN block, two different convolution kernels are used to extract time series features; The final feature extraction results are output to the spatial attention module and the feedforward neural network; the self-attention mechanism is introduced to propagate GAT. For a given node feature set R represents a set of real numbers, F is the number of features of each node, and the attention coefficient e between any two nodes ij It is expressed as: The softmax function is used to normalize the attention coefficient into a form that is easy to compare: a ij =softmax(LeakyReLU(e ij )) The Softmax function is an activation function, the Leaky ReLU function is an activation function used in neural networks, a uses a shared self-attention mechanism for all nodes, and e ij is the attention coefficient between any two nodes; Then use the GCN convolution rule to update the model features of these coefficients: Where h is the node feature set, σ is the activation function (such as ReLU), N is the number of nodes, a is the shared self-attention mechanism used by all nodes; W is the learned weight matrix; Pass the input data to the TCN layer, apply multiple 1D convolutional layers, restore it to the original size using a 1D deconvolutional layer, and output the predicted value; Where y(t) represents the predicted value, x(t) represents the input value, w(k) represents the convolution kernel, and K represents the number of convolution kernels. is a special form of the Kronecker product representing the tensor product, d k Represents delay.
6. The traffic flow prediction method based on attention-based long-short spatiotemporal graph neural network according to claim 1 is characterized by: In S4, the spatial attention module is used to extract the spatial correlation of traffic data. The Attention block in the module uses a score function to determine the size of the attention weight and obtains three matrices Q, K, and V; The attention mechanism scales the value of V based on the relationship between the key K and the query Q as follows: That(Q,K,V)=f A (Q,K)V Where Att represents the attention mechanism, f A () is the score function, Q, K, V are three matrices; To improve learning ability, Q, K, and V are linearly projected into different subspaces. By calculating the value of a single attention head, multiple attention heads are combined to obtain the input Y. MultiHead(Q,K,V)=[Head1·Head2·...·Head0]W0 in are the weight matrices of Q, K, and V respectively, W0 is the output of the linear combination, Softmax is the activation function, MultiHead represents the multi-head attention mechanism, and Head represents the attention head; The final output result is merged into LSTM, and the results of temporal convolution and spatial attention are fused through the gating mechanism. LSTM uses a cell state to store long-term memory, and then cooperates with the gating mechanism to filter information, thereby achieving control of long-term memory. Specifically: Input gate f t The role of is to determine the unit status that needs to be updated: f t =σ g (W f x t +U f h t-1 +b f ) The neural layer of the forget gate is responsible for filtering and processing the information in the joint state at the previous moment, as shown in the following formula: i t =σ g (W i x t +U t h t-1 +b i ) S t =tanh(W c ·[W c x t +U c h t-1 ]+b c The forget gate determines whether the information in the previous state is retained or discarded, and then updates the unit state; The functions of the input gate and the forget gate are further integrated to jointly affect the updating process of the unit state; c t =f t ·c t-1 +i t ·s t The function of the output value layer is to determine how much processed information is passed to the input threshold layer of the next time step and output the result; the t =s g (W o x t +U o h t-1 +b o ) The information update of the hidden state is shown below, which is achieved through a series of nonlinear transformations; h t =o t ·tanh(c t ) Where W f , b f , W i , b i , W c , b c , W o and b o are the weights and offsets of each threshold layer, f t is the input gate, i t It's the forget gate, t is the output gate, c t is the memory state, h t is the hidden state, σ and tanh are the activation functions.
7. The traffic flow prediction method based on attention-based long-short spatiotemporal graph neural network according to claim 4 is characterized by: In S32, GAT-TCN is used to extract temporal and spatial features respectively; In the TCN block, in order to effectively capture short-term and long-term dependencies, two convolution kernels of different sizes are used in the design of the model to achieve different modeling of these two types of temporal features to separate short-term and long-term dependencies; To capture short-term and long-term temporal dependencies, different sizes of convolution kernels are used in the design of short-term and long-term temporal convolutional networks. In the short TCN, convolution kernels of sizes 1×1, 1×2, and 1×3 are selected to simulate dependencies in shorter time scales; In the design of long TCN, in order to adapt to a wider range of time-dependent modeling needs, 1×1, 1×5 and 1×6 convolution kernels are used to extract long-term time-dependent features, and then the results of the three convolutions are merged respectively to output the predicted value.
8. The traffic flow prediction method based on attention-based long-short spatiotemporal graph neural network according to claim 1 is characterized by: In S3, in order to improve the generalization performance of a single task, a feature capture method based on a multi-layer feedforward neural network (FNN) is used; This method processes the single-task output of the spatiotemporal module through multiple fully connected hidden layers. The FNN structure of the framework consists of three fully connected hidden layers. The feedforward neural network obtains the final output of the network through layer-by-layer information transmission. (l) ; The entire network is considered as a composite function, with vector x as the input a of the first layer (0) , the output a of the Lth layer (l) As the output of the entire function; Let a (l) represents the activation vector of layer l, θ (l) Represents a matrix that maps the weights of layer l to layer l+1 Represents the weight from neuron K in layer l to neuron j in layer l+1; The FNN has a neural network with three fully connected hidden layers built with the following matrices: Where σ is the activation function, the value of σ ranges from 0 to 1, a (l) is the output, θ (l) is the matrix of corresponding weights.
9. The traffic flow prediction method based on attention-based long-short spatiotemporal graph neural network according to claim 1 is characterized by: In S4, the final prediction result is mapped back to the real space to obtain the predicted traffic flow data; The loss between the actual value and the predicted value is calculated using the mean absolute error MAE, mean absolute percentage error MAPE and root mean square error RMSE; Mean absolute error: Root Mean Square Error: Mean absolute percentage error: where x i,j and Represent the actual value and predicted value respectively, N and Q are the number of samples.
10. A computer-readable storage medium, characterized in that: A computer program or code is stored thereon, which, when loaded and executed by a processor, can implement the steps of the method described in any one of claims 1 to 9.
Citation Information
Cited By
Multi-source heterogeneous data driven space-time modeling method for scenic spot traffic flow prediction
CN121354344A