A traffic flow prediction method considering spatio-temporal similarity

By using dynamic graph convolutional neural network in traffic flow prediction, combining dynamic graph construction units and spatiotemporal layer units, the problem that existing methods are difficult to respond to dynamic changes in nodes is solved, and higher traffic flow prediction accuracy is achieved.

CN115762160BActive Publication Date: 2025-05-27CHONGQING CITY MANAGEMENT COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211443967.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2025-05-27
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

The existing traffic flow prediction method based on graph convolutional neural networks is difficult to effectively respond to the dynamic changes in node dependencies, resulting in low prediction accuracy.

Method used

A dynamic graph convolutional neural network traffic flow prediction method considering spatial and temporal similarity is proposed. By constructing dynamic graphs and multiple spatiotemporal layer units, the time dependence and spatial dependence in traffic flow data are modeled using the combination of dynamic graph construction units and spatiotemporal layer units.

Benefits of technology

By considering space-time similarity and dynamic changes, the accuracy and accuracy of traffic flow prediction are improved, and the problems of model degradation and slow convergence are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762160B_ABST
    Figure CN115762160B_ABST
Patent Text Reader

Abstract

The present invention discloses a traffic flow prediction method, which includes: defining three time scales of adjacent periods, days, and weeks according to the similarity shown by traffic flow data on different time scales; constructing three prediction modules with the same functions for the three time scales of adjacent periods, days, and weeks; constructing the inputs of the prediction modules for different time scales; and obtaining the prediction values from the outputs of the prediction modules for different time scales. The method of the present invention can improve the accuracy of traffic flow prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent transportation, and particularly relates to a dynamic graph convolutional neural network traffic flow prediction method considering the spatio-temporal similarity of non-adjacent nodes. Background Art

[0002] Traffic flow prediction is one of the most critical issues in the research of intelligent transportation and smart cities. High-precision traffic flow prediction helps to improve the dynamic traffic management and intelligent service allocation capabilities of traffic resources, and plays an important role in the traffic management and smart travel of large cities. Traffic flow prediction is a typical spatio-temporal data mining problem. The change of traffic flow state is a non-linear change process coupled with multiple factors, and is susceptible to various factors such as drivers, driving road conditions, travel habits, and the functional attributes of road network nodes. Therefore, traffic flow data has complex dynamic spatio-temporal characteristics. Therefore, how to better model spatio-temporal correlation is crucial for improving the accuracy of traffic flow prediction. According to different prediction principles, the methods for realizing distance detection can be mainly divided into the following categories: statistical methods, machine learning methods, deep learning methods, etc. At present, methods based on statistics and machine learning are restricted by the modeling principle, resulting in limited modeling and expression capabilities for the complex spatio-temporal relationships in traffic flow data. Therefore, the prediction effect is still not very ideal. While deep learning methods can model multi-dimensional features, realize the approximation of complex functions, and mine complex non-linear relationships in data, showing great application prospects. Among the deep learning prediction methods, they can be divided into time series analysis methods, convolutional neural network analysis methods, and graph convolutional neural networks according to the modeling principle. Among them, the prediction based on graph convolutional neural network has the advantages of strong irregular image modeling and mining capabilities and less consumption of computing resources. And dynamic graphs can better adapt to the dynamic spatio-temporal change characteristics of traffic flow. Therefore, traffic flow prediction based on dynamic graph convolutional neural network is the main research direction at present and has good application prospects.

[0003] Existing prediction methods based on graph convolutional neural networks generally extract spatial features by aggregating node information in the spatial domain, and combine recurrent neural networks to extract time features and then output prediction values. However, existing research, whether it is based on a graph predefined by heuristic rules or a graph adaptively generated according to a similarity measurement mechanism based on a predefined graph, cannot well respond to the dynamic changes of node dependence relationships, resulting in poor prediction accuracy.

[0004] Therefore, a new traffic flow prediction method is needed to improve the prediction accuracy. Summary of the Invention

[0005] According to the above analysis, aiming at the deficiencies of the existing technology, the present invention utilizes deep learning algorithms related to graph convolutional neural networks to build a deep learning prediction model, proposes corresponding modeling strategies, and realizes dynamic graph convolutional neural network traffic flow prediction considering spatio-temporal similarity.

[0006] The technical solution of the present invention is realized as follows:

[0007] A traffic flow prediction method considering spatio-temporal similarity, including:

[0008] Step 1: Conduct statistical analysis on traffic flow data, and determine three time scales according to the correlation of traffic flow data at different time scales: adjacent time periods, days, and weeks;

[0009] Step 2: For the three time scales, construct three prediction modules with the same function but different time scales: adjacent time period prediction module, daily prediction module, and weekly prediction module. The prediction module is used to predict traffic flow. Among them, each prediction module includes a dynamic graph construction unit and two or more spatio-temporal layer units. The two or more spatio-temporal layer units are connected through residual connections. The dynamic graph construction unit is used to construct a dynamic graph reflecting the spatio-temporal dynamic changes of road network nodes. The spatio-temporal layer unit is used to model time dependence and spatial dependence. Among them, the dynamic graph is represented by the adjacency matrix DA t indicating, the adjacency matrix DA t is related to the node correlation distance metric adjacency matrix A pre-distance , the non-adjacent node similarity metric adjacency matrix A pre-similarity , the adaptive embedding adjacency matrix A adp , and the dynamic correlation adjacency matrix A dynamic . Among them, the non-adjacent node similarity metric adjacency matrix A pre-similarity is related to the Wasserstein distance between nodes;

[0010] Step 3: Construct the inputs of the prediction modules with different time scales. The inputs include adjacent time period data X a , daily cycle data X b , and weekly cycle data X c . Among them,

[0011]

[0012]

[0013]

[0014] N represents the number of nodes in the road network, F represents the feature dimension, the sampling frequency of the detection node is q times per day, the length of the sampling time series in one day is q, and the current detection time is t0 , T p represents the future time interval to be predicted, T a represents the time interval of the adjacent past period T b represents the time interval composed of the same time of the past few days T c represents the time interval composed of the same time of the past few weeks where N a , N b , N c is a positive integer greater than 1, representing the multiple of the prediction step length at different time scales;

[0015] Step 4: Determine the output of the prediction modules at different time scales and the predicted value of the traffic flow data according to the prediction modules at different time scales and the input.

[0016] Furthermore, construct a dynamic graph reflecting the spatio-temporal dynamic changes of road network nodes, including:

[0017] S11: Determine the adjacency matrix A of the node correlation distance metric pre-distance , which is expressed by the formula:

[0018]

[0019]

[0020] where E represents the set of traffic flows between any two nodes represents the weight from node v i to node v j , (v i , v j ) represents the traffic flow from node v i to node v j , and (v i , v j ) ≠ (v j , v i ), dist(v i , v j ) represents the Euclidean distance from node v i to v j , k(i) represents the set of k neighborhood nodes of node v i when dist(v j ) ≤ γ, γ is the set threshold, when the distance between two nodes exceeds the threshold γ, the weight coefficient is 0, indicating no interaction relationship i , represents the variance of the distance, and n represents the number of nodes; ​

[0021] S12: Determine the adjacency matrix A of non - adjacent node similarity metrics pre-similarity , which is expressed by the formula:

[0022]

[0023]

[0024]

[0025] where S ij represents the similarity score between node v i and node v j , the value of S ii is zero, E s represents the edge combination selected through the top - k mechanism, F(E s ,v i ,v r ) takes a value of 0 or 1, indicating whether the edge between node v i and node v r is selected, R represents the number of the selected edge set, Wasserstein(v i ,v j ) represents the Wasserstein distance between node v i and node v j ;

[0026] S13: Determine the adaptive embedding adjacency matrix A adp , which is expressed by the formula:

[0027]

[0028] where SoftMax(·) is the activation function, whose role is to map the input into the interval (0,1), represents the randomly generated learnable embedding matrix, N is the number of nodes, d' is the embedding dimension of each node and is much smaller than N,

[0029] S14: Determine the dynamic correlation adjacency matrix A dynamic , which is expressed by the formula:

[0030]

[0031] where,

[0032]

[0033]

[0034]

[0035] FC(·) represents a fully connected network, represents the node attributes after mapping, and the input data N represents the number of nodes, F represents the dimension of features, T represents the sequence length of the input data. After passing through the fully connected layer, the feature dimension is increased from F dimensions to D dimensions. d represents the one-dimensional convolutional dilation coefficient, DC d (·) represents the one-dimensional dilated convolution operation. d takes d1 or d2, and the length of the convolution kernel is K, w k is the convolution kernel, DC d1 (·) and DC d2 (·) respectively represent two one-dimensional dilated convolution operations with different dilation coefficients, represents the data at time t,

[0036] S15: Determine the adjacency matrix DA representing the dynamic graph t , which is expressed by the formula:

[0037] DA t = ReLu(A pre-distance + A pre-similarity + A adp + A dynamic ).

[0038] Furthermore, the number of spatio-temporal layer units in each prediction module is four.

[0039] Furthermore, the spatio-temporal layer unit includes a gated dilated causal convolution sub-unit and a graph convolution sub-unit with bidirectional random walks. Among them, the gated dilated causal convolution sub-unit is used to capture the duration relationship by expanding the receptive field of the node, control the information retention rate through the gating mechanism, and extract the temporal characteristics in the traffic flow data through the duration relationship and the retention rate; the graph convolution sub-unit with bidirectional random walks is used to extract the spatial features in the traffic flow data using the graph convolution with bidirectional random walks.

[0040] Furthermore, the gated dilated causal convolution sub-unit includes two or more layers of dilated causal convolution operations. By splicing the results of each layer of dilated causal convolution, the short-term and long-term temporal features in the original input data sequence are extracted. Among them, at the time t of splicing the l-th layer of dilated causal convolution, for the node v i at p l the output result of the dilated causal convolution in the dimension is expressed by the formula:

[0041]

[0042] where d l is the dilation coefficient of the l-th layer, is the convolution kernel with length W The element in, P l-1 and P l are the dimensions of the input and output, represents the sequence features of the output of each layer. l = 0, 1, ……, L represents the number of layers of the dilated causal convolution,

[0043] For the result of each layer of the dilated causal convolution perform concatenation represents the 0th layer, that is, the input data. After passing through the feature mapping learning function h(·), the concatenation result is expressed by the formula:

[0044]

[0045] where: * represents the dilated causal convolution and concatenation operations, Θ represents the set of all parameters in the concatenation result obtained after the dilated causal convolution concatenation operation, represents the input of any prediction module.

[0046] Furthermore, the method further includes:

[0047] After determining the concatenation result, introduce a gating mechanism to determine the output result y of the gated dilated causal convolution sub-unit tcn , which is expressed by the formula:

[0048]

[0049] where σ(·) scales the value after convolution between 0 and 1. The closer the value is to 1, the more detailed the retained information is, and the closer it is to 0, the lower the retention rate is; ⊙ represents the Hadamard product, and tanh(·) represents the activation function.

[0050] Furthermore, the defined bidirectional walk graph convolution is: for the input at time t on graph G Let the state transition equation be D p =(diag(Aj)) -1 A, and its reverse state transition equation is D r =(diag(Aj)) -1 A T , then the bidirectional convolution is calculated as follows:

[0051]

[0052] where, are the parameters of the filter f(θ), and its output is the prediction result where O is the feature dimension of its output data,

[0053] For the input of the bidirectional walk graph convolution of the hth layer spatio-temporal layer unit is at time t After determining the adjacency matrix representing the dynamic graph through the dynamic graph construction unit, determine the output of the random bidirectional walk convolution at the h-th layer and the t-th moment. It is expressed by the formula:

[0054]

[0055] Concatenate the random bidirectional walk convolutions at each moment of the h-th layer to obtain h ranges from 1 to H, and H represents the total number of spatio-temporal layer units in each prediction module.

[0056] Furthermore, the method further includes:

[0057] Through the input of the spatio-temporal layer of the (h - 1)-th layer and the operation result after passing through the spatio-temporal layer units of the (h - 1)-th layer Determine the output of the spatio-temporal layer of the h-th layer It is expressed by the formula:

[0058]

[0059] where h ranges from 1 to H,

[0060] Furthermore, step four specifically includes:

[0061] By passing the data X of adjacent time periods a , the daily cycle data X b and the weekly cycle data X c through the prediction modules of the different time scales, the outputs of the different time scales and

[0062] Weight the outputs of the different time scales and to determine the predicted value of the traffic flow data It is expressed by the formula:

[0063]

[0064] where: W c , W b , W a represent the weight coefficients of the outputs under different time scales, indicating the influence magnitude of each time scale on the predicted value, and ⊙ represents the multiplication of the corresponding elements of the matrices.

[0065] The beneficial effects of the present invention are:

[0066] 1. The present invention fully considers the similarity of traffic flow data at different time scales. By predicting the traffic flow at different time scales respectively and then comprehensively considering its influence, the final prediction is made more accurate.

[0067] 2. The dynamic graph construction unit of the prediction module disclosed in the present invention designs four adjacency matrices constructed in different attribute forms to express different types of similarity of traffic flow data, making the constructed dynamic graph more reasonable and accurate.

[0068] 3. The present invention adopts multiple spatio-temporal layers connected by residual connections to deepen the network depth to capture high-dimensional spatio-temporal features, thereby avoiding the problems of model degradation and slow convergence.

[0069] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent description, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 is a schematic diagram of the dynamic graph construction process;

[0071] Figure 2 is a schematic diagram of the gated dilated causal convolution structure;

[0072] Figure 3 is a framework diagram of the network model. DETAILED IMPLEMENTATION MANNER

[0073] Traffic flow data has complex dynamic spatio-temporal characteristics, which also contain rich spatio-temporal similarity information. From the spatio-temporal perspective, traffic flow data not only shows a high degree of similarity at different time scales, but also the traffic flow changes of non-adjacent nodes are similar due to the similarity of the functional areas (such as business districts, factories, etc.) where the road network nodes are located. If spatio-temporal similarity can be fully considered in traffic flow spatio-temporal data mining, the accuracy of traffic flow prediction will be improved.

[0074] The following combines the framework diagram of the dynamic graph convolutional neural network prediction model considering spatio-temporal similarity, that is Figure 3 , to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0075] Figure 3 is the framework diagram of the network model, combined withFigure 3 , the input of each Block is a tensor, where: N is the number of detection nodes, F is the data feature dimension (traffic flow, speed, and occupancy), and T is the number of time steps in the input sequence. It first performs feature dimension elevation through a fully connected layer, and then is input into the dynamic graph construction and spatio-temporal layer in parallel. After dynamic graph construction, an adjacency matrix representing the dynamic graph is obtained. After the spatio-temporal layer operation with residual connection, the result of the Block is output, and the outputs of the 3 Blocks are fused and calculated to obtain the prediction value where S refers to the number of time steps to be predicted. The following is a detailed description of the above four steps:

[0076] Step 1: Define three time scales of adjacent period, day, and week according to the similarity shown by traffic flow data at different time scales. Through statistical and correlation analysis of traffic flow data, it is found through visualization and correlation coefficient that traffic flow data shows high similarity at different time scales such as different time periods, days, and weeks in the time dimension. Therefore, three time scales of adjacent period, day, and week can be defined.

[0077] Step 2: Build three prediction modules with the same functions for the three time scales of adjacent period, day, and week: adjacent period prediction module, day prediction module, and week prediction module. The three prediction modules are used to predict traffic flow. Among them, the adjacent period prediction module is used to predict the traffic flow in the adjacent period of a specific moment, the day prediction module is used to predict the daily traffic flow, and the week prediction module is used to predict the weekly traffic flow. Among them, each prediction module includes a dynamic graph construction unit and two or more spatio-temporal layer units (hereinafter referred to as "spatio-temporal layers"), and the two or more spatio-temporal layer units are connected by residual connection. The dynamic graph construction unit is used to construct a dynamic graph reflecting the spatio-temporal dynamic changes of road network nodes, and the spatio-temporal layer unit is used to model time dependence and space dependence. In some embodiments, the dynamic graph is represented by an adjacency matrix DA t indicating that the adjacency matrix DA t is related to the node correlation distance metric adjacency matrix A pre-distancr , the non-adjacent node similarity metric adjacency matrix A pre-similarity , the adaptive embedding adjacency matrix A adp , and the dynamic correlation adjacency matrix A dynamic . Among them, the non-adjacent node similarity metric adjacency matrix A pre-similarity is related to the Wasserstein distance between nodes. In some embodiments, the number of spatio-temporal layers included in each prediction module is four. In some embodiments, the spatio-temporal layer is composed of gated dilated causal convolution (GTCN) and bidirectional random walk graph convolution (KDGCN), as shown in Figure 3 the network structure shown.

[0078] In some embodiments, the dynamic graph construction unit includes four adjacency matrices constructed in different attribute forms to express the spatiotemporal similarities of different types of traffic flow data, and then constructs a dynamic graph reflecting the spatiotemporal dynamic changes of road network nodes based on these adjacency matrices. Specifically, the distance measurement adjacency matrix A can be calculated by using traffic flow data according to the formula pre-distance Matrix, which uses the distance of adjacent nodes to characterize the spatial correlation of adjacent nodes; then calculate the adjacency matrix A of non-adjacent node similarity measure pre-similarity Matrix, which is used to characterize the spatiotemporal similarity between non-adjacent nodes; then the adaptive embedding adjacency matrix A is calculated adp Matrix, used to learn some similar inherent attribute characteristics in the road network (such as POI distribution, road grade, etc.); and obtain the dynamic correlation adjacency matrix A according to the dynamic correlation calculation process dynamic Matrix, and finally the dynamic adjacency matrix expression DA is obtained t . Figure 1 It is the process of building a dynamic graph, combined with Figure 1 ,Constructing a dynamic graph that reflects the spatiotemporal dynamic changes of road network nodes includes the following steps:

[0079] S11: Determine the node correlation distance metric adjacency matrix A pre-distance , as shown in formula (1)-(2):

[0080]

[0081]

[0082] Among them, E represents the set of traffic flow directions between any two nodes, Represents node v i To node v j The weight of (v i ,v j ) represents the slave node v i To node v j The traffic flow direction, and (v i ,v j )≠(v j ,v i ), dist(v i ,v j ) represents node v i to v j Euclidean distance, k(i) represents the distance between dist(v i ,v j )≤γ when node v i The set of k neighboring nodes, γ is the set threshold. When the distance between two nodes exceeds the threshold γ, the weight coefficient is 0, indicating that there is no interaction relationship. Represents the variance of the distance, n = N, both representing the number of nodes.

[0083] S12: Determine the adjacency matrix A for non - adjacent node similarity measurement pre-similarity , in some embodiments, the Wasserstein distance can be used to measure similarity, as shown in Equation (3):

[0084] S ij = exp(-Wasserstein(v i , v j )) (3)

[0085] Among them, S ij represents the similarity score between node v i and node v j . If the value of S ii is zero, the adjacency matrix for non - adjacent node similarity measurement can be as shown in Equation (4). To make A pre-similarity sparse and asymmetric, we perform row normalization operations using the top - k mechanism.

[0086]

[0087]

[0088] Among them, E s represents the edge combination selected by the top - k mechanism, L(E s , v i , v r ) takes a value of 0 or 1, indicating whether the edge between node v i and node v r is selected. R represents the number of the selected edge set, and Wasserstein(v i , v j ) represents the Wasserstein distance between node v i and node v j .

[0089] S13: Determine the adaptive embedding adjacency matrix A adp . Its construction process is to first randomly generate a learnable embedding matrix (Embedding matrix) where N is the number of nodes, d' is the embedding dimension of each node and is much smaller than N, and then calculate through Equation (6)

[0090]

[0091] Among them, SoftMax(·) is the activation function, which maps the input to the interval (0, 1). represents a randomly generated learnable embedding matrix, N is the number of nodes, and d' is the embedding dimension of each node, which is much smaller than N.

[0092] S14: Determine the dynamic correlation adjacency matrix A dynamic .

[0093] Some external emergencies, such as traffic accidents and temporary traffic measures, will cause dynamic changes in the spatial correlation relationship between nodes. This application uses a dynamic correlation adjacency matrix to capture this dynamic spatial correlation. The construction process is as follows:

[0094] Suppose there is a set of input data containing N nodes, F-dimensional features, and a sequence length of T, denoted as:

[0095]

[0096] First, to better mine the features of the original data, input it into a fully connected network to increase the feature dimension from F dimensions to D dimensions, that is, feature mapping. The calculation method is as follows in formula (7):

[0097]

[0098] Among them, FC(·) represents the fully connected network, represents the node attributes after mapping, and the input data

[0099] Next, perform a one-dimensional convolution operation (Dilated Convolution) on along the time dimension. The one-dimensional convolution dilation coefficient is d, and the convolution kernel length is K. The calculation method is as in formula (8):

[0100]

[0101] Among them, DC d (·) represents the one-dimensional dilated convolution operation, d represents the one-dimensional convolution dilation coefficient, the convolution kernel length is K, and w k is the convolution kernel, represents the data at time t,

[0102] Next, transform the N×D×T matrix into an N×D matrix, denoted as M, by stacking several one-dimensional dilated convolution operations with different dilation coefficients, as in formula (9):

[0103]

[0104] Among them, DC d1 (·) and DC d2 (·) respectively represent two one-dimensional dilated convolution operations with different dilation coefficients d1 or d2. That is, by replacing d in formula (8) with d1 or d2 respectively, DC d1 (·) and DC d2 (·) can be obtained.

[0105] Furthermore, A is calculated through formula (10) dynamic .

[0106]

[0107] S15: Determine the adjacency matrix DA of the dynamic graph t , which can be obtained by fusing the above-defined adjacency matrix, as shown in formula (11):

[0108] DA t = ReLu(A pre-distance + A pre-similarity + A adp + A dynamic ). (11)

[0109] Among them,

[0110] In some embodiments, the spatio-temporal layer units with residual connections in each prediction module mainly include gated dilated causal convolution (GTCN), graph convolution with bidirectional random walk (KDGCN), and residual connection. Among them, the gated dilated causal convolution captures long-term relationships by expanding the receptive field of nodes, and at the same time controls the information retention rate through the gating mechanism. The two work together to model time dependence, that is, to extract the time characteristics in traffic flow data through the duration relationship and the retention rate; the graph convolution with bidirectional random walk converges node information by performing random walks on the generated graph to model spatial dependence, that is, to extract the spatial features in traffic flow data.

[0111] From the time dimension, traffic flow data is essentially a time series data. Therefore, this application uses gated dilated causal convolution to mine the time features in traffic flow data. Among them, the time features in traffic flow sequences are mined by concatenating dilated causal convolutions, and the gating mechanism is combined to adaptively determine the proportion of retained information. The modeling process is specifically introduced below from the concatenated dilated causal convolution and the gating mechanism. The specific network structure can be seen in Figure 2 .

[0112] The gated dilated causal convolution subunit includes two or more layers of dilated causal convolution operations. By concatenating the results of each layer of dilated causal convolution, short-term and long-term temporal features in the original input data sequence are extracted. Among them, at the l-th layer of dilated causal convolution at time t for node v i in the p l dimension, the output result of the dilated causal convolution can be expressed by formula (12) as:

[0113]

[0114] where d l is the dilation coefficient of the l-th layer, is the element in the convolution kernel of length W , P l-1 and P l are the input and output dimensions, represents the sequence features output by each layer, and l = 0, 1, ……, L represents the number of layers of dilated causal convolution.

[0115] Next, the results of each layer of the dilated causal convolution are concatenated where represents the 0-th layer, that is, the input data. After passing through the feature mapping learning function h(·), the concatenation result can be expressed by formula (13) as:

[0116]

[0117] where: * represents the dilated causal convolution and concatenation operations, Θ represents the set of all parameters in the concatenation result obtained after the dilated causal convolution concatenation operation, represents the input of any prediction module.

[0118] The gating mechanism can well control the influence of the hidden state and the current state on the future. Therefore, after determining the concatenation result, the gating mechanism can be introduced to determine the output result y tcn of the gated dilated causal convolution subunit, which can be expressed by formula (14) as:

[0119]

[0120] where σ(·) scales the value after convolution between 0 and 1, where the value closer to 1 indicates that the retained information is more detailed, and the value closer to 0 indicates a lower retention rate; ⊙ represents the Hadamard product, and tanh(·) represents the activation function.

[0121] Random bidirectional walk graph convolution is used to extract spatial features in traffic flow data. The bidirectional walk graph convolution is defined as: the input at time t on graph G Let the state transition equation be D p =(diag(A j)) -1 A, and its reverse state transition equation is D r =(diag(Aj)) -1 A T , then the bidirectional convolution is calculated as the following formula (15):

[0122]

[0123] where are the parameters of the filter f(θ), and its output is the prediction result where O is the feature dimension of its output data.

[0124] For the graph convolution input of the bidirectional walk of the spatio-temporal layer unit of the h-th layer at time t After determining the adjacency matrix representing the dynamic graph through the dynamic graph construction unit, the output of the random bidirectional walk convolution at the h-th layer and time t can be determined It can be expressed by formula (16) as:[[]]

[0125]

[0126] Concatenating the random bidirectional walk convolutions at each time of the h-th layer gives h takes values from 1 to H, and H represents the total number of spatio-temporal layer units in each prediction module.

[0127] Furthermore, stacking multiple spatio-temporal layers deepens the network depth to capture high-dimensional spatio-temporal features, but at the same time of deepening the depth, it also brings model degradation and slower convergence speed. Therefore, when stacking spatio-temporal layers, the residual connection method is used to alleviate these problems. Through the input of the (g - 1)-th spatio-temporal layer and the operation result after passing through the spatio-temporal layer unit of the (h - 1)-th layer Determine the output of the h-th spatio-temporal layer It can be expressed by formula (17) as:[[]]

[0128]

[0129] where h takes values from 1 to H,

[0130] Step 3: Construct the input of the prediction module with different time scales, and the input includes the adjacent period data X a , the daily cycle data X b and the weekly cycle data X c .

[0131] The adjacent period data Xa is the data directly adjacent to the prediction time period, reflecting the impact of the traffic state in the adjacent time period on the future, and is also the most effective data for predicting the future traffic state. From the spatio-temporal analysis of traffic flow data, it can be seen that the traffic state in the future time period is strongly correlated with the historical data in the directly adjacent time period, and the change trend of the historical data directly affects the traffic state in the future time period.

[0132] Daily cycle data X b is composed of the historical time series at the same time of a certain day in a week and the prediction time period. From the spatio-temporal analysis of traffic flow data, it can be seen that affected by people's work and rest patterns, the traffic flow data during the morning and evening rush hours and the rest of the day are similar. Therefore, the long-term periodicity in traffic flow data can be mined by modeling the daily cycle data. Because there are still slight differences in the daily change patterns, especially the change trends and morning and evening rush hours on weekdays and weekends are inconsistent, it is also necessary to construct weekly cycle data to learn the trend of long-term time.

[0133] Weekly cycle data X c refers to the historical time series composed of the same time in several weeks and the prediction time period. From the spatio-temporal analysis of traffic flow data, it can be seen that there is an obvious weekly cycle pattern in traffic data. For example, the traffic patterns on Mondays of each week are similar, but there are certain differences from the traffic patterns on Saturdays. It is necessary to use this weekly scale to learn this difference in changes.

[0134] Let the sampling frequency of the detection node be q times per day, the length of the sampling time series of one day be q, and the current detection time be t 0 , T p represents the future time interval to be predicted, T a represents the time interval of the past adjacent time period, T b represents the time interval composed of the same time in the past few days, T c represents the time interval composed of the same time in the past few weeks, where N a , N b , N c is a positive integer greater than 1, representing the multiple of the prediction step length at different time scales. Then the adjacent time period data X a , daily cycle data X b and weekly cycle data X c can be expressed as:

[0135]

[0136]

[0137]

[0138] Among them, N represents the number of nodes in the road network, and F represents the feature dimension.

[0139] Step 4: Determine the outputs of the prediction modules at different time scales and the predicted values of the traffic flow data according to the prediction modules at different time scales and the input, specifically including:

[0140] By passing the data X in adjacent time periods a , the daily cycle data X b and the weekly cycle data X c through the prediction modules at different time scales, the outputs at different time scales and

[0141] are weighted for the outputs at different time scales and to determine the predicted values of the traffic flow data As shown in formula (12):

[0142]

[0143] Where: W c , W b , W a represent the weight coefficients of the outputs at different time scales, which reflect the influence magnitudes of each time scale on the predicted value, and ⊙ represents the multiplication of corresponding elements of the matrices.

Claims

1. A traffic flow prediction method considering the spatio-temporal similarity of non-adjacent nodes, characterized in that, it includes: Step 1: Conduct statistical analysis on traffic flow data, and determine three time scales according to the correlation of traffic flow data at different time scales: adjacent time periods, days, and weeks; Step 2: For the three time scales, construct three prediction modules with the same function but different time scales: a short-term prediction module, a daily prediction module, and a weekly prediction module. The prediction modules are used to predict traffic flow. Each prediction module includes a dynamic graph construction unit and two or more spatio-temporal layer units. The two or more spatio-temporal layer units are connected by residual connections. The dynamic graph construction unit is used to construct a dynamic graph reflecting the spatio-temporal dynamic changes of road network nodes. The spatio-temporal layer unit is used to model time dependence and spatial dependence. The dynamic graph is represented by the adjacency matrix DA t denoted as, and the adjacency matrix DA t is related to the node correlation distance metric adjacency matrix A pre-distanc , the non-adjacent node similarity metric adjacency matrix A pre-similarity , the adaptive embedding adjacency matrix A adp and the dynamic correlation adjacency matrix A dynamic . Among them, the non-adjacent node similarity metric adjacency matrix A pre-similarity is related to the Wasserstein distance between nodes; Step 3: Construct the inputs of the prediction modules with different time scales, where the inputs include the adjacent period data X a , the daily cycle data X b , and the weekly cycle data X c , where Let \(N\) denote the number of nodes in the road network, and \(F\) denote the data feature dimension. The data features include traffic flow, speed, and occupancy. The sampling frequency of the detection nodes is \(q\) times per day, the length of the sampling time series for one day is \(q\), and the current detection time is \(t\). 0 , \(T\) p denotes the future time interval to be predicted, \(T\) a denotes the time interval of the past adjacent period, , \(T\) b denotes the time interval formed by the same time of the past few days, , \(T\) c denotes the time interval formed by the same time of the past few weeks, where \(N\) a , \(N\) b , \(N\) c are positive integers greater than 1, representing multiples of the prediction step length; Step 4: Determine the outputs of the prediction modules at different time scales and determine the predicted values of the traffic flow data according to the prediction modules at the different time scales and the input.

2. The traffic flow prediction method considering the spatio-temporal similarity of non-adjacent nodes according to claim 1, characterized in that, construct a dynamic graph reflecting the spatio-temporal dynamic changes of road network nodes, including: S11: Determine the node correlation distance metric adjacency matrix A pre-distance , which is expressed by the formula as follows: Among them, E represents the set of traffic flow directions between any two nodes, Represents node v i To node v j The weight of (v i , v j ) represents the slave node v i To node v j The traffic flow direction, and (v i , v j )≠(v j , v i ), dist(v i , v j ) represents node v i to v j Euclidean distance, k(i) represents the distance between dist(v i , v j )≤γ when node v i The set of k neighboring nodes, γ is the set threshold. When the distance between two nodes exceeds the threshold γ, the weight coefficient is 0, indicating that there is no interaction relationship. represents the variance of the distance, n = N represents the number of nodes; S12: Determine the adjacency matrix A of the non - adjacent node similarity metric pre-similarity , which is expressed by the formula as follows: s ij = exp(-Wasserstein(v i , v j )) Among them, S ij represents the similarity score between node v i and node v j . The value of S ii is zero. E s represents the edge combination selected through the top-k mechanism. L(E s , v i , v r ) takes a value of 0 or 1, indicating whether the edge between node v i and node v r is selected. R represents the number of the selected edge sets. Wasserstein(v i , v j ) represents the Wasserstein distance between node v i and node v j ; S13: Determine the adaptive embedding adjacency matrix A adp , which is expressed by the formula as follows: Among them, SoftMax(·) is an activation function, whose role is to map the input into the interval (0, 1). represents a randomly generated learnable embedding matrix, where d′ is the embedding dimension of each node and is much smaller than N. S14: Determine the dynamic correlation adjacency matrix A dynamic , which is expressed by the formula as follows: wherein, FC(·) represents a fully connected network, represents the node attributes after mapping, input data represents the number of nodes, F represents the dimension of features, T represents the sequence length of the input data. After passing through the fully connected layer, the feature dimension is increased from F to D. d represents the dilation coefficient of the 1D convolution, DC d (·) represents the 1D dilated convolution operation. d takes d1 or d2, and the convolution kernel length is K, w k is the convolution kernel, DC d1 (·) and DC d2 (·) respectively represent two 1D dilated convolution operations with different dilation coefficients, represents the data at time t, S15: Determine the adjacency matrix DA representing the dynamic graph t , which is expressed by the formula as follows: DAt = ReLu(A pre-distance + A pre-similarity + A adp + A dynamic ).

3. The traffic flow prediction method considering the spatio-temporal similarity of non-adjacent nodes according to claim 1, characterized in that, the number of spatio-temporal layer units of each prediction module is four.

4. The traffic flow prediction method considering the spatio-temporal similarity of non-adjacent nodes according to claim 1, characterized in that, the spatio-temporal layer unit includes a gated dilated causal convolutional subunit and a graph convolutional subunit with bidirectional random walks. Among them, the gated dilated causal convolutional subunit is used to capture the duration relationship by expanding the receptive field of the node, control the information retention rate through a gating mechanism, and extract the time characteristics in the traffic flow data through the duration relationship and the retention rate; the graph convolutional subunit with bidirectional random walks is used to extract the spatial features in the traffic flow data using graph convolution with bidirectional random walks.

5. The traffic flow prediction method considering the spatio-temporal similarity of non-adjacent nodes according to claim 4, characterized in that, The door expansion causal convolution sub-unit includes two or more layers of dilated causal convolution operations. By splicing the results of each layer of dilated causal convolution, short-term and long-term temporal features in the original input data sequence are extracted. Among them, at the l-th layer of dilated causal convolution, at time t, node v i at p l The output result of the dilated causal convolution in the p dimension is expressed by the formula: where d l is the expansion coefficient of the l-th layer, is an element in the convolutional kernel of length K , P l-1 and P l are the dimensions of the input and output, represents the sequence feature of the output of each layer, l = 0, 1,......, L, representing the number of layers of the dilated causal convolution For the results of each layer of the dilated causal convolution perform concatenation Denote the 0th layer, i.e., the input data, after passing through the feature mapping learning function h(·) to make the concatenated result It is expressed by the formula as: Wherein: * represents the dilated causal convolution and concatenation operations, Θ represents the set of all parameters in the concatenation result obtained after the dilated causal convolution concatenation operation, represents the input of any prediction module.

6. The traffic flow prediction method considering the spatio-temporal similarity of non-adjacent nodes according to claim 5, characterized in that, it further includes: After determining the splicing result, a gating mechanism is introduced to determine the output result y of the gated dilation causal convolutional subunit tcn , which is expressed by the formula as follows: wherein, σ(·) scales the value after convolution between 0 and 1, where the value closer to 1 indicates that the retained information is more detailed, and the value closer to 0 indicates a lower retention rate; ⊙ represents the Hadamard product, and tanh(·) represents the activation function.

7. The traffic flow prediction method considering the spatio-temporal similarity of non-adjacent nodes according to claim 6, characterized in that, The defined bidirectional walk graph convolution is: for the input at time t on graph G Let the state transition equation be D p =(diag(AI)) -1 A, and its reverse state transition equation is D r =(diag(AI)) -1 A T , then the bidirectional convolution is calculated as follows: Among them, is the parameter of the filter f(θ), and its output is the prediction result where O is the feature dimension of its output data, The input of the graph convolution for the two-way random walk of the spatio-temporal layer unit at the h-th layer is At time t After determining the adjacency matrix representing the dynamic graph through the dynamic graph construction unit, determine the output of the random two-way walk convolution at the h-th layer and time t It is expressed by the formula: Concatenate the random bidirectional walk convolutions at each moment of the h-th layer to obtain h ranges from 1 to H, where H represents the total number of spatio-temporal layer units in each prediction module.

8. The traffic flow prediction method considering the spatio-temporal similarity of non-adjacent nodes according to claim 7, characterized in that, it further includes: Input through the (h - 1)-th spatio-temporal layer and the operation result of the (h - 1)-th spatio-temporal layer unit to determine the output of the h-th spatio-temporal layer which is expressed by the formula as: Among them, 9. The traffic flow prediction method considering the spatio-temporal similarity of non-adjacent nodes according to claim 1, characterized in that, the specific content of Step 4 includes: By using the data X in adjacent time periods a , the daily cycle data X b and the weekly cycle data X c through the prediction modules of the different time scales, the outputs of the different time scales are output and The output of the different time scales and are weighted to determine the predicted value of the traffic flow data which is expressed by the formula as follows: Among them: W c , W b , W a represents the weight coefficients output at different time scales, indicating the influence of each time scale on the predicted value, and ⊙ represents the multiplication of corresponding elements of the matrix.