Traffic flow prediction method of graph convolutional network based on double-wavelet transform
By introducing dual wavelet transform and graph convolution networks into the traffic flow prediction model, the shortcomings of existing models in identifying traffic flow pattern mutations and non-periodic event predictions are solved, and higher prediction accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510052682.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
Existing traffic flow prediction models are difficult to accurately identify sudden traffic flow pattern mutations, resulting in poor prediction performance when non-periodic events occur.
A graph convolution network based on dual wavelet transformation is adopted, and the time-frequency characteristics of traffic flow are analyzed using different wavelet basis to enhance the model's capture ability of periodic changes and the reliability of the prediction results.
The model's sensitivity to traffic flow mutations is significantly improved, and it can predict traffic conditions more accurately, providing stronger support when periodic events occur.
Smart Images

Figure CN119992844A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic flow prediction, and in particular to a traffic flow prediction method based on a graph convolutional network using dual wavelet transform. Background Art
[0002] With the rapid development of social economy and science and technology, the development of transportation infrastructure has made great progress. Using modern computer technology, establishing a traffic flow prediction model has become an important part of smart city construction. Timely and accurate traffic flow prediction can alleviate urban traffic pressure and avoid traffic congestion. Constructing a reasonable prediction model, making full use of the potential value of traffic information, providing guidance for drivers, enabling them to better choose departure time and the best route, and improving the utilization rate of urban road resources.
[0003] At present, traffic flow prediction algorithms are divided into two categories: classical statistical models and machine learning models. Classical statistical models mainly use mathematical statistics methods to analyze and process historical traffic data. The most commonly used method in machine learning models is the neural network model. Neural networks can handle highly complex nonlinear problems and can automatically learn and adjust features. The current mainstream model is a hybrid model, in which different neural networks are combined to capture different characteristics of traffic flow. It combines the spatial model to capture spatial features with the temporal model to capture temporal features. For time series data, Fourier transform and wavelet transform are usually used to analyze them in the time domain. In the field of transportation, wavelet transform is usually used because wavelet transform has the characteristics of multi-scale analysis, can perform local analysis on non-stationary signals, and has better capture effect on non-stationary signals than Fourier transform. However, the analysis results of traditional single wavelet transform will be affected by the selected wavelet basis function. Different wavelet basis functions will produce different analysis results, which limits the analysis of complex time-varying signals. Secondly, for complex non-stationary signals containing multiple frequency components, single wavelet transform may not be able to obtain good time-frequency analysis results at the same time. In addition, due to the locality of the wavelet basis function, the wavelet coefficients at the signal boundary will be seriously disturbed, making it difficult to accurately characterize the characteristics of the signal. When using methods such as wavelet transform for signal analysis and reconstruction, inaccuracies or distortions will appear near the signal boundary. This error usually occurs near the starting and ending points of the signal, because the wavelet transform lacks sufficient context information when processing these boundary points, resulting in large differences between the reconstructed signal and the original signal in these areas. Summary of the invention
[0004] In view of the deficiencies of the above-mentioned prior art, the present invention provides a traffic flow prediction method based on a graph convolutional network of dual wavelet transform. The method adopts a combination of dual wavelet transform and hybrid graph convolution to solve the problem that the existing traffic flow prediction model is difficult to accurately identify sudden traffic flow pattern mutations, resulting in poor prediction performance when non-periodic events occur; the present invention proposes an innovative dual wavelet model for analyzing the periodic characteristics of traffic flow. The model analyzes the time-frequency characteristics of traffic flow by using different wavelet bases. In this way, the model can more fully analyze the data characteristics of traffic. This correction mechanism significantly improves the sensitivity of the model to traffic flow mutations, enabling it to more accurately predict traffic conditions when periodic events occur, thereby providing more powerful support for traffic management and planning. The introduction of the dual wavelet transform model not only enhances the model's ability to capture periodic changes, but also improves the reliability of the prediction results.
[0005] The object of the present invention is achieved by the following technical solution: a traffic flow prediction method based on a graph convolutional network with dual wavelet transform is designed, characterized in that the method specifically comprises the following steps:
[0006] The first step is to obtain model training data
[0007] Step 1.1: First, obtain the one-dimensional historical traffic flow data of a certain length of time on the target traffic network section through N observation points in the section area S; then preprocess the historical traffic flow data, and then regard each observation point in the section as a node in the section, and select data values of the preprocessed historical traffic flow data at a fixed time interval M, each data value corresponds to a monitoring time point, and obtain historical time series data;
[0008] Step 1.2: Get the training dataset
[0009] First, the historical time series data in step 1.1 is divided into a training set and a validation set in chronological order according to the number of monitoring time points in the ratio of P:(1-P); then the training set and the validation set are initialized; a window of T+K time steps is used to slide the training set and the validation set respectively, and a number of sample data of T+K monitoring time points are obtained on each training set and the validation set to complete the initialization; among them, in one sample data, the data of the first T monitoring time points are used as the input data of the model, and the data of the last K monitoring time points are used as the reference value of the model output; the input data of the model The reference value y∈R output by the model K×N , T and K are the number of time steps, and N is the number of nodes;
[0010] Step 1.3: The spatiotemporal traffic flow matrix obtained by transposing all uninitialized data of the training set Calculate the correlation graph adjacency matrix A Cor , similarity graph adjacency matrix A Wasserstein , and calculate the distance graph adjacency matrix A of the road segment area S Cost , Where T′ is the time series length of the uninitialized training set data sample, and N represents the number of nodes;
[0011] The second step is to build a traffic flow prediction model;
[0012] The traffic flow prediction network model includes a dual wavelet transform module for data preprocessing, a first space-time module, a second space-time module, an attention fusion module and a data post-processing module; the first space-time module has the same basic structure as the second space-time module, and the parameters are not shared; the first space-time module includes a first-layer space-time model and a second-layer space-time model, and each layer of the space-time model includes a time convolution module for extracting time features and an improved dynamic graph convolution module for extracting spatial features; the time convolution modules of the two-layer space-time models have the same structure and the parameters are not shared; the improved dynamic graph convolution modules of the two-layer space-time models have the same structure and the parameters are not shared;
[0013] Model input data First, it is input into the dual wavelet transform module, which first transposes the input data to obtain data X, X∈R N×T ;
[0014] Then the data X is transformed by two different wavelet basis functions:
[0015]
[0016] Then the separated high-frequency and low-frequency signals are subjected to inverse wavelet transform:
[0017]
[0018] Finally, the separated low-frequency signal Splice to X l , the high frequency signal Splice to X h :
[0019]
[0020] It represents the high-frequency component and low-frequency component of data X decomposed by wavelet basis 1 through discrete wavelet transform. It represents the high-frequency component and low-frequency component of data X decomposed by wavelet basis 2 through discrete wavelet transform; and Decompose them respectively through the corresponding inverse wavelet transform to obtain high-frequency features And low frequency characteristics Finally, the high-frequency features and Merge and transpose to get high-frequency output X h , X h ∈R T ×N×2 ; The low frequency features and Merge and transpose to get the low-frequency output X l , X l ∈R T×N×2 ; After that, the high frequency output X h , low frequency output X l Both expand the feature dimension through linear transformation and convert it into high-frequency signal Low frequency signal d is the number of dimensions set for the model;
[0021] Then the high frequency signal Low frequency signal They are input into the first space-time module and the second space-time module respectively; for high-frequency signals First, The result is input into the time convolution module in the first layer of the spatiotemporal model of the first spatiotemporal module, and then input into its improved dynamic graph convolution module. The output is connected through residual Get the result Then The input is sent to the second-layer spatiotemporal model of the first spatiotemporal module, and then processed by the temporal convolution module of the second-layer spatiotemporal model and the improved dynamic graph convolution module. The output of the improved dynamic graph convolution module is connected through a residual connection. Get the high frequency output For low frequency signals First, The result is input into the temporal convolution module in the first layer of the spatiotemporal model of the second spatiotemporal module, and then input into its improved dynamic graph convolution module. The output is connected through residual Get the result Then The input is sent to the second-layer spatiotemporal model of the second spatiotemporal module, and then processed by the temporal convolution module of the second-layer spatiotemporal model and the improved dynamic graph convolution module. The output of the improved dynamic graph convolution module is connected through a residual connection. Get the output of low frequency component
[0022] The improved dynamic graph convolution module uses improved graph convolution to capture spatial correlation; the improved graph convolution uses a combination of static directed graph convolution and dynamic directed graph convolution to embed the static graph as a graph filter into the dynamic graph node; the static directed graph captures the explicit relationship between nodes, and the dynamic directed graph is used as a supplement to learn the hidden relationship between nodes;
[0023] The attention fusion module uses the attention mechanism to integrate high-frequency components Low frequency component The specific process of fusion is as follows:
[0024] In order to generate the predicted values of the data at K monitoring time points, the input data is first convolved to generate and
[0025]
[0026] W h,conv , W l,conv is a learnable parameter matrix; ★ represents the convolution operation; and and Input to the attention module:
[0027]
[0028] W k , W V , W Q is the learnable parameter matrix, is the high frequency component, is the low frequency component, d is and The dimension of the fused data y and the low-frequency component Perform residual connection to obtain the output of the final attention fusion module
[0029] The output of the attention fusion module Input to the data post-processing module; the data post-processing module has two layers, the first layer is a linear layer and a ReLU activation function layer, and the second layer is a linear layer. The principle formula is:
[0030]
[0031] Where W 1 、b 1 and W 2 、b 2 are the learnable parameters of the first and second layers respectively, and ReLU(*) is the activation function; According to the data The predicted value of the data at the next K monitoring time points is obtained after being input into the traffic flow prediction network model; squeeze(*) means converting the result of the second-layer linear transformation into the same dimension as the data to be predicted;
[0032] The third step is to train the traffic flow prediction network model;
[0033] Use Python's pytorch to build the traffic flow prediction network model in step 2, use the normal distribution random function to initialize the training parameters in the model, set the batch training size, the model's learning rate, the number of model training rounds, and the hyperparameters of the wavelet basis function of the dual wavelet transform;
[0034] The data of the first T monitoring time points of a sample data in the training set As the input data of the model, the data of the last K monitoring time points are used as the reference value of the model output; at the same time, the correlation graph adjacency matrix A in step 1.3 is Cor , similarity graph adjacency matrix A Wasserstein And the distance graph adjacency matrix A Cost Input it into the initialized traffic flow prediction network model to obtain the predicted value of the data of N nodes at the subsequent K time points output by the network model; input a batch of sample data, and use the predicted value output by the network model and the reference value of the subsequent K time points in the batch of sample data to calculate the training loss value;
[0035] The prediction result obtained by the traffic flow prediction network model is recorded as The reference value is y∈R K×N , then the prediction loss is:
[0036]
[0037] K represents the number of predicted monitoring time points, N is the number of sample nodes, and B is the number of sample data in each batch; this prediction loss function is calculated by calculating the absolute error between each predicted value and the actual value;
[0038] The trend component of the unconverged forecast after forecasting The trend component y after decoupling from the true value l Do Loss loss:
[0039]
[0040] y l , They are the real low-frequency component after decoupling the reference value of the data sample and the low-frequency component predicted by the model; The decoupled real low-frequency component value of the data of the nth node at the kth monitoring time point representing the reference value of the bth data sample, represents the low-frequency component value predicted by the second spatiotemporal model of the data of the nth node at the kth monitoring time point of the bth data sample;
[0041] In addition, in order to improve the generalization ability of the model and prevent the model from overfitting, the L2 regularization term is introduced; the overall training loss function is therefore defined as:
[0042]
[0043] Where λ is the regularization parameter, which is used to control the influence of the regularization term on the overall loss, ω i represents the parameters of the model, Z represents the total number of parameters of the traffic flow prediction network model, represents the square of the L2 norm of the parameter as a measure of complexity;
[0044] Then, based on the overall training loss value, the Adma optimizer is used to feed-forward update the training parameters in the model to complete the training of a batch of data; the initial learning rate is set to 0.001; the value of the optimal MAE indicator min_MAE is initialized to +∞; after each batch of training, the training parameters of the model when the previous batch of data samples are trained are used as the initial parameters for the next batch of data samples. The training process of a batch of data samples is repeated, and it is continuously iterated until the last batch of data samples in the training set is trained, completing the training of a round of data samples;
[0045] After each training round, the validation set is used to evaluate the performance of the model trained this time, and the MAE index is calculated. If the MAE index of the validation set of the current round is lower than the minimum MAE value recorded previously, the MAE index of the validation set of the current round and the model parameters of this time are recorded and saved, indicating that the model parameters of this training are better than the previously saved model parameters; the training parameters of the model when the previous round of training is completed are used as the initial parameters for the next round of training, and the training process of one round is repeated, and it is continuously iterated. The ReduceLROnPlateau learning rate scheduler is used. When the MAE index on the validation set exceeds 20 epochs and no longer decreases, the learning rate is reduced to 0.1 times the original value until it reaches the minimum learning rate of 0.000002 and no longer decreases; the training is terminated when the training round reaches the preset epoch round, and the network parameters of the training round with the lowest MAE index of the validation set are saved to obtain the optimal traffic flow prediction network model;
[0046] The fourth step is to predict the traffic flow;
[0047] Obtain historical traffic flow data of the same dimension of not less than M×T time length before the time point to be predicted for N nodes in the target traffic network section area S in the first step, and then preprocess it using the method in the first step to select data X at T monitoring time points before the time point to be predicted. 0 As the output data of the network model; the data X 0 , the adjacency matrix A of the correlation graph in step 1.3 Cor , similarity graph adjacency matrix A Wasserstein And the distance graph adjacency matrix A Cost The data is input into the optimal traffic flow prediction network model in the third step. The network model outputs the traffic flow prediction values of N nodes at the time point to be predicted and the K-1 time points thereafter. The prediction values can be used as guiding information for future traffic flow of the road network in real scenarios.
[0048] Compared with the prior art, the beneficial effect of the present invention is that: the traffic flow prediction method of the graph convolution network based on dual wavelet transform designed by the present invention first analyzes the original traffic flow signal by using two different wavelet basis functions, overcomes some limitations of single wavelet transform, and alleviates the boundary effect problem of wavelet transform. Dual wavelet transform can take advantage of the advantages of different wavelet basis functions according to the characteristics of the signal, so as to better adapt to the analysis needs of complex signals. Different wavelet bases have different frequency, time interval and scale characteristics and different characteristics, such as parity. The use of dual wavelet transform can enrich the data feature representation and improve the accuracy of traffic flow prediction. And the method of embedding predefined graphs in dynamic graphs has stronger stability, better convergence effect and higher prediction accuracy compared with traditional dynamic graph convolution. Using the correlation graph sampling strategy, the computational complexity of graph convolution is reduced without much influence on the model prediction effect. Moreover, the prediction result loss, the prediction trend component loss and the L2 regularization term of the model parameters are used as the overall training loss to ensure that the model not only focuses on reducing the prediction error during training, but also effectively avoids the overfitting problem. This combination of loss functions also enables the model to maintain good performance when faced with new data, thereby achieving more reliable predictive capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0050] Figure 1It is a schematic diagram of the process of predicting traffic flow (i.e., the schematic diagram of the process of the fourth step) in an embodiment of a traffic flow prediction method of a graph convolutional network based on dual wavelet transform of the present invention.
[0051] Figure 2 It is a schematic diagram of the principle and process of dual wavelet transform of an embodiment of a traffic flow prediction method of a graph convolutional network based on dual wavelet transform of the present invention. DETAILED DESCRIPTION
[0052] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application provided in conjunction with the drawings below is not intended to limit the scope of protection of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0053] How the technical solution of the present invention solves the above technical problems is described in detail below with reference to the accompanying drawings.
[0054] The present invention relates to a traffic flow prediction method based on a graph convolutional network with dual wavelet transform, which specifically comprises the following steps:
[0055] The first step is to obtain model training data
[0056] Step 1.1: First, obtain one-dimensional historical traffic flow data of a certain length of time for the target traffic network section through N observation points in the section area S; then clean the historical traffic flow data, remove invalid data and use interpolation to fill in the invalid data, and perform normalization to complete the preprocessing of the historical traffic flow data; then regard each observation point in the section as a node in the section, and select data values from the preprocessed historical traffic flow data at a fixed time interval M (i.e., the interval between two adjacent time steps, which is 5 minutes in this embodiment), and each data value corresponds to a monitoring time point to obtain historical time series data.
[0057] Data normalization is performed to ensure the comparability of different indicator data. By limiting the data to a specific range, the adverse effects that may be caused by singular sample data can be eliminated. A common method of normalization is standard normalization:
[0058]
[0059] Among them, x i It represents a data sequence of a certain length before a certain feature data of node i is normalized. represents the data value at the kth time point in the data sequence, express The normalized value, E(x i ) represents the data sequence x i The mean of the data values, Represents the data sequence x i The standard deviation of the data values.
[0060] Step 1.2: Get the training dataset
[0061] First, the historical time series data in step 1.1 is divided into a training set and a validation set in chronological order according to the number of monitoring time points in the ratio of P:(1-P); P is not less than 60%, and P=60% in this embodiment. Then the training set and the validation set are initialized. A window of T+K time steps is used to slide slice the training set and the validation set respectively, and a number of sample data of T+K monitoring time points are obtained on each training set and the validation set to complete the initialization. Among them, in one sample data, the data of the first T monitoring time points are used as the input data of the model, and the data of the last K monitoring time points are used as the reference value of the model output.
[0062] Model input data The reference value y∈R output by the model K×N , T and K are the number of time steps, and N is the number of nodes.
[0063] Step 1.3: The spatiotemporal traffic flow matrix obtained by transposing all uninitialized data of the training set Calculate the correlation graph adjacency matrix A Cor , similarity graph adjacency matrix A Wasserstein , and calculate the distance graph adjacency matrix A of the road segment area S Cost , Where T′ is the time series length of the uninitialized training set data sample, and N represents the number of nodes.
[0064] Calculate A for the training set in step 1.2 Cor , A Wasserstein .set up represents the traffic flow data of node p at the monitoring time point t. The historical traffic flow of node p from time point 0 (i.e., the first time point) to time point T′ (i.e., the last time point of the training set time series data, for the convenience of expression, a simplified notation method is used) can be expressed as in[*] T It is the transposition operation of the vector [*]. It represents the traffic flow data at node p at the detection time point T′. Then the historical traffic flows of its adjacent nodes (including node p in the target traffic network section area S, a total of N nodes) are combined to form a spatiotemporal traffic flow matrix as follows:
[0065]
[0066] in It represents the traffic flow of the target traffic network section area S at the monitoring time point t, and N is the number of observation points.
[0067] The spatiotemporal traffic flow matrix obtained from the training set Calculate the adjacency matrix A Cor , A Wasserstein .
[0068] The Pearson correlation coefficient is used to calculate the similarity of historical data information between any two nodes. i,j , represents the adjacency matrix A cor The correlation coefficient of the coordinates (i, j) in o, that is, the correlation coefficient of node i and node j, is as follows:
[0069]
[0070] A i,j The value range of is between [-1,1]. The larger the value, the closer the connection between the two nodes. If it is a positive value, there is a positive correlation between the two nodes, and if it is a negative value, there is a negative correlation. is the covariance value of the historical time series data corresponding to the two nodes. and is the standard deviation, The expression formula is as follows:
[0071]
[0072] Where E represents the expectation. According to formula (3), set A Cor o is the relevant adjacency matrix, A i,j represents the correlation coefficient between node i and node j in the target traffic network section area S. The relevant adjacency matrix A of area S can be obtained Cor o:
[0073]
[0074] A Wassersteino represents the similarity of traffic distribution between sensors, because the Wasserstein distance between node q and node p represents the minimum average distance required to move data from distribution q to distribution p. The formula is as follows:
[0075]
[0076] Among them, Π(p,q) represents the set of all possible joint distributions of distributions p and q. For each possible joint distribution γ, we can sample γ~Π(x,y) to get a sample x and y, and calculate the distance between the two samples ||xy||. So we can calculate the expected value E of the distance between the two samples under the joint distribution γ. x,y~γ [||xy||]. inf represents the infimum. The lower bound that can be obtained for this expected value in all possible joint distributions is the Wasserstein distance.
[0077] We can get the adjacency matrix A Wasserstein o:
[0078]
[0079] Where W(i,j) represents the similarity between node i and node j in the target traffic network segment area S, and the formula is derived from (7).
[0080] Adjacency matrix A of road segment area S Cost o is obtained based on the true distance between nodes, and C(i,j) represents the actual measured distance between node i and node j:
[0081]
[0082] After that, the above three adjacency matrices are initialized to obtain the related graph adjacency matrix A Cor , similarity graph adjacency matrix A Wasserstein , distance graph adjacency matrix A Cost :
[0083]
[0084] Tanh is an activation function. The values in the matrix are scaled to between [-1, 1] and then the diagonal matrix I is added.
[0085] According to the above calculation method, we can get the A of the road section area S. Cor , A Wasserstein adjacency matrix, and A Cost Adjacency matrix.
[0086] The second step is to build a traffic flow prediction model.
[0087] The traffic flow prediction network model includes a dual wavelet transform module for data preprocessing, a first space-time module, a second space-time module, an attention fusion module and a data post-processing module. The basic structure of the first space-time module is the same as that of the second space-time module, and the parameters are not shared. The first space-time module includes a first-layer space-time model and a second-layer space-time model. Each space-time model contains a temporal convolution (Temporal Connection Network, TCN) module for extracting temporal features and an improved dynamic graph convolution module for extracting spatial features. The temporal convolution modules of the two-layer space-time models have the same structure and do not share parameters; the improved dynamic graph convolution modules of the two-layer space-time models have the same structure and do not share parameters;
[0088] Model input data First, it is input into the dual wavelet transform module, which first transposes the input data to obtain data X, X∈R N×T ;
[0089] The formula for the first-order discrete wavelet transform is:
[0090]
[0091] In the formula, g and h represent low-pass filter and high-pass filter respectively, ★ represents convolution operation, and ↓2 represents downsampling. Similarly, the formula for the first-order discrete inverse wavelet transform is:
[0092]
[0093] Then the data X is transformed by two different wavelet basis functions:
[0094]
[0095] Then the separated high-frequency and low-frequency signals are subjected to inverse wavelet transform:
[0096]
[0097] Finally, the separated low-frequency signal Splice to X l , the high frequency signal Splice to X h :
[0098]
[0099] Represents data X from wavelet basis 1 through discrete wavelet transform (DWT 1 ) decomposed to obtain high-frequency components and low-frequency components, The data X is represented by wavelet basis 2 through discrete wavelet transform (DWT 2) decomposed to obtain high-frequency components and low-frequency components; and Through the corresponding inverse wavelet transform (IDWT 1 and IDWT 2 ) to decompose and obtain high-frequency features And low frequency characteristics Finally, the high-frequency features and Merge and transpose to get high-frequency output X h , X h ∈R T×N×2 ; The low frequency features and Merge and transpose to get the low-frequency output X l , X l ∈R T×N×2 After that, the high frequency output X h , low frequency output X l Both expand the feature dimension through linear transformation and convert it into high-frequency signal Low frequency signal d is the number of dimensions set for the model.
[0100] Then the high frequency signal Low frequency signal They are input into the first space-time module and the second space-time module respectively. First, The result is input into the time convolution module in the first layer of the spatiotemporal model of the first spatiotemporal module, and then input into its improved dynamic graph convolution module. The output is connected through residual Get the result Then The input is sent to the second-layer spatiotemporal model of the first spatiotemporal module, and then processed by the temporal convolution module of the second-layer spatiotemporal model and the improved dynamic graph convolution module. The output of the improved dynamic graph convolution module is connected through a residual connection. Get the high frequency output For low frequency signals First, The result is input into the temporal convolution module in the first layer of the spatiotemporal model of the second spatiotemporal module, and then input into its improved dynamic graph convolution module. The output is connected through residual Get the result Then The input is sent to the second-layer spatiotemporal model of the second spatiotemporal module, and then processed by the temporal convolution module of the second-layer spatiotemporal model and the improved dynamic graph convolution module. The output of the improved dynamic graph convolution module is connected through a residual connection. Get the output of low frequency component
[0101] The temporal convolution module adopts causal convolution operation. Causal convolution can be regarded as a special one-dimensional convolution, which uses a local window filter to slide on the time slice. Given an input x(t), given a filter f∈R W , then the causal convolution formula of x(t) at a given time step t is:
[0102]
[0103] The symbol ★ represents the convolution operation. x(t) represents the input signal at time t, f(w) is the weight or coefficient of the convolution kernel, and W is the size of the convolution kernel.
[0104] The improved dynamic graph convolution module uses improved graph convolution to capture spatial correlation. The improved graph convolution uses a combination of static directed graph convolution and dynamic directed graph convolution to embed the static graph as a graph filter into the dynamic graph node. The static directed graph captures the explicit relationship between nodes, and the dynamic directed graph supplements it to learn the hidden relationship between nodes. The working principle of the improved dynamic graph convolution module is as follows:
[0105] First, randomly initialize the learnable nodes DE 1 and DE 2 , take the output X' of the temporal convolution module as input, X'∈R T×N×d , X emb ∈R N×T :
[0106] X emb = reshape(X′ T W e ) (18)
[0107] I 1 =G(X emb , A Wasserstein )(19)
[0108] I 2 =G(X emb , A Cost ) (20)
[0109]
[0110]
[0111] Where T refers to the number of time steps, N refers to the number of nodes, and d refers to the dimension size. e Refers to the weight matrix, and reshape refers to the operation of adjusting the dimension. emb ∈R N×T, Formulas (19) and (20) represent transferring traffic flow data features to dynamic graph nodes by means of graph convolution, and learning traffic flow data features by dynamic graph. Where G represents graph convolution operation. DE 1 and DE 2 Represents the nodes of the dynamic graph adjacency matrix, a e is a hyperparameter, representing the scaling factor. Wasserstein For the training set The similarity graph adjacency matrix calculated from the data, A Cost is the distance graph adjacency matrix of the road section area S. The adjacency matrix is determined by the node layout. A of each data sample Wasserstein , A Cost The same. 1 and I 2 is a graph filter with predefined graph information, ○ represents the Hadamard product, E 1 and E 2 For initialization node, E 1 、E 2 ∈R N×N :
[0112]
[0113] In the formula, Emb represents word vector embedding, and [1,2,3,…,N] is embedded into word vector and initialized to an N×N matrix.
[0114] A Cost It represents the actual distance between sensors, and generates graph filter I through graph convolution G 1 and I 2 , E 1 and E 2 To initialize the node, initialize the node E 1 、E 2 Through the image filter I 1 ,I 2 Embed predefined graph features.
[0115] For the graph convolution operation (G) in formulas (19) and (20), the correlation graph sampling strategy is adopted to reduce the computational complexity. Cor For sampling, A cor It represents the correlation between nodes. Given the number of nodes N, the graph sampling strategy is to select the k nodes that are most relevant to each node:
[0116] index=topk(A cor , k) (24)
[0117]
[0118] represents rounding down; index represents the index selected by the graph sampling strategy, topk(A cor , k) is based on A cor Select the top k most relevant nodes for each node; Indicates the original image A i The masked graph structure, A i It means A Wasserstein or A Cost According to the similarity graph adjacency matrix A Wasserstein The process for taking samples is similar.
[0119] The formula for the graph convolution G operation in formulas (19) and (20) is:
[0120]
[0121] α is a hyperparameter, X emb is the input data of the graph convolution operation, X emb ∈RT ×N , D i express The degree matrix of .
[0122] Then use the node similarity to construct the dynamic adjacency matrix graph A D :
[0123]
[0124] α s is a hyperparameter. In order to optimize the generation of the dynamic adjacency matrix, tanh and ReLU are activation functions.
[0125] Then A D , X emb , A cost Output to two layers of mixed graph convolution layer, the mixed graph convolution operation adopts the dynamic graph A D Convolution and Static Graphs A cost Convolution combination method:
[0126] GCN(X emb , A D )=δX emb +βA D D D -1 X emb +γA Cost D Cost -1 X emb (28)
[0127] X 1 =GCN(Xemb ,A D ) (29)
[0128] X 2 =GCN(X 1 ,A D ) (30)
[0129] Where δ, β, and γ represent hyperparameters; D D It means A D The degree matrix, D Cost It means A Cost The degree matrix of ; X1 and X2 represent the outputs of the first and second hybrid graph convolution layers, respectively. Static image A cost That is, the distance graph adjacency matrix A Cost .
[0130] Finally, the input X of the first mixed graph convolution layer emb , the output X of the first mixed graph convolutional layer 1 The output X of the second mixed graph convolution layer 2 Spliced together and linearly transformed, the three features are fused together to obtain the output of the improved dynamic graph convolution module
[0131]
[0132] W x represents the weight matrix, Represents the output result of the improved dynamic graph convolution module.
[0133] High frequency signal output by dual wavelet transform module Low frequency signal Input into the first space-time module and the second space-time module respectively to obtain the high-frequency component Low frequency component
[0134] The attention fusion module uses the attention mechanism to integrate high-frequency components Low frequency component In this fusion process, the model first calculates the importance of each component relative to the current prediction task through the attention mechanism, and dynamically adjusts their influence at different levels, so that the model can focus more specifically on the importance of signals of different frequencies. When there are significant trends or periodic behaviors in the input signal, the attention mechanism will automatically enhance the weight of the corresponding low-frequency components, and dynamically adjust the degree of attention according to the changes in high-frequency signals to ensure that the model does not miss important information during comprehensive analysis. The specific process is:
[0135] In order to generate the predicted values of the data at K monitoring time points, the input data is first convolved to generate and
[0136]
[0137] W h,conv , W l,conv is a learnable parameter matrix. ★ represents the convolution operation. and Input to the attention module:
[0138]
[0139] W k , W V , W Q is the learnable parameter matrix, is the high frequency component, is the low frequency component, d is and By combining the fused data y and the low-frequency component Perform residual connection to obtain the output of the final attention fusion module
[0140] The output of the attention fusion module Input to the data post-processing module. The data post-processing module has two layers. The first layer is a linear layer and a ReLU activation function layer. The second layer is a linear layer. The principle formula is:
[0141]
[0142] Where W 1 , b 1 and W 2 , b 2 are the learnable parameters of the first and second layers respectively, and ReLU(*) is the activation function; For data based The predicted values of the data at the next K monitoring time points are obtained after being input into the traffic flow prediction network model. squeeze(*) means converting the result of the second layer linear transformation to be consistent with the dimension of the data to be predicted.
[0143] The third step is to train the traffic flow prediction network model.
[0144] Use Python's pytorch to build the traffic flow prediction network model in step 2, use the normal distribution random function to initialize the training parameters in the model, set the batch training size, the learning rate of the model, the number of rounds of model training, and the hyperparameters of the wavelet basis function of the dual wavelet transform.
[0145] The data of the first T monitoring time points of a sample data in the training set As the input data of the model, the data of the last K monitoring time points are used as the reference value of the model output; at the same time, the correlation graph adjacency matrix A in step 1.3 is Cor , similarity graph adjacency matrix A Wasserstein And the distance graph adjacency matrix A Cost Input it into the initialized traffic flow prediction network model to obtain the predicted value of the data of N nodes at the next K time points output by the network model. Input a batch of sample data, and use the predicted value output by the network model and the reference value of the next K time points in the batch of sample data to calculate the training loss value.
[0146] The prediction result obtained by the traffic flow prediction network model is recorded as The reference value is y∈R K×N , then the prediction loss is:
[0147]
[0148] K represents the number of monitoring time points for prediction, N is the number of sample nodes, and B is the number of sample data in each batch. This prediction loss function is calculated by calculating the absolute error between each predicted value and the actual value.
[0149] In addition to the overall prediction loss, in order to quantitatively measure the difference between the prediction result and the true trend data, and thus judge the prediction ability of the model, the trend component of the prediction that is not fused after the prediction is The trend component y after decoupling from the true value l Calculate the Loss loss:
[0150]
[0151] y l , They are respectively the real low-frequency component after decoupling the reference value of the data sample and the low-frequency component predicted by the model (i.e., the low-frequency component output by the second spatiotemporal module). The decoupled real low-frequency component value of the data of the nth node at the kth monitoring time point representing the reference value of the bth data sample, Represents the low-frequency component value predicted by the second spatiotemporal model of the data of the nth node at the kth monitoring time point of the bth data sample.
[0152] The decoupling of the reference value of the data sample is also implemented by using a dual wavelet transform module, and the two different wavelet basis functions are also the same, except that the dimension of the input data is K×N, and the dimension of the output data is also K×N.
[0153] In addition, in order to improve the generalization ability of the model and prevent the model from overfitting, the L2 regularization term is introduced. This term penalizes the complexity of the model and encourages the model to learn simpler functions, that is, parameters with higher generalization ability. The overall training loss function is therefore defined as:
[0154]
[0155] Where λ is the regularization parameter, which is used to control the influence of the regularization term on the overall loss, ω i represents the parameters of the model, Z represents the total number of parameters of the traffic flow prediction network model, Represents the square of the L2 norm of the parameters as a measure of complexity.
[0156] Then, based on the overall training loss value, the Adma optimizer is used to feed-forward update the training parameters in the model to complete the training of a batch of data. The initial learning rate is set to 0.001. The value of the optimal MAE indicator min_MAE is initialized to +∞. After each batch of training is completed, the training parameters of the model when the previous batch of data samples are trained are used as the initial parameters for the next batch of data samples. The training process of a batch of data samples is repeated and iterated continuously until the last batch of data samples in the training set is trained and a round of data samples is completed.
[0157] After each training round, the validation set is used to evaluate the performance of the model trained this time, and the MAE index is calculated. If the MAE index of the validation set of the current round is lower than the previously recorded minimum MAE value (i.e., min_MAE), the MAE index of the validation set of the current round and the model parameters of this time are recorded and saved, indicating that the model parameters of this training are better than the previously saved model parameters. The training parameters of the model at the completion of the previous round of training are used as the initial parameters for the next round of training. The training process of one round is repeated, and it is continuously iterated. The ReduceLROnPlateau learning rate scheduler is used. When the MAE index on the validation set exceeds 20 epochs and no longer decreases, the learning rate is reduced to 0.1 times the original value until it reaches the minimum learning rate of 0.000002 and no longer decreases. When the training round reaches the preset epoch round, the training is terminated, and the network parameters of the training round with the lowest validation set MAE index are saved to obtain the optimal traffic flow prediction network model.
[0158] The fourth step is to predict traffic flow.
[0159] Obtain historical traffic flow data of the same dimension of not less than M×T time length before the time point to be predicted for N nodes in the target traffic network section area S in the first step, and then preprocess it using the method in the first step to select data X at T monitoring time points before the time point to be predicted. 0 As the output data of the network model; the data X 0 , the adjacency matrix A of the correlation graph in step 1.3 Cor , similarity graph adjacency matrix A Wasserstein And the distance graph adjacency matrix A Cost The input is then fed into the optimal traffic flow prediction network model in the third step, and the network model outputs the traffic flow prediction values of N nodes at the time point to be predicted and the K-1 time points thereafter. The prediction value can be used as guidance information for future traffic flow of the road network in real scenarios.
[0160] The method of the present invention is verified based on the PEMS data set.
[0161] The PEMS03 dataset contains the traffic flow on highways in the Los Angeles area of California, including 358 sensors and data records within one week, and each sensor records once every 5 minutes. The PEMS08 dataset contains the traffic flow on highways in the Los Angeles area of California, including 170 sensors and data records within two months, and each sensor records once every 5 minutes. The adjacency matrix graphs of the datasets are all dynamically constructed from the historical training data, where the number of feature data d′=1. The dataset is divided in chronological order, 60% for training, 20% for validation, and 20% for testing.
[0162] The present invention adopts three widely used evaluation indicators, namely, mean absolute error (MAE), root mean square error (RMSE) and mean absolute percentage error (MAPE), which are defined as:
[0163]
[0164] In the formula, and i They represent the predicted value and true value of the road traffic flow of node i respectively, and n is the number of sensor nodes in the road network.
[0165] The method of the present invention is compared with the existing methods based on three evaluation indicators:
[0166] The historical time step T is set to 12, and the predicted time step K is set to 12. The Adam optimizer with an initial learning rate of 0.001 is used to train this model. The batch size of the dataset is set to 32. The number of epochs is set to 250, and the parameters of the wavelet transform are set to coif1 and sym2 for the PEMS03 dataset and coif1 and db5 for the PEMS08 dataset. The feature size d is 128, and the experiments are conducted on a GPU configured with GeForce RTX 3090ti.
[0167] Table 1 Experimental comparison results
[0168]
[0169] As can be seen from Table 1, the model proposed in the present invention shows better performance. This is because the dual wavelet transform used in the proposed model can more comprehensively capture the frequency changes of the sequence, enrich the feature representation of the sequence, and alleviate the boundary effect problem of the traditional wavelet transform. In addition, when processing graph data with rich structures and relationships, the hybrid graph convolution can capture the complex relationships and features in the data from multiple perspectives or dimensions, and capture the spatial characteristics of traffic by combining static graphs and dynamic graphs, which improves the model's ability to process complex data and prediction accuracy.
[0170] Any matters not described in the present invention are applicable to the prior art.
Claims
1. A traffic flow prediction method based on graph convolutional network with dual wavelet transform, characterized in that: The method specifically comprises the following steps: The first step is to obtain model training data Step 1.1: First, obtain the one-dimensional historical traffic flow data of a certain length of time on the target traffic network section through N observation points in the section area S; then preprocess the historical traffic flow data, and then regard each observation point in the section as a node in the section, and select data values of the preprocessed historical traffic flow data at a fixed time interval M, each data value corresponds to a monitoring time point, and obtain historical time series data; Step 1.2: Get the training dataset First, the historical time series data in step 1.1 is divided into a training set and a validation set in chronological order according to the number of monitoring time points in the ratio of P:(1-P); then the training set and the validation set are initialized; a window of T+K time steps is used to slide the training set and the validation set respectively, and a number of sample data of T+K monitoring time points are obtained on each training set and the validation set to complete the initialization; among them, in one sample data, the data of the first T monitoring time points are used as the input data of the model, and the data of the last K monitoring time points are used as the reference value of the model output; the input data of the model The reference value y∈R output by the model K×N , T and K are the number of time steps, and N is the number of nodes; Step 1.3: The spatiotemporal traffic flow matrix obtained by transposing all uninitialized data of the training set Calculate the correlation graph adjacency matrix A Cor , similarity graph adjacency matrix A Wasserstein , and calculate the distance graph adjacency matrix A of the road segment area S Cost , Where T′ is the time series length of the uninitialized training set data sample, and N represents the number of nodes; The second step is to build a traffic flow prediction model; The traffic flow prediction network model includes a dual wavelet transform module for data preprocessing, a first space-time module, a second space-time module, an attention fusion module and a data post-processing module; the first space-time module has the same basic structure as the second space-time module, and the parameters are not shared; the first space-time module includes a first-layer space-time model and a second-layer space-time model, and each layer of the space-time model includes a time convolution module for extracting time features and an improved dynamic graph convolution module for extracting spatial features; the time convolution modules of the two-layer space-time models have the same structure and the parameters are not shared; the improved dynamic graph convolution modules of the two-layer space-time models have the same structure and the parameters are not shared; Model input data First, it is input into the dual wavelet transform module, which first transposes the input data to obtain data X, X∈R N×T ; Then the data X is transformed by two different wavelet basis functions: Then the separated high-frequency and low-frequency signals are subjected to inverse wavelet transform: Finally, the separated low-frequency signal Splice to X l , the high frequency signal Splice to X h : It represents the high-frequency component and low-frequency component of data X decomposed by wavelet basis 1 through discrete wavelet transform. It represents the high-frequency component and low-frequency component of data X decomposed by wavelet basis 2 through discrete wavelet transform; and Decompose them respectively through the corresponding inverse wavelet transform to obtain high-frequency features And low frequency characteristics Finally, the high-frequency features and Merge and transpose to get high-frequency output X h , X h ∈R T×N×2 ; The low frequency features and Merge and transpose to get the low-frequency output X l , X l ∈R T×N×2 ; After that, the high frequency output X h , low frequency output X l Both expand the feature dimension through linear transformation and convert it into high-frequency signal Low frequency signal d is the number of dimensions set for the model; Then the high frequency signal Low frequency signal They are input into the first space-time module and the second space-time module respectively; for high-frequency signals First, The result is input into the time convolution module in the first layer of the spatiotemporal model of the first spatiotemporal module, and then input into its improved dynamic graph convolution module. The output is connected through residual Get the result Then The input is sent to the second-layer spatiotemporal model of the first spatiotemporal module, and then processed by the temporal convolution module of the second-layer spatiotemporal model and the improved dynamic graph convolution module. The output of the improved dynamic graph convolution module is connected through a residual connection. Get the high frequency output For low frequency signals First, The result is input into the temporal convolution module in the first layer of the spatiotemporal model of the second spatiotemporal module, and then input into its improved dynamic graph convolution module. The output is connected through residual Get the result Then The input is sent to the second-layer spatiotemporal model of the second spatiotemporal module, and then processed by the temporal convolution module of the second-layer spatiotemporal model and the improved dynamic graph convolution module. The output of the improved dynamic graph convolution module is connected through a residual connection. Get the output of low frequency component The improved dynamic graph convolution module uses improved graph convolution to capture spatial correlation; the improved graph convolution uses a combination of static directed graph convolution and dynamic directed graph convolution to embed the static graph as a graph filter into the dynamic graph node; the static directed graph captures the explicit relationship between nodes, and the dynamic directed graph is used as a supplement to learn the hidden relationship between nodes; The attention fusion module uses the attention mechanism to integrate high-frequency components Low frequency component The specific process of fusion is as follows: In order to generate the predicted values of the data at K monitoring time points, the input data is first convolved to generate and W h,conv , W l,conv is a learnable parameter matrix; ★ represents the convolution operation; and and Input to the attention module: W k , W V , W Q is the learnable parameter matrix, is the high frequency component, is the low frequency component, d is and The dimension of the fused data y and the low-frequency component Perform residual connection to obtain the output of the final attention fusion module The output of the attention fusion module Input to the data post-processing module; the data post-processing module has two layers, the first layer is a linear layer and a ReLU activation function layer, and the second layer is a linear layer. The principle formula is: Where W1, b1 and W2, b2 are the learnable parameters of the first and second layers respectively, and ReLU(*) is the activation function; According to the data The predicted value of the data at the next K monitoring time points is obtained after being input into the traffic flow prediction network model; squeeze(*) means converting the result of the second-layer linear transformation into the same dimension as the data to be predicted; The third step is to train the traffic flow prediction network model; Use Python's pytorch to build the traffic flow prediction network model in step 2, use the normal distribution random function to initialize the training parameters in the model, set the batch training size, the model's learning rate, the number of model training rounds, and the hyperparameters of the wavelet basis function of the dual wavelet transform; The data of the first T monitoring time points of a sample data in the training set As the input data of the model, the data of the last K monitoring time points are used as the reference value of the model output; at the same time, the correlation graph adjacency matrix A in step 1.3 is Cor , similarity graph adjacency matrix A Wasserstein And the distance graph adjacency matrix A Cost Input it into the initialized traffic flow prediction network model to obtain the predicted value of the data of N nodes at the subsequent K time points output by the network model; input a batch of sample data, and use the predicted value output by the network model and the reference value of the subsequent K time points in the batch of sample data to calculate the training loss value; The prediction result obtained by the traffic flow prediction network model is recorded as The reference value is y∈R K×N , then the prediction loss is: K represents the number of predicted monitoring time points, N is the number of sample nodes, and B is the number of sample data in each batch; this prediction loss function is calculated by calculating the absolute error between each predicted value and the actual value; The trend component of the unconverged forecast after forecasting The trend component y after decoupling from the true value l Do Loss loss: y l , They are the real low-frequency component after decoupling the reference value of the data sample and the low-frequency component predicted by the model; The decoupled real low-frequency component value of the data of the nth node at the kth monitoring time point representing the reference value of the bth data sample, represents the low-frequency component value predicted by the second spatiotemporal model of the data of the nth node at the kth monitoring time point of the bth data sample; In addition, in order to improve the generalization ability of the model and prevent the model from overfitting, the L2 regularization term is introduced; the overall training loss function is therefore defined as: Where λ is the regularization parameter, which is used to control the influence of the regularization term on the overall loss, ω i represents the parameters of the model, Z represents the total number of parameters of the traffic flow prediction network model, represents the square of the L2 norm of the parameter as a measure of complexity; Then, based on the overall training loss value, the Adma optimizer is used to feed-forward update the training parameters in the model to complete the training of a batch of data; the initial learning rate is set to 0.001; the value of the optimal MAE indicator min_MAE is initialized to +∞; after each batch of training, the training parameters of the model when the previous batch of data samples are trained are used as the initial parameters for the next batch of data samples. The training process of a batch of data samples is repeated, and it is continuously iterated until the last batch of data samples in the training set is trained, completing the training of a round of data samples; After each training round, the validation set is used to evaluate the performance of the model trained this time, and the MAE index is calculated. If the MAE index of the validation set of the current round is lower than the minimum MAE value recorded previously, the MAE index of the validation set of the current round and the model parameters of this time are recorded and saved, indicating that the model parameters of this training are better than the previously saved model parameters; the training parameters of the model when the previous round of training is completed are used as the initial parameters for the next round of training, and the training process of one round is repeated, and it is continuously iterated. The ReduceLROnPlateau learning rate scheduler is used. When the MAE index on the validation set exceeds 20 epochs and no longer decreases, the learning rate is reduced to 0.1 times the original value until it reaches the minimum learning rate of 0.000002 and no longer decreases; the training is terminated when the training round reaches the preset epoch round, and the network parameters of the training round with the lowest MAE index of the validation set are saved to obtain the optimal traffic flow prediction network model; The fourth step is to predict the traffic flow; Obtain historical traffic flow data of the same dimension of not less than M×T time length before the time point to be predicted for N nodes in the target traffic network section area S in the first step, and then preprocess it using the method in the first step, and select the data X0 at T monitoring time points before the time point to be predicted as the output data of the network model; transform the data X0 and the correlation graph adjacency matrix A in step 1.3 into Cor , similarity graph adjacency matrix A Wasserstein And the distance graph adjacency matrix A Cost The data is input into the optimal traffic flow prediction network model in the third step. The network model outputs the traffic flow prediction values of N nodes at the time point to be predicted and the K-1 time points thereafter. The prediction values can be used as guiding information for future traffic flow of the road network in real scenarios.
2. According to claim 1, a traffic flow prediction method based on graph convolutional network with dual wavelet transform is characterized in that: In step 1.1 of the first step, the preprocessing of the historical traffic flow data specifically refers to cleaning the historical traffic flow data, eliminating invalid data, filling the invalid data with interpolation, and performing normalization.
3. According to claim 1, a traffic flow prediction method based on graph convolutional network with dual wavelet transform is characterized in that: In step 1.1 of the first step, M = 5 min.
4. The traffic flow prediction method based on graph convolutional network with dual wavelet transform according to claim 1 is characterized in that: In step 1.2 of the first step, P is not less than 60%.
5. The traffic flow prediction method based on graph convolutional network with dual wavelet transform according to claim 1 is characterized in that: In step 1.3 of the first step, the spatiotemporal traffic flow matrix obtained from the training set is The Pearson correlation coefficient is used to calculate the similarity of historical data information between any two nodes to obtain the relevant adjacency matrix A Cor o: A Wasserstein o represents the similarity of traffic distribution between sensors, because the Wasserstein distance between node q and node p represents the minimum average distance required to move data from distribution q to distribution p; The formula is as follows: Among them, Π(p, q) represents the set of all possible joint distributions of distributions p and q; for each possible joint distribution γ, we can sample γ~Π(x, y) to get a sample x and y, and calculate the distance between the two samples ||xy||; so we can calculate the expected value E of the distance between the two samples under the joint distribution γ x,y~γ [||xy||]; inf represents the infimum; the lower bound that can be obtained for this expected value in all possible joint distributions is the Wasserstein distance; We can get the adjacency matrix A Wasserstein o: Where W(i, j) represents the similarity between node i and node j in the target traffic network segment area S, which is obtained from (7); Adjacency matrix A of road segment area S Cost o is obtained based on the true distance between nodes, and C(i, j) represents the actual measured distance between node i and node j: After that, the above three adjacency matrices are initialized to obtain the related graph adjacency matrix A Cor , similarity graph adjacency matrix A Wasserstein , distance graph adjacency matrix A Cost : Tanh is an activation function. The values in the matrix are scaled to [-1, 1] and then the diagonal matrix I is added.
6. The traffic flow prediction method based on graph convolutional network with dual wavelet transform according to claim 1 is characterized in that: In the second step, the discrete wavelet transform is a first-order discrete wavelet transform, and the formula of the first-order discrete wavelet transform is: In the formula, g and h represent low-pass filter and high-pass filter respectively, ★ represents convolution operation, and ↓2 represents downsampling by 2; Similarly, the first-order discrete inverse wavelet transform formula is:
7. The traffic flow prediction method based on graph convolutional network with dual wavelet transform according to claim 1 is characterized in that: In the second step, the temporal convolution module uses a causal convolution operation; causal convolution can be seen as a special one-dimensional convolution that uses a local window filter to slide on a time slice; given an input x(t), given a filter f∈R W , then the causal convolution formula of x(t) at a given time step t is: Among them, the symbol ★ represents the convolution operation; x(t) represents the input signal at time t, f(w) is the weight or coefficient of the convolution kernel, and W is the size of the convolution kernel.
8. The traffic flow prediction method based on graph convolutional network with dual wavelet transform according to claim 1 is characterized in that: In the second step, the working principle of the improved dynamic graph convolution module is: First, the learnable nodes DE1 and DE2 are randomly initialized, and the output X′ of the temporal convolution module is used as input, X′∈R T ×N×d , X emb ∈R N×T : X emb =reshape(X′ T W e ) (18) I1=G(X emb ,A Wasserstein ) (19) I2=G(X emb ,A Cost ) (20) Where T refers to the number of time steps, N refers to the number of nodes, and d refers to the dimension size; W e refers to the weight matrix, reshape represents the operation of adjusting the dimension; G represents the graph convolution operation; DE1 and DE2 represent the nodes of the dynamic graph adjacency matrix, a e is a hyperparameter, indicating the scaling factor; I1 and I2 are graph filters with predefined graph information, It represents the Hadamard product, E1 and E2 are initialization nodes, E1, E2∈R N×N : Where Emb represents word vector embedding, [1, 2, 3, ..., N] is embedded into word vectors and initialized to an N×N matrix; For the graph convolution operations in formulas (19) and (20), the correlation graph sampling strategy is adopted to reduce the computational complexity; according to the correlation graph adjacency matrix A Cor For sampling, A cor Represents the correlation between nodes; given the number of nodes is N, the graph sampling strategy is to select the k nodes that are most relevant to each node: index=topk(A cor ,k) (24) represents rounding down; index represents the index selected by the graph sampling strategy, topk(A cor , k) is based on A cor Select the top k most relevant nodes for each node; Indicates the original image A i The masked graph structure, A i It means A Wasserstein or A Cost ; According to the similarity graph adjacency matrix A Wasserstein The process for taking samples is similar; The formula for the graph convolution G operation in formulas (19) and (20) is: α is a hyperparameter, X em b is the input data of the graph convolution operation, X emb ∈R T×N , D i express The degree matrix of Then use the node similarity to construct the dynamic adjacency matrix graph A D : α s is a hyperparameter. In order to optimize the generation of dynamic adjacency matrix, tanh and ReLU are activation functions; Then A D , X emb , A cost Output to two layers of mixed graph convolution layer, the mixed graph convolution operation adopts the dynamic graph A D Convolution and Static Graphs A cost Convolution combination method: GCN(X emb ,A D )=δX emb +βA D D D -1 X emb +YA Cost D Cost -1 X emb (28) X1=GCN(X emb ,A D ) (29) X2=GCN(X1,A D ) (30) Where δ, β, and γ represent hyperparameters; D D It means A D The degree matrix, D Cost It means A Cost degree matrix; X1 and X2 represent the outputs of the first and second mixed graph convolution layers, respectively; static graph A cost That is, the distance graph adjacency matrix A Cost ; Finally, the input X of the first mixed graph convolution layer emb The output X1 of the first hybrid graph convolution layer and the output X2 of the second hybrid graph convolution layer are concatenated together and linearly transformed to fuse the features of the three together to obtain the output of the improved dynamic graph convolution module. W x represents the weight matrix, Represents the output result of the improved dynamic graph convolution module.