A traffic flow prediction method and system with missing data

The traffic data is clustered and filled through the ONMF and GMF modules, combined with GCRNN learning the spatiotemporal characteristics of traffic flow, the problem of insufficient spatial correlation capture and missing data in the existing methods is solved, and a higher precision traffic flow prediction is achieved.

CN115700628BActive Publication Date: 2025-07-25HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211301263.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-24
Publication Date
2025-07-25
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

Existing traffic flow prediction methods fail to effectively capture the spatial correlation of traffic data, resulting in low prediction performance, especially in the absence of data, feature learning performance is affected, and lacks versatility and fine-grained characterization capabilities.

Method used

Orthogonal non-negative matrix decomposition (ONMF) cluster traffic data, generalized matrix decomposition (GMF) is used to fill in missing data, and combined with graph convolutional recurrent neural network (GCRNN) to learn the spatiotemporal characteristics of traffic flow, and irregular graph data are processed through graph Laplace decomposition, and GCN and GRU are fused to improve prediction accuracy.

Benefits of technology

It improves the accuracy and robustness of traffic flow prediction, can effectively learn spatiotemporal and spatial characteristics in the absence of data, improves the universality of the model and fine-grained characterization ability, and reduces the impact of missing data on prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115700628B_ABST
    Figure CN115700628B_ABST
Patent Text Reader

Abstract

The present invention discloses a traffic flow prediction method for data with missing values, including: obtaining a traffic data set of a certain area, which contains missing values, reconstructing the traffic data set into a traffic flow data matrix, inputting the traffic flow data matrix X into the Orthogonal Non-Negative Matrix Factorization (ONMF) module of a trained spatio-temporal prediction model to form K clusters, and filling the data in each cluster by using the Generalized Matrix Factorization (GMF) module of the spatio-temporal prediction model, so as to obtain a traffic flow data matrix after filling the data. Standardize the traffic flow data matrix after filling, and model the standardized traffic flow data matrix into a three-dimensional tensor according to the historical step length H and the prediction window W. Input the three-dimensional tensor into the Graph Convolutional Recurrent Neural Network (GCRNN) of the trained spatio-temporal prediction model to obtain the predicted data Y'. The present invention has universality in traffic prediction for data with missing values, and learns spatio-temporal features in a finer granularity to achieve more effective traffic flow prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of deep learning and intelligent transportation in artificial intelligence, and more specifically, relates to a traffic flow prediction method and system with missing data implemented using a Fine-grained Completion Graph Convolution Recurrent Network (abbreviated as FCGCRN). Background Art

[0002] In recent years, with the collection of massive data by sensors and monitoring systems, prediction tasks have been widely studied in various fields such as climate, finance, and transportation. As a classic application, traffic prediction is an essential part of an Intelligent Transportation System (ITS), and plays an important role in alleviating traffic congestion, reducing traffic accidents, and improving the quality of urban traffic services. Given historical traffic flow and existing road information, predicting the future state is crucial for traffic flow prediction. However, each future traffic flow data not only depends on the historical values of that traffic flow, but also depends on other traffic flows. At the same time, due to network jitter, equipment failures, etc., traffic data loss may occur when sensors collect traffic flow. Therefore, how to accurately predict the future state of traffic flow with missing data is a challenging problem.

[0003] Existing research on traffic flow prediction mainly includes three types of algorithms. The first is the statistical method, which assumes that each traffic flow is a stationary sequence and uses linear algorithms to fit traffic data, such as Historical Average (HA), Autoregressive Integrated Moving Average (ARIMA), and Gaussian process (GP); the second is the single neural network method, which uses the Recurrent Neural Network (RNN) and its variants Long Short Term Memory (LSTM) and Gated Recurrent Unit (GRU). This method can process long-range time series traffic data in less time; the third is the hybrid neural network method, which combines the Convolution Neural Network (CNN) or Graph Convolution Network (GCN) and the recurrent neural network RNN to capture the complex spatial dependence between traffic flows and the long-term dependence between single traffic sequences respectively.

[0004] However, the above existing traffic flow prediction methods all have some technical problems that cannot be ignored: First, traditional methods and single neural network methods only consider the characteristics of traffic flow data in the time dimension and do not explicitly model the interdependence between different time series, resulting in low prediction performance; Second, CNN in the hybrid neural network method encapsulates the interaction between traffic flows into a global hidden state and is limited to processing regular grid structures to capture spatial correlations, making its representation ability weak when dealing with non-grid structure spatial relationships, further affecting the prediction accuracy; Third, GCN in the hybrid neural network method depends on a predefined graph, making the model lack generality; Fourth, the hybrid method lacks a suitable parameter learning method, resulting in the inability to represent spatio-temporal correlations in a fine-grained manner, thus affecting the prediction accuracy; Fifth, the above three methods are highly sensitive to the missing of traffic data, resulting in the model being prone to introducing noise when learning features, thus reducing the prediction performance. Summary of the Invention

[0005] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a traffic flow prediction method and system with missing data, aiming to solve the technical problems that the existing traffic flow prediction methods cannot capture the spatial correlation of traffic data, resulting in the inability to truthfully reflect the future state of traffic flow; and the technical problems that the complex spatial correlation in non-Euclidean space cannot be characterized, resulting in the limitation of spatial representation to grid data; and the technical problems that the spatial representation is affected by the limitation of the predefined graph, resulting in the lack of generality of the hybrid neural network method; and the technical problems that the traffic flow cannot be characterized in fine granularity and the specific patterns of flow nodes cannot be captured, resulting in the impact on the accuracy of traffic flow prediction; and the technical problems that the data missing problem caused by network or device failures affects the feature learning performance, resulting in the impact on the learning of temporal and spatial correlation of the feature learning module.

[0006] To achieve the above object, according to one aspect of the present invention, a traffic flow prediction method with missing data is provided, including the following steps:

[0007] (1) Obtain a traffic data set of a certain area, which contains missing data, and reconstruct the traffic data set into a traffic flow data matrix.

[0008] (2) Input the traffic flow data matrix X obtained in step (1) into the orthogonal non-negative matrix factorization (ONMF) module of the trained spatio-temporal prediction model to form K clusters, and fill the data in each cluster using the generalized matrix factorization (GMF) module of the spatio-temporal prediction model to obtain the traffic flow data matrix after filling the data. For the traffic flow data matrix after filling perform normalization processing to obtain the normalized traffic flow data matrix and model the normalized traffic flow data matrix into a three-dimensional tensor according to the historical step H and the prediction window W.

[0009] (3) Input the three-dimensional tensor obtained in step (2) into the graph convolutional recurrent neural network (GCRNN) of the trained spatio-temporal prediction model to obtain the prediction data Y'.

[0010] Preferably, the traffic data is three-dimensional tensor data {time, node, traffic feature}, where time refers to the time when the traffic feature is collected at the node, the node refers to a single sensor, and the traffic features include vehicle speed feature, traffic flow feature, and number of people feature;

[0011] Preferably, in step (1), the traffic data set

[0012] where It represents the traffic flow data matrix of the nth node (i.e., sensor) set on all streets in this area at T moments, where n ∈ [1, N], t ∈ [1, T], T is any positive integer, and N is the total number of sensors deployed on all streets in this area, and there is indicating that the data is non - negative;

[0013] where is the c - th eigenvalue of the nth node at the t - th moment, where c ∈ [1, C], and C represents the types of traffic characteristics.

[0014] Preferably, in step (2), the process of inputting the traffic flow data matrix X obtained in step (1) into the ONMF module to form K clusters specifically includes:

[0015] (2 - 1) Initialize the matrix factors F and G of the traffic flow data matrix X to random values within (0, 1);

[0016] (2 - 2) According to the matrix factors F and G initialized in step (2 - 1) and using the update rules and update the observable data in the matrix. When the error of X - FG T converges, stop the iteration, thereby obtaining the updated matrix factor G;

[0017] (2 - 3) According to the updated matrix factor G obtained in step (2 - 2), cluster the traffic flow data matrix X into K clusters;

[0018] Preferably, in step (2), use the GMF module to fill the traffic flow data matrix X in each cluster to obtain the traffic flow data matrix after filling Perform normalization processing on the filled traffic flow data matrix to obtain the normalized traffic flow data matrix And according to the historical step H and the prediction window W, model the normalized traffic flow data matrix as a three - dimensional tensor This process includes the following sub - steps:

[0019] (2 - 4) In each cluster obtained in step (2 - 3), divide the traffic flow data matrix X into an observable data set and an unobservable data set, and reconstruct the observable data set to obtain the time vector v p ∈R m , node vector v q ∈R m and traffic flow vector v y ∈R m , and reconstruct the unobservable data set as the time vector v′ p∈R m′ and the node vector v' q ∈R m′ , where m represents the total number of observable data in the observable dataset, and m' represents the total number of unobservable data in the unobservable dataset;

[0020] (2-5) Input the two vectors v p and v q obtained in step (2-4) into the embedding layer of the GMF module to obtain the output of the embedding layer, namely the time matrix factor P and the node matrix factor Q:

[0021] P = e1(v p ),

[0022] Q = e2(v q ),

[0023] where P ∈ R m×a and Q ∈ R m×a are the time matrix factor and the node matrix factor respectively, a = 16 is the latent factor, and e1() and e2() represent the embedding functions, both of which use the torch.nn.Embedding() function in the Pytorch framework;

[0024] (2-6) Input the time matrix factor P and the node matrix factor Q obtained in step (2-5) into the decomposition layer of the GMF module to obtain the output result f(P, Q):

[0025] f(P, Q) = P ⊙ Q,

[0026] where ⊙ represents the element-wise product operation.

[0027] (2-7) Input the output result f(P, Q) obtained in step (2-6) into the filling layer of the GMF module, and the obtained output is the filling result g(P, Q):

[0028] g(P, Q) = a ott (W T (P ⊙ Q) + b),

[0029] where a out is the Relu activation function, and W and b represent the learnable weight and bias parameters in the GMF module respectively;

[0030] (2-8) Measure the filling result in step (2-7) using the mean squared error MSE to obtain the error value MSE between the traffic flow vector v y in step (2-4) and the filling result g(P, Q) in step (2-7). The calculation formula for this step is:

[0031]

[0032] (2-9) Update the learnable weights W and bias parameter b in the GMF module through the Adam optimizer;

[0033] (2-10) Repeat the above steps (2-8) to (2-9) until the MSE is less than the threshold or the number of training times reaches the preset number of epochs, so as to obtain the trained GMF module;

[0034] (2-11) Input the unobservable dataset obtained in step (2-4) into the GMF module trained in step (2-10) to obtain the traffic flow vector v′ y ∈R m′ ;

[0035] (2-12) According to the time vector v p under the observable dataset obtained in step (2-4), node vector v q , traffic flow vector v y and the time vector v′ p under the unobservable dataset, node vector v′ q and the traffic flow vector v′ y obtained in step (2-11), obtain the traffic flow data matrix after filling the data

[0036] (2-13) Standardize the traffic flow data matrix after filling the data in step (2-12) to obtain the standardized traffic flow data matrix

[0037]

[0038] where μ is the mean of the traffic flow data matrix , and σ is the standard deviation of ;

[0039] (2-14) Reconstruct the standardized traffic flow data matrix in step (2-13) into a three-dimensional tensor and a three-dimensional tensor Y∈R (T-H-W+1)×N×W .

[0040] Preferably, the GCRNN network is trained through the following steps:

[0041] (3-1) Divide the three-dimensional tensors and Y obtained in step (2-14) into a training set and a test set at a ratio of 6:4;

[0042] (3-2) Perform adaptive graph learning on the GCRNN through the parameter E to obtain the adjacency matrix

[0043] The calculation formula for this step is:

[0044]

[0045] where E ∈ R N×e is a learnable parameter matrix. The parameter E is initialized using the torch.FloatTensor() function in the Pytorch framework, and the optimal tuning results for the parameter e are 2 and 10;

[0046] (3-3) Input the data of the training set of the three-dimensional tensor obtained in step (3-1) at time t and the adjacency matrix obtained in step (3-2) into the graph convolutional neural network to obtain the graph convolutional result H of the Kth cluster K ∈ R N×h ;

[0047] The calculation formula for this step is:

[0048]

[0049] where G K ∈ R N×N is the Laplacian matrix of the Kth cluster, θ K ∈ R H×h and b K ∈ R h are the learnable parameters of the Kth cluster, and h is the number of neurons in the hidden layer of the graph convolutional recurrent neural network;

[0050] (3-4) Input the data of the training set of the three-dimensional tensor obtained in step (3-1) at time t into the recurrent neural network to obtain the representation result h at time t t ;

[0051] (3-5) Input the output h t obtained in step (3-4-4) into the two-dimensional convolutional layer of the GCRNN network (as shown in Figure 3 ) to obtain the final traffic flow prediction result Y' ∈ R N×W ;

[0052] The specific implementation of this step is as follows:

[0053] Y' = h t ★ f 1×h ,

[0054] Among them, the number of channels channel of the two-dimensional convolutional layer is 1, and the output W = 12, f 1×h indicates that the size of the convolutional kernel in the two-dimensional convolutional layer is (1, h), where h = 64;

[0055] (3 - 6) Use the L1 loss function to calculate the loss value O(Y, Y′) between the predicted result Y′ obtained in step (3 - 5) and the tensor Y obtained in step (2 - 14);

[0056] Specifically, the L1 loss function used in this step is:

[0057]

[0058] (3 - 7) Use the loss function in step (3 - 6) and the Adam optimizer in the Pytorch framework to iteratively update the learnable parameters E in step (3 - 2) and the learnable parameters from step (3 - 4 - 1) to (3 - 4 - 4) and ;

[0059] (3 - 8) Repeat the training process from step (3 - 6) to step (3 - 7) until the number of iterations in step (3 - 7) (100 in the present invention) or the loss value O(Y, Y′) in step (3 - 6) is less than the set threshold, and then end the training to obtain a preliminarily trained GCRNN model;

[0060] (3 - 9) Use the test set obtained in step (3 - 1) to verify the preliminarily trained GCRNN model in step (3 - 8) until the prediction error reaches the optimum, thereby obtaining a trained spatio-temporal prediction model.

[0061] Preferably, step (3 - 4) includes the following sub-steps:

[0062] (3 - 4 - 1) Input the data of the training set of the three-dimensional tensor obtained in step (3 - 1) at time t the adjacency matrix obtained in step (3 - 2) and the representation result h at the previous time t - 1 t-1 into the GCRNN to obtain the update gate z at time t t ∈R N×h ;

[0063] Specifically, the calculation formula for this step is:

[0064]

[0065] where h0 ∈ R N×hIndicates the initial state, which is a matrix composed entirely of 0s. [,] represents the concatenation operation. and is the learnable parameter of the update gate z at time t in the K-th cluster. t σ(·) is the sigmoid activation function.

[0066] (3-4-2) Input the data of the training set of the three-dimensional tensor obtained in step (3-1) at time t the adjacency matrix obtained in step (3-2) and the representation result h at the previous time t-1 t-1 into the GCRNN to obtain the reset gate r at time t t ∈R N×h ;

[0067] The calculation formula for this step is:

[0068]

[0069] where and are the learnable parameters of the reset gate r at time t in the K-th cluster. t

[0070] (3-4-3) Input the data of the training set of the three-dimensional tensor obtained in step (3-1) at time t the reset gate r obtained in step (3-4-2) t the adjacency matrix obtained in step (3-2) and the representation result h at the previous time t-1 t-1 into the GCRNN to obtain the transmission state at time t

[0071] The calculation formula for this step is:

[0072]

[0073] where ⊙ represents the element-wise product. and are the learnable parameters of the transmission state at time t in the K-th cluster.

[0074] (3-4-4) Input the update gate z t obtained in step (3-4-1), the representation result at the previous time t-1, and the transmission state obtained in step (3-4-3) into the GCRNN to obtain the representation result h at the current time t t ∈R N×h ;

[0075] The calculation formula for this step is:

[0076]

[0077] where z t ⊙h t-1 represents selective forgetting of the information h at the previous moment t - 1 t-1 and represents selective memory of the information containing the current moment t

[0078] According to another aspect of the present invention, a traffic flow prediction system with missing data is provided, including:

[0079] A first module for obtaining a traffic data set of a certain area, the traffic data set containing missing data, and reconstructing the traffic data set into a traffic flow data matrix;

[0080] A second module for inputting the traffic flow data matrix X obtained by the first module into the orthogonal non - negative matrix factorization (ONMF) module of the trained spatio - temporal prediction model to form K clusters, and filling the data in each cluster using the generalized matrix factorization (GMF) module of the spatio - temporal prediction model to obtain a traffic flow data matrix after filling the data Performing standardization processing on the traffic flow data matrix after filling to obtain a standardized traffic flow data matrix and modeling the standardized traffic flow data matrix into a three - dimensional tensor according to the historical step H and the prediction window W

[0081] A third module for inputting the three - dimensional tensor obtained by the second module into the graph convolutional recurrent neural network (GCRNN) of the trained spatio - temporal prediction model to obtain prediction data Y'.

[0082] Generally speaking, compared with the prior art through the above - mentioned technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0083] (1) Since the present invention adopts step (3), which integrates GCN and GRU to form a Graph Convolution Recurrent Neural Network (abbreviated as GCRNN). This module can use Chebyshev polynomials to approximate GCN to capture spatial correlations, which can improve computational performance. At the same time, the GRU model is adopted to selectively memorize key features to learn long-range temporal characteristics. Therefore, it can solve the technical problem that existing methods only represent temporal relationships and cannot truly reflect the future state of traffic flow.

[0084] (2) Since the present invention adopts step (3), when capturing spatial correlations, the GCN used depends on graph Laplacian decomposition to process irregular graph data. Therefore, it can solve the technical problem that spatial representation is limited to regular data.

[0085] (3) Since the present invention adopts step (3), it designs a graph learning network that can learn the graph structure in a data-driven manner rather than relying on a predefined graph structure. Therefore, it solves the technical problem of poor generality of hybrid neural network methods.

[0086] (4) Since the present invention adopts step (2), it designs a cluster parameter learning mechanism based on Orthogonal Nonnegative Matrix Factorization (abbreviated as ONMF), and jointly with step (3) learns cluster-internal shared-cluster-specific parameters and finely captures the spatio-temporal dependence relationship of traffic flow sequences. This mechanism enables traffic flow data from the same cluster to have a common parameter space, while traffic flows from different clusters have independent parameter spaces. Therefore, it solves the technical problem that the fine-grained representation problem affects the accuracy of traffic flow prediction.

[0087] (5) Since the present invention adopts step (2), it designs a Generalized Matrix Factorization (GMF) filling module, a simple and efficient deep learning method, to fill the missing traffic flow sequences by learning implicit interaction correlations. This module learns the time node interaction function in each traffic flow cluster by element-wise product rather than inner product, not only inheriting the advantages of matrix factorization but also fully exploring the non-linear intrinsic correlations of traffic flows in different clusters. Therefore, it makes up for the technical problem that the data missing problem caused by network or device failures affects the feature learning performance.

[0088] (6) Since the present invention adopts the cluster parameter learning mechanism in step (2), by controlling the number of clusters, the parameter learning mode greatly reduces the number of parameters compared with the node-specific method, that is, it balances the relationship between fine-grained feature learning and the number of parameters; the clustering method in this module better adapts to the missing traffic data and makes the number of nodes in each cluster relatively balanced.

[0089] (7) Since the present invention adopts step (2), it avoids the impact of missing data on the model by only updating the observable data; at the same time, the ONMF module in step (2) has the uniqueness of the solution and clustering interpretability: the uniqueness of the solution can ensure the stability of the algorithm; at the same time, due to the equivalence between the clustering module in step (2) and the K-means method, it has interpretability.

[0090] (8) The ONMF module, GMF module and GCRNN module of the present invention are independent and can be used alone or jointly to adapt to the existing spatio-temporal data prediction model to improve the prediction performance. Description of the Drawings

[0091] Figure 1 Represents the multi-modal features of the traffic dataset;

[0092] Figure 2 Is the framework of the spatio-temporal prediction model of the present invention;

[0093] Figure 3 Is the structural diagram of the GCRNN module designed by the present invention;

[0094] Figure 4 Is the ablation experiment of the spatio-temporal prediction model FCGCRN of the present invention on the PEMS04 dataset with a missing rate of 10%, where Figure 4 (a), Figure 4 (b) and Figure 4 (c) are the experimental results on three error metrics of mean absolute error MAE, root mean square error RMSE and mean absolute percentage error MAPE respectively;

[0095] Figure 5 Is the ablation experiment of the spatio-temporal prediction model FCGCRN of the present invention on the PEMS08 dataset with a missing rate of 10%, where Figure 5 (a), Figure 5 (b) and Figure 5 (c) are the experimental results on three error metrics of mean absolute error MAE, root mean square error RMSE and mean absolute percentage error MAPE respectively;

[0096] Figure 6 Is the ablation experiment of the spatio-temporal prediction model FCGCRN of the present invention on the PEMS04 dataset with a missing rate of 30%, whereFigure 6 (a), Figure 6 (b) and Figure 6 (c) are the experimental results on three error metrics: mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE), respectively;

[0097] Figure 7 is the ablation experiment of the spatio-temporal prediction model FCGCRN of the present invention on the PEMS08 dataset with a missing rate of 30%, where Figure 7 (a), Figure 7 (b) and Figure 7 (c) are the experimental results on three error metrics: mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE), respectively;

[0098] Figure 8 is the parameter analysis experiment of the spatio-temporal prediction model FCGCRN of the present invention on the cluster parameter K, where Figure 8 (a), Figure 8 (b), Figure 8 (c) and Figure 8 (d) are the analysis experiments of FCGCRN on the cluster parameter K on four datasets: PEMS04 (MR = 10), PEMS08 (MR = 10), PEMS04 (MR = 30), and PEMS08 (MR = 30), respectively. The measurement methods adopt mean absolute percentage error (MAPE) and root mean square error (RMSE);

[0099] Figure 9 is the flow chart of the traffic flow prediction method with missing data of the present invention. Detailed implementation manners

[0100] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0101] The basic idea of the present invention is to propose a traffic flow prediction method and system with missing data, which uses orthogonal non-negative matrix factorization to cluster traffic datasets, and uses generalized matrix factorization to fill in the missing traffic data in each cluster. It adopts an adaptive method to learn graphs and fuse graph convolutional recurrent neural networks to fine-grained represent the temporal and spatial characteristics in traffic data with a within-cluster sharing - between-cluster specific parameter learning mechanism, and finally uses a two-dimensional convolutional layer to complete the prediction of all traffic datasets.

[0102] As Figure 9As shown in the figure, the present invention provides a traffic flow prediction method for data with missing values, which specifically includes the following steps:

[0103] (1) Obtain a traffic data set of a certain area, which contains missing data, and reconstruct the traffic data set into a traffic flow data matrix;

[0104] Specifically, in this step, sensors arranged on each street in a certain area are used to obtain traffic data (which is three-dimensional tensor data) with missing values on that street. The traffic data with missing values on all streets constitutes a traffic data set, and then data dimensionality reduction processing is performed on this traffic data set to obtain a traffic flow data matrix.

[0105] The three-dimensional tensor data refers to {time, node, traffic feature}. Among them, time refers to the time when the traffic feature is collected at this node, node refers to a single sensor, and traffic features include vehicle speed feature, traffic flow feature, and number of people feature.

[0106] In this step, the traffic data set

[0107] where It represents the traffic flow data matrix of the nth node among all nodes (i.e., sensors) set on all streets in this area at T moments, n ∈ [1, N], t ∈ [1, T], T is any positive integer, N is the total number of sensors arranged on all streets in this area, and there is indicating that the data has non-negativity;

[0108] where is the cth eigenvalue of the nth node at the tth moment, c ∈ [1, C], where C represents the types of traffic features. There are 3 traffic features (vehicle speed feature, traffic flow feature, and number of people feature) in the present invention, so C = 3;

[0109] After obtaining the traffic data set Z, the present invention sets C to 1 and realizes data dimensionality reduction through the numpy.squeeze() function in the Numpy library of Python. At this time, the obtained traffic flow data matrix is

[0110] (2) Input the traffic flow data matrix X obtained in step (1) into the Orthogonal Nonnegative Matrix Factorization (ONMF) module of the trained spatio-temporal prediction model to form K clusters (see steps (2-1) to (2-3) below), and use the Generalized Matrix Factorization (GMF) module of the spatio-temporal prediction model to fill the data in each cluster to obtain the traffic flow data matrix after filling the data. (See steps (2-4) to (2-12) below) For the traffic flow data matrix after filling. Perform normalization processing to obtain the normalized traffic flow data matrix. (See step (2-13) below), and according to the historical step length H and the prediction window W, model the normalized traffic flow data matrix. As a three-dimensional tensor. (See step (2-14) below);

[0111] Specifically, the ONMF module mentioned in step (2) is Figure 2 The first part shown. This module is a clustering module that clusters the traffic flow data matrix obtained in step (1) into K clusters; ONMF adds an orthogonal constraint condition on the basis of matrix factorization, and is formally expressed as: min F≥0,G≥0 ||X - FG T || 2 , s.t. G T G = I. Where, Is the traffic flow data matrix (R represents real numbers, here it means X is a non-negative real matrix with T rows and N columns), And Are two matrix factors of the traffic flow data matrix X, Is the identity matrix, and the value range of K is from 2 to 5, preferably 2 or 3;

[0112] The GMF module is Figure 2 The second part shown, that is, the filling module of the present invention, including an embedding layer, a decomposition layer and a filling layer. According to the clustering result obtained in step (2-3), use the GMF module in each cluster and fill the missing data according to the observable data.

[0113] The process of inputting the traffic flow data matrix X obtained in step (1) into the ONMF module to form K clusters in this step (2) specifically includes:

[0114] (2-1) Initialize the matrix factors F and G of the traffic flow data matrix X to random values within (0, 1);

[0115] (2-2) Initialize the matrix factors F and G obtained in step (2-1), and use the update rules and to update the observable data in the matrix. When the error of X - FG T converges, stop the iteration to obtain the updated matrix factor G;

[0116] Since the traffic flow data matrix obtained in step (1) has missing values, the ONMF module only uses the observable data in the traffic flow data matrix.

[0117] The advantage of this step is that only using the observable data in the traffic flow data matrix can avoid the influence of missing data (i.e., unobservable data) on the clustering results. At the same time, the matrix factors F and G obtained by using the update rules of this step have the uniqueness of the solution, that is, under the same parameters, the clustering results are consistent each time.

[0118] (2-3) Cluster the traffic flow data matrix X into K clusters according to the updated matrix factor G obtained in step (2-2);

[0119] The advantage of this step is its interpretability, that is, there is an equivalence with the K-means clustering algorithm. Specifically, if the k-th probability of a certain node is the largest, it means that the node belongs to the k-th cluster; in this way, the traffic flow data matrix is clustered into K clusters. At the same time, the data volume ratios of the K clusters obtained in this step are balanced, providing favorable conditions for better characterizing spatio-temporal correlations in the follow-up. Steps (2-1) to (2-3) describe the detailed process of the ONMF module for clustering the traffic flow data matrix. Subsequently, by continuously adjusting the parameter K, it is fused with the GCRNN module to optimize the model, that is, to minimize the prediction error;

[0120] In this step (2), the GMF module is used to fill the traffic flow data matrix X in each cluster to obtain the traffic flow data matrix after filling the data Perform standardization processing on the filled traffic flow data matrix to obtain the standardized traffic flow data matrix and model the standardized traffic flow data matrix as a three-dimensional tensor according to the historical step H and the prediction window W This process includes the following sub-steps:

[0121] (2-4) In each cluster obtained in step (2-3), divide the traffic flow data matrix X into an observable data set (i.e., the set composed of non-missing data in the traffic flow data matrix X) and an unobservable data set (i.e., the set composed of missing data in the traffic flow data matrix X), and reconstruct the observable data set to obtain the time vector vp ∈R m and the node vector v q ∈R m and the traffic flow vector v y ∈R m , reconstruct the unobservable dataset into the time vector v' p ∈R m′ and the node vector v' q ∈R m′ , where m represents the total number of observable data in the observable dataset, and m' represents the total number of unobservable data in the unobservable dataset;

[0122] (2-5) Input the two vectors v p and v q obtained in step (2-4) into the embedding layer of the GMF module to obtain the output of the embedding layer, namely the time matrix factor P and the node matrix factor Q (as Figure 2 shown):

[0123] P = e1(v p ),

[0124] Q = e2(v q ),

[0125] where P ∈ R m×a and Q ∈ R m×a are the time matrix factor and the node matrix factor respectively, a = 16 is the latent factor, and e1() and e2() represent the embedding functions, both of which use the torch.nn.Embedding() function in the Pytorch framework. In the specific implementation process, the parameters of the two embedding functions are different;

[0126] (2-6) Input the time matrix factor P and the node matrix factor Q obtained in step (2-5) into the decomposition layer of the GMF module to obtain the output result f(P, Q):

[0127] f(P, Q) = P ⊙ Q,

[0128] where ⊙ represents the element-wise product operation.

[0129] The advantage of this step is that the way of using the element-wise product instead of the matrix inner product in the decomposition layer not only inherits the advantages of matrix decomposition but also fully explores the non-linear internal correlation of data in different clusters.

[0130] (2-7) Input the output result f(P, Q) obtained in step (2-6) into the filling layer of the GMF module, and the obtained output is the filling result g(P, Q):

[0131] g(P, Q) = a out (W T(P⊙Q)+b),

[0132] where a out is the Relu activation function, and W and b represent the learnable weights and bias parameters in the GMF module respectively;

[0133] The advantage of this step is that the filling layer is a simple and efficient neural network, and the missing traffic data can be filled by using this layer.

[0134] (2-8) Measure the filling result of step (2-7) using the Mean Square Error (MSE) to obtain the error value MSE between the traffic flow vector v y in step (2-4) and the filling result g(P, Q) in step (2-7). The calculation formula for this step is:

[0135]

[0136] (2-9) Update the learnable weights W and bias parameter b in the GMF module through the Adam optimizer;

[0137] (2-10) Repeat the above steps (2-8) to (2-9) until MSE is less than the threshold (10 in this invention -6 ) or the number of training times reaches the preset number of rounds (100 in this invention), so as to obtain the trained GMF module;

[0138] (2-11) Input the unobservable data set obtained in step (2-4) into the GMF module trained in step (2-10) to obtain the traffic flow vector v′ y ∈R m′ ;

[0139] At this point, the data filling work is completed.

[0140] (2-12) According to the time vector v p in the observable data set obtained in step (2-4), the node vector v q , the traffic flow vector v y , the time vector v′ p in the unobservable data set, the node vector v′ q , and the traffic flow vector v′ y obtained in step (2-11), obtain the traffic flow data matrix after filling the data

[0141] (2-13) Standardize the traffic flow data matrix after filling the data in step (2-12) to obtain the standardized traffic flow data matrix

[0142]

[0143] Among them, μ is the mean value of the traffic flow data matrix and σ is the standard deviation;

[0144] The advantages of this step are as follows: First, eliminate the data dimension to improve the convergence speed of the GCRNN module and reduce the computational cost; Second, prevent the occurrence of gradient explosion during the module training process.

[0145] (2-14) According to the historical step length H and the prediction window W, reconstruct the traffic flow data matrix normalized in step (2-13) into a three-dimensional tensor and a three-dimensional tensor Y∈R (T-H-W+1)×N×W ;

[0146] The spatio-temporal prediction model of the present invention is a Fine-grained Completion Graph Convolution Recurrent Network (abbreviated as FCGCRN), including an ONMF module, a GMF module, and a GCRNN module.

[0147] (3) Input the three-dimensional tensor modeled in step (2) into the Graph Convolution Recurrent Neural Network (abbreviated as GCRNN, as Figure 2 shown) of the trained spatio-temporal prediction model to obtain the predicted data Y', which has a low error rate.

[0148] As Figure 3 shown, the GCRNN of the present invention integrates a graph learning neural network, a graph convolutional neural network, and a recurrent neural network under the cluster parameter learning mechanism.

[0149] The GCRNN network in the present invention is trained through the following steps:

[0150] (3-1) Divide the three-dimensional tensor and Y obtained in step (2-14) into a training set and a test set at a ratio of 6:4;

[0151] (3-2) Perform adaptive graph learning on the GCRNN through the parameter E to obtain the adjacency matrix

[0152] The calculation formula for this step is:

[0153]

[0154] where \(E\in\mathbb{R}\) N×e is a learnable parameter matrix, and the parameter \(E\) is initialized using the torch.FloatTensor() function in the Pytorch framework; in the present invention, the optimal tuning results of the parameter \(e\) are 2 and 10;

[0155] (3-3) Input the data of the training set of the three-dimensional tensor at time \(t\) and the adjacency matrix obtained in step (3-2) into the graph convolutional neural network to obtain the graph convolutional result \(H\) of the \(K\)-th cluster K \(\in\mathbb{R}\) N×h ;

[0156] The calculation formula for this step is:

[0157]

[0158] where \(G\) K \(\in\mathbb{R}\) N×N is the Laplacian matrix of the \(K\)-th cluster, \(\theta\) K \(\in\mathbb{R}\) H×h and \(b\) K \(\in\mathbb{R}\) h are the learnable parameters of the \(K\)-th cluster, and \(h\) is the number of neurons in the hidden layer of the graph convolutional recurrent neural network.

[0159] Furthermore, the present invention uses Chebyshev polynomial expansion to approximate the Laplacian matrix \(G\) K . Among them, in the specific implementation process, \(T_0 = 1\), in the present invention, after parameter adjustment, it is finally determined that when the polynomial parameter \(d = 2\) and the cluster parameter \(K = 2\) or 3, the prediction effect is the best.

[0160] The advantages of this step are as follows: First, using Chebyshev polynomial to approximate the Laplacian matrix can represent high-dimensional features with low computational cost; Second, the adopted graph convolutional neural network can strongly represent spatial correlations (including Euclidean space and non-Euclidean space).

[0161] (3-4) Input the data of the training set of the three-dimensional tensor at time \(t\) into the recurrent neural network (which is used to represent temporal correlation) to obtain the representation result \(h\) at time \(t\) t ;

[0162] Specifically, in the present invention, a graph convolutional neural network is integrated into a recurrent neural network to form a GCRNN, so as to jointly represent temporal and spatial correlations, which is manifested as replacing the multi-layer perceptron MLP in the recurrent neural network GRU with a graph convolutional network GCN. The specific implementation steps are as follows:

[0163] (3-4-1) Input the data of the training set of the three-dimensional tensor obtained in step (3-1) at time t the adjacency matrix obtained in step (3-2) and the representation result h at the previous time t-1 t-1 (see step (3-4-4) below) into the GCRNN to obtain the update gate z at time t t ∈R N×h ;

[0164] Specifically, the calculation formula for this step is:

[0165]

[0166] where h0∈R N×h represents the initial state, which is a matrix composed entirely of 0s, [,] represents the concatenation operation, and are the learnable parameters of the update gate z at time t in the Kth cluster t , and σ(·) is the sigmoid activation function;

[0167] (3-4-2) Input the data of the training set of the three-dimensional tensor obtained in step (3-1) at time t the adjacency matrix obtained in step (3-2) and the representation result h at the previous time t-1 t-1 (see step (3-4-4) below) into the GCRNN to obtain the reset gate r at time t t ∈R N×h ;

[0168] The calculation formula for this step is:

[0169]

[0170] where, and are the learnable parameters of the reset gate r at time t in the Kth cluster t ;

[0171] (3-4-3) Input the data of the training set of the three-dimensional tensor obtained in step (3-1) at time t The reset gate r obtained in step (3-4-2) t , the adjacency matrix obtained in step (3-2) and the representation result h at the previous moment t-1 t-1 (see step (3-4-4) below) are input into the GCRNN to obtain the transmission state at time t

[0172] The calculation formula for this step is as follows:

[0173]

[0174] where, ⊙ represents element-wise product, and are the learnable parameters of the transmission state of the Kth cluster at time t ;

[0175] (3-4-4) Input the update gate z obtained in step (3-4-1) t , the representation result at the previous moment t-1 and the transmission state obtained in step (3-4-3) into the GCRNN to obtain the representation result h at the current moment t t ∈R N×h ;

[0176] The calculation formula for this step is as follows:

[0177]

[0178] where, z t ⊙h t-1 means selectively forgetting the information h at the previous moment t-1 t-1 , means selectively memorizing the information containing the current moment t ;

[0179] Steps (3-4-1) to (3-4-4) are all about feature learning at a certain moment t. In specific implementation, the information of T moments is iterated to represent long-term correlation, and the cumulative representation result at the last moment is used to predict future traffic flow.

[0180] In addition, due to the excellent representation ability of the graph convolutional neural network, only 2 layers of the graph neural network in the present invention are required to achieve a relatively low error result, that is, the number of layers of the GCRNN is 2 layers.

[0181] The advantages of the above steps (3-1) to (3-4) are as follows: First, it captures the temporal and spatial correlations in traffic prediction data in a fine-grained manner and achieves high-precision prediction. In the implementation of the prior art, the parameter matrix θ is shared by all nodes (i.e., roads). However, in traffic prediction, not all nodes adopt the same pattern. As Figure 1 shown, Road 1 exhibits the morning rush hour pattern, Roads 2 and 4 exhibit the evening rush hour pattern, and Road 3 exhibits both morning and evening rush hour patterns. Therefore, the present invention adopts a parameter learning mode of sharing within clusters - specific between clusters to characterize the spatio-temporal correlations in a fine-grained manner, that is, the implementation of the above GCRNN module is completed in K clusters. Second, the adopted recurrent neural network GRU can learn the characteristics within a long time range with fewer parameters and lower computational cost compared to other recurrent neural networks, and the learning ability of GRU is comparable to that of other recurrent neural networks.

[0182] (3-5) Input the output h t obtained in step (3-4-4) into the two-dimensional convolutional layer of the GCRNN network (as Figure 3 shown) to obtain the final traffic flow prediction result Y′∈R N×W ;

[0183] The specific implementation of this step is as follows:

[0184] Y′ = h t ★f 1×h ,

[0185] where the number of channels channel of the two-dimensional convolutional layer is 1, the output W = 12, and f 1×h represents that the convolution kernel size in the two-dimensional convolutional layer is (1, h), and in the present invention, the prediction effect is the best when h = 64;

[0186] (3-6) Use the L1 loss function to calculate the loss value O(Y, Y′) between the prediction result Y′ obtained in step (3-5) and the tensor Y (referring to the training set) obtained in step (2-14);

[0187] Specifically, the L1 loss function used in this step is:

[0188]

[0189] (3-7) Use the loss function in step (3-6) and the Adam optimizer in the Pytorch framework to iteratively update the learnable parameters E in step (3-2) and the learnable parameters and in steps (3-4-1) to (3-4-4);

[0190] (3 - 8) Repeat the training process of steps (3 - 6) to (3 - 7) until the number of iterations in step (3 - 7) (100 in the present invention) or the loss value O(Y, Y′) in step (3 - 6) is less than the set threshold (10 in the present invention) -6 to end the training, thereby obtaining a preliminarily trained GCRNN model;

[0191] (3 - 9) Use the test set obtained in step (3 - 1) to verify the GCRNN model preliminarily trained in step (3 - 8) until the prediction error reaches the optimum, thereby obtaining a trained spatio - temporal prediction model.

[0192] In summary, through the above description of the present invention, the main advantages of the present invention include:

[0193] 1. A traffic flow prediction method with missing data is proposed, which can fill in the missing traffic flow data and finely characterize the complex and long - range spatio - temporal correlation of historical traffic flow to achieve high - precision prediction of future traffic flow states.

[0194] 2. By dividing the traffic flow data, GCRNN can adopt a parameter learning mechanism of intra - cluster sharing - inter - cluster specificity to extract the features of traffic flow data. The clustering uses the orthogonal - constraint non - negative matrix factorization algorithm ONMF, which can better adapt to the missing traffic flow data and only iterates on the observable data. Moreover, the number of nodes in each cluster in the clustering result is evenly distributed, providing favorable conditions for better characterizing spatio - temporal relationships subsequently. In addition, clustering is implemented before data filling, making the clustering structure accurate and reliable.

[0195] 3. The generalized matrix factorization module GMF is used to implement missing value filling within each cluster, effectively utilizing the features with high intra - cluster correlation, and at the same time, a simple neural network is used to learn the non - linear relationship between data.

[0196] 4. The present invention uses a graph convolutional neural network to learn the spatial correlation between data. The graph structure is not limited to the Euclidean space and can better characterize the spatial relationship. At the same time, this graph convolutional neural network uses Chebyshev polynomials to approximate the traditional graph convolution, reducing the computational overhead while ensuring the representation effect. In the GCRNN module of the present invention, adaptive graph learning is used instead of predefined graphs, making the graph depend on data rather than on predefined structures.

[0197] 5. Modeling the traffic flow data into a tensor and performing standardization processing can effectively eliminate the differences between data scales and eliminate the influence on representation.

[0198] Experimental results

[0199] The present invention conducts experiments on two real traffic flow datasets, PEMS04 and PEMS08, with 10% and 30% missing rates (MR for short). The PEMS04 dataset has a total of 307 nodes, 16,992 time steps, and a total duration of 59 days; the PEMS08 dataset has 170 nodes, 17,856 time steps, and a total duration of 62 days. Both datasets use 5 minutes as a time step.

[0200] Through comparative experiments on real datasets, the effectiveness and accuracy of a traffic flow prediction method with missing data proposed by the present invention are verified, as shown in Tables 1 and 2 below, and Figures 4 to 7 as follows. The spatio-temporal prediction model FCGCRN of the present invention is compared with eight other baseline methods, including HA, ARIMA, Logistic Regression (LR for short), LSTM, DCRNN, MTGNN, DMSTGCN, and STGODE (these methods are directly cited by their English abbreviations in other people's papers without Chinese), and three metric indicators are adopted: Mean Absolute Error (MAE), Root Mean Square Error (RMSE for short), and Mean Absolute Percentage Error (MAPE for short). The smaller the error metric indicator, the better the prediction performance for missing data. Tables 1 - 2 show the experimental results under two missing rates and two real datasets. It can be seen from the tables that the prediction performance of the spatio-temporal prediction model FCGCRN is better than other baselines under the MAE and RMSE indicators, and better than most baseline methods under the MAPE indicator. Among them, the performance of the spatio-temporal prediction model FCGCRN is more prominent under high missing rates. As can be seen from Table 1, the prediction performance of deep learning methods (LR, LSTM, DCRNN, DMSTGCN, and STGODE) is better than that of traditional statistical methods (HA and ARIMA). Further, it can be seen that the adaptive graph learning models (MTGNN and STGODE) are more affected by data missing. In contrast, although the model of the present invention is also a graph learning model, its performance is more stable when there is data missing and the missing rate is large.

[0201] Figures 4 to 7 The ablation experiment of... shows the effectiveness of the key modules of the present invention. The Generalized Matrix Factorization module GMF ensures the integrity of its data by filling in missing data; the cluster parameter learning mechanism ensures the fine-grainedness of the model by sharing within clusters - specific parameters between clusters; the data-driven graph learning method in the GCRNN module makes the model more versatile in spatio-temporal sequence prediction tasks. Figure 8It represents the analysis of the key parameter cluster number K of the present invention. This parameter determines the diversity of the parameters in the GCRNN module and affects the number of parameters. As can be seen from the figure, the model has the best effect when K is 2 or 3, indicating that GCRNN learns 2 to 3 types of specific parameters. The experimental results correspond to the morning rush hour, evening rush hour, and morning and evening rush hour patterns of traffic sequences. In summary, the present invention is extremely stable, fine-grained, and general in spatio-temporal prediction tasks with missing data.

[0202] Table 1 Comparison experiment results of the spatio-temporal prediction model FCGCRN of the present invention and eight other baseline methods on PEMS04 and PEMS08 (missing rate is 10%) datasets

[0203]

[0204] Table 2 Comparison experiment results of the spatio-temporal prediction model FCGCRN of the present invention and eight other baseline methods on PEMS04 and PEMS08 (missing rate is 30%) datasets

[0205]

[0206] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A traffic flow prediction method with missing data, characterized in that, It includes the following steps: (1) Obtain the traffic data set of a certain area. This traffic data set contains missing data, and reconstruct the traffic data set into a traffic flow data matrix; (2)Input the traffic flow data matrix obtained in step (1) into the orthogonal non-negative matrix factorization (ONMF) module of the trained spatio-temporal prediction model to form K clusters, and use the generalized matrix factorization (GMF) module of the spatio-temporal prediction model to fill the data in each cluster to obtain the traffic flow data matrix after data filling , including: (2-4) In each cluster, divide the traffic flow data matrix into an observable data set and an unobservable data set, and reconstruct the observable data set to obtain a time vector , a node vector and a traffic flow vector , reconstruct the unobservable data set into a time vector and a node vector , where represents the total number of observable data in the observable data set, represents the total number of unobservable data in the unobservable data set; (2-5) Input the two vectors obtained in step (2-4) into the embedding layer of the GMF module to obtain the output of the embedding layer, namely the time matrix factor and and the node matrix factor : : , Among them, and are the time matrix factor and the node matrix factor respectively, = 16 is the latent factor, represents the embedding function, both of which use the torch.nn.Embedding() function in the Pytorch framework; (2-6) Input the time matrix factor P and the node matrix factor Q obtained in step (2-5) into the decomposition layer of the GMF module to obtain the output result : , Among them, represents the element-wise product operation; (2-7) Input the output result obtained in step (2-6) into the padding layer of the GMF module, and the obtained output is the padding result : , Among them, is the Relu activation function, and W and b respectively represent the learnable weights and bias parameters in the GMF module; (2-8) The filling result of step (2-7) is measured using the mean square error MSE to obtain the traffic flow vector of step (2-4). and the filling result of step (2-7). The error value between the two. , and the calculation formula for this step is: ; (2-9) Update the learnable weights W and bias parameters in the GMF module through the Adam optimizer ; (2-10) Repeat the above steps (2-8) to (2-9) until the MSE is less than the threshold or the number of training times reaches the preset number of rounds, so as to obtain a trained GMF module; (2-11) Input the unobservable dataset obtained in step (2-4) into the GMF module trained in step (2-10) to obtain the traffic flow vector ; The time vector under the observable data set obtained according to step (2-4) , the node vector , the traffic flow vector and the time vector under the unobservable data set , the node vector and the traffic flow vector obtained in step (2-11) , and obtain the traffic flow data matrix after filling the data ; For the filled traffic flow data matrix perform normalization to obtain the normalized traffic flow data matrix , and model the normalized traffic flow data matrix as a three-dimensional tensor ; (3) The three-dimensional tensor obtained by modeling in step (2) is input into the graph convolutional recurrent neural network GCRNN of the trained spatio-temporal prediction model to obtain prediction data .

2. The traffic flow prediction method with missing data according to claim 1, wherein The traffic data is three-dimensional tensor data {time, node, traffic feature}. Among them, time refers to the time when the traffic feature is collected at this node, node refers to a single sensor, and traffic features include vehicle speed feature, traffic flow feature and number of people feature.

3. The traffic flow prediction method with missing data according to claim 1 or 2, characterized in that In step (1), the traffic data set ; Among them , which represents the traffic flow data matrix of the nth node among all nodes set on all streets in this area at T moments, where n ∈ [1, N], t ∈ [1, T], T is any positive integer, and N is the total number of sensors deployed on all streets in this area, and there is , indicating that the data is non - negative; wherein is the c-th eigenvalue of the n-th node at the t-th moment, where c ∈ [1, C], and C represents the types of traffic features.

4. The traffic flow prediction method with missing data according to claim 3, wherein In step (2), the traffic flow data matrix obtained in step (1) The process of inputting into the ONMF module to form K clusters specifically includes: (2-1) Initialize the matrix factors F and G of the traffic flow data matrix to random values within (0, 1); (2-2) Initialize the matrix factors F and G obtained according to step (2-1), and adopt the update rules and update the observable data in the matrix. When the error converges, stop the iteration, thereby obtaining the updated matrix factor G; (2-3) Based on the updated matrix factor G obtained in step (2-2), cluster the traffic flow data matrix into K clusters.

5. The traffic flow prediction method with missing data according to claim 4, characterized in that In step (2), the traffic flow data matrix is filled in each cluster using the GMF module to obtain the traffic flow data matrix after filling the data The filled traffic flow data matrix is standardized to obtain the standardized traffic flow data matrix and the standardized traffic flow data matrix is modeled as a three-dimensional tensor according to the historical step H and the prediction window W This process includes the following sub-steps: This process includes the following sub-steps: (2-13) Standardize the traffic flow data matrix after filling data in step (2-12) to obtain a standardized traffic flow data matrix : , Among them, is the mean value of the traffic flow data matrix , and is 's standard deviation; (2-14) Reconstruct the traffic flow data matrix standardized in step (2-13) into a three-dimensional tensor according to the historical step length H and the prediction window W and a three-dimensional tensor and a three-dimensional tensor .

6. The traffic flow prediction method with missing data according to claim 5, characterized in that The GCRNN network is trained through the following steps: (3-1) Divide the three-dimensional tensor and Y into a training set and a test set at a ratio of 6:4; (3-2) Adaptive graph learning of the GCRNN is performed through the parameter E to obtain the adjacency matrix ; The calculation formula of this step is: , Among them, is a learnable parameter matrix. The parameter E is initialized using the torch.FloatTensor() function in the Pytorch framework, and the optimal tuning results of the parameter e are 2 and 10; (3-3) The data of the training set of the three-dimensional tensor at time t and the adjacency matrix obtained in step (3-2) are input into the graph convolutional neural network to obtain the graph convolutional result of the Kth cluster ; The calculation formula of this step is: , Among them, is the Laplacian matrix of the K-th cluster, and are the learnable parameters of the K-th cluster, and h is the number of neurons in the hidden layer of the graph convolutional recurrent neural network; (3 - 4) Input the data of the training set of the three - dimensional tensor at time t into the recurrent neural network to obtain the representation result at time t ; (3 - 5) Input the output obtained in step (3 - 4 - 4) into the two - dimensional convolutional layer of the GCRNN network to obtain the final traffic flow prediction result ; The specific implementation of this step is as follows: , Among them, the number of channels channel of the two-dimensional convolutional layer is 1, and the output W = 12. It means that the size of the convolutional kernel in the two-dimensional convolutional layer is (1, h), where h = 64. (3 - 6) Calculate the loss value between the prediction result obtained in step (3 - 5) and the tensor Y obtained in step (2 - 14) ; Specifically, the L1 loss function used in this step is: , (3-7) Use the loss function in step (3-6) and the Adam optimizer in the Pytorch framework to iteratively update the learnable parameters E in step (3-2) and the learnable parameters from (3-4-1) to (3-4-4). and for iterative update; (3 - 8) Repeat the training process of steps (3 - 6) to (3 - 7) until the number of iterations in step (3 - 7) or the loss value in step (3 - 6) is less than the set threshold, and then end the training to obtain a preliminarily trained GCRNN model; (3-9) Use the test set obtained in step (3-1) to verify the GCRNN model preliminarily trained in step (3-8) until the prediction error reaches the optimum, so as to obtain a trained spatio-temporal prediction model.

7. The traffic flow prediction method with missing data according to claim 6, characterized in that Step (3-4) includes the following sub-steps: (3-4-1) Input the data of the training set of the three-dimensional tensor at time t , the adjacency matrix obtained in step (3-2) and the representation result at the previous time t-1 into the GCRNN to obtain the update gate at time t ; Specifically, the calculation formula of this step is: ), Among them, represents the initial state, which is a matrix composed entirely of 0s. [,] represents the concatenation operation. and are the learnable parameters of the update gate at time t in the Kth cluster. is the sigmoid activation function. (3-4-2) The data of the training set of the three-dimensional tensor at time t , the adjacency matrix obtained in step (3-2) and the representation result at the previous time t-1 are input into the GCRNN to obtain the reset gate at time t ; The calculation formula of this step is: ), Among them, and are the learnable parameters of the reset gate at time t in the K-th cluster ; (3-4-3) Input the data of the training set of the three-dimensional tensor at time t , the reset gate obtained in step (3-4-2) , the adjacency matrix obtained in step (3-2) and the representation result at the previous time t-1 into the GCRNN to obtain the transmission state at time t ; The calculation formula of this step is: ), Among them, represents the element-wise product, and is the transmission state of the K-th cluster at time t is the learnable parameter; (3 - 4 - 4) Input the update gate obtained in step (3 - 4 - 1), the representation result at the previous moment t - 1, and the transmission state obtained in step (3 - 4 - 3) into the GCRNN to obtain the representation result at the current moment t ; The calculation formula of this step is: ; Among them, represents the information at the previous moment t - 1 for selective forgetting, represents the information including the current moment t for selective memory.

8. A traffic flow prediction system with missing data, which adopts the traffic flow prediction method with missing data according to claim 1, is characterized in that, The traffic flow prediction system includes: The first module is used to obtain the traffic data set of a certain area. This traffic data set contains missing data, and reconstruct the traffic data set into a traffic flow data matrix; The second module is used to input the traffic flow data matrix obtained by the first module into the orthogonal non - negative matrix factorization (ONMF) module of the trained spatio - temporal prediction model to form K clusters, and use the generalized matrix factorization (GMF) module of the spatio - temporal prediction model to fill the data in each cluster to obtain the traffic flow data matrix after filling the data , and perform standardization processing on the traffic flow data matrix after filling to obtain the standardized traffic flow data matrix , and model the standardized traffic flow data matrix as a three - dimensional tensor according to the historical step H and the prediction window W ; The third module is used to input the three-dimensional tensor modeled by the second module into the graph convolutional recurrent neural network GCRNN of the trained spatio-temporal prediction model to obtain prediction data .

Citation Information

Patent Citations

  • Multivariate time series prediction method and system, computer product and storage medium

    CN114493014A

  • Traffic network data prediction method based on graph space-time self-encoding network

    CN114565187A