Expressway traffic situation prediction method based on TCN-GCN
By constructing a TCN-GCN model based on time convolution network and graph convolution neural network, the existing highway traffic situation prediction methods in terms of accuracy and computational complexity are solved, and high-precision and low computational complexity traffic situation prediction are achieved.
Patent Information
- Application Number
- CN202510002601.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-30
AI Technical Summary
The existing highway traffic situation prediction methods have shortcomings in prediction accuracy and computational complexity, especially the limitation of spatial and temporal relationships and data feature processing capabilities.
The TCN-GCN model is adopted based on the time convolution network (TCN) and the graph convolution neural network (GCN). This model can extract temporal and spatial features at the same time, build and train it through the training set to achieve high-precision traffic situation prediction.
High-precision highway traffic situation prediction is achieved, the prediction process is simple and the calculation is small, and the traffic situation relationship under all time and space conditions can be fully considered.
Smart Images

Figure CN120071604A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the traffic situation of expressways, and more particularly to a method for predicting the traffic situation of expressways based on TCN-GCN. Background Art
[0002] With the increasing complexity of urban traffic systems and the continuous increase in traffic flow, traffic situation prediction has become increasingly important. Facing frequent problems such as traffic congestion and traffic accidents, accurate traffic situation prediction can optimize travel routes for travelers to improve traffic efficiency, and also facilitate traffic management departments to grasp road conditions in real time and formulate appropriate traffic management measures, thereby reducing the occurrence of congestion and accidents. Therefore, traffic situation prediction methods have important practical significance and are currently used for traffic situation prediction on various roads.
[0003] The economic development has made the construction of expressways in China increasingly perfect. Currently, traffic monitoring facilities such as gantries and toll stations have been fully popularized on expressways in China, and data collection devices including traffic sensors and cameras have been installed on most sections. These data also provide important data support for accurate traffic prediction. In addition, with the development of Internet technology, such as machine learning, especially deep learning, which can rely on data training and accurately reflect the traffic state by learning data features, provides new ideas for realizing high-precision expressway traffic situation prediction.
[0004] Traditional expressway traffic situation prediction methods mainly include the expressway traffic situation prediction method based on ARIMA, the expressway traffic situation prediction method based on linear regression, and the expressway traffic situation prediction method based on LSTM, etc. Although these traditional expressway traffic situation prediction methods have achieved certain results in some predictions, there are still many deficiencies. For example, the expressway traffic situation prediction method based on ARIMA and the expressway traffic situation prediction method based on linear regression only explore the linear relationship in the data and ignore some non-linear factors, resulting in low prediction accuracy. The expressway traffic situation prediction method based on LSTM, although it takes into account non-linear factors, only focuses on the time aspect and ignores the spatial relationship in the prediction process, and the prediction accuracy is also low.
[0005] In recent years, there has emerged an expressway traffic situation prediction method based on STGCN that combines spatial and temporal correlations. Although its prediction accuracy has been improved compared with traditional expressway traffic situation prediction methods, its ability to process data features is limited, the prediction accuracy is still not high, and at the same time, its prediction requires a large amount of computing power, with a large amount of calculation and a complex prediction process. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a highway traffic situation prediction method based on TCN-GCN, which has a simple prediction process, a small amount of calculation, and high prediction accuracy.
[0007] The technical solution adopted by the present invention to solve the above technical problems is as follows: A highway traffic situation prediction method based on TCN-GCN constructs a TCN-GCN model with time feature extraction function and space feature extraction function based on the time convolutional network TCN and the graph convolutional neural network GCN, obtains the historical traffic flow data on all gantries of the highway section to be predicted, and performs outlier exclusion and normalization processing on the historical traffic flow data, and constructs a training set X with the processed historical traffic flow data input and the corresponding label set Y out , and uses the training set X input and the label set Y out to train the TCN-GCN model to obtain the trained TCN-GCN model, and then obtains the traffic flow data of the highway section to be predicted in real time, and uses the TCN-GCN model for prediction to obtain the traffic situation prediction result of the highway section to be predicted.
[0008] Compared with the prior art, the advantages of the present invention are that considering that in highway traffic, the traffic situation in unknown time and space is not only restricted by upstream traffic flow and past traffic flow, but also affected by past spatial relationships, a TCN-GCN model with time feature extraction function and space feature extraction function is constructed based on the time convolutional network TCN and the graph convolutional neural network GCN, and the training set X input and the corresponding label set Y out are constructed after outlier exclusion and normalization processing of the historical traffic flow data to train the TCN-GCN model, so that the TCN-GCN model can simultaneously realize the time feature extraction function and the space feature extraction, so that when predicting, the different spatio-temporal change characteristics of the road can be obtained respectively, fully considering the traffic situation relationship under all spatio-temporal conditions, having high prediction accuracy, and directly predicting through the trained TCN-GCN model, the prediction process is simple and the amount of calculation is small.
[0009] Furthermore, the TCN-GCN model includes a time module, a space module, and a multi-layer perceptron module. The time module includes a temporal convolutional network (TCN), which is referred to as the first TCN. The space module includes a temporal convolutional network (TCN) and a graph convolutional neural network (GCN), where the TCN is referred to as the second TCN and the graph convolutional neural network (GCN) is referred to as the first GCN. The first TCN is used to access the traffic flow data at its input and extract temporal features from the traffic flow data, and the data with temporal features is output to the multi-layer perceptron module. The second TCN is used to access the traffic flow data at its input and extract historical spatial features from the traffic flow data, and the data with historical spatial features is output to the first GCN. The first GCN is used to extract spatial features from the data with historical spatial features output by the second TCN, and the data with spatial features is output to the multi-layer perceptron module. The multi-layer perceptron module is used to fuse the data with temporal features output by the first TCN and the data with spatial features output by the first GCN, and the traffic situation prediction result is obtained and output.
[0010] Further, the first TCN includes 5 convolutional layers, 4 residual connection layers, and an output layer. The 5 convolutional layers are sequentially referred to as the first convolutional layer to the fifth convolutional layer, the 4 residual connection layers are sequentially referred to as the first residual connection layer to the fourth residual connection layer, and the output layer is referred to as the first output layer. The first convolutional layer is used to access the traffic flow data at its input and perform convolutional processing on the traffic flow data to obtain the first data with a time series pattern and output it to the first residual connection layer. The first residual connection layer is used to access the first data with a time series pattern input from the first convolutional layer and the traffic flow data at its input, and sum the first data with a time series pattern and the traffic flow data at its input to obtain the first sum of traffic flow data and output it to the second convolutional layer. The second convolutional layer is used to access the first sum of traffic flow data input from the first residual connection layer and perform convolutional processing on the first sum of traffic flow data to obtain the second data with a time series pattern and output it to the second residual connection layer. The second residual connection layer is used to access the second data with a time series pattern input from the second convolutional layer and the traffic flow data at its input, and sum the second data with a time series pattern and the traffic flow data at its input to obtain the second sum of traffic flow data and output it to the third convolutional layer. The third convolutional layer is used to access the second sum of traffic flow data input from the second residual connection layer and perform convolutional processing on the second sum of traffic flow data to obtain the third data with a time series pattern and output it to the third residual connection layer. The third residual connection layer is used to access the third data with a time series pattern input from the third convolutional layer and the traffic flow data at its input, and sum the third data with a time series pattern and the traffic flow data at its input to obtain the third sum of traffic flow data and output it to the fourth convolutional layer. The fourth convolutional layer is used to access the third sum of traffic flow data input from the third residual connection layer and perform convolutional processing on the third sum of traffic flow data to obtain the fourth data with a time series pattern and output it to the fourth residual connection layer. The fourth residual connection layer is used to access the fourth data with a time series pattern input from the fourth convolutional layer and the traffic flow data at its input, and sum the fourth data with a time series pattern and the traffic flow data at its input to obtain the fourth sum of traffic flow data and output it to the fifth convolutional layer. The fifth convolutional layer is used to access the fourth sum of traffic flow data input from the fourth residual connection layer and perform convolutional processing on the fourth sum of traffic flow data to obtain the fifth data with a time series pattern and output it to the first output layer.When training the TCN-GCN model described above, the first output layer is used to input the fifth data with a time series pattern input to it by the fifth convolutional layer into the loss function pre-stored therein, perform backpropagation and gradient descent processing, and obtain data Y with time characteristics; tcn Output to the multi-layer perceptron module. When using the trained TCN-GCN model for prediction, the first output layer is used to directly perform forward propagation on the fifth data with a time series pattern input to it by the fifth convolutional layer, obtain the prediction result of the first TCN, and input it into the multi-layer perceptron module; the loss function in the first output layer uses the mean square error MSE, and the true value in the mean square error MSE uses the label set Y out .
[0011] Furthermore, the structure of the second TCN is the same as that of the first TCN, except that the first output layer of the second TCN outputs data to the first GCN instead of the multi-layer perceptron module; the first GCN includes one adjacency matrix layer, two spatial convolutional layers, one residual connection layer, and one output layer. The output layer is called the second output layer, and the two spatial convolutional layers are respectively called the first spatial convolutional layer and the second spatial convolutional layer; the adjacency matrix layer is used to extract the spatial features of the highway section to be predicted and perform normalization processing to obtain the adjacency matrix A with spatial features and input it into the first spatial convolutional layer and the second spatial convolutional layer respectively. Among them, the specific method for the adjacency matrix layer to extract the spatial features of the highway section to be predicted and perform normalization processing to obtain the adjacency matrix A with spatial features is as follows: Denote the number of gantries of the highway section to be predicted as P, and number the P gantries of the highway section to be predicted in sequence from 1 to P. The gantry numbered w is called gantry w, w = 1, 2,..., P. Gantry i and gantry j respectively represent the gantries of the highway section to be predicted, i = 1, 2,..., P, j = 1, 2,..., P. When i ≠ j and gantry i and gantry j are two adjacent gantries, use formula (1) to extract the spatial features of the highway section to be predicted and perform normalization processing. When i ≠ j and gantry i and gantry j are two non-adjacent gantries, use formula (2) to extract the spatial features of the highway section to be predicted and perform normalization processing. When i = j, use formula (3) to extract the spatial features of the highway section to be predicted and perform normalization processing:
[0012]
[0013] a ij = 0 (2)
[0014] a ij = 1 (3)
[0015] where \(e\) represents the base of the natural logarithm, and \(a\) ij represents the element in the \(i\)-th row and \(j\)-th column of the adjacency matrix \(A\) with spatial characteristics, and \(C\) ij represents the number of lanes of the section between gantry \(i\) and gantry \(j\) of the highway to be predicted, and \(D\) ij represents the distance between gantry \(i\) and gantry \(j\) of the highway to be predicted. represents the variance of the distance \(D\) ij which is obtained by the following method: first calculate the average value of the distances of all adjacent gantries of the highway section to be predicted, take this average value as the average distance, and then calculate the square value of the difference between \(D\) ij and this average distance, and divide this square value by \(P - 1\) to obtain the variance of \(D\) ij ; represents the variance of the number of lanes \(C\) ij which is obtained by the following method: first calculate the average value of the number of lanes between all adjacent gantries of the highway section to be predicted, take this average value as the average number of lanes, and then calculate the square value of the difference between \(C\) ij and this average number of lanes, and divide this square value by \(P - 1\) to obtain the variance of \(C\) ij ; The first-layer spatial convolutional layer is used to access the data output by the first output layer of the second TCN and the adjacency matrix \(A\) with spatial characteristics, and perform convolutional processing on the data output by the first output layer of the second TCN and the adjacency matrix \(A\) with spatial characteristics to obtain traffic flow data with spatial characteristics and output it to the residual connection layer. The residual connection layer is used to access the data output by the first output layer of the second TCN and the traffic flow data with spatial characteristics input to it by the first-layer spatial convolutional layer, and sum the data output by the first output layer of the second TCN and the traffic flow data with spatial characteristics input to it by the first-layer spatial convolutional layer to obtain the sum of the fifth traffic flow data and output it to the second-layer spatial convolutional layer; The second-layer spatial convolutional layer is used to access the sum of the fifth traffic flow data input to it by the residual connection layer and the adjacency matrix \(A\), and perform convolutional processing on the sum of the fifth traffic flow data input to it by the residual connection layer and the adjacency matrix \(A\) to obtain the final data \(Y\) with spatial characteristics space and output it to the second output layer; When training the TCN - GCN model, the second output layer is used to access the final data \(Y\) with spatial characteristics input to it by the second-layer spatial convolutional layer space , input the final data \(Y\) space into the pre-stored loss function at its place for backpropagation and gradient descent processing to obtain the data \(Y\) with spatial characteristics gcnInput the described multi-layer perceptron module; when using the trained TCN-GCN model for prediction, the second output layer is used to input the final data Y with spatial features at the second spatial convolution layer into it space Perform forward propagation to obtain the prediction result of the first GCN, and input it into the described multi-layer perceptron module; the loss function in the second output layer uses the mean squared error MSE, and the true value in the mean squared error MSE uses the label set Y out ;
[0016] Furthermore, when training the TCN-GCN model, the multi-layer perceptron module is used to access the data Y with temporal features input at the first output layer of the first TCN into it tcn and the data Y with spatial features input at the second output layer of the first GCN into it gcn , and input the data Y with temporal features tcn and the data Y with spatial features gcn into the pre-stored loss function at it, and after performing backpropagation and gradient descent processing, obtain the prediction result Y with spatio-temporal relationship and output it; when using the trained TCN-GCN model for prediction, the multi-layer perceptron module is used to perform forward propagation on the prediction results of the first TCN and the first GCN to obtain the final prediction result and output it, and the loss function in the multi-layer perceptron module uses the mean squared error MSE, and the true value in the mean squared error MSE uses the label set Y out 。
[0017] Furthermore, after excluding outliers and normalizing the historical traffic flow data of the highway section to be predicted, use the processed historical traffic flow data to construct the training set X input and the corresponding label set Y out The specific process is as follows:
[0018] S1. Obtain the traffic flow data of all gantries within the highway section to be predicted from T b to T c days from the highway management department, where T b represents the first T b days before the current date, and T c represents the first T c days before the current date, T b =1, 1≤T c -T b ≤15;
[0019] Each traffic flow data of each gantry consists of features such as timestamp, gantry ID, section flow, section speed, and section density. Among the traffic flow data of each gantry per day, the interval between the timestamps of two adjacent traffic flow data is five minutes, that is, one traffic flow data is recorded every five minutes. Each gantry has 288 traffic flow data, and the number of traffic flow data collected by each gantry is 288(T c -T b ). Denote the number of gantries within the highway section to be predicted as P. Concatenate all traffic flow data with the same timestamp to form a traffic flow data set of the highway section to be predicted, and obtain traffic flow data with a quantity of P * 288(T c -T b );
[0020] S2. Extract other features of each traffic flow data of the highway section to be predicted except for the timestamp and gantry ID, arrange them in ascending order of numerical value, and denote the first quartile and the third quartile after arrangement as Z 1 and Z 3 . Then, first delete the other feature values of each traffic flow data of the highway section to be predicted that do not satisfy being greater than Z 1 -1.5(Z 3 -Z 1 ) and less than Z 3 +1.5(Z 3 -Z 1 ). Then, use polynomial interpolation to interpolate the deleted features to obtain the interpolated traffic flow data. Split each interpolated traffic flow data into traffic flow data of each gantry, and arrange the traffic flow data of all gantries obtained after splitting in ascending order of gantry ID, and the traffic flow data of each gantry ID in ascending order of timestamp to form the data set X;
[0021] S3. Divide the training set and labels in the data set X to obtain the training set X inpu and the label set Y out . The division method is as follows:
[0022] S3.1. Extract the traffic flow data of each gantry in the dataset X in ascending order of gantry ID. The method of extracting the traffic flow data of each gantry is as follows: in the order of time stamps from the earliest to the latest, extract n + m pieces of traffic flow data each time. For the first time, start from the first piece of traffic flow data of the gantry and extract n + m pieces of traffic flow data. For the second time, start from the second piece of traffic flow data of the gantry and extract n + m pieces of traffic flow data, and so on, until the last piece of traffic flow data of the gantry is extracted. At this time, the extraction of the traffic flow data of this gantry is completed. Here, both m and n are positive integers, and n > m, n + m < 288(T c -T b ). For each time n + m pieces of traffic flow data of a certain gantry are extracted, take the first n pieces of traffic flow data as 1 input data in the order of time stamps from the earliest to the latest, and only keep the time stamp, gantry ID, and cross-section flow of the last m pieces of traffic flow data as 1 label (regard the time stamp, gantry ID, and cross-section flow as the output traffic situation). Thus, after the extraction of the traffic flow data of all gantries in the dataset X is completed, P * [288(T c -T b ) - n - m + 1] input data and P * [288(T c -T b ) - n - m + 1] labels are obtained;
[0023] S3.2. Use P * [288(T c -T b ) - n - m + 1] input data to form the training set X inpu , and use P * [288(T c -T b ) - n - m + 1] labels to form the label set Y out .
[0024] Furthermore, collect the traffic flow data of all gantries on the highway section to be predicted in real time, perform outlier and normalization processing to obtain the processed data, and input the processed data into the trained TCN - GCN model to output the future traffic situation of the section to be predicted. Description of the Drawings
[0025] Figure 1 is the flowchart of the highway traffic situation prediction method based on TCN - GCN of the present invention;
[0026] Figure 2 is the structural diagram of the TCN - GCN model in the highway traffic situation prediction method based on TCN - GCN of the present invention;
[0027] Figure 3It is the structure diagram of the first TCN in the TCN-GCN model of the highway traffic situation prediction method based on TCN-GCN of the present invention;
[0028] Figure 4 It is the structure diagram of the first GCN in the TCN-GCN model of the highway traffic situation prediction method based on TCN-GCN of the present invention;
[0029] Figure 5 It is the schematic structural diagram of the highway to be predicted in the example when predicting by using the highway traffic situation prediction method based on TCN-GCN of the present invention;
[0030] Figure 6 It is the average loss reduction diagram of the spatial module and the temporal module in the example when predicting by using the highway traffic situation prediction method based on TCN-GCN of the present invention. Detailed implementation manners
[0031] The present invention will be further described in detail below in conjunction with the embodiments with reference to the drawings.
[0032] Embodiment 1: As Figure 1 shown, a highway traffic situation prediction method based on TCN-GCN constructs a TCN-GCN model with time feature extraction function and spatial feature extraction function based on the temporal convolutional network TCN and the graph convolutional neural network GCN, obtains the historical traffic flow data on all gantries of the highway section to be predicted, and performs outlier exclusion and normalization processing on the historical traffic flow data, and constructs a training set X input and the corresponding label set Y out , and uses the training set X input and the label set Y out to train the TCN-GCN model, obtains the trained TCN-GCN model, then obtains the traffic flow data of the highway section to be predicted in real time, and uses the TCN-GCN model for prediction to obtain the traffic situation prediction result of the highway section to be predicted.
[0033] In this embodiment, considering that in highway traffic, the traffic situation in unknown time and space is not only restricted by upstream traffic flow and past traffic flow, but also affected by past spatial relationships. A TCN-GCN model with time feature extraction function and spatial feature extraction function is constructed based on the Temporal Convolutional Network (TCN) and the Graph Convolutional Network (GCN). After excluding outliers and normalizing the historical traffic flow data, a training set Xinput and a corresponding label set Yout are constructed to train the TCN-GCN model, enabling the TCN-GCN model to simultaneously achieve time feature extraction and spatial feature extraction. Therefore, during prediction, different spatio-temporal change features of the road can be obtained respectively, fully considering the traffic situation relationship under full spatio-temporal conditions, with high prediction accuracy. And the prediction is directly carried out through the trained TCN-GCN model, with a simple prediction process and small computational complexity.
[0034] Embodiment 2: This embodiment is basically the same as Embodiment 1, except that: in this embodiment, as Figure 2 shown, the TCN-GCN model includes a time module, a spatial module, and a multi-layer perceptron module. The time module includes a Temporal Convolutional Network (TCN), which is referred to as the first TCN; the spatial module includes a Temporal Convolutional Network (TCN) and a Graph Convolutional Network (GCN), the TCN is referred to as the second TCN, and the Graph Convolutional Network (GCN) is referred to as the first GCN; the first TCN is used to access the traffic flow data at its location and extract time features from the traffic flow data, and output the data with time features to the multi-layer perceptron module; the second TCN is used to access the traffic flow data at its location and extract historical spatial features from the traffic flow data, and output the data with historical spatial features to the first GCN; the first GCN is used to extract spatial features from the data with historical spatial features output by the second TCN, and output the data with spatial features to the multi-layer perceptron module; the multi-layer perceptron module is used to fuse the data with time features output by the first TCN and the data with spatial features output by the first GCN, and obtain and output the traffic situation prediction result.
[0035] In this embodiment, as Figure 3As shown, the first TCN includes 5 convolutional layers, 4 residual connection layers, and an output layer. The 5 convolutional layers are sequentially referred to as the first convolutional layer to the fifth convolutional layer, the 4 residual connection layers are sequentially referred to as the first residual connection layer to the fourth residual connection layer, and the output layer is referred to as the first output layer. The first convolutional layer is used to access the traffic flow data at its input from the outside and perform convolutional processing on the traffic flow data to obtain the first data with a time series pattern and output it to the first residual connection layer. The first residual connection layer is used to access the first data with a time series pattern input from the first convolutional layer and the traffic flow data at its input from the outside, and sum the first data with a time series pattern and the traffic flow data at its input from the outside to obtain the first sum of traffic flow data and output it to the second convolutional layer. The second convolutional layer is used to access the first sum of traffic flow data input from the first residual connection layer and perform convolutional processing on the first sum of traffic flow data to obtain the second data with a time series pattern and output it to the second residual connection layer. The second residual connection layer is used to access the second data with a time series pattern input from the second convolutional layer and the traffic flow data at its input from the outside, and sum the second data with a time series pattern and the traffic flow data at its input from the outside to obtain the second sum of traffic flow data and output it to the third convolutional layer. The third convolutional layer is used to access the second sum of traffic flow data input from the second residual connection layer and perform convolutional processing on the second sum of traffic flow data to obtain the third data with a time series pattern and output it to the third residual connection layer. The third residual connection layer is used to access the third data with a time series pattern input from the third convolutional layer and the traffic flow data at its input from the outside, and sum the third data with a time series pattern and the traffic flow data at its input from the outside to obtain the third sum of traffic flow data and output it to the fourth convolutional layer. The fourth convolutional layer is used to access the third sum of traffic flow data input from the third residual connection layer and perform convolutional processing on the third sum of traffic flow data to obtain the fourth data with a time series pattern and output it to the fourth residual connection layer. The fourth residual connection layer is used to access the fourth data with a time series pattern input from the fourth convolutional layer and the traffic flow data at its input from the outside, and sum the fourth data with a time series pattern and the traffic flow data at its input from the outside to obtain the fourth sum of traffic flow data and output it to the fifth convolutional layer. The fifth convolutional layer is used to access the fourth sum of traffic flow data input from the fourth residual connection layer and perform convolutional processing on the fourth sum of traffic flow data to obtain the fifth data with a time series pattern and output it to the first output layer. When training the TCN-GCN model, the first output layer is used to input the fifth data with a time series pattern input from the fifth convolutional layer into the pre-stored loss function at its input, perform backpropagation and gradient descent processing to obtain the data Y with time characteristicstcn Output to the multi-layer perceptron module. When using the trained TCN-GCN model for prediction, the first output layer is used to directly forward propagate the fifth data with time series patterns input at the fifth convolutional layer to obtain the prediction result of the first TCN, and input it into the multi-layer perceptron module; the loss function in the first output layer uses the mean square error MSE, and the true value in the mean square error MSE uses the label set Y out 。
[0036] In this embodiment, as Figure 4 shown, the structure of the second TCN is the same as that of the first TCN, the difference is that the first output layer of the second TCN outputs data to the first GCN instead of the multi-layer perceptron module; the first GCN includes one adjacency matrix layer, two spatial convolutional layers, one residual connection layer and one output layer, and its output layer is called the second output layer, and the two spatial convolutional layers are respectively called the first layer spatial convolutional layer and the second layer spatial convolutional layer; the adjacency matrix layer is used to extract the spatial features of the highway section to be predicted and perform normalization processing to obtain the adjacency matrix A with spatial features and input it into the first layer spatial convolutional layer and the second layer spatial convolutional layer respectively. Among them, the specific method for the adjacency matrix layer to extract the spatial features of the highway section to be predicted and perform normalization processing to obtain the adjacency matrix A with spatial features is as follows: Denote the number of gantries of the highway section to be predicted as P, and number the P gantries of the highway section to be predicted in sequence from 1 - P. The gantry numbered w is called gantry w, w = 1, 2,..., P. Gantry i and gantry j respectively represent the gantries of the highway section to be predicted, i = 1, 2,..., P, j = 1, 2,..., P. When i ≠ j and gantry i and gantry j are two adjacent gantries, use formula (1) to extract the spatial features of the highway section to be predicted and perform normalization processing. When i ≠ j and gantry i and gantry j are two non-adjacent gantries, use formula (2) to extract the spatial features of the highway section to be predicted and perform normalization processing. When i = j, use formula (3) to extract the spatial features of the highway section to be predicted and perform normalization processing:
[0037]
[0038] a ij =0 (2)
[0039]
[0040] Among them, e represents the base of the natural logarithm, a ij represents the element in the i-th row and j-th column of the adjacency matrix A with spatial features, C ij represents the number of lanes of the section between gantry i and gantry j of the highway to be predicted, D ijDenote the distance between gantry i and gantry j of the highway to be predicted. Denote the distance D ij The variance of which is obtained by the following method: First, calculate the average value of the distances between all adjacent gantries of the highway section to be predicted, take this average value as the average distance, and then calculate the square value of the difference between D ij and this average distance, and divide this square value by P - 1 to obtain the variance of D ij ; Denote the number of lanes C ij The variance of which is obtained by the following method: First, calculate the average value of the number of lanes between all adjacent gantries of the highway section to be predicted, take this average value as the average number of lanes, and then calculate the square value of the difference between C ij and this average number of lanes, and divide this square value by P - 1 to obtain the variance of C ij ; The first - layer spatial convolutional layer is used to access the data output by the first output layer of the second TCN and the adjacency matrix A with spatial features, and perform convolutional processing on the data output by the first output layer of the second TCN and the adjacency matrix A with spatial features to obtain traffic flow data with spatial features and output it to the residual connection layer. The residual connection layer is used to access the data output by the first output layer of the second TCN and the traffic flow data with spatial features input to it by the first - layer spatial convolutional layer, and sum the data output by the first output layer of the second TCN and the traffic flow data with spatial features input to it by the first - layer spatial convolutional layer to obtain the sum of the fifth traffic flow data and output it to the second - layer spatial convolutional layer; The second - layer spatial convolutional layer is used to access the sum of the fifth traffic flow data input to it by the residual connection layer and the adjacency matrix A, and perform convolutional processing on the sum of the fifth traffic flow data input to it by the residual connection layer and the adjacency matrix A to obtain the final data Y with spatial features space and output it to the second output layer; When training the TCN - GCN model, the second output layer is used to access the final data Y with spatial features input to it by the second - layer spatial convolutional layer space , input the final data Y space into the pre - stored loss function at its place for backpropagation and gradient descent processing to obtain the data Y with spatial features gcn and input it into the multi - layer perceptron module; When using the trained TCN - GCN model for prediction, the second output layer is used to perform forward propagation on the final data Y with spatial features input to it by the second - layer spatial convolutional layer space to obtain the prediction result of the first GCN and input it into the multi - layer perceptron module; The loss function in the second output layer uses the mean square error MSE, and the true value in the mean square error MSE uses the label set Y out ;
[0041] In this embodiment, when training the TCN-GCN model, the multi-layer perceptron module is used to access the first output layer of the first TCN and input the data Y with time features at this point tcn and the second output layer of the first GCN and input the data Y with spatial features at this point gcn , and input the data Y with time features tcn and the data Y with spatial features gcn into the pre-stored loss function at this point. After performing backpropagation and gradient descent processing, the prediction result Y with spatio-temporal relationship is obtained and output; when using the trained TCN-GCN model for prediction, the multi-layer perceptron module is used to perform forward propagation on the prediction results of the first TCN and the first GCN to obtain the final prediction result output. The loss function in the multi-layer perceptron module uses the mean square error MSE, and the true value in the mean square error MSE uses the label set Y out .
[0042] Since the traffic flow is also affected by conditions such as the geometric state of the road section and the type of traffic flow, in this embodiment, an adjacency matrix layer is set in the first GCN to form an adjacency matrix. Each element in the adjacency matrix is obtained based on the gantry-related information, presenting the states of each road in the highway section to be predicted. Thus, when predicting, the influence of conditions such as the geometric state of the road section and the type of traffic flow on the traffic flow is eliminated, and the prediction accuracy is further improved.
[0043] Embodiment 3: This embodiment is basically the same as Embodiment 2, except that: in this embodiment, after excluding outliers and normalizing the historical traffic flow data of the highway section to be predicted, the processed historical traffic flow data is used to construct the training set X input and the corresponding label set Y out . The specific process is as follows:
[0044] S1. Obtain the traffic flow data of all gantries within the highway section to be predicted from T b to T c days from the highway management department, where T b represents the first T b days before the current date, and T c represents the first T c days before the current date, T b =1, 1≤T c -T b ≤15;
[0045] The traffic flow data of each lane of each gantry consists of features such as timestamp, gantry ID, section flow, section speed, and section density. Among the traffic flow data of each gantry per day, the time interval between adjacent two traffic flow data timestamps is five minutes, that is, one traffic flow data is recorded every five minutes. Each gantry has 288 traffic flow data, and the number of traffic flow data collected by each gantry is 288(T c -T b ). Denote the number of gantries within the highway section to be predicted as P. Concatenate all traffic flow data with the same timestamp to form a traffic flow data set for the highway section to be predicted, and obtain traffic flow data with a quantity of P * 288(T c -T b ) for the highway section to be predicted;
[0046] S2. Extract other features of each traffic flow data of the highway section to be predicted except for the timestamp and gantry ID, arrange them in ascending order of numerical value, and denote the first quartile and the third quartile after arrangement as Z 1 and Z 3 . Then, first delete the other feature values of each traffic flow data of the highway section to be predicted except for the timestamp and gantry ID that do not satisfy being greater than Z 1 -1.5(Z 3 -Z 1 ) and less than Z 3 +1.5(Z 3 -Z 1 ). Then, use polynomial interpolation to interpolate the deleted features to obtain the interpolated traffic flow data. Split each interpolated traffic flow data into traffic flow data of each gantry, and arrange the traffic flow data of all gantries obtained after splitting in ascending order of gantry ID, and the traffic flow data of each gantry ID in ascending order of timestamp to form the data set X;
[0047] S3. Divide the training set and labels in the data set X to obtain the training set X inpu and the label set Y out . The division method is as follows:
[0048] S3.1. Extract the traffic flow data of each gantry in the dataset X in ascending order of gantry ID. The method of extracting the traffic flow data of each gantry is as follows: in the order of time stamps from earliest to latest, extract n + m traffic flow data each time. The first time, start from the first traffic flow data of the gantry and extract n + m traffic flow data. The second time, start from the second traffic flow data of the gantry and extract n + m traffic flow data, and so on, until the last traffic flow data of the gantry is extracted. At this time, the extraction of the traffic flow data of this gantry is completed. Here, both m and n are positive integers, and n > m, n + m < 288(T c -T b ). For every n + m traffic flow data extracted for a certain gantry, take the first n traffic flow data as 1 input data in the order of time stamps from earliest to latest, and only keep the time stamp, gantry ID, and section flow of the last m traffic flow data as 1 label (regard the time stamp, gantry ID, and section flow as the traffic situation to be output). Thus, after the extraction of the traffic flow data of all gantries in the dataset X is completed, P * [288(T c -T b ) - n - m + 1] input data and P * [288(T c -T b ) - n - m + 1] labels are obtained;
[0049] S3.2. Use P * [288(T c -T b ) - n - m + 1] input data to form the training set X inpu , and use P * [288(T c -T b ) - n - m + 1] labels to form the label set Y out .
[0050] In this embodiment, when making a prediction, collect the traffic flow data of all gantries on the highway section to be predicted in real time, and use the same method for outlier and normalization processing as the obtained training set X input and the corresponding label set Y out . Arrange the traffic flow data of each gantry in ascending order of time stamp to form a prediction dataset, that is, the processed data, and input the processed data into the trained TCN - GCN model to output the future traffic situation of the section to be predicted.
[0051] In this embodiment, build the highway section to be predicted in the sumo simulation software. The structure of the built highway section to be predicted is as Figure 5 shown, Figure 5 where the spatial distance of 16 km indicates that the total length of the highway section to be predicted is 16 km, and there is a gantry every 250 meters, Figure 5Among them, there is a gap between gantry 8 and gantry 11, indicating that the number of lanes narrows from three lanes to two lanes on the road section between gantry 8 and gantry 11. The longitudinal time interval of 5 min means that the data is recorded every 5 minutes. In this embodiment, the traffic flow data for 3 days is recorded in total. After processing the traffic flow data, it is input into the TCN-GCN model for training. The number of training times for both the time module and the spatial module is set to 20 times. The average losses of the loss functions of the time module and the spatial module are plotted, as Figure 6 shown. It can be found from Figure 6 that after 20 rounds of training, the average loss of the TCN-GCN model quickly converges to 0.02. Thus, it can be seen that in the freeway traffic situation prediction method based on TCN-GCN of the present invention, the TCN-GCN model can converge without multiple rounds of iteration, and has the characteristics of small computational complexity and easy training.
[0052] To verify the performance of the freeway traffic situation prediction method based on TCN-GCN of the present invention, two indicators, MAE (Mean Absolute Error) and RMSE (Root Mean Square Error), are used to evaluate the performance of the freeway traffic situation prediction method based on TCN-GCN of the present invention. At the same time, it is compared with the prediction method based on temporal convolution (using the temporal convolutional network TCN as the prediction model) and the prediction method based on graph convolution (using the graph convolutional network as the prediction model). By calculating and comparing the prediction effects of the graph convolutional network, the temporal convolutional network and the TCN-GCN model after 20 times of training, the prediction results are as follows: When using the graph convolutional network after 20 times of training for prediction, the MAE is 24.36 and the RMSE is 53.65. When using the temporal convolutional network after 20 times of training for prediction, the MAE is 26.34 and the RMSE is 58.49. When using the TCN-GCN model after 20 times of training for prediction, the MAE reaches 19.81 and the RMSE reaches 48.76. It can be judged from this that the freeway traffic situation prediction method based on TCN-GCN of the present invention has higher prediction accuracy compared with the prediction methods based on the traditional temporal convolutional network and graph convolutional network.
[0053] After the implementation of the freeway traffic situation prediction method based on TCN-GCN of the present invention, it is loaded into the edge computing devices deployed by the traffic management department or the powerful servers deployed in the data center or cloud platform. By establishing a connection between the edge computing devices or servers and the traffic flow data collection devices deployed by the traffic management department on the freeway, it can receive in real time the traffic flow data uploaded by the traffic flow data collection devices, process the traffic flow data, and make predictions, so as to achieve the purpose of predicting the traffic situation of different sections in the future in real time. The traffic management department can formulate appropriate traffic management decisions according to the prediction results, dynamically adjust traffic signals, issue traffic advice, etc., and realize functions such as traffic warning and congestion management. To sum up, the freeway traffic situation prediction method based on TCN-GCN of the present invention uses the TCN-GCN model that can simultaneously realize the time feature extraction function and the spatial feature extraction function for prediction, can respectively obtain different spatio-temporal change features of the road, fully considers the traffic situation relationship under the full spatio-temporal conditions, has a high prediction accuracy, and directly makes predictions through the trained TCN-GCN model. The prediction process is simple and the calculation amount is small.
Claims
1. A highway traffic situation prediction method based on TCN-GCN, characterized by Based on the time convolution network TCN and the graph convolution neural network GCN, a TCN-GCN model with time feature extraction function and spatial feature extraction function is constructed to obtain the historical traffic flow data on all gantries of the highway section to be predicted, and outliers are excluded and normalized for the historical traffic flow data. The processed historical traffic flow data is used to construct the training set X input and the corresponding label set Y out , using the training set X input and label set Y out The TCN-GCN model is trained to obtain the trained TCN-GCN model, and then the traffic flow data of the highway section to be predicted is obtained in real time. The TCN-GCN model is used for prediction to obtain the traffic situation prediction result of the highway section to be predicted.
2. The highway traffic situation prediction method based on TCN-GCN according to claim 1 is characterized in that The TCN-GCN model includes a time module, a space module and a multi-layer perceptron module. The time module includes a time convolution network TCN, which is called the first TCN; the space module includes a time convolution network TCN and a graph convolution neural network GCN, which is called the second TCN, and the graph convolution neural network GCN is called the first GCN; the first TCN is used to access the traffic flow data input from the outside, and extract the time features of the traffic flow data to obtain data with time features and output it to the multi-layer perceptron module; the second TCN is used to access the traffic flow data input from the outside, and extract the historical spatial features of the traffic flow data to obtain data with historical spatial features and output it to the first GCN; the first GCN is used to extract the spatial features of the data with historical spatial features output by the second TCN to obtain data with spatial features and output it to the multi-layer perceptron module; the multi-layer perceptron module is used to fuse the data with time features output by the first TCN and the data with spatial features output by the first GCN to obtain and output the traffic situation prediction result.
3. The highway traffic situation prediction method based on TCN-GCN according to claim 2 is characterized in that The first TCN includes 5 convolutional layers, 4 residual connection layers and an output layer; the 5 convolutional layers are sequentially referred to as the 1st convolutional layer to the 5th convolutional layer, the 4 residual connection layers are sequentially referred to as the 1st residual connection layer to the 4th residual connection layer, and the output layer is referred to as the first output layer; the 1st convolutional layer is used to access the traffic flow data inputted from the outside, and perform convolution processing on the traffic flow data to obtain the first data with a time series pattern and output it to the 1st residual connection layer; The first residual connection layer is used to access the first data with a time series pattern input to the first convolutional layer and the external traffic flow data input thereto, and sum the first data with a time series pattern and the external traffic flow data input thereto to obtain the sum of the first traffic flow data and output it to the second convolutional layer; the second convolutional layer is used to access the sum of the first traffic flow data input to the first residual connection layer, and convolve the sum of the first traffic flow data to obtain the second data with a time series pattern and output it to the second residual connection layer; the second residual connection layer is used to access the second convolutional layer input thereto The second data with a time series pattern and the external traffic flow data input thereto are summed up to obtain the sum of the second traffic flow data and output to the third convolution layer; the third convolution layer is used to access the second residual connection layer input to the second traffic flow data, and convolve the sum of the second traffic flow data to obtain the third data with a time series pattern and output to the third residual connection layer; the third residual connection layer is used to access the third convolution layer input to the third data with a time series pattern and the external input thereto The fourth convolutional layer is used to access the third traffic flow data inputted by the third convolutional layer, and convolve the third traffic flow data to obtain the fourth data with a time series pattern, and output it to the fourth residual connection layer; the fourth residual connection layer is used to access the fourth data with a time series pattern inputted by the fourth convolutional layer and the external traffic flow data, and convolve the fourth data with a time series pattern to obtain the fourth data with a time series pattern, and output it to the fourth residual connection layer; the fourth residual connection layer is used to access the fourth data with a time series pattern inputted by the fourth convolutional layer and the external traffic flow data, and The external traffic flow data input thereto is summed up to obtain the sum of the fourth traffic flow data and output it to the fifth convolution layer; the fifth convolution layer is used to access the fourth residual link layer to input the sum of the fourth traffic flow data thereto, and perform convolution processing on the sum of the fourth traffic flow data to obtain the fifth data with a time series pattern and output it to the first output layer; when training the TCN-GCN model, the first output layer is used to input the fifth data with a time series pattern input thereto by the fifth convolution layer into the pre-stored loss function thereto, perform back propagation and gradient descent processing, and obtain data Y with time characteristics tcn Output to the multi-layer perceptron module. When the trained TCN-GCN model is used for prediction, the first output layer is used to directly input the fifth data with time series mode at the fifth convolution layer for forward propagation to obtain the prediction result of the first TCN, and input it into the multi-layer perceptron module; the loss function in the first output layer adopts the mean square error MSE, and the true value in the mean square error MSE adopts the label set Y out .
4. The highway traffic situation prediction method based on TCN-GCN according to claim 3 is characterized in that The structure of the second TCN is the same as that of the first TCN, except that the first output layer of the second TCN outputs data to the first GCN instead of the multi-layer perceptron module; the first GCN includes an adjacency matrix layer, two spatial convolution layers, a residual connection layer and an output layer, wherein the output layer is referred to as the second output layer, and the two spatial convolution layers are referred to as the first spatial convolution layer and the second spatial convolution layer, respectively; the adjacency matrix layer is used to extract the spatial features of the highway section to be predicted and perform normalization processing, and the adjacency matrix A with spatial features is obtained and input into the first spatial convolution layer and the second spatial convolution layer, respectively, wherein the adjacency matrix layer extracts the spatial features of the highway section to be predicted and performs normalization processing, and obtains the adjacency matrix A with spatial features. The specific method of matrix A is as follows: the number of gantries of the highway section to be predicted is recorded as P, and the P gantries of the highway section to be predicted are numbered from 1-P in sequence. The gantry numbered w is called gantry w, w = 1, 2, ..., P, gantry i and gantry j represent the gantries of the highway section to be predicted, i = 1, 2, ..., P, j = 1, 2, ..., P, when i ≠ j, and gantry i and gantry j are two adjacent gantries, formula (1) is used to extract the spatial features of the highway section to be predicted and normalize them. When i ≠ j, and gantry i and gantry j are two non-adjacent gantries, formula (2) is used to extract the spatial features of the highway section to be predicted and normalize them. When i = j, formula (3) is used to extract the spatial features of the highway section to be predicted and normalize them: a ij =0 (2) a ij =1 (3) Among them, * is the multiplication symbol, e represents the base of natural logarithm, a ij Represents the element in the i-th row and j-th column of the adjacency matrix A with spatial features, C ij represents the number of lanes between gantry i and gantry j of the highway to be predicted, D ij represents the distance between gantry i and gantry j of the highway to be predicted, Indicates distance D ij The variance of D is obtained by the following method: first calculate the average value of the distances of all adjacent gantries of the highway section to be predicted, use the average value as the average distance, and then calculate D ij The square of the difference from the average distance is divided by P-1 to get D ij The variance of Indicates the number of lanes C ij The variance of C is obtained by the following method: first calculate the average number of lanes between all adjacent gantries of the highway section to be predicted, use the average value as the average number of lanes, and then calculate C ij The square of the difference between the average number of lanes and the average number of lanes is divided by P-1 to get C ij The variance of the first spatial convolution layer is used to access the data output by the first output layer of the second TCN and the adjacency matrix A with spatial features, and convolve the data output by the first output layer of the second TCN and the adjacency matrix A with spatial features to obtain traffic flow data with spatial features and output them to the residual connection layer. The residual connection layer is used to access the data output by the first output layer of the second TCN and the traffic flow data with spatial features input by the first spatial convolution layer, and sum the data output by the first output layer of the second TCN and the traffic flow data with spatial features input by the first spatial convolution layer to obtain the sum of the fifth traffic flow data and output it to the second spatial convolution layer; the second spatial convolution layer is used to access the sum of the fifth traffic flow data and the adjacency matrix A input by the residual connection layer, and convolve the sum of the fifth traffic flow data and the adjacency matrix A input by the residual connection layer to obtain the final data Y with spatial features. space Output to the second output layer; when training the TCN-GCN model, the second output layer is used to access the final data Y with spatial features input at the second spatial convolution layer space , the final data Y space Input the pre-stored loss function for back propagation and gradient descent to obtain data Y with spatial features gcn Input the multi-layer perceptron module; when the trained TCN-GCN model is used for prediction, the second output layer is used to input the final data Y with spatial features at the second spatial convolution layer space Perform forward propagation to obtain the prediction result of the first GCN, and input it into the multi-layer perceptron module; the loss function in the second output layer adopts the mean square error MSE, and the true value in the mean square error MSE adopts the label set Y out .
5. The highway traffic situation prediction method based on TCN-GCN according to claim 4 is characterized in that When training the TCN-GCN model, the multi-layer perceptron module is used to access the first output layer of the first TCN to input the data Y with time characteristics therein. tcn The second output layer of the first GCN inputs the data Y with spatial features gcn , and transform the data Y with time characteristics tcn and data Y with spatial characteristics gcn The pre-stored loss function is input, and after back propagation and gradient descent processing, a prediction result Y with a spatiotemporal relationship is obtained as an output; when the trained TCN-GCN model is used for prediction, the multi-layer perceptron module is used to forward propagate the prediction result of the first TCN and the prediction result of the first GCN to obtain the final prediction result output, and the loss function in the multi-layer perceptron module adopts the mean square error MSE, and the true value in the mean square error MSE adopts the label set Y out .
6. The highway traffic situation prediction method based on TCN-GCN according to claim 1 is characterized in that After outliers are eliminated and normalized for the historical traffic flow data of the highway section to be predicted, the processed historical traffic flow data are used to construct the training set X input and the corresponding label set Y out The specific process is: S1. Obtain all the gate frames in the highway section to be predicted from the highway management department. b to T c Traffic flow data of the day, where T b Indicates the first T of the current date b Day, T c Indicates the first T of the current date c Day, T b =1,1≤T c -T b ≤15; Each traffic flow data of each gantry is composed of the features of timestamp, gantry ID, cross-sectional flow, cross-sectional speed and cross-sectional density. In the traffic flow data of each gantry every day, the time interval between two adjacent traffic flow data is five minutes, that is, one traffic flow data is recorded every five minutes. Each gantry has 288 traffic flow data, and the number of traffic flow data collected by each gantry is 288 (T c -T b ), the number of gantries in the highway section to be predicted is recorded as P, and all traffic flow data with the same timestamp are spliced to form a traffic flow data set of the highway section to be predicted, and the number of highway sections to be predicted is P*288(T c -T b )’s traffic flow data; S2, extract the other features of each traffic flow data of the highway section to be predicted except the timestamp and the gantry ID, arrange them in ascending order according to the values, take the first quartile and the third quartile after the arrangement as Z1 and Z3 respectively, then first delete the other feature values of each traffic flow data of the highway section to be predicted except the timestamp and the gantry ID that do not satisfy the value greater than Z1-1.5 (Z3-Z1) and less than Z3+1.5 (Z3-Z1), and then use the polynomial interpolation method to interpolate the deleted features to obtain the interpolated traffic flow data, split each traffic flow data after interpolation into the traffic flow data of each gantry, and arrange the traffic flow data of all gantries obtained after the splitting in the order of the gantry ID from small to large, and the traffic flow data of each gantry ID is arranged in ascending order according to the timestamp, so as to form the data set X; S3. Divide the training set and the label in the data set X to obtain the training set X inpu and label set Y out , the division is as follows: S3.
1. Extract the traffic flow data of each gantry in the data set X in order from small to large according to the gantry ID, wherein the method of extracting the traffic flow data of each gantry is: extract n+m traffic flow data each time according to the time stamp from the earliest to the latest, wherein the first time, extract n+m traffic flow data from the first traffic flow data of the gantry, the second time, extract n+m traffic flow data from the second traffic flow data of the gantry, and so on, until the last traffic flow data of the gantry is extracted, at which time the traffic flow data extraction of the gantry is completed, wherein m and n are both positive integers, and n>m, n+m<288(T c -T b ), each time a certain gantry's n+m traffic flow data is extracted, the first n traffic flow data are used as 1 input data in the order of timestamp from the earliest to the latest, and the last m traffic flow data only retain the timestamp, gantry ID, and cross-sectional flow as 1 label (timestamp, gantry ID, and cross-sectional flow are used as the output traffic situation). After the traffic flow data of all gantry in the data set X are extracted, P*[288(T c -T b )-n-m+1] input data and P*[288(T c -T b )-n-m+1] labels; S3.2, using P*[288(T c -T b )-n-m+1] input data constitute the training set X inpu , using P*[288(T c -T b )-n-m+1] labels constitute the label set Y out .
7. The highway traffic situation prediction method based on TCN-GCN according to claim 1 is characterized in that The traffic flow data of all gantries on the highway section to be predicted are collected in real time, and outlier and normalization processing are performed to obtain the processed data. The processed data is then input into the trained TCN-GCN model to output the future traffic situation of the section to be predicted.