A traffic flow prediction method based on a residual network

By introducing new data division strategies, external data feature extraction networks and parallel convolutional networks into the traffic flow prediction method, the shortcomings of the existing methods in time and space characteristics are solved, and more efficient data utilization and model deployment are achieved, which is convenient for improving the traffic flow prediction effect.

CN117218841BActive Publication Date: 2025-06-20CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202311184661.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-14
Publication Date
2025-06-20
Estimated Expiration
2043-09-14

AI Technical Summary

Technical Problem

The existing traffic flow prediction methods have shortcomings when considering the time and spatial characteristics. Excessive model parameters are not conducive to deployment and insufficient utilization of external data.

Method used

A traffic flow prediction method based on residual network is proposed, through a new data division strategy and an external data feature extraction network, combining parallel convolutional network and a new residual element structure, temporal and spatial features are extracted, and external data is effectively utilized.

Benefits of technology

It realizes more effective use of historical vehicle traffic data and external data, reduces model parameters, improves prediction effect, and enhances the learning ability and deployment convenience of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117218841B_ABST
    Figure CN117218841B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of artificial intelligence, and particularly relates to a traffic flow prediction method based on a residual network, including: constructing a traffic flow prediction model, wherein the traffic flow prediction model includes a data input module, an external feature extraction module, and a spatial feature extraction module; obtaining information data of a location to be predicted, inputting the traffic data into the data input module for splitting to obtain traffic flow data at key times; inputting the external data into the external feature extraction module to obtain external features; after fusing the external features and the traffic flow data at key times, inputting the fused feature map into the spatial feature extraction module for traffic flow prediction to obtain a prediction result; the present invention uses a new data partitioning strategy to more effectively use historical traffic flow data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and particularly relates to a traffic flow prediction method based on a residual network. Background Art

[0002] Traffic flow prediction is to use the historical traffic flow data of different regions of the city to predict how much traffic flow there will be in these regions at the next moment. Traffic flow prediction plays an important role in the urban intelligent management system, which can help reduce traffic congestion and traffic accidents in the city. The traffic flow prediction task has begun to be widely studied. Traditional prediction methods use historical averages or ARIMA models to predict the traffic flow information at the next moment. Such methods ignore the spatio-temporal characteristics of traffic flow and only consider part of the time characteristics. The learning-based traffic flow prediction technology is a popular direction in recent years. By means of the convolutional neural network technology of end-to-end mapping in deep learning, the complex time and space characteristics in traffic flow are learned to achieve better prediction results.

[0003] Traditional traffic flow prediction methods include the traffic flow prediction model based on ARIMA. According to the number of time periods to be predicted, the short-term traffic flow data can be divided into corresponding dataset groups, and then the traffic flow in the next time period is predicted by each dataset group. However, the results of this model are relatively simple and do not consider the influence of spatial factors on traffic flow. Zhang et al. found that there is a certain periodicity in traffic flow, and the input data is divided, and the key time point data is extracted. And a CNN network is used to extract the influence of spatial factors on traffic flow, and a residual network is used to solve the problems of gradient disappearance and gradient explosion caused by too many network layers. However, such models require stacked residual networks, resulting in too large model parameters and being not conducive to deployment. Xie et al. proposed to use the VIT network to replace the CNN, which avoids the problem of too large model parameters caused by excessive stacking of residual layers. However, the VIT-based model requires a large amount of training data for training, and the model performance is not good when the training data is less. Guo et al. used the TCN network to enable the model to predict traffic flow from longer time inputs. However, this model only focuses on more time characteristics and lacks consideration of spatial characteristics. Lin et al. improved the prediction effect of the model by introducing new external data (POI). However, with the rapid development of the urban economy, various online celebrity scenic spots emerge in an endless stream, making the POI data need to be updated in time, resulting in low practicability of this method.

[0004] In summary, the prior art has the following technical problems:

[0005] 1. The consideration of time characteristics is insufficient, and not all useful traffic flow information in historical data is fully considered, which results in the insufficient utilization of the historical data input at one time.

[0006] 2. By stacking multiple layers of CNN residual networks to capture spatial characteristics, simply stacking more residual layers will not bring a huge improvement in the extraction of spatial characteristics. Instead, it will cause a huge increase in the number of model parameters, which is not conducive to the deployment of the model.

[0007] 3. The easily accessible external data (such as climate, holidays, etc.) is not fully utilized, and a good sub-network is not constructed to extract information from the external data. External data also has an important impact on the traffic flow of a city. Summary of the Invention

[0008] To solve the above problems existing in the prior art, the present invention proposes a traffic flow prediction method based on a residual network. The method includes: obtaining information data of a location to be predicted, and inputting the information data into a trained traffic flow prediction model to obtain a traffic flow prediction result of the location to be predicted; wherein, the traffic flow prediction model includes a data input module, an external feature extraction module, and a spatial feature extraction module.

[0009] Training the traffic flow prediction model includes:

[0010] S1: Obtain traffic data and external data of the location to be predicted, where the external data includes temperature, wind speed, weather, date, and whether it is a holiday.

[0011] S2: Input the traffic data into the data input module for splitting to obtain key time traffic flow data.

[0012] S3: Input the external data into the external feature extraction module to obtain external features.

[0013] S4: Fuse the external features and the key time traffic flow data, and input the fused feature map into the spatial feature extraction module for traffic flow prediction to obtain a prediction result.

[0014] S5: Calculate the loss function of the model according to the prediction result, adjust the model parameters, and complete the training of the model when the loss function converges.

[0015] Preferably, the formula for splitting the traffic data by inputting it into the data input module is:

[0016] F c ,F p ,F pn =Division(F t )

[0017] Among them, F c represents the traffic flow data of several moments before the current moment, F p represents the traffic flow data of the current moment in the past few days, F pn represents the traffic flow data of the next moment of the current moment in the past few days, F t represents all the traffic flow data of the past week of the current moment, and Division represents the data splitting operation.

[0018] Preferably, the external feature extraction module processes external data as follows: converting discrete data in the external data into one-hot labels, normalizing continuous data in the external data; splicing the one-hot labels with the normalized data; inputting the spliced data into the external feature embedding layer to obtain an external feature vector; performing convolution on the external feature vector, transforming the shape of the convolved features to obtain external data; performing an operation on the external data to extract the internal relationship of discrete features; inputting the extracted relationship and the external data into the external feature capture network to obtain external features; the formula for the operation of extracting the internal relationship of discrete features is:

[0019] Ext d = Ext D ×W

[0020] Among them, W is a matrix of learnable parameters, Ext D is the unprocessed discrete external data, and Ext d is the extracted discrete external data.

[0021] Furthermore, the external feature capture network extracts features from the external data, including:

[0022] Ext con = Concat(Normal(Ext c ), Ext d )

[0023] Among them, Ext c represents the continuous external data; Ext con represents the discontinuous external data; Vec represents the operation of converting to one-hot labels, Normal represents the normalization operation, and Concat represents the splicing operation; Ext con is the preliminarily processed external data.

[0024] Preferably, the spatial feature extraction module includes a dilated convolutional residual unit, an ordinary convolutional residual unit, a fusion module, and a fully connected layer. The dilated convolutional residual unit and the ordinary convolutional residual unit are arranged in parallel to form a parallel convolutional module.

[0025] Preferably, the process of processing the fused feature map by the spatial feature extraction module includes: inputting the fused feature map into a dilated convolutional residual unit for dilated convolutional operation, fusing the output result of the dilated convolution with the fused feature map to obtain a first feature map; inputting the fused feature map into a normal convolutional residual unit for convolutional operation, fusing the result of the convolutional operation with the fused feature map to obtain a second feature map; using a fusion module to fuse the first feature map and the second feature map to obtain a third fused feature map; inputting the third fused feature map into a fully connected layer to obtain a prediction result.

[0026] Further, the formula for the dilated convolutional residual unit to process the fused feature map is:

[0027]

[0028]

[0029] Among them, F represents the input data or the output of the previous residual network, represents the result of the first dilated convolutional operation in a residual element, represents the output of this residual element.

[0030] Further, the formula for the normal convolutional residual unit to process the fused feature map is:

[0031]

[0032]

[0033] Among them, F represents the input data or the output of the previous residual network, represents the result of the first convolutional operation in a residual element, represents the output of this residual element.

[0034] Further, the prediction result is:

[0035]

[0036] Among them, represents the output of the normal convolutional sub-network, represents the output of the dilated convolutional sub-network, f is a normal convolutional operation, is the final prediction result.

[0037] Preferably, the loss function of the model uses the root mean square error and the mean absolute error as evaluation indicators for the result; its expression is:

[0038]

[0039]

[0040] Among them, MSE represents the mean square error, MAX represents the maximum value among pixel values, F represents the true value of the traffic flow in the next moment of this area, represents the predicted value of the traffic flow in the next moment of this area.

[0041] Advantages of the present invention:

[0042] The present invention uses a new data partitioning strategy to more effectively use historical traffic flow data; the present invention proposes a brand-new external data feature extraction network, which can deeply excavate the mutual internal connection relationships between different external data, enabling more comprehensive utilization of external data. The present invention constructs a parallel convolutional network, enabling the model to separately extract the traffic flow impacts of adjacent and distant areas on the target area, and this structure can significantly reduce network parameters, enabling the model to be lightweight while obtaining good prediction results. The present invention constructs a new residual element structure. Compared with the previous serial residual element structure, the residual element of the present invention can incorporate the results of intermediate operations into the final result summation, making the learning ability of the model stronger. Description of the drawings

[0043] Figure 1 is the overall flowchart of the present invention;

[0044] Figure 2 is the structural diagram of the external feature extraction network of the present invention;

[0045] Figure 3 is the structural diagram of the parallel convolutional module of the present invention. Detailed implementation manners

[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0047] A traffic flow prediction method based on a residual network, as Figure 1 shown, the method includes: obtaining information data of a location to be predicted, inputting the information data into a trained traffic flow prediction model, and obtaining a traffic flow prediction result of the location to be predicted; wherein, the traffic flow prediction model includes a data input module, an external feature extraction module, and a spatial feature extraction module.

[0048] Training the traffic flow prediction model includes:

[0049] S1: Obtain the traffic data and external data of the location to be predicted, where the external data includes temperature, wind speed, weather, date, and whether it is a holiday;

[0050] S2: Input the traffic data into the data input module for splitting to obtain the traffic flow data at key times;

[0051] S3: Input the external data into the external feature extraction module to obtain external features;

[0052] S4: Integrate the external features and the traffic flow data at key times, input the integrated feature map into the spatial feature extraction module for traffic flow prediction to obtain the prediction result;

[0053] S5: Calculate the loss function of the model based on the prediction result, adjust the model parameters, and complete the training of the model when the loss function converges.

[0054] In this embodiment, the traffic data and external data of the location to be predicted are for the common open datasets such as TaxiBJ, TaxiCQ, BikeNYC, etc. to achieve traffic flow prediction. The entire design method includes an input module (data splitting), an external feature extraction module, and a spatial feature extraction module (parallel residual network structure).

[0055] The traditional data splitting method is to extract the adjacent moments in a time slice, the same moment in the past few days, and the same moment in the past few weeks. However, there are the following problems with such methods: (1) The data in the past few days will overlap with the weekly data; (2) All this data is historical data, and there will be a large prediction error at the turning point. Therefore, the data of the same moment in the past few weeks is replaced with the data of the next moment in the past few days, so that the model can handle the prediction problem at the turning point well; thus, the present invention uses a data input module to split the traffic data and extract the effective key time slice data. Specifically, it includes:

[0056] F c ,F p ,F pn =Division(F t )

[0057] where, F c represents the traffic flow data of several moments before the current moment, F p represents the traffic flow data of the current moment in the past few days, F pn represents the traffic flow data of the next moment of the current moment in the past few days, F t represents all the traffic flow data of the past week of the current moment, and Division represents the data splitting operation. For example, Figure 2As shown in the figure, the external feature extraction module processes external data as follows: converting discrete data in the external data into one-hot labels, normalizing continuous data in the external data; concatenating the one-hot labels with the normalized data; inputting the concatenated data into the external feature embedding layer to obtain an external feature vector; convolving the external feature vector and transforming the shape of the convolved features to obtain external data. The traditional external data processing method only performs one-hot processing on discrete data, but it can be found from real life that there is a certain internal connection between different external data. For example, today is a holiday, but if the weather is sunny, the traffic flow will definitely be higher than when the weather is rainy. In order to enable the model to discover this internal connection feature, an operation for extracting the internal relationship of discrete features is performed on the one-hot discrete data, and the formula is as follows: Ext d = Ext D × W

[0058] where W is a matrix of learnable parameters, Ext D is the unprocessed discrete external data, and Ext d is the extracted discrete external data.

[0059] The formula for concatenating the one-hot labels with the normalized data is:

[0060] Ext con = Concat(Normal(Ext c ), Vec(Ext d ))

[0061] Inputting the external data into the external feature capture network to obtain external features; the formula for obtaining external data is:

[0062] Ext con = Concat(Normal(Ext c ), Ext d )

[0063] where, Ext c represents continuous external data; Ext con represents discontinuous external data; Vec represents the operation of converting to one-hot labels, Normal represents the normalization operation, and Concat represents the concatenation operation; Ext con is the externally processed data after preliminary processing.

[0064] The traffic flow data at the selected key time points and the extracted external feature data are concatenated and then input into the parallel convolutional network. Using L1 as the loss function, a parameter model for traffic flow prediction is trained. Then, the TaxiBJ, TaxiCQ, and BikeNYC data sets are used to verify the effect of the model. The root mean square error (RMSE) and mean absolute error (MAE) are used as evaluation indicators. The lower the value of both, the better the model prediction effect.

[0065] like Figure 3 As shown in FIG. 1 , the spatial feature extraction module includes a dilated convolution residual element, a normal convolution residual element, a fusion module, and a fully connected layer. The dilated convolution residual element and the normal convolution residual element are arranged in parallel to form a parallel convolution module.

[0066] The process of using the spatial feature extraction module to process the fused feature map includes: inputting the fused feature map into the dilated convolution residual unit to perform a dilated convolution operation, fusing the output result of the dilated convolution with the fused feature map to obtain a first feature map; inputting the fused feature map into the ordinary convolution residual unit to perform a convolution operation, fusing the result of the convolution operation with the fused feature map to obtain a second feature map; using the fusion module to fuse the first feature map and the second feature map to obtain a third fused feature map; inputting the third fused feature map into the fully connected layer to obtain a prediction result.

[0067] The formula for processing the fused feature map by the dilated convolution residual unit is:

[0068]

[0069]

[0070] Among them, F represents the input data or the output of the previous residual network. represents the result of the first dilated convolution operation in a residual element, Represents the output of the residual element.

[0071] The formula for processing the fused feature map by the ordinary convolution residual unit is:

[0072]

[0073]

[0074] Among them, F represents the input data or the output of the previous residual network. represents the result of the first convolution operation in a residual element, Represents the output of the residual element.

[0075] The prediction results are:

[0076]

[0077] Among them, represents the output of the ordinary convolutional sub-network, represents the output of the dilated convolutional sub-network, and f is the ordinary convolutional operation, is the final prediction result.

[0078] The loss function of the model uses the root mean square error and the mean absolute error as evaluation metrics for the results; its expression is:

[0079]

[0080]

[0081] Among them, MSE represents the mean square error, MAX represents the maximum value in the pixel values, F represents the true value of the traffic flow in the next moment in this area, represents the predicted value of the traffic flow in the next moment in this area.

[0082] This embodiment uses the Python programming language and can run on mainstream computer platforms. The operating system used in this embodiment is CentOS 6.5, requiring an Intel i7 CPU, more than 16GB of memory, a hard disk space requirement of 32GB or more, an NVIDIA Tesla V100 GPU, and 32G of video memory.

[0083] This invention implements the content of this invention based on the PyTorch 0.4 framework and uses the Adam optimization algorithm to update the parameters of the model.

[0084] The datasets used are the TaxiBJ, TaxiCQ, and BikeNYC datasets. Among them, TaxiBJ contains 527 days of Beijing taxi traffic data, TaxiCQ contains 92 days of Chongqing taxi traffic data, and BikeNYC contains 183 days of New York bike traffic data. The first 80% of the data in each dataset is selected as the training set, the last 10% is used as the test set, and the remaining part is used as the validation set. The evaluation metrics are the traditional RMSE and MAE.

[0085] The model performing one pass of the gradient descent algorithm on all training data is called one epoch. Each epoch will update the parameters of the model, and the maximum number of epochs is set to 150. The learning rate is set to be updated every 10 iterations. During the 150 iterations of training the model, the model and its parameters that achieve the best results on the test dataset are saved.

[0086] The above-described embodiments further illustrate in detail the objectives, technical solutions, and advantages of the present invention. It should be understood that the above-described embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A traffic flow prediction method based on a residual network, characterized in that, Including: Obtain the information data of the location to be predicted, and input the information data into the trained traffic flow prediction model to obtain the traffic flow prediction result of the location to be predicted; among them, the traffic flow prediction model includes a data input module, an external feature extraction module, and a spatial feature extraction module; Training the traffic flow prediction model includes: S1: Obtain the traffic data and external data of the location to be predicted, where the external data includes temperature, wind speed, weather, date, and whether it is a holiday; S2: Input the traffic data into the data input module for splitting to obtain the traffic flow data at key times; S3: Input the external data into the external feature extraction module to obtain external features; specifically including: converting the discrete data in the external data into one-hot labels, normalizing the continuous data in the external data; splicing the one-hot labels with the normalized data; inputting the spliced data into the external feature embedding layer to obtain an external feature vector; performing convolution on the external feature vector, transforming the shape of the convolved feature to obtain external data; performing an operation on the external data to extract the internal relationship of discrete features; inputting the extracted relationship and external data into the external feature capture network to obtain external features; where the formula for the operation of extracting the internal relationship of discrete features is: Ext d = Ext D × W Among them, W is a matrix of learnable parameters, and Ext D is the unprocessed discrete external data, and Ext d is the extracted discrete external data; S4: Fuse the external features and the traffic flow data at key times, and input the fused feature map into the spatial feature extraction module for traffic flow prediction to obtain a prediction result; the process of using the spatial feature extraction module to process the fused feature map includes: inputting the fused feature map into the dilated convolutional residual unit for dilated convolutional operation, fusing the output result of the dilated convolution with the fused feature map to obtain a first feature map; inputting the fused feature map into the ordinary convolutional residual unit for convolutional operation, fusing the result of the convolutional operation with the fused feature map to obtain a second feature map; using a fusion module to fuse the first feature map and the second feature map to obtain a third fused feature map; inputting the third fused feature map into the fully connected layer to obtain a prediction result; S5: Calculate the loss function of the model according to the prediction result, adjust the model parameters, and complete the training of the model when the loss function converges.

2. The traffic flow prediction method based on a residual network according to claim 1, characterized in that, In the spatial feature extraction module, the dilated convolutional residual element and the ordinary convolutional residual element are juxtaposed to form a parallel convolutional module.

3. The traffic flow prediction method based on a residual network according to claim 1, characterized in that, The formula for the dilated convolutional residual element to process the fused feature map is: Among them, F represents the input data or the output of the previous residual network layer, represents the result of the first dilated convolution operation in a residual element, represents the output of this residual element.

4. The traffic flow prediction method based on a residual network according to claim 1, characterized in that, The formula for the ordinary convolutional residual element to process the fused feature map is: Among them, F represents the input data or the output of the previous residual network layer, represents the result of the first convolution operation in a residual element, represents the output of the residual element.

5. The traffic flow prediction method based on a residual network according to claim 1, characterized in that, The prediction result is: Among them, represents the output of the ordinary convolutional sub-network, represents the output of the dilated convolutional sub-network, and f is the ordinary convolutional operation, is the final prediction result.

6. The traffic flow prediction method based on a residual network according to claim 1, characterized in that, The loss function of the model uses the root mean square error RMSE and the mean absolute error MAE as evaluation indicators for the result; its expression is: Among them, F represents the true value of the traffic flow at the next moment at the location to be predicted, represents the predicted value of the traffic flow at the next moment at the location to be predicted.

Citation Information

Patent Citations

  • Traffic flow prediction method, system and device and readable storage medium

    CN115116227A

  • Traffic flow prediction method of space-time attention graph convolutional network based on multi-feature fusion

    CN116168548A

Cited By

  • Subway-bus transfer demand prediction method based on multi-graph feature fusion network

    CN122311542A