Traffic flow prediction method based on multi-dimensional collaborative space-time attention model
By constructing a multi-dimensional co-spatial attention model, combining local multi-channel cross-attention and cross-channel lightweight convolutional association modules, the existing traffic flow prediction technology has solved the problems of low accuracy and low efficiency in spatiotemporal dependency processing, and achieved more efficient traffic flow prediction.
Patent Information
- Application Number
- CN202510401391.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
The existing traffic flow prediction technology has low accuracy and low computational efficiency when dealing with complex space-time dependencies, and ignores the influence of multi-dimensional factors, resulting in inaccurate prediction results.
A multi-dimensional synchronous spatio-attention model (MCSTA) is constructed, including a Transformer self-attention module, a local multi-channel cross-attention (LMCEA) module and a cross-channel lightweight convolution association (NDLCCA) module. Through adaptive focus on multi-dimensional feature interactions of spatio-temporal data, one-dimensional convolution captures cross-channel correlation features to reduce the computational complexity.
It significantly improves the accuracy and efficiency of traffic flow prediction, can more accurately capture the multi-dimensional characteristics of traffic flow, and improves the prediction accuracy and calculation efficiency of the model.
Smart Images

Figure CN120258235A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent transportation systems, and particularly to a method for improving the accuracy and efficiency of traffic flow prediction by using an improved Transformer architecture. Background Art
[0002] Traffic flow prediction plays a crucial role in the efficient management and operation of modern urban transportation systems. With the rapid advancement of urbanization and the continuous increase in the number of road vehicles, accurate prediction of traffic flow has become a key challenge for urban planners and traffic management departments.
[0003] In addition, accurate traffic flow prediction is a core component of intelligent transportation systems, which can alleviate road congestion, improve fuel efficiency, and ultimately improve the quality of life of residents in large cities. Although existing traffic flow prediction technologies have made certain progress, there are still many problems.
[0004] On the one hand, traditional convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have limited capabilities in dealing with the complex spatio-temporal dependencies of traffic data and are difficult to capture the deep-level correlations in the data. On the other hand, existing attention mechanism models, such as Transformer, although can capture spatio-temporal dependencies to a certain extent, have a high computational complexity and face challenges in terms of efficiency in practical applications. In addition, these models often ignore the multi-dimensional characteristics of traffic data, such as the impact of factors like weather and time on traffic flow, resulting in inaccurate prediction results.
[0005] Therefore, there is an urgent need for an efficient prediction method that can predict traffic flow in multiple dimensions to improve the accuracy, precision, and efficiency of traffic flow prediction.
[0006] The information disclosed in this background art section is only intended to enhance the overall understanding of the present invention and should not be regarded as an admission or any form of suggestion that this information constitutes the prior art already known to those of ordinary skill in the art. Summary of the Invention
[0007] The object of the present invention is to provide a traffic flow prediction method based on a multi-dimensional collaborative spatio-temporal attention model, so as to solve the problems such as low accuracy of current traffic flow prediction, slow calculation efficiency of the prediction model, and large deviation between the prediction result and the true value.
[0008] To achieve the above object, the present invention provides a traffic flow prediction method based on a multi-dimensional collaborative spatio-temporal attention model, including:
[0009] Step 1, collect urban traffic data, and perform cleaning, normalization, and one-hot encoding processing on the data;
[0010] Step 2: Construct a multi-dimensional collaborative spatio-temporal attention model (MCSTA), which includes a Transformer self-attention module, an LMCEA module, and an NDLCCA module;
[0011] Step 3: Train the model. Divide the dataset into different time windows, and use the pre-training and fine-tuning parameter methods to train the model and evaluate the training results;
[0012] Step 4: Conduct ablation experiments.
[0013] Step 5: Implement deployment. Deploy the model into the intelligent transportation system to achieve real-time prediction and application of traffic flow.
[0014] Preferably, in the above technical solution, the data collected in Step 1 is the traffic flow data of shared bicycles and taxis in New York. After dividing the urban area into grids, record the inflow and outflow information of vehicles in each grid area. The urban area is evenly and regularly divided into I×J grids. For the grid (i,j) located in the i-th row and the j-th column, its inflow and outflow volumes within the time interval t are defined as follows:
[0015]
[0016] Preferably, in the above technical solution, in Step 1, the collected traffic data is cleaned to remove outliers and noise, and the Min-Max normalization method is used to scale the traffic flow values to the range of [-1,1], and one-hot encoding processing is performed on external factors such as weather and time.
[0017] Preferably, in the above technical solution, the construction of the multi-dimensional collaborative spatio-temporal attention model (MCSTA) in Step 2 includes:
[0018] (1) Integrate the proposed LMCEA module into the backbone network part of the Transformer self-attention module in parallel. The LMCEA module contains three parallel branches, which respectively process the features in the width (W), height (H), and channel (C) directions of the space. Each branch generates corresponding attention weights through steps such as rotation operation, squeezing transformation, and excitation transformation, so as to enhance the model's attention to features in different dimensions.
[0019] (2) Integrate the proposed NDLCCA module in parallel into the backbone network part of the Transformer self-attention module. The NDLCCA module uses one-dimensional convolution to capture local cross-channel interactions, avoiding dimensionality reduction operations. It aggregates spatial information through a global average pooling layer, and then adjusts the features through an operation layer controlled by parameters and an activation normalization layer. After that, the data is divided into two parallel branches. One branch passes through multiple alternating convolutional layers and a local cross-channel interaction (LCCI) module, and the other branch passes through an adaptive kernel size determination mechanism (AKSDM) module. The data from the two branches is merged at the fusion node to generate the final output.
[0020] 1) Local cross-channel interaction (LCCI) module:
[0021] The input features are first transformed by a convolutional layer. The transformed features are formed into multiple sub-feature sets through a splitting operation. Each sub-feature set uses one-dimensional convolution to extract in-channel feature information. The processed sub-feature sets are concatenated and merged, and then feature fusion and adjustment are completed through another convolutional layer to enhance the feature dependence ability and capture complex cross-channel relationships.
[0022] 2) Adaptive kernel size determination mechanism (AKSDM) module:
[0023] The input features first unify the feature representation form by a convolutional layer, and then enter a structure composed of N cascaded Blocks. Each Block contains two convolutional layers with a convolutional kernel of 3×3. The first convolutional layer is used for preliminary feature extraction, and then a normalization layer is set behind it to accelerate network convergence and reduce internal covariate shift. The second convolutional layer is used for further feature extraction, and a normalization layer is also set behind it. After connecting the features output by the N Blocks through the ReLU activation function, multi-scale features are integrated through a convolutional operation to output global features, realizing multi-scale hierarchical feature extraction.
[0024] Preferably, in the above technical solution, in step 3, the traffic flow data set is divided into three time windows of 30 minutes, 60 minutes, and 90 minutes. It is directly trained on the 30-minute data set, and the parameters are fine-tuned and trained on the 60-minute and 90-minute data sets.
[0025] Preferably, in the above technical solution, in step 3, the mean square error (MSE), root mean square error (RMSE), and mean absolute error (MAE) metrics are used to evaluate the training results.
[0026]
[0027] Preferably, in the above technical solution, the ablation experiment in step 4 on the traffic flow dataset includes removing the LMCEA module and the NDLCCA module respectively, and comparing the trained results with the unchanged MCSTA model.
[0028] Preferably, in the above technical solution, in step 5, the method will be deployed on the intelligent transportation system to predict the traffic flow of the city in real time, and the intelligent transportation construction will be rationally arranged according to the prediction results.
[0029] (1) The present invention is a traffic flow prediction method based on a multi-dimensional collaborative spatio-temporal attention model. The innovatively proposed MCSTA model adds the proposed LMCEA module and NDLCCA module in parallel on the basis of the original Transformer self-attention module. The proposed LMCEA module captures time and space dependence relationships from multiple dimensions, can adaptively focus on the local feature interactions between different influencing factors in traffic flow data, and extract more discriminative features. The proposed NDLCCA module uses one-dimensional convolution to capture local cross-channel interactions, while avoiding dimensionality reduction, which greatly reduces the computational complexity, and can more efficiently utilize the correlation between channels to accurately capture the associated features in traffic flow data. It improves the accuracy, precision and efficiency of the model for traffic flow prediction.
[0030] (2) The present invention is a traffic flow prediction method based on a multi-dimensional collaborative spatio-temporal attention model. When this method is deployed on the intelligent transportation system and applied to the construction of modern smart cities, the specific effect exceeds the theoretical practical application. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is the overall flowchart of the traffic flow prediction method based on the multi-dimensional collaborative spatio-temporal attention model according to the present invention;
[0032] Figure 2 is the structure diagram of the MCSTA model;
[0033] Figure 3 is the structure diagram of the proposed LMCEA module;
[0034] Figure 4 is the structure diagram of the proposed NDLCCA module;
[0035] Figure 5 is the structure diagram of the LCCI module and the AKSDM module in the added NDLCCA module;
[0036] Figure 6 is the comparison diagram of the traffic flow prediction value and the real value of the MCSTA model within 24 hours. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following will describe in detail the specific embodiments of the present invention in conjunction with the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific embodiments.
[0038] Unless otherwise explicitly stated, throughout the specification and claims, the term "comprise" or its variations such as "comprises" or "comprising" etc. will be understood to include the stated elements or components, without excluding other elements or other components.
[0039] As Figures 1 to 6 shown, a traffic flow prediction method based on a multi-dimensional collaborative spatio-temporal attention model according to a specific embodiment of the present invention includes the following steps:
[0040] Step 1: Collect urban traffic data and perform cleaning, normalization, and one-hot encoding on the data.
[0041] (1) The collected data is the traffic flow data of shared bicycles and taxis in New York. After dividing the urban area into grids, record the inflow and outflow information of vehicles in each grid area. The urban area is evenly and regularly divided into I×J grids. For the grid (i,j) located in the i-th row and the j-th column, its inflow and outflow within the time interval t are defined as follows:
[0042]
[0043] (2) Clean the collected traffic data, remove outliers and noise, use the Min-Max normalization method to scale the traffic flow values to the range of [-1,1], and perform one-hot encoding on external factors such as weather and time.
[0044] Step 2: Construct a multi-dimensional collaborative spatio-temporal attention model (MCSTA), which includes a Transformer self-attention module, an LMCEA module, and an NDLCCA module.
[0045] (1) Integrate the proposed LMCEA module in parallel into the backbone network part of the Transformer self-attention module. The LMCEA module contains three parallel branches, which respectively process the features in the width (W), height (H), and channel (C) directions of space. Each branch generates corresponding attention weights through steps such as rotation operation, squeezing transformation, and excitation transformation, so as to enhance the model's attention to features in different dimensions.
[0046] The LMCEA module captures time and space dependence relationships from multiple dimensions, can adaptively focus on the local feature interactions between different influencing factors in traffic flow data, and extract more discriminative features.
[0047] (2) Integrate the proposed NDLCCA module in parallel into the backbone network part of the Transformer self-attention module. The NDLCCA module uses one-dimensional convolution to capture local cross-channel interactions, avoiding dimensionality reduction operations. This module aggregates spatial information through a global average pooling layer, and then adjusts the features through an operation layer controlled by parameters and an activation normalization layer. Then, the data is divided into two parallel branches. One branch passes through multiple alternating convolutional layers and the local cross-channel interaction (LCCI) module, and the other branch passes through the adaptive kernel size determination mechanism (AKSDM) module. Finally, the data from the two branches is merged at the fusion node to generate the final output.
[0048] The NDLCCA module uses one-dimensional convolution to capture local cross-channel interactions, while avoiding dimensionality reduction, which greatly reduces the computational complexity, and can more efficiently utilize the correlation between channels, realizing the accurate capture of associated features in traffic flow data. It improves the accuracy, precision, and efficiency of the model for traffic flow prediction.
[0049] 1) Local cross-channel interaction (LCCI) module:
[0050] The input features are first transformed by a convolutional layer. The transformed features are formed into multiple sub-feature sets through a splitting operation. Each sub-feature set uses one-dimensional convolution to extract in-channel feature information. The processed sub-feature sets are concatenated and merged, and then feature fusion and adjustment are completed through another convolutional layer, so as to enhance the feature dependence ability and capture complex cross-channel relationships.
[0051] 2) Adaptive kernel size determination mechanism (AKSDM) module:
[0052] The input features first unify the feature expression form by a convolutional layer, and then enter a structure composed of N Blocks in cascade. Each Block contains two convolutional layers with a convolution kernel of 3×3. The first convolutional layer is used for preliminary feature extraction, and then a normalization layer is set to accelerate network convergence and reduce internal covariate shift. The second convolutional layer is used for further feature extraction, and then a normalization layer is also set. The features output by the N Blocks are connected through the ReLU activation function, and then the multi-scale features are integrated through a convolutional operation to output global features, realizing multi-scale hierarchical feature extraction.
[0053] Step 3: Train the model. Divide the dataset into different time windows, and use the pre-training and fine-tuning parameter methods to train the model, and evaluate the training results.
[0054] During the training process, the traffic flow dataset is divided into three time windows of 30 minutes, 60 minutes, and 90 minutes. It is directly trained on the 30-minute dataset, and the parameters are fine-tuned and trained on the 60-minute and 90-minute datasets.
[0055] The training results are evaluated using the Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and Mean Absolute Error (MAE) metrics.
[0056]
[0057] Step Four: Conduct ablation experiments
[0058] Ablation experiments are conducted on the improved method to verify the effectiveness of each module and analyze the reasons for the effects produced by each module. The LMCEA module and the NDLCCA module are removed respectively, and the trained results are compared with the unchanged MCSTA model. By comparing the ablation of each module, we can verify the effectiveness of each module. The results are as follows:
[0059] After removing the LMCEA module, both the RMSE value and the MAE value of the model increase, and the model performance decreases significantly, indicating that the LMCEA module plays a key role in mining the spatio-temporal correlation of traffic data and multi-dimensional collaborative prediction.
[0060] After removing the NDLCCA module, both the RMSE value and the MAE value of the model increase, and the model performance decreases significantly, indicating that the NDLCCA module plays a key role in local cross-channel interaction learning and improving prediction accuracy.
[0061] Step Five: Implement deployment, deploy the model into the intelligent transportation system to achieve real-time prediction and application of traffic flow
[0062] Deploy MCSTA in the intelligent transportation system and test the effect of the model in actual traffic flow prediction. We used MCSTA to predict the traffic flow in New York from 0h to 24h in a day. At the same time, we also statistically recorded the real traffic flow data in New York in real time. We plotted the predicted values and the real values on the same curve graph. By comparison, it can be found that the predicted values of our model are very close to the real values, and in terms of accuracy, precision, and efficiency, it has surpassed the current SOTA model.
[0063] The foregoing description of specific exemplary embodiments of the present invention is for purposes of illustration and exemplification. These descriptions are not intended to limit the invention to the precise forms disclosed, and obviously, many changes and variations are possible in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the invention and its practical applications, so that those skilled in the art can implement and utilize various different exemplary embodiments of the invention, as well as various different selections and changes. The scope of the present invention is intended to be defined by the claims and their equivalents.
Claims
1. A traffic flow prediction method based on a multi-dimensional collaborative spatio-temporal attention model, characterized in that Including: Step 1: Collect urban traffic data and perform cleaning, normalization, and one-hot encoding on the data; Step 2: Construct a multi-dimensional collaborative spatio-temporal attention model (MCSTA), which includes a Transformer self-attention module, an LMCEA module, and an NDLCCA module; Step 3: Train the model, divide the dataset into different time windows, use the pre-training and fine-tuning parameter methods to train the model, and evaluate the training results; Step 4: Conduct ablation experiments; Step 5: Implement deployment, deploy the model into the intelligent transportation system to achieve real-time prediction and application of traffic flow.
2. The traffic flow prediction method based on the multi-dimensional collaborative spatio-temporal attention model according to claim 1, wherein The data collected in Step 1 is the traffic flow data of shared bicycles and taxis in New York. After dividing the urban area into grids, record the inflow and outflow information of vehicles in each grid area. The urban area is evenly and regularly divided into I×J grids. For the grid (i,j) located in the i-th row and the j-th column, its inflow and outflow volumes within the time interval t are defined as follows:
3. The traffic flow prediction method based on the multi-dimensional collaborative spatio-temporal attention model according to claim 1, wherein In Step 1, the collected traffic data is cleaned to remove outliers and noise, and the Min-Max normalization method is used to scale the traffic flow values to the range of [-1,1]. One-hot encoding is performed on external factors such as weather and time.
4. The traffic flow prediction method based on the multi-dimensional collaborative spatio-temporal attention model according to claim 1, wherein Step 2 to construct the multi-dimensional collaborative spatio-temporal attention model (MCSTA) includes: (1) Integrate the proposed LMCEA module into the backbone network part of the Transformer self-attention module in parallel. The LMCEA module contains three parallel branches, which respectively process the features in the width (W), height (H), and channel (C) directions of the space. Each branch generates corresponding attention weights through steps such as rotation operation, squeezing transformation, and excitation transformation, so as to enhance the model's attention to features in different dimensions; (2) Integrate the proposed NDLCCA module into the backbone network part of the Transformer self-attention module in parallel. The NDLCCA module uses one-dimensional convolution to capture local cross-channel interactions, avoiding the dimensionality reduction operation. Aggregate the spatial information through the global average pooling layer, and then adjust the features through the parameter-controlled operation layer and the activation normalization layer. After that, the data is divided into two parallel branches. One branch passes through multiple alternating convolutional layers and the local cross-channel interaction (LCCI) module, and the other branch passes through the adaptive kernel size determination mechanism (AKSDM) module. The data of the two branches is merged at the fusion node with the initial data to generate the final output; 1) Local cross-channel interaction (LCCI) module: The input features are first converted by the convolutional layer. The converted features form multiple sub-feature sets through the splitting operation. Each sub-feature set extracts the in-channel feature information using one-dimensional convolution. The processed sub-feature sets are spliced and merged, and then the feature fusion and adjustment are completed through another convolutional layer, so as to enhance the feature dependence ability and capture complex cross-channel relationships; 2) Adaptive kernel size determination mechanism (AKSDM) module: The input features are first unified in their feature representation forms by the convolutional layer and then enter a structure composed of N cascaded Blocks. Each Block contains two convolutional layers with a convolutional kernel of 3×3. The first convolutional layer is used for preliminary feature extraction, and then a normalization layer is set up to accelerate network convergence and reduce internal covariate shift. The second convolutional layer is used for further feature extraction, and a normalization layer is also set up afterwards. After connecting the features output by the N Blocks through the ReLU activation function, multi-scale features are integrated through convolutional operations to output global features, realizing multi-scale hierarchical feature extraction.
5. The traffic flow prediction method based on the multi-dimensional collaborative spatio-temporal attention model according to claim 1, characterized in that, In step 3, the traffic flow dataset is divided into three time windows of 30 minutes, 60 minutes, and 90 minutes. It is directly trained on the 30-minute dataset, and the parameters are fine-tuned and trained on the 60-minute and 90-minute datasets.
6. The traffic flow prediction method based on the multi-dimensional collaborative spatio-temporal attention model according to claim 1, wherein, In step 3, the mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE) metrics are used to evaluate the training results.
7. The traffic flow prediction method based on the multi-dimensional collaborative spatio-temporal attention model according to claim 1, wherein, In step 4, ablation experiments are carried out on the traffic flow dataset, including removing the LMCEA module and the NDLCCA module respectively, and comparing the trained results with the unchanged MCSTA model.
8. The traffic flow prediction method based on the multi-dimensional collaborative spatio-temporal attention model according to claim 1, characterized in that, In step 5, this method will be deployed on the intelligent transportation system to predict the urban traffic flow in real time, and make a reasonable arrangement for the construction of the intelligent transportation according to the prediction results.