Traffic flow prediction method based on temporal-spatial feature fusion Transform in edge environment
By adopting the multi-head convolution low-rank decomposition attention mechanism and attention graph convolution of the Transformer model in the edge environment, combined with the gated unit, the problem of difficult space-time dependence and large resource overhead in traffic flow prediction is solved, and efficient and accurate traffic flow prediction is achieved.
Patent Information
- Application Number
- CN202510493358.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
In the edge environment, existing traffic flow prediction methods are difficult to effectively capture space-time dependencies, the prediction is insufficient real-time and the resource overhead is large, which affects the prediction accuracy and efficiency.
The multi-head convolution low-rank decomposition attention mechanism based on Transformer and attention graph convolution combined with gating unit are adopted to adaptively fuse space-time features to predict traffic flow through feedforward neural network.
It improves the accuracy and efficiency of traffic flow prediction, reduces resource overhead, adapts to complex space-time dependencies, and improves the real-time and robustness of the model.
Smart Images

Figure CN120356330A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of edge traffic flow prediction, and specifically relates to a traffic flow prediction method based on spatiotemporal feature fusion Transformer in an edge environment. Background Art
[0002] The popularity and development of the Internet of Things (IoT) has promoted the construction of smart cities. As an important part of smart cities, the Intelligent Traffic System (ITS) aims to provide intelligent management for urban traffic by using advanced technologies such as intelligent control and data communication. In ITS, traffic flow prediction is an indispensable link, which plays a vital role in planning travel routes, ensuring traffic safety, and alleviating traffic congestion. As an emerging computing paradigm in the IoT era, Mobile Edge Computing (MEC) provides a new solution to the traffic flow prediction problem in ITS by deploying edge servers near terminal devices to reduce data transmission delay and bandwidth consumption. By uploading traffic flow data to the edge server and then calling the prediction model to analyze and process historical and real-time traffic flow data, the traffic flow in the future period can be predicted. In the actual MEC environment, traffic flow is changing all the time, showing a high degree of dynamics and complex spatiotemporal dependencies, which makes it difficult to accurately and efficiently predict traffic flow. Therefore, how to accurately and efficiently predict traffic flow is a challenging problem that needs to be solved urgently.
[0003] Traffic flow data usually comes from various sources such as road sensors, cameras, and ultrasonic detection devices, and it contains characteristic information such as section flow, vehicle speed, and road congestion rate. With the development of ITS, the amount of traffic flow data has shown an explosive growth and is in various forms, with dynamic and complex spatio-temporal dependence and interactivity. To achieve accurate and efficient traffic flow prediction, it is necessary to extract effective spatio-temporal characteristic information from the original traffic flow data and design appropriate methods for prediction. Existing prediction methods include experimental simulation, statistics-based methods, classical machine learning-based methods, and deep learning-based methods. Specifically, experimental simulation has high requirements for domain knowledge reserve and low flexibility, and it is difficult to construct an accurate traffic flow prediction model in the face of a real traffic environment with complex situations and many influencing factors. In contrast, statistics-based methods based on the data itself have a lower dependence on domain knowledge but have high requirements for data quality, while real traffic flow data often fails to meet this requirement. In recent years, machine learning has been applied to traffic flow prediction, but when dealing with massive and complex traffic flow data, it cannot effectively extract the spatio-temporal characteristics and long-distance dependence relationships of traffic flow, seriously affecting the prediction accuracy and efficiency. With the rapid development of deep learning technology, it has shown strong capabilities in learning complex non-linear data patterns and is expected to address complex traffic flow prediction problems, but it also faces the following challenges:
[0004] (1) Spatio-temporal dependence is difficult to capture. Traffic flow data contains complex dependence relationships in the time and space dimensions, and the changes in both have a high spatio-temporal correlation. In addition, due to the interaction of spatio-temporal data, its data pattern has complex non-linear characteristics. Therefore, how to effectively capture the spatio-temporal characteristics and their dependence from traffic flow data is the key challenge to improve the accuracy of traffic flow prediction.
[0005] (2) It is difficult to ensure prediction real-time performance. In actual ITS, traffic flow often shows highly dynamic changes, and traffic flow prediction attaches great importance to real-time performance and timeliness. Therefore, it is required that the prediction model can quickly process massive data and give prediction results in a timely manner to respond to the real-time changes of traffic flow. Classic deep learning methods have a large delay in data processing and traffic flow prediction and are difficult to meet the high requirements for real-time performance in ITS traffic flow prediction.
[0006] (3) The resource overhead of the prediction model is large. With the wide application of edge intelligent devices and sensors, edge servers need to process massive traffic flow data. Existing deep learning-based traffic flow prediction methods usually have high computational complexity, which causes huge computational pressure on edge servers with limited resources.
[0007] Therefore, it is necessary to design an efficient traffic flow prediction method to effectively reduce resource overhead while ensuring prediction accuracy. Summary of the Invention
[0008] The purpose of the present invention is to provide a traffic flow prediction method based on spatio-temporal feature fusion Transformer in an edge environment, which can improve the prediction accuracy, effectively reduce the resource overhead, and improve the prediction efficiency.
[0009] To achieve the above purpose, the technical solution adopted by the present invention is: a traffic flow prediction method based on spatio-temporal feature fusion Transformer in an edge environment. First, perform feature embedding and encoding on the original traffic flow data; then, construct a multi-head convolutional low-rank decomposition attention mechanism to capture long-term time dependencies and obtain local context information; then, construct attention graph convolution to capture spatial dependencies; finally, adaptively fuse the spatio-temporal features through a gated unit, and then use a feed-forward neural network and a linear layer to achieve accurate prediction of future traffic flow.
[0010] Furthermore, the method includes:
[0011] S1. Construct an edge traffic flow prediction model and define the problem;
[0012] S2. Predict future traffic flow data.
[0013] Furthermore, the specific implementation method of step S1 is as follows:
[0014] First, construct an undirected graph G = {V, E, A} to represent the road topology of the prediction area; the vertex set V = {v1, v2,..., v N} represents the set of sensors, where N is the number of sensors; the edge set E = {e1, e2,..., e C} represents the distances between different sensors, where C is the number of edges; the adjacency matrix A ∈ R N×N represents the connection relationship between different sensors, which is defined as:
[0015]
[0016] where, A ij is the element in the i-th row and j-th column of the adjacency matrix A;
[0017] The feature matrix of the training samples is represented as X ∈ R N×F×T ; X = [X0, X t ,..., X T , where X t ∈ R N×F , F represents the number of traffic flow features, and T represents the length of the historical traffic flow sequence; based on the above definitions, the traffic flow prediction problem is formally expressed as:
[0018] [X T+1 ,X T+2 ,...,X T+P = f(G, [X0, X1,..., X T ) (2)
[0019] Among them, f represents the mapping function for prediction, and p represents the future traffic flow prediction length;
[0020] The goal of traffic flow prediction is to minimize the error between the predicted value and the true value; the optimization goal of the traffic flow prediction problem is expressed as:
[0021]
[0022] Among them, Y t and are the true value and the predicted value of the traffic flow respectively, and loss is the loss function; the result of traffic flow prediction is used to support the intelligent transportation system to make intelligent decisions.
[0023] Furthermore, to evaluate the accuracy of traffic flow prediction, MAE, RMSE, and MAPE are introduced as performance indicators, and their definitions are:
[0024]
[0025] Among them, n represents the total number of samples.
[0026] Furthermore, the specific implementation method of step S2 is as follows:
[0027] S21. Perform feature embedding and position encoding processing on historical traffic flow through the input layer;
[0028] S22. Use the encoding layer to extract and fuse complex spatio-temporal dependencies to achieve edge load prediction;
[0029] S23. Output the future traffic flow prediction result through the linear output layer.
[0030] Furthermore, in step S22, the encoding layer is mainly composed of a multi-head convolutional low-rank decomposition attention mechanism, an attention graph convolution, and a gated unit; the traffic flow features processed in step S21 are simultaneously input into the multi-head convolutional low-rank decomposition attention mechanism and the attention graph convolution to capture temporal and spatial dependencies respectively, and then are adaptively fused through the gated unit to obtain the spatio-temporal features of the traffic flow; subsequently, the fusion result is input into a feed-forward neural network to further extract effective spatio-temporal features, and two residual connections are introduced to avoid the problem of gradient disappearance.
[0031] Furthermore, the implementation method of the multi-head convolutional low-rank decomposition attention mechanism is:
[0032] First, to obtain local context information, before calculating the attention matrix, causal convolutions are performed on the input query Q and key K, which are defined as:
[0033]
[0034] Among them, Φ k represents the convolution kernel, m represents the size of the convolution kernel, represents the learnable parameter;
[0035] Next, for the input query Q, key K, and value V, calculate the attention matrix and perform low-rank factorization; this process is defined as:
[0036]
[0037] Among them, d k represents the scaling factor, and represent low-rank matrices;
[0038] Construct a multi-head self-attention mechanism, which is defined as:
[0039] H t = concact(head1,head2,...,head h )W O (9)
[0040] head h = Attention(ccon Q (Q),ccon Q (K),VW V ) (10)
[0041] Among them, concact(·) represents the concatenation operation, h represents the number of attention heads, W O and W V represent learnable parameters, and the final output H t ;
[0042] The implementation method of the attention graph convolution is as follows:
[0043] First, calculate the attention scores for the input features; subsequently, perform graph convolution operations, which are defined as:
[0044]
[0045] Among them, D is the degree matrix of A, and Θ is the linear projection function;
[0046] The gating unit is defined as:
[0047]
[0048] Among them, and represent learnable parameters.
[0049] The present invention also provides a traffic flow prediction system based on spatio-temporal feature fusion Transformer in an edge environment, including a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the above-mentioned method can be implemented.
[0050] The present invention also provides a computer-readable storage medium, on which computer program instructions executable by the processor are stored. When the processor runs the computer program instructions, the above-mentioned method can be implemented.
[0051] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a traffic flow prediction method based on spatio-temporal feature fusion Transformer in an edge environment. Based on the Transformer model, a parallel structure is adopted to effectively capture the spatio-temporal dependencies in traffic flow, thereby improving the accuracy and efficiency of traffic flow prediction. First, feature embedding and encoding are performed in the input layer; subsequently, the parallel structure in the encoding layer is used to capture temporal and spatial dependencies respectively and perform adaptive fusion; finally, the prediction result of future traffic flow is output through the output layer. Based on 4 real traffic datasets and comparison with 8 benchmark methods, the superiority of the present method is verified. Experimental results show that compared with other methods, the TFPformer method achieves the optimal prediction accuracy in different performance metrics. In addition, through ablation experiments, the effectiveness of each component in the TFPformer method in improving the accuracy and efficiency of the traffic flow prediction model is comprehensively verified. Description of the Drawings
[0052] Figure 1 It is a traffic flow prediction model diagram for the edge environment in the embodiment of the present invention;
[0053] Figure 2 It is an implementation architecture diagram of the traffic flow prediction method based on spatio-temporal feature fusion Transformer in the edge environment proposed in the embodiment of the present invention;
[0054] Figure 3 It is an ablation experiment result diagram of the embodiment of the present invention on the PEMS03 dataset;
[0055] Figure 4 It is an ablation experiment result diagram of the embodiment of the present invention on the PEMS04 dataset;
[0056] Figure 5This is the ablation experiment result graph of the embodiment of the present invention on the PEMS07 dataset;
[0057] Figure 6 This is the ablation experiment result graph of the embodiment of the present invention on the PEMS08 dataset;
[0058] Figure 7 This is the effectiveness result graph of low-rank decomposition in improving the training efficiency in the embodiment of the present invention. Detailed implementation manners
[0059] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0060] It should be noted that the following detailed description is exemplary and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0061] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0062] The present invention provides a traffic flow prediction method (TFPformer) based on spatio-temporal feature fusion Transformer in an edge environment, which adopts a parallel structure to capture the spatio-temporal dependence of traffic flow. First, a multi-head convolutional low-rank decomposition attention mechanism and an attention graph convolution are designed to capture the temporal and spatial features of traffic flow respectively. In particular, causal convolution is introduced to obtain local context information, and the attention matrix is decomposed into low rank, thereby effectively reducing the memory burden and greatly improving the model training efficiency. Subsequently, a gated unit is used to adaptively fuse the spatio-temporal features. Finally, a feed-forward neural network and a linear layer are used to output the prediction result. The prediction process fully extracts and retains the original spatio-temporal feature information, avoiding potential feature loss.
[0063] The following is the specific implementation process of this method.
[0064] 1. System model and problem definition
[0065] In order to better meet the actual application requirements of traffic flow prediction in ITS, the present invention proposes a traffic flow prediction model for the MEC environment, as Figure 1As shown below. First, the edge server collects the original traffic flow data from the sensors on the road and preprocesses it. Subsequently, the proposed TFPformer method is called to extract and fuse the spatio-temporal features, thereby achieving accurate and efficient traffic flow prediction. Finally, the prediction results are used to support the ITS for intelligent decision-making.
[0066] For the spatio-temporal prediction problem of traffic flow prediction, the present invention first constructs an undirected graph G = {V, E, A} to represent the road topology of the prediction area. The vertex set V = {v1, v2,..., v N} represents the set of sensors, where N is the number of sensors. The edge set E = {e1, e2,..., e C} represents the distances between different sensors, where C is the number of edges. The adjacency matrix A ∈ R N×N represents the connection relationship between different sensors, which is defined as:
[0067]
[0068] where, A ij is the element in the i-th row and j-th column of the adjacency matrix A.
[0069] Next, the feature matrix of the training samples can be represented as X ∈ R N×F×T . X = [X0,... X t ,..., X T , where X t ∈ R N ×F , F represents the number of traffic flow features, and T represents the length of the historical traffic flow sequence. Based on the above definitions, the traffic flow prediction problem can be formally represented as:
[0070] [X T+1 , X T+2 ,..., X T+P = f(G, [X0, X1,..., X T ) (2)
[0071] where, f represents the mapping function for prediction, and p represents the future traffic flow prediction length.
[0072] The goal of traffic flow prediction is to minimize the error between the predicted value and the true value. Specifically, the true value and the predicted value of the traffic flow are denoted as Y t and loss is the loss function. Therefore, the optimization goal of the traffic flow prediction problem can be represented as:
[0073]
[0074] The results of traffic flow prediction will be used to support intelligent decision-making in intelligent transportation systems.
[0075] To comprehensively evaluate the prediction performance, this embodiment uses metrics such as Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Mean Absolute Percentage Error (MAPE). The value ranges of the three metrics are all in [0, 1], and the closer to 0, the better the prediction performance. The specific definitions are as follows:
[0076]
[0077] where n represents the total number of samples.
[0078] 2. Overview of the technical solution
[0079] To solve the traffic flow prediction problem in the edge environment, the present invention proposes a novel traffic flow prediction method based on spatio-temporal feature fusion Transformer (TFPformer) in the edge environment. Its implementation is shown in Algorithm 1.
[0080]
[0081]
[0082] As Figure 2 shown, the proposed TFPformer method mainly consists of three parts, including an input layer, an encoding layer, and an output layer. The encoding layer is mainly composed of a multi-head convolutional low-rank decomposition attention mechanism, an attention graph convolution, and a gated unit. First, the input layer performs feature embedding and position encoding processing on historical traffic flows. Then, the processed traffic flow features are simultaneously input into the multi-head convolutional low-rank decomposition attention mechanism and the attention graph convolution to capture temporal and spatial dependencies respectively, and then are adaptively fused through the gated unit to obtain the spatio-temporal features of traffic flows. Subsequently, the fusion result is input into a feed-forward neural network to further extract effective spatio-temporal features, and two residual connections are introduced to avoid the problem of gradient disappearance. Finally, the future traffic flow prediction result is output through the output layer. The design of the encoding layer will be introduced in detail below.
[0083] 2.1. Multi-head convolutional low-rank decomposition
[0084] The present invention designs a multi - head convolutional low - rank decomposition attention mechanism for capturing the temporal dependencies of input features. The classical Transformer model adopts a global attention mechanism, which is difficult to effectively model local context information. To address this problem, the present invention combines causal convolution with the multi - head attention mechanism, enhancing the local modeling ability while maintaining the temporal causal relationship of data, thereby improving the model's learning ability for temporal dependencies. In the classical multi - head attention mechanism, the dot - product method is usually used to calculate the similarity between each input position and other positions, resulting in a large amount of memory consumption. To alleviate this problem, the present invention introduces low - rank decomposition for the attention matrix to reduce its computational complexity, thereby reducing the memory burden and improving the model training efficiency.
[0085] First, to obtain local context information, before calculating the attention matrix, causal convolution needs to be performed on the input query Q and key K, which is defined as:
[0086]
[0087] where, Φ k represents the convolutional kernel, m represents the size of the convolutional kernel, represents the learnable parameter.
[0088] Next, for the input query Q, key K, and value V, calculate the attention matrix and perform low - rank decomposition. This process is defined as:
[0089]
[0090] where, d k represents the scaling factor, and represent low - rank matrices.
[0091] For the single - head self - attention mechanism, its averaging suppresses the representation subspace information from different positions. In contrast, the multi - head self - attention mechanism can effectively solve this problem, which is defined as:
[0092] H t = concact(head1,head2,...,head h )W O (9)
[0093] head h = Attention(ccon Q (Q),ccon Q (K),VW V ) (10)
[0094] Among them, concact(·) represents the concatenation operation, h represents the number of attention heads, and W O and W V represent learnable parameters, and the final output H t is obtained.
[0095] 2.2. Attention Graph Convolution
[0096] The present invention designs an attention graph convolution for capturing the spatial dependencies of input features. Classical graph convolutions usually use a predefined graph adjacency matrix to aggregate neighboring nodes, which has certain limitations. Because it assumes that neighboring nodes have the same contribution, it is less flexible and cannot effectively represent complex spatial relationships in the data. To address this problem, based on classical graph convolutions, the present invention introduces scaled dot products to calculate the similarity between nodes. This design allows the model to adaptively adjust the weights of neighboring nodes, thereby effectively learning node features to capture the complex spatial dependencies of nodes.
[0097] First, for the input features, calculate their attention scores. Subsequently, perform a graph convolution operation, which is defined as:
[0098]
[0099] where D is the degree matrix of A, and Θ is the linear projection function.
[0100] 2.3. Gating Unit
[0101] In classical GRUs, gating mechanisms are used to control information propagation, which can be used to fuse multiple feature information and determine the contribution weights of each path to the final output. In view of this, the present invention introduces a gating unit to adaptively fuse the captured spatio-temporal dependencies to effectively balance the interaction between time and space information. By using the gating unit, it is possible to adaptively determine the contribution of each tuning path to the final output, thereby improving the model performance. This process is defined as:
[0102]
[0103] where and represent learnable parameters.
[0104] 3. Method Evaluation
[0105] The experiments of this invention were carried out on a high-performance workstation equipped with an Intel(R) Core(TM) i5-12600KF CPU and an NVIDIA GeForce RTX 4060 GPU, with the CUDA driver version being 11.6. Based on the deep learning framework PyTorch 2.3.1+cu121, the proposed TFPformer method was implemented. In this embodiment, a large number of experiments were conducted on 4 real traffic flow datasets to verify the effectiveness and superiority of the TFPformer method. The PEMS series datasets record the highway traffic flow in California, and the specific information is shown in Table 1. The original data was collected by road sensors every 5 minutes, and it contains feature information such as flow, occupancy, and speed. In the embodiment, graph data was constructed according to the flow characteristics and the geographical location information of road sensors.
[0106] Table 1 Dataset Description
[0107]
[0108] This method first preprocesses the original traffic flow data, including data standardization, smoothing processing, etc., to avoid the influence of some outliers or error values on the experiment. Subsequently, the dataset was divided into a training set, a validation set, and a test set, with a ratio of 6:2:2. In addition, the number of training epochs M was set to 100, the training batch size B was 8 (for PEMS03, PEMS04, PEMS08) and 4 (for PEMS07), the learning rate was 0.001, 4 encoding modules and 8 attention heads were used, the data embedding dimension was 64, and the traffic flow for the next 12 time steps was predicted using 12 historical time steps (1 hour).
[0109] To evaluate the superiority of the proposed TFPformer method, a large number of comparative experiments were conducted between this method and the following 8 benchmark methods:
[0110] (1) HA is a statistical method that predicts future values by calculating the historical average and is commonly used for predictions with periodic or seasonal patterns.
[0111] (2) LSTM is a classic variant of RNN and is specifically designed to handle time series prediction problems.
[0112] (3) STGCN does not use traditional CNNs and recurrent units, but describes the problem as a graph and uses a fully convolutional structure to build a prediction model.
[0113] (4) MTGNN is a general GNN framework specifically designed for multivariate time series data.
[0114] (5) The AGCRN introduces node - adaptive parameter learning, adaptive graph generation, and a recurrent network to capture spatio - temporal correlations in the sequence.
[0115] (6) The DCRNN uses bidirectional random walks and an encoder - decoder architecture with a predetermined sampling to capture spatial and temporal dependencies respectively.
[0116] (7) The Graph - WaveNet introduces an adaptive dependency matrix to combine node embeddings to capture hidden spatial dependencies in the data.
[0117] (8) The ASTGNN captures temporal and spatial dependencies in the data through self - attention mechanisms and dynamic graph convolutions respectively.
[0118] 4. Experimental Results and Analysis
[0119] To verify the superiority of the proposed TFPformer method, this embodiment compares the performance of different methods on four datasets, and the experimental results are shown in Tables 2-5. Specifically, the DCRNN method has almost the worst performance on each dataset. The MAE, RMSE, and MAPE on the PEMS03 dataset are 52.49, 38.82, and 74.91%, respectively. This is because the model structure of the DCRNN method is complex, and it models spatio-temporal dependencies through diffusion convolution. However, in real-world scenarios, the model is highly sensitive to new data and has weak adaptability, resulting in its inability to effectively handle the differences between training and test data. The remaining five methods consider spatio-temporal features simultaneously, and the prediction effects are significantly better than those of the classical HA and LSTM methods. Compared with the HA and LSTM methods, the STGCN method has achieved obvious performance improvement. However, it uses graph convolution and 1-D convolution and is difficult to fully capture complex spatio-temporal dependencies, especially when facing long-term predictions, there are performance bottlenecks. Both the AGCRN and Graph-WaveNet methods adopt an adaptive mechanism to extract spatio-temporal features from the data. Among them, the AGCRN method performs well on each dataset and achieves the best MAPE on the PEMS03 dataset. The Graph-WaveNet method has achieved two best indicators on the PEMS04, PEMS07, and PEMS08 datasets. This is because the adaptive mechanism helps to enhance the expressive ability of the model. However, these two methods extract spatio-temporal features in a sequential manner, which weakens or loses the association between space and time to a certain extent, thus affecting the quality of spatio-temporal features and reducing the prediction performance. The MAE and RMSE of the ASTGNN method on the PEMS03 dataset are 14.68 and 25.10, respectively, and it has achieved the best MAE or MAPE on the remaining three datasets. It also extracts spatio-temporal features from the data in stages and introduces a self-attention mechanism, so the computational complexity is relatively high, increasing the computational cost of model training. Compared with other methods, the proposed TFPformer method has achieved the best results in terms of three performance indicators on four datasets. Specifically, the MAE, RMSE, and MAPE of the TFPformer method are reduced by 30.25%, 23.78%, and 27.73% on average on the four datasets, respectively. This is due to the fact that the TFPformer method adopts a parallel multi-head convolutional low-rank decomposition attention mechanism and attention graph convolution, which can independently capture time and space dependencies and maintain the interactivity of spatio-temporal information, avoiding potential feature conflicts, and thus retaining important spatio-temporal information to the greatest extent. Further, the spatio-temporal information is adaptively fused through a gated unit to improve the robustness and adaptability of the model.
[0120] Comparison of Different Methods on the PEMS03 Dataset
[0121]
[0122] Table 3 Comparison of Different Methods on the PEMS04 Dataset
[0123]
[0124] Table 4 Comparison of Different Methods on the PEMS07 Dataset
[0125]
[0126] Table 5 Comparison of Different Methods on the PEMS08 Dataset
[0127]
[0128]
[0129] To verify the effectiveness of each component in the proposed TFPformer method, an ablation experiment was conducted in this embodiment. Compared with the TFPformer method, the TFPformer-A method does not include the attention graph convolution, the TFPformer-B method does not include the multi-head convolutional low-rank decomposition attention mechanism, and the TFPformer-C method does not include the attention graph convolution and the multi-head convolutional low-rank decomposition attention mechanism. The performance of each method on different datasets is as Figures 3 - 6 shown.
[0130] On the PEMS08 dataset, the MAE, RMSE, and MAPE of the TFPformer-A method are 11.91, 20.84, and 8.70%, respectively. The three metrics of the TFPformer-B method are 12.06, 21.02, and 8.72%, respectively. The three metrics of the TFPformer-C method are 12.11, 21.11, and 9.43%, respectively. Compared with the TFPformer-C method, the TFPformer-A and TFPformer-B methods are improved by 1.65%, 1.25%, 7.73% and 0.41%, 0.41%, 7.46% in the three metrics respectively, with a slight performance improvement. TFPformer effectively combines the attention graph convolution and the multi-head convolution low-rank decomposition attention mechanism. Compared with the TFPformer-C method, it is improved by 20.37%, 19.35%, and 34.91% in the three evaluation metrics respectively, with a significant performance improvement. This verifies the effectiveness of the two components in the TFPformer method and also reflects the effectiveness of the parallel structure in the TFPformer method. It is difficult to effectively capture spatio-temporal dependencies by simply relying on time or space information, so the improvement effect on the model performance is limited. By parallelly learning time and space dependencies, the modeling ability of the model for spatio-temporal features can be effectively improved, thereby significantly improving the prediction effect. On the PEMS03 dataset, compared with the TFPformer-C method, the TFPformer method is improved by 16.58%, 14.49%, and 4.23% in the three metrics respectively. On the PEMS04 dataset, the TFPformer method is improved by 17.07%, 16.49%, and 17.89% in the three metrics respectively. On the PEMS07 dataset, the TFPformer method is improved by 15.72%, 14.46%, and 18.46% in the three metrics respectively. The above experiments fully verify the effectiveness of each component in the proposed TFPformer method.
[0131] Furthermore, this embodiment evaluates the effectiveness of the multi-head convolution low-rank decomposition mechanism in the proposed TFPformer method in improving the model training efficiency. Specifically, the TFPformer-D method does not include the low-rank decomposition part in the multi-head convolution low-rank decomposition mechanism, and then compares the time used for training 100 rounds between it and the TFPformer method. As Figure 7As shown in the figure, compared with the TFPformer-D method, the training time of the TFPformer method on the PSEM08, PSEM04, PSEM03, and PSEM07 datasets is reduced by 61.01s, 159.65s, 234.98s, and 517.03s respectively, and the reduction ratios are 1.35%, 1.77%, 1.86%, and 1.12% respectively, which is roughly proportional to the number of nodes in the dataset. It can be seen from the experimental results that the low-rank decomposition designed in the TFPformer method helps to shorten the model training time and improve the model training efficiency.
[0132] This embodiment also provides a traffic flow prediction system based on spatio-temporal feature fusion Transformer in an edge environment, including a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the above method can be implemented.
[0133] This embodiment also provides a computer-readable storage medium, on which computer program instructions executable by the processor are stored. When the processor runs the computer program instructions, the above method can be implemented.
[0134] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0135] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0136] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 or blocks Figure 1 specified in one or more of the processes and / or blocks.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 or blocks Figure 1 specified in one or more of the processes and / or blocks.
[0138] As described above, the above are only preferred embodiments of the present invention, and are not limitations on the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A traffic flow prediction method based on spatio-temporal feature fusion Transformer in an edge environment, characterized in that First, perform feature embedding and encoding on the original traffic flow data; then, construct a multi-head convolutional low-rank decomposition attention mechanism to capture long-term temporal dependencies and obtain local context information; next, construct attention graph convolution to capture spatial dependencies; finally, adaptively fuse spatio-temporal features through a gated unit, and then use a feed-forward neural network and a linear layer to achieve accurate prediction of future traffic flow.
2. The traffic flow prediction method based on spatio-temporal feature fusion Transformer in the edge environment according to claim 1, wherein Including: S1. Construct an edge traffic flow prediction model and define the problem. S2. Predict future traffic flow data.
3. The traffic flow prediction method based on spatio-temporal feature fusion Transformer in the edge environment according to claim 2, wherein The specific implementation method of step S1 is as follows: First, construct an undirected graph $G = \{V, E, A\}$ to represent the road topology of the prediction area; the vertex set $V=\{v_1, v_2, \ldots, v$ N $\}$ represents the set of sensors, where $N$ is the number of sensors; the edge set $E = \{e_1, e_2, \ldots, e$ C $\}$ represents the distances between different sensors, where $C$ is the number of edges; the adjacency matrix $A\in\mathbb{R}$ N×N represents the connection relationships between different sensors, which is defined as: where A ij is the element in the i-th row and j-th column of the adjacency matrix A; The feature matrix of the training samples is represented as X ∈ R N×F×T ; X = [X0, X t ,..., X T , where X t ∈ R N×F , F represents the number of traffic flow features, and T represents the length of the historical traffic flow sequence; based on the above definitions, the traffic flow prediction problem is formally represented as: [X T+1 ,X T+2 ,...,X T+P = f(G,[X0,X1,...,X T ) (2) Where f represents the mapping function for prediction, and p represents the future traffic flow prediction length. The goal of traffic flow prediction is to minimize the error between the predicted value and the true value; the optimization goal of the traffic flow prediction problem is expressed as: Among them, Y t and are respectively the true value and the predicted value of the traffic flow, and loss is the loss function; the result of the traffic flow prediction is used to support the intelligent transportation system to make intelligent decisions.
4. The traffic flow prediction method based on spatio-temporal feature fusion Transformer in the edge environment according to claim 3, wherein, To evaluate the accuracy of traffic flow prediction, MAE, RMSE, and MAPE are introduced as performance metrics, and their definitions are: Where n represents the total number of samples.
5. The traffic flow prediction method based on spatio-temporal feature fusion Transformer in the edge environment according to claim 3, characterized in that, The specific implementation method of step S2 is as follows: S21. Perform feature embedding and positional encoding processing on historical traffic flow through the input layer. S22. Use the encoding layer to extract and fuse complex spatio-temporal dependencies to achieve edge load prediction. S23. Output the future traffic flow prediction result through the linear output layer.
6. The traffic flow prediction method based on spatio-temporal feature fusion Transformer in the edge environment according to claim 5, wherein, In step S22, the encoding layer mainly consists of a multi-head convolutional low-rank decomposition attention mechanism, attention graph convolution, and a gated unit; the traffic flow features processed in step S21 are simultaneously input into the multi-head convolutional low-rank decomposition attention mechanism and attention graph convolution to capture temporal and spatial dependencies respectively, and then are adaptively fused through the gated unit to obtain the spatio-temporal features of traffic flow; subsequently, the fusion result is input into a feed-forward neural network to further extract effective spatio-temporal features, and two residual connections are introduced to avoid the problem of gradient disappearance.
7. The traffic flow prediction method based on spatio-temporal feature fusion Transformer in the edge environment according to claim 6, characterized in that The implementation method of the multi-head convolutional low-rank decomposition attention mechanism is: First, to obtain local context information, before calculating the attention matrix, perform causal convolution on the input query Q and key K, and its definition is: Among them, Φ k represents the convolutional kernel, m represents the size of the convolutional kernel, represents the learnable parameter; Next, for the input query Q, key K, and value V, calculate the attention matrix and perform low-rank decomposition; this process is defined as: where d k represents a scaling factor, and represents a low-rank matrix; Construct a multi-head self-attention mechanism, and its definition is: H t = concact(head1, head2,..., head h )W O (9) head h = Attention(ccon Q (Q),ccon Q (K),VW V ) (10) Among them, concact(·) represents the concatenation operation, h represents the number of attention heads, W O and W V represent learnable parameters, and the final output H t ; The implementation method of the attention graph convolution is: First, calculate its attention score for the input feature; subsequently, perform a graph convolution operation, and its definition is: Among them, D is the degree matrix of A, and Θ is the linear projection function; The gated unit is defined as: Among them, and represent learnable parameters.
8. A traffic flow prediction system based on spatio-temporal feature fusion Transformer in an edge environment, characterized in that Including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the method described in any one of claims 1-7 can be implemented.
9. A computer-readable storage medium storing computer program instructions executable by a processor, characterized in that, When the processor runs the computer program instructions, the method described in any one of claims 1-7 can be implemented.
Citation Information
Cited By
Traffic flow prediction method and device based on multi-level space-time and perception fusion
CN121092937A
Cross-modal space-time event prediction method and system, computer and storage medium
CN121564661A
A method, system, computer, and storage medium for predicting cross-modal spatiotemporal events.
CN121564661B